Commit Graph
3683 Commits
Author SHA1 Message Date
Jinwoo-H 87f160dafa fix: route folder deletion through runtime cleanup 2026-09-01 14:08:49 -04:00
Brennan BensonandMerge Sim 5150c52045 fix(dev): skip blocking keychain diagnostic (#17877)
* fix(dev): skip blocking keychain diagnostic

* fix(dev): preserve forced secret protection report

---------

Co-authored-by: Merge Sim <sim@local>
2026-09-01 11:02:54 -07:00
Jinwoo Hong 69fff5eaa3 fix(runtime): defer websocket heartbeat startup probe (#17810)
* fix(runtime): defer websocket heartbeat startup probe

* fix(runtime): defer heartbeat probes until websocket auth
2026-09-01 14:00:11 -04:00
Neil 26dfa46aa9 fix(wsl): resolve sparse-checkout and diff-stamp gitdirs through the caller's distro (#17932)
Two main-process `resolveGitDir` call sites dropped the WSL distro their caller
already held, so they could only resolve a gitdir pointer whose spelling carries
its own translation: a `//wsl.localhost/<Distro>/...` base, or a `/mnt/<letter>`
drvfs pointer that maps to a drive letter on its own.

The layout that needs the third case is a repo inside the distro's filesystem
with its worktrees on the Windows drive. `git worktree list` reports the worktree
as `/mnt/c/wt/x`, which the listing translates to `C:\wt\x` — a base that no
longer names a distro — while the `.git` gitfile beside it points at
`/home/me/repo/.git/worktrees/x`, which has no drive to derive. Win32 then treats
that pointer as absolute and reads a path that names nothing:

- `detectSparseCheckout` stats `info/sparse-checkout` under the fabricated path,
  always misses, and reports the worktree as non-sparse — no sparse badge, and
  the file list claims files that are not on disk.
- `readWorktreeDiffStamp` reads HEAD under the same path, gets nothing, and
  returns null. Null is the safe answer ("cannot prove unchanged"), but it
  retires the settled-diff cache for every file in that worktree, so each diff
  respawns Git.

Both callers already have the distro: the listing threads its
`GitWorktreeExecOptions` to `annotateSparseCheckoutStatus`, on through
`detectSparseCheckoutCached` (the annotation cache added by #17859) and its
background revalidation probe, and finally to `detectSparseCheckout`; and
`file-diff` already forwards its `GitRuntimeOptions` to `readWorktreeDiffStamp`,
which now forwards it to `resolveGitDir` as well.

`resolveGitMetadataPath` still prefers a UNC base's distro and still tries drvfs
before the caller-named distro, so nothing that resolved before resolves
differently.

The cache hop matters twice over. It is the only remaining caller of
`detectSparseCheckout`, so without threading it the fix would not reach the
probe at all. And the cache is where the bug turns sticky. #17859 keyed entries
on `repoPath` + `worktreePath` alone, on the reasoning that the distro is a
property of the repo and so every read for a given `repoPath` carries the same
one. That invariant does not hold. `listRepoWorktrees(repo)` is called with no
options at all from the filesystem-auth root rebuild
(`registered-worktree-roots-cache.ts`, reached from `ensureAuthorizedRootsCache`
on any auth check with a dirty cache) and from the local worktree-ownership
check in `filesystem-worktree-helpers.ts`. Both land on the *same* key as the
distro-carrying listing, because `translateWslOutputPaths` derives the distro
from the cwd spelling before falling back to `options.wslDistro`, so a
UNC-spelled repo path yields the identical `C:\...` worktree row either way.
Measured on Windows in one process, branch build: a distro-less read followed by
a distro-carrying read reported the sparse worktree as non-sparse both times.

So `wslDistro` now joins the cache key -- trimmed and lowercased, matching how
the rest of the codebase compares distro names, and appended last so the
repo-scoped prefix delete still matches every variant. The per-path invalidate
becomes a prefix delete for the same reason, dropping every distro variant of a
removed or moved worktree.

Keying on it closes both halves of the defect. A correct caller can no longer be
served an answer derived without the distro it supplied. And because the entry a
reader reaches is now selected by the same distro it would re-probe with,
`revalidateInBackground` can no longer re-derive a warm entry under weaker
options -- which mattered on its own: a distro-less reader crossing the
five-minute window would otherwise flip a correct `true` to `false`, and the
resulting change notification runs the registered invalidator, clearing the
whole repo's cache and re-probing every worktree cold, on a five-minute loop.

Cost of the extra key dimension is bounded by the number of distinct distros a
given repo is actually read under: one where a distro is threaded everywhere,
two while the distro-less callers above still exist. Entries are still
repo-scoped, and both clears already sweep by prefix.

Per-platform delta:
- macOS/Linux: no change. Guest-pointer translation is gated to win32 and a
  caller-named distro is ignored off Windows; no caller supplies one there, so
  the cache keys and probes exactly as before.
- native Windows, no WSL: no change. `wslDistro` is undefined, so the resolver
  takes exactly the branches it took before and every read keys on the same
  empty distro component, so the cache behaves exactly as it did.
- Windows + WSL, UNC-spelled worktree: no change. The base already names the
  distro and outranks the caller's.
- Windows + WSL, drvfs-spelled worktree with a drvfs pointer: no change. The
  drive-letter derivation still runs first.
- Windows + WSL, drvfs-spelled worktree with a non-drvfs pointer: the sparse
  badge appears and the settled-diff cache starts hitting. Both previously
  failed toward "not sparse" / "do not cache", so neither can now serve a stale
  answer, and the distro-less listings no longer share the badge's cache entry.
- SSH/relay: none. Those paths return through the provider branch before
  reaching either function.
- folder workspaces, GitLab: none. Neither is on these code paths.

Not in this change:
- `readRepoCommonDirFromDisk` (worktree-listing). Passing the distro there is
  inert: a repo root's `.git` is a directory, so the gitfile-pointer branch never
  runs, and when `repoPath` itself is guest-spelled the preceding `stat` already
  fails — which no `resolveGitDir` option can fix.
- The two `findExistingWorktreeSymlinkPaths` calls on the removal paths. Both
  receive `registeredWorktree.path` from `listWorktreesStrict`, which already
  translates every row out of the guest namespace, so the distro would be a
  no-op. The `removeWorktreeLinkedPaths` unlink beside them is untranslated too,
  so a half-threaded fix would only move the refusal from Orca's preflight to
  `git worktree remove`.
- An absolute `commondir` payload, which `resolveGitCommonDir` still resolves
  untranslated. Git writes that file relative in the layouts above, and the
  failure direction is unchanged.
- Giving the two distro-less `listRepoWorktrees(repo)` callers a distro. Neither
  reads `isSparse` -- both use only `worktree.path` -- so the distro would buy
  them nothing they consume, while resolving a project runtime inside the
  filesystem-auth rebuild would put a call that throws on `repair-required`
  behind a catch that skips the whole repo's authorized roots. The cache key
  makes their reads harmless; skipping the annotation for callers that never
  read it is a separate, larger change. The third no-options call in
  `hosted-review.ts` is inside the `repo.connectionId` branch and returns
  through the SSH provider, so it never reaches this cache.
2026-09-01 04:25:19 -07:00
djegor315-sketch f27a30d1b6 fix(ai-vault): avoid large Codex scan timeouts (#17889)
Agent Session History exceeded its 130-second deadline on large local Codex
histories. Three costs combined: excluded worker transcripts were recognized on
their first line but still drained to EOF (1,011 files / ~18.3 GiB on the
reported corpus), large ignored records were fully decoded and JSON.parsed, and
the persisted parse cache was discarded on every app update.

- Stop resumable reads the moment `session_meta` marks a worker transcript.
- Skip decode + `JSON.parse` for records the parser only feeds to the timeline.
  The skip set is the complement of what `consumeCodexRecordLine` reads, and
  applies only above the bounded prefix limit, so a long opening prompt (which
  is the session title) still takes the exact parser.
- Prove cross-volume rollout aliases from a bounded `session_meta` read routed
  through the WSL transcript FS gate, carrying the scan's AbortSignal, fanned
  out across contested candidates with bounded concurrency.
- Make parse-cache schema 2 the semantic compatibility boundary so an update no
  longer forces a multi-gigabyte cold scan, fenced by a build-time ratchet on
  the persisted session shape.
- Report early-stopped transcripts as their own `aiVault.scan` attribute.

Verified on macOS, Ubuntu over SSH, and Windows: read volume drops 336 -> 49.5
MiB identically on all three; the Windows failing-test set is byte-identical to
main. Reported corpus: 130s timeout -> 58.6s, 244 sessions, 0 issues.

Fixes #17888.
2026-09-01 04:21:43 -07:00
Neil f176e49478 fix(git): narrow fork-remote fetch refspecs to tracked branches (#17887)
* fix(git): narrow fork-remote fetch refspecs to tracked branches

git remote add with no -t writes the wide +refs/heads/*:refs/remotes/<name>/*
refspec, so any later plain `git fetch` (user, agent, or Orca's own Fetch
action) re-imports a fork's entire branch set and its tags -- one real
machine had ~50 leaked/wide fork remotes producing 59,716 remote-tracking
refs. Mint and reuse now pin -t <branch> --no-tags; a rate-limited sweep
narrows and cleans up remotes minted before this fix; gitFetch self-heals
when a narrowed remote's tracked branch is later deleted upstream.

Refs #17828

* fix(git): soften narrow fork-remote refspec against deleted upstream branches

A bare `git fetch` in a worktree checked out on a fork-PR branch resolves to
the pr-* remote via branch.<name>.remote -- not origin -- making it the
dominant fetch shape in Orca's terminal-centric, agent-driven usage. The
previous literal-refspec design hard-failed that fetch ("couldn't find
remote ref") the moment the tracked branch was deleted/renamed upstream,
which is not the narrow edge case it was first described as.

Switch to a trailing-`*`-suffixed refspec source/destination
(refs/heads/<branch>*:refs/remotes/<name>/<branch>*). Verified against real
git: this restores wildcard zero-match tolerance (silent no-op instead of a
hard failure) and lets plain `git fetch --prune` reclaim the stale ref once
the branch disappears, at the cost of also matching sibling branches that
share the literal name as a prefix -- a materially smaller widening than the
original unbounded-import bug.

Also close a race with #17842's orphaned-pr-remote reconciliation sweep:
both sweeps read the same worktree-metadata store to pick candidate remotes,
so reconciliation can `remote remove` a remote this migration is
concurrently narrowing. `ensureRemoteTracksBranchNarrowly`'s plain `config
--add` would silently resurrect a url-less config section in that case;
re-check `remote.<name>.url` (via the new `remoteHasUrl`, plumbing rather
than porcelain `remote get-url`, which falls back to echoing the remote name
as a bogus URL) after the narrowing writes and remove the section if it's
gone.

* fix(git): update stale fork-remote mint assertions for -t/--no-tags and wildcard-suffix refspec

Four test files still asserted the pre-#17828 remote-add shape or the
literal (non-wildcard-suffixed) fetch refspec from before the
deleted-upstream-branch softening commit, so CI went red on that HEAD:

- worktree-push-target-refspec-real-git.test.ts: the migration fixture
  asserted a hardcoded tracked-ref count before narrowing. Under git
  >= 2.44, `followRemoteHEAD` auto-creates a `refs/remotes/<name>/HEAD`
  symref on the first fetch matching the full wildcard refspec, adding
  one untracked ref. Made the count/assertions robust to that ref's
  presence instead of hand-tuning the constant per git version.
- worktrees-wsl-runtime-routing.test.ts: assertions predated both the
  `-t <branch> --no-tags` mint change and the wildcard-suffix refspec
  change; updated to the full, correct call sequence and confirmed the
  WSL routing options (cwd, wslDistro) are threaded to every call.
- worktrees-create-metadata-persistence.test.ts and
  orca-runtime-tests/worktree-removal-and-reconciliation.spec.ts: same
  class of staleness, found via CI job log cross-referencing rather
  than being explicitly flagged.

Verified out of scope: the SSH fork-remote mint path
(prepareWorktreePushTargetSsh) is untouched by this PR -- it never
persists a `remote.<name>.fetch` refspec at all, using
provider.fetchRemoteTrackingRef for a targeted per-branch fetch
instead -- so worktrees-ssh-fork-push-target-remote.test.ts needed no
change.

* fix(git): migrate pr-* remotes with zero worktree-metadata trace too

The migration sweep's candidate discovery was purely metadata-driven
(store.getAllWorktreeMeta()), so a pr-* remote whose every referencing
worktree was removed outside preserve-on-delete (metadata purged, not
just the worktree) was permanently invisible to it and stayed on the
wide default forever.

Field data from a manual migration run against a real user's repo (31
pr-* remotes, 34,637 tracking refs, only 18 actually needed) found
exactly this: 15 of 31 remotes had no branch pinning them at all.

Widen discovery to every pr-* remote git reports on disk, in addition
to metadata-derived candidates. For a remote with no branch provenance
from either metadata or surviving branch.*.remote/.pushRemote config,
there's nothing to narrow *to* -- clear its fetch refspec entirely
instead (stays pushable, imports nothing on a plain fetch), gated on
it still carrying the untouched stock wide default so a user's own
custom pr-*-named remote isn't touched. Removing the remote outright
stays #17842's job.

Adds clearForkRemoteFetchRefspec (fork-remote-refspec.ts), 3 new
mocked-exec tests, and a real-git integration test proving a
subsequent plain `git fetch` on the cleared remote imports nothing.
2026-09-01 03:45:09 -07:00
Neil 84584b61d0 perf(git): cache sparse-checkout annotation on worktree listing (#17859)
* perf(git): cache sparse-checkout annotation on worktree listing

`git worktree list` never reports sparse-checkout state, so every listing paid a
per-worktree fs.stat + config read to detect it -- measured at ~9x the cost of
the `git worktree list` call it decorates on a 1000-worktree repo. Cache the
result per worktree path, invalidated by the existing worktree-change
invalidator registry plus explicit remove/move hooks, with a 5-minute
reconcile window bounding the one unwitnessed edge case (external
`git sparse-checkout` toggle with extensions.worktreeConfig off), matching the
precedent already accepted in readRepoWorktreeAdminFingerprint.

* perf(git): normalize/scope sparse-checkout cache keys, add SWR

Address independent-review follow-ups on the sparse-checkout annotation
cache (#17859):

- Extract canonicalWorktreePath() from areWorktreePathsEqual and key/invalidate
  the cache through it on both read and write, closing the disclosed
  path-spelling P2 outright instead of leaving it as a residual risk.
- Scope cache entries and clears by repo path (derived from the invalidator
  registry's repoId via a store lookup, falling back to a full clear when the
  repo can't be resolved), so churn in one repo no longer evicts a sibling
  repo's warm cache.
- Replace the hard 5-minute cutoff with stale-while-revalidate: past the
  window, callers get the cached value immediately while a deduplicated
  background probe corrects it and, on a flip, drives the existing
  worktrees-changed notification -- collapsing visible staleness from the
  full window to one refresh cycle at zero added listing latency.

Also corrects a stale claim in the original PR description: newer Git does
emit a `sparse` porcelain line (which annotateSparseCheckoutStatus already
skips), but Orca's Git 2.25 compatibility baseline predates it, so the
fallback detection this caches remains necessary.

* fix(git): stop background sparse-checkout revalidation resurrecting invalidated entries

Readiness-loop finding: a stale-while-revalidate probe in flight when a
worktree is removed/moved (or a repo's cache is cleared) would still write
its result back afterward, resurrecting an entry that was deliberately
dropped. Guard the write with a presence check so an invalidated key stays
absent until the next real read.

* fix(git): identity-check the sparse-checkout SWR write-back guard

The has()/presence guard from the previous commit only proved some
entry existed at the key, not that it was the one this revalidation
started from. A worktree removed and re-created at the same path while
a background re-detect was in flight would repopulate the key with a
fresh cold read, and the stale in-flight result would then overwrite
it -- exactly the race greptile (P1) and pullfrog both flagged as
still open. Compare the map's current entry by reference to the entry
captured when the revalidation began; a mismatch means something else
(invalidate, clear, or a fresh cold read) replaced it, and the stale
result must not be written back.

Added a regression test that fails against the old has() guard and
passes with the identity check: invalidate and repopulate the key with
a different value mid-flight, then let the stale revalidation settle
and assert the fresh value survives.
2026-09-01 03:44:50 -07:00
Neil 9542b45d99 fix(wsl): resolve conflict and working-tree probes in the host path namespace (#17895)
Git running inside a WSL distro writes `.git` gitdir pointers, and answers
`status --porcelain`, in the guest namespace. Node reads both back in the
Windows main process, where `/mnt/c/repo/.git` resolves to `C:\mnt\c\repo\.git`
and `/home/me/wt` names nothing at all. Four fs probes were built on those
fabricated paths and always came back "absent":

- `detectConflictOperation`'s four marker probes, so merge/rebase/cherry-pick
  badges silently went missing.
- `parseUnmergedEntry`'s compat existence check, so every `deleted_by_us` /
  `added_by_them` conflict rendered as 'deleted' regardless of the working tree.
- `findExistingWorktreeSymlinkPaths`' `lstat` from status, so Orca's own shared
  symlinks (node_modules and friends) showed as user changes.
- the same `lstat` from the hosted-review dirty preflight, which fails closed:
  an unreadable shared symlink read as uncommitted work and blocked PR/MR
  creation outright.

`resolveGitDir` computes the host spelling of the worktree once and uses it for
both the gitfile read and the pointer resolve, so a guest-spelled worktree path
is reached at all, and a relative pointer (`worktree.useRelativePaths`, git
2.48+) resolves against a spelling Win32 understands. The pointer itself now
goes through the already-landed `resolveGitMetadataPath`, and the function gains
an optional `{ wslDistro }` for a caller whose base path does not encode a
distro. `detectConflictOperation` forwards it, and the three callers that reach
it -- status-read, the runtime RPC, the `git:conflictOperation` IPC -- pass the
git options they already hold. The return type stays `Promise<string>`.

`resolveWorktreeHostPath` is the same rule applied to a worktree path, used by
status-read for the two working-tree probes and by the review preflight. Both it
and `resolveGitMetadataPath` now treat only a single-leading-slash path as guest
namespace: `//wsl.localhost/...` is already a host UNC spelling, and translating
it prepended a second share prefix.

`readWorktreeDiffStamp` needed the same one-namespace guarantee, since moving
translation inside `resolveGitDir` would otherwise make its HEAD and index real
while the working-tree stat stayed fabricated, letting a settled diff survive
every edit. #17896 landed that change first, so it is no longer in this diff;
its version is a superset and all four components already resolve from one
`hostWorktreePath`. What remains here is the `resolveGitDir` gitfile-pointer
fix that #17896 explicitly deferred, which `worktree-diff-stamp-host-paths.test.ts`
pins.

`getConflictCompatibilityStatus` moves from `existsSync` to async `access`, for
the same reason `detectConflictOperation` did: once these paths are real they
are `\\wsl.localhost\...` shares, and a sync probe per asymmetric conflict
blocks the Electron main thread for a 9p round trip on every status poll.

Per-platform delta:
- native Windows, no WSL: no behavioral change. Nothing here starts with a
  single `/`, so no path is translated. An absolute pointer is now returned
  verbatim rather than separator-normalized; every consumer re-joins or
  normalizes it before use.
- macOS/Linux: no change. Guest-pointer translation is gated to win32, and a
  caller-named distro is ignored off Windows.
- Windows + WSL: drvfs pointers and drvfs-spelled worktrees now resolve to their
  drive spelling instead of `C:\mnt\...`; a non-drvfs guest path resolves
  through the named distro's UNC share, or stays verbatim (ENOENT -> existing
  fail-safe) when none is named.
- SSH/relay: none. Those paths return before any of this via the provider
  branch; `src/relay/git-handler-status-ops.ts` keeps its own resolveGitDir.
- folder workspaces, GitLab: none. Neither is on these code paths.
2026-09-01 03:17:34 -07:00
Jinjing 899304d515 Increase artifact content size limit from 5 MiB to 10 MiB (#17910)
Doubles the maximum UTF-8 bytes accepted for manually shared artifacts,
enabling users to share larger content while maintaining recovery and
transport constraints.
2026-09-01 03:16:34 -07:00
Jinjing 2b6c14d4b5 Add startup delivery diagnostics and success announcements (#17814)
Terminal sessions now report startup command delivery details (whether written, presence, length, and delivery method) without logging the command text—preventing credential leakage and distinguishing missing commands from lost ones in diagnostics.

Setup scripts now announce completion on both POSIX and Windows before executing the startup command, so healthy setups don't appear stuck in the UI with "Waiting for setup..." as the last visible line.

Diagnostics failures are caught and ignored so they never break session creation.
2026-09-01 03:08:25 -07:00
Neil 1603810dde perf(worktree): make head-identity refresh incremental (#17843)
* perf(worktree): make head-identity refresh incremental

Head-identity refresh re-read `gitdir` + `HEAD` + a loose ref for every
linked worktree on every watcher burst. On a 973-worktree checkout that is
~2,800 metadata reads (~1.0s of main-process fs I/O) per event, and the
debounced pipeline fires on every commit in any worktree — so fleet-wide
agent activity degenerated into a continuous scan loop.

Watcher events already name the admin dir that changed. Classify each event
into a head-identity scope, memoize per-entry identities, and re-read only
the scoped entries. Refs resolved during a pass are replayed onto cached
entries that share the same branch, so `git worktree add --force` siblings
stay current without extra reads.

Invalidation stays conservative: an absent scope (watcher failure, event
overflow, cold start) means a full re-read, `packed-refs` writes invalidate
every entry, misses are never memoized, and one refresh per minute is
promoted back to a full re-read to bound the window where a ref moves with
no event under any admin dir.

Measured on the reported 989-entry checkout (macOS/APFS): one-worktree
commit 2,816 -> 2 file reads, 61ms -> 0.5ms p50 with an identical page
cache; a 20-worktree debounce burst costs 57 reads / 9.7ms; an external
`git worktree add`/`remove` costs one readdir / 1.0ms.

Refs #17828

* fix(worktree): harden incremental head-identity invalidation

Two holes found in self-review:

- An admin entry name removed and immediately reused inside one debounce
  window coalesced into a listing-only scope, so the reused entry kept
  serving the removed worktree's cached head. Name the entry alongside the
  listing on every `worktrees/<name>` create/delete.
- A non-ENOENT `readdir` failure on `worktrees/` collapsed the memo to the
  primary row, which then re-emitted every identity on recovery. Mirror
  worktree-git-common-polling: only a genuinely absent dir means empty; any
  other error keeps the previous listing.

* fix(worktree): let empty-scope bursts still take the head re-baseline

Adversarial review found the 60s full-rebaseline promotion was unreachable
whenever the triggering burst had an empty head-identity scope: the skip
guarded on the raw caller scope and returned before `resolveScope` ran, so
`lastFullReadAtMs` was never re-evaluated. A repo whose only churn is
`git worktree lock`/`unlock` or a sparse toggle — Orca's own prepared-checkout
flow locks and unlocks on every create — could starve the promotion forever
and hold a stale head indefinitely. Resolve the scope first and skip on the
effective scope.

Also stop deferring an add/remove that arrived while the `worktrees/` listing
was transiently unreadable: forget the memoized listing so the next refresh
re-enumerates whatever its scope, instead of waiting for another listing event.

Both fixes carry a test verified to fail without them.

* fix(worktree): return head-read completeness instead of sniffing the memo

Adversarial review round two. Six fixes, each with a test verified to fail
without it.

- `readGitCommonHeadIdentities` now returns `{ identities, listingComplete }`.
  The refresh layer was inferring "enumeration failed" from `cache.entryNames
  === null`, a reader-owned field whose null also means "cold start" — fragile
  in production and impossible to express in a mock.
- A read discarded by teardown, or one that could not enumerate `worktrees/`,
  no longer arms the 60s freshness clock.
- A queued refresh whose re-run met a destroyed window (macOS recreates the
  window while the watch lives on) was cleared and dropped. It now stays armed
  and is folded into the next request.
- An incomplete listing carries forward the baseline rows it could not observe,
  so recovery does not report every linked worktree as changed.
- The baseline advances after notifying, so a send into destroyed chrome leaves
  the move to be retried instead of diffing it away.
- A scope naming an entry the memoized listing does not know now forces a
  re-enumeration instead of resolving to zero work — this removes an unstated
  dependency on `diffGitCommon` emitting a dir-level create for new entries.
- Overflow states FULL at its construction site rather than relying on a
  downstream `?? FULL` for an absent field.

Also documents the load-bearing invariant behind the empty-scope skip (an empty
scope only reaches the refresh from a structural burst, which forces
`emit: false` and is always paired with a catalog notification for every repo
on the watch), and strengthens two tests that could not distinguish the
behaviour they claimed.

* fix(worktree): bound head-identity staleness with a one-shot catch-up

The previous re-baseline was opportunistic: it rode the next refresh, so a ref
that moves with no watched write (`git update-ref refs/heads/x` from a sibling
worktree) stayed stale until an event happened to arrive after the interval.
Pre-PR the very next event anywhere in the repo corrected it, so this was a
real narrowing of correctness, not just a pre-existing gap.

Arm a one-shot, unref'd timer when a SCOPED pass completes, firing one full
re-baseline an interval after the last full read, then disarming. A full pass
disarms instead of arming, so it never becomes a background poll, and the timer
only exists after an event — an idle repo still schedules nothing and reads
nothing. Cost is O(1) timer per active repo and at most one full read per
interval: the same operation the old code ran per event, 60x rarer.

This also converts "stale until some later event" into "stale at most one
interval, period", which is what bounds the blast radius of any invalidation
bug in the scoping itself.

Cleared on watch disposal. Three tests, each verified to fail without its fix:
the catch-up runs with no further events; a quiet repo issues no background
reads and the timer disarms after firing; disposal stops it.

* fix(worktree): treat an unreadable head as unknown, not absent

Reported independently by two PR reviewers. `readTrimmedFile` collapsed every
errno to `null`, so an EIO/EACCES/ENFILE on a `gitdir`, `HEAD`, loose ref, or
`packed-refs` read was indistinguishable from the file being absent — and the
caller deletes the cached identity on `null`. Same conflation AGENTS.md forbids
for the SSH verdict vocabulary: loss of contact is not evidence of absence.

Reads now report three outcomes, and an unknown:

- keeps the entry's last verified identity instead of evicting it,
- is never replayed onto siblings sharing the branch as "this ref is gone",
- marks the entry unverified so the very next pass re-reads it whatever its
  scope, and
- reports the pass incomplete, so it cannot arm the freshness clock.

The reviewers' stated consequence — that an evicted entry stays evicted until
the next full pass — did not hold, because `!cache.entries.has(name)` already
forced a re-read. The real cost was that one EMFILE evicted every entry it
touched and the next pass re-read all of them, which is exactly the full scan
this PR exists to remove, plus a spurious re-publish of every row.

Renames `listingComplete` to `complete`: it now covers entry reads too.
2026-09-01 02:53:05 -07:00
NeilandNeil dff2ff0ec3 fix(git): read the diff working tree and stamp through the host path spelling (#17896)
Git can execute inside a WSL distro against a raw Linux worktree path while Node,
on the Windows side, reads the same files back through Win32. `path.join(
'/home/me/repo/feature', 'src/file.ts')` on win32 produces the drive-relative
`\home\me\repo\feature\src\file.ts`, which resolves against whatever the current
drive happens to be and almost always ENOENTs. The same mis-spelling hits the
drvfs form, where `/mnt/c/repo` should read as `C:\repo`.

Two consequences, both on the Node side only (git already works, because it gets
the Linux path as its cwd and resolves it inside the distro):

- getDiff's unstaged working-tree read missed, `readWorkingTreeFile` mapped ENOENT
  to `exists: false`, and an existing file rendered as DELETED in the diff view.
- `readWorktreeDiffStamp` could not find `.git`, so the stamp was null, the settled
  diff cache neither hit nor stored, and every diff respawned `git show` - two
  `wsl.exe` spawns the cache exists specifically to avoid.

Both now spell the worktree directory for the reading host first, via a new
`resolveWorktreeHostPath` wrapper around the resolver that landed in #17804.
The wrapper exists because `resolveGitMetadataPath` trims: a gitfile payload
carries a trailing newline, but a directory name may legally begin or end with
whitespace on POSIX, so the wrapper keeps the caller's spelling whenever the
resolver only trimmed it. The stamp's opaque `value` still embeds the caller's
original `worktreePath`, so settled-cache identity is byte-identical and no cache
key moves.

`readWorktreeDiffStamp` was already `Promise<WorktreeDiffStamp | null>` with one
caller that treats null as a cache miss, so no new nullability enters the type
system and the resolver's never-null-for-a-non-empty-pointer contract is
untouched. The only unspellable input is an empty worktree path, handled locally
as "not provably unchanged" in the stamp and as a read *failure* (not a proven
deletion) in file-diff.

What changes for users

| Platform | Delta |
|---|---|
| macOS | No change. An absolute POSIX path is returned verbatim, including one whose directory name carries leading or trailing whitespace. |
| Linux | No change. Same reason. |
| Native Windows (no WSL) | No change. A `C:\...` or `\\server\share\...` path is already absolute for win32 and passes through verbatim. |
| Windows + WSL, UNC worktree path (`\\wsl.localhost\Ubuntu\...`) | No change. Already absolute for win32; passes through verbatim. This is today's common case. |
| Windows + WSL, drvfs worktree path (`/mnt/c/repo`) | Fixed. Reads as `C:\repo` instead of the drive-relative `\mnt\c\repo`. Needs no distro name. |
| Windows + WSL, Linux worktree path with a named distro (`/home/me/repo`) | Fixed. Reads as `\\wsl.localhost\Ubuntu\home\me\repo`. The deleted-file misrender goes away and the diff cache starts hitting. |
| Windows, POSIX path, no distro and not a drvfs mount | No change. Passes through verbatim, same ENOENT, same existing fallback. |
| SSH | No change. `runtime-git-diff-commands.ts` and the `git:diff` IPC both route to `provider.getDiff` for a connection, so this local code is never reached. |
| Relay / remote | No change. No RPC param, wire field, stream opcode, or published content is touched; the relay host runs the same local code and gets the same fix. |
| Folder workspace (non-git) | No change. `.git` is absent either way, `resolveGitDir` returns the same fallback, and the stamp stays null exactly as today. |
| GitLab / other providers | Not applicable. No provider-specific or review code is touched. |

What this does NOT do

- It does not fix `resolveGitDir` itself. For a drvfs repo whose worktree Orca
  already spells `C:\repo\feature`, the gitfile payload `gitdir: /mnt/c/repo/.git/
  worktrees/feature` is still mis-resolved by `path.resolve` to
  `C:\mnt\c\repo\.git\...`, so the stamp still returns null in that shape. Separate
  change, separate PR; this one neither fixes nor regresses it.
- It does not touch submodule path resolution. `resolveSubmoduleWorktreePath` is
  the path-escape guard and has a near-identical twin in the relay; changing it
  without escape tests on both is out of scope.
- It does not change `readHeadComponent`'s `commondir` resolution. The relative
  `../..` git actually writes takes the identical `path.resolve` branch, and an
  absolute POSIX `commondir` under a WSL UNC `gitDir` already resolves correctly
  because the UNC root is `\\wsl.localhost\<distro>\`.
- It does not reorder drvfs-before-UNC inside the shared resolver. That changes the
  identity of returned strings and needs a real Windows+WSL box.
- It does not add any Git command, option, or version dependency.

Costs and residual risk

- One extra pure function call per diff read. No I/O added or removed on the
  unaffected paths.
- Translation still trims. `resolveWorktreeHostPath` preserves whitespace only when
  no translation happened; a guest directory named `/home/me/repo ` loses its
  trailing space on a Windows reader. Reachable only on win32, where such a name is
  not addressable anyway, and the previous behavior for that shape was a
  drive-relative miss.
- A relative worktree path (no caller passes one) is now resolved against the
  process cwd instead of joined relative to it. Same file in every case except a
  relative name that itself ends in whitespace.
- `UNSPELLABLE_WORKING_TREE_READ`'s `exists`/`failed` fields are correct but not
  observable today: the stamp is null for the same input, so nothing can be cached
  and `reusable` cannot be read back. They are there so the branch stays right if
  `loadDiff` ever gains a second caller. The test pins the observable part - that no
  read lands on a cwd-relative path.
- Every test here mocks `node:fs/promises` and spoofs `process.platform`. They prove
  which path string reaches `stat`/`readFile`, which is the right assertion, but
  none of this has executed against a real 9p mount on a Windows+WSL box and this
  repo's CI has no such runner.
- Honest framing of the trigger: I could not demonstrate a mainline path that hands
  `getDiff` an untranslated POSIX worktree path on Windows today -
  `translateWslOutputPaths` UNC-translates worktree paths whenever a distro is
  known, `getWslHome` returns the UNC spelling, and `resolveWslRepoWorktreeBasePath`
  normalizes a configured Linux base. The drvfs case is the most plausible live one.
  Treat this as defense-in-depth that is a strict no-op on every configuration above
  except the two marked Fixed.

Verification

- `npx vitest run src/main/git src/shared/git-metadata-path.test.ts` -> 196 files /
  2241 tests passed, 2 files and 5 tests skipped. One failure,
  `git-admission-storm-measurement.test.ts > reports bounded-concurrency before and
  after measurements` (ENOENT scandir on its own temp state dir), is pre-existing
  and environmental: it fails identically in isolation and spawns real git children
  without touching any changed module.
- `npx vitest run src/main/git/status-diff-settled-cache.test.ts` -> 21/21 (16
  pre-existing, 5 new). `npx vitest run src/shared/git-metadata-path.test.ts` ->
  25/25 (19 pre-existing, 6 new cases across 3 new tests).
- `npx oxfmt --write` then `npx oxlint` on all five changed files -> clean.

Mutation checks - all eight production substitutions were reverted one at a time
and the suite re-run. Each fails at least one test, and no new test survives its
own mutation:

| Reverted | Failing test |
|---|---|
| file-diff working-tree read -> `worktreePath` | reads the working tree through the host spelling instead of reporting a deletion; invalidates when the working tree file is edited under the host spelling |
| stamp working-tree component -> `worktreePath` | invalidates when the working tree file is edited under the host spelling |
| stamp `.gitmodules` stat -> `worktreePath` | invalidates when .gitmodules appears under the host spelling |
| stamp `resolveGitDir` -> `worktreePath` | stamps through the host spelling so the second read does not respawn git |
| `options` threading at the `readWorktreeDiffStamp` call | stamps through the host spelling...; invalidates when .gitmodules appears... |
| wrapper's untrimmed preservation -> return the resolver's value | keeps whitespace that belongs to the directory name (both cases) |
| `UNSPELLABLE_WORKING_TREE_READ` -> a cwd-relative `readWorkingTreeFile` | reads nothing relative to the cwd when the worktree path has no host spelling |
| stamp's null early return -> `hostWorktreePath ?? worktreePath` | reads nothing relative to the cwd when the worktree path has no host spelling |

The settled-cache tests seed the fake filesystem through the platform-bound `path`
module rather than `path.win32`, so they assert real behavior on a POSIX CI host as
well as on Windows and are not gated on the host platform.

Co-authored-by: Neil <79079362+brennanb2025@users.noreply.github.com>
2026-09-01 02:39:48 -07:00
NeilandNeil e4a708da17 fix(wsl): read untracked line counts through the host's worktree spelling (#17897)
Untracked line counts, and the untracked share of the branch line total, come
from direct lstat/open calls rather than from git. When git executes inside a
WSL distro the worktree path can be a guest path, which Win32 reads as
drive-relative (`/home/me/repo`) or as a literal `C:\mnt\c\repo`. Every lstat
then fails, countFileAdditions swallows the error into `{}`, the file renders
with no +N, and the branch total silently undercounts it.

Route only those two filesystem reads through resolveWorktreeFilesystemPath,
which translates a guest-rooted path via the shared resolveGitMetadataPath.
Both call sites produce the same string, so the stat-keyed untracked cache
still hits across them. resolveGitMetadataPath itself is not modified, so its
other callers are untouched.

The wrapper translates only when the host is win32 and the path is a guest
spelling: exactly one leading slash, and unchanged by trim(). Each condition
is load-bearing.

- `//wsl.localhost/...` and `//wsl$/...` are UNC spellings that also start with
  `/`, and translating one prepends the distro root a second time, so a
  worktree that reads fine today would ENOENT on every untracked file. The same
  single-leading-slash guard is used, for the same reason, by
  resolveWslRepoWorktreeBasePath in src/shared/wsl-paths.ts.
- resolveGitMetadataPath returns `rawPath.trim()`, so a worktree directory name
  with leading or trailing whitespace (legal on ext4) would be re-spelt onto a
  different directory. Leaving those verbatim keeps them exactly as they read
  today.
- Off win32 the resolver is already an identity for these inputs; the platform
  check makes the macOS/Linux no-op structural rather than derived.

Per platform:
- Windows + WSL, worktree spelled as a guest path: untracked +N appears and the
  branch total includes it.
- Windows + WSL, worktree already spelled `\\wsl.localhost\...`,
  `//wsl.localhost/...` or `//wsl$/...`: unchanged, returned verbatim.
- Native Windows: `C:\...` is unchanged. One shape does change: a
  `/mnt/<drive>/...` worktree now resolves to `<Drive>:\...` instead of being
  passed through to a guaranteed lstat failure. Native Windows does not produce
  that spelling, and if it were reached the new result is the correct file.
- macOS / Linux: unchanged; not win32, returned verbatim, whitespace included.
- SSH / relay: unchanged. Remote status runs in the relay, which builds its own
  branch-total input and does not pass filesystemWorktreePath.
- Folder workspaces / GitLab: unaffected, no workspace-kind or provider
  behavior is touched.

No fail-closed degradation: attachLineStats still returns
`stagedStats !== null && unstagedStats !== null`, and createBranchLineTotalInput
still has no early return. The wrapper returns `string`, never null, so an
unmappable worktree keeps its old spelling and its old (missing) untracked
counts rather than dropping the staged/unstaged counts as well.

The `filesystemWorktreePath` field on computeGitBranchLineTotal is optional and
does not touch the lease key, so the coalescing/cooldown identity is unchanged.

Co-authored-by: Neil <neil@orca.local>
2026-09-01 02:39:31 -07:00
Neil f2db24bca0 fix(worktree): reclaim a prepared checkout whose discard failed in-process (#17899)
Speculative create-preparation evicts entries past the 3-entry limit and the
5-minute TTL, and both paths swallowed a failed `discardPreparedWorktree`.
The only other reclaim path, `cleanupStalePreparations`, skips any preparation
whose lock-reason pid is still alive, so a discard that failed inside the
running app stranded its scratch checkout and its locked worktree registration
until restart.

Record the failed discard keyed by host (repo path + WSL distro) and prepared
path, and retry it the next time a fresh preparation starts for that host.
The retry is kicked off before `listWorktreeGraph` but never awaited, so it
runs alongside the stale scan and the `worktree add` instead of sitting in
front of the user's create. It is capped at 3 attempts and warns when it gives
up. Enrolment is unconditional: `prepareWorktreeCreateCheckout` self-discards
on failure, but only best-effort, so a checkout that failed on a busy handle
can strand the same registration.
2026-09-01 02:39:03 -07:00
Neil c31e7b0d9a docs(main): restore rationale comments lost in the startup and cookie splits 2026-09-01 02:07:33 -07:00
Neil 15c4f1c447 fix(runtime): defer process incarnation lookup for plain sends
(cherry picked from commit d522a3e15a56c0d6d1b92b67414d51f6bf78e987)
2026-09-01 02:07:33 -07:00
Neil 83513b82d1 test(terminal): follow Hangul lifecycle extraction
(cherry picked from commit b62361fad271809819aa01711e437197065e3f75)
2026-09-01 02:07:33 -07:00
Neil d5db0bedfc fix(main): route browser cookie key commands through runner
(cherry picked from commit e321eccb41)
2026-09-01 02:07:33 -07:00
Neil de84a9f999 fix(main): resolve runtime lazily in observers
(cherry picked from commit 369496b27c)
2026-09-01 02:07:33 -07:00
Neil d462766cb0 refactor(main): split filesystem git remote handlers
(cherry picked from commit 2146cff06a)
2026-09-01 02:07:33 -07:00
Neil fbaa125a35 fix(main): merge split startup imports
(cherry picked from commit c0f38db5de)
2026-09-01 02:07:33 -07:00
Neil 5b4e7edb50 refactor(main): split backend services and startup
(cherry picked from commit a33573328b)
2026-09-01 02:07:33 -07:00
Neil 7e8337b155 test(preload): census split GitHub bridge owners
(cherry picked from commit cd08e4e91c)
2026-09-01 00:57:54 -07:00
Brennan BensonandMerge Sim d4db524ba7 fix(native-chat): stop unjournaled provider frames from killing the session (#17813)
A frame the classifier declines (status-chrome, suppressed-benign, stream-into-item)
is deliberately not journaled. #17720 turned that null translation into
`{accepted: false, reason: 'untranslated'}`, which is not `backpressure`, so the
notification retry queue treated it as unreplayable and escalated through fail() ->
forceCloseUnexpected -> connection.close(). The app-server latched `closing` and the
create path's next model/list rejected with "codex app-server connection is closed
(model/list)".

The provider emits `remoteControl/status/changed` right after initialize, so every
structured Codex session died on its first chrome frame. Admit the null translation
instead, before any bookkeeping or publish.

Also restore the error-frame exemption from the generic row cap, dropped by the same
PR: the cap now runs after the classification check, so a noisy turn can no longer
reduce provider errors to a suppression count. The test that pinned the capped
behavior is inverted to assert the exemption.

Co-authored-by: Merge Sim <sim@local>
2026-08-31 23:31:27 -07:00
Brennan BensonandMerge Sim 51bc2ec343 fix(native-chat): keep the attachments on a Claude turn that pasted images (#17801)
* fix(native-chat): keep the attachments on a Claude turn that pasted images

A Claude turn carrying pasted images reached native chat with no images at all —
no thumbnails on mobile, and not even an attachment chip on desktop. Nothing
showed that the message had any.

Both carriers were being dropped:

- Claude records the paths in a companion turn marked `isMeta`, holding one
  `[Image: source: <path>]` text block per image. The decoder treats an `isMeta`
  user row as injected, filters it down to tool-result blocks, and returns null
  when none remain — so the whole row went away.
- The prompt row's own `image` blocks are `{source: {type: 'base64'}}`, which
  carry no url or path, so `imageRefBlock` drops them too.

With the companion gone, `isImageSourceUserTurn` could never fire and the fold in
`normalizeImageTranscriptMessages` was unreachable on the Claude path.

Surveying every transcript under `~/.claude/projects`: 238 of 241 image-source
rows are `isMeta`, across every versioned release (2.1.220 through 2.1.237); the
3 that are not carry no version field at all. 38 of those rows hold more than one
content block, which also defeated the single-block rule in
`isImageSourceUserTurn`.

Let image-source text survive the injected-turn filter, and recognize a turn
whose blocks are *all* markers rather than only a lone one. An ordinary injected
turn (a skill preamble, a compact summary) is still dropped, and a turn that
mixes prose with a marker is still not an image-source turn.

Carrying the paths keeps the payload small; decoding the base64 instead would put
hundreds of KB per image on the wire to mobile.

* fix(native-chat): preserve image companion ordering

* fix(native-chat): keep image companions turn-local

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-31 23:27:56 -07:00
Brennan BensonandMerge Sim a9e6fb7eff fix(native-chat): stop rendering tool output as the agent's streaming reply (#17782)
* fix(native-chat): stop rendering tool output as the agent's streaming reply

A tool result could appear in native chat as a raw, un-collapsed "assistant"
bubble that never went away for the rest of the turn — on mobile it showed up
as a wall of a source file's contents, prefixed by "Exit code 1".

Providers publish a tool's stdout/error as `lastAssistantMessage` so status
cards and dashboard rows can preview what the agent just did. Native chat reuses
that same field as its live streaming bubble, so the preview rendered as prose.
For Claude the preview is *only ever* tool output mid-turn: claude-tool-fields
writes real prose exclusively at Stop, so the bubble could never contain an
actual streaming reply.

It also could not be retired. The bubble hides once a transcript assistant block
leads with the streamed text, and tool output never lands in one — so the only
remaining exit was the turn ending, which is why a long tool-heavy turn pinned it
on screen.

Carry provenance instead of changing what the status surfaces show: mark the
writes that come from a tool result/error, keep the flag in lockstep with the
value it describes through the listener merge, and have both native-chat
streaming paths ignore a flagged preview. Status cards, dashboard rows and
automation capture are untouched.

The wire field is optional, so an older host that never sends it keeps today's
behavior rather than silently suppressing previews.

* fix(native-chat): preserve tool output provenance through renderer sync

* fix(native-chat): retain preview provenance in Claude roster state

* test(native-chat): cover restored tool preview provenance

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-31 23:05:31 -07:00
Neil 9bc564c2aa fix(cleanup): follow WSL-written gitdir pointers on a Windows host (#17806)
`readLocalWorktreeGitDir` resolved a linked worktree's `.git` gitfile
pointer by hand, translating a POSIX-rooted pointer only when the
worktree path was itself `\\wsl.localhost\...`. On a Windows host
`path.isAbsolute('/mnt/c/repo/.git/worktrees/wt')` is true, so a
drive-path worktree (`C:\Users\me\wt`) whose gitfile was written by git
running inside WSL kept the guest spelling and was joined to
`\mnt\c\repo\.git\worktrees\wt\HEAD`. All four probes (HEAD,
COMMIT_EDITMSG, ORIG_HEAD, tail of logs/HEAD) missed, so the row's
lastActivityAt fell back to the worktree directory mtime and a worktree
with recent commits could read as stale in the cleanup browser.

Delegate to `resolveGitMetadataPath` (src/shared/git-metadata-path.ts,
landed in 7f63db7d7a, already used by repo-git-marker-scan.ts). Five
lines out, one in; the shared resolver is not modified.

Delta, enumerated over 10 base x 18 pointer x 3 platform combinations
(540 pairs) against a reimplementation of the removed branch: 28 differ,
every one win32 + non-UNC base + `/mnt/<lowercase-letter>` pointer.

- WSL-on-Windows: a drive-path worktree with a WSL-written pointer now
  probes its real git metadata.
- WSL UNC worktrees, native Windows, macOS, Linux: no change (0 deltas
  on darwin/linux, 0 for any `\\wsl.localhost\...` base).
- SSH / relay / folder workspaces: no change; remote repos return the
  persisted timestamp before any probe, and a folder workspace has no
  gitfile pointer.

Not strictly monotone: the old probe target `\mnt\c\...` is a real
drive-relative location, so if it existed with a newer mtime than the
genuine gitdir this lowers lastActivityAt for that row. The persisted
timestamp stays a floor via Math.max, and in every realistic case the
change only raises the value.

No signature changes, no options threading, no new call sites.
2026-08-31 22:37:34 -07:00
Neil d7d3114716 perf(wsl): single-flight the async WSL distro list (#17805)
On Windows, seven production call sites reach listWslDistrosAsync and on a cold
cache each spawned its own `wsl.exe --list --quiet` (5s timeout each): the
wsl:listDistros IPC behind the renderer capability read, the host.wsl.listDistros
RPC, the skill-install IPC, CLI registration reconciliation, the hook relay deps,
the kimi runtime home, plus relay preflight in the relay process. Concurrent
callers in one process now share one spawn.

Joining happens ahead of the negative cache, which also fixes a stranding bug: a
synchronous listWslDistros() landing an empty result mid-probe arms the 15s retry
window, and later async callers read that [] even though the pending probe is
about to see a distro that just finished provisioning. The non-empty-cache
short-circuit sits ahead of the join so a list already found synchronously is
still returned without waiting; that is main's existing behaviour preserved, not
a new fast path.

The shared promise cannot reject -- `catch` sits ahead of the stored promise, so
joiners get the same fail-safe [] the old per-caller catch returned -- and the
slot is cleared on settle, by the owning probe only.

wsl-directory-probe-command.ts is a verbatim move of the guest directory-probe
marker protocol and its parser out of wsl.ts, for oxlint max-lines headroom:
inlining it back makes wsl.ts 306 effective lines against a cap of 300. It takes
WslUncPathInfo from ../shared/wsl-paths -- the actual type of every value passed
at both call sites -- so it does not import from wsl.ts. _resetWslCachesForTests
and _setWslCachesForTests now share one resetWslDistroListState() instead of
repeating the same six assignments.

Per-platform delta:
- WSL on Windows: fewer wsl.exe spawns under startup fan-out, and a distro
  provisioned while a probe is pending is no longer hidden for the retry window.
- Native Windows without WSL: no behavioural change. The empty/failure retry
  windows, their backoff and the cache sequence guard are unchanged; N concurrent
  callers now cost one failed spawn instead of N.
- macOS, Linux, folder workspaces: no change. Both new early returns are
  unreachable off win32.
- SSH remote: no change for macOS/Linux hosts; a remote Windows host gets the
  Windows behaviour in its own process. No wire change -- host.wsl.listDistros
  keeps its string[] shape and its [] failure value.
- Relay: same single-flight inside the relay process. It stays per-process; the
  relay and main process still probe independently, as before.

Costs: a never-settling execFileUtf8 now pins the shared slot for the process
lifetime rather than only its own callers -- transient-to-permanent, not identical
exposure. And a joiner inherits the first probe's failure instead of making an
independent attempt.
2026-08-31 22:37:23 -07:00
Neil a9babde9a3 refactor(git): accept a caller-named WSL distro on the metadata path resolver (#17804)
Two small changes to the Git metadata read path. Neither has a user-visible
effect on any platform except for a malformed `.git` gitfile, described below.

1. resolveGitMetadataPath's third parameter becomes an options object
   `{ platform?, wslDistro? }`. A caller that knows which distro wrote a pointer
   can now say so, where previously only a WSL UNC base path could. The distro
   encoded in the base path still outranks the caller's, and translation only
   happens when the reading host is win32, so a caller-named distro cannot make
   a POSIX host fabricate a Windows path. The UNC-base branch is exempt from
   that gate because that spelling only exists on Windows. Main's other
   contracts are verbatim: never null for a non-empty pointer, and a drvfs
   pointer keeps its drive spelling even when a distro is named. Both production
   call sites (repo-git-marker-scan.ts) pass no options, so they are unchanged.

2. The `.git` gitfile marker parse moves into one shared function,
   parseGitdirMarkerPayload: `gitdir:` at the start of the file, payload
   trimmed, empty payload rejected — git's own read_gitfile_gently rule.
   resolve-git-dir.ts and repo-git-marker-scan.ts both call it; the latter had a
   near-identical private copy and is behaviorally identical after the swap
   (verified across twelve marker spellings; the only divergence, a
   whitespace-only payload, already resolved to null one call further down).
   Main's `/^gitdir:\s*(.+)\s*$/m` in resolve-git-dir captured trailing padding
   into the path and honored a `gitdir:` line anywhere in the file.

Per-platform delta: none on macOS, Linux, native Windows, WSL, SSH, relay, or
folder workspaces. The wslDistro option is inert; this change adds no caller.
For a malformed `.git` gitfile, padding is now stripped (strict improvement), a
whitespace-only payload falls back to `<worktree>/.git`, and a `gitdir:` line
that is not the first line is no longer honored — a narrowing, since main could
return a working gitdir there. All four resolveGitDir consumers already degrade
through a catch, so that case reports no sparse state / conflict operation /
diff stamp rather than failing.

Six other hand-rolled `gitdir:` parsers remain, including the relay's SSH copy;
converging them is its own change.
2026-08-31 22:37:09 -07:00
Neil 28373fcea7 fix(windows): repair the install-dir package ACL that blanks the window (#17740)
* fix(windows): repair the install-dir package ACL that blanks the window

An install tree carrying an orphan AppContainer ACE (S-1-15-2-<x>) with no
ALL RESTRICTED APPLICATION PACKAGES grant denies Chromium's LPAC children read
on the shipped modules; they die at init with 0x80000003 and the window stays
blank forever (electron/electron#51761).

- Tighten the probe verdict to require the S-1-15-2-2 grant specifically: an
  ALL APPLICATION PACKAGES (S-1-15-2-1) ACE, the Program Files default, does
  not appear in an LPAC token and cannot satisfy the orphan.
- Drop BUILTIN from the English-locale heuristic (fr-FR/es-ES print it
  verbatim) so a localized icacls is correctly reported as un-name-checkable.
- Add an additive, marker-guarded icacls self-repair: an inheritable root
  grant plus a flagless (RX) /T pass, never /grant:r.
- Route the crash-loop dialog through a testable prompt module that names the
  permission cause, offers Copy Commands without dismissing itself, and keeps
  the graphics-driver hint.

The repair only runs on win32, off serve mode, and only on the exact probe
verdict that reproduced the crash.

* docs(windows): correct the install-tree ACL walk cost model

* fix(windows): keep the install-ACL poison gate at the reproduced shape

An orphan package ACE alongside the Program Files ALL APPLICATION
PACKAGES default launches clean on win32 10.0.26200 / Electron 43.4.1,
so requiring S-1-15-2-2 specifically declared poison on healthy installs
- and this branch acts on that verdict with a tree-wide icacls write and
the crash-recovery dialog's primary cause. hasRestrictedPackageGrant
stays reported for triage; only the verdict reverts.
2026-08-31 21:15:19 -07:00
Neil e6257b6e32 perf(worktree): resolve the WSL workspace root off the main thread when preparing (#17792)
`computeWorkspaceRoot` resolves a WSL repo's mirror root through `getWslHome`,
which is a synchronous `execFileSync('wsl.exe', ...)` with a 5s timeout. Two
worktree preparation paths ran it on the Electron main thread:
`prepareLocalWorktreeRootForRepo` (repo registration, clone completion, repo
update, project host setup, folder->git upgrade) and `prepareWorktreeCreateForRepo`
(the speculative checkout started while the create composer is open). On a stopped
or cold distro that froze every window for up to 5s. Being fire-and-forget did not
help: only 3 of the 16 `prepareLocalWorktreeRootForRepo` call sites are `void`-ed,
the other 13 are awaited inside IPC handlers, and the sync probe blocks the main
thread either way. `prepareLocalWorktreeRootsForRepos` runs the same probe for
every repo from the settings-save handler.

Adopt the existing `computeWorkspaceRootAsync` (now exported) at those two call
sites, and give the two resolvers a shared mirror-distro decision and shared
root-from-home layout so they cannot drift apart.

Also thread the mirror distro into the prepare-side path settings.
`createLocalWorktree` passes `getWorktreeMirrorDistro(store, repo)` and
`prepareWorktreeCreateForRepo` did not, so a `C:\` repo on a WSL project runtime
prepared under `C:\workspaces` while the create click looked under the mirrored
WSL root: the keys never matched and every prepared checkout was discarded, after
paying for a full checkout that sat until the 5min TTL. Pre-existing on main;
included because it is the same line and the same resolver.

Scope of the win, stated precisely: only those two preparation paths stop
blocking. On a reachable distro `getWslHome` caches on success, so before this
change the first repo paid one blocking probe and the rest were cache hits -- the
change makes that one probe non-blocking, it does not remove N probes. Failed
probes are never cached, so on a stopped distro N repos did pay N sequential 5s
blocking probes and now share one in-flight async probe.

Costs: five sync `computeWorkspaceRoot` callers remain (allowed-roots resolution,
the create click in worktree-remote, CLI create, watch targets, worktree trash),
and `getWslHome` reads only `wslHomeCache` -- it cannot join an in-flight async
probe. The guaranteed synchronous cache warm-up therefore becomes a window in
which one of those callers can still block and can spawn a second concurrent
`wsl.exe`. Concretely: opening the create composer and clicking Create within a
few hundred ms on a cold distro now pays the freeze on the click instead of on the
background prep. Separately, the mirror-distro fix makes prepare spawn an async
`wsl.exe` home probe for `C:\` repos on a WSL runtime, where it previously spawned
none.

No race added: `prepareWorktreeCreateForRepo` computes the preparation key and
inserts the registry entry in one synchronous run after the await, so two
concurrent creates still dedupe to a single prepared checkout.

`worktree-create-preparation-wsl-root.test.ts` runs the real resolver through
prepare and then claims the entry with the production consume-side call shape
(including the mirror distro), so a divergence between the two resolvers fails a
test instead of silently discarding every prepared checkout.
2026-08-31 21:11:09 -07:00
Neil 5c831c1846 fix(crash-reporting): record Linux MemAvailable at process-gone (#17733)
Linux process-gone crash reports emitted only systemMemoryFreeMB, which is
/proc/meminfo MemFree — it excludes page cache and other reclaimable memory, so
an OOM-killed renderer could report gigabytes "free" and hide the pressure that
caused the kill. Electron 43 exposes MemAvailable as `available` on Linux, and
getSystemMemoryAtGoneDetails already had the getSystemMemoryInfo() result in
hand, so emit it as systemMemoryAvailableMB through the existing field table.

The field is omitted on macOS/Windows (where Electron does not report it) and
on a non-finite reading, matching every other bucket.
2026-08-31 21:02:19 -07:00
Neil a4cd00ed18 fix(wsl): drop the linked-worktree Git route cache after worktree mutations (#17791)
On Windows with a WSL distro configured, `prepareWslLinkedWorktreeGitRouting`
caches for 30s which Git owns a drive-letter checkout, by reading that checkout's
`.git` marker. `git worktree add/move/remove` rewrites exactly that marker, so
the verdict could stay authoritative for up to 30s after it stopped being true.

- Add `invalidateWslLinkedWorktreeGitRouting(cwd)`: drops the cached route and
  the probe retry backoff for that path and for anything under it (a submodule
  inside the worktree derived its route from the same marker walk). Eight calls
  at six sites: `worktree add`, `worktree move` (both paths), `worktree remove`,
  the prepared checkout's add, the finalize move (both paths), and the
  prepared-worktree discard. Five sites invalidate from a `finally`, because a
  Git failure can still have rewritten the marker; the prepared checkout's add
  invalidates on the success path only, since its failure path runs the discard,
  which has its own `finally`.
- Split the parent-directory marker walk into
  `wsl-linked-worktree-git-route-probe.ts` (the routing module was at the
  `max-lines` ceiling) and have it report whether the walk settled. A `.git`
  file with no `gitdir:` line is a half-written marker mid `worktree add`: it is
  now retried under the existing backoff instead of cached for 30s. The route it
  yields is unchanged (distro); only the number of parent walks changes.

An invalidation only drops cached state; a probe already in flight is left to
finish and cache normally, so a mutation landing mid-probe is no worse off than
main's 30s TTL. The cost is that a failed mutation also drops a still-correct
route: `gitExecFileAsync` and `gitStreamStdout` re-resolve the route after their
own `prepareWslLinkedWorktreeGitRouting`, and `gitSpawn` resolves again after the
git-admission wait, so a command already in flight for that path can take the
empty-cache default and run a host-owned checkout under `wsl.exe git`. It fails
once and self-heals on the next command.

Reachable only on win32 + configured WSL distro + drive-letter cwd; every other
platform and configuration never populates this cache, so the new calls scan two
empty maps.
2026-08-31 20:35:49 -07:00
Neil abc099e4c7 fix(worktree): run the create-base warm-up on the routed git host (#17794)
The speculative warm-up that runs while the create composer is open resolved
refs and fetched with host Git even when the project's runtime is a WSL distro,
while both the checkout preparation it feeds (`prepareWorktreeCreateForRepo`,
which already resolves `{ wslDistro }` itself) and the real create path run
inside the distro.

The concrete cost was a discarded fetch: `getCanonicalFetchKey` namespaces the
runtime's remote-fetch cache `wsl:<distro>` vs `local`, so the warm-up's fetch
landed in a namespace create never looks at, and create fetched again. On a
Windows host with no usable host-side Git the probes also failed outright, so
that cohort got no warm-up at all.

Thread the project's worktree Git options through the prefetch (resolved by a
non-throwing helper, because an optimistic warm-up must not surface a
repair-required runtime as a failure) so every probe and fetch runs where create
runs. `gitOptions` is a required argument, so a caller cannot drop the routing
silently. Host-routed calls keep their original arity, so macOS, Linux,
native-Windows-host projects, SSH repos and folder workspaces are unchanged.

Narrower than it looks: for a repo under \\wsl.localhost\<distro>\... the probes
were already routed by cwd, and for a repo on a Windows drive letter host Git
and WSL Git read the same on-disk repository, so the answers were already
correct there. What those cohorts gain is a fetch create can reuse; what they
pay is that the probes now run inside the distro (over /mnt/c for drive-letter
repos, which also newly arms the linked-worktree routing probe) and the
speculative fetch now shares create's per-remote fetch queue, as it always has
on native platforms.

Also collapse the three byte-equivalent copies of `hasLocalWorktreeBaseRef`
(create, prefetch, remote-repo create) into one in
git/worktree-base-ref-probe.ts, drop the host-only `hasLocalCommitObject` that
caused the routing bug, and add the first routing assertions on the create-path
consumers of the now-shared probe.
2026-08-31 20:32:59 -07:00
NeilandNeil 7f63db7d7a fix(git): resolve WSL drvfs Git metadata pointers on a Windows host (#17790)
When Orca's runtime is a WSL distro but the repo sits on a Windows drive, git
inside the distro writes `/mnt/c/...` into a worktree's `.git` gitfile and its
`commondir`, while Orca reads those files back through Win32.
`repo-git-marker-scan` returned the pointer verbatim, Windows read it as
drive-relative `C:\mnt\c\...`, and the worktree was reported `invalid`.

Move that resolver out of `repo-git-marker-scan` into
`src/shared/git-metadata-path.ts` and give it exactly one new case: on win32, a
drvfs pointer resolved against a base path that is not a WSL UNC path now gets
its drive spelling. Every other base/pointer/platform combination is
byte-identical to the deleted helper, verified differentially across a
base x pointer x platform matrix — macOS, Linux and native Windows are unchanged.

`toWindowsWslDrivePath` is factored out of `toWindowsWslPath` so the drvfs
matcher has one home; `toWindowsWslPath` itself is unchanged for all inputs,
including the line terminators JS `.` excludes (fuzzed 2M inputs, 0 divergences).

This changes the marker scan's verdict only. `resolve-git-dir.ts` and the relay's
own copy still `path.resolve` the same `/mnt/c/...` pointer in the Win32
namespace, so a worktree that is now accepted still degrades quietly in conflict
detection, sparse-checkout detection, the diff stamp and worktree listing. Those
parsers are deliberately untouched here; see the PR description.

Co-authored-by: Neil <neil@example.com>
2026-08-31 20:32:46 -07:00
Neil 8b2d72114b fix(gitlab): stop the native glab known-hosts probe waking an idle WSL distro (#17789)
On Windows with no host `glab.exe`, the cwd-less `glab auth status` known-hosts
probe fell through to `wsl.exe -d <default distro>`. Probe failures are never
cached, so when that WSL leg also fails (glab absent or logged out inside the
distro) every forge detection re-booted the distro; when it succeeds it cached
that distro's auth hosts under the 'native' execution key, which the comment two
lines above the call already forbids. gitlab-auth-and-rate-limit.ts already
passes allowDefaultWslFallback: false for this exact command; the known-hosts
probe now agrees with it.

Connection-keyed probes keep the fallback: glab has no SSH/relay dispatch, so the
`glab api` calls this gates run the same local CLI with no cwd and would
otherwise disagree with the probe. wsl:<distro>-keyed probes were already
unreachable by the fallback, so passing the flag there is inert.

Counted child_process.execFile calls over 3 sequential probes, process.platform
forced to 'win32', host glab mocked ENOENT (wsl.exe / glab.exe spawns):
  native key, glab absent in the distro too: 3/3 -> 0/3
  native key, glab logged in in the distro:  1/1 -> 0/3, and that distro's
    self-hosted hosts stop reaching the native known-hosts list
  connection key:                            1/1 -> 1/1, unchanged

Residual risk on that second config: a repo on a \\wsl$\ UNC path can be keyed
'native' (no project runtime match) while its own glab calls still route into
WSL by cwd, so it loses the seeded host and must re-derive it through
`glab auth status --hostname`. A co-resident Windows-path repo on the same host
can now write a shared `native\0<host>` unauthenticated negative that stalls that
recovery for one NEGATIVE_ENTRY_TTL_MS window. Fixing that properly means keying
the cache by the host that actually served the call, which needs an exec-layer
API change and is deliberately out of scope here.

Also splits the getGlabKnownHosts suite out of gl-utils.test.ts (796 counted
lines against the 800-line cap for tests) into gitlab-known-host-probe.test.ts.
2026-08-31 20:32:35 -07:00
NeilandNeil ad760c8b92 fix(worktree): reject Windows drive-qualified shared paths in the symlink guard (#17793)
getSafeRelativePath strips leading `/` and `\` before testing absoluteness, so
the only rooted spelling that can still reach the guard is a Windows drive
designator. It tested that with the host `path.isAbsolute`, which left two gaps:
the drive-absolute form `C:/payload` was refused on Windows but admitted as an
ordinary relative filename on macOS/Linux, and the drive-RELATIVE form
`C:payload` was admitted on every host including Windows, where
`win32.resolve(root, 'C:payload')` discards the worktree root and lands under
C:'s current directory. That holds for a drive root (`D:\wt`) and for the
`\\wsl.localhost\<Distro>\...` UNC root a WSL project uses, both verified.

The value reaches the guard from two configs: the per-user Worktree Shared Paths
setting alone on the create path (createWorktreeLinkedPaths, called with
`repo.symlinkPaths` from orca-runtime.ts:27820 and worktree-remote.ts:2636), and
that setting merged with the repo's checked-in `orca.yaml`
`worktree.sharedDirectories` on the removal and detection paths
(getWorktreeSharedLinkPaths). No repo config is required to reach it.

Replace both `isAbsolute` calls with a `/^[a-zA-Z]:/` test, verified by fuzz to
be a strict superset of `posix.isAbsolute || win32.isAbsolute` for every
post-strip input. No filesystem escape is closed on macOS/Linux, where such an
entry resolves to a literal in-worktree filename.

Cost: on POSIX, `:` is a legal filename character, so a shared/linked path whose
first segment is `<letter>:...` is now refused where it previously worked — it
stops being created, and if a worktree already holds an Orca-created symlink
there it stops being excluded from the untracked-file filters in all four
findExistingWorktreeSymlinkPaths callers, which means a refused non-force
worktree removal (remove-registered-local-worktree.ts:91, orca-runtime.ts:30205),
a phantom untracked row in Source Control (status-read.ts:90), and a blocked
hosted-review creation (hosted-review-creation-git-state.ts:290) — and
removeWorktreeLinkedPaths no longer unlinks it, so nothing cleans it up.
Accepted because a per-host verdict would defeat the point of judging the same
config identically on every host it is evaluated on.

Co-authored-by: Neil <neil@example.com>
2026-08-31 20:32:23 -07:00
Neil a5796ec8eb refactor(runtime): split OrcaRuntimeService and compatibility tests (#17605)
* refactor(runtime): split OrcaRuntimeService into focused modules

* test(runtime): cover admission tiers and strict worktree reconciliation

* fix(runtime): preserve owner and structured session visibility

* fix(runtime): port post-extraction compatibility fixes

* fix(runtime): preserve skill-share cancellation barrier

* test(runtime): update identity inventory after extraction

* fix(runtime): preserve hook transport environment cleanup

* fix(runtime): consolidate idle probe imports

* test(runtime): retire split file process allowlist entry

* fix(runtime): route child process types through shared boundary

* test(runtime): preserve worktree host metadata precedence

* fix(runtime): update extracted test seams

* fix(runtime): gate the split's ts-nocheck set and restore the stop-confirmed contract

Audit follow-ups for the OrcaRuntimeService split:

- Freeze the 171 @ts-nocheck files behind a ratchet so no new file can disable
  type checking. The split's linear mixin chain cannot express forward
  references yet, so the existing suppressions are grandfathered; the baseline
  may only shrink.
- Drop the stray @ts-nocheck at the end of orca-runtime-get-status.ts. It sat
  after the first statement, where TypeScript ignores it, so the module was
  already checked.
- Restore `retireRejectedPty(ptyId, stopConfirmed: boolean)` as a required
  argument. The split widened it to optional and patched the resulting error
  with `stopConfirmed === true`; an omitted argument would have silently taken
  the unverified-stop path instead of failing to compile.
- Guard that every orca-runtime-tests fragment is imported by the compatibility
  entrypoint. The fragments are .spec.ts, which no Vitest include glob matches,
  so one left out of the list would silently stop running.

* fix(runtime): restore four behaviors the OrcaRuntimeService split dropped

Audit findings against the refactor's true base (ad5ba2572e):

- retirePtyAgentLaunchAuthority collected pane keys after deleting the
  restored-authority receipt instead of before it. collectPaneKeysForPty reads
  that receipt, so a receipt-only pane lost its key and never had its agent-hook
  compatibility authority retired. on-pty-exit.ts already carried a comment
  naming this exact invariant.
- The PTY-exit path kept orchestrationMailboxNotifications.retirePty but lost
  the loop that schedules a debounced mail-pointer repoint for the dead pty's
  terminal handle and any run bound to its panes. Restores the schedule call
  count to 7, matching base.
- subscribeToPtyExit lost isPtyKnownExited's leaf fallback and its
  post-registration lifecycle-generation recheck. leavesByPtyId is rebuilt from
  the renderer graph independently of ptysById, so a leaf can outlive its pty
  record; without the fallback a caller waiting on an already-dead pty never
  gets released.
- The chain root declared `[key: string]: unknown`, which base had nowhere. It
  leaked through the exported runtime type into every consumer, so any misspelled
  member access typechecked as unknown instead of erroring, and it accounted for
  957 of the suppressed errors. Removing it costs zero type errors.

* fix(runtime): restore escalation prose and unscoped automation publication

Two more behaviors the split dropped, each with a regression test that fails
against the pre-fix code:

- The worker-exit escalation stopped deriving its title through
  buildOrchestrationTaskDisplayMetadata and inlined `task.spec` instead. That
  ignored an explicit task_title, dropped the single-line normalization and the
  80-character bound, and turned the no-spec case into a quoted, duplicated id.
  A multi-paragraph spec landed verbatim in the coordinator's banner. The
  existing 11 tests all use short single-line specs, where the derived title and
  the raw spec are identical, so none of them could see it.
  Also reverts an added `if (!handle) return` guard: the dispatch lookup is
  deliberately keyed on the pane as well, because a reminted handle no longer
  matches the row while the pane identity outlives the remint.
- updateAutomation stopped going through automationChangePublications and
  published `source` unconditionally while gating the fallback on a non-null
  destination. A destination the store can no longer name then published only
  the stale source, so subscribers scoped elsewhere kept rendering a row that
  had left them — the exact case the helper documents. The helper had been left
  with zero callers; all three sites use it again.

* fix(skills): stop swallowing lookup errors and hard-erroring on non-ssh hosts

Follow-ups from auditing the skill install path against the refactor's base:

- resolveWorktree wrapped showManagedWorktree in `.catch(() => null)`, so a
  transient git or IO failure surfaced to the user as
  skill-install-workspace-not-found with the real cause discarded. Errors
  propagate again; a genuine id mismatch still returns null.
- resolveSkillSshTarget threw skill-install-workspace-host-unavailable when the
  execution host was neither local nor ssh, on both the repo and folder
  branches. Base gated these on connectionId, so a runtime-owned repo simply
  was not an SSH install and fell through to the local path. Both return null
  again, and the error code the split invented is now unreferenced.
- listManagedSkillInstalls awaited the receipt walk and the worktree resolve in
  sequence. They are independent and either can hit disk, WSL, or an SSH scan,
  so Promise.all is restored.

Deliberately unchanged: resolving the worktree through listResolvedWorktrees
rather than showManagedWorktree, which disambiguates a worktree id colliding
across hosts and is covered by its own test, and the SSH-folder
skill-install-ssh-dispatch-required throw, which matches the repo branch.

* fix(runtime): merge duplicate worktree-logic imports

The #17448 port added a third import from ../ipc/worktree-logic, which the
code-quality oxlint config rejects under --deny-warnings. Plain oxlint does not
flag it, so it only surfaced in CI's static analysis job.

* ci: run the ts-nocheck ratchet in PR checks

pr-workflow-lint-parity requires every leaf command in `pnpm lint` to have a
matching step in pr.yml. The ratchet was wired into lint but not the workflow,
so PR CI would not have enforced it.

* Merge remote-tracking branch 'origin/main' and retry the paired-host launch evaluate

main advanced 9 commits; none touch the orca-runtime.ts this branch splits, so
nothing needed porting.

CI failed twice on `Execution context was destroyed` thrown from
headless-paired-runtime-host's first `evaluate` after launch — a different spec
each run, which is the signature of the flake #17780 describes rather than a
regression. That commit added retryTransientMainEvaluate and adopted it in five
helpers but not this call site, even though its docblock names exactly this
case: the first evaluate after electron.launch() resolves, before the app is
ready. Wrapped it the same way.
2026-08-31 19:34:55 -07:00
Jinwoo Hong 12be5aed9e fix(browser): refuse devtools for offscreen guests (#17485) 2026-08-31 22:15:47 -04:00
Neil f116d2ca2a test(ci): retry Windows teardown EPERM and restart evaluate misses (#17780)
Restart-survival polls treated a recycled renderer as a hard failure.
Wrap those evaluates so "Execution context was destroyed" is a pending
miss. Windows package-lane teardowns after a force-kill used rmSync
with force:true only, which does not absorb EPERM; put them on the
shared maxRetries:8 policy.
2026-08-31 18:53:01 -07:00
Jinjing d2aab68ae7 Automations ux improvement (#17626)
* Add keyboard navigation to automations UI

Improves workflow efficiency by enabling keyboard-driven navigation
across automations list, run history, and detail pane tabs.

* Add Escape key support to automations detail pane

Pressing Escape now clears external and automation run page views,
then returns to the automations list. Also improves cross-browser
compatibility of keyboard event handling by using Element checks and
getAttribute instead of dataset access.

* Fix keyboard navigation to let Enter key reach focused controls

- Enter key now passes through to focused buttons, links, and other interactive controls
- Arrow key navigation through automation run history still works
- Prevents intercepting native keyboard behavior of interactive elements

* improve test

* Move keyboard focus to follow row selection

When navigating automation runs with arrow keys, focus must follow the selection so Enter key acts on the newly selected row rather than the previously focused one.
2026-08-31 18:37:19 -07:00
Jinjing 50938b2dbd Serialize filesystem watcher batch flush operations (#17602)
* Serialize filesystem watcher batch flush operations

- Prevent dropped events during rapid concurrent file changes
- Queue and drain follow-up batches to preserve event ordering
- Cancel pending batch work when watchers are torn down

* Prevent queued batch drain while debounce timer is armed

An armed timer means the debounce window is still open. Drain only after
the window closes to avoid splitting related filesystem events across
separate payloads.

* Remove redundant batch timer cleanup

Rely on cancelLocalBatchFlush to handle the batch timer
teardown, eliminating duplicate logic in the watcher
cleanup path.
2026-08-31 18:28:26 -07:00
Neil 2222e54754 refactor(test): organize SSH and terminal recovery fixtures (#17751) 2026-08-31 18:18:15 -07:00
Neil ae35e044f2 fix(terminal): keep restored OSC-8 ranges across a no-op resize (#17759)
Cold restore seeded a checkpoint's OSC-8 link ranges and then replayed records
that resize, so any resize record after the checkpoint dropped them and
restored hyperlinks in scrollback lost clickability. Same-size resize records
reach the durable log routinely, because every attach re-asserts the pane's
dimensions and session-output-plane records each one without a same-size
dedupe — so an ordinary reattach was enough to lose the links.

Restored ranges are row-indexed, so clearing them on a reflow is right; a
resize to the size already applied is not a reflow. Gate on the dimensions
actually changing.

Introduced in d46349ce82 ("fix: improve mobile link modifier handling",
#5597), which added setRestoredOscLinks along with unconditional clearing in
both resize() and clearScrollback(). clearScrollback's clearing is correct and
is unchanged, with a test pinning it.

Found during adversarial review of #17752 and filed as #17756. Not a
regression from that PR: #17667 had incidentally masked it by gating no-op
resizes to protect a snapshot cache, and removing the cache removed the gate.
The same gate returns here on its own terms — as a correctness fix with tests,
rather than as a side effect of a cache.
2026-08-31 18:06:08 -07:00
Neil 704167197a perf(relay): serve one ps capture per window and pin the batched inventory path (#17763)
Follow-up defect fixes for the batched PTY-inventory evidence path (#17525),
now on main.

- One memoized `ps` capture serves both the lenient and strict views. The two
  readers ran byte-identical argv behind separate caches, so a relay serving
  both forked `ps` twice per 500ms window — the doubling issue #6288 removed.
- Drop the `byPgid`/`byTpgid` indexes no resolver reads, plus the zero-caller
  `parseProcessTableRowsStrict` and `getFreshStrictProcessTableSnapshot`; the
  batch resolver now reuses the shared index lookup and candidate score instead
  of private copies.
- Restore `getForegroundProcessName`'s ladder contract: the extracted table scan
  answers null again, so an unconfirmed wrapper fallback publishes the
  recognized (normalized) name rather than node-pty's raw one.
- Pin the SHIPPED `pty.listProcesses` path: one capture and one linear row pass
  for N panes, and node-pty's own name (never "shell") when the capture cannot
  disambiguate a `node`/`python` wrapper.
- Pin the hidden-pane cadence gate in the production option shape, and move the
  strict-parser coverage next to the parser it tests.
2026-08-31 17:45:55 -07:00
Jinwoo Hong 26031ca317 fix(browser): scroll oversized viewport presets (#17569)
* fix(browser): scroll oversized viewport presets

* fix(browser): preserve guest wheel scrolling at viewport edges

* fix(browser): keep viewport scroll state synchronized

* test: assert partial viewport wheel forwarding
2026-08-31 20:39:58 -04:00
Neil 4ac8a8912c fix(startup): install the app environment with the userData decision (#17755)
src/main/index.ts decided where userData lives at module scope, then installed the
AppEnvironment port ~180 lines later inside the single-instance-lock block. Every
statement in that gap was a latent failure: a path resolve there either threw
'AppEnvironment not initialized' and killed the process, or — with the accessor
installed but the decision not yet run — would have memoized the pre-override
directory in getCanonicalUserDataPath() for the whole session.

The first outcome shipped. #16761/#16698/#17509 were one statement landing in that
gap and killing every macOS `orca serve` across 1.4.190-1.4.192; #16762 moved that
call but left the gap.

Install the port and capture the canonical path immediately after the two calls
that decide them, so the window is zero rather than small. Both are inert at this
point — ElectronAppEnvironment holds no state and calls `app` lazily per accessor,
and initDataPath only joins strings — so nothing that depended on the old position
moves with them. The secret store stays where its pre-ready Keychain note applies.

The throw is kept and still covers the case it should: resolving a path before the
decision has run.

Guarded by a source-level assertion that the decision, the install and the capture
stay adjacent.

Fixes #17750
2026-08-31 16:40:29 -07:00
Neil 8ac1c6e2ac perf(git): bound ref and worktree scans (#17655)
* perf(git): bound ref and worktree scans

* fix(repo-search): clamp oversized ref limits

* fix(worktree): keep strict worktree listing unshared

The shared-scan re-export flipped every `listWorktreesStrict` caller from an
isolated subprocess to the coalesced scan. `git worktree prune` in the removal
recovery path does not bump the scan generation, so a post-prune verification
could join a pre-prune scan, see the stale row, and report a successful removal
as a stale registration. The same gap defeats the post-archive-hook rechecks
that exist to catch an external Git client locking the row.

Restore the unshared export and make coalescing opt-in via
`listWorktreesSharedStrict`, which existing callers already use deliberately.

* fix(git): separate a proven absent ref from a failed probe

`show-ref --verify --quiet` exits 1 for a missing ref, but so does `wsl.exe`
when its own launch fails, so reading any exit 1 as absence collapsed
`unverifiable` into `exited`. A genuine miss prints nothing while a wrapper
failure always explains itself, so require empty stderr alongside the exit
code; a runner that reports no stderr at all keeps its exit-code contract.

That same signal removes a spawn regression: `show-ref` is a direct-git read
under WSL, and the runner retried any numeric exit through the user's
interactive login shell. The replaced `for-each-ref` exited 0 on a miss, so
absence never retried; every absent probe now would. Treat a quiet exit 1 as
Git control flow and skip the fallback.

Also narrow the hosted-review suffix fallback: the replaced
`refs/remotes/*/<base>` could not cross a slash, but `show-ref -- <base>`
matches at any depth, so `origin/feature/main` answered a query for `main`
and submitted a review against a base the provider rejects.

Refresh the real-binary compatibility contract to the shipped excludes, and
assert exact probe concurrency rather than an upper bound so a regression to
serial probing fails.
2026-08-31 16:27:28 -07:00
Neil a5be10815c perf(terminal): drop the headless snapshot cache, keep the spawn lock (#17752)
Removes the snapshot memoization added in #17667 and everything that served
it: the epoch, markMutated/markWritten, the no-op resize gate, and the
HeadlessSnapshotCache module. The shared/exclusive spawn lock from the same PR
stays — it is the half that carries the measured win.

Why: a controlled A/B on merged main could not show the cache paying for its
memory. Same worktree, same sessions, reattach stable_adoption per session:

  cache + resize gate (as merged)   16 / 67 / 78 / 87 ms
  cache off, gate on                17 / 55 / 68 / 80 ms
  cache off, gate off               13 / 62 / 73 / 83 ms

Indistinguishable. Cold 4-tab activation was likewise unchanged (217-296ms
without the cache vs 226-309ms with), which is expected — a first attach is
always a miss.

The reason it under-delivers is a design fact the original PR missed: attach
does not request the full buffer. terminal-host-session-create.ts passes
resolveDaemonSessionScrollbackRows() — a deliberate 1000-row live window,
capped because unbounded retention once OOM-killed a host. Serializing 1000
rows is cheap, so there was little for a cache to save on that path.

Against that, the cache retained up to MAX_CACHED_SNAPSHOT_BYTES (4MB) per
entry across MAX_CACHED_SNAPSHOT_WINDOWS (2) entries per emulator, for the
session's lifetime, with no aggregate budget across sessions. It also carried
an invalidation contract that produced three separate over-invalidation bugs
during review (the parse fence, setCwd/setLastTitle, and the no-op resize).

The lock fix is unaffected and independently measured: the `options` phase,
which is pure queueing, went 0/125/212/291ms -> 1/1/1/1ms across a 4-tab
worktree activation and stays there.
2026-08-31 16:08:47 -07:00
Jinwoo Hong 8f15f217a2 Preserve user-set workspace names across branch changes (#17448)
* fix(worktrees): preserve user workspace names across branch changes

* test(worktrees): cover pinned rename metadata

* fix(workspaces): address display-name review edge cases

* fix(workspaces): keep automatic names fresh across refreshes

* fix(workspaces): preserve legacy CLI labels

* fix(workspaces): preserve display-name provenance across hosts

* fix(workspaces): honor legacy display-name provenance

* fix(workspaces): fence display-name refresh races

* fix(workspaces): accept peer renames from provenance-less hosts

The old-host preserve fence kept a pinned local label on every refresh,
which also suppressed a legitimate rename another client persisted
through the same host until app restart. Narrow it to labels the host
re-derived itself (branch short name, or path basename when detached);
any other changed label in a mode-less response is explicit meta a peer
wrote there. Stale prior-label responses stay covered by the downstream
staleness fence, in-flight writes by the pending fence.

* refactor(workspaces): unify display-name pin derivation

Three call sites (renderer optimistic update, local IPC updateMeta
handler, remote worktree.set handler) each restated the same formula;
a future edit to one would silently skew provenance between paths.
2026-08-31 19:08:05 -04:00