mirror of
https://github.com/stablyai/orca.git
synced 2026-09-29 16:02:50 +00:00
71f3700aecbf2694b54dbe6dbb0bc50fb6d480a3
1405
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
71f3700aec |
perf(remote): read the repo catalog once per publish, not once per worktree
`remoteWorkspace:setForConnectedTargets` costs 13 ms of main-thread time per call at 0.48 calls/sec — 0.63% of wall on a real session, the second most expensive IPC handler in the main process. Almost all of it is one line. `exportRemoteWorkspaceSession` asks `isTargetWorktree(worktreeId)` once per worktree in the session, and that callback called `targetForWorktree(store, ...)`, which called `store.getRepos()` — and `getRepos()` maps `hydrateRepo` over every repo row. So publishing to one SSH target re-hydrated the whole repo catalog once per worktree, then threw a fresh `createRepoRowExecutionHostLookup` (which itself `filter`s the catalog per lookup) away each time. The lookup is now built once per handler invocation and shared across targets: the rows cannot change inside one synchronous projection, and they are the same for every target. On the session that surfaced this — 413 worktrees, 13 repos, 1 connected target — that is 413 catalog hydrations (5369 `hydrateRepo` calls) per publish reduced to 1 (13 calls). The repo normaliser reached through `hydrateRepo` was the #2 self-time function in a 30 s main-process CPU profile at 0.25%. No user-facing trade-off: identical ownership resolution, identical exported session, identical stale-revision handling. |
||
|
|
7ed86a98ae |
perf(ipc): index worktree owners instead of rescanning the repo list per lookup (#18416)
Two hot lookups rescanned a whole table once per repo. `getLocalRepoForRegisteredWorktree` (59 IPC call sites, including Quick Open keystrokes and every File Explorer expand) walked the entire worktree-meta table once per repo. One pass now collects the owning repo ids, built lazily so a repo whose own path matches still never touches the table. `createRepoRowExecutionHostLookup` re-filtered the repo array on every `byId` / `byHost` call. Rows are grouped into a Map once at construction, preserving repo-list order so `rows[0]` still picks the same owner. |
||
|
|
e42c60e8a3 |
fix(ssh): resolve a pane's binding from the target partition, not the stale local copy (#18546)
One SSH pane accumulated one extra reattachable lease per relay restart (2, 3, 4, 5, 6 across five), and every one of them costs a `pty.attach` round trip on every later connect, forever. Nothing prunes `sshRemotePtyLeases`, so the fan-out only grows. `supersedeSiblingLeasesForPane` is fenced on the PTY the pane is durably bound to, and `durablyBoundPtyIdForPane` read `state.workspaceSession` (local) before `workspaceSessionsByHostId['ssh:<target>']`. But `persistPtyBinding(binding, hostId)` updates ONLY the host partition: AFTER-PERSIST local= ssh:t@@pty2:old:1 host= ssh:t@@pty2:new:1 So for the length of a reconnect the local copy still names the predecessor, the fence resolved to it, supersession took an already-`expired` lease as its winner, and returned having marked nothing. Both partitions agree again once the renderer republishes its layout, which is why the settled store looks consistent and hid this. Read both partitions as an ordered list, target's own first, and test the fence by membership rather than by equality with whichever was read first. Pick the winner preferring a lease this client still has a route to, since the stale partition names an expired one. Never retire a lease that is both bound and live, so a partition disagreement can't strand a running remote process. Superseded predecessors stay `expired` and are never `terminated`: losing a lease is not evidence the shell died (docs/reference/ssh-execution-boundary.md). A pane with no binding is skipped rather than pruned, so a genuine orphan stays askable. Also re-runs supersession from the binding side after each spawn commit's binding write, so the lease/binding order at a call site no longer decides, and reconciles every pane for a target immediately before `reattachKnownPtys` reads the set it feeds to `pty.attach` — that repairs stores which already accumulated these rows. The guard suite could not catch this: every assertion bound the pane BEFORE upserting the lease, an order no caller uses. Rewritten to the spawn commits' real order (lease, then binding, then the binding-side trigger); it fails 8 assertions without this change. Added a suite that drives the real `persistPtyIpcSpawnCommit` rather than the store primitives, including the exact stale-partition state written by production's own binding writer. Verified on the Docker SSH lane: five `relay.js` SIGKILLs with recovery between each, reattachable leases flat at one per pane. Note: this bounds the reattach SET, not the store. `sshRemotePtyLeases` still has no cap or TTL and rows still accumulate; pruning is left alone deliberately, since an `expired` row without `supersededBy` is a genuine orphan and must not be dropped on age. |
||
|
|
5d8532f6d3 |
fix(worktrees): resolve the execution host at both worktree-create entry points (#18545)
Two entry points create the same workspace and disagreed about how to read its host. `orca-runtime-create-managed-worktree.ts:63` resolved through `getRepoSshConnectionId` and then normalized the row; the `worktrees:create` IPC handler branched on raw `repo.connectionId` (`register-worktree-create-handlers.ts:66-69`). So a repo naming its owner only as `executionHostId: 'ssh:<target>'` created remotely through the runtime and ran `git worktree add` on the client against a remote path through IPC (#11163). Same repo, two entry points, different answers. Both now take one route, resolved through the existing layer (`getRepoExecutionHostId` -> #18296's `resolveGitRouteForHost`). No new resolver. The row normalization on the `ssh` variant is kept, and it is a **workaround, not the pattern**. `createRemoteWorktree` and its callees re-read `repo.connectionId!` at five depths in `ipc/worktree-remote.ts` (1627, 1847, 1848, 1865, 2029), so the resolved connection has to reach them through the field they already read. It travels only as far as that object does — anything downstream that re-reads the row from the store still sees the unnormalized one, and it cannot express the `runtime:` refusal on its own. Proper fix, deliberately not done here: give that pipeline an explicit connection parameter and delete `repo.connectionId!` from it so every reader becomes a compile error, the technique #18307/#18325 used. That is a change inside a 2800-line module plus its callers, and it wants its own PR. Three answers that used to collapse into one, now distinct at both entry points: - `executionHostId: 'ssh:*'` with no `connectionId` -> that SSH host (IPC used to create locally); - `executionHostId: 'local'` with a surviving `connectionId` -> local, since a local row cannot nest an SSH namespace. This is what `getRepoSshConnectionId` and therefore the runtime sibling already answered; IPC used to go remote; - `runtime:<env>` -> refused. Its worktree is created by that environment's own server and the SSH target on its repo row is that server's nested one, addressable only as (environmentId, targetId). The renderer already routes runtime-environment creates over `worktree.create` RPC rather than this IPC channel, so reaching either entry point with one is a routing mistake. Matches `workspace-cleanup-git-route` and `runtime-git-command-target`. Folder-workspace creation is untouched on both sides: it is a registration, not a filesystem create, so the route is resolved after that branch on the IPC side, and on the runtime side only the agent trust write consumes it — where a `runtime:` host now yields `null` instead of the nested target, so the write stops going to a same-named target in this client's table. No wire or persistence change: the normalized row is a local value passed to the create pipeline, never stored, and `CreateWorktreeResult` is untouched. |
||
|
|
a5d6114baf |
fix(ssh): stop pane adoption certifying a death from the relay's not-found union (#18531)
* fix(ssh): stop pane adoption certifying a death from the relay's not-found union
`attachStablePaneOwner` was the last reader that synthesised a runtime exit
from a reattach refusal, and it published code `0` — which
`orca-runtime-on-pty-exit` records as `rememberPtyLivenessVerdict(exited)`, a
death certificate whose only legitimate writer is a host-delivered exit frame.
The refusal it acted on is a union. `pty.attach` answers `PTY "<id>" not found`
both for a pid the relay probed with `isProcessAlive` and for an id its session
map simply never had — which, because ids carry a per-start mint epoch, is every
id minted before a relay restart, checked against nothing. So a relay restart
plus a reconnect certified a shell that was still running under the old daemon's
orphaned process tree, retired the pane binding, and cold-started a second agent
onto the same transcript. The sibling `handlePtyReattachFailure` has always
refused to certify from that union; this path did not.
- The relay marks the one refusal it backed with a liveness check
(`PTY_ATTACH_PROVEN_EXITED_MARKER`). The marker is additive, so an unmarked
answer — including an older relay's — stays ambiguous, which is the safe
direction.
- The client mints that half as `SshPtyProvenExitedOnRelayError`, a subclass so
every existing `isSshPtyAbsentFromRelayError` consumer is unchanged.
- Pane adoption publishes `UNVERIFIED_PROCESS_EXIT_CODE` (-1), the sentinel its
sibling publishes, and passes `hostExitConfirmed` only for evidence that
observed the process: the marked relay refusal, or `SessionNotFoundError` from
the registry that owns the PTY. The ambiguous half now records `unverifiable`
instead of `exited`.
- The gone-branch keys on the error type rather than the bare `PTY ".+" not
found` text, so an untyped string can no longer authorise abandoning a
binding — the discriminator `pty-connect-limits.ts` already documented.
Refs docs/reference/ssh-execution-boundary.md
* test(pty): make the pane-adoption fixtures throw what real providers throw
These four fixtures rejected with bare `new Error('Session not found: ...')` and
`new Error('PTY "..." not found')`. No provider produces either untyped:
`local-pty-spawn` and `decodeDaemonResponseError` both mint
`SessionNotFoundError`, and the SSH reattach path types the relay's wire text
before any pane sees it. Fixtures that skip the type were the reason a
message-shaped gate looked adequate.
The exit-code expectations move with it: the pane path now publishes the -1
stop sentinel plus `hostExitConfirmed`, so a certificate follows the evidence
rather than a synthesized zero.
|
||
|
|
4cc0b8de61 |
perf(hot-paths): delete allocation-only work in sort, explorer, monaco, rpc, snapshots (#18372)
* perf(hot-paths): delete allocation-only work in sort, explorer, monaco, rpc, snapshots * fix(perf): revert snapshot revision fast-path — same revision can carry a new session * perf(hot-paths): drop the unproven rpc buffer rewrite, dedupe the equality helpers - Revert the unix-socket chunk-carry change. Its comment claimed it avoided O(n^2) rescans, but chunks is reset to [remainder] every data event, so the join plus the tail byteLength is two passes where the old code did one; benchmarks showed no win. It also moved consumed-frame bookkeeping out of the closure, so a synchronous throw from the handler would re-dispatch frames. - project-host-compatibility: fold the two byte-identical array comparators into one generic arraysEqualByJson. - smart-attention: drop the leftover byTab.size === 0 branch that returned the same value as the line after it. |
||
|
|
3a32e084dd |
perf(renderer): index diff comments, skip no-op hydration, drop duplicate normalizes (#18375)
* perf(renderer): index diff comments, skip no-op hydration, drop duplicate normalizes * fix(perf): keep tree-path stability hook render-pure for react-doctor * fix(perf): publish the returned array from the tree-path stability hook The ref was written with the raw input but read in render to pick the return value, so it trailed one commit and a wave of content-equal arrays flipped identity every render — re-firing the uncancellable full-tree git check-ignore it exists to prevent. Publish `stable` instead, keyed on `[stable]`. Also drops the hydrateOverrides no-op skip: notifyChange is not a bare wakeup (it drives getPanesNeedingOverrideFit -> safeFit and the remote viewport re-claim), and the branch never fires in production anyway. |
||
|
|
a9f2fbb684 | chore(workspaces): drop the dead workspaceCleanup:hasKillableLocalProcesses IPC (#18386) | ||
|
|
b8da193b7a | fix(ssh): route the remaining expired-lease readers through the reattach predicate (#18378) | ||
|
|
d05dd8ef50 |
fix(source-control): route hosted reviews by resolved execution host (#18382)
`ForgeProvider.createReview(repoPath, input, connectionId, options)` and the `connectionId` on `ForgeProviderRepositoryContext` carried the same collapse the five prior migrations closed: `string | null` spells "genuinely local", "runtime host" and "could not resolve" with one value. Because it was decided two layers up -- `repo.connectionId ?? null` at the `hostedReview:*` IPC handlers and in `RuntimeHostedReviewCommands` -- a row naming its owner only as `executionHostId: ssh:<target>` ran the whole review path against this machine's copy of a remote path (#11163): `git rev-parse`, `git status`, the base-on-remote ref probe, the upstream divergence read, and `gh`/`glab` with no host flags. Replace it with a required `ExecutionHostId` threaded from the decision point through the contract, routed by #18296's `resolveGitRouteForHost`. The parameter is removed rather than added beside, so all five implementations -- GitLab, GitHub, Bitbucket, Azure DevOps, Gitea -- and every caller became a compile error. None of these families carries `@ts-nocheck`, so unlike #18325 that guarantee is real here; `orca-runtime-file-commands.ts` does, but it only constructs `RuntimeHostedReviewCommands` with unchanged deps. Also fixed at the sites: - The branch cache scoped entries on `connectionId ?? ''`, so two rows at one path on different hosts shared one cached review, one backoff deadline and one invalidation. Keyed on the resolved host now, as #18377 did for its probe key. - `hostedReview:create` resolved shared symlink paths and normalized worktree paths off the raw field, so an `executionHostId`-only SSH row read `orca.yaml` and `resolve()`d a remote POSIX path on the client. Those ask the file-holder question -- `getRepoSshConnectionId` -- not the dialable one. - An SSH host with no provider now refuses inside the git-state layer instead of reaching the local branch, keeping "remote and unreachable" distinct from "local" (docs/reference/ssh-execution-boundary.md). `runtime:` is a routing mistake inside `hostedReviewSshConnectionId` -- that environment's server runs its own git, and the SSH target on its repo row is nested in that server's namespace, so dialing it here reaches a same-named box of ours. But store-backed callers ask `getRepoHostedReviewExecutionHostId` first, which is "what may this client dial" and answers `local` for a `runtime:` row. That is deliberate and matches #18377: the runtime registration controller only adopts a `runtime:` stamp onto a row with no `connectionId` (`runtimeRepoMatchesExecutionHost` refuses to match an SSH row), so the checkout really is in this process and refusing would regress a runtime server creating reviews for its own rows. No wire change. `connectionId` on `CreateHostedReviewArgs`, `CreateStackedHostedReviewArgs` and `HostedReviewCreationEligibilityArgs` in src/shared/hosted-review.ts is untouched -- every host already ignores it in favor of the repo row, and removing it from the request types would only churn the schema older clients still populate. The main-side eligibility input `Omit`s it so nothing on this side can read the ambiguous field again. |
||
|
|
316ec38f67 |
fix(repos): route icon and remote-identity probes on a resolved execution host (#18377)
`detectRepoIcon`, `detectRepoIconAndUpstream`, `detectGitHubAvatarIcon`, `detectRepoFileIcon` and `probeGitRemoteIdentity` took a `connectionId`-shaped parameter threaded down from their callers. That shape spells "runtime host", "unresolved" and "genuinely local" all as one falsy value, and because it is a *parameter* each caller decided independently what to pass — a wrong answer was invisible at the boundary. Replace it with a required `ExecutionHostId` and route through #18296's `resolveGitRouteForHost` / `resolveFilesystemRouteForHost`. The parameter is removed rather than added beside, so every caller became a compile error. No new resolver, no wire change: nothing these modules return carries a host id. Fixed at the call sites: - `repo-git-remote-identity-enrichment` read `repo.connectionId` raw, so a row minted with only `executionHostId: ssh:<t>` ran `git remote -v` against this machine's copy of the path (#11163), and a `runtime:` row handed its *nested* SSH target to this client's dispatch table — a same-named box of ours. - Its location key had the same collapse, so two rows at one path on different hosts shared a probe, an abort controller and a backoff deadline. - `runtime-repository-fork-backfill` guarded on `repo.connectionId`, so an `executionHostId`-only SSH row had its upstream read off the client. `runtime:` is refused inside the modules (this process does not execute another environment's git or filesystem), but store-backed callers ask `getSshTargetIdForExecutionHost` — "what may this client dial" — so a `runtime:` row keeps the probe this process has always run for it. Registering and cloning stay `local` on purpose: those controllers do the filesystem work here, whatever host id is stamped on the row (see `assertCloneHostIsSupported`). |
||
|
|
d5803bdbc4 |
feat(ssh): host-stamped remote foreground identity (#18078)
* docs: add SSH agent identity implementation plan * feat(ssh): host-stamped remote foreground identity * fix(runtime): preserve unfenced inspect call shape * perf(ssh): traverse foreground descendants linearly * fix(ssh): bound retired PTY evidence records * test(ssh): cover retired incarnation retention * fix(ssh): make remote process inspection total * Split SSH identity build hot spots * Fix process table snapshot module split * test(ssh): update process inspection expectations * docs: drop the SSH identity plan from the PR The design doc does not belong in the product repo; it stays out of the shipped tree while the implementation carries its own comments. --------- Co-authored-by: Merge Sim <sim@local> |
||
|
|
fe5efb24c8 |
fix(cleanup): route workspace cleanup by resolved execution host, not repo connectionId (#18358)
The workspace-cleanup scan threaded `provider: IGitProvider | null`, derived from a raw `repo.connectionId` read, through listing, activity and git evidence. That `null` spelled "this is local", "the host is remote but unreachable" and "the host is a runtime environment" with one value, so a row naming its owner only as `executionHostId: 'ssh:<target>'` listed worktrees, statted paths and ran `git status` for a *remote* checkout on this client (#11163). The three sites had to move together: the `provider!` assertions in workspace-cleanup-git-evidence.ts were sound only because they re-read the same field that workspace-cleanup-worktree-listing.ts used to decide whether `provider` was populated. Migrating one alone turns them into crashes. Routing now goes through the shared resolution layer -- `getRepoExecutionHostId` for the repo that produces the listing, `getWorktreeExecutionHostId` for the workspace's own host -- into `resolveGitRouteForHost` from #18296's host-keyed dispatch. The ambiguous carrier is removed rather than supplemented, so every reader became a compile error; unlike #18325's family, workspace-cleanup carries no `@ts-nocheck`, so that guarantee is real here. `runtime:<env>` is not a route variant. Its Git runs on that environment's own server and the SSH target on its repo row is that server's nested one, addressable only as (environmentId, targetId); handing it to this client's SSH table dials a same-named target in the wrong namespace. It throws, matching workspace-space-repo-scan and repos:listForExecutionHost. No wire change: `WorkspaceCleanupCandidate` (including `connectionId` and `executionHostId`) and the workspaceCleanup RPC UI-state schema are untouched. |
||
|
|
94be54d16c |
perf(file-explorer): stop rebuilding the whole visible tree twice per directory refresh (#18319)
* perf(file-explorer): stop rebuilding the whole visible tree twice per directory refresh The per-directory loading flag moves out of `dirCache` into a sibling `Set<string>`, so a `dirCache` identity change now means "children changed". Every identity change re-ran `getFileExplorerIgnoredQueryRelativePaths` (full recursive walk) and `createVisibleFileExplorerRowProjection` (full flatten, new Map, new array identity cascading into virtual rows, selection, keyboard nav and the name filter) over the whole visible tree — and half of those rebuilds produced a byte-identical row set. Also in this change: - `refreshFileExplorerExpandedDirs` no longer pre-marks every expanded dir in `dirCache`; the 13 progressive commits stay. - `flushBatch` paces its `fs.stat` fanout at 8 (was up to 5,000 concurrent onto libuv's 4-thread pool), matching parcel-watcher-event-delivery.ts. - The editor external-watch loop bails before allocating a notification for a path no open file matches. - `createCachedDirPathIndex` is built lazily, only when a direct `dirPath in cache` lookup misses. * fix(file-explorer): keep the loading-dirs ref out of the render body React Doctor's no-ref-current-in-render flagged the render-body mirror, and it was right: a render React discards would still have mutated the ref. The ref is now authoritative and written only from callbacks, with one updater that moves the ref and the state together. Side effect, in the safe direction: loadDir's in-flight guard now sees a mark the moment it is made instead of one commit later, so a second non-forced read of a directory already being read is deduped rather than started and then superseded. Forced reads (refreshDir, refreshTree) bypass the guard and are unaffected. Also moves the in-flight check out of decideExpandedDirLoad and into the expansion effect that owns the fan-out, restoring the two-argument signature. This clears the no-pass-data-to-parent warning the three-argument call had dragged onto a changed line, and it keeps the pure staleness decision pure. |
||
|
|
57681ecd09 |
fix(remote): resolve the spawn cwd, the node manager dir, the vault host and the scrollback seed (#17952)
* fix(remote): resolve workspace cwd, mise Node, host scope, and TUI scrollback honestly #15296 relay: a folder workspace id (`folder:<uuid>`) carries no path, so the worktree-id split yielded nothing and $HOME silently won. Resolve the spawn cwd through worktreeId -> ORCA_WORKSPACE_ROOT -> host default, and refuse an agent spawn outright when a folder workspace names a root this host cannot resolve. #11733 ssh: generalize the NVM dotfile scrape into `orca_dotfile_dirs` and drive mise off `MISE_DATA_DIR` / `XDG_DATA_HOME` instead of a hardcoded `$HOME/.local/share/mise`. #13713 ai-vault: an unresolvable workspace host is `unverifiable`, not local. Widen the default scope to every host rather than scanning the client's own history and reporting "No agent sessions found". #6106 terminal: hydration asked the renderer for `scrollback: 0` while an alt-screen TUI was up, which drops the normal buffer's shell history rather than the TUI bytes. Drop the flag; readers already split the two buffers apart. * fix(remote): stop the relay answering host questions for a guest execution host Three findings from review of the spawn-cwd resolver, all the same shape: a path question answered against the wrong host, or with the wrong key. - resolveRelaySpawnCwd refused an agent launch whenever a folder workspace named a root that did not stat on the relay. But relayHostDirectoryExists stats the relay's *own* filesystem, and the relay supports WSL shells, so a folder workspace on a Windows relay launching into WSL now threw where it previously spawned -- contradicting the function's own doc comment, which says an absent path for that exact host pair is a miss, not a refusal. Thread the shell's execution host in and demote the refusal to a miss when the spawn does not run on the relay's filesystem. - requireRelaySpawnCwd's doc claims both call sites route through one resolver so the fence can never be keyed on a directory the spawn won't use, but the fence key was still computed with the non-stripping splitWorktreeId while the cwd used splitWorktreeIdForFilesystem. For a `::workspace:<uuid>` id those disagree by construction, in adjacent lines: the removal fence guarded a path no spawn ever enters. Same defect in shutdownForWorktreePath and the revive path; all three now use the filesystem split. - The remote Node probe expanded `$HOME` and `~/` prefixes out of a dotfile assignment but not `$XDG_DATA_HOME`, so `MISE_DATA_DIR=$XDG_DATA_HOME/...` was used as a literal directory name. Add the case arm, defaulting to the POSIX `$HOME/.local/share` the seed value already uses -- sshd's exec channel usually has no XDG_DATA_HOME at all. |
||
|
|
9e9b80cb37 |
perf(relay): stop two unbounded growth terms behind the long-session SSH slowdown (#17818)
Two costs grew for the life of an SSH session and never came back down. 1. The relay port scan walked every process in /proc and readlink'd every fd even after every listening socket already had an owner. Cost was O(host processes x fds) per scan, repeating for the session's life. Exit as soon as every inode is attributed. 2. SshPtyModelAdmission kept closed provider generations in a Set<number>. Provider generations are a process-global monotonic counter shared by every SSH target, so the set gained one entry per relay reconnect forever. After 500k reconnects main retains ~10,234 KB / 500,000 entries; with this change, 18 KB / 1 range. Closed generations now live in SshPtyClosedGenerationRanges, which collapses contiguous closed runs. Membership stays exact -- a generation below the high-water mark can still be live on another host, so a high-water approximation would reject a healthy target's output. The range container's has()/add() were a linear scan; both are now binary search. has() is on the per-output-chunk admission path, so a scan would have traded a bounded Set lookup for one that degrades with fragmentation. This also speeds up ssh-pty-output-generation-guard.ts, which already uses this container on main. Known limitation, deliberately not addressed here: the closed-generation set is bounded in the healthy case (one range) but unbounded when generations leak, since each leaked generation leaves a permanent gap. Sublinear is not bounded. A live-generation set would be bounded by construction and is the better long-term design; that is a follow-up. |
||
|
|
9cda5a9dc0 |
fix(worktrees): stop a resolved-worktree snapshot answering for repos it never saw (#18295)
* fix(worktrees): stop a resolved-worktree snapshot answering for repos it never saw `listResolvedWorktrees` caches one fleet-wide snapshot for RESOLVED_WORKTREE_CACHE_TTL_MS (1s) and reuses it on time alone. Nothing invalidates it when a repo is registered, so for up to a second after a repo row lands, every caller reads a snapshot computed before that repo existed -- and reads the gap as a verdict. The visible failure is the SSH skill install. `resolveSkillSshTarget` resolves a workspace-scope destination through that snapshot, so installing into a worktree on a host connected moments earlier threw `skill-install-workspace-not-found`: the client asserting a remote workspace is absent on the strength of client-side bookkeeping that had never looked at the host. That is the shape `docs/reference/ssh-execution-boundary.md` rules out -- absence from a client-side set is not evidence about the execution host. It made `tests/e2e/ssh-skill-installation.spec.ts:108` fail 3 runs in 4 locally and deterministically in the Docker SSH lane, where connect-then-install lands inside the one-second window every time. The snapshot now carries the repo-registration revision it was computed under and is only reused while that revision still holds. The counter is the one `bumpLocalWorktreeScanGeneration` already advances on every repo add, removal and update, so the check is O(1) and cannot drift from the mutation sites. * fix(worktrees): key the snapshot on repo mutations only, not on generation reads Two things the headless-reattach lane surfaced. The revision I keyed the snapshot on was `generationSequence`, which `getLocalWorktreeScanGeneration` also advances when it mints a key for a repo id nothing has scanned yet. That is a read, not a mutation, so a read path could discard a snapshot that was still perfectly valid -- the mirror image of the staleness this fixes, and a way to make a lookup fail that would otherwise have succeeded. The counter now advances only where the scan generation is actually bumped: repo add, removal, update, and scan-cache invalidation. Separately, `pty-restore-record-seeding.test.ts` primed the cache by writing its private `resolved` field with a literal spelling out `worktrees`, `platformByRepoId` and `expiresAt`. That literal is a second copy of the cache's freshness contract, so adding a field to the real entry left the fake one failing the check: the primed snapshot was rejected, resolution fell through to a real scan, and the headless fixture -- which has no git -- got `selector_not_found`. It now primes through `getSnapshot` so the cache stamps its own entry and the two cannot drift again. The revision never moved during that test (0 before and after), so nothing was being invalidated; the fake entry simply never satisfied the contract. |
||
|
|
7c94d12190 |
fix(ssh): route four host-blind seams through the resolved execution host (#17919)
* fix(host-routing): resolve the execution host before reading a connection Three issues in one defect class: a resolver reads one spelling of one arbitrarily chosen row instead of resolving the worktree's execution host, so something local answers a question about a remote. returned that row's connectionId. With duplicate repo rows for one repo id it could pair a runtime owner with a client-owned SSH connection. It now resolves through the same ambiguity-aware index getRuntimeEnvironmentIdForWorktree uses, prefers the repo row for the host the worktree names, and derives the connection from the resolved host. Conflicting rows return `undefined` (this module's documented "cannot determine the host"), never `null`. `store.getRepo(worktree.repoId)?.connectionId ?? null`. `getRepo` is host-blind and the same repo id can exist on local, SSH and runtime hosts, so a remote worktree could spawn its PTY on the client with the remote cwd. resolveWorktreeLaunchHost picks the row for the worktree's host and reads the connection off that host; conflicting rows are unresolved, not local. session-partition owner maps that contradict each other. Both now compute through one shared function whose argument records the divergence. No behaviour change on either side: converging needs a read-both migration, since both partitions hold real data written by shipping builds. * fix(host-routing): keep nested SSH connections resolvable under a runtime host getRepoSshConnectionId read only the resolved execution host, so a repo row owned by a runtime that reaches a nested SSH target (connectionId: ssh-*, executionHostId: runtime:*) resolved to no connection — answering 'local' for a remote worktree, the same defect #17909 fixed in the other direction. * fix(host-routing): resolve both sides of the execution host through one rule The renderer resolver leaked between two different SSH hosts: a worktree on `ssh:m4air` whose only indexed repo row belonged to `openclaw` answered 'openclaw', because the host-scoped lookup missing fell through to an id-only one. Main's resolver, in the same change, answered 'm4air' — two resolvers, one right and one wrong, on identical input. Both sides now adapt one shared rule (`worktree-execution-host-resolution.ts`): the worktree's own host outranks every repo row, and a row on a different host is never evidence about this one. The renderer's WeakMap index becomes the memoizing adapter it always was; `resolveWorktreeLaunchHost` becomes main's mapping of unresolved onto its throw. Settles the rule the change previously answered two ways. `getRepoSshConnectionId` and `getSshTargetIdForExecutionHost` disagreed for a runtime host carrying a nested `connectionId`; they now compose, so the execution host is the single authority. On a `runtime:*` row that field is a paired HUB's private SSH target, spread through by `repoWithFetchedOwner` and unaddressable from this client — the project-first successor of the row nulls it for exactly that reason. That also fixes the `kind !== 'ssh'` fallback, which fired for `local`: a row declaring itself local handed out an SSH connection. * fix(ssh): resolve the execution host in the worktree scan and managed create The worktree scan and createManagedWorktree both picked remote-vs-local from repo.connectionId, so a row stamped only executionHostId: 'ssh:*' was scanned and created on the client against a remote path. The folder branch returns before the check, so its agent-trust write landed locally too. Refs #11163 * fix(ssh): stop over-rejecting and refusing SSH hosts the process owns runtimeRepoMatchesExecutionHost rejected an unstamped SSH repo against its own ssh:<connectionId>, so repo-add/clone dedupe could register a second row for a path the host already owns. assertHostIsSupported made the CLI/runtime RPC refuse --host ssh:* while the same process's IPC handler routed it correctly; setupExistingFolder now shares that registration. Clone still refuses, because nothing in this process clones onto an SSH host. Refs #11163 * test(ssh): retarget the SSH host-setup guard spec at the substitution it prevents setupProjectExistingFolder now registers the remote path through the same addRemoteRepoFromPath the desktop IPC uses, so it fails on the host's terms (connection not registered) rather than a categorical refusal. The local clone/probe side effects it exists to catch are still asserted absent. Refs #11163 * fix(cli): require an absolute path when setting a project up on an SSH host Routing --host ssh:* to the remote registration made relative paths newly reachable there, and they were resolved against the client cwd — registering a path that names the wrong machine. Refs #11163 * fix(repos): read the SSH registry directly so the runtime stays Node-bootable Routing runtime project setup through addRemoteRepoFromPath dragged ipc/ssh -- and its 25-module electron graph -- into the runtime bundle. ssh-target-registry already exists for exactly this; ipc/ssh only re-exports it. * fix(ssh): close the agent-launch and session-export host-blind twins Three sites left on the legacy spelling, all the same shape as the ones this branch already fixed: - `launchAgentTerminal` did `getRepo(worktree.repoId)` then wrote agent trust with that row's `connectionId`. Host-blind, so a repo id carried by two SSH hosts wrote a remote path into the *client's* Codex/Cursor/Copilot config and the agent on the host never saw the trust. Every sibling call site already passes the resolved `workspace.connectionId`; this was the last that did not. - `targetForWorktree` (workspace-session export) fell back to the same host-blind read, so a session could be published to a machine that never owned the worktree. Unresolvable ownership now exports to nobody. - `addRemoteRepoFromPath` minted `connectionId`-only rows while being the routing path this branch adds, so it kept creating rows in exactly the spelling the branch works around. It now stamps `toSshExecutionHostId(connectionId)` at creation; `reassignSshTargetId` already migrates both spellings, so target rename stays correct. Tests cover two *different* SSH hosts throughout — the case none of the earlier duplicate-row tests had, all of which were local-vs-ssh or runtime-vs-ssh. |
||
|
|
510305e574 |
fix(relay): signal capacity loss instead of dropping, hanging, or truncating (#17870)
Three failures with one shape: a payload past a fixed capacity was met with silence, with a wait that never ends, or with a prefix presented as a whole. **The workspace snapshot was silently dropped.** `workspace.changed` carries the tab/session list, and a snapshot past the producer frame capacity (12288 B on a Node <=21 remote) was dropped with only a relay stderr line, so the client kept a stale list forever. The relay now publishes per client and, for a client whose sink refused the frame, sends a compact `workspace.stale` marker on the control lane; the client re-reads through `workspace.get`, whose lane is budgeted in megabytes rather than in one producer frame. A new JSON-RPC notification rather than a new field on `workspace.changed`: `normalizeSnapshot(undefined, ns)` yields revision 0 and an empty session, so a Rule-1 field would make an old client replace its tab list with nothing — worse than the drop. An old client ignores the unknown method and is exactly where it is today. The marker retention/retry machinery is extracted from the `fs.changed` overflow path and shared by both. **The Windows upload hung, and the fix for it could truncate.** `#16432` was attributed to `[Console]::In.ReadToEnd()` materializing the base64 bundle. That is not what the reporter measured: he also measured `new IO.StreamReader([Console]::OpenStandardInput())` — an incremental reader — hanging at 1 MB. The limit is in the stdin the host hands PowerShell over a non-pty ssh exec, not in the string the script builds. - `uploadFileViaSystemSsh` — the user file-import path — was piping a whole file into one Windows stdin, unchunked and untimed. That is the path large files take; it now chunks into 32 KB writes and bounds each wait. - The Windows directory upload reuses that single-file path rather than repeating a weaker copy of chunk-read + write-buffer; the `ino`/`dev` TOCTOU verification comes with it. - A Windows write needing more than one exec lands on a `.orca-partial` staging path and is published by rename, so a failed chunk cannot leave a truncated artifact under the real name. `exclusive` is enforced once at the rename, not on the first chunk, where a retry met its own leftovers. - The mkdir batch reads stdin through the stream reader the reporter measured surviving 50 KB, not `[Console]::In`, which he measured wedging at that size. - `waitForChannelClose` takes an optional bound. A wedged PowerShell stays alive at idle CPU and never closes, so without one the promise is simply never settled and the caller waits forever with no error to show. **Quick Open showed a prefix as the whole workspace.** The mechanism "a full page means there is more" only works if the caller named the cap, and the failing UI named none — it hardcoded `truncated: false`. Quick Open now names `QUICK_OPEN_LISTING_MAX_RESULTS` on both the Electron IPC hop and the runtime-RPC hop (the field #17954 added to `files.listAll`), and reads a full page as truncation. The local hop honours the cap too, which it previously ignored. Rebase note on `fs.listFiles`: an earlier revision of this work also clamped the host unconditionally, and #17934 escalated an uncapped request to an explicit error. #17954 has since landed and made an oversized reply streamable, which removes the premise — the host no longer has to choose between a prefix and a refusal, so it returns the whole listing when no limit is named and only clamps a limit it was given. Keeping either would have regressed #17954 and hard-failed three in-tree callers that deliberately pass no options (`runtime-file-commands-search-runtime-files.ts:81`, `filesystem-read-handlers.ts:125`, `runtime-file-commands-constructor.ts:41`). |
||
|
|
7104056984 |
fix(watcher): route relay watch-root capacity refusals off the fast ladder (#17950)
* fix(ssh): stop two unrecoverable relay refusal loops A relay refusal that is a pure function of state the client cannot change was being retried forever, on two different paths. - pty.openClient: a superseded owner proof is refuted evidence, not a transient fault. The client kept re-presenting the identical proof, so every reconnect reproduced the same refusal until the relay was redeployed (#12895, #12931). It is now dropped exactly as a stale lease already is, and the claim re-asked without it. - fs.watch: the relay's watch-root capacity refusal was classified 'unavailable' and retried at 1 Hz per root for 60s, re-armed indefinitely. A folder workspace with more repos than the cap turns that into a permanent install storm scaled by the excess root count (#11196). It is now its own 'capacity' result that goes straight to the existing dormant backoff, mirroring what the local watcher path already does. * fix(watcher): route relay watch-root capacity refusals off the fast ladder A full watch-root cap is a decision, not a fault, so a 1 Hz reinstall per refused root only bills the relay the load that keeps the cap busy (#11196). Capacity refusals now go straight to the dormant backoff. The relay side no longer refuses on a slot it is about to hand back: an over-cap caused by roots still unsubscribing waits once on the teardowns settling — the release event, mirroring WatcherSupervisorCapacityWait — before it answers. A parked waiter is excluded from the accounting so it cannot take a slot from the root already reclaiming one. Drops the SSH owner-recovery half of this branch. Its premise — that a -32043 SUPERSEDED refusal is permanent — is false: the refusal fires only while the incumbent is 'active', and assertPtyConsumerOwnerRecovery explicitly admits the identical lower-generation proof once the incumbent flips to 'disconnected' (relay-pty-consumer-owner-displacement.test.ts proves it). The remedy could not work either: the proofless re-ask routes into refuseHeldPtyConsumerOwner, which is declared `: never` and, with sameClient true by construction, always throws. It would have traded one refusal loop for another, minus the checkpoints and minus the proof that resumes the claim once the relay reaps the incumbent. * fix(i18n): restore the activity-options key the rebase dropped * fix(i18n): union en.json with main so the rebase cannot drop keys |
||
|
|
31007c0d86 |
fix(ssh): reclaim relay PTYs the client has provably lost, on host attestation only (#17831)
* fix(ssh): reclaim relay PTYs the host attests this client orphaned (#9819) Orca could lose track of terminals running on an SSH relay until the 50-slot cap refused to open any more. This reclaims them, and the whole design is built around the fact that getting it wrong destroys a user's running process on their remote machine: the failure mode is leak, never kill. A stop requires all nine of: 1. the relay published an `ownerClientInstanceId` read from the live authenticated consumer grant of the connection that requested the spawn — never from a spawn parameter, since an echoed claim is no evidence; absent means skip 2. that id equals this client's persisted consumer identity 3. this connection holds the negotiated `session-owner` grant 4. `paneBound === true`, host-published 5. no `agentSessionOwners` — the host still advertises it as adoptable 6. `hostAgeMs >= 30s`, measured on the host's clock 7. this client has no route: not reattached, no lease outside terminated/expired, no pending kill, and no `expired` lease either — an expired lease is the record of a process deliberately left running, never a licence to kill it 8. every stop is fenced on the incarnation the same listing published, and on the owner identity, both re-checked by the host 9. a pass wanting to stop more than 8 refuses entirely Absence from a client-side set is `unverifiable` by construction (docs/reference/ssh-execution-boundary.md): a second machine attaches to the same relay and displaces the session owner, and its live agents are missing from this client's store for exactly the reason a genuine orphan is. So the host has to attest ownership, and the host has to attest that nothing is running. That second attestation is measured over the pane's whole tty, not its foreground process group. `tpgid == pgid` is foreground-only: on a real `bash -i` on a real pty, a shell holding `sleep 300 &` and a shell holding a Ctrl-Z'd job both read `pgid == tpgid`, `Ss+` — byte-identical to an idle prompt, with only the job's own row differing. A foreground-only gate therefore attests `pnpm build &` and a suspended editor as idle, and the stop that follows SIGKILLs every process group on the tty. `shellOwnsEveryTtyProcessGroup` is measured over that same set of groups, so the evidence and the kill describe the same thing. No new probe: `tpgid` already identifies the terminal, because a process group belongs to one session and a session to at most one controlling terminal. The freshness field is real rather than decorative. `capturedAgeMs` is stamped from when the capture was taken, deliberately as an upper bound since the process table is TTL-shared, and the sweep refuses an observation older than its own pass budget, counting its own elapsed time since the listing arrived. Stale evidence degrades to "do not sweep", never to "sweep". The display consumer of the same measurement keeps no age budget, as a stated decision: a stale pane title costs a redraw and self-corrects. `pty.shutdown` is authorized on the host that owns the process. `pty.spawn` and `pty.attach` both take a request context and check it; the one irreversible call took none, so the rule above lived entirely on the client that decided to make the call. It gains an optional `expectedOwnerClientInstanceId` and refuses unless the connection still authenticates as that identity AND this host recorded it at spawn. Finally, a reattach refusal now says whether it observed the process. Three refusals carry the same `SSH_SESSION_EXPIRED` text and only one is absence; `restoreRequired` means the PTY is live and only its source stream is not. Testing that text with `.includes()` expired the lease and deleted ownership for a running process, erasing this client's only record of it — and a PTY with no record is one the sweep may stop. Wire compatibility: four new optional fields and one new optional param on existing methods, no new method and no new stream opcode (Rule 1, and Rule 2 does not apply). Rule 1's caveat is discharged explicitly — no reader requires any of them, each absence is a named skip reason, and an ordinary pane teardown must omit the owner fence because a revived PTY carries no attested owner at all. New client plus old relay stops zero PTYs; old client plus new relay never reads the fields. Windows relay hosts publish no evidence and therefore never sweep. Verified by joining the real publisher to the real client reader over `ps` captured verbatim from a Linux container, and by driving a real group-for-group SIGKILL against a real pty: backgrounded and suspended jobs survive by pid, and an idle shell is still reclaimed, so the narrowed predicate is not a silent no-op. Squashed deliberately. The sweep is unsafe at every intermediate commit of its own history — before the foreground gate it reaps a hand-launched `claude`, and with a foreground-only gate it reaps a backgrounded build — so this ships as one commit with no bisectable state that kills live work. Refs #9819. Folds in #17939. * fix(i18n): restore the activity-options key the rebase dropped * fix(i18n): union en.json with main so the rebase cannot drop keys |
||
|
|
104f9655e4 |
perf(git): answer remote-URL questions from one subprocess, not one per remote (#18158)
Four copies of the same loop ran `git remote` and then a serial `git remote get-url <name>` per remote to answer "which remote has this URL". On a repo with 58 remotes that is 59 subprocesses -- measured at 1083 ms -- for one question, and worktree create asks it several times. `git remote -v` answers for every remote from one child, reporting the same insteadOf-expanded first fetch URL `get-url` prints. The batched `cat-file --batch-check` branch-conflict probe decides from stdout, but its WSL route was unfenced, so a login-shell fallback printed the distro banner onto the stream it parses. That broke the one-line-per-ref contract, made every batch undecided, and fell straight back to one `show-ref` per remote -- the cost the batch exists to remove. Measured at 58 remotes / 4346 branches, spawns and wall time: push-target remote scan 59 -> 1 (1083 ms -> 8 ms) branch-conflict probe 60 -> 3 (984 ms -> 43 ms) configured push target 123 -> 6 (2707 ms -> 157 ms) |
||
|
|
61e010079f |
New agent dashboard (#18222)
* more obvious toggle
* more obvious toggle
* feat(activity): redesign thread rows and add child agent filtering
- Emphasize task title and last activity in row layout over metadata
- Add child agent toggle; hide orchestration workers by default
- Support collapsible groups and ungrouped view mode
- Improve orchestration worker message handling to surface replies
- Add sidebar search and filter controls for agent activity
* periodic checkin
* feat(activity): add "Clear completed" action and performance improvement
- Add "Clear completed" action for activity threads with undo window; clears completed and interrupted rows from view, persists across restart
- Virtualize activity thread list to render only viewport-bounded rows
- Cache activity thread search text to prevent recomputation on every keystroke
- Cache dashboard bucket counts per-worktree for selective invalidation on unrelated changes
- Use useDeferredValue for activity search filtering to keep input responsive
- Make compact mode the default display for activity threads
- Add activity-cleared-at persisted state tracking (per-pane cutoff timestamps)
* improve style
* minor change
* feat(activity): add persisted host and project filters to agents view
Agents scope filters are deliberately separate from workspace-nav filters so a monitoring surface never inherits workspace context silently. Filters survive restarts and always display an active-filter chips row with hidden count, making filtering visible and reversible.
* Graduate Agents view from experimental, refine activity handling
- Agents Dashboard moves from experimental to standard feature with showAgentsSidebar setting controlling visibility
- Add identity-checked cache eviction (dropPersisted IPC) to prevent newer runs from being evicted when UI clears older status, fixing clear-completed safety
- Extract ActivityThreadHoverCardSummary and ActivityThreadListToolbar components for better organization and reusability
- Implement mark-thread-read as separate action from select with clickable bell icon
- Add hasActivityThreadWorkspace helper for checking workspace availability across hosts (SSH/runtime targets)
- Preserve scope filter array identity during hydration for memo optimization
- Track manually-unread turns in auto-ack to prevent re-acknowledgement
- Clean up activity cleared-at cutoffs on pane retirement
- Remove activity-thread-hover-card max-lines lint override (code refactored below threshold)
* Refactor agent cache identity to use timing fields only
- Simplify AgentStatusCacheIdentity: keep only paneKey, receivedAt, stateStartedAt
- This fixes silent no-ops where renderer-enriched fields diverged from main's cache
- Add worktree-jump-navigation for navigating activity to workspaces
- Add manual mark-unread protection separate from auto-ack
- Optimize activity owner resolution with per-build memoization
- Optimize detected worktree lookup with indexed search
* Remove sticky header, add scroll position persistence
Replace the floating sticky header overlay with scroll position memory via
a ref. This preserves the user's scroll location when switching between
threads or remounting the agents list, improving UX without requiring
React state.
* Implement sticky group headers in activity thread list
Keep group headers visible at the top while scrolling when threads are grouped. Headers stick to the viewport while their section is in view, then unstick as the next header approaches.
* add blue flash
* update settings appearnce
* Extracted activity acknowledgement/clearance actions from the oversized UI slice.
- Removed dead sidebar search/menu props and the unused search ref.
- Removed the unnecessary sidebar visibility bitmask.
- Replaced hardcoded sidebar toggle colors with design-system tokens.
- Removed duplicate “mark all read / clear completed” controls in the sidebar.
- Preserved manual-unread state correctly across pane retire, transfer, and drop.
- Made clear-completed cutoffs monotonic so clock skew cannot resurrect old activity.
- Fixed blank workspace names in hover cards with the existing fallback helper.
- Added missing localization entries and stabilized hydrated filter array identity.
- Updated misleading Agents setting copy to describe both sidebar surfaces.
* add onboarding guide for the new agents panel
* Add activity clearance tracking and synced agent view settings
Agent view filters and presentation settings now sync across paired clients.
Preserves per-pane activity clearance cutoffs in persistent state. Improves
activity thread row accessibility with proper ARIA roles, and preserves
terminal host ownership after pane teardown via retained terminal handle.
* rm html
* Graduate Agents from experimental and improve activity visibility
- Migrate `showAgentsSidebar` setting from legacy experimental flags; default new profiles to the agents sidebar
- Replace scoped-thread filtering with visible-thread filtering so bulk actions (mark all read, clear completed) only affect rendered rows
- Rewrite child agent classification as a set of visible pane keys to fix orphan promotion and parent-cycle handling
- Improve activity cleared-at cutoff lifecycle: preserve on row dismissal (pane may still be live) but clear on pane removal
- Add pagehide flush for pending clear-completed evictions so quit/reload cannot replay cleared activity
- Polish agents sidebar: unread count badge, expand button, onboarding intro for migrated/new users
- Extract shared time-ago formatting to a library module
- Fix scroll restoration to defer until content can contain the saved offset
- Improve stable message hold for compact agent rows using state instead of refs
- Add worktree filter-visibility check to distinguish collapsed-but-unfiltered from filtered-hidden
* Graduate Agents from experimental and improve activity visibility
- Remove the deprecated full-page Agents view; fix settings navigation fallback
- Refactor bulk action bindings and separate mark-all-read from visible threads
- Preserve sidebar collapse state across remounts; fix child-agent badge filtering
- Add safety window for scroll-restore and improve worktree host-qualified filtering
* Graduate Agents from experimental and add manual unread tracking
- Move Agents sidebar from experimental settings to standard feature with intro flow
- Add persistent manual unread turn tracking for activity feed
- Consolidate workspace activation through activateAndRevealWorkspace dispatcher
- Improve sidebar view toggle with radio semantics and arrow-key navigation
* Graduate Agents sidebar and separate dashboard experiment
The Agents tab now has its own `showAgentsSidebar` setting (defaults on) independent from the dashboard popout experiment. Activity unread counting is simplified to count all events uniformly without mode-specific filtering. Dashboard visibility is now controlled solely by `experimentalAgentDashboardPopout`, with its own UI in the Experimental settings pane. Migration path updated: only `experimentalActivity=true` graduates to the sidebar; the dashboard experiment remains separate.
* Add agent-session tab support to activity tracking
Build activity event contexts from structured agent-session tabs and
worktree-attributed status entries. When activating a thread, try
agent-session tab activation before falling back to terminal pane.
* • The workspace sidebar tab is now a static Spaces
label—no grouping-based “Projects” label or hidden
width-reservation span.
* Show unread count badge and prioritize attention-needing agent threads
Activity group order now surfaces threads needing attention (blocked,
waiting, interrupted) before working/done so they're never buried. The
Agents tab shows an unread count badge while viewing Spaces, since the
open Agents list already highlights unread rows.
Also improves UX text ("Hide Agents" vs "Maybe later"), accessibility
with proper ARIA labels, and handles edge cases: preserves read state
for retained panes on SSH reconnect and handles deleted worktrees
gracefully in navigation.
* Batch agent-status evictions and optimize activity pane rebuilds
- Add dropPersistedStatusEntries batch API; consolidate evictions into one persist
- Implement fallback timeout in clear-completed for unseen toast callbacks
- Project only activity-relevant tabs; memoize terminal tab derivations
- Stabilize activity virtualizer key to prevent unnecessary item measurements
* Remove unread count badge from Agents sidebar tab
Simplify useActivityUnreadCount by removing the enabled parameter and
conditional logic, as the badge is no longer displayed in the UI.
* Deduplicate activity unread counts across source overlaps
Live pane status is the primary source; retained and migration entries
serve as fallback caches that may briefly overlap it during lifecycle
transitions. Count each pane only once by tracking seen keys, prioritizing
the live status as the canonical source.
Also fix monitoring state display: it's a distinct agent state, not a
tool-running row state, so exclude it from tool preview checks.
* Update activity pane tests to remove unread badge assertions
- Remove ActivityPaneVisibility type and readActivityPaneVisibility() helper
- Update agentsSidebarButton selector to match badge-less state
- Simplify assertions to check pane focus instead of visibility isolation
- Remove test for unread badge acknowledgement flow
* Fix activity pane workspace resolution and localization handling
- Thread defaultHostId through activity operations for correct host resolution
- Add language-aware caching for standalone terminal names with cache invalidation
- Fix scroll restoration bounds calculation for tall viewports
- Add focus management to sidebar radio group keyboard navigation
- Refresh localized sidebar content on language changes
- Preserve activity state across heartbeats to prevent history loss
- Improve host-id strictness in worktree jump navigation
* Preserve activity view when settings fetch fails
A failed window.api.settings.get() leaves settings null, which was
incorrectly treated as opt-out. Add the missing null check so the
activity-view gate only applies when settings are available.
Includes tests for this scenario and related edge cases in keyboard
navigation, worktree jumping, and session state handling.
|
||
|
|
4bc20cb842 |
fix(wsl): name an explicit Windows cwd for wsl.exe spawns (#17834)
* fix(wsl): name an explicit Windows cwd for wsl.exe spawns Removing the worktree Orca was launched from broke every wsl.exe spawn for the rest of the session. The WSL command builders passed `cwd: undefined` meaning "the directory is inside the command" -- but CreateProcessW reads NULL as "inherit the parent's", and the parent's was a \\wsl.localhost path Linux had just deleted. Fixes #16463 * fix(wsl): name the spawn directory at the six remaining wsl.exe sites The first commit fixed the WSL command builders. Six spawn sites were left inheriting the process cwd, which is the same deletable `\\wsl.localhost` worktree: `wsl-availability` (both probes), the WSL filesystem watcher, the agent-hook relay launch, the UNC delete, and the local worktree filesystem. `wsl-availability` is the one that matters most, and it turns the bug into a latching false negative. `isRetryableWslProbeFailure` returns false for ENOENT, so a spawn that failed only because the inherited cwd was gone is cached as "WSL is not installed" on the 10-minute definitive TTL with exponential backoff up to 30 minutes. Git keeps working and Orca reports WSL unavailable -- worse than the bug being fixed. ENOENT stays non-retryable. It is answer-shaped for the reason it is meant to be -- wsl.exe is not on PATH -- and naming the directory is what removes the one cause that was not. Making it retryable would instead re-probe every non-WSL Windows machine on the short window, and would leave the false ENOENT in place for the other five sites, which have no cache to correct. Three of these are also on the `runWslProcess` W3 migration allowlist; this is the interim until they move, and matches what #17837 does inside the runner. |
||
|
|
f9db653e14 |
perf(worktrees): gate worktree metadata hygiene on evidence, not on every listing (#18034)
* perf(worktrees): gate worktree metadata hygiene on evidence, not on every listing Dangling `worktreeMeta` pruning rode the detected-worktree listing, a polled read path. Each pass captured a prune expectation over the repo's whole metadata table (a JSON.stringify per row) and then stat'd every path-missing candidate. Both are O(all rows), and most rows are refused anyway — pinned by a persisted session, or structurally unremovable on this host — so the work repeated forever without converging, pinning the main process in fs completion callbacks (#17775). Three changes, no behavior lost: - Probe only rows a delete could still accept. Session ownership and structural removability are pure functions of persisted state, so deciding them before the filesystem inverts the cheap and expensive halves. The filter is advisory; the authoritative checks are unchanged, so it can only shrink the stat fan-out. - Extract `isLocallyRemovableWorktreeMetadataRow` so probe-avoidance and the delete share one definition of removability. - Gate the metadata + lineage prune on evidence instead of the listing: a worktree lifecycle event, a mutation that can make a row more removable (session-owner release, metadata removal, SSH lease release, automation run finishing or deletion, repo deregistration), or a git listing that differs from the one the last pass ran against. With none of those the pass is a provable repeat and is skipped, so a quiescent app does no hygiene work at all. The gate deliberately ignores metadata writes that only add or update a claim: the listing path itself stamps metadata, so re-arming on those would restore the storm. A missed signal leaves a row in place until the next one; nothing is deleted that would not have been deleted anyway. * refactor(worktrees): fold repo prune-gate teardown behind one call Merging both import blocks during the rebase pushed the file past the 300-line budget. The two calls are one intention -- retire this repo's gate state on a full removal, and re-arm the shared inputs either way -- so name that in the module that owns the gate. |
||
|
|
b4ba3e97ff |
perf(worktree): defer fork-PR remote creation from create-time to first use (#17922)
* perf(worktree): defer fork-PR remote creation from create-time to first use Fork-PR review worktrees eagerly ran `git remote add` + `git fetch` for the contributor's fork (and pinned branch.<x>.remote) at create time, even for a read-only review. That grows remote count unboundedly with review volume and pays a network fetch nobody asked for yet. Defer prepareWorktreePushTarget(Ssh) and the --set-upstream-to configure step at create time (local + SSH, IPC + runtime create paths); persist the pushTarget metadata untouched. Materialize the remote on demand the first time push/pull/fetch/fast-forward actually needs it, via two shared functions (materializeWorktreePushTargetRemote(Ssh)) reused across the legacy IPC handlers and the RPC runtime sync commands. A cheap `remote get-url <name>` probe keeps steady-state calls down to one extra subprocess once materialized, instead of repeating the O(remotes) scan. Add repo-local `remote.<name>.orca-created` config provenance, written when the remote is added, so cleanup can recognize ownership of a remote that was lazily materialized (and therefore never round-tripped through the store's `remoteCreated` flag). Refs #17828 * perf(worktree): materialize a deferred fork-PR remote on terminal spawn An agent running raw git in a freshly opened fork-PR review terminal has no usable upstream until an Orca-driven sync happens -- "sync through Orca first" isn't available mid-task, and git pull/log @{u}.. hard-fail without one (verified against real git). Fire the same on-demand materialization used by push/pull/fetch/fast-forward from the single terminal-spawn resolver (resolveTerminalWorkspaceLaunchTarget), fire-and-forget, so a newly opened terminal gets a working upstream without blocking spawn. * fix(worktree): retest deferred fork-remote CI failures, fix SSH provenance-marker RPC Rewrites the 5 CI failures on the deferred fork-remote change (#17828) as evidence, not fixtures: the SSH relay-upgrade/rollback/sibling-ownership tests move to materializeWorktreePushTargetRemoteSsh, where that unchanged logic now actually runs (create defers it to first sync). While writing a stricter test that routes its mock exec through the relay's real validateGitExecArgs, found that the SSH provenance-marker write (`git config remote.<name>.orca-created true`) was unconditionally rejected by the relay's generic git.exec (it blocks all non-read-only config writes) -- a real bug that would break every SSH fork-remote materialization against a live relay. Fixes it with a narrow git.markRemoteOrcaCreated RPC, mirroring renameCurrentBranch, with a graceful no-op fallback for relays that predate it. * fix(worktree): scope post-#17887 test assertions past narrow-refspec config calls Rebasing onto #17887's narrow-refspec `remote add` broke two broad `['config']` call-filters into false positives/negatives, and the local materialize test still asserted the pre-#17887 wide `remote add`/fetch-refspec forms. * fix(worktree): restructure upstream restore, persist provenance, widen short-circuit refspec (#17828 review) - Move upstream restoration to the materializer level so it runs on both the remoteAlreadyMatchesUrl short-circuit and the full-prepare path, not just buried inside prepare*. - Persist {remoteCreated, remoteName} to the store on materialize so #17842's orphan sweep can see a lazily-created remote, including via desktop IPC, terminal-spawn, and the RPC host-callback paths. - Widen the refspec on the local short-circuit path too (SSH's bare `remote add` refspec gap remains a documented, pre-existing limitation). - Fetch the branch's tracking ref before restoring upstream when the short-circuit widens onto a *new* branch on an already-existing remote -- a bare refspec-config widen never itself imports anything, so `branch --set-upstream-to` was hard-failing for a sibling worktree's first materialize (found via a real-git fixture, not just mocked unit tests). Skipped when the ref already exists so the common repeat-call case stays a local-only probe with no network round-trip. * fix(worktree): merge duplicate shared/worktree/types import oxlint --deny-warnings flags the split import as no-duplicates; full pnpm lint was failing on it after the #17828 review restructuring. * fix(worktree): scope the deferred fetch timeout to fetch calls, retarget stale create-time assertions CI on the previous push failed 3 shards, all argument-shape mismatches: - worktrees-wsl-runtime-routing.test.ts: the "restructure upstream restore" commit wrapped every call `prepareWorktreePushTarget` makes (remote, remote add, config, fetch) with DEFERRED_PUSH_TARGET_FETCH_TIMEOUT_MS, not just the network fetch. Local git subprocesses never need a timeout; scope it to `args[0] === 'fetch'` only, matching the short-circuit path's existing pattern. Updated the test to expect the timeout on the fetch call specifically (point 5 legitimately adds it there), while every other call stays untimed. - worktrees-create-metadata-persistence.test.ts (2 tests): stale from before this session -- create no longer mints a fork remote at all (#17828 deferred that to first sync), so asserting `remote add`/`fetch`/`remoteCreated: true` at create time no longer matches reality. Retargeted both tests to assert the deferred contract (no remote add at create, pushTarget persisted unmaterialized); minting itself stays covered by worktree-remote-push-target-materialization.test.ts and worktree-push-target-setup.test.ts. Re-verified all 5 fixture points (mint upstream, store persistence, single-flight, short-circuit refspec widen + fetch-missing-ref for local and SSH, finite timeout) against a real git fixture after this fix -- all still pass. * fix(worktree): hook pty:spawn into deferred push-target materialization (#17828) triggerTerminalSpawnPushTargetMaterialization only fired for agent/background/ mobile terminals; the desktop GUI's own pty:spawn path (new tab, split, reattach) never materialized a deferred fork-PR remote before raw git commands could run there. Add a small wrapper that resolves the worktree's push target and owning repo from args.worktreeId via the store, and fire-and-forget delegates to the existing materializer, wired as the first statement of runPtyIpcSpawn. Degrades silently (optional chaining + catch) so a partial/fake Store in existing spawn tests can't turn this into a spawn-blocking throw. * test(worktree): retarget stale editor-remote-branch assertions for worktreeId threading runtime-git-sync-client's local-path fetch/pull/fastForward/push calls now forward context.worktreeId (needed by the main-process handlers to key deferred push-target materialization). Update the 17 call-site mocks across 15 tests in editor-remote-branch-actions.test.ts to expect worktreeId: 'wt-1', matching the already-correct source behavior -- no assertion was loosened. * fix(worktree): give a materialize joiner its own branch wiring The materialize single flight is keyed on the remote, but everything after the remote add is per-branch. A sibling worktree joining an in-flight mint for a different branch received the minter's target and skipped its own refspec widen, tracking-ref fetch, and upstream link, so its branch ended with no upstream at all. Wait for the remote, then run the per-branch work against the joiner's own target -- the same path the already-exists short-circuit takes, now shared rather than duplicated. Adopting a remote a sibling minted also stamps ownership, so removing the minter cannot strand the survivor's metadata outside the orphan sweep's reach. * fix(worktree): stop a failed mint from leaving a config-only fork remote Review of the joiner fix found it made things worse in three ways. Swallowing the mint's rejection let a joiner adopt a remote the rollback had already removed, writing remote.<name>.fetch with no URL. Verified on real git: that ghost section breaks `git fetch --all`, forces every later mint to a `-2` name, and cannot be removed by `git remote remove`. Propagate instead; the in-flight map is already cleared, so a retry re-mints. The SSH twin still returned the minter's target to a joiner, so the original per-branch bug survived there. It now adopts against its own target through a twin helper. The ownership stamp was unreachable: it required both a store and a repo id, and no caller passes both. Derive the repo id from the worktree id. Adopters also write remote config, and concurrent `git config --add` has no lock retry -- 135 of 160 writes failed at 8-way concurrency, and equal values duplicate the refspec. Chain adoptions per remote. |
||
|
|
7dd2ff586a |
fix(ssh): stop expiring relay-reset leases when the force-stop threw (#17962)
A force-stop that rejected never observed the remote shells, so bulk-expiring their leases in the finally block recorded a verdict Orca does not hold. Mirror ssh:terminateSessions: only a fulfilled stop retires a lease. Local PTY handles are still cleared, so nothing is stranded — the next connect reattaches the survivors or expires them on host evidence. |
||
|
|
058e618bb4 |
fix(ssh): stop a failed worktree scan from publishing authoritative emptiness (#17833)
* fix(ssh): keep an unreadable worktree catalog from authorizing teardown #14004: the relay's worktree-list fallback caught every failure and returned `[]`, so `SshGitProvider.listWorktrees` resolved as a success with an empty list. Downstream reconciliation treats a resolved listing as authoritative, which reaches `teardownMissingWorktreeTerminalsBestEffort` and the unregistered-worktree removal paths — a data-loss path from a failed scan. - relay: the `-z`-unsupported fallback lane propagates its failure instead of swallowing it to `[]`. - provider: an empty or malformed `git.listWorktrees` response is refused as `WorktreeCatalogUnavailableError`. A Git repo always lists its own checkout, so a zero-row listing can only be a scan that never answered — this is the mixed-version guard against relays that still swallow. - `listRepoWorktrees`: an unreachable SSH host reports unavailable instead of an empty catalog. #12661: `ssh:terminateSessions` now returns `{ terminated, unverifiable }`, so an offline sweep that only tore down local transport cannot be mistaken for a remote kill. The Manage-hosts toast warns instead of claiming success. * chore(i18n): register the unreachable-terminal terminate message |
||
|
|
bed9734a9d |
Prevent deleted workspace browser snapshot resurrection (#17779)
* Prevent deleted workspace browser snapshot resurrection * fix: tear down folder workspace browser tabs * fix: fence pre-publication browser snapshots * fix: route folder deletion through runtime cleanup * chore: retrigger CI * fix: sweep folder PTYs on runtime deletion * fix: restore deletion fences after runtime refactor * test: cover deleted renderer snapshot after recreation * fix: avoid publishing ambiguous worktree snapshots * fix: preserve optional worktree index state * fix: fence paired PTYs on worktree removal * fix: harden deletion fence and folder-delete teardown - Folder-group delete no longer fails on a mixed-host group: an ambiguous connection skips the PTY sweep instead of rejecting the delete. - Share one folder-workspace PTY teardown helper between the runtime removal path and the project-group controller. - Simplify the mobile snapshot fence: identity-carrying frames are judged against the live catalog instanceId and clear the fence once the successor is accepted; identity-less frames are fenced by renderer generation. Drops the unbounded epoch bookkeeping. - A fenced frame no longer triggers a resync request on every sync while the renderer still lists it as unchanged. - Cross-host id collisions publish without an instanceId rather than blanking the mobile session for that workspace. - Folder delete IPC always routes through the runtime; the store-only fallback and double notify are gone. - Drop the redundant rescue-path tombstone check; ownership is purged at removal. - Fence tests drive removeWorktreeMetadataAndHistory + syncWindowGraph instead of seeding the fence map, and add accept-after-recreate, no-resync, and ambiguous-host folder delete cases. |
||
|
|
6c8eea5ebe | perf(worktree): fix the prepared-checkout hit rate and make misses visible (#17863) | ||
|
|
7a69357856 | fix(worktree): widen git-common watch on event-batch overflow (#17916) | ||
|
|
a7db6c336b |
perf(git): skip the sparse probe for worktree listings that never read it (#18050)
Three main-process call sites list a repo's worktrees to read `worktree.path` and nothing else, but went through the annotated listing, so each one paid a sparse-checkout probe per worktree and cached the result nobody consumed: - `registered-worktree-roots-cache.ts` rebuilds the filesystem-auth authorized roots. `invalidateAuthorizedRootsCache()` fires on every worktree create and remove, plus repo add/clone/settings changes, so this reruns constantly. - `filesystem-source-control-ai-targets.ts` checks whether a local repo owns a worktree path. - `hosted-review.ts` verifies a worktree belongs to the repo before granting access. The probe is an `fs.stat` of the per-worktree `info/sparse-checkout` plus, when that file is non-empty, a git config read. On a WSL-hosted repo both cross 9p. #17859 cached it and #17932 keyed that cache on the distro, which fixed a wrong answer but also meant the distro-less callers above populate a second entry per worktree — probed cold, revalidated on their own five-minute loop, and read by nobody. Worktree create/remove clears the sparse cache and dirties the roots cache together, so both variants go cold at once and the discarded half is re-probed in full on the next auth check. `listRepoWorktreeGraph` routes those callers to `listWorktreeGraph`, which already existed as the annotation-free listing (#17655). Doing only that would have cost a second `git worktree list`. The scan cache keys in-flight scans on a `kind`, and graph and lenient were separate kinds, so a roots rebuild overlapping a sidebar refresh would spawn its own subprocess where the two previously coalesced. That is a real regression on macOS, Linux and native Windows, where `getLocalProjectWorktreeGitOptions` returns `{}` and both callers land on the identical key; on WSL they already differ by distro and never shared. So the annotated listing is now the graph listing plus annotation, rather than a parallel scan of its own: `listWorktrees` awaits `listWorktreeGraph` and annotates the rows it returns. Both soften a Git failure to `[]`, so they can share one listing; strict keeps its own because it must be able to reject. The two kinds ran Git twice before and now run it once, so the overlap case gets strictly faster instead of paying for the opt-out. An annotated scan holds two in-flight entries now (its own, plus the graph listing it shares). Keeping its own entry matters: `detectSparseCheckoutCached` dedupes revalidation but not the initial fill, so two concurrent badge readers sharing only the graph scan would both probe. Per-platform delta: - macOS/Linux: fewer probes on the three call sites; one `git worktree list` instead of two when a graph and an annotated scan overlap. - native Windows, no WSL: same, and the saved subprocess is the expensive half. - Windows + WSL: the largest win. The discarded probes were 9p round-trips re-paid cold after every worktree create/remove. - SSH/relay: none. `listRepoWorktreeGraph` returns through the same provider branch as `listRepoWorktrees` before reaching local Git. - folder workspaces: none. Both return the same synthetic folder worktree. Not in this change: - The badge listing itself. It still probes, still annotates, and still keys on the distro exactly as #17932 left it. - The remaining `listRepoWorktrees` callers. They read `isSparse`, or feed rows to something that does. |
||
|
|
a7fda48fe3 |
feat(telemetry): measure macOS stale-daemon adoption and cwd denials (#18043)
* feat(telemetry): measure macOS stale-daemon adoption and cwd denials Adds two enum-only PostHog events so #17696 can be sized instead of guessed at: - daemon_adopted: once per macOS launch that keeps a daemon an earlier app launch forked (invisible to daemon_lifecycle, which only sees replacements). Carries app-version match, spawner-path class (installed app / Squirrel ShipIt cache / other / missing), the existing TCC attribution verdict, and the bucketed live-session count. - daemon_pty_cwd_denied: the symptom itself. The daemon probes the requested cwd in its own process (only its TCC context counts) and returns an additive cwdReadableByDaemon field; the app emits only when the daemon was denied AND the app can read the same path, so a missing or genuinely unreadable cwd never counts. Non-permission errors read as readable on purpose. Both emitters swallow every failure; nothing here can delay or fail daemon startup or a PTY spawn. Off macOS neither event fires. The new wire field is optional, so older daemons and clients are unaffected. * fix(telemetry): keep cwd-denial classification inside the swallow guard Read the pid record at emit time (inside the try) rather than passing the adapter's startup snapshot: a throwing app-environment read can no longer escape spawn(), and a denial after a respawn is billed to the daemon that actually spawned the PTY. |
||
|
|
d7123591ce |
perf(git): pack the loose refs Orca's own fetches leave behind (#17857)
* perf(git): pack the loose refs Orca's own fetches leave behind Orca strips git's auto-maintenance off every fetch it issues (GIT_FETCH_SKIP_AUTO_MAINTENANCE_CONFIG_ARGS) and never compensated, so nothing in an Orca-driven checkout ever packs refs. One real machine reached 36,574 loose refs, where `git show-ref -- main` costs 5.2s and every worktree create pays for it. Add an idle-time, per-repo `git pack-refs --all --prune`, armed by the fetches that create the debt. It runs only after ten minutes of quiet on that repo, only above 1000 loose refs (probed with a walk bounded by that threshold, not by the backlog), one at a time across the whole app, at the background admission tier, and never while an agent is working, a create is prepared or in flight, a worktree removal is deleting refs, the app is quitting, or the machine is on battery. A user who set `maintenance.auto=false` or `gc.auto=0` has opted out. Measured on a 36,001-loose-ref fixture (macOS/APFS, git 2.44): `show-ref` 5.5-12.2s -> 30-49ms, `for-each-ref` 4.0-10.8s -> 43-48ms. Also fixes a pre-existing bug the split exposed: `--path-format=absolute` is ignored before git 2.31, and taking rev-parse's stdout raw collapsed every repo on such a host onto one fetch-serialization key. Refs #17828 * perf(git): make idle ref maintenance preemptible and cheaper to probe The idle veto was one-directional: it stopped a pack from starting during a create, removal, or agent work, but nothing stopped those from starting during a pack. A user-clicked Fetch, a branch delete, or a worktree removal that needed `packed-refs.lock` mid-rewrite could fail with `unable to create packed-refs.lock` -- a git error with no visible cause. Make the pack cancellable end to end. An AbortSignal now reaches the `pack-refs` child and both pre-pack probes, and `pause()` aborts what is running, waits for it to actually stop, and holds a suspension count so nothing new starts until the caller releases. Every entry point that deletes a ref takes that pause: gitFetch, gitPull, gitFastForward, removeWorktree, forceDeleteLocalBranch, prepareWorktreeCreateCheckout, addWorktree. Five more triggers close the rest of the window: battery drop, window focus, quit, the attempt deadline, and any other git command queueing for an admission slot. Judge a pack by re-probing the backlog rather than by the child's exit code. Measured in the field: another Orca session moved a branch mid-pack, git reported `cannot lock ref`, skipped that ref and packed the rest -- 36,688 loose refs down to 3. On a machine running several sessions that is the normal case, and retrying it would be wrong. Probe with one batched `readdir` per directory instead of streaming `opendir`, which issues a thread-pool round trip every 32 entries: 177ms -> 23ms on a real 36,600-ref repository, with half the event-loop lag. The walk stays strictly sequential so it can never occupy more than one of libuv's four filesystem threads. `PackRefsLockOwnership` makes a lock left by SIGKILL attributable, and only reclaims one when a marker exists, the lock is older than any pack-refs could run for, and the recorded process is gone. Refs #17828 * fix(git): wait out the packed-refs lock instead of killing the pack Measured on Git 2.55/APFS with 37k loose refs: a full `pack-refs --all --prune` takes 23-32s but holds `packed-refs.lock` for only 0.03-1.37s of it. The other ~95% is the prune phase, during which a concurrent `fetch --prune`, `branch -D` or `update-ref` succeeds every time -- per-ref locks last microseconds and git retries for `core.filesRefLockTimeout`. So the abort-on-everything design was strictly harmful. SIGTERM into the prune loop strands an empty `refs/**/*.lock` about one time in five (9/30, 5/40, 6/30 kills): `tempfile.c` opens the lock O_EXCL before `activate_tempfile()` links it into the list the signal handler walks, and a pack does ~36k lock cycles. Afterwards `update-ref -d` on that ref fails with `cannot lock ref ... File exists`, permanently. On Windows `taskkill /f` never runs git's handlers at all, so an abort inside the rewrite strands `packed-refs.lock` every time. Never signal the child. `packRefs` no longer takes an abort signal; it polls `packed-refs.lock` and reports the window through a `PackedRefsLockReporter`. `pause()` resolves when the lock is released -- bounded, and free during the prune -- while the suspension counter still blocks new attempts. Battery and window-focus become do-not-start rather than stop-what-is-running, and quit waits for the lock and lets the child finish orphaned. For strands that already exist, `PackRefsLockOwnership` now also reclaims `refs/**/*.lock` under the same three conditions plus a 0-byte check, and a lock carrying our own not-yet-reclaimable marker records `locked` with a 30min retry instead of the 6h failure cooldown -- so a Windows strand self-heals in half an hour rather than six. Reverts the git admission-scheduler event bus, which existed only to drive the abort this removes. Refs #17828 * test(git): make the ref-maintenance waits survive a loaded runner CI shard 4/8 failed on `restarts every armed countdown when the user does ref work themselves`, which passes locally. The `until()` helper spun a fixed 200 event-loop turns and then returned silently, so on a contended runner the filesystem probe had not finished and the assertion that followed failed with an unrelated message. Bound the wait by wall clock instead and throw a named error, which immediately exposed a second latent bug: the single-flight test's second wait could never succeed, because the deferred repo's retry is on a faked `setTimeout` that spinning the real loop never advances. It had been passing only because the old helper gave up quietly. Add a timer-aware variant for those, and have the countdown test await a signal the fake pack resolves rather than polling at all. Verified stable across five sequential runs and once under load average 32 with six concurrent suites. Refs #17828 |
||
|
|
8b7d778a2e |
perf(git-common): bound the fs-stat fan-out in the worktree pollers (#17839)
* perf(git-common): bound the fs-stat fan-out in the worktree pollers snapshotGitCommon and snapshotBase issued one fs op per candidate via Promise.all/a serial loop, unbounded by worktree count. At 973 live worktrees this queued ~6,800 concurrent stat calls (measured peak 6000 in a 1000-entry synthetic benchmark) onto libuv's 4-thread default pool, starving every other main-process fs operation for the scan's duration (~1s). Bound both to concurrency 8 via the existing forEachWithConcurrency helper, matching the precedent in exact-ref-probe.ts and worktree-head-identity-reader.ts. Peak concurrent stats dropped 6000 -> 48 in the benchmark; wall time was essentially unchanged (495ms -> 541ms), since the real bottleneck was never total scan time but pool starvation of unrelated work. Also make the no-native-watch and crash-fuse polling fallbacks in worktree-git-common-watch.ts / worktree-git-common-narrow-watch.ts self-calibrate their cadence: on platforms/paths where this poller is the sole change signal, a fixed 2s cadence at hundreds of worktrees approaches a permanent scan loop. Stretch the interval so a scan stays a bounded fraction (10%) of its own cadence, capped at 30s, floored at the configured base interval. Left the reconciliation backstop (fixed 30s cadence, already accepted) and checkPendingMarkers (bounded by concurrent-worktree-creation count, not total count) untouched. Fixes #17828 * perf(git-common): split the tripwire from the per-entry sweep cadence Review on #17839 found a real staleness trade-off: adaptiveCadence gated ALL detection (worktree add/remove, HEAD, dirty refs, AND per-entry commit signals) behind one stretched interval, so on the crash-fuse polling fallback the reviewer measured cadence sitting at 5.4-10s sustained and hitting the 30s cap once a single scan reached 3s at 973 worktrees -- worse than the pre-#17828 fixed ~2s+250ms baseline for signals users notice immediately (sidebar worktree list, branch labels). Split snapshotGitCommon into a cheap structural "tripwire" (readdir, worktreesDir signature, primary-file signatures, newly-appeared entries -- ~5-6 fs ops, O(1) in worktree count) that always runs on the fixed pollIntervalMs, and the O(n) per-entry sweep (commit/dirty detection) that alone is gated by the adaptive cadence via a nextSweepDueAt deadline. Existing, unchanged entries are carried over by reference on a tripwire-only tick (no re-stat), so diffing produces no spurious events; genuinely new entries are still stat'd immediately so worktree add remains real-time. This keeps everything on one ticking-flag-guarded loop (no new concurrency/race surface) -- scheduling stays fixed at pollIntervalMs; only nextSweepDueAt stretches. Also drop the adaptive-cadence seed heuristic entirely: nextSweepDueAt starts at 0, so the first regular tick after bootstrap sweeps unconditionally on its own schedule instead of guessing an initial interval from the bootstrap snapshot's duration (which could stretch the very first tick to 10-30s on a slow disk). Documented that worktree-git-common-watch.ts's adaptiveCadence call site is unreachable in production (Electron only ships darwin/linux/win32, both covered by NARROW_WATCH_PLATFORMS) rather than implying it protects real users. The reachable path is the narrow-watch crash-fuse fallback in worktree-git-common-narrow-watch.ts. Filed #17878 to track the real long-term fix: periodically retrying the upgrade back to the narrow watch after a crash-fuse trip, so the degraded/polling state doesn't need to be tuned at all once the underlying failure clears. * perf(git-common): gate per-entry structural stats on the entry-dir signature Every real git write inside a worktree admin entry (HEAD, index, config.worktree, locked) goes through a lock file + rename, which moves the entry directory's own mtime/ctime/size signature. Only `gitdir` (worktree move/repair) is rewritten in place, and that's already covered by the periodic ungated backstop (INDEX_BACKSTOP_TICKS). The previous comment claiming structural leaves "change in place every tick" was wrong; verified against git 2.55 across checkout, commit, amend, reset, ref updates, stash, worktree lock/unlock, config --worktree, and index writes. Gate all six per-entry stats behind the entry dir's own signature instead of stat-ing every leaf unconditionally every tick: an unchanged entry now costs one stat per tick instead of six, and a changed one still costs six (bounded by change rate, not worktree count). This also fixes the actual in-flight fan-out: forEachWithConcurrency(entries, 8) previously still issued 6 stats per in-flight entry (48 real concurrent ops); with the gate, warm ticks issue ~1 stat per entry, so true in-flight tracks the concurrency limit directly. This makes the follow-up adaptive-cadence machinery from the prior commit unnecessary: the crash-fuse and no-narrow-watch polling fallbacks no longer need to stretch their own cadence, since a warm sweep across hundreds of worktrees is now cheap regardless of interval. Revert both call sites to a fixed pollIntervalMs and delete the adaptive-cadence option, the split tripwire/sweep cadence, and the seed heuristic — none of it earns its complexity once the real per-entry cost is fixed at the source. Per-entry staleness on the crash-fuse path returns to a fixed 2s + 250ms debounce instead of the previous 5.4-30s adaptive stretch. Refs #17828 |
||
|
|
e89321192a |
perf(worktree): batch remote conflict probes, re-arm the prepared checkout (#17829)
* perf(worktree): batch remote conflict probes, re-arm the prepared checkout A repo with many remotes paid one `git show-ref --verify` subprocess per remote on every branch-conflict check during create. Ask one `git cat-file --batch-check` over stdin instead; it reports a missing ref as data rather than a failed exit, so a batch stays as decidable as the per-ref probe. Hosts that cannot feed stdin, and undecided batches, still fall back to the per-ref path. The prepared checkout was single-use, so the second create in a row paid the full cold `git worktree add`. Re-arm it in the background after one is consumed; the existing TTL and preparation limit still bound it. The create timing recorder existed but its phases were never emitted and did not cover preflight, leaving a multi-second gap in the trace with no attribution. Add `resolve_name`/`prepare_push_target` phases and record the breakdown, plus the unattributed remainder, on the create span. * fix(worktree): format the conflicting review number eagerly for the create error * perf(worktree): re-arm a prepared checkout only for a burst of creates Re-arming after every consumed preparation spends a full checkout and ~200MB of disk on a user who created one worktree and stopped, then pays an unexplained delete when the TTL expires five minutes later. Track when each preparation key was last consumed and only replace it when a second create lands inside the burst window, so the warm second create is still free and an isolated create costs nothing. * fix(worktree): address review findings on the create-path batching Three findings from PR review: The `batched.found` fallback in the remote-conflict probe was unreachable — a present ref is decisive, so `found` never survives with `unknown` set, and the guard above already returns that case. `rearmPreparation` checked for an existing preparation before recording the consume, so a prefetch that re-armed the key while create finalized swallowed the timestamp and made the next create look isolated when it was really mid-burst. Create runs some phases concurrently, so summing phase durations double-counted overlap and understated `unattributed_ms` — the one number that matters when a create is slow for no visible reason. Measure the union of the phase intervals instead. * refactor(worktree): move stale-preparation cleanup into its own module The preparation module crossed the 300-line budget. Crash recovery is a separate concern from the pool itself — it discards preparations another process left registered, single-flighted per repo and runtime so a burst of arming calls shares one worktree listing. * test(worktree): make the re-arm test able to fail The burst test armed a preparation manually after the second consume, so the third checkout appeared whether or not the re-arm produced it — the assertion passed with re-arming disabled. Drop that arming call so the third checkout can only come from the re-arm, and assert the consume results rather than discarding them. |
||
|
|
f2fa4a7754 |
fix(worktrees): drop an unreachable runtime arm from the retirement gate
`findExactRepoOwner` already refuses a repo carrying both a runtime `executionHostId` and a `connectionId` -- `resolveRepoOwnershipEvidence` calls that pair contradictory, and one non-owned candidate voids the whole lookup. There is also no way for a `connectionId` to yield a `runtime:` host id, since `toSshExecutionHostId` always emits `ssh:`. The runtime arm of `connectionMatchesHost` could therefore never decide anything, and the test meant to pin it was passing through the contradiction gate instead. Keep the SSH arm, which does gate, and record where the runtime refusal actually comes from. Unreachable code on a destructive path reads as a guarantee it is not making. Refs #17776 |
||
|
|
398aeccdfe |
fix(worktrees): retire runtime-host metadata a scan proved gone
A paired client's WorktreeMeta for a runtime host is exempt from gcStaleWorktreeMeta -- that GC skips any row that is not local on both the repo and the meta's hostId -- so a scan-proven removal is the only thing that ever retires one. Both halves of that path were gated to `ssh:`, so the client kept a row for every remote worktree it had ever seen and dropped none. The renderer already computed the removals for runtime hosts and purged its own in-memory state with them; only the persisted half bailed. Widen it, and the matching main-side handler, to runtime hosts. `OffHostExecutionHostId` names the set precisely: the hosts the local-only GC skips. Also require `source === 'git'` before retiring anything. `session-fallback` reports `authoritative: true` but is the truncated, visibility-filtered `worktree.list` reply from a host too old for `worktree.detectedList`; its omissions are no evidence a checkout is gone. That guard did not matter while this only ran the in-memory purge, and does now that it deletes rows. A repo that reaches its checkouts over a connection is still never condemned under a runtime host id -- the host that executes owns that verdict. Refs #17776 |
||
|
|
93a258c81d |
fix(worktrees): reclaim orphaned pr-* fork remotes (#17842)
* fix(worktrees): reclaim orphaned pr-* fork remotes pr-* remotes Orca adds for fork-PR worktrees were only ever pruned by a single worktree's own removal, and only when that removal had complete provenance metadata, no branch pinning it, and actually ran through Orca. Legacy metadata missing remoteCreated, "preserve branch on delete" pinning the remote via branch.*.remote config after the worktree is gone, and worktrees removed outside Orca entirely all left the remote behind forever -- one real user accumulated ~50 leaked remotes this way. Add a repo-scoped reconciliation sweep that inverts the existing cleanup predicates over every pr-* remote instead of one removal, reusing sameGitHubRemoteUrl/hasBranchConfigUsingRemote so no new safety logic is introduced. It only touches a remote some worktree's persisted pushTarget explicitly recorded Orca creating (remoteCreated: true) -- naming and URL shape alone are not proof of provenance. Runs opportunistically alongside existing single-target cleanup (including RuntimePreservedBranchCleanup's force-delete path), rate-limited per repo, and fire-and-forget so it never adds latency to the worktree-removal path a user is waiting on. Fixes #17828 * test(worktrees): set a local git identity in the pr-remote fixture CI runners have no global git identity, so `git commit` in the fixture repos failed with "Author identity unknown" -- only passed locally because dev machines have one. Set user.name/user.email (plus commit.gpgSign and core.hooksPath, matching src/main/git/repo-remote-drift-real.test.ts) as local repo config in both the main and cloned "fork" fixture repos, so the test is independent of the runner's global config, signing setup, or hooks. |
||
|
|
4c24a28df0 | refactor(linux): trim AppImage CLI registration seams | ||
|
|
da4a83bd22 | fix(linux): give the CLI one entrypoint by extracting the AppImage once | ||
|
|
8cc7634051 |
refactor: name modules for their domain instead of 'helpers'
Renames seven -helpers modules for the concept their functions operate on, and splits three that were genuine grab-bags -- each had a clean cleavage along its importers, which is the signal AGENTS.md describes for a file holding more than one responsibility. Leaves keybindings/definitions-core-1..4 alone: definitions.ts spreads them in order, so their concatenation order is the command palette order and regrouping them thematically would be a user-visible change. Records that reasoning in a comment so it is not re-litigated. |
||
|
|
fc68d2c3a2 |
refactor(preload): name bridge modules for what they expose
The split named these -part-N, which says nothing. Renames each for the group of bridge methods it actually exposes and folds the single-method window-reveal module into the window-controls module it belongs with. Verified by walking the composed contextBridge surface before and after: 1060 keys, identical nesting and value types, zero delta. The bridge modules carry no satisfies annotation, so a dropped key here is a runtime error in the renderer rather than a typecheck failure. |
||
|
|
f176e49478 |
fix(git): narrow fork-remote fetch refspecs to tracked branches (#17887)
* fix(git): narrow fork-remote fetch refspecs to tracked branches git remote add with no -t writes the wide +refs/heads/*:refs/remotes/<name>/* refspec, so any later plain `git fetch` (user, agent, or Orca's own Fetch action) re-imports a fork's entire branch set and its tags -- one real machine had ~50 leaked/wide fork remotes producing 59,716 remote-tracking refs. Mint and reuse now pin -t <branch> --no-tags; a rate-limited sweep narrows and cleans up remotes minted before this fix; gitFetch self-heals when a narrowed remote's tracked branch is later deleted upstream. Refs #17828 * fix(git): soften narrow fork-remote refspec against deleted upstream branches A bare `git fetch` in a worktree checked out on a fork-PR branch resolves to the pr-* remote via branch.<name>.remote -- not origin -- making it the dominant fetch shape in Orca's terminal-centric, agent-driven usage. The previous literal-refspec design hard-failed that fetch ("couldn't find remote ref") the moment the tracked branch was deleted/renamed upstream, which is not the narrow edge case it was first described as. Switch to a trailing-`*`-suffixed refspec source/destination (refs/heads/<branch>*:refs/remotes/<name>/<branch>*). Verified against real git: this restores wildcard zero-match tolerance (silent no-op instead of a hard failure) and lets plain `git fetch --prune` reclaim the stale ref once the branch disappears, at the cost of also matching sibling branches that share the literal name as a prefix -- a materially smaller widening than the original unbounded-import bug. Also close a race with #17842's orphaned-pr-remote reconciliation sweep: both sweeps read the same worktree-metadata store to pick candidate remotes, so reconciliation can `remote remove` a remote this migration is concurrently narrowing. `ensureRemoteTracksBranchNarrowly`'s plain `config --add` would silently resurrect a url-less config section in that case; re-check `remote.<name>.url` (via the new `remoteHasUrl`, plumbing rather than porcelain `remote get-url`, which falls back to echoing the remote name as a bogus URL) after the narrowing writes and remove the section if it's gone. * fix(git): update stale fork-remote mint assertions for -t/--no-tags and wildcard-suffix refspec Four test files still asserted the pre-#17828 remote-add shape or the literal (non-wildcard-suffixed) fetch refspec from before the deleted-upstream-branch softening commit, so CI went red on that HEAD: - worktree-push-target-refspec-real-git.test.ts: the migration fixture asserted a hardcoded tracked-ref count before narrowing. Under git >= 2.44, `followRemoteHEAD` auto-creates a `refs/remotes/<name>/HEAD` symref on the first fetch matching the full wildcard refspec, adding one untracked ref. Made the count/assertions robust to that ref's presence instead of hand-tuning the constant per git version. - worktrees-wsl-runtime-routing.test.ts: assertions predated both the `-t <branch> --no-tags` mint change and the wildcard-suffix refspec change; updated to the full, correct call sequence and confirmed the WSL routing options (cwd, wslDistro) are threaded to every call. - worktrees-create-metadata-persistence.test.ts and orca-runtime-tests/worktree-removal-and-reconciliation.spec.ts: same class of staleness, found via CI job log cross-referencing rather than being explicitly flagged. Verified out of scope: the SSH fork-remote mint path (prepareWorktreePushTargetSsh) is untouched by this PR -- it never persists a `remote.<name>.fetch` refspec at all, using provider.fetchRemoteTrackingRef for a targeted per-branch fetch instead -- so worktrees-ssh-fork-push-target-remote.test.ts needed no change. * fix(git): migrate pr-* remotes with zero worktree-metadata trace too The migration sweep's candidate discovery was purely metadata-driven (store.getAllWorktreeMeta()), so a pr-* remote whose every referencing worktree was removed outside preserve-on-delete (metadata purged, not just the worktree) was permanently invisible to it and stayed on the wide default forever. Field data from a manual migration run against a real user's repo (31 pr-* remotes, 34,637 tracking refs, only 18 actually needed) found exactly this: 15 of 31 remotes had no branch pinning them at all. Widen discovery to every pr-* remote git reports on disk, in addition to metadata-derived candidates. For a remote with no branch provenance from either metadata or surviving branch.*.remote/.pushRemote config, there's nothing to narrow *to* -- clear its fetch refspec entirely instead (stays pushable, imports nothing on a plain fetch), gated on it still carrying the untouched stock wide default so a user's own custom pr-*-named remote isn't touched. Removing the remote outright stays #17842's job. Adds clearForkRemoteFetchRefspec (fork-remote-refspec.ts), 3 new mocked-exec tests, and a real-git integration test proving a subsequent plain `git fetch` on the cleared remote imports nothing. |
||
|
|
84584b61d0 |
perf(git): cache sparse-checkout annotation on worktree listing (#17859)
* perf(git): cache sparse-checkout annotation on worktree listing `git worktree list` never reports sparse-checkout state, so every listing paid a per-worktree fs.stat + config read to detect it -- measured at ~9x the cost of the `git worktree list` call it decorates on a 1000-worktree repo. Cache the result per worktree path, invalidated by the existing worktree-change invalidator registry plus explicit remove/move hooks, with a 5-minute reconcile window bounding the one unwitnessed edge case (external `git sparse-checkout` toggle with extensions.worktreeConfig off), matching the precedent already accepted in readRepoWorktreeAdminFingerprint. * perf(git): normalize/scope sparse-checkout cache keys, add SWR Address independent-review follow-ups on the sparse-checkout annotation cache (#17859): - Extract canonicalWorktreePath() from areWorktreePathsEqual and key/invalidate the cache through it on both read and write, closing the disclosed path-spelling P2 outright instead of leaving it as a residual risk. - Scope cache entries and clears by repo path (derived from the invalidator registry's repoId via a store lookup, falling back to a full clear when the repo can't be resolved), so churn in one repo no longer evicts a sibling repo's warm cache. - Replace the hard 5-minute cutoff with stale-while-revalidate: past the window, callers get the cached value immediately while a deduplicated background probe corrects it and, on a flip, drives the existing worktrees-changed notification -- collapsing visible staleness from the full window to one refresh cycle at zero added listing latency. Also corrects a stale claim in the original PR description: newer Git does emit a `sparse` porcelain line (which annotateSparseCheckoutStatus already skips), but Orca's Git 2.25 compatibility baseline predates it, so the fallback detection this caches remains necessary. * fix(git): stop background sparse-checkout revalidation resurrecting invalidated entries Readiness-loop finding: a stale-while-revalidate probe in flight when a worktree is removed/moved (or a repo's cache is cleared) would still write its result back afterward, resurrecting an entry that was deliberately dropped. Guard the write with a presence check so an invalidated key stays absent until the next real read. * fix(git): identity-check the sparse-checkout SWR write-back guard The has()/presence guard from the previous commit only proved some entry existed at the key, not that it was the one this revalidation started from. A worktree removed and re-created at the same path while a background re-detect was in flight would repopulate the key with a fresh cold read, and the stale in-flight result would then overwrite it -- exactly the race greptile (P1) and pullfrog both flagged as still open. Compare the map's current entry by reference to the entry captured when the revalidation began; a mismatch means something else (invalidate, clear, or a fresh cold read) replaced it, and the stale result must not be written back. Added a regression test that fails against the old has() guard and passes with the identity check: invalidate and repopulate the key with a different value mid-flight, then let the stale revalidation settle and assert the fresh value survives. |
||
|
|
9542b45d99 |
fix(wsl): resolve conflict and working-tree probes in the host path namespace (#17895)
Git running inside a WSL distro writes `.git` gitdir pointers, and answers
`status --porcelain`, in the guest namespace. Node reads both back in the
Windows main process, where `/mnt/c/repo/.git` resolves to `C:\mnt\c\repo\.git`
and `/home/me/wt` names nothing at all. Four fs probes were built on those
fabricated paths and always came back "absent":
- `detectConflictOperation`'s four marker probes, so merge/rebase/cherry-pick
badges silently went missing.
- `parseUnmergedEntry`'s compat existence check, so every `deleted_by_us` /
`added_by_them` conflict rendered as 'deleted' regardless of the working tree.
- `findExistingWorktreeSymlinkPaths`' `lstat` from status, so Orca's own shared
symlinks (node_modules and friends) showed as user changes.
- the same `lstat` from the hosted-review dirty preflight, which fails closed:
an unreadable shared symlink read as uncommitted work and blocked PR/MR
creation outright.
`resolveGitDir` computes the host spelling of the worktree once and uses it for
both the gitfile read and the pointer resolve, so a guest-spelled worktree path
is reached at all, and a relative pointer (`worktree.useRelativePaths`, git
2.48+) resolves against a spelling Win32 understands. The pointer itself now
goes through the already-landed `resolveGitMetadataPath`, and the function gains
an optional `{ wslDistro }` for a caller whose base path does not encode a
distro. `detectConflictOperation` forwards it, and the three callers that reach
it -- status-read, the runtime RPC, the `git:conflictOperation` IPC -- pass the
git options they already hold. The return type stays `Promise<string>`.
`resolveWorktreeHostPath` is the same rule applied to a worktree path, used by
status-read for the two working-tree probes and by the review preflight. Both it
and `resolveGitMetadataPath` now treat only a single-leading-slash path as guest
namespace: `//wsl.localhost/...` is already a host UNC spelling, and translating
it prepended a second share prefix.
`readWorktreeDiffStamp` needed the same one-namespace guarantee, since moving
translation inside `resolveGitDir` would otherwise make its HEAD and index real
while the working-tree stat stayed fabricated, letting a settled diff survive
every edit. #17896 landed that change first, so it is no longer in this diff;
its version is a superset and all four components already resolve from one
`hostWorktreePath`. What remains here is the `resolveGitDir` gitfile-pointer
fix that #17896 explicitly deferred, which `worktree-diff-stamp-host-paths.test.ts`
pins.
`getConflictCompatibilityStatus` moves from `existsSync` to async `access`, for
the same reason `detectConflictOperation` did: once these paths are real they
are `\\wsl.localhost\...` shares, and a sync probe per asymmetric conflict
blocks the Electron main thread for a 9p round trip on every status poll.
Per-platform delta:
- native Windows, no WSL: no behavioral change. Nothing here starts with a
single `/`, so no path is translated. An absolute pointer is now returned
verbatim rather than separator-normalized; every consumer re-joins or
normalizes it before use.
- macOS/Linux: no change. Guest-pointer translation is gated to win32, and a
caller-named distro is ignored off Windows.
- Windows + WSL: drvfs pointers and drvfs-spelled worktrees now resolve to their
drive spelling instead of `C:\mnt\...`; a non-drvfs guest path resolves
through the named distro's UNC share, or stays verbatim (ENOENT -> existing
fail-safe) when none is named.
- SSH/relay: none. Those paths return before any of this via the provider
branch; `src/relay/git-handler-status-ops.ts` keeps its own resolveGitDir.
- folder workspaces, GitLab: none. Neither is on these code paths.
|
||
|
|
1603810dde |
perf(worktree): make head-identity refresh incremental (#17843)
* perf(worktree): make head-identity refresh incremental Head-identity refresh re-read `gitdir` + `HEAD` + a loose ref for every linked worktree on every watcher burst. On a 973-worktree checkout that is ~2,800 metadata reads (~1.0s of main-process fs I/O) per event, and the debounced pipeline fires on every commit in any worktree — so fleet-wide agent activity degenerated into a continuous scan loop. Watcher events already name the admin dir that changed. Classify each event into a head-identity scope, memoize per-entry identities, and re-read only the scoped entries. Refs resolved during a pass are replayed onto cached entries that share the same branch, so `git worktree add --force` siblings stay current without extra reads. Invalidation stays conservative: an absent scope (watcher failure, event overflow, cold start) means a full re-read, `packed-refs` writes invalidate every entry, misses are never memoized, and one refresh per minute is promoted back to a full re-read to bound the window where a ref moves with no event under any admin dir. Measured on the reported 989-entry checkout (macOS/APFS): one-worktree commit 2,816 -> 2 file reads, 61ms -> 0.5ms p50 with an identical page cache; a 20-worktree debounce burst costs 57 reads / 9.7ms; an external `git worktree add`/`remove` costs one readdir / 1.0ms. Refs #17828 * fix(worktree): harden incremental head-identity invalidation Two holes found in self-review: - An admin entry name removed and immediately reused inside one debounce window coalesced into a listing-only scope, so the reused entry kept serving the removed worktree's cached head. Name the entry alongside the listing on every `worktrees/<name>` create/delete. - A non-ENOENT `readdir` failure on `worktrees/` collapsed the memo to the primary row, which then re-emitted every identity on recovery. Mirror worktree-git-common-polling: only a genuinely absent dir means empty; any other error keeps the previous listing. * fix(worktree): let empty-scope bursts still take the head re-baseline Adversarial review found the 60s full-rebaseline promotion was unreachable whenever the triggering burst had an empty head-identity scope: the skip guarded on the raw caller scope and returned before `resolveScope` ran, so `lastFullReadAtMs` was never re-evaluated. A repo whose only churn is `git worktree lock`/`unlock` or a sparse toggle — Orca's own prepared-checkout flow locks and unlocks on every create — could starve the promotion forever and hold a stale head indefinitely. Resolve the scope first and skip on the effective scope. Also stop deferring an add/remove that arrived while the `worktrees/` listing was transiently unreadable: forget the memoized listing so the next refresh re-enumerates whatever its scope, instead of waiting for another listing event. Both fixes carry a test verified to fail without them. * fix(worktree): return head-read completeness instead of sniffing the memo Adversarial review round two. Six fixes, each with a test verified to fail without it. - `readGitCommonHeadIdentities` now returns `{ identities, listingComplete }`. The refresh layer was inferring "enumeration failed" from `cache.entryNames === null`, a reader-owned field whose null also means "cold start" — fragile in production and impossible to express in a mock. - A read discarded by teardown, or one that could not enumerate `worktrees/`, no longer arms the 60s freshness clock. - A queued refresh whose re-run met a destroyed window (macOS recreates the window while the watch lives on) was cleared and dropped. It now stays armed and is folded into the next request. - An incomplete listing carries forward the baseline rows it could not observe, so recovery does not report every linked worktree as changed. - The baseline advances after notifying, so a send into destroyed chrome leaves the move to be retried instead of diffing it away. - A scope naming an entry the memoized listing does not know now forces a re-enumeration instead of resolving to zero work — this removes an unstated dependency on `diffGitCommon` emitting a dir-level create for new entries. - Overflow states FULL at its construction site rather than relying on a downstream `?? FULL` for an absent field. Also documents the load-bearing invariant behind the empty-scope skip (an empty scope only reaches the refresh from a structural burst, which forces `emit: false` and is always paired with a catalog notification for every repo on the watch), and strengthens two tests that could not distinguish the behaviour they claimed. * fix(worktree): bound head-identity staleness with a one-shot catch-up The previous re-baseline was opportunistic: it rode the next refresh, so a ref that moves with no watched write (`git update-ref refs/heads/x` from a sibling worktree) stayed stale until an event happened to arrive after the interval. Pre-PR the very next event anywhere in the repo corrected it, so this was a real narrowing of correctness, not just a pre-existing gap. Arm a one-shot, unref'd timer when a SCOPED pass completes, firing one full re-baseline an interval after the last full read, then disarming. A full pass disarms instead of arming, so it never becomes a background poll, and the timer only exists after an event — an idle repo still schedules nothing and reads nothing. Cost is O(1) timer per active repo and at most one full read per interval: the same operation the old code ran per event, 60x rarer. This also converts "stale until some later event" into "stale at most one interval, period", which is what bounds the blast radius of any invalidation bug in the scoping itself. Cleared on watch disposal. Three tests, each verified to fail without its fix: the catch-up runs with no further events; a quiet repo issues no background reads and the timer disarms after firing; disposal stops it. * fix(worktree): treat an unreadable head as unknown, not absent Reported independently by two PR reviewers. `readTrimmedFile` collapsed every errno to `null`, so an EIO/EACCES/ENFILE on a `gitdir`, `HEAD`, loose ref, or `packed-refs` read was indistinguishable from the file being absent — and the caller deletes the cached identity on `null`. Same conflation AGENTS.md forbids for the SSH verdict vocabulary: loss of contact is not evidence of absence. Reads now report three outcomes, and an unknown: - keeps the entry's last verified identity instead of evicting it, - is never replayed onto siblings sharing the branch as "this ref is gone", - marks the entry unverified so the very next pass re-reads it whatever its scope, and - reports the pass incomplete, so it cannot arm the freshness clock. The reviewers' stated consequence — that an evicted entry stays evicted until the next full pass — did not hold, because `!cache.entries.has(name)` already forced a re-read. The real cost was that one EMFILE evicted every entry it touched and the next pass re-read all of them, which is exactly the full scan this PR exists to remove, plus a spurious re-publish of every row. Renames `listingComplete` to `complete`: it now covers entry reads too. |
||
|
|
d462766cb0 |
refactor(main): split filesystem git remote handlers
(cherry picked from commit
|
||
|
|
5b4e7edb50 |
refactor(main): split backend services and startup
(cherry picked from commit
|
||
|
|
7e8337b155 |
test(preload): census split GitHub bridge owners
(cherry picked from commit
|