mirror of
https://github.com/stablyai/orca.git
synced 2026-09-30 00:03:15 +00:00
ae426bfbeb024fc4d0f3c7ac0dafc84dfcdacc4e
7
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6c8eea5ebe | perf(worktree): fix the prepared-checkout hit rate and make misses visible (#17863) | ||
|
|
3777070eaf |
fix(worktree): restore the stale-cleanup signal after the module split (#18058)
* fix(worktree): restore the stale-cleanup signal after the module split Moving stale-preparation cleanup into its own module took `staleCleanupInFlight` with it, but `hasPendingWorktreeCreatePreparations` still read it directly. Both sides were green in isolation — the reference arrived on main while the split was in review — so the break only appeared once they merged, and it fails typecheck for every branch built on main. Expose the predicate from the module that owns the map, and cover the signal with a test so the idle gate's "a create is imminent" answer cannot silently regress again. * test(worktree): anchor the pending-signal test on the scan, not on await depth |
||
|
|
d7123591ce |
perf(git): pack the loose refs Orca's own fetches leave behind (#17857)
* perf(git): pack the loose refs Orca's own fetches leave behind Orca strips git's auto-maintenance off every fetch it issues (GIT_FETCH_SKIP_AUTO_MAINTENANCE_CONFIG_ARGS) and never compensated, so nothing in an Orca-driven checkout ever packs refs. One real machine reached 36,574 loose refs, where `git show-ref -- main` costs 5.2s and every worktree create pays for it. Add an idle-time, per-repo `git pack-refs --all --prune`, armed by the fetches that create the debt. It runs only after ten minutes of quiet on that repo, only above 1000 loose refs (probed with a walk bounded by that threshold, not by the backlog), one at a time across the whole app, at the background admission tier, and never while an agent is working, a create is prepared or in flight, a worktree removal is deleting refs, the app is quitting, or the machine is on battery. A user who set `maintenance.auto=false` or `gc.auto=0` has opted out. Measured on a 36,001-loose-ref fixture (macOS/APFS, git 2.44): `show-ref` 5.5-12.2s -> 30-49ms, `for-each-ref` 4.0-10.8s -> 43-48ms. Also fixes a pre-existing bug the split exposed: `--path-format=absolute` is ignored before git 2.31, and taking rev-parse's stdout raw collapsed every repo on such a host onto one fetch-serialization key. Refs #17828 * perf(git): make idle ref maintenance preemptible and cheaper to probe The idle veto was one-directional: it stopped a pack from starting during a create, removal, or agent work, but nothing stopped those from starting during a pack. A user-clicked Fetch, a branch delete, or a worktree removal that needed `packed-refs.lock` mid-rewrite could fail with `unable to create packed-refs.lock` -- a git error with no visible cause. Make the pack cancellable end to end. An AbortSignal now reaches the `pack-refs` child and both pre-pack probes, and `pause()` aborts what is running, waits for it to actually stop, and holds a suspension count so nothing new starts until the caller releases. Every entry point that deletes a ref takes that pause: gitFetch, gitPull, gitFastForward, removeWorktree, forceDeleteLocalBranch, prepareWorktreeCreateCheckout, addWorktree. Five more triggers close the rest of the window: battery drop, window focus, quit, the attempt deadline, and any other git command queueing for an admission slot. Judge a pack by re-probing the backlog rather than by the child's exit code. Measured in the field: another Orca session moved a branch mid-pack, git reported `cannot lock ref`, skipped that ref and packed the rest -- 36,688 loose refs down to 3. On a machine running several sessions that is the normal case, and retrying it would be wrong. Probe with one batched `readdir` per directory instead of streaming `opendir`, which issues a thread-pool round trip every 32 entries: 177ms -> 23ms on a real 36,600-ref repository, with half the event-loop lag. The walk stays strictly sequential so it can never occupy more than one of libuv's four filesystem threads. `PackRefsLockOwnership` makes a lock left by SIGKILL attributable, and only reclaims one when a marker exists, the lock is older than any pack-refs could run for, and the recorded process is gone. Refs #17828 * fix(git): wait out the packed-refs lock instead of killing the pack Measured on Git 2.55/APFS with 37k loose refs: a full `pack-refs --all --prune` takes 23-32s but holds `packed-refs.lock` for only 0.03-1.37s of it. The other ~95% is the prune phase, during which a concurrent `fetch --prune`, `branch -D` or `update-ref` succeeds every time -- per-ref locks last microseconds and git retries for `core.filesRefLockTimeout`. So the abort-on-everything design was strictly harmful. SIGTERM into the prune loop strands an empty `refs/**/*.lock` about one time in five (9/30, 5/40, 6/30 kills): `tempfile.c` opens the lock O_EXCL before `activate_tempfile()` links it into the list the signal handler walks, and a pack does ~36k lock cycles. Afterwards `update-ref -d` on that ref fails with `cannot lock ref ... File exists`, permanently. On Windows `taskkill /f` never runs git's handlers at all, so an abort inside the rewrite strands `packed-refs.lock` every time. Never signal the child. `packRefs` no longer takes an abort signal; it polls `packed-refs.lock` and reports the window through a `PackedRefsLockReporter`. `pause()` resolves when the lock is released -- bounded, and free during the prune -- while the suspension counter still blocks new attempts. Battery and window-focus become do-not-start rather than stop-what-is-running, and quit waits for the lock and lets the child finish orphaned. For strands that already exist, `PackRefsLockOwnership` now also reclaims `refs/**/*.lock` under the same three conditions plus a 0-byte check, and a lock carrying our own not-yet-reclaimable marker records `locked` with a 30min retry instead of the 6h failure cooldown -- so a Windows strand self-heals in half an hour rather than six. Reverts the git admission-scheduler event bus, which existed only to drive the abort this removes. Refs #17828 * test(git): make the ref-maintenance waits survive a loaded runner CI shard 4/8 failed on `restarts every armed countdown when the user does ref work themselves`, which passes locally. The `until()` helper spun a fixed 200 event-loop turns and then returned silently, so on a contended runner the filesystem probe had not finished and the assertion that followed failed with an unrelated message. Bound the wait by wall clock instead and throw a named error, which immediately exposed a second latent bug: the single-flight test's second wait could never succeed, because the deferred repo's retry is on a faked `setTimeout` that spinning the real loop never advances. It had been passing only because the old helper gave up quietly. Add a timer-aware variant for those, and have the countdown test await a signal the fake pack resolves rather than polling at all. Verified stable across five sequential runs and once under load average 32 with six concurrent suites. Refs #17828 |
||
|
|
e89321192a |
perf(worktree): batch remote conflict probes, re-arm the prepared checkout (#17829)
* perf(worktree): batch remote conflict probes, re-arm the prepared checkout A repo with many remotes paid one `git show-ref --verify` subprocess per remote on every branch-conflict check during create. Ask one `git cat-file --batch-check` over stdin instead; it reports a missing ref as data rather than a failed exit, so a batch stays as decidable as the per-ref probe. Hosts that cannot feed stdin, and undecided batches, still fall back to the per-ref path. The prepared checkout was single-use, so the second create in a row paid the full cold `git worktree add`. Re-arm it in the background after one is consumed; the existing TTL and preparation limit still bound it. The create timing recorder existed but its phases were never emitted and did not cover preflight, leaving a multi-second gap in the trace with no attribution. Add `resolve_name`/`prepare_push_target` phases and record the breakdown, plus the unattributed remainder, on the create span. * fix(worktree): format the conflicting review number eagerly for the create error * perf(worktree): re-arm a prepared checkout only for a burst of creates Re-arming after every consumed preparation spends a full checkout and ~200MB of disk on a user who created one worktree and stopped, then pays an unexplained delete when the TTL expires five minutes later. Track when each preparation key was last consumed and only replace it when a second create lands inside the burst window, so the warm second create is still free and an isolated create costs nothing. * fix(worktree): address review findings on the create-path batching Three findings from PR review: The `batched.found` fallback in the remote-conflict probe was unreachable — a present ref is decisive, so `found` never survives with `unknown` set, and the guard above already returns that case. `rearmPreparation` checked for an existing preparation before recording the consume, so a prefetch that re-armed the key while create finalized swallowed the timestamp and made the next create look isolated when it was really mid-burst. Create runs some phases concurrently, so summing phase durations double-counted overlap and understated `unattributed_ms` — the one number that matters when a create is slow for no visible reason. Measure the union of the phase intervals instead. * refactor(worktree): move stale-preparation cleanup into its own module The preparation module crossed the 300-line budget. Crash recovery is a separate concern from the pool itself — it discards preparations another process left registered, single-flighted per repo and runtime so a burst of arming calls shares one worktree listing. * test(worktree): make the re-arm test able to fail The burst test armed a preparation manually after the second consume, so the third checkout appeared whether or not the re-arm produced it — the assertion passed with re-arming disabled. Drop that arming call so the third checkout can only come from the re-arm, and assert the consume results rather than discarding them. |
||
|
|
f2db24bca0 |
fix(worktree): reclaim a prepared checkout whose discard failed in-process (#17899)
Speculative create-preparation evicts entries past the 3-entry limit and the 5-minute TTL, and both paths swallowed a failed `discardPreparedWorktree`. The only other reclaim path, `cleanupStalePreparations`, skips any preparation whose lock-reason pid is still alive, so a discard that failed inside the running app stranded its scratch checkout and its locked worktree registration until restart. Record the failed discard keyed by host (repo path + WSL distro) and prepared path, and retry it the next time a fresh preparation starts for that host. The retry is kicked off before `listWorktreeGraph` but never awaited, so it runs alongside the stale scan and the `worktree add` instead of sitting in front of the user's create. It is capped at 3 attempts and warns when it gives up. Enrolment is unconditional: `prepareWorktreeCreateCheckout` self-discards on failure, but only best-effort, so a checkout that failed on a busy handle can strand the same registration. |
||
|
|
e6257b6e32 |
perf(worktree): resolve the WSL workspace root off the main thread when preparing (#17792)
`computeWorkspaceRoot` resolves a WSL repo's mirror root through `getWslHome`,
which is a synchronous `execFileSync('wsl.exe', ...)` with a 5s timeout. Two
worktree preparation paths ran it on the Electron main thread:
`prepareLocalWorktreeRootForRepo` (repo registration, clone completion, repo
update, project host setup, folder->git upgrade) and `prepareWorktreeCreateForRepo`
(the speculative checkout started while the create composer is open). On a stopped
or cold distro that froze every window for up to 5s. Being fire-and-forget did not
help: only 3 of the 16 `prepareLocalWorktreeRootForRepo` call sites are `void`-ed,
the other 13 are awaited inside IPC handlers, and the sync probe blocks the main
thread either way. `prepareLocalWorktreeRootsForRepos` runs the same probe for
every repo from the settings-save handler.
Adopt the existing `computeWorkspaceRootAsync` (now exported) at those two call
sites, and give the two resolvers a shared mirror-distro decision and shared
root-from-home layout so they cannot drift apart.
Also thread the mirror distro into the prepare-side path settings.
`createLocalWorktree` passes `getWorktreeMirrorDistro(store, repo)` and
`prepareWorktreeCreateForRepo` did not, so a `C:\` repo on a WSL project runtime
prepared under `C:\workspaces` while the create click looked under the mirrored
WSL root: the keys never matched and every prepared checkout was discarded, after
paying for a full checkout that sat until the 5min TTL. Pre-existing on main;
included because it is the same line and the same resolver.
Scope of the win, stated precisely: only those two preparation paths stop
blocking. On a reachable distro `getWslHome` caches on success, so before this
change the first repo paid one blocking probe and the rest were cache hits -- the
change makes that one probe non-blocking, it does not remove N probes. Failed
probes are never cached, so on a stopped distro N repos did pay N sequential 5s
blocking probes and now share one in-flight async probe.
Costs: five sync `computeWorkspaceRoot` callers remain (allowed-roots resolution,
the create click in worktree-remote, CLI create, watch targets, worktree trash),
and `getWslHome` reads only `wslHomeCache` -- it cannot join an in-flight async
probe. The guaranteed synchronous cache warm-up therefore becomes a window in
which one of those callers can still block and can spawn a second concurrent
`wsl.exe`. Concretely: opening the create composer and clicking Create within a
few hundred ms on a cold distro now pays the freeze on the click instead of on the
background prep. Separately, the mirror-distro fix makes prepare spawn an async
`wsl.exe` home probe for `C:\` repos on a WSL runtime, where it previously spawned
none.
No race added: `prepareWorktreeCreateForRepo` computes the preparation key and
inserts the registry entry in one synchronous run after the await, so two
concurrent creates still dedupe to a single prepared checkout.
`worktree-create-preparation-wsl-root.test.ts` runs the real resolver through
prepare and then claims the entry with the production consume-side call shape
(including the mirror distro), so a divergence between the two resolvers fails a
test instead of silently discarding every prepared checkout.
|
||
|
|
3ab9766e38 |
perf(worktree): prepare checkouts while the composer is open
Squashed merge of PR #17290. |