Files
orca/src/main/ipc/worktree-include-copy-budget.ts
T
Brennan Benson b31e9bb03d fix(worktree): bound the .worktreeinclude copy so a huge include can't freeze workspace creation (#10540)
* fix(worktree): bound the .worktreeinclude copy so a huge include can't freeze creation

`.worktreeinclude` copying was bounded in entry count (1000) but unbounded in
bytes and files, and awaited inline during worktree creation. A repo listing
`node_modules` froze creation for minutes behind the create dialog on Linux and
Windows, where the fallback is a full `fs.cp` (macOS gets a cheap APFS clone).

Measure each copy-mode source against a cumulative budget (2 GB / 50k files)
before the first byte is written, and refuse the entries that bust it. Refused
entries ride the existing `CreateWorktreeResult.warning` channel so a workspace
never silently comes up missing its included files.

Pre-measurement rather than mid-copy abort: `fs.cp` ignores its `signal`
option, so a started copy cannot be cancelled and would strand a partial tree.
Refusing up front means there is no partial state to clean up.

* fix(worktree): don't charge bytes for copy-on-write clones, and bound the sizing walk

Two defects in the copy budget, both found by review:

- The byte limit was applied on macOS, where the copy is an APFS clonefile.
  Measured: a 2.7 GB tree clones in 22 ms and consumes no disk. Refusing it on
  a 2 GB byte ceiling denied work that was already free — a regression on the
  one platform this bound was never meant to touch. Bytes are now charged only
  when a byte-for-byte copy will actually run; the volume probe that decides
  this is the same cached df+diskutil pair the clone runs, and writes nothing,
  so the "refuse before the first byte" invariant holds. The entry limit still
  applies everywhere: inodes are real work even on the clone path.

- A refused entry consumed no budget, so a `.worktreeinclude` listing many
  over-budget directories paid a fresh full-limit walk for each one — up to
  1000 x 50,000 lstat calls, re-creating the stall this bounds. The walk is now
  charged against its own ceiling whatever the verdict.

Also documents that `admit()` must be awaited sequentially (CodeRabbit).

* fix(worktree): give the sizing walk headroom so one huge entry can't starve the rest

The walk ceiling added in the previous commit was seeded with maxEntries, the
same number the entry limit uses. Sizing an entry that busts the file-count
limit walks maxEntries + 1, driving the ceiling negative, so every later
`.worktreeinclude` entry was refused without being measured at all.

That regressed the common case: a repo listing `node_modules` plus `.env` used
to get `.env`; it silently got nothing. Reproduced, and now covered by a test
that fails when the headroom is removed.

The walk now gets 5x the entry budget, so total sizing work stays bounded
(<=250k lstat per materialization, vs the 1000 x 50k this ceiling exists to
prevent) while ordinary lists never reach it. Entries refused because earlier
ones exhausted the walk report a distinct 'sizing' reason, so the warning stops
quoting size limits at a 4-byte file that was never measured.

* fix(worktree): bill a failed clone's bytes, and blame the right ceiling

Two follow-on defects from the copy-on-write fix:

- A predicted APFS clone that then failed mid-copy (EPERM, ENOSPC) fell through
  to a real `fs.cp` whose bytes were never charged, because the entry had been
  admitted on the premise that cloning is free. That reopened the unbounded
  copy on macOS. The measured size is already known, so the fallback now bills
  it and refuses if it no longer fits, reporting the entry as skipped instead
  of silently copying gigabytes. A clone that was never viable
  (ApfsCloneUnavailableError) was already charged as a real copy, so that path
  keeps falling back as before.

- The walk ceiling is also applied inside the measurement via
  min(remainingEntries, remainingWalk), and when the walk term bound, the
  refusal was still reported as 'entries' — telling the user a 3-file directory
  busted a 4-file limit. It now attributes to whichever ceiling actually bound.

Also fixes the singular warning text, which said "entry X was not copied ...
copying them would exceed ... Copy them in manually".

* fix(worktree): flag a partial clone leftover, cap the warning, cover two branches

- A clone that fails partway only removes an *empty* reservation, so leftovers
  can survive at the target. Reporting that entry as simply "not copied" sent
  the user to copy it in manually, straight into a half-populated directory.
  Those skips now carry mayBePartial and the warning says to check the path
  first. Cleaning up the leftovers stays the deferred follow-up it already was.

- The warning enumerated every skipped path. `.worktreeinclude` allows 1000
  entries and all of them can be skipped, so it now names five and counts the
  rest — an unbounded string is a poor look in a PR about bounds.

- Two load-bearing branches had no test, both proven by surviving mutants: the
  `bytesAreCopied` short-circuit (reachable when a wedged df/diskutil makes the
  volume probe answer "no clone", so bytes are charged up front and must not be
  billed twice), and chargeBytes actually consuming budget for later entries.

* fix(worktree): only flag directory clones as partial, and cap that list too

- mayBePartial was set for every refused clone fallback, but only a *directory*
  clone can leave anything behind: the file path clones into a temp name and
  publishes with link(2), so a failure leaves nothing at the target. Sending
  the user to inspect a path that does not exist is its own small lie.

- The partial-copy sentence sliced to five names without the "and N more" that
  the other sentence appends, so entries past the fifth were surfaced nowhere.
  Both sentences now share one nameList helper.
2026-07-25 05:11:03 -07:00

233 lines
9.3 KiB
TypeScript

import { lstat, readdir } from 'node:fs/promises'
import { join } from 'node:path'
/** Ceiling on what one worktree materialization may copy, measured before any
* bytes are written. Both limits are cumulative across the whole run, so a
* hundred medium entries trip the same guard one huge entry does. */
export type WorktreeCopyBudget = {
maxBytes: number
maxEntries: number
}
// Why: `.worktreeinclude` is a repo-authored list, and a repo that lists
// `node_modules` freezes worktree creation for minutes behind an inline copy
// (macOS gets a cheap APFS clone; Linux/Windows get a full `fs.cp`). These
// limits clear real payloads — `.env` files, `.vscode/`, small build caches —
// and refuse dependency trees. The entry limit matters as much as the byte
// limit: 200k tiny files are slow to copy even though they weigh little.
export const DEFAULT_WORKTREE_COPY_BUDGET: WorktreeCopyBudget = {
maxBytes: 2 * 1024 * 1024 * 1024,
maxEntries: 50_000
}
// Why: the sizing walk gets headroom over the copy budget so one refused
// `node_modules` cannot starve the small entries listed after it — it burns
// maxEntries+1 measuring, and without headroom nothing else would be sized.
const WORKTREE_COPY_SIZING_HEADROOM = 5
export type WorktreeCopyBudgetExceededReason =
| 'bytes'
| 'entries'
/** Not this entry's fault: earlier entries used up the total sizing walk. */
| 'sizing'
export type WorktreeCopySizeVerdict =
| { withinBudget: true; bytes: number; entries: number }
| { withinBudget: false; reason: WorktreeCopyBudgetExceededReason }
export type SkippedWorktreeCopyPath = {
path: string
reason: WorktreeCopyBudgetExceededReason
/** The copy was abandoned after it had started, so leftovers may remain —
* "copy it in manually" would then merge into a half-populated directory. */
mayBePartial?: boolean
}
export type WorktreeCopyAdmitOptions = {
/** False when the backend clones copy-on-write (APFS `clonefile`), where
* bytes cost nothing and only inode count is real work. */
bytesAreCopied?: boolean
}
export type WorktreeCopyBudgetTracker = {
/** Measure `source` against what is left of the budget. A `withinBudget`
* verdict consumes the measured size; an over-budget verdict consumes
* nothing, so later, smaller entries still get their chance.
*
* Await each call before the next: the remaining pool is read before the
* measurement walk and written after it, so concurrent callers would both
* size against the same stale pool and could jointly bust the budget. */
admit: (source: string, options?: WorktreeCopyAdmitOptions) => Promise<WorktreeCopySizeVerdict>
/** Bill bytes that were measured but not charged, because the copy was
* expected to clone and then didn't. Returns false if they no longer fit,
* in which case the caller must not run the copy. */
chargeBytes: (bytes: number) => boolean
}
type MeasuredCopySize = {
verdict: WorktreeCopySizeVerdict
/** Entries actually walked, whatever the verdict — this is the measurement's
* own cost, which the tracker charges so a long list of over-budget entries
* cannot re-freeze creation by re-walking for each one. */
walked: number
}
async function measureCopySize(
source: string,
remainingBytes: number,
remainingEntries: number,
remainingWalk: number
): Promise<MeasuredCopySize> {
let bytes = 0
let entries = 0
const pending: string[] = [source]
while (pending.length > 0) {
const current = pending.pop() as string
let stats: Awaited<ReturnType<typeof lstat>>
try {
stats = await lstat(current)
} catch {
// Raced away between the walk and now — the copy will skip it too.
continue
}
entries += 1
if (entries > Math.min(remainingEntries, remainingWalk)) {
// Why: attribute to whichever ceiling actually bound. Blaming the file
// limit for a walk that earlier entries used up would quote the user a
// limit this entry never approached.
const reason = remainingWalk < remainingEntries ? 'sizing' : 'entries'
return { verdict: { withinBudget: false, reason }, walked: entries }
}
// Why: both copy backends reproduce a nested symlink as a symlink rather
// than following it, so walking through one would double-count a shared
// target and could loop forever on a cycle.
if (stats.isSymbolicLink()) {
continue
}
if (stats.isDirectory()) {
try {
for (const name of await readdir(current)) {
pending.push(join(current, name))
}
} catch {
// Unreadable directory — nothing measurable, and the copy will report it.
}
continue
}
bytes += stats.size
if (bytes > remainingBytes) {
return { verdict: { withinBudget: false, reason: 'bytes' }, walked: entries }
}
}
return { verdict: { withinBudget: true, bytes, entries }, walked: entries }
}
/** Why a pre-measurement pass rather than aborting mid-copy: `fs.cp` ignores
* its `signal` option, so a copy that has started cannot be cancelled and
* would leave a partial tree behind. Refusing before the first byte is
* written keeps the worktree in a state the user can reason about. The walk
* is itself bounded — it returns the moment either limit is crossed. */
export function createWorktreeCopyBudgetTracker(
budget: WorktreeCopyBudget = DEFAULT_WORKTREE_COPY_BUDGET
): WorktreeCopyBudgetTracker {
let remainingBytes = budget.maxBytes
let remainingEntries = budget.maxEntries
// Why: refused entries consume no copy budget, so without a separate ceiling
// on walking itself a `.worktreeinclude` listing 1000 over-budget directories
// would pay a fresh full-limit walk for each one — the very stall this bounds.
let remainingWalk = budget.maxEntries * WORKTREE_COPY_SIZING_HEADROOM
return {
admit: async (source, { bytesAreCopied = true } = {}) => {
if (remainingWalk <= 0) {
return { withinBudget: false, reason: 'sizing' }
}
const { verdict, walked } = await measureCopySize(
source,
bytesAreCopied ? remainingBytes : Number.POSITIVE_INFINITY,
remainingEntries,
remainingWalk
)
remainingWalk -= walked
if (verdict.withinBudget) {
if (bytesAreCopied) {
remainingBytes -= verdict.bytes
}
remainingEntries -= verdict.entries
}
return verdict
},
chargeBytes: (bytes) => {
if (bytes > remainingBytes) {
return false
}
remainingBytes -= bytes
return true
}
}
}
function formatByteLimit(maxBytes: number): string {
const gigabytes = maxBytes / (1024 * 1024 * 1024)
if (gigabytes >= 1) {
return `${Number(gigabytes.toFixed(1))} GB`
}
return `${Math.max(1, Math.round(maxBytes / (1024 * 1024)))} MB`
}
const MAX_NAMED_SKIPPED_ENTRIES = 5
/** User-facing warning for entries the budget refused. Returns undefined when
* nothing was skipped so callers can spread it conditionally. */
export function formatWorktreeIncludeCopyWarning(
skipped: readonly SkippedWorktreeCopyPath[],
budget: WorktreeCopyBudget = DEFAULT_WORKTREE_COPY_BUDGET
): string | undefined {
if (skipped.length === 0) {
return undefined
}
// Why: `.worktreeinclude` allows 1000 entries and every one can be skipped,
// so enumerating them all would put a multi-kilobyte sentence in a warning.
const nameList = (entries: readonly SkippedWorktreeCopyPath[]): string => {
const shown = entries.slice(0, MAX_NAMED_SKIPPED_ENTRIES)
const names = shown.map((entry) => `"${entry.path}"`).join(', ')
const rest = entries.length - shown.length
return rest > 0 ? `${names} and ${rest.toLocaleString('en-US')} more` : names
}
const describe = (entries: readonly SkippedWorktreeCopyPath[]): string => {
const subject = entries.length === 1 ? 'entry' : 'entries'
const verb = entries.length === 1 ? 'was' : 'were'
return `.worktreeinclude ${subject} ${nameList(entries)} ${verb} not copied into the new workspace`
}
const pronoun = (count: number): string => (count === 1 ? 'it' : 'them')
// Why: an entry refused because earlier ones exhausted the sizing walk never
// approached the limits itself, so quoting them at the user would be a lie.
const overBudget = skipped.filter((entry) => entry.reason !== 'sizing')
const unsized = skipped.filter((entry) => entry.reason === 'sizing')
const sentences: string[] = []
if (overBudget.length > 0) {
sentences.push(
`${describe(overBudget)}: copying ${pronoun(overBudget.length)} would exceed the ` +
`${formatByteLimit(budget.maxBytes)} / ${budget.maxEntries.toLocaleString('en-US')} ` +
`file limit that keeps workspace creation responsive.`
)
}
const partial = skipped.filter((entry) => entry.mayBePartial)
if (unsized.length > 0) {
sentences.push(
`${describe(unsized)}: earlier entries used up the budget for measuring what to copy.`
)
}
if (partial.length > 0) {
// Why: the copy was abandoned after it started, so "copy it in manually"
// would merge into whatever the interrupted run already left behind.
sentences.push(
`${nameList(partial)} may hold a partial copy from the interrupted attempt — check ` +
`${pronoun(partial.length)} before reusing this workspace.`
)
}
sentences.push(
`Copy ${pronoun(skipped.length)} in manually if this workspace needs ${pronoun(skipped.length)}.`
)
return sentences.join(' ')
}