Commit Graph
4 Commits
Author SHA1 Message Date
Jinwoo Hong 52b6851b6e fix(worktree-create): prioritize creation Git and defer background preparation (#20722)
* fix(worktree-create): run create git commands at interactive tier, defer pool side jobs, bound queue wait by timeout

Creating a worktree on a busy machine stalled for minutes because the create's
own git competed for the same admission budget as everything else.

- The create path never set an admission tier, so it defaulted to 'status' and
  could never use the scheduler's headroom slots. It now tags the option objects
  that reach git directly: the add, the post-add listing, the base-ref probes and
  the prepared-checkout finalize. The speculative warm-up and the SSH path are
  unchanged.
- The prepared-pool re-arm is a full `reset --hard`; it ran mid-create and held a
  general slot. `consumePreparedWorktreeCreate` now returns it as a thunk the
  create runs after the startup terminal is spawned. Stale-preparation
  reclamation (`worktree unlock` / `worktree remove`) drops to 'background'.
- A command's timeout only armed once its child spawned, so a saturated queue
  could hold a 1s command indefinitely. Admission now takes the same deadline and
  raises GitCommandTimeoutError without spawning; a caller abort still reports as
  an abort.

The tier is kept off the `{ wslDistro }` routing objects: several callers test
those for emptiness to decide whether a repo has local git routing at all.

* fix(worktree-create): keep a bounded queue wait from reading as an absent base ref

The admission deadline added in the previous commit made every create-path probe's
15s/120s budget cover the queue wait. The default-base and worktree-base probes answer
`false`/`null` for any failure, so a saturated queue reported a repo that has origin/main
as having no default base and the create refused to start. Both probe families now let
`GitCommandTimeoutError` through, and the branch-name resolution loop, the push-target
configuration and the post-add listing run at the create's interactive tier so they reach
the headroom the rest of the create already uses.

Also: the deferred pool re-arm re-checks the pool inside the thunk, since `startPreparation`
replaces a map entry outright and would strand a prefetch's locked checkout with no owner;
the shared worktree scan keys on the tier so an interactive listing cannot inherit a queued
status scan's wait; and the deadline's microtask hop is gone, along with two fake-timer
`vi.waitFor` calls that jumped the clock past a 10ms budget before the grant settled.

* fix(worktree-create): preserve probe fallbacks and defer runtime replenishment

* fix(worktree-create): preserve interactive priority through prepared claims

* fix(worktree-create): prioritize CLI creation and preserve SHA probe timeouts

* fix(worktree-create): scope Git execution policy at creation boundaries

* fix(worktree-create): preserve inconclusive Git probe timeouts

* test(runtime): align creation fixtures with scoped Git execution

* test(native-chat): extract windowing layout fixture to satisfy file limit

* fix(git): restore execution-only timeouts while queued

* refactor(worktree-create): remove unrelated error-handling changes

* chore: narrow review scope and clarify preparation timing

* test(native-chat): restore fixture extraction to fix CI lint

* refactor(git): keep the admission scheduler in its original module

Reverts a move-only extraction. Inlines the single-use command-class
wrapper so the tier-resolution import fits the file's line budget.

* fix(worktrees): re-arm the prepared pool after CLI create launches terminals

The runtime create fired the pool re-arm right after materialization, so its
`reset --hard` competed with the startup agent's first git reads. Return the
thunk to the caller and fire it last, matching the desktop path.

* fix(worktrees): skip a preparation whose checkout is still running

An interactive create that claimed an in-flight preparation awaited a checkout
queued at background, so on a saturated budget it yielded to every arriving
status poller until aging promoted it. The create now misses with not_ready and
does its own add at interactive; the preparation stays armed for the next one.

Also drops the one-field policy object from the Git operation executor.

* fix(worktrees): report repo_mismatch before not_ready when selecting a preparation

The readiness filter ran before the same-repo check, so another repo's
in-flight preparation was labeled not_ready instead of repo_mismatch, hiding
the cap-thrash signal for multi-project users. The hit/miss decision is
unchanged.

* test(runtime): type the worktree-meta stub against WorktreeMeta

Main now rejects bare object parameters, and the merge picked that rule up.

* fix(worktrees): wait on in-flight preparations and re-arm the pool on failed creates

A create landing mid-checkout now claims the in-flight preparation and awaits it, as main
did. The `checkoutFinished` filter and its `not_ready` miss reason made the create skip a
prepared checkout that was seconds from done and pay a full cold add instead; on a 40k-file
repo that turned a 0.2-1.5s create into 2.4-4.3s. The preparation's own git also runs at
`status` again rather than `background`, so awaiting it does not park behind status pollers.
Only the stale reclaim stays `background`, which no create waits on.

The deferred pool re-arm now fires on every path, not just the success path. Main armed the
replacement synchronously inside the consume, so a later failure in include copy, push-target
setup, or terminal startup still left one warming. The thunk stays deferred until after
terminal startup for admission ordering, but a `finally` on the desktop create and matching
failure-path fires on the runtime create restore that guarantee. It fires exactly once.

* refactor(runtime): carry the pool re-arm in one holder

The runtime create used three mechanisms to guarantee the deferred pool re-arm fires: a
catch in the git create, a catch on materialization, and a holder fired in the managed
create's finally. The desktop create already used one holder for the same guarantee.

The holder now threads down through the create args, so the git create arms it at the point
it consumes a prepared checkout and nothing below has to handle the failure case. The thunk
already re-checks the pool before arming, so a single fire point in the outermost finally
covers every failure after the consume. Behavior is unchanged; both flipped failure-path
tests still assert exactly one fire, and each fails without the production change.
2026-09-15 19:02:50 -04:00
Neil d05dd8ef50 fix(source-control): route hosted reviews by resolved execution host (#18382)
`ForgeProvider.createReview(repoPath, input, connectionId, options)` and the
`connectionId` on `ForgeProviderRepositoryContext` carried the same collapse the
five prior migrations closed: `string | null` spells "genuinely local", "runtime
host" and "could not resolve" with one value. Because it was decided two layers
up -- `repo.connectionId ?? null` at the `hostedReview:*` IPC handlers and in
`RuntimeHostedReviewCommands` -- a row naming its owner only as
`executionHostId: ssh:<target>` ran the whole review path against this machine's
copy of a remote path (#11163): `git rev-parse`, `git status`, the base-on-remote
ref probe, the upstream divergence read, and `gh`/`glab` with no host flags.

Replace it with a required `ExecutionHostId` threaded from the decision point
through the contract, routed by #18296's `resolveGitRouteForHost`. The parameter
is removed rather than added beside, so all five implementations -- GitLab,
GitHub, Bitbucket, Azure DevOps, Gitea -- and every caller became a compile
error. None of these families carries `@ts-nocheck`, so unlike #18325 that
guarantee is real here; `orca-runtime-file-commands.ts` does, but it only
constructs `RuntimeHostedReviewCommands` with unchanged deps.

Also fixed at the sites:

- The branch cache scoped entries on `connectionId ?? ''`, so two rows at one
  path on different hosts shared one cached review, one backoff deadline and one
  invalidation. Keyed on the resolved host now, as #18377 did for its probe key.
- `hostedReview:create` resolved shared symlink paths and normalized worktree
  paths off the raw field, so an `executionHostId`-only SSH row read `orca.yaml`
  and `resolve()`d a remote POSIX path on the client. Those ask the file-holder
  question -- `getRepoSshConnectionId` -- not the dialable one.
- An SSH host with no provider now refuses inside the git-state layer instead of
  reaching the local branch, keeping "remote and unreachable" distinct from
  "local" (docs/reference/ssh-execution-boundary.md).

`runtime:` is a routing mistake inside `hostedReviewSshConnectionId` -- that
environment's server runs its own git, and the SSH target on its repo row is
nested in that server's namespace, so dialing it here reaches a same-named box of
ours. But store-backed callers ask `getRepoHostedReviewExecutionHostId` first,
which is "what may this client dial" and answers `local` for a `runtime:` row.
That is deliberate and matches #18377: the runtime registration controller only
adopts a `runtime:` stamp onto a row with no `connectionId`
(`runtimeRepoMatchesExecutionHost` refuses to match an SSH row), so the checkout
really is in this process and refusing would regress a runtime server creating
reviews for its own rows.

No wire change. `connectionId` on `CreateHostedReviewArgs`,
`CreateStackedHostedReviewArgs` and `HostedReviewCreationEligibilityArgs` in
src/shared/hosted-review.ts is untouched -- every host already ignores it in
favor of the repo row, and removing it from the request types would only churn
the schema older clients still populate. The main-side eligibility input `Omit`s
it so nothing on this side can read the ambiguous field again.
2026-09-03 01:32:46 -07:00
Neil abc099e4c7 fix(worktree): run the create-base warm-up on the routed git host (#17794)
The speculative warm-up that runs while the create composer is open resolved
refs and fetched with host Git even when the project's runtime is a WSL distro,
while both the checkout preparation it feeds (`prepareWorktreeCreateForRepo`,
which already resolves `{ wslDistro }` itself) and the real create path run
inside the distro.

The concrete cost was a discarded fetch: `getCanonicalFetchKey` namespaces the
runtime's remote-fetch cache `wsl:<distro>` vs `local`, so the warm-up's fetch
landed in a namespace create never looks at, and create fetched again. On a
Windows host with no usable host-side Git the probes also failed outright, so
that cohort got no warm-up at all.

Thread the project's worktree Git options through the prefetch (resolved by a
non-throwing helper, because an optimistic warm-up must not surface a
repair-required runtime as a failure) so every probe and fetch runs where create
runs. `gitOptions` is a required argument, so a caller cannot drop the routing
silently. Host-routed calls keep their original arity, so macOS, Linux,
native-Windows-host projects, SSH repos and folder workspaces are unchanged.

Narrower than it looks: for a repo under \\wsl.localhost\<distro>\... the probes
were already routed by cwd, and for a repo on a Windows drive letter host Git
and WSL Git read the same on-disk repository, so the answers were already
correct there. What those cohorts gain is a fetch create can reuse; what they
pay is that the probes now run inside the distro (over /mnt/c for drive-letter
repos, which also newly arms the linked-worktree routing probe) and the
speculative fetch now shares create's per-remote fetch queue, as it always has
on native platforms.

Also collapse the three byte-equivalent copies of `hasLocalWorktreeBaseRef`
(create, prefetch, remote-repo create) into one in
git/worktree-base-ref-probe.ts, drop the host-only `hasLocalCommitObject` that
caused the routing bug, and add the first routing assertions on the create-path
consumers of the now-shared probe.
2026-08-31 20:32:59 -07:00
Neil a5796ec8eb refactor(runtime): split OrcaRuntimeService and compatibility tests (#17605)
* refactor(runtime): split OrcaRuntimeService into focused modules

* test(runtime): cover admission tiers and strict worktree reconciliation

* fix(runtime): preserve owner and structured session visibility

* fix(runtime): port post-extraction compatibility fixes

* fix(runtime): preserve skill-share cancellation barrier

* test(runtime): update identity inventory after extraction

* fix(runtime): preserve hook transport environment cleanup

* fix(runtime): consolidate idle probe imports

* test(runtime): retire split file process allowlist entry

* fix(runtime): route child process types through shared boundary

* test(runtime): preserve worktree host metadata precedence

* fix(runtime): update extracted test seams

* fix(runtime): gate the split's ts-nocheck set and restore the stop-confirmed contract

Audit follow-ups for the OrcaRuntimeService split:

- Freeze the 171 @ts-nocheck files behind a ratchet so no new file can disable
  type checking. The split's linear mixin chain cannot express forward
  references yet, so the existing suppressions are grandfathered; the baseline
  may only shrink.
- Drop the stray @ts-nocheck at the end of orca-runtime-get-status.ts. It sat
  after the first statement, where TypeScript ignores it, so the module was
  already checked.
- Restore `retireRejectedPty(ptyId, stopConfirmed: boolean)` as a required
  argument. The split widened it to optional and patched the resulting error
  with `stopConfirmed === true`; an omitted argument would have silently taken
  the unverified-stop path instead of failing to compile.
- Guard that every orca-runtime-tests fragment is imported by the compatibility
  entrypoint. The fragments are .spec.ts, which no Vitest include glob matches,
  so one left out of the list would silently stop running.

* fix(runtime): restore four behaviors the OrcaRuntimeService split dropped

Audit findings against the refactor's true base (ad5ba2572e):

- retirePtyAgentLaunchAuthority collected pane keys after deleting the
  restored-authority receipt instead of before it. collectPaneKeysForPty reads
  that receipt, so a receipt-only pane lost its key and never had its agent-hook
  compatibility authority retired. on-pty-exit.ts already carried a comment
  naming this exact invariant.
- The PTY-exit path kept orchestrationMailboxNotifications.retirePty but lost
  the loop that schedules a debounced mail-pointer repoint for the dead pty's
  terminal handle and any run bound to its panes. Restores the schedule call
  count to 7, matching base.
- subscribeToPtyExit lost isPtyKnownExited's leaf fallback and its
  post-registration lifecycle-generation recheck. leavesByPtyId is rebuilt from
  the renderer graph independently of ptysById, so a leaf can outlive its pty
  record; without the fallback a caller waiting on an already-dead pty never
  gets released.
- The chain root declared `[key: string]: unknown`, which base had nowhere. It
  leaked through the exported runtime type into every consumer, so any misspelled
  member access typechecked as unknown instead of erroring, and it accounted for
  957 of the suppressed errors. Removing it costs zero type errors.

* fix(runtime): restore escalation prose and unscoped automation publication

Two more behaviors the split dropped, each with a regression test that fails
against the pre-fix code:

- The worker-exit escalation stopped deriving its title through
  buildOrchestrationTaskDisplayMetadata and inlined `task.spec` instead. That
  ignored an explicit task_title, dropped the single-line normalization and the
  80-character bound, and turned the no-spec case into a quoted, duplicated id.
  A multi-paragraph spec landed verbatim in the coordinator's banner. The
  existing 11 tests all use short single-line specs, where the derived title and
  the raw spec are identical, so none of them could see it.
  Also reverts an added `if (!handle) return` guard: the dispatch lookup is
  deliberately keyed on the pane as well, because a reminted handle no longer
  matches the row while the pane identity outlives the remint.
- updateAutomation stopped going through automationChangePublications and
  published `source` unconditionally while gating the fallback on a non-null
  destination. A destination the store can no longer name then published only
  the stale source, so subscribers scoped elsewhere kept rendering a row that
  had left them — the exact case the helper documents. The helper had been left
  with zero callers; all three sites use it again.

* fix(skills): stop swallowing lookup errors and hard-erroring on non-ssh hosts

Follow-ups from auditing the skill install path against the refactor's base:

- resolveWorktree wrapped showManagedWorktree in `.catch(() => null)`, so a
  transient git or IO failure surfaced to the user as
  skill-install-workspace-not-found with the real cause discarded. Errors
  propagate again; a genuine id mismatch still returns null.
- resolveSkillSshTarget threw skill-install-workspace-host-unavailable when the
  execution host was neither local nor ssh, on both the repo and folder
  branches. Base gated these on connectionId, so a runtime-owned repo simply
  was not an SSH install and fell through to the local path. Both return null
  again, and the error code the split invented is now unreferenced.
- listManagedSkillInstalls awaited the receipt walk and the worktree resolve in
  sequence. They are independent and either can hit disk, WSL, or an SSH scan,
  so Promise.all is restored.

Deliberately unchanged: resolving the worktree through listResolvedWorktrees
rather than showManagedWorktree, which disambiguates a worktree id colliding
across hosts and is covered by its own test, and the SSH-folder
skill-install-ssh-dispatch-required throw, which matches the repo branch.

* fix(runtime): merge duplicate worktree-logic imports

The #17448 port added a third import from ../ipc/worktree-logic, which the
code-quality oxlint config rejects under --deny-warnings. Plain oxlint does not
flag it, so it only surfaced in CI's static analysis job.

* ci: run the ts-nocheck ratchet in PR checks

pr-workflow-lint-parity requires every leaf command in `pnpm lint` to have a
matching step in pr.yml. The ratchet was wired into lint but not the workflow,
so PR CI would not have enforced it.

* Merge remote-tracking branch 'origin/main' and retry the paired-host launch evaluate

main advanced 9 commits; none touch the orca-runtime.ts this branch splits, so
nothing needed porting.

CI failed twice on `Execution context was destroyed` thrown from
headless-paired-runtime-host's first `evaluate` after launch — a different spec
each run, which is the signature of the flake #17780 describes rather than a
regression. That commit added retryTransientMainEvaluate and adopted it in five
helpers but not this call site, even though its docblock names exactly this
case: the first evaluate after electron.launch() resolves, before the app is
ready. Wrapped it the same way.
2026-08-31 19:34:55 -07:00