mirror of
https://github.com/stablyai/orca.git
synced 2026-10-07 16:02:29 +00:00
cf71ae4cb6f8202a7cc8a424aa84d127c845cbe2
335
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
328caa2160 |
fix(git): reduce queries and preserve data across execution hosts (#24602)
* fix(git): reduce queries and preserve data across execution hosts * fix(ci): exercise pinned Git and serialize mobile dependency entrypoints * fix(relay): preserve fresh diff retries after hung shared reads * test(git): wait for fetch barrier before canceling preparation * fix(i18n): describe index-preserving discard in every locale * fix(git): retain clone diagnostics and allow WSL policy startup * test(git): refresh default-base and branch-safety fixtures |
||
|
|
b5869eeaae |
fix(worktrees): a worktree delete git fails partway stays listed and can be retried (#23952)
* fix(worktrees): delete removed checkouts in git, not in Orca's file pool Local worktree removal renamed the checkout into a sibling trash root and deleted it in the background with a recursive fs.rm in the main process. That queued one request per entry on libuv's shared 4-thread file pool, so for minutes every other async fs call in the main process (the agent-session store behind chat sends, file explorer reads) waited behind the delete. `git worktree remove` now deletes the checkout inline in git's own process again, so the card stays in its Deleting state for the length of the delete while Orca's file pool stays free. No timeout applies to the call, so a large delete is never killed halfway. If git reports success but the path still exists (Git for Windows leaves junctions and their parent directories in place), the leftover is deleted with the existing removeHostTree; WSL checkouts stay with the distro. Nothing creates trash any more: the scheduling queue, rename/restore helpers and the trash_rename span are gone. The startup sweep stays to drain entries older releases left behind, and now removes each emptied trash root so the obligation ends. * fix(worktrees): let Git delete Windows checkouts with long paths enabled Removal now always runs Git's own recursive delete, and worktree creation checks out with core.longpaths on Windows, so a deep checkout Orca created could fail to delete with "Filename too long" (#6433). The Windows recovery then finishes the delete but keeps the branch. Pass the same command-scoped core.longpaths option to `git worktree remove` so Git can delete what it created. Also point the CI shard timing entry at the renamed real-git removal suite. * fix(worktrees): keep an inherited GIT_ASK_YESNO out of the worktree delete Git for Windows asks $GIT_ASK_YESNO whether to retry when a file stays locked during a recursive delete. Orca's git env inherits the user's environment, so an inherited value would run an arbitrary prompt program in the middle of a removal. Drop it for the removal call only. * perf(worktrees): run worktree deletes under their own limit, outside git admission `git worktree remove` now deletes the whole checkout in Git's own process, which takes 20-35 s on a large tree. It took a general git admission slot at status tier for that whole time, and that cap is as small as two slots on a machine with six or fewer cores, so two deletes blocked every status read. Deletes now skip general admission and queue under their own limit of two per host instead: two concurrent deletes already saturate one disk, and more only slow each other down. Leftover cleanup runs inside the same slot. * fix(worktrees): delete removed checkouts in the background and mark them removing Since the checkout is deleted by `git worktree remove` in Git's own process, a large delete takes 20-35 s. Answering the request only after that made web and mobile (30 s), paired desktop (60/180 s) and the CLI (60 s) report a failure for a delete that was still going, and mobile silently re-showed the row. The request now does everything that can refuse (lock, cleanliness, archive hook, watcher/terminal gate, terminal stop, shared-link unlink), records the removal in an in-memory table on the host and answers `removing: true`. The delete, branch cleanup and metadata purge run after it in the same order as before, and the watcher/terminal gate stays held until they finish. - Listings mark rows in the table `removing` for clients that advertise `worktree.background-removal.v1` (the desktop renderer, paired desktop and web), and leave them out for everyone else (older clients, mobile, the CLI), which already dropped the row when the request answered. - The outcome (removed, with any preserved branch, or the error) rides the existing worktrees-changed event as an optional field, sent after the row has left the table. - A repeat delete while Git runs joins it. A create at the same path or with the same branch is refused with "Cleanup is pending; try again shortly"; create's name search skips the path, so generated names move on. - Nothing is persisted: after a quit or crash Git still lists the checkout and it can be deleted again. WSL checkouts still delete inline. - `orca worktree rm` says the checkout is still being deleted. * fix(worktrees): keep the existing Deleting card until the host's Git finishes The host now answers a local worktree delete on acceptance and deletes in the background. The renderer keeps the existing delete state set until the host publishes how it ended: - The delete that asked waits for the outcome on the worktrees-changed event (local IPC or the paired runtime's client event), then runs the same teardown, preserved-branch toast and card error an inline delete did. If that event is lost to a dropped connection, a listing that shows the row gone after it was marked removing finishes the wait, and one that shows it back without the marker fails it. - Any other renderer (a reload, a paired desktop, web) sets the same delete state from the host's `removing` marker and clears it when the marker goes. A failure the host publishes lands on that card's existing error. - Web advertises `worktree.background-removal.v1` so the host sends it the marker; paired desktop does through the Electron capability list. No new component, style or state: the card reads the delete state it always did. A host that predates this answers when done without `removing`, and the renderer takes that as finished, as before. * test(worktrees): type the removal harness and projection for the node typecheck * fix(worktrees): don't fail a delete retry with an earlier attempt's buffered failure A background removal's outcome that reached this renderer with no waiter (another client's delete, a host-marked card, or one already settled from listings) was buffered for 60 s and consumed by the next delete of the same workspace, so retrying a failed delete failed at once with the old error while the host was deleting. Drop the buffered outcome before sending the request; only an outcome that arrives after it can belong to it. * fix(worktrees): let only a gap in host events settle a background delete from listings Git unlists the checkout before the host deletes the branch, cleans the push target and purges metadata, and the worktree-directory watcher refetches within 250 ms. The renderer read the missing row as a finished delete, so the waiter resolved without the preserved branch (no toast) and a failure in those last steps showed as success; the real outcome was then dropped. The listing fallback exists only for a lost outcome event, so it now applies only after this host's event stream had a gap: a new subscription or a replay after reconnect. * perf(worktrees): let a bulk delete start each same-repo checkout delete once the host accepts the last A bulk delete ran one worktree at a time per repo (#2259, for packed-refs and ref-lock races in branch cleanup). With Git now deleting each checkout for 20-35 s before the request settles, N worktrees in one repo took N times that. The renderer now queues same-repo deletes only until the host accepts each one; a parent still waits for its nested children to finish. The host serializes the branch cleanup step per repo itself, which also covers removals started by different clients. * test(worktrees): pin the host platform in the mocked removal suites so they pass on Windows Removal now passes -c core.longpaths=true on Windows, so the exact-argv assertions and command-keyed mocks never matched there (17 failures on a Windows host). Pin darwin as the add-worktree suites already do, and drive the one Windows-specific case through the same spy. * test(worktrees): type the blocked git remove result instead of a broad object The anti-slop static-analysis gate rejects `object` parameters. * test(worktrees): clear the changed-code quality gate in the removal suites Merge the duplicate node:fs import, build the mock child without a cast, read worktrees:list rows through one typed helper, and give the remaining casts a SAFETY line. * fix(worktrees): record each background delete durably and finish it after a quit or crash A quit mid-delete left git to finish the checkout on its own while the branch delete and metadata purge never ran; a crash left a normal-looking row. Each accepted local removal now writes a record beside the profile state before git starts, clears it on success or failure, and the host runs the same delete again for any record left at startup, re-deriving what remains from git and disk. An orderly quit stops the checkout delete without waiting for it. * test(worktrees): type the interrupted-removal assertions for the node typecheck * fix(worktrees): finish an interrupted delete that already removed the checkout's .git file Quit stops git worktree remove mid-delete, and Git deletes the checkout's .git file wherever it falls in directory order. Git then refuses the checkout ("validation failed ... .git does not exist") on every retry, so the startup finish failed and the row could never be deleted from Orca. A registered checkout this record owns that has lost its .git file now finishes like an unregistered one: leftover files, prune, then the branch. * fix(worktrees): let Git finish an interrupted delete, and never take a different checkout A quit or crash that stops `git worktree remove` after it deleted the checkout's .git file left a registered checkout Git refuses to remove. The previous fix deleted that leftover inside Orca's process, which is the bulk delete this change exists to avoid (and on Windows the leftover can be most of the checkout). The startup finish now rewrites the missing .git file from Git's own admin entry for that path and lets `git worktree remove --force` delete it. `git worktree repair` is not used: it also re-points every other registered path, including a checkout another repository now owns there. Orca deletes the leftover itself only when no admin entry claims the path. The startup finish forces, so it now leaves the path alone when the checkout there is not the one recorded: a registered worktree on a different branch or head, or a `.git` at a path Git already unregistered. The record is dropped and the card shows why. The record write before Git starts is now bounded (2 s, logged when exceeded) so a stalled disk cannot hold the delete, and the outcome is published before the record's clear reaches disk. * test(worktrees): compare worktree paths by value and tear down with Windows lock retries Git prints forward slashes in `git worktree list` on Windows, so the real-Git removal suites never found a joined path there: positive checks failed and negative ones passed without proving anything. They now compare Git's parsed rows by value. Teardown uses the shared retrying removeTree, since Windows can hold the deleted checkout busy for a moment after Git exits. Adds a relative-path worktree case for the .git restore (skipped before Git 2.48). * fix(worktrees): reply to a worktree delete when it has finished, not on a broadcast event A current client's delete request now waits for the host's background delete and gets its real result (removed, a preserved branch, or the error) as the reply, the way it did before the delete moved off the request. A request that arrives while the delete runs joins it and gets the same result. Every other view keeps reading the host's `removing` marker: the row leaving means the delete finished, and the row listed again without the marker shows "The delete did not finish. Try again." on a card that view had marked Deleting. A request whose reply is lost (a timeout or a dropped connection) settles the same way from a fresh listing instead of reporting a failure. Clients without the background-removal capability (mobile, the CLI, older desktops) are still answered on acceptance and have rows under removal left out of their listings. This removes the outcome on worktreesChanged and everything it needed: the renderer's outcome waiters, early-outcome buffer and TTL, per-host event-gap generations, the request pre-registration, and the accept callback bulk delete used. Bulk delete runs same-repo deletes in parallel only on this machine, whose host serializes branch cleanup per repo; SSH and paired hosts stay serialized. * test(worktrees): type the pending-removal host id in the background-removal suite * fix(worktrees): answer a delete request even when a concurrent removal of the same worktree replaced its record The desktop app's removal and the runtime removal (CLI, paired clients) coalesce separately, so both can be accepted for one worktree. The second replaced the first's record, and the first delete then finished without resolving the request waiting on it, leaving the desktop card on Deleting indefinitely. Each delete now settles the request it was started for. * fix(worktrees): run same-repo removal archive hooks and teardown one at a time on the host Local bulk delete now sends same-repo removals in parallel, so their archive hooks, terminal teardown and preflight ran at once; a hook that writes refs can race the repo's ref locks (#2259). The host now serializes each local removal up to acceptance per repo, for every client; Git's checkout delete still runs in parallel under the delete limit. * fix(runtime): keep waiting worktree deletes out of a host's foreground call slots worktree.rm now replies only after Git deletes the checkout (up to minutes), so on paired desktop and web each waiting delete held one of the host's 8 foreground call slots, and a bulk delete queued listing refreshes and every other foreground call behind it. Deletes now run in their own lane with the same bound; the 2-slot background lane stays for status polls. * fix(worktrees): join a same-worktree delete accepted while a removal waited its repo turn The desktop app and the runtime (CLI, paired clients, web) check for a running delete before they queue for the repo's acceptance turn. A delete of the same worktree from the other path, accepted while this one queued, was missed: this request re-ran the archive hook, stopped the terminals again and started a second `git worktree remove` on the directory Git was deleting. The queued acceptance now re-checks and joins the running delete. * fix(worktrees): fence a resumed delete's checkout from startup, and drop rows a listing read before the delete finished A delete a quit or crash interrupted took its terminal and file-watcher gate only when the resume job ran, after the first window was shown; session restore could open a shell or watcher inside the half-deleted checkout first, and on Windows that handle can fail the resumed git delete. Loading the records now fences each recorded path, and the resumed job takes the fence over in the same tick it takes its own gate. A listing that read git's registration before a delete finished, and replied after the removal record cleared, returned the row unmarked, so other views briefly showed "The delete did not finish". Listings now capture the pending removals before reading git and leave out a row whose delete finished successfully since; a row whose delete failed stays listed as before. * test(worktrees): keep git's auto-maintenance out of the real-git removal suite CI's Git 2.55 failed the file-pool test in teardown with ENOTEMPTY on the scratch repo's objects/pack after the test body passed: the 3,000-file commit's detached auto-maintenance was still writing a pack. The scratch repo now disables auto-maintenance and auto-gc. * fix(worktrees): one archive-hook approval covers a same-repo bulk delete again Local same-repo deletes now start together, so each queued its trust prompt with a state snapshot taken before the first prompt was answered; approving the first still showed the same prompt once per remaining worktree. The queued check now reads the store when its turn comes. * fix(worktrees): a delete Git fails partway stays listed with its error; Delete retries it `git worktree remove --force` drops the checkout's registration even when it cannot delete a file (root-owned files, `chflags uchg`, a read-only Windows directory). Orca lists workspaces from Git, so the row vanished after the error, leaving the checkout, the branch and Orca's metadata with no way to retry. - A background delete that fails with the checkout still on disk, unregistered, and still the removed checkout's own leftover keeps its durable removal record with the error (`failure`) instead of clearing it. Every other failure clears it as before. - Local listings (desktop list/list-all/detected, runtime list/ps/detected) add a row for each such record, carrying `removalError`, and for a pending removal whose checkout Git no longer lists (shown as removing). - Delete on that row (desktop IPC and runtime worktree.rm) runs the recorded removal again: terminal teardown, then the leftover, prune, branch and metadata, under the per-host delete limit. - The record ends on a successful retry, when the checkout is gone (listing or startup), when a different checkout takes the path, or on forget-local. Startup never retries a failed record. - The finish's unregistered-path rule accepts a `.git` file naming the admin entry Git removed (the leftover's own) and still refuses any other `.git`. The removal table and listing projection move out of the background removal module into worktree-removal-table.ts and worktree-removal-listing.ts. * fix(renderer): show a failed delete's host error on its card A row the host lists with removalError gets the existing delete-state error (no new element), cleared when the host stops listing it failed. A row this view marked Deleting that comes back failed, and a lost delete reply settled from the listing, report the host's error instead of the generic one. * test(worktrees): type the failed-removal listing and refresh mocks * fix(worktrees): a failed delete's retry never removes a checkout Git registers at the path again The retry replays the recorded choices (force, branch deletion) that were made for the unregistered leftover. If the user removed the leftover and `git worktree add`ed the same branch at the path, the new checkout matched the record's branch and head, so Delete force-removed it with its uncommitted files, skipping the normal delete's cleanliness check. The retry now refuses a registered checkout and lets the record go, so the next Delete takes the normal path. * fix(worktrees): a failed delete's record ends at startup once its repo is removed from Orca * fix(renderer): a failed delete's row offers Remove from Orca * test(worktrees): type the failed-removal IPC test's module mocks without casts * fix(worktrees): Delete picks retry or a normal delete from Git's current listing * fix(worktrees): a failed delete's retry checks the leftover again right before deleting it * fix(worktrees): Remove from Orca reaches paired clients and matches the failed row's own host * fix(worktrees): a failed delete's retry re-lists only its own repo's authorized roots * fix(worktrees): a second Delete joins a retry already running, and the startup finish keeps its last-resort delete * fix(renderer): a failed delete's card says it failed, and Delete keeps the error for its dialog * fix(renderer): a failed delete's dialog shows the host's error, and its card label keeps the full error a hover away * style(worktrees): import the removal result types in one statement * test(worktrees): a runtime listing right after a failed delete shows the failed row, not the cached scan * fix(worktrees): pass the runtime retry's PTY-stop waiver in the shape the waiver invariant pins * fix(worktrees): drop Remove from Orca from failed-delete rows A failed delete stays listed with its error and Delete retries it; the separate forget item, its dialog copy and forget's failed-record clearing are removed. The startup clear for repos no longer in Orca keeps matching the local copy only. * test(worktrees): wait for the dropped record's write before the failed-removal suite tears down |
||
|
|
de77c4b065 |
fix(worktrees): list a folder once when git reports it twice (#24357)
When git lists the same folder twice (a leftover worktree registration that points at the main checkout), Orca's runtime listing turned each line into its own worktree with the same id, so `orca worktree current`, `active` and `branch:` failed with selector_ambiguous, and paired clients saw a duplicate row. The runtime scan now keeps git's first row per folder, the rule the desktop sidebar already uses. Separately, for a bare or separate-git-dir repo added through a linked worktree, the scan no longer relabels the main row with that worktree's folder (it relabels only when the folder's git dir is the common git dir), so the worktree keeps its own row and branch in the CLI and the sidebar. No extra git command runs. Part of #23631: the "Profile state writer command timed out" toast in that issue has a separate cause. |
||
|
|
d6d2795da5 |
chore(worktree): include create timing and spare outcome in the workspace create events (#24483)
* chore(worktree): include create timing and spare outcome in the workspace-created event The workspace_created and workspace_create_failed events gain optional, numbers-and-enums-only fields built from what the create already measured: total and per-phase durations, the prepared-checkout hit/miss and miss reason, the execution host (local/WSL/SSH), a worktree count bucket, how many other creates were in flight, whether the repo has a post-checkout hook (file existence only, probed after the create returns), and for a failure the phase it died in plus elapsed time. No new git process runs; consent and opt-out are unchanged. * fix(worktree): attribute failed_phase by error, label WSL-path repos, skip the hook check with telemetry off - failed_phase now names the outermost timed step the thrown error (or its cause) left, so a caught failure or a concurrent sibling step can no longer be misattributed; the old-relay SSH error keeps its cause so it still reads as git_worktree_add. - execution_host follows the same rule Git routing uses, so a \\wsl.localhost repo reads wsl. - The post-checkout hook check does not read the repo when telemetry is disabled. - Privacy page mentions the miss reason code and the failed step. * test(worktree): pin the old-relay SSH add error to git_worktree_add through its cause * fix(worktree): name the create event field sets for their role, and type the old-relay test's caught error * fix(worktree): send create events from runtime creates and record what the spare checkout did Runtime creates (CLI, agents, phone app, paired clients, orchestration, server automations) reuse prepared checkouts like the app's own creates, but recorded no timing and sent no events. Both entry points now start one shared sender (workspace-create-telemetry.ts), so every create sends exactly one event with the same fields, plus create_entry_point (app | runtime). Spare-checkout fields: - concurrent_preparations: peak prepared-checkout builds and background discards running during the create, excluding the one it used; the window closes before the create's own re-arm starts. - prepared_checkout_claim / prepared_checkout_discard phases, so on a miss git_worktree_add minus the prepared_checkout_* phases is the plain checkout. - prepared_checkout_reset (none | base_moved | retargeted) replaces the retargeted flag; prepared_checkout_origin (prefetch | rearm) on hits. - workspace_create_failed carries the spare outcome and its wait. - repo_index_size_bucket from one stat of .git/index in the existing post-create probe (telemetry on, local/WSL only, 2 s cap). * fix(worktree): add spare build and idle time, the re-arm prefetch origin, and a tracked-file count - prepared_checkout_build_ms / prepared_checkout_idle_ms on hits: from arming the spare to ready, and how long it sat ready before the claim (0 when the create waited). readyAt is recorded in the pool's existing ready handler. - prepared_checkout_origin gains rearm_then_prefetch: an automatic re-arm that the dialog prefetch then asked for too, so rearm means the re-arm alone. - repo_file_count_bucket replaces the index byte size: the entry count from the 12-byte index header, which is the same in every index version; left out for a split or sparse index. - The shared sender never lets a failed send change the create's result or error; it logs instead and still ends the create's concurrency membership. - Tests pin the runtime SSH create's timing hand-off and the throwing-send cases on both entry points. * fix(worktree): leave out the file count under any sparse checkout and time spare builds monotonically - The repo probe also reads .git/config.worktree, where git sparse-checkout --sparse-index writes index.sparse, and omits the file count whenever sparse checkout or a sparse index is on in either file, with Git's boolean spellings. core.hooksPath there is honoured too. - prepared_checkout_build_ms / _idle_ms use performance.now(), like every other duration; the build is timed from its own start (buildStartedAt). - The origin field comment names all three values. * fix(worktree): keep the spare's build time on its first build and count worktrees by lock reason - prepared_checkout_build_ms runs from the first build's start (including any wait for the base fetch it is built on) to its first ready; a later tip refresh no longer restarts it, though it still counts as new preparation work. prepared_checkout_idle_ms runs from the latest ready (build or refresh) to the claim. - The worktree count reads each .git/worktrees entry's locked file and leaves out entries whose lock reason names an Orca preparation, the way the listing does, instead of subtracting this process's spares. That covers spares from other processes, crash leftovers and spares being discarded, and cannot run one low while a spare's admin dir does not exist yet. It has its own 1.5 s cap inside the probe. |
||
|
|
34ae0933e4 |
fix(worktrees): keep creation fast in large repositories (#24346)
* fix(worktrees): remove repeated scans and keep prepared checkouts fresh * fix(worktrees): reclaim unlocked fallback preparations safely * refactor(worktrees): simplify creation ownership and idle maintenance * fix(git): keep ref maintenance armed after an index-only pass An idle attempt that found the pack index due but refs still cooling down returned without rescheduling, so loose refs from the arming fetch waited for the next write instead of the ref cooldown. |
||
|
|
12b8ef8c0b |
fix(worktree): update local main safely, once per branch, alongside the checkout (#23698)
* fix(worktree): retry local main refresh through git lock contention and skip false alarms * fix(worktree): overlap the local main refresh with the checkout and run one refresh per repo at a time * fix(worktree): skip the local base refresh when the create makes that branch itself Creating a workspace named feature-x from origin/feature-x runs `worktree add -b feature-x`, which now overlaps the refresh. The refresh's drift probe could see refs/heads/feature-x missing and its presence probe then see it (the add just wrote it), which reported "not fast-forward" and showed a sticky "Local feature-x was not refreshed" warning. `-b` refuses an existing branch, so there is nothing to refresh in that case: skip it on the local, prepared-checkout and SSH create paths. The SSH overlap tests move to their own file so the existing suite stays under the line limit. * fix(worktree): say plainly what happens after the local base refresh queue wait expires * test(worktree): prove SSH local base refreshes of one repo run one at a time * test(worktree): drop type assertions from the SSH refresh overlap test mocks * fix(worktree): fast-forward local main with one host-owned merge --ff-only per branch Moves the whole local base refresh into one shared routine that runs on the execution host (main process for local and WSL repos, the relay for SSH), so the app no longer keeps a second copy of the checks, queue and retry. A checked-out branch now moves with merge --ff-only (hooks, auto-gc and autostash off) instead of status-then-reset --hard, which silently overwrote an untracked file the new commit adds and could discard an edit or a commit made after the check. A free branch moves with a compare-and-swap update-ref that writes a reflog message. Status reads no longer take index.lock. Creates of one branch share one run plus at most one trailing run; a create waits at most 30 s and never starts a competing mutation. The failure toast is keyed by repo and branch because every create that joined a run reports the same fact. * fix(worktree): fast-forward local main even when the repo requires signed merges With merge.verifySignatures=true, the owner-checkout fast-forward refused an unsigned origin/main tip, so every create warned "Local main was not refreshed" where the old reset moved main. The new workspace is already created from that same unsigned commit, and a branch that is not checked out moves without a signature check, so the refusal protected nothing. Turn the setting off for this one merge, like the hooks, gc and autostash overrides. * fix(worktree): clear git read caches when a shared local main update lands late The update of local main can finish after a create stopped waiting for it, so the shared run now invalidates git read caches itself. The index.lock real-git test also no longer reads the developer's global git config. * fix(worktree): never overwrite an ignored file when fast-forwarding local main A plain `git merge --ff-only` silently replaces an ignored file (for example a local `.env`) at a path the new commit starts tracking. Pass `--no-overwrite-ignore` so git refuses instead and the create reports the checkout as having local changes. Supported on the fast-forward path since well before Git 2.25. Also make the relay test for one-refresh-per-branch hold the first merge until the second request has reached the relay, so it fails without the coalescing. * fix(worktree): keep the local main update a plain fast-forward whatever the user's merge settings say A per-branch mergeOptions such as '-s ours' or '--squash', or pull.twohead=ours, made the update create a merge commit that dropped upstream, or stage upstream without moving main, while reporting success. The command now clears the branch's mergeOptions and passes the strategy and signature choice on the command line, which beats any config. After the move Orca confirms local main is exactly the target before reporting it updated. The exact command also runs in the Git 2.25 compatibility suite. * fix(worktree): make the Git 2.25 fast-forward contract pass in CI and rerun on every change to it The new real-Git contract for the local main fast-forward wrote a post-merge hook into .git/hooks, which does not exist when the repo is created by the uninstalled Git 2.25.5 build CI uses (no templates), so the Git compatibility check failed. Create the directory first. The Git compatibility check also did not run when only the fast-forward module changed, so a later edit to its merge arguments (for example a flag Git 2.25 lacks) would skip the one check that tests them. Add the module to the check's paths. * fix(worktree): answer every create from a local main update toward its own base Creates from different remotes' main (origin/main and upstream/main) shared one queued update per repo and branch, which ran only the latest caller's target: a create could get no result for its own base, or a false "not refreshed" warning computed for another remote's main. The per-branch runner now queues one run per distinct target, still one at a time per branch, and only callers toward the same target share a queued run. Applied in the app and the relay. * test(worktree): record the third create's result in the mixed-remote burst tests and update the toast id rationale * chore(worktree): correct the toast id rationale |
||
|
|
6f2a7d05c9 |
fix(worktrees): let git delete removed checkouts so chat sends never wait behind them (#23837)
* fix(worktrees): delete removed checkouts in git, not in Orca's file pool Local worktree removal renamed the checkout into a sibling trash root and deleted it in the background with a recursive fs.rm in the main process. That queued one request per entry on libuv's shared 4-thread file pool, so for minutes every other async fs call in the main process (the agent-session store behind chat sends, file explorer reads) waited behind the delete. `git worktree remove` now deletes the checkout inline in git's own process again, so the card stays in its Deleting state for the length of the delete while Orca's file pool stays free. No timeout applies to the call, so a large delete is never killed halfway. If git reports success but the path still exists (Git for Windows leaves junctions and their parent directories in place), the leftover is deleted with the existing removeHostTree; WSL checkouts stay with the distro. Nothing creates trash any more: the scheduling queue, rename/restore helpers and the trash_rename span are gone. The startup sweep stays to drain entries older releases left behind, and now removes each emptied trash root so the obligation ends. * fix(worktrees): let Git delete Windows checkouts with long paths enabled Removal now always runs Git's own recursive delete, and worktree creation checks out with core.longpaths on Windows, so a deep checkout Orca created could fail to delete with "Filename too long" (#6433). The Windows recovery then finishes the delete but keeps the branch. Pass the same command-scoped core.longpaths option to `git worktree remove` so Git can delete what it created. Also point the CI shard timing entry at the renamed real-git removal suite. * fix(worktrees): keep an inherited GIT_ASK_YESNO out of the worktree delete Git for Windows asks $GIT_ASK_YESNO whether to retry when a file stays locked during a recursive delete. Orca's git env inherits the user's environment, so an inherited value would run an arbitrary prompt program in the middle of a removal. Drop it for the removal call only. * perf(worktrees): run worktree deletes under their own limit, outside git admission `git worktree remove` now deletes the whole checkout in Git's own process, which takes 20-35 s on a large tree. It took a general git admission slot at status tier for that whole time, and that cap is as small as two slots on a machine with six or fewer cores, so two deletes blocked every status read. Deletes now skip general admission and queue under their own limit of two per host instead: two concurrent deletes already saturate one disk, and more only slow each other down. Leftover cleanup runs inside the same slot. * fix(worktrees): delete removed checkouts in the background and mark them removing Since the checkout is deleted by `git worktree remove` in Git's own process, a large delete takes 20-35 s. Answering the request only after that made web and mobile (30 s), paired desktop (60/180 s) and the CLI (60 s) report a failure for a delete that was still going, and mobile silently re-showed the row. The request now does everything that can refuse (lock, cleanliness, archive hook, watcher/terminal gate, terminal stop, shared-link unlink), records the removal in an in-memory table on the host and answers `removing: true`. The delete, branch cleanup and metadata purge run after it in the same order as before, and the watcher/terminal gate stays held until they finish. - Listings mark rows in the table `removing` for clients that advertise `worktree.background-removal.v1` (the desktop renderer, paired desktop and web), and leave them out for everyone else (older clients, mobile, the CLI), which already dropped the row when the request answered. - The outcome (removed, with any preserved branch, or the error) rides the existing worktrees-changed event as an optional field, sent after the row has left the table. - A repeat delete while Git runs joins it. A create at the same path or with the same branch is refused with "Cleanup is pending; try again shortly"; create's name search skips the path, so generated names move on. - Nothing is persisted: after a quit or crash Git still lists the checkout and it can be deleted again. WSL checkouts still delete inline. - `orca worktree rm` says the checkout is still being deleted. * fix(worktrees): keep the existing Deleting card until the host's Git finishes The host now answers a local worktree delete on acceptance and deletes in the background. The renderer keeps the existing delete state set until the host publishes how it ended: - The delete that asked waits for the outcome on the worktrees-changed event (local IPC or the paired runtime's client event), then runs the same teardown, preserved-branch toast and card error an inline delete did. If that event is lost to a dropped connection, a listing that shows the row gone after it was marked removing finishes the wait, and one that shows it back without the marker fails it. - Any other renderer (a reload, a paired desktop, web) sets the same delete state from the host's `removing` marker and clears it when the marker goes. A failure the host publishes lands on that card's existing error. - Web advertises `worktree.background-removal.v1` so the host sends it the marker; paired desktop does through the Electron capability list. No new component, style or state: the card reads the delete state it always did. A host that predates this answers when done without `removing`, and the renderer takes that as finished, as before. * test(worktrees): type the removal harness and projection for the node typecheck * fix(worktrees): don't fail a delete retry with an earlier attempt's buffered failure A background removal's outcome that reached this renderer with no waiter (another client's delete, a host-marked card, or one already settled from listings) was buffered for 60 s and consumed by the next delete of the same workspace, so retrying a failed delete failed at once with the old error while the host was deleting. Drop the buffered outcome before sending the request; only an outcome that arrives after it can belong to it. * fix(worktrees): let only a gap in host events settle a background delete from listings Git unlists the checkout before the host deletes the branch, cleans the push target and purges metadata, and the worktree-directory watcher refetches within 250 ms. The renderer read the missing row as a finished delete, so the waiter resolved without the preserved branch (no toast) and a failure in those last steps showed as success; the real outcome was then dropped. The listing fallback exists only for a lost outcome event, so it now applies only after this host's event stream had a gap: a new subscription or a replay after reconnect. * perf(worktrees): let a bulk delete start each same-repo checkout delete once the host accepts the last A bulk delete ran one worktree at a time per repo (#2259, for packed-refs and ref-lock races in branch cleanup). With Git now deleting each checkout for 20-35 s before the request settles, N worktrees in one repo took N times that. The renderer now queues same-repo deletes only until the host accepts each one; a parent still waits for its nested children to finish. The host serializes the branch cleanup step per repo itself, which also covers removals started by different clients. * test(worktrees): pin the host platform in the mocked removal suites so they pass on Windows Removal now passes -c core.longpaths=true on Windows, so the exact-argv assertions and command-keyed mocks never matched there (17 failures on a Windows host). Pin darwin as the add-worktree suites already do, and drive the one Windows-specific case through the same spy. * test(worktrees): type the blocked git remove result instead of a broad object The anti-slop static-analysis gate rejects `object` parameters. * test(worktrees): clear the changed-code quality gate in the removal suites Merge the duplicate node:fs import, build the mock child without a cast, read worktrees:list rows through one typed helper, and give the remaining casts a SAFETY line. * fix(worktrees): record each background delete durably and finish it after a quit or crash A quit mid-delete left git to finish the checkout on its own while the branch delete and metadata purge never ran; a crash left a normal-looking row. Each accepted local removal now writes a record beside the profile state before git starts, clears it on success or failure, and the host runs the same delete again for any record left at startup, re-deriving what remains from git and disk. An orderly quit stops the checkout delete without waiting for it. * test(worktrees): type the interrupted-removal assertions for the node typecheck * fix(worktrees): finish an interrupted delete that already removed the checkout's .git file Quit stops git worktree remove mid-delete, and Git deletes the checkout's .git file wherever it falls in directory order. Git then refuses the checkout ("validation failed ... .git does not exist") on every retry, so the startup finish failed and the row could never be deleted from Orca. A registered checkout this record owns that has lost its .git file now finishes like an unregistered one: leftover files, prune, then the branch. * fix(worktrees): let Git finish an interrupted delete, and never take a different checkout A quit or crash that stops `git worktree remove` after it deleted the checkout's .git file left a registered checkout Git refuses to remove. The previous fix deleted that leftover inside Orca's process, which is the bulk delete this change exists to avoid (and on Windows the leftover can be most of the checkout). The startup finish now rewrites the missing .git file from Git's own admin entry for that path and lets `git worktree remove --force` delete it. `git worktree repair` is not used: it also re-points every other registered path, including a checkout another repository now owns there. Orca deletes the leftover itself only when no admin entry claims the path. The startup finish forces, so it now leaves the path alone when the checkout there is not the one recorded: a registered worktree on a different branch or head, or a `.git` at a path Git already unregistered. The record is dropped and the card shows why. The record write before Git starts is now bounded (2 s, logged when exceeded) so a stalled disk cannot hold the delete, and the outcome is published before the record's clear reaches disk. * test(worktrees): compare worktree paths by value and tear down with Windows lock retries Git prints forward slashes in `git worktree list` on Windows, so the real-Git removal suites never found a joined path there: positive checks failed and negative ones passed without proving anything. They now compare Git's parsed rows by value. Teardown uses the shared retrying removeTree, since Windows can hold the deleted checkout busy for a moment after Git exits. Adds a relative-path worktree case for the .git restore (skipped before Git 2.48). * fix(worktrees): reply to a worktree delete when it has finished, not on a broadcast event A current client's delete request now waits for the host's background delete and gets its real result (removed, a preserved branch, or the error) as the reply, the way it did before the delete moved off the request. A request that arrives while the delete runs joins it and gets the same result. Every other view keeps reading the host's `removing` marker: the row leaving means the delete finished, and the row listed again without the marker shows "The delete did not finish. Try again." on a card that view had marked Deleting. A request whose reply is lost (a timeout or a dropped connection) settles the same way from a fresh listing instead of reporting a failure. Clients without the background-removal capability (mobile, the CLI, older desktops) are still answered on acceptance and have rows under removal left out of their listings. This removes the outcome on worktreesChanged and everything it needed: the renderer's outcome waiters, early-outcome buffer and TTL, per-host event-gap generations, the request pre-registration, and the accept callback bulk delete used. Bulk delete runs same-repo deletes in parallel only on this machine, whose host serializes branch cleanup per repo; SSH and paired hosts stay serialized. * test(worktrees): type the pending-removal host id in the background-removal suite * fix(worktrees): answer a delete request even when a concurrent removal of the same worktree replaced its record The desktop app's removal and the runtime removal (CLI, paired clients) coalesce separately, so both can be accepted for one worktree. The second replaced the first's record, and the first delete then finished without resolving the request waiting on it, leaving the desktop card on Deleting indefinitely. Each delete now settles the request it was started for. * fix(worktrees): run same-repo removal archive hooks and teardown one at a time on the host Local bulk delete now sends same-repo removals in parallel, so their archive hooks, terminal teardown and preflight ran at once; a hook that writes refs can race the repo's ref locks (#2259). The host now serializes each local removal up to acceptance per repo, for every client; Git's checkout delete still runs in parallel under the delete limit. * fix(runtime): keep waiting worktree deletes out of a host's foreground call slots worktree.rm now replies only after Git deletes the checkout (up to minutes), so on paired desktop and web each waiting delete held one of the host's 8 foreground call slots, and a bulk delete queued listing refreshes and every other foreground call behind it. Deletes now run in their own lane with the same bound; the 2-slot background lane stays for status polls. * fix(worktrees): join a same-worktree delete accepted while a removal waited its repo turn The desktop app and the runtime (CLI, paired clients, web) check for a running delete before they queue for the repo's acceptance turn. A delete of the same worktree from the other path, accepted while this one queued, was missed: this request re-ran the archive hook, stopped the terminals again and started a second `git worktree remove` on the directory Git was deleting. The queued acceptance now re-checks and joins the running delete. * fix(worktrees): fence a resumed delete's checkout from startup, and drop rows a listing read before the delete finished A delete a quit or crash interrupted took its terminal and file-watcher gate only when the resume job ran, after the first window was shown; session restore could open a shell or watcher inside the half-deleted checkout first, and on Windows that handle can fail the resumed git delete. Loading the records now fences each recorded path, and the resumed job takes the fence over in the same tick it takes its own gate. A listing that read git's registration before a delete finished, and replied after the removal record cleared, returned the row unmarked, so other views briefly showed "The delete did not finish". Listings now capture the pending removals before reading git and leave out a row whose delete finished successfully since; a row whose delete failed stays listed as before. * test(worktrees): keep git's auto-maintenance out of the real-git removal suite CI's Git 2.55 failed the file-pool test in teardown with ENOTEMPTY on the scratch repo's objects/pack after the test body passed: the 3,000-file commit's detached auto-maintenance was still writing a pack. The scratch repo now disables auto-maintenance and auto-gc. * fix(worktrees): one archive-hook approval covers a same-repo bulk delete again Local same-repo deletes now start together, so each queued its trust prompt with a state snapshot taken before the first prompt was answered; approving the first still showed the same prompt once per remaining worktree. The queued check now reads the store when its turn comes. |
||
|
|
a781a602a8 |
test: retire duplicate cases that replay an owner across a re-export or provider shim (#24114)
Resolves 208 candidate pairs where the same case title appears verbatim in two or more
files, produced by a repo-wide scan calibrated against a known positive. 46 case
declarations removed across 32 files, 798 lines gone. No file deleted whole, no
production code touched.
The headline result is the measurement, not the deletions: across the three buckets that
reported in detail, the signal ran roughly 86% false-positive (3/42, 9/42, and the rest).
It has good recall and poor precision, and it reorders a reading queue rather than
replacing one. Calibrating a detector against a known positive proves recall, not
precision.
What the deletions were:
- Duplicate invocation through a re-export shim. `native-chat-tool-summary.ts` is a
ten-line `export {...} from '../../../../shared/native-chat-tool-summary'`, and
`agent-status.ts:161` is `export { isExplicitAgentStatusFresh } from
'./pane-agent-evidence'`. Cases on the shim side were byte-equivalent to the owner's
with no rendering or transport hop.
- Provider-local replays of a shared helper: three `repository-ref` providers that are
each `createRemoteRefProbeCache(parseXRef)` and contribute nothing to transient
handling; two `local-pty` and `daemon/session` tables replaying
`shell-startup-output-scanner`, whose owner additionally checks every split point.
- A reader-side replay of store policy. `runtime-worktree-agent-rows-structured.test.ts`
asserted an attention-to-blocked mapping; the reader contains zero `attention` or
`blocked` tokens and copies `state` through. The mapping lives in
`structuredAgentSessionAgentStatus`. Consistent with
`docs/reference/agent-status-store.md`: readers keep only presentation policy.
- Constructor-only subclass duplication: the shared capability-cache case is covered by
`codex-app-server-capability-cache.test.ts`, whose ten cases include the identical
title plus all four risks `docs/reference/git-compatibility.md` names — first fallback,
later cached call, concurrent probes, per-host isolation.
- A private predicate duplicated at a real boundary, varying only a path passed straight
into the shared predicate.
Why most pairs were KEPT, because the false positives are principled rather than noise:
- Two independent execution hosts. `src/relay/git-handler-*` and `src/main/git/*` are
separate Git implementations that cannot import each other and hold separate capability
caches, exactly as the compatibility doc requires; the repo already ships
`status-branch-line-total-relay-parity.test.ts` to pin the duality deliberately. Neither
side's argv, timeout or cache regression is visible to the other.
- Deliberately duplicated production siblings: Codex vs Claude (different account fields,
different CLIs, different wire protocols), gitea vs bitbucket (`/pulls/42` vs
`/pullrequests/42`), gl-utils vs gh-utils (separate in-flight maps). Same contract
shape, different implementations — an identical title is the correct naming.
- Shared-predicate consumers: one side tests the predicate, the other tests a caller's
wiring to it. A caller that forgot to call the predicate passes the shared test.
In a codebase with intentional provider and host symmetry, identical test titles are
expected, and the signal cannot distinguish "copied" from "parallel by design" because
both produce the same prose. Only reading both bodies separates them.
Verified: 6,968 desktop test files pass; the three modified mobile files pass (39 cases);
`check-reliability-gates.mjs` 140 gates; nothing under
`mobile/src/test-support/rpc-recording/` or `mobile/rpc-foundation/goldens/` touched.
62 local failures across 12 files were each accounted for and none is caused by this
change: `browser-manager-tab-identity`, `browser-manager-viewport-ownership`,
`session-scanner-codex-workers` and `managed-hook-script-refresh` all fail identically in
a pristine `origin/main` worktree; five `mobile-web-app-*-render` tests need Playwright
browsers this machine lacks; `structured-agent-session-restart-ownership` and
`ssh-remote-commands` pass in isolation and fail only under concurrent load.
|
||
|
|
b7209b5ae9 |
perf(git): relist only the repo whose worktrees changed, and stop blocking main on sync git (#23998)
* perf(git): stop blocking main on the open-on-remote git cascade
`getRemoteFileUrl` ran up to 6 sequential `gitExecFileSync` calls on the Electron
main thread — `remote get-url`, then `getDefaultBaseRef`'s `symbolic-ref` plus up
to four `rev-parse --verify` probes — each with its own 15s timeout and no yield
between them.
A complete async twin already existed (`getDefaultBaseRefAsync` ->
`resolveDefaultBaseRefViaExec`, sharing DEFAULT_BASE_REF_PROBES), so the sync
cascade is deleted rather than converted. `getRemoteUrl`, `getRemoteFileUrl` and
`getRemoteCommitUrl` become async; all four downstream callers were already async
(`filesystem-git-url-handlers` inside `ipcMain.handle`, `runtime-git-diff-commands`
async methods) and the provider contract already typed both wrappers
`Promise<string | null>`, so no new async plumbing was needed.
Removes 3 of the 10 `gitExecFileSync` sites and the confusing name collision with
the unrelated async `getDefaultBaseRef` in hosted-review-creation-git-state.
The base-ref regression tests keep their coverage, repointed at the public async
`getBaseRefDefault`.
* perf(git): resolve the repo root in one sync spawn instead of two
getGitRepoRoot ran `rev-parse --is-inside-work-tree` and then `rev-parse
--show-toplevel` as separate blocking spawns. Each sync git call holds the main
thread for up to its whole 15s timeout, so the spawn count is the cost — and this
function is called twice per "Add Project" on a linked worktree, once directly and
once through getLinkedWorktreeMainRepoRoot's self-recursion.
Combined into one invocation. Safe only here: in a bare repo the combined form
exits non-zero, and both that throw and the plain `false` already land on the same
marker-scan fallback. probeGitRepo deliberately does NOT combine — it has to read
`false` cleanly to go on and detect a bare repo, which the combined form's exit 128
would misread as indeterminate.
* perf(git): rebuild only the repos whose authorized roots actually changed
One worktree create called `invalidateAuthorizedRootsCache()`, which dirties every
registered owner. The next authorization-requiring IPC then rebuilt by listing EVERY
repo — and the rebuild never consulted `dirty` when choosing what to list, so `dirty`
gated only whether a rebuild ran, not its scope. At 58 repos that is 58
`git worktree list` spawns, roughly ten seconds of git wall-clock through an
admission budget of four, to rediscover roots one repo changed.
Both halves were needed; scoping the invalidation alone changed nothing.
- `markAuthorizedRootsOwnerDirty` dirties a single owner, reusing the per-owner
primitives `registerWorktreeRootsForRepo` already used. It leaves `baseRevision`
and the per-repo revision map alone — that pair is the global side-effect-token
fence, and bumping it would retire in-flight tokens for untouched repos.
- `rebuildAuthorizedRootsCache(store, onlyDirty)` re-lists only owners that are
dirty, have no listing yet, or still hold recovered roots (those are retired by
comparison against a fresh listing, so skipping them would strand them as
authorized). Only `ensureAuthorizedRootsCache` passes `onlyDirty`; an explicit
rebuild keeps re-listing everything because callers use it to force a refresh —
`filesystem-auth.test.ts` pins that contract.
`invalidateAuthorizedRootsCacheForRepo` wraps the primitive and falls back to the
global form for an unknown owner or a missing store, rather than silently skipping an
invalidation and leaving a stale allowlist. Applied to the worktree-create path.
Changes that can alter the owner SET (store swap, host/WSL re-routing, nested-repo
import, folder->git upgrade) stay global. Removal paths are not converted yet.
The allowlist contents are unchanged and the failure direction is a false denial
rather than a false allow. The relist predicate is split into its own module so it is
testable alone and the cache file stays inside its line budget without a suppression.
* test(perf): measure what git orchestration actually costs the main thread
The existing churn probe (ORCA_MAIN_THREAD_DIAGNOSTICS=1) reported spawn-initiation
cost for git/gh/glab only — its 7 call sites all sit inside git/command-runner — so
it was blind to `spawnProcess`/`runProcess`, the repo's own mandated wrapper, and to
the blocking `execFileSync('ps')` per PTY resize. That understated total churn across
115 main call sites.
- `spawn-observer.ts`: a settable seam, since shared code cannot import src/main.
Unregistered in the daemon/relay/CLI, where it costs one boolean check.
- `spawnProcess` brackets `nodeSpawn` and reports; exec-file-capture's own report is
removed because it routes through runProcess and would double-count.
- `posix-pty-foreground-group` now reports its full blocking duration. Note this
lands on the daemon, not main, whenever the daemon hosts the PTY.
- `ORCA_UNMINIFIED_MAIN=1` build flag, because a minified main bundle cannot
attribute CPU-profile self time to real function names. Defaults unchanged.
- `main-thread-git-cost.spec.ts` + `analyze-main-cpuprofile.mjs`: sweeps concurrency
against real registered repos, captures the churn lines and a V8 CPU profile of
main per phase.
What it found, which is why this is worth keeping: at the width-4 admission ceiling
(~90 git:status/s) main sees ZERO event-loop gaps over 50ms and a worst gap of 23ms,
and is 85% idle. Git orchestration does not stall the main thread. Of the cost it
does incur, spawn-init is 58%, parse 5%, stdout drain 4%.
* test(perf): name the inspector params type the anti-slop gate requires
The broad `object` parameter trips anti-slop(no-object-parameters); the only
Profiler call that passes params sends `{ interval }`.
|
||
|
|
fed1eca486 |
test: stop restating internal tuning constants, keep the ones that are contracts (#23950)
Removes ~74 assertions of the form `expect(SOME_CONSTANT).toBe(<literal>)` where the literal is an internal tuning value — a timeout, retry count, debounce interval, cache TTL, circuit-breaker window, Tailwind class string. Those cannot fail for any reason a user would notice: they fail only when someone deliberately changes the number, and then the test is simply updated. They are copies of the declaration. The same pattern is NOT junk when the exact value is observable outside this process, so those were deliberately kept: - terminal byte contracts: `\r`, `\x03` ETX, Kitty escapes, `\x1b[?1;2c`; - wire and capability values: `agent.launch.v2`, protocol 3 / min-compatible 2, daemon per-feature boundary versions (a daemon survives app updates, so those pin what an old field daemon may be trusted with), relay header tokens; - security invariants: the `127.0.0.1` bind default, an empty iframe `sandbox`; - values external processes read: exit code 78 (EX_CONFIG) and exit code 3 (systemd `RestartPreventExitStatus`), `ORCA_AGENT_SESSION_SPAWN_TOKEN`, `npx skills …` commands users paste, on-disk journal schema versions, the `orca_<hash>` filename prefix the fish sweeper matches; - third-party names: expo-router's `unstable_settings` / `ErrorBoundary`, iOS Safari's 16px zoom threshold. Where a case asserted a relation rather than a literal — `A < B`, a sum of parts, a cap compared against a sibling budget — the relation stays and only the literal went. Test-only changes: no production file is touched and no test file is deleted. |
||
|
|
31012aeb09 |
test: remove assertion-free probes, copied inventories and export-shape checks (#23816)
Second audit wave, targeting three more junk patterns: - assertion-free cases that run code and assert nothing, so they pass no matter what the code does; - inventory literals re-typed from a production declaration, where the only way the assertion can fail is someone editing one of the two copies; - export key-set and export-shape loops (`typeof x === 'function'` over every export) that restate what TypeScript already enforces. Yield is much smaller than wave 1 on purpose: the assertion-free scanner has a high false-positive rate, because many flagged blocks assert through a shared helper or their oracle is "this must not throw". Those were kept. `mobileWebCheckArgs` in `config/scripts/run-mobile-web-app-checks.mjs` is de-exported — after the inventory comparison went away, nothing outside the module read it. |
||
|
|
3976ad4c59 | perf(test): remove obsolete structural snapshots (#23777) | ||
|
|
e13631ee53 |
Prioritize workspace opening over replacement checkout preparation (#23013)
* Prioritize workspace opening over replacement checkout preparation * Preserve Git hook semantics and exercise preparation edge cases |
||
|
|
7ea01279cd |
feat(search): bundle ripgrep for local, WSL, and SSH search (#22396)
* feat(search): bundle ripgrep for local, WSL, and SSH search Ship @vscode/ripgrep-universal's prebuilt rg for all six relay platforms in every desktop artifact. Local and WSL searches spawn the bundled binary and drop the git ls-files / git grep fallbacks; SSH deploys upload the remote's binary once per ripgrep version and the relay prefers it over PATH rg. * fix(search): address bundled ripgrep review findings - Key the SSH ripgrep cache on the binary's content hash; a package bump is the only update step - glibc verifier: read arch tokens below the slice root and accept static ELFs (arm64 release blocker) - Ship ripgrep/PCRE2/musl license notices; bundle rg with orcad - Packaged builds never spawn a bare rg; report fd pressure as transient - SSH: install rg before sweep/GC, size-validate installs, back off instead of disabling on launch failure - Scope Dependabot to @vscode/ripgrep-universal; revert unrelated lockfile churn * chore(search): drop bundled-ripgrep reference doc; assert full packaging layout parity * refactor(search): one entry point for spawning the bundled ripgrep Local Quick Open, Quick Open path search, the Explorer name filter, and runtime text search each repeated the same three steps: resolve the bundled command, spread in the WSL distro, spread in the WSL shell expression. Fold that into spawnBundledRipgrep so one place owns the rule that a bare 'rg' must never reach spawn, and simplify the resolver's command/packaged checks. Restore the AGENTS.md ripgrep rule dropped alongside its reference doc in |
||
|
|
7da9788c83 |
fix(updater): send the gh token and cache the release picker's build list (#21902)
* fix(updater): send the gh token and cache the release picker's build list The dev build picker listed releases through api.github.com with no Authorization header, so it spent GitHub's 60/hour per-IP bucket that every unauthenticated caller on the same network shares, and it refetched on every settings mount and channel click. When that bucket ran dry the picker showed "No builds found" with a rate-limit line even though GitHub was healthy and the user's own token had its full quota. Attach the local `gh auth token` when there is one so the request draws from the user's 5000/hour bucket, fall back to unauthenticated on a rejected token or a spent token bucket, cache the list per channel for five minutes in the main process (the refresh button forces a reload), classify 403 by the rate-limit headers, and say when the limit resets. Fixes #21898 * fix(updater): don't trip breaker for secondary rate limits GitHub sends x-ratelimit-remaining: 0 on both primary and secondary limits. Secondary limits carry Retry-After and shouldn't block all core gh commands — only the primary limit should trip the shared breaker. * Scope gh rate limits to execution environment * Add build list cache hint to release channel settings Inform users that build lists are cached for 5 minutes and they can refresh to check for new builds immediately. This makes the cache behavior visible and explains why a manual refresh is necessary to bypass the cache. |
||
|
|
4085e1cf60 |
fix(memory): release stale session registries (#21734)
* fix(memory): bound session and lifecycle registries * fix(memory): bound transient filesystem registries * fix(memory): cap path and locale caches * fix(memory): bound runtime recovery registries * fix(memory): bound host mirror gap verdicts * fix(memory): bound shell startup env cache * fix(memory): bound gitlab host context cache * fix(memory): release removed ssh generations * fix(memory): expire cloud refresh replay guards * fix(memory): release retired plugin generations * fix(memory): bound plugin log key retention * fix(memory): bound automation authority generations * fix(memory): bound native chat enrichment cache * fix(memory): bound web session tracking generations * fix(memory): bound codex credential absence paths * fix(memory): bound WSL canonical path cache * fix(memory): bound sparse checkout cache * fix(memory): bound shared directory cache * fix(memory): bound advertised URL scan snapshots * fix(memory): bound automation manager cache * fix(memory): bound web session reorder intents * fix(memory): bound web session focus intents * fix(memory): bound web session handoffs * fix(memory): bound automation dispatch tokens * fix(memory): bound host mirror waiters * fix(memory): bound retained session activity * fix(memory): bound retained session activity * fix(memory): bound web session close intents * fix(memory): bound cloud session cache * fix(memory): bound WSL home cache * fix(memory): bound SSH capability cache * fix(memory): bound trust grant cooldowns * fix(memory): bound WSL auth drain state * fix(memory): bound Linear workspace credential cache * fix(memory): bound local Git capability cache * fix(memory): bound WSL Git environment cache * fix(memory): bound WSL Git environment cache * fix(memory): bound WSL preflight cache * fix(memory): keep hot cache entries warm * fix(memory): preserve generation fences across eviction * fix(memory): close remaining eviction fences * fix(memory): align evicted upstream generations * fix(memory): trim successful capability probes * fix(auth): retain expired refresh replay evidence --------- Co-authored-by: m4air <m4air@Mac.localdomain> |
||
|
|
c34b944136 |
feat(github): bind projects to a specific gh account (#13664)
* feat(github): bind projects to a specific gh account Adds per-project `Repo.ghAccount` so repo-scoped gh calls (create-worktree issue/PR search, work items, hosted-review reads and mutations) run as the bound account via ephemeral child-env token injection instead of the globally active gh login. Multi-account resolution is capability-gated (gh >= 2.40) and fails closed when the bound account or host is unavailable; Project View stays ambient by design. Repository settings gains a section for selecting or clearing a keyring-backed account (shadcn `Select`), with mixed-version "not enforced" handling for older remote runtimes. Attached `-Rhost/owner/repo` forms are covered by the host-drift guard and its tests; es/ja/ko/zh catalogs carry the section's strings. `getLocalProjectGhExecOptions` centralizes the binding lookup so every gh execution path picks it up, including the Electron `hostedReview:*` handlers that previously stayed on the ambient login. `gh auth token` (a keyring read) is exempt from the rate-limit breaker gate so a tripped bucket cannot turn a bound-token resolve into a false "unavailable". The `ghAccount` update field and the two binding RPC methods live in the shared RPC params contract; the generated catalog is regenerated. Fixes #13612 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012B3QEP5iP4WGGEpPLtkHqA * fix(settings): make GitHub account refresh secondary * fix(github): satisfy strict casting quality checks * test(rpc): use runtime fixture for repo binding * fix(github): preserve project account for PR worktree lookups * test(rpc): avoid incomplete runtime settings fixture * fix(i18n): add GitHub account refresh label * fix(i18n): refresh runtime required catalog * fix(windows): preserve mobile patch bytes --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Co-authored-by: Neil <neil@stably.ai> |
||
|
|
85d1ffc072 |
fix: accept enterprise managed GitHub owner logins (#20450)
Unify owner validation across project pickers and repository overrides. Preserve EMU usernames in API and auth-status branch-prefix resolution, with regression coverage. Co-authored-by: Neil <neil@stably.ai> |
||
|
|
3e5eb0329a |
feat(cli): orca search over the agent session index (#20514)
* feat(cli): orca search over the agent session index `orca search <query>` calls PR 5's `aiVault.searchSessions` over the CLI's existing runtime RPC, against the host `--environment` / `--pairing-code` selects and no other. `orca search --index-status` calls `aiVault.searchStatus`. It is the proof the contract works with no panel. Every flag maps onto a contract field and nothing else: `--scope`, `--fresh`, `--limit`, `--cursor`, repeatable `--agent` and `--path`, `--since`, `--sort`, `--debug`, `--json`. No fan-out, no merged output, no `--host`. One command rather than a `search status` subcommand: the query is a bare positional, so `orca search status` could not be told apart from searching for the word "status". `--status` is unavailable because `orchestration task-list --status <state>` already owns the name as a valued flag. No new runtime capability. PR 5 decided an explicit `method_not_found` refusal maps to `unavailable/no-service`, so reusing `createSessionSearchClient` gives an old host a plain "this host runs no session search service" answer at exit 0 instead of a raw JSON-RPC error. `CommandSpec.repeatableFlags` scopes repeatability per command, because `--agent` must repeat for search and stay single-valued for `worktree create`. `help.ts` sat exactly at max-lines, so `skills-command-flag-help.ts` becomes `command-scoped-flag-help.ts` carrying both tables at the same call-site size. * refactor(cli): drop the search type assertions main's casting gate now rejects Main gained a `consistent-type-assertions: never` scan in the changed-code gate after this branch was cut, and it reported twelve assertions in the new files. The four in the argument parser were avoidable. `readEnum` now keeps the value `find` returns, which already carries the narrow type, and the agent filter goes through an `isAiVaultAgent` predicate over a `Set<string>` instead of widening the agent tuple. The test now narrows the printed envelope by shape and re-reads the printed result through `AiVaultSearchResponseSchema`, so the JSON assertions are checked rather than claimed, and the flag table is typed so its callback needs no cast. One assertion is left, for the structural fake client, with the SAFETY rationale AGENTS.md requires. * fix(cli): sanitize host strings and scope pre-command repeatable flags Route every host-supplied string the search formatter prints through the escape stripper, and resolve the repeatable-flag set from the command tokens ahead when a flag sits before the command. * refactor(cli): resolve repeatable flag rules once per command * fix(cli): clarify session search availability and SSH scope * feat(cli): hide orca search until the settings toggle ships `orca search` stays dispatchable but leaves every discovery surface: root help, group help, unknown-command suggestions, and `agent-context --json`. `buildAgentContext` did not filter hidden specs, so it also stops leaking the hidden `terminal stop`. |
||
|
|
52b6851b6e |
fix(worktree-create): prioritize creation Git and defer background preparation (#20722)
* fix(worktree-create): run create git commands at interactive tier, defer pool side jobs, bound queue wait by timeout
Creating a worktree on a busy machine stalled for minutes because the create's
own git competed for the same admission budget as everything else.
- The create path never set an admission tier, so it defaulted to 'status' and
could never use the scheduler's headroom slots. It now tags the option objects
that reach git directly: the add, the post-add listing, the base-ref probes and
the prepared-checkout finalize. The speculative warm-up and the SSH path are
unchanged.
- The prepared-pool re-arm is a full `reset --hard`; it ran mid-create and held a
general slot. `consumePreparedWorktreeCreate` now returns it as a thunk the
create runs after the startup terminal is spawned. Stale-preparation
reclamation (`worktree unlock` / `worktree remove`) drops to 'background'.
- A command's timeout only armed once its child spawned, so a saturated queue
could hold a 1s command indefinitely. Admission now takes the same deadline and
raises GitCommandTimeoutError without spawning; a caller abort still reports as
an abort.
The tier is kept off the `{ wslDistro }` routing objects: several callers test
those for emptiness to decide whether a repo has local git routing at all.
* fix(worktree-create): keep a bounded queue wait from reading as an absent base ref
The admission deadline added in the previous commit made every create-path probe's
15s/120s budget cover the queue wait. The default-base and worktree-base probes answer
`false`/`null` for any failure, so a saturated queue reported a repo that has origin/main
as having no default base and the create refused to start. Both probe families now let
`GitCommandTimeoutError` through, and the branch-name resolution loop, the push-target
configuration and the post-add listing run at the create's interactive tier so they reach
the headroom the rest of the create already uses.
Also: the deferred pool re-arm re-checks the pool inside the thunk, since `startPreparation`
replaces a map entry outright and would strand a prefetch's locked checkout with no owner;
the shared worktree scan keys on the tier so an interactive listing cannot inherit a queued
status scan's wait; and the deadline's microtask hop is gone, along with two fake-timer
`vi.waitFor` calls that jumped the clock past a 10ms budget before the grant settled.
* fix(worktree-create): preserve probe fallbacks and defer runtime replenishment
* fix(worktree-create): preserve interactive priority through prepared claims
* fix(worktree-create): prioritize CLI creation and preserve SHA probe timeouts
* fix(worktree-create): scope Git execution policy at creation boundaries
* fix(worktree-create): preserve inconclusive Git probe timeouts
* test(runtime): align creation fixtures with scoped Git execution
* test(native-chat): extract windowing layout fixture to satisfy file limit
* fix(git): restore execution-only timeouts while queued
* refactor(worktree-create): remove unrelated error-handling changes
* chore: narrow review scope and clarify preparation timing
* test(native-chat): restore fixture extraction to fix CI lint
* refactor(git): keep the admission scheduler in its original module
Reverts a move-only extraction. Inlines the single-use command-class
wrapper so the tier-resolution import fits the file's line budget.
* fix(worktrees): re-arm the prepared pool after CLI create launches terminals
The runtime create fired the pool re-arm right after materialization, so its
`reset --hard` competed with the startup agent's first git reads. Return the
thunk to the caller and fire it last, matching the desktop path.
* fix(worktrees): skip a preparation whose checkout is still running
An interactive create that claimed an in-flight preparation awaited a checkout
queued at background, so on a saturated budget it yielded to every arriving
status poller until aging promoted it. The create now misses with not_ready and
does its own add at interactive; the preparation stays armed for the next one.
Also drops the one-field policy object from the Git operation executor.
* fix(worktrees): report repo_mismatch before not_ready when selecting a preparation
The readiness filter ran before the same-repo check, so another repo's
in-flight preparation was labeled not_ready instead of repo_mismatch, hiding
the cap-thrash signal for multi-project users. The hit/miss decision is
unchanged.
* test(runtime): type the worktree-meta stub against WorktreeMeta
Main now rejects bare object parameters, and the merge picked that rule up.
* fix(worktrees): wait on in-flight preparations and re-arm the pool on failed creates
A create landing mid-checkout now claims the in-flight preparation and awaits it, as main
did. The `checkoutFinished` filter and its `not_ready` miss reason made the create skip a
prepared checkout that was seconds from done and pay a full cold add instead; on a 40k-file
repo that turned a 0.2-1.5s create into 2.4-4.3s. The preparation's own git also runs at
`status` again rather than `background`, so awaiting it does not park behind status pollers.
Only the stale reclaim stays `background`, which no create waits on.
The deferred pool re-arm now fires on every path, not just the success path. Main armed the
replacement synchronously inside the consume, so a later failure in include copy, push-target
setup, or terminal startup still left one warming. The thunk stays deferred until after
terminal startup for admission ordering, but a `finally` on the desktop create and matching
failure-path fires on the runtime create restore that guarantee. It fires exactly once.
* refactor(runtime): carry the pool re-arm in one holder
The runtime create used three mechanisms to guarantee the deferred pool re-arm fires: a
catch in the git create, a catch on materialization, and a holder fired in the managed
create's finally. The desktop create already used one holder for the same guarantee.
The holder now threads down through the create args, so the git create arms it at the point
it consumes a prepared checkout and nothing below has to handle the failure case. The thunk
already re-checks the pool before arming, so a single fire point in the outermost finally
covers every failure after the consume. Behavior is unchanged; both flipped failure-path
tests still assert exactly one fire, and each fails without the production change.
|
||
|
|
8edec28a55 |
fix(worktree): keep a WSL checkout case so delete cannot take the twin branch (#20273)
* fix(worktree): let a POSIX path keep its case on a Windows desktop `canonicalWorktreePath` folded case whenever `process.platform` was win32, without asking what the path itself was. A WSL or SSH checkout is spelled `/home/alice/ws/feature` on a Windows desktop too, and ext4 is case-sensitive, so `/home/alice/ws/Feature` and `/home/alice/ws/feature` — two real checkouts on two real branches — collapsed into one row. `removeWorktree` picks the row it is about to remove with that comparison and reads the branch off it. Requesting `/home/alice/ws/feature` removed the right directory (the path rides in argv) and then ran `git branch -d -- Feature`. The same wrong row feeds `assertWorktreeUnlockedForRemoval`, so a locked twin blocks an unlocked delete and an unlocked twin lets a locked one through. Whose filesystem a path names is a property of the path, not of the desktop reading it, so a POSIX-absolute path now takes POSIX rules at any platform and a POSIX/Windows pair is never equal — `win32.resolve` would otherwise give the POSIX path a drive root and manufacture the equality. Windows drive and UNC paths, including WSL UNC aliases, keep folding case as before. Two call sites already carried private copies of this rule (`isSameCommonDirPath`, `ipc/worktree-path-comparison`); this is the same rule at the source. The removal path is the one that never got one. * fix(worktree): keep a WSL checkout's case through the UNC spelling too The first commit gave POSIX-absolute paths POSIX case rules, which is right but does not reach the WSL case it claimed. `listWorktreesStrict` runs every listed path through `translateWorktreePath`, so git-in-the-distro's `/home/alice/ws/Feature` arrives as `\\wsl.localhost\Ubuntu\home\alice\ws\Feature` and the POSIX branch never sees it. The removal suite mocks `translateWslOutputPaths` to identity, which is why the end-to-end test passed without exercising the translation production always applies. Driving the real translator, the original defect survived unchanged: a request naming `...\ws\feature` ran `git worktree remove --force ...\ws\feature` and then `git branch -d -- Feature`. The filesystem behind `\\wsl.localhost` is ext4, so the UNC spelling is case-sensitive for the same reason the Linux spelling is — except where Windows genuinely folds: the `\\wsl$` share alias, the distro name, and a drvfs `/mnt/<letter>` tail, which really is a Windows volume. `foldWslUncPathCaseInsensitiveParts` already draws exactly that line and `git-fetch-head-lock` already depends on it, so this reuses it rather than writing a fourth copy of the rule. Windows drive paths keep folding whole. The end-to-end case now drives the real translator instead of the mock, so the translation cannot go missing again without the test noticing. |
||
|
|
231e805b1e |
fix(lint): enable anti-slop/no-shape-in-symbol-names (#20785)
Flip `anti-slop/no-shape-in-symbol-names` from "off" to "error" and clear
every violation under src, config, tests and mobile.
What the rule bans
------------------
The case-insensitive substring "shape" in any JS/TS identifier: variables,
functions, parameters, types, type parameters, class members, private names,
object-literal keys and JSX identifiers. The one exemption is a statically
accessed member read owned by another value (`zodObject.shape` is fine), so
third-party APIs stay readable without a suppression.
"Shape" names a value's structure rather than its domain role. `UserShape`,
`validateArgShape` and `errorShape` all tell you the symbol is "an object
with some fields" -- which is already what a type says -- while saying
nothing about what the value is for or who owns it. The rule forces the
name to carry the domain instead.
Violations fixed
----------------
689 violations across 109 files at baseline (verified by re-running the
audit against the pre-change tree with the rule set to "error").
Fix pattern
-----------
Rename for the domain role, not the structure:
-type FieldShape = 'list' | 'map' | 'whole'
-const FIELD_SHAPES = { ... } satisfies Record<keyof Observation, FieldShape>
+type FieldEncoding = 'list' | 'map' | 'whole'
+const FIELD_ENCODINGS = { ... } satisfies Record<keyof Observation, FieldEncoding>
-function assertGitPushTargetShape(target: unknown): void
+function assertValidGitPushTarget(target: unknown): void
-function describeReadDirPathShape(p: string): ReadDirPathKind
+function classifyReadDirPath(p: string): ReadDirPathKind
Predicates became statements about the value (`isDeltaShapedProviderFrameKind`
-> `isDeltaProviderFrameKind`, `isDeleteShapedDiscardEntry` ->
`discardDeletesEntryFile`, `isSkillsCliAgentKeyShaped` ->
`isUsableSkillsCliAgentKey`). Type aliases dropped the suffix where the
remaining name was already unambiguous (`GhGraphqlErrorShape` ->
`GhGraphqlError`).
No wire-visible name was renamed: no IPC or RPC channel, stream opcode,
request/response param, persisted field, or i18n key. The `--shape=symlink|copy`
CLI flag read by .github/workflows/skill-update-roundtrip.yml is unchanged --
only the local variable holding it was renamed.
Exemptions
----------
They are file-scoped entries in config/oxlint-anti-slop.json, not inline
`oxlint-disable` comments. An inline directive naming an anti-slop rule reads
back as an UNUSED directive under the root lint scan, which does not load this
plugin -- the changed-code quality gate counts that warning, so the comment form
cannot be used for a rule that lives only in this config.
* src/renderer/src/components/browser-pane/annotate/**:
in the screenshot annotator a "shape" is the drawn geometry -- pen, arrow,
rect, ellipse, highlight. That is a genuine domain noun, and it pervades
every symbol in the module.
* repo-icon.tsx, repo-header-project-actions.tsx, mobile MobileRepoIcon.tsx:
lucide exports the icon component as `Shapes`. The name is theirs, and the
matching REPO_LUCIDE_ICONS key is the persisted icon name shared with the
desktop picker -- renaming it would orphan saved repo icons.
* src/shared/onboarding-state-types.ts, src/shared/constants.ts:
`shapedSidebar` is a persisted onboarding-checklist field and a telemetry
enum member; renaming it would orphan saved state.
* src/shared/rpc-contract/rpc-send-params.ts: matching zod's own literal `shape`
property is what selects the ZodObject branch of the conditional type.
No exemption was added merely to avoid a rename. Eight symbols initially
suppressed as "a cross-module refactor outside this change" were proven to have
zero non-TypeScript references repo-wide and renamed instead.
Zod's `ZodRawShape` needed no exemption at all: `Readonly<Record<string,
z.ZodType>>` is its definition, so repo-update-params.ts and
ui-update-value-tolerance-params.ts spell it out instead. Likewise
telemetry-event-classification.ts now reads `.shape` through an `in` narrowing,
which also retires two pre-existing type assertions; three more assertions the
rename had dragged onto changed lines (two `JSON.parse` sites, one node:sqlite
row read) became annotations and an explicit row mapping.
Verified
--------
* Audit reports zero violations; confirmed the rule genuinely fires by
planting a probe violation.
* node config/scripts/run-typecheck-projects-in-parallel.mjs exits 0.
* Vitest over src/shared, src/main/github/project-view, the annotate module,
the repo-icon components and the Chromium SameSite electron spec: all green.
* All 66 removed "shape" identifiers grepped repo-wide across every file type;
none survive.
* node config/scripts/generate-rpc-params-catalog.mjs --check exits 0.
* node --check on every changed .mjs; oxfmt clean on all changed files.
* `pnpm run check:code-quality:changed` reports 0 findings.
Not machine-verified: the 3 mobile/ files (its Vitest run cannot resolve
`expo/tsconfig.base.json` in this worktree), and the WSL- and Playwright-gated
specs. All are rename- or comment-only hunks, read in full.
|
||
|
|
bfdec26352 |
fix(lint): enable anti-slop/no-object-parameters (#20781)
The rule rejects the broad `object` type on any function input (declarations, expressions, arrows, methods, call/construct signatures, function types), plus local aliases and unions that resolve to `object`. `object` accepts every non-primitive while exposing no properties, so it documents nothing and pushes callers into assertions at the boundary. Fixes all 185 violations across src, config, tests and mobile, and flips the rule from "off" to "error" in config/oxlint-anti-slop.json. Approach: replace each `object` input with the type its owner already has. Most sites took an existing domain type or a type-only import (36 added); 40 new aliases name shapes that had none. Where a value is genuinely only compared by reference, it gets a named identity token instead of a shape -- `Record<string, never>`, the built-in `WeakKey`, or a `unique symbol` brand, matching the branding already used in src/shared. Same treatment for WeakMap and Map key parameters. Two `as unknown as` casts became unnecessary once the parameter carried a real type and were removed; no new casts were added. Suppressions added: none. No `oxlint-disable` for this rule anywhere, and no max-lines disable or per-file bump. Three files sat exactly at their max-lines cap, so the added type imports were made line-neutral rather than suppressed: - src/main/ipc/browser.ts exports the existing guest-registration args type (renamed BrowserGuestArgs) so browser.test.ts reuses it on one line. - pane-scroll.ts takes TerminalScrollIntentTarget through the existing pane-manager-types import via a type-only re-export. - direct-rpc-client.ts drops the identity parameter entirely: the session check moved into the sendProbe callback that owns the token. Verified: anti-slop config reports zero violations over src config tests mobile; run-typecheck-projects-in-parallel exits 0; 144 affected test files pass (1749 tests); oxlint and oxfmt clean on all changed files. Mobile has no runnable test/typecheck target in this worktree (expo is not installed), so its 6 files were typechecked against a standalone config and diffed against the base branch -- error sets are byte-identical, including test files. |
||
|
|
775a932651 |
fix(git): distinguish binary absence from missing cwd on spawn ENOENT (#20798)
* fix(repos): preserve unknown Git availability * fix(git): distinguish binary absence from missing cwd on spawn ENOENT Node reports ENOENT for both a missing git binary and a missing working directory during spawn. The fix checks specifically for spawn syscall, then verifies the cwd exists to disambiguate. This prevents reporting "no Git" when the error is actually a missing working directory. Centralizes probe logic in a reusable function; other failures cause rejection so callers preserve the unknown status instead of collapsing to false. |
||
|
|
99062ed80b |
fix(worktrees): preserve unverifiable disk witness (#20713)
* fix(worktrees): preserve unverifiable disk witness * fix(worktrees): follow gitdir/commondir markers in disk witness The disk witness validates created worktrees by reading the repo's common directory from disk. Previously it only checked for a direct .git directory and returned a status object that conflated different failure modes. Now it properly follows .gitdir and commondir pointer files to locate the true common directory, fixing detection on repos with linked git directories (worktrees, submodules) and WSL scenarios. Error handling is simplified: definitive absence returns undefined, other read failures throw with proper cause chains, eliminating the ambiguous "unverifiable" state that would mask real errors. * fix: validate gitdir marker targets are directories When a .git marker points to a missing or non-directory path, that's unverifiable—not the same as an absent .git file (bare repo). Validate accessibility before reading commondir to catch these errors clearly. |
||
|
|
41e42beab4 |
fix(worktrees): safely remove prunable git-file registrations (#20617)
Preserve checkout files and the named branch when removing a positively attested malformed Git-file registration. Reject file/symlink targets in deferred directory deletion. Verified exact head with 75 focused tests including actual Git malformation, preserved marker/file bytes and branch HEAD. Independent review and complete product CI passed. WSL routing is covered by unit tests; direct SSH fails safely without local recovery. Fixes #17316 |
||
|
|
ee1a0a4e2d |
fix(git): avoid Windows tree kills after the command has exited (#20606)
Validated and independently reviewed OMP integration fix. |
||
|
|
31db2774f8 |
fix(git): skip upstream remote probes when the remote is absent (#18455)
* fix(git): skip upstream remote probes when the remote is absent Issue and PR resolvers listed remotes by probing `git remote get-url upstream` on every poll, including origin-only clones where that remote cannot exist. List remotes once, cache against git config, and skip the probe unless `upstream` is present. * fix(git): avoid stale remote probe cache entries * fix(github): observe origin repository probe failures * fix(github): observe verified origin probe failures * fix(github): skip missing upstream probe for PR lists * test(github): scope the #9171 lazy-resolution guard to default-branch commands The guard asserted that no git command runs for an open PR, using "no git at all" as a proxy for "no default-branch resolution". Remote-name listing is a separate concern, so allow it and keep every other command forbidden; the symbolic-ref/rev-parse resolution this issue is about stays unreachable. --------- Co-authored-by: Neil <neil@stably.ai> |
||
|
|
471463f4ce |
perf(git): normalize tracked discard paths once per operation (#20299)
* perf(git): normalize tracked discard paths once per operation * chore(git): drop the now-dead tracked-pathspec re-export |
||
|
|
490937a043 |
perf(git): classify bulk discard paths once (#20274)
* perf(git): classify bulk discard paths once * style: format bulk discard partition |
||
|
|
afb43ea914 |
perf: avoid duplicate repository directory probes (#20253)
Co-authored-by: Orca Worker <orca-worker@localhost> |
||
|
|
c84007c541 | feat(rpc): generate a shared params catalog from the host registry, gated on parse parity (#19961) | ||
|
|
75c1f32f81 | fix(gh): log when gh/glab is killed at its deadline (#18555) | ||
|
|
15dabf8d0b |
perf(worktree): overlap base refresh with prepared checkout (#18998)
* Stop obsolete worktree preparations when evicted or expired * Let worktree preparation proceed during stale reclamation * Verify creation during stalled stale worktree reclamation * Preserve preparation ownership until Git removal starts * test: keep artifact share fixtures unexpired across calendar dates (#18955) * perf(worktree): overlap base refresh with prepared checkout |
||
|
|
a63a4579cf |
Let worktree creation proceed during stale preparation reclamation (#18967)
* Stop obsolete worktree preparations when evicted or expired * Let worktree preparation proceed during stale reclamation * Verify creation during stalled stale worktree reclamation * Preserve preparation ownership until Git removal starts * test: keep artifact share fixtures unexpired across calendar dates (#18955) |
||
|
|
e2270fe94d |
Stop evicted and expired worktree preparations (#18951)
* Stop obsolete worktree preparations when evicted or expired * fix(worktree): skip discard retries for registrations an aborted checkout already removed An evicted or expired preparation now aborts its checkout, which self-discards the registration before the pool's own discard runs. That second discard failed with "is not a working tree" and was enrolled for up to three retries on later preparations for the same host, spawning Git only to fail again and warning that the path stays registered when it was already gone. Also accept fs.watch events without a filename in the abort real-Git test, and add an opt-in bench (ORCA_WORKTREE_PREPARATION_CANCEL_BENCH=1) that measures a fresh checkout's wall time with obsolete checkouts left running versus aborted. |
||
|
|
295684dc6d | perf: skip unrelated shared symlink probes during Git status (#18918) | ||
|
|
d7767fb196 |
perf(worktree): remove redundant creation and terminal startup work (#18793)
* perf(worktree): remove redundant creation and terminal startup work * test(worktree): cover optimized creation call signatures Preserve explicit branch adoption, WSL callback routing and sparse cleanup expectations. * perf: preserve user Git checkout worker settings * perf(git): skip malformed remote base probes * perf(cli): avoid loading other agent hooks for Codex preflight * fix(build): retain Codex preflight entry for packaged CLI * test(ssh): wait for replacement PTY before lease recovery input * test(ssh): verify recovered shell execution and lease ownership * test(electron): reap isolated macOS crash reporters on teardown * test: allow either observed self-exit snapshot ordering * test: capture frozen-host input recovery evidence |
||
|
|
a823f97d63 | perf(worktree): skip remote probes with no possible result (#18821) | ||
|
|
06ca54ae7c |
fix(test): stop detached git maintenance racing the divergence fixture teardown (#18810)
`worktree-base-divergence-real-git.test.ts` builds cap-sized histories (100 and 101 commits). Every `git commit` detaches `git maintenance run --auto`, whose commit-graph task arms at 100 new commits, so the fixture reliably spawns a background `git commit-graph write --split` that keeps creating `.git/objects/info/commit-graphs` entries after the synchronous exec returns. The `afterEach` recursive remove is then deleting `.git/objects` underneath a live writer and dies with ENOTEMPTY — which is how "counts drift in both directions" failed on main. Reuse the existing `GIT_FETCH_SKIP_AUTO_MAINTENANCE_CONFIG_ARGS` (it already covers modern maintenance and legacy auto-gc, so it holds at the Git 2.25 baseline) in the fixture's git helper. Traced spawns of `git commit-graph write` over a full run of this file: 4 before, 0 after. The production path under test only runs `rev-list` and `merge-base`, neither of which triggers auto-maintenance, so there is nothing to fix outside the fixture. |
||
|
|
0a821e5bc8 |
fix(crash-reporting): make the own-Chromium gate a real choke point, and stop a refusal leaking the root (#18459)
* fix(crash-reporting): make the own-Chromium gate a real choke point
Round-3 review found the guard was not the choke point its own comments
claimed: six pid-addressed `taskkill /pid <pid> /t /f` families in main were
ungated and uninstrumented, so the stale-pid shape stayed producible and a
`selfInitiatedTreeKillCount: 0` could read as exculpatory when it was not.
- Gate the remaining main-process families: the git command-runner abort, the
notebook-cell and automation-precheck timeouts.
- Turn the `src/shared` seam into the gate itself (`process-tree-kill-gate`), so
the runProcess choke point, the codex app-server deadline kill and the
ephemeral-VM recipe kill ask the same decision. Those three are compiled into
the CLI/relay too and cannot import main; main installs the guard at preflight.
- Ratchet (`main-process-tree-kill-gate.test.ts`): a new pid-addressed taskkill
in main that skips the gate fails, and the allowlist entries must still exist.
- Give pid-addressed kills eviction priority in the 32-entry ring: 32 routine
`win-pty-job` teardowns from a window-close burst no longer evict the one
entry that discriminates a self-kill from an external one.
- Correct the coverage doc, which described the uninstrumented Windows sites as
POSIX `process.kill(-pid)` group kills and omitted the git and codex paths.
* fix(crash-reporting): keep a refused tree-kill from leaking the root it owns
A refusal must block the pid-addressed tree walk, not the termination. Five of
the six gated sites returned on refusal with no fallback, so a refused
`taskkill /pid /t /f` left git.exe, a timed-out notebook cell, an automation
precheck or an ephemeral-VM recipe running while the caller reported it stopped.
The root kill is addressed by the child handle, which cannot reach the recycled
pid the refusal is about, so it stays correct and required on that path.
Also fixes the ring eviction the scope preference introduced: with the ring
saturated by pid-addressed kills, the only non-pid-addressed entry is the one
just pushed, so the splice evicted itself and the detail came back `{}` --
byte-identical to the external-kill arm, in the window-close case the guard
exists for. Eviction now excludes the newest entry and falls back to FIFO.
Tests: refusal now asserts the root kill at all six sites, and the ring covers
the saturated-pid ordering as well as round 3's group-burst ordering.
* fix(crash-reporting): stop a refused tree-kill leaking the commit-message agent, and count call sites
Two round-5 blocking findings, both open on main and on both branches.
`killSourceControlAgentProcess` had no root-kill fallback on its win32 arm: the
taskkill was the only termination, so once the own-Chromium gate could refuse it
the promise resolved having killed nothing. Both callers do
`terminationComplete ??= killSourceControlAgentProcess(child)` and then release
the managed-home lock on that promise, so a refusal left the local Codex/Claude
commit-message agent running while the caller reported it stopped -- the
lock-contention failure the taskkill was added for. Same fix as the six sibling
sites: the handle-addressed root kill cannot reach the recycled pid the refusal
is about, so it stays correct and required on that path.
The ratchet was file-granular, not call-site granular: one gate mention anywhere
in a file exempted every taskkill in it, which left the six files that now ask
the gate ratchet-blind -- the inverse of what it is for. It now counts `/pid`
call sites against gate admissions per file, so a second ungated kill inside an
existing family fails. Keying on the `/pid` argument rather than a quoted
`taskkill` also catches a kill whose program name comes from a constant. The
three comments that claimed more than the old scan enforced now state the rule
and its two remaining blind spots.
Also: the recording in `admitSelfInitiatedTreeKill` is now wrapped the way the
`admitProcessTreeKill` seam already wraps it, with the refusal decision taken
before anything that can throw so a diagnostics failure cannot flip it; and
`orca-chromium-process-pids` documents the false-positive direction (a stale
`getAppMetrics()` entry plus pid reuse refuses a live unrelated child), which is
the mechanism the root-kill fallback exists to bound.
Tests: refusal now asserts the root kill at all seven sites; the ratchet asserts
call-site counting and the constant-program form.
* test(crash-reporting): run the own-Chromium gate against real Windows trees
Nothing on this branch had ever executed on Windows. The unit tests pin the
gate's decision against a mocked taskkill, which cannot show that the decision
does anything to a real process: that `/T /F` reaps a detached grandchild, that
a refusal leaves that tree standing, or that the handle-addressed root kill the
refusal path falls back to reaps the root while orphaning descendants.
Adds a win32-gated live test covering all four, registered in both the
`package_windows` CI lane and `WINDOWS_PACKAGE_TESTS` as
`win32-test-lane-registration` requires.
Also completes the coverage doc's "never instrumented" list, which omitted the
macOS keyboard-input-source probe's POSIX group kill in `ipc/app.ts`.
* fix(crash-reporting): pin the commit-message root kill on the Windows arm
The first Windows run of this branch found nine failures the macOS suite
cannot see: `commit-message-text-generation-test-harness` asserts
`expect(child.kill).not.toHaveBeenCalled()` on `process.platform === 'win32'`,
which is the contract the previous commit deliberately replaced — and it
branches on the real platform, so it is dead code everywhere CI runs today.
The harness now asserts the handle-addressed root kill on every platform. On
win32 it lands after the tree walk, so the expectation waits rather than reading
one tick early, and its ten call sites await it. Red against the pre-fix arm at
all seven sites; the production code is unchanged.
* test(crash-reporting): remove the Windows lane marker tree through the retrying helper
The new win32 spec teardown used a raw rmSync, which the windows-lane-tree-removal
boundary ratchet rejects — and which is exactly the EPERM the ratchet exists to
prevent, since this spec's marker directory is written by processes it has just
force-killed.
* fix(crash-reporting): only refuse pid-addressed tree walks, disclose the handle-less codex site
The own-Chromium gate refused the POSIX process-group arm of
signalProcessTree as well, which was new macOS/Linux behaviour: a stale
getAppMetrics() entry plus pid reuse would orphan a group that main reaps
today. A POSIX group only holds what Orca put in it, so the refusal is now
scoped to win-taskkill-tree and the POSIX arm is recorded and admitted like
the other group kills in main. That also drops the synchronous
getAppMetrics() read from every POSIX termination.
codex-turn-added-roots kills roots found by a table walk, so a refusal has
no handle to fall back to. Pin that the refusal is visible - crumb written,
turn reported as not cancelled - rather than fixing what cannot be fixed.
* test(crash-reporting): detach the Windows survival fixture and observe real spawns
|
||
|
|
d3501f7ad6 |
fix(worktrees): stop a failed worktree scan from being recorded as an authoritative empty listing (#18456)
* fix(worktrees): stop a failed worktree scan from pruning as an authoritative empty listing A `git worktree list` that could not run at all — a WSL distro that stopped resolving, a hung mount, a git binary that errored — was softened to `[]` by the lenient listing path, so the detected scan published it as `authoritative: true, worktrees: []`. That disables the #1158 retention guard, drops every persisted tab for the repo, and the pruned session is written back to disk on the next launch. The loss is permanent, not a transient glitch. Route the detected scan through a strict listing that still reports the two genuinely empty states (repo path gone, not a Git repo) as `[]`. Everything else rejects, so the existing catch answers `authoritative: false` and the destructive halves (`rememberLocalWorktreeRoots`, `pruneLineageForMissingRepoWorktrees`) never see a failed scan. Also make the failure readable: wsl.exe reports its own launch failures as exit 0xFFFFFFFF with an EMPTY stderr and the `Wsl/Service/WSL_E_*` line on stdout as UTF-16LE, which is why the field bundle carried a git error with no text. Set WSL_UTF8 for WSL-routed git (matching the wsl runner, #9010) and attach that stdout diagnostic to the error so `git.exec` spans name the cause. * fix(worktrees): surface a failed worktree scan's cause on the repo header with a retry A failed scan now travels with its reason (optional unavailableReason on DetectedWorktreeListResult), the repo header shows it with click-to-retry, and the WSL deleted-guest-directory shape measured on a real Windows host is pinned as retained-not-pruned. |
||
|
|
c5ed53a1f5 |
perf(git): overlap the three independent submodule status reads (#18623)
getSubmoduleStatus awaited the inner status, the parent gitlink oid and the submodule HEAD one after another, though none depends on another's result. The Source Control panel re-runs this on every parent status poll while a submodule row is expanded, and the same function backs the git.submoduleStatus RPC, so over SSH or WSL each read was a separate round trip. Total latency drops from the sum of the round trips to the longest one. |
||
|
|
949c9d3353 |
perf(worktrees): classify each worktree once, defer the SSH meta index, drop the conflict-path probe (#18433)
* perf(worktrees): classify each worktree once, defer the SSH meta index, unserialise conflict probes Three redundancies on the worktree-catalog and git-status read paths: - buildDetectedGitWorktrees ran mergeWorktree + toDetectedWorktree twice for every visible row. Discovery backfill returns the same meta object when it wrote nothing, and both builders are pure over it, so skip the second pass on identity. - The SSH worktree-meta index parsed every worktree id on the host, then threw it away whenever the provider was connected. Build it lazily, memoised. - Unmerged `u` records were resolved one fs.access at a time. Resolve the prefix the cap can reach with 8-way concurrency, keyed by record index so Git's output order and error precedence are unchanged. * perf(git): read the porcelain worktree mode instead of probing conflicted paths Every porcelain-v2 `u` record already carries `mW`, the working-tree mode Git stat'ed for that row: `000000` means the conflicted path is absent. Reading it replaces the per-conflict `fs.access`, so the bounded-concurrency resolver, its `= 8` cap, and the order/error-precedence invariant are unnecessary rather than cheaper. `access()` stays only as a fallback for a malformed `mW`, so `parseUnmergedEntry` keeps its signature and neither status-read.ts nor the relay loop changes. Also corrects two fixtures that encoded `mW=100644` for a file that does not exist, which real Git never emits. * fix(test): import the conflict parser statically so the CJS cli project compiles |
||
|
|
5a626dcdf4 |
refactor(git): share push-target resolution between local and the SSH relay (#18406)
`src/relay/git-handler-push-target.ts` and `src/main/git/remote.ts` carried
identical ~160-line copies of the resolver that decides which remote a plain
`git push` hits. Identical today is exactly when to share it: the cost of a
future divergence is pushing to the wrong remote, which retrying does not undo.
Move the resolver to src/shared/git-push-target-resolution.ts, parameterized on
a `(args) => Promise<{ stdout }>` runner — the only thing the two hosts actually
differ in — and delete both copies. The relay entry point keeps only the work
that is genuinely relay-side: re-validating an explicit target that arrived over
the wire and running `check-ref-format` on it.
No behavior change on either path, and nothing new or different is published, so
this engages no rule in remote-wire-compatibility. No git command changes.
src/relay/git-push-target-local-parity.test.ts scripts one repository's config
and requires `git.push` over the real relay dispatcher and the desktop's
`gitPush` to emit the same push argv, plus the argv each case should produce.
|
||
|
|
53adf5e2e6 |
fix(git): share one failed-command error-text reader between local and the SSH relay (#18398)
* fix(git): share one error-text reader between the local and relay branch-delete fallbacks The relay and the desktop each carried their own `getErrorText`, and they had drifted: the relay read `message` + `stderr` + `stdout`, the desktop only `message` + `stderr`. A `git branch -d` refusal arriving on `stdout` therefore routed the SSH removal through prune-and-retry while the local removal gave up and preserved the branch. Against a real binary the two agree, because Git prints the refusal through `error()` on every supported version — verified on 2.25.1, 2.38.1, 2.49.1 and 2.55.0, none of which put a byte of it on stdout. What the desktop copy actually missed is that Orca classifies errors it built itself, with the Git output on `.stdout`: `worktree remove`'s submodule retry attaches `git status --porcelain` that way on both paths. The stdout-reading form is also already the shared spelling — `isSubmoduleWorktreeRemovalRefusal` uses it for both hosts — so this converges on it rather than on the shorter one. Move the reader to src/shared/git-command-failure-text.ts and the predicate it feeds to src/shared/git-branch-delete-refusal.ts, and delete all three copies. The predicate carries both refusal wordings live in the supported range: Git through 2.40 says "checked out at", 2.43+ says "used by worktree at". The real-binary contract now pins that boundary: the refusal is recognized, it lands on stderr, and stdout stays empty on every Git in the matrix. * fix(test): consolidate the duplicate worktree import in the parity test |
||
|
|
860ee73a11 |
fix(git): parse sparse and cquoted paths on the SSH relay (#18389)
The relay carried its own copies of the worktree-list and unmerged-entry porcelain parsers, and both had drifted from the desktop originals: the relay copy had no `sparse` branch, so SSH sparse checkouts were never marked, and it never C-quote-decoded a conflict path, so a conflicted file with a space or non-ASCII byte was published under its raw quoted name and probed as missing. Move both parsers into src/shared and delete the relay copies, so there is one implementation each. Type the relay's worktree-list plumbing on GitWorktreeInfo instead of Record<string, unknown> so a field-copying step can no longer silently drop a newly parsed field. `isSparse` is a new optional field on the git.listWorktrees result (remote-wire-compatibility Rule 1); Git <2.28 omits the porcelain line and the field stays absent. No new git subcommand or option. Closes #18280 |
||
|
|
fb48a9771b |
fix(gh): reap the whole gh/glab process tree at the deadline on POSIX (#18258)
`gh` and `glab` on PATH are routinely shims — mise, asdf, volta, or a hand-written wrapper — so a timed-out invocation has a chain to stop, not one process. `execFileCapture`'s POSIX kill path signals only the direct child; the descendants are orphaned to init and keep running. #18234 is exactly that shape: `bash ~/.local/bin/gh` -> `mise x gh` -> `gh`, where the reporter found the tail reparented to `systemd --user` and still at 100% CPU nearly two hours later. The 15s deadline #18239 added bounds Orca's semaphore slot and its promise; it does not bound the CPU burn. Route both CLIs through `execFileCaptureToTermination`, the primitive git's barrier path already uses: POSIX children spawn `detached`, the deadline signals `-pgid` and escalates to SIGKILL, and the promise waits for verified termination. Windows behaviour is unchanged (`taskkill /t` either way). Switching primitives also swapped execFile's hard maxBuffer failure for `runProcess`'s silent clipping, which would have turned an oversized gh response into a shorter valid-looking one. `ProcessResult` now reports truncation and the capture rejects on it, restoring the old contract and closing the same latent gap on git's barrier path. |
||
|
|
02a417c04b |
perf(renderer): stop six timers from ticking behind a hidden window (#18134)
* perf(renderer): stop six timers from ticking behind a hidden window
IntensiveWakeUpThrottling is disabled in this app, so a renderer interval
really does fire at full rate with the window hidden. Six of them had
nothing to observe them:
- NativeChatWorkingStatus ran a 1s interval + setState per in-flight turn
purely to advance an elapsed-seconds counter. Deleted the effect and
derived elapsed during render from the shared, visibility-gated
useNow(1_000) clock, so N turns collapse onto one tick.
- The chromium-error fallback poll (250ms) kept probing a stuck-loading
guest to write a loadError nobody could see.
- The contextual-tour full-pass interval (500ms) woke twice a second to
queue a rAF a hidden window never paints.
- Three feature-wall animation timers (3600/2400/2400ms) kept committing
React renders for animations nobody was watching.
- The landing preflight poll (30s) kept forcing IPC refreshes.
All five gated timers reuse installWindowVisibilityInterval. Each either
resumes where it left off (animations) or re-derives from durable state on
the becoming-visible run, so hiding and re-showing is observationally
identical to never hiding.
* test(git): stop two empty commits in the divergence fixture from hashing alike
`counts drift in both directions` builds 100 empty commits, resets to the fork
point, then adds one more — expecting 100 ahead + 1 behind to clear the cap of
100. An empty commit's hash covers only parent, tree, message and a
one-second-granularity timestamp, and every commit in the fixture reuses
`commit ${index}` starting from 0. On a runner fast enough to finish the whole
build inside one wall-clock second (CI: 1059ms for the case, ~7ms per commit),
the post-reset `commit 0` hashed identically to the first `commit 0` of the
chain, so Git handed back that same object and left the branch 99/0 apart
instead of 100/1 — `within`, not `exceeded`.
Numbering the empty commits across calls makes the fixture build the 101
distinct commits it already claimed to. Reproduced deterministically by pinning
GIT_AUTHOR_DATE/GIT_COMMITTER_DATE, which forces the timestamp collision the
fast runner hits by chance: fails with the exact CI assertion before, passes
after.
|
||
|
|
1910ab9c9c |
fix(github): bound and coalesce the Orca star check so gh children cannot pile up (#18239)
`checkOrcaStarred`, `starOrca` and `getAuthenticatedViewer` were the only gh call sites that reached for the legacy `execFileAsync` instead of `ghExecFileAsync`, so they ran with no deadline, no process-tree kill and no coalescing. A `gh` that never exits therefore ran forever and never released its slot in the 4-wide GitHub semaphore in gh-utils. Route all three through `ghExecFileAsync`, coalesce concurrent star checks onto one child, and hoist the Landing star-state effect out of the conditionally rendered footer so a repo-catalog rewrite no longer remounts it and re-forks gh. Adds a ratchet test asserting no file outside the command runner names `gh` as a spawned program. Fixes #18234 |