Commit Graph
7647 Commits
Author SHA1 Message Date
Merge Sim e5b751e4af fix(claude): don't clear a managed selection when the outgoing oauth identity is unreadable
Item 18 of the enumeration, which I had deferred as "clears nothing durable".
That was wrong, and the review's reading of runtime-auth-sync is right.

`readManagedOauthAccount`'s null is the third disjunct of the test that decides
whether Orca restores the user's system default before dropping a managed
selection. `runtimeOauthAccountMatches(null)` is false, so a failed read did not
cause a wrong restore -- it caused NO restore, while the
`updateSettings({ activeClaudeManagedAccountId: null })` a few lines later ran
regardless. The runtime kept holding the managed account's credentials with
nothing in Orca pointing at them: not a wrong undo, a missing one.

It is also not a rare corner. The disjunct only decides when the account is
unchanged AND `hasMaterializedRuntimeAuth` is false, and the service constructor
seeds `lastSyncedAccountId` from persisted settings without materializing -- so
that is the state of the FIRST sync after every app start with a selected host
managed account, and the oauth identity is the only evidence available.

The read is now tri-state, and an indeterminate one defers the whole transition:
neither the restore nor the clear. Deferring rather than restoring-anyway is
deliberate -- `restoreSystemDefaultSnapshotForMissingManagedCredentials` consumes
the same oauth value for its own ownership checks, so running it on a failed read
would degrade exactly the checks that decide what is safe to overwrite. Malformed
JSON stays dispositive: that is a completed observation of the file.

The two teardown blocks were byte-identical once both used the hoisted read, so
they are now one method, and the selection-clearing helpers moved to their own
module to stay under the line cap.

Refs STA-5674.
2026-09-01 17:07:03 -07:00
Merge Sim 466a745cb9 fix(claude): make every dispositive verdict prove itself, not just the probe
Round 2 of review on #17993. The tri-state had reached the ownership probe and
one caller; five other places still turned a failed observation into an answer,
and one turned a user's request into a refusal.

Enumerated every path in the lane that can produce a dispositive verdict — path
checks, marker reads, credential reads, error classification, and the four
decisions that clear or delete durable state — and fixed them together rather
than one at a time.

- The credentials read is now a result, not `string | null`. Its null said "there
  are no credentials", which `doSyncForCurrentSelection` acts on by clearing the
  user's account; a locked file, an unsearchable directory, or a failed realpath
  produced the same null. Both sync branches now leave the selection alone on an
  indeterminate read and clear only on a completed absence or invalid content.
  The darwin keychain read is classified the same way: its helper resolves null
  only for a genuine not-found, and every other failure throws.
- The account path check requires exactly one segment between the managed root
  and `auth`. A shell `*` matches `/`, so the guest accepted
  `<root>/other/acct/auth` as account `acct`, and the host-side `endsWith` agreed.
  Orca only ever creates `<root>/<accountId>/auth`.
- `isDefinitiveAbsence` guards its errno read again. Reverting it last commit was
  wrong: the sibling STA-5616 fix is on another branch, so this branch shipped a
  fail-closed predicate that throws. Written byte-identical to `349510a18e` so the
  merge is a no-op rather than a conflict.
- The unproven-error predicate is total. `instanceof` runs a prototype lookup a
  Proxy can trap and throw from, so it moves inside the guard; a chain longer
  than the inspection depth counts as unproven, because those links were never
  looked at. A cycle still answers `false` — every reachable link was seen.
- Explicit removal surfaces a failed unlink instead of reporting success. The
  caller already rolls its settings change back and rethrows, so the account
  returns and the user can retry; telling them it is gone while the credentials
  are still on disk is the same silent retention this path was fixed to stop. A
  root spelling that could not be canonicalised is reported too: with one
  spelling to compare against, "not ours" was never established.

Three tests asserted a failure scenario without injecting the failure — the
removal test armed a probe that no longer runs, a darwin test used a one-shot
rejection an earlier call consumed, and keychain call counts accumulated across
tests. Each now proves the fault reached the code under test.

Refs STA-5674.
2026-09-01 16:25:59 -07:00
Merge Sim e5858f1026 test(claude): merge the duplicate node:fs import
The code-quality lint pass runs a rule the pre-commit hook does not, so a
duplicated `node:fs` import passed locally and failed static analysis in CI.
2026-09-01 14:24:48 -07:00
Merge Sim d08089071b fix(claude): guard the adoption errno read in-lane and leave the shared predicate to STA-5616
Reverts this branch's edit to `src/shared/definitive-filesystem-absence.ts`.
`349510a18e` on `brennanb2025/sta-5616-codex-wsl-ownership` already guards
`isDefinitiveAbsence`'s errno read, with the reasoning written out there; two
branches rewriting the same shared file only produces a conflict. This branch
inherits that fix when the two land, and the reasoning is not restated here.

Note the sibling keeps its `readErrnoCode` private, so the import this branch had
added would not have resolved even after that change landed.

What remains in-lane is the Claude-specific read: the legacy-marker adoption path
classifies its own write failure, and `EEXIST` was being read straight off a
caught value. A `catch` receives whatever was thrown, so a throwing `code`
accessor escaped the branch meant to classify it. Only `EEXIST` is proof of a
marker that is already there; anything unreadable is indeterminate.

Also spelling out what the previous commit understated: replacing the ownership
probe in explicit removal with a lexically derived directory is a security
hardening, not a refactor. The path it replaces deleted
`resolve(canonicalAuthPath, '..')` — the parent of the *canonicalised* path — so
a symlink at the auth directory redirected a recursive delete to the link
target's parent, outside the managed root entirely. Deriving
`<managedAccountsRoot>/<accountId>` from the account ID and requiring the
persisted path to spell `<root>/<accountId>/auth` bounds the delete to the
directory Orca created, and recursive removal unlinks a symlink rather than
following it.

Refs STA-5674.
2026-09-01 13:50:53 -07:00
Merge Sim c518a6e9c3 fix(claude): stop over-refusing on removal and under-refusing in the second probe
Follow-up to the #17993 review. The tri-state landed in one lane and one caller;
five places still turned a failed observation into a verdict, and one turned a
user's explicit request into a refusal.

- Explicit removal no longer probes at all. The user's request is the authority,
  so a gate that cannot complete must not leave credentials on disk with nothing
  in the UI still pointing at them. The directory is derived lexically from the
  account ID instead, which is stricter than the probe it replaces: the old path
  deleted `resolve(canonicalAuthPath, '..')` and therefore followed a symlink out
  of the managed root. Automatic failed-add cleanup keeps requiring a dispositive
  verdict, and the two authorities are now wired to two different methods rather
  than sharing one.
- runtime-auth had a second, separate WSL probe whose null cleared the user's
  active account. It now shares the verdict resolver, and sync leaves the
  selection alone on an indeterminate one. Once ownership is proven, the
  credential read reuses that path: a second probe could fail where the first
  succeeded and reach sync as "missing credentials", which clears the selection
  just the same.
- The guest script proved a directory was readable and searchable before
  believing anything it did not find. `test -f` reports an unsearchable
  directory exactly as it reports an empty one, so an EACCES was arriving as a
  dispositive `missing-marker`. It also rejects symlinked candidates and markers
  now, matching the host gate.
- `isDefinitiveAbsence` and the unproven-error predicate are total. Reading
  `.code` or `.cause` off a caught value can itself throw, which escaped the very
  path meant to fail closed. A shape that cannot be inspected counts as unproven,
  because deletion requires proof.
- The host-visible WSL branch validates marker contents against the account ID
  and rejects symlinks, instead of accepting any directory with a marker file.

The guest script's tagged-line behaviour is now covered by executing it under
bash against a real temp tree, including the mode-000 cases.
2026-09-01 13:29:28 -07:00
Merge Sim b0fe91a713 fix(claude): stop a fail-once ownership fault from deleting the credentials it just wrote
During managed-account add, `writeOauthAccount` re-runs the ownership gate after
`writeCredentials` has already written `.credentials.json`. A transient fault
there (a cold WSL distro, a probe timeout, an antivirus lock) threw out of
`persist` into `cleanupFailedAdd`, whose re-gate then succeeded — the fault was
transient — and `rmSync`'d the whole account directory, credentials included.

Three collapses made that unavoidable, and Claude had no typed errors at all, so
no caller could branch on any of them:

- `assertOwnedWsl` reported `timedOut` and every non-zero exit as definitive
  absence, and its catch-all relabelled everything (including that absence) as a
  trust verdict.
- `resolveOwnedClaudeManagedAuthPath` returned the same `null` for a stranger's
  directory and for a `realpathSync` that failed.
- `cleanupFailedAdd` deleted on any add failure.

Adds the `owned | untrusted | indeterminate` verdict and the two error classes
the Codex lane already has (mirrored, not extracted — merging the lanes is
STA-5616), classifies both probes against it, and stops cleanup from deleting on
an unproven observation: it leaks the directory and rethrows the original
failure instead. The WSL guest now states its observation on a tagged line
rather than encoding it in an exit status no caller could tell apart from
`wsl.exe` failing to start.

User-requested removal is unchanged and still deletes everything, keychain
credentials included.

Fixes STA-5674.
2026-09-01 11:54:50 -07:00
Jinwoo Hong 02a7742406 fix(artifacts): raise desktop sharing limit to 5 MiB (#17708)
* fix(artifacts): raise desktop sharing limit to 5 MiB

* fix(artifacts): enforce recovery content limit

* fix(artifacts): bound recovery request envelopes

* fix(artifacts): clarify oversized request error
2026-08-31 16:51:11 -04:00
Neil aa658d28e3 test(updater): stop a slow module import from failing the next test (#17726)
* test(updater): stop a slow module import from failing the next test

`updater.ts` is 2.4k lines. Its first transform in a worker costs ~1.4s idle
but 45s+ when the machine is oversubscribed, which is past the 30s
`testTimeout`. Vitest cannot cancel the timed-out test body, so the abandoned
continuation went on to call `setupAutoUpdater` during the *next* test — with
the harness already reset — and failed it with:

    AssertionError: expected "vi.fn()" to be called 1 times, but got 2 times

That is the exact signature of the abandoned-instance timer flake fixed in
#17649/#17663, so a machine-load timeout reads as that regression returning and
sends the reader hunting in the wrong place.

Two changes, in `updater-test-module-loader.ts`:

- `loadUpdaterModule()` replaces every `await import('./updater')` in the suite.
  It records the test that asked for the module and throws if the import
  resolves after that test ended, stranding the continuation so the timeout
  stays the only reported failure. This removes the trap.
- `warmUpdaterModule()` imports the module once in `beforeAll`. The transform is
  cached across `vi.resetModules()` — only a file's first import pays it — so
  warming moves that one slow import onto the 60s `hookTimeout` and leaves every
  in-test import at re-evaluation cost (~25ms idle).

Measured on a 16-core mac, first vs later import in one file: 1439ms / 25ms
idle, 8339ms / 149ms under 40 CPU hogs, 45521ms / 15182ms under 400.

Under 400 hogs the suite went from 15 files and 22 tests failing (15 timeouts
plus 7 misleading assertion failures) to 23/23 files and 269/269 passing. Under
900 hogs it degrades into 14 plain `Hook timed out in 60000ms` failures and zero
assertion failures.

* fix: tighten the fence, surface its warning, stop patching timers on warm-up

Review findings on the loader:

Drop trackRealTimers() from warmUpdaterModule(). It was inert — updater.ts
arms no timers at module scope — and actively harmful for the 5 files that
build their own mocks and never call clearTrackedRealTimers(). Those files
previously had pristine timer globals; the warm-up installed a wrapper that
was never restored and whose armed-handle set grew unbounded.

Key the fence on TestRunner.getCurrentTest() instead of currentTestName.
Nothing ever clears currentTestName, so the fence only fired once the *next*
test had started; a continuation resolving during the timed-out test's own
teardown, or after the file's last test, was still handed the module. The
last-test case mattered: the harness afterAll has already cleared timer
tracking by then.

Emit the diagnostic through process.emitWarning. The throw lands on a promise
vitest already settled, so the message explaining why the continuation was
stranded was discarded and reached nobody — which was the entire payoff.

Widen the loader test's race margin 50ms -> 500ms. It gated on the
test-to-test transition completing in 50ms, so the regression test for a
contention bug could itself fail under contention.
2026-08-31 13:22:27 -07:00
Neil 08d95b9979 test(native-chat): remove initial-snapshot recovery race in watch-error test (#17722)
Root cause: the test cleared the injected tail-reader failure *before*
writing the recovered transcript line. The capped rotation retry loop is
still firing at that point, so a retry drain could succeed against the
still-empty file, consume the pending initial drain, and emit an empty
initial snapshot (`[], false, 0, undefined, undefined`). The later manual
watch callback then took the append path, and `u-recovered` never reached
onInitialSnapshot -- producing the CI failure
`expected [ false, +0, ...(5) ] to deeply equal ArrayContaining{...}`.

That empty-snapshot-then-append sequence is correct product behavior, so
this is a test bug: write the content first, then clear the failure, so no
drain can ever observe a readable-but-empty transcript. The assertion now
checks the exact recovered snapshot instead of a flattened
arrayContaining, so an empty recovery snapshot fails loudly.
2026-08-31 13:22:20 -07:00
Neil a381b47437 test(cursor): widen Windows hook spawn budget to fix ETIMEDOUT flake (#17721)
* test(cursor): widen Windows hook spawn budget to fix ETIMEDOUT flake

`package (windows)` failed once on an unrelated packaging PR with
`spawnSync cmd.exe ETIMEDOUT` at hook-service.test.ts:78. This is an
infrastructure-timing flake, not a logic race: the assertion is
`expect(result.error).toBeUndefined()` and `ETIMEDOUT` only means the
spawnSync `timeout` elapsed.

Cursor is the heaviest of the hook-service suites on Windows. Its managed
command is the PowerShell encoded launcher, so one hook run is
cmd.exe -> powershell.exe -> cursor-hook.cmd -> curl.exe: four process
creations, one of them a CLR start that installer-utils.ts itself
documents as ~300ms warm and "visibly slow". The sibling suites (codex,
grok, agent-hooks/installer-utils) spawn the .cmd directly and set no
per-spawn timeout at all, so 15s here was a one-off, not a convention.

Raise the per-spawn budget 15s -> 30s to match
WINDOWS_PROCESS_TEST_TIMEOUT_MS in src/shared/setup-agent-sequencing*.test.ts
and the 30-90s used by the real-subprocess tests in src/main/browser. 30s
is ~30x the warm cost of the chain, which leaves room for CPU contention
and Defender scanning of the freshly written .cmd on a packaging runner.

The default vitest testTimeout is also 30s, which would have become the
new binding constraint (the protocol case runs 16 chains back to back), so
give the four spawning cases 120s. That keeps ETIMEDOUT - which names the
stuck process - as the failure you see, instead of an opaque case timeout.

No product behavior changes and no end-to-end coverage of the Windows
launcher is removed.

* fix: halve the case timeout and correct the contention rationale

Review found the stated cause wrong. pr.yml runs "Test Windows-specific
boundaries" before "Build package inputs", so electron-builder is not
running. The real contender is that vitest invocation itself: ~25 files at
maxWorkers 4, including five real-Electron suites and two node-pty tests.

120s was over-provisioned. windows-hook-payload-delivery.test.ts drives the
identical PowerShell chain on the same job with a 60s case budget; 60s gives
the same property here (16 warm spawns plus one 30s outlier) and halves
time-to-signal on a genuinely stuck chain.

Also record that 30s deliberately exceeds the product's own
MANAGED_HOOK_TIMEOUT_SECONDS (10s) — this test gates launcher correctness,
not user latency, so the SLA is not the right bound. Left
windows-hook-payload-delivery.test.ts at 15s: its value is deliberate, set
to mirror Claude Code abandoning a hook at 10s.
2026-08-31 13:22:12 -07:00
Brennan BensonandMerge Sim 872bd51d47 fix(native-chat): reland large structured command results (#17720)
* fix(native-chat): preserve large structured command results (#17707)

* fix(native-chat): preserve large structured command results

* chore: place native chat validation artifacts under docs

* chore: drop stale root package config

* fix(native-chat): enforce rebuilt lifecycle append slots

---------

Co-authored-by: Merge Sim <sim@local>

* chore: omit native-chat reland planning docs

* fix(native-chat): remove journal store import cycle

* fix(native-chat): keep journal factory acyclic

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-31 13:12:19 -07:00
Neil dc5db4b01c fix(lint): merge duplicate process-table-snapshot imports (#17724)
main is red on `static analysis`: oxlint's code-quality pass runs with
--deny-warnings, and agent-foreground-process-batch.test.ts imports
'../../shared/process-table-snapshot' twice (lines 5 and 13), tripping
"Modules should not be imported multiple times in the same file".

Introduced by #17525. It blocks every open PR, none of which can go green
until this lands.
2026-08-31 13:04:40 -07:00
Brennan BensonandMerge Sim 9477b5fcbb feat(ssh): batch process evidence in PTY inventory (#17525)
* feat(ssh): batch process evidence in PTY inventory

* fix(ssh): accept Linux kernel process rows and make no-evidence polling push-driven

* fix(ssh): preserve process evidence polling semantics

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-31 12:44:41 -07:00
Brennan Benson 894ed75abb Revert "fix(native-chat): preserve large structured command results (#17707)" (#17719)
This reverts commit 5fe37729ea.
2026-08-31 12:34:49 -07:00
Brennan BensonandMerge Sim 5fe37729ea fix(native-chat): preserve large structured command results (#17707)
* fix(native-chat): preserve large structured command results

* chore: place native chat validation artifacts under docs

* chore: drop stale root package config

* fix(native-chat): enforce rebuilt lifecycle append slots

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-31 12:34:03 -07:00
Jinjing 68de1be517 fix: satisfy GitLab hook and test lint gates (#17694)
* fix: satisfy GitLab hook and test lint gates

* Rename electron-vite target config to .cts

The .cts extension keeps the config as CommonJS, allowing electron-vite
to load each parallel target without sharing its timestamp-named ESM
temp file.
2026-08-31 12:33:50 -07:00
Brennan BensonandMerge Sim aabcc57366 fix(runtime): publish remote control outages to host surfaces (#17531)
* fix(runtime): publish remote control diagnostics to renderer

* test(runtime): account for diagnostics bridge listener

* fix(i18n): add runtime connection state labels

* test(runtime): clean up shared control connection

* fix(runtime): fence diagnostics by shared-control capability

* fix(runtime): preserve authoritative transport state

* fix(runtime): preserve diagnostic overlay lifecycle

* fix(runtime): avoid publishing unchanged diagnostics state

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-31 12:25:17 -07:00
Jinwoo Hong db84894eef fix(runtime): isolate paired terminal creates from host focus (#17713) 2026-08-31 15:21:42 -04:00
Jinwoo Hong 63d60e0ef0 fix(mobile): unblock targeted SSH session tab refresh (#17486)
* Fix targeted mobile SSH session tab refresh

* Preserve fail-open explicit workspace resolution

* Strengthen SSH session refresh oracle
2026-08-31 15:18:21 -04:00
Jinwoo Hong bbbb59e18c test: cover quick commands, catalog links, and long discard dialogs (#17489) 2026-08-31 14:56:15 -04:00
Jinwoo Hong d1350735ef fix(orchestration): explain invalid send message types (#17487) 2026-08-31 14:34:55 -04:00
Jinwoo Hong faaf38ac45 fix(orchestration): submit staged mail pointer while working (#17470) 2026-08-31 14:25:34 -04:00
Brennan BensonandMerge Sim 5ff1aa540e fix(codex): re-land WSL direct-home cutover with counsel findings fixed (#16854)
* fix(codex): safely re-land WSL direct homes

* fix(codex): finish WSL direct-home cutover

* fix(codex): coalesce WSL launch hook installs

* perf(codex): avoid duplicate retired WSL session scan

* fix(codex): retain canonical WSL retired-home path

* fix(codex): fail closed before retiring WSL auth

* fix(codex): reopen WSL drain after rollback

* fix(codex): preserve WSL source on unknown panes

* fix(codex): harden repeated WSL runtime drains

* perf(codex): bound pending WSL session scans

* fix(codex): recover invalid WSL session watermarks

* fix(codex): validate retained WSL scan state

* fix(codex): accept durable WSL scan state

* test(codex): cover the drain's inode-identity guard against destination replacement

Removing the four `target_auth -ef temporary_destination_auth` assertions left
all 33 apply-script tests passing, so a regression deleting them would have
shipped silently. Reproduced before writing this.

A hash check cannot catch the case. The pinned hard link keeps the original
inode, so it still hashes correctly after another writer atomically renames a
different file over the destination path; only inode identity sees it. Without
the guard the script exits 0 and retires the source, leaving the user holding
bytes nothing validated. The new case asserts the source survives.

The harness is split by responsibility so no file exceeds its max-lines budget:
fixtures, the coreutils interference shims, the run types, the apply runner, and
the recovery/absent runners. The atomic-rename hook is deliberately separate
from the in-place rewrite shim because different guards catch them.

* fix(codex): keep the split drain harness inside the child-process boundaries

Extracting the harness into non-test modules moved it out of the exemptions the
single test file had: three new files import child_process, and two spawned
without windowsHide.

Adds the three to the import allowlist, and sets windowsHide on the spawns
rather than exempting them - the flag is correct for these calls regardless of
the ratchet, and they are skipped on win32 anyway.

---------

Co-authored-by: Merge Sim <sim@local>
2026-08-31 11:20:22 -07:00
Jinwoo Hong 3cb9b5d87f fix(serve): keep Chromium switches out of CLI redirect (#17633) 2026-08-31 14:11:33 -04:00
Jinjing 393992cbe7 Add show-more button for recent tabs in empty-query palette (#17693)
* feat: add show-more button for recent tabs in empty-query palette

* Reveal all recent tabs in one show-more click, limit badges to 9

The show-more button now expands the entire recent tabs list instead of
paging through it. Badges are limited to the first 9 rows since only
those digit positions are addressable in the keyboard chord.
2026-08-31 09:23:47 -07:00
Jinjing d5d3c4898a perf(diff): defer large diffs until user loads them (#17521)
* perf(diff): defer large diffs until user loads them

Rendering very large diffs would freeze the UI. Diffs exceeding
MAX_AUTOMATIC_DIFF_CHANGED_LINES now show a prompt allowing users
to load them on demand instead of automatically rendering.

* perf(diff): defer large diffs until user loads them

Diffs with >10,000 changed lines are now deferred and only rendered when
the user explicitly clicks "Load diff" in a prompt. This improves initial
render performance for large file changes while maintaining full access
when needed.

* perf(diff): defer large diffs until user loads them

Prevents UI freeze when opening files with very large diffs by
deferring render until the user explicitly loads them.

* fix(diff-view): defer loading large untracked files and refactor fallbac

Split on-demand load decision logic to distinguish tracked vs untracked files — large untracked files now properly defer loading while untracked images remain automatic. Extract fallback height computation into a dedicated function to centralize the logic for render-limited and in-flight-loading states, reducing code duplication and clarifying when to use bounded fallback heights.

* fix(diff-view): defer loading large SVG files

SVG renders as source text in the diff view rather than a preview, so should defer like other text files. Also fix Windows e2e test cleanup by using post-Electron shutdown.
2026-08-31 09:20:12 -07:00
NeilandBrennan Benson fbe94ceff6 fix: close readiness gaps found by merged-change audit (#17159)
* fix(ssh): fence stale kills and retired pane replay

* fix(ssh): support cancellable interactive authentication

* fix(ssh): await remote catalog before snapshot adoption

* fix(pty): contain Windows ConPTY input failures

* fix(power): avoid redundant macOS display blocking

* perf(editor): narrow markdown override subscriptions

* fix(quick-open): close directory handles after reads

* refactor(linux): remove unused proc socket scanner

* fix(usage): apply flat Sonnet 4.6 pricing

* ci: prime Node next native test cache

* docs(skills): resolve snapshot cleanup data path

* fix(ssh): recover install locks after host reboot

* test(ssh): recognize boot-aware install locks

* test(ssh): prove previous-boot lock recovery live

* test(wire): pin pre-metadata release coverage

* fix(terminal): preserve remote tab ownership through recovery races

* test(runtime): fence replaced terminal handles in agent guard

* fix(ssh): preserve remote snapshot authority across polls

* fix(pty): contain late ConPTY output EPIPE

* test(pty): register Windows exit watcher before kill

* fix: close SSH and tab readiness race gaps

* fix(tabs): retain headless order and placeholder titles

* fix(build): avoid parallel electron-vite config race

* test(windows): avoid MSYS temp path rewriting

* test(windows): avoid killing exited PTY

* fix(pty): avoid late ConPTY input teardown race

* fix(terminal): sync reconnect error ownership after commit

* fix(runtime): use canonical worktree identity comparison

* test(ssh): assert complete cold-hydration baseline

* test(windows): invoke quoted retention fixture via PowerShell

* test(windows): read ConPTY grid through mode con

* fix(terminal): publish PTY replacements atomically

* fix(terminal): infer stale identity on reattach

* fix(terminal): fence stale pane PTY callbacks

* fix(terminal): fence stale pane binds after rebind

* fix(terminal): reject stale pane transport callbacks

* fix(terminal): fence mirrored reattach spawn callbacks

* fix(terminal): replace stale pane PTYs on remount

* fix(ci): size the Windows launcher-compile test budget from measurement

`native-smoke (windows-latest)` fails ~4.5% of runs on
`preserves a multiline argument through the compiled remote launcher`
with "Test timed out in 15000ms" — on unrelated PRs, for reasons that
have nothing to do with them. Across 176 sampled attempts it is the only
red that job produced, and it hit seven different PRs in two days:
#16900, #16904, #16915, #16955 (twice), #16979, #17014, #17085.

The test is six process creations: powershell.exe forks csc.exe, then
the freshly compiled orca.exe forks node.exe, twice. Hosted Windows
runners periodically slow process creation down, and this test amplifies
that far harder than anything else in the job. Comparing the 80 attempts
where it ran under 3s against the 12 where it ran over 12s, its own
median goes 2198ms -> 15917ms (7.2x) while the same file's
powershell-only test moves 556 -> 686ms (1.2x), the cmd.exe and Git Bash
process tests in the neighbouring file move 1.4x, and the other 35 files
put together move 1.5x.

Measured across those 176 attempts: 1881ms to 35438ms, p50 4264ms,
correlation +0.881 with the job's total Vitest duration. 8 of 176 (4.5%)
exceeded the 15s cap; 2 of 176 (1.1%) also exceeded the shared 30s
testTimeout, so deleting the override and inheriting the config is not
enough on its own. 60s clears all 176 with 1.7x headroom on the worst.

This is slow, not hung. Every body here is synchronous spawnSync, so
Vitest cannot interrupt one — the timer fires only after the body
returns and the reported duration is real elapsed time. That is why a
failure reads `× ... 22464ms` under `Test timed out in 15000ms`. The
work finished; the stopwatch was short. Seven reruns at one identical
head measured 2053 / 4680 / 5551 / 8732 / 13506 / 14868 / 21937ms — the
last of those would have been red on code that had not changed.

The 15s came from #8897, which raised this test off Vitest's built-in 5s
default because the job then ran bare `pnpm vitest run`. #8909 landed
3h27m later and pointed the job at config/vitest.config.ts, which is the
real fix for that. The constant stayed behind and has been the binding
budget ever since.

* fix(terminal): fence stale remount reattach ownership

* fix(terminal): reconcile mounted pane identity after replacement

* fix(terminal): fence stale reattach fallback ownership

* fix(terminal): fence deferred SSH reattach ownership

* fix(terminal): fence stale split pane ownership callbacks

* fix(terminal): keep stale spawns from consuming startup

---------

Co-authored-by: Brennan Benson <79079362+brennanb2025@users.noreply.github.com>
2026-08-31 08:17:40 -07:00
Neil 75e5c996c1 perf(relay): stop ACK boundary scans at first pending boundary (#17491)
* perf(relay): stop ACK boundary scans at first pending boundary

* test(relay): pin PTY source boundary cleanup and guard ascending sends

The early-`break` in advanceCredit is only correct while sentBoundaries is
inserted in ascending sentEndSu order. Turn that implicit invariant into a
throw at the sole live write site (commitPtySourceSend), and assert the
post-state directly instead of inferring it from an iteration budget:

- assert the surviving boundary set after the 1,023-ACK benchmark
- cover the jump-ahead cumulative ACK that must delete many boundaries in
  one pass (the case an over-eager `break` would get wrong)
- cover the settleReservedPtySourceAck -> advanceCredit entry point
- drop an arithmetically-implied assertion and CI benchmark log noise

* perf(relay): reclaim ACK boundaries with a monotone cursor

The early-break Set scan still rebuilt a Set iterator per ACK, so V8 walked
delete tombstones and the drain stayed superlinear; the visit-count test could
not see it because it stubbed sentBoundaries with a generator over a private
Set. Replace the Set with an ascending boundary list plus a monotone cursor,
assert the real structure, and add a benchmark over the shipped code.

* test(relay): enforce ascending sent-boundary inserts in the collection

Move the ascending-order precondition into PtySourceSentBoundaries.add so
both insert sites are covered, and assert per-ACK span reclamation in the drain.

* test(relay): collapse ledger test record accessors into getDeliveryRecord

Rebase onto #17490 left two structurally identical internals accessors
(getCursorRecord, getBoundaryRecord); one typed accessor covers both.
2026-08-31 03:10:44 -07:00
Neil ae2eeff55d perf(relay): index PTY source-credit send spans (#17490)
* perf(relay): index PTY source-credit send spans

* perf(relay): maintain PTY source-credit retention totals

* test(relay): pin PTY send-cursor rebase across ACK reclaim

Cover the Math.max clamp branch in reclaimCreditedSpans where reclaim
removes spans at or past the send cursor, and widen the seeded fuzz case
to 20 spans per seed so the cursor actually traverses spans; assert the
cursor never overshoots the span containing sentEndSu.

* refactor(relay): drop dead retained-total helpers and pin retention counters

The incremental PtySourceCreditRetention counters replaced the recompute-from-records
helpers; delete the now-unreferenced exports and recompute the totals from the live
records inside the ledger tests so the counters have an independent oracle.

* test(relay): bound send-span reads instead of pinning the read pattern

Address review feedback on the send-span cursor coverage:
- replace the exact indexed-read pin and the tautological naive-visit
  assertion with a linear bound that still fails on the old Array.find path
- drop the per-run bench console.log
- assert retention totals immediately after rotate(), the only path that
  removes and re-adds a record in one call

Also count the replacement delivery in retention as it enters the delivery
map so the "in deliveries <=> counted" invariant never has a hole.
2026-08-31 02:41:15 -07:00
Neil e17c98d425 fix(daemon): bound the whole boot-recovery sequence with one budget (STA-5732) (#17427)
* fix(daemon): bound the whole boot-recovery sequence with one budget (STA-5732)

* fix(daemon): keep socket probes inside recovery budget

* fix(daemon): size the recovery budget against the real post-kill tail

The 24s budget reserved only 9s for everything after the deadline, leaving
27s of the startup PTY gate's fail-open cap unused — and every unused second
is one where a daemon that would have drained gets killed with its live PTYs
instead. Reserve each post-deadline stage's actual hard cap (kill 10.5s, fork
10s, lease 5s) and spend the rest: 24s -> 32s of adopt window.

* fix(daemon): keep the last-resort endpoint rescue outside the recovery budget

The rescue probe in the launcher's outer catch was clamped to the recovery
budget's remainder, but it runs *after* that budget by construction — past
prepareDaemonReplacement, killStaleDaemon, the fork and the adoption lease.
The remainder is therefore essentially always negative, so Math.max(1, ...)
handed a live socket a 1ms connect window. On the loaded machine this path
exists for the probe loses to its own timer, the launcher rethrows, and a
recoverable degraded adoption becomes total daemon loss for the whole run —
the outcome the comment above it exists to prevent. Restore the 1s default
and pin the window with a test that drives the launcher to that catch with
the budget already spent.

Also make the deliberate narrowing legible instead of implicit:

- daemon-recovery-budget.ts: TRANSIENT_WEDGE_DRAIN_MS documented 20s as the
  grace #8697 sized, but #8697's merged second commit (840d3277d1) widened
  it to 11 retries ~= 60s. Record that 20s is the drain estimate and that the
  budget deliberately sits under #8697's shipped grace.
- daemon-init-wedged-daemon-grace.test.ts: pin the trade directly — a wedge
  draining after the budget is replaced and loses its live sessions.
- Rewrite 'preserves a daemon that stays wedged until the LAST allowed grace
  retry' onto the simulated clock. It never mocked Date.now, so its 12 probes
  elapsed ~0ms and asserted a retry grace the wall clock can no longer
  deliver; it now pins the last drain the budget still adopts.

* fix(daemon): name the socket probe default and correct the grace-retry rationale

Answers the review round on the budget accounting: the outer-catch endpoint
rescue is deliberately outside it, and the preflight clamp no longer duplicates
probeDaemonSocket's default as a bare literal.
2026-08-31 02:41:11 -07:00
Neil 7cb1db63db test(updater): cancel the real timers an abandoned updater instance leaks (#17663)
* test(updater): cancel the real timers an abandoned updater instance leaks

#17649 stamped `loadElectronAutoUpdater()` with a generation so an abandoned `updater`
module instance could no longer drive the shared `autoUpdater` spies. That fenced one spy
graph but left the leak channel itself open: `resetUpdaterMocks()` still cannot cancel the
real timers the previous instance armed, so the stale instance keeps running and keeps
reaching every shared spy the fence does not cover.

Exposed chains, all with exact call-count assertions on them:

- 1s `updateCheckSilentSettleTimer` -> `completeSilentUpdateCheck()` ->
  `scheduleAutomaticUpdateCheck()` on the next test's fake clock -> `runBackgroundUpdateCheck()`
  -> `pinDefaultReleaseFeed()` -> `fetchNewerReleaseTagsWithReadiness` -> `fetchNewerReleaseTagsMock`
  (updater.check-preflight.test.ts:59,309,528; updater.publishing-window-feed.test.ts:382,458)
- `scheduleUpdateNudgeCheck()` -> `fetchNudgeMock` / `shouldApplyNudgeMock`
  (updater.nudge-campaign.test.ts:168,175)
- the previous test's `webContents.send` mock, which still receives a stale 'not-available'
- `completeSilentUpdateCheck()`'s 1h retry, which several files straddle with 59min + 1min

Close the channel instead of ignoring its effects. The harness now wraps the real
`setTimeout`/`setInterval`/`clearTimeout`/`clearInterval` globals while a test file is using
it, and `resetUpdaterMocks()` cancels every real handle armed since the last reset. Fake
handles are already discarded by `vi.useRealTimers()`, so real handles were the only leak
channel left.

The patch installs only after `vi.useRealTimers()` (never over a fake clock, so it cannot
capture fake handles), restores only the globals still holding its wrappers, hands back
untouched Node `Timeout` objects so `unref()` keeps working, and is removed in `afterAll` so
no unrelated file in the same worker sees it. Vitest arms its own test timeouts through
`getSafeTimers()`, snapshotted at worker setup, so nothing here can capture or cancel them.

The #17649 generation fence stays in place — this is additive defense in depth.

* fix: drop fake clocks before handing the timer globals back

The afterAll uninstall silently no-opped in 4 of the 10 harness files. Its
identity guard (globalThis.setTimeout === wrapper) fails whenever a file's
last test leaves a fake clock installed, and no updater test calls
vi.useRealTimers() — the only restore is the next beforeEach, which never
runs after the last test. Affected: check-settlement, publishing-window-feed,
quit-and-install, and this PR's own leaked-timers test.

Nothing broke because vitest defaults isolate:true, so the stranded wrapper
died with the per-file process. Under --no-isolate it would have been a real
leak: the wrapper stays installed for every later file in the worker, the
armed-handle sets retain every Timeout forever, and a later updater file's
reset would cancel live timers belonging to unrelated suites.

Also scope the module docstring — node:timers/promises and util.promisify
bypass the globals entirely, so a future `await setTimeout(...)` in
updater.ts would reopen the leak with no failing test.
2026-08-31 02:17:00 -07:00
Neil 6bbed15a11 fix(worktree): gate agent activation on the live surface census, not renderer state (STA-5701) (#17428)
* fix(worktree): gate agent activation on the live surface census, not renderer state (STA-5701)

* fix(worktree): seed a pane when the surface census cannot prove ownership (STA-5701)

Failing closed must not also fail silent. When the census is unverifiable
the sweep adopts nothing and mints nothing, yet the gate still reported
'adopted' — and both callers suppress their own seeding on any outcome but
'empty', so the workspace ended with zero surfaces. The sweep now reports
whether any live PTY holds a surface and the gate hands the caller its seed
when none does. Also folds equivalent workspace-path spellings in the census
index and in exact-surface binding, so a host row spelled differently is
neither dropped (mint a duplicate) nor unbindable (no pane).

* fix(worktree): name the live PTYs the surface census declined (STA-5701)

The adoption sweep can leave a live PTY without a surface — an unreadable
census, two host surfaces claiming one PTY, or a host-named leaf the
persisted layout does not have. The gate already stops reporting 'adopted'
in that case so the caller seeds a shell, but the decline itself was mute.

- adoptLiveWorkspacePtySurfaces now returns { surfaced, declinedPtyIds }
  and the gate warns with the workspace and the PTY ids left unsurfaced.
- Pin the host-named-leaf decline, which had no test either way.
- Pin the superseded-inventory race in terminal.list: a concurrent refresh
  makes hostScope.hostIds empty, which is what makes the renderer's
  'unverifiable' verdict reachable on a plain local machine.
2026-08-31 01:40:54 -07:00
Neil b5746724d4 perf(relay): account pending PTY output incrementally (#17639) 2026-08-31 01:28:29 -07:00
Neil 22a9e30ba1 test(updater): detach stale updater module instances from the shared autoUpdater mock (#17649)
Root cause of the `updater.startup-scheduling` flake: `resetUpdaterMocks()` calls
`vi.resetModules()`, which abandons the previous test's `updater` module instance but
cannot cancel the real timers that instance already armed. The earlier real-timer tests
leave a 1s `updateCheckSilentSettleTimer` pending; it fires a second or so later, i.e.
during a *later* test that has since installed a fake clock. The abandoned instance then
runs `completeSilentUpdateCheck()` -> `scheduleAutomaticUpdateCheck(24h)`, arming that
timer on the running test's fake clock at its epoch. `reschedules the next automatic
check 24 hours after finding an available update` advances 1h + 23h, so the stale 24h
timer lands exactly at the end of the 23h window, and the stale instance calls the
shared `autoUpdaterMock.checkForUpdates` spy -> 2 calls instead of 1.

Whether the leaked real timer fires before or after the next test installs its fake
timers is real-clock dependent, which is why it reproduced ~1 in 12 runs and only when
the whole file runs (30/30 pass with `-t` filtering to the single test).

This is a test-isolation bug, not a product bug: production has exactly one updater
module instance and one clock, so no stale instance can exist.

Fix: the harness already detaches abandoned instances on the event side (it clears the
`app`/`autoUpdater` handler maps on reset); extend the same idea to the call side.
`loadElectronAutoUpdater()` now hands each module instance a generation-stamped view of
`autoUpdaterMock`, and `reset()` bumps the generation, so a stale instance's calls and
property writes are dropped instead of driving the spies the running test asserts on.

Verified: 40/40 clean runs of `pnpm test src/main/updater.startup-scheduling.test.ts`
(0 failures), plus all 22 `src/main/updater*` files (264 tests) green.
2026-08-31 00:55:01 -07:00
Neil c09810b641 perf(rpc): restore compiled Zod request schemas without override (#17374)
* perf(rpc): compile Zod request schemas lazily

* chore(deps): pin zod 4.5.4 and except it from the release-age gate

4.5.4 is the first release fixing isRecursiveSchema (upstream 84e416f, #6500),
which compile() calls on every schema — on 4.5.0 it fired .default() factories
at compile time. Verified: compile-time factory calls 0 on 4.5.4, 1 on 4.5.0.
2026-08-31 00:47:22 -07:00
Neil ca516a4306 test(git): pin FETCH_HEAD lock order in the admission lifetime test (#17641)
`serializes FETCH_HEAD callers before they enter admission` assumed that two
same-repo fetches join the FETCH_HEAD lock lane in call order. They do not.

`runWithGitFetchHeadLock` first `await`s `fetchLockPath`, which walks the
filesystem (`realpath`, `stat` per parent directory, `readFile` of `commondir`,
`realpath` again) before it calls `runWithGitOperationLock`, and the lane is
registered only after that walk resolves. For a non-existent `/repo` that is
five libuv threadpool round-trips per caller. Two callers issued back to back
run their chains concurrently, so lane order is threadpool completion order,
not call order.

When the `interactive` fetch won that race it entered the lane ahead of the
`background` fetch. On the first caller's release it reached admission
immediately and, being interactive, took the free network headroom slot instead
of queueing, while the background fetch stayed parked on the lock. `queued`
therefore settled at 0 and never reached the asserted 1. Measured inversion
rate for the bare lock-path walk was 54/500 on an idle machine; the test itself
failed 5/20 locally, always at the same assertion, matching the two CI failures
on unrelated PRs (#17530, #17630) at the same line.

Fix the premise rather than the symptom: stub only the key derivation, keeping
the real FIFO `runWithGitOperationLock` that the test actually exercises, so the
lane is registered synchronously with the call. Key derivation keeps its own
coverage in `src/shared/git-fetch-head-lock.test.ts`. This also stops the fetch
tests in this file from sharing one global `/.git/FETCH_HEAD` lane with each
other and from touching the real filesystem.

Verified deterministic: 40/40, then 30/30 clean runs, plus 25/25 with twelve CPU
hogs and a concurrent `src/main/git/command-runner/` run saturating the box.
2026-08-31 00:09:07 -07:00
Neil 6cc9319d87 perf(wsl): warn at project-add when the tree sits on a Windows drive (#17636)
* perf(wsl): warn at project-add when the tree sits on a Windows drive

Worktree placement now puts new workspaces inside the distro, but a project
whose own tree is on C:\ still pays the 9p/drvfs crossing on every git command
it runs — measured at ~20x for a clean `git status` against the same tree on
ext4. Nothing in the UI says so, so the project just feels slow.

Warn once, right after the add succeeds, naming the distro the project's git
actually runs in. The advisory is wrapped so it can never fail the add.

Two path shapes cross the boundary and both warn: a Windows drive path under a
WSL project runtime, and the UNC spelling of a distro's own drvfs mount
(\\wsl.localhost\Ubuntu\mnt\c\...), which crosses it however the runtime is set.
A tree already inside the distro, a drive path under Windows-host git, a plain
UNC share, and every POSIX/SSH path stay silent.

* chore(i18n): register the WSL filesystem boundary advisory keys in en.json
2026-08-30 23:48:50 -07:00
Jinwoo Hong 46fa1a98d0 fix(browser): show reload loading feedback (#17635) 2026-08-31 02:46:17 -04:00
Neil 017811f0de perf(renderer): memoize active worktree editor files (#17484)
Reuse the filtered editor-file projection while Terminal rerenders without changing its inputs.
2026-08-30 23:40:45 -07:00
2sumtech d641d87905 fix(browser): accept an empty --value in cookie set and --pass in set credentials (#17226) 2026-08-30 23:30:49 -07:00
Neil c55121231a fix(remote): re-activate the pending host surface on paired PTY attach (STA-5291) (#17429) 2026-08-30 23:21:18 -07:00
Neil 6a65d8406a perf(ssh): index source spans by ID (#17504) 2026-08-30 23:05:37 -07:00
Neil 337b729f82 perf(main): summarize workspace space rows in one pass (#17483) 2026-08-30 23:05:32 -07:00
Neil 900fce25cd perf(renderer): cache active terminal chrome projection (#17480) 2026-08-30 23:05:28 -07:00
Neil bc6ed60e83 perf(browser): group client-hosted row publication (#17479)
Reuse one registry snapshot when publishing rows for every workspace.
2026-08-30 23:05:24 -07:00
Neil 8c6cf4769e perf(main): index structured TUI process rows (#17478)
Build a PID index once per process snapshot so descendant ownership checks
avoid rescanning the full row list for every PID.
2026-08-30 23:05:20 -07:00
Neil 490a7de5fa perf(shared): index git history merge parents lazily (#17469) 2026-08-30 23:05:16 -07:00
Neil 435a5e2490 perf(renderer): index folder workspace host ownership (#17468) 2026-08-30 23:05:12 -07:00
Neil 4d19b3382c perf(shared): project automation list in one pass (#17466) 2026-08-30 23:05:08 -07:00
Neil 824dc89e97 perf(shared): classify folder workspace repos in one pass (#17464) 2026-08-30 23:05:03 -07:00