Files
orca/config/build-plugins
OrcaWinandOrcaWin 4ec6bbf588 Kill hung WSL transcript filesystem operations via child process with route quarantine (#15381)
* fix(native-chat): kill hung WSL operations via child process

Stalled UNC file operations hold libuv permits even after the gate
timeout expires, blocking Chat tab recovery. Two stalled operations
fill both permits and freeze all WSL access until restart.

Fork file I/O for UNC paths into a separate child process. On deadline
expiry, kill the process to force the hung syscall to exit. This frees
the permit for the affected tab's next read. Temporarily quarantine the
stalled route to avoid retry storms.

* chore: drop internal review artifact from the repo root

* fix(native-chat): harden the WSL transcript fs sidecar

Review follow-ups on the sidecar isolation change:

- Only the deadline may abort running gate work. The sole waiter's
  same-duration timeout fired first, killed healthy children on caller
  abandonment, and settled the task before the deadline could quarantine
  a stalled route - leaving the back-off dead for every dedupe:false op.
- Resolve the fork entry from out/main/chunks too: the resolver compiles
  into a shared chunk, and the scanner service child has no
  process.resourcesPath, so packaged WSL vault scans threw entry-not-found
  (masked as an empty tree).
- Allowlist the fork env instead of spreading process.env; ambient
  NODE_OPTIONS would halt or --require code into every child.
- Wrap transport faults (spawn failure, child death) in
  WslTranscriptFsError('unavailable') so discovery reports them as scan
  issues instead of misreading them as missing paths or empty trees.
- Gate the vitest in-process fallback on the vitest worker global so a
  leaked VITEST=true cannot revert production to in-process UNC syscalls.
- Reap idle sidecar processes after 60s instead of holding them for the
  app session.
- Split 'open' into its own protocol union member so the reusable-call
  Exclude actually strips it from the pooled-process API.
- Guard kill('SIGKILL') against the teardown race where an exiting child
  emits an unlistened 'error', and dispatch reads by handle kind before
  path spelling.

* fix(native-chat): probe stalled WSL routes instead of a fixed quarantine

Remaining review follow-ups:

- Escalating route quarantine: first strike lifts after 5s so a distro
  that was cold-booting when its op hit the deadline recovers on the
  next poll (~35s total instead of ~90s); repeat stalls double the
  back-off toward the prior 2x-timeout cap, and any settle the deadline
  did not force clears the strikes. Queued same-route tasks fail fast
  at quarantine instead of stranding one waiter deadline per file in
  sequential scans.
- Single request implementation: the vitest in-process fallback now runs
  the child's own dispatcher (WslTranscriptFsProcessOperations + decode),
  so unit suites exercise exactly what the forked process executes and
  the per-call-site fallback closures are gone. Dirent fixtures gained
  the full kind-flag set the serializer reads.
- Dropped the production-dead per-route close queue; UNC FileHandles
  (test fallback only) mirror the process-handle close contract.
- Error class, messages, and factories move to wsl-transcript-fs-error
  (re-exported from the gate) to keep the gate under the lines budget.

* fix(native-chat): harden WSL transcript fs with route quarantine strike

Extract quarantine logic into a dedicated module with strike decay: stalls older
than 5 minutes restart from base back-off, and concurrent-lane timeouts count as
one incident. Allow joining live in-flight tasks on quarantined routes (they cost
no new I/O). Preserve quarantine across transport faults (child death). Handle
file shrinking during tail reads by detecting short reads and returning empty.
Defer file closes that arrive mid-read instead of refusing, preventing slot
leaks. Separate process slot and boundary-finding concerns into focused modules.

* fix(native-chat): enforce route quarantine windows and isolate lanes per

A late result arriving after the deadline was incorrectly lifting the route
quarantine, allowing subsequent work to start before the back-off period
expired. Now late results are correctly recognized as stale and never cut
the quarantine short.

Process work is now isolated per (route, priority) lane so a scan stall
cannot block exact reads on the same distro. Each lane gets its own client
and process pool; late results and handle faults stay scoped to their lane.

Tests now fake performance.now() alongside timers (the quarantine clock
depends on it) and wait for the full back-off window to expire rather than
advancing by 0. Gate state is reset between test cases since late releases
never lift the quarantine.

---------

Co-authored-by: OrcaWin <293788423+OrcaWin@users.noreply.github.com>
2026-08-18 16:57:35 -07:00
..