mirror of
https://github.com/stablyai/orca.git
synced 2026-10-08 00:02:38 +00:00
7faa9f7cd3bda68189646b1e9def6cc8aae3b1cb
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f0b5b8566c |
Bound OpenCode history reads and repeated worker failures (#25292)
* fix(opencode): keep scan budgets across queue waits and batches Reuse the scan-owned lifetime proposed in #10708 by @AmethystLiang with the existing shared worker queue. * test(opencode): check nonempty session fixtures and lint scoped controls * Derive OpenCode scan deadline message from its budget --------- Co-authored-by: Neil Parker <nwparker@MacBook-Pro-3.localdomain> Co-authored-by: OpenCode issue campaign <codex@localhost> |
||
|
|
f8a31edd6f |
refactor(cursor): move the desktop-login read onto the shared foreign SQLite reader worker (#24603)
* refactor(sqlite): rename the OpenCode SQLite worker entry to foreign-sqlite-reader (STA-9122) The worker thread that reads OpenCode's database off the main thread is about to read other apps' databases too, so its entry is renamed to what it is: src/main/foreign-sqlite-readers/foreign-sqlite-reader-entry.ts, built as out/main/foreign-sqlite-reader-entry.js. Why now: #24572 fixed the Cursor focus freeze with a second, dedicated worker. Rather than grow one worker per foreign app, the next commit moves Cursor onto this entry and deletes that worker. This commit is the rename only; #24572's cursor-desktop-profile-worker-entry lines stay until then. It moves out of ai-vault/ into a new foreign-sqlite-readers/ module because it will no longer be session-scanner code; the module will own the readers, their dispatch, protocol and main-process client. The entry still routes only OpenCode kinds in this commit. The OpenCode dispatch, protocol and process entry stay in ai-vault/ and stay OpenCode-only, because the SSH/WSL relay reader bundles them (build-relay.mjs). Every reference is updated: electron.vite.config.ts input key, knip entry, the plain-node entry guard and its test, the asarUnpack list (the scanner service still spawns this entry under ELECTRON_RUN_AS_NODE), and the electron-builder test that reads the filename. The filename and the beside-or-one-up (Rollup chunks) lookup now live in foreign-sqlite-reader-entry-path.ts, which the OpenCode spawn reuses, plus an Electron-main resolver that uses the packaged app.asar path. * fix(cursor): move the desktop-login read from its dedicated worker onto the foreign SQLite reader (STA-9122) #24572 fixed the Cursor focus freeze (#24360) with a dedicated worker (rate-limits/cursor-desktop-profile-worker*.ts). Orca already runs OpenCode's database reads on a worker, and more foreign-app SQLite reads are coming, so keeping one worker per app means one entry, build input, asarUnpack line, knip entry and guard line each. This keeps one pattern instead: Cursor's state.vscdb read runs on the shared foreign SQLite reader entry, and the dedicated worker, its entry and its config lines are deleted. What moves: - The read itself is a pure cursorProfile reader in foreign-sqlite-readers/readers/ (was rate-limits/cursor-desktop-state-db.ts), run only on the worker. A separate dispatch owns the new kinds and refuses an unknown kind. The entry routes OpenCode kinds to the untouched OpenCode dispatch, so the relay's OpenCode reader stays byte-identical. - ForeignSqliteReaderClient gives each reader its own WorkerThreadRequestQueue lane (own lazily started, idle-torn-down thread; one shared factory) with in-flight dedupe per database path. Any failure resolves to the reader's existing failure value and never falls back to the main thread. Kept from #24572, so every reader gets them: - Await worker retirement before respawning. Worker.terminate() cannot interrupt a native SQLite call (e.g. a WAL-index rebuild), so the old thread lives on until that call returns; respawning at once stacked a new thread on the same work for every timed-out read (#24572 measured three live workers). This belongs in the shared host, which fire-and-forgot terminate(): LazyWorkerThreadHost now takes awaitRetirement and refuses to spawn until the terminated worker settles, and the queue fails calls closed meanwhile. Opt-in, because pure-JS clients (session scanner abort, port scan) respawn right after an abort. A rejected terminate() also ends retirement, so it cannot latch the reader off (raised in #24572's review). - 10 s Cursor timeout, 2 consecutive deaths, a queue cap of 8. - #24572's worker tests, rewritten against the shared client: responsive caller plus coalesced probes, unavailable worker without path leaks, stalled-worker recovery, no respawn before retirement, dispose settles. Tests: reader, dispatch, client (timeout, 10 s default, crash, malformed, unavailable without a main-thread read, dedupe, queue cap, own thread per reader), queue retirement (stalled and rejected terminate), import boundary, and an event-loop test reading a ~50 MB WAL with no -shm on a real worker. * test(sqlite): walk the reader import boundary with the shared source-tree scan (STA-9122) * fix(sqlite): key reader dedupe on a caller-supplied key, not the path alone (STA-9122) |
||
|
|
da6d483ab9 |
fix(vault): read OpenCode SQLite inside WSL and SSH hosts (#23128)
* fix(vault): read OpenCode SQLite on WSL and SSH execution hosts * fix(vault): bound host setup and preserve cancellation across readers * fix(vault): keep WSL discovery visible and isolate probe tests * fix: retry local Vault runtime downloads without reserving remote stages Preserve verified remote cache reuse and latch only unresolved host work. Update WSL source-guard and remote dedup test integration. * fix(queue): discard aborted requests before respawn * fix(vault): preserve host paths and recover setup after reconnect * fix(ssh): fence every runtime platform probe across reconnects * fix(vault): reject setup results from superseded SSH connections --------- Co-authored-by: m4air <m4air@m4airs-Air.localdomain> |
||
|
|
631b51f508 |
perf(usage): run the Claude/Codex/OpenCode usage scans on a worker thread (#21114)
* perf(codex-usage): resume rollout scans at the last parsed byte Codex rollout files are append-only and grow all day, but any append changed both mtime and size, so `canReuse` discarded the cached entry and the scanner re-read the whole file from byte 0 on the Electron main process. On one real corpus that was 6.59 GB re-read per cycle across 26.63 GB / 21,110 files. Each parsed file now persists a resume point: the offset just past the last newline-terminated line, the parse context at that offset (session id, cwd, model, running totals), a sha256 of the 4 KiB before it, and the file's dev:ino. A grown file resumes there and merges the appended rollup into the cached one; anything unproven falls back to a full reparse — truncation, an in-place rewrite, rotation, a counted tail with no trailing newline, a legacy copied-session suffix offset, or a file that must reclaim deferred fork claims. Resume never depends on mtime equality, so a coarse-mtime filesystem cannot hide an append. Fixture: a 75,737-byte rollout with a 758-byte append re-read 76,495 bytes before and 8,950 after (the append plus two bounded 4 KiB boundary windows). Also bounds the automation-attribution force predicate for both Codex and Claude: it keyed on `lastScanError`, so a persistently failing scan forced a fresh full rescan on every single lookup. It now keys on the most recent scan attempt, which is one forced scan per run regardless of outcome. * perf(usage): run the Claude/Codex/OpenCode usage scans on a worker thread The three first-party usage scans walk whole rollout and transcript corpora and read OpenCode's SQLite synchronously, all on the Electron main process. They rarely produce a long stall — the JSONL reader streams, so it yields to the loop between chunks — but they pin the main-process event loop at ~95% utilization for the scan's whole duration, which is what every IPC message, timer and window event then queues behind. Move that work to one lazily-spawned, unref'd worker thread shared by all three providers, following the OpenCode SQLite scanner precedent (#8864). Measured on a synthetic 4,000-rollout corpus (25.8 MB cache): a cold scan drops from 2,147 ms of main-thread time to 31 ms, and a steady-state incremental scan from 165 ms to 64 ms. The worker is stateless and the cache crosses the boundary both ways. That costs ~64 ms of structured clone at this corpus size, against 2,147 ms saved on the cold path, and it keeps the persisted cache the single source of truth — a worker-owned copy would need an invalidation protocol and a second resident copy of the same multi-MB array. Failure is closed, never a silent empty result: a worker that cannot spawn, times out, or crash-loops rejects, and the store records the scan error and keeps the previous projection. Two clients already carried the same FIFO/timeout/crash-cap machinery, so extract it once as WorkerThreadRequestQueue (with the packaged entry-path resolver as worker-thread-entry-path) and move all three onto it, rather than adding a third copy. Their existing tests pass unchanged. The oracle is event-loop utilization on the calling thread, not a stopwatch: usage-scan-worker-event-loop.test.ts runs the same scan both ways and asserts the worker leg leaves the caller idle while the main-thread leg does not, so CI load moves both legs together (#18788). * test(usage): compare the two scan arms instead of two fixed thresholds The event-loop oracle claimed to be self-calibrating — its header said "the ratio is self-calibrating, so CI load moves both legs together (#18788) instead of tipping a fixed millisecond threshold." It computed no ratio. Two separate `it()` blocks each asserted an absolute threshold against its own arm, run separately, so load moved them independently. The comment described a test nobody wrote, and the flake it promised was impossible is the one that landed: `activeRatio > 0.8` on the calling-thread arm measured 0.764 on an ubuntu runner. Fixing the comment is not enough, because the fraction is the wrong quantity. CPU contention drags the calling-thread arm's active/wall fraction *down* toward the worker's, since the loop parks waiting on a contended libuv pool. A 4-vCPU Linux container measured that arm at 0.175-0.756 across twenty runs, idle and loaded — never once above 0.8. Active *milliseconds* move the other way: contention stretches the caller's JS time far more than it stretches the worker arm's fixed post-and-deserialize cost, so the gap widens under load. Merge the two arms into one case over one corpus and assert the worker arm costs the caller under a fifth of the inline arm's active milliseconds. Same twenty Linux runs: 10.9x-83.6x, passing throughout. Keep the presence preconditions on both arms — an arm that silently scanned nothing satisfies the comparison trivially — and extend them to the calling-thread arm, which previously checked only file and session counts. * fix(ports): name the dropped command when the probe queue is full The shared-queue extraction turned `Port scan command queue is full; dropped ${command}.` into a constant string, because `describeFull` was given no way to see the request. Pile-up is per-probe, so the name is the only thing in that log that identifies which of lsof/ps/netstat was shed. Pass the rejected request to `describeFull` and restore the name. The request is built before the cap check so it exists to be named; the id it burns is a correlation token, so a gap costs nothing. The existing overflow test asserted only the error class, which is why the regression escaped a 29-test suite. It now dispatches the overflow under a different command than the accepted ones and asserts the message text, so a message that names the wrong request fails too. Also add a direct WorkerThreadRequestQueue test. Three subsystems share the queue and each client test only sees the parts its own protocol exercises, with `queueCap` reachable from port-scan alone. Covers one-at-a-time FIFO dispatch, the deadline starting at dispatch rather than enqueue, the consecutive-death cap, and both points where that count clears. And record the child-process hazard at the usage worker entry. `terminate()` reaps nothing the thread spawned, and OpenCode discovery reaches a fork today: `wslGated*` forks the WSL transcript sidecar for a `\\wsl$\...` path, which a Windows `OPENCODE_DB` or `XDG_DATA_HOME` can be. One scan through that entry with a UNC `OPENCODE_DB` forked a sidecar that outlived `terminate()`. * test(ai-vault): assert the OpenCode worker messages exactly, not by fragment Checked every message string in the two clients the shared-queue extraction rewrote against origin/main. Only the port-scan queue-full one regressed (fixed in the previous commit); the OpenCode SQLite client's four messages render identically, the remaining source diffs being renames — `error.message` to `lastError`, `call.timeoutMs` and `CALL_DEADLINE_MS` to `timeoutMs`. `session-scanner-worker-client.ts` was not touched by the extraction. But its suite could not have caught it either. `/timed out/`, `/exited with code/` and a bare `rejects.toThrow()` all still match a message that has lost its interpolated value, which is the same blind spot that let the port-scan regression through. Assert the rendered text instead: the timeout names its deadline, the exit names its code, and the crash-loop drain still carries the text of the fault that killed the run. * fix(usage): correct the worker entry's child-process note The previous note said `worker.terminate()` leaves a forked sidecar orphaned. It does not, and the reproduction that appeared to show it used a stub sidecar missing the `process.on('disconnect', () => process.exit(0))` the real entry has. With a faithful one: the sidecar lives exactly as long as the thread and is gone within 2s of `terminate()`, because tearing the thread down closes the IPC channel it owned. Two worker lifecycles forked two sidecars and leaked neither, and the pre-worker main-thread path reaps its sidecar the same way, on host exit. What is true and worth recording: a fork is reachable from this bundle at all, which is easy to miss; it survives only as long as the channel does; and the sidecar is now re-forked per worker lifecycle instead of pooled for the app's life. State those, and warn that a future child which does not exit on channel close would not get the same free cleanup. * fix(usage): kill a wedged scan worker on no progress, not on wall clock `USAGE_SCAN_TIMEOUT_MS` was a 10-minute deadline on the whole scan. A cold scan of a real history is legitimately minutes — 637 s measured on a 30 GB corpus with 300 worktrees before the per-cwd memo, ~51 s after — so a larger corpus or a slower disk crosses it. Crossing it killed the worker, recorded a scan error and left the cache unadvanced, so the next refresh started cold and died at the same point, forever. The deadline is now a no-progress window. The worker posts a file counter as it walks the corpus (`UsageScanWorkerProgress`, rate-limited to one message a second), and `WorkerThreadRequestQueue` re-arms the active call's timer on each one via the new optional `isProgress`. Clients that do not pass it keep the plain wall-clock deadline. `MAX_CONSECUTIVE_DEATHS` and idle teardown are unchanged. * refactor(usage): report scan progress as a file count, not one call per file Claude's scanner walks batches, so a per-file callback made it loop just to bump a counter. |