Files
orca/config/scripts/session-search-write-benchmark.ts
T
Jinwoo-H 0f93c6d06b fix(ai-vault-search): close the regressions the second review round found
Second adversarial review of the consolidation pass, six areas, and the
fixes for what it turned up.

Providers: a degraded OpenCode capture read no longer publishes the file
cursor. The read marks the capture scope incomplete, the flag rides back
across the worker hop, the writer takes the not-applied path so the store
marks the candidate stale and the next scan re-parses; no red badge for a
retryable read. The parse deadline no longer counts time queued behind
unrelated index writes (separate worker and consumer bounds). Codex metadata
refresh returns the same reference whenever nothing changed. The OpenCode
worker host is generalised to LazyWorkerThreadHost and the port scanner
adopts it, dropping its duplicate lifecycle.

Renderer: the poll latch is gone. Main emits aiVault:searchIndexingChanged
after every apply, and the store does one read, subscribes to the push, and
polls only while a phase genuinely advances; a failed read is an unconfirmed
reading that keeps polling with a 4/8/16/32 s backoff, cleared by success or
push. The web client, which has no push channel, keeps its visibility-gated
interval. Coverage is observed only from a current result and never
overwrites a newer run. idle is a resting phase. Unverified sources surface
in the panel status line.

Service and store: a store that is transiently null during clear no longer
answers "search is off"; coverage falls back to consent. A superseded
transient write failure rebuilds the progress batch so a stale failure count
does not linger. Consent is persisted before it is applied. The WAL sampling
comment names the real second connection. messages_batch is a partial index
(schema 11). The benchmarks drive the real store. Database removal uses the
shared Windows retry loop.

Query: operator-only pages go through the ranking owner, so forks collapse
the same way for `repo:app` and `needle repo:app`. Typo repair falls through
to the best visible candidate instead of abandoning the prefix. The grouped
CTE's measured cost is recorded next to it. The outbound projection validates
against strict enums while the received schema keeps its fallbacks, and the
result type is pinned by equality rather than mutual extension.

Remote and CLI: search operations share one relay lane again, so a status or
configure issued during a query waits unsent and survives a sidecar fault. A
query gets the scan budget rather than the title budget. A rejected backfill
releases the owner lease. The configure path no longer refuses on an
in-flight `applied`. The search boolean vocabulary has one owner in the spec.
The method record moves into the contract module and gains searchCoverage.
The CLI status formatter treats applied:false as a caveat, not as unavailable.
An aggregate is not partial merely because a host reported unverifiable
sources.

Every fix carries a test that fails when it is reverted.
2026-09-07 22:56:33 -04:00

70 lines
2.5 KiB
TypeScript

import assert from 'node:assert/strict'
import { mkdtemp, rm, stat } from 'node:fs/promises'
import { join } from 'node:path'
import { tmpdir } from 'node:os'
import { setImmediate as yieldToEventLoop } from 'node:timers/promises'
import { SessionSearchStore } from '../../src/main/ai-vault-search/session-search-store'
import { stagedWriteUpdate } from '../../src/main/ai-vault-search/session-search-staged-write-test-fixture'
// Bundle with esbuild --bundle --platform=node, then run on the host under test.
// Everything runs through SessionSearchStore, so the numbers include the query log,
// the post-write cleanup and the compaction that a real index pays for.
/**
* Staging yields with `setImmediate` between chunks, so a peer chain samples the gap
* each chunk leaves. Measuring from outside keeps the owner's own step hook untouched.
*/
async function sampleLoopStalls(running: () => boolean, stalls: number[]): Promise<void> {
let previous = performance.now()
while (running()) {
await yieldToEventLoop()
const now = performance.now()
stalls.push(now - previous)
previous = now
}
}
const root = await mkdtemp(join(tmpdir(), 'orca-search-write-bench-'))
try {
const path = join(root, 'index.sqlite')
const errors: unknown[] = []
const store = new SessionSearchStore(path, (error) => errors.push(error))
try {
for (const mode of ['replace', 'append', 'replace'] as const) {
const update = stagedWriteUpdate(
`benchmarkneedle ${'synthetic coding context src/example.ts '.repeat(5)}`,
60000,
mode
)
const stalls: number[] = []
let writing = true
const start = performance.now()
const sampler = sampleLoopStalls(() => writing, stalls)
await store.apply(update)
writing = false
await sampler
const wallMs = performance.now() - start
assert.deepEqual(errors, [])
assert.equal(store.search({ query: 'benchmarkneedle' }).hits.length, 1)
console.log(
JSON.stringify({
platform: process.platform,
node: process.version,
mode,
rows: 60000,
wallMs,
maxLoopStallMs: Math.max(...stalls),
samples: stalls.length,
walBytes: (await stat(`${path}-wal`)).size
})
)
// Drains the tombstones the write left and compacts, exactly as a live purge does.
await store.purgeOlderThan(null)
}
} finally {
store.close()
}
} finally {
await rm(root, { recursive: true, force: true })
}