mirror of
https://github.com/stablyai/orca.git
synced 2026-10-04 08:02:09 +00:00
Agent Session History exceeded its 130-second deadline on large local Codex histories. Three costs combined: excluded worker transcripts were recognized on their first line but still drained to EOF (1,011 files / ~18.3 GiB on the reported corpus), large ignored records were fully decoded and JSON.parsed, and the persisted parse cache was discarded on every app update. - Stop resumable reads the moment `session_meta` marks a worker transcript. - Skip decode + `JSON.parse` for records the parser only feeds to the timeline. The skip set is the complement of what `consumeCodexRecordLine` reads, and applies only above the bounded prefix limit, so a long opening prompt (which is the session title) still takes the exact parser. - Prove cross-volume rollout aliases from a bounded `session_meta` read routed through the WSL transcript FS gate, carrying the scan's AbortSignal, fanned out across contested candidates with bounded concurrency. - Make parse-cache schema 2 the semantic compatibility boundary so an update no longer forces a multi-gigabyte cold scan, fenced by a build-time ratchet on the persisted session shape. - Report early-stopped transcripts as their own `aiVault.scan` attribute. Verified on macOS, Ubuntu over SSH, and Windows: read volume drops 336 -> 49.5 MiB identically on all three; the Windows failing-test set is byte-identical to main. Reported corpus: 130s timeout -> 58.6s, 244 sessions, 0 issues. Fixes #17888.
59 lines
1.9 KiB
TypeScript
59 lines
1.9 KiB
TypeScript
import { readTranscriptSlice } from '../native-chat/wsl-transcript-fs-access'
|
|
import type { WslTranscriptFsTaskPriority } from '../native-chat/wsl-transcript-fs-gate'
|
|
|
|
const ROLLOUT_READ_LIMIT = 64 * 1024
|
|
|
|
/**
|
|
* Read the Codex session id without streaming the full rollout transcript.
|
|
* Defaults to `exact` because the live-resume proof is interactive; the AI
|
|
* Vault scan passes `scan` so a bulk sweep cannot starve it.
|
|
*/
|
|
export async function readCodexRolloutSessionMetaId(
|
|
filePath: string,
|
|
signal?: AbortSignal,
|
|
priority: WslTranscriptFsTaskPriority = 'exact'
|
|
): Promise<string | null> {
|
|
let head: Buffer
|
|
try {
|
|
// Gated, not raw fs: AI Vault hands this every scan candidate, including
|
|
// `\\wsl.localhost\...` rollouts, so a stalled distro must fail on the
|
|
// transcript gate's deadline instead of blocking the scan (#15453). A
|
|
// listed rollout may also vanish before it is read — Codex prunes and
|
|
// rewrites these files — and one missing file must not abort the scan.
|
|
head = await readTranscriptSlice(filePath, 0, ROLLOUT_READ_LIMIT, priority, signal)
|
|
} catch (error) {
|
|
// A cancelled scan must surface as cancellation, not as one more rollout
|
|
// that could not prove its id.
|
|
if (signal?.aborted) {
|
|
throw error
|
|
}
|
|
return null
|
|
}
|
|
const firstLine = head.toString('utf8').split(/\r?\n/, 1)[0]?.trim()
|
|
if (!firstLine) {
|
|
return null
|
|
}
|
|
try {
|
|
const record = JSON.parse(firstLine) as {
|
|
type?: unknown
|
|
id?: unknown
|
|
session_id?: unknown
|
|
thread_id?: unknown
|
|
payload?: { id?: unknown; session_id?: unknown; thread_id?: unknown }
|
|
}
|
|
if (record.type !== 'session_meta') {
|
|
return null
|
|
}
|
|
const id =
|
|
record.payload?.id ??
|
|
record.payload?.session_id ??
|
|
record.payload?.thread_id ??
|
|
record.id ??
|
|
record.session_id ??
|
|
record.thread_id
|
|
return typeof id === 'string' && id.length > 0 ? id : null
|
|
} catch {
|
|
return null
|
|
}
|
|
}
|