mirror of
https://github.com/stablyai/orca.git
synced 2026-09-28 16:02:45 +00:00
Two corpora, because they answer different questions. The 10.5 MB / 40-session corpus is what the scope split costs a reader: conversation is about 1.6x faster at p50 and 3.4x at p95 than the full corpus, which is the argument for the second FTS table being the one a keystroke can afford. The candidate limit needs more sessions than the limit before it costs anything, so it is swept over 2,500 one-turn transcripts with every session matching. Limits are interleaved sample by sample: run back to back, the first configuration pays for every page the OS cache had not seen and the ordering alone moved p95 further than the limit did. The doc says plainly what these numbers do not cover. They are cost, not relevance; the MRR figures quoted beside the BM25 weights and the identifier shadow column come from a shoot-out over real transcripts and cannot be reproduced from this repository.