mirror of
https://github.com/quickwit-oss/tantivy.git
synced 2026-10-06 20:02:45 +00:00
`Streamer::term_ord()` derived the ordinal by counting `delta_reader.advance()` calls from a seed taken only when the stream had a key lower bound. An automaton search has no key bounds, so the seed was 0 — but the reader it is handed *is* block-pruned by that automaton (`get_block_iterator_for_range_and_automaton`). Every block skipped ahead of the first match went uncounted, so `term_ord()` returned the term's position among the blocks actually scanned rather than its ordinal in the dictionary. The error is silent and grows with how deep the first match sits: on a 262k-term dictionary, a regex matching only the last key reported ordinal 999 instead of 262142. Callers that resolve those ordinals back to terms therefore act on a different term entirely — `build_allowed_term_ids_for_str` builds the terms aggregation's allowed-ordinal bitset this way, so an `include` regex could make the aggregation count unrelated terms while returning the expected bucket count. Broad patterns matching from the start of the dictionary hid it, since nothing is pruned ahead of the first match. Each slice handed to the reader now carries the ordinal of its first term, and the streamer resets to it on entering a slice instead of incrementing across the gap. Blocks merged into one slice stay contiguous, so counting within a slice is still correct. `Dictionary::sorted_ords_to_term_cb` already tracked `BlockAddr::first_ordinal` explicitly, which is why ordinal->term resolution was unaffected.