* Changing the way aggregation access their value.
They now get values via a ValueSource abstraction.
The aggregation collector also gets the possibility to
register ValueSourceProvider describing value columns that are computed
on the fly.
Finally, segment aggregation that require a full column
now manipulates a Arc<dyn ColumnValue> directly.
* CR comments
* Clippy
* Fixing regression
---------
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
Sorted segment merges copied every document's positions for a term into a vector before sorting by mapped document ID. Common terms with many positions therefore consumed memory proportional to documents × positions, outside the indexing writer budget.
Stream shuffled postings through a min-heap with one cursor per input segment and one reusable positions buffer. The mapping preserves each segment's document order; deleted/filtered postings are skipped. The stacked path continues to stream directly.
RegexQuery can be built from a compiled Regex, but RegexPhraseQuery always compiled its patterns with Regex::new and its default DFA state limit. from_regexes takes the compiled regexes directly, so a caller can build them with its own limits, or reuse ones it already compiled.
* Expose the Levenshtein automaton of FuzzyTermQuery
Callers that need the terms a fuzzy query expands to (e.g. to highlight them) had to copy the private DfaWrapper and depend on the same levenshtein_automata version as tantivy, or their expansion could silently diverge from the query's. FuzzyTermQuery::automaton returns the automaton the query matches with, built from the same cached LevenshteinAutomatonBuilder.
* Export DfaWrapper outside of tests
The re-export was still behind #[cfg(test)], so FuzzyTermQuery::automaton returned a type other crates could not name. A doctest now uses it from outside the crate.
Snippet only kept a copy of the fragment text, so callers that need to extend the fragment or map it back to the source (e.g. to keep punctuation that the tokenizer leaves outside the last token) had to search for it, which picks the wrong occurrence when the same text appears earlier.
Following up on #3135, the regexes are compiled on first use and kept in the query, so a query reused across searchers, or cloned, determinizes each pattern once. RegexPhraseQuery::regexes exposes them for callers that inspect the patterns before searching, e.g. to count a phrase's per-segment expansions against max_expansions, instead of compiling them a second time.
* jitexpr
* Added possible necessary conditions to jitexpr.
A function makes it possible to infer a necessary query from an expression to match.
We can then accelerate queries involving a calculated field by not even evaluating the expression
on docs that trivially do not match.
* CR comment
* CR comment
* Fixing unit tests
---------
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
RegexPhraseWeight::phrase_scorer compiled every phrase term's regex again for each segment. Determinizing a wide pattern (e.g. an alternation of typo or morphology expansions) costs milliseconds to tens of milliseconds, so on a multi-segment index the query was dominated by recompiling the same automata. The regexes are now compiled in RegexPhraseQuery::regex_phrase_weight and shared through Arc.
An invalid pattern is now reported when the weight is built, instead of when the first segment is scored.
Avoid formatting keys and finalizing sub-aggregations for buckets
discarded by the final size limit. Preallocate the retained bucket vector
and reuse the cutoff helper with intermediate tuple keys.
Preserve final-key ordering and finalize the ordering metric when sorting
by a sub-aggregation. Cover mixed key types, empty sums, and discarded
document counts in tests.
Resolve retained string ordinals in sorted batches per field instead
of decoding dictionary blocks separately for each tuple component.
Reuse decoding for repeated ordinals while preserving tuple order
and existing numeric and missing-value handling.
Probe the accumulated map only for incoming keys rather than rehashing every accumulated key on each merge. This removes quadratic work when folding many disjoint multi-terms results. The same helper serves range and composite buckets; existing bucket values still merge left-to-right and the wire representation and pruning rules are unchanged.
Cover overlapping/disjoint tuple keys, empty inputs, recursive range subaggregations, postcard round trips, error/count bookkeeping, and merging after pruning.
Same benchmark and configuration as 0062d2f0d (Apple M4 Max, rustc 1.98.0):
Median milliseconds for 10 / 100 / 1000 inputs:
shared 8: 0.006365 / 0.0656 / 0.6449
disjoint 16: 0.0126 / 0.1486 / 1.6882
disjoint 160: 0.1336 / 1.5558 / 20.4019
At 1000 inputs this is 52.4x and 61.6x faster for the disjoint cases, with unchanged measured peak allocation. Shared-key control remains within approximately 2% of baseline.
Validation: 308 aggregation tests pass with default features and 308 with quickwit; changed-file rustfmt and git diff checks pass. Clippy --lib --bench agg_bench passes with the pre-existing clippy::drop_non_drop warning allowed (unchanged drop(add_document) in the benchmark).
The Str path and the non string path are very different.
This PR isolates them as different segment aggregation collector.
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
Intermediate writes are an implementation detail. From Tantivy side we only decide when a file is finished. This change is done to remove flush overhead for VecWriter. On finalization we just move the Vec to the directory.
Rename TerminatingWrite to FinishableWrite to reflect these semantics.
Treat a fieldless exists leaf as match-all and make the lenient set parser
stop when parsing no longer consumes input. Add public QueryParser proptests
covering both regressions.
Findings and original fixes by Oleksii Syniakov in osyniakov/tantivy#4.
Co-authored-by: Oleksii Syniakov <1282756+osyniakov@users.noreply.github.com>
Generate one random-offset linear tie-breaker sequence per segment and
preserve its values through normal columnar merging. Enable the linear
codec to keep initial segments compact.
Silence the 7 warnings reported by `cargo clippy --all --tests`:
- Replace manual `% n == 0` parity checks with `is_multiple_of`
(`manual_is_multiple_of`) in index_writer, doc_predicate_query and
seek_danger tests. `is_multiple_of` is stable for unsigned integers as
of 1.87, which matches the crate MSRV.
- Prefix the intentionally unused `other_segment_meta` binding with an
underscore. The binding is kept rather than dropped because the
inventory only tracks live `SegmentMeta`s, so dropping it would
invalidate the `inventory.all()` assertion in the same test.
No behavior change.
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
Use the default seek-based implementation until the optimized path can handle invalid child states safely. A child miss followed by another child hit previously allowed refill to consume an unaligned scorer.
Add shared seek_danger contract proptests, a regression for #3086, and a source inventory guard for future overrides.
Serialize vector-backed term entries as a map so Quickwit nodes running
different versions can exchange Postcard aggregation results. Add a legacy
fixture covering nested terms and histogram aggregations.