Commit Graph
2709 Commits
Author SHA1 Message Date
Paul Masurel 89e66ca5ca blop 2026-08-17 17:42:04 +02:00
Paul Masurel 64fe39a659 added support for calculated predicate query 2026-08-17 12:11:25 +02:00
Pascal Seitz 7373d54c2d fix clippy 2026-08-12 21:08:06 +08:00
Ming 0401b45781 feat: Extensible segment components via plugin trait (#2993)
## Motivation

Today every segment component — postings, fast fields, field norms, store — is hardcoded into several places. Adding a new per-segment data structure means forking Tantivy and editing each of those sites.

This PR introduces a `SegmentPlugin` trait that lets a custom component participate in the full segment lifecycle — write, serialize, merge, garbage collection, space usage — through the same interface the built-ins use, without touching Tantivy internals. The four built-in components are themselves reimplemented as plugins.

We (ParadeDB) plan on using this trait for 1) additional segment metadata for partitioning 2) custom vector index.

## Plugin Trait

Two traits. The first is the `SegmentPlugin` factory:

```rust
pub trait SegmentPlugin: Send + Sync + 'static {
    /// File extensions this component owns, e.g. ["idx", "pos", "term"] for postings.
    fn extensions(&self) -> &[&str];

    /// Create a writer for the indexing path.
    fn create_writer(&self, ctx: &PluginWriterContext) -> crate::Result<Box<dyn PluginWriter>>;

    /// Merge this component across several source segments into the target segment.
    fn merge(&self, ctx: PluginMergeContext) -> crate::Result<()>;

    /// Report on-disk space usage, keyed by component name. Has a default impl.
    fn space_usage(&self, reader: &SegmentReader)
        -> crate::Result<BTreeMap<String, ComponentSpaceUsage>>;
}
```

A `SegmentPlugin` owns one or more file extensions and knows how to (a) build a writer for the indexing path and (b) merge itself across segments.

The second trait is the segment writer:

```rust
pub trait PluginWriter: Send + Any {
    /// Called once per document, in doc-id order, for every plugin writer.
    fn add_document(&mut self, doc_id: DocId, doc: &TantivyDocument, schema: &Schema)
        -> crate::Result<()> { Ok(()) }

    /// Serialize accumulated data to segment files (honoring an optional doc-id remap).
    fn serialize(&mut self, segment: &Segment, doc_id_map: Option<&DocIdMapping>) -> crate::Result<()>;

    fn close(self: Box<Self>) -> crate::Result<()>;
    fn mem_usage(&self) -> usize;

    fn as_any(&self) -> &dyn Any;       // downcast support, Rust 1.86
    fn as_any_mut(&mut self) -> &mut dyn Any;
}
```

The write path no longer has any by-name wiring: `SegmentWriter` hands every document to every plugin writer's `add_document`, and `finalize()` calls `serialize` then close on each. 

## Key Design Decisions

1. The index — not the segment — owns the plugin set. The set of custom plugins is recorded once, at index creation, in `IndexMeta`. `#[serde(default)]` makes this backward compatible.

2. Plugins are registered, like tokenizers — and re-registration is enforced fail-closed. Plugins are not serialized; they're re-attached on every `Index::open` via `register_plugin`, exactly like custom tokenizers. To prevent consumers from accidentally forgetting to register a plugin, we validate the registered plugin set against the persisted set when the index is first used for a write/merge/GC operation.

3. Registration order is the write/merge order. Built-ins come first (field norms → postings → fast fields → store), then custom plugins.

4. The read side needs no plugin hook. Custom component data is read back through the existing public surface `SegmentReader::open_read`.

5. Backwards compatibility — behavior for existing indexes is unchanged.
2026-08-10 16:06:11 -07:00
PSeitz-dd 52b607d7cd Merge pull request #3018 from PSeitz/low_card_histogram_lanes
Up to 2x faster fused term-histogram aggregation
2026-08-10 09:39:48 +02:00
Pascal Seitz d212a9c532 Address fused term-histogram review feedback
Decouple the low-cardinality sub-aggregation threshold from the fused
count-lane threshold, rely on the full-column invariant in the linear
resolver, and explicitly test the single-bucket resolver.
2026-08-10 15:30:57 +08:00
Pascal Seitz 6797d8c577 const assert, unreachable 2026-08-10 15:03:46 +08:00
mmustafasenoglu cac4ba9404 chore: downgrade all info! logs to debug! (#3013)
* chore: downgrade info! logs to debug! in managed_directory.rs

* chore: downgrade info! logs to debug! in file_watcher.rs

* chore: downgrade info! logs to debug! in index_writer.rs

* chore: downgrade info! logs to debug! in prepared_commit.rs

* chore: downgrade info! logs to debug! in segment_updater.rs
2026-08-07 11:58:56 +02:00
Marc Bachmann 52a95b159b fix: do not drive a non-summing score combiner with Block-WAND (#3022)
`BooleanWeight::for_each_pruning` handed a union of term scorers straight
to `block_wand`, whatever the weight's score combiner. Block-WAND prunes on
the sum of the per-term block maxima and scores the survivors the same way,
so it silently replaces the combiner with a sum. A `DisjunctionMaxQuery`
collected through `TopDocs` therefore scored a document matching in two
fields as the sum of both instead of the better of the two, and ignored the
tie breaker.

The union only reaches that path when every clause reads term frequencies,
which is why the bug hid: a query asking for `Basic` postings falls back to
`BufferedUnionScorer`, which does honor the combiner.

`ScoreCombiner` now declares whether Block-WAND may drive it. Only
`SumCombiner` opts in; everything else falls back to the plain union, which
still prunes on the threshold but makes no assumption about how the scores
combine.

The intersection specialization needs no such fallback: `Intersection::score`
sums its children whatever the combiner (the combiner only shapes how should
clauses combine), so `block_wand_intersection`'s summing already matches the
unpruned scorer for every combiner. A second test pins that equivalence by
comparing `for_each_pruning` against a manual scorer walk on a must+must
weight built with `DisjunctionMaxCombiner`.
2026-08-07 11:57:22 +02:00
Xuanwo 3a55bc986d Replace rust-stemmers with frostem for stemming (#3028)
rust-stemmers is unmaintained and lags behind upstream Snowball.
frostem tracks snowball main and regenerates algorithms automatically.
2026-08-06 10:54:56 +02:00
Pascal Seitz 3bd8a6b4d1 Optimize fused term-histogram aggregation
Avoid histogram decoding for single-bucket ranges and use precomputed
boundaries for small ranges. Split hot counters into lanes to reduce
write dependencies on low-cardinality grids.
2026-08-06 08:41:07 +02:00
trinity-1686a 0743fe7d67 reintroduce get_fast_field_tokenizer_name 2026-08-04 18:24:15 +02:00
Pascal Seitz 83ffe96189 Reuse existing term ordinals for multi-terms missing values
Map a configured string missing value to its dictionary ordinal when the
term is already present. This coalesces real and missing documents before
the segment-level cutoff, preserving counts and sub-aggregation results.

Keep synthetic sentinels for absent terms and non-string columns, and add
regression coverage with a segment size of one.
2026-08-04 13:23:19 +02:00
Pascal Seitz 92fd5c5cde Refactor multi-term key collection paths
Extract full and non-full field key construction into dedicated helpers.
Rename the surviving document buffer and avoid copying IDs on the full-field path.
2026-08-04 13:23:19 +02:00
Pascal Seitz c9bf057162 Skip unreachable multi-term missing handling
Omit missing accessors when any physical column is full because the logical
field cannot be absent.
2026-08-04 13:23:19 +02:00
Pascal Seitz e77c256f01 simpler missing algorithm 2026-08-04 13:23:19 +02:00
Pascal Seitz 21bb20ff1a better clustering algorithm 2026-08-04 13:23:19 +02:00
Pascal Seitz 16f38f6a64 refactor 2026-08-04 13:23:19 +02:00
Pascal Seitz d3abc9a0b3 simplify 2026-08-04 13:23:19 +02:00
Pascal Seitz 17de33d8f0 merge multi-terms code 2026-08-04 13:23:19 +02:00
Pascal Seitz 295d5e4aa8 fix merge 2026-08-04 13:23:19 +02:00
Pascal Seitz 90ff347cc9 move impl 2026-08-04 13:23:19 +02:00
Pascal Seitz 766454d32a improve naming 2026-08-04 13:23:19 +02:00
Pascal Seitz 3010dbb8ea Optimize sparse multi-terms collection
Filter documents without values after each single-valued field so later
decoding and key construction only process viable documents. Use exact
block lengths to validate aligned results and keep unsupported missing-value
combinations on the general path.
2026-08-04 13:23:19 +02:00
Pascal Seitz 864a1d4b2d improve missing performance 2026-08-04 13:23:19 +02:00
Pascal Seitz 2d1ddca54d fmt 2026-08-04 13:23:19 +02:00
Pascal Seitz f1efd748fa cleanup 2026-08-04 13:23:19 +02:00
Pascal Seitz 164f62c4cc Simplify optional multi-terms document filtering
Track valid document positions in a reusable bitset instead of compacting
parallel document and index buffers after each optional field. Decode each
field against the original block and consult the mask before writing keys or
populating buckets.

This favors simpler collection logic over shrinking later sparse decodes.
2026-08-04 13:23:19 +02:00
Pascal Seitz 035fbe4d47 Refactor multi-terms collection around typed key codecs
Fan out dynamic/JSON physical column combinations during aggregation tree
construction so each segment collector handles one typed accessor per
field and applies missing values exactly once.

Use a generic key codec and shared term bucket maps for both packed and
unpacked keys. Keep optional and multivalued fields eligible for packed
storage, short-circuit sparse documents, and account for spilled composite
keys.
2026-08-04 13:23:19 +02:00
Pascal Seitz 2d91c31468 rename vars 2026-08-04 13:23:19 +02:00
Pascal Seitz 59b6859c16 move term req data to collector 2026-08-04 13:23:19 +02:00
trinity-1686a 9bcd99c396 rustfmt 2026-08-04 13:23:19 +02:00
trinity.pointard d6666d09f3 extract per-field fast-path compatiblity check 2026-08-04 13:23:19 +02:00
trinity.pointard 8af115f802 decode column once per doc max 2026-08-04 13:23:19 +02:00
trinity.pointard 2f477e9414 more codex review 2026-08-04 13:23:19 +02:00
trinity.pointard 1fcf935f5c fix handling of bytes for cardinality agg 2026-08-04 13:23:19 +02:00
trinity.pointard 73010f4d86 codex review 2026-08-04 13:23:19 +02:00
trinity.pointard 53d7af7b50 rustfmt 2026-08-04 13:23:19 +02:00
trinity.pointard 901af02fe0 fix bool serialization 2026-08-04 13:23:19 +02:00
trinity.pointard 9a3b4f3a8a first pass of optimising multi-terms agg 2026-08-04 13:23:19 +02:00
trinity.pointard 261c0a11d4 allow extracting values from IntermediateMultiTermsBucketResult 2026-08-04 13:23:19 +02:00
trinity.pointard ebd37278ca first swab at multi-terms impl 2026-08-04 13:23:19 +02:00
Pascal Seitz 2b9ba006b8 test: exercise parsed exists queries against documents 2026-08-03 12:14:51 +02:00
Pascal Seitz e1bc31e27e docs: document exists query syntax 2026-08-03 12:14:51 +02:00
Darkheir 276740bed4 feat: PR review suggestion
Signed-off-by: Darkheir <raphael.cohen@sekoia.io>
2026-08-03 12:14:51 +02:00
Darkheir f981e6e1f1 feat: Add support for Exists leaf in query parser
Signed-off-by: Darkheir <raphael.cohen@sekoia.io>
2026-08-03 12:14:51 +02:00
Paul MasurelandPaul Masurel 3bb9a430dd Refactoring TextFastFieldOptions (#3020)
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-08-02 17:17:02 +02:00
Pascal Seitz 667132fa7a rename to WithLanes 2026-07-29 08:57:54 +02:00
PSeitz 95f2c0c5b6 Optimize low-cardinality term aggregation counters 2026-07-29 08:57:54 +02:00
Pascal Seitz 277e59776b Extend Vec term storage to 20k with eager/lazy bucket ids
Raise MAX_NUM_TERMS_FOR_VEC to 20_000 so the dense Vec term storage
(direct-indexed, no hashing/paging) is used for low/moderate cardinality
both with and without sub-aggregations, not just the <100 low-card case.
Split the old single threshold into MAX_NUM_TERMS_FOR_LOWCARD_SUBAGG (100,
still gates the Vec + LowCard-buffer pairing) and MAX_NUM_TERMS_FOR_VEC.

VecTermBuckets is now generic over a compile-time `const LAZY: bool`:
- eager (default): assigns all ids up front in `new`, branchless term_entry
- lazy: assigns on first occurrence, keeping the sub-agg bucket range down
  to the terms that actually occur (no blowup for sparse ordinal spaces)

A new BucketIdSlot::ASSIGNS_ID const plus the LAZY const fully gate the
first-seen branch out at monomorphization, so the only branchy path is the
deliberately-lazy <BucketId, true> instantiation. The with-sub-agg Vec path
picks eager below MAX_NUM_TERMS_FOR_EAGER_BUCKET_IDS (4096) and lazy above,
avoiding the ~6-11% hot-loop regression eager fixes while keeping large
sparse cases compact.

Biggest impact on terms_zipf_1000_only with -40%
```
full
terms_7                                               Memory: 46.5 KB              Avg: 2.3566ms (+4.29%)      Median: 2.3566ms (+4.29%)      [2.3566ms .. 2.3566ms]
terms_all_unique                                      Memory: 11.5 MB              Avg: 4.8532ms (-5.74%)      Median: 4.8532ms (-5.74%)      [4.8532ms .. 4.8532ms]
terms_all_unique_order_by_key                         Memory: 11.5 MB              Avg: 4.8448ms (-0.53%)      Median: 4.8448ms (-0.53%)      [4.8448ms .. 4.8448ms]
terms_150_000                                         Memory: 2.7 MB               Avg: 5.5295ms (+0.64%)      Median: 5.5295ms (+0.64%)      [5.5295ms .. 5.5295ms]
terms_many_top_1000                                   Memory: 5.3 MB               Avg: 8.4738ms (+0.58%)      Median: 8.4738ms (+0.58%)      [8.4738ms .. 8.4738ms]
terms_many_order_by_term                              Memory: 2.7 MB (-0.01%)      Avg: 4.5959ms (-0.15%)      Median: 4.5959ms (-0.15%)      [4.5959ms .. 4.5959ms]
terms_many_with_top_hits                              Memory: 48.9 MB (-0.00%)     Avg: 85.8206ms (-5.46%)     Median: 85.8206ms (-5.46%)     [85.8206ms .. 85.8206ms]
terms_all_unique_with_avg_sub_agg                     Memory: 54.7 MB (-0.00%)     Avg: 17.8747ms (+8.16%)     Median: 17.8747ms (+8.16%)     [17.8747ms .. 17.8747ms]
terms_many_with_avg_sub_agg                           Memory: 13.5 MB (-0.00%)     Avg: 14.7797ms (+2.54%)     Median: 14.7797ms (+2.54%)     [14.7797ms .. 14.7797ms]
terms_status_with_avg_sub_agg                         Memory: 92.1 KB (-0.35%)     Avg: 5.1467ms (+0.72%)      Median: 5.1467ms (+0.72%)      [5.1467ms .. 5.1467ms]
terms_status_with_terms_zipf_1000_sub_agg             Memory: 213.2 KB             Avg: 3.9576ms (-2.73%)      Median: 3.9576ms (-2.73%)      [3.9576ms .. 3.9576ms]
terms_zipf_1000_with_terms_status_sub_agg             Memory: 726.5 KB (+0.28%)    Avg: 12.1432ms (+3.30%)     Median: 12.1432ms (+3.30%)     [12.1432ms .. 12.1432ms]
terms_status_with_histogram                           Memory: 139.0 KB             Avg: 2.4717ms (+2.45%)      Median: 2.4717ms (+2.45%)      [2.4717ms .. 2.4717ms]
terms_status_with_date_histogram                      Memory: 145.9 KB             Avg: 2.3644ms (-0.44%)      Median: 2.3644ms (-0.44%)      [2.3644ms .. 2.3644ms]
terms_status_with_date_histogram_hard_bounds          Memory: 145.4 KB             Avg: 2.5705ms (+4.70%)      Median: 2.5705ms (+4.70%)      [2.5705ms .. 2.5705ms]
terms_status_with_date_histogram_and_sibling_terms    Memory: 143.3 KB             Avg: 3.8946ms (+1.61%)      Median: 3.8946ms (+1.61%)      [3.8946ms .. 3.8946ms]
terms_zipf_1000_only                                  Memory: 75.8 KB (-0.02%)     Avg: 1.2684ms (-41.50%)     Median: 1.2684ms (-41.50%)     [1.2684ms .. 1.2684ms]
terms_zipf_1000_with_histogram                        Memory: 1.2 MB (+0.17%)      Avg: 20.9323ms (+4.08%)     Median: 20.9323ms (+4.08%)     [20.9323ms .. 20.9323ms]
terms_zipf_1000_with_avg_sub_agg                      Memory: 486.7 KB (+2.95%)    Avg: 8.5297ms (-1.66%)      Median: 8.5297ms (-1.66%)      [8.5297ms .. 8.5297ms]
terms_many_json_mixed_type_with_avg_sub_agg           Memory: 17.9 MB              Avg: 27.0210ms (+1.81%)     Median: 27.0210ms (+1.81%)     [27.0210ms .. 27.0210ms]
terms_status_with_cardinality_agg                     Memory: 93.9 KB              Avg: 3.2022ms (-0.64%)      Median: 3.2022ms (-0.64%)      [3.2022ms .. 3.2022ms]
terms_100_buckets_with_cardinality_agg                Memory: 9.8 MB (+0.31%)      Avg: 49.3188ms (-1.70%)     Median: 49.3188ms (-1.70%)     [49.3188ms .. 49.3188ms]
terms_many_with_single_term_order_by_card             Memory: 48.9 MB              Avg: 77.2127ms (-5.02%)     Median: 77.2127ms (-5.02%)     [77.2127ms .. 77.2127ms]
terms_many_with_single_term_2_order_by_card           Memory: 40.6 MB (-0.02%)     Avg: 45.8004ms (-17.57%)    Median: 45.8004ms (-17.57%)    [45.8004ms .. 45.8004ms]
```
2026-07-27 18:00:37 +02:00