Commit Graph
3717 Commits
Author SHA1 Message Date
Pascal Seitz 906fb8ef87 update comment 2026-09-07 18:26:37 +08:00
Pascal Seitz b586087392 fmt 2026-09-07 18:26:37 +08:00
Pascal Seitz ed1287a07f Preserve intermediate aggregation wire compatibility
Serialize vector-backed term entries as a map so Quickwit nodes running
different versions can exchange Postcard aggregation results. Add a legacy
fixture covering nested terms and histogram aggregations.
2026-09-07 18:26:37 +08:00
Pascal Seitz 921e98c63c Make count tie pruning deterministic at segment level 2026-09-07 18:26:37 +08:00
Pascal Seitz af2e6fac49 Remove equality from intermediate aggregation results 2026-09-07 18:26:37 +08:00
Pascal Seitz 71c7b78f32 Preserve term aggregation semantics during optimized merges
Merge normalized floating-point buckets and their sub-aggregations, and use
the final key ordering to resolve count ties while pruning results.
2026-09-07 18:26:37 +08:00
Pascal Seitz afe1666d6f Make aggregation key ordering total
Use total floating-point ordering for aggregation keys so ordering is
consistent with bitwise equality for NaNs and signed zero. This also
removes fallible key comparisons from term aggregation sorting.
2026-09-07 18:26:37 +08:00
Pascal Seitz 44a403d446 faster term agg merges 2026-09-07 18:26:37 +08:00
Pascal Seitz 797543f3ca add agg bench with 100 segments 2026-09-07 18:26:37 +08:00
Paul Masurel 1ae2da1859 Add DocPredicateQuery and a generic FunctionPredicate implementation (#3072)
* Add DocPredicateQuery and a generic FunctionPredicate implementation

Introduces the DocPredicate/SegmentDocPredicate abstraction and
DocPredicateQuery, a query that walks a segment's documents by
repeatedly evaluating a per-segment predicate (with a scorer_danger
fast path for boolean intersections). FunctionPredicate is a first,
generic implementation built from a plain per-segment factory
function, with no dependency on fast fields or any other segment
data structure.

* CR comments

* Better explain
2026-09-05 18:54:32 +02:00
Paul MasurelandPaul Masurel b5d8deb80c Fix fractional JSON range bounds on integer columns (#3074)
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-04 18:54:07 +02:00
Paul MasurelandPaul Masurel 79eedf4a81 Added the capacity to build scorer and seek_danger in a joint manner. (#3070)
* Added the capacity to build scorer and seek_danger in a joint manner.

* Added some partial support of scorer_danger for PhraseWeight

* Bugfix following CR

* CR: introducing OccurWeights

* CR comment

---------

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-03 14:07:38 +02:00
Paul MasurelandPaul Masurel ed6a90e840 Minor change: removing BooleanWeight::new as it is quite the footgun (#3071)
BooleanWeight::new just default min_should_match to 1.
This is misleading because I think someone who passes a single must clause would expect the result
to match that clause.

Due to min_should_match, it actually returns nothing.

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-02 13:47:51 +02:00
Pascal Seitz c3b91b617a use next_doc in mixed type columns
Remove full and empty variants from exists docsets
Simplify exists docset advancement
fix size hint in exist query
2026-09-02 14:14:45 +08:00
Pascal Seitz 271635b5a6 Use default seek behavior for exists docsets 2026-09-02 14:14:45 +08:00
Pascal Seitz a03135ac4f Simplify column index collection 2026-09-02 14:14:45 +08:00
Pascal Seitz 14c119745b Clarify fast-field exists weight name 2026-09-02 14:14:45 +08:00
Pascal Seitz a03870d41b Up to 500x Faster exists queries on columns (hehe)
Use specialized optional and multivalued column indexes for single-column
exists queries, while retaining a generic fallback for dynamic column unions.
Implement seek_danger with direct value checks to avoid unnecessary scans.

```
exists
column populated in 0.1% of docs (sparse blocks)
optional           Avg: 0.0406ms (-99.83%)    Median: 0.0395ms (-99.84%)    [0.0367ms .. 0.0496ms]    Output: 4_836
multivalued        Avg: 0.0432ms (-99.85%)    Median: 0.0424ms (-99.86%)    [0.0401ms .. 0.0503ms]    Output: 4_836
column populated in 10% of docs (dense blocks)
optional           Avg: 6.1052ms (-52.67%)    Median: 6.1089ms (-51.50%)    [5.9257ms .. 6.2702ms]    Output: 499_966
multivalued        Avg: 6.5485ms (-64.36%)    Median: 6.4573ms (-64.12%)    [6.3324ms .. 8.0991ms]    Output: 499_966
column populated in 90% of docs (dense blocks)
optional           Avg: 12.9950ms (+14.89%)    Median: 12.8676ms (+14.74%)    [12.6985ms .. 14.3235ms]    Output: 4_498_850
multivalued        Avg: 15.7813ms (-39.05%)    Median: 15.6722ms (-39.20%)    [15.5501ms .. 16.6724ms]    Output: 4_498_850
term_AND_exists
term matches 0.01% of docs, column populated in 0.1% (sparse blocks)
optional           Avg: 7381ns (-99.70%)    Median: 7030ns (-99.72%)    [6798ns .. 0.0144ms]    Output: 1
multivalued        Avg: 8435ns (-99.74%)    Median: 7372ns (-99.75%)    [7152ns .. 0.0313ms]    Output: 1
term matches 1% of docs, column populated in 0.1% (sparse blocks)
optional           Avg: 0.1047ms (-99.52%)    Median: 0.1045ms (-99.52%)    [0.1040ms .. 0.1072ms]    Output: 51
multivalued        Avg: 0.1070ms (-99.59%)    Median: 0.1069ms (-99.58%)    [0.1063ms .. 0.1097ms]    Output: 51
term matches 50% of docs, column populated in 0.1% (sparse blocks)
optional           Avg: 0.2970ms (-98.72%)    Median: 0.2839ms (-98.76%)    [0.2648ms .. 0.3446ms]    Output: 2_405
multivalued        Avg: 0.2952ms (-98.95%)    Median: 0.2923ms (-98.96%)    [0.2641ms .. 0.3283ms]    Output: 2_405
term matches 0.01% of docs, column populated in 10% (dense blocks)
optional           Avg: 0.0124ms (-48.17%)    Median: 0.0124ms (-47.62%)    [0.0120ms .. 0.0131ms]    Output: 50
multivalued        Avg: 0.0196ms (-46.61%)    Median: 0.0197ms (-46.21%)    [0.0184ms .. 0.0205ms]    Output: 50
term matches 1% of docs, column populated in 10% (dense blocks)
optional           Avg: 0.8094ms (-44.69%)    Median: 0.8021ms (-44.60%)    [0.7787ms .. 0.9247ms]    Output: 5_023
multivalued        Avg: 0.8439ms (-58.79%)    Median: 0.8325ms (-58.70%)    [0.8099ms .. 0.9976ms]    Output: 5_023
term matches 50% of docs, column populated in 10% (dense blocks)
optional           Avg: 12.3175ms (-32.65%)    Median: 12.3049ms (-32.54%)    [12.1053ms .. 12.8036ms]    Output: 249_853
multivalued        Avg: 12.7152ms (-46.31%)    Median: 12.6789ms (-46.23%)    [12.5517ms .. 13.0544ms]    Output: 249_853
term matches 0.01% of docs, column populated in 90% (dense blocks)
optional           Avg: 0.0627ms (+3.59%)    Median: 0.0640ms (+5.17%)    [0.0557ms .. 0.0680ms]    Output: 439
multivalued        Avg: 0.1168ms (+1.16%)    Median: 0.1173ms (+1.90%)    [0.1050ms .. 0.1226ms]    Output: 439
term matches 1% of docs, column populated in 90% (dense blocks)
optional           Avg: 0.5377ms (-0.19%)     Median: 0.5346ms (-0.26%)     [0.5308ms .. 0.5550ms]    Output: 45_342
multivalued        Avg: 0.6365ms (-31.29%)    Median: 0.6351ms (-30.75%)    [0.6245ms .. 0.6572ms]    Output: 45_342
term matches 50% of docs, column populated in 90% (dense blocks)
optional           Avg: 27.3888ms (-1.86%)     Median: 27.3356ms (-1.48%)     [27.1377ms .. 28.7410ms]    Output: 2_249_485
multivalued        Avg: 29.4214ms (-25.34%)    Median: 29.3960ms (-25.17%)    [29.2107ms .. 30.2541ms]    Output: 2_249_485
```
2026-09-02 14:14:45 +08:00
Pascal Seitz c85e428e0f Benchmark exists queries on columns 2026-09-02 14:14:45 +08:00
Paul MasurelandPaul Masurel 21ae09c0b4 Refactoring to introduce calculated fields. (#3066)
* Make aggregation column block accessor private

Move ColumnBlockAccessor out of tantivy-columnar and into the aggregation implementation. Keep its behavior and tests intact while removing it from the public columnar API.

* Refactor aggregation value block access

* Following comment

Removing Fused from the element that are private to the module.
Re-added optimization in stats for single doc requests to avoid perf regression
Renamed Fused -> Flattened

---------

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-01 15:10:19 +02:00
Paul MasurelandPaul Masurel 35a6187ed1 fix formatting (#3062)
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-08-31 14:40:27 +02:00
jinhelin 7b0e97a0b6 fix(query): disable child scoring in ConstScoreQuery (#3059) 2026-08-27 17:52:17 +02:00
Divyesh Kakadiya cce9950a9e Clarify that OnCommitWithDelay reload is not synchronized with commit() (#3057)
commit() only guarantees documents are persisted to the Directory; an
already-open IndexReader only picks up the change according to its
ReloadPolicy. With the default OnCommitWithDelay, the reload happens
asynchronously on a background thread and is not guaranteed to be done
by the time commit() returns, on any Directory implementation
(including RamDirectory), so a search immediately after commit() can
still miss the documents just committed.

Document this explicitly on ReloadPolicy::OnCommitWithDelay,
IndexReader::reload(), and IndexWriter::commit(), and add a runnable
example showing the ReloadPolicy::Manual + explicit reader.reload()
pattern for deterministic read-your-writes.

See #1824
2026-08-26 14:46:02 +02:00
Paul MasurelandPaul Masurel 266a6c48f2 Fix numeric term key normalization across segments (#3051)
Normalize numerical values before constructing intermediate term keys so equal values merge even when their physical column types differ between segments.

Add a regression test covering u64 and i64 columns with a shared value.

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-08-19 23:09:39 +02:00
Paul MasurelandPaul Masurel fa904b3a72 Minor refactoring of find_missing_docs (#3049)
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-08-19 18:11:14 +02:00
Pavlos Rontidis 039a72958e Merge pull request #3041 from quickwit-oss/cose-sync-security-policy-20260813224347
chore: sync security-policy
2026-08-14 09:09:08 -04:00
DavIvek a1aa571905 test(query-parser): cover json phrase prefix and slop against indexed documents
Adds an end-to-end test that indexes real JSON documents and asserts hit
counts, complementing the existing logical-AST assertions. The test fails
without the fix (phrase prefix returns 0 hits instead of 2).
2026-08-14 16:10:31 +08:00
DavIvek 51d3000ce6 fix(query-parser): honor phrase prefix and slop on JSON fields
Phrase prefix (`"..."*`) and slop (`"..."~N`) were silently dropped on JSON
fields. `generate_literals_for_json_object` hard-coded `slop: 0` and
`prefix: false`, so a query like `data.name:"foo bar"*` degraded to an exact
phrase, even though the underlying `PhrasePrefixQuery` already supports
JSON-path terms.

Thread `slop`/`prefix` through the JSON branch, exactly as the `Str` branch
already does, and add the "phrase prefix requires at least two terms" guard for
the single-token case (mirroring `generate_literals_for_str`).

Tests added as JSON analogues of the existing text-field tests:
- test_phrase_prefix_on_json_field
- test_phrase_prefix_too_short_on_json_field
- test_phrase_slop_on_json_field

This aligns JSON fields with the phrase `~`/`*` behavior already documented for
`QueryParser`.
2026-08-14 16:10:31 +08:00
David Yaffe afb3aed299 chore: sync security-policy from cose 2026-08-13 18:43:47 -04:00
Pascal Seitz 7373d54c2d fix clippy 2026-08-12 21:08:06 +08:00
Paul MasurelandPaul Masurel 1f32c1a8af Removing TODO.txt (#3037)
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-08-11 10:30:15 +02:00
Ming 0401b45781 feat: Extensible segment components via plugin trait (#2993)
## Motivation

Today every segment component — postings, fast fields, field norms, store — is hardcoded into several places. Adding a new per-segment data structure means forking Tantivy and editing each of those sites.

This PR introduces a `SegmentPlugin` trait that lets a custom component participate in the full segment lifecycle — write, serialize, merge, garbage collection, space usage — through the same interface the built-ins use, without touching Tantivy internals. The four built-in components are themselves reimplemented as plugins.

We (ParadeDB) plan on using this trait for 1) additional segment metadata for partitioning 2) custom vector index.

## Plugin Trait

Two traits. The first is the `SegmentPlugin` factory:

```rust
pub trait SegmentPlugin: Send + Sync + 'static {
    /// File extensions this component owns, e.g. ["idx", "pos", "term"] for postings.
    fn extensions(&self) -> &[&str];

    /// Create a writer for the indexing path.
    fn create_writer(&self, ctx: &PluginWriterContext) -> crate::Result<Box<dyn PluginWriter>>;

    /// Merge this component across several source segments into the target segment.
    fn merge(&self, ctx: PluginMergeContext) -> crate::Result<()>;

    /// Report on-disk space usage, keyed by component name. Has a default impl.
    fn space_usage(&self, reader: &SegmentReader)
        -> crate::Result<BTreeMap<String, ComponentSpaceUsage>>;
}
```

A `SegmentPlugin` owns one or more file extensions and knows how to (a) build a writer for the indexing path and (b) merge itself across segments.

The second trait is the segment writer:

```rust
pub trait PluginWriter: Send + Any {
    /// Called once per document, in doc-id order, for every plugin writer.
    fn add_document(&mut self, doc_id: DocId, doc: &TantivyDocument, schema: &Schema)
        -> crate::Result<()> { Ok(()) }

    /// Serialize accumulated data to segment files (honoring an optional doc-id remap).
    fn serialize(&mut self, segment: &Segment, doc_id_map: Option<&DocIdMapping>) -> crate::Result<()>;

    fn close(self: Box<Self>) -> crate::Result<()>;
    fn mem_usage(&self) -> usize;

    fn as_any(&self) -> &dyn Any;       // downcast support, Rust 1.86
    fn as_any_mut(&mut self) -> &mut dyn Any;
}
```

The write path no longer has any by-name wiring: `SegmentWriter` hands every document to every plugin writer's `add_document`, and `finalize()` calls `serialize` then close on each. 

## Key Design Decisions

1. The index — not the segment — owns the plugin set. The set of custom plugins is recorded once, at index creation, in `IndexMeta`. `#[serde(default)]` makes this backward compatible.

2. Plugins are registered, like tokenizers — and re-registration is enforced fail-closed. Plugins are not serialized; they're re-attached on every `Index::open` via `register_plugin`, exactly like custom tokenizers. To prevent consumers from accidentally forgetting to register a plugin, we validate the registered plugin set against the persisted set when the index is first used for a write/merge/GC operation.

3. Registration order is the write/merge order. Built-ins come first (field norms → postings → fast fields → store), then custom plugins.

4. The read side needs no plugin hook. Custom component data is read back through the existing public surface `SegmentReader::open_read`.

5. Backwards compatibility — behavior for existing indexes is unchanged.
2026-08-10 16:06:11 -07:00
PSeitz-dd 52b607d7cd Merge pull request #3018 from PSeitz/low_card_histogram_lanes
Up to 2x faster fused term-histogram aggregation
2026-08-10 09:39:48 +02:00
Pascal Seitz d212a9c532 Address fused term-histogram review feedback
Decouple the low-cardinality sub-aggregation threshold from the fused
count-lane threshold, rely on the full-column invariant in the linear
resolver, and explicitly test the single-bucket resolver.
2026-08-10 15:30:57 +08:00
Pascal Seitz 6797d8c577 const assert, unreachable 2026-08-10 15:03:46 +08:00
trinity-1686a 5fc9fccb61 Merge pull request #3030 from foundational-io/fix/streamer-term-ord-skipped-blocks
fix(sstable): report the real term ordinal when an automaton prunes blocks
2026-08-07 17:14:53 +02:00
dependabot[bot] 1990205ecf Bump actions/checkout from 6.0.3 to 7.0.1 (#3006)
Bumps [actions/checkout](https://github.com/actions/checkout) from 6.0.3 to 7.0.1.
- [Release notes](https://github.com/actions/checkout/releases)
- [Changelog](https://github.com/actions/checkout/blob/main/CHANGELOG.md)
- [Commits](https://github.com/actions/checkout/compare/df4cb1c069e1874edd31b4311f1884172cec0e10...3d3c42e5aac5ba805825da76410c181273ba90b1)

---
updated-dependencies:
- dependency-name: actions/checkout
  dependency-version: 7.0.1
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-07 12:03:20 +02:00
mmustafasenoglu cac4ba9404 chore: downgrade all info! logs to debug! (#3013)
* chore: downgrade info! logs to debug! in managed_directory.rs

* chore: downgrade info! logs to debug! in file_watcher.rs

* chore: downgrade info! logs to debug! in index_writer.rs

* chore: downgrade info! logs to debug! in prepared_commit.rs

* chore: downgrade info! logs to debug! in segment_updater.rs
2026-08-07 11:58:56 +02:00
dependabot[bot] a91476fd56 build(deps): bump ossf/scorecard-action from 2.4.3 to 2.4.4 (#3015)
Bumps [ossf/scorecard-action](https://github.com/ossf/scorecard-action) from 2.4.3 to 2.4.4.
- [Release notes](https://github.com/ossf/scorecard-action/releases)
- [Changelog](https://github.com/ossf/scorecard-action/blob/main/RELEASE.md)
- [Commits](https://github.com/ossf/scorecard-action/compare/4eaacf0543bb3f2c246792bd56e8cdeffafb205a...2d1146689b8cda280b9bc96326124645441f03bc)

---
updated-dependencies:
- dependency-name: ossf/scorecard-action
  dependency-version: 2.4.4
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-07 11:58:11 +02:00
dependabot[bot] 60a982c6e1 build(deps): update base64 requirement from 0.22.0 to 0.23.0 (#3014)
Updates the requirements on [base64](https://github.com/marshallpierce/rust-base64) to permit the latest version.
- [Changelog](https://github.com/marshallpierce/rust-base64/blob/master/RELEASE-NOTES.md)
- [Commits](https://github.com/marshallpierce/rust-base64/compare/v0.22.0...v0.23.0)

---
updated-dependencies:
- dependency-name: base64
  dependency-version: 0.23.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-07 11:57:54 +02:00
Marc Bachmann 52a95b159b fix: do not drive a non-summing score combiner with Block-WAND (#3022)
`BooleanWeight::for_each_pruning` handed a union of term scorers straight
to `block_wand`, whatever the weight's score combiner. Block-WAND prunes on
the sum of the per-term block maxima and scores the survivors the same way,
so it silently replaces the combiner with a sum. A `DisjunctionMaxQuery`
collected through `TopDocs` therefore scored a document matching in two
fields as the sum of both instead of the better of the two, and ignored the
tie breaker.

The union only reaches that path when every clause reads term frequencies,
which is why the bug hid: a query asking for `Basic` postings falls back to
`BufferedUnionScorer`, which does honor the combiner.

`ScoreCombiner` now declares whether Block-WAND may drive it. Only
`SumCombiner` opts in; everything else falls back to the plain union, which
still prunes on the threshold but makes no assumption about how the scores
combine.

The intersection specialization needs no such fallback: `Intersection::score`
sums its children whatever the combiner (the combiner only shapes how should
clauses combine), so `block_wand_intersection`'s summing already matches the
unpruned scorer for every combiner. A second test pins that equivalence by
comparing `for_each_pruning` against a manual scorer walk on a must+must
weight built with `DisjunctionMaxCombiner`.
2026-08-07 11:57:22 +02:00
Xuanwo 3a55bc986d Replace rust-stemmers with frostem for stemming (#3028)
rust-stemmers is unmaintained and lags behind upstream Snowball.
frostem tracks snowball main and regenerates algorithms automatically.
2026-08-06 10:54:56 +02:00
dependabot[bot] 5ca3933200 Update lru requirement from 0.16.3 to 0.18.2 (#3034)
Updates the requirements on [lru](https://github.com/jeromefroe/lru-rs) to permit the latest version.
- [Changelog](https://github.com/jeromefroe/lru-rs/blob/master/CHANGELOG.md)
- [Commits](https://github.com/jeromefroe/lru-rs/compare/0.16.3...0.18.2)

---
updated-dependencies:
- dependency-name: lru
  dependency-version: 0.18.2
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-06 10:18:16 +02:00
Pascal Seitz 3bd8a6b4d1 Optimize fused term-histogram aggregation
Avoid histogram decoding for single-bucket ranges and use precomputed
boundaries for small ranges. Split hot counters into lanes to reduce
write dependencies on low-cardinality grids.
2026-08-06 08:41:07 +02:00
Pascal Seitz 62ca2f3608 add termhistogram bench for few buckets 2026-08-06 08:41:07 +02:00
trinity-1686a 86641f72df Merge pull request #3029 from quickwit-oss/trinity.pointard/get-ff-tokenizer-comeback
reintroduce get_fast_field_tokenizer_name
2026-08-04 21:04:46 +02:00
ildis 7f6685f5d6 fix(sstable): report the real term ordinal when an automaton prunes blocks
`Streamer::term_ord()` derived the ordinal by counting `delta_reader.advance()`
calls from a seed taken only when the stream had a key lower bound. An automaton
search has no key bounds, so the seed was 0 — but the reader it is handed *is*
block-pruned by that automaton (`get_block_iterator_for_range_and_automaton`).
Every block skipped ahead of the first match went uncounted, so `term_ord()`
returned the term's position among the blocks actually scanned rather than its
ordinal in the dictionary.

The error is silent and grows with how deep the first match sits: on a 262k-term
dictionary, a regex matching only the last key reported ordinal 999 instead of
262142. Callers that resolve those ordinals back to terms therefore act on a
different term entirely — `build_allowed_term_ids_for_str` builds the terms
aggregation's allowed-ordinal bitset this way, so an `include` regex could make
the aggregation count unrelated terms while returning the expected bucket count.
Broad patterns matching from the start of the dictionary hid it, since nothing is
pruned ahead of the first match.

Each slice handed to the reader now carries the ordinal of its first term, and
the streamer resets to it on entering a slice instead of incrementing across the
gap. Blocks merged into one slice stay contiguous, so counting within a slice is
still correct.

`Dictionary::sorted_ords_to_term_cb` already tracked `BlockAddr::first_ordinal`
explicitly, which is why ordinal->term resolution was unaffected.
2026-08-04 21:04:17 +03:00
trinity-1686a 0743fe7d67 reintroduce get_fast_field_tokenizer_name 2026-08-04 18:24:15 +02:00
Pascal Seitz bc225eecdf rename bench name 2026-08-04 13:23:19 +02:00
Pascal Seitz 83ffe96189 Reuse existing term ordinals for multi-terms missing values
Map a configured string missing value to its dictionary ordinal when the
term is already present. This coalesces real and missing documents before
the segment-level cutoff, preserving counts and sub-aggregation results.

Keep synthetic sentinels for absent terms and non-string columns, and add
regression coverage with a segment size of one.
2026-08-04 13:23:19 +02:00