Commit Graph
3742 Commits
Author SHA1 Message Date
Pascal Seitz dfe904667e Minimize unrelated changes 2026-09-16 23:39:47 +08:00
Pascal Seitz a9d3f1a026 Run nightly fmt 2026-09-16 23:39:47 +08:00
Pascal Seitz bbc1b9e1d1 Restore TerminatingWrite API 2026-09-16 23:39:47 +08:00
Pascal Seitz 1c8a4801b3 Shrink RamDirectory files before storing 2026-09-16 23:39:47 +08:00
Pascal Seitz 2ab4fb80f0 Buffer RamDirectory writes
Use the standard BufWriter capacity to avoid forwarding every small write
directly to the underlying VecWriter.
2026-09-16 23:39:47 +08:00
Pascal Seitz e46d44c388 Finalize mmap writers in directory tests
Use finish instead of flush before reading completed files and clarify
when file data is synchronized.
2026-09-16 23:39:47 +08:00
Pascal Seitz f2a106df1c Prevent concurrent RamDirectory writers
Reserve paths when writers are opened and release reservations when
writers finish or drop.
2026-09-16 23:39:47 +08:00
Pascal Seitz 98d91f0436 Remove intermediate flush from directory contract
Intermediate writes are an implementation detail. From Tantivy side we only decide when a file is finished. This change is done to remove flush overhead for VecWriter. On finalization we just move the Vec to the directory.

Rename TerminatingWrite to FinishableWrite to reflect these semantics.
2026-09-16 23:39:47 +08:00
Pascal Seitz b06d8e9a04 Track active RamDirectory writer memory
Include VecWriter allocations in total_mem_usage so callers can account
for memory before writers are flushed or terminated.
2026-09-16 23:39:47 +08:00
Pascal Seitz 54b82e582f Report unsupported whitespace as a syntax error 2026-09-16 22:53:35 +08:00
Pascal SeitzandOleksii Syniakov 34edbf7821 Fix query parser panics on malformed input
Treat a fieldless exists leaf as match-all and make the lenient set parser
stop when parsing no longer consumes input. Add public QueryParser proptests
covering both regressions.

Findings and original fixes by Oleksii Syniakov in osyniakov/tantivy#4.

Co-authored-by: Oleksii Syniakov <1282756+osyniakov@users.noreply.github.com>
2026-09-16 22:53:35 +08:00
Paul Masurel 94a116f103 (Calculated fields) Add feature-gated JIT expression document predicate (#3081) 2026-09-15 16:28:53 +02:00
Paul MasurelandPaul Masurel 99043285f9 Set jitexpr license to MIT (#3101)
Add MIT license metadata to the jitexpr crate manifest, matching
tantivy and the other workspace subcrates, and include a copy of
the MIT license text.

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-14 15:08:55 +02:00
Paul MasurelandPaul Masurel 4abd98fa19 Bump coverage nightly pin to satisfy cranelift MSRV (#3094)
The coverage job pinned nightly-2025-12-01, which is rustc 1.93. Since
`--all-features --workspace` pulls in jitexpr -> cranelift 0.134.4, and
those crates declare rust-version = "1.94.0", cargo refused to build:

    error: rustc 1.93.1 is not supported by the following packages:
      cranelift@0.134.4 requires rustc 1.94.0

Bump the pin to nightly-2026-08-31 (1.100.0-nightly), which builds
jitexpr cleanly.

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-11 09:32:00 +02:00
Paul MasurelandPaul Masurel 3bad51518a Fix clippy warnings in test code (#3091)
Silence the 7 warnings reported by `cargo clippy --all --tests`:

- Replace manual `% n == 0` parity checks with `is_multiple_of`
  (`manual_is_multiple_of`) in index_writer, doc_predicate_query and
  seek_danger tests. `is_multiple_of` is stable for unsigned integers as
  of 1.87, which matches the crate MSRV.
- Prefix the intentionally unused `other_segment_meta` binding with an
  underscore. The binding is kept rather than dropped because the
  inventory only tracks live `SegmentMeta`s, so dropping it would
  invalidate the `inventory.all()` assertion in the same test.

No behavior change.

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-10 21:26:45 +02:00
Paul Masurel 0c0530eeb4 Free JITModule executable memory when CompiledFn is dropped (#3090)
JITModule never releases its executable allocation on drop; the only
way to reclaim it is the consuming, unsafe JITModule::free_memory.
CompiledFn previously stored the module in a plain field, so every
compiled expression's executable memory leaked for the remainder of
the process once the CompiledFn was dropped.

Wrap the module in ManuallyDrop and add an explicit Drop impl for
CompiledFn that calls free_memory. This is safe because CompiledFn is
only ever constructed behind an Arc and never exposes its entry point
or other raw pointers into the module outside &self-bounded calls, so
drop only runs once no call into the module can be in flight or happen
afterward.

Verified empirically: 200k compile+drop cycles peaked at ~3.3 GB RSS
before this fix and ~6 MB after.
2026-09-10 21:22:01 +02:00
dependabot[bot] dd0bc23360 Update zstd requirement from 0.13 to 0.14 (#3085)
Updates the requirements on [zstd](https://github.com/gyscos/zstd-rs) to permit the latest version.
- [Release notes](https://github.com/gyscos/zstd-rs/releases)
- [Commits](https://github.com/gyscos/zstd-rs/compare/v0.13.0...v0.14.0)

---
updated-dependencies:
- dependency-name: zstd
  dependency-version: 0.14.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-10 18:36:24 +02:00
Paul Masurel 5a82f19f5d (Calculated fields) Implements more functions (#3083)
* Add remaining expression functions

Currently experimental (!)
2026-09-10 18:35:51 +02:00
Paul Masurel f8ec38fd9f (Calculated fields) Add JIT expression framework with ADD, IS_NULL, and REGEXP_EXTRACT (#3078)
Add JIT expression framework with ADD, IS_NULL, and REGEXP_EXTRACT
2026-09-10 15:30:02 +02:00
Pascal Seitz bc8a21cf12 Document 0.26.2 bugfixes and mark 0.26 as released 2026-09-08 23:36:16 +08:00
PSeitz-dd 9d10bf1c6c Merge pull request #3088 from PSeitz/check_seek_danger_panic
fix: disable buffered union seek_danger override
2026-09-08 17:05:24 +02:00
Pascal Seitz 66ff8466d5 fix: disable buffered union seek_danger override
Use the default seek-based implementation until the optimized path can handle invalid child states safely. A child miss followed by another child hit previously allowed refill to consume an unaligned scorer.

Add shared seek_danger contract proptests, a regression for #3086, and a source inventory guard for future overrides.
2026-09-08 16:34:05 +02:00
pascal 780395c1db improve BufferedUnionScorer::advance 2026-09-08 18:12:08 +08:00
pascal 6929bae13f fix union performance regression 2026-09-08 18:12:08 +08:00
Paul MasurelandPaul Masurel 1020eba2b9 Expose segment file listing on IndexMeta (#3080)
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-07 13:55:06 +02:00
Pascal Seitz 906fb8ef87 update comment 2026-09-07 18:26:37 +08:00
Pascal Seitz b586087392 fmt 2026-09-07 18:26:37 +08:00
Pascal Seitz ed1287a07f Preserve intermediate aggregation wire compatibility
Serialize vector-backed term entries as a map so Quickwit nodes running
different versions can exchange Postcard aggregation results. Add a legacy
fixture covering nested terms and histogram aggregations.
2026-09-07 18:26:37 +08:00
Pascal Seitz 921e98c63c Make count tie pruning deterministic at segment level 2026-09-07 18:26:37 +08:00
Pascal Seitz af2e6fac49 Remove equality from intermediate aggregation results 2026-09-07 18:26:37 +08:00
Pascal Seitz 71c7b78f32 Preserve term aggregation semantics during optimized merges
Merge normalized floating-point buckets and their sub-aggregations, and use
the final key ordering to resolve count ties while pruning results.
2026-09-07 18:26:37 +08:00
Pascal Seitz afe1666d6f Make aggregation key ordering total
Use total floating-point ordering for aggregation keys so ordering is
consistent with bitwise equality for NaNs and signed zero. This also
removes fallible key comparisons from term aggregation sorting.
2026-09-07 18:26:37 +08:00
Pascal Seitz 44a403d446 faster term agg merges 2026-09-07 18:26:37 +08:00
Pascal Seitz 797543f3ca add agg bench with 100 segments 2026-09-07 18:26:37 +08:00
Paul Masurel 1ae2da1859 Add DocPredicateQuery and a generic FunctionPredicate implementation (#3072)
* Add DocPredicateQuery and a generic FunctionPredicate implementation

Introduces the DocPredicate/SegmentDocPredicate abstraction and
DocPredicateQuery, a query that walks a segment's documents by
repeatedly evaluating a per-segment predicate (with a scorer_danger
fast path for boolean intersections). FunctionPredicate is a first,
generic implementation built from a plain per-segment factory
function, with no dependency on fast fields or any other segment
data structure.

* CR comments

* Better explain
2026-09-05 18:54:32 +02:00
Paul MasurelandPaul Masurel b5d8deb80c Fix fractional JSON range bounds on integer columns (#3074)
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-04 18:54:07 +02:00
Paul MasurelandPaul Masurel 79eedf4a81 Added the capacity to build scorer and seek_danger in a joint manner. (#3070)
* Added the capacity to build scorer and seek_danger in a joint manner.

* Added some partial support of scorer_danger for PhraseWeight

* Bugfix following CR

* CR: introducing OccurWeights

* CR comment

---------

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-03 14:07:38 +02:00
Paul MasurelandPaul Masurel ed6a90e840 Minor change: removing BooleanWeight::new as it is quite the footgun (#3071)
BooleanWeight::new just default min_should_match to 1.
This is misleading because I think someone who passes a single must clause would expect the result
to match that clause.

Due to min_should_match, it actually returns nothing.

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-02 13:47:51 +02:00
Pascal Seitz c3b91b617a use next_doc in mixed type columns
Remove full and empty variants from exists docsets
Simplify exists docset advancement
fix size hint in exist query
2026-09-02 14:14:45 +08:00
Pascal Seitz 271635b5a6 Use default seek behavior for exists docsets 2026-09-02 14:14:45 +08:00
Pascal Seitz a03135ac4f Simplify column index collection 2026-09-02 14:14:45 +08:00
Pascal Seitz 14c119745b Clarify fast-field exists weight name 2026-09-02 14:14:45 +08:00
Pascal Seitz a03870d41b Up to 500x Faster exists queries on columns (hehe)
Use specialized optional and multivalued column indexes for single-column
exists queries, while retaining a generic fallback for dynamic column unions.
Implement seek_danger with direct value checks to avoid unnecessary scans.

```
exists
column populated in 0.1% of docs (sparse blocks)
optional           Avg: 0.0406ms (-99.83%)    Median: 0.0395ms (-99.84%)    [0.0367ms .. 0.0496ms]    Output: 4_836
multivalued        Avg: 0.0432ms (-99.85%)    Median: 0.0424ms (-99.86%)    [0.0401ms .. 0.0503ms]    Output: 4_836
column populated in 10% of docs (dense blocks)
optional           Avg: 6.1052ms (-52.67%)    Median: 6.1089ms (-51.50%)    [5.9257ms .. 6.2702ms]    Output: 499_966
multivalued        Avg: 6.5485ms (-64.36%)    Median: 6.4573ms (-64.12%)    [6.3324ms .. 8.0991ms]    Output: 499_966
column populated in 90% of docs (dense blocks)
optional           Avg: 12.9950ms (+14.89%)    Median: 12.8676ms (+14.74%)    [12.6985ms .. 14.3235ms]    Output: 4_498_850
multivalued        Avg: 15.7813ms (-39.05%)    Median: 15.6722ms (-39.20%)    [15.5501ms .. 16.6724ms]    Output: 4_498_850
term_AND_exists
term matches 0.01% of docs, column populated in 0.1% (sparse blocks)
optional           Avg: 7381ns (-99.70%)    Median: 7030ns (-99.72%)    [6798ns .. 0.0144ms]    Output: 1
multivalued        Avg: 8435ns (-99.74%)    Median: 7372ns (-99.75%)    [7152ns .. 0.0313ms]    Output: 1
term matches 1% of docs, column populated in 0.1% (sparse blocks)
optional           Avg: 0.1047ms (-99.52%)    Median: 0.1045ms (-99.52%)    [0.1040ms .. 0.1072ms]    Output: 51
multivalued        Avg: 0.1070ms (-99.59%)    Median: 0.1069ms (-99.58%)    [0.1063ms .. 0.1097ms]    Output: 51
term matches 50% of docs, column populated in 0.1% (sparse blocks)
optional           Avg: 0.2970ms (-98.72%)    Median: 0.2839ms (-98.76%)    [0.2648ms .. 0.3446ms]    Output: 2_405
multivalued        Avg: 0.2952ms (-98.95%)    Median: 0.2923ms (-98.96%)    [0.2641ms .. 0.3283ms]    Output: 2_405
term matches 0.01% of docs, column populated in 10% (dense blocks)
optional           Avg: 0.0124ms (-48.17%)    Median: 0.0124ms (-47.62%)    [0.0120ms .. 0.0131ms]    Output: 50
multivalued        Avg: 0.0196ms (-46.61%)    Median: 0.0197ms (-46.21%)    [0.0184ms .. 0.0205ms]    Output: 50
term matches 1% of docs, column populated in 10% (dense blocks)
optional           Avg: 0.8094ms (-44.69%)    Median: 0.8021ms (-44.60%)    [0.7787ms .. 0.9247ms]    Output: 5_023
multivalued        Avg: 0.8439ms (-58.79%)    Median: 0.8325ms (-58.70%)    [0.8099ms .. 0.9976ms]    Output: 5_023
term matches 50% of docs, column populated in 10% (dense blocks)
optional           Avg: 12.3175ms (-32.65%)    Median: 12.3049ms (-32.54%)    [12.1053ms .. 12.8036ms]    Output: 249_853
multivalued        Avg: 12.7152ms (-46.31%)    Median: 12.6789ms (-46.23%)    [12.5517ms .. 13.0544ms]    Output: 249_853
term matches 0.01% of docs, column populated in 90% (dense blocks)
optional           Avg: 0.0627ms (+3.59%)    Median: 0.0640ms (+5.17%)    [0.0557ms .. 0.0680ms]    Output: 439
multivalued        Avg: 0.1168ms (+1.16%)    Median: 0.1173ms (+1.90%)    [0.1050ms .. 0.1226ms]    Output: 439
term matches 1% of docs, column populated in 90% (dense blocks)
optional           Avg: 0.5377ms (-0.19%)     Median: 0.5346ms (-0.26%)     [0.5308ms .. 0.5550ms]    Output: 45_342
multivalued        Avg: 0.6365ms (-31.29%)    Median: 0.6351ms (-30.75%)    [0.6245ms .. 0.6572ms]    Output: 45_342
term matches 50% of docs, column populated in 90% (dense blocks)
optional           Avg: 27.3888ms (-1.86%)     Median: 27.3356ms (-1.48%)     [27.1377ms .. 28.7410ms]    Output: 2_249_485
multivalued        Avg: 29.4214ms (-25.34%)    Median: 29.3960ms (-25.17%)    [29.2107ms .. 30.2541ms]    Output: 2_249_485
```
2026-09-02 14:14:45 +08:00
Pascal Seitz c85e428e0f Benchmark exists queries on columns 2026-09-02 14:14:45 +08:00
Paul MasurelandPaul Masurel 21ae09c0b4 Refactoring to introduce calculated fields. (#3066)
* Make aggregation column block accessor private

Move ColumnBlockAccessor out of tantivy-columnar and into the aggregation implementation. Keep its behavior and tests intact while removing it from the public columnar API.

* Refactor aggregation value block access

* Following comment

Removing Fused from the element that are private to the module.
Re-added optimization in stats for single doc requests to avoid perf regression
Renamed Fused -> Flattened

---------

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-01 15:10:19 +02:00
Paul MasurelandPaul Masurel 35a6187ed1 fix formatting (#3062)
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-08-31 14:40:27 +02:00
jinhelin 7b0e97a0b6 fix(query): disable child scoring in ConstScoreQuery (#3059) 2026-08-27 17:52:17 +02:00
Divyesh Kakadiya cce9950a9e Clarify that OnCommitWithDelay reload is not synchronized with commit() (#3057)
commit() only guarantees documents are persisted to the Directory; an
already-open IndexReader only picks up the change according to its
ReloadPolicy. With the default OnCommitWithDelay, the reload happens
asynchronously on a background thread and is not guaranteed to be done
by the time commit() returns, on any Directory implementation
(including RamDirectory), so a search immediately after commit() can
still miss the documents just committed.

Document this explicitly on ReloadPolicy::OnCommitWithDelay,
IndexReader::reload(), and IndexWriter::commit(), and add a runnable
example showing the ReloadPolicy::Manual + explicit reader.reload()
pattern for deterministic read-your-writes.

See #1824
2026-08-26 14:46:02 +02:00
Paul MasurelandPaul Masurel 266a6c48f2 Fix numeric term key normalization across segments (#3051)
Normalize numerical values before constructing intermediate term keys so equal values merge even when their physical column types differ between segments.

Add a regression test covering u64 and i64 columns with a shared value.

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-08-19 23:09:39 +02:00
Paul MasurelandPaul Masurel fa904b3a72 Minor refactoring of find_missing_docs (#3049)
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-08-19 18:11:14 +02:00