Commit Graph
3769 Commits
Author SHA1 Message Date
dependabot[bot] e99f3e58a5 Bump github/codeql-action/upload-sarif from 4.36.1 to 4.38.2
Bumps [github/codeql-action/upload-sarif](https://github.com/github/codeql-action) from 4.36.1 to 4.38.2.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/87557b9c84dde89fdd9b10e88954ac2f4248e463...2892aa5e19bbd11bc0cff5427e3b750a04d9e3c2)

---
updated-dependencies:
- dependency-name: github/codeql-action/upload-sarif
  dependency-version: 4.38.2
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-09-28 20:03:25 +00:00
Paul MasurelandPaul Masurel 047464cf92 Accelerating jitexpr conditions using necessary conditions (#3129)
* jitexpr

* Added possible necessary conditions to jitexpr.

A function makes it possible to infer a necessary query from an expression to match.
We can then accelerate queries involving a calculated field by not even evaluating the expression
on docs that trivially do not match.

* CR comment

* CR comment

* Fixing unit tests

---------

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-28 18:46:37 +02:00
palmoni5 b9125aad55 Compile RegexPhraseQuery regexes once per weight (#3135)
RegexPhraseWeight::phrase_scorer compiled every phrase term's regex again for each segment. Determinizing a wide pattern (e.g. an alternation of typo or morphology expansions) costs milliseconds to tens of milliseconds, so on a multi-segment index the query was dominated by recompiling the same automata. The regexes are now compiled in RegexPhraseQuery::regex_phrase_weight and shared through Arc.

An invalid pattern is now reported when the weight is built, instead of when the first segment is scored.
2026-09-28 11:37:50 +02:00
David Yaffe 6182c6062f Merge pull request #3084 from quickwit-oss/dependabot/cargo/datasketches-0.5.0
Update datasketches requirement from 0.3.0 to 0.5.0
2026-09-23 15:53:13 -04:00
dayaffe 53e90ab192 Merge main into datasketches 0.5 update 2026-09-23 18:39:37 +00:00
Pascal Seitz 08d5836b8f Format with nightly rustfmt 2026-09-23 20:21:08 +08:00
Pascal Seitz c321b170a8 Clarify chunk size and packed value type in bit unpacking fast path 2026-09-23 20:21:08 +08:00
Pascal Seitz 6023163e37 Extract bit unpacking load helper with explicit safety contract
Replace the closure with an inline(always) function taking the data
slice and bit address, making its inputs and safety requirements explicit.
2026-09-23 20:21:08 +08:00
PSeitzandPaul Masurel f5fb40950a Update bitpacker/src/bitpacker.rs
Co-authored-by: Paul Masurel <paul@quickwit.io>
2026-09-23 20:21:08 +08:00
Pascal Seitz 21a6913217 Simplify bitpacked range decoding
Use scalar decoding for ranges overlapping the final partial load.
Clarify the 64-value aggregation block specialization and name the
generic decoding chunk size.

Apply nightly formatting to the touched benchmark and columnar code.
2026-09-23 20:21:08 +08:00
Pascal Seitz 4ede8d64f2 Optimize low-bit-width block decoding
Decode eight values per load for 64-value blocks and specialize bit
widths 1 through 7 to enable constant folding in the hot path.
2026-09-23 20:21:08 +08:00
Pascal Seitz 4a660f59d2 Faster columnar data fetching 2026-09-23 20:21:08 +08:00
Pascal Seitz be8a1d0986 Add 26-bit date histogram aggregation benchmark 2026-09-23 20:21:08 +08:00
Pascal Seitz e229de6db7 Collect decoded multi-term keys before assigning bucket positions 2026-09-22 19:53:52 +08:00
Pascal Seitz 7d3a3d6bfe Use unwrap_or(false) for missing bucket value check 2026-09-22 19:53:52 +08:00
Pascal Seitz 62f3d4b4dd Apply multi-terms cutoff before final bucket conversion
Avoid formatting keys and finalizing sub-aggregations for buckets
discarded by the final size limit. Preallocate the retained bucket vector
and reuse the cutoff helper with intermediate tuple keys.

Preserve final-key ordering and finalize the ordering metric when sorting
by a sub-aggregation. Cover mixed key types, empty sums, and discarded
document counts in tests.
2026-09-22 19:53:52 +08:00
Pascal Seitz 3655e6f8cb Batch multi-terms dictionary lookups after pruning
Resolve retained string ordinals in sorted batches per field instead
of decoding dictionary blocks separately for each tuple component.
Reuse decoding for repeated ordinals while preserving tuple order
and existing numeric and missing-value handling.
2026-09-22 19:53:52 +08:00
Pascal Seitz 085a8327d3 Merge incoming aggregation buckets into the accumulator
Probe the accumulated map only for incoming keys rather than rehashing every accumulated key on each merge. This removes quadratic work when folding many disjoint multi-terms results. The same helper serves range and composite buckets; existing bucket values still merge left-to-right and the wire representation and pruning rules are unchanged.

Cover overlapping/disjoint tuple keys, empty inputs, recursive range subaggregations, postcard round trips, error/count bookkeeping, and merging after pruning.

Same benchmark and configuration as 0062d2f0d (Apple M4 Max, rustc 1.98.0):
Median milliseconds for 10 / 100 / 1000 inputs:
shared 8:     0.006365 / 0.0656 / 0.6449
disjoint 16:  0.0126   / 0.1486 / 1.6882
disjoint 160: 0.1336   / 1.5558 / 20.4019
At 1000 inputs this is 52.4x and 61.6x faster for the disjoint cases, with unchanged measured peak allocation. Shared-key control remains within approximately 2% of baseline.

Validation: 308 aggregation tests pass with default features and 308 with quickwit; changed-file rustfmt and git diff checks pass. Clippy --lib --bench agg_bench passes with the pre-existing clippy::drop_non_drop warning allowed (unchanged drop(add_document) in the benchmark).
2026-09-22 19:53:52 +08:00
Pascal Seitz 0f512fc3e2 Benchmark multi-terms aggregations across many segments
Run the existing many-segment aggregation benchmarks at 100 and 1,000
segments with one million total documents. Add multi-terms and nested
terms cases for status/Zipf and high-cardinality/Zipf combinations,
including top-500 requests, to expose merge costs across many segments.

The nested top-500 case limits outer buckets, while multi-terms limits
tuples globally.

Run both groups with: cargo bench --bench agg_bench -- _segments
2026-09-22 19:53:52 +08:00
Paul MasurelandPaul Masurel 507e1760aa Minor refactor of the cardinality aggregation. (#3120)
The Str path and the non string path are very different.
This PR isolates them as different segment aggregation collector.

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-22 10:47:49 +02:00
Paul Masurel b91a405ee9 Make aggregation request-data types crate-internal (#3110)
Hiding the per segment *AggReqData members

FilterAggReqData::reqandMultiTermsAggReqData::sub_aggregations` were dead code so they were removed.
2026-09-18 07:27:01 +02:00
Paul MasurelandPaul Masurel 20d7f72f2c Addressing -0.0 problem in SafeF64 (#3109)
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-17 12:29:31 +02:00
Paul Masurel 1496f58963 Added jitexpr compilation cache (#3107)
* Added jitexpr compilation cache
* Added jitexpr to the CI test matrix
2026-09-17 10:34:15 +02:00
David Yaffe e69acca220 Fix cardinality calls for datasketches 0.5 2026-09-16 14:55:28 -04:00
Pascal Seitz 1680885e68 Report active RamDirectory writers as existing 2026-09-16 23:39:47 +08:00
Pascal Seitz e56587f3df Rename memory usage tracker finish to release 2026-09-16 23:39:47 +08:00
Pascal Seitz dfe904667e Minimize unrelated changes 2026-09-16 23:39:47 +08:00
Pascal Seitz a9d3f1a026 Run nightly fmt 2026-09-16 23:39:47 +08:00
Pascal Seitz bbc1b9e1d1 Restore TerminatingWrite API 2026-09-16 23:39:47 +08:00
Pascal Seitz 1c8a4801b3 Shrink RamDirectory files before storing 2026-09-16 23:39:47 +08:00
Pascal Seitz 2ab4fb80f0 Buffer RamDirectory writes
Use the standard BufWriter capacity to avoid forwarding every small write
directly to the underlying VecWriter.
2026-09-16 23:39:47 +08:00
Pascal Seitz e46d44c388 Finalize mmap writers in directory tests
Use finish instead of flush before reading completed files and clarify
when file data is synchronized.
2026-09-16 23:39:47 +08:00
Pascal Seitz f2a106df1c Prevent concurrent RamDirectory writers
Reserve paths when writers are opened and release reservations when
writers finish or drop.
2026-09-16 23:39:47 +08:00
Pascal Seitz 98d91f0436 Remove intermediate flush from directory contract
Intermediate writes are an implementation detail. From Tantivy side we only decide when a file is finished. This change is done to remove flush overhead for VecWriter. On finalization we just move the Vec to the directory.

Rename TerminatingWrite to FinishableWrite to reflect these semantics.
2026-09-16 23:39:47 +08:00
Pascal Seitz b06d8e9a04 Track active RamDirectory writer memory
Include VecWriter allocations in total_mem_usage so callers can account
for memory before writers are flushed or terminated.
2026-09-16 23:39:47 +08:00
Pascal Seitz 54b82e582f Report unsupported whitespace as a syntax error 2026-09-16 22:53:35 +08:00
Pascal SeitzandOleksii Syniakov 34edbf7821 Fix query parser panics on malformed input
Treat a fieldless exists leaf as match-all and make the lenient set parser
stop when parsing no longer consumes input. Add public QueryParser proptests
covering both regressions.

Findings and original fixes by Oleksii Syniakov in osyniakov/tantivy#4.

Co-authored-by: Oleksii Syniakov <1282756+osyniakov@users.noreply.github.com>
2026-09-16 22:53:35 +08:00
Paul Masurel 94a116f103 (Calculated fields) Add feature-gated JIT expression document predicate (#3081) 2026-09-15 16:28:53 +02:00
Paul MasurelandPaul Masurel 99043285f9 Set jitexpr license to MIT (#3101)
Add MIT license metadata to the jitexpr crate manifest, matching
tantivy and the other workspace subcrates, and include a copy of
the MIT license text.

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-14 15:08:55 +02:00
Paul MasurelandPaul Masurel 4abd98fa19 Bump coverage nightly pin to satisfy cranelift MSRV (#3094)
The coverage job pinned nightly-2025-12-01, which is rustc 1.93. Since
`--all-features --workspace` pulls in jitexpr -> cranelift 0.134.4, and
those crates declare rust-version = "1.94.0", cargo refused to build:

    error: rustc 1.93.1 is not supported by the following packages:
      cranelift@0.134.4 requires rustc 1.94.0

Bump the pin to nightly-2026-08-31 (1.100.0-nightly), which builds
jitexpr cleanly.

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-11 09:32:00 +02:00
Paul MasurelandPaul Masurel 3bad51518a Fix clippy warnings in test code (#3091)
Silence the 7 warnings reported by `cargo clippy --all --tests`:

- Replace manual `% n == 0` parity checks with `is_multiple_of`
  (`manual_is_multiple_of`) in index_writer, doc_predicate_query and
  seek_danger tests. `is_multiple_of` is stable for unsigned integers as
  of 1.87, which matches the crate MSRV.
- Prefix the intentionally unused `other_segment_meta` binding with an
  underscore. The binding is kept rather than dropped because the
  inventory only tracks live `SegmentMeta`s, so dropping it would
  invalidate the `inventory.all()` assertion in the same test.

No behavior change.

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-10 21:26:45 +02:00
Paul Masurel 0c0530eeb4 Free JITModule executable memory when CompiledFn is dropped (#3090)
JITModule never releases its executable allocation on drop; the only
way to reclaim it is the consuming, unsafe JITModule::free_memory.
CompiledFn previously stored the module in a plain field, so every
compiled expression's executable memory leaked for the remainder of
the process once the CompiledFn was dropped.

Wrap the module in ManuallyDrop and add an explicit Drop impl for
CompiledFn that calls free_memory. This is safe because CompiledFn is
only ever constructed behind an Arc and never exposes its entry point
or other raw pointers into the module outside &self-bounded calls, so
drop only runs once no call into the module can be in flight or happen
afterward.

Verified empirically: 200k compile+drop cycles peaked at ~3.3 GB RSS
before this fix and ~6 MB after.
2026-09-10 21:22:01 +02:00
dependabot[bot] dd0bc23360 Update zstd requirement from 0.13 to 0.14 (#3085)
Updates the requirements on [zstd](https://github.com/gyscos/zstd-rs) to permit the latest version.
- [Release notes](https://github.com/gyscos/zstd-rs/releases)
- [Commits](https://github.com/gyscos/zstd-rs/compare/v0.13.0...v0.14.0)

---
updated-dependencies:
- dependency-name: zstd
  dependency-version: 0.14.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-09-10 18:36:24 +02:00
Paul Masurel 5a82f19f5d (Calculated fields) Implements more functions (#3083)
* Add remaining expression functions

Currently experimental (!)
2026-09-10 18:35:51 +02:00
Paul Masurel f8ec38fd9f (Calculated fields) Add JIT expression framework with ADD, IS_NULL, and REGEXP_EXTRACT (#3078)
Add JIT expression framework with ADD, IS_NULL, and REGEXP_EXTRACT
2026-09-10 15:30:02 +02:00
Pascal Seitz bc8a21cf12 Document 0.26.2 bugfixes and mark 0.26 as released 2026-09-08 23:36:16 +08:00
PSeitz-dd 9d10bf1c6c Merge pull request #3088 from PSeitz/check_seek_danger_panic
fix: disable buffered union seek_danger override
2026-09-08 17:05:24 +02:00
Pascal Seitz 66ff8466d5 fix: disable buffered union seek_danger override
Use the default seek-based implementation until the optimized path can handle invalid child states safely. A child miss followed by another child hit previously allowed refill to consume an unaligned scorer.

Add shared seek_danger contract proptests, a regression for #3086, and a source inventory guard for future overrides.
2026-09-08 16:34:05 +02:00
pascal 780395c1db improve BufferedUnionScorer::advance 2026-09-08 18:12:08 +08:00
pascal 6929bae13f fix union performance regression 2026-09-08 18:12:08 +08:00