Intermediate writes are an implementation detail. From Tantivy side we only decide when a file is finished. This change is done to remove flush overhead for VecWriter. On finalization we just move the Vec to the directory.
Rename TerminatingWrite to FinishableWrite to reflect these semantics.
Treat a fieldless exists leaf as match-all and make the lenient set parser
stop when parsing no longer consumes input. Add public QueryParser proptests
covering both regressions.
Findings and original fixes by Oleksii Syniakov in osyniakov/tantivy#4.
Co-authored-by: Oleksii Syniakov <1282756+osyniakov@users.noreply.github.com>
Add MIT license metadata to the jitexpr crate manifest, matching
tantivy and the other workspace subcrates, and include a copy of
the MIT license text.
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
The coverage job pinned nightly-2025-12-01, which is rustc 1.93. Since
`--all-features --workspace` pulls in jitexpr -> cranelift 0.134.4, and
those crates declare rust-version = "1.94.0", cargo refused to build:
error: rustc 1.93.1 is not supported by the following packages:
cranelift@0.134.4 requires rustc 1.94.0
Bump the pin to nightly-2026-08-31 (1.100.0-nightly), which builds
jitexpr cleanly.
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
Silence the 7 warnings reported by `cargo clippy --all --tests`:
- Replace manual `% n == 0` parity checks with `is_multiple_of`
(`manual_is_multiple_of`) in index_writer, doc_predicate_query and
seek_danger tests. `is_multiple_of` is stable for unsigned integers as
of 1.87, which matches the crate MSRV.
- Prefix the intentionally unused `other_segment_meta` binding with an
underscore. The binding is kept rather than dropped because the
inventory only tracks live `SegmentMeta`s, so dropping it would
invalidate the `inventory.all()` assertion in the same test.
No behavior change.
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
JITModule never releases its executable allocation on drop; the only
way to reclaim it is the consuming, unsafe JITModule::free_memory.
CompiledFn previously stored the module in a plain field, so every
compiled expression's executable memory leaked for the remainder of
the process once the CompiledFn was dropped.
Wrap the module in ManuallyDrop and add an explicit Drop impl for
CompiledFn that calls free_memory. This is safe because CompiledFn is
only ever constructed behind an Arc and never exposes its entry point
or other raw pointers into the module outside &self-bounded calls, so
drop only runs once no call into the module can be in flight or happen
afterward.
Verified empirically: 200k compile+drop cycles peaked at ~3.3 GB RSS
before this fix and ~6 MB after.
Use the default seek-based implementation until the optimized path can handle invalid child states safely. A child miss followed by another child hit previously allowed refill to consume an unaligned scorer.
Add shared seek_danger contract proptests, a regression for #3086, and a source inventory guard for future overrides.
Serialize vector-backed term entries as a map so Quickwit nodes running
different versions can exchange Postcard aggregation results. Add a legacy
fixture covering nested terms and histogram aggregations.
Use total floating-point ordering for aggregation keys so ordering is
consistent with bitwise equality for NaNs and signed zero. This also
removes fallible key comparisons from term aggregation sorting.
* Add DocPredicateQuery and a generic FunctionPredicate implementation
Introduces the DocPredicate/SegmentDocPredicate abstraction and
DocPredicateQuery, a query that walks a segment's documents by
repeatedly evaluating a per-segment predicate (with a scorer_danger
fast path for boolean intersections). FunctionPredicate is a first,
generic implementation built from a plain per-segment factory
function, with no dependency on fast fields or any other segment
data structure.
* CR comments
* Better explain
* Added the capacity to build scorer and seek_danger in a joint manner.
* Added some partial support of scorer_danger for PhraseWeight
* Bugfix following CR
* CR: introducing OccurWeights
* CR comment
---------
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
BooleanWeight::new just default min_should_match to 1.
This is misleading because I think someone who passes a single must clause would expect the result
to match that clause.
Due to min_should_match, it actually returns nothing.
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
* Make aggregation column block accessor private
Move ColumnBlockAccessor out of tantivy-columnar and into the aggregation implementation. Keep its behavior and tests intact while removing it from the public columnar API.
* Refactor aggregation value block access
* Following comment
Removing Fused from the element that are private to the module.
Re-added optimization in stats for single doc requests to avoid perf regression
Renamed Fused -> Flattened
---------
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
commit() only guarantees documents are persisted to the Directory; an
already-open IndexReader only picks up the change according to its
ReloadPolicy. With the default OnCommitWithDelay, the reload happens
asynchronously on a background thread and is not guaranteed to be done
by the time commit() returns, on any Directory implementation
(including RamDirectory), so a search immediately after commit() can
still miss the documents just committed.
Document this explicitly on ReloadPolicy::OnCommitWithDelay,
IndexReader::reload(), and IndexWriter::commit(), and add a runnable
example showing the ReloadPolicy::Manual + explicit reader.reload()
pattern for deterministic read-your-writes.
See #1824
Normalize numerical values before constructing intermediate term keys so equal values merge even when their physical column types differ between segments.
Add a regression test covering u64 and i64 columns with a shared value.
Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>