Pascal Seitz
fc84c0f299
fix REUSE_AGG_BENCH_INDEX
2026-07-27 18:00:37 +02:00
PSeitz-dd
70f0b039f5
Merge pull request #2978 from quickwit-oss/mallets/finalize-docidmapping
...
feat: add custom doc id mapping finalization
2026-07-22 15:20:19 +02:00
Luca Cominardi
2554072e8f
fix: address manual mapping review comments
...
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-07-22 14:50:29 +02:00
Luca Cominardi
fe1bdf4a54
fix: make fmt
2026-07-22 11:48:17 +02:00
Luca Cominardi
2c967e5d20
refactor: share doc id mapping construction
...
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-07-22 11:46:39 +02:00
Luca Cominardi
2350c8e645
Update src/indexer/segment_writer.rs
...
Co-authored-by: PSeitz <PSeitz@users.noreply.github.com >
2026-07-22 10:22:51 +02:00
Luca Cominardi
0a071c55f7
fix: clarify manual doc mapping serialization test
...
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-07-21 12:41:32 +02:00
Luca Cominardi
56bd23bafe
fix: move doc mapping tests to doc mapping module
...
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-07-21 12:40:08 +02:00
dependabot[bot]
31ca1a8ba2
Update lz4_flex requirement from 0.13 to 0.14
...
Updates the requirements on [lz4_flex](https://github.com/pseitz/lz4_flex ) to permit the latest version.
- [Release notes](https://github.com/pseitz/lz4_flex/releases )
- [Changelog](https://github.com/PSeitz/lz4_flex/blob/main/CHANGELOG.md )
- [Commits](https://github.com/pseitz/lz4_flex/compare/0.13.0...0.14.0 )
---
updated-dependencies:
- dependency-name: lz4_flex
dependency-version: 0.14.0
dependency-type: direct:production
...
Signed-off-by: dependabot[bot] <support@github.com >
2026-07-20 09:51:23 +02:00
trinity-1686a
057458bf14
use enum PruneMode instead of bool
2026-07-13 12:16:18 +02:00
trinity-1686a
f05ef0c4cc
cr
2026-07-13 12:16:18 +02:00
trinity-1686a
16dfddf31a
add method to prune intermediate agg results
2026-07-13 12:16:18 +02:00
Pascal Seitz
61693134be
fix cache flush in aggregations
...
fixes #2992
```
full
terms_7 Memory: 37.2 KB Avg: 2.3958ms (+0.31%) Median: 2.3896ms (+0.18%) [2.3573ms .. 2.5176ms]
terms_all_unique Memory: 10.8 MB Avg: 5.5144ms (-1.07%) Median: 5.4625ms (-1.98%) [5.3364ms .. 5.9712ms]
terms_all_unique_order_by_key Memory: 10.8 MB Avg: 5.2614ms (-0.85%) Median: 5.2177ms (-1.21%) [5.0823ms .. 5.6316ms]
terms_150_000 Memory: 2.7 MB Avg: 5.5335ms (-1.07%) Median: 5.5152ms (-1.06%) [5.4151ms .. 5.9654ms]
terms_many_top_1000 Memory: 5.2 MB Avg: 8.3579ms (-1.53%) Median: 8.3604ms (-0.95%) [8.2184ms .. 8.5421ms]
terms_many_order_by_term Memory: 2.7 MB Avg: 4.6713ms (-0.07%) Median: 4.6569ms (-0.15%) [4.5994ms .. 4.9115ms]
terms_all_unique_with_avg_sub_agg Memory: 54.0 MB Avg: 17.4981ms (-2.43%) Median: 17.6075ms (-1.75%) [15.8166ms .. 18.9250ms]
terms_status_with_avg_sub_agg Memory: 90.3 KB Avg: 5.6365ms (+7.77%) Median: 5.6255ms (+7.97%) [5.5489ms .. 5.8254ms]
terms_status_with_terms_zipf_1000_sub_agg Memory: 318.5 KB (+56.52%) Avg: 4.4504ms (+11.55%) Median: 4.4436ms (+11.59%) [4.3858ms .. 4.5692ms]
terms_zipf_1000_with_terms_status_sub_agg Memory: 684.9 KB Avg: 11.8606ms (+0.19%) Median: 11.8360ms (-0.02%) [11.7478ms .. 12.0609ms]
terms_status_with_histogram Memory: 139.5 KB Avg: 2.4524ms (-1.09%) Median: 2.4521ms (-0.23%) [2.4179ms .. 2.5049ms]
terms_status_with_date_histogram Memory: 136.7 KB Avg: 2.3407ms (-1.28%) Median: 2.3359ms (-1.06%) [2.3001ms .. 2.4310ms]
terms_status_with_date_histogram_hard_bounds Memory: 136.1 KB Avg: 2.5113ms (-2.04%) Median: 2.5073ms (-0.97%) [2.4455ms .. 2.7280ms]
terms_status_with_date_histogram_and_sibling_terms Memory: 137.3 KB Avg: 3.8695ms (-0.48%) Median: 3.8653ms (+0.11%) [3.8093ms .. 4.0528ms]
terms_zipf_1000 Memory: 69.8 KB Avg: 2.2022ms (-1.95%) Median: 2.2026ms (-1.12%) [2.1705ms .. 2.2859ms]
terms_zipf_1000_with_histogram Memory: 1.2 MB Avg: 20.4087ms (-0.02%) Median: 20.3665ms (+0.11%) [20.1912ms .. 20.7933ms]
terms_zipf_1000_with_avg_sub_agg Memory: 472.0 KB Avg: 8.7387ms (-3.48%) Median: 8.7043ms (-3.44%) [8.6466ms .. 9.1396ms]
terms_zipf_90 Memory: 55.3 KB Avg: 1.3784ms (-2.04%) Median: 1.3787ms (-1.61%) [1.3484ms .. 1.4611ms]
terms_zipf_90_with_sum_sub_agg Memory: 367.6 KB Avg: 4.8520ms (+8.43%) Median: 4.8326ms (+8.94%) [4.8058ms .. 5.1278ms]
terms_many_json_mixed_type_with_avg_sub_agg Memory: 17.8 MB Avg: 25.0853ms (-7.70%) Median: 25.0591ms (-7.12%) [24.8103ms .. 25.4936ms]
terms_status_with_cardinality_agg Memory: 91.8 KB Avg: 3.3667ms (+1.47%) Median: 3.3690ms (+1.66%) [3.3311ms .. 3.4070ms]
terms_100_buckets_with_cardinality_agg Memory: 9.9 MB Avg: 48.4768ms (-3.07%) Median: 48.3745ms (-3.38%) [48.1425ms .. 49.5503ms]
```
2026-07-12 16:40:05 +02:00
Pascal Seitz
7152d53182
clippy
2026-07-10 12:33:33 +02:00
Luca Cominardi
1c5af9489a
fix: cargo fmt
2026-07-07 15:25:35 +02:00
Luca Cominardi
6f7cc10a9b
fix: delete remap temp store after single segment finalize
...
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-07-07 15:16:24 +02:00
Pascal Seitz
6b8bd7b884
reorder if block
2026-07-03 17:44:44 +02:00
Pascal Seitz
9db05b660e
fix overflow issue
2026-07-03 17:44:44 +02:00
Pascal Seitz
057e9d6618
make BucketId optional in aggregations
...
If term aggregations don't have sub-aggregations, we don't need to carry
BucketId(u32).
```
full
terms_7 Memory: 46.5 KB Avg: 2.4024ms (+1.96%) Median: 2.4024ms (+1.96%) [2.4024ms .. 2.4024ms]
terms_all_unique Memory: 11.5 MB (-9.93%) Avg: 4.7910ms (-22.62%) Median: 4.7910ms (-22.62%) [4.7910ms .. 4.7910ms]
terms_all_unique_order_by_key Memory: 11.5 MB (-9.94%) Avg: 4.8056ms (-21.31%) Median: 4.8056ms (-21.31%) [4.8056ms .. 4.8056ms]
terms_150_000 Memory: 2.7 MB (-9.90%) Avg: 5.7312ms (-5.96%) Median: 5.7312ms (-5.96%) [5.7312ms .. 5.7312ms]
terms_many_top_1000 Memory: 5.3 MB Avg: 8.5912ms (-3.34%) Median: 8.5912ms (-3.34%) [8.5912ms .. 8.5912ms]
terms_many_order_by_term Memory: 2.7 MB (-9.95%) Avg: 4.5581ms (-10.79%) Median: 4.5581ms (-10.79%) [4.5581ms .. 4.5581ms]
terms_many_with_top_hits Memory: 48.9 MB Avg: 102.3672ms (-2.30%) Median: 102.3672ms (-2.30%) [102.3672ms .. 102.3672ms]
terms_all_unique_with_avg_sub_agg Memory: 54.7 MB Avg: 18.4644ms (-0.11%) Median: 18.4644ms (-0.11%) [18.4644ms .. 18.4644ms]
terms_many_with_avg_sub_agg Memory: 13.5 MB Avg: 16.7345ms (-4.09%) Median: 16.7345ms (-4.09%) [16.7345ms .. 16.7345ms]
terms_status_with_avg_sub_agg Memory: 92.1 KB Avg: 5.1590ms (+0.33%) Median: 5.1590ms (+0.33%) [5.1590ms .. 5.1590ms]
terms_status_with_terms_zipf_1000_sub_agg Memory: 213.2 KB Avg: 3.9562ms (+1.49%) Median: 3.9562ms (+1.49%) [3.9562ms .. 3.9562ms]
terms_zipf_1000_with_terms_status_sub_agg Memory: 724.4 KB Avg: 11.7206ms (-4.68%) Median: 11.7206ms (-4.68%) [11.7206ms .. 11.7206ms]
terms_status_with_histogram Memory: 139.0 KB Avg: 2.3985ms (-1.07%) Median: 2.3985ms (-1.07%) [2.3985ms .. 2.3985ms]
terms_status_with_date_histogram Memory: 145.9 KB Avg: 2.2963ms (-1.86%) Median: 2.2963ms (-1.86%) [2.2963ms .. 2.2963ms]
terms_status_with_date_histogram_hard_bounds Memory: 145.4 KB Avg: 2.4312ms (-1.37%) Median: 2.4312ms (-1.37%) [2.4312ms .. 2.4312ms]
terms_status_with_date_histogram_and_sibling_terms Memory: 143.3 KB Avg: 3.8401ms (+0.44%) Median: 3.8401ms (+0.44%) [3.8401ms .. 3.8401ms]
terms_zipf_1000 Memory: 75.8 KB Avg: 2.1611ms (-2.70%) Median: 2.1611ms (-2.70%) [2.1611ms .. 2.1611ms]
terms_zipf_1000_with_histogram Memory: 1.2 MB Avg: 20.1369ms (-1.16%) Median: 20.1369ms (-1.16%) [20.1369ms .. 20.1369ms]
terms_zipf_1000_with_avg_sub_agg Memory: 472.7 KB Avg: 8.5851ms (-4.93%) Median: 8.5851ms (-4.93%) [8.5851ms .. 8.5851ms]
terms_many_json_mixed_type_with_avg_sub_agg Memory: 17.9 MB Avg: 26.3988ms (-6.61%) Median: 26.3988ms (-6.61%) [26.3988ms .. 26.3988ms]
```
2026-07-03 17:44:44 +02:00
Pascal Seitz
95172089e7
cleanup benchmark
2026-07-03 13:42:33 +02:00
Pascal Seitz
f313f22df7
Speed up range-query intersections via seek_danger on RangeDocSet (up to ~50x faster)
...
A regular seek on RangeDocSet is costly: on a miss it fetches blocks and
scans the column forward to materialize the next matching doc. As a
non-leading docset in an intersection that work is wasted — the driver only
asks "does this candidate match?". seek_danger answers that with a cheap
point lookup via Column::values_for_doc, returning a lower bound on a miss
and leaving forward progress to the caller.
Forward seek_danger through ConstScorer.
Benchmarks (bool_queries_with_range, _all_results / DocSetCollector):
```
dense and 0.1% a
a_AND_num_rand:[0_TO_9]_all_results Avg: 0.0827ms (-4.60%) Median: 0.0825ms (-4.82%) [0.0809ms .. 0.0891ms] Output: 43
a_AND_num_asc:[0_TO_9]_all_results Avg: 0.1937ms (-3.70%) Median: 0.1930ms (-3.59%) [0.1806ms .. 0.2044ms] Output: 100
a_AND_num_rand_fast:[0_TO_9]_all_results Avg: 0.0367ms (-92.67%) Median: 0.0365ms (-92.65%) [0.0340ms .. 0.0398ms] Output: 43
a_AND_num_asc_fast:[0_TO_9]_all_results Avg: 0.1052ms (-98.05%) Median: 0.1050ms (-97.98%) [0.1009ms .. 0.1117ms] Output: 100
num_rand_fast:[0_TO_9]_AND_num_asc_fast:[0_TO_9]_all_results Avg: 2.7147ms (-51.42%) Median: 2.7075ms (-49.58%) [2.6806ms .. 2.7799ms] Output: 968
dense and 1% a
a_AND_num_rand:[0_TO_9]_all_results Avg: 0.4373ms (-9.71%) Median: 0.4357ms (-10.12%) [0.4117ms .. 0.4711ms] Output: 463
a_AND_num_asc:[0_TO_9]_all_results Avg: 0.2342ms (-2.50%) Median: 0.2338ms (-2.56%) [0.2247ms .. 0.2452ms] Output: 1_054
a_AND_num_rand_fast:[0_TO_9]_all_results Avg: 0.3956ms (-82.86%) Median: 0.3943ms (-82.90%) [0.3815ms .. 0.4119ms] Output: 463
a_AND_num_asc_fast:[0_TO_9]_all_results Avg: 0.4896ms (-91.16%) Median: 0.4862ms (-90.81%) [0.4797ms .. 0.5084ms] Output: 1_054
num_rand_fast:[0_TO_9]_AND_num_asc_fast:[0_TO_9]_all_results Avg: 2.7108ms (-50.81%) Median: 2.6925ms (-49.51%) [2.6688ms .. 2.7868ms] Output: 968
dense and 10% a
a_AND_num_rand:[0_TO_9]_all_results Avg: 0.9869ms (-3.71%) Median: 0.9833ms (-3.83%) [0.9518ms .. 1.1218ms] Output: 4_914
a_AND_num_asc:[0_TO_9]_all_results Avg: 0.6352ms (-3.74%) Median: 0.6363ms (-3.32%) [0.6158ms .. 0.6488ms] Output: 10_152
a_AND_num_rand_fast:[0_TO_9]_all_results Avg: 3.1264ms (+0.39%) Median: 3.1466ms (+1.34%) [3.0261ms .. 3.2051ms] Output: 4_914
a_AND_num_asc_fast:[0_TO_9]_all_results Avg: 4.1547ms (-31.12%) Median: 4.0933ms (-28.55%) [3.7648ms .. 4.7600ms] Output: 10_152
num_rand_fast:[0_TO_9]_AND_num_asc_fast:[0_TO_9]_all_results Avg: 2.6973ms (-52.30%) Median: 2.6901ms (-49.86%) [2.6689ms .. 2.7677ms] Output: 968
```
Gains are largest when the range query is the non-leading docset of a low-cardinality intersection.
2026-07-03 13:42:33 +02:00
Pascal Seitz
82bee54a00
block_search: drop unsafe indexing, remove K=64
...
For K∈{2,4,8,16,32} LLVM proves the index bounds and elides the checks,
so get_unchecked buys nothing on the production K=8 path. K=64 was the
only value that defeated bounds-check elision (one check in the tail
scan) and it was instantiated in tests only — drop it.
block_search: cite the k-ary search paper
Document that kary_search is the 'k-ary search on a sorted array' variant
from Schlegel, Gemulla & Lehner (DaMoN 2009), specialized to a lower-bound.
2026-07-02 18:27:40 +02:00
Pascal Seitz
6892995d02
10% faster intersections: Use k-ary in block search
2026-07-02 18:27:40 +02:00
trinity-1686a
9704e6c0e3
Merge pull request #2983 from quickwit-oss/trinity.pointard/quickselect-term-agg
...
use select-nth instead of full sort in segment level agg top-k selection
2026-07-02 15:38:14 +02:00
trinity.pointard
715590b357
rename local var
2026-07-02 12:00:00 +00:00
trinity.pointard
d496e402ca
rustfmt
2026-07-02 07:12:04 +00:00
trinity.pointard
348ca1e309
don't count matching doc twice
2026-06-30 16:09:11 +00:00
trinity.pointard
5e4fe3520c
better handle sorted buckets
2026-06-30 14:56:24 +00:00
Luca Cominardi
b7234c153e
fix: clean up manual mapping finalization invariants
...
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-30 13:12:23 +02:00
Luca Cominardi
e3be28814e
fix: reject default finalize with manual doc id mapping
...
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-30 10:56:46 +02:00
pascal
a9733ba8c2
Keep buffered union refill out of line
...
BufferedUnionScorer is the hot path for full union traversal, including (TopDocs, Count) where Count forces all matches to be visited. After the block-wand intersection changes, LLVM started inlining the refill helper into the advance path, which regressed TOP_100_COUNT union queries even though the union algorithm did not change.
Force the refill helper out of line so the advance loop stays small and stable while pruning collectors continue to use Block-WAND.
Benchmark on search-benchmark-game TOP_100_COUNT union query set (301 queries, sum of per-query medians):
- tantivy 0.26: 0.853646s
- main before: 0.918605s
- this change: 0.841659s
2026-06-29 19:33:50 +02:00
pascal
874d54a63a
Remove union wrapping for single-terms
...
search-benchmark-game shows TOP_100_COUNT regression on queries tagged intersection_union.
The regression came from allowing single-term boolean unions to become TermUnion for Block-WAND. https://github.com/quickwit-oss/tantivy/pull/2915
When such a scorer is used as the optional side of RequiredOptionalScorer, boxing converted the lone term into BufferedUnionScorer.
Keep the TermUnion representation available for pruning collection, but unwrap one-term unions when boxing or doing non-pruning iteration.
2026-06-29 19:33:50 +02:00
Luca Cominardi
3c6eb92c7e
fix: restore clone bound on doc id iterator
...
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-29 11:31:38 +02:00
Luca Cominardi
9f73d3fe54
docs: clarify manual doc id mapping serialization
...
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-29 11:15:53 +02:00
Luca Cominardi
50272240c6
test: move doc id mapping validation coverage
...
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-29 11:13:35 +02:00
trinity.pointard
74a510cb56
try to use select-nth instead of full sort in segment level agg top-k selection
2026-06-29 09:13:21 +00:00
Luca Cominardi
a6fd070e3d
fix: validate manual doc id mappings at construction
...
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-29 10:52:41 +02:00
Luca Cominardi
06c046bdc9
fix: enable manual doc id mapping in single segment test
...
The regression test calls finalize_with_doc_id_mapping, so the index must opt into the temporary docstore path before constructing the segment writer.
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-26 17:56:54 +02:00
Luca Cominardi
63f65dcd71
fix: add missing manual_doc_id_mapping field in zstd test IndexSettings initializer
...
Struct literal was missing the new field introduced in the manual doc id
mapping refactor, causing a compile error under --all-features.
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-26 17:47:33 +02:00
Luca Cominardi
8a4c5b9013
refactor: gate manual doc id mapping via settings
...
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-26 17:37:59 +02:00
Luca Cominardi
f7355e60cd
feat: add single segment doc id mapping finalization
...
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-26 16:47:14 +02:00
Luca Cominardi
90603d2396
refactor custom doc id mapping finalization
...
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-26 16:24:19 +02:00
Luca Cominardi
910861a3e9
feat: add custom doc id mapping finalization
...
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-06-26 14:45:57 +02:00
trinity-1686a
02e34508e2
Merge pull request #2971 from quickwit-oss/trinity.pointard/fix-slop-overflow
...
fix overflow on large jumps in linear sequence
2026-06-23 10:18:29 +02:00
trinity-1686a
4031d97bac
fix overflow on large jumps in linear sequence
...
new limit prevent an overflow in eval which caused the residual to be 64b when a slop of zero would give a smaller one
2026-06-23 00:13:27 +02:00
Ming
384f31d350
feat: Restore index sorting ( #2959 )
...
We ([ParadeDB](https://github.com/paradedb/paradedb )) have restored and been using the removed [index sorting](https://github.com/quickwit-oss/tantivy/issues/2352 ) feature in our Tantivy fork.
Our use case is sorting the index by Postgres' internal `ctid` identifier. Results returned from Tantivy must be checked against Postgres' visibility map, and checking them in ctid order is much more cache friendly, resulting in up to 80% speedups for certain queries.
This PR is split into 5 commits, corresponding to the index sorting reversal plus bug fixes we uncovered during our usage of index sorting.
| Commit | Maps to | What it does |
|---|---|---|
| `2aea0ad9f` | foundation ([#104 ](https://github.com/paradedb/tantivy/pull/104 )) | Restore `SegmentComponent::TempStore` (revert of upstream #2815 ). Subsumes fork PR [#104 ](https://github.com/paradedb/tantivy/pull/104 )'s CI fix. |
| `9205bcb0c` | [#92 ](https://github.com/paradedb/tantivy/pull/92 ) | Restore sort-by-field (single-segment + merge paths). |
| `39c790f0f` | [#101 ](https://github.com/paradedb/tantivy/pull/101 ) | Enable `sort_by` for `Str`/`Bytes` fast fields. |
| `9c4341a87` | [#105 ](https://github.com/paradedb/tantivy/pull/105 ) | Native typed numeric sort-key comparison (precision/NULL fix). |
| `2d9ba2418` | [#106 ](https://github.com/paradedb/tantivy/pull/106 ) | Preserve NULL ordering in numeric segment merges. |
We have discussed with the Tantivy maintainers and they indicated they would be open to this PR. Another motivation for landing this PR is we are planning on contributing a significant refactor that makes Tantivy's segment components extensible, and landing that without index sorting leads to too many conflicts.
2026-06-22 11:22:25 -07:00
Pascal Seitz
1e859fd78d
fix term aggregation u32::MAX overflow issue
2026-06-18 17:07:43 +08:00
Pascal Seitz
f451fa938f
explain why naive scorer must accumulate scores in WAND order
2026-06-17 18:58:58 +08:00
Pascal Seitz
2a82dd6f64
fix flaky test
2026-06-17 18:58:58 +08:00
Pascal Seitz
c096b2ad89
aggregation/terms: charge fused term_counts to the memory limit
...
term_counts (one u32/term) was allocated but not charged to
AggregationLimitsGuard, so a memory limit could be exceeded silently.
Charge it, skip allocating it when unbounded, and add a regression test.
2026-06-16 21:23:23 +08:00