Commit Graph

3622 Commits

Author SHA1 Message Date
Pascal Seitz fc84c0f299 fix REUSE_AGG_BENCH_INDEX 2026-07-27 18:00:37 +02:00
PSeitz-dd 70f0b039f5 Merge pull request #2978 from quickwit-oss/mallets/finalize-docidmapping
feat: add custom doc id mapping finalization
2026-07-22 15:20:19 +02:00
Luca Cominardi 2554072e8f fix: address manual mapping review comments
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-22 14:50:29 +02:00
Luca Cominardi fe1bdf4a54 fix: make fmt 2026-07-22 11:48:17 +02:00
Luca Cominardi 2c967e5d20 refactor: share doc id mapping construction
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-22 11:46:39 +02:00
Luca Cominardi 2350c8e645 Update src/indexer/segment_writer.rs
Co-authored-by: PSeitz <PSeitz@users.noreply.github.com>
2026-07-22 10:22:51 +02:00
Luca Cominardi 0a071c55f7 fix: clarify manual doc mapping serialization test
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-21 12:41:32 +02:00
Luca Cominardi 56bd23bafe fix: move doc mapping tests to doc mapping module
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-21 12:40:08 +02:00
dependabot[bot] 31ca1a8ba2 Update lz4_flex requirement from 0.13 to 0.14
Updates the requirements on [lz4_flex](https://github.com/pseitz/lz4_flex) to permit the latest version.
- [Release notes](https://github.com/pseitz/lz4_flex/releases)
- [Changelog](https://github.com/PSeitz/lz4_flex/blob/main/CHANGELOG.md)
- [Commits](https://github.com/pseitz/lz4_flex/compare/0.13.0...0.14.0)

---
updated-dependencies:
- dependency-name: lz4_flex
  dependency-version: 0.14.0
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-07-20 09:51:23 +02:00
trinity-1686a 057458bf14 use enum PruneMode instead of bool 2026-07-13 12:16:18 +02:00
trinity-1686a f05ef0c4cc cr 2026-07-13 12:16:18 +02:00
trinity-1686a 16dfddf31a add method to prune intermediate agg results 2026-07-13 12:16:18 +02:00
Pascal Seitz 61693134be fix cache flush in aggregations
fixes #2992

```
full
terms_7                                               Memory: 37.2 KB               Avg: 2.3958ms (+0.31%)     Median: 2.3896ms (+0.18%)     [2.3573ms .. 2.5176ms]
terms_all_unique                                      Memory: 10.8 MB               Avg: 5.5144ms (-1.07%)     Median: 5.4625ms (-1.98%)     [5.3364ms .. 5.9712ms]
terms_all_unique_order_by_key                         Memory: 10.8 MB               Avg: 5.2614ms (-0.85%)     Median: 5.2177ms (-1.21%)     [5.0823ms .. 5.6316ms]
terms_150_000                                         Memory: 2.7 MB                Avg: 5.5335ms (-1.07%)     Median: 5.5152ms (-1.06%)     [5.4151ms .. 5.9654ms]
terms_many_top_1000                                   Memory: 5.2 MB                Avg: 8.3579ms (-1.53%)     Median: 8.3604ms (-0.95%)     [8.2184ms .. 8.5421ms]
terms_many_order_by_term                              Memory: 2.7 MB                Avg: 4.6713ms (-0.07%)     Median: 4.6569ms (-0.15%)     [4.5994ms .. 4.9115ms]
terms_all_unique_with_avg_sub_agg                     Memory: 54.0 MB               Avg: 17.4981ms (-2.43%)    Median: 17.6075ms (-1.75%)    [15.8166ms .. 18.9250ms]
terms_status_with_avg_sub_agg                         Memory: 90.3 KB               Avg: 5.6365ms (+7.77%)     Median: 5.6255ms (+7.97%)     [5.5489ms .. 5.8254ms]
terms_status_with_terms_zipf_1000_sub_agg             Memory: 318.5 KB (+56.52%)    Avg: 4.4504ms (+11.55%)    Median: 4.4436ms (+11.59%)    [4.3858ms .. 4.5692ms]
terms_zipf_1000_with_terms_status_sub_agg             Memory: 684.9 KB              Avg: 11.8606ms (+0.19%)    Median: 11.8360ms (-0.02%)    [11.7478ms .. 12.0609ms]
terms_status_with_histogram                           Memory: 139.5 KB              Avg: 2.4524ms (-1.09%)     Median: 2.4521ms (-0.23%)     [2.4179ms .. 2.5049ms]
terms_status_with_date_histogram                      Memory: 136.7 KB              Avg: 2.3407ms (-1.28%)     Median: 2.3359ms (-1.06%)     [2.3001ms .. 2.4310ms]
terms_status_with_date_histogram_hard_bounds          Memory: 136.1 KB              Avg: 2.5113ms (-2.04%)     Median: 2.5073ms (-0.97%)     [2.4455ms .. 2.7280ms]
terms_status_with_date_histogram_and_sibling_terms    Memory: 137.3 KB              Avg: 3.8695ms (-0.48%)     Median: 3.8653ms (+0.11%)     [3.8093ms .. 4.0528ms]
terms_zipf_1000                                       Memory: 69.8 KB               Avg: 2.2022ms (-1.95%)     Median: 2.2026ms (-1.12%)     [2.1705ms .. 2.2859ms]
terms_zipf_1000_with_histogram                        Memory: 1.2 MB                Avg: 20.4087ms (-0.02%)    Median: 20.3665ms (+0.11%)    [20.1912ms .. 20.7933ms]
terms_zipf_1000_with_avg_sub_agg                      Memory: 472.0 KB              Avg: 8.7387ms (-3.48%)     Median: 8.7043ms (-3.44%)     [8.6466ms .. 9.1396ms]
terms_zipf_90                                         Memory: 55.3 KB               Avg: 1.3784ms (-2.04%)     Median: 1.3787ms (-1.61%)     [1.3484ms .. 1.4611ms]
terms_zipf_90_with_sum_sub_agg                        Memory: 367.6 KB              Avg: 4.8520ms (+8.43%)     Median: 4.8326ms (+8.94%)     [4.8058ms .. 5.1278ms]
terms_many_json_mixed_type_with_avg_sub_agg           Memory: 17.8 MB               Avg: 25.0853ms (-7.70%)    Median: 25.0591ms (-7.12%)    [24.8103ms .. 25.4936ms]
terms_status_with_cardinality_agg                     Memory: 91.8 KB               Avg: 3.3667ms (+1.47%)     Median: 3.3690ms (+1.66%)     [3.3311ms .. 3.4070ms]
terms_100_buckets_with_cardinality_agg                Memory: 9.9 MB                Avg: 48.4768ms (-3.07%)    Median: 48.3745ms (-3.38%)    [48.1425ms .. 49.5503ms]
```
2026-07-12 16:40:05 +02:00
Pascal Seitz 7152d53182 clippy 2026-07-10 12:33:33 +02:00
Luca Cominardi 1c5af9489a fix: cargo fmt 2026-07-07 15:25:35 +02:00
Luca Cominardi 6f7cc10a9b fix: delete remap temp store after single segment finalize
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-07 15:16:24 +02:00
Pascal Seitz 6b8bd7b884 reorder if block 2026-07-03 17:44:44 +02:00
Pascal Seitz 9db05b660e fix overflow issue 2026-07-03 17:44:44 +02:00
Pascal Seitz 057e9d6618 make BucketId optional in aggregations
If term aggregations don't have sub-aggregations, we don't need to carry
BucketId(u32).

```
full
terms_7                                               Memory: 46.5 KB             Avg: 2.4024ms (+1.96%)      Median: 2.4024ms (+1.96%)      [2.4024ms .. 2.4024ms]
terms_all_unique                                      Memory: 11.5 MB (-9.93%)    Avg: 4.7910ms (-22.62%)     Median: 4.7910ms (-22.62%)     [4.7910ms .. 4.7910ms]
terms_all_unique_order_by_key                         Memory: 11.5 MB (-9.94%)    Avg: 4.8056ms (-21.31%)     Median: 4.8056ms (-21.31%)     [4.8056ms .. 4.8056ms]
terms_150_000                                         Memory: 2.7 MB (-9.90%)     Avg: 5.7312ms (-5.96%)      Median: 5.7312ms (-5.96%)      [5.7312ms .. 5.7312ms]
terms_many_top_1000                                   Memory: 5.3 MB              Avg: 8.5912ms (-3.34%)      Median: 8.5912ms (-3.34%)      [8.5912ms .. 8.5912ms]
terms_many_order_by_term                              Memory: 2.7 MB (-9.95%)     Avg: 4.5581ms (-10.79%)     Median: 4.5581ms (-10.79%)     [4.5581ms .. 4.5581ms]
terms_many_with_top_hits                              Memory: 48.9 MB             Avg: 102.3672ms (-2.30%)    Median: 102.3672ms (-2.30%)    [102.3672ms .. 102.3672ms]
terms_all_unique_with_avg_sub_agg                     Memory: 54.7 MB             Avg: 18.4644ms (-0.11%)     Median: 18.4644ms (-0.11%)     [18.4644ms .. 18.4644ms]
terms_many_with_avg_sub_agg                           Memory: 13.5 MB             Avg: 16.7345ms (-4.09%)     Median: 16.7345ms (-4.09%)     [16.7345ms .. 16.7345ms]
terms_status_with_avg_sub_agg                         Memory: 92.1 KB             Avg: 5.1590ms (+0.33%)      Median: 5.1590ms (+0.33%)      [5.1590ms .. 5.1590ms]
terms_status_with_terms_zipf_1000_sub_agg             Memory: 213.2 KB            Avg: 3.9562ms (+1.49%)      Median: 3.9562ms (+1.49%)      [3.9562ms .. 3.9562ms]
terms_zipf_1000_with_terms_status_sub_agg             Memory: 724.4 KB            Avg: 11.7206ms (-4.68%)     Median: 11.7206ms (-4.68%)     [11.7206ms .. 11.7206ms]
terms_status_with_histogram                           Memory: 139.0 KB            Avg: 2.3985ms (-1.07%)      Median: 2.3985ms (-1.07%)      [2.3985ms .. 2.3985ms]
terms_status_with_date_histogram                      Memory: 145.9 KB            Avg: 2.2963ms (-1.86%)      Median: 2.2963ms (-1.86%)      [2.2963ms .. 2.2963ms]
terms_status_with_date_histogram_hard_bounds          Memory: 145.4 KB            Avg: 2.4312ms (-1.37%)      Median: 2.4312ms (-1.37%)      [2.4312ms .. 2.4312ms]
terms_status_with_date_histogram_and_sibling_terms    Memory: 143.3 KB            Avg: 3.8401ms (+0.44%)      Median: 3.8401ms (+0.44%)      [3.8401ms .. 3.8401ms]
terms_zipf_1000                                       Memory: 75.8 KB             Avg: 2.1611ms (-2.70%)      Median: 2.1611ms (-2.70%)      [2.1611ms .. 2.1611ms]
terms_zipf_1000_with_histogram                        Memory: 1.2 MB              Avg: 20.1369ms (-1.16%)     Median: 20.1369ms (-1.16%)     [20.1369ms .. 20.1369ms]
terms_zipf_1000_with_avg_sub_agg                      Memory: 472.7 KB            Avg: 8.5851ms (-4.93%)      Median: 8.5851ms (-4.93%)      [8.5851ms .. 8.5851ms]
terms_many_json_mixed_type_with_avg_sub_agg           Memory: 17.9 MB             Avg: 26.3988ms (-6.61%)     Median: 26.3988ms (-6.61%)     [26.3988ms .. 26.3988ms]
```
2026-07-03 17:44:44 +02:00
Pascal Seitz 95172089e7 cleanup benchmark 2026-07-03 13:42:33 +02:00
Pascal Seitz f313f22df7 Speed up range-query intersections via seek_danger on RangeDocSet (up to ~50x faster)
A regular seek on RangeDocSet is costly: on a miss it fetches blocks and
scans the column forward to materialize the next matching doc. As a
non-leading docset in an intersection that work is wasted — the driver only
asks "does this candidate match?". seek_danger answers that with a cheap
point lookup via Column::values_for_doc, returning a lower bound on a miss
and leaving forward progress to the caller.

Forward seek_danger through ConstScorer.

Benchmarks (bool_queries_with_range, _all_results / DocSetCollector):

```
dense and 0.1% a
a_AND_num_rand:[0_TO_9]_all_results                                 Avg: 0.0827ms (-4.60%)     Median: 0.0825ms (-4.82%)     [0.0809ms .. 0.0891ms]    Output: 43
a_AND_num_asc:[0_TO_9]_all_results                                  Avg: 0.1937ms (-3.70%)     Median: 0.1930ms (-3.59%)     [0.1806ms .. 0.2044ms]    Output: 100
a_AND_num_rand_fast:[0_TO_9]_all_results                            Avg: 0.0367ms (-92.67%)    Median: 0.0365ms (-92.65%)    [0.0340ms .. 0.0398ms]    Output: 43
a_AND_num_asc_fast:[0_TO_9]_all_results                             Avg: 0.1052ms (-98.05%)    Median: 0.1050ms (-97.98%)    [0.1009ms .. 0.1117ms]    Output: 100
num_rand_fast:[0_TO_9]_AND_num_asc_fast:[0_TO_9]_all_results        Avg: 2.7147ms (-51.42%)    Median: 2.7075ms (-49.58%)    [2.6806ms .. 2.7799ms]    Output: 968
dense and 1% a
a_AND_num_rand:[0_TO_9]_all_results                                 Avg: 0.4373ms (-9.71%)     Median: 0.4357ms (-10.12%)    [0.4117ms .. 0.4711ms]    Output: 463
a_AND_num_asc:[0_TO_9]_all_results                                  Avg: 0.2342ms (-2.50%)     Median: 0.2338ms (-2.56%)     [0.2247ms .. 0.2452ms]    Output: 1_054
a_AND_num_rand_fast:[0_TO_9]_all_results                            Avg: 0.3956ms (-82.86%)    Median: 0.3943ms (-82.90%)    [0.3815ms .. 0.4119ms]    Output: 463
a_AND_num_asc_fast:[0_TO_9]_all_results                             Avg: 0.4896ms (-91.16%)    Median: 0.4862ms (-90.81%)    [0.4797ms .. 0.5084ms]    Output: 1_054
num_rand_fast:[0_TO_9]_AND_num_asc_fast:[0_TO_9]_all_results        Avg: 2.7108ms (-50.81%)    Median: 2.6925ms (-49.51%)    [2.6688ms .. 2.7868ms]    Output: 968
dense and 10% a
a_AND_num_rand:[0_TO_9]_all_results                                 Avg: 0.9869ms (-3.71%)     Median: 0.9833ms (-3.83%)     [0.9518ms .. 1.1218ms]    Output: 4_914
a_AND_num_asc:[0_TO_9]_all_results                                  Avg: 0.6352ms (-3.74%)     Median: 0.6363ms (-3.32%)     [0.6158ms .. 0.6488ms]    Output: 10_152
a_AND_num_rand_fast:[0_TO_9]_all_results                            Avg: 3.1264ms (+0.39%)     Median: 3.1466ms (+1.34%)     [3.0261ms .. 3.2051ms]    Output: 4_914
a_AND_num_asc_fast:[0_TO_9]_all_results                             Avg: 4.1547ms (-31.12%)    Median: 4.0933ms (-28.55%)    [3.7648ms .. 4.7600ms]    Output: 10_152
num_rand_fast:[0_TO_9]_AND_num_asc_fast:[0_TO_9]_all_results        Avg: 2.6973ms (-52.30%)    Median: 2.6901ms (-49.86%)    [2.6689ms .. 2.7677ms]    Output: 968
```

Gains are largest when the range query is the non-leading docset of a low-cardinality intersection.
2026-07-03 13:42:33 +02:00
Pascal Seitz 82bee54a00 block_search: drop unsafe indexing, remove K=64
For K∈{2,4,8,16,32} LLVM proves the index bounds and elides the checks,
so get_unchecked buys nothing on the production K=8 path. K=64 was the
only value that defeated bounds-check elision (one check in the tail
scan) and it was instantiated in tests only — drop it.

block_search: cite the k-ary search paper

Document that kary_search is the 'k-ary search on a sorted array' variant
from Schlegel, Gemulla & Lehner (DaMoN 2009), specialized to a lower-bound.
2026-07-02 18:27:40 +02:00
Pascal Seitz 6892995d02 10% faster intersections: Use k-ary in block search 2026-07-02 18:27:40 +02:00
trinity-1686a 9704e6c0e3 Merge pull request #2983 from quickwit-oss/trinity.pointard/quickselect-term-agg
use select-nth instead of full sort in segment level agg top-k selection
2026-07-02 15:38:14 +02:00
trinity.pointard 715590b357 rename local var 2026-07-02 12:00:00 +00:00
trinity.pointard d496e402ca rustfmt 2026-07-02 07:12:04 +00:00
trinity.pointard 348ca1e309 don't count matching doc twice 2026-06-30 16:09:11 +00:00
trinity.pointard 5e4fe3520c better handle sorted buckets 2026-06-30 14:56:24 +00:00
Luca Cominardi b7234c153e fix: clean up manual mapping finalization invariants
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-30 13:12:23 +02:00
Luca Cominardi e3be28814e fix: reject default finalize with manual doc id mapping
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-30 10:56:46 +02:00
pascal a9733ba8c2 Keep buffered union refill out of line
BufferedUnionScorer is the hot path for full union traversal, including (TopDocs, Count) where Count forces all matches to be visited. After the block-wand intersection changes, LLVM started inlining the refill helper into the advance path, which regressed TOP_100_COUNT union queries even though the union algorithm did not change.

Force the refill helper out of line so the advance loop stays small and stable while pruning collectors continue to use Block-WAND.

Benchmark on search-benchmark-game TOP_100_COUNT union query set (301 queries, sum of per-query medians):
- tantivy 0.26: 0.853646s
- main before: 0.918605s
- this change: 0.841659s
2026-06-29 19:33:50 +02:00
pascal 874d54a63a Remove union wrapping for single-terms
search-benchmark-game shows TOP_100_COUNT regression on queries tagged intersection_union.

The regression came from allowing single-term boolean unions to become TermUnion for Block-WAND. https://github.com/quickwit-oss/tantivy/pull/2915
When such a scorer is used as the optional side of RequiredOptionalScorer, boxing converted the lone term into BufferedUnionScorer.

Keep the TermUnion representation available for pruning collection, but unwrap one-term unions when boxing or doing non-pruning iteration.
2026-06-29 19:33:50 +02:00
Luca Cominardi 3c6eb92c7e fix: restore clone bound on doc id iterator
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 11:31:38 +02:00
Luca Cominardi 9f73d3fe54 docs: clarify manual doc id mapping serialization
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 11:15:53 +02:00
Luca Cominardi 50272240c6 test: move doc id mapping validation coverage
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 11:13:35 +02:00
trinity.pointard 74a510cb56 try to use select-nth instead of full sort in segment level agg top-k selection 2026-06-29 09:13:21 +00:00
Luca Cominardi a6fd070e3d fix: validate manual doc id mappings at construction
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-29 10:52:41 +02:00
Luca Cominardi 06c046bdc9 fix: enable manual doc id mapping in single segment test
The regression test calls finalize_with_doc_id_mapping, so the index must opt into the temporary docstore path before constructing the segment writer.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-26 17:56:54 +02:00
Luca Cominardi 63f65dcd71 fix: add missing manual_doc_id_mapping field in zstd test IndexSettings initializer
Struct literal was missing the new field introduced in the manual doc id
mapping refactor, causing a compile error under --all-features.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-26 17:47:33 +02:00
Luca Cominardi 8a4c5b9013 refactor: gate manual doc id mapping via settings
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-26 17:37:59 +02:00
Luca Cominardi f7355e60cd feat: add single segment doc id mapping finalization
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-26 16:47:14 +02:00
Luca Cominardi 90603d2396 refactor custom doc id mapping finalization
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-26 16:24:19 +02:00
Luca Cominardi 910861a3e9 feat: add custom doc id mapping finalization
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-26 14:45:57 +02:00
trinity-1686a 02e34508e2 Merge pull request #2971 from quickwit-oss/trinity.pointard/fix-slop-overflow
fix overflow on large jumps in linear sequence
2026-06-23 10:18:29 +02:00
trinity-1686a 4031d97bac fix overflow on large jumps in linear sequence
new limit prevent an overflow in eval which caused the residual to be 64b when a slop of zero would give a smaller one
2026-06-23 00:13:27 +02:00
Ming 384f31d350 feat: Restore index sorting (#2959)
We ([ParadeDB](https://github.com/paradedb/paradedb)) have restored and been using the removed [index sorting](https://github.com/quickwit-oss/tantivy/issues/2352) feature in our Tantivy fork.

Our use case is sorting the index by Postgres' internal `ctid` identifier. Results returned from Tantivy must be checked against Postgres' visibility map, and checking them in ctid order is much more cache friendly, resulting in up to 80% speedups for certain queries.

This PR is split into 5 commits, corresponding to the index sorting reversal plus bug fixes we uncovered during our usage of index sorting.

| Commit | Maps to | What it does |
|---|---|---|
| `2aea0ad9f` | foundation ([#104](https://github.com/paradedb/tantivy/pull/104)) | Restore `SegmentComponent::TempStore` (revert of upstream #2815). Subsumes fork PR [#104](https://github.com/paradedb/tantivy/pull/104)'s CI fix. |
| `9205bcb0c` | [#92](https://github.com/paradedb/tantivy/pull/92) | Restore sort-by-field (single-segment + merge paths). |
| `39c790f0f` | [#101](https://github.com/paradedb/tantivy/pull/101) | Enable `sort_by` for `Str`/`Bytes` fast fields. |
| `9c4341a87` | [#105](https://github.com/paradedb/tantivy/pull/105) | Native typed numeric sort-key comparison (precision/NULL fix). |
| `2d9ba2418` | [#106](https://github.com/paradedb/tantivy/pull/106) | Preserve NULL ordering in numeric segment merges. |

We have discussed with the Tantivy maintainers and they indicated they would be open to this PR. Another motivation for landing this PR is we are planning on contributing a significant refactor that makes Tantivy's segment components extensible, and landing that without index sorting leads to too many conflicts.
2026-06-22 11:22:25 -07:00
Pascal Seitz 1e859fd78d fix term aggregation u32::MAX overflow issue 2026-06-18 17:07:43 +08:00
Pascal Seitz f451fa938f explain why naive scorer must accumulate scores in WAND order 2026-06-17 18:58:58 +08:00
Pascal Seitz 2a82dd6f64 fix flaky test 2026-06-17 18:58:58 +08:00
Pascal Seitz c096b2ad89 aggregation/terms: charge fused term_counts to the memory limit
term_counts (one u32/term) was allocated but not charged to
AggregationLimitsGuard, so a memory limit could be exceeded silently.
Charge it, skip allocating it when unbounded, and add a regression test.
2026-06-16 21:23:23 +08:00