Commit Graph
14 Commits
Author SHA1 Message Date
Paul MasurelandPaul Masurel 047464cf92 Accelerating jitexpr conditions using necessary conditions (#3129)
* jitexpr

* Added possible necessary conditions to jitexpr.

A function makes it possible to infer a necessary query from an expression to match.
We can then accelerate queries involving a calculated field by not even evaluating the expression
on docs that trivially do not match.

* CR comment

* CR comment

* Fixing unit tests

---------

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2026-09-28 18:46:37 +02:00
Pascal Seitz c3b91b617a use next_doc in mixed type columns
Remove full and empty variants from exists docsets
Simplify exists docset advancement
fix size hint in exist query
2026-09-02 14:14:45 +08:00
Pascal Seitz 271635b5a6 Use default seek behavior for exists docsets 2026-09-02 14:14:45 +08:00
Pascal Seitz a03135ac4f Simplify column index collection 2026-09-02 14:14:45 +08:00
Pascal Seitz 14c119745b Clarify fast-field exists weight name 2026-09-02 14:14:45 +08:00
Pascal Seitz a03870d41b Up to 500x Faster exists queries on columns (hehe)
Use specialized optional and multivalued column indexes for single-column
exists queries, while retaining a generic fallback for dynamic column unions.
Implement seek_danger with direct value checks to avoid unnecessary scans.

```
exists
column populated in 0.1% of docs (sparse blocks)
optional           Avg: 0.0406ms (-99.83%)    Median: 0.0395ms (-99.84%)    [0.0367ms .. 0.0496ms]    Output: 4_836
multivalued        Avg: 0.0432ms (-99.85%)    Median: 0.0424ms (-99.86%)    [0.0401ms .. 0.0503ms]    Output: 4_836
column populated in 10% of docs (dense blocks)
optional           Avg: 6.1052ms (-52.67%)    Median: 6.1089ms (-51.50%)    [5.9257ms .. 6.2702ms]    Output: 499_966
multivalued        Avg: 6.5485ms (-64.36%)    Median: 6.4573ms (-64.12%)    [6.3324ms .. 8.0991ms]    Output: 499_966
column populated in 90% of docs (dense blocks)
optional           Avg: 12.9950ms (+14.89%)    Median: 12.8676ms (+14.74%)    [12.6985ms .. 14.3235ms]    Output: 4_498_850
multivalued        Avg: 15.7813ms (-39.05%)    Median: 15.6722ms (-39.20%)    [15.5501ms .. 16.6724ms]    Output: 4_498_850
term_AND_exists
term matches 0.01% of docs, column populated in 0.1% (sparse blocks)
optional           Avg: 7381ns (-99.70%)    Median: 7030ns (-99.72%)    [6798ns .. 0.0144ms]    Output: 1
multivalued        Avg: 8435ns (-99.74%)    Median: 7372ns (-99.75%)    [7152ns .. 0.0313ms]    Output: 1
term matches 1% of docs, column populated in 0.1% (sparse blocks)
optional           Avg: 0.1047ms (-99.52%)    Median: 0.1045ms (-99.52%)    [0.1040ms .. 0.1072ms]    Output: 51
multivalued        Avg: 0.1070ms (-99.59%)    Median: 0.1069ms (-99.58%)    [0.1063ms .. 0.1097ms]    Output: 51
term matches 50% of docs, column populated in 0.1% (sparse blocks)
optional           Avg: 0.2970ms (-98.72%)    Median: 0.2839ms (-98.76%)    [0.2648ms .. 0.3446ms]    Output: 2_405
multivalued        Avg: 0.2952ms (-98.95%)    Median: 0.2923ms (-98.96%)    [0.2641ms .. 0.3283ms]    Output: 2_405
term matches 0.01% of docs, column populated in 10% (dense blocks)
optional           Avg: 0.0124ms (-48.17%)    Median: 0.0124ms (-47.62%)    [0.0120ms .. 0.0131ms]    Output: 50
multivalued        Avg: 0.0196ms (-46.61%)    Median: 0.0197ms (-46.21%)    [0.0184ms .. 0.0205ms]    Output: 50
term matches 1% of docs, column populated in 10% (dense blocks)
optional           Avg: 0.8094ms (-44.69%)    Median: 0.8021ms (-44.60%)    [0.7787ms .. 0.9247ms]    Output: 5_023
multivalued        Avg: 0.8439ms (-58.79%)    Median: 0.8325ms (-58.70%)    [0.8099ms .. 0.9976ms]    Output: 5_023
term matches 50% of docs, column populated in 10% (dense blocks)
optional           Avg: 12.3175ms (-32.65%)    Median: 12.3049ms (-32.54%)    [12.1053ms .. 12.8036ms]    Output: 249_853
multivalued        Avg: 12.7152ms (-46.31%)    Median: 12.6789ms (-46.23%)    [12.5517ms .. 13.0544ms]    Output: 249_853
term matches 0.01% of docs, column populated in 90% (dense blocks)
optional           Avg: 0.0627ms (+3.59%)    Median: 0.0640ms (+5.17%)    [0.0557ms .. 0.0680ms]    Output: 439
multivalued        Avg: 0.1168ms (+1.16%)    Median: 0.1173ms (+1.90%)    [0.1050ms .. 0.1226ms]    Output: 439
term matches 1% of docs, column populated in 90% (dense blocks)
optional           Avg: 0.5377ms (-0.19%)     Median: 0.5346ms (-0.26%)     [0.5308ms .. 0.5550ms]    Output: 45_342
multivalued        Avg: 0.6365ms (-31.29%)    Median: 0.6351ms (-30.75%)    [0.6245ms .. 0.6572ms]    Output: 45_342
term matches 50% of docs, column populated in 90% (dense blocks)
optional           Avg: 27.3888ms (-1.86%)     Median: 27.3356ms (-1.48%)     [27.1377ms .. 28.7410ms]    Output: 2_249_485
multivalued        Avg: 29.4214ms (-25.34%)    Median: 29.3960ms (-25.17%)    [29.2107ms .. 30.2541ms]    Output: 2_249_485
```
2026-09-02 14:14:45 +08:00
Paul MasurelandPaul Masurel f88b7200b2 Optimization when posting list are saturated. (#2745)
* Optimization when posting list are saturated.

If a posting list doc freq is the segment reader's
max_doc, and if scoring does not matter, we can replace it
by a AllScorer.

In turn, in a boolean query, we can dismiss  all scorers and
empty scorers, to accelerate the request.

* Added range query optimization

* CR comment

* CR comments

* CR comment

---------

Co-authored-by: Paul Masurel <paul.masurel@datadoghq.com>
2025-11-26 15:50:57 +01:00
PSeitz-dd 203751f2fe Optimize ExistsQuery for a high number of dynamic columns (#2694)
* Optimize ExistsQuery for a high number of dynamic columns

The previous algorithm checked _each_ doc in _each_ column for
existence. This causes huge cost on JSON fields with e.g. 100k columns.
Compute a bitset instead if we have more than one column.

add `iter_docs` to the multivalued_index

* add benchmark

subfields=1
exists_json_union    Memory: 89.3 KB (+2.01%)    Avg: 0.4865ms (-26.03%)    Median: 0.4865ms (-26.03%)    [0.4865ms .. 0.4865ms]
subfields=2
exists_json_union    Memory: 68.1 KB     Avg: 1.7048ms (-0.46%)    Median: 1.7048ms (-0.46%)    [1.7048ms .. 1.7048ms]
subfields=3
exists_json_union    Memory: 61.8 KB     Avg: 2.0742ms (-2.22%)    Median: 2.0742ms (-2.22%)    [2.0742ms .. 2.0742ms]
subfields=4
exists_json_union    Memory: 119.8 KB (+103.44%)    Avg: 3.9500ms (+42.62%)    Median: 3.9500ms (+42.62%)    [3.9500ms .. 3.9500ms]
subfields=5
exists_json_union    Memory: 120.4 KB (+107.65%)    Avg: 3.9610ms (+20.65%)    Median: 3.9610ms (+20.65%)    [3.9610ms .. 3.9610ms]
subfields=6
exists_json_union    Memory: 120.6 KB (+107.49%)    Avg: 3.8903ms (+3.11%)    Median: 3.8903ms (+3.11%)    [3.8903ms .. 3.8903ms]
subfields=7
exists_json_union    Memory: 120.9 KB (+106.93%)    Avg: 3.6220ms (-16.22%)    Median: 3.6220ms (-16.22%)    [3.6220ms .. 3.6220ms]
subfields=8
exists_json_union    Memory: 121.3 KB (+106.23%)    Avg: 4.0981ms (-15.97%)    Median: 4.0981ms (-15.97%)    [4.0981ms .. 4.0981ms]
subfields=16
exists_json_union    Memory: 123.1 KB (+103.09%)    Avg: 4.3483ms (-92.26%)    Median: 4.3483ms (-92.26%)    [4.3483ms .. 4.3483ms]
subfields=256
exists_json_union    Memory: 204.6 KB (+19.85%)    Avg: 3.8874ms (-99.01%)    Median: 3.8874ms (-99.01%)    [3.8874ms .. 3.8874ms]
subfields=4096
exists_json_union    Memory: 2.0 MB     Avg: 3.5571ms (-99.90%)    Median: 3.5571ms (-99.90%)    [3.5571ms .. 3.5571ms]
subfields=65536
exists_json_union    Memory: 28.3 MB     Avg: 14.4417ms (-99.97%)    Median: 14.4417ms (-99.97%)    [14.4417ms .. 14.4417ms]
subfields=262144
exists_json_union    Memory: 113.3 MB     Avg: 66.2860ms (-99.95%)    Median: 66.2860ms (-99.95%)    [66.2860ms .. 66.2860ms]

* rename methods
2025-09-16 18:21:03 +02:00
PSeitzandPascal Seitz 945af922d1 clippy (#2661)
* clippy

* use readable version

---------

Co-authored-by: Pascal Seitz <pascal.seitz@datadoghq.com>
2025-07-02 11:25:03 +02:00
Remi Dettai 71cf19870b Exist queries match subpath fields (#2558)
* Exist queries match subpath fields

* Make subpath check optional

* Add async subpath listing
2025-01-06 10:17:39 +01:00
PSeitz 1b4076691f refactor fast field query (#2452)
As preparation of #2023 and #1709

* Use Term to pass parameters
* merge u64 and ip fast field range query

Side note: I did not rename range_query_u64_fastfield, because then git can't track the changes.
2024-07-15 18:08:05 +08:00
PSeitz 74940e9345 clippy (#2349)
* fix clippy

* fix clippy

* fix duplicate imports
2024-04-09 07:54:44 +02:00
PSeitz 48630ceec9 move into new index module (#2259)
move core modules to index module
2024-01-31 10:30:04 +01:00
Igor Motov 19325132b7 Fast-field based implementation of ExistsQuery (#2160)
Adds an implementation of ExistsQuery that takes advantage of fast fields.

Fixes #2159
2023-09-07 11:51:49 +09:00