Files
Pascal Seitz a03870d41b Up to 500x Faster exists queries on columns (hehe)
Use specialized optional and multivalued column indexes for single-column
exists queries, while retaining a generic fallback for dynamic column unions.
Implement seek_danger with direct value checks to avoid unnecessary scans.

```
exists
column populated in 0.1% of docs (sparse blocks)
optional           Avg: 0.0406ms (-99.83%)    Median: 0.0395ms (-99.84%)    [0.0367ms .. 0.0496ms]    Output: 4_836
multivalued        Avg: 0.0432ms (-99.85%)    Median: 0.0424ms (-99.86%)    [0.0401ms .. 0.0503ms]    Output: 4_836
column populated in 10% of docs (dense blocks)
optional           Avg: 6.1052ms (-52.67%)    Median: 6.1089ms (-51.50%)    [5.9257ms .. 6.2702ms]    Output: 499_966
multivalued        Avg: 6.5485ms (-64.36%)    Median: 6.4573ms (-64.12%)    [6.3324ms .. 8.0991ms]    Output: 499_966
column populated in 90% of docs (dense blocks)
optional           Avg: 12.9950ms (+14.89%)    Median: 12.8676ms (+14.74%)    [12.6985ms .. 14.3235ms]    Output: 4_498_850
multivalued        Avg: 15.7813ms (-39.05%)    Median: 15.6722ms (-39.20%)    [15.5501ms .. 16.6724ms]    Output: 4_498_850
term_AND_exists
term matches 0.01% of docs, column populated in 0.1% (sparse blocks)
optional           Avg: 7381ns (-99.70%)    Median: 7030ns (-99.72%)    [6798ns .. 0.0144ms]    Output: 1
multivalued        Avg: 8435ns (-99.74%)    Median: 7372ns (-99.75%)    [7152ns .. 0.0313ms]    Output: 1
term matches 1% of docs, column populated in 0.1% (sparse blocks)
optional           Avg: 0.1047ms (-99.52%)    Median: 0.1045ms (-99.52%)    [0.1040ms .. 0.1072ms]    Output: 51
multivalued        Avg: 0.1070ms (-99.59%)    Median: 0.1069ms (-99.58%)    [0.1063ms .. 0.1097ms]    Output: 51
term matches 50% of docs, column populated in 0.1% (sparse blocks)
optional           Avg: 0.2970ms (-98.72%)    Median: 0.2839ms (-98.76%)    [0.2648ms .. 0.3446ms]    Output: 2_405
multivalued        Avg: 0.2952ms (-98.95%)    Median: 0.2923ms (-98.96%)    [0.2641ms .. 0.3283ms]    Output: 2_405
term matches 0.01% of docs, column populated in 10% (dense blocks)
optional           Avg: 0.0124ms (-48.17%)    Median: 0.0124ms (-47.62%)    [0.0120ms .. 0.0131ms]    Output: 50
multivalued        Avg: 0.0196ms (-46.61%)    Median: 0.0197ms (-46.21%)    [0.0184ms .. 0.0205ms]    Output: 50
term matches 1% of docs, column populated in 10% (dense blocks)
optional           Avg: 0.8094ms (-44.69%)    Median: 0.8021ms (-44.60%)    [0.7787ms .. 0.9247ms]    Output: 5_023
multivalued        Avg: 0.8439ms (-58.79%)    Median: 0.8325ms (-58.70%)    [0.8099ms .. 0.9976ms]    Output: 5_023
term matches 50% of docs, column populated in 10% (dense blocks)
optional           Avg: 12.3175ms (-32.65%)    Median: 12.3049ms (-32.54%)    [12.1053ms .. 12.8036ms]    Output: 249_853
multivalued        Avg: 12.7152ms (-46.31%)    Median: 12.6789ms (-46.23%)    [12.5517ms .. 13.0544ms]    Output: 249_853
term matches 0.01% of docs, column populated in 90% (dense blocks)
optional           Avg: 0.0627ms (+3.59%)    Median: 0.0640ms (+5.17%)    [0.0557ms .. 0.0680ms]    Output: 439
multivalued        Avg: 0.1168ms (+1.16%)    Median: 0.1173ms (+1.90%)    [0.1050ms .. 0.1226ms]    Output: 439
term matches 1% of docs, column populated in 90% (dense blocks)
optional           Avg: 0.5377ms (-0.19%)     Median: 0.5346ms (-0.26%)     [0.5308ms .. 0.5550ms]    Output: 45_342
multivalued        Avg: 0.6365ms (-31.29%)    Median: 0.6351ms (-30.75%)    [0.6245ms .. 0.6572ms]    Output: 45_342
term matches 50% of docs, column populated in 90% (dense blocks)
optional           Avg: 27.3888ms (-1.86%)     Median: 27.3356ms (-1.48%)     [27.1377ms .. 28.7410ms]    Output: 2_249_485
multivalued        Avg: 29.4214ms (-25.34%)    Median: 29.3960ms (-25.17%)    [29.2107ms .. 30.2541ms]    Output: 2_249_485
```
2026-09-02 14:14:45 +08:00
..
2026-07-10 12:33:33 +02:00
2023-11-20 02:59:59 +01:00
2025-12-01 12:15:41 +01:00

Columnar format

This crate describes columnar format used in tantivy.

Goals

This format is special in the following way.

  • it needs to be compact
  • accessing a specific column does not require to load the entire columnar. It can be done in 2 to 3 random access.
  • columns of several types can be associated with the same column name.
  • it needs to support columns with different types (str, u64, i64, f64) and different cardinality (required, optional, multivalued).
  • columns, once loaded, offer cheap random access.
  • it is designed to allow range queries.

Coercion rules

Users can create a columnar by inserting rows to a ColumnarWriter, and serializing it into a Write object. Nothing prevents a user from recording values with different type to the same column_name.

In that case, tantivy-columnar's behavior is as follows:

  • JsonValues are grouped into 3 types (String, Number, bool). Values that corresponds to different groups are mapped to different columns. For instance, String values are treated independently from Number or boolean values. tantivy-columnar will simply emit several columns associated to a given column_name.
  • Only one column for a given json value type is emitted. If number values with different number types are recorded (e.g. u64, i64, f64), tantivy-columnar will pick the first type that can represents the set of appended value, with the following prioriy order (i64, u64, f64). i64 is picked over u64 as it is likely to yield less change of types. Most use cases strictly requiring u64 show the restriction on 50% of the values (e.g. a 64-bit hash). On the other hand, a lot of use cases can show rare negative value.

Columnar format

This columnar format may have more than one column (with different types) associated to the same column_name (see Coercion rules above). The (column_name, column_type) couple however uniquely identifies a column. That couple is serialized as a column column_key. The format of that key is: [column_name][ZERO_BYTE][column_type_header: u8]

COLUMNAR:=
    [COLUMNAR_DATA]
    [COLUMNAR_KEY_TO_DATA_INDEX]
    [COLUMNAR_FOOTER];


# Columns are sorted by their column key.
COLUMNAR_DATA:=
    [COLUMN_DATA]+;

COLUMNAR_FOOTER := [RANGE_SSTABLE_BYTES_LEN: 8 bytes little endian]

The columnar file starts by the actual column data, concatenated one after the other, sorted by column key.

A sstable associates `(column name, column_cardinality, column_type) to range of bytes.

Column name may not contain the zero byte \0.

Listing all columns associated to column_name can therefore be done by listing all keys prefixed by [column_name][ZERO_BYTE]

The associated range of bytes refer to a range of bytes

This crate exposes a columnar format for tantivy. This format is described in README.md

The crate introduces the following concepts.

Columnar is an equivalent of a dataframe. It maps column_key to Column.

A Column<T> associates a RowId (u32) to any number of values.

This is made possible by wrapping a ColumnIndex and a ColumnValue object. The ColumnValue<T> represents a mapping that associates each RowId to exactly one single value.

The ColumnIndex then maps each RowId to a set of RowId in the ColumnValue.

For optimization, and compression purposes, the ColumnIndex has three possible representation, each for different cardinalities.

  • Full

All RowId have exactly one value. The ColumnIndex is the trivial mapping.

  • Optional

All RowIds can have at most one value. The ColumnIndex is the trivial mapping ColumnRowId -> Option<ColumnValueRowId>.

  • Multivalued

All RowIds can have any number of values. The column index is mapping values to a range.

All these objects are implemented an unit tested independently in their own module:

  • columnar
  • column_index
  • column_values
  • column