bump to 0.22.1

Fix TopNComputer for reverse order
update CHANGELOG, use github API in cliff (#2354 )
2026-01-08 10:02:55 +00:00 · 2025-07-17 11:22:19 +08:00 · 2025-07-17 10:54:21 +08:00 · 2024-04-15 10:07:20 +02:00 · 2024-04-10 08:20:52 +02:00 · 2024-04-10 08:09:09 +02:00
24 changed files with 278 additions and 109 deletions
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -1,3 +1,65 @@
+Tantivy 0.22
+================================
+
+Tantivy 0.22 will be able to read indices created with Tantivy 0.21.
+
+#### Bugfixes
+- Fix null byte handling in JSON paths (null bytes in json keys caused panic during indexing) [#2345](https://github.com/quickwit-oss/tantivy/pull/2345)(@PSeitz)
+- Fix bug that can cause `get_docids_for_value_range` to panic. [#2295](https://github.com/quickwit-oss/tantivy/pull/2295)(@fulmicoton)
+- Avoid 1 document indices by increase min memory to 15MB for indexing [#2176](https://github.com/quickwit-oss/tantivy/pull/2176)(@PSeitz)
+- Fix merge panic for JSON fields [#2284](https://github.com/quickwit-oss/tantivy/pull/2284)(@PSeitz)
+- Fix bug occuring when merging JSON object indexed with positions. [#2253](https://github.com/quickwit-oss/tantivy/pull/2253)(@fulmicoton)
+- Fix empty DateHistogram gap bug [#2183](https://github.com/quickwit-oss/tantivy/pull/2183)(@PSeitz)
+- Fix range query end check (fields with less than 1 value per doc are affected) [#2226](https://github.com/quickwit-oss/tantivy/pull/2226)(@PSeitz)
+- Handle exclusive out of bounds ranges on fastfield range queries [#2174](https://github.com/quickwit-oss/tantivy/pull/2174)(@PSeitz)
+
+#### Breaking API Changes
+- rename ReloadPolicy onCommit to onCommitWithDelay [#2235](https://github.com/quickwit-oss/tantivy/pull/2235)(@giovannicuccu)
+- Move exports from the root into modules [#2220](https://github.com/quickwit-oss/tantivy/pull/2220)(@PSeitz)
+- Accept field name instead of `Field` in FilterCollector [#2196](https://github.com/quickwit-oss/tantivy/pull/2196)(@PSeitz)
+- remove deprecated IntOptions and DateTime [#2353](https://github.com/quickwit-oss/tantivy/pull/2353)(@PSeitz)
+
+#### Features/Improvements
+- Tantivy documents as a trait: Index data directly without converting to tantivy types first [#2071](https://github.com/quickwit-oss/tantivy/pull/2071)(@ChillFish8)
+- encode some part of posting list as -1 instead of direct values (smaller inverted indices) [#2185](https://github.com/quickwit-oss/tantivy/pull/2185)(@trinity-1686a)
+- **Aggregation**
+  - Support to deserialize f64 from string [#2311](https://github.com/quickwit-oss/tantivy/pull/2311)(@PSeitz)
+  - Add a top_hits aggregator [#2198](https://github.com/quickwit-oss/tantivy/pull/2198)(@ditsuke)
+  - Support bool type in term aggregation [#2318](https://github.com/quickwit-oss/tantivy/pull/2318)(@PSeitz)
+  - Support ip adresses in term aggregation [#2319](https://github.com/quickwit-oss/tantivy/pull/2319)(@PSeitz)
+  - Support date type in term aggregation [#2172](https://github.com/quickwit-oss/tantivy/pull/2172)(@PSeitz)
+  - Support escaped dot when addressing field [#2250](https://github.com/quickwit-oss/tantivy/pull/2250)(@PSeitz)
+
+- Add ExistsQuery to check documents that have a value [#2160](https://github.com/quickwit-oss/tantivy/pull/2160)(@imotov)
+- Expose TopDocs::order_by_u64_field again [#2282](https://github.com/quickwit-oss/tantivy/pull/2282)(@ditsuke)
+
+- **Memory/Performance**
+  - Faster TopN: replace BinaryHeap with TopNComputer [#2186](https://github.com/quickwit-oss/tantivy/pull/2186)(@PSeitz)
+  - reduce number of allocations during indexing [#2257](https://github.com/quickwit-oss/tantivy/pull/2257)(@PSeitz)
+  - Less Memory while indexing: docid deltas while indexing [#2249](https://github.com/quickwit-oss/tantivy/pull/2249)(@PSeitz)
+  - Faster indexing: use term hashmap in fastfield [#2243](https://github.com/quickwit-oss/tantivy/pull/2243)(@PSeitz)
+  - term hashmap remove copy in is_empty, unused unordered_id [#2229](https://github.com/quickwit-oss/tantivy/pull/2229)(@PSeitz)
+  - add method to fetch block of first values in columnar [#2330](https://github.com/quickwit-oss/tantivy/pull/2330)(@PSeitz)
+  - Faster aggregations: add fast path for full columns in fetch_block [#2328](https://github.com/quickwit-oss/tantivy/pull/2328)(@PSeitz)
+  - Faster sstable loading: use fst for sstable index [#2268](https://github.com/quickwit-oss/tantivy/pull/2268)(@trinity-1686a)
+
+- **QueryParser**
+  - allow newline where we allow space in query parser [#2302](https://github.com/quickwit-oss/tantivy/pull/2302)(@trinity-1686a)
+  - allow some mixing of occur and bool in strict query parser [#2323](https://github.com/quickwit-oss/tantivy/pull/2323)(@trinity-1686a)
+  - handle * inside term in lenient query parser [#2228](https://github.com/quickwit-oss/tantivy/pull/2228)(@trinity-1686a)
+  - add support for exists query syntax in query parser [#2170](https://github.com/quickwit-oss/tantivy/pull/2170)(@trinity-1686a)
+- Add shared search executor [#2312](https://github.com/quickwit-oss/tantivy/pull/2312)(@MochiXu)
+- Truncate keys to u16::MAX in term hashmap [#2299](https://github.com/quickwit-oss/tantivy/pull/2299)(@PSeitz)
+- report if a term matched when warming up posting list [#2309](https://github.com/quickwit-oss/tantivy/pull/2309)(@trinity-1686a)
+- Support json fields in FuzzyTermQuery [#2173](https://github.com/quickwit-oss/tantivy/pull/2173)(@PingXia-at)
+- Read list of fields encoded in term dictionary for JSON fields [#2184](https://github.com/quickwit-oss/tantivy/pull/2184)(@PSeitz)
+- add collect_block to BoxableSegmentCollector [#2331](https://github.com/quickwit-oss/tantivy/pull/2331)(@PSeitz)
+- expose collect_block buffer size [#2326](https://github.com/quickwit-oss/tantivy/pull/2326)(@PSeitz)
+- Forward regex parser errors [#2288](https://github.com/quickwit-oss/tantivy/pull/2288)(@adamreichold)
+- Make FacetCounts defaultable and cloneable. [#2322](https://github.com/quickwit-oss/tantivy/pull/2322)(@adamreichold)
+- Derive Debug for SchemaBuilder [#2254](https://github.com/quickwit-oss/tantivy/pull/2254)(@GodTamIt)
+- add missing inlines to tantivy options [#2245](https://github.com/quickwit-oss/tantivy/pull/2245)(@PSeitz)
+
 Tantivy 0.21.1
 ================================
 #### Bugfixes
--- a/Cargo.toml
+++ b/Cargo.toml
@@ -1,6 +1,6 @@
 [package]
 name = "tantivy"
-version = "0.22.0-dev"
+version = "0.22.1"
 authors = ["Paul Masurel <paul.masurel@gmail.com>"]
 license = "MIT"
 categories = ["database-implementations", "data-structures"]
@@ -52,13 +52,13 @@ itertools = "0.12.0"
 measure_time = "0.8.2"
 arc-swap = "1.5.0"

-columnar = { version= "0.2", path="./columnar", package ="tantivy-columnar" }
-sstable = { version= "0.2", path="./sstable", package ="tantivy-sstable", optional = true }
-stacker = { version= "0.2", path="./stacker", package ="tantivy-stacker" }
-query-grammar = { version= "0.21.0", path="./query-grammar", package = "tantivy-query-grammar" }
-tantivy-bitpacker = { version= "0.5", path="./bitpacker" }
-common = { version= "0.6", path = "./common/", package = "tantivy-common" }
-tokenizer-api = { version= "0.2", path="./tokenizer-api", package="tantivy-tokenizer-api" }
+columnar = { version= "0.3", path="./columnar", package ="tantivy-columnar" }
+sstable = { version= "0.3", path="./sstable", package ="tantivy-sstable", optional = true }
+stacker = { version= "0.3", path="./stacker", package ="tantivy-stacker" }
+query-grammar = { version= "0.22.0", path="./query-grammar", package = "tantivy-query-grammar" }
+tantivy-bitpacker = { version= "0.6", path="./bitpacker" }
+common = { version= "0.7", path = "./common/", package = "tantivy-common" }
+tokenizer-api = { version= "0.3", path="./tokenizer-api", package="tantivy-tokenizer-api" }
 sketches-ddsketch = { version = "0.2.1", features = ["use_serde"] }
 futures-util = { version = "0.3.28", optional = true }
 fnv = "1.0.7"
--- a/bitpacker/Cargo.toml
+++ b/bitpacker/Cargo.toml
@@ -1,6 +1,6 @@
 [package]
 name = "tantivy-bitpacker"
-version = "0.5.0"
+version = "0.6.0"
 edition = "2021"
 authors = ["Paul Masurel <paul.masurel@gmail.com>"]
 license = "MIT"
--- a/cliff.toml
+++ b/cliff.toml
@@ -1,6 +1,10 @@
 # configuration file for git-cliff{ pattern = "foo", replace = "bar"}
 # see https://github.com/orhun/git-cliff#configuration-file

+[remote.github]
+owner = "quickwit-oss"
+repo = "tantivy"
+
 [changelog]
 # changelog header
 header = """
@@ -8,15 +12,43 @@ header = """
 # template for the changelog body
 # https://tera.netlify.app/docs/#introduction
 body = """
-{% if version %}\
-    {{ version | trim_start_matches(pat="v") }} ({{ timestamp | date(format="%Y-%m-%d") }})
-    ==================
-{% else %}\
-    ## [unreleased]
-{% endif %}\
+## What's Changed
+
+{%- if version %} in {{ version }}{%- endif -%}
 {% for commit in commits %}
-    - {% if commit.breaking %}[**breaking**] {% endif %}{{ commit.message | split(pat="\n") | first | trim | upper_first }}(@{{ commit.author.name }})\
-{% endfor %}
+  {% if commit.github.pr_title -%}
+    {%- set commit_message = commit.github.pr_title -%}
+  {%- else -%}
+    {%- set commit_message = commit.message -%}
+  {%- endif -%}
+  - {{ commit_message | split(pat="\n") | first | trim }}\
+    {% if commit.github.pr_number %} \
+      [#{{ commit.github.pr_number }}]({{ self::remote_url() }}/pull/{{ commit.github.pr_number }}){% if commit.github.username %}(@{{ commit.github.username }}){%- endif -%} \
+    {%- endif %}
+{%- endfor -%}
+
+{% if github.contributors | filter(attribute="is_first_time", value=true) | length != 0 %}
+  {% raw %}\n{% endraw -%}
+  ## New Contributors
+{%- endif %}\
+{% for contributor in github.contributors | filter(attribute="is_first_time", value=true) %}
+  * @{{ contributor.username }} made their first contribution
+    {%- if contributor.pr_number %} in \
+      [#{{ contributor.pr_number }}]({{ self::remote_url() }}/pull/{{ contributor.pr_number }}) \
+    {%- endif %}
+{%- endfor -%}
+
+{% if version %}
+    {% if previous.version %}
+      **Full Changelog**: {{ self::remote_url() }}/compare/{{ previous.version }}...{{ version }}
+    {% endif %}
+{% else -%}
+  {% raw %}\n{% endraw %}
+{% endif %}
+
+{%- macro remote_url() -%}
+  https://github.com/{{ remote.github.owner }}/{{ remote.github.repo }}
+{%- endmacro -%}
 """
 # remove the leading and trailing whitespace from the template
 trim = true
@@ -25,53 +57,24 @@ footer = """
 """

 postprocessors = [
-    { pattern = 'Paul Masurel', replace = "fulmicoton"}, # replace with github user
-    { pattern = 'PSeitz', replace = "PSeitz"}, # replace with github user
-    { pattern = 'Adam Reichold', replace = "adamreichold"}, # replace with github user
-    { pattern = 'trinity-1686a', replace = "trinity-1686a"}, # replace with github user
-    { pattern = 'Michael Kleen', replace = "mkleen"}, # replace with github user
-    { pattern = 'Adrien Guillo', replace = "guilload"}, # replace with github user
-    { pattern = 'François Massot', replace = "fmassot"}, # replace with github user
-    { pattern = 'Naveen Aiathurai', replace = "naveenann"}, # replace with github user
-    { pattern = '', replace = ""}, # replace with github user
 ]

 [git]
 # parse the commits based on https://www.conventionalcommits.org
 # This is required or commit.message contains the whole commit message and not just the title
-conventional_commits = true
+conventional_commits = false
 # filter out the commits that are not conventional
-filter_unconventional = false
+filter_unconventional = true
 # process each line of a commit as an individual commit
 split_commits = false
 # regex for preprocessing the commit messages
 commit_preprocessors = [
-    { pattern = '\((\w+\s)?#([0-9]+)\)', replace = "[#${2}](https://github.com/quickwit-oss/tantivy/issues/${2})"}, # replace issue numbers
+    { pattern = '\((\w+\s)?#([0-9]+)\)', replace = ""},
 ]
 #link_parsers = [
    #{ pattern = "#(\\d+)", href = "https://github.com/quickwit-oss/tantivy/pulls/$1"},
 #]
 # regex for parsing and grouping commits
-commit_parsers = [
-    { message = "^feat", group = "Features"},
-    { message = "^fix", group = "Bug Fixes"},
-    { message = "^doc", group = "Documentation"},
-    { message = "^perf", group = "Performance"},
-    { message = "^refactor", group = "Refactor"},
-    { message = "^style", group = "Styling"},
-    { message = "^test", group = "Testing"},
-    { message = "^chore\\(release\\): prepare for", skip = true},
-    { message = "(?i)clippy", skip = true},
-    { message = "(?i)dependabot", skip = true},
-    { message = "(?i)fmt", skip = true},
-    { message = "(?i)bump", skip = true},
-    { message = "(?i)readme", skip = true},
-    { message = "(?i)comment", skip = true},
-    { message = "(?i)spelling", skip = true},
-    { message = "^chore", group = "Miscellaneous Tasks"},
-    { body = ".*security", group = "Security"},
-    { message = ".*", group = "Other", default_scope = "other"},
-]
 # protect breaking changes from being skipped due to matching a skipping commit_parser
 protect_breaking_commits = false
 # filter out the commits that are not matched by commit parsers
--- a/columnar/Cargo.toml
+++ b/columnar/Cargo.toml
@@ -1,6 +1,6 @@
 [package]
 name = "tantivy-columnar"
-version = "0.2.0"
+version = "0.3.0"
 edition = "2021"
 license = "MIT"
 homepage = "https://github.com/quickwit-oss/tantivy"
@@ -12,10 +12,10 @@ categories = ["database-implementations", "data-structures", "compression"]
 itertools = "0.12.0"
 fastdivide = "0.4.0"

-stacker = { version= "0.2", path = "../stacker", package="tantivy-stacker"}
-sstable = { version= "0.2", path = "../sstable", package = "tantivy-sstable" }
-common = { version= "0.6", path = "../common", package = "tantivy-common" }
-tantivy-bitpacker = { version= "0.5", path = "../bitpacker/" }
+stacker = { version= "0.3", path = "../stacker", package="tantivy-stacker"}
+sstable = { version= "0.3", path = "../sstable", package = "tantivy-sstable" }
+common = { version= "0.7", path = "../common", package = "tantivy-common" }
+tantivy-bitpacker = { version= "0.6", path = "../bitpacker/" }
 serde = "1.0.152"
 downcast-rs = "1.2.0"

--- a/common/Cargo.toml
+++ b/common/Cargo.toml
@@ -1,6 +1,6 @@
 [package]
 name = "tantivy-common"
-version = "0.6.0"
+version = "0.7.0"
 authors = ["Paul Masurel <paul@quickwit.io>", "Pascal Seitz <pascal@quickwit.io>"]
 license = "MIT"
 edition = "2021"
@@ -14,7 +14,7 @@ repository = "https://github.com/quickwit-oss/tantivy"

 [dependencies]
 byteorder = "1.4.3"
-ownedbytes = { version= "0.6", path="../ownedbytes" }
+ownedbytes = { version= "0.7", path="../ownedbytes" }
 async-trait = "0.1"
 time = { version = "0.3.10", features = ["serde-well-known"] }
 serde = { version = "1.0.136", features = ["derive"] }
--- a/common/src/datetime.rs
+++ b/common/src/datetime.rs
@@ -1,5 +1,3 @@
-#![allow(deprecated)]
-
 use std::fmt;
 use std::io::{Read, Write};

@@ -27,9 +25,6 @@ pub enum DateTimePrecision {
    Nanoseconds,
 }

-#[deprecated(since = "0.20.0", note = "Use `DateTimePrecision` instead")]
-pub type DatePrecision = DateTimePrecision;
-
 /// A date/time value with nanoseconds precision.
 ///
 /// This timestamp does not carry any explicit time zone information.
--- a/common/src/lib.rs
+++ b/common/src/lib.rs
@@ -15,8 +15,6 @@ mod vint;
 mod writer;
 pub use bitset::*;
 pub use byte_count::ByteCount;
-#[allow(deprecated)]
-pub use datetime::DatePrecision;
 pub use datetime::{DateTime, DateTimePrecision};
 pub use group_by::GroupByIteratorExtended;
 pub use json_path_writer::JsonPathWriter;
--- a/ownedbytes/Cargo.toml
+++ b/ownedbytes/Cargo.toml
@@ -1,7 +1,7 @@
 [package]
 authors = ["Paul Masurel <paul@quickwit.io>", "Pascal Seitz <pascal@quickwit.io>"]
 name = "ownedbytes"
-version = "0.6.0"
+version = "0.7.0"
 edition = "2021"
 description = "Expose data as static slice"
 license = "MIT"
--- a/query-grammar/Cargo.toml
+++ b/query-grammar/Cargo.toml
@@ -1,6 +1,6 @@
 [package]
 name = "tantivy-query-grammar"
-version = "0.21.0"
+version = "0.22.0"
 authors = ["Paul Masurel <paul.masurel@gmail.com>"]
 license = "MIT"
 categories = ["database-implementations", "data-structures"]
--- a/src/aggregation/metric/top_hits.rs
+++ b/src/aggregation/metric/top_hits.rs
@@ -13,6 +13,7 @@ use crate::aggregation::intermediate_agg_result::{
    IntermediateAggregationResult, IntermediateMetricResult,
 };
 use crate::aggregation::segment_agg_result::SegmentAggregationCollector;
+use crate::aggregation::AggregationError;
 use crate::collector::TopNComputer;
 use crate::schema::term::JSON_PATH_SEGMENT_SEP_STR;
 use crate::schema::OwnedValue;
@@ -96,6 +97,14 @@ pub struct TopHitsAggregation {
    #[serde(rename = "docvalue_fields")]
    #[serde(default)]
    doc_value_fields: Vec<String>,
+
+    // Not supported
+    _source: Option<serde_json::Value>,
+    fields: Option<serde_json::Value>,
+    script_fields: Option<serde_json::Value>,
+    highlight: Option<serde_json::Value>,
+    explain: Option<serde_json::Value>,
+    version: Option<serde_json::Value>,
 }

 #[derive(Debug, Clone, PartialEq, Default)]
@@ -142,12 +151,49 @@ fn globbed_string_to_regex(glob: &str) -> Result<Regex, crate::TantivyError> {
    })
 }

+fn use_doc_value_fields_err(parameter: &str) -> crate::Result<()> {
+    Err(crate::TantivyError::AggregationError(
+        AggregationError::InvalidRequest(format!(
+            "The `{}` parameter is not supported, only `docvalue_fields` is supported in \
+             `top_hits` aggregation",
+            parameter
+        )),
+    ))
+}
+fn unsupported_err(parameter: &str) -> crate::Result<()> {
+    Err(crate::TantivyError::AggregationError(
+        AggregationError::InvalidRequest(format!(
+            "The `{}` parameter is not supported in the `top_hits` aggregation",
+            parameter
+        )),
+    ))
+}
+
 impl TopHitsAggregation {
    /// Validate and resolve field retrieval parameters
    pub fn validate_and_resolve_field_names(
        &mut self,
        reader: &ColumnarReader,
    ) -> crate::Result<()> {
+        if self._source.is_some() {
+            use_doc_value_fields_err("_source")?;
+        }
+        if self.fields.is_some() {
+            use_doc_value_fields_err("fields")?;
+        }
+        if self.script_fields.is_some() {
+            use_doc_value_fields_err("script_fields")?;
+        }
+        if self.explain.is_some() {
+            unsupported_err("explain")?;
+        }
+        if self.highlight.is_some() {
+            unsupported_err("highlight")?;
+        }
+        if self.version.is_some() {
+            unsupported_err("version")?;
+        }
+
        self.doc_value_fields = self
            .doc_value_fields
            .iter()
--- a/src/aggregation/mod.rs
+++ b/src/aggregation/mod.rs
@@ -159,6 +159,10 @@ use itertools::Itertools;
 use serde::de::{self, Visitor};
 use serde::{Deserialize, Deserializer, Serialize};

+pub(crate) fn invalid_agg_request(message: String) -> crate::TantivyError {
+    crate::TantivyError::AggregationError(AggregationError::InvalidRequest(message))
+}
+
 fn parse_str_into_f64<E: de::Error>(value: &str) -> Result<f64, E> {
    let parsed = value.parse::<f64>().map_err(|_err| {
        de::Error::custom(format!("Failed to parse f64 from string: {:?}", value))
--- a/src/collector/top_score_collector.rs
+++ b/src/collector/top_score_collector.rs
@@ -787,7 +787,7 @@ impl<Score, D, const R: bool> From<TopNComputerDeser<Score, D, R>> for TopNCompu
    }
 }

-impl<Score, D, const R: bool> TopNComputer<Score, D, R>
+impl<Score, D, const REVERSE_ORDER: bool> TopNComputer<Score, D, REVERSE_ORDER>
 where
    Score: PartialOrd + Clone,
    D: Serialize + DeserializeOwned + Ord + Clone,
@@ -808,7 +808,10 @@ where
    #[inline]
    pub fn push(&mut self, feature: Score, doc: D) {
        if let Some(last_median) = self.threshold.clone() {
-            if feature < last_median {
+            if !REVERSE_ORDER && feature > last_median {
+                return;
+            }
+            if REVERSE_ORDER && feature < last_median {
                return;
            }
        }
@@ -843,7 +846,7 @@ where
    }

    /// Returns the top n elements in sorted order.
-    pub fn into_sorted_vec(mut self) -> Vec<ComparableDoc<Score, D, R>> {
+    pub fn into_sorted_vec(mut self) -> Vec<ComparableDoc<Score, D, REVERSE_ORDER>> {
        if self.buffer.len() > self.top_n {
            self.truncate_top_n();
        }
@@ -854,7 +857,7 @@ where
    /// Returns the top n elements in stored order.
    /// Useful if you do not need the elements in sorted order,
    /// for example when merging the results of multiple segments.
-    pub fn into_vec(mut self) -> Vec<ComparableDoc<Score, D, R>> {
+    pub fn into_vec(mut self) -> Vec<ComparableDoc<Score, D, REVERSE_ORDER>> {
        if self.buffer.len() > self.top_n {
            self.truncate_top_n();
        }
@@ -864,9 +867,11 @@ where

 #[cfg(test)]
 mod tests {
+    use proptest::prelude::*;
+
    use super::{TopDocs, TopNComputer};
    use crate::collector::top_collector::ComparableDoc;
-    use crate::collector::Collector;
+    use crate::collector::{Collector, DocSetCollector};
    use crate::query::{AllQuery, Query, QueryParser};
    use crate::schema::{Field, Schema, FAST, STORED, TEXT};
    use crate::time::format_description::well_known::Rfc3339;
@@ -958,6 +963,44 @@ mod tests {
        }
    }

+    proptest! {
+        #[test]
+        fn test_topn_computer_asc_prop(
+          limit in 0..10_usize,
+          docs in proptest::collection::vec((0..100_u64, 0..100_u64), 0..100_usize),
+        ) {
+            let mut computer: TopNComputer<_, _, false> = TopNComputer::new(limit);
+            for (feature, doc) in &docs {
+                computer.push(*feature, *doc);
+            }
+            let mut comparable_docs = docs.into_iter().map(|(feature, doc)| ComparableDoc { feature, doc }).collect::<Vec<_>>();
+            comparable_docs.sort();
+            comparable_docs.truncate(limit);
+            prop_assert_eq!(
+                computer.into_sorted_vec(),
+                comparable_docs,
+            );
+        }
+
+        #[test]
+        fn test_topn_computer_desc_prop(
+          limit in 0..10_usize,
+          docs in proptest::collection::vec((0..100_u64, 0..100_u64), 0..100_usize),
+        ) {
+            let mut computer: TopNComputer<_, _, true> = TopNComputer::new(limit);
+            for (feature, doc) in &docs {
+                computer.push(*feature, *doc);
+            }
+            let mut comparable_docs = docs.into_iter().map(|(feature, doc)| ComparableDoc { feature, doc }).collect::<Vec<_>>();
+            comparable_docs.sort();
+            comparable_docs.truncate(limit);
+            prop_assert_eq!(
+                computer.into_sorted_vec(),
+                comparable_docs,
+            );
+        }
+    }
+
    #[test]
    fn test_top_collector_not_at_capacity_without_offset() -> crate::Result<()> {
        let index = make_index()?;
@@ -1371,4 +1414,29 @@ mod tests {
        );
        Ok(())
    }
+
+    #[test]
+    fn test_topn_computer_asc() {
+        let mut computer: TopNComputer<u32, u32, false> = TopNComputer::new(2);
+
+        computer.push(1u32, 1u32);
+        computer.push(2u32, 2u32);
+        computer.push(3u32, 3u32);
+        computer.push(2u32, 4u32);
+        computer.push(4u32, 5u32);
+        computer.push(1u32, 6u32);
+        assert_eq!(
+            computer.into_sorted_vec(),
+            &[
+                ComparableDoc {
+                    feature: 1u32,
+                    doc: 1u32,
+                },
+                ComparableDoc {
+                    feature: 1u32,
+                    doc: 6u32,
+                }
+            ]
+        );
+    }
 }
--- a/src/index/index.rs
+++ b/src/index/index.rs
@@ -83,7 +83,7 @@ fn save_new_metas(
 ///
 /// ```
 /// use tantivy::schema::*;
-/// use tantivy::{Index, IndexSettings, IndexSortByField, Order};
+/// use tantivy::{Index, IndexSettings};
 ///
 /// let mut schema_builder = Schema::builder();
 /// let id_field = schema_builder.add_text_field("id", STRING);
@@ -96,10 +96,7 @@ fn save_new_metas(
 ///
 /// let schema = schema_builder.build();
 /// let settings = IndexSettings{
-///     sort_by_field: Some(IndexSortByField{
-///         field: "number".to_string(),
-///         order: Order::Asc
-///     }),
+///     docstore_blocksize: 100_000,
 ///     ..Default::default()
 /// };
 /// let index = Index::builder().schema(schema).settings(settings).create_in_ram();
--- a/src/index/index_meta.rs
+++ b/src/index/index_meta.rs
@@ -288,6 +288,10 @@ impl Default for IndexSettings {
 /// Presorting documents can greatly improve performance
 /// in some scenarios, by applying top n
 /// optimizations.
+#[deprecated(
+    since = "0.22.0",
+    note = "We plan to remove index sorting in `0.23`. If you need index sorting, please comment on the related issue https://github.com/quickwit-oss/tantivy/issues/2352 and explain your use case."
+)]
 #[derive(Clone, Debug, Serialize, Deserialize, Eq, PartialEq)]
 pub struct IndexSortByField {
    /// The field to sort the documents by
--- a/src/lib.rs
+++ b/src/lib.rs
@@ -178,6 +178,7 @@ pub use crate::future_result::FutureResult;
 pub type Result<T> = std::result::Result<T, TantivyError>;

 mod core;
+#[allow(deprecated)] // Remove with index sorting
 pub mod indexer;

 #[allow(unused_doc_comments)]
@@ -189,6 +190,7 @@ pub mod collector;
 pub mod directory;
 pub mod fastfield;
 pub mod fieldnorm;
+#[allow(deprecated)] // Remove with index sorting
 pub mod index;
 pub mod positions;
 pub mod postings;
@@ -223,6 +225,7 @@ pub use self::snippet::{Snippet, SnippetGenerator};
 pub use crate::core::json_utils;
 pub use crate::core::{Executor, Searcher, SearcherGeneration};
 pub use crate::directory::Directory;
+#[allow(deprecated)] // Remove with index sorting
 pub use crate::index::{
    Index, IndexBuilder, IndexMeta, IndexSettings, IndexSortByField, InvertedIndexReader, Order,
    Segment, SegmentComponent, SegmentId, SegmentMeta, SegmentReader,
@@ -234,8 +237,6 @@ pub use crate::index::{
 pub use crate::indexer::PreparedCommit;
 pub use crate::indexer::{IndexWriter, SingleSegmentIndexWriter};
 pub use crate::postings::Postings;
-#[allow(deprecated)]
-pub use crate::schema::DatePrecision;
 pub use crate::schema::{DateOptions, DateTimePrecision, Document, TantivyDocument, Term};

 /// Index format version.
--- a/src/schema/date_time_options.rs
+++ b/src/schema/date_time_options.rs
@@ -1,7 +1,5 @@
 use std::ops::BitOr;

-#[allow(deprecated)]
-pub use common::DatePrecision;
 pub use common::DateTimePrecision;
 use serde::{Deserialize, Serialize};

--- a/src/schema/document/de.rs
+++ b/src/schema/document/de.rs
@@ -160,7 +160,7 @@ pub enum ValueType {
    /// A dynamic object value.
    Object,
    /// A JSON object value. Deprecated.
-    #[deprecated]
+    #[deprecated(note = "We keep this for backwards compatibility, use Object instead")]
    JSONObject,
 }

--- a/src/schema/document/owned_value.rs
+++ b/src/schema/document/owned_value.rs
@@ -1,4 +1,4 @@
-use std::collections::BTreeMap;
+use std::collections::{btree_map, BTreeMap};
 use std::fmt;
 use std::net::Ipv6Addr;

@@ -45,7 +45,7 @@ pub enum OwnedValue {
    /// A set of values.
    Array(Vec<Self>),
    /// Dynamic object value.
-    Object(Vec<(String, Self)>),
+    Object(BTreeMap<String, Self>),
    /// IpV6 Address. Internally there is no IpV4, it needs to be converted to `Ipv6Addr`.
    IpAddr(Ipv6Addr),
 }
@@ -148,10 +148,10 @@ impl ValueDeserialize for OwnedValue {

            fn visit_object<'de, A>(&self, mut access: A) -> Result<Self::Value, DeserializeError>
            where A: ObjectAccess<'de> {
-                let mut elements = Vec::new();
+                let mut elements = BTreeMap::new();

                while let Some((key, value)) = access.next_entry()? {
-                    elements.push((key, value));
+                    elements.insert(key, value);
                }

                Ok(OwnedValue::Object(elements))
@@ -248,13 +248,12 @@ impl<'de> serde::Deserialize<'de> for OwnedValue {

            fn visit_map<A>(self, mut map: A) -> Result<Self::Value, A::Error>
            where A: MapAccess<'de> {
-                let mut object =
-                    map.size_hint()
-                       .map(Vec::with_capacity)
-                       .unwrap_or_default();
+                let mut object = BTreeMap::new();
+
                while let Some((key, value)) = map.next_entry()? {
-                    object.push((key, value));
+                    object.insert(key, value);
                }
+
                Ok(OwnedValue::Object(object))
            }
        }
@@ -364,8 +363,7 @@ impl From<PreTokenizedString> for OwnedValue {

 impl From<BTreeMap<String, OwnedValue>> for OwnedValue {
    fn from(object: BTreeMap<String, OwnedValue>) -> OwnedValue {
-        let key_values = object.into_iter().collect();
-        OwnedValue::Object(key_values)
+        OwnedValue::Object(object)
    }
 }

@@ -419,15 +417,18 @@ impl From<serde_json::Value> for OwnedValue {

 impl From<serde_json::Map<String, serde_json::Value>> for OwnedValue {
    fn from(map: serde_json::Map<String, serde_json::Value>) -> Self {
-        let object: Vec<(String, OwnedValue)> = map.into_iter()
-            .map(|(key, value)| (key, OwnedValue::from(value)))
-            .collect();
+        let mut object = BTreeMap::new();
+
+        for (key, value) in map {
+            object.insert(key, OwnedValue::from(value));
+        }
+
        OwnedValue::Object(object)
    }
 }

 /// A wrapper type for iterating over a serde_json object producing reference values.
-pub struct ObjectMapIter<'a>(std::slice::Iter<'a, (String, OwnedValue)>);
+pub struct ObjectMapIter<'a>(btree_map::Iter<'a, String, OwnedValue>);

 impl<'a> Iterator for ObjectMapIter<'a> {
    type Item = (&'a str, &'a OwnedValue);
--- a/src/schema/mod.rs
+++ b/src/schema/mod.rs
@@ -130,8 +130,6 @@ mod text_options;
 use columnar::ColumnType;

 pub use self::bytes_options::BytesOptions;
-#[allow(deprecated)]
-pub use self::date_time_options::DatePrecision;
 pub use self::date_time_options::{DateOptions, DateTimePrecision, DATE_TIME_PRECISION_INDEXED};
 pub use self::document::{DocParsingError, Document, OwnedValue, TantivyDocument, Value};
 pub(crate) use self::facet::FACET_SEP_BYTE;
@@ -146,8 +144,6 @@ pub use self::index_record_option::IndexRecordOption;
 pub use self::ip_options::{IntoIpv6Addr, IpAddrOptions};
 pub use self::json_object_options::JsonObjectOptions;
 pub use self::named_field_document::NamedFieldDocument;
-#[allow(deprecated)]
-pub use self::numeric_options::IntOptions;
 pub use self::numeric_options::NumericOptions;
 pub use self::schema::{Schema, SchemaBuilder};
 pub use self::term::{Term, ValueBytes, JSON_END_OF_PATH};
--- a/src/schema/numeric_options.rs
+++ b/src/schema/numeric_options.rs
@@ -5,10 +5,6 @@ use serde::{Deserialize, Serialize};
 use super::flags::CoerceFlag;
 use crate::schema::flags::{FastFlag, IndexedFlag, SchemaFlagList, StoredFlag};

-#[deprecated(since = "0.17.0", note = "Use NumericOptions instead.")]
-/// Deprecated use [`NumericOptions`] instead.
-pub type IntOptions = NumericOptions;
-
 /// Define how an `u64`, `i64`, or `f64` field should be handled by tantivy.
 #[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize, Default)]
 #[serde(from = "NumericOptionsDeser")]
--- a/sstable/Cargo.toml
+++ b/sstable/Cargo.toml
@@ -1,6 +1,6 @@
 [package]
 name = "tantivy-sstable"
-version = "0.2.0"
+version = "0.3.0"
 edition = "2021"
 license = "MIT"
 homepage = "https://github.com/quickwit-oss/tantivy"
@@ -10,8 +10,8 @@ categories = ["database-implementations", "data-structures", "compression"]
 description = "sstables for tantivy"

 [dependencies]
-common = {version= "0.6", path="../common", package="tantivy-common"}
-tantivy-bitpacker = { version= "0.5", path="../bitpacker" }
+common = {version= "0.7", path="../common", package="tantivy-common"}
+tantivy-bitpacker = { version= "0.6", path="../bitpacker" }
 tantivy-fst = "0.5"
 # experimental gives us access to Decompressor::upper_bound
 zstd = { version = "0.13", features = ["experimental"] }
--- a/stacker/Cargo.toml
+++ b/stacker/Cargo.toml
@@ -1,6 +1,6 @@
 [package]
 name = "tantivy-stacker"
-version = "0.2.0"
+version = "0.3.0"
 edition = "2021"
 license = "MIT"
 homepage = "https://github.com/quickwit-oss/tantivy"
@@ -9,7 +9,7 @@ description = "term hashmap used for indexing"

 [dependencies]
 murmurhash32 = "0.3"
-common = { version = "0.6", path = "../common/", package = "tantivy-common" }
+common = { version = "0.7", path = "../common/", package = "tantivy-common" }
 ahash = { version = "0.8.11", default-features = false, optional = true }
 rand_distr = "0.4.3"

--- a/tokenizer-api/Cargo.toml
+++ b/tokenizer-api/Cargo.toml
@@ -1,6 +1,6 @@
 [package]
 name = "tantivy-tokenizer-api"
-version = "0.2.0"
+version = "0.3.0"
 license = "MIT"
 edition = "2021"
 description = "Tokenizer API of tantivy"
Author	SHA1	Message	Date
Pascal Seitz	2ef142ece3	bump to 0.22.1	2025-07-17 11:22:19 +08:00
Pascal Seitz	8791660ef5	Fix TopNComputer for reverse order	2025-07-17 10:54:21 +08:00
PSeitz	17d5869ad6	update CHANGELOG, use github API in cliff (#2354 ) * update CHANGELOG, use github API in cliff * reset version to 0.21.1, before release * chore: Release * remove unreleased from CHANGELOG	2024-04-15 10:07:20 +02:00
PSeitz	dfa3aed32d	check unsupported parameters top_hits (#2351 ) * check unsupported parameters top_hits * move to function	2024-04-10 08:20:52 +02:00
PSeitz	398817ce7b	add index sorting deprecation warning (#2353 ) * add index sorting deprecation warning * remove deprecated IntOptions and DatePrecision	2024-04-10 08:09:09 +02:00