fixing compilation

Removed 'static in compression_lz4.
Reorganized and added termdict unit tests.
2026-01-07 09:32:54 +00:00 · 2020-12-09 17:14:41 +09:00 · 2020-12-09 16:57:01 +09:00 · 2020-12-09 16:57:01 +09:00 · 2020-12-09 16:57:01 +09:00 · 2020-12-09 16:57:01 +09:00
41 changed files with 546 additions and 4386 deletions
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -1,23 +1,18 @@
 Tantivy 0.14.0
 =========================
- Remove dependency to atomicwrites #833 .Implemented by @fulmicoton upon suggestion and research from @asafigan).
+- Remove dependency to atomicwrites #833 .Implemented by @pmasurel upon suggestion and research from @asafigan). 
 - Migrated tantivy error from the now deprecated `failure` crate to `thiserror` #760. (@hirevo)
- API Change. Accessing the typed value off a `Schema::Value` now returns an Option instead of panicking if the type does not match.
+- API Change. Accessing the typed value off a `Schema::Value` now returns an Option instead of panicking if the type does not match. 
 - Large API Change in the Directory API. Tantivy used to assume that all files could be somehow memory mapped. After this change, Directory return a `FileSlice` that can be reduced and eventually read into an `OwnedBytes` object. Long and blocking io operation are still required by they do not span over the entire file.
 - Added support for Brotli compression in the DocStore. (@ppodolsky)
 - Added helper for building intersections and unions in BooleanQuery (@guilload)
 - Bugfix in `Query::explain`
 - Removed dependency on `notify` #924. Replaced with `FileWatcher` struct that polls meta file every 500ms in background thread. (@halvorboe @guilload)
 - Added `FilterCollector`, which wraps another collector and filters docs using a predicate over a fast field (@barrotsteindev)
- Simplified the encoding of the skip reader struct. BlockWAND max tf is now encoded over a single byte. (@fulmicoton)
- `FilterCollector` now supports all Fast Field value types (@barrotsteindev)
- FastField are not all loaded when opening the segment reader. (@fulmicoton)
-
-This version breaks compatibility and requires users to reindex everything.

 Tantivy 0.13.2
 ===================
-Bugfix. Acquiring a facet reader on a segment that does not contain any
+Bugfix. Acquiring a facet reader on a segment that does not contain any 
 doc with this facet returns `None`. (#896)

 Tantivy 0.13.1
@@ -28,7 +23,7 @@ Updated misc dependency versions.
 Tantivy 0.13.0
 ======================
 Tantivy 0.13 introduce a change in the index format that will require
-you to reindex your index (BlockWAND information are added in the skiplist).
+you to reindex your index (BlockWAND information are added in the skiplist). 
 The index size increase is minor as this information is only added for
 full blocks.
 If you have a massive index for which reindexing is not an option, please contact me
@@ -37,7 +32,7 @@ so that we can discuss possible solutions.
 - Bugfix in `FuzzyTermQuery` not matching terms by prefix when it should (@Peachball)
 - Relaxed constraints on the custom/tweak score functions. At the segment level, they can be mut, and they are not required to be Sync + Send.
 - `MMapDirectory::open` does not return a `Result` anymore.
- Change in the DocSet and Scorer API. (@fulmicoton).
+- Change in the DocSet and Scorer API. (@fulmicoton). 
 A freshly created DocSet point directly to their first doc. A sentinel value called TERMINATED marks the end of a DocSet.
 `.advance()` returns the new DocId. `Scorer::skip(target)` has been replaced by `Scorer::seek(target)` and returns the resulting DocId.
 As a result, iterating through DocSet now looks as follows
@@ -51,7 +46,7 @@ while doc != TERMINATED {
 The change made it possible to greatly simplify a lot of the docset's code.
 - Misc internal optimization and introduction of the `Scorer::for_each_pruning` function. (@fulmicoton)
 - Added an offset option to the Top(.*)Collectors. (@robyoung)
- Added Block WAND. Performance on TOP-K on term-unions should be greatly increased. (@fulmicoton, and special thanks
+- Added Block WAND. Performance on TOP-K on term-unions should be greatly increased. (@fulmicoton, and special thanks 
 to the PISA team for answering all my questions!)

 Tantivy 0.12.0
@@ -59,14 +54,14 @@ Tantivy 0.12.0
 - Removing static dispatch in tokenizers for simplicity. (#762)
 - Added backward iteration for `TermDictionary` stream. (@halvorboe)
 - Fixed a performance issue when searching for the posting lists of a missing term (@audunhalland)
- Added a configurable maximum number of docs (10M by default) for a segment to be considered for merge (@hntd187, landed by @halvorboe #713)
+- Added a configurable maximum number of docs (10M by default) for a segment to be considered for merge (@hntd187, landed by @halvorboe #713) 
 - Important Bugfix #777, causing tantivy to retain memory mapping. (diagnosed by @poljar)
 - Added support for field boosting. (#547, @fulmicoton)

 ## How to update?

-Crates relying on custom tokenizer, or registering tokenizer in the manager will require some
-minor changes. Check https://github.com/tantivy-search/tantivy/blob/main/examples/custom_tokenizer.rs
+Crates relying on custom tokenizer, or registering tokenizer in the manager will require some 
+minor changes. Check https://github.com/tantivy-search/tantivy/blob/master/examples/custom_tokenizer.rs
 to check for some code sample.

 Tantivy 0.11.3
@@ -102,7 +97,7 @@ Tantivy 0.11.0

 ## How to update?

- The index format is changed. You are required to reindex your data to use tantivy 0.11.
+- The index format is changed. You are required to reindex your data to use tantivy 0.11. 
 - `Box<dyn BoxableTokenizer>` has been replaced by a `BoxedTokenizer` struct.
 - Regex are now compiled when the `RegexQuery` instance is built. As a result, it can now return
 an error and handling the `Result` is required.
@@ -126,26 +121,26 @@ Tantivy 0.10.0

 *Tantivy 0.10.0 index format is compatible with the index format in 0.9.0.*

- Added an API to easily tweak or entirely replace the
- default score. See `TopDocs::tweak_score`and `TopScore::custom_score` (@fulmicoton)
+- Added an API to easily tweak or entirely replace the 
+ default score. See `TopDocs::tweak_score`and `TopScore::custom_score` (@pmasurel)
 - Added an ASCII folding filter (@drusellers)
- Bugfix in `query.count` in presence of deletes (@fulmicoton)
- Added `.explain(...)` in `Query` and `Weight` to (@fulmicoton)
- Added an efficient way to `delete_all_documents` in `IndexWriter` (@petr-tik).
+- Bugfix in `query.count` in presence of deletes (@pmasurel)
+- Added `.explain(...)` in `Query` and `Weight` to (@pmasurel)
+- Added an efficient way to `delete_all_documents` in `IndexWriter` (@petr-tik). 
  All segments are simply removed.

 Minor
 ---------
 - Switched to Rust 2018 (@uvd)
- Small simplification of the code.
+- Small simplification of the code. 
 Calling .freq() or .doc() when .advance() has never been called
 on segment postings should panic from now on.
 - Tokens exceeding `u16::max_value() - 4` chars are discarded silently instead of panicking.
 - Fast fields are now preloaded when the `SegmentReader` is created.
 - `IndexMeta` is now public.  (@hntd187)
 - `IndexWriter` `add_document`, `delete_term`. `IndexWriter` is `Sync`, making it possible to use it with a `
-Arc<RwLock<IndexWriter>>`. `add_document` and `delete_term` can
-only require a read lock. (@fulmicoton)
+Arc<RwLock<IndexWriter>>`. `add_document` and `delete_term` can 
+only require a read lock. (@pmasurel)
 - Introducing `Opstamp` as an expressive type alias for `u64`. (@petr-tik)
 - Stamper now relies on `AtomicU64` on all platforms (@petr-tik)
 - Bugfix - Files get deleted slightly earlier
@@ -159,7 +154,7 @@ Your program should be usable as is.

 Fast fields used to be accessed directly from the `SegmentReader`.
 The API changed, you are now required to acquire your fast field reader via the
-`segment_reader.fast_fields()`, and use one of the typed method:
+`segment_reader.fast_fields()`, and use one of the typed method: 
 - `.u64()`, `.i64()` if your field is single-valued ;
 - `.u64s()`, `.i64s()` if your field is multi-valued ;
 - `.bytes()` if your field is bytes fast field.
@@ -168,16 +163,16 @@ The API changed, you are now required to acquire your fast field reader via the

 Tantivy 0.9.0
 =====================
-*0.9.0 index format is not compatible with the
+*0.9.0 index format is not compatible with the 
 previous index format.*
- MAJOR BUGFIX :
+- MAJOR BUGFIX : 
  Some `Mmap` objects were being leaked, and would never get released. (@fulmicoton)
 - Removed most unsafe (@fulmicoton)
 - Indexer memory footprint improved. (VInt comp, inlining the first block. (@fulmicoton)
 - Stemming in other language possible (@pentlander)
 - Segments with no docs are deleted earlier (@barrotsteindev)
- Added grouped add and delete operations.
-  They are guaranteed to happen together (i.e. they cannot be split by a commit).
+- Added grouped add and delete operations. 
+  They are guaranteed to happen together (i.e. they cannot be split by a commit). 
  In addition, adds are guaranteed to happen on the same segment. (@elbow-jason)
 - Removed `INT_STORED` and `INT_INDEXED`. It is now possible to use `STORED` and `INDEXED`
  for int fields. (@fulmicoton)
@@ -191,26 +186,26 @@ tantivy 0.9 brought some API breaking change.
 To update from tantivy 0.8, you will need to go through the following steps.

 - `schema::INT_INDEXED` and `schema::INT_STORED`  should be replaced by `schema::INDEXED` and `schema::INT_STORED`.
- The index now does not hold the pool of searcher anymore. You are required to create an intermediary object called
-`IndexReader` for this.
-
+- The index now does not hold the pool of searcher anymore. You are required to create an intermediary object called 
+`IndexReader` for this. 
+    
    ```rust
    // create the reader. You typically need to create 1 reader for the entire
    // lifetime of you program.
    let reader = index.reader()?;
-
+    
    // Acquire a searcher (previously `index.searcher()`) is now written:
    let searcher = reader.searcher();
-
-    // With the default setting of the reader, you are not required to
+    
+    // With the default setting of the reader, you are not required to 
    // call `index.load_searchers()` anymore.
    //
    // The IndexReader will pick up that change automatically, regardless
    // of whether the update was done in a different process or not.
-    // If this behavior is not wanted, you can create your reader with
+    // If this behavior is not wanted, you can create your reader with 
    // the `ReloadPolicy::Manual`, and manually decide when to reload the index
    // by calling `reader.reload()?`.
-
+  
    ```


@@ -225,7 +220,7 @@ Tantivy 0.8.1
 =====================
 Hotfix of #476.

-Merge was reflecting deletes before commit was passed.
+Merge was reflecting deletes before commit was passed. 
 Thanks @barrotsteindev  for reporting the bug.


@@ -233,7 +228,7 @@ Tantivy 0.8.0
 =====================
 *No change in the index format*
 - API Breaking change in the collector API. (@jwolfe, @fulmicoton)
- Multithreaded search (@jwolfe, @fulmicoton)
+- Multithreaded search (@jwolfe, @fulmicoton) 


 Tantivy 0.7.1
@@ -261,7 +256,7 @@ Tantivy 0.6.1
        - Exclusive `field:{startExcl to endExcl}`
        - Mixed `field:[startIncl to endExcl}` and vice versa
        - Unbounded `field:[start to *]`, `field:[* to end]`
-
+ 

 Tantivy 0.6
 ==========================
@@ -269,10 +264,10 @@ Tantivy 0.6
 Special thanks to @drusellers and @jason-wolfe for their contributions
 to this release!

- Removed C code. Tantivy is now pure Rust. (@fulmicoton)
- BM25 (@fulmicoton)
- Approximate field norms encoded over 1 byte. (@fulmicoton)
- Compiles on stable rust (@fulmicoton)
+- Removed C code. Tantivy is now pure Rust. (@pmasurel)
+- BM25 (@pmasurel)
+- Approximate field norms encoded over 1 byte. (@pmasurel)
+- Compiles on stable rust (@pmasurel)
 - Add &[u8] fastfield for associating arbitrary bytes to each document (@jason-wolfe) (#270)
    - Completely uncompressed
    - Internally: One u64 fast field for indexes, one fast field for the bytes themselves.
@@ -280,7 +275,7 @@ to this release!
 - Add Stopword Filter support (@drusellers)
 - Add a FuzzyTermQuery (@drusellers)
 - Add a RegexQuery (@drusellers)
- Various performance improvements (@fulmicoton)_
+- Various performance improvements (@pmasurel)_


 Tantivy 0.5.2
--- a/Cargo.toml
+++ b/Cargo.toml
@@ -1,6 +1,6 @@
 [package]
 name = "tantivy"
-version = "0.14.0"
+version = "0.14.0-dev"
 authors = ["Paul Masurel <paul.masurel@gmail.com>"]
 license = "MIT"
 categories = ["database-implementations", "data-structures"]
@@ -33,7 +33,7 @@ levenshtein_automata = "0.2"
 uuid = { version = "0.8", features = ["v4", "serde"] }
 crossbeam = "0.8"
 futures = {version = "0.3",  features=["thread-pool"] }
-tantivy-query-grammar = { version="0.14.0", path="./query-grammar" }
+tantivy-query-grammar = { version="0.14.0-dev", path="./query-grammar" }
 stable_deref_trait = "1"
 rust-stemmers = "1"
 downcast-rs = "1"
@@ -53,11 +53,10 @@ lru = "0.6"
 winapi = "0.3"

 [dev-dependencies]
-rand = "0.8"
+rand = "0.7"
 maplit = "1"
 matches = "0.1.8"
 proptest = "0.10"
-criterion = "0.3"

 [dev-dependencies.fail]
 version = "0.4"
@@ -98,7 +97,3 @@ travis-ci = { repository = "tantivy-search/tantivy" }
 name = "failpoints"
 path = "tests/failpoints/mod.rs"
 required-features = ["fail/failpoints"]
-
-[[bench]]
-name = "analyzer"
-harness = false
--- a/README.md
+++ b/README.md
@@ -1,9 +1,9 @@

-[![Build Status](https://travis-ci.org/tantivy-search/tantivy.svg?branch=main)](https://travis-ci.org/tantivy-search/tantivy)
-[![codecov](https://codecov.io/gh/tantivy-search/tantivy/branch/main/graph/badge.svg)](https://codecov.io/gh/tantivy-search/tantivy)
+[![Build Status](https://travis-ci.org/tantivy-search/tantivy.svg?branch=master)](https://travis-ci.org/tantivy-search/tantivy)
+[![codecov](https://codecov.io/gh/tantivy-search/tantivy/branch/master/graph/badge.svg)](https://codecov.io/gh/tantivy-search/tantivy)
 [![Join the chat at https://gitter.im/tantivy-search/tantivy](https://badges.gitter.im/tantivy-search/tantivy.svg)](https://gitter.im/tantivy-search/tantivy?utm_source=badge&utm_medium=badge&utm_campaign=pr-badge&utm_content=badge)
 [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
-[![Build status](https://ci.appveyor.com/api/projects/status/r7nb13kj23u8m9pj/branch/main?svg=true)](https://ci.appveyor.com/project/fulmicoton/tantivy/branch/main)
+[![Build status](https://ci.appveyor.com/api/projects/status/r7nb13kj23u8m9pj/branch/master?svg=true)](https://ci.appveyor.com/project/fulmicoton/tantivy/branch/master)
 [![Crates.io](https://img.shields.io/crates/v/tantivy.svg)](https://crates.io/crates/tantivy)

 ![Tantivy](https://tantivy-search.github.io/logo/tantivy-logo.png)
--- a/benches/alice.txt
+++ b/benches/alice.txt
--- a/benches/analyzer.rs
+++ b/benches/analyzer.rs
@@ -1,22 +0,0 @@
-use criterion::{criterion_group, criterion_main, Criterion};
-use tantivy::tokenizer::TokenizerManager;
-
-const ALICE_TXT: &'static str = include_str!("alice.txt");
-
-pub fn criterion_benchmark(c: &mut Criterion) {
-    let tokenizer_manager = TokenizerManager::default();
-    let tokenizer = tokenizer_manager.get("default").unwrap();
-    c.bench_function("default-tokenize-alice", |b| {
-        b.iter(|| {
-            let mut word_count = 0;
-            let mut token_stream = tokenizer.token_stream(ALICE_TXT);
-            while token_stream.advance() {
-                word_count += 1;
-            }
-            assert_eq!(word_count, 30_731);
-        })
-    });
-}
-
-criterion_group!(benches, criterion_benchmark);
-criterion_main!(benches);
--- a/examples/custom_collector.rs
+++ b/examples/custom_collector.rs
@@ -14,7 +14,7 @@ use tantivy::fastfield::FastFieldReader;
 use tantivy::query::QueryParser;
 use tantivy::schema::Field;
 use tantivy::schema::{Schema, FAST, INDEXED, TEXT};
-use tantivy::{doc, Index, Score, SegmentReader};
+use tantivy::{doc, Index, Score, SegmentReader, TantivyError};

 #[derive(Default)]
 struct Stats {
@@ -72,7 +72,16 @@ impl Collector for StatsCollector {
        _segment_local_id: u32,
        segment_reader: &SegmentReader,
    ) -> tantivy::Result<StatsSegmentCollector> {
-        let fast_field_reader = segment_reader.fast_fields().u64(self.field)?;
+        let fast_field_reader = segment_reader
+            .fast_fields()
+            .u64(self.field)
+            .ok_or_else(|| {
+                let field_name = segment_reader.schema().get_field_name(self.field);
+                TantivyError::SchemaError(format!(
+                    "Field {:?} is not a u64 fast field.",
+                    field_name
+                ))
+            })?;
        Ok(StatsSegmentCollector {
            fast_field_reader,
            stats: Stats::default(),
--- a/query-grammar/Cargo.toml
+++ b/query-grammar/Cargo.toml
@@ -1,6 +1,6 @@
 [package]
 name = "tantivy-query-grammar"
-version = "0.14.0"
+version = "0.14.0-dev"
 authors = ["Paul Masurel <paul.masurel@gmail.com>"]
 license = "MIT"
 categories = ["database-implementations", "data-structures"]
--- a/src/collector/facet_collector.rs
+++ b/src/collector/facet_collector.rs
@@ -398,8 +398,6 @@ impl<'a> Iterator for FacetChildIterator<'a> {
 }

 impl FacetCounts {
-    /// Returns an iterator over all of the facet count pairs inside this result.
-    /// See the documentation for `FacetCollector` for a usage example.
    pub fn get<T>(&self, facet_from: T) -> FacetChildIterator<'_>
    where
        Facet: From<T>,
@@ -419,8 +417,6 @@ impl FacetCounts {
        FacetChildIterator { underlying }
    }

-    /// Returns a vector of top `k` facets with their counts, sorted highest-to-lowest by counts.
-    /// See the documentation for `FacetCollector` for a usage example.
    pub fn top_k<T>(&self, facet: T, k: usize) -> Vec<(&Facet, u64)>
    where
        Facet: From<T>,
--- a/src/collector/filter_collector_wrapper.rs
+++ b/src/collector/filter_collector_wrapper.rs
@@ -11,9 +11,9 @@
 // Importing tantivy...
 use std::marker::PhantomData;

-use crate::collector::{Collector, SegmentCollector};
-use crate::fastfield::{FastFieldReader, FastValue};
+use crate::fastfield::{FastValue, FastFieldReader};
 use crate::schema::Field;
+use crate::collector::{Collector, SegmentCollector};
 use crate::{Score, SegmentReader, TantivyError};

 /// The `FilterCollector` collector filters docs using a u64 fast field value and a predicate.
@@ -22,20 +22,23 @@ use crate::{Score, SegmentReader, TantivyError};
 /// ```rust
 /// use tantivy::collector::{TopDocs, FilterCollector};
 /// use tantivy::query::QueryParser;
-/// use tantivy::schema::{Schema, TEXT, INDEXED, FAST};
+/// use tantivy::schema::{Schema, FAST, TEXT};
+/// use tantivy::DateTime;
+/// use std::str::FromStr;
 /// use tantivy::{doc, DocAddress, Index};
 ///
 /// let mut schema_builder = Schema::builder();
 /// let title = schema_builder.add_text_field("title", TEXT);
-/// let price = schema_builder.add_u64_field("price", INDEXED | FAST);
+/// let price = schema_builder.add_u64_field("price", FAST);
+/// let date = schema_builder.add_date_field("date", FAST);
 /// let schema = schema_builder.build();
 /// let index = Index::create_in_ram(schema);
 ///
 /// let mut index_writer = index.writer_with_num_threads(1, 10_000_000).unwrap();
-/// index_writer.add_document(doc!(title => "The Name of the Wind", price => 30_200u64));
-/// index_writer.add_document(doc!(title => "The Diary of Muadib", price => 29_240u64));
-/// index_writer.add_document(doc!(title => "A Dairy Cow", price => 21_240u64));
-/// index_writer.add_document(doc!(title => "The Diary of a Young Girl", price => 20_120u64));
+/// index_writer.add_document(doc!(title => "The Name of the Wind", price => 30_200u64, date => DateTime::from_str("1898-04-09T00:00:00+00:00").unwrap()));
+/// index_writer.add_document(doc!(title => "The Diary of Muadib", price => 29_240u64, date => DateTime::from_str("2020-04-09T00:00:00+00:00").unwrap()));
+/// index_writer.add_document(doc!(title => "A Dairy Cow", price => 21_240u64, date => DateTime::from_str("2019-04-09T00:00:00+00:00").unwrap()));
+/// index_writer.add_document(doc!(title => "The Diary of a Young Girl", price => 20_120u64, date => DateTime::from_str("2018-04-09T00:00:00+00:00").unwrap()));
 /// assert!(index_writer.commit().is_ok());
 ///
 /// let reader = index.reader().unwrap();
@@ -43,8 +46,8 @@ use crate::{Score, SegmentReader, TantivyError};
 ///
 /// let query_parser = QueryParser::for_index(&index, vec![title]);
 /// let query = query_parser.parse_query("diary").unwrap();
-/// let no_filter_collector = FilterCollector::new(price, &|value: u64| value > 20_120u64, TopDocs::with_limit(2));
-/// let top_docs = searcher.search(&query, &no_filter_collector).unwrap();
+/// let filter_some_collector = FilterCollector::new(price, &|value: u64| value > 20_120u64, TopDocs::with_limit(2));
+/// let top_docs = searcher.search(&query, &filter_some_collector).unwrap();
 ///
 /// assert_eq!(top_docs.len(), 1);
 /// assert_eq!(top_docs[0].1, DocAddress(0, 1));
@@ -53,6 +56,17 @@ use crate::{Score, SegmentReader, TantivyError};
 /// let filtered_top_docs = searcher.search(&query, &filter_all_collector).unwrap();
 ///
 /// assert_eq!(filtered_top_docs.len(), 0);
+/// 
+/// fn date_debug(value: DateTime) -> bool {
+///     println!("date: {:?}", value);
+///     assert_eq!(value, DateTime::from_str("1000-04-09T00:00:00+00:00").unwrap());
+///     (value - DateTime::from_str("2019-04-09T00:00:00+00:00").unwrap()).num_weeks() > 0
+/// }
+/// 
+/// let filter_dates_collector = FilterCollector::new(date, &date_debug, TopDocs::with_limit(2));
+/// let filtered_date_docs = searcher.search(&query, &filter_all_collector).unwrap();
+///
+/// assert_eq!(filtered_date_docs.len(), 5);
 /// ```
 pub struct FilterCollector<TCollector, TPredicate, TPredicateValue: FastValue>
 where
@@ -111,28 +125,39 @@ where
                field_entry.name()
            )));
        }
-        let requested_type = TPredicateValue::to_type();
-        let field_schema_type = field_entry.field_type().value_type();
-        if requested_type != field_schema_type {
+        let schema_type = TPredicateValue::to_type();
+        let requested_type = field_entry.field_type().value_type();
+        if schema_type != requested_type {
            return Err(TantivyError::SchemaError(format!(
                "Field {:?} is of type {:?}!={:?}",
                field_entry.name(),
-                requested_type,
-                field_schema_type
+                schema_type,
+                requested_type
            )));
        }

-        let fast_field_reader = segment_reader
-            .fast_fields()
-            .typed_fast_field_reader(self.field)?;
+        let err_closure = || {
+            let field_name = segment_reader.schema().get_field_name(self.field);
+            TantivyError::SchemaError(format!(
+                "Field {:?} is not a u64 fast field.",
+                field_name
+            ))
+        };
+        let fast_fields = segment_reader.fast_fields();

+        let fast_value_type = TPredicateValue::to_type();
+        // TODO  do a runtime check of `fast_value_type` against the schema.
+
+        let fast_field_reader_opt = fast_fields.typed_fast_field_reader(self.field);
+        let fast_field_reader = fast_field_reader_opt
+            .ok_or_else(|| TantivyError::SchemaError(format!("{:?} is not declared as a fast field in the schema.", self.field)))?;
        let segment_collector = self
            .collector
            .for_segment(segment_local_id, segment_reader)?;

        Ok(FilterSegmentCollector {
-            fast_field_reader,
-            segment_collector,
+            fast_field_reader ,
+            segment_collector: segment_collector,
            predicate: self.predicate,
            t_predicate_value: PhantomData,
        })
--- a/src/collector/mod.rs
+++ b/src/collector/mod.rs
@@ -109,7 +109,6 @@ pub use self::tweak_score_top_collector::{ScoreSegmentTweaker, ScoreTweaker};

 mod facet_collector;
 pub use self::facet_collector::FacetCollector;
-pub use self::facet_collector::FacetCounts;
 use crate::query::Weight;

 mod docset_collector;
--- a/src/collector/tests.rs
+++ b/src/collector/tests.rs
@@ -8,12 +8,12 @@ use crate::DocId;
 use crate::Score;
 use crate::SegmentLocalId;

-use crate::collector::{FilterCollector, TopDocs};
+use crate::collector::{TopDocs, FilterCollector};
 use crate::query::QueryParser;
 use crate::schema::{Schema, FAST, TEXT};
 use crate::DateTime;
-use crate::{doc, Index};
 use std::str::FromStr;
+use crate::{doc, Index};

 pub const TEST_COLLECTOR_WITH_SCORE: TestCollector = TestCollector {
    compute_score: true,
@@ -25,6 +25,7 @@ pub const TEST_COLLECTOR_WITHOUT_SCORE: TestCollector = TestCollector {

 #[test]
 pub fn test_filter_collector() {
+
    let mut schema_builder = Schema::builder();
    let title = schema_builder.add_text_field("title", TEXT);
    let price = schema_builder.add_u64_field("price", FAST);
@@ -35,7 +36,6 @@ pub fn test_filter_collector() {
    let mut index_writer = index.writer_with_num_threads(1, 10_000_000).unwrap();
    index_writer.add_document(doc!(title => "The Name of the Wind", price => 30_200u64, date => DateTime::from_str("1898-04-09T00:00:00+00:00").unwrap()));
    index_writer.add_document(doc!(title => "The Diary of Muadib", price => 29_240u64, date => DateTime::from_str("2020-04-09T00:00:00+00:00").unwrap()));
-    index_writer.add_document(doc!(title => "The Diary of Anne Frank", price => 18_240u64, date => DateTime::from_str("2019-04-20T00:00:00+00:00").unwrap()));
    index_writer.add_document(doc!(title => "A Dairy Cow", price => 21_240u64, date => DateTime::from_str("2019-04-09T00:00:00+00:00").unwrap()));
    index_writer.add_document(doc!(title => "The Diary of a Young Girl", price => 20_120u64, date => DateTime::from_str("2018-04-09T00:00:00+00:00").unwrap()));
    assert!(index_writer.commit().is_ok());
@@ -45,30 +45,27 @@ pub fn test_filter_collector() {

    let query_parser = QueryParser::for_index(&index, vec![title]);
    let query = query_parser.parse_query("diary").unwrap();
-    let filter_some_collector = FilterCollector::new(
-        price,
-        &|value: u64| value > 20_120u64,
-        TopDocs::with_limit(2),
-    );
+    let filter_some_collector = FilterCollector::new(price, &|value: u64| value > 20_120u64, TopDocs::with_limit(2));
    let top_docs = searcher.search(&query, &filter_some_collector).unwrap();

    assert_eq!(top_docs.len(), 1);
    assert_eq!(top_docs[0].1, DocAddress(0, 1));

-    let filter_all_collector: FilterCollector<_, _, u64> =
-        FilterCollector::new(price, &|value| value < 5u64, TopDocs::with_limit(2));
+    let filter_all_collector: FilterCollector<_, _, u64> = FilterCollector::new(price, &|value| value < 5u64, TopDocs::with_limit(2));
    let filtered_top_docs = searcher.search(&query, &filter_all_collector).unwrap();

    assert_eq!(filtered_top_docs.len(), 0);

-    fn date_filter(value: DateTime) -> bool {
-        (value - DateTime::from_str("2019-04-09T00:00:00+00:00").unwrap()).num_weeks() > 0
+    fn date_debug(value: DateTime) -> bool {
+    println!("date: {:?}", value);
+    assert_eq!(value, DateTime::from_str("1000-04-09T00:00:00+00:00").unwrap());
+    (value - DateTime::from_str("2019-04-09T00:00:00+00:00").unwrap()).num_weeks() > 0
    }

-    let filter_dates_collector = FilterCollector::new(date, &date_filter, TopDocs::with_limit(5));
+    let filter_dates_collector = FilterCollector::new(date, &date_debug, TopDocs::with_limit(2));
    let filtered_date_docs = searcher.search(&query, &filter_dates_collector).unwrap();

-    assert_eq!(filtered_date_docs.len(), 2);
+    assert_eq!(filtered_date_docs.len(), 5);
 }

 /// Stores all of the doc ids.
@@ -240,7 +237,12 @@ impl Collector for BytesFastFieldTestCollector {
        _segment_local_id: u32,
        segment_reader: &SegmentReader,
    ) -> crate::Result<BytesFastFieldSegmentCollector> {
-        let reader = segment_reader.fast_fields().bytes(self.field)?;
+        let reader = segment_reader
+            .fast_fields()
+            .bytes(self.field)
+            .ok_or_else(|| {
+                crate::TantivyError::InvalidArgument("Field is not a bytes fast field.".to_string())
+            })?;
        Ok(BytesFastFieldSegmentCollector {
            vals: Vec::new(),
            reader,
--- a/src/collector/top_collector.rs
+++ b/src/collector/top_collector.rs
@@ -2,9 +2,9 @@ use crate::DocAddress;
 use crate::DocId;
 use crate::SegmentLocalId;
 use crate::SegmentReader;
+use serde::export::PhantomData;
 use std::cmp::Ordering;
 use std::collections::BinaryHeap;
-use std::marker::PhantomData;

 /// Contains a feature (field, score, etc.) of a document along with the document address.
 ///
--- a/src/collector/top_score_collector.rs
+++ b/src/collector/top_score_collector.rs
@@ -146,14 +146,15 @@ impl CustomScorer<u64> for ScorerByField {
    type Child = ScorerByFastFieldReader;

    fn segment_scorer(&self, segment_reader: &SegmentReader) -> crate::Result<Self::Child> {
-        // We interpret this field as u64, regardless of its type, that way,
-        // we avoid needless conversion. Regardless of the fast field type, the
-        // mapping is monotonic, so it is sufficient to compute our top-K docs.
-        //
-        // The conversion will then happen only on the top-K docs.
-        let ff_reader: FastFieldReader<u64> = segment_reader
+        let ff_reader = segment_reader
            .fast_fields()
-            .typed_fast_field_reader(self.field)?;
+            .u64_lenient(self.field)
+            .ok_or_else(|| {
+                crate::TantivyError::SchemaError(format!(
+                    "Field requested ({:?}) is not a fast field.",
+                    self.field
+                ))
+            })?;
        Ok(ScorerByFastFieldReader { ff_reader })
    }
 }
@@ -231,7 +232,7 @@ impl TopDocs {
    /// #   let title = schema_builder.add_text_field("title", TEXT);
    /// #   let rating = schema_builder.add_u64_field("rating", FAST);
    /// #   let schema = schema_builder.build();
-    /// #
+    /// #  
    /// #   let index = Index::create_in_ram(schema);
    /// #   let mut index_writer = index.writer_with_num_threads(1, 10_000_000)?;
    /// #   index_writer.add_document(doc!(title => "The Name of the Wind", rating => 92u64));
@@ -261,7 +262,7 @@ impl TopDocs {
    ///     let top_books_by_rating = TopDocs
    ///                 ::with_limit(10)
    ///                  .order_by_u64_field(rating_field);
-    ///
+    ///     
    ///     // ... and here are our documents. Note this is a simple vec.
    ///     // The `u64` in the pair is the value of our fast field for
    ///     // each documents.
@@ -271,13 +272,13 @@ impl TopDocs {
    ///     // query.
    ///     let resulting_docs: Vec<(u64, DocAddress)> =
    ///          searcher.search(query, &top_books_by_rating)?;
-    ///
+    ///     
    ///     Ok(resulting_docs)
    /// }
    /// ```
    ///
    /// # See also
-    ///
+    ///  
    /// To confortably work with `u64`s, `i64`s, `f64`s, or `date`s, please refer to
    /// [.order_by_fast_field(...)](#method.order_by_fast_field) method.
    pub fn order_by_u64_field(
@@ -289,7 +290,7 @@ impl TopDocs {

    /// Set top-K to rank documents by a given fast field.
    ///
-    /// If the field is not a fast field, or its field type does not match the generic type, this method does not panic,
+    /// If the field is not a fast field, or its field type does not match the generic type, this method does not panic,  
    /// but an explicit error will be returned at the moment of collection.
    ///
    /// Note that this method is a generic. The requested fast field type will be often
@@ -313,7 +314,7 @@ impl TopDocs {
    /// #   let title = schema_builder.add_text_field("company", TEXT);
    /// #   let rating = schema_builder.add_i64_field("revenue", FAST);
    /// #   let schema = schema_builder.build();
-    /// #
+    /// #  
    /// #   let index = Index::create_in_ram(schema);
    /// #   let mut index_writer = index.writer_with_num_threads(1, 10_000_000)?;
    /// #   index_writer.add_document(doc!(title => "MadCow Inc.", rating => 92_000_000i64));
@@ -342,7 +343,7 @@ impl TopDocs {
    ///     let top_company_by_revenue = TopDocs
    ///                 ::with_limit(2)
    ///                  .order_by_fast_field(revenue_field);
-    ///
+    ///     
    ///     // ... and here are our documents. Note this is a simple vec.
    ///     // The `i64` in the pair is the value of our fast field for
    ///     // each documents.
@@ -352,7 +353,7 @@ impl TopDocs {
    ///     // query.
    ///     let resulting_docs: Vec<(i64, DocAddress)> =
    ///          searcher.search(query, &top_company_by_revenue)?;
-    ///
+    ///     
    ///     Ok(resulting_docs)
    /// }
    /// ```
@@ -391,7 +392,7 @@ impl TopDocs {
    ///
    /// In the following example will will tweak our ranking a bit by
    /// boosting popular products a notch.
-    ///
+    ///  
    /// In more serious application, this tweaking could involved running a
    /// learning-to-rank model over various features
    ///
@@ -522,7 +523,7 @@ impl TopDocs {
    /// #   let index = Index::create_in_ram(schema);
    /// #   let mut index_writer = index.writer_with_num_threads(1, 10_000_000)?;
    /// #   let product_name = index.schema().get_field("product_name").unwrap();
-    /// #
+    /// #   
    /// let popularity: Field = index.schema().get_field("popularity").unwrap();
    /// let boosted: Field = index.schema().get_field("boosted").unwrap();
    /// #   index_writer.add_document(doc!(boosted=>1u64, product_name => "The Diary of Muadib", popularity => 1u64));
@@ -556,7 +557,7 @@ impl TopDocs {
    ///                 segment_reader.fast_fields().u64(popularity).unwrap();
    ///             let boosted_reader =
    ///                 segment_reader.fast_fields().u64(boosted).unwrap();
-    ///
+    ///    
    ///             // We can now define our actual scoring function
    ///             move |doc: DocId| {
    ///                 let popularity: u64 = popularity_reader.get(doc);
@@ -993,7 +994,9 @@ mod tests {
        let segment = searcher.segment_reader(0);
        let top_collector = TopDocs::with_limit(4).order_by_u64_field(size);
        let err = top_collector.for_segment(0, segment).err().unwrap();
-        assert!(matches!(err, crate::TantivyError::SchemaError(_)));
+        assert!(
+            matches!(err, crate::TantivyError::SchemaError(msg) if msg == "Field requested (Field(0)) is not a fast field.")
+        );
        Ok(())
    }

--- a/src/common/mod.rs
+++ b/src/common/mod.rs
@@ -115,16 +115,11 @@ pub fn u64_to_i64(val: u64) -> i64 {
 /// For simplicity, tantivy internally handles `f64` as `u64`.
 /// The mapping is defined by this function.
 ///
-/// Maps `f64` to `u64` in a monotonic manner, so that bytes lexical order is preserved.
+/// Maps `f64` to `u64` so that lexical order is preserved.
 ///
 /// This is more suited than simply casting (`val as u64`)
 /// which would truncate the result
 ///
-/// # Reference
-///
-/// Daniel Lemire's [blog post](https://lemire.me/blog/2020/12/14/converting-floating-point-numbers-to-integers-while-preserving-order/)
-/// explains the mapping in a clear manner.
-///
 /// # See also
 /// The [reverse mapping is `u64_to_f64`](./fn.u64_to_f64.html).
 #[inline(always)]
@@ -153,7 +148,6 @@ pub(crate) mod test {
    pub use super::minmax;
    pub use super::serialize::test::fixed_size_test;
    use super::{compute_num_bits, f64_to_u64, i64_to_u64, u64_to_f64, u64_to_i64};
-    use proptest::prelude::*;
    use std::f64;

    fn test_i64_converter_helper(val: i64) {
@@ -164,15 +158,6 @@ pub(crate) mod test {
        assert_eq!(u64_to_f64(f64_to_u64(val)), val);
    }

-    proptest! {
-        #[test]
-        fn test_f64_converter_monotonicity_proptest((left, right) in (proptest::num::f64::NORMAL, proptest::num::f64::NORMAL)) {
-            let left_u64 = f64_to_u64(left);
-            let right_u64 = f64_to_u64(right);
-            assert_eq!(left_u64 < right_u64,  left < right);
-        }
-    }
-
    #[test]
    fn test_i64_converter() {
        assert_eq!(i64_to_u64(i64::min_value()), u64::min_value());
--- a/src/core/index.rs
+++ b/src/core/index.rs
@@ -35,21 +35,12 @@ fn load_metas(
    inventory: &SegmentMetaInventory,
 ) -> crate::Result<IndexMeta> {
    let meta_data = directory.atomic_read(&META_FILEPATH)?;
-    let meta_string = String::from_utf8(meta_data).map_err(|_utf8_err| {
-        error!("Meta data is not valid utf8.");
-        DataCorruption::new(
-            META_FILEPATH.to_path_buf(),
-            "Meta file does not contain valid utf8 file.".to_string(),
-        )
-    })?;
+    let meta_string = String::from_utf8_lossy(&meta_data);
    IndexMeta::deserialize(&meta_string, &inventory)
        .map_err(|e| {
            DataCorruption::new(
                META_FILEPATH.to_path_buf(),
-                format!(
-                    "Meta file cannot be deserialized. {:?}. Content: {:?}",
-                    e, meta_string
-                ),
+                format!("Meta file cannot be deserialized. {:?}.", e),
            )
        })
        .map_err(From::from)
--- a/src/core/segment_reader.rs
+++ b/src/core/segment_reader.rs
@@ -114,7 +114,12 @@ impl SegmentReader {
                field_entry.name()
            )));
        }
-        let term_ords_reader = self.fast_fields().u64s(field)?;
+        let term_ords_reader = self.fast_fields().u64s(field).ok_or_else(|| {
+            DataCorruption::comment_only(format!(
+                "Cannot find data for hierarchical facet {:?}",
+                field_entry.name()
+            ))
+        })?;
        let termdict = self
            .termdict_composite
            .open_read(field)
@@ -178,10 +183,8 @@ impl SegmentReader {

        let fast_fields_data = segment.open_read(SegmentComponent::FASTFIELDS)?;
        let fast_fields_composite = CompositeFile::open(&fast_fields_data)?;
-        let fast_field_readers = Arc::new(FastFieldReaders::new(
-            schema.clone(),
-            fast_fields_composite,
-        )?);
+        let fast_field_readers =
+            Arc::new(FastFieldReaders::load_all(&schema, &fast_fields_composite)?);

        let fieldnorm_data = segment.open_read(SegmentComponent::FIELDNORMS)?;
        let fieldnorm_readers = FieldNormReaders::open(fieldnorm_data)?;
@@ -307,7 +310,7 @@ impl SegmentReader {
    }

    /// Returns an iterator that will iterate over the alive document ids
-    pub fn doc_ids_alive(&self) -> impl Iterator<Item = DocId> + '_ {
+    pub fn doc_ids_alive<'a>(&'a self) -> impl Iterator<Item = DocId> + 'a {
        (0u32..self.max_doc).filter(move |doc| !self.is_deleted(*doc))
    }

--- a/src/directory/footer.rs
+++ b/src/directory/footer.rs
@@ -115,18 +115,6 @@ impl Footer {
                }
                Ok(())
            }
-            VersionedFooter::V3 {
-                crc32: _crc,
-                store_compression,
-            } => {
-                if &library_version.store_compression != store_compression {
-                    return Err(Incompatibility::CompressionMismatch {
-                        library_compression_format: library_version.store_compression.to_string(),
-                        index_compression_format: store_compression.to_string(),
-                    });
-                }
-                Ok(())
-            }
            VersionedFooter::UnknownVersion => Err(Incompatibility::IndexMismatch {
                library_version: library_version.clone(),
                index_version: self.version.clone(),
@@ -148,31 +136,24 @@ pub enum VersionedFooter {
        crc32: CrcHashU32,
        store_compression: String,
    },
-    // Block wand max termfred on 1 byte
-    V3 {
-        crc32: CrcHashU32,
-        store_compression: String,
-    },
 }

 impl BinarySerializable for VersionedFooter {
    fn serialize<W: io::Write>(&self, writer: &mut W) -> io::Result<()> {
        let mut buf = Vec::new();
        match self {
-            VersionedFooter::V3 {
+            VersionedFooter::V2 {
                crc32,
                store_compression: compression,
            } => {
                // Serializes a valid `VersionedFooter` or panics if the version is unknown
                // [   version    |   crc_hash  | compression_mode ]
                // [    0..4      |     4..8    |     variable     ]
-                BinarySerializable::serialize(&3u32, &mut buf)?;
+                BinarySerializable::serialize(&2u32, &mut buf)?;
                BinarySerializable::serialize(crc32, &mut buf)?;
                BinarySerializable::serialize(compression, &mut buf)?;
            }
-            VersionedFooter::V2 { .. }
-            | VersionedFooter::V1 { .. }
-            | VersionedFooter::UnknownVersion => {
+            VersionedFooter::V1 { .. } | VersionedFooter::UnknownVersion => {
                return Err(io::Error::new(
                    io::ErrorKind::InvalidInput,
                    "Cannot serialize an unknown versioned footer ",
@@ -201,7 +182,7 @@ impl BinarySerializable for VersionedFooter {
        reader.read_exact(&mut buf[..])?;
        let mut cursor = &buf[..];
        let version = u32::deserialize(&mut cursor)?;
-        if version > 3 {
+        if version != 1 && version != 2 {
            return Ok(VersionedFooter::UnknownVersion);
        }
        let crc32 = u32::deserialize(&mut cursor)?;
@@ -211,14 +192,9 @@ impl BinarySerializable for VersionedFooter {
                crc32,
                store_compression,
            }
-        } else if version == 2 {
-            VersionedFooter::V2 {
-                crc32,
-                store_compression,
-            }
        } else {
-            assert_eq!(version, 3);
-            VersionedFooter::V3 {
+            assert_eq!(version, 2);
+            VersionedFooter::V2 {
                crc32,
                store_compression,
            }
@@ -229,7 +205,6 @@ impl BinarySerializable for VersionedFooter {
 impl VersionedFooter {
    pub fn crc(&self) -> Option<CrcHashU32> {
        match self {
-            VersionedFooter::V3 { crc32, .. } => Some(*crc32),
            VersionedFooter::V2 { crc32, .. } => Some(*crc32),
            VersionedFooter::V1 { crc32, .. } => Some(*crc32),
            VersionedFooter::UnknownVersion { .. } => None,
@@ -268,7 +243,7 @@ impl<W: TerminatingWrite> Write for FooterProxy<W> {
 impl<W: TerminatingWrite> TerminatingWrite for FooterProxy<W> {
    fn terminate_ref(&mut self, _: AntiCallToken) -> io::Result<()> {
        let crc32 = self.hasher.take().unwrap().finalize();
-        let footer = Footer::new(VersionedFooter::V3 {
+        let footer = Footer::new(VersionedFooter::V2 {
            crc32,
            store_compression: crate::store::COMPRESSION.to_string(),
        });
@@ -303,7 +278,7 @@ mod tests {
        let footer = Footer::deserialize(&mut &vec[..]).unwrap();
        assert!(matches!(
           footer.versioned_footer,
-           VersionedFooter::V3 { store_compression, .. }
+           VersionedFooter::V2 { store_compression, .. }
           if store_compression == crate::store::COMPRESSION
        ));
        assert_eq!(&footer.version, crate::version());
@@ -313,7 +288,7 @@ mod tests {
    fn test_serialize_deserialize_footer() {
        let mut buffer = Vec::new();
        let crc32 = 123456u32;
-        let footer: Footer = Footer::new(VersionedFooter::V3 {
+        let footer: Footer = Footer::new(VersionedFooter::V2 {
            crc32,
            store_compression: "lz4".to_string(),
        });
@@ -325,7 +300,7 @@ mod tests {
    #[test]
    fn footer_length() {
        let crc32 = 1111111u32;
-        let versioned_footer = VersionedFooter::V3 {
+        let versioned_footer = VersionedFooter::V2 {
            crc32,
            store_compression: "lz4".to_string(),
        };
@@ -346,7 +321,7 @@ mod tests {
            // versionned footer length
            12 | 128,
            // index format version
-            3,
+            2,
            0,
            0,
            0,
@@ -365,7 +340,7 @@ mod tests {
        let versioned_footer = VersionedFooter::deserialize(&mut cursor).unwrap();
        assert!(cursor.is_empty());
        let expected_crc: u32 = LittleEndian::read_u32(&v_footer_bytes[5..9]) as CrcHashU32;
-        let expected_versioned_footer: VersionedFooter = VersionedFooter::V3 {
+        let expected_versioned_footer: VersionedFooter = VersionedFooter::V2 {
            crc32: expected_crc,
            store_compression: "lz4".to_string(),
        };
--- a/src/directory/ram_directory.rs
+++ b/src/directory/ram_directory.rs
@@ -226,9 +226,13 @@ impl Directory for RAMDirectory {
        )));
        let path_buf = PathBuf::from(path);

-        self.fs.write().unwrap().write(path_buf, data);
+        // Reserve the path to prevent calls to .write() to succeed.
+        self.fs.write().unwrap().write(path_buf.clone(), &[]);

-        if path == *META_FILEPATH {
+        let mut vec_writer = VecWriter::new(path_buf, self.clone());
+        vec_writer.write_all(data)?;
+        vec_writer.flush()?;
+        if path == Path::new(&*META_FILEPATH) {
            let _ = self.fs.write().unwrap().watch_router.broadcast();
        }
        Ok(())
--- a/src/fastfield/facet_reader.rs
+++ b/src/fastfield/facet_reader.rs
@@ -1,4 +1,4 @@
-use super::MultiValuedFastFieldReader;
+use super::MultiValueIntFastFieldReader;
 use crate::error::DataCorruption;
 use crate::schema::Facet;
 use crate::termdict::TermDictionary;
@@ -20,7 +20,7 @@ use std::str;
 /// list of facets. This ordinal is segment local and
 /// only makes sense for a given segment.
 pub struct FacetReader {
-    term_ords: MultiValuedFastFieldReader<u64>,
+    term_ords: MultiValueIntFastFieldReader<u64>,
    term_dict: TermDictionary,
    buffer: Vec<u8>,
 }
@@ -29,12 +29,12 @@ impl FacetReader {
    /// Creates a new `FacetReader`.
    ///
    /// A facet reader just wraps :
-    /// - a `MultiValuedFastFieldReader` that makes it possible to
+    /// - a `MultiValueIntFastFieldReader` that makes it possible to
    /// access the list of facet ords for a given document.
    /// - a `TermDictionary` that helps associating a facet to
    /// an ordinal and vice versa.
    pub fn new(
-        term_ords: MultiValuedFastFieldReader<u64>,
+        term_ords: MultiValueIntFastFieldReader<u64>,
        term_dict: TermDictionary,
    ) -> FacetReader {
        FacetReader {
--- a/src/fastfield/mod.rs
+++ b/src/fastfield/mod.rs
@@ -28,7 +28,7 @@ pub use self::delete::write_delete_bitset;
 pub use self::delete::DeleteBitSet;
 pub use self::error::{FastFieldNotAvailableError, Result};
 pub use self::facet_reader::FacetReader;
-pub use self::multivalued::{MultiValuedFastFieldReader, MultiValuedFastFieldWriter};
+pub use self::multivalued::{MultiValueIntFastFieldReader, MultiValueIntFastFieldWriter};
 pub use self::reader::FastFieldReader;
 pub use self::readers::FastFieldReaders;
 pub use self::serializer::FastFieldSerializer;
--- a/src/fastfield/multivalued/mod.rs
+++ b/src/fastfield/multivalued/mod.rs
@@ -1,8 +1,8 @@
 mod reader;
 mod writer;

-pub use self::reader::MultiValuedFastFieldReader;
-pub use self::writer::MultiValuedFastFieldWriter;
+pub use self::reader::MultiValueIntFastFieldReader;
+pub use self::writer::MultiValueIntFastFieldWriter;

 #[cfg(test)]
 mod tests {
--- a/src/fastfield/multivalued/reader.rs
+++ b/src/fastfield/multivalued/reader.rs
@@ -10,22 +10,29 @@ use crate::DocId;
 /// The `idx_reader` associated, for each document, the index of its first value.
 ///
 #[derive(Clone)]
-pub struct MultiValuedFastFieldReader<Item: FastValue> {
+pub struct MultiValueIntFastFieldReader<Item: FastValue> {
    idx_reader: FastFieldReader<u64>,
    vals_reader: FastFieldReader<Item>,
 }

-impl<Item: FastValue> MultiValuedFastFieldReader<Item> {
+impl<Item: FastValue> MultiValueIntFastFieldReader<Item> {
    pub(crate) fn open(
        idx_reader: FastFieldReader<u64>,
        vals_reader: FastFieldReader<Item>,
-    ) -> MultiValuedFastFieldReader<Item> {
-        MultiValuedFastFieldReader {
+    ) -> MultiValueIntFastFieldReader<Item> {
+        MultiValueIntFastFieldReader {
            idx_reader,
            vals_reader,
        }
    }

+    pub(crate) fn into_u64s_reader(self) -> MultiValueIntFastFieldReader<u64> {
+        MultiValueIntFastFieldReader {
+            idx_reader: self.idx_reader,
+            vals_reader: self.vals_reader.into_u64_reader(),
+        }
+    }
+
    /// Returns `(start, stop)`, such that the values associated
    /// to the given document are `start..stop`.
    fn range(&self, doc: DocId) -> (u64, u64) {
--- a/src/fastfield/multivalued/writer.rs
+++ b/src/fastfield/multivalued/writer.rs
@@ -18,7 +18,7 @@ use std::io;
 /// in your schema
 /// - add your document simply by calling `.add_document(...)`.
 ///
-/// The `MultiValuedFastFieldWriter` can be acquired from the
+/// The `MultiValueIntFastFieldWriter` can be acquired from the
 /// fastfield writer, by calling [`.get_multivalue_writer(...)`](./struct.FastFieldsWriter.html#method.get_multivalue_writer).
 ///
 /// Once acquired, writing is done by calling calls to
@@ -29,17 +29,17 @@ use std::io;
 /// This makes it possible to push unordered term ids,
 /// during indexing and remap them to their respective
 /// term ids when the segment is getting serialized.
-pub struct MultiValuedFastFieldWriter {
+pub struct MultiValueIntFastFieldWriter {
    field: Field,
    vals: Vec<UnorderedTermId>,
    doc_index: Vec<u64>,
    is_facet: bool,
 }

-impl MultiValuedFastFieldWriter {
+impl MultiValueIntFastFieldWriter {
    /// Creates a new `IntFastFieldWriter`
    pub(crate) fn new(field: Field, is_facet: bool) -> Self {
-        MultiValuedFastFieldWriter {
+        MultiValueIntFastFieldWriter {
            field,
            vals: Vec::new(),
            doc_index: Vec::new(),
@@ -47,7 +47,7 @@ impl MultiValuedFastFieldWriter {
        }
    }

-    /// Access the field associated to the `MultiValuedFastFieldWriter`
+    /// Access the field associated to the `MultiValueIntFastFieldWriter`
    pub fn field(&self) -> Field {
        self.field
    }
--- a/src/fastfield/reader.rs
+++ b/src/fastfield/reader.rs
@@ -42,6 +42,24 @@ impl<Item: FastValue> FastFieldReader<Item> {
        })
    }

+    pub(crate) fn into_u64_reader(self) -> FastFieldReader<u64> {
+        FastFieldReader {
+            bit_unpacker: self.bit_unpacker,
+            min_value_u64: self.min_value_u64,
+            max_value_u64: self.max_value_u64,
+            _phantom: PhantomData,
+        }
+    }
+
+    pub(crate) fn cast<TFastValue: FastValue>(self) -> FastFieldReader<TFastValue> {
+        FastFieldReader {
+            bit_unpacker: self.bit_unpacker,
+            min_value_u64: self.min_value_u64,
+            max_value_u64: self.max_value_u64,
+            _phantom: PhantomData,
+        }
+    }
+
    /// Return the value associated to the given document.
    ///
    /// This accessor should return as fast as possible.
--- a/src/fastfield/readers.rs
+++ b/src/fastfield/readers.rs
@@ -1,22 +1,28 @@
 use crate::common::CompositeFile;
-use crate::directory::FileSlice;
-use crate::fastfield::MultiValuedFastFieldReader;
 use crate::fastfield::{BytesFastFieldReader, FastValue};
+use crate::fastfield::MultiValueIntFastFieldReader;
 use crate::fastfield::{FastFieldNotAvailableError, FastFieldReader};
 use crate::schema::{Cardinality, Field, FieldType, Schema};
 use crate::space_usage::PerFieldSpaceUsage;
-use crate::TantivyError;
+use std::collections::HashMap;

 /// Provides access to all of the FastFieldReader.
 ///
 /// Internally, `FastFieldReaders` have preloaded fast field readers,
 /// and just wraps several `HashMap`.
-#[derive(Clone)]
 pub struct FastFieldReaders {
-    schema: Schema,
+    fast_field_i64: HashMap<Field, FastFieldReader<i64>>,
+    fast_field_u64: HashMap<Field, FastFieldReader<u64>>,
+    fast_field_f64: HashMap<Field, FastFieldReader<f64>>,
+    fast_field_date: HashMap<Field, FastFieldReader<crate::DateTime>>,
+    fast_field_i64s: HashMap<Field, MultiValueIntFastFieldReader<i64>>,
+    fast_field_u64s: HashMap<Field, MultiValueIntFastFieldReader<u64>>,
+    fast_field_f64s: HashMap<Field, MultiValueIntFastFieldReader<f64>>,
+    fast_field_dates: HashMap<Field, MultiValueIntFastFieldReader<crate::DateTime>>,
+    fast_bytes: HashMap<Field, BytesFastFieldReader>,
    fast_fields_composite: CompositeFile,
 }
-#[derive(Eq, PartialEq, Debug)]
+
 enum FastType {
    I64,
    U64,
@@ -44,167 +50,232 @@ fn type_and_cardinality(field_type: &FieldType) -> Option<(FastType, Cardinality
 }

 impl FastFieldReaders {
-    pub(crate) fn new(
-        schema: Schema,
-        fast_fields_composite: CompositeFile,
+    pub(crate) fn load_all(
+        schema: &Schema,
+        fast_fields_composite: &CompositeFile,
    ) -> crate::Result<FastFieldReaders> {
-        Ok(FastFieldReaders {
-            fast_fields_composite,
-            schema,
-        })
+        let mut fast_field_readers = FastFieldReaders {
+            fast_field_i64: Default::default(),
+            fast_field_u64: Default::default(),
+            fast_field_f64: Default::default(),
+            fast_field_date: Default::default(),
+            fast_field_i64s: Default::default(),
+            fast_field_u64s: Default::default(),
+            fast_field_f64s: Default::default(),
+            fast_field_dates: Default::default(),
+            fast_bytes: Default::default(),
+            fast_fields_composite: fast_fields_composite.clone(),
+        };
+        for (field, field_entry) in schema.fields() {
+            let field_type = field_entry.field_type();
+            if let FieldType::Bytes(bytes_option) = field_type {
+                if !bytes_option.is_fast() {
+                    continue;
+                }
+                let fast_field_idx_file = fast_fields_composite
+                    .open_read_with_idx(field, 0)
+                    .ok_or_else(|| FastFieldNotAvailableError::new(field_entry))?;
+                let idx_reader = FastFieldReader::open(fast_field_idx_file)?;
+                let data = fast_fields_composite
+                    .open_read_with_idx(field, 1)
+                    .ok_or_else(|| FastFieldNotAvailableError::new(field_entry))?;
+                let bytes_fast_field_reader = BytesFastFieldReader::open(idx_reader, data)?;
+                fast_field_readers
+                    .fast_bytes
+                    .insert(field, bytes_fast_field_reader);
+            } else if let Some((fast_type, cardinality)) = type_and_cardinality(field_type) {
+                match cardinality {
+                    Cardinality::SingleValue => {
+                        if let Some(fast_field_data) = fast_fields_composite.open_read(field) {
+                            match fast_type {
+                                FastType::U64 => {
+                                    let fast_field_reader = FastFieldReader::open(fast_field_data)?;
+                                    fast_field_readers
+                                        .fast_field_u64
+                                        .insert(field, fast_field_reader);
+                                }
+                                FastType::I64 => {
+                                    let fast_field_reader =
+                                        FastFieldReader::open(fast_field_data.clone())?;
+                                    fast_field_readers
+                                        .fast_field_i64
+                                        .insert(field, fast_field_reader);
+                                }
+                                FastType::F64 => {
+                                    let fast_field_reader =
+                                        FastFieldReader::open(fast_field_data.clone())?;
+                                    fast_field_readers
+                                        .fast_field_f64
+                                        .insert(field, fast_field_reader);
+                                }
+                                FastType::Date => {
+                                    let fast_field_reader =
+                                        FastFieldReader::open(fast_field_data.clone())?;
+                                    fast_field_readers
+                                        .fast_field_date
+                                        .insert(field, fast_field_reader);
+                                }
+                            }
+                        } else {
+                            return Err(From::from(FastFieldNotAvailableError::new(field_entry)));
+                        }
+                    }
+                    Cardinality::MultiValues => {
+                        let idx_opt = fast_fields_composite.open_read_with_idx(field, 0);
+                        let data_opt = fast_fields_composite.open_read_with_idx(field, 1);
+                        if let (Some(fast_field_idx), Some(fast_field_data)) = (idx_opt, data_opt) {
+                            let idx_reader = FastFieldReader::open(fast_field_idx)?;
+                            match fast_type {
+                                FastType::I64 => {
+                                    let vals_reader = FastFieldReader::open(fast_field_data)?;
+                                    let multivalued_int_fast_field =
+                                        MultiValueIntFastFieldReader::open(idx_reader, vals_reader);
+                                    fast_field_readers
+                                        .fast_field_i64s
+                                        .insert(field, multivalued_int_fast_field);
+                                }
+                                FastType::U64 => {
+                                    let vals_reader = FastFieldReader::open(fast_field_data)?;
+                                    let multivalued_int_fast_field =
+                                        MultiValueIntFastFieldReader::open(idx_reader, vals_reader);
+                                    fast_field_readers
+                                        .fast_field_u64s
+                                        .insert(field, multivalued_int_fast_field);
+                                }
+                                FastType::F64 => {
+                                    let vals_reader = FastFieldReader::open(fast_field_data)?;
+                                    let multivalued_int_fast_field =
+                                        MultiValueIntFastFieldReader::open(idx_reader, vals_reader);
+                                    fast_field_readers
+                                        .fast_field_f64s
+                                        .insert(field, multivalued_int_fast_field);
+                                }
+                                FastType::Date => {
+                                    let vals_reader = FastFieldReader::open(fast_field_data)?;
+                                    let multivalued_int_fast_field =
+                                        MultiValueIntFastFieldReader::open(idx_reader, vals_reader);
+                                    fast_field_readers
+                                        .fast_field_dates
+                                        .insert(field, multivalued_int_fast_field);
+                                }
+                            }
+                        } else {
+                            return Err(From::from(FastFieldNotAvailableError::new(field_entry)));
+                        }
+                    }
+                }
+            }
+        }
+        Ok(fast_field_readers)
    }

    pub(crate) fn space_usage(&self) -> PerFieldSpaceUsage {
        self.fast_fields_composite.space_usage()
    }

-    fn fast_field_data(&self, field: Field, idx: usize) -> crate::Result<FileSlice> {
-        self.fast_fields_composite
-            .open_read_with_idx(field, idx)
-            .ok_or_else(|| {
-                let field_name = self.schema.get_field_entry(field).name();
-                TantivyError::SchemaError(format!("Field({}) data was not found", field_name))
-            })
-    }
-
-    fn check_type(
-        &self,
-        field: Field,
-        expected_fast_type: FastType,
-        expected_cardinality: Cardinality,
-    ) -> crate::Result<()> {
-        let field_entry = self.schema.get_field_entry(field);
-        let (fast_type, cardinality) =
-            type_and_cardinality(field_entry.field_type()).ok_or_else(|| {
-                crate::TantivyError::SchemaError(format!(
-                    "Field {:?} is not a fast field.",
-                    field_entry.name()
-                ))
-            })?;
-        if fast_type != expected_fast_type {
-            return Err(crate::TantivyError::SchemaError(format!(
-                "Field {:?} is of type {:?}, expected {:?}.",
-                field_entry.name(),
-                fast_type,
-                expected_fast_type
-            )));
-        }
-        if cardinality != expected_cardinality {
-            return Err(crate::TantivyError::SchemaError(format!(
-                "Field {:?} is of cardinality {:?}, expected {:?}.",
-                field_entry.name(),
-                cardinality,
-                expected_cardinality
-            )));
-        }
-        Ok(())
-    }
-
-    pub(crate) fn typed_fast_field_reader<TFastValue: FastValue>(
-        &self,
-        field: Field,
-    ) -> crate::Result<FastFieldReader<TFastValue>> {
-        let fast_field_slice = self.fast_field_data(field, 0)?;
-        FastFieldReader::open(fast_field_slice)
-    }
-
-    pub(crate) fn typed_fast_field_multi_reader<TFastValue: FastValue>(
-        &self,
-        field: Field,
-    ) -> crate::Result<MultiValuedFastFieldReader<TFastValue>> {
-        let fast_field_slice_idx = self.fast_field_data(field, 0)?;
-        let fast_field_slice_vals = self.fast_field_data(field, 1)?;
-        let idx_reader = FastFieldReader::open(fast_field_slice_idx)?;
-        let vals_reader: FastFieldReader<TFastValue> =
-            FastFieldReader::open(fast_field_slice_vals)?;
-        Ok(MultiValuedFastFieldReader::open(idx_reader, vals_reader))
-    }
-
    /// Returns the `u64` fast field reader reader associated to `field`.
    ///
    /// If `field` is not a u64 fast field, this method returns `None`.
-    pub fn u64(&self, field: Field) -> crate::Result<FastFieldReader<u64>> {
-        self.check_type(field, FastType::U64, Cardinality::SingleValue)?;
-        self.typed_fast_field_reader(field)
+    pub fn u64(&self, field: Field) -> Option<FastFieldReader<u64>> {
+        self.fast_field_u64.get(&field).cloned()
+    }
+
+    /// If the field is a u64-fast field return the associated reader.
+    /// If the field is a i64-fast field, return the associated u64 reader. Values are
+    /// mapped from i64 to u64 using a (well the, it is unique) monotonic mapping.    ///
+    ///
+    /// This method is useful when merging segment reader.
+    pub(crate) fn u64_lenient(&self, field: Field) -> Option<FastFieldReader<u64>> {
+        if let Some(u64_ff_reader) = self.u64(field) {
+            return Some(u64_ff_reader);
+        }
+        if let Some(i64_ff_reader) = self.i64(field) {
+            return Some(i64_ff_reader.into_u64_reader());
+        }
+        if let Some(f64_ff_reader) = self.f64(field) {
+            return Some(f64_ff_reader.into_u64_reader());
+        }
+        if let Some(date_ff_reader) = self.date(field) {
+            return Some(date_ff_reader.into_u64_reader());
+        }
+        None
+    }
+
+    pub(crate) fn typed_fast_field_reader<TFastValue: FastValue>(&self, field: Field) -> Option<FastFieldReader<TFastValue>> {
+        self.u64_lenient(field).map(|fast_field_reader| fast_field_reader.cast())
    }

    /// Returns the `i64` fast field reader reader associated to `field`.
    ///
    /// If `field` is not a i64 fast field, this method returns `None`.
-    pub fn i64(&self, field: Field) -> crate::Result<FastFieldReader<i64>> {
-        self.check_type(field, FastType::I64, Cardinality::SingleValue)?;
-        self.typed_fast_field_reader(field)
+    pub fn i64(&self, field: Field) -> Option<FastFieldReader<i64>> {
+        self.fast_field_i64.get(&field).cloned()
    }

    /// Returns the `i64` fast field reader reader associated to `field`.
    ///
    /// If `field` is not a i64 fast field, this method returns `None`.
-    pub fn date(&self, field: Field) -> crate::Result<FastFieldReader<crate::DateTime>> {
-        self.check_type(field, FastType::Date, Cardinality::SingleValue)?;
-        self.typed_fast_field_reader(field)
+    pub fn date(&self, field: Field) -> Option<FastFieldReader<crate::DateTime>> {
+        self.fast_field_date.get(&field).cloned()
    }

    /// Returns the `f64` fast field reader reader associated to `field`.
    ///
    /// If `field` is not a f64 fast field, this method returns `None`.
-    pub fn f64(&self, field: Field) -> crate::Result<FastFieldReader<f64>> {
-        self.check_type(field, FastType::F64, Cardinality::SingleValue)?;
-        self.typed_fast_field_reader(field)
+    pub fn f64(&self, field: Field) -> Option<FastFieldReader<f64>> {
+        self.fast_field_f64.get(&field).cloned()
    }

    /// Returns a `u64s` multi-valued fast field reader reader associated to `field`.
    ///
    /// If `field` is not a u64 multi-valued fast field, this method returns `None`.
-    pub fn u64s(&self, field: Field) -> crate::Result<MultiValuedFastFieldReader<u64>> {
-        self.check_type(field, FastType::U64, Cardinality::MultiValues)?;
-        self.typed_fast_field_multi_reader(field)
+    pub fn u64s(&self, field: Field) -> Option<MultiValueIntFastFieldReader<u64>> {
+        self.fast_field_u64s.get(&field).cloned()
+    }
+
+    /// If the field is a u64s-fast field return the associated reader.
+    /// If the field is a i64s-fast field, return the associated u64s reader. Values are
+    /// mapped from i64 to u64 using a (well the, it is unique) monotonic mapping.
+    ///
+    /// This method is useful when merging segment reader.
+    pub(crate) fn u64s_lenient(&self, field: Field) -> Option<MultiValueIntFastFieldReader<u64>> {
+        if let Some(u64s_ff_reader) = self.u64s(field) {
+            return Some(u64s_ff_reader);
+        }
+        if let Some(i64s_ff_reader) = self.i64s(field) {
+            return Some(i64s_ff_reader.into_u64s_reader());
+        }
+        if let Some(f64s_ff_reader) = self.f64s(field) {
+            return Some(f64s_ff_reader.into_u64s_reader());
+        }
+        None
    }

    /// Returns a `i64s` multi-valued fast field reader reader associated to `field`.
    ///
    /// If `field` is not a i64 multi-valued fast field, this method returns `None`.
-    pub fn i64s(&self, field: Field) -> crate::Result<MultiValuedFastFieldReader<i64>> {
-        self.check_type(field, FastType::I64, Cardinality::MultiValues)?;
-        self.typed_fast_field_multi_reader(field)
+    pub fn i64s(&self, field: Field) -> Option<MultiValueIntFastFieldReader<i64>> {
+        self.fast_field_i64s.get(&field).cloned()
    }

    /// Returns a `f64s` multi-valued fast field reader reader associated to `field`.
    ///
    /// If `field` is not a f64 multi-valued fast field, this method returns `None`.
-    pub fn f64s(&self, field: Field) -> crate::Result<MultiValuedFastFieldReader<f64>> {
-        self.check_type(field, FastType::F64, Cardinality::MultiValues)?;
-        self.typed_fast_field_multi_reader(field)
+    pub fn f64s(&self, field: Field) -> Option<MultiValueIntFastFieldReader<f64>> {
+        self.fast_field_f64s.get(&field).cloned()
    }

    /// Returns a `crate::DateTime` multi-valued fast field reader reader associated to `field`.
    ///
    /// If `field` is not a `crate::DateTime` multi-valued fast field, this method returns `None`.
-    pub fn dates(
-        &self,
-        field: Field,
-    ) -> crate::Result<MultiValuedFastFieldReader<crate::DateTime>> {
-        self.check_type(field, FastType::Date, Cardinality::MultiValues)?;
-        self.typed_fast_field_multi_reader(field)
+    pub fn dates(&self, field: Field) -> Option<MultiValueIntFastFieldReader<crate::DateTime>> {
+        self.fast_field_dates.get(&field).cloned()
    }

    /// Returns the `bytes` fast field reader associated to `field`.
    ///
    /// If `field` is not a bytes fast field, returns `None`.
-    pub fn bytes(&self, field: Field) -> crate::Result<BytesFastFieldReader> {
-        let field_entry = self.schema.get_field_entry(field);
-        if let FieldType::Bytes(bytes_option) = field_entry.field_type() {
-            if !bytes_option.is_fast() {
-                return Err(crate::TantivyError::SchemaError(format!(
-                    "Field {:?} is not a fast field.",
-                    field_entry.name()
-                )));
-            }
-            let fast_field_idx_file = self.fast_field_data(field, 0)?;
-            let idx_reader = FastFieldReader::open(fast_field_idx_file)?;
-            let data = self.fast_field_data(field, 1)?;
-            BytesFastFieldReader::open(idx_reader, data)
-        } else {
-            Err(FastFieldNotAvailableError::new(field_entry).into())
-        }
+    pub fn bytes(&self, field: Field) -> Option<BytesFastFieldReader> {
+        self.fast_bytes.get(&field).cloned()
    }
 }
--- a/src/fastfield/writer.rs
+++ b/src/fastfield/writer.rs
@@ -1,4 +1,4 @@
-use super::multivalued::MultiValuedFastFieldWriter;
+use super::multivalued::MultiValueIntFastFieldWriter;
 use crate::common;
 use crate::common::BinarySerializable;
 use crate::common::VInt;
@@ -13,7 +13,7 @@ use std::io;
 /// The fastfieldswriter regroup all of the fast field writers.
 pub struct FastFieldsWriter {
    single_value_writers: Vec<IntFastFieldWriter>,
-    multi_values_writers: Vec<MultiValuedFastFieldWriter>,
+    multi_values_writers: Vec<MultiValueIntFastFieldWriter>,
    bytes_value_writers: Vec<BytesFastFieldWriter>,
 }

@@ -46,14 +46,14 @@ impl FastFieldsWriter {
                            single_value_writers.push(fast_field_writer);
                        }
                        Some(Cardinality::MultiValues) => {
-                            let fast_field_writer = MultiValuedFastFieldWriter::new(field, false);
+                            let fast_field_writer = MultiValueIntFastFieldWriter::new(field, false);
                            multi_values_writers.push(fast_field_writer);
                        }
                        None => {}
                    }
                }
                FieldType::HierarchicalFacet => {
-                    let fast_field_writer = MultiValuedFastFieldWriter::new(field, true);
+                    let fast_field_writer = MultiValueIntFastFieldWriter::new(field, true);
                    multi_values_writers.push(fast_field_writer);
                }
                FieldType::Bytes(bytes_option) => {
@@ -87,7 +87,7 @@ impl FastFieldsWriter {
    pub fn get_multivalue_writer(
        &mut self,
        field: Field,
-    ) -> Option<&mut MultiValuedFastFieldWriter> {
+    ) -> Option<&mut MultiValueIntFastFieldWriter> {
        // TODO optimize
        self.multi_values_writers
            .iter_mut()
--- a/src/functional_test.rs
+++ b/src/functional_test.rs
@@ -1,93 +1,45 @@
-use crate::Index;
-use crate::Searcher;
-use crate::{doc, schema::*};
 use rand::thread_rng;
-use rand::Rng;
 use std::collections::HashSet;

-fn check_index_content(searcher: &Searcher, vals: &[u64]) -> crate::Result<()> {
+use crate::schema::*;
+use crate::Index;
+use crate::Searcher;
+use rand::Rng;
+
+fn check_index_content(searcher: &Searcher, vals: &HashSet<u64>) {
    assert!(searcher.segment_readers().len() < 20);
    assert_eq!(searcher.num_docs() as usize, vals.len());
-    for segment_reader in searcher.segment_readers() {
-        let store_reader = segment_reader.get_store_reader()?;
-        for doc_id in 0..segment_reader.max_doc() {
-            let _doc = store_reader.get(doc_id)?;
-        }
-    }
-    Ok(())
 }

 #[test]
 #[ignore]
-fn test_functional_store() -> crate::Result<()> {
-    let mut schema_builder = Schema::builder();
-
-    let id_field = schema_builder.add_u64_field("id", INDEXED | STORED);
-    let schema = schema_builder.build();
-
-    let index = Index::create_in_ram(schema);
-    let reader = index.reader()?;
-
-    let mut rng = thread_rng();
-
-    let mut index_writer = index.writer_with_num_threads(3, 12_000_000)?;
-
-    let mut doc_set: Vec<u64> = Vec::new();
-
-    let mut doc_id = 0u64;
-    for iteration in 0..500 {
-        dbg!(iteration);
-        let num_docs: usize = rng.gen_range(0..4);
-        if doc_set.len() >= 1 {
-            let doc_to_remove_id = rng.gen_range(0..doc_set.len());
-            let removed_doc_id = doc_set.swap_remove(doc_to_remove_id);
-            index_writer.delete_term(Term::from_field_u64(id_field, removed_doc_id));
-        }
-        for _ in 0..num_docs {
-            doc_set.push(doc_id);
-            index_writer.add_document(doc!(id_field=>doc_id));
-            doc_id += 1;
-        }
-        index_writer.commit()?;
-        reader.reload()?;
-        let searcher = reader.searcher();
-        check_index_content(&searcher, &doc_set)?;
-    }
-    Ok(())
-}
-
-#[test]
-#[ignore]
-fn test_functional_indexing() -> crate::Result<()> {
+fn test_indexing() {
    let mut schema_builder = Schema::builder();

    let id_field = schema_builder.add_u64_field("id", INDEXED);
    let multiples_field = schema_builder.add_u64_field("multiples", INDEXED);
    let schema = schema_builder.build();

-    let index = Index::create_from_tempdir(schema)?;
-    let reader = index.reader()?;
+    let index = Index::create_from_tempdir(schema).unwrap();
+    let reader = index.reader().unwrap();

    let mut rng = thread_rng();

-    let mut index_writer = index.writer_with_num_threads(3, 120_000_000)?;
+    let mut index_writer = index.writer_with_num_threads(3, 120_000_000).unwrap();

    let mut committed_docs: HashSet<u64> = HashSet::new();
    let mut uncommitted_docs: HashSet<u64> = HashSet::new();

    for _ in 0..200 {
-        let random_val = rng.gen_range(0..20);
+        let random_val = rng.gen_range(0, 20);
        if random_val == 0 {
-            index_writer.commit()?;
+            index_writer.commit().expect("Commit failed");
            committed_docs.extend(&uncommitted_docs);
            uncommitted_docs.clear();
-            reader.reload()?;
+            reader.reload().unwrap();
            let searcher = reader.searcher();
            // check that everything is correct.
-            check_index_content(
-                &searcher,
-                &committed_docs.iter().cloned().collect::<Vec<u64>>(),
-            )?;
+            check_index_content(&searcher, &committed_docs);
        } else {
            if committed_docs.remove(&random_val) || uncommitted_docs.remove(&random_val) {
                let doc_id_term = Term::from_field_u64(id_field, random_val);
@@ -103,5 +55,4 @@ fn test_functional_indexing() -> crate::Result<()> {
            }
        }
    }
-    Ok(())
 }
--- a/src/indexer/log_merge_policy.rs
+++ b/src/indexer/log_merge_policy.rs
@@ -8,7 +8,7 @@ const DEFAULT_MIN_LAYER_SIZE: u32 = 10_000;
 const DEFAULT_MIN_MERGE_SIZE: usize = 8;
 const DEFAULT_MAX_MERGE_SIZE: usize = 10_000_000;

-/// `LogMergePolicy` tries to merge segments that have a similar number of
+/// `LogMergePolicy` tries tries to merge segments that have a similar number of
 /// documents.
 #[derive(Debug, Clone)]
 pub struct LogMergePolicy {
--- a/src/indexer/merger.rs
+++ b/src/indexer/merger.rs
@@ -7,7 +7,7 @@ use crate::fastfield::BytesFastFieldReader;
 use crate::fastfield::DeleteBitSet;
 use crate::fastfield::FastFieldReader;
 use crate::fastfield::FastFieldSerializer;
-use crate::fastfield::MultiValuedFastFieldReader;
+use crate::fastfield::MultiValueIntFastFieldReader;
 use crate::fieldnorm::FieldNormsSerializer;
 use crate::fieldnorm::FieldNormsWriter;
 use crate::fieldnorm::{FieldNormReader, FieldNormReaders};
@@ -246,7 +246,7 @@ impl IndexMerger {
        for reader in &self.readers {
            let u64_reader: FastFieldReader<u64> = reader
                .fast_fields()
-                .typed_fast_field_reader(field)
+                .u64_lenient(field)
                .expect("Failed to find a reader for single fast field. This is a tantivy bug and it should never happen.");
            if let Some((seg_min_val, seg_max_val)) =
                compute_min_max_val(&u64_reader, reader.max_doc(), reader.delete_bitset())
@@ -290,7 +290,7 @@ impl IndexMerger {
        fast_field_serializer: &mut FastFieldSerializer,
    ) -> crate::Result<()> {
        let mut total_num_vals = 0u64;
-        let mut u64s_readers: Vec<MultiValuedFastFieldReader<u64>> = Vec::new();
+        let mut u64s_readers: Vec<MultiValueIntFastFieldReader<u64>> = Vec::new();

        // In the first pass, we compute the total number of vals.
        //
@@ -298,8 +298,9 @@ impl IndexMerger {
        // what should be the bit length use for bitpacking.
        for reader in &self.readers {
            let u64s_reader = reader.fast_fields()
-                .typed_fast_field_multi_reader(field)
+                .u64s_lenient(field)
                .expect("Failed to find index for multivalued field. This is a bug in tantivy, please report.");
+
            if let Some(delete_bitset) = reader.delete_bitset() {
                for doc in 0u32..reader.max_doc() {
                    if delete_bitset.is_alive(doc) {
@@ -352,7 +353,7 @@ impl IndexMerger {
            for (segment_ord, segment_reader) in self.readers.iter().enumerate() {
                let term_ordinal_mapping: &[TermOrdinal] =
                    term_ordinal_mappings.get_segment(segment_ord);
-                let ff_reader: MultiValuedFastFieldReader<u64> = segment_reader
+                let ff_reader: MultiValueIntFastFieldReader<u64> = segment_reader
                    .fast_fields()
                    .u64s(field)
                    .expect("Could not find multivalued u64 fast value reader.");
@@ -396,10 +397,8 @@ impl IndexMerger {
        // We go through a complete first pass to compute the minimum and the
        // maximum value and initialize our Serializer.
        for reader in &self.readers {
-            let ff_reader: MultiValuedFastFieldReader<u64> = reader
-                .fast_fields()
-                .typed_fast_field_multi_reader(field)
-                .expect(
+            let ff_reader: MultiValueIntFastFieldReader<u64> =
+                reader.fast_fields().u64s_lenient(field).expect(
                    "Failed to find multivalued fast field reader. This is a bug in \
                     tantivy. Please report.",
                );
@@ -446,7 +445,11 @@ impl IndexMerger {
        let mut bytes_readers: Vec<BytesFastFieldReader> = Vec::new();

        for reader in &self.readers {
-            let bytes_reader = reader.fast_fields().bytes(field)?;
+            let bytes_reader = reader.fast_fields().bytes(field).ok_or_else(|| {
+                crate::TantivyError::InvalidArgument(
+                    "Bytes fast field {:?} not found in segment.".to_string(),
+                )
+            })?;
            if let Some(delete_bitset) = reader.delete_bitset() {
                for doc in 0u32..reader.max_doc() {
                    if delete_bitset.is_alive(doc) {
--- a/src/lib.rs
+++ b/src/lib.rs
@@ -96,7 +96,7 @@
 //! A good place for you to get started is to check out
 //! the example code (
 //! [literate programming](https://tantivy-search.github.io/examples/basic_search.html) /
-//! [source code](https://github.com/tantivy-search/tantivy/blob/main/examples/basic_search.rs))
+//! [source code](https://github.com/tantivy-search/tantivy/blob/master/examples/basic_search.rs))

 #[cfg_attr(test, macro_use)]
 extern crate serde_json;
@@ -174,7 +174,7 @@ use once_cell::sync::Lazy;
 use serde::{Deserialize, Serialize};

 /// Index format version.
-const INDEX_FORMAT_VERSION: u32 = 3;
+const INDEX_FORMAT_VERSION: u32 = 2;

 /// Structure version for the index.
 #[derive(Clone, PartialEq, Eq, Serialize, Deserialize)]
@@ -866,39 +866,39 @@ mod tests {
        let searcher = reader.searcher();
        let segment_reader: &SegmentReader = searcher.segment_reader(0);
        {
-            let fast_field_reader_res = segment_reader.fast_fields().u64(text_field);
-            assert!(fast_field_reader_res.is_err());
+            let fast_field_reader_opt = segment_reader.fast_fields().u64(text_field);
+            assert!(fast_field_reader_opt.is_none());
        }
        {
            let fast_field_reader_opt = segment_reader.fast_fields().u64(stored_int_field);
-            assert!(fast_field_reader_opt.is_err());
+            assert!(fast_field_reader_opt.is_none());
        }
        {
            let fast_field_reader_opt = segment_reader.fast_fields().u64(fast_field_signed);
-            assert!(fast_field_reader_opt.is_err());
+            assert!(fast_field_reader_opt.is_none());
        }
        {
            let fast_field_reader_opt = segment_reader.fast_fields().u64(fast_field_float);
-            assert!(fast_field_reader_opt.is_err());
+            assert!(fast_field_reader_opt.is_none());
        }
        {
            let fast_field_reader_opt = segment_reader.fast_fields().u64(fast_field_unsigned);
-            assert!(fast_field_reader_opt.is_ok());
+            assert!(fast_field_reader_opt.is_some());
            let fast_field_reader = fast_field_reader_opt.unwrap();
            assert_eq!(fast_field_reader.get(0), 4u64)
        }

        {
-            let fast_field_reader_res = segment_reader.fast_fields().i64(fast_field_signed);
-            assert!(fast_field_reader_res.is_ok());
-            let fast_field_reader = fast_field_reader_res.unwrap();
+            let fast_field_reader_opt = segment_reader.fast_fields().i64(fast_field_signed);
+            assert!(fast_field_reader_opt.is_some());
+            let fast_field_reader = fast_field_reader_opt.unwrap();
            assert_eq!(fast_field_reader.get(0), 4i64)
        }

        {
-            let fast_field_reader_res = segment_reader.fast_fields().f64(fast_field_float);
-            assert!(fast_field_reader_res.is_ok());
-            let fast_field_reader = fast_field_reader_res.unwrap();
+            let fast_field_reader_opt = segment_reader.fast_fields().f64(fast_field_float);
+            assert!(fast_field_reader_opt.is_some());
+            let fast_field_reader = fast_field_reader_opt.unwrap();
            assert_eq!(fast_field_reader.get(0), 4f64)
        }
        Ok(())
--- a/src/positions/reader.rs
+++ b/src/positions/reader.rs
@@ -132,7 +132,7 @@ impl PositionReader {
            "offset arguments should be increasing."
        );
        let delta_to_block_offset = offset as i64 - self.block_offset as i64;
-        if !(0..128).contains(&delta_to_block_offset) {
+        if delta_to_block_offset < 0 || delta_to_block_offset >= 128 {
            // The first position is not within the first block.
            // We need to decompress the first block.
            let delta_to_anchor_offset = offset - self.anchor_offset;
--- a/src/postings/segment_postings.rs
+++ b/src/postings/segment_postings.rs
@@ -1,11 +1,14 @@
 use crate::common::HasLen;
+use crate::directory::FileSlice;
 use crate::docset::DocSet;
 use crate::fastfield::DeleteBitSet;
 use crate::positions::PositionReader;
 use crate::postings::compression::COMPRESSION_BLOCK_SIZE;
+use crate::postings::serializer::PostingsSerializer;
 use crate::postings::BlockSearcher;
 use crate::postings::BlockSegmentPostings;
 use crate::postings::Postings;
+use crate::schema::IndexRecordOption;
 use crate::{DocId, TERMINATED};

 /// `SegmentPostings` represents the inverted list or postings associated to
@@ -65,11 +68,7 @@ impl SegmentPostings {
    /// It serializes the doc ids using tantivy's codec
    /// and returns a `SegmentPostings` object that embeds a
    /// buffer with the serialized data.
-    #[cfg(test)]
    pub fn create_from_docs(docs: &[u32]) -> SegmentPostings {
-        use crate::directory::FileSlice;
-        use crate::postings::serializer::PostingsSerializer;
-        use crate::schema::IndexRecordOption;
        let mut buffer = Vec::new();
        {
            let mut postings_serializer =
@@ -98,9 +97,6 @@ impl SegmentPostings {
        doc_and_tfs: &[(u32, u32)],
        fieldnorms: Option<&[u32]>,
    ) -> SegmentPostings {
-        use crate::directory::FileSlice;
-        use crate::postings::serializer::PostingsSerializer;
-        use crate::schema::IndexRecordOption;
        use crate::fieldnorm::FieldNormReader;
        use crate::Score;
        let mut buffer: Vec<u8> = Vec::new();
--- a/src/postings/skip.rs
+++ b/src/postings/skip.rs
@@ -1,46 +1,32 @@
-use std::convert::TryInto;
-
+use crate::common::{read_u32_vint_no_advance, serialize_vint_u32, BinarySerializable};
 use crate::directory::OwnedBytes;
 use crate::postings::compression::{compressed_block_size, COMPRESSION_BLOCK_SIZE};
 use crate::query::BM25Weight;
 use crate::schema::IndexRecordOption;
 use crate::{DocId, Score, TERMINATED};

-#[inline(always)]
-fn encode_block_wand_max_tf(max_tf: u32) -> u8 {
-    max_tf.min(u8::MAX as u32) as u8
-}
-
-#[inline(always)]
-fn decode_block_wand_max_tf(max_tf_code: u8) -> u32 {
-    if max_tf_code == u8::MAX {
-        u32::MAX
-    } else {
-        max_tf_code as u32
-    }
-}
-
-#[inline(always)]
-fn read_u32(data: &[u8]) -> u32 {
-    u32::from_le_bytes(data[..4].try_into().unwrap())
-}
-
-#[inline(always)]
-fn write_u32(val: u32, buf: &mut Vec<u8>) {
-    buf.extend_from_slice(&val.to_le_bytes());
-}
-
 pub struct SkipSerializer {
    buffer: Vec<u8>,
+    prev_doc: DocId,
 }

 impl SkipSerializer {
    pub fn new() -> SkipSerializer {
-        SkipSerializer { buffer: Vec::new() }
+        SkipSerializer {
+            buffer: Vec::new(),
+            prev_doc: 0u32,
+        }
    }

    pub fn write_doc(&mut self, last_doc: DocId, doc_num_bits: u8) {
-        write_u32(last_doc, &mut self.buffer);
+        assert!(
+            last_doc > self.prev_doc,
+            "write_doc(...) called with non-increasing doc ids. \
+             Did you forget to call clear maybe?"
+        );
+        let delta_doc = last_doc - self.prev_doc;
+        self.prev_doc = last_doc;
+        delta_doc.serialize(&mut self.buffer).unwrap();
        self.buffer.push(doc_num_bits);
    }

@@ -49,13 +35,16 @@ impl SkipSerializer {
    }

    pub fn write_total_term_freq(&mut self, tf_sum: u32) {
-        write_u32(tf_sum, &mut self.buffer);
+        tf_sum
+            .serialize(&mut self.buffer)
+            .expect("Should never fail");
    }

    pub fn write_blockwand_max(&mut self, fieldnorm_id: u8, term_freq: u32) {
-        let block_wand_tf = encode_block_wand_max_tf(term_freq);
-        self.buffer
-            .extend_from_slice(&[fieldnorm_id, block_wand_tf]);
+        self.buffer.push(fieldnorm_id);
+        let mut buf = [0u8; 8];
+        let bytes = serialize_vint_u32(term_freq, &mut buf);
+        self.buffer.extend_from_slice(bytes);
    }

    pub fn data(&self) -> &[u8] {
@@ -63,6 +52,7 @@ impl SkipSerializer {
    }

    pub fn clear(&mut self) {
+        self.prev_doc = 0u32;
        self.buffer.clear();
    }
 }
@@ -169,13 +159,18 @@ impl SkipReader {
    }

    fn read_block_info(&mut self) {
-        let bytes = self.owned_read.as_slice();
-        let advance_len: usize;
-        self.last_doc_in_block = read_u32(bytes);
-        let doc_num_bits = bytes[4];
+        let doc_delta = {
+            let bytes = self.owned_read.as_slice();
+            let mut buf = [0; 4];
+            buf.copy_from_slice(&bytes[..4]);
+            u32::from_le_bytes(buf)
+        };
+        self.last_doc_in_block += doc_delta as DocId;
+        let doc_num_bits = self.owned_read.as_slice()[4];
+
        match self.skip_info {
            IndexRecordOption::Basic => {
-                advance_len = 5;
+                self.owned_read.advance(5);
                self.block_info = BlockInfo::BitPacked {
                    doc_num_bits,
                    tf_num_bits: 0,
@@ -185,10 +180,11 @@ impl SkipReader {
                };
            }
            IndexRecordOption::WithFreqs => {
+                let bytes = self.owned_read.as_slice();
                let tf_num_bits = bytes[5];
                let block_wand_fieldnorm_id = bytes[6];
-                let block_wand_term_freq = decode_block_wand_max_tf(bytes[7]);
-                advance_len = 8;
+                let (block_wand_term_freq, num_bytes) = read_u32_vint_no_advance(&bytes[7..]);
+                self.owned_read.advance(7 + num_bytes);
                self.block_info = BlockInfo::BitPacked {
                    doc_num_bits,
                    tf_num_bits,
@@ -198,11 +194,16 @@ impl SkipReader {
                };
            }
            IndexRecordOption::WithFreqsAndPositions => {
+                let bytes = self.owned_read.as_slice();
                let tf_num_bits = bytes[5];
-                let tf_sum = read_u32(&bytes[6..10]);
+                let tf_sum = {
+                    let mut buf = [0; 4];
+                    buf.copy_from_slice(&bytes[6..10]);
+                    u32::from_le_bytes(buf)
+                };
                let block_wand_fieldnorm_id = bytes[10];
-                let block_wand_term_freq = decode_block_wand_max_tf(bytes[11]);
-                advance_len = 12;
+                let (block_wand_term_freq, num_bytes) = read_u32_vint_no_advance(&bytes[11..]);
+                self.owned_read.advance(11 + num_bytes);
                self.block_info = BlockInfo::BitPacked {
                    doc_num_bits,
                    tf_num_bits,
@@ -212,7 +213,6 @@ impl SkipReader {
                };
            }
        }
-        self.owned_read.advance(advance_len);
    }

    pub fn block_info(&self) -> BlockInfo {
@@ -274,24 +274,6 @@ mod tests {
    use crate::directory::OwnedBytes;
    use crate::postings::compression::COMPRESSION_BLOCK_SIZE;

-    #[test]
-    fn test_encode_block_wand_max_tf() {
-        for tf in 0..255 {
-            assert_eq!(super::encode_block_wand_max_tf(tf), tf as u8);
-        }
-        for &tf in &[255, 256, 1_000_000, u32::MAX] {
-            assert_eq!(super::encode_block_wand_max_tf(tf), 255);
-        }
-    }
-
-    #[test]
-    fn test_decode_block_wand_max_tf() {
-        for tf in 0..255 {
-            assert_eq!(super::decode_block_wand_max_tf(tf), tf as u32);
-        }
-        assert_eq!(super::decode_block_wand_max_tf(255), u32::MAX);
-    }
-
    #[test]
    fn test_skip_with_freq() {
        let buf = {
--- a/src/query/term_query/term_scorer.rs
+++ b/src/query/term_query/term_scorer.rs
@@ -302,7 +302,7 @@ mod tests {
        let mut rng = rand::thread_rng();
        writer.set_merge_policy(Box::new(NoMergePolicy));
        for _ in 0..3_000 {
-            let term_freq = rng.gen_range(1..10000);
+            let term_freq = rng.gen_range(1, 10000);
            let words: Vec<&str> = std::iter::repeat("bbbb").take(term_freq).collect();
            let text = words.join(" ");
            writer.add_document(doc!(text_field=>text));
--- a/src/schema/named_field_document.rs
+++ b/src/schema/named_field_document.rs
@@ -1,5 +1,5 @@
 use crate::schema::Value;
-use serde::{Deserialize, Serialize};
+use serde::Serialize;
 use std::collections::BTreeMap;

 /// Internal representation of a document used for JSON
@@ -8,5 +8,5 @@ use std::collections::BTreeMap;
 /// A `NamedFieldDocument` is a simple representation of a document
 /// as a `BTreeMap<String, Vec<Value>>`.
 ///
-#[derive(Debug, Deserialize, Serialize)]
+#[derive(Serialize)]
 pub struct NamedFieldDocument(pub BTreeMap<String, Vec<Value>>);
--- a/src/store/index/block.rs
+++ b/src/store/index/block.rs
@@ -43,9 +43,6 @@ impl CheckpointBlock {

    /// Adding another checkpoint in the block.
    pub fn push(&mut self, checkpoint: Checkpoint) {
-        if let Some(prev_checkpoint) = self.checkpoints.last() {
-            assert!(checkpoint.follows(prev_checkpoint));
-        }
        self.checkpoints.push(checkpoint);
    }

--- a/src/store/index/mod.rs
+++ b/src/store/index/mod.rs
@@ -26,12 +26,6 @@ pub struct Checkpoint {
    pub end_offset: u64,
 }

-impl Checkpoint {
-    pub(crate) fn follows(&self, other: &Checkpoint) -> bool {
-        (self.start_doc == other.end_doc) && (self.start_offset == other.end_offset)
-    }
-}
-
 impl fmt::Debug for Checkpoint {
    fn fmt(&self, f: &mut fmt::Formatter) -> fmt::Result {
        write!(
@@ -45,16 +39,13 @@ impl fmt::Debug for Checkpoint {
 #[cfg(test)]
 mod tests {

-    use std::{io, iter};
+    use std::io;

-    use futures::executor::block_on;
    use proptest::strategy::{BoxedStrategy, Strategy};

    use crate::directory::OwnedBytes;
-    use crate::indexer::NoMergePolicy;
-    use crate::schema::{SchemaBuilder, STORED, STRING};
    use crate::store::index::Checkpoint;
-    use crate::{DocAddress, DocId, Index, Term};
+    use crate::DocId;

    use super::{SkipIndex, SkipIndexBuilder};

@@ -63,7 +54,7 @@ mod tests {
        let mut output: Vec<u8> = Vec::new();
        let skip_index_builder: SkipIndexBuilder = SkipIndexBuilder::new();
        skip_index_builder.write(&mut output)?;
-        let skip_index: SkipIndex = SkipIndex::open(OwnedBytes::new(output));
+        let skip_index: SkipIndex = SkipIndex::from(OwnedBytes::new(output));
        let mut skip_cursor = skip_index.checkpoints();
        assert!(skip_cursor.next().is_none());
        Ok(())
@@ -81,7 +72,7 @@ mod tests {
        };
        skip_index_builder.insert(checkpoint);
        skip_index_builder.write(&mut output)?;
-        let skip_index: SkipIndex = SkipIndex::open(OwnedBytes::new(output));
+        let skip_index: SkipIndex = SkipIndex::from(OwnedBytes::new(output));
        let mut skip_cursor = skip_index.checkpoints();
        assert_eq!(skip_cursor.next(), Some(checkpoint));
        assert_eq!(skip_cursor.next(), None);
@@ -95,7 +86,7 @@ mod tests {
            Checkpoint {
                start_doc: 0,
                end_doc: 3,
-                start_offset: 0,
+                start_offset: 4,
                end_offset: 9,
            },
            Checkpoint {
@@ -130,7 +121,7 @@ mod tests {
        }
        skip_index_builder.write(&mut output)?;

-        let skip_index: SkipIndex = SkipIndex::open(OwnedBytes::new(output));
+        let skip_index: SkipIndex = SkipIndex::from(OwnedBytes::new(output));
        assert_eq!(
            &skip_index.checkpoints().collect::<Vec<_>>()[..],
            &checkpoints[..]
@@ -142,40 +133,6 @@ mod tests {
        (doc as u64) * (doc as u64)
    }

-    #[test]
-    fn test_merge_store_with_stacking_reproducing_issue969() -> crate::Result<()> {
-        let mut schema_builder = SchemaBuilder::default();
-        let text = schema_builder.add_text_field("text", STORED | STRING);
-        let body = schema_builder.add_text_field("body", STORED);
-        let schema = schema_builder.build();
-        let index = Index::create_in_ram(schema);
-        let mut index_writer = index.writer_for_tests()?;
-        index_writer.set_merge_policy(Box::new(NoMergePolicy));
-        let long_text: String = iter::repeat("abcdefghijklmnopqrstuvwxyz")
-            .take(1_000)
-            .collect();
-        for _ in 0..20 {
-            index_writer.add_document(doc!(body=>long_text.clone()));
-        }
-        index_writer.commit()?;
-        index_writer.add_document(doc!(text=>"testb"));
-        for _ in 0..10 {
-            index_writer.add_document(doc!(text=>"testd", body=>long_text.clone()));
-        }
-        index_writer.commit()?;
-        index_writer.delete_term(Term::from_field_text(text, "testb"));
-        index_writer.commit()?;
-        let segment_ids = index.searchable_segment_ids()?;
-        block_on(index_writer.merge(&segment_ids))?;
-        let reader = index.reader()?;
-        let searcher = reader.searcher();
-        assert_eq!(searcher.num_docs(), 30);
-        for i in 0..searcher.num_docs() as u32 {
-            let _doc = searcher.doc(DocAddress(0u32, i))?;
-        }
-        Ok(())
-    }
-
    #[test]
    fn test_skip_index_long() -> io::Result<()> {
        let mut output: Vec<u8> = Vec::new();
@@ -193,28 +150,26 @@ mod tests {
        }
        skip_index_builder.write(&mut output)?;
        assert_eq!(output.len(), 4035);
-        let resulting_checkpoints: Vec<Checkpoint> = SkipIndex::open(OwnedBytes::new(output))
+        let resulting_checkpoints: Vec<Checkpoint> = SkipIndex::from(OwnedBytes::new(output))
            .checkpoints()
            .collect();
        assert_eq!(&resulting_checkpoints, &checkpoints);
        Ok(())
    }

-    fn integrate_delta(vals: Vec<u64>) -> Vec<u64> {
-        let mut output = Vec::with_capacity(vals.len() + 1);
-        output.push(0u64);
+    fn integrate_delta(mut vals: Vec<u64>) -> Vec<u64> {
        let mut prev = 0u64;
-        for val in vals {
-            let new_val = val + prev;
+        for val in vals.iter_mut() {
+            let new_val = *val + prev;
            prev = new_val;
-            output.push(new_val);
+            *val = new_val;
        }
-        output
+        vals
    }

    // Generates a sequence of n valid checkpoints, with n < max_len.
    fn monotonic_checkpoints(max_len: usize) -> BoxedStrategy<Vec<Checkpoint>> {
-        (0..max_len)
+        (1..max_len)
            .prop_flat_map(move |len: usize| {
                (
                    proptest::collection::vec(1u64..20u64, len as usize).prop_map(integrate_delta),
@@ -266,7 +221,7 @@ mod tests {
             }
             let mut buffer = Vec::new();
             skip_index_builder.write(&mut buffer).unwrap();
-             let skip_index = SkipIndex::open(OwnedBytes::new(buffer));
+             let skip_index = SkipIndex::from(OwnedBytes::new(buffer));
             let iter_checkpoints: Vec<Checkpoint> = skip_index.checkpoints().collect();
             assert_eq!(&checkpoints[..], &iter_checkpoints[..]);
             test_skip_index_aux(skip_index, &checkpoints[..]);
--- a/src/store/index/skip_index.rs
+++ b/src/store/index/skip_index.rs
@@ -35,11 +35,11 @@ struct Layer {
 }

 impl Layer {
-    fn cursor(&self) -> impl Iterator<Item = Checkpoint> + '_ {
+    fn cursor<'a>(&'a self) -> impl Iterator<Item = Checkpoint> + 'a {
        self.cursor_at_offset(0u64)
    }

-    fn cursor_at_offset(&self, start_offset: u64) -> impl Iterator<Item = Checkpoint> + '_ {
+    fn cursor_at_offset<'a>(&'a self, start_offset: u64) -> impl Iterator<Item = Checkpoint> + 'a {
        let data = &self.data.as_slice();
        LayerCursor {
            remaining: &data[start_offset as usize..],
@@ -59,25 +59,7 @@ pub struct SkipIndex {
 }

 impl SkipIndex {
-    pub fn open(mut data: OwnedBytes) -> SkipIndex {
-        let offsets: Vec<u64> = Vec::<VInt>::deserialize(&mut data)
-            .unwrap()
-            .into_iter()
-            .map(|el| el.0)
-            .collect();
-        let mut start_offset = 0;
-        let mut layers = Vec::new();
-        for end_offset in offsets {
-            let layer = Layer {
-                data: data.slice(start_offset as usize, end_offset as usize),
-            };
-            layers.push(layer);
-            start_offset = end_offset;
-        }
-        SkipIndex { layers }
-    }
-
-    pub(crate) fn checkpoints(&self) -> impl Iterator<Item = Checkpoint> + '_ {
+    pub(crate) fn checkpoints<'a>(&'a self) -> impl Iterator<Item = Checkpoint> + 'a {
        self.layers
            .last()
            .into_iter()
@@ -108,3 +90,22 @@ impl SkipIndex {
        Some(cur_checkpoint)
    }
 }
+
+impl From<OwnedBytes> for SkipIndex {
+    fn from(mut data: OwnedBytes) -> SkipIndex {
+        let offsets: Vec<u64> = Vec::<VInt>::deserialize(&mut data)
+            .unwrap()
+            .into_iter()
+            .map(|el| el.0)
+            .collect();
+        let mut start_offset = 0;
+        let mut layers = Vec::new();
+        for end_offset in offsets {
+            layers.push(Layer {
+                data: data.slice(start_offset as usize, end_offset as usize),
+            });
+            start_offset = end_offset;
+        }
+        SkipIndex { layers }
+    }
+}
--- a/src/store/index/skip_index_builder.rs
+++ b/src/store/index/skip_index_builder.rs
@@ -28,20 +28,18 @@ impl LayerBuilder {
    ///
    /// If the block was empty to begin with, simply return None.
    fn flush_block(&mut self) -> Option<Checkpoint> {
-        if let Some((start_doc, end_doc)) = self.block.doc_interval() {
+        self.block.doc_interval().map(|(start_doc, end_doc)| {
            let start_offset = self.buffer.len() as u64;
            self.block.serialize(&mut self.buffer);
            let end_offset = self.buffer.len() as u64;
            self.block.clear();
-            Some(Checkpoint {
+            Checkpoint {
                start_doc,
                end_doc,
                start_offset,
                end_offset,
-            })
-        } else {
-            None
-        }
+            }
+        })
    }

    fn push(&mut self, checkpoint: Checkpoint) {
@@ -50,7 +48,7 @@ impl LayerBuilder {

    fn insert(&mut self, checkpoint: Checkpoint) -> Option<Checkpoint> {
        self.push(checkpoint);
-        let emit_skip_info = self.block.len() >= CHECKPOINT_PERIOD;
+        let emit_skip_info = (self.block.len() % CHECKPOINT_PERIOD) == 0;
        if emit_skip_info {
            self.flush_block()
        } else {
--- a/src/store/reader.rs
+++ b/src/store/reader.rs
@@ -35,7 +35,7 @@ impl StoreReader {
        let (data_file, offset_index_file) = split_file(store_file)?;
        let index_data = offset_index_file.read_bytes()?;
        let space_usage = StoreSpaceUsage::new(data_file.len(), offset_index_file.len());
-        let skip_index = SkipIndex::open(index_data);
+        let skip_index = SkipIndex::from(index_data);
        Ok(StoreReader {
            data: data_file,
            cache: Arc::new(Mutex::new(LruCache::new(LRU_CACHE_CAPACITY))),
@@ -46,7 +46,7 @@ impl StoreReader {
        })
    }

-    pub(crate) fn block_checkpoints(&self) -> impl Iterator<Item = Checkpoint> + '_ {
+    pub(crate) fn block_checkpoints<'a>(&'a self) -> impl Iterator<Item = Checkpoint> + 'a {
        self.skip_index.checkpoints()
    }

--- a/src/store/writer.rs
+++ b/src/store/writer.rs
@@ -72,7 +72,6 @@ impl StoreWriter {
        if !self.current_block.is_empty() {
            self.write_and_compress_block()?;
        }
-        assert_eq!(self.first_doc_in_block, self.doc);
        let doc_shift = self.doc;
        let start_shift = self.writer.written_bytes() as u64;

@@ -87,17 +86,12 @@ impl StoreWriter {
            checkpoint.end_doc += doc_shift;
            checkpoint.start_offset += start_shift;
            checkpoint.end_offset += start_shift;
-            self.register_checkpoint(checkpoint);
+            self.offset_index_writer.insert(checkpoint);
+            self.doc = checkpoint.end_doc;
        }
        Ok(())
    }

-    fn register_checkpoint(&mut self, checkpoint: Checkpoint) {
-        self.offset_index_writer.insert(checkpoint);
-        self.first_doc_in_block = checkpoint.end_doc;
-        self.doc = checkpoint.end_doc;
-    }
-
    fn write_and_compress_block(&mut self) -> io::Result<()> {
        assert!(self.doc > 0);
        self.intermediary_buffer.clear();
@@ -106,13 +100,14 @@ impl StoreWriter {
        self.writer.write_all(&self.intermediary_buffer)?;
        let end_offset = self.writer.written_bytes();
        let end_doc = self.doc;
-        self.register_checkpoint(Checkpoint {
+        self.offset_index_writer.insert(Checkpoint {
            start_doc: self.first_doc_in_block,
            end_doc,
            start_offset,
            end_offset,
        });
        self.current_block.clear();
+        self.first_doc_in_block = self.doc;
        Ok(())
    }
Author	SHA1	Message	Date
Paul Masurel	c5d50e2138	fixing compilation	2020-12-09 17:14:41 +09:00
Paul Masurel	af2f067b69	Removed 'static in compression_lz4.	2020-12-09 16:57:01 +09:00
Paul Masurel	c056140d81	Reorganized and added termdict unit tests.	2020-12-09 16:57:01 +09:00
Paul Masurel	c6204f49d3	Minor changes - Open{Write,Read}Error::wrap_io_error made public - Arc<PathBuf> -> Arc<Path> in file_watcher.	2020-12-09 16:57:01 +09:00
Paul Masurel	738a1a0188	Small refactoring	2020-12-09 16:57:01 +09:00
Paul Masurel	294e8b4659	TermDictionary.finish does not flush	2020-12-09 16:57:01 +09:00
Paul Masurel	2f342257d3	Several TermDict operation now returns an io::Result	2020-12-09 16:57:01 +09:00
Paul Masurel	e9e2984ac2	Simplified counting writer and removed flush	2020-12-09 16:57:01 +09:00
Paul Masurel	e1f9271be3	Moved the term merger	2020-12-09 16:57:01 +09:00
Paul Masurel	883eb92df9	Cargo fmt	2020-12-09 16:57:01 +09:00
Paul Masurel	590654ceb8	Isolated fst impl of termdictionary in a specific module.	2020-12-09 16:57:01 +09:00
Paul Masurel	7367eb5455	DocSet is send	2020-12-09 16:55:59 +09:00
Paul Masurel	e22330c7e0	Attempt to fix bug surfacing sometimes in test. Recently, `test_index_manual_policy_mmap` has been failing on Windows. The idea addressed by this patch is that we forget to sync the parent directory with the current implementation of atomic writes. This was done correctly when we were relying the atomicwrites crate. crossing fingers	2020-12-09 16:55:59 +09:00
Paul Masurel	ad4c2be21b	Fix perf regression in the benchmark for the Count collector. In order to reduce IO, we introduced a way to instanciate a dummy constant FieldnormReader which worked by allocating a buffer with as many bytes as there are docs in the segments. This allocation is not a negligible by any mean. This PR works by offering two implementation for the FieldnormReader. The const field norm reader simply returns the same value all of the time, while the array based one does the same as the current one.	2020-12-09 16:55:59 +09:00
Paul Masurel	412c83c336	Added specialized implementation for count/count_including... in &mut DocSet	2020-12-09 16:55:59 +09:00
Paul Masurel	3006db17c1	Avoid computing the BM25 weight if scoring is disabled	2020-12-09 16:55:59 +09:00
Paul Masurel	b4fc185dc5	Applied CR comments	2020-12-09 16:55:59 +09:00
Adrien Guillo	20a314093f	Replace some `Arc<Box<dyn...` with `Arc<dyn...`	2020-12-09 16:55:59 +09:00
Paul Masurel	9b3eb59e9b	No filelen problem.	2020-12-09 16:55:59 +09:00
Adrien Guillo	6d33ae307a	Add helper methods for reading u8 and u64 to `OwnedBytes`	2020-12-09 16:55:59 +09:00