Change Footer version handling, Make compression dynamic
Change Footer version handling
Simplify version handling by switching to JSON instead of binary serialization.
fixes#1058
Make compression dynamic
Instead of choosing the compression during compile time via a feature flag, you can now have multiple compression algorithms enabled and decide during runtime which one to choose via IndexSettings. Changing the compression algorithm on an index is also supported. The information which algorithm was used in the doc store is stored in the DocStoreFooter. The default is the lz4 block format.
fixes#904
Handle merging of different compressors
Fix feature flag names
Add doc store test for all compressors
* sort index by field
add sort info to IndexSettings
generate docid mapping for sorted field (only fastfield)
remap singlevalue fastfield
* support docid mapping in multivalue fastfield
move docid mapping to serialization step (less intermediate data for mapping)
add support for docid mapping in multivalue fastfield
* handle docid map in bytes fastfield
* forward docid mapping, remap postings
* fix merge conflicts
* move test to index_sorter
* add docid index mapping old->new
add docid mapping for both directions old->new (used in postings) and new->old (used in fast field)
handle mapping in postings recorder
warn instead of info for MAX_TOKEN_LEN
* remap docid in fielnorm
* resort docids in recorder, more extensive tests
* handle index sorting in docstore
handle index sort in docstore, by saving all the docs in a temp docstore file (SegmentComponent::TempStore). On serialization the docid mapping is used to create a docstore in the correct order by reader the old docstore.
add docstore sort tests
refactor tests
* refactor
rename docid doc_id
rename docid_map doc_id_map
rename DocidMapping DocIdMapping
fix typo
* u32 to DocId
* better doc_id_map creation
remove unstable sort
* add non mut method to FastFieldWriters
add _mut prefix to &mut methods
* remove sort_index
* fix clippy issues
* fix SegmentComponent iterator
use std::mem::replace
* fix test
* fmt
* handle indexsettings deserialize
* add reading, writing bytes to doc store
get bytes of document in doc store
add store_bytes method doc writer to accept serialized document
add serialization index settings test
* rename index_sorter to doc_id_mapping
use bufferlender in recorder
* fix compile issue, make sort_by_field optional
* fix test compile
* validate index settings on merge
validate index settings on merge
forward merge info to SegmentSerializer (for TempStore)
* fix doctest
* add itertools, use kmerge
add itertools, use kmerge
push because rustfmt fails
* implement/test merge for fastfield
implement/test merge for fastfield
rename len to num_deleted in DeleteBitSet
* Use precalculated docid mapping in merger
Use precalculated docid mapping in merger for sorted indices instead of on the fly calculation
Add index creation macro benchmark, but commented out for now, since it is not really usable due to long runtimes, and extreme fluctuations. May be better suited in criterion or an external bench bin
* fix fast field reader docs
fix fast field reader docs, Error instead of None returned
add u64s_lenient to fastreader
add create docid mapping benchmark
* add test for multifast field merge
refactor test
add test for multifast field merge
* add num_bytes to BytesFastFieldReader
equivalent to num_vals in MultiValuedFastFieldReader
* add MultiValueLength trait
add MultiValueLength trait in order to unify index creation for BytesFastFieldReader and MultiValuedFastFieldReader in merger
* Add ReaderWithOrdinal, fix
Add ReaderWithOrdinal to associate data to a reader in merger
Fix bytes offset index creation in merger
* add test for merging bytes with sorted docids
* Merge fieldnorm for sorted index
* handle posting list in merge in sorted index
handle posting list in merge in sorted index by using doc id mapping for sorting
reuse SegmentOrdinal type
* handle doc store order in merge in sorted index
* fix typo, cleanup
* make IndexSetting non-optional
* fix type, rename test file
fix type
rename test file
add type
* remove SegmentReaderWithOrdinal accessors
* cargo fmt
* add index sort & merge test to include deletes
* Fix posting list merge issue
Fix posting list merge issue - ensure serializer always gets monotonically increasing doc ids
handle sorting and merging for facets field
* performance: cache field readers, use bytes for doc store merge
* change facet merge test to cover index sorting
* add RawDocument abstraction to access bytes in doc store
* fix deserialization, update changelog
fix deserialization
update changelog
forward error on merge failed
* cache store readers to utilize lru cache (4x performance)
cache store readers, to utilize lru cache (4x faster performance, due to less decompress calls on the block)
* add include_temp_doc_store flag in InnerSegmentMeta
unset flag on deserialization and after finalize of a segment
set flag when creating new instances
Tantivy used to assume that all files could be somehow memory mapped. After this change, Directory return a `FileSlice` that can be reduced and eventually read into an `OwnedBytes` object. Long and blocking io operation are still required by they do not span over the entire file.
* add support for indexed bytes fast field
* remove backup code file
* refine test cases
* Simplified unit test. Renamed it as it is testing the storable part. Not the indexed part.
* Small refactoring and added unit test. If multivalued we only retain the first FAST value.
Co-authored-by: Raul <raul.tang.lc@gmail.com>
- Change in the DocSet and Scorer API. (@fulmicoton).
A freshly created DocSet point directly to their first doc. A sentinel value called TERMINATED marks the end of a DocSet.
`.advance()` returns the new DocId. `Scorer::skip(target)` has been replaced by `Scorer::seek(target)` and returns the resulting DocId.
As a result, iterating through DocSet now looks as follows
```rust
let mut doc = docset.doc();
while doc != TERMINATED {
// ...
doc = docset.advance();
}
```
The change made it possible to greatly simplify a lot of the docset's code.
- Misc internal optimization and introduction of the `Scorer::for_each_pruning` function. (@fulmicoton)
* Alternative take on boosted queries
* Fixing unit test
* Added boosting to the query grammar.
* Made BoostQuery public.
* Added support for boosting field in QueryParser
Closes#547
* Added backwards iteration to termdict
* Ran formatter
* Updated fst dependency
* Updated dependency
* Changelog and version
* Fixed version
* Made it part of 12.0