tantivy

mirror of https://github.com/quickwit-oss/tantivy.git synced 2026-01-08 01:52:54 +00:00

Author	SHA1	Message	Date
Mathias Svensson	6a8a8557d2	Use `slice::iter` instead of `into_iter` to avoid future breakage (#679 ) * Use `slice::iter` instead of `into_iter` to avoid future breakage `an_array.into_iter()` currently just works because of the autoref feature, which then calls `<[T] as IntoIterator>::into_iter`. But in the future, arrays will implement `IntoIterator`, too. In order to avoid problems in the future, the call is replaced by `iter()` which is shorter and more explicit. * cargo fmt	2019-10-31 20:59:50 +09:00
Alberto Piai	3a65dc84c8	TopDocs: ensure stable sorting on equal score (#675 ) * TopDocs: ensure stable sorting on equal score When selecting the top K documents by score, we need to ensure stable sorting. Until now, for documents with the same score, we were relying on the (arbitrary) order returned by the BinaryHeap used to implement the collectors. This patch fixes the problem by explicitly using the doc address when harvesting the `TopSegmentCollector` and when merging the results in `TopCollector::merge_fruits()`. This is important (for example) to implement pagination correctly using the TopDocs collector. If sorting isn't stable, documents that have the same score might be ranked in different positions depending on the specific K that was used, thus appearing in two different pages, or in none at all. Fixes gh-671 * TMP: alternative solution (see previous commit) If we add the constrait that D is also PartialOrd in ComparableDoc<T, D>, then we can move the comparison by doc address directly in the cmp implementation of ComparableDoc. * TMP rebase as first commit: add benchmarks for TopSegmentCollector * fixup! TMP: alternative solution (see previous commit) * TMP add changelog entry * TMP run cargo fmt	2019-10-26 15:27:25 +09:00
dependabot-preview[bot]	ce42bbf5c9	Update base64 requirement from 0.10.0 to 0.11.0 (#676 ) Updates the requirements on [base64](https://github.com/marshallpierce/rust-base64) to permit the latest version. - [Release notes](https://github.com/marshallpierce/rust-base64/releases) - [Changelog](https://github.com/marshallpierce/rust-base64/blob/master/RELEASE-NOTES.md) - [Commits](https://github.com/marshallpierce/rust-base64/compare/v0.10.0...v0.11.0) Signed-off-by: dependabot-preview[bot] <support@dependabot.com>	2019-10-26 15:24:47 +09:00
Paul Masurel	7b21b3f25a	Refactoring around Field (#673 ) * Refactoring around Field Removing the contract about the order of the field, and the field id allocation. * Update delete_queue.rs * Update field.rs	2019-10-25 09:06:44 +09:00
Paul Masurel	46caec1040	Updating uuid to 0.8 (#674 )	2019-10-25 09:02:00 +09:00
petr-tik	1187a02a3e	Fixed #664 (#667 ) Removed references to u8 and old documentation	2019-10-22 09:34:10 +09:00
Andrew Banchich	f6c525b19e	Fix grammar / punctuation (#668 )	2019-10-21 10:50:53 +09:00
petr-tik	4a8f7712f3	Add a doctest to BooleanQuery (#630 ) * Add a doctest to BooleanQuery Closes #446 Mark a function that is only used in tests to be compiled for tests only Fix doc-comments in a couple of related files * Minor corrections remove whitespace, fix typos, add explicit dyn marker * WIP: BooleanQuery doc test Trying to nest several BooleanQueries together * Addressed old review rust 2018 edition + make function available to everyone * Box the previous query to resolve the type error * Rework wording in DocAdress document strings * Reworded and restructured the docstring	2019-10-07 10:05:12 +09:00
Paul Masurel	2f867aad17	Fix bench (#663 ) * fmt * Fixing bench compilation	2019-10-04 17:07:49 +09:00
Paul Masurel	5c6580eb15	fmt (#661 )	2019-10-04 12:10:01 +09:00
Paul Masurel	4c3941750b	Waiting potentially longer on watch	2019-10-01 09:50:46 +09:00
Paul Masurel	2ea8e618f2	Merge branch 'hotfix-656'	2019-10-01 09:44:56 +09:00
Paul Masurel	94f27f990b	Address #656 Broke the reference loop to make sure that the watch_router can be dropped, and the thread exits.	2019-10-01 09:34:22 +09:00
Paul Masurel	349e8aa348	Removed enum variants on type alias	2019-09-26 18:43:29 +09:00
Paul Masurel	cde9b78b8d	Fixing the issue associated with the Regex performance change	2019-09-18 18:29:27 +09:00
fdb-hiroshima	d8894f0bd2	add checksum check in ManagedDirectory (#605 ) * add checksum check in ManagedDirectory fix #400 * flush after writing checksum * don't checksum atomic file access and clone managed_paths * implement a footer storing metadata about a file this is more of a poc, it require some refactoring into multiple files `terminate(self)` is implemented, but not used anywhere yet * address comments and simplify things with new contract use BitOrder for integer to raw byte conversion consider atomic write imply atomic read, which might not actually be true use some indirection to have a boxable terminating writer * implement TerminatingWrite and make terminate() be called where it should add dependancy to drop_bomb to help find where terminate() should be called implement TerminatingWrite for wrapper writers make tests pass /!\ some tests seems to pass where they shouldn't * remove usage of drop_bomb * fmt * add test for checksum * address some review comments * update changelog * fmt	2019-09-18 18:26:25 +09:00
fdb-hiroshima	7e08e0047b	fix Term documentation (#655 ) u64-based fields are actually 4+8=12 bytes long	2019-09-11 18:49:35 +09:00
fdb-hiroshima	1a817f117f	fix documentation error (#654 ) Union missdocumented as doing an intersection Union and Intersection can hold more than 2 DocSets	2019-09-11 17:12:08 +09:00
petr-tik	2ec19b21ae	Remove unnecessary duplicate methods (#650 ) Closes #649 Spotted by @imor	2019-09-09 06:36:04 +09:00
Raminder Singh	141f5a93f7	Using FnvHashMap for mapping UnorderedTermId to TermOrdinal. Fixes #507 (#647 ) * Using FnvHashMap for mapping UnorderedTermId to TermOrdinal. Fixes #507 * Fixed cargo fmt errors	2019-09-07 19:40:21 +09:00
Paul Masurel	df47d55cd2	Occur debug interface (#648 )	2019-09-07 15:08:45 +09:00
Raminder Singh	5e579fd6b7	Fixed clippy warning: unneeded return statement (#646 )	2019-09-07 10:14:37 +09:00
Paul Masurel	4b9c1dce69	Moving queyr grammar to a different crate. (#645 )	2019-09-05 09:37:28 +09:00
Paul Masurel	d74f71bbef	Lighter regex dependency. (#644 ) Detail on https://github.com/rust-lang/regex/pull/613	2019-09-04 13:10:12 +09:00
Paul Masurel	5196ca41d8	Small code clean up	2019-09-03 09:22:32 +09:00
dependabot-preview[bot]	4959e06151	Update once_cell requirement from 0.2 to 1.0 (#643 ) Updates the requirements on [once_cell](https://github.com/matklad/once_cell) to permit the latest version. - [Release notes](https://github.com/matklad/once_cell/releases) - [Changelog](https://github.com/matklad/once_cell/blob/master/CHANGELOG.md) - [Commits](https://github.com/matklad/once_cell/compare/v0.2.0...v1.0.2) Signed-off-by: dependabot-preview[bot] <support@dependabot.com>	2019-09-03 07:00:45 +09:00
Paul Masurel	c1635c13f6	RegexQuery performance: make it possible to cache Regexes - remastered by fulmicoton (Closes #639 ) (#641 ) * small docs cleanup * only compile a regex once per RegexQuery Building a `Regex` is an expensive operation. Users of `RegexQuery` need to cache and reuse regexes when searching across multiple fields. This is the first step towards allowing that: we can store the `Regex` directly in the `RegexQuery`, instead of the string pattern. * RegexQuery: account for possible failure in the constructor When building a regex from a str pattern, we have to account for the possibility that the pattern is invalid. Before the previous commit, the failure would happen in the `specialized_weight` method. Now that we store a compiled `Regex` in `RegexQuery`, `specialized_weight` doesn't fail anymore, and we can fail early while constructing `RegexQuery` if the pattern is invalid. This is a breaking change for users of `RegexQuery::new`. * add RegexQuery::from_regex method This builds a `RegexQuery` from an already compiled `Regex`. The use of `Into<Arc<Regex>>` is to allow the caller to either simply pass a `Regex`, or an `Arc<Regex>`, in case it needs to be cached and shared on the caller's side. * Using an Arc in AutomatonWeight Closes #639	2019-08-22 16:14:01 +09:00
Paul Masurel	135e0ea2e9	Expose new segment meta from Index (#637 )	2019-08-19 10:39:15 +09:00
Paul Masurel	f283bfd7ab	Added segmentid_from_string (#636 )	2019-08-19 10:37:30 +09:00
Joshua Dutton	9f74786db2	Update import statements in examples, doctests (#633 ) Update import statements to edition 2018, including removing `extern crate` and `#[macro_use]`. Alphabetize the statements.	2019-08-19 07:26:35 +09:00
Joshua Dutton	32e5d7a0c7	Fix trait object in doctest (#635 )	2019-08-19 07:25:00 +09:00
Joshua Dutton	84c615cff1	Fixing typos (#634 )	2019-08-19 07:24:05 +09:00
Paul Masurel	039c0a0863	Introducing a wrapper struct instead of Boxed<BoxableTokenizer> (#631 ) Closes #629	2019-08-15 16:37:04 +09:00
Paul Masurel	b3b0138b82	Change for tantivy-py Schema.convert_named_doc Better Debug string for Terms and TermQueries	2019-08-14 17:44:25 +09:00
petr-tik	ea56160cdc	Added cargo-fmt to CI runs (#627 ) * Added cargo-fmt to CI runs Closes #625 * Remove fmt from appveyor builds Windows seems to have issues with install components through rustup. Formatting should be equally informative regardless of the OS, so best to keep it in Linux on Travis	2019-08-12 08:25:47 +09:00
petr-tik	028b0a749c	Elastic unbounded range query (#624 ) * Tidy up fmt remove unneccessary -> Result<()> followed by run.unwrap() in a test * Adding support for elasticsearch-style unbounded queries Extend the UserInputBound to include Unbounded, so we can reuse formatting and internal query format * Still working on elastic-style range queries Fixes #498 Merge the elastic_range into range Reformat to make code easier to follow, use optional() macro to return Some * Fixed bugs Made the range parser insensitive to whitespace between the ":" and the range. Removed optional parsing of field. Added a unit test for the range parser. Derived PartialEq to compare the results of parsing as structs, instead of strings. Found a bug with that unit test - "}" was parsed as an UserInputBound::Exclusive, instead of UserInputBound::Unbounded. Added an early detection-and-return for in the original range parser * Correct failing test Assume that we will use "{" for Unbounded ranges Add a note in the changelog cargo-fmt * Moved parenthesis to a newline to make nested if-else more visible	2019-08-12 08:24:47 +09:00
Paul Masurel	941f06eb9f	Added Schema.from_named_doc	2019-08-11 16:50:32 +09:00
Paul Masurel	04832a86eb	WTF is this file doing here (#622 )	2019-08-08 21:54:10 +09:00
fdb-hiroshima	beb8e990cd	fix parsing neg float in range query (#621 ) fix #620	2019-08-08 20:41:04 +09:00
Paul Masurel	001af3876f	cargo fmt	2019-08-08 18:07:19 +09:00
Paul Masurel	f428f344da	Various bugfix in the query parser (#619 )	2019-08-08 17:48:21 +09:00
Paul Masurel	143f78eced	Trying to fix #609 (#616 )	2019-08-06 20:33:30 +09:00
Kornel	754b55eee5	Bump deps (#613 ) * Bump crossbeam * Warnings-- * Remove outdated tempdir	2019-08-05 22:21:22 +09:00
Paul Masurel	280ea1209c	Changes required for python binding (#610 )	2019-08-01 17:26:21 +09:00
petr-tik	0154dbe477	Replace unwrap with match and proper Error handling (#606 ) * Replace unwrap with match and proper Error handling * Replaced 'magic' values with a documented variable Didn't like the unexplained 0..3 range, thought it was best as a variable Calculating Levenshtein distance is expensive, so best explain why we should keep it low	2019-07-31 08:16:02 +09:00
Paul Masurel	efd1af1325	Closes #544 . (#607 ) Prepare for release 0.10.1	2019-07-30 13:38:06 +09:00
fdb-hiroshima	c91eb7fba7	add to_path for Facet (#604 ) fix #580 0.10.1	2019-07-27 17:58:43 +09:00
fdb-hiroshima	6eb4e08636	add support for float (#603 ) * add basic support for float as for i64, they are mapped to u64 for indexing query parser don't work yet * Update value.rs * implement support for float in query parser * Update README.md	2019-07-27 17:57:33 +09:00
Paul Masurel	c3231ca252	Added phrase query tests (#601 )	2019-07-22 13:43:00 +09:00
Paul Masurel	7211df6719	Failrs (#600 ) * Single thread tests * Isolating fail tests into a different binary	2019-07-22 13:17:21 +09:00

1 2 3 4 5 ...

1475 Commits