lancedb

mirror of https://github.com/lancedb/lancedb.git synced 2025-12-24 22:09:58 +00:00

Author	SHA1	Message	Date
QianZhu	1b990983b3	Qian/make vector col optional (#950 ) remote SDK tests were completed through lancedb_integtest	2024-02-12 16:35:44 -08:00
Weston Pace	a9727eb318	feat: add support for filter during merge insert when matched (#948 ) Closes #940	2024-02-09 10:26:14 -08:00
Ayush Chaurasia	d982ee934a	feat(python): Reranker DX improvements (#904 ) - Most users might not know how to use `QueryBuilder` object. Instead we should just pass the string query. - Add new rerankers: Colbert, openai	2024-02-06 13:59:31 +05:30
Will Jones	57605a2d86	feat(python): add `read_consistency_interval` argument (#828 ) This PR refactors how we handle read consistency: does the `LanceTable` class always pick up modifications to the table made by other instance or processes. Users have three options they can set at the connection level: 1. (Default) `read_consistency_interval=None` means it will not check at all. Users can call `table.checkout_latest()` to manually check for updates. 2. `read_consistency_interval=timedelta(0)` means always check for updates, giving strong read consistency. 3. `read_consistency_interval=timedelta(seconds=20)` means check for updates every 20 seconds. This is eventual consistency, a compromise between the two options above. ## Table reference state There is now an explicit difference between a `LanceTable` that tracks the current version and one that is fixed at a historical version. We now enforce that users cannot write if they have checked out an old version. They are instructed to call `checkout_latest()` before calling the write methods. Since `conn.open_table()` doesn't have a parameter for version, users will only get fixed references if they call `table.checkout()`. The difference between these two can be seen in the repr: Table that are fixed at a particular version will have a `version` displayed in the repr. Otherwise, the version will not be shown. ```python >>> table LanceTable(connection=..., name="my_table") >>> table.checkout(1) >>> table LanceTable(connection=..., name="my_table", version=1) ``` I decided to not create different classes for these states, because I think we already have enough complexity with the Cloud vs OSS table references. Based on #812	2024-02-05 08:12:19 -08:00
Ayush Chaurasia	738511c5f2	feat(python): add support new openai embedding functions (#912 ) @PrashantDixit0 --------- Co-authored-by: Chang She <759245+changhiskhan@users.noreply.github.com>	2024-02-04 18:19:42 -08:00
Bert	a248d7feec	fix: add request retry to python client (#917 ) Adds capability to the remote python SDK to retry requests (fixes #911) This can be configured through environment: - `LANCE_CLIENT_MAX_RETRIES`= total number of retries. Set to 0 to disable retries. default = 3 - `LANCE_CLIENT_CONNECT_RETRIES` = number of times to retry request in case of TCP connect failure. default = 3 - `LANCE_CLIENT_READ_RETRIES` = number of times to retry request in case of HTTP request failure. default = 3 - `LANCE_CLIENT_RETRY_STATUSES` = http statuses for which the request will be retried. passed as comma separated list of ints. default `500, 502, 503` - `LANCE_CLIENT_RETRY_BACKOFF_FACTOR` = controls time between retry requests. see [here](`23f2287eb5/src/urllib3/util/retry.py (L141-L146)`). default = 0.25 Only read requests will be retried: - list table names - query - describe table - list table indices This does not add retry capabilities for writes as it could possibly cause issues in the case where the retried write isn't idempotent. For example, in the case where the LB times-out the request but the server completes the request anyway, we might not want to blindly retry an insert request.	2024-02-02 11:27:29 -05:00
Weston Pace	d77e95a4f4	feat: upgrade to lance 0.9.11 and expose merge_insert (#906 ) This adds the python bindings requested in #870 The javascript/rust bindings will be added in a future PR.	2024-02-01 11:36:29 -08:00
Ayush Chaurasia	3ffed89793	feat(python): Hybrid search & Reranker API (#824 ) based on https://github.com/lancedb/lancedb/pull/713 - The Reranker api can be plugged into vector only or fts only search but this PR doesn't do that (see example - https://txt.cohere.com/rerank/) ### Default reranker -- `LinearCombinationReranker(weight=0.7, fill=1.0)` ``` table.search("hello", query_type="hybrid").rerank(normalize="score").to_pandas() ``` ### Available rerankers LinearCombinationReranker ``` from lancedb.rerankers import LinearCombinationReranker # Same as default table.search("hello", query_type="hybrid").rerank( normalize="score", reranker=LinearCombinationReranker() ).to_pandas() # with custom params reranker = LinearCombinationReranker(weight=0.3, fill=1.0) table.search("hello", query_type="hybrid").rerank( normalize="score", reranker=reranker ).to_pandas() ``` Cohere Reranker ``` from lancedb.rerankers import CohereReranker # default model.. English and multi-lingual supported. See docstring for available custom params table.search("hello", query_type="hybrid").rerank( normalize="rank", # score or rank reranker=CohereReranker() ).to_pandas() ``` CrossEncoderReranker ``` from lancedb.rerankers import CrossEncoderReranker table.search("hello", query_type="hybrid").rerank( normalize="rank", reranker=CrossEncoderReranker() ).to_pandas() ``` ## Using custom Reranker ``` from lancedb.reranker import Reranker class CustomReranker(Reranker): def rerank_hybrid(self, vector_result, fts_result): combined_res = self.merge_results(vector_results, fts_results) # or use custom combination logic # Custom rerank logic here return combined_res ``` - [x] Expand testing - [x] Make sure usage makes sense - [x] Run simple benchmarks for correctness (Seeing weird result from cohere reranker in the toy example) - Support diverse rerankers by default: - [x] Cross encoding - [x] Cohere - [x] Reciprocal Rank Fusion --------- Co-authored-by: Chang She <759245+changhiskhan@users.noreply.github.com> Co-authored-by: Prashanth Rao <35005448+prrao87@users.noreply.github.com>	2024-01-30 19:10:33 +05:30
Raghav Dixit	d1a7257810	feat(python): Embedding fn support for gte-mlx/gte-large (#873 ) have added testing and an example in the docstring, will be pushing a separate PR in recipe repo for rag example --------- Co-authored-by: Ayush Chaurasia <ayush.chaurarsia@gmail.com>	2024-01-30 11:21:57 +05:30
Ayush Chaurasia	d84e0d1db8	feat(python): Aws Bedrock embeddings integration (#822 ) Supports amazon titan, cohere english & cohere multi-lingual base models.	2024-01-28 02:04:15 +05:30
Lei Xu	83ed8d1e49	bug: add a test for fp16 (#837 ) Add test to ingest fp16 to a database	2024-01-20 16:23:28 -08:00
Will Jones	d012db24c2	ci: lint and enforce linting (#829 ) @eddyxu added instructions for linting here: `7af213801a/python/README.md (L45-L50)` However, we had a lot of failures and weren't checking this in CI. This PR fixes all lints and adds a check to CI to keep us in compliance with the lints.	2024-01-19 13:09:14 -08:00
Chang She	af8263af94	feat(python): allow the entire table to be converted a polars dataframe (#814 )	2024-01-15 15:49:16 -08:00
Chang She	be4ab9eef3	feat(python): add exist_ok option to create table (#813 ) This mimics CREATE TABLE IF NOT EXISTS behavior. We add `db.create_table(..., exist_ok=True)` parameter. By default it is set to False, so trying to create a table with the same name will raise an exception. If set to True, then it only opens the table if it already exists. If you pass in a schema, it will be checked against the existing table to make sure you get what you want. If you pass in data, it will NOT be added to the existing table.	2024-01-15 11:09:18 -08:00
Chang She	49333e522c	feat(python): basic polars integration (#811 ) We should now be able to directly ingest polars dataframes and return results as polars dataframes ![image](https://github.com/lancedb/lancedb/assets/759245/828b1260-c791-45f1-a047-aa649575e798)	2024-01-13 16:38:16 -08:00
Ayush Chaurasia	4568df422d	feat(python): Add gemini text embedding function (#806 ) Named it Gemini-text for now. Not sure how complicated it will be to support both text and multimodal embeddings under the same class "gemini"..But its not something to worry about for now I guess.	2024-01-12 22:38:55 -08:00
Sebastian Law	99adfe065a	use requests instead of aiohttp for underlying http client (#803 ) instead of starting and stopping the current thread's event loop on every http call, just make an http call.	2024-01-10 00:07:50 -05:00
Chang She	277406509e	chore(python): add docstring for limit behavior (#800 ) Closes #796	2024-01-09 20:20:13 -08:00
Chang She	63411b4d8b	feat(python): add phrase query option for fts (#798 ) addresses #797 Problem: tantivy does not expose option to explicitly Proposed solution here: 1. Add a `.phrase_query()` option 2. Under the hood, LanceDB takes care of wrapping the input in quotes and replace nested double quotes with single quotes I've also filed an upstream issue, if they support phrase queries natively then we can get rid of our manual custom processing here.	2024-01-09 19:41:31 -08:00
Chang She	d998f80b04	feat(python): add count_rows with filter option (#801 ) Closes #795	2024-01-09 19:33:03 -08:00
Chang She	99ba5331f0	feat(python): support new style optional syntax (#793 )	2024-01-09 07:03:29 -08:00
Chang She	121687231c	chore(python): document phrase queries in fts (#788 ) closes #769 Add unit test and documentation on using quotes to perform a phrase query	2024-01-08 21:49:31 -08:00
Chang She	60b22d84bf	chore(python): handle NaN input in fts ingestion (#763 ) If the input text is None, Tantivy raises an error complaining it cannot add a NoneType. We handle this upstream so None's are not added to the document. If all of the indexed fields are None then we skip this document.	2024-01-04 11:45:12 -08:00
Chang She	7e75e50d3a	chore(python): update embedding API to use openai 1.6.1 (#751 ) API has changed significantly, namely `openai.Embedding.create` no longer exists. https://github.com/openai/openai-python/discussions/742 Update the OpenAI embedding function and put a minimum on the openai sdk version.	2023-12-28 15:05:57 -08:00
Chang She	4b8af261a3	feat: add timezone handling for datetime in pydantic (#578 ) If you add timezone information in the Field annotation for a datetime then that will now be passed to the pyarrow data type. I'm not sure how pyarrow enforces timezones, right now, it silently coerces to the timezone given in the column regardless of whether the input had the matching timezone or not. This is probably not the right behavior. Though we could just make it so the user has to make the pydantic model do the validation instead of doing that at the pyarrow conversion layer.	2023-12-28 11:02:56 -08:00
Chang She	c8728d4ca1	feat(python): add post filtering for full text search (#739 ) Closes #721 fts will return results as a pyarrow table. Pyarrow tables has a `filter` method but it does not take sql filter strings (only pyarrow compute expressions). Instead, we do one of two things to support `tbl.search("keywords").where("foo=5").limit(10).to_arrow()`: Default path: If duckdb is available then use duckdb to execute the sql filter string on the pyarrow table. Backup path: Otherwise, write the pyarrow table to a lance dataset and then do `to_table(filter=<filter>)` Neither is ideal. Default path has two issues: 1. requires installing an extra library (duckdb) 2. duckdb mangles some fields (like fixed size list => list) Backup path incurs a latency penalty (~20ms on ssd) to write the resultset to disk. In the short term, once #676 is addressed, we can write the dataset to "memory://" instead of disk, this makes the post filter evaluate much quicker (ETA next week). In the longer term, we'd like to be able to evaluate the filter string on the pyarrow Table directly, one possibility being that we use Substrait to generate pyarrow compute expressions from sql string. Or if there's enough progress on pyarrow, it could support Substrait expressions directly (no ETA) --------- Co-authored-by: Will Jones <willjones127@gmail.com>	2023-12-27 09:31:04 -08:00
Chang She	8f9ad978f5	feat(python): support list of list fields from pydantic schema (#747 ) For object detection, each row may correspond to an image and each image can have multiple bounding boxes of x-y coordinates. This means that a `bbox` field is potentially "list of list of float". This adds support in our pydantic-pyarrow conversion for nested lists.	2023-12-27 09:10:09 -08:00
Weston Pace	dc5126d8d1	feat: add the ability to create scalar indices (#679 ) This is a pretty direct binding to the underlying lance capability	2023-12-21 09:50:10 -08:00
Chang She	7bbb2872de	bug(python): fix path handling in windows (#724 ) Use pathlib for local paths so that pathlib can handle the correct separator on windows. Closes #703 --------- Co-authored-by: Will Jones <willjones127@gmail.com>	2023-12-20 15:41:36 -08:00
Will Jones	f9dd7a5d8a	fix: prevent duplicate data in FTS index (#728 ) This forces the user to replace the whole FTS directory when re-creating the index, prevent duplicate data from being created. Previously, the whole dataset was re-added to the existing index, duplicating existing rows in the index. This (in combination with lancedb/lance#1707) caused #726, since the duplicate data emitted duplicate indices for `take()` and an upstream issue caused those queries to fail. This solution isn't ideal, since it makes the FTS index temporarily unavailable while the index is built. In the future, we should have multiple FTS index directories, which would allow atomic commits of new indexes (as well as multiple indexes for different columns). Fixes #498. Fixes #726. --------- Co-authored-by: Chang She <759245+changhiskhan@users.noreply.github.com>	2023-12-20 13:07:07 -08:00
Will Jones	1d4943688d	upgrade lance to v0.9.1 (#727 ) This brings in some important bugfixes related to take and aarch64 Linux. See changes at: https://github.com/lancedb/lance/releases/tag/v0.9.1	2023-12-20 13:06:54 -08:00
Chang She	7856a94d2c	feat(python): support nested reference for fts (#723 ) https://github.com/lancedb/lance/issues/1739 Support nested field reference in full text search --------- Co-authored-by: Will Jones <willjones127@gmail.com>	2023-12-20 12:28:53 -08:00
Chang She	371d2f979e	feat(python): add option to flatten output in to_pandas (#722 ) Closes https://github.com/lancedb/lance/issues/1738 We add a `flatten` parameter to the signature of `to_pandas`. By default this is None and does nothing. If set to True or -1, then LanceDB will flatten structs before converting to a pandas dataframe. All nested structs are also flattened. If set to any positive integer, then LanceDB will flatten structs up to the specified level of nesting. --------- Co-authored-by: Weston Pace <weston.pace@gmail.com>	2023-12-20 12:23:07 -08:00
Chang She	bd0034a157	feat: support nested pydantic schema (#707 )	2023-12-14 18:20:45 -08:00
Will Jones	d087e7891d	feat(python): add update query support for Python (#654 ) Closes #69 Will not pass until https://github.com/lancedb/lance/pull/1585 is released	2023-12-14 11:28:32 -08:00
Bert	6eb662de9b	fix: python remote correct open_table error message (#659 )	2023-11-24 19:28:33 -05:00
Rok Mihevc	d8e3e54226	feat(python): expose index cache size (#655 ) This is to enable https://github.com/lancedb/lancedb/issues/641. Should be merged after https://github.com/lancedb/lance/pull/1587 is released.	2023-11-18 14:17:40 -08:00
Ayush Chaurasia	1e8678f11a	Multi-task instructor model with quantization support & weak_lru cache for embedding function models (#612 ) resolves #608	2023-11-09 12:34:18 +05:30
Lei Xu	554e068917	chore: improve create_table API consistency between local and remote SDK (#627 )	2023-11-03 13:15:11 -07:00
Ayush Chaurasia	1589499f89	Exponential standoff retry support for handling rate limited embedding functions (#614 ) Users ingesting data using rate limited apis don't need to manually make the process sleep for counter rate limits resolves #579	2023-11-02 19:20:10 +05:30
Bert	24111d543a	fix!: sort table names (#619 ) https://github.com/lancedb/lance/issues/1385	2023-11-01 10:50:09 -04:00
Weston Pace	b517134309	feat: allow prefiltering with index (#610 ) Support for prefiltering with an index was added in lance version 0.8.7. We can remove the lancedb check that prevents this. Closes #261	2023-10-31 13:11:03 -07:00
Chang She	cd9debc3b7	fix(python): fix multiple embedding functions bug (#597 ) Closes #594 The embedding functions are pydantic models so multiple instances with the same parameters are considered ==, which means that if you have multiple embedding columns it's possible for the embeddings to get overwritten. Instead we use `is` instead of == to avoid this problem. testing: modified unit test to include this case	2023-10-24 13:05:05 -04:00
Ayush Chaurasia	0293bbe142	[Python]Embeddings API refactor (#580 ) Sets things up for this -> https://github.com/lancedb/lancedb/issues/579 - Just separates out the registry/ingestion code from the function implementation code - adds a `get_registry` util - package name "open-clip" -> "open-clip-torch"	2023-10-17 22:32:19 -07:00
Prashanth Rao	86efb11572	Add pyarrow date and timestamp type conversion from pydantic (#576 )	2023-10-16 19:42:24 -07:00
Lei Xu	eff94ecea8	chore: bump lance to 0.8.5 (#561 ) Bump lance to 0.5.8	2023-10-14 12:38:43 -07:00
Ayush Chaurasia	683824f1e9	Add cohere embedding function (#550 )	2023-10-13 16:27:34 +05:30
Will Jones	db7bdefe77	feat: cleanup and compaction (#518 ) #488	2023-10-11 12:49:12 -07:00
Chang She	e1ae2bcbd8	feat: add to_list and to_pandas api's (#556 ) Add `to_list` to return query results as list of python dict (so we're not too pandas-centric). Closes #555 Add `to_pandas` API and add deprecation warning on `to_df`. Closes #545 Co-authored-by: Chang She <chang@lancedb.com>	2023-10-11 12:18:55 -07:00
Ayush Chaurasia	a1377afcaa	feat: telemetry, error tracking, CLI & config manager (#538 ) Co-authored-by: Lance Release <lance-dev@lancedb.com> Co-authored-by: Rob Meng <rob.xu.meng@gmail.com> Co-authored-by: Will Jones <willjones127@gmail.com> Co-authored-by: Chang She <759245+changhiskhan@users.noreply.github.com> Co-authored-by: rmeng <rob@lancedb.com> Co-authored-by: Chang She <chang@lancedb.com> Co-authored-by: Rok Mihevc <rok@mihevc.org>	2023-10-08 23:11:39 +05:30

1 2 3

113 Commits