lancedb

mirror of https://github.com/lancedb/lancedb.git synced 2025-12-26 14:49:57 +00:00

Author	SHA1	Message	Date
Johannes Kolbe	1f4ac71fa3	apply fixes for notebook (#989 )	2024-02-19 15:36:52 +05:30
Ayush Chaurasia	b5aad2d856	docs: Add meta tag for image preview (#988 ) I think this should work. Need to deploy it to be sure as it can be tested locally. Can be tested here. 2 things about this solution: * All pages have a same meta tag, i.e, lancedb banner * If needed, we can automatically use the first image of each page and generate meta tags using the ultralytics mkdocs plugin that we did for this purpose - https://github.com/ultralytics/mkdocs	2024-02-19 14:07:31 +05:30
Chang She	ca6f55b160	doc: update navigation links for embedding functions (#986 )	2024-02-17 12:12:11 -08:00
Chang She	6f8cf1e068	doc: improve embedding functions documentation (#983 ) Got some user feedback that the `implicit` / `explicit` distinction is confusing. Instead I was thinking we would just deprecate the `with_embeddings` API and then organize working with embeddings into 3 buckets: 1. manually generate embeddings 2. use a provided embedding function 3. define your own custom embedding function	2024-02-17 10:39:28 -08:00
Will Jones	2447372c1f	docs: show DuckDB with dataset, not table (#974 ) Using datasets is preferred way to allow filter and projection pushdown, as well as aggregated larger-than-memory tables.	2024-02-16 09:18:18 -08:00
Ayush Chaurasia	f0298d8372	docs: Minimal reranking evaluation benchmarks (#977 )	2024-02-15 22:16:53 +05:30
Prashanth Rao	78e5fb5451	[docs]: Fix typos and clarity in hybrid search docs (#966 ) - Fixed typos and added some clarity to the hybrid search docs - Changed "Airbnb" case to be as per the [official company name](https://en.wikipedia.org/wiki/Airbnb) (the "bnb" shouldn't be capitalized", and the text in the document aligns with this - Fixed headers in nav bar	2024-02-13 23:25:59 +05:30
Ayush Chaurasia	4fbabdeec3	docs: Add setup cell for colab example (#965 )	2024-02-13 20:42:01 +05:30
Ayush Chaurasia	eb31d95fef	feat(python): hybrid search updates, examples, & latency benchmarks (#964 ) - Rename safe_import -> attempt_import_or_raise (closes https://github.com/lancedb/lancedb/pull/923) - Update docs - Add Notebook example (@changhiskhan you can use it for the talk. Comes with "open in colab" button) - Latency benchmark & results comparison, sanity check on real-world data - Updates the default openai model to gpt-4	2024-02-13 17:58:39 +05:30
Nitish Sharma	f53aace89c	Minor updates to FAQ (#935 ) Based on discussion over discord, adding minor updates to the FAQ section about benchmarks, practical data size and concurrency in LanceDB	2024-02-07 20:49:25 -08:00
Ayush Chaurasia	d982ee934a	feat(python): Reranker DX improvements (#904 ) - Most users might not know how to use `QueryBuilder` object. Instead we should just pass the string query. - Add new rerankers: Colbert, openai	2024-02-06 13:59:31 +05:30
QianZhu	e412194008	fix hybrid search example (#922 )	2024-02-03 09:26:32 +05:30
QianZhu	09cd08222d	make it explicit about the vector column data type (#916 ) <img width="837" alt="Screenshot 2024-02-01 at 4 23 34 PM" src="https://github.com/lancedb/lancedb/assets/1305083/4f0f5c5a-2a24-4b00-aad1-ef80a593d964"> [ <img width="838" alt="Screenshot 2024-02-01 at 4 26 03 PM" src="https://github.com/lancedb/lancedb/assets/1305083/ca073bc8-b518-4be3-811d-8a7184416f07"> ](url) --------- Co-authored-by: Weston Pace <weston.pace@gmail.com>	2024-02-02 09:02:02 -08:00
Weston Pace	d77e95a4f4	feat: upgrade to lance 0.9.11 and expose merge_insert (#906 ) This adds the python bindings requested in #870 The javascript/rust bindings will be added in a future PR.	2024-02-01 11:36:29 -08:00
QianZhu	f5726e2d0c	arrow table/f16 example (#907 )	2024-01-31 14:41:28 -08:00
Will Jones	8d0ea29f89	docs: provide AWS S3 cleanup and permissions advice (#903 ) Adding some more quick advice for how to setup AWS S3 with LanceDB. --------- Co-authored-by: Prashanth Rao <35005448+prrao87@users.noreply.github.com>	2024-01-31 09:24:54 -08:00
Ayush Chaurasia	3ffed89793	feat(python): Hybrid search & Reranker API (#824 ) based on https://github.com/lancedb/lancedb/pull/713 - The Reranker api can be plugged into vector only or fts only search but this PR doesn't do that (see example - https://txt.cohere.com/rerank/) ### Default reranker -- `LinearCombinationReranker(weight=0.7, fill=1.0)` ``` table.search("hello", query_type="hybrid").rerank(normalize="score").to_pandas() ``` ### Available rerankers LinearCombinationReranker ``` from lancedb.rerankers import LinearCombinationReranker # Same as default table.search("hello", query_type="hybrid").rerank( normalize="score", reranker=LinearCombinationReranker() ).to_pandas() # with custom params reranker = LinearCombinationReranker(weight=0.3, fill=1.0) table.search("hello", query_type="hybrid").rerank( normalize="score", reranker=reranker ).to_pandas() ``` Cohere Reranker ``` from lancedb.rerankers import CohereReranker # default model.. English and multi-lingual supported. See docstring for available custom params table.search("hello", query_type="hybrid").rerank( normalize="rank", # score or rank reranker=CohereReranker() ).to_pandas() ``` CrossEncoderReranker ``` from lancedb.rerankers import CrossEncoderReranker table.search("hello", query_type="hybrid").rerank( normalize="rank", reranker=CrossEncoderReranker() ).to_pandas() ``` ## Using custom Reranker ``` from lancedb.reranker import Reranker class CustomReranker(Reranker): def rerank_hybrid(self, vector_result, fts_result): combined_res = self.merge_results(vector_results, fts_results) # or use custom combination logic # Custom rerank logic here return combined_res ``` - [x] Expand testing - [x] Make sure usage makes sense - [x] Run simple benchmarks for correctness (Seeing weird result from cohere reranker in the toy example) - Support diverse rerankers by default: - [x] Cross encoding - [x] Cohere - [x] Reciprocal Rank Fusion --------- Co-authored-by: Chang She <759245+changhiskhan@users.noreply.github.com> Co-authored-by: Prashanth Rao <35005448+prrao87@users.noreply.github.com>	2024-01-30 19:10:33 +05:30
Prashanth Rao	f150768739	Fix image bgcolor (#891 ) Minor fix to change the background color for an image in the docs. It's now readable in both light and dark modes (earlier version made it impossible to read in dark mode).	2024-01-30 16:50:29 +05:30
Ayush Chaurasia	b432ecf2f6	doc: Add documentation chatbot for LanceDB (#890 ) <img width="1258" alt="Screenshot 2024-01-29 at 10 05 52 PM" src="https://github.com/lancedb/lancedb/assets/15766192/7c108fde-e993-415c-ad01-72010fd5fe31">	2024-01-30 11:24:57 +05:30
Raghav Dixit	d1a7257810	feat(python): Embedding fn support for gte-mlx/gte-large (#873 ) have added testing and an example in the docstring, will be pushing a separate PR in recipe repo for rag example --------- Co-authored-by: Ayush Chaurasia <ayush.chaurarsia@gmail.com>	2024-01-30 11:21:57 +05:30
Lei Xu	e5796a4836	doc: fix js example of create index (#886 )	2024-01-28 17:02:36 -08:00
Lei Xu	b9c5323265	doc: use snippet for rust code example and make sure rust examples run through CI (#885 )	2024-01-28 14:30:30 -08:00
Chang She	13acc8a480	doc(rust): minor fixes for Rust quick start. (#878 )	2024-01-28 11:40:52 -08:00
Lei Xu	22b9eceb12	chore: convert all js doc test to use snippet. (#881 )	2024-01-28 11:39:25 -08:00
Lei Xu	5f62302614	doc: use code snippet for typescript examples (#880 ) The typescript code is in a fully function file, that will be run via the CI.	2024-01-27 22:52:37 -08:00
Ayush Chaurasia	d84e0d1db8	feat(python): Aws Bedrock embeddings integration (#822 ) Supports amazon titan, cohere english & cohere multi-lingual base models.	2024-01-28 02:04:15 +05:30
Lei Xu	b49bc113c4	chore: add one rust SDK e2e example (#876 ) Co-authored-by: Chang She <759245+changhiskhan@users.noreply.github.com>	2024-01-26 22:41:20 -08:00
Lei Xu	77b5b1cf0e	doc: update quick start for full rust example (#872 )	2024-01-26 16:19:43 -08:00
Will Jones	4e1ed2b139	docs: document basics of configuring object storage (#832 ) Created based on upstream PR https://github.com/lancedb/lance/pull/1849 Closes #681 --------- Co-authored-by: Prashanth Rao <35005448+prrao87@users.noreply.github.com>	2024-01-24 15:27:22 -08:00
Lei Xu	9a9fc77a95	doc: improve docs for nodejs connect functions (#833 ) * improve the docstring for NodeJS connect functions and `ConnectOptions` parameters. * Simplify `npm run build` steps.	2024-01-19 16:07:53 -08:00
Prashanth Rao	8f54cfcde9	Docs updates incl. Polars (#827 ) This PR makes the following aesthetic and content updates to the docs. - [x] Fix max width issue on mobile: Content should now render more cleanly and be more readable on smaller devices - [x] Improve image quality of flowchart in data management page - [x] Fix syntax highlighting in text at the bottom of the IVF-PQ concepts page - [x] Add example of Polars LazyFrames to docs (Integrations) - [x] Add example of adding data to tables using Polars (guides)	2024-01-18 20:43:59 -08:00
Prashanth Rao	119b928a52	docs: Updates and refactor (#683 ) This PR makes incremental changes to the documentation. * Closes #697 * Closes #698 ## Chores - [x] Add dark mode - [x] Fix headers in navbar - [x] Add `extra.css` to customize navbar styles - [x] Customize fonts for prose/code blocks, navbar and admonitions - [x] Inspect all admonition boxes (remove redundant dropdowns) and improve clarity and readability - [x] Ensure that all images in the docs have white background (not transparent) to be viewable in dark mode - [x] Improve code formatting in code blocks to make them consistent with autoformatters (eslint/ruff) - [x] Add bolder weight to h1 headers - [x] Add diagram showing the difference between embedded (OSS) and serverless (Cloud) - [x] Fix [Creating an empty table](https://lancedb.github.io/lancedb/guides/tables/#creating-empty-table) section: right now, the subheaders are not clickable. - [x] In critical data ingestion methods like `table.add` (among others), the type signature often does not match the actual code - [x] Proof-read each documentation section and rewrite as necessary to provide more context, use cases, and explanations so it reads less like reference documentation. This is especially important for CRUD and search sections since those are so central to the user experience. ## Restructure/new content - [x] The section for [Adding data](https://lancedb.github.io/lancedb/guides/tables/#adding-to-a-table) only shows examples for pandas and iterables. We should include pydantic models, arrow tables, etc. - [x] Add conceptual tutorial for IVF-PQ index - [x] Clearly separate vector search, FTS and filtering sections so that these are easier to find - [x] Add docs on refine factor to explain its importance for recall. Closes #716 - [x] Add an FAQ page showing answers to commonly asked questions about LanceDB. Closes #746 - [x] Add simple polars example to the integrations section. Closes #756 and closes #153 - [ ] Add basic docs for the Rust API (more detailed API docs can come later). Closes #781 - [x] Add a section on the various storage options on local vs. cloud (S3, EBS, EFS, local disk, etc.) and the tradeoffs involved. Closes #782 - [x] Revamp filtering docs: add pre-filtering examples and redo headers and update content for SQL filters. Closes #783 and closes #784. - [x] Add docs for data management: compaction, cleaning up old versions and incremental indexing. Closes #785 - [ ] Add a benchmark section that also discusses some best practices. Closes #787 --------- Co-authored-by: Ayush Chaurasia <ayush.chaurarsia@gmail.com> Co-authored-by: Will Jones <willjones127@gmail.com>	2024-01-19 00:18:37 +05:30
Chang She	af8263af94	feat(python): allow the entire table to be converted a polars dataframe (#814 )	2024-01-15 15:49:16 -08:00
Chang She	be4ab9eef3	feat(python): add exist_ok option to create table (#813 ) This mimics CREATE TABLE IF NOT EXISTS behavior. We add `db.create_table(..., exist_ok=True)` parameter. By default it is set to False, so trying to create a table with the same name will raise an exception. If set to True, then it only opens the table if it already exists. If you pass in a schema, it will be checked against the existing table to make sure you get what you want. If you pass in data, it will NOT be added to the existing table.	2024-01-15 11:09:18 -08:00
Ayush Chaurasia	4568df422d	feat(python): Add gemini text embedding function (#806 ) Named it Gemini-text for now. Not sure how complicated it will be to support both text and multimodal embeddings under the same class "gemini"..But its not something to worry about for now I guess.	2024-01-12 22:38:55 -08:00
Chang She	121687231c	chore(python): document phrase queries in fts (#788 ) closes #769 Add unit test and documentation on using quotes to perform a phrase query	2024-01-08 21:49:31 -08:00
Chang She	b0a88a7286	feat(python): Set heap size to get faster fts indexing performance (#762 ) By default tantivy-py uses 128MB heapsize. We change the default to 1GB and we allow the user to customize this locally this makes `test_fts.py` run 10x faster	2024-01-07 15:15:13 -08:00
sudhir	bf5202f196	Make examples work with current version of Openai api's (#779 ) These examples don't work because of changes in openai api from version 1+	2024-01-07 14:27:56 -08:00
Chris	8be2861061	Minor Fixes to Ingest Embedding Functions Docs (#777 ) Addressed minor typos and grammatical issues to improve readability --------- Co-authored-by: Christopher Correa <chris.correa@gmail.com>	2024-01-07 14:27:40 -08:00
Vladimir Varankin	0560e3a0e5	Minor corrections for docs of embedding_functions (#780 ) In addition to #777, this pull request fixes more typos in the documentation for "Ingest Embedding Functions".	2024-01-07 14:26:35 -08:00
QianZhu	b83fbfc344	small bug fix for example code in SaaS JS doc (#770 )	2024-01-04 14:30:34 -08:00
Bengsoon Chuah	7d55a94efd	Add relevant imports for each step (#764 ) I found that it was quite incoherent to have to read through the documentation and having to search which submodule that each class should be imported from. For example, it is cumbersome to have to navigate to another documentation page to find out that `EmbeddingFunctionRegistry` is from `lancedb.embeddings`	2024-01-04 11:15:42 -08:00
QianZhu	4d8e401d34	SaaS JS API sdk doc (#740 ) Co-authored-by: Aidan <64613310+aidangomar@users.noreply.github.com>	2024-01-03 16:24:21 -08:00
Xin Hao	8411c36b96	docs: fix link (#752 )	2023-12-29 15:33:24 -08:00
Chang She	4b8af261a3	feat: add timezone handling for datetime in pydantic (#578 ) If you add timezone information in the Field annotation for a datetime then that will now be passed to the pyarrow data type. I'm not sure how pyarrow enforces timezones, right now, it silently coerces to the timezone given in the column regardless of whether the input had the matching timezone or not. This is probably not the right behavior. Though we could just make it so the user has to make the pydantic model do the validation instead of doing that at the pyarrow conversion layer.	2023-12-28 11:02:56 -08:00
Chang She	c8728d4ca1	feat(python): add post filtering for full text search (#739 ) Closes #721 fts will return results as a pyarrow table. Pyarrow tables has a `filter` method but it does not take sql filter strings (only pyarrow compute expressions). Instead, we do one of two things to support `tbl.search("keywords").where("foo=5").limit(10).to_arrow()`: Default path: If duckdb is available then use duckdb to execute the sql filter string on the pyarrow table. Backup path: Otherwise, write the pyarrow table to a lance dataset and then do `to_table(filter=<filter>)` Neither is ideal. Default path has two issues: 1. requires installing an extra library (duckdb) 2. duckdb mangles some fields (like fixed size list => list) Backup path incurs a latency penalty (~20ms on ssd) to write the resultset to disk. In the short term, once #676 is addressed, we can write the dataset to "memory://" instead of disk, this makes the post filter evaluate much quicker (ETA next week). In the longer term, we'd like to be able to evaluate the filter string on the pyarrow Table directly, one possibility being that we use Substrait to generate pyarrow compute expressions from sql string. Or if there's enough progress on pyarrow, it could support Substrait expressions directly (no ETA) --------- Co-authored-by: Will Jones <willjones127@gmail.com>	2023-12-27 09:31:04 -08:00
elliottRobinson	eab9072bb5	Update default_embedding_functions.md (#744 ) Modify some grammar, punctuation, and spelling errors.	2023-12-26 19:24:22 +05:30
Will Jones	ee0f0611d9	docs: update node API reference (#734 ) This command hasn't been run for a while...	2023-12-22 10:14:31 -08:00
Will Jones	34966312cb	docs: enhance Update user guide (#735 ) Closes #705	2023-12-22 10:14:21 -08:00
Chang She	0965d7dd5a	doc(javascript): minor improvement on docs for working with tables (#736 ) Closes #639 Closes #638	2023-12-20 20:05:22 -08:00

1 2 3 4

191 Commits