lancedb

mirror of https://github.com/lancedb/lancedb.git synced 2025-12-27 23:12:58 +00:00

Author	SHA1	Message	Date
V	d999d72c8d	docs: pandas example (#2044 ) Fix example for section ## From pandas DataFrame	2025-01-24 11:37:47 -08:00
Will Jones	bcfc93cc88	fix(python): various fixes for async query builders (#2048 ) This includes several improvements and fixes to the Python Async query builders: 1. The API reference docs show all the methods for each builder 2. The hybrid query builder now has all the same setter methods as the vector search one, so you can now set things like `.distance_type()` on a hybrid query. 3. Re-rankers are now properly hooked up and tested for FTS and vector search. Previously the re-rankers were accidentally bypassed in unit tests, because the builders overrode `.to_arrow()`, but the unit test called `.to_batches()` which was only defined in the base class. Now all builders implement `.to_batches()` and leave `.to_arrow()` to the base class. 4. The `AsyncQueryBase` and `AsyncVectoryQueryBase` setter methods now return `Self`, which provides the appropriate subclass as the type hint return value. Previously, `AsyncQueryBase` had them all hard-coded to `AsyncQuery`, which was unfortunate. (This required bringing in `typing-extensions` for older Python version, but I think it's worth it.)	2025-01-20 16:14:34 -08:00
BubbleCal	214d0debf5	docs: claim LanceDB supports float16/float32/float64 for multivector (#2040 )	2025-01-21 07:04:15 +08:00
Will Jones	f059372137	feat: add `drop_index()` method (#2039 ) Closes #1665	2025-01-20 10:08:51 -08:00
BubbleCal	66cbf6b6c5	feat: support multivector type (#2005 ) Signed-off-by: BubbleCal <bubble-cal@outlook.com>	2025-01-13 14:10:40 -08:00
Keming	ce9506db71	docs(hnsw): fix markdown list style (#2015 )	2025-01-13 08:53:13 -08:00
Josef Gugglberger	55ffc96e56	docs: update storage.md, fix Azure Sync connect example (#2010 ) In the sync code example there was also an `await`. ![image](https://github.com/user-attachments/assets/4e1a1bd9-f2fb-4dbe-a9a6-1384ab63edbb)	2025-01-10 09:01:19 -08:00
BubbleCal	3c0a64be8f	feat: support distance range in queries (#1999 ) this also updates the docs --------- Signed-off-by: BubbleCal <bubble-cal@outlook.com>	2025-01-08 11:03:27 +08:00
Will Jones	0e496ed3b5	docs: contributing guide (#1970 ) * Adds basic contributing guides. * Simplifies Python development with a Makefile.	2025-01-07 15:11:16 -08:00
QianZhu	17c9e9afea	docs: add async examples to doc (#1941 ) - added sync and async tabs for python examples - moved python code to tests/docs --------- Co-authored-by: Will Jones <willjones127@gmail.com>	2025-01-07 15:10:25 -08:00
Wyatt Alt	0b45ef93c0	docs: assorted copyedits (#1998 ) This includes a handful of minor edits I made while reading the docs. In addition to a few spelling fixes, * standardize on "rerank" over "re-rank" in prose * terminate sentences with periods or colons as appropriate * replace some usage of dashes with colons, such as in "Try it yourself - <link>" All changes are surface-level. No changes to semantics or structure. --------- Co-authored-by: Will Jones <willjones127@gmail.com>	2025-01-06 15:04:48 -08:00
ahaapple	164ce397c2	docs: fix full-text search (Native FTS) TypeScript doc error (#1992 ) Fix ``` Cannot find name 'queryType'.ts(2304) any ```	2025-01-03 13:36:10 -05:00
Renato Marroquin	0cb6da6b7e	docs: add new indexes to python docs (#1945 ) closes issue #1855 Co-authored-by: Renato Marroquin <renato.marroquin@oracle.com>	2024-12-28 15:35:10 -08:00
BubbleCal	e70fd4fecc	feat: support IVF_FLAT, binary vectors and hamming distance (#1955 ) binary vectors and hamming distance can work on only IVF_FLAT, so introduce them all in this PR. --------- Signed-off-by: BubbleCal <bubble-cal@outlook.com>	2024-12-24 10:36:20 -08:00
Will Jones	61a714a459	docs: improve optimization docs (#1957 ) * Add `See Also` section to `cleanup_old_files` and `compact_files` so they know it's linked to `optimize`. * Fixes link to `compact_files` arguments * Improves formatting of note.	2024-12-19 10:55:11 -08:00
Will Jones	980aa70e2d	feat(python): async-sync feature parity on Table (#1914 ) ### Changes to sync API * Updated `LanceTable` and `LanceDBConnection` reprs * Add `storage_options`, `data_storage_version`, and `enable_v2_manifest_paths` to sync create table API. * Add `storage_options` to `open_table` in sync API. * Add `list_indices()` and `index_stats()` to sync API * `create_table()` will now create only 1 version when data is passed. Previously it would always create two versions: 1 to create an empty table and 1 to add data to it. ### Changes to async API * Add `embedding_functions` to async `create_table()` API. * Added `head()` to async API ### Refactors * Refactor index parameters into dataclasses so they are easier to use from Python * Moved most tests to use an in-memory DB so we don't need to create so many temp directories Closes #1792 Closes #1932 --------- Co-authored-by: Weston Pace <weston.pace@gmail.com>	2024-12-13 12:56:44 -08:00
QianZhu	c0ee370f83	docs: improve schema evolution api examples (#1929 )	2024-12-12 10:52:06 -08:00
QianZhu	17e4022045	docs: add faq to cloud doc (#1907 ) Co-authored-by: Will Jones <willjones127@gmail.com>	2024-12-12 10:07:03 -08:00
BubbleCal	3324e7d525	feat: support 4bit PQ (#1916 )	2024-12-10 10:36:03 +08:00
Will Jones	db125013fc	docs: better formatting for Node API docs (#1892 ) * Sets `"useCodeBlocks": true` * Adds a post-processing script `nodejs/typedoc_post_process.js` that puts the parameter description on the same line as the parameter name, like it is in our Python docs. This makes the text hierarchy clearer in those sections and also makes the sections shorter.	2024-12-09 17:04:09 -08:00
Bert	239f725b32	feat(python)!: async-sync feature parity on Connections (#1905 ) Closes #1791 Closes #1764 Closes #1897 (Makes this unnecessary) BREAKING CHANGE: when using azure connection string `az://...` the call to connect will fail if the azure storage credentials are not set. this is breaking from the previous behaviour where the call would fail after connect, when user invokes methods on the connection.	2024-12-05 14:54:39 -05:00
Will Jones	79eaa52184	feat: schema evolution APIs in all SDKs (#1851 ) * Support `add_columns`, `alter_columns`, `drop_columns` in Remote SDK and async Python * Add `data_type` parameter to node * Docs updates	2024-12-04 14:47:50 -08:00
Lei Xu	bd82e1f66d	feat(python): add support for Azure OpenAPI SDK (#1906 ) Closes #1699	2024-12-04 13:09:38 -08:00
Will Jones	69d9beebc7	docs: improve style and introduction to Python API docs (#1885 ) I found the signatures difficult to read and the parameter section not very space efficient.	2024-11-26 09:17:35 -08:00
QianZhu	3e9321fc40	docs: improve scalar index and filtering (#1874 ) improved the docs on build a scalar index and pre-/post-filtering --------- Co-authored-by: Weston Pace <weston.pace@gmail.com>	2024-11-25 11:30:57 -08:00
QianZhu	285071e5c8	docs: full-text search doc update (#1861 ) Co-authored-by: BubbleCal <bubble-cal@outlook.com>	2024-11-20 21:07:30 -08:00
QianZhu	114866fbcf	docs: OSS doc improvement (#1859 ) OSS doc improvement - HNSW index parameter explanation and others. --------- Co-authored-by: BubbleCal <bubble-cal@outlook.com>	2024-11-20 17:51:11 -08:00
Frank Liu	5387c0e243	docs: add Voyage models to sidebar (#1858 )	2024-11-20 14:20:14 -08:00
fzowl	f2e3989831	docs: voyageai embedding in the index (#1813 ) The code to support VoyageAI embedding and rerank models was added in the https://github.com/lancedb/lancedb/pull/1799 PR. Some of the documentation changes was also made, here adding the VoyageAI embedding doc link to the index page. These are my first PRs in lancedb and while i checked the documentation/code structure, i might missed something important. Please let me know if any changes required!	2024-11-18 14:34:16 -08:00
Emmanuel Ferdman	83ae52938a	docs: update migration reference (#1837 ) # PR Summary PR fixes the `migration.md` reference in `docs/src/guides/tables.md`. On the way, it also fixes some typos found in that document. Signed-off-by: Emmanuel Ferdman <emmanuelferdman@gmail.com>	2024-11-18 14:33:32 -08:00
BubbleCal	b23d8abcdd	docs: introduce incremental indexing for FTS (#1789 ) don't merge it before https://github.com/lancedb/lancedb/pull/1769 merged --------- Signed-off-by: BubbleCal <bubble-cal@outlook.com>	2024-11-18 20:21:28 +08:00
Will Jones	587c0824af	feat: flexible null handling and insert subschemas in Python (#1827 ) * Test that we can insert subschemas (omit nullable columns) in Python. * More work is needed to support this in Node. See: https://github.com/lancedb/lancedb/issues/1832 * Test that we can insert data with nullable schema but no nulls in non-nullable schema. * Add `"null"` option for `on_bad_vectors` where we fill with null if the vector is bad. * Make null values not considered bad if the field itself is nullable.	2024-11-15 11:33:00 -08:00
Will Jones	0fd8a50bd7	ci(node): run examples in CI (#1796 ) This is done as setup for a PR that will fix the OpenAI dependency issue. * [x] FTS examples * [x] Setup mock openai * [x] Ran `npm audit fix` * [x] sentences embeddings test * [x] Double check formatting of docs examples	2024-11-13 11:10:56 -08:00
Ayush Chaurasia	90e9c52d0a	docs: update hybrid search example to latest langchain (#1824 ) Co-authored-by: qzhu <qian@lancedb.com>	2024-11-12 20:06:25 -08:00
QianZhu	5117aecc38	docs: search param explanation for OSS doc (#1815 ) ![Screenshot 2024-11-09 at 11 09 14 AM](https://github.com/user-attachments/assets/2aeba016-aeff-4658-85c6-8640285ba0c9)	2024-11-11 11:57:17 -08:00
fzowl	cbbc07d0f5	feat: voyageai support (#1799 ) Adding VoyageAI embedding and rerank support	2024-11-09 00:51:20 +05:30
Will Jones	a324f4ad7a	feat(node): enable logging and show full errors (#1775 ) This exposes the `LANCEDB_LOG` environment variable in node, so that users can now turn on logging. In addition, fixes a bug where only the top-level error from Rust was being shown. This PR makes sure the full error chain is included in the error message. In the future, will improve this so the error chain is set on the [cause](https://nodejs.org/api/errors.html#errorcause) property of JS errors https://github.com/lancedb/lancedb/issues/1779 Fixes #1774	2024-10-29 15:13:34 -07:00
Rithik Kumar	d71df4572e	docs: revamp langchain integration page (#1773 ) Before - <img width="1030" alt="Screenshot 2024-10-28 132932" src="https://github.com/user-attachments/assets/63f78bfa-949e-473e-ab22-0c692577fa3e"> After - <img width="1037" alt="Screenshot 2024-10-28 132727" src="https://github.com/user-attachments/assets/85a12f6c-74f0-49ba-9f1a-fe77ad125704">	2024-10-29 22:55:50 +05:30
Rithik Kumar	aa269199ad	docs: fix archived examples links (#1751 )	2024-10-29 22:55:27 +05:30
BubbleCal	32fdcf97db	feat!: upgrade lance to 0.19.1 (#1762 ) BREAKING CHANGE: default tokenizer no longer does stemming or stop-word removal. Users should explicitly turn that option on in the future. - upgrade lance to 0.19.1 - update the FTS docs - update the FTS API Upstream change notes: https://github.com/lancedb/lance/releases/tag/v0.19.1 --------- Signed-off-by: BubbleCal <bubble-cal@outlook.com> Co-authored-by: Will Jones <willjones127@gmail.com>	2024-10-29 09:03:52 -07:00
Will Jones	48f46d4751	docs(node): update `indexStats` signature and regenerate docs (#1742 ) `indexStats` still referenced UUID even though in https://github.com/lancedb/lancedb/pull/1702 we changed it to take name instead.	2024-10-18 10:53:28 -07:00
Dominik Weckmüller	e7b56b7b2a	docs: add permanent link chain icon to headings without impacting SEO (#1746 ) I noted that there are no permanent links in the docs. Adapted the current best solution from https://github.com/squidfunk/mkdocs-material/discussions/3535. It adds a GitHub-like chain icon to the left of each heading (right on mobile) and does not impact SEO unlike the default solution with pilcrow char `¶` that might show up on google search results. <img alt="image" src="https://user-images.githubusercontent.com/182589/153004627-6df3f8e9-c747-4f43-bd62-a8dabaa96c3f.gif">	2024-10-14 11:58:23 -07:00
Olzhas Alexandrov	5ccd0edec2	docs: clarify infrastructure requirements for S3 Express One Zone (#1745 )	2024-10-11 14:06:28 -06:00
Rithik Kumar	6ceaf8b06e	docs: add langchainjs writing assistant (#1719 )	2024-10-03 00:55:00 +05:30
Prashant Dixit	e2ca8daee1	docs: saleforce's sfr rag (#1717 ) This PR adds Salesforce's newly released SFR RAG	2024-10-02 21:15:24 +05:30
Rithik Kumar	7b2cdd2269	docs: revamp Voxel51 v1 (#1714 ) Revamp Voxel51 ![image](https://github.com/user-attachments/assets/7ac34457-74ec-4654-b1d1-556e3d7357f5)	2024-10-01 11:59:03 +05:30
Akash Saravanan	d6b5054778	feat(python): add support for trust_remote_code in hf embeddings (#1712 ) Resovles #1709. Adds `trust_remote_code` as a parameter to the `TransformersEmbeddingFunction` class with a default of False. Updated relevant documentation with the same.	2024-10-01 01:06:28 +05:30
Ayush Chaurasia	86978e7588	feat!: enforce all rerankers always return relevance score & deprecate linear combination fixes (#1687 ) - Enforce all rerankers always return _relevance_score. This was already loosely done in tests before but based on user feedback its better to always have _relevance_score present in all reranked results - Deprecate LinearCombinationReranker in docs. And also fix a case where it would not return _relevance_score if one result set was missing	2024-09-23 12:12:02 +05:30
Rithik Kumar	11072b9edc	docs: phidata integration page (#1678 ) Added new integration page for phidata : ![image](https://github.com/user-attachments/assets/8cd9b420-f249-4eac-ac13-ae53983822be)	2024-09-21 00:40:47 +05:30
Rithik Kumar	dcd5f51036	docs: add understand embeddings v1 (#1643 ) Before getting started with managing embeddings. Let's understand embeddings (LanceDB way) ![Screenshot 2024-09-14 012144](https://github.com/user-attachments/assets/7c5435dc-5316-47e9-8d7d-9994ab13b93d)	2024-09-14 02:07:00 +05:30

1 2 3 4 5 ...

384 Commits