mirror of
https://github.com/lancedb/lancedb.git
synced 2026-08-16 11:08:24 +00:00
python-v0.35.0-beta.2
2678 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3fd322a93a | Bump version: 0.35.0-beta.1 → 0.35.0-beta.2 python-v0.35.0-beta.2 | ||
|
|
d8f0982ee8 |
chore: update lance dependency to v9.0.0-beta.23 (#3665)
Updates the Rust workspace Lance dependencies and Java lance-core from v9.0.0-beta.19 to v9.0.0-beta.23. No compatibility fixes were required; strict workspace Clippy and Rust formatting pass. Lance tag: https://github.com/lance-format/lance/releases/tag/v9.0.0-beta.23 --------- Co-authored-by: Jack Ye <yezhaoqin@gmail.com> |
||
|
|
7276c34c51 |
chore(deps): bump the rust-minor-patch group across 1 directory with 6 updates (#3658)
Bumps the rust-minor-patch group with 6 updates in the / directory: | Package | From | To | | --- | --- | --- | | [regex](https://github.com/rust-lang/regex) | `1.12.4` | `1.13.0` | | [bytes](https://github.com/tokio-rs/bytes) | `1.12.0` | `1.12.1` | | [uuid](https://github.com/uuid-rs/uuid) | `1.23.4` | `1.23.5` | | [http-body](https://github.com/hyperium/http-body) | `1.0.1` | `1.1.0` | | [napi](https://github.com/napi-rs/napi-rs) | `3.10.3` | `3.10.5` | | [napi-derive](https://github.com/napi-rs/napi-rs) | `3.5.9` | `3.5.10` | Updates `regex` from 1.12.4 to 1.13.0 <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/rust-lang/regex/blob/master/CHANGELOG.md">regex's changelog</a>.</em></p> <blockquote> <h1>1.13.0 (2026-07-09)</h1> <p>This release includes a new API, a <code>regex!</code> macro, for lazy compilation of a regex from a string literal. If you use regexes a lot, it's likely you've already written one exactly like it. The new macro can be used like this:</p> <pre lang="rust"><code>use regex::regex; <p>fn is_match(line: &str) -> bool {<br /> // The regex will be compiled approximately once and reused automatically.<br /> // This avoids the footgun of using <code>Regex::new</code> here, which would<br /> // guarantee that it would be compiled every time this routine is called.<br /> // This would likely make this routine much slower than it needs to be.<br /> regex!(r"bar|baz").is_match(line)<br /> }</p> <p>let hay = "<br /> path/to/foo:54:Blue Harvest<br /> path/to/bar:90:Something, Something, Something, Dark Side<br /> path/to/baz:3:It's a Trap!<br /> ";</p> <p>let matches = hay.lines().filter(|line| is_match(line)).count();<br /> assert_eq!(matches, 2);<br /> </code></pre></p> <p>Improvements:</p> <ul> <li><a href="https://redirect.github.com/rust-lang/regex/issues/709">#709</a>: Add a new <code>regex!</code> macro for efficient and automatic reuse of a compiled regex.</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/rust-lang/regex/commit/926af2e68eca3ce089815790541cf50759ba2c59"><code>926af2e</code></a> 1.13.0</li> <li><a href="https://github.com/rust-lang/regex/commit/7d941a93561430cd259bb9ceb84cc66f33ae7be8"><code>7d941a9</code></a> regex-automata-0.4.15</li> <li><a href="https://github.com/rust-lang/regex/commit/e358341229ebd5feb9a78d8cc85b459c3c7b6600"><code>e358341</code></a> api: add <code>regex!</code> macro for lazy compilation</li> <li><a href="https://github.com/rust-lang/regex/commit/c42033379c8760105ef90287f319de73d1572242"><code>c420333</code></a> automata: disable miri on a couple doc tests</li> <li><a href="https://github.com/rust-lang/regex/commit/b9d2cf724f89754ea879b6c223d2292c4d3e2dd3"><code>b9d2cf7</code></a> github: add FUNDING link</li> <li><a href="https://github.com/rust-lang/regex/commit/0858006b1460ba781deda54b8d2b01b3f9f949f7"><code>0858006</code></a> docs: add AI policy for contributors</li> <li><a href="https://github.com/rust-lang/regex/commit/468fc64ecd6493caaca40dbe8319c31c5c08a83d"><code>468fc64</code></a> automata: reject dense DFA start states that are match states</li> <li>See full diff in <a href="https://github.com/rust-lang/regex/compare/1.12.4...1.13.0">compare view</a></li> </ul> </details> <br /> Updates `bytes` from 1.12.0 to 1.12.1 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/tokio-rs/bytes/releases">bytes's releases</a>.</em></p> <blockquote> <h2>Bytes v1.12.1</h2> <h1>1.12.1 (July 8th, 2026)</h1> <h3>Fixed</h3> <ul> <li>Properly handle when <code>Box::new</code> panics (<a href="https://redirect.github.com/tokio-rs/bytes/issues/837">#837</a>)</li> </ul> </blockquote> </details> <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/tokio-rs/bytes/blob/master/CHANGELOG.md">bytes's changelog</a>.</em></p> <blockquote> <h1>1.12.1 (July 8th, 2026)</h1> <h3>Fixed</h3> <ul> <li>Properly handle when <code>Box::new</code> panics (<a href="https://redirect.github.com/tokio-rs/bytes/issues/837">#837</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/tokio-rs/bytes/commit/76c0fbb54ed4336caf9d2311658a2f4a5627c21d"><code>76c0fbb</code></a> Release bytes v1.12.1 (<a href="https://redirect.github.com/tokio-rs/bytes/issues/838">#838</a>)</li> <li><a href="https://github.com/tokio-rs/bytes/commit/924c82bf0053cb13a0fb5165925d564622b2092f"><code>924c82b</code></a> Handle unwinding from Box::new (<a href="https://redirect.github.com/tokio-rs/bytes/issues/837">#837</a>)</li> <li>See full diff in <a href="https://github.com/tokio-rs/bytes/compare/v1.12.0...v1.12.1">compare view</a></li> </ul> </details> <br /> Updates `uuid` from 1.23.4 to 1.23.5 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/uuid-rs/uuid/releases">uuid's releases</a>.</em></p> <blockquote> <h2>v1.23.5</h2> <h2>What's Changed</h2> <ul> <li>doc: Fix broken link by <a href="https://github.com/frostyplanet"><code>@frostyplanet</code></a> in <a href="https://redirect.github.com/uuid-rs/uuid/pull/891">uuid-rs/uuid#891</a></li> <li>perf: Optimize UUID hex parsing and formatting by <a href="https://github.com/geeknoid"><code>@geeknoid</code></a> in <a href="https://redirect.github.com/uuid-rs/uuid/pull/894">uuid-rs/uuid#894</a></li> <li>Prepare for 1.23.5 release by <a href="https://github.com/KodrAus"><code>@KodrAus</code></a> in <a href="https://redirect.github.com/uuid-rs/uuid/pull/895">uuid-rs/uuid#895</a></li> </ul> <h2>New Contributors</h2> <ul> <li><a href="https://github.com/geeknoid"><code>@geeknoid</code></a> made their first contribution in <a href="https://redirect.github.com/uuid-rs/uuid/pull/894">uuid-rs/uuid#894</a></li> </ul> <p><strong>Full Changelog</strong>: <a href="https://github.com/uuid-rs/uuid/compare/v1.23.4...v1.23.5">https://github.com/uuid-rs/uuid/compare/v1.23.4...v1.23.5</a></p> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/uuid-rs/uuid/commit/5dc6b3d1a995e6244a386740588c8d094ca30690"><code>5dc6b3d</code></a> Merge pull request <a href="https://redirect.github.com/uuid-rs/uuid/issues/895">#895</a> from uuid-rs/cargo/v1.23.5</li> <li><a href="https://github.com/uuid-rs/uuid/commit/5a7dfe50e2a2cf41a9d4330e00971e891bcb990f"><code>5a7dfe5</code></a> prepare for 1.23.5 release</li> <li><a href="https://github.com/uuid-rs/uuid/commit/9b4bfc8fe359e24638eccf6c6be424c25ad6ba8c"><code>9b4bfc8</code></a> Merge pull request <a href="https://redirect.github.com/uuid-rs/uuid/issues/894">#894</a> from geeknoid/main</li> <li><a href="https://github.com/uuid-rs/uuid/commit/5acc5a550ef1ccec951f1d2618b33e1171a88b9e"><code>5acc5a5</code></a> perf: Optimize UUID hex parsing and formatting</li> <li><a href="https://github.com/uuid-rs/uuid/commit/1e5d8679542d2bb15412a86839006dc01f680a51"><code>1e5d867</code></a> Merge pull request <a href="https://redirect.github.com/uuid-rs/uuid/issues/891">#891</a> from frostyplanet/doc</li> <li><a href="https://github.com/uuid-rs/uuid/commit/49310f04afd83b7d7667c1e6d7f26f93f46cedda"><code>49310f0</code></a> doc: Fix broken link</li> <li>See full diff in <a href="https://github.com/uuid-rs/uuid/compare/v1.23.4...v1.23.5">compare view</a></li> </ul> </details> <br /> Updates `http-body` from 1.0.1 to 1.1.0 <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/hyperium/http-body/commit/3396328602f7b147ae7b13f022c2b94dff9434e3"><code>3396328</code></a> http-body v1.1.0</li> <li><a href="https://github.com/hyperium/http-body/commit/2fb78de9c875c364b7eb1a1a117acc3b83ffb13a"><code>2fb78de</code></a> chore: bump license year (<a href="https://redirect.github.com/hyperium/http-body/issues/170">#170</a>)</li> <li><a href="https://github.com/hyperium/http-body/commit/b16554b604e598466f6ae5a2689d637230d56d3e"><code>b16554b</code></a> chore(ci): bump checkout to v7</li> <li><a href="https://github.com/hyperium/http-body/commit/c0c53caee7b5192e83cd2bcd273f66419b8acedc"><code>c0c53ca</code></a> chore(ci): use msrv aware update for msrv job</li> <li><a href="https://github.com/hyperium/http-body/commit/5ed15d2c3d10592c82c4bab30c2cda060831bc47"><code>5ed15d2</code></a> tests: fix clippy::double_parens</li> <li><a href="https://github.com/hyperium/http-body/commit/c8cb37f9ce2f8723b25e1ef1a9f6cb63ef1f9c54"><code>c8cb37f</code></a> Derive <code>Copy</code> trait to <code>SizeHint</code> struct (<a href="https://redirect.github.com/hyperium/http-body/issues/164">#164</a>)</li> <li><a href="https://github.com/hyperium/http-body/commit/915d6d5cbb5406b09f1d95978096094a1d35d5bf"><code>915d6d5</code></a> feat(util): add <code>InspectErr</code>, <code>InspectFrame</code> combinators (<a href="https://redirect.github.com/hyperium/http-body/issues/161">#161</a>)</li> <li><a href="https://github.com/hyperium/http-body/commit/0fc0a9415cff00df921c2e8b5b6bbcb9e1a34263"><code>0fc0a94</code></a> docs: fix broken intradoc links (<a href="https://redirect.github.com/hyperium/http-body/issues/162">#162</a>)</li> <li><a href="https://github.com/hyperium/http-body/commit/5a849d49dc8ddba3382cead6d0368264fae5d827"><code>5a849d4</code></a> chore: add FUNDING.yml</li> <li><a href="https://github.com/hyperium/http-body/commit/1a91851246be2ed913d6ace3f5cc18acf0d1d332"><code>1a91851</code></a> feat: impl <code>Add</code> for <code>SizeHint</code>'s (<a href="https://redirect.github.com/hyperium/http-body/issues/156">#156</a>)</li> <li>Additional commits viewable in <a href="https://github.com/hyperium/http-body/compare/v1.0.1...v1.1.0">compare view</a></li> </ul> </details> <br /> Updates `napi` from 3.10.3 to 3.10.5 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/napi-rs/napi-rs/releases">napi's releases</a>.</em></p> <blockquote> <h2>napi-v3.10.5</h2> <h3>Fixed</h3> <ul> <li><em>(napi)</em> release FunctionRef off the JS thread via the custom-GC TSFN (<a href="https://redirect.github.com/napi-rs/napi-rs/pull/3394">#3394</a>)</li> </ul> <h2>napi-v3.10.4</h2> <h3>Fixed</h3> <ul> <li><em>(cli)</em> align build and project configuration (<a href="https://redirect.github.com/napi-rs/napi-rs/pull/3387">#3387</a>)</li> </ul> <h3>Other</h3> <ul> <li><em>(readme)</em> point sponsors image at napi.rs/sponsors.svg (<a href="https://redirect.github.com/napi-rs/napi-rs/pull/3379">#3379</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/napi-rs/napi-rs/commit/970988341eb7f859d2df6da1fb7b12f404a2123e"><code>9709883</code></a> chore(napi): release v3.10.5 (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3395">#3395</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/c931c97a82ad9da42e86c141ce92cbe322930585"><code>c931c97</code></a> fix(napi): release FunctionRef off the JS thread via the custom-GC TSFN (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3394">#3394</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/3812aa748caeb1fdb72d773564827a23307b81d8"><code>3812aa7</code></a> chore: release (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3380">#3380</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/ce5677944b8e66e44396b435dcb154122b2b8732"><code>ce56779</code></a> chore(release): publish</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/b9825c713ff4f871a47c8be897db9859508f4bd5"><code>b9825c7</code></a> fix(derive): defer receiver borrow until argument conversion (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3392">#3392</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/aa49714ed8a5619d65407ceb4ad9e79a1ee5b332"><code>aa49714</code></a> fix(cli): align build and project configuration (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3387">#3387</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/68cbb8d63a73d4c740c4c1c9b61b82c88e13f8b7"><code>68cbb8d</code></a> chore(deps): update yarn to v4.17.1 (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3385">#3385</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/3069f442c30ce3d02e218a29a865ae89d3f50847"><code>3069f44</code></a> fix(sys): fall back to libnode.dll for symbol loading on MSVC targets (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3384">#3384</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/b0157131dc4086debffd321db318eb2c6c905401"><code>b015713</code></a> fix(cli): validate cross-compilation flags upfront and document them accurate...</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/81a35ce09c67765cdfdc06b909318e10d1345193"><code>81a35ce</code></a> chore(deps): update dependency oxc-parser to ^0.139.0 (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3382">#3382</a>)</li> <li>Additional commits viewable in <a href="https://github.com/napi-rs/napi-rs/compare/napi-v3.10.3...napi-v3.10.5">compare view</a></li> </ul> </details> <br /> Updates `napi-derive` from 3.5.9 to 3.5.10 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/napi-rs/napi-rs/releases">napi-derive's releases</a>.</em></p> <blockquote> <h2>napi-derive-v3.5.10</h2> <h3>Other</h3> <ul> <li>updated the following local packages: napi-derive-backend</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/napi-rs/napi-rs/commit/3812aa748caeb1fdb72d773564827a23307b81d8"><code>3812aa7</code></a> chore: release (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3380">#3380</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/ce5677944b8e66e44396b435dcb154122b2b8732"><code>ce56779</code></a> chore(release): publish</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/b9825c713ff4f871a47c8be897db9859508f4bd5"><code>b9825c7</code></a> fix(derive): defer receiver borrow until argument conversion (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3392">#3392</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/aa49714ed8a5619d65407ceb4ad9e79a1ee5b332"><code>aa49714</code></a> fix(cli): align build and project configuration (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3387">#3387</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/68cbb8d63a73d4c740c4c1c9b61b82c88e13f8b7"><code>68cbb8d</code></a> chore(deps): update yarn to v4.17.1 (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3385">#3385</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/3069f442c30ce3d02e218a29a865ae89d3f50847"><code>3069f44</code></a> fix(sys): fall back to libnode.dll for symbol loading on MSVC targets (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3384">#3384</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/b0157131dc4086debffd321db318eb2c6c905401"><code>b015713</code></a> fix(cli): validate cross-compilation flags upfront and document them accurate...</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/81a35ce09c67765cdfdc06b909318e10d1345193"><code>81a35ce</code></a> chore(deps): update dependency oxc-parser to ^0.139.0 (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3382">#3382</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/4bff1272b0c045117c74f541afe9d7b47852181e"><code>4bff127</code></a> docs(readme): point sponsors image at napi.rs/sponsors.svg (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3379">#3379</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/1ac467e06e71f78b983630926c7908894d08e496"><code>1ac467e</code></a> chore(napi): release v3.10.3 (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3376">#3376</a>)</li> <li>Additional commits viewable in <a href="https://github.com/napi-rs/napi-rs/compare/napi-derive-v3.5.9...napi-derive-v3.5.10">compare view</a></li> </ul> </details> <br /> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
1918d1a3b6 |
fix(rust): skip embedding functions for empty batches (#3646)
Fixes #3174 Also fixes #3645 Empty record batches now append correctly typed empty embedding arrays without invoking embedding providers. This avoids OpenAI requests with an invalid empty input while preserving source-column validation and the non-empty execution paths. As a small cleanup, the single- and multi-embedding code paths now share a single upfront lookup of their source columns ("input_columns") instead of each path looking them up independently. Also moves `lance-testing` from regular dependencies to dev-dependencies where it belongs. Tests run: - `cargo fmt --all -- --check` - `cargo test --quiet -p lancedb --lib empty_batch_skips_embedding_functions` - `cargo test --quiet -p lancedb --lib empty_batch_still_validates_source_column` - `cargo test --quiet -p lancedb --lib test_create_empty_table_with_embeddings` - `cargo check --quiet -p lancedb --features remote --tests --examples` - `cargo clippy --quiet -p lancedb --features remote --tests --examples` - `cargo test --quiet -p lancedb --lib` - `cargo test --quiet --features remote --tests` |
||
|
|
3b626efa47 |
fix(python): fill bad vector values element-wise (#3613)
## Summary Fix `on_bad_vectors="fill"` so it replaces only invalid or missing vector values instead of replacing the entire vector row. Fixes #3026. ## Reasoning The old Python sanitizer detected whether a vector row was bad at row granularity. For `fill`, it then used that row-level flag to replace the whole vector with `[fill_value] * dim`. That meant an input like `[1.0, NaN, 3.0]` became `[0.0, 0.0, 0.0]`, even though the documented and more useful behavior is to preserve valid values and fill only the bad element. I checked whether this should be a Rust-side fix so TypeScript users would benefit too. Today, Rust core exposes `NaNVectorBehavior::{Error, Keep}` for rejecting or keeping NaN vectors, while the Python `on_bad_vectors` API (`error`, `drop`, `fill`, `null`) is implemented in the Python ingestion sanitizer before data reaches Rust. TypeScript does not expose the Python `on_bad_vectors="fill"` behavior today. Moving this exact behavior to Rust would be a broader cross-language API change, so this PR keeps the fix scoped to the currently affected Python API. ## What changed - Added a small helper that fills bad vector rows by preserving valid elements, replacing NaN elements with `fill_value`, truncating vectors longer than the expected dimension, and padding short vectors with `fill_value`. - Kept the existing fast path unchanged: the helper only runs after bad vectors are detected and `on_bad_vectors="fill"` is selected. - Updated sanitizer and table tests to assert element-wise NaN replacement and short-vector padding for both `create_table` and `add`. ## Validation - `uv run ruff format .` - `uv run ruff check .` - `cd python && uv run --no-sync pytest python/tests/test_util.py::test_handle_bad_vectors_jagged python/tests/test_util.py::test_handle_bad_vectors_nan python/tests/test_table.py::test_create_with_nans python/tests/test_table.py::test_add_with_nans -vv` Targeted pytest result: `10 passed`. ## Why this fix is Python-side (and not Rust) The problematic behavior lives in Python’s `on_bad_vectors` sanitizer, before data is handed off to Rust. Rust currently only exposes `NaNVectorBehavior::{Error, Keep}` for add operations, while Python has the richer `on_bad_vectors={"error","drop","fill","null"}` API. TypeScript does not currently expose the Python-style fill behavior, so moving this exact fix into Rust would require designing a broader cross-language bad-vector handling API. This PR keeps the change scoped to the existing affected surface: Python’s `on_bad_vectors="fill"` path. This way, Python users immediately benefit. |
||
|
|
137eac9b50 |
docs: add LanceDB agent skill for portable pipelines (#3662)
## What the new agent skill covers We want to help users _easily_ write LanceDB pipelines to bring their data in from other places, no matter whether they use LanceDB OSS or Enterprise. The `lancedb` set of skills contains guidance for agents on the following: - Distinguishes local and remote table capabilities. - Promotes bounded reads using `select()` and `limit()`. - Prevents accidental full-table materialization. - Documents correct Python sync/async scan APIs. - Recommends validated Python schemas and batched ingestion. - Provides indexing, query-tuning, diagnostics, and maintenance guidance. - Documents the Enterprise table-name cache issue: avoid immediately reusing a dropped or overwritten table name; write to a fresh name and rename after propagation. - Adds Python and TypeScript API, pattern, and performance references. - Adds a heuristic scanner for potentially unsafe Python and TypeScript materialization patterns. This change only adds agent documentation and tooling: no LanceDB runtime code, Rust code, SDK APIs, dependencies, or CI configuration are modified. ## Context The LanceDB agent skill was accidentally pushed directly to `main` in `8ea78e3fbcb26718112ab4ddec55a91804b869d3`, bypassing the normal review workflow. That commit was reverted on `main` by `c12a6dce` so the protected branch is back to its prior content. |
||
|
|
06b53c97d6 |
feat: add table FTS query tokenization (#3659)
## Summary - add table-level FTS query tokenization returning token text and position - use the native index tokenizer for local tables and remote index metadata for remote tables - expose sync and async Python table wrappers with focused coverage |
||
|
|
711e05619b |
perf: skip Dataset::index_statistics() for all index types (#3346)
`Dataset::index_statistics()` loads index files and does meaningful CPU work to serialize low-level info. Most fields `NativeTable::index_stats()` needs are available from manifest metadata via `Dataset::describe_indices()`, which is much cheaper. `NativeTable::index_stats()` now: - Calls `describe_indices()` filtered by name; returns `Ok(None)` if no match. - Parses `distance_type` from `description.details()` JSON (the `VectorIndexDetails` proto stored in the manifest by recent Lance versions). - Falls back to `index_statistics()` only for vector indices where `details()` returns no `distance_type` — this handles older Lance datasets that didn't write `VectorIndexDetails`. - `Unknown` index types (e.g. Lance's internal `FragReuseIndex`) are explicitly filtered out of `list_indices` rather than erroring. ## Test plan - [x] `test_create_scalar_index` — asserts `index_type`, `distance_type`, and `num_unindexed_rows > 0` after adding rows post-index - [x] `test_create_fm_index`, `test_create_bitmap_index`, `test_create_label_list_index` — added `index_stats` assertions - [x] IvfPq, IvfHnswPq, IvfHnswSq, IvfHnswFlat tests assert `distance_type == Some(L2)` - [x] `test_list_indices_skip_frag_reuse` — FragReuseIndex is filtered by the Unknown guard in `list_indices` --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
afc0e5f497 | chore: upgrade spin dependency in lock file to avoid yanked version (#3663) | ||
|
|
c12a6dce9f |
Revert "add LanceDB agent skill for portable pipelines"
This reverts commit
|
||
|
|
8ea78e3fbc | add LanceDB agent skill for portable pipelines | ||
|
|
40238d240a |
fix(python): preserve phrase semantics in sync queries (#3654)
## Summary - serialize sync phrase queries consistently for execution and query plans - restore the documented no-argument hybrid `phrase_query()` behavior - keep reranker input as the original user text without mutating the builder Fixes #3653. ## Testing - `python/.venv/bin/python -m pytest <8 focused test nodes> -q` (`8 passed`) - `python/.venv/bin/python -m ruff format --check python/python/lancedb/query.py python/python/tests/test_fts.py python/python/tests/test_hybrid_query.py` - `python/.venv/bin/python -m ruff check .` - `git diff --check origin/main...HEAD` The complete hybrid module and the real native FTS phrase test were not completed in the current PyO3 runtime environment: both stalled in the native `lancedb.connect()` fixture and were interrupted without an assertion failure. |
||
|
|
60428e1a32 |
chore(deps): bump rand from 0.9.4 to 0.10.1 (#3648)
Bumps [rand](https://github.com/rust-random/rand) from 0.9.4 to 0.10.1. <details> <summary>Changelog</summary> <p><em>Sourced from <a href="https://github.com/rust-random/rand/blob/master/CHANGELOG.md">rand's changelog</a>.</em></p> <blockquote> <h2>[0.10.1] — 2026-02-11</h2> <p>This release includes a fix for a soundness bug; see <a href="https://redirect.github.com/rust-random/rand/issues/1763">#1763</a>.</p> <h3>Changes</h3> <ul> <li>Document panic behavior of <code>make_rng</code> and add <code>#[track_caller]</code> (<a href="https://redirect.github.com/rust-random/rand/issues/1761">#1761</a>)</li> <li>Deprecate feature <code>log</code> (<a href="https://redirect.github.com/rust-random/rand/issues/1763">#1763</a>)</li> </ul> <p><a href="https://redirect.github.com/rust-random/rand/issues/1761">#1761</a>: <a href="https://redirect.github.com/rust-random/rand/pull/1761">rust-random/rand#1761</a> <a href="https://redirect.github.com/rust-random/rand/issues/1763">#1763</a>: <a href="https://redirect.github.com/rust-random/rand/pull/1763">rust-random/rand#1763</a></p> <h2>[0.10.0] - 2026-02-08</h2> <h3>Changes</h3> <ul> <li>The dependency on <code>rand_chacha</code> has been replaced with a dependency on <code>chacha20</code>. This changes the implementation behind <code>StdRng</code>, but the output remains the same. There may be some API breakage when using the ChaCha-types directly as these are now the ones in <code>chacha20</code> instead of <code>rand_chacha</code> (<a href="https://redirect.github.com/rust-random/rand/issues/1642">#1642</a>).</li> <li>Rename fns <code>IndexedRandom::choose_multiple</code> -> <code>sample</code>, <code>choose_multiple_array</code> -> <code>sample_array</code>, <code>choose_multiple_weighted</code> -> <code>sample_weighted</code>, struct <code>SliceChooseIter</code> -> <code>IndexedSamples</code> and fns <code>IteratorRandom::choose_multiple</code> -> <code>sample</code>, <code>choose_multiple_fill</code> -> <code>sample_fill</code> (<a href="https://redirect.github.com/rust-random/rand/issues/1632">#1632</a>)</li> <li>Use Edition 2024 and MSRV 1.85 (<a href="https://redirect.github.com/rust-random/rand/issues/1653">#1653</a>)</li> <li>Let <code>Fill</code> be implemented for element types, not sliceable types (<a href="https://redirect.github.com/rust-random/rand/issues/1652">#1652</a>)</li> <li>Fix <code>OsError::raw_os_error</code> on UEFI targets by returning <code>Option<usize></code> (<a href="https://redirect.github.com/rust-random/rand/issues/1665">#1665</a>)</li> <li>Replace fn <code>TryRngCore::read_adapter(..) -> RngReadAdapter</code> with simpler struct <code>RngReader</code> (<a href="https://redirect.github.com/rust-random/rand/issues/1669">#1669</a>)</li> <li>Remove fns <code>SeedableRng::from_os_rng</code>, <code>try_from_os_rng</code> (<a href="https://redirect.github.com/rust-random/rand/issues/1674">#1674</a>)</li> <li>Remove <code>Clone</code> support for <code>StdRng</code>, <code>ReseedingRng</code> (<a href="https://redirect.github.com/rust-random/rand/issues/1677">#1677</a>)</li> <li>Use <code>postcard</code> instead of <code>bincode</code> to test the serde feature (<a href="https://redirect.github.com/rust-random/rand/issues/1693">#1693</a>)</li> <li>Avoid excessive allocation in <code>IteratorRandom::sample</code> when <code>amount</code> is much larger than iterator size (<a href="https://redirect.github.com/rust-random/rand/issues/1695">#1695</a>)</li> <li>Rename <code>os_rng</code> -> <code>sys_rng</code>, <code>OsRng</code> -> <code>SysRng</code>, <code>OsError</code> -> <code>SysError</code> (<a href="https://redirect.github.com/rust-random/rand/issues/1697">#1697</a>)</li> <li>Rename <code>Rng</code> -> <code>RngExt</code> as upstream <code>rand_core</code> has renamed <code>RngCore</code> -> <code>Rng</code> (<a href="https://redirect.github.com/rust-random/rand/issues/1717">#1717</a>)</li> </ul> <h3>Additions</h3> <ul> <li>Add fns <code>IndexedRandom::choose_iter</code>, <code>choose_weighted_iter</code> (<a href="https://redirect.github.com/rust-random/rand/issues/1632">#1632</a>)</li> <li>Pub export <code>Xoshiro128PlusPlus</code>, <code>Xoshiro256PlusPlus</code> prngs (<a href="https://redirect.github.com/rust-random/rand/issues/1649">#1649</a>)</li> <li>Pub export <code>ChaCha8Rng</code>, <code>ChaCha12Rng</code>, <code>ChaCha20Rng</code> behind <code>chacha</code> feature (<a href="https://redirect.github.com/rust-random/rand/issues/1659">#1659</a>)</li> <li>Fn <code>rand::make_rng() -> R where R: SeedableRng</code> (<a href="https://redirect.github.com/rust-random/rand/issues/1734">#1734</a>)</li> </ul> <h3>Removals</h3> <ul> <li>Removed <code>ReseedingRng</code> (<a href="https://redirect.github.com/rust-random/rand/issues/1722">#1722</a>)</li> <li>Removed unused feature "nightly" (<a href="https://redirect.github.com/rust-random/rand/issues/1732">#1732</a>)</li> <li>Removed feature <code>small_rng</code> (<a href="https://redirect.github.com/rust-random/rand/issues/1732">#1732</a>)</li> </ul> <p><a href="https://redirect.github.com/rust-random/rand/issues/1632">#1632</a>: <a href="https://redirect.github.com/rust-random/rand/pull/1632">rust-random/rand#1632</a> <a href="https://redirect.github.com/rust-random/rand/issues/1642">#1642</a>: <a href="https://redirect.github.com/rust-random/rand/pull/1642">rust-random/rand#1642</a> <a href="https://redirect.github.com/rust-random/rand/issues/1649">#1649</a>: <a href="https://redirect.github.com/rust-random/rand/pull/1649">rust-random/rand#1649</a> <a href="https://redirect.github.com/rust-random/rand/issues/1652">#1652</a>: <a href="https://redirect.github.com/rust-random/rand/pull/1652">rust-random/rand#1652</a> <a href="https://redirect.github.com/rust-random/rand/issues/1653">#1653</a>: <a href="https://redirect.github.com/rust-random/rand/pull/1653">rust-random/rand#1653</a> <a href="https://redirect.github.com/rust-random/rand/issues/1659">#1659</a>: <a href="https://redirect.github.com/rust-random/rand/pull/1659">rust-random/rand#1659</a> <a href="https://redirect.github.com/rust-random/rand/issues/1665">#1665</a>: <a href="https://redirect.github.com/rust-random/rand/pull/1665">rust-random/rand#1665</a> <a href="https://redirect.github.com/rust-random/rand/issues/1669">#1669</a>: <a href="https://redirect.github.com/rust-random/rand/pull/1669">rust-random/rand#1669</a> <a href="https://redirect.github.com/rust-random/rand/issues/1674">#1674</a>: <a href="https://redirect.github.com/rust-random/rand/pull/1674">rust-random/rand#1674</a> <a href="https://redirect.github.com/rust-random/rand/issues/1677">#1677</a>: <a href="https://redirect.github.com/rust-random/rand/pull/1677">rust-random/rand#1677</a> <a href="https://redirect.github.com/rust-random/rand/issues/1693">#1693</a>: <a href="https://redirect.github.com/rust-random/rand/pull/1693">rust-random/rand#1693</a> <a href="https://redirect.github.com/rust-random/rand/issues/1695">#1695</a>: <a href="https://redirect.github.com/rust-random/rand/pull/1695">rust-random/rand#1695</a> <a href="https://redirect.github.com/rust-random/rand/issues/1697">#1697</a>: <a href="https://redirect.github.com/rust-random/rand/pull/1697">rust-random/rand#1697</a></p> <!-- raw HTML omitted --> </blockquote> <p>... (truncated)</p> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/rust-random/rand/commit/27ff4cb7ced3122a1f677fc248c1a07e59ddc8cd"><code>27ff4cb</code></a> Prepare v0.10.1: deprecate feature <code>log</code> (<a href="https://redirect.github.com/rust-random/rand/issues/1763">#1763</a>)</li> <li><a href="https://github.com/rust-random/rand/commit/98d06386dc4e1d1c89a91f4e483d571921c29ecf"><code>98d0638</code></a> make_rng: document panic and add #[track_caller] (<a href="https://redirect.github.com/rust-random/rand/issues/1761">#1761</a>)</li> <li><a href="https://github.com/rust-random/rand/commit/54e5eaaa7ac11af3aa60b5ccc486182189e6f9ef"><code>54e5eaa</code></a> Fix doc error (<a href="https://redirect.github.com/rust-random/rand/issues/1758">#1758</a>)</li> <li><a href="https://github.com/rust-random/rand/commit/1ce4c080186730595a8d464591d17aac22a42252"><code>1ce4c08</code></a> Bump itoa from 1.0.17 to 1.0.18 in the all-deps group (<a href="https://redirect.github.com/rust-random/rand/issues/1756">#1756</a>)</li> <li><a href="https://github.com/rust-random/rand/commit/ccb734b9c22891a19f11be125c2f09a43809b08e"><code>ccb734b</code></a> docs: fix typo in doc comment (<a href="https://redirect.github.com/rust-random/rand/issues/1754">#1754</a>)</li> <li><a href="https://github.com/rust-random/rand/commit/357eb7de9c9c80184449e8b515c821e48cf4df74"><code>357eb7d</code></a> Bump libc from 0.2.182 to 0.2.183 in the all-deps group (<a href="https://redirect.github.com/rust-random/rand/issues/1753">#1753</a>)</li> <li><a href="https://github.com/rust-random/rand/commit/5e77fe5d61b886988cae67b6d8fb09e405845c63"><code>5e77fe5</code></a> Fix trait references in documentation (<a href="https://redirect.github.com/rust-random/rand/issues/1752">#1752</a>)</li> <li><a href="https://github.com/rust-random/rand/commit/da891850ab2b38f4322ec140ae29d305dfb162c3"><code>da89185</code></a> Bump the all-deps group with 3 updates (<a href="https://redirect.github.com/rust-random/rand/issues/1751">#1751</a>)</li> <li><a href="https://github.com/rust-random/rand/commit/50516ff45c3675d9c2d247e70bc8db691ed8366d"><code>50516ff</code></a> Bump the all-deps group with 2 updates (<a href="https://redirect.github.com/rust-random/rand/issues/1749">#1749</a>)</li> <li><a href="https://github.com/rust-random/rand/commit/fd71de97fdc7050b9a2d8384f5f8afce7d991ca3"><code>fd71de9</code></a> Bump the all-deps group with 2 updates (<a href="https://redirect.github.com/rust-random/rand/issues/1747">#1747</a>)</li> <li>Additional commits viewable in <a href="https://github.com/rust-random/rand/compare/0.9.4...0.10.1">compare view</a></li> </ul> </details> <br /> [](https://docs.github.com/en/github/managing-security-vulnerabilities/about-dependabot-security-updates#about-compatibility-scores) Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore this major version` will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this minor version` will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself) - `@dependabot ignore this dependency` will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself) </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
5b982f2f05 |
feat(python): added support for WatsonxReranker component (#3642)
## Summary Adds `WatsonxReranker` to the Python bindings, integrating the [IBM watsonx.ai text rerank API](https://cloud.ibm.com/docs/apis/watsonx-ai#text-rerank) via the `ibm_watsonx_ai` SDK (`pip install ibm-watsonx-ai`). ## Parameters | Parameter | Default | Description | |---|---|---| | `model_name` | `"cross-encoder/ms-marco-minilm-l-12-v2"` | Rerank model ID | | `column` | `"text"` | Table column used as document input | | `top_n` | `None` | Return only the top-n results | | `return_score` | `"relevance"` | `"relevance"` or `"all"` | | `api_key` | `None` | Falls back to `WATSONX_API_KEY` env var | | `project_id` | `None` | Falls back to `WATSONX_PROJECT_ID` env var — mutually exclusive with `space_id` | | `space_id` | `None` | Falls back to `WATSONX_SPACE_ID` env var — mutually exclusive with `project_id` | | `url` | `None` | Defaults to `https://us-south.ml.cloud.ibm.com` | | `truncate_input_tokens` | `None` | Token truncation limit | ## Usage ```python from lancedb.rerankers import WatsonxReranker # credentials from environment variables reranker = WatsonxReranker() # or passed explicitly reranker = WatsonxReranker( api_key="<key>", project_id="<project-id>", # or space_id="<space-id>" top_n=5, ) ``` ## Testing Integration test added in `test_rerankers.py`, skipped unless `WATSONX_API_KEY` and one of `WATSONX_PROJECT_ID` / `WATSONX_SPACE_ID` are set. |
||
|
|
cde48fad95 |
ci: remove CODEOWNERS file (#3655)
The CODEOWNERS file added in #3312 automatically requests reviewers on every PR — the `*` default owner routes all changes to two reviewers. This is mostly noise for contributors, and we prefer a single requested reviewer per PR. Remove the file. Reverts #3312. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
1f2068b9fe |
fix(python): gemini batching, user agent and variable dims (#3618)
Carrying over from #2915, this patch introduces: * Single-API call batching support for Gemini embeddings (up to 100 at a time, the API limit) * A versioned user agent header for Gemini API calls * Support for [variable embedding dimension size](https://ai.google.dev/gemini-api/docs/embeddings#control-embedding-size) (Gemini is MRL trained) |
||
|
|
7527890607 |
fix(python): preserve zero distance bounds in hybrid search (#3652)
## Summary - preserve explicit `0.0` distance bounds in synchronous hybrid search - distinguish omitted `None` endpoints from zero-valued endpoints when configuring the vector child query - add a public end-to-end regression test for a zero upper bound ## Testing - `cd python && uv run --extra tests pytest python/tests/test_hybrid_query.py -q` - `uv run --project python ruff format --check python/python/lancedb/query.py python/python/tests/test_hybrid_query.py` - `uv run --project python ruff check .` Fixes #3651 |
||
|
|
a548e59d49 |
feat(python): blob v2 fetch API (#3578)
Python bindings for blob v2 read on **local** tables. Rust read APIs landed in #3562. This PR wires `fetch_blob_files`, `fetch_blobs`, v2 query/`to_pandas(blob_mode="bytes")`, and hidden `_rowid` metadata so `fetch_*` works from query hits without exposing `_rowid` in the column list. **Cloud:** `RemoteTable.fetch_blobs` / `fetch_blob_files` raise `NotImplementedError` until Phalanx ships the server route (separate track; not blocking local merge). ### Primary path: lazy file handles ```python table = db.create_table("videos", schema=pa.schema([ pa.field("id", pa.int64()), lancedb.blob("video"), ])) table.add([{"id": 1, "video": open("clip.mp4", "rb").read()}]) hits = table.search().select(["id", "video"]).to_arrow() handle = table.fetch_blob_files("video", hits)[0] # seek + partial read — PyAV / decoders can use the handle handle.seek(frame_offset) chunk = handle.read_range(0, 65536) ``` `BlobFile` exposes `seek`, `read`, `read_range`, `read_up_to`, and works with `BufferedReader`. ### When you want full bytes ```python blobs = table.fetch_blobs("video", hits) # eager materialize, null-aligned df = table.to_pandas(blob_mode="bytes") # descriptors → bytes in pandas ``` ### `_rowid` (join key, not user `id`) Fetch needs Lance row ids. For v2 blob queries we auto-inject `_rowid`, stash it in Arrow schema metadata on `to_arrow()`, and drop the visible column unless you pass `.with_row_id(True)`. v1 legacy blobs (`lance-encoding:blob`) unchanged; fetch on v1 raises the migration error. ## Test plan - [x] `./scripts/test-blob.sh python` (105 passed in worktree) - [x] `fetch_blob_files` lazy read, seek, partial read, null alignment, cross-fragment dups - [x] hybrid query → `fetch_blobs` / `fetch_blob_files` - [ ] Will re-review after seek/`BlobFile` commit (`d77ab1a6`) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Cursor <cursoragent@cursor.com> |
||
|
|
104fc5a08e | Bump version: 0.32.0-beta.0 → 0.32.0-beta.1 | ||
|
|
715be580d0 | Bump version: 0.35.0-beta.0 → 0.35.0-beta.1 python-v0.35.0-beta.1 | ||
|
|
0d9c87a079 |
ci(nodejs): move Windows builds to larger runner and use ThinLTO (#3634)
The `build - aarch64-pc-windows-msvc` node build job (and, marginally, the x86_64 one) had started hitting `rustc-LLVM ERROR: out of memory` while linking the `lancedb-nodejs` cdylib — most recently surfaced by #3526, which adds the goosefs backend (and its tonic/prost gRPC subtree) to the default node binary. The peak-memory step is the fat-LTO codegen (`lto=fat`, `codegen-units=1` from `.cargo/config.toml`), which merges the whole crate graph into a single LLVM module and runs single-threaded. It therefore neither parallelizes across cores nor fits in the 16 GB of the standard `windows-latest` runner as the dependency graph grows. This PR: - Moves both `*-pc-windows-msvc` node build jobs to `windows-2025-8x-x64` (more memory + cores). - Overrides the release profile to ThinLTO for just these jobs, via `CARGO_PROFILE_RELEASE_LTO=thin` / `CARGO_PROFILE_RELEASE_CODEGEN_UNITS=16` in `pre_build`. ThinLTO parallelizes the cross-module optimization across the runner's cores and keeps peak memory well under the limit. Scoped so Python wheels and Rust release builds keep fat LTO. The larger runner alone would clear the OOM but waste the added cores on the single-threaded fat-LTO tail; ThinLTO is what makes the extra cores actually reduce wall-clock and gives durable memory headroom for future dependency growth. Tradeoff: ThinLTO can leave a small runtime-perf gap vs fat LTO for the node native binary, but it recovers most of it and is a common release configuration. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8e364e6812 | Bump version: 0.31.0-beta.6 → 0.32.0-beta.0 | ||
|
|
32a2776446 | Bump version: 0.34.0-beta.6 → 0.35.0-beta.0 python-v0.35.0-beta.0 | ||
|
|
285add40dd |
feat: expose Lance metrics via OpenTelemetry in Python and Node (#3609)
Bridges Lance's internal `metrics`-crate instrumentation (object store request counts, bytes, latency, errors, and throttles) into OpenTelemetry, in both the Python and Node bindings, with a shared adapter in the Rust core. This is the LanceDB counterpart to lance-format/lance#7537. ## Rust core (`rust/lancedb`) Two new, **off-by-default** features: - `metrics` — re-exports the [`metrics`](https://docs.rs/metrics) crate as `lancedb::metrics` and turns on Lance's object-store instrumentation. Install any `metrics`-compatible recorder to collect them. - `metrics-otel` — adds `lancedb::metrics_otel`, a pull-based adapter that installs a process-global recorder aggregating into lock-free cumulative storage and exposes a snapshot/catalog API (`register_metrics_recorder`, `metrics_catalog`, `snapshot_metrics`, `MetricPoint`/`MetricValue`/`MetricKind`/`MetricDescription`). Both bindings build on this. ## Python `lancedb.otel.instrument_lancedb_metrics()` registers each metric as an OpenTelemetry observable instrument on the given (or global) `MeterProvider`. Available via the `otel` extra (`pip install lancedb[otel]`), which pulls in only `opentelemetry-api` — the application supplies and configures the SDK. ## Node `instrumentLanceDbMetrics()` provides the equivalent wiring against `@opentelemetry/api`. This is the only public entry point; the underlying recorder/catalog/snapshot functions stay internal. Because OpenTelemetry has no asynchronous histogram instrument, histograms are exported Prometheus-style as `<name>_bucket` (with an `le` attribute), `<name>_count`, and `<name>_sum`. Only `_sum` carries the histogram's unit; `_bucket` and `_count` observe cumulative counts and are unitless. The adapter is enabled by default in the Python and Node builds, and off by default in the Rust crate. ## Notes - Requires Lance ≥ `v9.0.0-beta.19`, which ships the object-store metrics APIs (upstream lance-format/lance#7537, now merged). `main` is already on beta.19, so this is a single feature commit with no dependency bump. - Tests: 8 Rust unit tests, 3 Python tests, 2 Node tests, all covering the end-to-end object-store-metrics → OpenTelemetry path. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
22bf091de1 |
fix: avoid manifest writes for read-only directory namespace opens (#3635)
Bumps Lance to v9.0.0-beta.19, which includes lance-format/lance#7687 for side-effect-free DirectoryNamespace read paths. This fixes root-level read-only table opens that previously could trigger `__manifest` creation through directory namespace construction, including Hugging Face bucket reads with read-only tokens. A LanceDB regression test now covers root listing operations without creating `__manifest`. Fixes #3633. |
||
|
|
ff81428a9c |
fix(python): flatten_columns raises when flatten=False (#3629)
### Summary
`flatten_columns` raises `ValueError` when called with `flatten=False`,
even though `False` should mean "do not flatten". This is reachable from
the public API — `Query.to_pandas(flatten=...)` and
`to_batches(flatten=...)` type their `flatten` param as
`Optional[Union[int, bool]]` and pass it straight to `flatten_columns`.
### Cause
`bool` is a subclass of `int`, so `isinstance(False, int)` is `True`.
`flatten=False` skips the `flatten is True` check, falls into the
integer branch, and `False <= 0` evaluates to `True`, raising:
```
ValueError: Please specify a positive integer for flatten or the boolean value `True`
```
### Reproduction
```python
import lancedb
db = lancedb.connect("/tmp/db")
t = db.create_table("t", data=[{"id": 1, "vector": [0.1, 0.2]}])
t.search([0.1, 0.2]).to_pandas(flatten=False) # -> ValueError
```
### Fix
Guard the integer branch with `not isinstance(flatten, bool)` so that
`flatten=False` (and `None`) mean "do not flatten". Behavior is
otherwise unchanged:
- `flatten=True` → flatten all nested levels
- positive `int` → flatten to that depth
- non-positive `int` (e.g. `0`) → still rejected with `ValueError`
Added a regression test in `tests/test_util.py` covering `None`,
`False`, `True`, a positive depth, and `0`.
|
||
|
|
75c5c83f12 |
fix(python): resolve Ollama embedding serialization error in create_table (#3583)
This PR fixes a serialization error when using Ollama embeddings in `create_table`. The use of `@cached_property` for the Ollama client was causing issues during serialization/pickling, which is required by certain LanceDB operations (like when using multiprocessing or certain storage backends). Switching to a standard `@property` ensures the client is instantiated when needed without being stored in a way that breaks serialization. Verified with the following script: ```python import lancedb from lancedb.embeddings import get_registry import pickle registry = get_registry().get(\"ollama\") model = registry(name=\"llama3\") # This would fail before the fix pickled = pickle.dumps(model) unpickled = pickle.loads(pickled) ``` Fixes #2629 (or similar serialization issues reported). --------- Co-authored-by: Unmilan Mukherjee <Missing-Identity@users.noreply.github.com> |
||
|
|
291e9e37be |
feat: add Tencent COS and GooseFS object store support via new feature flags (#3526)
## Summary Closes #3525 This PR wires up two new optional object-store backends at the LanceDB layer, exposing capabilities that already exist upstream in `lance` / `lance-io`: | Backend | Cargo feature | Default in Rust crate | Default in Python wheel | Default in Node binding | | --- | --- | --- | --- | --- | | **Tencent COS** | `cos` | ❌ off | ✅ on | ❌ off | | **GooseFS** | `goosefs` | ❌ off | ✅ on | ✅ on | Both backends are additive and do not affect existing users who don't opt in. ## Motivation - **Tencent COS** is the dominant object storage in the China region. Tencent Cloud users currently need an S3-compatible proxy or a private fork to use LanceDB against COS buckets. - **GooseFS** is Tencent Cloud's distributed cache acceleration layer that sits in front of COS/S3, a common pattern for vector search / AI training where the same hot dataset is read repeatedly. - This brings COS / GooseFS to feature parity with the existing first-class backends (`aws`, `gcs`, `azure`, `oss`, `huggingface`). See the linked issue #3525 for the full discussion. ## Changes ### `rust/lancedb/Cargo.toml` Add two new optional features that pull through the corresponding upstream feature flags: ```toml cos = ["lance/tencent", "lance-io/tencent"] goosefs = [ "lance/goosefs", "lance-io/goosefs", "lance-namespace-impls/dir-goosefs", ] ``` ### `python/Cargo.toml` Enable both `cos` and `goosefs` by default for the Python wheels, so `pip install lancedb` works against COS / GooseFS out of the box (consistent with how `aws` / `gcs` / `azure` / `oss` are bundled today): ```diff -default = ["remote", "lancedb/aws", "lancedb/gcs", "lancedb/azure", "lancedb/dynamodb", "lancedb/oss", "lancedb/huggingface"] +default = ["remote", "lancedb/aws", "lancedb/gcs", "lancedb/azure", "lancedb/dynamodb", "lancedb/oss", "lancedb/huggingface", "lancedb/cos", "lancedb/goosefs"] ``` ### `nodejs/Cargo.toml` Enable `goosefs` by default for the Node binding (COS kept opt-in to limit the default native binary size; can be revisited based on demand): ```diff -default = ["remote", "lancedb/aws", "lancedb/gcs", "lancedb/azure", "lancedb/dynamodb", "lancedb/oss", "lancedb/huggingface"] +default = ["remote", "lancedb/aws", "lancedb/gcs", "lancedb/azure", "lancedb/dynamodb", "lancedb/oss", "lancedb/huggingface", "lancedb/goosefs"] ``` ### `Cargo.lock` Regenerated to reflect the transitive dependencies brought in by the new upstream features. No manual edits. ## Example Usage ### Rust ```toml # Cargo.toml lancedb = { version = "0.30", features = ["cos", "goosefs"] } ``` ```rust // Tencent COS let db = lancedb::connect("cos://my-bucket/my-db").execute().await?; // GooseFS let db = lancedb::connect("goosefs://my-namespace/my-db").execute().await?; ``` ### Python ```python import lancedb db = lancedb.connect( "cos://my-bucket/my-db", storage_options={ "secret_id": "...", "secret_key": "...", "region": "ap-guangzhou", }, ) ``` ## Backwards Compatibility - All new features are **opt-in** at the Rust crate level (`default = []` for `lancedb` itself is unchanged). - The Python wheel gains both backends by default, increasing wheel size slightly but matching the existing pattern of bundling all major cloud backends. - Node binding only adds `goosefs` to defaults; existing users see no behavior change. ## Testing - `cargo check --all-features` ✅ - `cargo check -p lancedb --features cos` ✅ - `cargo check -p lancedb --features goosefs` ✅ - End-to-end COS / GooseFS smoke tests require Tencent Cloud credentials and are intentionally not added to CI in this PR (same approach used for `s3-test`). Happy to add a gated test feature in a follow-up if reviewers prefer. ## Checklist - [x] Added `cos` and `goosefs` features to `rust/lancedb/Cargo.toml` - [x] Updated `python/Cargo.toml` default features - [x] Updated `nodejs/Cargo.toml` default features - [x] Regenerated `Cargo.lock` - [x] Verified build with `--all-features` - [ ] Documentation update (can be done in a follow-up PR once API stabilizes) ## Related - Issue: #3525 - Upstream support: [`lance/tencent`](https://github.com/lance-format/lance), [`lance/goosefs`](https://github.com/lance-format/lance) |
||
|
|
6c066530e5 |
feat: add get_lsm_write_spec to read the installed LSM write spec (#3631)
## Summary Adds `Table::get_lsm_write_spec` returning `Option<LsmWriteSpec>` — the read counterpart to the existing `set_lsm_write_spec` / `unset_lsm_write_spec`. Returns `None` when the MemWAL LSM write path is not enabled; otherwise reconstructs the spec (mode, shard column, `num_buckets`, `maintained_indexes`, `writer_config_defaults`) exactly as installed. ## Changes - **Rust core (`NativeTable`)** — reconstructs the spec from `mem_wal_index_details()`, resolving the shard column from its Lance field id via the dataset schema. This is a raw metadata read, so it is unaffected by `describe_indices` system-index filtering. - **Remote (`RemoteTable`)** — reads the `__lance_mem_wal` system index through `index/list` with `include_system: true` (so the curated `list_indices` surface stays unchanged), then parses the index `details` JSON. It matches the index by name and ignores `index_type`, so no client `IndexType` variant is needed. It uses the **server-resolved `column` name** from the details (Lance field ids do not travel to the remote client). - **Python + TypeScript bindings** — sync and async, mirroring `set`/`unset`, with round-trip tests (bucket / identity / unsharded, plus `None` when unset). ## Tests - Rust: native round-trip unit test + remote mock-endpoint tests (present + absent). All green (`cargo test --features remote -p lancedb`). - Python/TS: round-trip tests added; binding-runtime execution runs in CI. ## Dependencies for the remote path The remote path is complete on the client side but depends on two out-of-repo pieces to work end-to-end: 1. **lance** — emit the server-resolved shard **`column`** name in the MemWAL index `details` JSON (field ids can't reach the client). See lance-format/lance#7667. 2. **server** — honor `include_system` on `index/list` so the `__lance_mem_wal` entry is returned for this read. Against an older server (no `include_system`), the remote getter degrades gracefully to `Ok(None)` rather than erroring. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
f428c6a76c |
chore: update lance dependency to v9.0.0-beta.18 (#3632)
Updates LanceDB's Lance dependencies to v9.0.0-beta.18.\n\nThis refreshes the Rust workspace lockfile and Java lance-core version using the repository update script. Triggering Lance tag: https://github.com/lancedb/lance/releases/tag/v9.0.0-beta.18 |
||
|
|
df89c133ca |
feat(python)!: align Permutation.with_format("torch") with HuggingFace set_format("torch") (#3369)
Closes #3245. > **BREAKING CHANGE:** `with_format("torch")` no longer returns a list of stacked row tensors. It now returns per-row dicts so PyTorch's default `DataLoader` collate stacks them into `{col: tensor(B,)}`. Switch to `with_format("torch_row")` to keep the old shape. ### What changed `"torch"` now returns a list of per-row dicts (`[{col: tensor}, ...]`) at every indexed access path. The default `DataLoader` collate stacks them into a column-keyed batched dict, no custom `collate_fn` needed. The old shape is preserved under a new `"torch_row"` literal. `"torch_col"` is unchanged. The unbatching lives inside the transform (`batch_to_tensor_dict`), not `__getitems__`, so the shape survives pickling and works under `DataLoader(num_workers>0, multiprocessing_context="spawn")`. ### Format comparison | Format | `iter(batch_size=N)` | `__getitems__([0,1,2])` | `DataLoader` default collate | |---|---|---|---| | `"torch"` (new) | `list[{col: tensor}]` length N | `list[{col: tensor}]` length 3 | `{col: tensor(B,)}` | | `"torch_row"` (old `"torch"` behavior) | `list[tensor(n_cols,)]` length N | `list[tensor(n_cols,)]` length 3 | `tensor(B, n_cols)` | | `"torch_col"` (unchanged) | `tensor(n_cols, N)` | `tensor(n_cols, 3)` | needs `collate_fn=lambda x: x` | Output matches HuggingFace `Dataset.set_format("torch")` on container shape, keys, and values at every access path. The only divergence: HuggingFace downcasts `float64` to `torch.float32` by default, LanceDB preserves dtype. Verified by `scripts/verify_torch_format.py`. ### Migration ```python # Old default — column names lost, shape was tensor(B, n_cols) DataLoader(Permutation.identity(table).with_format("torch")) # New default — column names preserved DataLoader(Permutation.identity(table).with_format("torch")) # {col: tensor(B,)} # Keep old behavior DataLoader(Permutation.identity(table).with_format("torch_row")) # tensor(B, n_cols) ``` |
||
|
|
ec763521d4 |
chore: update lance dependency to v9.0.0-beta.17 (#3627)
Updates Lance Rust workspace dependencies and Java lance-core to v9.0.0-beta.17. Includes the required PyO3 compatibility fix for the newer dependency set. Triggering Lance tag: https://github.com/lance-format/lance/releases/tag/v9.0.0-beta.17 --------- Co-authored-by: Jack Ye <yezhaoqin@gmail.com> |
||
|
|
f8dc2f78ee |
ci: add CODEOWNERS file for sensitive paths (#3312)
Fixes #3296 ## Problem The repository has no `CODEOWNERS` file, so there is no enforced review routing for sensitive areas such as release workflows, auth code, and FFI boundaries. This means changes to critical paths can be merged without an explicit codeowner review. ## Solution Add `.github/CODEOWNERS` covering: - `/.github/workflows/` — release/publish workflows (supply chain risk) - `/rust/lancedb/src/remote/` — remote client & auth code - `/python/src/` and `/nodejs/src/` — FFI language boundaries The listed owners (`@jackye1995`, `@wjones127`, `@Xuanwo`, `@AyushExel`) are based on recent merge activity. Feel free to adjust to match the actual team structure or replace with GitHub team handles if preferred. ## Testing No code change — only adds a metadata file. GitHub will start routing review requests automatically once this is merged and branch protection is configured to require codeowner approval. Co-authored-by: octo-patch <octo-patch@github.com> |
||
|
|
3bcff0165e |
feat: support date, datetime, bytes, and Decimal literals in expr builder (#3235)
### **Summary** Closes #3212 Extends the Python `lit()` helper to natively support three additional types (`date`, `datetime`, and `Decimal`) and implements reflexive operators for the `Expr` class. This implementation specifically addresses the blocking feedback regarding precision loss, CI discovery, and query engine limitations: * **Logic Refactoring**: Simplified `lit()` by combining `date` and `datetime` normalization into ISO-8601 strings, ensuring stable SQL parsing across different engine locales. * **Precision Preservation**: `decimal.Decimal` objects are now passed as high-precision strings to the Rust bridge, bypassing intermediate float conversions and preserving full 128-bit decimal precision for DataFusion. * **Averted CI Failures**: Temporarily deferred `bytes` literal support to a future PR to resolve a known DataFusion `expr_to_sql` limitation that was crashing the `Doctest` runner. * **Reflexive Operators**: Added support for "literal-first" arithmetic and logical operations (e.g., `10 + col('a')` or `True & col('active')`). Redundant reflexive comparisons (e.g., `__rlt__`) were pruned as Python's data model handles them automatically. * **Integration Verification**: Added dedicated integration tests in the official test directory to ensure the query engine correctly handles the new types and preserves bit-perfect fidelity. ### **Changes** #### [python/python/lancedb/expr.py](file:///c:/Users/Laksh/Documents/lancedb/python/python/lancedb/expr.py) * Updated `lit()` to handle `date`, `datetime`, and `Decimal` natively. * Implemented reflexive operators (`__radd__`, `__rand__`, `__rmul__`, etc.) to support literals on the left-hand side. * Removed the problematic `bytes` doctest example and `lit()` type support to unblock CI. #### [python/src/expr.rs](file:///c:/Users/Laksh/Documents/lancedb/python/src/expr.rs) * Modified the Rust FFI bridge to extract `Decimal` objects as strings. * Ensured the `expr_lit` handler is ready to receive normalized temporal strings. * Consolidated imports and added missing operator documentation. #### [python/python/lancedb/_lancedb.pyi](file:///c:/Users/Laksh/Documents/lancedb/python/python/lancedb/_lancedb.pyi) * Updated type stubs for `expr_lit` to include `Any` (allowing for `Decimal`). ### **Testing** Added several new advanced test cases in [python/python/tests/test_expr.py](file:///c:/Users/Laksh/Documents/lancedb/python/python/tests/test_expr.py) covering: * **High-precision Decimal preservation**: Verified against 128-bit boundaries with a "one point off" test case (`1.234567890123456789 < 1.234567890123456790`). * **Reflexive operator positioning**: Verified successful query construction with literals on the left. * **Timezone-aware normalization**: Confirmed stable behavior for `datetime` objects. * **Integration Testing**: Confirmed Date32 and Decimal columns return the correct Python types and values from the engine during `.to_arrow()` calls. --------- Co-authored-by: Will Jones <willjones127@gmail.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
c6db80dd0b |
feat: add an elastic dataloader as an iterable dataset (#3509)
# Elastic Streaming Dataloader ## Motivation Training large models on LanceDB tables today requires loading the entire dataset into memory or writing bespoke batching logic. This PR introduces `StreamingDataset`, a PyTorch `IterableDataset` that streams directly from a LanceDB table with two hard guarantees that are difficult to achieve together: **elastic determinism** and **resumability**. ## Goals ### Elastic determinism The dataset partitions the table into a fixed number of *splits* (controlled by `num_splits`, `shuffle_seed`, and `epoch`). Samples are yielded by round-robining over splits one sample per split per cycle. Because the split structure is fixed, the set of samples that makes up each global training step is identical regardless of `world_size` or `num_workers`. You can scale your cluster up or down between runs and the model sees the same data in the same order — no re-sharding, no gradient variance from topology changes. ### Resumability `state_dict()` / `load_state_dict()` capture how many samples each split has consumed. Because all splits are the same size and the round-robin design keeps them in lockstep, the state reduces to a single scalar (`samples_consumed_per_split`) that is topology-independent. A checkpoint saved with 8 GPUs can resume correctly on 4 GPUs or 16 GPUs without any adjustment. ### PyTorch `IterableDataset` / streaming `StreamingDataset` implements the standard PyTorch `IterableDataset` interface, so it drops into any existing `DataLoader` pipeline without modification. Data is fetched lazily from Lance in chunks — only the rows needed for the current batch are ever in memory. Compared to the map dataset this takes more work from pytorch and puts it into the dataset itself (e.g. shuffling, filtering, etc.). We do this because we cannot achieve things like elastic determinism or prefiltering otherwise. ### Multi-worker support DataLoader workers are automatically assigned contiguous sub-blocks of splits (the rank's splits are divided evenly across workers). Each worker is independent: no shared state, no inter-process coordination. The only constraint is that `num_splits` must be divisible by `world_size * num_workers`. That being said, multi-worker is highly discouraged as it relies on multiprocessing which is inefficient. Still, we want to support it. ### Filters as prefilters Filters are applied at *permutation-build time* via `PermutationBuilder.filter()`, not re-evaluated on every fetch. The filtered row IDs are stored in the permutation table so that subsequent reads see only the matching rows. This allows us to avoid loading rows that don't match the filter (which is the default pytorch behavior) ### Prefetching Two parameters control the I/O pipeline: - `read_batch_size` (default 64) — number of rows fetched per `take_offsets` call. Larger values amortise per-request overhead, which is critical on object storage where a single round-trip can cost ~100 ms. - `prefetch_batches` (default 4) — number of batches prefetched in parallel per split via a `ThreadPoolExecutor`. While the model processes the current batch, the next several batches are already in flight, hiding storage latency behind compute. If set correctly then you can get good performance even with num_workers=0 (unless you are bottlenecked on transform). ### Transform parallelism The underlying `Permutation` API supports a `with_transform()` callback for decoding, augmentation, and format conversion. Unfortunately, this is not parallelized. Pytorch typically parallelizes this with num_workers which is multiprocessing which is highly inefficient. For simple transforms we should be able to utilize multithreading and Rust based UDFs. For complex python UDFs we could have a dedicated multiprocessing pipeline for just the transform. Or we could just utilize multithreading. In both cases we would exclude the I/O stage from the multiprocessing because that ends up being very memory hungry and inefficient. --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
f84190fe12 |
chore(deps): bump the rust-minor-patch group with 2 updates (#3621)
Bumps the rust-minor-patch group with 2 updates: [napi](https://github.com/napi-rs/napi-rs) and [napi-derive](https://github.com/napi-rs/napi-rs). Updates `napi` from 3.9.4 to 3.10.3 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/napi-rs/napi-rs/releases">napi's releases</a>.</em></p> <blockquote> <h2>napi-v3.10.3</h2> <h3>Fixed</h3> <ul> <li><em>(napi)</em> preserve the JS error object when cloning an Error off-thread (<a href="https://redirect.github.com/napi-rs/napi-rs/pull/3375">#3375</a>)</li> </ul> <h2>napi-v3.10.2</h2> <h3>Fixed</h3> <ul> <li><em>(napi)</em> keep message and cause when cloning a JS-exception Error off-thread (<a href="https://redirect.github.com/napi-rs/napi-rs/pull/3373">#3373</a>)</li> </ul> <h2>napi-v3.10.1</h2> <h3>Fixed</h3> <ul> <li><em>(napi)</em> release Error's exception reference via the custom GC when dropped off-thread. (<a href="https://redirect.github.com/napi-rs/napi-rs/pull/3370">#3370</a>)</li> <li><em>(napi)</em> stop ref exception object in ThreadsafeFunction sync-throw path on wasm targets (<a href="https://redirect.github.com/napi-rs/napi-rs/pull/3369">#3369</a>)</li> </ul> <h3>Other</h3> <ul> <li><em>(napi)</em> share class accessor trampolines (<a href="https://redirect.github.com/napi-rs/napi-rs/pull/3364">#3364</a>)</li> <li>optimize object field raw property access (<a href="https://redirect.github.com/napi-rs/napi-rs/pull/3365">#3365</a>)</li> </ul> <h2>napi-v3.10.0</h2> <h3>Added</h3> <ul> <li><em>(napi)</em> implement <code>To</code>/<code>FromNapiValue</code> for <code>OsString</code>, <code>OsStr</code>, <code>Path</code> and <code>PathBuf</code> (<a href="https://redirect.github.com/napi-rs/napi-rs/pull/3339">#3339</a>)</li> </ul> <h3>Fixed</h3> <ul> <li><em>(napi)</em> route custom-GC Buffer/TypedArray cross-thread drops through the owning isolate (<a href="https://redirect.github.com/napi-rs/napi-rs/pull/3357">#3357</a>) (<a href="https://redirect.github.com/napi-rs/napi-rs/pull/3360">#3360</a>)</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/napi-rs/napi-rs/commit/1ac467e06e71f78b983630926c7908894d08e496"><code>1ac467e</code></a> chore(napi): release v3.10.3 (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3376">#3376</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/9d672f9f9ac4784364548cac55c15444f4d2b1f8"><code>9d672f9</code></a> fix(napi): preserve the JS error object when cloning an Error off-thread (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3375">#3375</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/35476aebcc774a33b7e79e79d6c476db88a50215"><code>35476ae</code></a> chore(napi): release v3.10.2 (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3374">#3374</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/7844c7343f92f3ef45f2756dbba224c342a8467e"><code>7844c73</code></a> ci: dogfood script-jail <a href="https://github.com/v0"><code>@v0</code></a>.2.10 (lifecycle audit gate + safe install) (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3343">#3343</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/d449ccd8c50ad2268b051458ef919848e72b40a5"><code>d449ccd</code></a> fix(napi): keep message and cause when cloning a JS-exception Error off-threa...</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/2ec02a67a0ffdbe8dcbe93f7f24d1d79b861216b"><code>2ec02a6</code></a> chore(deps): update dependency electron to v43 (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3361">#3361</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/fd0a99f83015d4b67a591641d9ce66edf08d9740"><code>fd0a99f</code></a> chore(deps): update dependency <code>@types/sinon</code> to v22 (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3366">#3366</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/745cd8561f9be2781cc04c8ba4564c8f436792c1"><code>745cd85</code></a> fix: de-flake Windows CI (ava import-from-project EPERM race + cli e2e timeou...</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/2785de583a97e49adea8194090fca2ee12f067c8"><code>2785de5</code></a> chore: release (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3367">#3367</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/441ae7a7b6ddb06a2682a7dd27cf186a8afca9e8"><code>441ae7a</code></a> fix(napi): release Error's exception reference via the custom GC when dropped...</li> <li>Additional commits viewable in <a href="https://github.com/napi-rs/napi-rs/compare/napi-v3.9.4...napi-v3.10.3">compare view</a></li> </ul> </details> <br /> Updates `napi-derive` from 3.5.7 to 3.5.9 <details> <summary>Release notes</summary> <p><em>Sourced from <a href="https://github.com/napi-rs/napi-rs/releases">napi-derive's releases</a>.</em></p> <blockquote> <h2>napi-derive-v3.5.9</h2> <h3>Other</h3> <ul> <li>updated the following local packages: napi-derive-backend</li> </ul> <h2>napi-derive-v3.5.8</h2> <h3>Other</h3> <ul> <li>updated the following local packages: napi-derive-backend</li> </ul> </blockquote> </details> <details> <summary>Commits</summary> <ul> <li><a href="https://github.com/napi-rs/napi-rs/commit/2785de583a97e49adea8194090fca2ee12f067c8"><code>2785de5</code></a> chore: release (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3367">#3367</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/441ae7a7b6ddb06a2682a7dd27cf186a8afca9e8"><code>441ae7a</code></a> fix(napi): release Error's exception reference via the custom GC when dropped...</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/cfa3b77ed50dd3639278b219f5d0f630c596cfac"><code>cfa3b77</code></a> fix(deps): update emnapi to v1.11.2 (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3371">#3371</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/65918a6d195fa007985c83baf97a9ce82a95c2cf"><code>65918a6</code></a> fix(napi): stop ref exception object in ThreadsafeFunction sync-throw path on...</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/324c5502fb4deaabd6d76253e8a8e380c5a2bbb5"><code>324c550</code></a> perf(napi): share class accessor trampolines (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3364">#3364</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/80caf6063deb42468f2742bee02cc43ecb2e111d"><code>80caf60</code></a> perf: optimize object field raw property access (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3365">#3365</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/f72afd58976a83bb0776c6a71171673d94e82226"><code>f72afd5</code></a> chore: release (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3354">#3354</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/4effa4da6247a91048ca3462f2ff8eccdcfabfa4"><code>4effa4d</code></a> chore(deps): lock file maintenance (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3363">#3363</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/f2bf197f629e491362d1911578c57c33be2e561f"><code>f2bf197</code></a> chore(deps): lock file maintenance (<a href="https://redirect.github.com/napi-rs/napi-rs/issues/3362">#3362</a>)</li> <li><a href="https://github.com/napi-rs/napi-rs/commit/962a2f0504517c0f83ff7357100c8b5fc26203af"><code>962a2f0</code></a> fix(napi): route custom-GC Buffer/TypedArray cross-thread drops through the o...</li> <li>Additional commits viewable in <a href="https://github.com/napi-rs/napi-rs/compare/napi-derive-v3.5.7...napi-derive-v3.5.9">compare view</a></li> </ul> </details> <br /> Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting `@dependabot rebase`. [//]: # (dependabot-automerge-start) [//]: # (dependabot-automerge-end) --- <details> <summary>Dependabot commands and options</summary> <br /> You can trigger Dependabot actions by commenting on this PR: - `@dependabot rebase` will rebase this PR - `@dependabot recreate` will recreate this PR, overwriting any edits that have been made to it - `@dependabot show <dependency name> ignore conditions` will show all of the ignore conditions of the specified dependency - `@dependabot ignore <dependency name> major version` will close this group update PR and stop Dependabot creating any more for the specific dependency's major version (unless you unignore this specific dependency's major version or upgrade to it yourself) - `@dependabot ignore <dependency name> minor version` will close this group update PR and stop Dependabot creating any more for the specific dependency's minor version (unless you unignore this specific dependency's minor version or upgrade to it yourself) - `@dependabot ignore <dependency name>` will close this group update PR and stop Dependabot creating any more for the specific dependency (unless you unignore this specific dependency or upgrade to it yourself) - `@dependabot unignore <dependency name>` will remove all of the ignore conditions of the specified dependency - `@dependabot unignore <dependency name> <ignore condition>` will remove the ignore condition of the specified dependency and ignore conditions </details> Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> |
||
|
|
122dcd0f66 |
chore: ignore RUSTSEC-2026-0194 and RUSTSEC-2026-0195 in cargo deny (#3616)
quick-xml < 0.41.0 has two DoS advisories (quadratic attribute-name check and unbounded namespace allocation in NsReader). All three versions in our lockfile (0.26.0, 0.38.4, 0.39.4) are below the patched threshold. These are pulled in transitively by inferno (dev-only flame-graph dep), lance-namespace-impls (git dep from lance), and opendal/reqsign (cloud storage XML parsing). None of these paths expose attacker- controlled XML; clearing them requires upstream to upgrade to quick-xml >= 0.41.0. Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
e6661a7285 |
fix: handle empty/wrong-length vectors returned by embedding functions (#3192)
## Summary - When an embedding function returns an empty list (e.g. `[]`) for an input row — as can happen when a model produces no output for a blank string — `_append_vector_columns` crashed with `ArrowInvalid: Length of item not correct: expected N but got array of size 0` because PyArrow cannot fit a zero-length value into a fixed-size list element. - The fix adds a validation step in `gen()`, inside `_append_vector_columns`, that replaces any vector whose length does not match the expected `ndims` (including empty lists and `None`) with `None` before `pa.array()` is called. - `None` is a valid null in a PyArrow fixed-size list array, so the bad entry flows into `_handle_bad_vectors` and is handled according to the caller-supplied `on_bad_vectors` policy (`error` / `drop` / `fill` / `null`) instead of causing an unconditional crash. ## Test plan - [ ] Added `test_embedding_with_empty_output_vectors` in `python/python/tests/test_embeddings.py` that uses an embedding function returning `[]` for empty-string inputs, calls `table.add(..., on_bad_vectors="drop")`, and asserts no crash and that bad rows are correctly dropped. - [ ] Existing `test_embedding_with_bad_results` continues to pass (NaN vectors still handled correctly). - [ ] Verified manually that `pa.array([[1.,2.,3.,4.], []], type=pa.list_(pa.float32(), 4))` raises `ArrowInvalid` without the fix, and succeeds with `None` in place of `[]`. Fixes #1672 --------- Co-authored-by: Will Jones <willjones127@gmail.com> Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|
|
37466a0390 | Bump version: 0.31.0-beta.5 → 0.31.0-beta.6 | ||
|
|
bfce8a510d | Bump version: 0.34.0-beta.5 → 0.34.0-beta.6 python-v0.34.0-beta.6 | ||
|
|
a1261e6299 |
fix(python): average MRR reciprocal ranks over all rankings (#3599)
## What
`MRRReranker.rerank_multivector` averages each document's reciprocal
ranks over the wrong denominator. It divides by the number of rankings
the document *happens to appear in*, instead of the total number of
rankings being fused.
```python
# python/python/lancedb/rerankers/mrr.py
for result_id, reciprocal_ranks in mrr_score_map.items():
mean_rr = np.mean(reciprocal_ranks) # divides by len(present systems)
```
`mrr_score_map[doc]` only accumulates a reciprocal rank for the systems
in which the document was returned, so `np.mean` never accounts for the
systems that missed it.
## Why it's wrong
Mean Reciprocal Rank fusion treats a system that didn't return a
document as a reciprocal rank of `0` and averages across **all**
systems. That's the exact mechanism by which it rewards cross-system
consensus. Dividing by the appearance count removes that, so a document
liked by a single ranking can beat one ranked highly by every ranking.
Concretely, fusing 3 vector rankings:
| Doc | Ranks | Current score | Correct score |
|-----|-------|---------------|---------------|
| A | #1 in 1 system only | `mean([1.0]) = 1.000` | `1.0 / 3 = 0.333` |
| B | #1, #1, #2 across all 3 | `mean([1, 1, .5]) = 0.833` | `2.5 / 3 =
0.833` |
The current code ranks **A above B** - a document two of three rankings
ignored outranks one all three ranked at or near the top.
This also makes `rerank_multivector` inconsistent with `rerank_hybrid`
in the same file, which already treats a missing system as `0`
(`vector_rr = 0.0` / `fts_rr = 0.0`), and with the class docstring
("average of reciprocal ranks across different search results").
## Fix
Divide the summed reciprocal ranks by the total number of rankings:
```python
num_systems = len(vector_results)
...
mean_rr = float(np.sum(reciprocal_ranks)) / num_systems
```
## Tests
Adds `test_mrr_multivector_rewards_consensus`, which asserts the exact
MRR scores and that the consensus document ranks first. It fails on
`main` and passes with this change. Existing reranker tests are
unaffected.
|
||
|
|
17c499177f |
docs(python): add missing parameter documentation for when_matched_update_all (#3536)
Fixes #2493 Added target. prefix requirement to where parameter docstring. |
||
|
|
d889321b5e |
fix!: combine repeated where filters with AND instead of replacing (#3585)
BREAKING CHANGE: When passing multiple where clauses to a query, they now stack instead of replacing the previous filter. Previously, calling `where`/`only_if` more than once on a query silently replaced the previous filter, so only the last filter was applied. This was surprising and could return rows that an earlier filter should have excluded. This implements the alternative suggested in https://github.com/lancedb/lancedb/pull/3514#issuecomment-4664901580: instead of rejecting a second filter, repeated filters are combined with a logical AND (`(previous) AND (new)`). The combination happens in the Rust core (`QueryBase::only_if` and `only_if_expr`), so it applies to all SDKs at once (Rust, Python async, and TypeScript). The Python sync query builder keeps its own filter state, so it combines filters in the binding layer as well. SQL string and expression filters are combined within their own representation. When the two representations are mixed, the expression is lowered to SQL (via `expr_to_sql_string`) and the filters are combined as SQL strings, so chaining `where` works regardless of which form each filter takes. Fixes #2649 ## Tests - Rust: `cargo test --features remote -p lancedb --lib query` - Python: `uv run --extra tests pytest python/tests/test_query.py` - TypeScript: `pnpm test __test__/query.test.ts` 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
8a37f2ad77 |
feat(rust): re-export arrow and datafusion crates from lancedb (#3576)
lancedb's public API forces downstream crates to construct foreign types
— `RecordBatch`/arrays/builders for `Table::add(...)` (arrow), and
`datafusion_expr::Expr` for `only_if_expr`/`expr_projection`/merge
filters. The required version must exactly match lancedb's internal
arrow/datafusion line, but nothing on the API surface makes that
visible. Drift surfaces only as confusing trait/type errors:
```text
error[E0277]: the trait bound `RecordBatch: Scannable` is not satisfied
= note: there are multiple different versions of crate `arrow_array` in the dependency graph
```
This re-exports the crates lancedb already pins, so consumers can rely
on a single, guaranteed-matching line via a discoverable import path
instead of declaring their own (potentially mismatched) direct
dependency.
- `lancedb::arrow::{arrow, arrow_array, arrow_buffer, arrow_cast,
arrow_data, arrow_ipc, arrow_ord, arrow_schema, arrow_select}` —
previously only `arrow_schema` was re-exported. `arrow-buffer` is
promoted from a transitive to a direct dependency.
- `lancedb::datafusion` — `Expr` is a first-class part of the query and
merge APIs (`only_if_expr`, `expr_projection`,
`QueryFilter::Datafusion`, `when_matched_update_all_expr`), and
`ExecutionPlan` is returned from `create_plan`.
This follows DataFusion's own precedent of re-exporting `arrow`. The
coupling already exists via the trait/impl bounds — this surfaces it
rather than hiding it behind an `E0277`.
Closes #3575
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
f94673ae5e |
ci: update deprecated GitHub Actions to latest versions (Fixes #3577) (#3608)
Fixes #3577 ## Problem GitHub Actions is deprecating Node.js 20 on its runners. Multiple workflows in lancedb use action versions that target Node.js 20 (`actions/checkout@v4`, `actions/setup-node@v4`, `actions/cache@v4`, `actions/upload-artifact@v4`, `actions/download-artifact@v4`, `pnpm/action-setup@v4`). These are being force-run on Node.js 24, generating deprecation warnings. ## Solution Updated all deprecated actions to their latest major versions that support Node.js 24: | Action | Old Version | New Version | |--------|------------|-------------| | `actions/checkout` | @v4 | @v6 | | `actions/setup-node` | @v4 | @v6 | | `actions/cache` | @v4 | @v5 | | `actions/upload-artifact` | @v4 | @v7 | | `actions/download-artifact` | @v4 | @v8 | | `pnpm/action-setup` | @v4 | @v6 | Note: `actions/checkout@v6` and `actions/upload-artifact@v7` are already used in `pypi-publish.yml` — this PR extends the same versions to all remaining workflows. ### Files Changed - `.github/workflows/npm-publish.yml` — Updated checkout, setup-node, cache, upload-artifact, download-artifact, pnpm - `.github/workflows/nodejs.yml` — Updated checkout, setup-node, pnpm - `.github/workflows/python.yml` — Updated checkout - `.github/workflows/rust.yml` — Updated checkout - `.github/workflows/java.yml` — Updated checkout - `.github/workflows/java-publish.yml` — Updated checkout - `.github/workflows/cargo-publish.yml` — Updated checkout - `.github/workflows/docs.yml` — Updated checkout, setup-node - `.github/workflows/dev.yml` — Updated setup-node - `.github/workflows/codex-fix-ci.yml` — Updated checkout, setup-node, pnpm - `.github/workflows/codex-update-lance-dependency.yml` — Updated checkout, setup-node - `.github/workflows/license-header-check.yml` — Updated checkout - `.github/workflows/make-release-commit.yml` — Updated checkout - `.github/workflows/update_package_lock_run.yml` — Updated checkout - `.github/workflows/update_package_lock_run_nodejs.yml` — Updated checkout ## Verification - All 20 YAML files validated with `yaml.safe_load()` — no syntax errors - GitHub Actions CI will validate the actual action versions at runtime ## Changelog | Date | Change | Author | |------|--------|--------| | 2026-07-01 | Updated all deprecated Node 20 actions to latest versions across 15 workflow files | rtmalikian | --- **Disclosure:** This code was developed with assistance from DeepSeek-v4-pro (DeepSeek) via Hermes Agent (Nous Research). All changes were reviewed and verified for correctness. Signed-off-by: rtmalikian <rtmalikian@gmail.com> |
||
|
|
3b70fc4c9d |
fix(python): route async namespace connections through rust (#3603)
Summary: - Route built-in async namespace-backed connections through the Rust namespace connector. - Delegate async namespace/table management methods to the inner AsyncConnection while keeping the custom implementation Python-client fallback. - Add regressions for the native async dir path and lazy namespace_client() construction. Validated locally with targeted namespace/db/table pytest, full test_namespace.py, ruff, cargo fmt/check/clippy, and cargo test -p lancedb-python. |
||
|
|
3a7b02119b | Bump version: 0.31.0-beta.4 → 0.31.0-beta.5 | ||
|
|
bcbc0da090 | Bump version: 0.34.0-beta.4 → 0.34.0-beta.5 python-v0.34.0-beta.5 | ||
|
|
9bead9f53d |
fix(python): route sync namespace connections through rust (#3598)
Summary: - Route built-in sync namespace connections through the Rust namespace connector. - Keep custom namespace clients on the existing Python fallback. - Preserve namespace-backed to_lance compatibility with lazy Python client construction and add regressions. |
||
|
|
0351b77984 |
feat(remote): monotonic reads via x-lancedb-min-read-version watermark (#3597)
## Summary Adds per-session monotonic reads for remote (LanceDB Cloud/Enterprise) tables, preventing successive reads on a handle from moving *backward* in dataset version when a load balancer routes them to query nodes with differently-cached views. Each `RemoteTable` handle tracks the highest dataset version it has observed in a read response — surfaced by the server via a new `x-lancedb-version` response header — and sends it back as `x-lancedb-min-read-version` on subsequent reads (`count_rows`, `query`). A query node whose cache is behind that version refreshes before serving; a node already at/beyond it serves from cache at no extra cost. The watermark is sourced only from reads (always committed dataset versions), so unlike the retired `x-lancedb-min-version` it is unaffected by WAL writes returning WAL entry ids. It is reset on `checkout_latest()`. Both headers are optional and ignored by older peers. Server-side enforcement lives in LanceDB Enterprise. Targets the `codex/update-lance-9-0-0-beta-8` integration branch to match the Enterprise submodule pin. |