Compare commits

...

20 Commits

Author SHA1 Message Date
Gatefixer 1673176f16 fix(python): type structured searches as fts 2026-08-06 03:27:12 +00:00
Gatefixer 5df01b96f4 fix(python): preserve search query builder types 2026-08-05 18:10:46 +00:00
Justin Miller c7ea91f3ea test: cover blob null/empty preservation across Table::optimize (#3774)
## Description

`Table::optimize()` compacts through
`lance::dataset::optimize::compact_files`
(`rust/lancedb/src/table/optimize.rs:155`). Until
lance-format/lance#7965 that rewrite corrupted blob columns holding null
or empty values, which is what #3744 reports:

- **storage 2.0** (legacy v1 `lance-encoding:blob` descriptors): every
payload following a null or empty row in the same fragment was rewritten
as `{position: 0, size: 0}`, so it read back as `b""` and the new
fragment no longer referenced the bytes — silent payload loss,
unrecoverable once the pre-optimize versions are pruned.
- **storage 2.2** (blob v2): a valid empty value was rewritten as null,
destroying the null-vs-empty distinction.

Both manifestations share one root cause: `is_inline_null_blob`
classified any inline blob with `position == 0 && size == 0` as null,
which is also exactly what a *valid empty value* looks like. Such rows
were dropped from `blob_read_addrs`, misaligning every payload that
followed.

The behaviour is already correct on `main`: the vendored lance crate
first carried the fix at `v10.0.0-beta.3` (#3710) and is now
`v10.1.0-beta.1` (#3757). What was missing is coverage — nothing in this
repo exercised a blob column containing a null or empty value through
`optimize()`, which is why this shipped unnoticed. This PR adds that
guard.

## Tests

Two tests in `rust/lancedb/tests/blob_integration.rs`, reusing the
file's existing 64 KiB dedicated-blob helpers and a delete-triggered
fragment rewrite. After `id IN (1, 4)` is deleted the surviving rows are
`2` (null), `3` (valid empty), `5` and `6` (payloads) — payloads sit
immediately after the null/empty, which is where the misalignment
landed.

- `optimize_preserves_v1_blob_payloads_with_null_and_empty` — storage
2.0; asserts the **payload bytes** are unchanged across
`OptimizeAction::All` (what the Python/Node `optimize()` bindings
invoke). Payloads are read through `lance::Dataset::take_blobs`, since
`Table::fetch_blobs` rejects legacy v1 columns. The before/after
descriptors are reported on failure but deliberately *not* asserted:
compaction repacks the blob file, so they shift legitimately (id 5
`(131072, 65536)` → `(0, 65536)`, id 6 `(196608, 65536)` → `(65536,
65536)`). Note that a post-compaction `position: 0` is both the
legitimate first-payload offset and the bug's signature, so asserting
descriptors would be actively misleading.
- `optimize_preserves_blob_v2_null_and_empty_distinction` — storage >=
2.2; asserts a null stays null and a valid empty value stays non-null
empty.

Both assert the pre-optimize state first, so a setup change that stops
producing the null/empty/payload mix fails loudly instead of passing
vacuously.

Both also assert the returned `CompactionMetrics` show a fragment was
actually rewritten. These tests depend on `delete("id IN (1, 4)")`
pushing the fragment past lance's `materialize_deletions_threshold` (0.1
by default; 2 of 6 rows here). That coupling is invisible and unasserted
otherwise: against a forced no-op (`materialize_deletions_threshold:
1.5`) the metrics come back all zeroes and *every payload assertion
still passes*. Since the whole point of these tests is to survive
dependency changes, they check that the rewrite happened rather than
trusting the planner to keep selecting the fragment.

Guard verified against a pre-fix lance: with the published
`lancedb==0.36.0` wheel (vendors lance 9.0.0), `Table.optimize()` on the
same data rewrites the descriptors of the two rows following the
null/empty from `(131072, 65536)` and `(196608, 65536)` to `(0, 0)`, and
the payloads read back empty. Against the pinned `v10.1.0-beta.1`, all
39 tests in the file pass, adding roughly 10–20 ms to the file's
runtime.

## Not addressed here

- **No released artifact has the fix yet.** PyPI `lancedb` 0.36.0
(2026-07-29) vendors lance 9.0.0; npm `@lancedb/lancedb` 0.37.1-beta.0
predates the bump. No 9.x lance tag carries the fix: `v10.0.0-beta.3` is
the first tag containing it, every `v9.1.0-beta.1`…`beta.8` is behind
it, and `v9.0.0` / `v9.0.1-rc.1` sit on a diverged branch without it. A
stable lancedb release needs a stable lance >= 10.
- **The version skew #3744 flagged is still live.**
`python/pyproject.toml` pins `pylance==9.0.0rc1` for the `tests` extra
against a vendored `10.1.0-beta.1`, so Python CI still cannot observe
this class of divergence.
- **Only the single-fragment rewrite shape is covered.** Both tests
rewrite one fragment by materializing deletions. lance's own
`test_compact_blob_v1/v2_preserves_null_empty_and_payload_order` cover
the multi-fragment merge shape (3 fragments → 1) at unit level, so this
PR is complementary rather than redundant — it covers the binding-level
path through `Table::optimize` — but it would not catch a regression
that only appears when *merging* fragments.
`multi_fragment_dedicated_blob_table` in the same file makes that a
cheap follow-up.

Closes #3744

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 12:23:03 -07:00
Wyatt Alt 8e24dd3828 feat(rust)!: make add_columns a builder (#3778)
Table::add_columns now takes no arguments and returns AddColumnsBuilder,
so calls become .add_columns().transform(t).execute().

read_columns was the second positional argument but reaches only one of
the five transform variants. In lance's add_columns_to_fragments only
BatchUDF receives the caller's value: SqlExpressions replaces it with
the columns its expressions reference, Stream and Reader pass None, and
AllNulls reads nothing. So it was mandatory on every call -- all
eighteen call sites here passed None -- and silently discarded four
times out of five. As a builder method it is optional, and setting it
where lance would discard it is now an error, which does reject a call
that previously succeeded while ignoring the argument.

Matches the builders add, update, and merge_insert already use.
2026-08-04 11:18:22 -07:00
Adityaj0 f79dc017c4 fix: when_not_matched_by_source_delete() doesn't reset a previously-set condition (#3771)
## Summary

`LanceMergeInsertBuilder.when_not_matched_by_source_delete()` didn't
clear a previously-set condition when called again with no argument (or
a different condition type). Per the docstring, `condition=None` means
"delete all unmatched rows," but if the builder had already been
configured with a string/Expr condition, a later no-arg call left the
stale condition in place instead of widening the delete to
unconditional.

Fixes #3767

## Change

Each call now unconditionally sets both
`_when_not_matched_by_source_condition` and
`_when_not_matched_by_source_condition_expr` (one to the new value, the
other to `None`), so the latest call always wins — consistent with every
other setter on this builder (e.g.
`when_matched_update_all(where=...)`).

## Test plan

- [x] New regression test
`test_merge_insert_by_source_delete_reconfigure` in
`python/python/tests/test_table.py`
- [x] `uv run --extra tests pytest
python/tests/test_table.py::test_merge_insert_by_source_delete_reconfigure
python/tests/test_table.py::test_merge_insert_by_source_delete_expr
python/tests/test_table.py::test_merge_insert_by_source_delete_expr_async
-vv` — 3 passed
- [x] `uv run --extra dev ruff format` / `ruff check` — clean

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-03 15:49:52 -07:00
Adityaj0 e6ae93f52a fix: hybrid search minimum_nprobes(0) silently no-ops instead of raising (#3770)
## Summary

`LanceHybridQueryBuilder._create_query_builders()` checked
`self._minimum_nprobes` for truthiness instead of `is not None` — the
very next line correctly checks `is not None` for
`self._maximum_nprobes`. Since `0` is falsy in Python,
`.minimum_nprobes(0)` on a hybrid query silently dropped the value
instead of forwarding it to the vector sub-query, where it would raise
the same `ValueError` a plain vector query raises for the same input
(`minimum_nprobes must be greater than 0`, validated in
`rust/lancedb/src/query.rs` and covered for the plain-query path by
`test_invalid_nprobes_sync`).

Fixes #3766

## Change

One-line fix: `if self._minimum_nprobes:` → `if self._minimum_nprobes is
not None:`, matching the existing `maximum_nprobes` check right below
it.

## Test plan

- [x] New regression test
`test_hybrid_query_minimum_nprobes_zero_raises` in
`python/python/tests/test_hybrid_query.py`
- [x] `uv run --extra tests pytest python/tests/test_hybrid_query.py
-vv` — 13 passed
- [x] `uv run --extra dev ruff format` / `ruff check` — clean

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-03 12:54:41 -07:00
Drew Gallardo 3dd9c598e9 feat(remote): add seekable blob range reads (#3750)
## Summary

- Implements Cloud `fetch_blob_files`: returns real seekable `BlobFile`
handles over HTTP Range instead of `NotSupported`.
- Completes the second Cloud blob read verb after #3684 (`fetch_blobs` =
eager whole bytes; this = lazy / partial / sequential reads).
- Same public handle API as local (`read_range`, `read_up_to`, `seek`,
`tell`, `close`), so one code path works for local and Cloud.

Large blobs (video, audio, PDFs) should not require downloading the
whole object to inspect a header or stream a slice. After search,
callers open a handle and read only what they need:

```python
hits = table.search(vec).select(["id", "video"]).limit(5).to_arrow()

with table.fetch_blob_files("video", hits)[0] as f:
    header = f.read_range(0, 256)
    f.seek(keyframe_offset)
    chunk = f.read_up_to(1 << 20)
```

### Behavior

- Handle creation probes size with `bytes=0-0` (bounded concurrency,
input order preserved).
- `204` → null (`None`); `416` with `bytes */0` → valid empty blob;
other `416` → error.
- `read_range` validates `Content-Range` and body length; OOB ranges
fail with `invalid_input` before the request (aligned with Lance).
- `read_up_to` reuses one open-ended Range response across sequential
reads; `seek` drops it.
- Servers older than 0.5.0 get a clear `NotSupported` (does not suggest
`fetch_blobs`, which they also lack).

## Testing

- `cargo test --features remote -p lancedb remote_blob`
- `cargo test --features remote -p lancedb test_blob`
- `cargo clippy --features remote --tests --examples` (no new warnings
from this change)

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-03 08:38:08 -07:00
dependabot[bot] 9e26bf3fba chore(deps): bump the rust-minor-patch group with 3 updates (#3758)
Bumps the rust-minor-patch group with 3 updates:
[http](https://github.com/hyperium/http),
[napi-derive](https://github.com/napi-rs/napi-rs) and
[napi-build](https://github.com/napi-rs/napi-rs).

Updates `http` from 1.4.2 to 1.5.0
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/hyperium/http/releases">http's
releases</a>.</em></p>
<blockquote>
<h2>v1.5.0</h2>
<h2>What's Changed</h2>
<ul>
<li>feat(method): add QUERY method by <a
href="https://github.com/seanmonstar"><code>@​seanmonstar</code></a> in
<a
href="https://redirect.github.com/hyperium/http/pull/798">hyperium/http#798</a></li>
<li>fix(uri): allow empty paths in uri::Builder by <a
href="https://github.com/seanmonstar"><code>@​seanmonstar</code></a> in
<a
href="https://redirect.github.com/hyperium/http/pull/853">hyperium/http#853</a></li>
<li>perf(header,uri): faster value validation, URI parse/format, map
inserts by <a
href="https://github.com/geeknoid"><code>@​geeknoid</code></a> in <a
href="https://redirect.github.com/hyperium/http/pull/852">hyperium/http#852</a></li>
<li>fix(uri): enforce max length in PathAndQuery by <a
href="https://github.com/seanmonstar"><code>@​seanmonstar</code></a> in
<a
href="https://redirect.github.com/hyperium/http/pull/856">hyperium/http#856</a></li>
</ul>
<h2>New Contributors</h2>
<ul>
<li><a href="https://github.com/geeknoid"><code>@​geeknoid</code></a>
made their first contribution in <a
href="https://redirect.github.com/hyperium/http/pull/852">hyperium/http#852</a></li>
</ul>
<p><strong>Full Changelog</strong>: <a
href="https://github.com/hyperium/http/compare/v1.4.2...v1.5.0">https://github.com/hyperium/http/compare/v1.4.2...v1.5.0</a></p>
</blockquote>
</details>
<details>
<summary>Changelog</summary>
<p><em>Sourced from <a
href="https://github.com/hyperium/http/blob/master/CHANGELOG.md">http's
changelog</a>.</em></p>
<blockquote>
<h1>1.5.0 (July 29, 2026)</h1>
<ul>
<li>Add <code>Method::QUERY</code> constant for the new QUERY method
defined in RFC 10008.</li>
<li>Fix <code>uri::Builder::path_and_query()</code> to allow empty
strings to mean no path.</li>
<li>Fix <code>uri::PathAndQuery</code> parsing to enforce URI max
length.</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/hyperium/http/commit/16fc9a7b840c2181e7f8b37397c107b0ffcd050d"><code>16fc9a7</code></a>
v1.5.0</li>
<li><a
href="https://github.com/hyperium/http/commit/e559023f67e3fad6ecc3ee91307be178e0f13626"><code>e559023</code></a>
fix(uri): enforce max length in PathAndQuery (<a
href="https://redirect.github.com/hyperium/http/issues/856">#856</a>)</li>
<li><a
href="https://github.com/hyperium/http/commit/2178e175c4e247a33ba5f6ca3503afb1afbaabba"><code>2178e17</code></a>
perf(header,uri): faster value validation, URI parse/format, map inserts
(<a
href="https://redirect.github.com/hyperium/http/issues/852">#852</a>)</li>
<li><a
href="https://github.com/hyperium/http/commit/03c8cd7faeddfad00873b4d58a45ecdf74ebebe6"><code>03c8cd7</code></a>
fix(uri): allow empty paths in uri::Builder (<a
href="https://redirect.github.com/hyperium/http/issues/853">#853</a>)</li>
<li><a
href="https://github.com/hyperium/http/commit/bb8705b25cdb6e29081edf9ade2ea124f6783e18"><code>bb8705b</code></a>
feat(method): add QUERY method (<a
href="https://redirect.github.com/hyperium/http/issues/798">#798</a>)</li>
<li>See full diff in <a
href="https://github.com/hyperium/http/compare/v1.4.2...v1.5.0">compare
view</a></li>
</ul>
</details>
<br />

Updates `napi-derive` from 3.6.0 to 3.6.1
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/napi-rs/napi-rs/releases">napi-derive's
releases</a>.</em></p>
<blockquote>
<h2>napi-derive-v3.6.1</h2>
<h3>Other</h3>
<ul>
<li>updated the following local packages: napi-derive-backend</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/58bd87fa524a837a7c962ab4103e5588557ccd81"><code>58bd87f</code></a>
chore: release (<a
href="https://redirect.github.com/napi-rs/napi-rs/issues/3414">#3414</a>)</li>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/9da87236dbc4fef99f066b7a130f4d0377308d44"><code>9da8723</code></a>
chore(release): publish</li>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/8d22196aa98a1e6e70584561f5446d117d9c802c"><code>8d22196</code></a>
chore(deps): update dependency oxc-parser to ^0.142.0 (<a
href="https://redirect.github.com/napi-rs/napi-rs/issues/3422">#3422</a>)</li>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/abc30fbafc2e3967d499cef970c68b3edfefd850"><code>abc30fb</code></a>
build(deps): bump postcss from 8.5.17 to 8.5.23 (<a
href="https://redirect.github.com/napi-rs/napi-rs/issues/3421">#3421</a>)</li>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/55421392cbaa24d4df69419e4c6d4958fbcb6a12"><code>5542139</code></a>
build(deps): bump fast-xml-parser from 5.9.3 to 5.10.1 (<a
href="https://redirect.github.com/napi-rs/napi-rs/issues/3418">#3418</a>)</li>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/dc4ee8c89cc27ce30e239482199b3b3d786bf8b6"><code>dc4ee8c</code></a>
build(deps): bump fast-uri from 3.1.3 to 3.1.4 (<a
href="https://redirect.github.com/napi-rs/napi-rs/issues/3419">#3419</a>)</li>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/050d985196174b4be830cdb813d09e2705258455"><code>050d985</code></a>
feat(async-runtime): drain-linger surface + lock-free scheduler
internals (<a
href="https://redirect.github.com/napi-rs/napi-rs/issues/3">#3</a>...</li>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/e0b87086eefe0e7efeea6d269e9403c4be4ba9aa"><code>e0b8708</code></a>
chore(deps): update dependency oxc-parser to ^0.141.0 (<a
href="https://redirect.github.com/napi-rs/napi-rs/issues/3417">#3417</a>)</li>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/fc8494010697d078a93a528c3180271f6f187504"><code>fc84940</code></a>
chore(deps): update actions/setup-node action to v7 (<a
href="https://redirect.github.com/napi-rs/napi-rs/issues/3413">#3413</a>)</li>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/ee598db45985ef11e18c7340801c28bb2452b688"><code>ee598db</code></a>
build(deps): bump protobufjs from 7.6.4 to 7.6.5 (<a
href="https://redirect.github.com/napi-rs/napi-rs/issues/3410">#3410</a>)</li>
<li>Additional commits viewable in <a
href="https://github.com/napi-rs/napi-rs/compare/napi-derive-v3.6.0...napi-derive-v3.6.1">compare
view</a></li>
</ul>
</details>
<br />

Updates `napi-build` from 2.3.2 to 2.4.0
<details>
<summary>Release notes</summary>
<p><em>Sourced from <a
href="https://github.com/napi-rs/napi-rs/releases">napi-build's
releases</a>.</em></p>
<blockquote>
<h2>napi-build-v2.4.0</h2>
<h3>Added</h3>
<ul>
<li><em>(cli)</em> support non-threaded WASI targets (<a
href="https://redirect.github.com/napi-rs/napi-rs/pull/3353">#3353</a>)</li>
</ul>
</blockquote>
</details>
<details>
<summary>Commits</summary>
<ul>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/58bd87fa524a837a7c962ab4103e5588557ccd81"><code>58bd87f</code></a>
chore: release (<a
href="https://redirect.github.com/napi-rs/napi-rs/issues/3414">#3414</a>)</li>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/9da87236dbc4fef99f066b7a130f4d0377308d44"><code>9da8723</code></a>
chore(release): publish</li>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/8d22196aa98a1e6e70584561f5446d117d9c802c"><code>8d22196</code></a>
chore(deps): update dependency oxc-parser to ^0.142.0 (<a
href="https://redirect.github.com/napi-rs/napi-rs/issues/3422">#3422</a>)</li>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/abc30fbafc2e3967d499cef970c68b3edfefd850"><code>abc30fb</code></a>
build(deps): bump postcss from 8.5.17 to 8.5.23 (<a
href="https://redirect.github.com/napi-rs/napi-rs/issues/3421">#3421</a>)</li>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/55421392cbaa24d4df69419e4c6d4958fbcb6a12"><code>5542139</code></a>
build(deps): bump fast-xml-parser from 5.9.3 to 5.10.1 (<a
href="https://redirect.github.com/napi-rs/napi-rs/issues/3418">#3418</a>)</li>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/dc4ee8c89cc27ce30e239482199b3b3d786bf8b6"><code>dc4ee8c</code></a>
build(deps): bump fast-uri from 3.1.3 to 3.1.4 (<a
href="https://redirect.github.com/napi-rs/napi-rs/issues/3419">#3419</a>)</li>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/050d985196174b4be830cdb813d09e2705258455"><code>050d985</code></a>
feat(async-runtime): drain-linger surface + lock-free scheduler
internals (<a
href="https://redirect.github.com/napi-rs/napi-rs/issues/3">#3</a>...</li>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/e0b87086eefe0e7efeea6d269e9403c4be4ba9aa"><code>e0b8708</code></a>
chore(deps): update dependency oxc-parser to ^0.141.0 (<a
href="https://redirect.github.com/napi-rs/napi-rs/issues/3417">#3417</a>)</li>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/fc8494010697d078a93a528c3180271f6f187504"><code>fc84940</code></a>
chore(deps): update actions/setup-node action to v7 (<a
href="https://redirect.github.com/napi-rs/napi-rs/issues/3413">#3413</a>)</li>
<li><a
href="https://github.com/napi-rs/napi-rs/commit/ee598db45985ef11e18c7340801c28bb2452b688"><code>ee598db</code></a>
build(deps): bump protobufjs from 7.6.4 to 7.6.5 (<a
href="https://redirect.github.com/napi-rs/napi-rs/issues/3410">#3410</a>)</li>
<li>Additional commits viewable in <a
href="https://github.com/napi-rs/napi-rs/compare/napi-build-v2.3.2...napi-build-v2.4.0">compare
view</a></li>
</ul>
</details>
<br />


Dependabot will resolve any conflicts with this PR as long as you don't
alter it yourself. You can also trigger a rebase manually by commenting
`@dependabot rebase`.

[//]: # (dependabot-automerge-start)
[//]: # (dependabot-automerge-end)

---

<details>
<summary>Dependabot commands and options</summary>
<br />

You can trigger Dependabot actions by commenting on this PR:
- `@dependabot rebase` will rebase this PR
- `@dependabot recreate` will recreate this PR, overwriting any edits
that have been made to it
- `@dependabot show <dependency name> ignore conditions` will show all
of the ignore conditions of the specified dependency
- `@dependabot ignore <dependency name> major version` will close this
group update PR and stop Dependabot creating any more for the specific
dependency's major version (unless you unignore this specific
dependency's major version or upgrade to it yourself)
- `@dependabot ignore <dependency name> minor version` will close this
group update PR and stop Dependabot creating any more for the specific
dependency's minor version (unless you unignore this specific
dependency's minor version or upgrade to it yourself)
- `@dependabot ignore <dependency name>` will close this group update PR
and stop Dependabot creating any more for the specific dependency
(unless you unignore this specific dependency or upgrade to it yourself)
- `@dependabot unignore <dependency name>` will remove all of the ignore
conditions of the specified dependency
- `@dependabot unignore <dependency name> <ignore condition>` will
remove the ignore condition of the specified dependency and ignore
conditions


</details>

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-02 09:53:09 -07:00
Will Jones 93354baf34 chore: upgrade rust toolchain to 1.97.0 (#3643)
Bumps the pinned Rust toolchain from 1.95.0 to the latest stable
(1.97.0).

Rust 1.97's clippy adds `useless_borrows_in_formatting`, which flags a
redundant `&` in `format!`/`debug!` arguments in a few places. This PR
removes those to keep `cargo clippy` clean.

No behavior change; the MSRV (`rust-version = "1.91.0"`) is unchanged.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-31 16:40:58 -07:00
LuQQiu 05602ec7d5 chore: update lance dependency to v10.1.0-beta.1 (#3757)
Updates the Lance Rust workspace dependencies and Java lance-core
version to v10.1.0-beta.1.

Includes a compatibility fix for the Lance file writer API by using the
explicit V2_1 writer creation path for permutation shuffle spill files.

Triggered by
https://github.com/lance-format/lance/releases/tag/v10.1.0-beta.1
2026-07-31 16:16:17 -07:00
Wyatt Alt e3b472c212 feat: connection-level job operations (#3755)
Adds job operations to the connection surface, building on the Job
handle from #3742: job(id), list_jobs, get_job, cancel_job, and
job_history, plus a non-blocking Job.status(). Implemented on the
Database trait (defaulting to NotSupported), the remote backend
(/v1/jobs), and the Python and Node bindings; job_history returns Arrow
batches.

errors() and progress() are not included.

Tested with mocked endpoints in all three languages.

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 12:51:43 -07:00
Wyatt Alt a6418b6cb9 feat: create_index returns a Job handle (#3742)
IndexBuilder::execute now returns a Job with wait and cancel methods.
Local tables build the index synchronously and return an already-done
job. Remote tables read the job id the server returns from create_index
and track it through the /v1/jobs API: wait polls describe until the job
reaches a terminal state and cancel posts a cancellation. Servers that
return no job id yield a done job, so behavior against older servers is
unchanged. The job id is not exposed on the handle.

The Python and TypeScript bindings keep their current signatures and
discard the handle; exposing Job there is left to follow-ups.

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 07:32:28 -07:00
Cohen Karnell dd2b11eda2 fix(python): log when storage_options is ignored in RemoteDBConnection.open_table (#3743)
`RemoteDBConnection.open_table` accepts `storage_options` and never uses
it:

```python
def open_table(
    self,
    name: str,
    *,
    namespace_path: Optional[List[str]] = None,
    storage_options: Optional[Dict[str, str]] = None,
    index_cache_size: Optional[int] = None,
    ...
) -> Table:
    ...
    if index_cache_size is not None:
        logging.info("index_cache_size is ignored in LanceDb Cloud ...")

    table = LOOP.run(self._conn.open_table(name, namespace_path=namespace_path))
```

The value is never passed down and never mentioned. `index_cache_size`
is ignored on Cloud in the
same way, but it says so.

I checked this at runtime on 0.34.0, not just by reading it: swapping
the inner connection for a
recorder, `open_table("t", storage_options={...})` hands the layer below
`['namespace_path']` and
nothing else, no log record is emitted, and the same probe shows
`index_cache_size` producing its
message as expected.

This adds the matching log line, so the two ignored parameters behave
the same way. `ruff check` and
`ruff format --check` are clean on the file.

A note on severity. This is not a security hole and nothing is exposed.
Someone passing credentials
there gets silence instead of an error, and finds out later.

One thing I am unsure about, and it changes the fix. I have assumed
per-table storage options are
meaningless on Cloud, which is what the `index_cache_size` line next to
it implies about managed
storage. If they are supposed to work, then the right change is to pass
them through to
`self._conn.open_table` instead and this patch is the wrong one. Happy
to redo it that way.

I did not check whether `create_table` or the async connection have the
same gap.
2026-07-30 19:32:31 -07:00
Will Jones 5a1015ba72 docs(python): fill gaps in the Python API reference (#3746)
`docs/src/python/python.md` is the whole Python API reference, but it is
maintained by hand and had drifted from the public API. Anything not
listed there simply doesn't get rendered, so a number of public,
documented, tested APIs were invisible to users — most notably branch
management, where `diff` and `merge` live.

I audited every public symbol reachable from `lancedb` and its
subpackages against the `:::` directives on the page. This adds the
missing ones:

- **Branching** — `Branches`, `AsyncBranches` (`list` / `create` /
`checkout` / `delete` / `diff` / `merge`)
- **Tables** — `TableStatistics` (returned by `Table.stats()`; the
fragment-level stats classes were already listed)
- **Full text queries** — `FullTextQuery`, `MatchQuery`, `PhraseQuery`,
`BoostQuery`, `MultiMatchQuery`, `BooleanQuery`, `FullTextOperator`,
`Occur`
- **Querying** — `LanceEmptyQueryBuilder`, `LanceTakeQueryBuilder`,
`AsyncTakeQuery`
- **Indices** — `Fm` (the FM-index for substring search), `IndexConfig`
- **Blobs** — `blob`, `BlobType`, `BlobFile`
- **Namespaces** — `connect_namespace`, `connect_namespace_async`, and
both namespace connection classes
- **Remote config** — `TlsConfig`, `HeaderProvider`, `OAuthConfig`,
`OAuthFlowType`
- **Rerankers** — the `Reranker` base class plus `JinaReranker`,
`RRFReranker`, `MRRReranker`, `AnswerdotaiRerankers`,
`VoyageAIReranker`, `WatsonxReranker` (5 of 12 were listed)
- **Embeddings** — `get_registry`, `register`, and the 14 embedding
functions that were missing (3 of 17 were listed)
- **PyTorch** — `StreamingDataset` and the permutation API it is built
on
- **Misc** — `Session`, `tokenize`, `FtsToken`, `pydantic.Vector`,
`pydantic.MultiVector`, `instrument_lancedb_metrics`, and the two
exception types

It also repairs cross-references in docstrings that no longer resolve:
links into guide pages that have since moved to lancedb.com
(`querying-an-ann-index`, `experimental-full-text-search`),
`lance.dataset` references with no inventory behind them, and the
relative targets `[Table](Table)` and `[PyArrow Table](pyarrow.Table)`.

Deliberately left out: concrete implementation classes reached through
their abstract base (`LanceTable`, `LanceDBConnection`,
`RemoteDBConnection`), query base classes already covered by
`inherited_members: true`, and internal plumbing such as
`FullTextSearchQuery` and `ColumnOrdering`.

## Testing

The docs job only runs on pushes to `main`, so I built the site locally
and compared against a build of `upstream/main`: every added entry
resolves, and no symbol that was rendered before stopped being rendered
when the four packages moved to automodule. `mkdocs build --strict`
exits 0 on this branch, against 61 warnings on `main`.

## Also in this PR

`lancedb.index`, `lancedb.embeddings`, `lancedb.remote` and
`lancedb.rerankers` are now rendered by a single mkdocstrings directive
each, driven by the module's `__all__`, rather than a hand-maintained
list. These four are where most of the drift was, and `__all__` is
harder to forget than a docs page. `lancedb.embeddings` had no
`__all__`; without one mkdocstrings renders no members at all for a
re-export package, so one is added. AGENTS.md gains a section on how the
page is wired up and how to build the docs locally.

Rendering all that code for the first time surfaced ~100 more build
warnings, which would have made #3707 (turning on `mkdocs build
--strict`) harder to land, so the warning backlog is cleared here too.
97 of the 158 warnings were one systematic false positive — griffe
cannot see the generated `__init__` of a pydantic dataclass, so every
documented parameter looks unknown — switched off via
`warn_unknown_params`. The remaining 61 came from 15 docstrings with
real bugs: prose trailing a `Parameters` section (we were rendering
parameters called `The`, `you` and `To`), types dropped because numpydoc
needs spaces around the colon, `num_partitions, default sqrt(num_rows)`
parsing as a list of names and inventing a `default` parameter, and one
parameter indented five spaces. `mkdocs build --strict` now exits 0.

---

#3747 (the coverage test that keeps this from happening again) is
stacked on this branch, so review it after this one.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 16:50:05 -07:00
Farmer.Chillax 48945d0658 feat(python): add namespace/table exist support (#3460)
In the current LanceDB usage implementation, there is no way to check
whether a table or namespace already exists. This PR introduces the
namespace_exists and table_exists methods to determine the existence of
tables and namespaces.

useage like this:
```
# check table exists
db.table_exists(table_id=['xxx'])

# check namespace exists
db.namespace_exists(namespace_id=['xxx'])
```

fixes: #3419

---------

Signed-off-by: farmer <farmerchillax@outlook.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-30 15:20:47 -07:00
Drew Gallardo 77208fd464 feat(remote): add RemoteTable fetch_blobs HTTP client (#3684)
Remote half of the blob read path. #3578 did local Python. This makes
`RemoteTable` hit the server.

- `fetch_blobs(column, row_ids or hits)` → bytes over `POST
/v1/table/{id}/fetch_blobs/`
- `blob_columns()` from the cached schema (describe already has the
metadata, no extra route)
- search then `fetch_blobs` works. row identity rides inside the blob
descriptor so you do not need a public `_rowid`
- `fetch_blob_files` still `NotSupported` on remote. use `fetch_blobs`
for full bytes for now. Range is a follow up

Accepts Binary / LargeBinary / BinaryView on the way back. Empty
`row_ids` short-circuits. Version + branch go in the request body same
as other read calls.

### Example

```python
db = lancedb.connect(uri="db://my-project", api_key=...)
table = db.open_table("clips")

hits = table.search(query_vec).select(["id", "video"]).limit(10).to_arrow()
# hits is just id + video. row ids are stashed on the descriptor
blobs = table.fetch_blobs("video", hits)  # null-aligned, same length as hits
```

Or pass ids yourself:

```python
blobs = table.fetch_blobs("video", [10, 20, 30])
```

### Testing

- `cargo test -p lancedb --features remote --lib`
- `cargo test -p lancedb --features remote --test blob_integration`
- `pytest python/tests/test_remote_db.py -k remote_blob`
- live e2e against a local 0.5.0 remote server (search → fetch, nulls,
nested path, old server gate)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 12:16:01 -07:00
Joaquin Hui b505dc1315 fix: distinguish corrupt table from missing in open_table (#3731)
`table_names()` lists any `*.lance` directory, but `open_table()` maps
every `DatasetNotFound` to `TableNotFound`, so a corrupt or
partially-written table looks identical to one that never existed
(#3127). This takes the issue's Option 2: on `DatasetNotFound`, check
the parent listing for the table's `.lance` entry — the same predicate
`table_names()` uses — and return a new `TableCorrupted` error when the
directory is present. The check runs only on the error path, and any
failure in the recheck falls back to the previous `TableNotFound`
behavior.

Tests cover the reporter's empty-dir repro, a deleted-manifest case,
true absence (still `TableNotFound`), and an end-to-end list-then-open
assertion; the three new corrupt-case tests fail without the src change.
`cargo test -p lancedb --lib` 732 passed, clippy/fmt clean, `cargo check
--workspace --all-targets` clean (both language bindings end in wildcard
error arms).

Two notes for review: `Error` isn't `#[non_exhaustive]`, so the new
variant is technically semver-breaking for exhaustive matchers (pre-1.0,
and the alternative — changing `TableNotFound`'s shape — breaks more);
and on the Python side corrupt tables now surface as `RuntimeError`
rather than `ValueError`, which is the intended distinction but worth a
maintainer's eye. `open_from_namespace` was left unchanged since
namespace listings come from a server-side registry, not directory
globbing.

Closes #3127
2026-07-30 07:56:43 -07:00
sanskar singh bhardwaj 7dfdfe6401 fix(remote): surface masked merge_insert stream errors under HTTP2 (#2339) (#3711)
## Summary

Fixes #2339. `merge_insert()` on the remote client could mask the real
cause of a mid-stream input error, reporting only:

> stream error sent by user: unexpected internal error

## Root cause

There were two divergent streaming-write code paths in the remote
client:

- `add()` uses `RemoteInsertExec`, which streams the request body
through a `tokio::sync::oneshot` error side-channel and drains it before
reporting the HTTP result. If the input stream errors mid-body, the
original error is recovered.
- `merge_insert()` used a legacy path (`send_streaming` ->
`reader_as_body`) that piped arrow `Some(Err(e))` straight into the
HTTP2 request body. Hyper swallows body-stream errors under HTTP2 (see
hyperium/hyper#2547), so the original error was lost and only the
generic transport error surfaced.

## Fix

Consolidate both write paths onto the side-channel mechanism instead of
patching the legacy path:

- Generalize `RemoteInsertExec` into `RemoteWriteExec`, carrying a
`WriteOp` enum (`Insert { overwrite }` | `MergeInsert { query, timeout
}`) that selects the endpoint, query params, request-timeout header, and
response parsing. The executor returns a `WriteResult` enum (`Add` |
`Merge`) with typed accessors, and `with_new_children` still resets the
result so the rescannable retry loop is unaffected.
- Route `merge_insert()` through `RemoteWriteExec`. The public API only
accepts a `RecordBatchReader` (not rescannable), so the reader is
buffered into a `Vec<RecordBatch>` before the retry loop to preserve the
previous retry-on-retryable-status behaviour. This mirrors what the old
`send_streaming(with_retry=true)` path already did.
- Remove the now-unused `send_streaming` / `reader_as_body` /
`buffer_reader` / `make_reader` helpers. Multipart stays insert-only
(the server has no multipart merge_insert endpoint), so that hot path is
behaviorally unchanged.

## Testing

- Added `test_merge_insert_input_error_surfaces_original`, which drives
an erroring input through the single-request `merge_insert` path and
asserts the original error (`boom`) is surfaced rather than the masked
HTTP error. Confirmed it fails without the side-channel drain (it then
reports a masked `500 ... request or response body error`).
- Full suite green: `cargo test -p lancedb --lib --features remote` ->
694 passed, 0 failed. Includes the existing
`test_merge_insert_retries_on_409`, confirming retry behaviour is
preserved.
2026-07-30 07:56:25 -07:00
heart 4dc2d9a0f2 fix(python): avoid async work in sync reprs (#3620)
## Summary

- keep the existing synchronous `connect()` path unchanged
- make `LanceDBConnection.__repr__` and `LanceTable.__repr__`
side-effect-free
- add a regression test that verifies sync reprs do not call the Python
background loop

## Root cause

The freeze is caused by debugger rendering, not by `connect()` itself:

1. debugpy stops at a breakpoint and suspends all Python threads.
2. The debugger renders the new `db_connection` local by calling
`repr()`.
3. `LanceDBConnection.__repr__` reads `read_consistency_interval`.
4. That property calls `LOOP.run(...).result()`.
5. The `LanceDBBackgroundEventLoop` thread is suspended by the debugger,
so `repr()` waits for a thread that cannot run.

This explains why the symptom appears immediately after `connect()`: it
is the first point where a connection object exists in locals and is
automatically rendered. `LanceTable.__repr__` had the same problem
because it also read the connection's consistency interval.

This follows the same principle as #3411: `__repr__` must not trigger
async work or I/O that a debugger assumes is lightweight.

## Evidence

I reproduced the behavior with the real LanceDB classes and debugpy
1.8.21 using a DAP client:

- latest `main` (`ff6ff099`): the debugger reported `allThreadsStopped:
true`, and evaluating `repr(db_connection)` timed out
- this branch (`5755a5ba`): the same evaluation returned
`LanceDBConnection(uri='/tmp/lancedb-debug-repro')` immediately
- setting `PYDEVD_UNBLOCK_THREADS_TIMEOUT=0` also allowed the original
repr path to complete, independently confirming that it was waiting on a
suspended thread

The regression test creates a connection and table, replaces `LOOP.run`
with a function that fails, and verifies that both reprs still work.

## Validation

- `maturin develop --manifest-path python/Cargo.toml`
- `python -m pytest
python/python/tests/test_db.py::test_sync_repr_does_not_use_background_loop
python/python/tests/test_table.py::test_consistency -q` (`4 passed`)
- `ruff check .`
- `ruff format --check python/python/lancedb/db.py
python/python/lancedb/table.py python/python/tests/test_db.py
python/python/tests/test_table.py`
- `git diff --check`

Refs #3611.
2026-07-30 07:56:13 -07:00
LanceDB Robot 1ad6ce3a4e chore: update lance dependency to v10.0.0-beta.7 (#3745)
Updates the Rust workspace Lance dependencies and Java lance-core
dependency to v10.0.0-beta.7. No compatibility fixes were required; full
workspace clippy passed with warnings denied. Lance tag:
https://github.com/lance-format/lance/releases/tag/v10.0.0-beta.7

---------

Co-authored-by: Lu Qiu <luqiujob@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 18:14:27 -07:00
89 changed files with 8700 additions and 1853 deletions
+29
View File
@@ -92,6 +92,8 @@ Python bindings changes:
* Should use `LOOP.run()` to call the corresponding `AsyncTable` method. * Should use `LOOP.run()` to call the corresponding `AsyncTable` method.
6. Add concrete sync method to `RemoteTable` class in `python/python/lancedb/remote/table.py`. 6. Add concrete sync method to `RemoteTable` class in `python/python/lancedb/remote/table.py`.
7. Add unit test in `python/tests/test_table.py`. 7. Add unit test in `python/tests/test_table.py`.
8. If you added a new public class or module-level function (not just a method on an
existing class), expose it in the API reference. See "Python API reference" below.
TypeScript bindings changes: TypeScript bindings changes:
@@ -103,6 +105,33 @@ TypeScript bindings changes:
5. Add test in `nodejs/__test__/table.test.ts`. 5. Add test in `nodejs/__test__/table.test.ts`.
6. Run `npm run docs` to generate TypeScript documentation. 6. Run `npm run docs` to generate TypeScript documentation.
## Python API reference
`docs/src/python/python.md` is the entire Python API reference. It is maintained by
hand, and anything not listed there is not rendered at all, so new public classes and
module-level functions have to be added explicitly. How depends on the module:
* `lancedb.index`, `lancedb.embeddings`, `lancedb.remote`, and `lancedb.rerankers` are
rendered by a single directive each, driven by the module's `__all__`. Add the new
name to `__all__` and it appears; forget, and it is silently omitted.
* Everything else (`lancedb`, `lancedb.table`, `lancedb.query`, `lancedb.db`, ...) is
listed symbol by symbol. Add a `::: lancedb.<module>.<Name>` line to the matching
section, and remember that the page separates synchronous and asynchronous APIs.
Deliberately undocumented: concrete implementations reached through an abstract base
(`LanceTable`, `LanceDBConnection`, `RemoteDBConnection`), query base classes already
covered by `inherited_members`, and internal helpers.
Cross-references in docstrings use mkdocstrings syntax, `[text][lancedb.table.Table]`.
Plain relative links such as `[Table](Table)` do not resolve. To check your work:
```shell
pip install -r docs/requirements.txt
cd docs && PYTHONPATH=. mkdocs build
```
The docs site only builds on pushes to `main`, so this is not covered by PR CI.
## Review Guidelines ## Review Guidelines
Please consider the following when reviewing code contributions. Please consider the following when reviewing code contributions.
Generated
+114 -118
View File
@@ -601,7 +601,7 @@ dependencies = [
"bytes", "bytes",
"fastrand", "fastrand",
"hex", "hex",
"http 1.4.2", "http 1.5.0",
"sha1 0.10.6", "sha1 0.10.6",
"time", "time",
"tokio", "tokio",
@@ -664,7 +664,7 @@ dependencies = [
"bytes-utils", "bytes-utils",
"fastrand", "fastrand",
"http 0.2.12", "http 0.2.12",
"http 1.4.2", "http 1.5.0",
"http-body 0.4.6", "http-body 0.4.6",
"http-body 1.1.0", "http-body 1.1.0",
"percent-encoding", "percent-encoding",
@@ -694,7 +694,7 @@ dependencies = [
"bytes", "bytes",
"fastrand", "fastrand",
"http 0.2.12", "http 0.2.12",
"http 1.4.2", "http 1.5.0",
"http-body-util", "http-body-util",
"regex-lite", "regex-lite",
"tracing", "tracing",
@@ -719,7 +719,7 @@ dependencies = [
"bytes", "bytes",
"fastrand", "fastrand",
"http 0.2.12", "http 0.2.12",
"http 1.4.2", "http 1.5.0",
"regex-lite", "regex-lite",
"tracing", "tracing",
] ]
@@ -743,7 +743,7 @@ dependencies = [
"bytes", "bytes",
"fastrand", "fastrand",
"http 0.2.12", "http 0.2.12",
"http 1.4.2", "http 1.5.0",
"regex-lite", "regex-lite",
"tracing", "tracing",
] ]
@@ -773,7 +773,7 @@ dependencies = [
"hex", "hex",
"hmac 0.13.0", "hmac 0.13.0",
"http 0.2.12", "http 0.2.12",
"http 1.4.2", "http 1.5.0",
"http-body 1.1.0", "http-body 1.1.0",
"lru", "lru",
"percent-encoding", "percent-encoding",
@@ -802,7 +802,7 @@ dependencies = [
"bytes", "bytes",
"fastrand", "fastrand",
"http 0.2.12", "http 0.2.12",
"http 1.4.2", "http 1.5.0",
"regex-lite", "regex-lite",
"tracing", "tracing",
] ]
@@ -826,7 +826,7 @@ dependencies = [
"bytes", "bytes",
"fastrand", "fastrand",
"http 0.2.12", "http 0.2.12",
"http 1.4.2", "http 1.5.0",
"regex-lite", "regex-lite",
"tracing", "tracing",
] ]
@@ -851,7 +851,7 @@ dependencies = [
"aws-types", "aws-types",
"fastrand", "fastrand",
"http 0.2.12", "http 0.2.12",
"http 1.4.2", "http 1.5.0",
"regex-lite", "regex-lite",
"tracing", "tracing",
] ]
@@ -873,7 +873,7 @@ dependencies = [
"hex", "hex",
"hmac 0.13.0", "hmac 0.13.0",
"http 0.2.12", "http 0.2.12",
"http 1.4.2", "http 1.5.0",
"p256", "p256",
"percent-encoding", "percent-encoding",
"ring", "ring",
@@ -906,7 +906,7 @@ dependencies = [
"bytes", "bytes",
"crc-fast", "crc-fast",
"hex", "hex",
"http 1.4.2", "http 1.5.0",
"http-body 1.1.0", "http-body 1.1.0",
"http-body-util", "http-body-util",
"md-5 0.11.0", "md-5 0.11.0",
@@ -940,7 +940,7 @@ dependencies = [
"bytes-utils", "bytes-utils",
"futures-core", "futures-core",
"futures-util", "futures-util",
"http 1.4.2", "http 1.5.0",
"http-body 1.1.0", "http-body 1.1.0",
"http-body-util", "http-body-util",
"percent-encoding", "percent-encoding",
@@ -961,7 +961,7 @@ dependencies = [
"h2 0.3.27", "h2 0.3.27",
"h2 0.4.14", "h2 0.4.14",
"http 0.2.12", "http 0.2.12",
"http 1.4.2", "http 1.5.0",
"http-body 0.4.6", "http-body 0.4.6",
"hyper 0.14.32", "hyper 0.14.32",
"hyper 1.9.0", "hyper 1.9.0",
@@ -1023,7 +1023,7 @@ dependencies = [
"bytes", "bytes",
"fastrand", "fastrand",
"http 0.2.12", "http 0.2.12",
"http 1.4.2", "http 1.5.0",
"http-body 0.4.6", "http-body 0.4.6",
"http-body 1.1.0", "http-body 1.1.0",
"http-body-util", "http-body-util",
@@ -1044,7 +1044,7 @@ dependencies = [
"aws-smithy-types", "aws-smithy-types",
"bytes", "bytes",
"http 0.2.12", "http 0.2.12",
"http 1.4.2", "http 1.5.0",
"pin-project-lite", "pin-project-lite",
"tokio", "tokio",
"tracing", "tracing",
@@ -1070,7 +1070,7 @@ checksum = "7442cb268338f0eb8278140a107c046756aa01093d8ef5e99628d34ae09c94f5"
dependencies = [ dependencies = [
"aws-smithy-runtime-api", "aws-smithy-runtime-api",
"aws-smithy-types", "aws-smithy-types",
"http 1.4.2", "http 1.5.0",
] ]
[[package]] [[package]]
@@ -1084,7 +1084,7 @@ dependencies = [
"bytes-utils", "bytes-utils",
"futures-core", "futures-core",
"http 0.2.12", "http 0.2.12",
"http 1.4.2", "http 1.5.0",
"http-body 0.4.6", "http-body 0.4.6",
"http-body 1.1.0", "http-body 1.1.0",
"http-body-util", "http-body-util",
@@ -1132,7 +1132,7 @@ dependencies = [
"axum-core", "axum-core",
"bytes", "bytes",
"futures-util", "futures-util",
"http 1.4.2", "http 1.5.0",
"http-body 1.1.0", "http-body 1.1.0",
"http-body-util", "http-body-util",
"hyper 1.9.0", "hyper 1.9.0",
@@ -1165,7 +1165,7 @@ dependencies = [
"async-trait", "async-trait",
"bytes", "bytes",
"futures-util", "futures-util",
"http 1.4.2", "http 1.5.0",
"http-body 1.1.0", "http-body 1.1.0",
"http-body-util", "http-body-util",
"mime", "mime",
@@ -3421,8 +3421,8 @@ checksum = "42703706b716c37f96a77aea830392ad231f44c9e9a67872fa5548707e11b11c"
[[package]] [[package]]
name = "fsst" name = "fsst"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arrow-array", "arrow-array",
"rand 0.9.5", "rand 0.9.5",
@@ -3842,7 +3842,7 @@ dependencies = [
"fnv", "fnv",
"futures-core", "futures-core",
"futures-sink", "futures-sink",
"http 1.4.2", "http 1.5.0",
"indexmap 2.14.0", "indexmap 2.14.0",
"slab", "slab",
"tokio", "tokio",
@@ -3953,7 +3953,7 @@ checksum = "629d8f3bbeda9d148036d6b0de0a3ab947abd08ce90626327fc3547a49d59d97"
dependencies = [ dependencies = [
"dirs", "dirs",
"futures", "futures",
"http 1.4.2", "http 1.5.0",
"indicatif", "indicatif",
"libc", "libc",
"log", "log",
@@ -3976,7 +3976,7 @@ checksum = "430b33fa84f92796d4d263070b6c0d3ca219df7b9a0e1853ee431029b1612bcd"
dependencies = [ dependencies = [
"async-trait", "async-trait",
"bytes", "bytes",
"http 1.4.2", "http 1.5.0",
"more-asserts", "more-asserts",
"serde", "serde",
"thiserror 2.0.18", "thiserror 2.0.18",
@@ -4041,9 +4041,9 @@ dependencies = [
[[package]] [[package]]
name = "http" name = "http"
version = "1.4.2" version = "1.5.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "6970f50e31d6fc17d3fa27329444bfa74e196cf62e95052a3f6fee181dba6425" checksum = "918d3568bebf352712bc2ef3d46a8bcf1a75b373be6539de198e9105cbbf9ce0"
dependencies = [ dependencies = [
"bytes", "bytes",
"itoa", "itoa",
@@ -4067,7 +4067,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ca2a8f2913ee65f60facd6a5905613afaa448497a0230cc41ce022d93290bc2c" checksum = "ca2a8f2913ee65f60facd6a5905613afaa448497a0230cc41ce022d93290bc2c"
dependencies = [ dependencies = [
"bytes", "bytes",
"http 1.4.2", "http 1.5.0",
] ]
[[package]] [[package]]
@@ -4078,7 +4078,7 @@ checksum = "b021d93e26becf5dc7e1b75b1bed1fd93124b374ceb73f43d4d4eafec896a64a"
dependencies = [ dependencies = [
"bytes", "bytes",
"futures-core", "futures-core",
"http 1.4.2", "http 1.5.0",
"http-body 1.1.0", "http-body 1.1.0",
"pin-project-lite", "pin-project-lite",
] ]
@@ -4145,7 +4145,7 @@ dependencies = [
"futures-channel", "futures-channel",
"futures-core", "futures-core",
"h2 0.4.14", "h2 0.4.14",
"http 1.4.2", "http 1.5.0",
"http-body 1.1.0", "http-body 1.1.0",
"httparse", "httparse",
"httpdate", "httpdate",
@@ -4177,7 +4177,7 @@ version = "0.27.9"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "33ca68d021ef39cf6463ab54c1d0f5daf03377b70561305bb89a8f83aab66e0f" checksum = "33ca68d021ef39cf6463ab54c1d0f5daf03377b70561305bb89a8f83aab66e0f"
dependencies = [ dependencies = [
"http 1.4.2", "http 1.5.0",
"hyper 1.9.0", "hyper 1.9.0",
"hyper-util", "hyper-util",
"rustls 0.23.40", "rustls 0.23.40",
@@ -4211,7 +4211,7 @@ dependencies = [
"bytes", "bytes",
"futures-channel", "futures-channel",
"futures-util", "futures-util",
"http 1.4.2", "http 1.5.0",
"http-body 1.1.0", "http-body 1.1.0",
"hyper 1.9.0", "hyper 1.9.0",
"ipnet", "ipnet",
@@ -4777,8 +4777,8 @@ checksum = "e037a2e1d8d5fdbd49b16a4ea09d5d6401c1f29eca5ff29d03d3824dba16256a"
[[package]] [[package]]
name = "lance" name = "lance"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arc-swap", "arc-swap",
"arrow", "arrow",
@@ -4852,8 +4852,8 @@ dependencies = [
[[package]] [[package]]
name = "lance-arrow" name = "lance-arrow"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arrow-array", "arrow-array",
"arrow-buffer", "arrow-buffer",
@@ -4875,7 +4875,7 @@ dependencies = [
[[package]] [[package]]
name = "lance-arrow-scalar" name = "lance-arrow-scalar"
version = "58.0.0" version = "58.0.0"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arrow-array", "arrow-array",
"arrow-buffer", "arrow-buffer",
@@ -4889,7 +4889,7 @@ dependencies = [
[[package]] [[package]]
name = "lance-arrow-stats" name = "lance-arrow-stats"
version = "58.0.0" version = "58.0.0"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arrow-array", "arrow-array",
"arrow-schema", "arrow-schema",
@@ -4898,8 +4898,8 @@ dependencies = [
[[package]] [[package]]
name = "lance-bitpacking" name = "lance-bitpacking"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arrayref", "arrayref",
"crunchy", "crunchy",
@@ -4909,14 +4909,15 @@ dependencies = [
[[package]] [[package]]
name = "lance-core" name = "lance-core"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arrow-array", "arrow-array",
"arrow-buffer", "arrow-buffer",
"arrow-data", "arrow-data",
"arrow-schema", "arrow-schema",
"async-trait", "async-trait",
"blake3",
"byteorder", "byteorder",
"bytes", "bytes",
"datafusion-common", "datafusion-common",
@@ -4949,8 +4950,8 @@ dependencies = [
[[package]] [[package]]
name = "lance-datafusion" name = "lance-datafusion"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arrow", "arrow",
"arrow-array", "arrow-array",
@@ -4980,8 +4981,8 @@ dependencies = [
[[package]] [[package]]
name = "lance-datagen" name = "lance-datagen"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arrow", "arrow",
"arrow-array", "arrow-array",
@@ -4998,8 +4999,8 @@ dependencies = [
[[package]] [[package]]
name = "lance-derive" name = "lance-derive"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"proc-macro2", "proc-macro2",
"quote", "quote",
@@ -5008,8 +5009,8 @@ dependencies = [
[[package]] [[package]]
name = "lance-encoding" name = "lance-encoding"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arrow-arith", "arrow-arith",
"arrow-array", "arrow-array",
@@ -5044,12 +5045,13 @@ dependencies = [
[[package]] [[package]]
name = "lance-file" name = "lance-file"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arrow-arith", "arrow-arith",
"arrow-array", "arrow-array",
"arrow-buffer", "arrow-buffer",
"arrow-cast",
"arrow-data", "arrow-data",
"arrow-schema", "arrow-schema",
"arrow-select", "arrow-select",
@@ -5075,8 +5077,8 @@ dependencies = [
[[package]] [[package]]
name = "lance-index" name = "lance-index"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arc-swap", "arc-swap",
"arrow", "arrow",
@@ -5143,8 +5145,8 @@ dependencies = [
[[package]] [[package]]
name = "lance-index-core" name = "lance-index-core"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arrow-array", "arrow-array",
"arrow-schema", "arrow-schema",
@@ -5166,18 +5168,12 @@ dependencies = [
[[package]] [[package]]
name = "lance-io" name = "lance-io"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arrow", "arrow",
"arrow-arith",
"arrow-array", "arrow-array",
"arrow-buffer",
"arrow-cast",
"arrow-data",
"arrow-schema", "arrow-schema",
"arrow-select",
"async-recursion",
"async-trait", "async-trait",
"aws-config", "aws-config",
"aws-credential-types", "aws-credential-types",
@@ -5186,7 +5182,7 @@ dependencies = [
"chrono", "chrono",
"futures", "futures",
"goosefs-sdk", "goosefs-sdk",
"http 1.4.2", "http 1.5.0",
"io-uring", "io-uring",
"lance-arrow", "lance-arrow",
"lance-core", "lance-core",
@@ -5210,8 +5206,8 @@ dependencies = [
[[package]] [[package]]
name = "lance-linalg" name = "lance-linalg"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arrow-array", "arrow-array",
"arrow-buffer", "arrow-buffer",
@@ -5227,8 +5223,8 @@ dependencies = [
[[package]] [[package]]
name = "lance-namespace" name = "lance-namespace"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arrow", "arrow",
"async-trait", "async-trait",
@@ -5240,8 +5236,8 @@ dependencies = [
[[package]] [[package]]
name = "lance-namespace-impls" name = "lance-namespace-impls"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arrow", "arrow",
"arrow-ipc", "arrow-ipc",
@@ -5295,8 +5291,8 @@ dependencies = [
[[package]] [[package]]
name = "lance-select" name = "lance-select"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arrow-array", "arrow-array",
"arrow-buffer", "arrow-buffer",
@@ -5311,8 +5307,8 @@ dependencies = [
[[package]] [[package]]
name = "lance-table" name = "lance-table"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arrow", "arrow",
"arrow-array", "arrow-array",
@@ -5351,8 +5347,8 @@ dependencies = [
[[package]] [[package]]
name = "lance-testing" name = "lance-testing"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"arrow-array", "arrow-array",
"arrow-schema", "arrow-schema",
@@ -5365,8 +5361,8 @@ dependencies = [
[[package]] [[package]]
name = "lance-tokenizer" name = "lance-tokenizer"
version = "10.0.0-beta.5" version = "10.1.0-beta.1"
source = "git+https://github.com/lance-format/lance.git?tag=v10.0.0-beta.5#ddb8e28ca238f29628b8e1795ddccbb7bf75e5c8" source = "git+https://github.com/lance-format/lance.git?tag=v10.1.0-beta.1#68f4d4c1d0c4871b067557c61fc405078f1ab3b7"
dependencies = [ dependencies = [
"icu_segmenter", "icu_segmenter",
"jieba-rs", "jieba-rs",
@@ -5418,7 +5414,7 @@ dependencies = [
"goosefs-sdk", "goosefs-sdk",
"half", "half",
"hf-hub", "hf-hub",
"http 1.4.2", "http 1.5.0",
"http-body 1.1.0", "http-body 1.1.0",
"lance", "lance",
"lance-arrow", "lance-arrow",
@@ -6084,15 +6080,15 @@ dependencies = [
[[package]] [[package]]
name = "napi-build" name = "napi-build"
version = "2.3.2" version = "2.4.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c9c366d2c8c60b86fa632df75f745509b52f9128f91a6bad4c796e44abb505e1" checksum = "5282704fbe8d49b0cf8b08e3f33233416a528658f205c7e5ace63b582de0b11c"
[[package]] [[package]]
name = "napi-derive" name = "napi-derive"
version = "3.6.0" version = "3.6.1"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "a49c513341a61a16a10af6efcce46b30d0822ba2d4fb197d24d33dfc199c78d5" checksum = "4d5c9c02556ea6dc99dffd36c1ce60141411657438501a125b675776d011ce92"
dependencies = [ dependencies = [
"convert_case", "convert_case",
"ctor 1.0.5", "ctor 1.0.5",
@@ -6104,9 +6100,9 @@ dependencies = [
[[package]] [[package]]
name = "napi-derive-backend" name = "napi-derive-backend"
version = "6.0.0" version = "6.1.1"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "4747005fa3e2c9989ac45a723a514c5db2411238b72981a3cda4c701a9dfea17" checksum = "d60b5d773ad46c698c8cc2cd9fde0b283d39cbb7f71c04bee633c7bdba4423bd"
dependencies = [ dependencies = [
"convert_case", "convert_case",
"proc-macro2", "proc-macro2",
@@ -6360,7 +6356,7 @@ dependencies = [
"futures-channel", "futures-channel",
"futures-core", "futures-core",
"futures-util", "futures-util",
"http 1.4.2", "http 1.5.0",
"http-body-util", "http-body-util",
"httparse", "httparse",
"humantime", "humantime",
@@ -6481,7 +6477,7 @@ dependencies = [
"base64 0.22.1", "base64 0.22.1",
"bytes", "bytes",
"futures", "futures",
"http 1.4.2", "http 1.5.0",
"http-body 1.1.0", "http-body 1.1.0",
"jiff", "jiff",
"log", "log",
@@ -6506,7 +6502,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0d6f81ba6960e3fae1882f253b114b21d7e444e1534f209c7737a79f6243eb6f" checksum = "0d6f81ba6960e3fae1882f253b114b21d7e444e1534f209c7737a79f6243eb6f"
dependencies = [ dependencies = [
"futures", "futures",
"http 1.4.2", "http 1.5.0",
"mea", "mea",
"opendal-core", "opendal-core",
] ]
@@ -6550,7 +6546,7 @@ checksum = "0030644366ef5d8cbe3a4a5822bf99a4aafddc1666e9d24b44d158d9062fc76a"
dependencies = [ dependencies = [
"base64 0.22.1", "base64 0.22.1",
"bytes", "bytes",
"http 1.4.2", "http 1.5.0",
"log", "log",
"opendal-core", "opendal-core",
"opendal-service-azure-common", "opendal-service-azure-common",
@@ -6571,7 +6567,7 @@ checksum = "6dea4908d490143a9b0b7f7a790e139ff829b06a023f670455ed3d44f664b361"
dependencies = [ dependencies = [
"base64 0.22.1", "base64 0.22.1",
"bytes", "bytes",
"http 1.4.2", "http 1.5.0",
"log", "log",
"opendal-core", "opendal-core",
"opendal-service-azure-common", "opendal-service-azure-common",
@@ -6589,7 +6585,7 @@ version = "0.57.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9b489f13c42e69d69bdd72952b634356ec43a7881a20259b38b540fcecdf4051" checksum = "9b489f13c42e69d69bdd72952b634356ec43a7881a20259b38b540fcecdf4051"
dependencies = [ dependencies = [
"http 1.4.2", "http 1.5.0",
"opendal-core", "opendal-core",
] ]
@@ -6600,7 +6596,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "aa8cafe9729213375c7331019b0cb756ad3e1aff7f45cd32c45eae91ebde8901" checksum = "aa8cafe9729213375c7331019b0cb756ad3e1aff7f45cd32c45eae91ebde8901"
dependencies = [ dependencies = [
"bytes", "bytes",
"http 1.4.2", "http 1.5.0",
"log", "log",
"opendal-core", "opendal-core",
"quick-xml 0.39.4", "quick-xml 0.39.4",
@@ -6618,7 +6614,7 @@ checksum = "48de101aac565ed06af4b47903c24eafd249075553ec1fb18256751c45148d47"
dependencies = [ dependencies = [
"async-trait", "async-trait",
"bytes", "bytes",
"http 1.4.2", "http 1.5.0",
"log", "log",
"opendal-core", "opendal-core",
"percent-encoding", "percent-encoding",
@@ -6653,7 +6649,7 @@ checksum = "c4922661976a1d40794a2adfbdb888cc3c23097690f825a92f773af38908a848"
dependencies = [ dependencies = [
"bytes", "bytes",
"hf-xet", "hf-xet",
"http 1.4.2", "http 1.5.0",
"log", "log",
"opendal-core", "opendal-core",
"percent-encoding", "percent-encoding",
@@ -6669,7 +6665,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "328fa55e8888cbdfe00826bfea2a79042422b720e8369e9e021e46121dea5ace" checksum = "328fa55e8888cbdfe00826bfea2a79042422b720e8369e9e021e46121dea5ace"
dependencies = [ dependencies = [
"bytes", "bytes",
"http 1.4.2", "http 1.5.0",
"log", "log",
"opendal-core", "opendal-core",
"quick-xml 0.39.4", "quick-xml 0.39.4",
@@ -6688,7 +6684,7 @@ dependencies = [
"base64 0.22.1", "base64 0.22.1",
"bytes", "bytes",
"crc32c", "crc32c",
"http 1.4.2", "http 1.5.0",
"log", "log",
"md-5 0.11.0", "md-5 0.11.0",
"opendal-core", "opendal-core",
@@ -7587,7 +7583,7 @@ version = "0.14.3"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "343d3bd7056eda839b03204e68deff7d1b13aba7af2b2fd16890697274262ee7" checksum = "343d3bd7056eda839b03204e68deff7d1b13aba7af2b2fd16890697274262ee7"
dependencies = [ dependencies = [
"heck 0.5.0", "heck 0.4.1",
"itertools 0.14.0", "itertools 0.14.0",
"log", "log",
"multimap", "multimap",
@@ -8230,7 +8226,7 @@ checksum = "57ac2757f3140aa2e213b554148ae0b52733e624fc6723f0cc6bb3d440176c95"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"form_urlencoded", "form_urlencoded",
"http 1.4.2", "http 1.5.0",
"log", "log",
"percent-encoding", "percent-encoding",
"reqsign-core", "reqsign-core",
@@ -8248,7 +8244,7 @@ dependencies = [
"anyhow", "anyhow",
"bytes", "bytes",
"form_urlencoded", "form_urlencoded",
"http 1.4.2", "http 1.5.0",
"log", "log",
"percent-encoding", "percent-encoding",
"quick-xml 0.39.4", "quick-xml 0.39.4",
@@ -8270,7 +8266,7 @@ dependencies = [
"base64 0.22.1", "base64 0.22.1",
"bytes", "bytes",
"form_urlencoded", "form_urlencoded",
"http 1.4.2", "http 1.5.0",
"jsonwebtoken", "jsonwebtoken",
"log", "log",
"pem", "pem",
@@ -8295,7 +8291,7 @@ dependencies = [
"futures", "futures",
"hex", "hex",
"hmac 0.12.1", "hmac 0.12.1",
"http 1.4.2", "http 1.5.0",
"jiff", "jiff",
"log", "log",
"percent-encoding", "percent-encoding",
@@ -8322,7 +8318,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "35cc609b49c69e76ecaceb775a03f792d1ed3e7755ab3548d4534fd801e3242e" checksum = "35cc609b49c69e76ecaceb775a03f792d1ed3e7755ab3548d4534fd801e3242e"
dependencies = [ dependencies = [
"form_urlencoded", "form_urlencoded",
"http 1.4.2", "http 1.5.0",
"jsonwebtoken", "jsonwebtoken",
"log", "log",
"percent-encoding", "percent-encoding",
@@ -8342,7 +8338,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e128f19525861dbded59e1e7c17653a8ed63d573ca04aed708d552dbef5bb32a" checksum = "e128f19525861dbded59e1e7c17653a8ed63d573ca04aed708d552dbef5bb32a"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"http 1.4.2", "http 1.5.0",
"log", "log",
"percent-encoding", "percent-encoding",
"reqsign-core", "reqsign-core",
@@ -8362,7 +8358,7 @@ dependencies = [
"futures-core", "futures-core",
"futures-util", "futures-util",
"h2 0.4.14", "h2 0.4.14",
"http 1.4.2", "http 1.5.0",
"http-body 1.1.0", "http-body 1.1.0",
"http-body-util", "http-body-util",
"hyper 1.9.0", "hyper 1.9.0",
@@ -8406,7 +8402,7 @@ dependencies = [
"bytes", "bytes",
"futures-core", "futures-core",
"futures-util", "futures-util",
"http 1.4.2", "http 1.5.0",
"http-body 1.1.0", "http-body 1.1.0",
"http-body-util", "http-body-util",
"hyper 1.9.0", "hyper 1.9.0",
@@ -8460,7 +8456,7 @@ checksum = "199dda04a536b532d0cc04d7979e39b1c763ea749bf91507017069c00b96056f"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"async-trait", "async-trait",
"http 1.4.2", "http 1.5.0",
"reqwest 0.13.3", "reqwest 0.13.3",
"thiserror 2.0.18", "thiserror 2.0.18",
"tower-service", "tower-service",
@@ -8502,9 +8498,9 @@ dependencies = [
[[package]] [[package]]
name = "rkyv" name = "rkyv"
version = "0.8.16" version = "0.8.17"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "73389e0c99e664f919275ab5b5b0471391fe9a8de61e1dff9b1eaf56a90f16e3" checksum = "815cc8a37159a463064825246cadb07961e25cd9885908606f6d08a98d8f8874"
dependencies = [ dependencies = [
"bytecheck", "bytecheck",
"bytes", "bytes",
@@ -8521,9 +8517,9 @@ dependencies = [
[[package]] [[package]]
name = "rkyv_derive" name = "rkyv_derive"
version = "0.8.16" version = "0.8.17"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "5d2ed0b54125315fb36bd021e82d314d1c126548f871634b483f46b31d13cac6" checksum = "c0ed1a78a1b19d184b0daa629dd9a024573173ec7d485b287cb369fb3607cc1c"
dependencies = [ dependencies = [
"proc-macro2", "proc-macro2",
"quote", "quote",
@@ -9281,7 +9277,7 @@ version = "0.8.9"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c1c97747dbf44bb1ca44a561ece23508e99cb592e862f22222dcf42f51d1e451" checksum = "c1c97747dbf44bb1ca44a561ece23508e99cb592e862f22222dcf42f51d1e451"
dependencies = [ dependencies = [
"heck 0.5.0", "heck 0.4.1",
"proc-macro2", "proc-macro2",
"quote", "quote",
"syn 2.0.117", "syn 2.0.117",
@@ -9293,7 +9289,7 @@ version = "0.9.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "54254b8531cafa275c5e096f62d48c81435d1015405a91198ddb11e967301d40" checksum = "54254b8531cafa275c5e096f62d48c81435d1015405a91198ddb11e967301d40"
dependencies = [ dependencies = [
"heck 0.5.0", "heck 0.4.1",
"proc-macro2", "proc-macro2",
"quote", "quote",
"syn 2.0.117", "syn 2.0.117",
@@ -9734,7 +9730,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "32497e9a4c7b38532efcdebeef879707aa9f794296a4f0244f6f69e9bc8574bd" checksum = "32497e9a4c7b38532efcdebeef879707aa9f794296a4f0244f6f69e9bc8574bd"
dependencies = [ dependencies = [
"fastrand", "fastrand",
"getrandom 0.4.2", "getrandom 0.3.4",
"once_cell", "once_cell",
"rustix", "rustix",
"windows-sys 0.61.2", "windows-sys 0.61.2",
@@ -10062,7 +10058,7 @@ dependencies = [
"base64 0.22.1", "base64 0.22.1",
"bytes", "bytes",
"h2 0.4.14", "h2 0.4.14",
"http 1.4.2", "http 1.5.0",
"http-body 1.1.0", "http-body 1.1.0",
"http-body-util", "http-body-util",
"hyper 1.9.0", "hyper 1.9.0",
@@ -10118,7 +10114,7 @@ checksum = "1e9cd434a998747dd2c4276bc96ee2e0c7a2eadf3cae88e52be55a05fa9053f5"
dependencies = [ dependencies = [
"bitflags 2.11.1", "bitflags 2.11.1",
"bytes", "bytes",
"http 1.4.2", "http 1.5.0",
"http-body 1.1.0", "http-body 1.1.0",
"http-body-util", "http-body-util",
"pin-project-lite", "pin-project-lite",
@@ -10138,7 +10134,7 @@ dependencies = [
"bytes", "bytes",
"futures-core", "futures-core",
"futures-util", "futures-util",
"http 1.4.2", "http 1.5.0",
"http-body 1.1.0", "http-body 1.1.0",
"http-body-util", "http-body-util",
"pin-project-lite", "pin-project-lite",
@@ -11159,7 +11155,7 @@ dependencies = [
"clap", "clap",
"crc32fast", "crc32fast",
"futures", "futures",
"http 1.4.2", "http 1.5.0",
"hyper 1.9.0", "hyper 1.9.0",
"lazy_static", "lazy_static",
"more-asserts", "more-asserts",
@@ -11233,7 +11229,7 @@ dependencies = [
"chrono", "chrono",
"clap", "clap",
"gearhash", "gearhash",
"http 1.4.2", "http 1.5.0",
"itertools 0.14.0", "itertools 0.14.0",
"lazy_static", "lazy_static",
"more-asserts", "more-asserts",
+14 -14
View File
@@ -13,20 +13,20 @@ categories = ["database-implementations"]
rust-version = "1.91.0" rust-version = "1.91.0"
[workspace.dependencies] [workspace.dependencies]
lance = { "version" = "=10.0.0-beta.5", default-features = false, "tag" = "v10.0.0-beta.5", "git" = "https://github.com/lance-format/lance.git" } lance = { "version" = "=10.1.0-beta.1", default-features = false, "tag" = "v10.1.0-beta.1", "git" = "https://github.com/lance-format/lance.git" }
lance-core = { "version" = "=10.0.0-beta.5", "tag" = "v10.0.0-beta.5", "git" = "https://github.com/lance-format/lance.git" } lance-core = { "version" = "=10.1.0-beta.1", "tag" = "v10.1.0-beta.1", "git" = "https://github.com/lance-format/lance.git" }
lance-datagen = { "version" = "=10.0.0-beta.5", "tag" = "v10.0.0-beta.5", "git" = "https://github.com/lance-format/lance.git" } lance-datagen = { "version" = "=10.1.0-beta.1", "tag" = "v10.1.0-beta.1", "git" = "https://github.com/lance-format/lance.git" }
lance-file = { "version" = "=10.0.0-beta.5", "tag" = "v10.0.0-beta.5", "git" = "https://github.com/lance-format/lance.git" } lance-file = { "version" = "=10.1.0-beta.1", "tag" = "v10.1.0-beta.1", "git" = "https://github.com/lance-format/lance.git" }
lance-io = { "version" = "=10.0.0-beta.5", default-features = false, "tag" = "v10.0.0-beta.5", "git" = "https://github.com/lance-format/lance.git" } lance-io = { "version" = "=10.1.0-beta.1", default-features = false, "tag" = "v10.1.0-beta.1", "git" = "https://github.com/lance-format/lance.git" }
lance-index = { "version" = "=10.0.0-beta.5", "tag" = "v10.0.0-beta.5", "git" = "https://github.com/lance-format/lance.git" } lance-index = { "version" = "=10.1.0-beta.1", "tag" = "v10.1.0-beta.1", "git" = "https://github.com/lance-format/lance.git" }
lance-linalg = { "version" = "=10.0.0-beta.5", "tag" = "v10.0.0-beta.5", "git" = "https://github.com/lance-format/lance.git" } lance-linalg = { "version" = "=10.1.0-beta.1", "tag" = "v10.1.0-beta.1", "git" = "https://github.com/lance-format/lance.git" }
lance-namespace = { "version" = "=10.0.0-beta.5", "tag" = "v10.0.0-beta.5", "git" = "https://github.com/lance-format/lance.git" } lance-namespace = { "version" = "=10.1.0-beta.1", "tag" = "v10.1.0-beta.1", "git" = "https://github.com/lance-format/lance.git" }
lance-namespace-impls = { "version" = "=10.0.0-beta.5", default-features = false, "tag" = "v10.0.0-beta.5", "git" = "https://github.com/lance-format/lance.git" } lance-namespace-impls = { "version" = "=10.1.0-beta.1", default-features = false, "tag" = "v10.1.0-beta.1", "git" = "https://github.com/lance-format/lance.git" }
lance-table = { "version" = "=10.0.0-beta.5", "tag" = "v10.0.0-beta.5", "git" = "https://github.com/lance-format/lance.git" } lance-table = { "version" = "=10.1.0-beta.1", "tag" = "v10.1.0-beta.1", "git" = "https://github.com/lance-format/lance.git" }
lance-testing = { "version" = "=10.0.0-beta.5", "tag" = "v10.0.0-beta.5", "git" = "https://github.com/lance-format/lance.git" } lance-testing = { "version" = "=10.1.0-beta.1", "tag" = "v10.1.0-beta.1", "git" = "https://github.com/lance-format/lance.git" }
lance-datafusion = { "version" = "=10.0.0-beta.5", "tag" = "v10.0.0-beta.5", "git" = "https://github.com/lance-format/lance.git" } lance-datafusion = { "version" = "=10.1.0-beta.1", "tag" = "v10.1.0-beta.1", "git" = "https://github.com/lance-format/lance.git" }
lance-encoding = { "version" = "=10.0.0-beta.5", "tag" = "v10.0.0-beta.5", "git" = "https://github.com/lance-format/lance.git" } lance-encoding = { "version" = "=10.1.0-beta.1", "tag" = "v10.1.0-beta.1", "git" = "https://github.com/lance-format/lance.git" }
lance-arrow = { "version" = "=10.0.0-beta.5", "tag" = "v10.0.0-beta.5", "git" = "https://github.com/lance-format/lance.git" } lance-arrow = { "version" = "=10.1.0-beta.1", "tag" = "v10.1.0-beta.1", "git" = "https://github.com/lance-format/lance.git" }
ahash = "0.8" ahash = "0.8"
# Note that this one does not include pyarrow # Note that this one does not include pyarrow
arrow = { version = "58.0.0", optional = false } arrow = { version = "58.0.0", optional = false }
+5
View File
@@ -51,6 +51,11 @@ plugins:
paths: [../python/python] paths: [../python/python]
options: options:
docstring_style: numpy docstring_style: numpy
docstring_options:
# Attributes documented in a `Parameters` section, and pydantic
# dataclasses whose `__init__` griffe cannot see statically, both
# trip this check. It reports nothing actionable here.
warn_unknown_params: false
heading_level: 3 heading_level: 3
show_signature_annotations: true show_signature_annotations: true
show_root_heading: true show_root_heading: true
+1 -1
View File
@@ -1,7 +1,7 @@
# Contributing to LanceDB Typescript # Contributing to LanceDB Typescript
This document outlines the process for contributing to LanceDB Typescript. This document outlines the process for contributing to LanceDB Typescript.
For general contribution guidelines, see [CONTRIBUTING.md](../CONTRIBUTING.md). For general contribution guidelines, see [CONTRIBUTING.md](https://github.com/lancedb/lancedb/blob/main/CONTRIBUTING.md).
## Project layout ## Project layout
+97
View File
@@ -25,6 +25,27 @@ the underlying connection has been closed.
## Methods ## Methods
### cancelJob()
```ts
abstract cancelJob(jobId): Promise<boolean>
```
Request cancellation of a server-side job by id.
Resolves to true if the server accepted the cancellation, false if no
such job exists. Cancelling an already-terminal job is a no-op success.
#### Parameters
* **jobId**: `string`
#### Returns
`Promise`&lt;`boolean`&gt;
***
### cloneTable() ### cloneTable()
```ts ```ts
@@ -365,6 +386,26 @@ Drop an existing table.
*** ***
### getJob()
```ts
abstract getJob(jobId): Promise<null | JobDescription>
```
Describe a single server-side job by id.
Resolves to `null` when the server has no such job.
#### Parameters
* **jobId**: `string`
#### Returns
`Promise`&lt;`null` \| [`JobDescription`](../interfaces/JobDescription.md)&gt;
***
### isOpen() ### isOpen()
```ts ```ts
@@ -379,6 +420,62 @@ Return true if the connection has not been closed
*** ***
### job()
```ts
abstract job(jobId): Job
```
A [Job](Job.md) handle for a server-side job by id.
The handle is constructed without a server round trip; an unknown id
surfaces when the handle is used. Dropping the handle has no effect on
the job itself.
#### Parameters
* **jobId**: `string`
#### Returns
[`Job`](Job.md)
***
### jobHistory()
```ts
abstract jobHistory(jobId?): Promise<Table<any>>
```
The lifecycle event history of a server-side job, as an Arrow table.
Lists history across all jobs when `jobId` is omitted.
#### Parameters
* **jobId?**: `string`
#### Returns
`Promise`&lt;`Table`&lt;`any`&gt;&gt;
***
### listJobs()
```ts
abstract listJobs(): Promise<JobInfo[]>
```
List server-side jobs across the database's tables.
#### Returns
`Promise`&lt;[`JobInfo`](../interfaces/JobInfo.md)[]&gt;
***
### listNamespaces() ### listNamespaces()
```ts ```ts
+83
View File
@@ -0,0 +1,83 @@
[**@lancedb/lancedb**](../README.md) • **Docs**
***
[@lancedb/lancedb](../globals.md) / Job
# Class: Job
A handle to an operation that may still be running.
## Constructors
### new Job()
```ts
new Job(): Job
```
#### Returns
[`Job`](Job.md)
## Accessors
### id
```ts
get id(): null | string
```
Identifies the operation on the server that is running it. Operations
that run in this process have no server id. The value is opaque.
#### Returns
`null` \| `string`
## Methods
### cancel()
```ts
cancel(): Promise<void>
```
Request cancellation. Cancelling a finished operation is a no-op.
#### Returns
`Promise`&lt;`void`&gt;
***
### status()
```ts
status(): Promise<string>
```
The operation's current lifecycle state: "running", "finished",
"failed", or "cancelled".
A point snapshot; unlike [Job.wait](Job.md#wait) it does not block or reject
on a terminal failure state. States a newer server reports that this
client version does not know pass through as-is.
#### Returns
`Promise`&lt;`string`&gt;
***
### wait()
```ts
wait(): Promise<void>
```
Wait until the operation reaches a terminal state.
#### Returns
`Promise`&lt;`void`&gt;
+23
View File
@@ -295,6 +295,29 @@ await table.createIndex("my_float_col");
*** ***
### createIndexAsync()
```ts
abstract createIndexAsync(column, options?): Promise<Job>
```
Create an index, returning a handle to the indexing job.
The job may already be complete when returned; callers must not assume
the index exists until [Job.wait](Job.md#wait) resolves.
#### Parameters
* **column**: `string`
* **options?**: `Partial`&lt;[`IndexOptions`](../interfaces/IndexOptions.md)&gt;
#### Returns
`Promise`&lt;[`Job`](Job.md)&gt;
***
### currentBranch() ### currentBranch()
```ts ```ts
+4
View File
@@ -25,6 +25,7 @@
- [Connection](classes/Connection.md) - [Connection](classes/Connection.md)
- [HeaderProvider](classes/HeaderProvider.md) - [HeaderProvider](classes/HeaderProvider.md)
- [Index](classes/Index.md) - [Index](classes/Index.md)
- [Job](classes/Job.md)
- [MakeArrowTableOptions](classes/MakeArrowTableOptions.md) - [MakeArrowTableOptions](classes/MakeArrowTableOptions.md)
- [MatchQuery](classes/MatchQuery.md) - [MatchQuery](classes/MatchQuery.md)
- [MergeInsertBuilder](classes/MergeInsertBuilder.md) - [MergeInsertBuilder](classes/MergeInsertBuilder.md)
@@ -88,6 +89,9 @@
- [IvfFlatOptions](interfaces/IvfFlatOptions.md) - [IvfFlatOptions](interfaces/IvfFlatOptions.md)
- [IvfPqOptions](interfaces/IvfPqOptions.md) - [IvfPqOptions](interfaces/IvfPqOptions.md)
- [IvfRqOptions](interfaces/IvfRqOptions.md) - [IvfRqOptions](interfaces/IvfRqOptions.md)
- [JobDescription](interfaces/JobDescription.md)
- [JobFailureInfo](interfaces/JobFailureInfo.md)
- [JobInfo](interfaces/JobInfo.md)
- [ListNamespacesOptions](interfaces/ListNamespacesOptions.md) - [ListNamespacesOptions](interfaces/ListNamespacesOptions.md)
- [ListNamespacesResponse](interfaces/ListNamespacesResponse.md) - [ListNamespacesResponse](interfaces/ListNamespacesResponse.md)
- [LsmWriteSpec](interfaces/LsmWriteSpec.md) - [LsmWriteSpec](interfaces/LsmWriteSpec.md)
+66
View File
@@ -0,0 +1,66 @@
[**@lancedb/lancedb**](../README.md) • **Docs**
***
[@lancedb/lancedb](../globals.md) / JobDescription
# Interface: JobDescription
A described job from `Connection.getJob`.
## Properties
### creationMs
```ts
creationMs: number;
```
When the job was created, in milliseconds since the epoch.
***
### failure?
```ts
optional failure: JobFailureInfo;
```
Why the job failed, when the job is failed and the server reports a
reason.
***
### jobId
```ts
jobId: string;
```
***
### jobType
```ts
jobType: string;
```
***
### specJson?
```ts
optional specJson: string;
```
The job-type-specific specification as a JSON string, when present.
***
### state
```ts
state: string;
```
Lifecycle state: "running", "finished", "failed", or "cancelled".
+33
View File
@@ -0,0 +1,33 @@
[**@lancedb/lancedb**](../README.md) • **Docs**
***
[@lancedb/lancedb](../globals.md) / JobFailureInfo
# Interface: JobFailureInfo
The server's account of why a job failed.
## Properties
### message?
```ts
optional message: string;
```
***
### phase?
```ts
optional phase: string;
```
***
### retryable?
```ts
optional retryable: boolean;
```
+58
View File
@@ -0,0 +1,58 @@
[**@lancedb/lancedb**](../README.md) • **Docs**
***
[@lancedb/lancedb](../globals.md) / JobInfo
# Interface: JobInfo
A row from `Connection.listJobs`: one server-side job.
## Properties
### createdAtMillis
```ts
createdAtMillis: number;
```
When the job was created, in milliseconds since the epoch.
***
### jobId
```ts
jobId: string;
```
The job id -- what `Connection.getJob` and `Connection.cancelJob`
accept.
***
### jobType
```ts
jobType: string;
```
***
### state
```ts
state: string;
```
Lifecycle state: "running", "finished", "failed", or "cancelled".
***
### table
```ts
table: string;
```
The table the job runs against, without URI or namespace.
+114 -49
View File
@@ -26,6 +26,18 @@ is also an [asynchronous API client](#connections-asynchronous).
::: lancedb.db.DBConnection ::: lancedb.db.DBConnection
::: lancedb.Session
## Namespaces (Synchronous)
A namespace-backed connection resolves tables through a
[Lance namespace](https://lancedb.github.io/lance-namespace/) service instead of
listing a storage directory.
::: lancedb.connect_namespace
::: lancedb.namespace.LanceNamespaceDBConnection
## Tables (Synchronous) ## Tables (Synchronous)
::: lancedb.table.Table ::: lancedb.table.Table
@@ -34,8 +46,12 @@ is also an [asynchronous API client](#connections-asynchronous).
::: lancedb.table.FragmentSummaryStats ::: lancedb.table.FragmentSummaryStats
::: lancedb.table.TableStatistics
::: lancedb.table.Tags ::: lancedb.table.Tags
::: lancedb.table.Branches
## Expressions ## Expressions
Type-safe expression builder for filters and projections. Use these instead Type-safe expression builder for filters and projections. Use these instead
@@ -62,29 +78,46 @@ of raw SQL strings with [where][lancedb.query.LanceQueryBuilder.where] and
::: lancedb.query.LanceHybridQueryBuilder ::: lancedb.query.LanceHybridQueryBuilder
::: lancedb.query.LanceEmptyQueryBuilder
::: lancedb.query.LanceTakeQueryBuilder
## Full text queries
Structured full text queries can be passed to
[Table.search][lancedb.table.Table.search] or
[AsyncTable.search][lancedb.table.AsyncTable.search] in place of a query string,
and combined with [BooleanQuery][lancedb.query.BooleanQuery].
::: lancedb.query.FullTextQuery
::: lancedb.query.MatchQuery
::: lancedb.query.PhraseQuery
::: lancedb.query.BoostQuery
::: lancedb.query.MultiMatchQuery
::: lancedb.query.BooleanQuery
::: lancedb.query.FullTextOperator
::: lancedb.query.Occur
## Embeddings ## Embeddings
::: lancedb.embeddings.registry.EmbeddingFunctionRegistry ::: lancedb.embeddings
options:
::: lancedb.embeddings.base.EmbeddingFunctionConfig show_root_heading: false
show_root_toc_entry: false
::: lancedb.embeddings.base.EmbeddingFunction
::: lancedb.embeddings.base.TextEmbeddingFunction
::: lancedb.embeddings.sentence_transformers.SentenceTransformerEmbeddings
::: lancedb.embeddings.openai.OpenAIEmbeddings
::: lancedb.embeddings.open_clip.OpenClipEmbeddings
## Remote configuration ## Remote configuration
::: lancedb.remote.ClientConfig ::: lancedb.remote
options:
::: lancedb.remote.TimeoutConfig show_root_heading: false
show_root_toc_entry: false
::: lancedb.remote.RetryConfig
## Context ## Context
@@ -122,7 +155,22 @@ tokens = list(lancedb.tokenize("acme makes searchable data",
custom_stop_words=["acme"])) custom_stop_words=["acme"]))
``` ```
::: lancedb.index.FTS ::: lancedb.tokenize
::: lancedb.FtsToken
## Blobs
Blob columns store large binary values out of line so they can be read lazily
instead of being materialized with the rest of the row.
::: lancedb.blob
::: lancedb.BlobType
::: lancedb._blob.BlobFile
options:
show_root_full_path: false
## Utilities ## Utilities
@@ -130,6 +178,14 @@ tokens = list(lancedb.tokenize("acme makes searchable data",
::: lancedb.merge.LanceMergeInsertBuilder ::: lancedb.merge.LanceMergeInsertBuilder
::: lancedb.otel.instrument_lancedb_metrics
## Exceptions
::: lancedb.exceptions.MissingValueError
::: lancedb.exceptions.MissingColumnError
## Integrations ## Integrations
## Pydantic ## Pydantic
@@ -138,19 +194,30 @@ tokens = list(lancedb.tokenize("acme makes searchable data",
::: lancedb.pydantic.vector ::: lancedb.pydantic.vector
::: lancedb.pydantic.Vector
::: lancedb.pydantic.MultiVector
::: lancedb.pydantic.LanceModel ::: lancedb.pydantic.LanceModel
## PyTorch
::: lancedb.streaming.StreamingDataset
::: lancedb.permutation.permutation_builder
::: lancedb.permutation.PermutationBuilder
::: lancedb.permutation.Permutation
::: lancedb.permutation.Transforms
## Reranking ## Reranking
::: lancedb.rerankers.linear_combination.LinearCombinationReranker ::: lancedb.rerankers
options:
::: lancedb.rerankers.cohere.CohereReranker show_root_heading: false
show_root_toc_entry: false
::: lancedb.rerankers.colbert.ColbertReranker
::: lancedb.rerankers.cross_encoder.CrossEncoderReranker
::: lancedb.rerankers.openai.OpenaiReranker
## Connections (Asynchronous) ## Connections (Asynchronous)
@@ -161,6 +228,12 @@ can be used to create, list, or open tables.
::: lancedb.db.AsyncConnection ::: lancedb.db.AsyncConnection
## Namespaces (Asynchronous)
::: lancedb.connect_namespace_async
::: lancedb.namespace.AsyncLanceNamespaceDBConnection
## Tables (Asynchronous) ## Tables (Asynchronous)
Table hold your actual data as a collection of records / rows. Table hold your actual data as a collection of records / rows.
@@ -169,32 +242,20 @@ Table hold your actual data as a collection of records / rows.
::: lancedb.table.AsyncTags ::: lancedb.table.AsyncTags
::: lancedb.table.AsyncBranches
## Indices (Asynchronous) ## Indices (Asynchronous)
Indices can be created on a table to speed up queries. This section Indices can be created on a table to speed up queries. This section
lists the indices that LanceDb supports. lists the indices that LanceDb supports.
::: lancedb.index.BTree ::: lancedb.index
options:
::: lancedb.index.Bitmap show_root_heading: false
show_root_toc_entry: false
::: lancedb.index.LabelList # `lang_mapping` is defined in the module rather than imported, so it is
# picked up despite not being in `__all__`. It is an internal lookup table.
::: lancedb.index.FTS filters: ["!^_", "!^lang_mapping$"]
::: lancedb.index.IvfPq
::: lancedb.index.HnswPq
::: lancedb.index.HnswSq
::: lancedb.index.IvfFlat
::: lancedb.index.IvfSq
::: lancedb.index.IvfRq
::: lancedb.index.HnswFlat
::: lancedb.table.IndexStatistics ::: lancedb.table.IndexStatistics
@@ -222,3 +283,7 @@ rows nearest to a query vector and can be created with the
::: lancedb.query.AsyncHybridQuery ::: lancedb.query.AsyncHybridQuery
options: options:
inherited_members: true inherited_members: true
::: lancedb.query.AsyncTakeQuery
options:
inherited_members: true
+1 -1
View File
@@ -28,7 +28,7 @@
<properties> <properties>
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding> <project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
<arrow.version>15.0.0</arrow.version> <arrow.version>15.0.0</arrow.version>
<lance-core.version>10.0.0-beta.5</lance-core.version> <lance-core.version>10.1.0-beta.1</lance-core.version>
<spotless.skip>false</spotless.skip> <spotless.skip>false</spotless.skip>
<spotless.version>2.30.0</spotless.version> <spotless.version>2.30.0</spotless.version>
<spotless.java.googlejavaformat.version>1.7</spotless.java.googlejavaformat.version> <spotless.java.googlejavaformat.version>1.7</spotless.java.googlejavaformat.version>
+1 -1
View File
@@ -1,7 +1,7 @@
# Contributing to LanceDB Typescript # Contributing to LanceDB Typescript
This document outlines the process for contributing to LanceDB Typescript. This document outlines the process for contributing to LanceDB Typescript.
For general contribution guidelines, see [CONTRIBUTING.md](../CONTRIBUTING.md). For general contribution guidelines, see [CONTRIBUTING.md](https://github.com/lancedb/lancedb/blob/main/CONTRIBUTING.md).
## Project layout ## Project layout
+93
View File
@@ -877,3 +877,96 @@ describe("remote connection", () => {
}); });
}); });
}); });
describe("remote connection jobs surface", () => {
it("lists, describes, cancels, and reads history", async () => {
const { tableFromArrays, tableToIPC } = await import("apache-arrow");
const eventsTable = tableFromArrays({ state: ["created", "succeeded"] });
const eventsBody = Buffer.from(tableToIPC(eventsTable, "stream"));
await withMockDatabase(
(req, res) => {
let body = "";
req.on("data", (chunk) => {
body += chunk;
});
req.on("end", () => {
const payload = body.length > 0 ? JSON.parse(body) : {};
if (req.url === "/v1/jobs/list") {
if (payload["page_token"] === undefined) {
res
.writeHead(200, { "Content-Type": "application/json" })
.end(
'{"jobs": [{"job_id": "job-1", "table": "t1", ' +
'"job_type": "create_index", "state": "in_progress", ' +
'"created_at_millis": 1000}], "page_token": "next"}',
);
} else {
res
.writeHead(200, { "Content-Type": "application/json" })
.end(
'{"jobs": [{"job_id": "job-2", "table": "t2", ' +
'"job_type": "create_index", "state": "succeeded", ' +
'"created_at_millis": 2000}]}',
);
}
} else if (req.url === "/v1/jobs/describe") {
if (payload["job_id"] !== "job-1") {
res.writeHead(404).end("no such job");
return;
}
res
.writeHead(200, { "Content-Type": "application/json" })
.end(
'{"job_id": "job-1", "job_type": "create_index", ' +
'"job_state": "FAILED", "creation_ms": 1000, ' +
'"spec": {"column": "vec"}, "failure": {"phase": "execute", ' +
'"message": "worker died", "retryable": true}}',
);
} else if (req.url === "/v1/jobs/cancel") {
if (payload["job_id"] !== "job-1") {
res.writeHead(404).end("no such job");
return;
}
res
.writeHead(200, { "Content-Type": "application/json" })
.end('{"job_id": "job-1"}');
} else if (req.url === "/v1/jobs/query_events") {
res
.writeHead(200, {
"Content-Type": "application/vnd.apache.arrow.stream",
})
.end(eventsBody);
} else {
res.writeHead(404).end();
}
});
},
async (db) => {
const jobs = await db.listJobs();
expect(jobs.map((job) => job.jobId)).toEqual(["job-1", "job-2"]);
expect(jobs[0].state).toEqual("running");
expect(jobs[1].state).toEqual("finished");
const description = await db.getJob("job-1");
expect(description?.state).toEqual("failed");
expect(JSON.parse(description?.specJson ?? "")).toEqual({
column: "vec",
});
expect(description?.failure?.message).toEqual("worker died");
expect(await db.getJob("missing")).toBeNull();
expect(await db.cancelJob("job-1")).toBe(true);
expect(await db.cancelJob("missing")).toBe(false);
const history = await db.jobHistory("job-1");
expect(history.numRows).toEqual(2);
const job = db.job("job-1");
expect(job.id).toEqual("job-1");
expect(await job.status()).toEqual("failed");
await expect(job.wait()).rejects.toThrow("worker died");
},
);
});
});
+5 -1
View File
@@ -851,7 +851,11 @@ describe("When creating an index", () => {
afterEach(() => tmpDir.removeCallback()); afterEach(() => tmpDir.removeCallback());
it("should create a vector index on vector columns", async () => { it("should create a vector index on vector columns", async () => {
await tbl.createIndex("vec"); const job = await tbl.createIndexAsync("vec");
expect(job.id).toBeNull();
await job.wait();
// Cancelling a job that already finished succeeds and does nothing.
await job.cancel();
// check index directory // check index directory
const indexDir = path.join(tmpDir.name, "test.lance", "_indices"); const indexDir = path.join(tmpDir.name, "test.lance", "_indices");
+62
View File
@@ -1,6 +1,7 @@
// SPDX-License-Identifier: Apache-2.0 // SPDX-License-Identifier: Apache-2.0
// SPDX-FileCopyrightText: Copyright The LanceDB Authors // SPDX-FileCopyrightText: Copyright The LanceDB Authors
import { tableFromIPC } from "apache-arrow";
import { import {
Data, Data,
SchemaLike, SchemaLike,
@@ -20,6 +21,9 @@ import type {
CreateNamespaceResponse, CreateNamespaceResponse,
DescribeNamespaceResponse, DescribeNamespaceResponse,
DropNamespaceResponse, DropNamespaceResponse,
Job,
JobDescription,
JobInfo,
ListNamespacesResponse, ListNamespacesResponse,
} from "./native"; } from "./native";
export type { export type {
@@ -436,6 +440,40 @@ export abstract class Connection {
newName: string, newName: string,
options?: RenameTableOptions, options?: RenameTableOptions,
): Promise<void>; ): Promise<void>;
/**
* A {@link Job} handle for a server-side job by id.
*
* The handle is constructed without a server round trip; an unknown id
* surfaces when the handle is used. Dropping the handle has no effect on
* the job itself.
*/
abstract job(jobId: string): Job;
/** List server-side jobs across the database's tables. */
abstract listJobs(): Promise<JobInfo[]>;
/**
* Describe a single server-side job by id.
*
* Resolves to `null` when the server has no such job.
*/
abstract getJob(jobId: string): Promise<JobDescription | null>;
/**
* Request cancellation of a server-side job by id.
*
* Resolves to true if the server accepted the cancellation, false if no
* such job exists. Cancelling an already-terminal job is a no-op success.
*/
abstract cancelJob(jobId: string): Promise<boolean>;
/**
* The lifecycle event history of a server-side job, as an Arrow table.
*
* Lists history across all jobs when `jobId` is omitted.
*/
abstract jobHistory(jobId?: string): Promise<ArrowTable>;
} }
/** @hideconstructor */ /** @hideconstructor */
@@ -722,6 +760,30 @@ export class LocalConnection extends Connection {
options?.newNamespacePath, options?.newNamespacePath,
); );
} }
job(jobId: string): Job {
return this.inner.job(jobId);
}
async listJobs(): Promise<JobInfo[]> {
return this.inner.listJobs();
}
async getJob(jobId: string): Promise<JobDescription | null> {
return this.inner.getJob(jobId);
}
async cancelJob(jobId: string): Promise<boolean> {
return this.inner.cancelJob(jobId);
}
async jobHistory(jobId?: string): Promise<ArrowTable> {
const buf = await this.inner.jobHistory(jobId);
if (buf.length === 0) {
return new ArrowTable();
}
return tableFromIPC(buf);
}
} }
/** /**
+7 -1
View File
@@ -85,7 +85,13 @@ export {
RenameTableOptions, RenameTableOptions,
} from "./connection"; } from "./connection";
export { Session } from "./native.js"; export {
Job,
JobDescription,
JobFailureInfo,
JobInfo,
Session,
} from "./native.js";
export { export {
ExecutableQuery, ExecutableQuery,
+28
View File
@@ -30,6 +30,7 @@ import {
DropColumnsResult, DropColumnsResult,
IndexConfig, IndexConfig,
IndexStatistics, IndexStatistics,
Job,
Branches as NativeBranches, Branches as NativeBranches,
OptimizeStats, OptimizeStats,
TableStatistics, TableStatistics,
@@ -358,6 +359,17 @@ export abstract class Table {
options?: Partial<IndexOptions>, options?: Partial<IndexOptions>,
): Promise<void>; ): Promise<void>;
/**
* Create an index, returning a handle to the indexing job.
*
* The job may already be complete when returned; callers must not assume
* the index exists until {@link Job.wait} resolves.
*/
abstract createIndexAsync(
column: string,
options?: Partial<IndexOptions>,
): Promise<Job>;
/** /**
* Drop an index from the table. * Drop an index from the table.
* *
@@ -940,6 +952,22 @@ export class LocalTable extends Table {
); );
} }
async createIndexAsync(
column: string,
options?: Partial<IndexOptions>,
): Promise<Job> {
// biome-ignore lint/suspicious/noExplicitAny: skip
const nativeIndex = (options?.config as any)?.inner;
return await this.inner.createIndexAsync(
nativeIndex,
column,
options?.replace,
options?.waitTimeoutSeconds,
options?.name,
options?.train,
);
}
async dropIndex(name: string): Promise<void> { async dropIndex(name: string): Promise<void> {
await this.inner.dropIndex(name); await this.inner.dropIndex(name);
} }
+63
View File
@@ -340,6 +340,69 @@ impl Connection {
self.get_inner()?.drop_all_tables(&ns).await.default_error() self.get_inner()?.drop_all_tables(&ns).await.default_error()
} }
/// A `Job` handle for a server-side job by id.
///
/// The handle is constructed without a server round trip; an unknown id
/// surfaces when the handle is used.
#[napi]
pub fn job(&self, job_id: String) -> napi::Result<crate::job::Job> {
let job = self.get_inner()?.job(job_id).default_error()?;
Ok(crate::job::Job::new(job))
}
/// List server-side jobs across the database's tables.
#[napi(catch_unwind)]
pub async fn list_jobs(&self) -> napi::Result<Vec<crate::job::JobInfo>> {
let jobs = self.get_inner()?.list_jobs().await.default_error()?;
Ok(jobs.into_iter().map(Into::into).collect())
}
/// Describe a single server-side job by id. `null` when the server has
/// no such job.
#[napi(catch_unwind)]
pub async fn get_job(
&self,
job_id: String,
) -> napi::Result<Option<crate::job::JobDescription>> {
let description = self.get_inner()?.get_job(&job_id).await.default_error()?;
Ok(description.map(Into::into))
}
/// Request cancellation of a server-side job by id. Returns true if the
/// server accepted the cancellation, false if no such job exists.
#[napi(catch_unwind)]
pub async fn cancel_job(&self, job_id: String) -> napi::Result<bool> {
self.get_inner()?.cancel_job(&job_id).await.default_error()
}
/// The lifecycle event history of a server-side job (all jobs when
/// `job_id` is null), as an Arrow IPC stream buffer. Empty when there is
/// no history.
#[napi(catch_unwind)]
pub async fn job_history(&self, job_id: Option<String>) -> napi::Result<Buffer> {
let batches = self
.get_inner()?
.job_history(job_id.as_deref())
.await
.default_error()?;
let Some(first) = batches.first() else {
return Ok(Buffer::from(Vec::<u8>::new()));
};
let mut out = Vec::new();
let mut writer = arrow_ipc::writer::StreamWriter::try_new(&mut out, &first.schema())
.map_err(|e| napi::Error::from_reason(e.to_string()))?;
for batch in &batches {
writer
.write(batch)
.map_err(|e| napi::Error::from_reason(e.to_string()))?;
}
writer
.finish()
.map_err(|e| napi::Error::from_reason(e.to_string()))?;
drop(writer);
Ok(Buffer::from(out))
}
#[napi(catch_unwind)] #[napi(catch_unwind)]
/// Describe a namespace and return its properties. /// Describe a namespace and return its properties.
pub async fn describe_namespace( pub async fn describe_namespace(
+123
View File
@@ -0,0 +1,123 @@
// SPDX-License-Identifier: Apache-2.0
// SPDX-FileCopyrightText: Copyright The LanceDB Authors
use std::sync::Arc;
use napi_derive::napi;
use crate::error::NapiErrorExt;
/// A handle to an operation that may still be running.
#[napi]
pub struct Job {
inner: Arc<lancedb::Job>,
}
impl Job {
pub(crate) fn new(inner: lancedb::Job) -> Self {
Self {
inner: Arc::new(inner),
}
}
}
#[napi]
impl Job {
/// Identifies the operation on the server that is running it. Operations
/// that run in this process have no server id. The value is opaque.
#[napi(getter)]
pub fn id(&self) -> Option<String> {
self.inner.id().map(str::to_string)
}
/// The operation's current lifecycle state: "running", "finished",
/// "failed", or "cancelled".
///
/// A point snapshot; unlike {@link Job.wait} it does not block or reject
/// on a terminal failure state. States a newer server reports that this
/// client version does not know pass through as-is.
#[napi(catch_unwind)]
pub async fn status(&self) -> napi::Result<String> {
self.inner.status().await.default_error()
}
/// Wait until the operation reaches a terminal state.
#[napi(catch_unwind)]
pub async fn wait(&self) -> napi::Result<()> {
self.inner.wait().await.default_error()
}
/// Request cancellation. Cancelling a finished operation is a no-op.
#[napi(catch_unwind)]
pub async fn cancel(&self) -> napi::Result<()> {
self.inner.cancel().await.default_error()
}
}
/// A row from `Connection.listJobs`: one server-side job.
#[napi(object)]
pub struct JobInfo {
/// The job id -- what `Connection.getJob` and `Connection.cancelJob`
/// accept.
pub job_id: String,
/// The table the job runs against, without URI or namespace.
pub table: String,
pub job_type: String,
/// Lifecycle state: "running", "finished", "failed", or "cancelled".
pub state: String,
/// When the job was created, in milliseconds since the epoch.
pub created_at_millis: i64,
}
impl From<lancedb::database::JobInfo> for JobInfo {
fn from(info: lancedb::database::JobInfo) -> Self {
Self {
job_id: info.job_id,
table: info.table,
job_type: info.job_type,
state: info.state,
created_at_millis: info.created_at_millis,
}
}
}
/// The server's account of why a job failed.
#[napi(object)]
pub struct JobFailureInfo {
pub phase: Option<String>,
pub message: Option<String>,
pub retryable: Option<bool>,
}
/// A described job from `Connection.getJob`.
#[napi(object)]
pub struct JobDescription {
pub job_id: String,
pub job_type: String,
/// Lifecycle state: "running", "finished", "failed", or "cancelled".
pub state: String,
/// When the job was created, in milliseconds since the epoch.
pub creation_ms: i64,
/// The job-type-specific specification as a JSON string, when present.
pub spec_json: Option<String>,
/// Why the job failed, when the job is failed and the server reports a
/// reason.
pub failure: Option<JobFailureInfo>,
}
impl From<lancedb::database::JobDescription> for JobDescription {
fn from(description: lancedb::database::JobDescription) -> Self {
Self {
job_id: description.job_id,
job_type: description.job_type,
state: description.state,
creation_ms: description.creation_ms,
spec_json: (!description.spec.is_null()).then(|| description.spec.to_string()),
failure: description.failure.map(|failure| JobFailureInfo {
phase: failure.phase,
message: failure.message,
retryable: failure.retryable,
}),
}
}
}
+1
View File
@@ -11,6 +11,7 @@ mod error;
mod header; mod header;
mod index; mod index;
mod iterator; mod iterator;
mod job;
pub mod merge; pub mod merge;
pub mod otel; pub mod otel;
pub mod permutation; pub mod permutation;
+39 -2
View File
@@ -168,6 +168,39 @@ impl Table {
builder.execute().await.default_error() builder.execute().await.default_error()
} }
#[napi(catch_unwind)]
pub async fn create_index_async(
&self,
index: Option<&Index>,
column: String,
replace: Option<bool>,
wait_timeout_s: Option<i64>,
name: Option<String>,
train: Option<bool>,
) -> napi::Result<crate::job::Job> {
let lancedb_index = if let Some(index) = index {
index.consume()?
} else {
lancedb::index::Index::Auto
};
let mut builder = self.inner_ref()?.create_index(&[column], lancedb_index);
if let Some(replace) = replace {
builder = builder.replace(replace);
}
if let Some(timeout) = wait_timeout_s {
builder =
builder.wait_timeout(std::time::Duration::from_secs(timeout.try_into().unwrap()));
}
if let Some(name) = name {
builder = builder.name(name);
}
if let Some(train) = train {
builder = builder.train(train);
}
let job = builder.execute_async().await.default_error()?;
Ok(crate::job::Job::new(job))
}
#[napi(catch_unwind)] #[napi(catch_unwind)]
pub async fn drop_index(&self, index_name: String) -> napi::Result<()> { pub async fn drop_index(&self, index_name: String) -> napi::Result<()> {
self.inner_ref()? self.inner_ref()?
@@ -306,7 +339,9 @@ impl Table {
let transforms = NewColumnTransform::SqlExpressions(transforms); let transforms = NewColumnTransform::SqlExpressions(transforms);
let res = self let res = self
.inner_ref()? .inner_ref()?
.add_columns(transforms, None) .add_columns()
.transform(transforms)
.execute()
.await .await
.default_error()?; .default_error()?;
Ok(res.into()) Ok(res.into())
@@ -323,7 +358,9 @@ impl Table {
let transforms = NewColumnTransform::AllNulls(schema); let transforms = NewColumnTransform::AllNulls(schema);
let res = self let res = self
.inner_ref()? .inner_ref()?
.add_columns(transforms, None) .add_columns()
.transform(transforms)
.execute()
.await .await
.default_error()?; .default_error()?;
Ok(res.into()) Ok(res.into())
+1
View File
@@ -140,6 +140,7 @@ include = [
"python/lancedb/remote/errors.py", "python/lancedb/remote/errors.py",
"python/lancedb/embeddings/__init__.py", "python/lancedb/embeddings/__init__.py",
"python/lancedb/_lancedb.pyi", "python/lancedb/_lancedb.pyi",
"python/typing_tests/table_search.py",
] ]
exclude = ["python/tests/"] exclude = ["python/tests/"]
pythonVersion = "3.13" pythonVersion = "3.13"
+3
View File
@@ -20,6 +20,7 @@ from .remote import ClientConfig
from .remote.db import RemoteDBConnection from .remote.db import RemoteDBConnection
from .expr import Expr, col, lit, func from .expr import Expr, col, lit, func
from .schema import blob, vector, BlobType from .schema import blob, vector, BlobType
from .job import AsyncJob, Job
from .table import AsyncTable, Table from .table import AsyncTable, Table
from .types import BaseTokenizerType from .types import BaseTokenizerType
from ._lancedb import Session from ._lancedb import Session
@@ -500,6 +501,7 @@ __all__ = [
"connect_namespace", "connect_namespace",
"connect_namespace_async", "connect_namespace_async",
"AsyncConnection", "AsyncConnection",
"AsyncJob",
"AsyncLanceNamespaceDBConnection", "AsyncLanceNamespaceDBConnection",
"AsyncTable", "AsyncTable",
"FtsToken", "FtsToken",
@@ -513,6 +515,7 @@ __all__ = [
"BlobType", "BlobType",
"vector", "vector",
"DBConnection", "DBConnection",
"Job",
"LanceDBConnection", "LanceDBConnection",
"LanceNamespaceDBConnection", "LanceNamespaceDBConnection",
"RemoteDBConnection", "RemoteDBConnection",
+6 -23
View File
@@ -14,14 +14,10 @@ import pyarrow as pa
from .expr import Expr from .expr import Expr
from .schema import blob_v2_column_paths from .schema import blob_v2_column_paths
from .types import BlobMode, QueryProjection, QueryProjectionSpec from .types import BlobMode, QueryProjection, QueryProjectionSpec
from .util import get_uri_scheme
if TYPE_CHECKING: if TYPE_CHECKING:
from _typeshed import WriteableBuffer from _typeshed import WriteableBuffer
from .remote.table import RemoteTable
from .table import AsyncTable, Table
BLOB_MODE_TO_HANDLING = { BLOB_MODE_TO_HANDLING = {
"lazy": "blobs_descriptions", "lazy": "blobs_descriptions",
"bytes": "all_binary", "bytes": "all_binary",
@@ -104,22 +100,6 @@ def validate_blob_mode(blob_mode: BlobMode) -> None:
raise ValueError(f"blob_mode must be one of {modes}, got {blob_mode!r}") raise ValueError(f"blob_mode must be one of {modes}, got {blob_mode!r}")
def supports_blob_auto_row_id(table: Table | AsyncTable | RemoteTable) -> bool:
"""Blob auto row-id applies to native tables, not LanceDB Cloud."""
from .remote.table import RemoteTable
if isinstance(table, RemoteTable):
return False
inner = getattr(table, "_inner", None)
if inner is not None:
uri = inner.database().uri
if isinstance(uri, str) and get_uri_scheme(uri) == "db":
return False
return True
def projection_includes_blob_column( def projection_includes_blob_column(
projection: QueryProjection, projection: QueryProjection,
blob_columns: Iterable[str], blob_columns: Iterable[str],
@@ -164,16 +144,14 @@ def v2_projection_needs_row_id(
def blob_auto_row_id_for_scan( def blob_auto_row_id_for_scan(
table: Table | AsyncTable | RemoteTable,
schema: pa.Schema, schema: pa.Schema,
projection: QueryProjection, projection: QueryProjection,
*, *,
with_row_id: bool | None, with_row_id: bool | None,
) -> bool: ) -> bool:
"""Auto row-id only applies when the caller said nothing about row ids."""
if with_row_id is not None: if with_row_id is not None:
return False return False
if not supports_blob_auto_row_id(table):
return False
return v2_projection_needs_row_id(schema, projection, with_row_id=False) return v2_projection_needs_row_id(schema, projection, with_row_id=False)
@@ -186,6 +164,11 @@ def finalize_blob_query_table(
) -> pa.Table: ) -> pa.Table:
if user_requested_row_id or not blob_auto_row_id: if user_requested_row_id or not blob_auto_row_id:
return tbl return tbl
if "_rowid" not in tbl.column_names:
# A backend that ignores the row-id request leaves nothing to stash. Hand
# back the projection as-is so fetch_blobs raises the error that names the
# ways to supply row ids, rather than failing here about a hidden column.
return tbl
return stash_auto_row_ids(tbl, blob_paths) return stash_auto_row_ids(tbl, blob_paths)
+70
View File
@@ -146,6 +146,13 @@ class Connection(object):
start_after: Optional[str], start_after: Optional[str],
limit: Optional[int], limit: Optional[int],
) -> list[str]: ... # Deprecated: Use list_tables instead ) -> list[str]: ... # Deprecated: Use list_tables instead
def job(self, job_id: str) -> Job: ...
async def list_jobs(self) -> List[JobInfo]: ...
async def get_job(self, job_id: str) -> Optional[JobDescription]: ...
async def cancel_job(self, job_id: str) -> bool: ...
async def job_history(
self, job_id: Optional[str] = None
) -> List[pa.RecordBatch]: ...
async def create_table( async def create_table(
self, self,
name: str, name: str,
@@ -209,6 +216,47 @@ class BlobFile:
def read_range(self, offset: int, length: int) -> bytes: ... def read_range(self, offset: int, length: int) -> bytes: ...
def read_up_to(self, length: int) -> bytes: ... def read_up_to(self, length: int) -> bytes: ...
class Job:
@property
def id(self) -> Optional[str]: ...
async def status(self) -> str: ...
async def wait(self) -> None: ...
async def cancel(self) -> None: ...
class JobInfo:
@property
def job_id(self) -> str: ...
@property
def table(self) -> str: ...
@property
def job_type(self) -> str: ...
@property
def state(self) -> str: ...
@property
def created_at_millis(self) -> int: ...
class JobFailureInfo:
@property
def phase(self) -> Optional[str]: ...
@property
def message(self) -> Optional[str]: ...
@property
def retryable(self) -> Optional[bool]: ...
class JobDescription:
@property
def job_id(self) -> str: ...
@property
def job_type(self) -> str: ...
@property
def state(self) -> str: ...
@property
def creation_ms(self) -> int: ...
@property
def spec_json(self) -> Optional[str]: ...
@property
def failure(self) -> Optional[JobFailureInfo]: ...
class Table: class Table:
def name(self) -> str: ... def name(self) -> str: ...
def __repr__(self) -> str: ... def __repr__(self) -> str: ...
@@ -248,6 +296,28 @@ class Table:
name: Optional[str], name: Optional[str],
train: Optional[bool], train: Optional[bool],
): ... ): ...
async def create_index_async(
self,
column: str,
index: Union[
IvfFlat,
IvfSq,
IvfPq,
HnswPq,
HnswSq,
HnswFlat,
BTree,
Bitmap,
LabelList,
Fm,
FTS,
],
replace: Optional[bool],
wait_timeout: Optional[object],
*,
name: Optional[str],
train: Optional[bool],
) -> Job: ...
async def list_versions(self) -> List[Dict[str, Any]]: ... async def list_versions(self) -> List[Dict[str, Any]]: ...
async def version(self) -> int: ... async def version(self) -> int: ...
async def checkout(self, version: Union[int, str]): ... async def checkout(self, version: Union[int, str]): ...
+168 -7
View File
@@ -45,6 +45,7 @@ from lance_namespace.errors import NamespaceNotEmptyError, TableNotFoundError
from . import __version__ from . import __version__
from ._lancedb import connect as lancedb_connect # type: ignore from ._lancedb import connect as lancedb_connect # type: ignore
from .job import AsyncJob, Job
from .table import ( from .table import (
AsyncTable, AsyncTable,
LanceTable, LanceTable,
@@ -63,6 +64,7 @@ if TYPE_CHECKING:
from .pydantic import LanceModel from .pydantic import LanceModel
from ._lancedb import Connection as LanceDbConnection from ._lancedb import Connection as LanceDbConnection
from ._lancedb import JobDescription, JobInfo
from .common import DATA, URI from .common import DATA, URI
from .embeddings import EmbeddingFunctionConfig from .embeddings import EmbeddingFunctionConfig
from ._lancedb import Session from ._lancedb import Session
@@ -178,6 +180,51 @@ class DBConnection(EnforceOverrides):
"Namespace operations are not supported for this connection type" "Namespace operations are not supported for this connection type"
) )
def namespace_exists(self, namespace_id: List[str]) -> bool:
"""Check if a namespace exists.
Parameters
----------
namespace_id: List[str]
The namespace identifier to check.
Returns
-------
bool
True if the namespace exists, False otherwise.
Raises
------
NotImplementedError
If the connection type does not support namespace operations.
"""
raise NotImplementedError(
"Namespace operations are not supported for this connection type"
)
def table_exists(self, table_id: List[str]) -> bool:
"""Check if a table exists.
Parameters
----------
table_id: List[str]
The table identifier to check (full path including namespace
segments and table name).
Returns
-------
bool
True if the table exists, False otherwise.
Raises
------
NotImplementedError
If the connection type does not support namespace operations.
"""
raise NotImplementedError(
"Namespace operations are not supported for this connection type"
)
def list_tables( def list_tables(
self, self,
namespace_path: Optional[List[str]] = None, namespace_path: Optional[List[str]] = None,
@@ -359,7 +406,7 @@ class DBConnection(EnforceOverrides):
Data is converted to Arrow before being written to disk. For maximum Data is converted to Arrow before being written to disk. For maximum
control over how data is saved, either provide the PyArrow schema to control over how data is saved, either provide the PyArrow schema to
convert to or else provide a [PyArrow Table](pyarrow.Table) directly. convert to or else provide a [PyArrow Table][pyarrow.Table] directly.
>>> import pyarrow as pa >>> import pyarrow as pa
>>> custom_schema = pa.schema([ >>> custom_schema = pa.schema([
@@ -563,6 +610,46 @@ class DBConnection(EnforceOverrides):
""" """
raise NotImplementedError("serialize is not supported for this connection type") raise NotImplementedError("serialize is not supported for this connection type")
def job(self, job_id: str) -> Job:
"""A [Job][lancedb.job.Job] handle for a server-side job by id.
The handle is constructed without a server round trip; an unknown id
surfaces when the handle is used. Dropping the handle has no effect
on the job itself.
"""
raise NotImplementedError("job is not supported for this connection type")
def list_jobs(self) -> List[JobInfo]:
"""List server-side jobs across the database's tables."""
raise NotImplementedError("list_jobs is not supported for this connection type")
def get_job(self, job_id: str) -> Optional[JobDescription]:
"""Describe a single server-side job by id.
Returns None when the server has no such job.
"""
raise NotImplementedError("get_job is not supported for this connection type")
def cancel_job(self, job_id: str) -> bool:
"""Request cancellation of a server-side job by id.
Returns True if the server accepted the cancellation, False if no
such job exists. Cancelling an already-terminal job is a no-op
success.
"""
raise NotImplementedError(
"cancel_job is not supported for this connection type"
)
def job_history(self, job_id: Optional[str] = None) -> List[pa.RecordBatch]:
"""The lifecycle event history of a server-side job, as Arrow batches.
Lists history across all jobs when `job_id` is None.
"""
raise NotImplementedError(
"job_history is not supported for this connection type"
)
class LanceDBConnection(DBConnection): class LanceDBConnection(DBConnection):
""" """
@@ -688,11 +775,7 @@ class LanceDBConnection(DBConnection):
return cls(None, _inner=inner) return cls(None, _inner=inner)
def __repr__(self) -> str: def __repr__(self) -> str:
val = f"{self.__class__.__name__}(uri={self._conn.uri!r}" return f"{self.__class__.__name__}(uri={self._conn.uri!r})"
if self.read_consistency_interval is not None:
val += f", read_consistency_interval={repr(self.read_consistency_interval)}"
val += ")"
return val
@override @override
def serialize(self) -> str: def serialize(self) -> str:
@@ -1129,6 +1212,47 @@ class LanceDBConnection(DBConnection):
) )
) )
@override
def job(self, job_id: str) -> Job:
"""A [Job][lancedb.job.Job] handle for a server-side job by id.
The handle is constructed without a server round trip; an unknown id
surfaces when the handle is used. Dropping the handle has no effect
on the job itself.
"""
return Job(self._conn.job(job_id))
@override
def list_jobs(self) -> List[JobInfo]:
"""List server-side jobs across the database's tables."""
return LOOP.run(self._conn.list_jobs())
@override
def get_job(self, job_id: str) -> Optional[JobDescription]:
"""Describe a single server-side job by id.
Returns None when the server has no such job.
"""
return LOOP.run(self._conn.get_job(job_id))
@override
def cancel_job(self, job_id: str) -> bool:
"""Request cancellation of a server-side job by id.
Returns True if the server accepted the cancellation, False if no
such job exists. Cancelling an already-terminal job is a no-op
success.
"""
return LOOP.run(self._conn.cancel_job(job_id))
@override
def job_history(self, job_id: Optional[str] = None) -> List[pa.RecordBatch]:
"""The lifecycle event history of a server-side job, as Arrow batches.
Lists history across all jobs when `job_id` is None.
"""
return LOOP.run(self._conn.job_history(job_id))
@override @override
def namespace_client(self) -> LanceNamespace: def namespace_client(self) -> LanceNamespace:
"""Get the equivalent namespace client for this connection. """Get the equivalent namespace client for this connection.
@@ -1529,7 +1653,7 @@ class AsyncConnection(object):
Data is converted to Arrow before being written to disk. For maximum Data is converted to Arrow before being written to disk. For maximum
control over how data is saved, either provide the PyArrow schema to control over how data is saved, either provide the PyArrow schema to
convert to or else provide a [PyArrow Table](pyarrow.Table) directly. convert to or else provide a [PyArrow Table][pyarrow.Table] directly.
>>> import pyarrow as pa >>> import pyarrow as pa
>>> custom_schema = pa.schema([ >>> custom_schema = pa.schema([
@@ -1838,6 +1962,43 @@ class AsyncConnection(object):
namespace_path = [] namespace_path = []
await self._inner.drop_all_tables(namespace_path=namespace_path) await self._inner.drop_all_tables(namespace_path=namespace_path)
def job(self, job_id: str) -> AsyncJob:
"""An [AsyncJob][lancedb.job.AsyncJob] handle for a server-side job
by id.
The handle is constructed without a server round trip; an unknown id
surfaces when the handle is used. Dropping the handle has no effect
on the job itself.
"""
return AsyncJob(self._inner.job(job_id))
async def list_jobs(self) -> List[JobInfo]:
"""List server-side jobs across the database's tables."""
return await self._inner.list_jobs()
async def get_job(self, job_id: str) -> Optional[JobDescription]:
"""Describe a single server-side job by id.
Returns None when the server has no such job.
"""
return await self._inner.get_job(job_id)
async def cancel_job(self, job_id: str) -> bool:
"""Request cancellation of a server-side job by id.
Returns True if the server accepted the cancellation, False if no
such job exists. Cancelling an already-terminal job is a no-op
success.
"""
return await self._inner.cancel_job(job_id)
async def job_history(self, job_id: Optional[str] = None) -> List[pa.RecordBatch]:
"""The lifecycle event history of a server-side job, as Arrow batches.
Lists history across all jobs when `job_id` is None.
"""
return await self._inner.job_history(job_id)
async def namespace_client(self) -> LanceNamespace: async def namespace_client(self) -> LanceNamespace:
"""Get the equivalent namespace client for this connection. """Get the equivalent namespace client for this connection.
@@ -21,3 +21,32 @@ from .watsonx import WatsonxEmbeddings
from .voyageai import VoyageAIEmbeddingFunction from .voyageai import VoyageAIEmbeddingFunction
from .colpali import ColPaliEmbeddings from .colpali import ColPaliEmbeddings
from .siglip import SigLipEmbeddings from .siglip import SigLipEmbeddings
# The API reference renders this package with a single mkdocstrings directive,
# which only picks up names listed here. New embedding functions must be added
# to both the imports above and this list, or they will silently go undocumented.
__all__ = [
"EmbeddingFunction",
"EmbeddingFunctionConfig",
"TextEmbeddingFunction",
"EmbeddingFunctionRegistry",
"get_registry",
"register",
"SentenceTransformerEmbeddings",
"OpenAIEmbeddings",
"OpenClipEmbeddings",
"BedRockText",
"CohereEmbeddingFunction",
"GeminiText",
"GteEmbeddings",
"InstructorEmbeddingFunction",
"JinaEmbeddings",
"OllamaEmbeddings",
"TransformersEmbeddingFunction",
"ColbertEmbeddings",
"VoyageAIEmbeddingFunction",
"WatsonxEmbeddings",
"ColPaliEmbeddings",
"ImageBindEmbeddings",
"SigLipEmbeddings",
]
+5 -5
View File
@@ -21,20 +21,20 @@ class BedRockText(TextEmbeddingFunction):
""" """
Parameters Parameters
---------- ----------
name: str, default "amazon.titan-embed-text-v1" name : str, default "amazon.titan-embed-text-v1"
The model ID of the bedrock model to use. Supported models for are: The model ID of the bedrock model to use. Supported models for are:
- amazon.titan-embed-text-v1 - amazon.titan-embed-text-v1
- cohere.embed-english-v3 - cohere.embed-english-v3
- cohere.embed-multilingual-v3 - cohere.embed-multilingual-v3
region: str, default "us-east-1" region : str, default "us-east-1"
Optional name of the AWS Region in which the service should be called. Optional name of the AWS Region in which the service should be called.
profile_name: str, default None profile_name : str, default None
Optional name of the AWS profile to use for calling the Bedrock service. Optional name of the AWS profile to use for calling the Bedrock service.
If not specified, the default profile will be used. If not specified, the default profile will be used.
assumed_role: str, default None assumed_role : str, default None
Optional ARN of an AWS IAM role to assume for calling the Bedrock service. Optional ARN of an AWS IAM role to assume for calling the Bedrock service.
If not specified, the current active credentials will be used. If not specified, the current active credentials will be used.
role_session_name: str, default "lancedb-embeddings" role_session_name : str, default "lancedb-embeddings"
Optional name of the AWS IAM role session to use for calling the Bedrock Optional name of the AWS IAM role session to use for calling the Bedrock
service. If not specified, "lancedb-embeddings" name will be used. service. If not specified, "lancedb-embeddings" name will be used.
+5 -3
View File
@@ -22,7 +22,7 @@ class CohereEmbeddingFunction(TextEmbeddingFunction):
Parameters Parameters
---------- ----------
name: str, default "embed-multilingual-v2.0" name : str, default "embed-multilingual-v2.0"
The name of the model to use. List of acceptable models: The name of the model to use. List of acceptable models:
* embed-english-v3.0 * embed-english-v3.0
@@ -33,12 +33,14 @@ class CohereEmbeddingFunction(TextEmbeddingFunction):
* embed-english-light-v2.0 * embed-english-light-v2.0
* embed-multilingual-v2.0 * embed-multilingual-v2.0
source_input_type: str, default "search_document" source_input_type : str, default "search_document"
The input type for the source column in the database The input type for the source column in the database
query_input_type: str, default "search_query" query_input_type : str, default "search_query"
The input type for the query column in the database The input type for the query column in the database
Notes
-----
Cohere supports following input types: Cohere supports following input types:
| Input Type | Description | | Input Type | Description |
+2 -2
View File
@@ -44,7 +44,7 @@ class ColPaliEmbeddings(EmbeddingFunction):
The token pooling strategy to use, by default "hierarchical". The token pooling strategy to use, by default "hierarchical".
- "hierarchical": Progressively pools tokens to reduce sequence length. - "hierarchical": Progressively pools tokens to reduce sequence length.
- "lambda": A simpler pooling that uses a custom `pooling_func`. - "lambda": A simpler pooling that uses a custom `pooling_func`.
pooling_func: typing.Callable, optional pooling_func : typing.Callable, optional
A function to use for pooling when `pooling_strategy` is "lambda". A function to use for pooling when `pooling_strategy` is "lambda".
pool_factor : int pool_factor : int
Factor to reduce sequence length if token pooling is enabled (default 2). Factor to reduce sequence length if token pooling is enabled (default 2).
@@ -52,7 +52,7 @@ class ColPaliEmbeddings(EmbeddingFunction):
Quantization configuration for the model. (default None, bitsandbytes needed) Quantization configuration for the model. (default None, bitsandbytes needed)
batch_size : int batch_size : int
Batch size for processing inputs (default 2). Batch size for processing inputs (default 2).
offload_folder: str, optional offload_folder : str, optional
Folder to offload model weights if using CPU offloading (default None). This is Folder to offload model weights if using CPU offloading (default None). This is
useful for large models that do not fit in memory. useful for large models that do not fit in memory.
""" """
@@ -48,16 +48,16 @@ class GeminiText(TextEmbeddingFunction):
Parameters Parameters
---------- ----------
name: str, default "gemini-embedding-001" name : str, default "gemini-embedding-001"
The name of the model to use. Supported models include: The name of the model to use. Supported models include:
- "gemini-embedding-001" (768 dimensions) - "gemini-embedding-001" (768 dimensions)
Note: The legacy "models/embedding-001" format is also supported but Note: The legacy "models/embedding-001" format is also supported but
"gemini-embedding-001" is recommended. "gemini-embedding-001" is recommended.
query_task_type: str, default "retrieval_query" query_task_type : str, default "retrieval_query"
Sets the task type for the queries. Sets the task type for the queries.
source_task_type: str, default "retrieval_document" source_task_type : str, default "retrieval_document"
Sets the task type for ingestion. Sets the task type for ingestion.
Examples Examples
+4 -4
View File
@@ -26,13 +26,13 @@ class GteEmbeddings(TextEmbeddingFunction):
Parameters Parameters
---------- ----------
name: str, default "thenlper/gte-large" name : str, default "thenlper/gte-large"
The name of the model to use. The name of the model to use.
device: str, default "cpu" device : str, default "cpu"
Sets the device type for the model. Sets the device type for the model.
normalize: str, default "True" normalize : str, default "True"
Controls normalize param in encode function for the transformer. Controls normalize param in encode function for the transformer.
mlx: bool, default False mlx : bool, default False
Controls which model to use. False for gte-large,True for the mlx version. Controls which model to use. False for gte-large,True for the mlx version.
Examples Examples
@@ -35,23 +35,23 @@ class InstructorEmbeddingFunction(TextEmbeddingFunction):
Parameters Parameters
---------- ----------
name: str name : str
The name of the model to use. Available models are listed at The name of the model to use. Available models are listed at
https://github.com/xlang-ai/instructor-embedding#model-list; https://github.com/xlang-ai/instructor-embedding#model-list;
The default model is hkunlp/instructor-base The default model is hkunlp/instructor-base
batch_size: int, default 32 batch_size : int, default 32
The batch size to use when generating embeddings The batch size to use when generating embeddings
device: str, default "cpu" device : str, default "cpu"
The device to use when generating embeddings The device to use when generating embeddings
show_progress_bar: bool, default True show_progress_bar : bool, default True
Whether to show a progress bar when generating embeddings Whether to show a progress bar when generating embeddings
normalize_embeddings: bool, default True normalize_embeddings : bool, default True
Whether to normalize the embeddings Whether to normalize the embeddings
quantize: bool, default False quantize : bool, default False
Whether to quantize the model Whether to quantize the model
source_instruction: str, default "represent the document for retrieval" source_instruction : str, default "represent the document for retrieval"
The instruction for the source column The instruction for the source column
query_instruction: str, default "represent the document for retrieving the most query_instruction : str, default "represent the document for retrieving the most
similar documents" similar documents"
The instruction for the query The instruction for the query
+2 -2
View File
@@ -40,10 +40,10 @@ class JinaEmbeddings(EmbeddingFunction):
Parameters Parameters
---------- ----------
name: str, default "jina-clip-v1". Note that some models support both image name : str, default "jina-clip-v1". Note that some models support both image
and text embeddings and some just text embedding and text embeddings and some just text embedding
api_key: str, default None api_key : str, default None
The api key to access Jina API. If you pass None, you can set JINA_API_KEY The api key to access Jina API. If you pass None, you can set JINA_API_KEY
environment variable environment variable
@@ -21,13 +21,13 @@ class SentenceTransformerEmbeddings(TextEmbeddingFunction):
Parameters Parameters
---------- ----------
name: str, default "all-MiniLM-L6-v2" name : str, default "all-MiniLM-L6-v2"
The name of the model to use. The name of the model to use.
device: str, default "cpu" device : str, default "cpu"
The device to use for the model The device to use for the model
normalize: bool, default True normalize : bool, default True
Whether to normalize the embeddings Whether to normalize the embeddings
trust_remote_code: bool, default True trust_remote_code : bool, default True
Whether to trust the remote code Whether to trust the remote code
""" """
+2 -2
View File
@@ -167,7 +167,7 @@ class VoyageAIEmbeddingFunction(EmbeddingFunction):
Parameters Parameters
---------- ----------
name: str name : str
The name of the model to use. List of acceptable models: The name of the model to use. List of acceptable models:
* voyage-4 (1024 dims, general-purpose and multilingual retrieval) * voyage-4 (1024 dims, general-purpose and multilingual retrieval)
@@ -185,7 +185,7 @@ class VoyageAIEmbeddingFunction(EmbeddingFunction):
* voyage-law-2 * voyage-law-2
* voyage-code-2 * voyage-code-2
output_dimension: int, optional output_dimension : int, optional
The output dimension for models that support flexible dimensions. The output dimension for models that support flexible dimensions.
Currently only voyage-multimodal-3.5 supports this feature. Currently only voyage-multimodal-3.5 supports this feature.
Valid options: 256, 512, 1024 (default), 2048. Valid options: 256, 512, 1024 (default), 2048.
+12
View File
@@ -23,3 +23,15 @@ class MissingColumnError(KeyError):
return ( return (
f"Error: Column '{self.column_name}' does not exist in the DataFrame object" f"Error: Column '{self.column_name}' does not exist in the DataFrame object"
) )
class JobFailedError(RuntimeError):
"""Exception raised when an asynchronous job reaches the failed state."""
pass
class JobCancelledError(RuntimeError):
"""Exception raised when an asynchronous job was cancelled."""
pass
+26 -23
View File
@@ -219,7 +219,7 @@ class HnswPq:
distance has a range of (-, ). If the vectors are normalized (i.e. their distance has a range of (-, ). If the vectors are normalized (i.e. their
l2 norm is 1), then dot distance is equivalent to the cosine distance. l2 norm is 1), then dot distance is equivalent to the cosine distance.
num_partitions, default sqrt(num_rows) num_partitions: int, default sqrt(num_rows)
The number of IVF partitions to create. The number of IVF partitions to create.
@@ -228,7 +228,7 @@ class HnswPq:
will require too much memory. Each partition becomes its own HNSW graph, so will require too much memory. Each partition becomes its own HNSW graph, so
setting this value higher reduces the peak memory use of training. setting this value higher reduces the peak memory use of training.
num_sub_vectors, default is vector dimension / 16 num_sub_vectors: int, default is vector dimension / 16
Number of sub-vectors of PQ. Number of sub-vectors of PQ.
@@ -244,13 +244,13 @@ class HnswPq:
If the dimension is not visible by 8 then we use 1 subvector. This is not If the dimension is not visible by 8 then we use 1 subvector. This is not
ideal and will likely result in poor performance. ideal and will likely result in poor performance.
num_bits: int, default 8 num_bits: int, default 8
Number of bits to encode each sub-vector. Number of bits to encode each sub-vector.
This value controls how much the sub-vectors are compressed. The more bits This value controls how much the sub-vectors are compressed. The more bits
the more accurate the index but the slower search. Only 4 and 8 are supported. the more accurate the index but the slower search. Only 4 and 8 are supported.
max_iterations, default 50 max_iterations: int, default 50
Max iterations to train kmeans. Max iterations to train kmeans.
@@ -263,7 +263,7 @@ class HnswPq:
those cases it is unlikely that setting this larger will lead to the index those cases it is unlikely that setting this larger will lead to the index
converging anyways. converging anyways.
sample_rate, default 256 sample_rate: int, default 256
The rate used to calculate the number of training vectors for kmeans. The rate used to calculate the number of training vectors for kmeans.
@@ -279,14 +279,14 @@ class HnswPq:
Increasing this value might improve the quality of the index but in Increasing this value might improve the quality of the index but in
most cases the default should be sufficient. most cases the default should be sufficient.
m, default 20 m: int, default 20
The number of neighbors to select for each vector in the HNSW graph. The number of neighbors to select for each vector in the HNSW graph.
This value controls the tradeoff between search speed and accuracy. This value controls the tradeoff between search speed and accuracy.
The higher the value the more accurate the search but the slower it will be. The higher the value the more accurate the search but the slower it will be.
ef_construction, default 300 ef_construction: int, default 300
The number of candidates to evaluate during the construction of the HNSW graph. The number of candidates to evaluate during the construction of the HNSW graph.
@@ -297,7 +297,7 @@ class HnswPq:
This value should be set to a value that is not less than `ef` in the This value should be set to a value that is not less than `ef` in the
search phase. search phase.
target_partition_size, default is 1,048,576 target_partition_size: int, default is 1,048,576
The target size of each partition. The target size of each partition.
@@ -351,7 +351,7 @@ class HnswSq:
distance has a range of (-, ). If the vectors are normalized (i.e. their distance has a range of (-, ). If the vectors are normalized (i.e. their
l2 norm is 1), then dot distance is equivalent to the cosine distance. l2 norm is 1), then dot distance is equivalent to the cosine distance.
num_partitions, default sqrt(num_rows) num_partitions: int, default sqrt(num_rows)
The number of IVF partitions to create. The number of IVF partitions to create.
@@ -360,7 +360,7 @@ class HnswSq:
will require too much memory. Each partition becomes its own HNSW graph, so will require too much memory. Each partition becomes its own HNSW graph, so
setting this value higher reduces the peak memory use of training. setting this value higher reduces the peak memory use of training.
max_iterations, default 50 max_iterations: int, default 50
Max iterations to train kmeans. Max iterations to train kmeans.
@@ -373,7 +373,7 @@ class HnswSq:
In those cases it is unlikely that setting this larger will lead to In those cases it is unlikely that setting this larger will lead to
the index converging anyways. the index converging anyways.
sample_rate, default 256 sample_rate: int, default 256
The rate used to calculate the number of training vectors for kmeans. The rate used to calculate the number of training vectors for kmeans.
@@ -389,14 +389,14 @@ class HnswSq:
Increasing this value might improve the quality of the index but in Increasing this value might improve the quality of the index but in
most cases the default should be sufficient. most cases the default should be sufficient.
m, default 20 m: int, default 20
The number of neighbors to select for each vector in the HNSW graph. The number of neighbors to select for each vector in the HNSW graph.
This value controls the tradeoff between search speed and accuracy. This value controls the tradeoff between search speed and accuracy.
The higher the value the more accurate the search but the slower it will be. The higher the value the more accurate the search but the slower it will be.
ef_construction, default 300 ef_construction: int, default 300
The number of candidates to evaluate during the construction of the HNSW graph. The number of candidates to evaluate during the construction of the HNSW graph.
@@ -407,7 +407,7 @@ class HnswSq:
This value should be set to a value that is not less than `ef` in the search This value should be set to a value that is not less than `ef` in the search
phase. phase.
target_partition_size, default is 1,048,576 target_partition_size: int, default is 1,048,576
The target size of each partition. The target size of each partition.
@@ -460,7 +460,7 @@ class HnswFlat:
distance has a range of (-, ). If the vectors are normalized (i.e. their distance has a range of (-, ). If the vectors are normalized (i.e. their
l2 norm is 1), then dot distance is equivalent to the cosine distance. l2 norm is 1), then dot distance is equivalent to the cosine distance.
num_partitions, default sqrt(num_rows) num_partitions: int, default sqrt(num_rows)
The number of IVF partitions to create. The number of IVF partitions to create.
@@ -470,18 +470,18 @@ class HnswFlat:
graph, so setting this value higher reduces the peak memory use of graph, so setting this value higher reduces the peak memory use of
training. training.
max_iterations, default 50 max_iterations: int, default 50
Max iterations to train kmeans. Max iterations to train kmeans.
When training an IVF index we use kmeans to calculate the partitions. When training an IVF index we use kmeans to calculate the partitions.
This parameter controls how many iterations of kmeans to run. This parameter controls how many iterations of kmeans to run.
sample_rate, default 256 sample_rate: int, default 256
The rate used to calculate the number of training vectors for kmeans. The rate used to calculate the number of training vectors for kmeans.
m, default 20 m: int, default 20
The number of neighbors to select for each vector in the HNSW graph. The number of neighbors to select for each vector in the HNSW graph.
@@ -489,7 +489,7 @@ class HnswFlat:
The higher the value the more accurate the search but the slower it The higher the value the more accurate the search but the slower it
will be. will be.
ef_construction, default 300 ef_construction: int, default 300
The number of candidates to evaluate during the construction of the HNSW The number of candidates to evaluate during the construction of the HNSW
graph. graph.
@@ -501,7 +501,7 @@ class HnswFlat:
than 500. This value should be set to a value that is not less than `ef` than 500. This value should be set to a value that is not less than `ef`
in the search phase. in the search phase.
target_partition_size, default is 1,048,576 target_partition_size: int, default is 1,048,576
The target size of each partition. The target size of each partition.
""" """
@@ -605,7 +605,7 @@ class IvfFlat:
The default value is 256. The default value is 256.
target_partition_size, default is 8192 target_partition_size: int, default is 8192
The target size of each partition. The target size of each partition.
@@ -769,7 +769,7 @@ class IvfPq:
The default value is 256. The default value is 256.
target_partition_size, default is 8192 target_partition_size: int, default is 8192
The target size of each partition. The target size of each partition.
@@ -830,7 +830,7 @@ class IvfRq:
sample_rate: int, default 256 sample_rate: int, default 256
Controls the number of training vectors: sample_rate * num_partitions. Controls the number of training vectors: sample_rate * num_partitions.
target_partition_size, default is 8192 target_partition_size: int, default is 8192
Target size of each partition. Target size of each partition.
""" """
@@ -845,6 +845,9 @@ class IvfRq:
accelerator: Optional[str] = None accelerator: Optional[str] = None
# The API reference renders this module with a single mkdocstrings directive,
# which only picks up names listed here. New public names must be added to this
# list, or they will silently go undocumented.
__all__ = [ __all__ = [
"BTree", "BTree",
"IvfPq", "IvfPq",
+105
View File
@@ -0,0 +1,105 @@
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright The LanceDB Authors
"""Handles to operations a server may run asynchronously."""
import asyncio
from datetime import timedelta
from typing import Optional
from lancedb.background_loop import LOOP
from . import _lancedb
class AsyncJob:
"""A handle to an operation that may still be running.
The operation may already be complete when the handle is created.
"""
def __init__(self, inner: Optional["_lancedb.Job"]):
self._inner = inner
@property
def id(self) -> Optional[str]:
"""Identifies the operation on the server that is running it.
Returned for correlating with server logs or the jobs API. Operations
that run in this process have no server id and return `None`. The value
is opaque: parsing it or storing it to resume the job later is not
supported.
"""
return self._inner.id if self._inner is not None else None
async def status(self) -> str:
"""The operation's current lifecycle state: "running", "finished",
"failed", or "cancelled".
A point snapshot; unlike `wait` it does not block or raise on a
terminal failure state. States a newer server reports that this
client version does not know pass through as-is.
"""
if self._inner is None:
return "finished"
return await self._inner.status()
async def wait(self, timeout: Optional[timedelta] = None):
"""Wait until the operation reaches a terminal state.
Raises `JobFailedError` if the operation failed, `JobCancelledError`
if it was cancelled, and `TimeoutError` if `timeout` elapses first.
"""
if self._inner is None:
return
if timeout is None:
await self._inner.wait()
else:
await asyncio.wait_for(self._inner.wait(), timeout.total_seconds())
async def cancel(self):
"""Request cancellation. Cancelling a finished operation is a no-op."""
if self._inner is None:
return
await self._inner.cancel()
class Job:
"""Synchronous counterpart of `AsyncJob`."""
def __init__(self, inner: Optional[AsyncJob]):
self._inner = inner
@property
def id(self) -> Optional[str]:
"""Identifies the operation on the server that is running it.
See :attr:`AsyncJob.id`.
"""
return self._inner.id if self._inner is not None else None
def status(self) -> str:
"""The operation's current lifecycle state: "running", "finished",
"failed", or "cancelled".
See :meth:`AsyncJob.status`.
"""
if self._inner is None:
return "finished"
return LOOP.run(self._inner.status())
def wait(self, timeout: Optional[timedelta] = None):
"""Block until the operation reaches a terminal state.
Raises `JobFailedError` if the operation failed, `JobCancelledError`
if it was cancelled, and `TimeoutError` if `timeout` elapses first.
"""
if self._inner is None:
return
LOOP.run(self._inner.wait(timeout))
def cancel(self):
"""Request cancellation. Cancelling a finished operation is a no-op."""
if self._inner is None:
return
LOOP.run(self._inner.cancel())
+3 -1
View File
@@ -92,8 +92,10 @@ class LanceMergeInsertBuilder(object):
self._when_not_matched_by_source_delete = True self._when_not_matched_by_source_delete = True
if isinstance(condition, Expr): if isinstance(condition, Expr):
self._when_not_matched_by_source_condition_expr = condition._inner self._when_not_matched_by_source_condition_expr = condition._inner
elif condition is not None: self._when_not_matched_by_source_condition = None
else:
self._when_not_matched_by_source_condition = condition self._when_not_matched_by_source_condition = condition
self._when_not_matched_by_source_condition_expr = None
return self return self
def use_index(self, use_index: bool) -> LanceMergeInsertBuilder: def use_index(self, use_index: bool) -> LanceMergeInsertBuilder:
+95 -1
View File
@@ -38,7 +38,11 @@ from lance_namespace_urllib3_client.models.query_table_request_vector import (
QueryTableRequestVector, QueryTableRequestVector,
) )
from lance_namespace_urllib3_client.models.string_fts_query import StringFtsQuery from lance_namespace_urllib3_client.models.string_fts_query import StringFtsQuery
from lance_namespace.errors import NamespaceNotEmptyError, TableNotFoundError from lance_namespace.errors import (
NamespaceNotEmptyError,
NamespaceNotFoundError,
TableNotFoundError,
)
from lancedb._lancedb import ( from lancedb._lancedb import (
connect_namespace as _connect_namespace, connect_namespace as _connect_namespace,
connect_namespace_client as _connect_namespace_client, connect_namespace_client as _connect_namespace_client,
@@ -53,6 +57,8 @@ from lance_namespace import (
DropNamespaceResponse, DropNamespaceResponse,
ListNamespacesResponse, ListNamespacesResponse,
ListTablesResponse, ListTablesResponse,
NamespaceExistsRequest,
TableExistsRequest,
) )
from lancedb.table import AsyncTable, LanceTable, Table from lancedb.table import AsyncTable, LanceTable, Table
from lancedb.util import validate_table_name from lancedb.util import validate_table_name
@@ -780,6 +786,51 @@ class LanceNamespaceDBConnection(DBConnection):
""" """
return LOOP.run(self._inner.describe_namespace(namespace_path)) return LOOP.run(self._inner.describe_namespace(namespace_path))
@override
def namespace_exists(self, namespace_id: List[str]) -> bool:
"""
Check if a namespace exists.
Parameters
----------
namespace_id : List[str]
The namespace identifier to check.
Returns
-------
bool
True if the namespace exists, False otherwise.
"""
request = NamespaceExistsRequest(id=namespace_id)
try:
self._namespace_client.namespace_exists(request)
return True
except NamespaceNotFoundError:
return False
@override
def table_exists(self, table_id: List[str]) -> bool:
"""
Check if a table exists.
Parameters
----------
table_id : List[str]
The table identifier to check (full path including namespace
segments and table name).
Returns
-------
bool
True if the table exists, False otherwise.
"""
request = TableExistsRequest(id=table_id)
try:
self._namespace_client.table_exists(request)
return True
except TableNotFoundError:
return False
@override @override
def list_tables( def list_tables(
self, self,
@@ -1233,6 +1284,49 @@ class AsyncLanceNamespaceDBConnection:
""" """
return await self._inner.describe_namespace(namespace_path) return await self._inner.describe_namespace(namespace_path)
async def namespace_exists(self, namespace_id: List[str]) -> bool:
"""
Check if a namespace exists.
Parameters
----------
namespace_id : List[str]
The namespace identifier to check.
Returns
-------
bool
True if the namespace exists, False otherwise.
"""
request = NamespaceExistsRequest(id=namespace_id)
try:
self._namespace_client.namespace_exists(request)
return True
except NamespaceNotFoundError:
return False
async def table_exists(self, table_id: List[str]) -> bool:
"""
Check if a table exists.
Parameters
----------
table_id : List[str]
The table identifier to check (full path including namespace
segments and table name).
Returns
-------
bool
True if the table exists, False otherwise.
"""
request = TableExistsRequest(id=table_id)
try:
self._namespace_client.table_exists(request)
return True
except TableNotFoundError:
return False
async def list_tables( async def list_tables(
self, self,
namespace_path: Optional[List[str]] = None, namespace_path: Optional[List[str]] = None,
+11 -6
View File
@@ -438,7 +438,8 @@ class Permutation:
_reader: Optional[PermutationReader] = None, _reader: Optional[PermutationReader] = None,
): ):
""" """
Internal constructor. Use [from_tables](#from_tables) instead. Internal constructor. Use
[from_tables][lancedb.permutation.Permutation.from_tables] instead.
""" """
assert base_table is not None, "base_table is required" assert base_table is not None, "base_table is required"
assert selection is not None, "selection is required" assert selection is not None, "selection is required"
@@ -985,8 +986,9 @@ class Permutation:
types. Conversion of strings, lists, and structs will require creating python types. Conversion of strings, lists, and structs will require creating python
objects and this is not zero-copy. objects and this is not zero-copy.
For custom formatting, use [with_transform](#with_transform) which overrides For custom formatting, use
this method. [with_transform][lancedb.permutation.Permutation.with_transform] which
overrides this method.
""" """
assert format is not None, "format is required" assert format is not None, "format is required"
if format == "python": if format == "python":
@@ -1061,7 +1063,8 @@ class Permutation:
Note: this method returns a new permutation and does not modify `self` Note: this method returns a new permutation and does not modify `self`
It is provided for compatibility with the huggingface Dataset API. It is provided for compatibility with the huggingface Dataset API.
Use [with_skip](#with_skip) instead to avoid confusion. Use [with_skip][lancedb.permutation.Permutation.with_skip] instead to
avoid confusion.
""" """
return self.with_skip(skip) return self.with_skip(skip)
@@ -1084,7 +1087,8 @@ class Permutation:
Note: this method returns a new permutation and does not modify `self` Note: this method returns a new permutation and does not modify `self`
It is provided for compatibility with the huggingface Dataset API. It is provided for compatibility with the huggingface Dataset API.
Use [with_take](#with_take) instead to avoid confusion. Use [with_take][lancedb.permutation.Permutation.with_take] instead to
avoid confusion.
""" """
return self.with_take(limit) return self.with_take(limit)
@@ -1107,7 +1111,8 @@ class Permutation:
Note: this method returns a new permutation and does not modify `self` Note: this method returns a new permutation and does not modify `self`
It is provided for compatibility with the huggingface Dataset API. It is provided for compatibility with the huggingface Dataset API.
Use [with_repeat](#with_repeat) instead to avoid confusion. Use [with_repeat][lancedb.permutation.Permutation.with_repeat] instead
to avoid confusion.
""" """
return self.with_repeat(times) return self.with_repeat(times)
+21 -23
View File
@@ -52,7 +52,6 @@ from ._blob import (
finalize_blob_query_table, finalize_blob_query_table,
replace_v2_blob_columns_with_bytes, replace_v2_blob_columns_with_bytes,
replace_v2_blob_columns_with_bytes_sync, replace_v2_blob_columns_with_bytes_sync,
supports_blob_auto_row_id,
validate_blob_mode, validate_blob_mode,
) )
from .types import BlobMode, QueryProjection from .types import BlobMode, QueryProjection
@@ -651,7 +650,8 @@ class Query(pydantic.BaseModel):
distance_type : Optional[str] distance_type : Optional[str]
the distance type to use for vector search the distance type to use for vector search
This can be l2 (default), cosine and dot. See [metric definitions][search] for This can be l2 (default), cosine and dot. See
[metric definitions](https://lancedb.com/docs/search/vector-search/) for
more details. more details.
If this is not a vector search this will be None. If this is not a vector search this will be None.
@@ -664,8 +664,9 @@ class Query(pydantic.BaseModel):
- A higher number makes search more accurate but also slower. - A higher number makes search more accurate but also slower.
- See discussion in [Querying an ANN Index][querying-an-ann-index] for - See discussion in
tuning advice. [Querying an ANN Index](https://lancedb.com/docs/indexing/)
for tuning advice.
Will be None if this is not a vector search. Will be None if this is not a vector search.
refine_factor : Optional[int] refine_factor : Optional[int]
@@ -673,8 +674,9 @@ class Query(pydantic.BaseModel):
- A higher number makes search more accurate but also slower. - A higher number makes search more accurate but also slower.
- See discussion in [Querying an ANN Index][querying-an-ann-index] for - See discussion in
tuning advice. [Querying an ANN Index](https://lancedb.com/docs/indexing/)
for tuning advice.
Will be None if this is not a vector search. Will be None if this is not a vector search.
lower_bound : Optional[float] lower_bound : Optional[float]
@@ -1277,10 +1279,7 @@ class LanceQueryBuilder(ABC):
return self._with_row_id is True return self._with_row_id is True
def _blob_auto_row_id_enabled(self) -> bool: def _blob_auto_row_id_enabled(self) -> bool:
if not supports_blob_auto_row_id(self._table):
return False
return blob_auto_row_id_for_scan( return blob_auto_row_id_for_scan(
self._table,
self._table.schema, self._table.schema,
self._columns, self._columns,
with_row_id=self._with_row_id, with_row_id=self._with_row_id,
@@ -1651,8 +1650,8 @@ class LanceVectorQueryBuilder(LanceQueryBuilder):
Higher values will yield better recall (more likely to find vectors if Higher values will yield better recall (more likely to find vectors if
they exist) at the expense of latency. they exist) at the expense of latency.
See discussion in [Querying an ANN Index][querying-an-ann-index] for See discussion in [Querying an ANN Index](https://lancedb.com/docs/indexing/)
tuning advice. for tuning advice.
This method sets both the minimum and maximum number of probes to the same This method sets both the minimum and maximum number of probes to the same
value. See `minimum_nprobes` and `maximum_nprobes` for more fine-grained value. See `minimum_nprobes` and `maximum_nprobes` for more fine-grained
@@ -1752,8 +1751,8 @@ class LanceVectorQueryBuilder(LanceQueryBuilder):
As an example, a refine factor of 2 will sample 2x as many vectors as As an example, a refine factor of 2 will sample 2x as many vectors as
requested, re-ranks them, and returns the top half most relevant results. requested, re-ranks them, and returns the top half most relevant results.
See discussion in [Querying an ANN Index][querying-an-ann-index] for See discussion in [Querying an ANN Index](https://lancedb.com/docs/indexing/)
tuning advice. for tuning advice.
Parameters Parameters
---------- ----------
@@ -2698,7 +2697,7 @@ class LanceHybridQueryBuilder(LanceQueryBuilder):
self._fts_query.phrase_query(True) self._fts_query.phrase_query(True)
if self._distance_type: if self._distance_type:
self._vector_query.metric(self._distance_type) self._vector_query.metric(self._distance_type)
if self._minimum_nprobes: if self._minimum_nprobes is not None:
self._vector_query.minimum_nprobes(self._minimum_nprobes) self._vector_query.minimum_nprobes(self._minimum_nprobes)
if self._maximum_nprobes is not None: if self._maximum_nprobes is not None:
self._vector_query.maximum_nprobes(self._maximum_nprobes) self._vector_query.maximum_nprobes(self._maximum_nprobes)
@@ -2771,7 +2770,7 @@ class AsyncQueryBase(object):
) )
async def _maybe_add_blob_row_id(self) -> None: async def _maybe_add_blob_row_id(self) -> None:
if self._table is None or not supports_blob_auto_row_id(self._table): if self._table is None:
self._blob_auto_row_id = False self._blob_auto_row_id = False
self._blob_paths = () self._blob_paths = ()
return return
@@ -2779,7 +2778,6 @@ class AsyncQueryBase(object):
req = self._inner.to_query_request() req = self._inner.to_query_request()
schema = await self._table.schema() schema = await self._table.schema()
self._blob_auto_row_id = blob_auto_row_id_for_scan( self._blob_auto_row_id = blob_auto_row_id_for_scan(
self._table,
schema, schema,
req.select, req.select,
with_row_id=self._with_row_id, with_row_id=self._with_row_id,
@@ -3031,7 +3029,6 @@ class AsyncQueryBase(object):
schema = await self._table.schema() schema = await self._table.schema()
blob_auto_row_id = blob_auto_row_id_for_scan( blob_auto_row_id = blob_auto_row_id_for_scan(
self._table,
schema, schema,
query.columns, query.columns,
with_row_id=self._with_row_id, with_row_id=self._with_row_id,
@@ -3379,8 +3376,9 @@ class AsyncQuery(AsyncStandardQuery):
are various ANN search parameters that will let you fine tune your recall are various ANN search parameters that will let you fine tune your recall
accuracy vs search latency. accuracy vs search latency.
Vector searches always have a [limit][]. If `limit` has not been called then Vector searches always have a
a default `limit` of 10 will be used. [limit][lancedb.query.AsyncVectorQuery.limit]. If `limit` has not been
called then a default `limit` of 10 will be used.
Typically, a single vector is passed in as the query. However, you can also Typically, a single vector is passed in as the query. However, you can also
pass in multiple vectors. When multiple vectors are passed in, if the vector pass in multiple vectors. When multiple vectors are passed in, if the vector
@@ -3511,8 +3509,9 @@ class AsyncFTSQuery(AsyncStandardQuery):
are various ANN search parameters that will let you fine tune your recall are various ANN search parameters that will let you fine tune your recall
accuracy vs search latency. accuracy vs search latency.
Hybrid searches always have a [limit][]. If `limit` has not been called then Hybrid searches always have a
a default `limit` of 10 will be used. [limit][lancedb.query.AsyncHybridQuery.limit]. If `limit` has not been
called then a default `limit` of 10 will be used.
Typically, a single vector is passed in as the query. However, you can also Typically, a single vector is passed in as the query. However, you can also
pass in multiple vectors. This can be useful if you want to find the nearest pass in multiple vectors. This can be useful if you want to find the nearest
@@ -3875,10 +3874,9 @@ class AsyncHybridQuery(AsyncStandardQuery, AsyncVectorQueryBase):
req = fts_query._inner.to_query_request() req = fts_query._inner.to_query_request()
blob_auto_row_id = False blob_auto_row_id = False
blob_paths: tuple[str, ...] = () blob_paths: tuple[str, ...] = ()
if self._table is not None and supports_blob_auto_row_id(self._table): if self._table is not None:
schema = await self._table.schema() schema = await self._table.schema()
blob_auto_row_id = blob_auto_row_id_for_scan( blob_auto_row_id = blob_auto_row_id_for_scan(
self._table,
schema, schema,
req.select, req.select,
with_row_id=self._with_row_id, with_row_id=self._with_row_id,
+3
View File
@@ -11,6 +11,9 @@ from lancedb import __version__
from .header import HeaderProvider from .header import HeaderProvider
from .oauth import OAuthConfig, OAuthFlowType from .oauth import OAuthConfig, OAuthFlowType
# The API reference renders this module with a single mkdocstrings directive,
# which only picks up names listed here. New public names must be added to this
# list, or they will silently go undocumented.
__all__ = [ __all__ = [
"TimeoutConfig", "TimeoutConfig",
"RetryConfig", "RetryConfig",
+51 -1
View File
@@ -7,7 +7,7 @@ import json
import logging import logging
from concurrent.futures import ThreadPoolExecutor from concurrent.futures import ThreadPoolExecutor
import sys import sys
from typing import Any, Dict, Iterable, List, Optional, Union from typing import TYPE_CHECKING, Any, Dict, Iterable, List, Optional, Union
from urllib.parse import urlparse from urllib.parse import urlparse
import warnings import warnings
@@ -23,6 +23,10 @@ import pyarrow as pa
from ..common import DATA from ..common import DATA
from ..db import DBConnection, LOOP from ..db import DBConnection, LOOP
from ..job import Job
if TYPE_CHECKING:
from .._lancedb import JobDescription, JobInfo
from ..embeddings import EmbeddingFunctionConfig from ..embeddings import EmbeddingFunctionConfig
from lance_namespace import ( from lance_namespace import (
LanceNamespace, LanceNamespace,
@@ -415,6 +419,11 @@ class RemoteDBConnection(DBConnection):
if namespace_path is None: if namespace_path is None:
namespace_path = [] namespace_path = []
if storage_options is not None:
logging.info(
"storage_options is ignored in LanceDb Cloud"
" (storage is managed; set storage_options on connect() instead)"
)
if index_cache_size is not None: if index_cache_size is not None:
logging.info( logging.info(
"index_cache_size is ignored in LanceDb Cloud" "index_cache_size is ignored in LanceDb Cloud"
@@ -684,6 +693,47 @@ class RemoteDBConnection(DBConnection):
) )
) )
@override
def job(self, job_id: str) -> Job:
"""A [Job][lancedb.job.Job] handle for a server-side job by id.
The handle is constructed without a server round trip; an unknown id
surfaces when the handle is used. Dropping the handle has no effect
on the job itself.
"""
return Job(self._conn.job(job_id))
@override
def list_jobs(self) -> List["JobInfo"]:
"""List server-side jobs across the database's tables."""
return LOOP.run(self._conn.list_jobs())
@override
def get_job(self, job_id: str) -> Optional["JobDescription"]:
"""Describe a single server-side job by id.
Returns None when the server has no such job.
"""
return LOOP.run(self._conn.get_job(job_id))
@override
def cancel_job(self, job_id: str) -> bool:
"""Request cancellation of a server-side job by id.
Returns True if the server accepted the cancellation, False if no
such job exists. Cancelling an already-terminal job is a no-op
success.
"""
return LOOP.run(self._conn.cancel_job(job_id))
@override
def job_history(self, job_id: Optional[str] = None) -> List[pa.RecordBatch]:
"""The lifecycle event history of a server-side job, as Arrow batches.
Lists history across all jobs when `job_id` is None.
"""
return LOOP.run(self._conn.job_history(job_id))
@override @override
def namespace_client(self) -> LanceNamespace: def namespace_client(self) -> LanceNamespace:
"""Get the equivalent namespace client for this connection. """Get the equivalent namespace client for this connection.
+2 -2
View File
@@ -53,9 +53,9 @@ class RetryError(LanceDBClientError):
"""An error that occurs when the client has exceeded the maximum number of retries. """An error that occurs when the client has exceeded the maximum number of retries.
The retry strategy can be adjusted by setting the The retry strategy can be adjusted by setting the
[retry_config](lancedb.remote.ClientConfig.retry_config) in the client [retry_config][lancedb.remote.ClientConfig.retry_config] in the client
configuration. This is passed in the `client_config` argument of configuration. This is passed in the `client_config` argument of
[connect](lancedb.connect) and [connect_async](lancedb.connect_async). [connect][lancedb.connect] and [connect_async][lancedb.connect_async].
The __cause__ attribute of this exception will be the last exception that The __cause__ attribute of this exception will be the last exception that
caused the retry to fail. It will be an caused the retry to fail. It will be an
+44 -12
View File
@@ -20,6 +20,7 @@ from typing import (
import warnings import warnings
from lancedb import __version__ from lancedb import __version__
from lancedb._blob import BlobFile
from lancedb._lancedb import ( from lancedb._lancedb import (
AddColumnsResult, AddColumnsResult,
@@ -47,6 +48,7 @@ from lancedb.index import (
IvfSq, IvfSq,
LabelList, LabelList,
) )
from lancedb.job import Job
from lancedb.remote.db import LOOP from lancedb.remote.db import LOOP
from lancedb.table import IndexConfigType, KNOWN_METRICS from lancedb.table import IndexConfigType, KNOWN_METRICS
import pyarrow as pa import pyarrow as pa
@@ -540,6 +542,34 @@ class RemoteTable(Table):
) )
) )
def create_index_async(
self,
column: str,
*,
config: IndexConfigType,
replace: Optional[bool] = None,
wait_timeout: Optional[timedelta] = None,
name: Optional[str] = None,
train: bool = True,
) -> Job:
"""Create an index, returning a handle to the indexing job.
The job may already be complete when returned; callers must not assume
the index exists until :meth:`Job.wait` returns.
"""
return Job(
LOOP.run(
self._table.create_index_async(
column,
replace=replace,
config=config,
wait_timeout=wait_timeout,
name=name,
train=train,
)
)
)
def _is_legacy_create_index_call( def _is_legacy_create_index_call(
self, self,
first_arg: str, first_arg: str,
@@ -580,8 +610,9 @@ class RemoteTable(Table):
progress: Optional[Union[bool, Callable, Any]] = None, progress: Optional[Union[bool, Callable, Any]] = None,
write_parallelism: Optional[int] = None, write_parallelism: Optional[int] = None,
) -> AddResult: ) -> AddResult:
"""Add more data to the [Table](Table). It has the same API signature as """Add more data to the [Table][lancedb.table.Table].
the OSS version.
It has the same API signature as the OSS version.
Parameters Parameters
---------- ----------
@@ -641,7 +672,8 @@ class RemoteTable(Table):
fast_search: bool = False, fast_search: bool = False,
) -> LanceVectorQueryBuilder: ) -> LanceVectorQueryBuilder:
"""Create a search query to find the nearest neighbors """Create a search query to find the nearest neighbors
of the given query vector. We currently support [vector search][search] of the given query vector. We currently support
[vector search](https://lancedb.com/docs/search/vector-search/)
All query options are defined in All query options are defined in
[LanceVectorQueryBuilder][lancedb.query.LanceVectorQueryBuilder]. [LanceVectorQueryBuilder][lancedb.query.LanceVectorQueryBuilder].
@@ -1037,22 +1069,22 @@ class RemoteTable(Table):
) )
def blob_columns(self) -> list[str]: def blob_columns(self) -> list[str]:
raise NotImplementedError( return LOOP.run(self._table.blob_columns())
"blob_columns() is not yet supported on the LanceDB Cloud"
)
def fetch_blobs(self, column: str, row_ids) -> pa.LargeBinaryArray: def fetch_blobs(
raise NotImplementedError("fetch_blobs() is not supported on LanceDB Cloud") self, column: str, row_ids: Union[list[int], pa.Table]
) -> pa.LargeBinaryArray:
return LOOP.run(self._table.fetch_blobs(column, row_ids))
def fetch_blob_ranges(self, column: str, requests) -> pa.LargeBinaryArray: def fetch_blob_ranges(self, column: str, requests) -> pa.LargeBinaryArray:
raise NotImplementedError( raise NotImplementedError(
"fetch_blob_ranges() is not supported on LanceDB Cloud" "fetch_blob_ranges() is not supported on LanceDB Cloud"
) )
def fetch_blob_files(self, column: str, row_ids): def fetch_blob_files(
raise NotImplementedError( self, column: str, row_ids: Union[list[int], pa.Table]
"fetch_blob_files() is not supported on LanceDB Cloud" ) -> "list[Optional[BlobFile]]":
) return LOOP.run(self._table.fetch_blob_files(column, row_ids))
def head(self, n=5) -> pa.Table: def head(self, n=5) -> pa.Table:
""" """
@@ -14,6 +14,9 @@ from .answerdotai import AnswerdotaiRerankers
from .voyageai import VoyageAIReranker from .voyageai import VoyageAIReranker
from .watsonx import WatsonxReranker from .watsonx import WatsonxReranker
# The API reference renders this module with a single mkdocstrings directive,
# which only picks up names listed here. New public names must be added to this
# list, or they will silently go undocumented.
__all__ = [ __all__ = [
"Reranker", "Reranker",
"CrossEncoderReranker", "CrossEncoderReranker",
+252 -34
View File
@@ -40,6 +40,7 @@ from ._blob import (
from .types import BlobMode from .types import BlobMode
from lancedb.arrow import peek_reader from lancedb.arrow import peek_reader
from lancedb.background_loop import LOOP, embedding_executor from lancedb.background_loop import LOOP, embedding_executor
from lancedb.job import AsyncJob, Job
from .dependencies import ( from .dependencies import (
_check_for_hugging_face, _check_for_hugging_face,
_check_for_lance, _check_for_lance,
@@ -977,6 +978,24 @@ class Table(ABC):
""" """
raise NotImplementedError raise NotImplementedError
def create_index_async(
self,
column: str,
*,
config: IndexConfigType,
replace: Optional[bool] = None,
wait_timeout: Optional[timedelta] = None,
name: Optional[str] = None,
train: bool = True,
) -> Job:
"""Create an index, returning a handle to the indexing job.
Takes the same arguments as :meth:`create_index`. The job may already
be complete when returned; callers must not assume the index exists
until :meth:`Job.wait` returns.
"""
raise NotImplementedError
def drop_index(self, name: str) -> None: def drop_index(self, name: str) -> None:
""" """
Drop an index from the table. Drop an index from the table.
@@ -1211,7 +1230,7 @@ class Table(ABC):
progress: Optional[Union[bool, Callable, Any]] = None, progress: Optional[Union[bool, Callable, Any]] = None,
write_parallelism: Optional[int] = None, write_parallelism: Optional[int] = None,
) -> AddResult: ) -> AddResult:
"""Add more data to the [Table](Table). """Add more data to the [Table][lancedb.table.Table].
Parameters Parameters
---------- ----------
@@ -1331,6 +1350,91 @@ class Table(ABC):
return LanceMergeInsertBuilder(self, on) return LanceMergeInsertBuilder(self, on)
@overload
def search(
self,
query: None = None,
vector_column_name: Optional[str] = None,
query_type: Literal["auto", "vector", "fts"] = "auto",
ordering_field_name: Optional[str] = None,
fts_columns: Optional[Union[str, List[str]]] = None,
) -> LanceEmptyQueryBuilder: ...
@overload
def search(
self,
query: str,
vector_column_name: Optional[str] = None,
query_type: Literal["auto"] = "auto",
ordering_field_name: Optional[str] = None,
fts_columns: Optional[Union[str, List[str]]] = None,
) -> Union[LanceFtsQueryBuilder, LanceVectorQueryBuilder]: ...
@overload
def search(
self,
query: FullTextQuery,
vector_column_name: Optional[str] = None,
query_type: QueryType = "auto",
ordering_field_name: Optional[str] = None,
fts_columns: Optional[Union[str, List[str]]] = None,
) -> LanceFtsQueryBuilder: ...
@overload
def search(
self,
query: Union[VEC, "PIL.Image.Image", Tuple],
vector_column_name: Optional[str] = None,
query_type: Literal["auto"] = "auto",
ordering_field_name: Optional[str] = None,
fts_columns: Optional[Union[str, List[str]]] = None,
) -> LanceVectorQueryBuilder: ...
@overload
def search(
self,
query: Optional[Union[VEC, str, "PIL.Image.Image", Tuple]] = None,
vector_column_name: Optional[str] = None,
query_type: Literal["vector"] = "vector",
ordering_field_name: Optional[str] = None,
fts_columns: Optional[Union[str, List[str]]] = None,
) -> LanceVectorQueryBuilder: ...
@overload
def search(
self,
query: Optional[Union[str, FullTextQuery]] = None,
vector_column_name: Optional[str] = None,
query_type: Literal["fts"] = "fts",
ordering_field_name: Optional[str] = None,
fts_columns: Optional[Union[str, List[str]]] = None,
) -> LanceFtsQueryBuilder: ...
@overload
def search(
self,
query: Optional[Union[VEC, str, "PIL.Image.Image", Tuple]] = None,
vector_column_name: Optional[str] = None,
query_type: Literal["hybrid"] = "hybrid",
ordering_field_name: Optional[str] = None,
fts_columns: Optional[Union[str, List[str]]] = None,
) -> LanceHybridQueryBuilder: ...
@overload
def search(
self,
query: Optional[Union[VEC, str, "PIL.Image.Image", Tuple]] = None,
vector_column_name: Optional[str] = None,
query_type: QueryType = "auto",
ordering_field_name: Optional[str] = None,
fts_columns: Optional[Union[str, List[str]]] = None,
) -> Union[
LanceEmptyQueryBuilder,
LanceFtsQueryBuilder,
LanceHybridQueryBuilder,
LanceVectorQueryBuilder,
]: ...
@abstractmethod @abstractmethod
def search( def search(
self, self,
@@ -1343,8 +1447,8 @@ class Table(ABC):
fts_columns: Optional[Union[str, List[str]]] = None, fts_columns: Optional[Union[str, List[str]]] = None,
) -> LanceQueryBuilder: ) -> LanceQueryBuilder:
"""Create a search query to find the nearest neighbors """Create a search query to find the nearest neighbors
of the given query vector. We currently support [vector search][search] of the given query vector. We currently support [vector search](https://lancedb.com/docs/search/vector-search/)
and [full-text search][experimental-full-text-search]. and [full-text search](https://lancedb.com/docs/search/full-text-search/).
All query options are defined in All query options are defined in
[LanceQueryBuilder][lancedb.query.LanceQueryBuilder]. [LanceQueryBuilder][lancedb.query.LanceQueryBuilder].
@@ -1574,8 +1678,10 @@ class Table(ABC):
"""Open lazy, seekable :class:`~lancedb._blob.BlobFile` handles. """Open lazy, seekable :class:`~lancedb._blob.BlobFile` handles.
Prefer this over :meth:`fetch_blobs` for large payloads. ``row_ids`` is Prefer this over :meth:`fetch_blobs` for large payloads. ``row_ids`` is
a ``list[int]`` or query ``pyarrow.Table`` with ``_rowid`` (or stashed a ``list[int]`` or a query ``pyarrow.Table`` carrying row identity via
row-id metadata). Null rows are ``None``. Local tables only. ``_rowid`` or a ``_lance_row_id`` field on the blob descriptor. Null
rows are ``None``. Remote tables require LanceDB Cloud server 0.5.0 or
newer.
""" """
@abstractmethod @abstractmethod
@@ -1778,7 +1884,7 @@ class Table(ABC):
for faster reads. for faster reads.
Arguments are passed onto Lance's Arguments are passed onto Lance's
[compact_files][lance.dataset.DatasetOptimizer.compact_files]. `lance.dataset.DatasetOptimizer.compact_files`.
For most cases, the default should be fine. For most cases, the default should be fine.
See Also See Also
@@ -1832,6 +1938,8 @@ class Table(ABC):
retrain: bool, default False retrain: bool, default False
This parameter is no longer used and is deprecated. This parameter is no longer used and is deprecated.
Notes
-----
The frequency an application should call optimize is based on the frequency of The frequency an application should call optimize is based on the frequency of
data modifications. If data is frequently added, deleted, or updated then data modifications. If data is frequently added, deleted, or updated then
optimize should be run frequently. A good rule of thumb is to run optimize if optimize should be run frequently. A good rule of thumb is to run optimize if
@@ -1986,15 +2094,14 @@ class Table(ABC):
change permanent you can use the `[Self::restore]` method. change permanent you can use the `[Self::restore]` method.
Any operation that modifies the table will fail while the table is in a checked Any operation that modifies the table will fail while the table is in a checked
out state. out state. To return the table to a normal state use
`[Self::checkout_latest]`.
Parameters Parameters
---------- ----------
version: int | str, version: int | str,
The version to check out. A version number (`int`) or a tag The version to check out. A version number (`int`) or a tag
(`str`) can be provided. (`str`) can be provided.
To return the table to a normal state use `[Self::checkout_latest]`
""" """
@abstractmethod @abstractmethod
@@ -2468,13 +2575,7 @@ class LanceTable(Table):
return LOOP.run(self._table.count_rows(filter)) return LOOP.run(self._table.count_rows(filter))
def __repr__(self) -> str: def __repr__(self) -> str:
val = f"{self.__class__.__name__}(name={self.name!r}" return f"{self.__class__.__name__}(name={self.name!r}, _conn={self._conn!r})"
if self._conn.read_consistency_interval is not None:
val += ", read_consistency_interval={!r}".format(
self._conn.read_consistency_interval
)
val += f", _conn={self._conn!r})"
return val
def __str__(self) -> str: def __str__(self) -> str:
return self.__repr__() return self.__repr__()
@@ -2783,6 +2884,34 @@ class LanceTable(Table):
) )
) )
def create_index_async(
self,
column: str,
*,
config: IndexConfigType,
replace: Optional[bool] = None,
wait_timeout: Optional[timedelta] = None,
name: Optional[str] = None,
train: bool = True,
) -> Job:
"""Create an index, returning a handle to the indexing job.
The job may already be complete when returned; callers must not assume
the index exists until :meth:`Job.wait` returns.
"""
return Job(
LOOP.run(
self._table.create_index_async(
column,
replace=replace,
config=config,
wait_timeout=wait_timeout,
name=name,
train=train,
)
)
)
def _is_legacy_create_index_call( def _is_legacy_create_index_call(
self, self,
first_arg: str, first_arg: str,
@@ -3335,7 +3464,47 @@ class LanceTable(Table):
) )
@overload @overload
def search( # type: ignore def search(
self,
query: None = None,
vector_column_name: Optional[str] = None,
query_type: Literal["auto", "vector", "fts"] = "auto",
ordering_field_name: Optional[str] = None,
fts_columns: Optional[Union[str, List[str]]] = None,
) -> LanceEmptyQueryBuilder: ...
@overload
def search(
self,
query: str,
vector_column_name: Optional[str] = None,
query_type: Literal["auto"] = "auto",
ordering_field_name: Optional[str] = None,
fts_columns: Optional[Union[str, List[str]]] = None,
) -> Union[LanceFtsQueryBuilder, LanceVectorQueryBuilder]: ...
@overload
def search(
self,
query: FullTextQuery,
vector_column_name: Optional[str] = None,
query_type: QueryType = "auto",
ordering_field_name: Optional[str] = None,
fts_columns: Optional[Union[str, List[str]]] = None,
) -> LanceFtsQueryBuilder: ...
@overload
def search(
self,
query: Union[VEC, "PIL.Image.Image", Tuple],
vector_column_name: Optional[str] = None,
query_type: Literal["auto"] = "auto",
ordering_field_name: Optional[str] = None,
fts_columns: Optional[Union[str, List[str]]] = None,
) -> LanceVectorQueryBuilder: ...
@overload
def search(
self, self,
query: Optional[Union[VEC, str, "PIL.Image.Image", Tuple]] = None, query: Optional[Union[VEC, str, "PIL.Image.Image", Tuple]] = None,
vector_column_name: Optional[str] = None, vector_column_name: Optional[str] = None,
@@ -3347,7 +3516,7 @@ class LanceTable(Table):
@overload @overload
def search( def search(
self, self,
query: Optional[Union[VEC, str, "PIL.Image.Image", Tuple]] = None, query: Optional[Union[str, FullTextQuery]] = None,
vector_column_name: Optional[str] = None, vector_column_name: Optional[str] = None,
query_type: Literal["fts"] = "fts", query_type: Literal["fts"] = "fts",
ordering_field_name: Optional[str] = None, ordering_field_name: Optional[str] = None,
@@ -3357,9 +3526,7 @@ class LanceTable(Table):
@overload @overload
def search( def search(
self, self,
query: Optional[ query: Optional[Union[VEC, str, "PIL.Image.Image", Tuple]] = None,
Union[VEC, str, "PIL.Image.Image", Tuple, FullTextQuery]
] = None,
vector_column_name: Optional[str] = None, vector_column_name: Optional[str] = None,
query_type: Literal["hybrid"] = "hybrid", query_type: Literal["hybrid"] = "hybrid",
ordering_field_name: Optional[str] = None, ordering_field_name: Optional[str] = None,
@@ -3369,12 +3536,17 @@ class LanceTable(Table):
@overload @overload
def search( def search(
self, self,
query: None = None, query: Optional[Union[VEC, str, "PIL.Image.Image", Tuple]] = None,
vector_column_name: Optional[str] = None, vector_column_name: Optional[str] = None,
query_type: QueryType = "auto", query_type: QueryType = "auto",
ordering_field_name: Optional[str] = None, ordering_field_name: Optional[str] = None,
fts_columns: Optional[Union[str, List[str]]] = None, fts_columns: Optional[Union[str, List[str]]] = None,
) -> LanceEmptyQueryBuilder: ... ) -> Union[
LanceEmptyQueryBuilder,
LanceFtsQueryBuilder,
LanceHybridQueryBuilder,
LanceVectorQueryBuilder,
]: ...
def search( def search(
self, self,
@@ -3387,8 +3559,8 @@ class LanceTable(Table):
fts_columns: Optional[Union[str, List[str]]] = None, fts_columns: Optional[Union[str, List[str]]] = None,
) -> LanceQueryBuilder: ) -> LanceQueryBuilder:
"""Create a search query to find the nearest neighbors """Create a search query to find the nearest neighbors
of the given query vector. We currently support [vector search][search] of the given query vector. We currently support [vector search](https://lancedb.com/docs/search/vector-search/)
and [full-text search][search]. and [full-text search](https://lancedb.com/docs/search/full-text-search/).
Examples Examples
-------- --------
@@ -3418,8 +3590,9 @@ class LanceTable(Table):
- *default None*. - *default None*.
Acceptable types are: list, np.ndarray, PIL.Image.Image Acceptable types are: list, np.ndarray, PIL.Image.Image
- If None then the select/[where][sql]/limit clauses are applied - If None then the
to filter the table select/[where][lancedb.query.LanceQueryBuilder.where]/limit clauses
are applied to filter the table
vector_column_name: str, optional vector_column_name: str, optional
The name of the vector column to search. The name of the vector column to search.
@@ -3813,6 +3986,8 @@ class LanceTable(Table):
retrain: bool, default False retrain: bool, default False
This parameter is no longer used and is deprecated. This parameter is no longer used and is deprecated.
Notes
-----
The frequency an application should call optimize is based on the frequency of The frequency an application should call optimize is based on the frequency of
data modifications. If data is frequently added, deleted, or updated then data modifications. If data is frequently added, deleted, or updated then
optimize should be run frequently. A good rule of thumb is to run optimize if optimize should be run frequently. A good rule of thumb is to run optimize if
@@ -4691,7 +4866,7 @@ class AsyncTable:
Parameters Parameters
---------- ----------
**kwargs **kwargs
Forwarded to [`lance.dataset`][lance.dataset]. Forwarded to `lance.dataset`.
Returns Returns
------- -------
@@ -4867,6 +5042,46 @@ class AsyncTable:
) )
raise e raise e
async def create_index_async(
self,
column: str,
*,
replace: Optional[bool] = None,
config: Optional[
Union[
IvfFlat,
IvfPq,
IvfRq,
HnswPq,
HnswSq,
HnswFlat,
BTree,
Bitmap,
LabelList,
Fm,
FTS,
]
] = None,
wait_timeout: Optional[timedelta] = None,
name: Optional[str] = None,
train: bool = True,
) -> AsyncJob:
"""Create an index, returning a handle to the indexing job.
Takes the same arguments as :meth:`create_index`. The job may already
be complete when returned; callers must not assume the index exists
until :meth:`AsyncJob.wait` resolves.
"""
job = await self._inner.create_index_async(
column,
index=config,
replace=replace,
wait_timeout=wait_timeout,
name=name,
train=train,
)
return AsyncJob(job)
async def drop_index(self, name: str) -> None: async def drop_index(self, name: str) -> None:
""" """
Drop an index from the table. Drop an index from the table.
@@ -5010,7 +5225,7 @@ class AsyncTable:
progress: Optional[Union[bool, Callable, Any]] = None, progress: Optional[Union[bool, Callable, Any]] = None,
write_parallelism: Optional[int] = None, write_parallelism: Optional[int] = None,
) -> AddResult: ) -> AddResult:
"""Add more data to the [Table](Table). """Add more data to the [AsyncTable][lancedb.table.AsyncTable].
Parameters Parameters
---------- ----------
@@ -5212,8 +5427,8 @@ class AsyncTable:
fts_columns: Optional[Union[str, List[str]]] = None, fts_columns: Optional[Union[str, List[str]]] = None,
) -> Union[AsyncHybridQuery, AsyncFTSQuery, AsyncVectorQuery]: ) -> Union[AsyncHybridQuery, AsyncFTSQuery, AsyncVectorQuery]:
"""Create a search query to find the nearest neighbors """Create a search query to find the nearest neighbors
of the given query vector. We currently support [vector search][search] of the given query vector. We currently support [vector search](https://lancedb.com/docs/search/vector-search/)
and [full-text search][experimental-full-text-search]. and [full-text search](https://lancedb.com/docs/search/full-text-search/).
All query options are defined in [AsyncQuery][lancedb.query.AsyncQuery]. All query options are defined in [AsyncQuery][lancedb.query.AsyncQuery].
@@ -5774,15 +5989,14 @@ class AsyncTable:
change permanent you can use the `[Self::restore]` method. change permanent you can use the `[Self::restore]` method.
Any operation that modifies the table will fail while the table is in a checked Any operation that modifies the table will fail while the table is in a checked
out state. out state. To return the table to a normal state use
`[Self::checkout_latest]`.
Parameters Parameters
---------- ----------
version: int | str, version: int | str,
The version to check out. A version number (`int`) or a tag The version to check out. A version number (`int`) or a tag
(`str`) can be provided. (`str`) can be provided.
To return the table to a normal state use `[Self::checkout_latest]`
""" """
try: try:
await self._inner.checkout(version) await self._inner.checkout(version)
@@ -5966,6 +6180,8 @@ class AsyncTable:
retrain: bool, default False retrain: bool, default False
This parameter is no longer used and is deprecated. This parameter is no longer used and is deprecated.
Notes
-----
The frequency an application should call optimize is based on the frequency of The frequency an application should call optimize is based on the frequency of
data modifications. If data is frequently added, deleted, or updated then data modifications. If data is frequently added, deleted, or updated then
optimize should be run frequently. A good rule of thumb is to run optimize if optimize should be run frequently. A good rule of thumb is to run optimize if
@@ -6346,6 +6562,8 @@ class Branches:
dry_run: bool, default False dry_run: bool, default False
When True, only preview. When False, attempt the merge. When True, only preview. When False, attempt the merge.
Notes
-----
A rejected merge returns ``status="rejected"`` instead of raising. A rejected merge returns ``status="rejected"`` instead of raising.
""" """
return LOOP.run(self._table.branches.merge(from_branch, dry_run)) return LOOP.run(self._table.branches.merge(from_branch, dry_run))
+3 -3
View File
@@ -226,13 +226,13 @@ def test_fetch_blob_ranges_validates_requests():
table = _blob_table("range_validation", [{"id": 1, "image": b"abc"}]) table = _blob_table("range_validation", [{"id": 1, "image": b"abc"}])
row_id = _row_ids_by_id(table)[1] row_id = _row_ids_by_id(table)[1]
with pytest.raises(RuntimeError, match="exceeds blob size"): with pytest.raises(ValueError, match="exceeds blob size"):
table.fetch_blob_ranges("image", [(row_id, 2, 2)]) table.fetch_blob_ranges("image", [(row_id, 2, 2)])
with pytest.raises(RuntimeError, match="offset \\+ length overflowed"): with pytest.raises(ValueError, match="offset \\+ length overflowed"):
table.fetch_blob_ranges("image", [(row_id, 2**64 - 1, 1)]) table.fetch_blob_ranges("image", [(row_id, 2**64 - 1, 1)])
with pytest.raises(ValueError, match="row ids"): with pytest.raises(ValueError, match="row IDs"):
table.fetch_blob_ranges("image", [(2**64 - 1, 0, 1)]) table.fetch_blob_ranges("image", [(2**64 - 1, 0, 1)])
+15
View File
@@ -62,6 +62,21 @@ def test_basic(tmp_path):
assert db.open_table("test").name == db["test"].name assert db.open_table("test").name == db["test"].name
def test_sync_repr_does_not_use_background_loop(tmp_path, monkeypatch):
from lancedb.background_loop import LOOP
db = lancedb.connect(tmp_path)
table = db.create_table("test", data=[{"id": 1}])
def fail_run(*args, **kwargs):
raise AssertionError("repr should not use the Python background loop")
monkeypatch.setattr(LOOP, "run", fail_run)
assert repr(db) == f"LanceDBConnection(uri={str(tmp_path)!r})"
assert repr(table) == f"LanceTable(name='test', _conn={db!r})"
def test_ingest_pd(tmp_path): def test_ingest_pd(tmp_path):
db = lancedb.connect(tmp_path) db = lancedb.connect(tmp_path)
+13
View File
@@ -123,6 +123,19 @@ async def test_async_hybrid_query_default_limit(table: AsyncTable):
assert texts.count("a") == 1 assert texts.count("a") == 1
def test_hybrid_query_minimum_nprobes_zero_raises(sync_table: Table):
# minimum_nprobes(0) must raise the same validation error a plain vector
# query raises, not silently no-op because 0 is falsy.
with pytest.raises(ValueError, match="minimum_nprobes must be greater than 0"):
(
sync_table.search(query_type="hybrid")
.vector([0.0, 0.4])
.text("dog")
.minimum_nprobes(0)
.to_arrow()
)
def test_hybrid_query_distance_range(sync_table: Table): def test_hybrid_query_distance_range(sync_table: Table):
reranker = RRFReranker(return_score="all") reranker = RRFReranker(return_score="all")
result = ( result = (
+9
View File
@@ -84,6 +84,15 @@ async def binary_table(db_async):
) )
@pytest.mark.asyncio
async def test_create_index_async_returns_done_job(some_table: AsyncTable):
job = await some_table.create_index_async("id", config=BTree())
assert job.id is None
await job.wait()
assert len(await some_table.list_indices()) == 1
await job.cancel()
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_create_scalar_index(some_table: AsyncTable): async def test_create_scalar_index(some_table: AsyncTable):
# Can create # Can create
@@ -18,6 +18,7 @@ Tests verify:
""" """
import copy import copy
import os
import shutil import shutil
import sys import sys
import tempfile import tempfile
@@ -239,7 +240,7 @@ def create_tracking_namespace(
dir_props = {f"storage.{k}": v for k, v in storage_options_with_refresh.items()} dir_props = {f"storage.{k}": v for k, v in storage_options_with_refresh.items()}
if bucket_name.startswith("/") or bucket_name.startswith("file://"): if os.path.isabs(bucket_name) or bucket_name.startswith("file://"):
dir_props["root"] = f"{bucket_name}/namespace_root" dir_props["root"] = f"{bucket_name}/namespace_root"
else: else:
dir_props["root"] = f"s3://{bucket_name}/namespace_root" dir_props["root"] = f"s3://{bucket_name}/namespace_root"
@@ -767,3 +768,70 @@ def test_namespace_with_schema_only(s3_bucket: str, use_custom: bool):
# Verify data was added # Verify data was added
assert table.count_rows() == 2 assert table.count_rows() == 2
@pytest.mark.parametrize("use_custom", [False, True], ids=["DirectoryNS", "CustomNS"])
def test_namespace_exists(use_custom: bool):
"""
Test namespace_exists returns True for existing and False for non-existent.
"""
temp_dir = tempfile.mkdtemp()
try:
ns_client, _ = create_tracking_namespace(
bucket_name=temp_dir,
storage_options={},
credential_expires_in_seconds=3600,
use_custom=use_custom,
)
db = LanceNamespaceDBConnection(ns_client)
namespace_name = f"test_ns_{uuid.uuid4().hex[:8]}"
db.create_namespace([namespace_name])
# Existing namespace should return True
assert db.namespace_exists(namespace_id=[namespace_name]) is True
# Non-existent namespace should return False
assert db.namespace_exists(namespace_id=["nonexistent_ns"]) is False
finally:
shutil.rmtree(temp_dir, ignore_errors=True)
@pytest.mark.parametrize("use_custom", [False, True], ids=["DirectoryNS", "CustomNS"])
def test_table_exists(use_custom: bool):
"""
Test table_exists returns True for existing table and False for non-existent.
"""
temp_dir = tempfile.mkdtemp()
try:
ns_client, _ = create_tracking_namespace(
bucket_name=temp_dir,
storage_options={},
credential_expires_in_seconds=3600,
use_custom=use_custom,
)
db = LanceNamespaceDBConnection(ns_client)
namespace_name = f"test_ns_{uuid.uuid4().hex[:8]}"
db.create_namespace([namespace_name])
table_name = f"test_table_{uuid.uuid4().hex}"
namespace_path = [namespace_name]
schema = pa.schema(
[
pa.field("id", pa.int64()),
pa.field("vector", pa.list_(pa.float32(), 2)),
pa.field("text", pa.string()),
]
)
db.create_table(table_name, schema=schema, namespace_path=namespace_path)
# Existing table should return True
table_id = namespace_path + [table_name]
assert db.table_exists(table_id=table_id) is True
# Non-existent table should return False
assert db.table_exists(table_id=namespace_path + ["nonexistent_table"]) is False
finally:
shutil.rmtree(temp_dir, ignore_errors=True)
+443 -1
View File
@@ -812,6 +812,121 @@ def test_table_create_indices():
table.drop_index("custom_fts_idx") table.drop_index("custom_fts_idx")
def test_remote_create_index_async_returns_job():
from lancedb.index import BTree
describe_calls = []
def handler(request):
content_len = int(request.headers.get("Content-Length", 0))
body = request.rfile.read(content_len) if content_len > 0 else b""
if request.path == "/v1/table/test/create_index/":
request.send_response(200)
request.send_header("Content-Type", "application/json")
request.end_headers()
request.wfile.write(b'{"job_id": "job-1"}')
elif request.path == "/v1/jobs/describe":
assert json.loads(body)["job_id"] == "job-1"
describe_calls.append(1)
state = "IN_PROGRESS" if len(describe_calls) == 1 else "DONE"
request.send_response(200)
request.send_header("Content-Type", "application/json")
request.end_headers()
request.wfile.write(
json.dumps(dict(job_id="job-1", job_state=state)).encode()
)
elif request.path == "/v1/jobs/cancel":
assert json.loads(body)["job_id"] == "job-1"
request.send_response(200)
request.send_header("Content-Type", "application/json")
request.end_headers()
request.wfile.write(b"{}")
elif request.path == "/v1/table/test/create/?mode=create":
request.send_response(200)
request.send_header("Content-Type", "application/json")
request.end_headers()
request.wfile.write(b"{}")
elif request.path == "/v1/table/test/describe/":
request.send_response(200)
request.send_header("Content-Type", "application/json")
request.end_headers()
request.wfile.write(
json.dumps(
dict(
version=1,
schema=dict(
fields=[
dict(name="id", type={"type": "int64"}, nullable=False),
]
),
)
).encode()
)
else:
request.send_response(404)
request.end_headers()
with mock_lancedb_connection(handler) as db:
table = db.create_table("test", [{"id": 1}])
job = table.create_index_async("id", config=BTree())
assert job.id == "job-1"
job.wait(timeout=timedelta(seconds=30))
assert len(describe_calls) == 2
job.cancel()
def test_remote_job_wait_raises_on_failure():
from lancedb.exceptions import JobFailedError
from lancedb.index import BTree
def handler(request):
content_len = int(request.headers.get("Content-Length", 0))
body = request.rfile.read(content_len) if content_len > 0 else b""
if request.path == "/v1/table/test/create_index/":
request.send_response(200)
request.send_header("Content-Type", "application/json")
request.end_headers()
request.wfile.write(b'{"job_id": "job-2"}')
elif request.path == "/v1/jobs/describe":
assert json.loads(body)["job_id"] == "job-2"
request.send_response(200)
request.send_header("Content-Type", "application/json")
request.end_headers()
request.wfile.write(
json.dumps(dict(job_id="job-2", job_state="FAILED")).encode()
)
elif request.path == "/v1/table/test/create/?mode=create":
request.send_response(200)
request.send_header("Content-Type", "application/json")
request.end_headers()
request.wfile.write(b"{}")
elif request.path == "/v1/table/test/describe/":
request.send_response(200)
request.send_header("Content-Type", "application/json")
request.end_headers()
request.wfile.write(
json.dumps(
dict(
version=1,
schema=dict(
fields=[
dict(name="id", type={"type": "int64"}, nullable=False),
]
),
)
).encode()
)
else:
request.send_response(404)
request.end_headers()
with mock_lancedb_connection(handler) as db:
table = db.create_table("test", [{"id": 1}])
job = table.create_index_async("id", config=BTree())
with pytest.raises(JobFailedError, match="job-2"):
job.wait()
def test_remote_create_index_new_api(): def test_remote_create_index_new_api():
received_requests = [] received_requests = []
@@ -1020,7 +1135,7 @@ def query_test_table(query_handler, *, server_version=Version("0.1.0")):
request.send_header("Content-Type", "application/json") request.send_header("Content-Type", "application/json")
request.send_header("phalanx-version", str(server_version)) request.send_header("phalanx-version", str(server_version))
request.end_headers() request.end_headers()
request.wfile.write(b"{}") request.wfile.write(b'{"version": 1, "schema": {"fields": []}}')
elif request.path == "/v1/table/test/query/": elif request.path == "/v1/table/test/query/":
content_len = int(request.headers.get("Content-Length")) content_len = int(request.headers.get("Content-Length"))
body = request.rfile.read(content_len) body = request.rfile.read(content_len)
@@ -1858,3 +1973,330 @@ def test_inherited_remote_table_reopens_after_fork():
finally: finally:
server.shutdown() server.shutdown()
server_thread.join() server_thread.join()
BLOB_DESCRIBE_RESPONSE = {
"table": "test",
"version": 1,
"schema": {
"fields": [
{"name": "id", "type": {"type": "int64"}, "nullable": False},
{
"name": "image",
"type": {
"type": "struct",
"fields": [
{
"name": "data",
"type": {"type": "large_binary"},
"nullable": True,
},
{"name": "uri", "type": {"type": "string"}, "nullable": True},
],
},
"nullable": True,
"metadata": {
"ARROW:extension:name": "lance.blob.v2",
"ARROW:extension:metadata": "",
},
},
]
},
}
def blob_query_response_table():
image_field = pa.field(
"image",
pa.struct(
[
pa.field("kind", pa.uint8(), nullable=False),
pa.field("position", pa.uint64(), nullable=False),
pa.field("size", pa.uint64(), nullable=False),
pa.field("blob_id", pa.uint32(), nullable=False),
pa.field("blob_uri", pa.string(), nullable=False),
]
),
metadata={"lance-encoding:blob": "true"},
)
images = pa.StructArray.from_arrays(
[
pa.array([1, 0, 0], type=pa.uint8()),
pa.array([0, 0, 0], type=pa.uint64()),
pa.array([5, 0, 5], type=pa.uint64()),
pa.array([1, 0, 2], type=pa.uint32()),
pa.array(["", "", ""], type=pa.string()),
],
fields=image_field.type,
mask=pa.array([False, True, False]),
)
return pa.Table.from_arrays(
[
pa.array([1, 2, 3], type=pa.int64()),
images,
pa.array([10, 20, 30], type=pa.uint64()),
],
schema=pa.schema(
[
pa.field("id", pa.int64(), nullable=False),
image_field,
pa.field("_rowid", pa.uint64()),
]
),
)
@contextlib.contextmanager
def blob_remote_table(*, server_version=Version("0.5.0")):
def handler(request):
if request.path == "/v1/table/test/describe/":
request.send_response(200)
request.send_header("Content-Type", "application/json")
request.send_header("phalanx-version", str(server_version))
request.end_headers()
request.wfile.write(json.dumps(BLOB_DESCRIBE_RESPONSE).encode())
elif request.path.startswith("/v1/table/test/blob/image/"):
path = request.path.partition("?")[0]
row_id = int(path.split("/")[-2])
payload = {10: b"alpha", 20: None, 30: b"gamma"}[row_id]
if payload is None:
request.send_response(204)
request.end_headers()
return
byte_range = request.headers["Range"].removeprefix("bytes=")
start_text, end_text = byte_range.split("-", maxsplit=1)
start = int(start_text)
end = int(end_text) if end_text else len(payload) - 1
chunk = payload[start : end + 1]
request.send_response(206)
request.send_header("Content-Range", f"bytes {start}-{end}/{len(payload)}")
request.send_header("Content-Length", str(len(chunk)))
request.end_headers()
request.wfile.write(chunk)
elif request.path == "/v1/table/test/query/":
content_len = int(request.headers.get("Content-Length", 0))
body = json.loads(request.rfile.read(content_len))
assert body["columns"] == ["id", "image"]
assert body["with_row_id"] is True
response_table = blob_query_response_table()
request.send_response(200)
request.send_header("Content-Type", "application/vnd.apache.arrow.file")
request.end_headers()
with pa.ipc.new_file(request.wfile, response_table.schema) as writer:
writer.write_table(response_table)
elif request.path == "/v1/table/test/fetch_blobs/":
content_len = int(request.headers.get("Content-Length", 0))
body = json.loads(request.rfile.read(content_len))
assert body["column"] == "image"
assert body["row_ids"] == [10, 20, 30]
response_table = pa.table(
{"image": pa.array([b"alpha", None, b"gamma"], type=pa.large_binary())}
)
request.send_response(200)
request.send_header("Content-Type", "application/vnd.apache.arrow.stream")
request.end_headers()
with pa.ipc.new_stream(request.wfile, response_table.schema) as writer:
writer.write_table(response_table)
else:
request.send_response(404)
request.end_headers()
with mock_lancedb_connection(handler) as db:
yield db.open_table("test")
def test_remote_blob_columns_and_fetch():
with blob_remote_table() as table:
assert table.blob_columns() == ["image"]
blobs = table.fetch_blobs("image", [10, 20, 30])
assert blobs.to_pylist() == [b"alpha", None, b"gamma"]
def test_remote_blob_files_are_lazy_seekable_handles():
with blob_remote_table() as table:
files = table.fetch_blob_files("image", [10, 20, 30])
assert len(files) == 3
alpha, null_row, gamma = files
assert null_row is None
assert alpha is not None
assert gamma is not None
assert alpha.size() == 5
assert alpha.read_range(1, 3) == b"lph"
gamma.seek(2)
assert gamma.read() == b"mma"
def test_remote_blob_fetch_accepts_query_table():
hits = pa.table({"_rowid": pa.array([10, 20, 30], type=pa.uint64())})
with blob_remote_table() as table:
blobs = table.fetch_blobs("image", hits)
assert blobs.to_pylist() == [b"alpha", None, b"gamma"]
def test_remote_blob_query_stashes_row_ids_for_fetch():
with blob_remote_table() as table:
hits = table.search().select(["id", "image"]).limit(3).to_arrow()
assert "_rowid" not in hits.column_names
assert "_lance_row_id" in hits.schema.field("image").type.names
blobs = table.fetch_blobs("image", hits)
assert blobs.to_pylist() == [b"alpha", None, b"gamma"]
def test_remote_blob_query_survives_a_server_that_ignores_the_row_id_request():
def handler(request):
if request.path == "/v1/table/test/describe/":
request.send_response(200)
request.send_header("Content-Type", "application/json")
request.send_header("phalanx-version", "0.5.0")
request.end_headers()
request.wfile.write(json.dumps(BLOB_DESCRIBE_RESPONSE).encode())
elif request.path == "/v1/table/test/query/":
content_len = int(request.headers.get("Content-Length", 0))
assert json.loads(request.rfile.read(content_len))["with_row_id"] is True
response_table = blob_query_response_table().drop_columns(["_rowid"])
request.send_response(200)
request.send_header("Content-Type", "application/vnd.apache.arrow.file")
request.end_headers()
with pa.ipc.new_file(request.wfile, response_table.schema) as writer:
writer.write_table(response_table)
else:
request.send_response(404)
request.end_headers()
with mock_lancedb_connection(handler) as db:
table = db.open_table("test")
hits = table.search().select(["id", "image"]).limit(3).to_arrow()
assert hits.column_names == ["id", "image"]
assert "_lance_row_id" not in hits.schema.field("image").type.names
with pytest.raises(ValueError, match="pass a list of row ids"):
table.fetch_blobs("image", hits)
def test_remote_blob_byte_apis_not_supported_on_old_server():
with blob_remote_table(server_version=Version("0.1.0")) as table:
assert table.blob_columns() == ["image"]
with pytest.raises(NotImplementedError, match="not supported"):
table.fetch_blobs("image", [1])
with pytest.raises(NotImplementedError, match="not supported"):
table.fetch_blob_files("image", [1])
def test_remote_connection_jobs_surface():
from lancedb.exceptions import JobFailedError
schema = pa.schema([("state", pa.string())])
batch = pa.record_batch([pa.array(["created", "done"])], schema=schema)
sink = pa.BufferOutputStream()
with pa.ipc.new_stream(sink, schema) as writer:
writer.write_batch(batch)
events_body = sink.getvalue().to_pybytes()
def handler(request):
content_len = int(request.headers.get("Content-Length", 0))
body = request.rfile.read(content_len) if content_len > 0 else b""
payload = json.loads(body) if body else {}
if request.path == "/v1/jobs/list":
if payload.get("page_token") is None:
rsp = dict(
jobs=[
dict(
job_id="job-1",
table="t1",
job_type="create_index",
state="in_progress",
created_at_millis=1000,
)
],
page_token="next",
)
else:
assert payload["page_token"] == "next"
rsp = dict(
jobs=[
dict(
job_id="job-2",
table="t2",
job_type="create_index",
state="succeeded",
created_at_millis=2000,
)
]
)
request.send_response(200)
request.send_header("Content-Type", "application/json")
request.end_headers()
request.wfile.write(json.dumps(rsp).encode())
elif request.path == "/v1/jobs/describe":
if payload["job_id"] != "job-1":
request.send_response(404)
request.end_headers()
return
request.send_response(200)
request.send_header("Content-Type", "application/json")
request.end_headers()
request.wfile.write(
json.dumps(
dict(
job_id="job-1",
job_type="create_index",
job_state="FAILED",
creation_ms=1000,
spec=dict(column="vec"),
failure=dict(
phase="execute", message="worker died", retryable=True
),
)
).encode()
)
elif request.path == "/v1/jobs/cancel":
if payload["job_id"] != "job-1":
request.send_response(404)
request.end_headers()
return
request.send_response(200)
request.send_header("Content-Type", "application/json")
request.end_headers()
request.wfile.write(b'{"job_id": "job-1"}')
elif request.path == "/v1/jobs/query_events":
assert payload["job_id"] == "job-1"
request.send_response(200)
request.send_header("Content-Type", "application/vnd.apache.arrow.stream")
request.end_headers()
request.wfile.write(events_body)
else:
request.send_response(404)
request.end_headers()
with mock_lancedb_connection(handler) as db:
jobs = db.list_jobs()
assert [job.job_id for job in jobs] == ["job-1", "job-2"]
assert jobs[0].state == "running"
assert jobs[0].table == "t1"
assert jobs[1].state == "finished"
description = db.get_job("job-1")
assert description.job_type == "create_index"
assert description.state == "failed"
assert json.loads(description.spec_json) == {"column": "vec"}
assert description.failure.message == "worker died"
assert description.failure.retryable is True
assert db.get_job("missing") is None
assert db.cancel_job("job-1") is True
assert db.cancel_job("missing") is False
batches = db.job_history("job-1")
assert len(batches) == 1
assert batches[0].num_rows == 2
assert batches[0].column("state").to_pylist() == ["created", "done"]
job = db.job("job-1")
assert job.id == "job-1"
assert job.status() == "failed"
with pytest.raises(JobFailedError, match="worker died"):
job.wait(timeout=timedelta(seconds=5))
+32 -3
View File
@@ -1402,6 +1402,15 @@ async def test_async_open_table_with_branch_version(tmp_path):
assert await pinned.count_rows() == 4 # writable again assert await pinned.count_rows() == 4 # writable again
def test_create_index_async_returns_done_job(mem_db: DBConnection):
table = mem_db.create_table("job_test", [{"id": i} for i in range(10)])
job = table.create_index_async("id", config=BTree())
assert job.id is None
job.wait()
assert len(table.list_indices()) == 1
job.cancel()
@patch("lancedb.table.AsyncTable.create_index") @patch("lancedb.table.AsyncTable.create_index")
def test_create_index_method(mock_create_index, mem_db: DBConnection): def test_create_index_method(mock_create_index, mem_db: DBConnection):
table = mem_db.create_table( table = mem_db.create_table(
@@ -2355,6 +2364,29 @@ def test_merge_insert_by_source_delete_expr(mem_db: DBConnection):
assert table.to_arrow().sort_by("a") == expected assert table.to_arrow().sort_by("a") == expected
def test_merge_insert_by_source_delete_reconfigure(mem_db: DBConnection):
# Calling when_not_matched_by_source_delete() again with no condition must
# widen the delete to unconditional, not keep the earlier condition around.
table = mem_db.create_table(
"my_table",
data=pa.table({"a": [1, 2, 3], "b": ["a", "b", "c"]}),
)
new_data = pa.table({"a": [2, 4], "b": ["x", "z"]})
merge_insert_res = (
table.merge_insert("a")
.when_matched_update_all()
.when_not_matched_insert_all()
.when_not_matched_by_source_delete("a > 2")
.when_not_matched_by_source_delete()
.execute(new_data)
)
assert merge_insert_res.num_deleted_rows == 2
expected = pa.table({"a": [2, 4], "b": ["x", "z"]})
assert table.to_arrow().sort_by("a") == expected
@pytest.mark.asyncio @pytest.mark.asyncio
async def test_merge_insert_by_source_delete_expr_async( async def test_merge_insert_by_source_delete_expr_async(
mem_db_async: AsyncConnection, mem_db_async: AsyncConnection,
@@ -3087,9 +3119,6 @@ def test_consistency(tmp_path, consistency_interval):
db2 = lancedb.connect(tmp_path, read_consistency_interval=consistency_interval) db2 = lancedb.connect(tmp_path, read_consistency_interval=consistency_interval)
table2 = db2.open_table("my_table") table2 = db2.open_table("my_table")
if consistency_interval is not None:
assert "read_consistency_interval=datetime.timedelta(" in repr(db2)
assert "read_consistency_interval=datetime.timedelta(" in repr(table2)
assert table2.version == table.version assert table2.version == table.version
table.add([{"id": 1}]) table.add([{"id": 1}])
@@ -0,0 +1,75 @@
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright The LanceDB Authors
from typing import assert_type
from lancedb.db import DBConnection
from lancedb.query import (
FullTextQuery,
LanceEmptyQueryBuilder,
LanceFtsQueryBuilder,
LanceHybridQueryBuilder,
LanceVectorQueryBuilder,
)
from lancedb.table import LanceTable
from lancedb.types import QueryType
def check_table_search_types(
connection: DBConnection,
lance_table: LanceTable,
full_text_query: FullTextQuery,
query_type: QueryType,
) -> None:
table = connection.open_table("table")
assert_type(table.search(), LanceEmptyQueryBuilder)
assert_type(table.search([1.0, 2.0]), LanceVectorQueryBuilder)
assert_type(
table.search("query"),
LanceFtsQueryBuilder | LanceVectorQueryBuilder,
)
assert_type(table.search("query", query_type="vector"), LanceVectorQueryBuilder)
assert_type(table.search("query", query_type="fts"), LanceFtsQueryBuilder)
assert_type(table.search("query", query_type="hybrid"), LanceHybridQueryBuilder)
assert_type(table.search("query", None, "vector"), LanceVectorQueryBuilder)
assert_type(table.search("query", None, "fts"), LanceFtsQueryBuilder)
assert_type(table.search("query", None, "hybrid"), LanceHybridQueryBuilder)
assert_type(
lance_table.search("query"),
LanceFtsQueryBuilder | LanceVectorQueryBuilder,
)
assert_type(table.search(full_text_query), LanceFtsQueryBuilder)
assert_type(table.search(full_text_query, query_type="auto"), LanceFtsQueryBuilder)
assert_type(
table.search(full_text_query, query_type="vector"), LanceFtsQueryBuilder
)
assert_type(table.search(full_text_query, query_type="fts"), LanceFtsQueryBuilder)
assert_type(
table.search(full_text_query, query_type="hybrid"), LanceFtsQueryBuilder
)
assert_type(
table.search(full_text_query, query_type=query_type), LanceFtsQueryBuilder
)
assert_type(lance_table.search(full_text_query), LanceFtsQueryBuilder)
assert_type(
lance_table.search(full_text_query, query_type="auto"), LanceFtsQueryBuilder
)
assert_type(
lance_table.search(full_text_query, query_type="vector"),
LanceFtsQueryBuilder,
)
assert_type(
lance_table.search(full_text_query, query_type="fts"), LanceFtsQueryBuilder
)
assert_type(
lance_table.search(full_text_query, query_type="hybrid"),
LanceFtsQueryBuilder,
)
assert_type(
lance_table.search(full_text_query, query_type=query_type),
LanceFtsQueryBuilder,
)
+55 -2
View File
@@ -13,7 +13,11 @@ use crate::{
runtime::future_into_py, runtime::future_into_py,
table::Table, table::Table,
}; };
use arrow::{datatypes::Schema, ffi_stream::ArrowArrayStreamReader, pyarrow::FromPyArrow}; use arrow::{
datatypes::Schema,
ffi_stream::ArrowArrayStreamReader,
pyarrow::{FromPyArrow, ToPyArrow},
};
use lancedb::{ use lancedb::{
connection::Connection as LanceConnection, connection::Connection as LanceConnection,
connection::NamespaceClientPushdownOperation, connection::NamespaceClientPushdownOperation,
@@ -24,7 +28,7 @@ use pyo3::{
Bound, FromPyObject, Py, PyAny, PyRef, PyResult, Python, Bound, FromPyObject, Py, PyAny, PyRef, PyResult, Python,
exceptions::{PyRuntimeError, PyValueError}, exceptions::{PyRuntimeError, PyValueError},
pyclass, pyfunction, pymethods, pyclass, pyfunction, pymethods,
types::{PyDict, PyDictMethods}, types::{PyDict, PyDictMethods, PyList, PyListMethods},
}; };
#[pyclass] #[pyclass]
@@ -536,6 +540,55 @@ impl Connection {
}) })
}) })
} }
pub fn job(&self, job_id: String) -> PyResult<crate::job::Job> {
let inner = self.get_inner()?.clone();
Ok(crate::job::Job::new(inner.job(job_id).infer_error()?))
}
pub fn list_jobs(self_: PyRef<'_, Self>) -> PyResult<Bound<'_, PyAny>> {
let inner = self_.get_inner()?.clone();
future_into_py(self_.py(), async move {
let jobs = inner.list_jobs().await.infer_error()?;
Ok(jobs
.into_iter()
.map(crate::job::JobInfo::from)
.collect::<Vec<_>>())
})
}
pub fn get_job(self_: PyRef<'_, Self>, job_id: String) -> PyResult<Bound<'_, PyAny>> {
let inner = self_.get_inner()?.clone();
future_into_py(self_.py(), async move {
let description = inner.get_job(&job_id).await.infer_error()?;
Ok(description.map(crate::job::JobDescription::from))
})
}
pub fn cancel_job(self_: PyRef<'_, Self>, job_id: String) -> PyResult<Bound<'_, PyAny>> {
let inner = self_.get_inner()?.clone();
future_into_py(self_.py(), async move {
inner.cancel_job(&job_id).await.infer_error()
})
}
#[pyo3(signature = (job_id=None))]
pub fn job_history(
self_: PyRef<'_, Self>,
job_id: Option<String>,
) -> PyResult<Bound<'_, PyAny>> {
let inner = self_.get_inner()?.clone();
future_into_py(self_.py(), async move {
let batches = inner.job_history(job_id.as_deref()).await.infer_error()?;
Python::attach(|py| {
let list = PyList::empty(py);
for batch in batches {
list.append(batch.to_pyarrow(py)?)?;
}
Ok(list.unbind())
})
})
}
} }
#[pyfunction] #[pyfunction]
+12
View File
@@ -102,6 +102,18 @@ impl<T> PythonErrorExt<T> for std::result::Result<T, LanceError> {
err.setattr(intern!(py, "__cause__"), cause_err)?; err.setattr(intern!(py, "__cause__"), cause_err)?;
Err(PyErr::from_value(err)) Err(PyErr::from_value(err))
}), }),
LanceError::JobFailed { .. } => Python::attach(|py| {
let cls = py
.import(intern!(py, "lancedb.exceptions"))?
.getattr(intern!(py, "JobFailedError"))?;
Err(PyErr::from_value(cls.call1((err.to_string(),))?))
}),
LanceError::JobCancelled { .. } => Python::attach(|py| {
let cls = py
.import(intern!(py, "lancedb.exceptions"))?
.getattr(intern!(py, "JobCancelledError"))?;
Err(PyErr::from_value(cls.call1((err.to_string(),))?))
}),
_ => self.runtime_error(), _ => self.runtime_error(),
}, },
} }
+145
View File
@@ -0,0 +1,145 @@
// SPDX-License-Identifier: Apache-2.0
// SPDX-FileCopyrightText: Copyright The LanceDB Authors
use std::sync::Arc;
use crate::runtime::future_into_py;
use pyo3::{Bound, PyAny, PyRef, PyResult, pyclass, pymethods};
use crate::error::PythonErrorExt;
#[pyclass]
pub struct Job {
inner: Arc<lancedb::Job>,
}
impl Job {
pub(crate) fn new(inner: lancedb::Job) -> Self {
Self {
inner: Arc::new(inner),
}
}
}
#[pymethods]
impl Job {
#[getter]
pub fn id(&self) -> Option<String> {
self.inner.id().map(str::to_string)
}
pub fn status(self_: PyRef<'_, Self>) -> PyResult<Bound<'_, PyAny>> {
let inner = self_.inner.clone();
future_into_py(
self_.py(),
async move { inner.status().await.infer_error() },
)
}
pub fn wait(self_: PyRef<'_, Self>) -> PyResult<Bound<'_, PyAny>> {
let inner = self_.inner.clone();
future_into_py(self_.py(), async move {
inner.wait().await.infer_error()?;
Ok(())
})
}
pub fn cancel(self_: PyRef<'_, Self>) -> PyResult<Bound<'_, PyAny>> {
let inner = self_.inner.clone();
future_into_py(self_.py(), async move {
inner.cancel().await.infer_error()?;
Ok(())
})
}
}
/// A row from `Connection.list_jobs`: one server-side job.
#[pyclass(get_all, skip_from_py_object)]
#[derive(Clone)]
pub struct JobInfo {
job_id: String,
table: String,
job_type: String,
state: String,
created_at_millis: i64,
}
#[pymethods]
impl JobInfo {
fn __repr__(&self) -> String {
format!(
"JobInfo(job_id={:?}, table={:?}, job_type={:?}, state={:?}, created_at_millis={})",
self.job_id, self.table, self.job_type, self.state, self.created_at_millis
)
}
}
impl From<lancedb::database::JobInfo> for JobInfo {
fn from(info: lancedb::database::JobInfo) -> Self {
Self {
job_id: info.job_id,
table: info.table,
job_type: info.job_type,
state: info.state,
created_at_millis: info.created_at_millis,
}
}
}
/// The server's account of why a job failed.
#[pyclass(get_all, skip_from_py_object)]
#[derive(Clone)]
pub struct JobFailureInfo {
phase: Option<String>,
message: Option<String>,
retryable: Option<bool>,
}
#[pymethods]
impl JobFailureInfo {
fn __repr__(&self) -> String {
format!(
"JobFailureInfo(phase={:?}, message={:?}, retryable={:?})",
self.phase, self.message, self.retryable
)
}
}
/// A described job from `Connection.get_job`.
#[pyclass(get_all, skip_from_py_object)]
#[derive(Clone)]
pub struct JobDescription {
job_id: String,
job_type: String,
state: String,
creation_ms: i64,
spec_json: Option<String>,
failure: Option<JobFailureInfo>,
}
#[pymethods]
impl JobDescription {
fn __repr__(&self) -> String {
format!(
"JobDescription(job_id={:?}, job_type={:?}, state={:?}, creation_ms={})",
self.job_id, self.job_type, self.state, self.creation_ms
)
}
}
impl From<lancedb::database::JobDescription> for JobDescription {
fn from(description: lancedb::database::JobDescription) -> Self {
Self {
job_id: description.job_id,
job_type: description.job_type,
state: description.state,
creation_ms: description.creation_ms,
spec_json: (!description.spec.is_null()).then(|| description.spec.to_string()),
failure: description.failure.map(|failure| JobFailureInfo {
phase: failure.phase,
message: failure.message,
retryable: failure.retryable,
}),
}
}
}
+5
View File
@@ -25,6 +25,7 @@ pub mod error;
pub mod expr; pub mod expr;
pub mod header; pub mod header;
pub mod index; pub mod index;
pub mod job;
pub mod namespace; pub mod namespace;
pub mod oauth; pub mod oauth;
pub mod otel; pub mod otel;
@@ -44,6 +45,10 @@ pub fn _lancedb(_py: Python, m: &Bound<'_, PyModule>) -> PyResult<()> {
m.add_class::<Connection>()?; m.add_class::<Connection>()?;
m.add_class::<Session>()?; m.add_class::<Session>()?;
m.add_class::<Table>()?; m.add_class::<Table>()?;
m.add_class::<crate::job::Job>()?;
m.add_class::<crate::job::JobInfo>()?;
m.add_class::<crate::job::JobDescription>()?;
m.add_class::<crate::job::JobFailureInfo>()?;
m.add_class::<PyBlobFile>()?; m.add_class::<PyBlobFile>()?;
m.add_class::<IndexConfig>()?; m.add_class::<IndexConfig>()?;
m.add_class::<Query>()?; m.add_class::<Query>()?;
+68 -13
View File
@@ -426,9 +426,11 @@ pub struct PyBlobFile {
impl PyBlobFile { impl PyBlobFile {
fn read_bytes(self_: PyRef<'_, Self>) -> PyResult<Py<PyBytes>> { fn read_bytes(self_: PyRef<'_, Self>) -> PyResult<Py<PyBytes>> {
let inner = self_.inner.clone(); let inner = self_.inner.clone();
let bytes = block_on(async move { inner.read().await }) let py = self_.py();
let bytes = py
.detach(move || block_on(async move { inner.read().await }))
.map_err(|e| PyRuntimeError::new_err(format!("blob read failed: {e}")))?; .map_err(|e| PyRuntimeError::new_err(format!("blob read failed: {e}")))?;
Ok(PyBytes::new(self_.py(), bytes.as_ref()).unbind()) Ok(PyBytes::new(py, bytes.as_ref()).unbind())
} }
pub fn read(self_: PyRef<'_, Self>) -> PyResult<Bound<'_, PyAny>> { pub fn read(self_: PyRef<'_, Self>) -> PyResult<Bound<'_, PyAny>> {
@@ -444,24 +446,32 @@ impl PyBlobFile {
fn close(self_: PyRef<'_, Self>) -> PyResult<()> { fn close(self_: PyRef<'_, Self>) -> PyResult<()> {
let inner = self_.inner.clone(); let inner = self_.inner.clone();
block_on(async move { inner.close().await }) self_
.py()
.detach(move || block_on(async move { inner.close().await }))
.map_err(|e| PyRuntimeError::new_err(format!("blob close failed: {e}"))) .map_err(|e| PyRuntimeError::new_err(format!("blob close failed: {e}")))
} }
fn is_closed(self_: PyRef<'_, Self>) -> bool { fn is_closed(self_: PyRef<'_, Self>) -> bool {
let inner = self_.inner.clone(); let inner = self_.inner.clone();
block_on(async move { inner.is_closed().await }) self_
.py()
.detach(move || block_on(async move { inner.is_closed().await }))
} }
fn seek(self_: PyRef<'_, Self>, position: u64) -> PyResult<()> { fn seek(self_: PyRef<'_, Self>, position: u64) -> PyResult<()> {
let inner = self_.inner.clone(); let inner = self_.inner.clone();
block_on(async move { inner.seek(position).await }) self_
.py()
.detach(move || block_on(async move { inner.seek(position).await }))
.map_err(|e| PyRuntimeError::new_err(format!("blob seek failed: {e}"))) .map_err(|e| PyRuntimeError::new_err(format!("blob seek failed: {e}")))
} }
fn tell(self_: PyRef<'_, Self>) -> PyResult<u64> { fn tell(self_: PyRef<'_, Self>) -> PyResult<u64> {
let inner = self_.inner.clone(); let inner = self_.inner.clone();
block_on(async move { inner.tell().await }) self_
.py()
.detach(move || block_on(async move { inner.tell().await }))
.map_err(|e| PyRuntimeError::new_err(format!("blob tell failed: {e}"))) .map_err(|e| PyRuntimeError::new_err(format!("blob tell failed: {e}")))
} }
@@ -475,16 +485,20 @@ impl PyBlobFile {
.checked_add(length as u64) .checked_add(length as u64)
.ok_or_else(|| PyValueError::new_err("offset + length overflowed"))?; .ok_or_else(|| PyValueError::new_err("offset + length overflowed"))?;
let inner = self_.inner.clone(); let inner = self_.inner.clone();
let bytes = block_on(async move { inner.read_range(offset..end).await }) let py = self_.py();
let bytes = py
.detach(move || block_on(async move { inner.read_range(offset..end).await }))
.map_err(|e| PyRuntimeError::new_err(format!("blob read_range failed: {e}")))?; .map_err(|e| PyRuntimeError::new_err(format!("blob read_range failed: {e}")))?;
Ok(PyBytes::new(self_.py(), bytes.as_ref()).unbind()) Ok(PyBytes::new(py, bytes.as_ref()).unbind())
} }
fn read_up_to(self_: PyRef<'_, Self>, length: usize) -> PyResult<Py<PyBytes>> { fn read_up_to(self_: PyRef<'_, Self>, length: usize) -> PyResult<Py<PyBytes>> {
let inner = self_.inner.clone(); let inner = self_.inner.clone();
let bytes = block_on(async move { inner.read_up_to(length).await }) let py = self_.py();
.map_err(|e| PyRuntimeError::new_err(format!("blob read failed: {e}")))?; let bytes = py
Ok(PyBytes::new(self_.py(), bytes.as_ref()).unbind()) .detach(move || block_on(async move { inner.read_up_to(length).await }))
.map_err(|e| PyRuntimeError::new_err(format!("blob read_up_to failed: {e}")))?;
Ok(PyBytes::new(py, bytes.as_ref()).unbind())
} }
} }
@@ -805,6 +819,37 @@ impl Table {
}) })
} }
#[pyo3(signature = (column, index=None, replace=None, wait_timeout=None, *, name=None, train=None))]
pub fn create_index_async<'a>(
self_: PyRef<'a, Self>,
column: String,
index: Option<Bound<'_, PyAny>>,
replace: Option<bool>,
wait_timeout: Option<Bound<'_, PyAny>>,
name: Option<String>,
train: Option<bool>,
) -> PyResult<Bound<'a, PyAny>> {
let index = extract_index_params(&index)?;
let timeout = wait_timeout.map(|t| t.extract::<std::time::Duration>().unwrap());
let mut op = self_
.inner_ref()?
.create_index_with_timeout(&[column], index, timeout);
if let Some(replace) = replace {
op = op.replace(replace);
}
if let Some(name) = name {
op = op.name(name);
}
if let Some(train) = train {
op = op.train(train);
}
future_into_py(self_.py(), async move {
let job = op.execute_async().await.infer_error()?;
Ok(crate::job::Job::new(job))
})
}
pub fn drop_index(self_: PyRef<'_, Self>, index_name: String) -> PyResult<Bound<'_, PyAny>> { pub fn drop_index(self_: PyRef<'_, Self>, index_name: String) -> PyResult<Bound<'_, PyAny>> {
let inner = self_.inner_ref()?.clone(); let inner = self_.inner_ref()?.clone();
future_into_py(self_.py(), async move { future_into_py(self_.py(), async move {
@@ -1330,7 +1375,12 @@ impl Table {
let inner = self_.inner_ref()?.clone(); let inner = self_.inner_ref()?.clone();
future_into_py(self_.py(), async move { future_into_py(self_.py(), async move {
let result = inner.add_columns(definitions, None).await.infer_error()?; let result = inner
.add_columns()
.transform(definitions)
.execute()
.await
.infer_error()?;
Ok(AddColumnsResult::from(result)) Ok(AddColumnsResult::from(result))
}) })
} }
@@ -1344,7 +1394,12 @@ impl Table {
let inner = self_.inner_ref()?.clone(); let inner = self_.inner_ref()?.clone();
future_into_py(self_.py(), async move { future_into_py(self_.py(), async move {
let result = inner.add_columns(transform, None).await.infer_error()?; let result = inner
.add_columns()
.transform(transform)
.execute()
.await
.infer_error()?;
Ok(AddColumnsResult::from(result)) Ok(AddColumnsResult::from(result))
}) })
} }
+1168 -1066
View File
File diff suppressed because it is too large Load Diff
+1 -1
View File
@@ -1,2 +1,2 @@
[toolchain] [toolchain]
channel = "1.95.0" channel = "1.97.0"
+198 -2
View File
@@ -9,6 +9,7 @@
//! //!
//! Blob tables require Lance file format >= 2.2 and stable row ids at create. //! Blob tables require Lance file format >= 2.2 and stable row ids at create.
use std::ops::Range;
use std::sync::Arc; use std::sync::Arc;
use arrow_array::LargeBinaryArray; use arrow_array::LargeBinaryArray;
@@ -17,10 +18,202 @@ use arrow_schema::{DataType, Field, Schema};
use lance::dataset::{BlobRangeRequest as LanceBlobRangeRequest, Dataset, WriteParams}; use lance::dataset::{BlobRangeRequest as LanceBlobRangeRequest, Dataset, WriteParams};
use lance_arrow::FieldExt; use lance_arrow::FieldExt;
use lance_encoding::version::LanceFileVersion; use lance_encoding::version::LanceFileVersion;
use lance_io::object_store::ObjectStore;
use object_store::path::Path;
use crate::error::{Error, Result}; use crate::error::{Error, Result};
pub use lance::dataset::BlobFile; /// Seekable handle for one blob value, backed by local storage or a remote
/// HTTP byte-range endpoint.
#[derive(Debug)]
pub struct BlobFile {
inner: BlobFileInner,
}
#[derive(Debug)]
enum BlobFileInner {
Native(lance::dataset::BlobFile),
#[cfg(feature = "remote")]
Remote(Box<crate::remote::table::blobs::RemoteBlobFile>),
}
impl From<lance::dataset::BlobFile> for BlobFile {
fn from(value: lance::dataset::BlobFile) -> Self {
Self {
inner: BlobFileInner::Native(value),
}
}
}
#[cfg(feature = "remote")]
impl From<crate::remote::table::blobs::RemoteBlobFile> for BlobFile {
fn from(value: crate::remote::table::blobs::RemoteBlobFile) -> Self {
Self {
inner: BlobFileInner::Remote(Box::new(value)),
}
}
}
impl BlobFile {
/// Inline reader over a data-file slice.
pub fn new_inline(
object_store: Arc<ObjectStore>,
path: Path,
position: u64,
size: u64,
) -> Self {
lance::dataset::BlobFile::new_inline(object_store, path, position, size).into()
}
/// Dedicated sidecar-file reader.
pub fn new_dedicated(object_store: Arc<ObjectStore>, path: Path, size: u64) -> Self {
lance::dataset::BlobFile::new_dedicated(object_store, path, size).into()
}
/// Packed reader for a slice in a shared sidecar.
pub fn new_packed(
object_store: Arc<ObjectStore>,
path: Path,
position: u64,
size: u64,
) -> Self {
lance::dataset::BlobFile::new_packed(object_store, path, position, size).into()
}
/// External reader at a resolved object location.
pub fn new_external(
object_store: Arc<ObjectStore>,
path: Path,
uri: String,
position: u64,
size: u64,
) -> Self {
lance::dataset::BlobFile::new_external(object_store, path, uri, position, size).into()
}
/// Close the handle.
pub async fn close(&self) -> lance_core::Result<()> {
match &self.inner {
BlobFileInner::Native(file) => file.close().await,
#[cfg(feature = "remote")]
BlobFileInner::Remote(file) => file.close().await,
}
}
/// Whether the handle is closed.
pub async fn is_closed(&self) -> bool {
match &self.inner {
BlobFileInner::Native(file) => file.is_closed().await,
#[cfg(feature = "remote")]
BlobFileInner::Remote(file) => file.is_closed(),
}
}
/// Read a range without moving the cursor.
pub async fn read_range(&self, range: Range<u64>) -> lance_core::Result<bytes::Bytes> {
match &self.inner {
BlobFileInner::Native(file) => file.read_range(range).await,
#[cfg(feature = "remote")]
BlobFileInner::Remote(file) => file.read_range(range).await,
}
}
/// Read ranges without moving the cursor.
pub async fn read_ranges(
&self,
ranges: &[Range<u64>],
) -> lance_core::Result<Vec<bytes::Bytes>> {
match &self.inner {
BlobFileInner::Native(file) => file.read_ranges(ranges).await,
#[cfg(feature = "remote")]
BlobFileInner::Remote(file) => file.read_ranges(ranges).await,
}
}
/// Read from the cursor to the end.
pub async fn read(&self) -> lance_core::Result<bytes::Bytes> {
match &self.inner {
BlobFileInner::Native(file) => file.read().await,
#[cfg(feature = "remote")]
BlobFileInner::Remote(file) => file.read().await,
}
}
/// Read up to `len` bytes and advance the cursor.
pub async fn read_up_to(&self, len: usize) -> lance_core::Result<bytes::Bytes> {
match &self.inner {
BlobFileInner::Native(file) => file.read_up_to(len).await,
#[cfg(feature = "remote")]
BlobFileInner::Remote(file) => file.read_up_to(len).await,
}
}
/// Move the cursor to `new_cursor`.
pub async fn seek(&self, new_cursor: u64) -> lance_core::Result<()> {
match &self.inner {
BlobFileInner::Native(file) => file.seek(new_cursor).await,
#[cfg(feature = "remote")]
BlobFileInner::Remote(file) => file.seek(new_cursor).await,
}
}
/// Current cursor position.
pub async fn tell(&self) -> lance_core::Result<u64> {
match &self.inner {
BlobFileInner::Native(file) => file.tell().await,
#[cfg(feature = "remote")]
BlobFileInner::Remote(file) => file.tell().await,
}
}
/// Blob length in bytes.
pub fn size(&self) -> u64 {
match &self.inner {
BlobFileInner::Native(file) => file.size(),
#[cfg(feature = "remote")]
BlobFileInner::Remote(file) => file.size(),
}
}
/// Physical byte offset in the data file. `None` on remote handles. The
/// Cloud byte-range route does not expose storage layout.
pub fn position(&self) -> Option<u64> {
match &self.inner {
BlobFileInner::Native(file) => Some(file.position()),
#[cfg(feature = "remote")]
BlobFileInner::Remote(_) => None,
}
}
/// Path of the data file holding the blob. `None` on remote handles. The
/// Cloud byte-range route does not expose storage layout.
pub fn data_path(&self) -> Option<&Path> {
match &self.inner {
BlobFileInner::Native(file) => Some(file.data_path()),
#[cfg(feature = "remote")]
BlobFileInner::Remote(_) => None,
}
}
/// Native storage layout. `None` on remote handles. The Cloud byte-range
/// route does not expose layout.
pub fn kind(&self) -> Option<lance_core::datatypes::BlobKind> {
match &self.inner {
BlobFileInner::Native(file) => Some(file.kind()),
#[cfg(feature = "remote")]
BlobFileInner::Remote(_) => None,
}
}
/// External URI for native handles. Remote handles do not expose storage URIs.
pub fn uri(&self) -> Option<&str> {
match &self.inner {
BlobFileInner::Native(file) => file.uri(),
#[cfg(feature = "remote")]
BlobFileInner::Remote(_) => None,
}
}
}
/// One row-specific blob range read request. /// One row-specific blob range read request.
/// ///
@@ -264,7 +457,10 @@ pub(crate) async fn take_blob_files_aligned(
let handles = dataset.take_blobs(row_ids, column).await?; let handles = dataset.take_blobs(row_ids, column).await?;
ensure_all_row_ids_resolved(column, row_ids.len(), handles.len())?; ensure_all_row_ids_resolved(column, row_ids.len(), handles.len())?;
Ok(handles) Ok(handles
.into_iter()
.map(|handle| handle.map(Into::into))
.collect())
} }
#[cfg(test)] #[cfg(test)]
+39 -2
View File
@@ -23,8 +23,8 @@ use crate::connection::create_table::CreateTableBuilder;
use crate::data::scannable::Scannable; use crate::data::scannable::Scannable;
use crate::database::listing::ListingDatabase; use crate::database::listing::ListingDatabase;
use crate::database::{ use crate::database::{
CloneTableRequest, Database, DatabaseOptions, OpenTableRequest, ReadConsistency, CloneTableRequest, Database, DatabaseOptions, JobDescription, JobInfo, OpenTableRequest,
TableNamesRequest, ReadConsistency, TableNamesRequest,
}; };
use crate::embeddings::{EmbeddingRegistry, MemoryRegistry}; use crate::embeddings::{EmbeddingRegistry, MemoryRegistry};
use crate::error::{Error, Result}; use crate::error::{Error, Result};
@@ -456,6 +456,10 @@ impl Connection {
/// ///
/// # Returns /// # Returns
/// Created [`TableRef`], or [`Error::TableNotFound`] if the table does not exist. /// Created [`TableRef`], or [`Error::TableNotFound`] if the table does not exist.
/// If the table's storage is present but holds no readable dataset (for example a
/// `<name>.lance` directory left behind by an interrupted drop and re-create, which
/// [`Self::table_names`] still lists) this returns [`Error::TableCorrupted`]
/// instead.
pub fn open_table(&self, name: impl Into<String>) -> OpenTableBuilder { pub fn open_table(&self, name: impl Into<String>) -> OpenTableBuilder {
OpenTableBuilder::new( OpenTableBuilder::new(
self.internal.clone(), self.internal.clone(),
@@ -513,6 +517,39 @@ impl Connection {
self.internal.read_consistency().await self.internal.read_consistency().await
} }
/// A [`crate::job::Job`] handle for a server-side job by id, suitable for
/// waiting on or cancelling the job.
///
/// The handle is constructed without a server round trip; an unknown id
/// surfaces when the handle is used. Only server-backed databases support
/// job handles by id.
pub fn job(&self, job_id: impl AsRef<str>) -> Result<crate::job::Job> {
self.internal.job(job_id.as_ref())
}
/// List server-side jobs across the database's tables.
pub async fn list_jobs(&self) -> Result<Vec<JobInfo>> {
self.internal.list_jobs().await
}
/// Describe a single server-side job by id. `None` when the server has no
/// such job.
pub async fn get_job(&self, job_id: impl AsRef<str>) -> Result<Option<JobDescription>> {
self.internal.get_job(job_id.as_ref()).await
}
/// Request cancellation of a server-side job by id. Returns true if the
/// server accepted the cancellation, false if no such job exists.
pub async fn cancel_job(&self, job_id: impl AsRef<str>) -> Result<bool> {
self.internal.cancel_job(job_id.as_ref()).await
}
/// The lifecycle event history of a server-side job (all jobs when
/// `job_id` is `None`), as recorded Arrow batches.
pub async fn job_history(&self, job_id: Option<&str>) -> Result<Vec<RecordBatch>> {
self.internal.job_history(job_id).await
}
/// Drop a table in the database. /// Drop a table in the database.
/// ///
/// # Arguments /// # Arguments
+66
View File
@@ -18,6 +18,8 @@ use std::collections::HashMap;
use std::sync::Arc; use std::sync::Arc;
use std::time::Duration; use std::time::Duration;
use arrow_array::RecordBatch;
use lance::dataset::ReadParams; use lance::dataset::ReadParams;
use lance_namespace::LanceNamespace; use lance_namespace::LanceNamespace;
use lance_namespace::models::{ use lance_namespace::models::{
@@ -200,6 +202,45 @@ pub enum ReadConsistency {
Strong, Strong,
} }
/// A row from [`Database::list_jobs`]: one server-side job (index build,
/// compaction, column refresh, ...).
#[derive(Debug, Clone)]
pub struct JobInfo {
/// The job id -- what [`Database::get_job`] and [`Database::cancel_job`]
/// accept.
pub job_id: String,
/// The table the job runs against, without URI or namespace.
pub table: String,
pub job_type: String,
/// Lifecycle state: "running", "finished", "failed", or "cancelled".
pub state: String,
/// When the job was created, in milliseconds since the epoch.
pub created_at_millis: i64,
}
/// A described job from [`Database::get_job`]: lifecycle state plus the
/// job-type-specific specification.
#[derive(Debug, Clone)]
pub struct JobDescription {
pub job_id: String,
pub job_type: String,
/// Lifecycle state: "running", "finished", "failed", or "cancelled".
pub state: String,
/// When the job was created, in milliseconds since the epoch.
pub creation_ms: i64,
/// The job-type-specific specification. Null when the server omits it.
pub spec: serde_json::Value,
/// Why the job failed, when the job is failed and the server reports a
/// reason.
pub failure: Option<crate::error::JobFailure>,
}
fn job_op_not_supported<T>(what: &str) -> Result<T> {
Err(crate::error::Error::NotSupported {
message: format!("{} is not supported by this database", what),
})
}
/// The `Database` trait defines the interface for database implementations. /// The `Database` trait defines the interface for database implementations.
/// ///
/// A database is responsible for managing tables and their metadata. /// A database is responsible for managing tables and their metadata.
@@ -245,6 +286,31 @@ pub trait Database:
/// ///
/// See [`CloneTableRequest`] for detailed documentation and examples. /// See [`CloneTableRequest`] for detailed documentation and examples.
async fn clone_table(&self, request: CloneTableRequest) -> Result<Arc<dyn BaseTable>>; async fn clone_table(&self, request: CloneTableRequest) -> Result<Arc<dyn BaseTable>>;
/// A [`crate::job::Job`] handle for a server-side job by id, suitable for
/// waiting on or cancelling the job. The handle is constructed without a
/// server round trip; an unknown id surfaces when the handle is used.
fn job(&self, _job_id: &str) -> Result<crate::job::Job> {
job_op_not_supported("job")
}
/// List server-side jobs across the database's tables.
async fn list_jobs(&self) -> Result<Vec<JobInfo>> {
job_op_not_supported("list_jobs")
}
/// Describe a single job by id. `None` when the server has no such job.
async fn get_job(&self, _job_id: &str) -> Result<Option<JobDescription>> {
job_op_not_supported("get_job")
}
/// Request cancellation of a job by id. Returns true if the server
/// accepted the cancellation, false if no such job exists. Cancelling an
/// already-terminal job is a no-op success.
async fn cancel_job(&self, _job_id: &str) -> Result<bool> {
job_op_not_supported("cancel_job")
}
/// The lifecycle event history of a job (all jobs when `job_id` is
/// `None`), as recorded Arrow batches.
async fn job_history(&self, _job_id: Option<&str>) -> Result<Vec<RecordBatch>> {
job_op_not_supported("job_history")
}
/// Open a table in the database /// Open a table in the database
async fn open_table(&self, request: OpenTableRequest) -> Result<Arc<dyn BaseTable>>; async fn open_table(&self, request: OpenTableRequest) -> Result<Arc<dyn BaseTable>>;
/// Rename a table in the database /// Rename a table in the database
@@ -11,7 +11,9 @@ use lance_core::{cache::LanceCache, utils::futures::FinallyStreamExt};
use lance_encoding::decoder::{DecoderPlugins, FilterExpression}; use lance_encoding::decoder::{DecoderPlugins, FilterExpression};
use lance_file::{ use lance_file::{
reader::{FileReader, FileReaderOptions}, reader::{FileReader, FileReaderOptions},
writer::{FileWriter, FileWriterOptions}, version::ConcreteFileVersion,
versions,
writer::FileWriterOptions,
}; };
use lance_io::{ use lance_io::{
ReadBatchParams, ReadBatchParams,
@@ -152,8 +154,12 @@ impl Shuffler {
source: None, source: None,
})?; })?;
let object_writer = object_store.create(&path).await?; let object_writer = object_store.create(&path).await?;
let writer = let writer = versions::create_writer(
FileWriter::try_new(object_writer, schema.clone(), FileWriterOptions::default())?; ConcreteFileVersion::V2_1,
object_writer,
schema.clone(),
FileWriterOptions::default(),
)?;
file_writers.push(writer); file_writers.push(writer);
} }
+3 -3
View File
@@ -264,7 +264,7 @@ pub fn compute_output_schema(
let field_name = ed let field_name = ed
.dest_column .dest_column
.clone() .clone()
.unwrap_or_else(|| format!("{}_embedding", &ed.source_column)); .unwrap_or_else(|| format!("{}_embedding", ed.source_column));
sb.push(Field::new( sb.push(Field::new(
field_name, field_name,
@@ -291,7 +291,7 @@ pub fn compute_embeddings_for_batch(
let dst_field_name = fld let dst_field_name = fld
.dest_column .dest_column
.clone() .clone()
.unwrap_or_else(|| format!("{}_embedding", &fld.source_column)); .unwrap_or_else(|| format!("{}_embedding", fld.source_column));
let dst_field = Field::new( let dst_field = Field::new(
dst_field_name, dst_field_name,
@@ -315,7 +315,7 @@ impl<R: RecordBatchReader> WithEmbeddings<R> {
let field_name = ed let field_name = ed
.dest_column .dest_column
.clone() .clone()
.unwrap_or_else(|| format!("{}_embedding", &ed.source_column)); .unwrap_or_else(|| format!("{}_embedding", ed.source_column));
Ok(Field::new( Ok(Field::new(
field_name, field_name,
func.dest_type()?.into_owned(), func.dest_type()?.into_owned(),
+56 -1
View File
@@ -1,7 +1,8 @@
// SPDX-License-Identifier: Apache-2.0 // SPDX-License-Identifier: Apache-2.0
// SPDX-FileCopyrightText: Copyright The LanceDB Authors // SPDX-FileCopyrightText: Copyright The LanceDB Authors
use std::sync::PoisonError; use std::fmt::{self, Display, Formatter};
use std::sync::{Arc, PoisonError};
use arrow_schema::ArrowError; use arrow_schema::ArrowError;
use datafusion_common::DataFusionError; use datafusion_common::DataFusionError;
@@ -9,6 +10,46 @@ use snafu::Snafu;
pub(crate) type BoxError = Box<dyn std::error::Error + Send + Sync>; pub(crate) type BoxError = Box<dyn std::error::Error + Send + Sync>;
/// Why a job failed, to whatever precision the backend provides.
///
/// A job run in this process carries the error it failed with in [`Self::source`].
/// A job run remotely carries whatever the server reported, which older servers
/// do not report at all. Every field is absent rather than invented when the
/// backend does not supply it.
#[derive(Debug, Clone, Default)]
pub struct JobFailure {
/// The stage the job was in, when known.
pub phase: Option<String>,
/// A human-readable reason, when known.
pub message: Option<String>,
/// Whether a retry could clear the failure, when known.
pub retryable: Option<bool>,
/// The error the job failed with, when it ran in this process.
pub source: Option<Arc<Error>>,
}
impl JobFailure {
/// A failure whose only known detail is the error that caused it.
pub(crate) fn from_source(source: Arc<Error>) -> Self {
Self {
message: Some(source.to_string()),
source: Some(source),
..Default::default()
}
}
}
impl Display for JobFailure {
fn fmt(&self, f: &mut Formatter<'_>) -> fmt::Result {
match (&self.message, &self.phase) {
(Some(message), Some(phase)) => write!(f, ": {message} (in {phase})"),
(Some(message), None) => write!(f, ": {message}"),
(None, Some(phase)) => write!(f, " in {phase}"),
(None, None) => Ok(()),
}
}
}
#[derive(Debug, Snafu)] #[derive(Debug, Snafu)]
#[snafu(visibility(pub(crate)))] #[snafu(visibility(pub(crate)))]
pub enum Error { pub enum Error {
@@ -18,6 +59,10 @@ pub enum Error {
InvalidInput { message: String }, InvalidInput { message: String },
#[snafu(display("Table '{name}' was not found"))] #[snafu(display("Table '{name}' was not found"))]
TableNotFound { name: String, source: BoxError }, TableNotFound { name: String, source: BoxError },
#[snafu(display(
"Table '{name}' exists but could not be loaded (it may be corrupt or incomplete): {source}"
))]
TableCorrupted { name: String, source: BoxError },
#[snafu(display("Database '{name}' was not found"))] #[snafu(display("Database '{name}' was not found"))]
DatabaseNotFound { name: String }, DatabaseNotFound { name: String },
#[snafu(display("Database '{name}' already exists."))] #[snafu(display("Database '{name}' already exists."))]
@@ -40,6 +85,13 @@ pub enum Error {
Runtime { message: String }, Runtime { message: String },
#[snafu(display("Timeout error: {message}"))] #[snafu(display("Timeout error: {message}"))]
Timeout { message: String }, Timeout { message: String },
#[snafu(display("Job{} failed{failure}", job_id.as_ref().map(|id| format!(" {id}")).unwrap_or_default()))]
JobFailed {
job_id: Option<String>,
failure: JobFailure,
},
#[snafu(display("Job{} was cancelled", job_id.as_ref().map(|id| format!(" {id}")).unwrap_or_default()))]
JobCancelled { job_id: Option<String> },
// 3rd party / external errors // 3rd party / external errors
#[snafu(display("object_store error: {source}"))] #[snafu(display("object_store error: {source}"))]
@@ -121,6 +173,9 @@ impl From<lance::Error> for Error {
match source { match source {
lance::Error::Wrapped { error, .. } => Self::from_box_error(error), lance::Error::Wrapped { error, .. } => Self::from_box_error(error),
lance::Error::External { source } => Self::from_box_error(source), lance::Error::External { source } => Self::from_box_error(source),
lance::Error::InvalidInput { source, .. } => Self::InvalidInput {
message: source.to_string(),
},
_ => Self::Lance { source }, _ => Self::Lance { source },
} }
} }
+9 -1
View File
@@ -10,7 +10,7 @@ use std::time::Duration;
use vector::IvfFlatIndexBuilder; use vector::IvfFlatIndexBuilder;
use crate::index::vector::IvfRqIndexBuilder; use crate::index::vector::IvfRqIndexBuilder;
use crate::{DistanceType, Error, Result, table::BaseTable}; use crate::{DistanceType, Error, Result, job::Job, table::BaseTable};
use self::{ use self::{
scalar::{BTreeIndexBuilder, BitmapIndexBuilder, FmIndexBuilder, LabelListIndexBuilder}, scalar::{BTreeIndexBuilder, BitmapIndexBuilder, FmIndexBuilder, LabelListIndexBuilder},
@@ -305,6 +305,14 @@ impl IndexBuilder {
pub async fn execute(self) -> Result<()> { pub async fn execute(self) -> Result<()> {
self.parent.clone().create_index(self).await self.parent.clone().create_index(self).await
} }
/// Creates the index, returning a [`Job`] tracking the operation.
///
/// The job may already be complete when returned, and callers must not
/// assume the index exists until [`Job::wait`] resolves.
pub async fn execute_async(self) -> Result<Job> {
self.parent.clone().create_index_async(self).await
}
} }
#[derive(Debug, Clone, PartialEq, Deserialize)] #[derive(Debug, Clone, PartialEq, Deserialize)]
+182
View File
@@ -0,0 +1,182 @@
// SPDX-License-Identifier: Apache-2.0
// SPDX-FileCopyrightText: Copyright The LanceDB Authors
//! Handles to operations a server may run asynchronously.
use std::sync::Arc;
use async_trait::async_trait;
use tokio::sync::watch;
use tokio::task::{AbortHandle, JoinHandle};
use crate::error::{Error, JobFailure, Result};
/// Backend-specific tracking for an asynchronous operation.
#[async_trait]
pub(crate) trait JobHandle: Send + Sync {
/// Server-assigned id, when the backend has one.
fn id(&self) -> Option<&str> {
None
}
async fn status(&self) -> Result<String>;
async fn wait(&self) -> Result<()>;
async fn cancel(&self) -> Result<()>;
}
/// A handle to an operation that may still be running.
///
/// The operation may already be complete when the handle is created.
pub struct Job {
handle: Option<Box<dyn JobHandle>>,
}
impl std::fmt::Debug for Job {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.debug_struct("Job")
.field("id", &self.id())
.field("done", &self.handle.is_none())
.finish()
}
}
impl Job {
/// A job whose operation finished before the handle was created.
pub(crate) fn new_done() -> Self {
Self { handle: None }
}
pub(crate) fn new(handle: Box<dyn JobHandle>) -> Self {
Self {
handle: Some(handle),
}
}
/// A job running as a task in this process.
pub(crate) fn spawned(task: JoinHandle<Result<()>>) -> Self {
Self::new(Box::new(SpawnedJob::new(task)))
}
/// Identifies the operation on the server that is running it.
///
/// Returned for correlating with server logs or the jobs API. Operations
/// that run in this process have no server id and return `None`. The
/// value is opaque: parsing it or storing it to resume the job later is
/// not supported.
pub fn id(&self) -> Option<&str> {
self.handle.as_ref().and_then(|handle| handle.id())
}
/// The operation's current lifecycle state: "running", "finished",
/// "failed", or "cancelled".
///
/// A point snapshot; unlike [`Job::wait`] it does not block, raise on a
/// terminal failure state, or retry. States a newer server reports that
/// this client version does not know pass through as-is.
pub async fn status(&self) -> Result<String> {
match &self.handle {
None => Ok("finished".to_string()),
Some(handle) => handle.status().await,
}
}
/// Waits until the operation reaches a terminal state.
///
/// Returns [`crate::Error::JobFailed`] if the operation failed and
/// [`crate::Error::JobCancelled`] if it was cancelled.
pub async fn wait(&self) -> Result<()> {
match &self.handle {
None => Ok(()),
Some(handle) => handle.wait().await,
}
}
/// Requests cancellation of the operation.
///
/// Cancelling an operation that already finished is a no-op.
pub async fn cancel(&self) -> Result<()> {
match &self.handle {
None => Ok(()),
Some(handle) => handle.cancel().await,
}
}
}
/// How an in-process operation ended. Cloneable so every waiter can be given
/// the outcome; [`Error`] is not, so failures share one behind an [`Arc`].
#[derive(Clone)]
enum Outcome {
Succeeded,
Failed(Arc<Error>),
Cancelled,
}
impl Outcome {
fn into_result(self) -> Result<()> {
match self {
Self::Succeeded => Ok(()),
Self::Failed(source) => Err(Error::JobFailed {
job_id: None,
failure: JobFailure::from_source(source),
}),
Self::Cancelled => Err(Error::JobCancelled { job_id: None }),
}
}
}
/// Tracks an operation running as a task in this process. A second task
/// watches the first so that aborting it still produces an outcome, and so
/// that every caller of `wait` observes the same one.
struct SpawnedJob {
outcome: watch::Receiver<Option<Outcome>>,
abort: AbortHandle,
}
impl SpawnedJob {
fn new(task: JoinHandle<Result<()>>) -> Self {
let abort = task.abort_handle();
let (tx, outcome) = watch::channel(None);
tokio::spawn(async move {
let outcome = match task.await {
Ok(Ok(())) => Outcome::Succeeded,
Ok(Err(err)) => Outcome::Failed(Arc::new(err)),
Err(err) if err.is_cancelled() => Outcome::Cancelled,
Err(err) => Outcome::Failed(Arc::new(Error::Runtime {
message: format!("index job task failed: {err}"),
})),
};
let _ = tx.send(Some(outcome));
});
Self { outcome, abort }
}
}
#[async_trait]
impl JobHandle for SpawnedJob {
async fn status(&self) -> Result<String> {
let label = match &*self.outcome.borrow() {
None => "running",
Some(Outcome::Succeeded) => "finished",
Some(Outcome::Failed(_)) => "failed",
Some(Outcome::Cancelled) => "cancelled",
};
Ok(label.to_string())
}
async fn wait(&self) -> Result<()> {
let mut outcome = self.outcome.clone();
let settled = outcome
.wait_for(|outcome| outcome.is_some())
.await
.map_err(|_| Error::Runtime {
message: "index job outcome was dropped before it completed".to_string(),
})?
.clone()
.expect("wait_for returns once an outcome is set");
settled.into_result()
}
async fn cancel(&self) -> Result<()> {
self.abort.abort();
Ok(())
}
}
+3 -1
View File
@@ -184,6 +184,7 @@ pub mod expr;
pub mod index; pub mod index;
pub mod io; pub mod io;
pub mod ipc; pub mod ipc;
pub mod job;
#[cfg(feature = "metrics-otel")] #[cfg(feature = "metrics-otel")]
pub mod metrics_otel; pub mod metrics_otel;
#[cfg(feature = "polars")] #[cfg(feature = "polars")]
@@ -203,7 +204,8 @@ use serde::{Deserialize, Serialize};
pub use blob::{BlobRangeRequest, blob, is_blob}; pub use blob::{BlobRangeRequest, blob, is_blob};
pub use connection::{ConnectNamespaceBuilder, Connection}; pub use connection::{ConnectNamespaceBuilder, Connection};
pub use error::{Error, Result}; pub use error::{Error, JobFailure, Result};
pub use job::Job;
use lance_index::vector::ApproxMode as LanceApproxMode; use lance_index::vector::ApproxMode as LanceApproxMode;
use lance_linalg::distance::DistanceType as LanceDistanceType; use lance_linalg::distance::DistanceType as LanceDistanceType;
/// Re-export of the [`metrics`](https://docs.rs/metrics) crate facade. Enable /// Re-export of the [`metrics`](https://docs.rs/metrics) crate facade. Enable
+1 -1
View File
@@ -8,13 +8,13 @@
pub(crate) mod client; pub(crate) mod client;
pub(crate) mod db; pub(crate) mod db;
pub(crate) mod job;
pub mod oauth; pub mod oauth;
mod retry; mod retry;
pub(crate) mod table; pub(crate) mod table;
pub(crate) mod util; pub(crate) mod util;
const ARROW_STREAM_CONTENT_TYPE: &str = "application/vnd.apache.arrow.stream"; const ARROW_STREAM_CONTENT_TYPE: &str = "application/vnd.apache.arrow.stream";
#[cfg(test)]
const ARROW_FILE_CONTENT_TYPE: &str = "application/vnd.apache.arrow.file"; const ARROW_FILE_CONTENT_TYPE: &str = "application/vnd.apache.arrow.file";
#[cfg(test)] #[cfg(test)]
const JSON_CONTENT_TYPE: &str = "application/json"; const JSON_CONTENT_TYPE: &str = "application/json";
+2 -2
View File
@@ -706,7 +706,7 @@ impl<S: HttpSend> RestfulLanceDbClient<S> {
.err_to_http(request_id.clone())?; .err_to_http(request_id.clone())?;
debug!( debug!(
"Received response for request_id={}: {:?}", "Received response for request_id={}: {:?}",
request_id, &response request_id, response
); );
Ok((request_id, response)) Ok((request_id, response))
} }
@@ -768,7 +768,7 @@ impl<S: HttpSend> RestfulLanceDbClient<S> {
Ok((status, response)) if status.is_success() => { Ok((status, response)) if status.is_success() => {
debug!( debug!(
"Received response for request_id={}: {:?}", "Received response for request_id={}: {:?}",
retry_counter.request_id, &response retry_counter.request_id, response
); );
return Ok((retry_counter.request_id, response)); return Ok((retry_counter.request_id, response));
} }
+383 -1
View File
@@ -20,7 +20,7 @@ use lance_namespace::models::{
use crate::Error; use crate::Error;
use crate::database::{ use crate::database::{
CloneTableRequest, CreateTableMode, CreateTableRequest, Database, DatabaseOptions, CloneTableRequest, CreateTableMode, CreateTableRequest, Database, DatabaseOptions,
OpenTableRequest, ReadConsistency, TableNamesRequest, JobDescription, JobInfo, OpenTableRequest, ReadConsistency, TableNamesRequest,
}; };
use crate::error::Result; use crate::error::Result;
use crate::remote::util::stream_as_body; use crate::remote::util::stream_as_body;
@@ -79,6 +79,10 @@ impl ServerVersion {
pub fn support_multipart_write(&self) -> bool { pub fn support_multipart_write(&self) -> bool {
self.0 >= semver::Version::new(0, 4, 0) self.0 >= semver::Version::new(0, 4, 0)
} }
pub fn support_blobs(&self) -> bool {
self.0 >= semver::Version::new(0, 5, 0)
}
} }
pub const OPT_REMOTE_PREFIX: &str = "remote_database_"; pub const OPT_REMOTE_PREFIX: &str = "remote_database_";
@@ -428,6 +432,73 @@ fn build_cache_key(name: &str, namespace: &[String]) -> String {
key.iter().map(|b| format!("{:02x}", b)).collect() key.iter().map(|b| format!("{:02x}", b)).collect()
} }
#[derive(serde::Deserialize)]
struct RemoteListJobRow {
job_id: String,
#[serde(default)]
table: String,
#[serde(default)]
job_type: String,
#[serde(default)]
state: String,
#[serde(default)]
created_at_millis: i64,
}
#[derive(serde::Deserialize)]
struct RemoteListJobsResponse {
#[serde(default)]
jobs: Vec<RemoteListJobRow>,
#[serde(default)]
page_token: Option<String>,
}
/// The server's account of why a job failed. Absent from older servers,
/// which report only the terminal state.
#[derive(serde::Deserialize)]
struct RemoteReportedFailure {
#[serde(default)]
phase: Option<String>,
#[serde(default)]
message: Option<String>,
#[serde(default)]
retryable: Option<bool>,
}
#[derive(serde::Deserialize)]
struct RemoteDescribeJobResponse {
job_id: String,
#[serde(default)]
job_type: String,
job_state: String,
#[serde(default)]
creation_ms: i64,
#[serde(default)]
spec: serde_json::Value,
#[serde(default)]
failure: Option<RemoteReportedFailure>,
}
/// Server job states -> the client vocabulary ("running" / "finished" /
/// "failed" / "cancelled"). Covers both the describe enum (IN_PROGRESS /
/// DONE / FAILED / CANCELLED) and the registry's lowercase list-row states
/// (in_progress / succeeded / failed / canceled / timed_out). States this
/// client version does not know (e.g. created, queued) pass through as-is.
fn job_state_to_client(state: &str) -> String {
match state {
"IN_PROGRESS" | "in_progress" => "running",
"DONE" | "done" | "succeeded" => "finished",
"FAILED" | "failed" | "TIMED_OUT" | "timed_out" => "failed",
"CANCELLED" | "cancelled" | "canceled" => "cancelled",
other => other,
}
.to_string()
}
/// Bound on `list_jobs` page walking; a warning is logged when the listing
/// is truncated at this many pages.
const MAX_LIST_JOBS_PAGES: usize = 100;
#[async_trait] #[async_trait]
impl<S: HttpSend> Database for RemoteDatabase<S> { impl<S: HttpSend> Database for RemoteDatabase<S> {
fn uri(&self) -> &str { fn uri(&self) -> &str {
@@ -441,6 +512,108 @@ impl<S: HttpSend> Database for RemoteDatabase<S> {
}) })
} }
fn job(&self, job_id: &str) -> Result<crate::job::Job> {
Ok(crate::job::Job::new(Box::new(super::job::RemoteJob::new(
self.client.clone(),
job_id.to_string(),
))))
}
async fn list_jobs(&self) -> Result<Vec<JobInfo>> {
let mut out = Vec::new();
let mut page_token: Option<String> = None;
for page in 0..MAX_LIST_JOBS_PAGES {
let mut body = serde_json::json!({});
if let Some(token) = &page_token {
body["page_token"] = serde_json::Value::String(token.clone());
}
let req = self.client.post("/v1/jobs/list").json(&body);
let (request_id, rsp) = self.client.send(req).await?;
let rsp = self.client.check_response(&request_id, rsp).await?;
let body: RemoteListJobsResponse = rsp.json().await.err_to_http(request_id)?;
out.extend(body.jobs.into_iter().map(|row| JobInfo {
job_id: row.job_id,
table: row.table,
job_type: row.job_type,
state: job_state_to_client(&row.state),
created_at_millis: row.created_at_millis,
}));
page_token = body.page_token;
if page_token.is_none() {
break;
}
if page + 1 == MAX_LIST_JOBS_PAGES {
log::warn!(
"list_jobs truncated after {} pages ({} jobs)",
MAX_LIST_JOBS_PAGES,
out.len()
);
}
}
Ok(out)
}
async fn get_job(&self, job_id: &str) -> Result<Option<JobDescription>> {
let req = self
.client
.post("/v1/jobs/describe")
.json(&serde_json::json!({ "job_id": job_id }));
let (request_id, rsp) = self.client.send(req).await?;
let rsp = match self.client.check_response(&request_id, rsp).await {
Ok(rsp) => rsp,
Err(Error::Http {
status_code: Some(StatusCode::NOT_FOUND),
..
}) => return Ok(None),
Err(err) => return Err(err),
};
let body: RemoteDescribeJobResponse = rsp.json().await.err_to_http(request_id)?;
Ok(Some(JobDescription {
job_id: body.job_id,
job_type: body.job_type,
state: job_state_to_client(&body.job_state),
creation_ms: body.creation_ms,
spec: body.spec,
failure: body.failure.map(|reported| crate::error::JobFailure {
phase: reported.phase,
message: reported.message,
retryable: reported.retryable,
source: None,
}),
}))
}
async fn cancel_job(&self, job_id: &str) -> Result<bool> {
let req = self
.client
.post("/v1/jobs/cancel")
.json(&serde_json::json!({ "job_id": job_id }));
let (request_id, rsp) = self.client.send(req).await?;
match self.client.check_response(&request_id, rsp).await {
Ok(_) => Ok(true),
Err(Error::Http {
status_code: Some(StatusCode::NOT_FOUND),
..
}) => Ok(false),
Err(err) => Err(err),
}
}
async fn job_history(&self, job_id: Option<&str>) -> Result<Vec<arrow_array::RecordBatch>> {
let mut body = serde_json::json!({});
if let Some(job_id) = job_id {
body["job_id"] = serde_json::Value::String(job_id.to_string());
}
let req = self.client.post("/v1/jobs/query_events").json(&body);
let (request_id, rsp) = self.client.send(req).await?;
let rsp = self.client.check_response(&request_id, rsp).await?;
let bytes = rsp.bytes().await.err_to_http(request_id)?;
let reader = arrow_ipc::reader::StreamReader::try_new(std::io::Cursor::new(bytes), None)?;
reader
.collect::<std::result::Result<Vec<_>, _>>()
.map_err(Into::into)
}
async fn table_names(&self, request: TableNamesRequest) -> Result<Vec<String>> { async fn table_names(&self, request: TableNamesRequest) -> Result<Vec<String>> {
let mut req = if !request.namespace_path.is_empty() { let mut req = if !request.namespace_path.is_empty() {
let namespace_id = let namespace_id =
@@ -661,6 +834,7 @@ impl<S: HttpSend> Database for RemoteDatabase<S> {
RemoteTable::<S>::handle_table_not_found(&request.name, rsp, &request_id).await?; RemoteTable::<S>::handle_table_not_found(&request.name, rsp, &request_id).await?;
let rsp = self.client.check_response(&request_id, rsp).await?; let rsp = self.client.check_response(&request_id, rsp).await?;
let version = parse_server_version(&request_id, &rsp)?; let version = parse_server_version(&request_id, &rsp)?;
let describe_body = rsp.text().await.ok();
let table_identifier = build_table_identifier( let table_identifier = build_table_identifier(
&request.name, &request.name,
&request.namespace_path, &request.namespace_path,
@@ -673,6 +847,12 @@ impl<S: HttpSend> Database for RemoteDatabase<S> {
table_identifier, table_identifier,
version, version,
)); ));
// This describe already carries the schema, so hand it to the table
// instead of making the first schema read fetch it again. A version or
// branch pin applied after this invalidates the cache.
if let Some(body) = &describe_body {
table.seed_schema(body);
}
let cache_key = build_cache_key(&request.name, &request.namespace_path); let cache_key = build_cache_key(&request.name, &request.namespace_path);
self.table_cache.insert(cache_key, table.clone()).await; self.table_cache.insert(cache_key, table.clone()).await;
Ok(table) Ok(table)
@@ -923,6 +1103,7 @@ impl From<StorageOptions> for RemoteOptions {
mod tests { mod tests {
use super::{NamespaceHeaderProviderContext, build_cache_key}; use super::{NamespaceHeaderProviderContext, build_cache_key};
use std::collections::HashMap; use std::collections::HashMap;
use std::sync::atomic::{AtomicUsize, Ordering};
use std::sync::{Arc, OnceLock}; use std::sync::{Arc, OnceLock};
use arrow_array::{Int32Array, RecordBatch}; use arrow_array::{Int32Array, RecordBatch};
@@ -1073,6 +1254,46 @@ mod tests {
assert_eq!(table.name(), "table1"); assert_eq!(table.name(), "table1");
} }
#[tokio::test]
async fn test_open_table_seeds_the_schema_from_its_describe() {
let describe_calls = Arc::new(AtomicUsize::new(0));
let counted = describe_calls.clone();
let conn = Connection::new_with_handler(move |request| {
assert_eq!(request.url().path(), "/v1/table/table1/describe/");
counted.fetch_add(1, Ordering::SeqCst);
http::Response::builder()
.status(200)
.body(
r#"{"version": 1, "schema": {"fields": [
{"name": "id", "type": {"type": "int64"}, "nullable": false}
]}}"#
.to_string(),
)
.unwrap()
});
let table = conn.open_table("table1").execute().await.unwrap();
let schema = table.schema().await.unwrap();
assert_eq!(schema.field(0).name(), "id");
assert_eq!(describe_calls.load(Ordering::SeqCst), 1);
}
#[tokio::test]
async fn test_open_table_survives_a_describe_body_it_cannot_parse() {
let conn = Connection::new_with_handler(|request| {
assert_eq!(request.url().path(), "/v1/table/table1/describe/");
http::Response::builder()
.status(200)
.body(r#"{"table": "table1"}"#.to_string())
.unwrap()
});
let table = conn.open_table("table1").execute().await.unwrap();
assert_eq!(table.name(), "table1");
}
#[tokio::test] #[tokio::test]
async fn test_open_table_branch_and_version() { async fn test_open_table_branch_and_version() {
let conn = Connection::new_with_handler(|request| { let conn = Connection::new_with_handler(|request| {
@@ -2042,4 +2263,165 @@ mod tests {
assert!(list_response.tables.contains(&"table3".to_string())); assert!(list_response.tables.contains(&"table3".to_string()));
} }
} }
#[tokio::test]
async fn test_list_jobs_paginates() {
let page = Arc::new(AtomicUsize::new(0));
let conn = Connection::new_with_handler(move |request| {
assert_eq!(request.method(), &reqwest::Method::POST);
assert_eq!(request.url().path(), "/v1/jobs/list");
let body: serde_json::Value =
serde_json::from_slice(request.body().unwrap().as_bytes().unwrap()).unwrap();
match page.fetch_add(1, Ordering::SeqCst) {
0 => {
assert!(body.get("page_token").is_none());
http::Response::builder()
.status(200)
.body(
r#"{"jobs": [{"job_id": "job-1", "table": "t1", "job_type": "create_index", "state": "in_progress", "created_at_millis": 1000}], "page_token": "next"}"#,
)
.unwrap()
}
_ => {
assert_eq!(body["page_token"], "next");
http::Response::builder()
.status(200)
.body(
r#"{"jobs": [{"job_id": "job-2", "table": "t2", "job_type": "create_index", "state": "succeeded", "created_at_millis": 2000}, {"job_id": "job-3", "table": "t3", "job_type": "create_index", "state": "timed_out", "created_at_millis": 3000}]}"#,
)
.unwrap()
}
}
});
let jobs = conn.list_jobs().await.unwrap();
assert_eq!(jobs.len(), 3);
assert_eq!(jobs[0].job_id, "job-1");
assert_eq!(jobs[0].table, "t1");
assert_eq!(jobs[0].state, "running");
assert_eq!(jobs[1].job_id, "job-2");
assert_eq!(jobs[1].state, "finished");
assert_eq!(jobs[1].created_at_millis, 2000);
assert_eq!(jobs[2].job_id, "job-3");
assert_eq!(jobs[2].state, "failed");
}
#[tokio::test]
async fn test_get_job() {
let conn = Connection::new_with_handler(|request| {
assert_eq!(request.method(), &reqwest::Method::POST);
assert_eq!(request.url().path(), "/v1/jobs/describe");
let body: serde_json::Value =
serde_json::from_slice(request.body().unwrap().as_bytes().unwrap()).unwrap();
assert_eq!(body["job_id"], "job-1");
http::Response::builder()
.status(200)
.body(
r#"{"job_id": "job-1", "job_type": "create_index", "job_state": "FAILED", "creation_ms": 1000, "spec": {"column": "vec"}, "failure": {"phase": "execute", "message": "worker died", "retryable": true}}"#,
)
.unwrap()
});
let job = conn.get_job("job-1").await.unwrap().unwrap();
assert_eq!(job.job_id, "job-1");
assert_eq!(job.job_type, "create_index");
assert_eq!(job.state, "failed");
assert_eq!(job.creation_ms, 1000);
assert_eq!(job.spec["column"], "vec");
let failure = job.failure.unwrap();
assert_eq!(failure.phase.as_deref(), Some("execute"));
assert_eq!(failure.message.as_deref(), Some("worker died"));
assert_eq!(failure.retryable, Some(true));
}
#[tokio::test]
async fn test_get_job_missing_is_none() {
let conn = Connection::new_with_handler(|_| {
http::Response::builder()
.status(404)
.body("no such job")
.unwrap()
});
assert!(conn.get_job("nope").await.unwrap().is_none());
}
#[tokio::test]
async fn test_cancel_job() {
let conn = Connection::new_with_handler(|request| {
assert_eq!(request.url().path(), "/v1/jobs/cancel");
http::Response::builder()
.status(200)
.body(r#"{"job_id": "job-1"}"#)
.unwrap()
});
assert!(conn.cancel_job("job-1").await.unwrap());
let conn = Connection::new_with_handler(|_| {
http::Response::builder()
.status(404)
.body("no such job")
.unwrap()
});
assert!(!conn.cancel_job("nope").await.unwrap());
}
#[tokio::test]
async fn test_job_history_parses_arrow_stream() {
let schema = Arc::new(Schema::new(vec![Field::new(
"state",
DataType::Utf8,
false,
)]));
let batch = RecordBatch::try_new(
schema.clone(),
vec![Arc::new(arrow_array::StringArray::from(vec![
"created", "done",
]))],
)
.unwrap();
let mut body = Vec::new();
{
let mut writer = arrow_ipc::writer::StreamWriter::try_new(&mut body, &schema).unwrap();
writer.write(&batch).unwrap();
writer.finish().unwrap();
}
let conn = Connection::new_with_handler(move |request| {
assert_eq!(request.url().path(), "/v1/jobs/query_events");
let req_body: serde_json::Value =
serde_json::from_slice(request.body().unwrap().as_bytes().unwrap()).unwrap();
assert_eq!(req_body["job_id"], "job-1");
http::Response::builder()
.status(200)
.body(body.clone())
.unwrap()
});
let batches = conn.job_history(Some("job-1")).await.unwrap();
assert_eq!(batches.len(), 1);
assert_eq!(batches[0].num_rows(), 2);
}
#[tokio::test]
async fn test_conn_job_waits_to_done() {
let polls = Arc::new(AtomicUsize::new(0));
let polls_ref = polls.clone();
let conn = Connection::new_with_handler(move |request| {
assert_eq!(request.url().path(), "/v1/jobs/describe");
let state = if polls_ref.fetch_add(1, Ordering::SeqCst) == 0 {
"IN_PROGRESS"
} else {
"DONE"
};
http::Response::builder()
.status(200)
.body(format!(
r#"{{"job_id": "job-1", "job_type": "create_index", "job_state": "{}", "creation_ms": 1}}"#,
state
))
.unwrap()
});
let job = conn.job("job-1").unwrap();
assert_eq!(job.id(), Some("job-1"));
assert_eq!(job.status().await.unwrap(), "running");
job.wait().await.unwrap();
assert_eq!(job.status().await.unwrap(), "finished");
assert!(polls.load(Ordering::SeqCst) >= 3);
}
} }
+170
View File
@@ -0,0 +1,170 @@
// SPDX-License-Identifier: Apache-2.0
// SPDX-FileCopyrightText: Copyright The LanceDB Authors
//! Tracking for server-side jobs through the `/v1/jobs` API.
use std::time::Duration;
use async_trait::async_trait;
use tokio::time::sleep;
use serde::{Deserialize, Deserializer};
use crate::error::{Error, JobFailure, Result};
use crate::job::JobHandle;
use crate::remote::client::{HttpSend, RequestResultExt, RestfulLanceDbClient};
/// Delay before the second job-state poll; doubles up to [`MAX_POLL_INTERVAL`].
const INITIAL_POLL_INTERVAL: Duration = Duration::from_millis(200);
const MAX_POLL_INTERVAL: Duration = Duration::from_secs(5);
#[derive(Debug, Clone, PartialEq, Eq)]
enum JobState {
InProgress,
Cancelled,
Failed,
Done,
/// A state this client version does not know; treated as still running
/// and reported as-is if the job never settles.
Other(String),
}
impl<'de> Deserialize<'de> for JobState {
fn deserialize<D: Deserializer<'de>>(deserializer: D) -> std::result::Result<Self, D::Error> {
Ok(Self::from(String::deserialize(deserializer)?.as_str()))
}
}
impl JobState {
/// The client vocabulary label for this state.
fn client_label(&self) -> String {
match self {
Self::InProgress => "running".to_string(),
Self::Done => "finished".to_string(),
Self::Failed => "failed".to_string(),
Self::Cancelled => "cancelled".to_string(),
Self::Other(state) => state.clone(),
}
}
}
impl From<&str> for JobState {
fn from(state: &str) -> Self {
match state {
"IN_PROGRESS" => Self::InProgress,
"CANCELLED" => Self::Cancelled,
// The server reports a timed-out job as FAILED on describe;
// accept the raw registry state too in case a future server
// stops folding it.
"FAILED" | "TIMED_OUT" => Self::Failed,
"DONE" => Self::Done,
other => Self::Other(other.to_string()),
}
}
}
/// The server's account of why a job failed. Absent from older servers, which
/// report only the terminal state.
#[derive(Deserialize)]
struct ReportedFailure {
#[serde(default)]
phase: Option<String>,
#[serde(default)]
message: Option<String>,
#[serde(default)]
retryable: Option<bool>,
}
#[derive(Deserialize)]
struct DescribeJobResponse {
job_state: JobState,
#[serde(default)]
failure: Option<ReportedFailure>,
}
pub struct RemoteJob<S: HttpSend> {
client: RestfulLanceDbClient<S>,
job_id: String,
}
impl<S: HttpSend> RemoteJob<S> {
pub fn new(client: RestfulLanceDbClient<S>, job_id: String) -> Self {
Self { client, job_id }
}
/// One `/v1/jobs/describe` round trip.
async fn describe(&self) -> Result<DescribeJobResponse> {
let request = self
.client
.post("/v1/jobs/describe")
.json(&serde_json::json!({ "job_id": self.job_id }));
let (request_id, response) = self.client.send(request).await?;
let response = self.client.check_response(&request_id, response).await?;
let body = response.text().await.err_to_http(request_id.clone())?;
let description: DescribeJobResponse =
serde_json::from_str(&body).map_err(|e| Error::Http {
source: format!("failed to parse job description: {}", e).into(),
request_id,
status_code: None,
})?;
Ok(description)
}
}
#[async_trait]
impl<S: HttpSend> JobHandle for RemoteJob<S> {
fn id(&self) -> Option<&str> {
Some(&self.job_id)
}
async fn status(&self) -> Result<String> {
Ok(self.describe().await?.job_state.client_label())
}
async fn wait(&self) -> Result<()> {
let mut interval = INITIAL_POLL_INTERVAL;
loop {
let description = self.describe().await?;
match description.job_state {
JobState::Done => return Ok(()),
JobState::Failed => {
return Err(Error::JobFailed {
job_id: Some(self.job_id.clone()),
failure: description
.failure
.map(|reported| JobFailure {
phase: reported.phase,
message: reported.message,
retryable: reported.retryable,
source: None,
})
.unwrap_or_default(),
});
}
JobState::Cancelled => {
return Err(Error::JobCancelled {
job_id: Some(self.job_id.clone()),
});
}
JobState::InProgress => {}
JobState::Other(ref state) => {
log::debug!("job {} is in unrecognized state {state}", self.job_id)
}
}
sleep(interval).await;
interval = (interval * 2).min(MAX_POLL_INTERVAL);
}
}
async fn cancel(&self) -> Result<()> {
let request = self
.client
.post("/v1/jobs/cancel")
.json(&serde_json::json!({ "job_id": self.job_id }));
let (request_id, response) = self.client.send(request).await?;
self.client
.check_response(&request_id, response)
.await
.map(|_| ())
}
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+237 -70
View File
@@ -1,7 +1,8 @@
// SPDX-License-Identifier: Apache-2.0 // SPDX-License-Identifier: Apache-2.0
// SPDX-FileCopyrightText: Copyright The LanceDB Authors // SPDX-FileCopyrightText: Copyright The LanceDB Authors
//! DataFusion ExecutionPlan for inserting data into remote LanceDB tables. //! DataFusion ExecutionPlan for streaming writes (add / merge_insert) to
//! remote LanceDB tables.
use std::sync::{Arc, Mutex}; use std::sync::{Arc, Mutex};
use std::time::{Duration, Instant}; use std::time::{Duration, Instant};
@@ -23,28 +24,57 @@ use lance::io::exec::utils::InstrumentedRecordBatchStreamAdapter;
use crate::Error; use crate::Error;
use crate::remote::ARROW_STREAM_CONTENT_TYPE; use crate::remote::ARROW_STREAM_CONTENT_TYPE;
use crate::remote::client::{HttpSend, RestfulLanceDbClient, Sender}; use crate::remote::client::{HttpSend, RestfulLanceDbClient, Sender};
use crate::remote::table::RemoteTable; use crate::remote::table::{MergeInsertRequest, REQUEST_TIMEOUT_HEADER, RemoteTable};
use crate::table::AddResult;
use crate::table::datafusion::insert::COUNT_SCHEMA; use crate::table::datafusion::insert::COUNT_SCHEMA;
use crate::table::write_progress::WriteProgressTracker; use crate::table::write_progress::WriteProgressTracker;
use crate::table::{AddResult, MergeResult};
/// ExecutionPlan for inserting data into a remote LanceDB table. /// The write operation a [`RemoteWriteExec`] performs. Both variants share the
/// same Arrow-IPC streaming body and error side-channel; only the target
/// endpoint, query parameters, and parsed result type differ.
#[derive(Debug, Clone)]
pub(crate) enum WriteOp {
/// `add`: stream to `/v1/table/{id}/insert/`, optionally overwriting.
Insert { overwrite: bool },
/// `merge_insert`: stream to `/v1/table/{id}/merge_insert/` with the merge
/// parameters carried as query params. Multipart is not supported for this
/// operation (the server has no multipart merge_insert endpoint), so an
/// `upload_id` combined with this op is a programming error.
MergeInsert {
query: MergeInsertRequest,
timeout: Option<Duration>,
},
}
/// The parsed server response for a completed write, discriminated by the
/// operation that produced it.
#[derive(Debug, Clone)]
pub(crate) enum WriteResult {
Add(AddResult),
Merge(MergeResult),
}
/// ExecutionPlan for streaming a write (add or merge_insert) to a remote
/// LanceDB table.
/// ///
/// Streams data as Arrow IPC to `/v1/table/{id}/insert/` endpoint. /// Streams data as Arrow IPC to the endpoint selected by [`WriteOp`]. Both
/// operations reuse the same error side-channel so an input stream error (e.g.
/// NaN rejection) surfaces with its original message rather than the masked
/// HTTP error Hyper produces when a request body stream fails under HTTP2.
/// ///
/// When `upload_id` is set, inserts are staged as part of a multipart write /// When `upload_id` is set, inserts are staged as part of a multipart write
/// session and the plan supports multiple partitions for parallel uploads. /// session and the plan supports multiple partitions for parallel uploads.
/// Without `upload_id`, the plan requires a single partition and commits /// Without `upload_id`, the plan requires a single partition and commits
/// immediately. /// immediately. Multipart applies to `add` only.
#[derive(Debug)] #[derive(Debug)]
pub struct RemoteInsertExec<S: HttpSend = Sender> { pub struct RemoteWriteExec<S: HttpSend = Sender> {
table_name: String, table_name: String,
identifier: String, identifier: String,
client: RestfulLanceDbClient<S>, client: RestfulLanceDbClient<S>,
input: Arc<dyn ExecutionPlan>, input: Arc<dyn ExecutionPlan>,
overwrite: bool, op: WriteOp,
properties: Arc<PlanProperties>, properties: Arc<PlanProperties>,
add_result: Arc<Mutex<Option<AddResult>>>, result: Arc<Mutex<Option<WriteResult>>>,
metrics: ExecutionPlanMetricsSet, metrics: ExecutionPlanMetricsSet,
upload_id: Option<String>, upload_id: Option<String>,
tracker: Option<Arc<WriteProgressTracker>>, tracker: Option<Arc<WriteProgressTracker>>,
@@ -61,27 +91,28 @@ pub struct RemoteInsertExec<S: HttpSend = Sender> {
max_request_duration: Option<Duration>, max_request_duration: Option<Duration>,
} }
impl<S: HttpSend + 'static> RemoteInsertExec<S> { impl<S: HttpSend + 'static> RemoteWriteExec<S> {
/// Create a new single-partition RemoteInsertExec. /// Create a new single-partition RemoteWriteExec.
pub fn new( pub fn new(
table_name: String, table_name: String,
identifier: String, identifier: String,
client: RestfulLanceDbClient<S>, client: RestfulLanceDbClient<S>,
input: Arc<dyn ExecutionPlan>, input: Arc<dyn ExecutionPlan>,
overwrite: bool, op: WriteOp,
tracker: Option<Arc<WriteProgressTracker>>, tracker: Option<Arc<WriteProgressTracker>>,
branch: Option<String>, branch: Option<String>,
) -> Self { ) -> Self {
Self::new_inner( Self::new_inner(
table_name, identifier, client, input, overwrite, None, tracker, branch, None, None, table_name, identifier, client, input, op, None, tracker, branch, None, None,
) )
} }
/// Create a multi-partition RemoteInsertExec for use with multipart writes. /// Create a multi-partition RemoteWriteExec for use with multipart writes.
/// ///
/// Each partition's insert is staged under the given `upload_id` without /// Each partition's insert is staged under the given `upload_id` without
/// committing. The caller is responsible for calling the complete (or abort) /// committing. The caller is responsible for calling the complete (or abort)
/// endpoint after all partitions finish. /// endpoint after all partitions finish. Multipart is insert-only, so the
/// op is fixed to [`WriteOp::Insert`].
#[allow(clippy::too_many_arguments)] #[allow(clippy::too_many_arguments)]
pub fn new_multipart( pub fn new_multipart(
table_name: String, table_name: String,
@@ -100,7 +131,7 @@ impl<S: HttpSend + 'static> RemoteInsertExec<S> {
identifier, identifier,
client, client,
input, input,
overwrite, WriteOp::Insert { overwrite },
Some(upload_id), Some(upload_id),
tracker, tracker,
branch, branch,
@@ -115,7 +146,7 @@ impl<S: HttpSend + 'static> RemoteInsertExec<S> {
identifier: String, identifier: String,
client: RestfulLanceDbClient<S>, client: RestfulLanceDbClient<S>,
input: Arc<dyn ExecutionPlan>, input: Arc<dyn ExecutionPlan>,
overwrite: bool, op: WriteOp,
upload_id: Option<String>, upload_id: Option<String>,
tracker: Option<Arc<WriteProgressTracker>>, tracker: Option<Arc<WriteProgressTracker>>,
branch: Option<String>, branch: Option<String>,
@@ -140,9 +171,9 @@ impl<S: HttpSend + 'static> RemoteInsertExec<S> {
identifier, identifier,
client, client,
input, input,
overwrite, op,
properties: Arc::new(properties), properties: Arc::new(properties),
add_result: Arc::new(Mutex::new(None)), result: Arc::new(Mutex::new(None)),
metrics: ExecutionPlanMetricsSet::new(), metrics: ExecutionPlanMetricsSet::new(),
upload_id, upload_id,
tracker, tracker,
@@ -152,14 +183,30 @@ impl<S: HttpSend + 'static> RemoteInsertExec<S> {
} }
} }
/// Get the add result after execution. /// Get the add result after execution, if this exec ran an insert.
// TODO: this will be used when we wire this up to Table::add().
#[allow(dead_code)]
pub fn add_result(&self) -> Option<AddResult> { pub fn add_result(&self) -> Option<AddResult> {
self.add_result match self
.result
.lock() .lock()
.unwrap_or_else(|e| e.into_inner()) .unwrap_or_else(|e| e.into_inner())
.clone() .clone()
{
Some(WriteResult::Add(r)) => Some(r),
_ => None,
}
}
/// Get the merge result after execution, if this exec ran a merge_insert.
pub fn merge_result(&self) -> Option<MergeResult> {
match self
.result
.lock()
.unwrap_or_else(|e| e.into_inner())
.clone()
{
Some(WriteResult::Merge(r)) => Some(r),
_ => None,
}
} }
/// Stream the input into an HTTP body as an Arrow IPC stream, capturing any /// Stream the input into an HTTP body as an Arrow IPC stream, capturing any
@@ -464,24 +511,24 @@ impl<S: HttpSend + 'static> PartRequestCtx<'_, S> {
} }
} }
impl<S: HttpSend + 'static> DisplayAs for RemoteInsertExec<S> { impl<S: HttpSend + 'static> DisplayAs for RemoteWriteExec<S> {
fn fmt_as(&self, t: DisplayFormatType, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { fn fmt_as(&self, t: DisplayFormatType, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match t { match t {
DisplayFormatType::Default | DisplayFormatType::Verbose => { DisplayFormatType::Default | DisplayFormatType::Verbose => {
write!( write!(f, "RemoteWriteExec: table={}, op=", self.table_name)?;
f, match &self.op {
"RemoteInsertExec: table={}, overwrite={}", WriteOp::Insert { overwrite } => write!(f, "insert, overwrite={}", overwrite),
self.table_name, self.overwrite WriteOp::MergeInsert { .. } => write!(f, "merge_insert"),
) }
} }
DisplayFormatType::TreeRender => { DisplayFormatType::TreeRender => {
write!(f, "RemoteInsertExec") write!(f, "RemoteWriteExec")
} }
} }
} }
} }
impl<S: HttpSend + 'static> ExecutionPlan for RemoteInsertExec<S> { impl<S: HttpSend + 'static> ExecutionPlan for RemoteWriteExec<S> {
fn name(&self) -> &str { fn name(&self) -> &str {
Self::static_name() Self::static_name()
} }
@@ -516,15 +563,18 @@ impl<S: HttpSend + 'static> ExecutionPlan for RemoteInsertExec<S> {
) -> DataFusionResult<Arc<dyn ExecutionPlan>> { ) -> DataFusionResult<Arc<dyn ExecutionPlan>> {
if children.len() != 1 { if children.len() != 1 {
return Err(DataFusionError::Internal( return Err(DataFusionError::Internal(
"RemoteInsertExec requires exactly one child".to_string(), "RemoteWriteExec requires exactly one child".to_string(),
)); ));
} }
// Building a fresh exec (with a new, empty `result`) is what makes the
// outer rescannable retry loop work: `reset_state()` clears the captured
// result so a re-execution starts clean.
Ok(Arc::new(Self::new_inner( Ok(Arc::new(Self::new_inner(
self.table_name.clone(), self.table_name.clone(),
self.identifier.clone(), self.identifier.clone(),
self.client.clone(), self.client.clone(),
children[0].clone(), children[0].clone(),
self.overwrite, self.op.clone(),
self.upload_id.clone(), self.upload_id.clone(),
self.tracker.clone(), self.tracker.clone(),
self.branch.clone(), self.branch.clone(),
@@ -540,11 +590,19 @@ impl<S: HttpSend + 'static> ExecutionPlan for RemoteInsertExec<S> {
) -> DataFusionResult<SendableRecordBatchStream> { ) -> DataFusionResult<SendableRecordBatchStream> {
if self.upload_id.is_none() && partition != 0 { if self.upload_id.is_none() && partition != 0 {
return Err(DataFusionError::Internal( return Err(DataFusionError::Internal(
"RemoteInsertExec only supports single partition execution without upload_id" "RemoteWriteExec only supports single partition execution without upload_id"
.to_string(), .to_string(),
)); ));
} }
// Multipart is insert-only: the server has no multipart merge_insert
// endpoint, so a merge_insert with an upload_id is a programming error.
if self.upload_id.is_some() && matches!(self.op, WriteOp::MergeInsert { .. }) {
return Err(DataFusionError::Internal(
"merge_insert does not support multipart (upload_id) writes".to_string(),
));
}
let input_stream = self.input.execute(partition, context)?; let input_stream = self.input.execute(partition, context)?;
let input_schema = input_stream.schema(); let input_schema = input_stream.schema();
let input_stream: SendableRecordBatchStream = let input_stream: SendableRecordBatchStream =
@@ -556,8 +614,8 @@ impl<S: HttpSend + 'static> ExecutionPlan for RemoteInsertExec<S> {
)); ));
let client = self.client.clone(); let client = self.client.clone();
let identifier = self.identifier.clone(); let identifier = self.identifier.clone();
let overwrite = self.overwrite; let op = self.op.clone();
let add_result = self.add_result.clone(); let result_slot = self.result.clone();
let table_name = self.table_name.clone(); let table_name = self.table_name.clone();
let upload_id = self.upload_id.clone(); let upload_id = self.upload_id.clone();
let tracker = self.tracker.clone(); let tracker = self.tracker.clone();
@@ -568,10 +626,12 @@ impl<S: HttpSend + 'static> ExecutionPlan for RemoteInsertExec<S> {
let stream = futures::stream::once(async move { let stream = futures::stream::once(async move {
// Multipart writes with a byte budget split the partition into // Multipart writes with a byte budget split the partition into
// several bounded, still-streamed requests so no single request // several bounded, still-streamed requests so no single request
// stays open long enough to hit the client read timeout. // stays open long enough to hit the client read timeout. This path
// is insert-only (guarded above).
if let (Some(upload_id), Some(max_bytes)) = if let (Some(upload_id), Some(max_bytes)) =
(upload_id.as_deref(), max_bytes_per_request) (upload_id.as_deref(), max_bytes_per_request)
{ {
let overwrite = matches!(op, WriteOp::Insert { overwrite: true });
let ctx = PartRequestCtx { let ctx = PartRequestCtx {
client: &client, client: &client,
identifier: &identifier, identifier: &identifier,
@@ -592,16 +652,36 @@ impl<S: HttpSend + 'static> ExecutionPlan for RemoteInsertExec<S> {
)?); )?);
} }
let mut request = client // Build the request for the selected operation. Both endpoints take
.post(&format!("/v1/table/{}/insert/", identifier)) // an Arrow-IPC streaming body and reuse the same error side-channel.
.header(CONTENT_TYPE, ARROW_STREAM_CONTENT_TYPE); let mut request = match &op {
WriteOp::Insert { overwrite } => {
let mut request = client
.post(&format!("/v1/table/{}/insert/", identifier))
.header(CONTENT_TYPE, ARROW_STREAM_CONTENT_TYPE);
if *overwrite {
request = request.query(&[("mode", "overwrite")]);
}
if let Some(ref uid) = upload_id {
request = request.query(&[("upload_id", uid.as_str())]);
}
request
}
WriteOp::MergeInsert { query, timeout } => {
let mut request = client
.post(&format!("/v1/table/{}/merge_insert/", identifier))
.query(query)
.header(CONTENT_TYPE, ARROW_STREAM_CONTENT_TYPE);
if let Some(timeout) = timeout {
// (If it doesn't fit into u64, it's not worth sending anyways.)
if let Ok(timeout_ms) = u64::try_from(timeout.as_millis()) {
request = request.header(REQUEST_TIMEOUT_HEADER, timeout_ms);
}
}
request
}
};
if overwrite {
request = request.query(&[("mode", "overwrite")]);
}
if let Some(ref uid) = upload_id {
request = request.query(&[("upload_id", uid.as_str())]);
}
if let Some(ref b) = branch { if let Some(ref b) = branch {
request = request.query(&[("branch", b.as_str())]); request = request.query(&[("branch", b.as_str())]);
} }
@@ -635,6 +715,8 @@ impl<S: HttpSend + 'static> ExecutionPlan for RemoteInsertExec<S> {
// If the request failed due to an input stream error, surface the // If the request failed due to an input stream error, surface the
// original error (e.g. NaN rejection) instead of the HTTP error. // original error (e.g. NaN rejection) instead of the HTTP error.
// This is the crux of the #2339 fix: Hyper silently swallows body
// stream errors under HTTP2, so we recover the original here.
if let Ok(stream_err) = error_rx.try_recv() { if let Ok(stream_err) = error_rx.try_recv() {
return Err(stream_err); return Err(stream_err);
} }
@@ -642,7 +724,7 @@ impl<S: HttpSend + 'static> ExecutionPlan for RemoteInsertExec<S> {
let (request_id, response) = result?; let (request_id, response) = result?;
// For multipart writes, the staging response is not the final // For multipart writes, the staging response is not the final
// version. Only parse AddResult for non-multipart inserts. // version. Only parse the result for non-multipart writes.
if upload_id.is_none() { if upload_id.is_none() {
let body_text = response.text().await.map_err(|e| { let body_text = response.text().await.map_err(|e| {
DataFusionError::External(Box::new(Error::Http { DataFusionError::External(Box::new(Error::Http {
@@ -652,21 +734,44 @@ impl<S: HttpSend + 'static> ExecutionPlan for RemoteInsertExec<S> {
})) }))
})?; })?;
let parsed_result = if body_text.trim().is_empty() { let parsed_result = match &op {
// Backward compatible with old servers WriteOp::Insert { .. } => {
AddResult { version: 0 } let add = if body_text.trim().is_empty() {
} else { // Backward compatible with old servers
serde_json::from_str(&body_text).map_err(|e| { AddResult { version: 0 }
DataFusionError::External(Box::new(Error::Http { } else {
source: format!("Failed to parse add response: {}", e).into(), serde_json::from_str(&body_text).map_err(|e| {
request_id: request_id.clone(), DataFusionError::External(Box::new(Error::Http {
status_code: None, source: format!("Failed to parse add response: {}", e).into(),
})) request_id: request_id.clone(),
})? status_code: None,
}))
})?
};
WriteResult::Add(add)
}
WriteOp::MergeInsert { .. } => {
let merge = if body_text.trim().is_empty() {
// Backward compatible with old servers
MergeResult::default()
} else {
serde_json::from_str(&body_text).map_err(|e| {
DataFusionError::External(Box::new(Error::Http {
source: format!("Failed to parse merge_insert response: {}", e)
.into(),
request_id: request_id.clone(),
status_code: None,
}))
})?
};
WriteResult::Merge(merge)
}
}; };
let mut res_lock = add_result.lock().map_err(|_| { let mut res_lock = result_slot.lock().map_err(|_| {
DataFusionError::Execution("Failed to acquire lock for add_result".to_string()) DataFusionError::Execution(
"Failed to acquire lock for write result".to_string(),
)
})?; })?;
*res_lock = Some(parsed_result); *res_lock = Some(parsed_result);
} else { } else {
@@ -680,7 +785,7 @@ impl<S: HttpSend + 'static> ExecutionPlan for RemoteInsertExec<S> {
})?; })?;
} }
// Return a single batch with count 0 (actual count is tracked in add_result) // Return a single batch with count 0 (actual count is tracked in result)
let count_array: ArrayRef = Arc::new(UInt64Array::from(vec![0u64])); let count_array: ArrayRef = Arc::new(UInt64Array::from(vec![0u64]));
let batch = RecordBatch::try_new(COUNT_SCHEMA.clone(), vec![count_array])?; let batch = RecordBatch::try_new(COUNT_SCHEMA.clone(), vec![count_array])?;
Ok::<_, DataFusionError>(batch) Ok::<_, DataFusionError>(batch)
@@ -711,9 +816,11 @@ mod tests {
use std::sync::atomic::{AtomicUsize, Ordering}; use std::sync::atomic::{AtomicUsize, Ordering};
use std::sync::{Arc, Mutex}; use std::sync::{Arc, Mutex};
use super::RemoteInsertExec; use super::RemoteWriteExec;
use super::WriteOp;
use crate::Table; use crate::Table;
use crate::remote::ARROW_STREAM_CONTENT_TYPE; use crate::remote::ARROW_STREAM_CONTENT_TYPE;
use crate::remote::table::MergeInsertRequest;
use crate::table::datafusion::BaseTableAdapter; use crate::table::datafusion::BaseTableAdapter;
fn schema_json() -> &'static str { fn schema_json() -> &'static str {
@@ -1028,7 +1135,7 @@ mod tests {
let input = input_plan_from_batches(schema, batches).await; let input = input_plan_from_batches(schema, batches).await;
// A 1-byte budget forces every batch into its own part. // A 1-byte budget forces every batch into its own part.
let exec = RemoteInsertExec::new_multipart( let exec = RemoteWriteExec::new_multipart(
"my_table".to_string(), "my_table".to_string(),
"my_table".to_string(), "my_table".to_string(),
client, client,
@@ -1068,7 +1175,7 @@ mod tests {
// A large byte budget and no time limit keep the whole partition in a // A large byte budget and no time limit keep the whole partition in a
// single part. // single part.
let exec = RemoteInsertExec::new_multipart( let exec = RemoteWriteExec::new_multipart(
"my_table".to_string(), "my_table".to_string(),
"my_table".to_string(), "my_table".to_string(),
client, client,
@@ -1109,7 +1216,7 @@ mod tests {
// A large byte budget but a tiny duration budget: writing and sending // A large byte budget but a tiny duration budget: writing and sending
// one batch already takes longer than the limit, so each batch is cut // one batch already takes longer than the limit, so each batch is cut
// into its own part on the time check rather than the byte check. // into its own part on the time check rather than the byte check.
let exec = RemoteInsertExec::new_multipart( let exec = RemoteWriteExec::new_multipart(
"my_table".to_string(), "my_table".to_string(),
"my_table".to_string(), "my_table".to_string(),
client, client,
@@ -1144,7 +1251,7 @@ mod tests {
// write relies on another partition having data to commit. // write relies on another partition having data to commit.
let input = input_plan_from_batches(schema, vec![]).await; let input = input_plan_from_batches(schema, vec![]).await;
let exec = RemoteInsertExec::new_multipart( let exec = RemoteWriteExec::new_multipart(
"my_table".to_string(), "my_table".to_string(),
"my_table".to_string(), "my_table".to_string(),
client, client,
@@ -1184,7 +1291,7 @@ mod tests {
let input = input_plan_from_batches(schema, batches).await; let input = input_plan_from_batches(schema, batches).await;
// A 1-byte budget forces every batch into its own part. // A 1-byte budget forces every batch into its own part.
let exec = RemoteInsertExec::new_multipart( let exec = RemoteWriteExec::new_multipart(
"my_table".to_string(), "my_table".to_string(),
"my_table".to_string(), "my_table".to_string(),
client, client,
@@ -1233,7 +1340,7 @@ mod tests {
]; ];
let input = input_plan_from_partitions(schema, partitions).await; let input = input_plan_from_partitions(schema, partitions).await;
let exec = RemoteInsertExec::new_multipart( let exec = RemoteWriteExec::new_multipart(
"my_table".to_string(), "my_table".to_string(),
"my_table".to_string(), "my_table".to_string(),
client, client,
@@ -1267,7 +1374,7 @@ mod tests {
// A large byte budget keeps the good batch and the following error in // A large byte budget keeps the good batch and the following error in
// the same part, exercising the mid-part abort path. // the same part, exercising the mid-part abort path.
let input: Arc<dyn ExecutionPlan> = Arc::new(ErroringExec::new()); let input: Arc<dyn ExecutionPlan> = Arc::new(ErroringExec::new());
let exec = RemoteInsertExec::new_multipart( let exec = RemoteWriteExec::new_multipart(
"my_table".to_string(), "my_table".to_string(),
"my_table".to_string(), "my_table".to_string(),
client, client,
@@ -1297,6 +1404,66 @@ mod tests {
); );
} }
#[tokio::test]
async fn test_merge_insert_input_error_surfaces_original() {
// Regression test for #2339 on the single-request merge_insert path.
// When the input stream errors mid-body, Hyper masks it under HTTP2 as a
// generic "stream error sent by user" message. The error side-channel
// must recover and surface the original DataFusion error instead.
use futures::StreamExt;
let client = crate::remote::client::test_utils::client_with_handler(|request| {
assert_eq!(request.url().path(), "/v1/table/my_table/merge_insert/");
http::Response::builder()
.status(200)
.body(
r#"{"version": 2, "num_updated_rows": 0, "num_inserted_rows": 0, "num_deleted_rows": 0}"#
.to_string(),
)
.unwrap()
});
let query = MergeInsertRequest {
on: "id".to_string(),
when_matched_update_all: false,
when_matched_update_all_filt: None,
when_not_matched_insert_all: false,
when_not_matched_by_source_delete: false,
when_not_matched_by_source_delete_filt: None,
use_index: true,
use_lsm: None,
};
let input: Arc<dyn ExecutionPlan> = Arc::new(ErroringExec::new());
let exec = RemoteWriteExec::new(
"my_table".to_string(),
"my_table".to_string(),
client,
input,
WriteOp::MergeInsert {
query,
timeout: None,
},
None,
None,
);
let mut stream = exec.execute(0, Arc::new(TaskContext::default())).unwrap();
let mut err = None;
while let Some(item) = stream.next().await {
if let Err(e) = item {
err = Some(e);
break;
}
}
let err = err.expect("expected the input stream error to surface");
assert!(
err.to_string().contains("boom"),
"expected original input error, got: {err}"
);
}
#[tokio::test] #[tokio::test]
async fn test_multipart_records_progress_within_a_part() { async fn test_multipart_records_progress_within_a_part() {
use crate::table::write_progress::{ProgressCallback, WriteProgress, WriteProgressTracker}; use crate::table::write_progress::{ProgressCallback, WriteProgress, WriteProgressTracker};
@@ -1327,7 +1494,7 @@ mod tests {
// A large byte budget keeps all three batches in one part; smooth // A large byte budget keeps all three batches in one part; smooth
// progress therefore requires bytes to be reported per chunk rather than // progress therefore requires bytes to be reported per chunk rather than
// once when the part completes. // once when the part completes.
let exec = RemoteInsertExec::new_multipart( let exec = RemoteWriteExec::new_multipart(
"my_table".to_string(), "my_table".to_string(),
"my_table".to_string(), "my_table".to_string(),
client, client,
+177 -37
View File
@@ -3,6 +3,7 @@
//! LanceDB Table APIs //! LanceDB Table APIs
use crate::blob::BlobFile;
use arrow_array::{LargeBinaryArray, RecordBatch, RecordBatchReader}; use arrow_array::{LargeBinaryArray, RecordBatch, RecordBatchReader};
use arrow_schema::{Schema, SchemaRef}; use arrow_schema::{Schema, SchemaRef};
use async_trait::async_trait; use async_trait::async_trait;
@@ -12,7 +13,6 @@ use datafusion_physical_plan::ExecutionPlan;
use datafusion_physical_plan::display::DisplayableExecutionPlan; use datafusion_physical_plan::display::DisplayableExecutionPlan;
use futures::StreamExt; use futures::StreamExt;
use futures::stream::FuturesUnordered; use futures::stream::FuturesUnordered;
use lance::dataset::BlobFile;
pub use lance::dataset::ColumnAlteration; pub use lance::dataset::ColumnAlteration;
pub use lance::dataset::NewColumnTransform; pub use lance::dataset::NewColumnTransform;
pub use lance::dataset::ReadParams; pub use lance::dataset::ReadParams;
@@ -50,12 +50,14 @@ use crate::DistanceType;
use crate::blob::BlobRangeRequest; use crate::blob::BlobRangeRequest;
use crate::data::scannable::{PeekedScannable, Scannable, estimate_write_partitions}; use crate::data::scannable::{PeekedScannable, Scannable, estimate_write_partitions};
use crate::database::Database; use crate::database::Database;
use crate::database::listing::LANCE_FILE_EXTENSION;
use crate::database::read_freshness::TableFreshness; use crate::database::read_freshness::TableFreshness;
use crate::embeddings::{EmbeddingDefinition, EmbeddingRegistry, MemoryRegistry}; use crate::embeddings::{EmbeddingDefinition, EmbeddingRegistry, MemoryRegistry};
use crate::error::{Error, Result}; use crate::error::{Error, Result};
use crate::index::IndexStatistics; use crate::index::IndexStatistics;
use crate::index::{Index, IndexBuilder}; use crate::index::{Index, IndexBuilder};
use crate::index::{IndexConfig, IndexStatisticsImpl, IndexType}; use crate::index::{IndexConfig, IndexStatisticsImpl, IndexType};
use crate::job::Job;
use crate::query::{IntoQueryVector, Query, QueryExecutionOptions, TakeQuery, VectorQuery}; use crate::query::{IntoQueryVector, Query, QueryExecutionOptions, TakeQuery, VectorQuery};
use crate::table::datafusion::insert::InsertExec; use crate::table::datafusion::insert::InsertExec;
use crate::utils::{PatchReadParam, PatchWriteParam, resolve_arrow_field_path}; use crate::utils::{PatchReadParam, PatchWriteParam, resolve_arrow_field_path};
@@ -63,6 +65,7 @@ use crate::utils::{PatchReadParam, PatchWriteParam, resolve_arrow_field_path};
use self::dataset::DatasetConsistencyWrapper; use self::dataset::DatasetConsistencyWrapper;
use self::merge::MergeInsertBuilder; use self::merge::MergeInsertBuilder;
pub mod add_columns;
mod add_data; mod add_data;
pub mod branch_merge; pub mod branch_merge;
mod create_index; mod create_index;
@@ -77,6 +80,7 @@ pub mod schema_evolution;
pub mod update; pub mod update;
pub mod write_progress; pub mod write_progress;
use crate::index::waiter::wait_for_index; use crate::index::waiter::wait_for_index;
pub use add_columns::AddColumnsBuilder;
#[cfg(feature = "remote")] #[cfg(feature = "remote")]
pub(crate) use add_data::PreprocessingOutput; pub(crate) use add_data::PreprocessingOutput;
pub use add_data::{AddDataBuilder, AddDataMode, AddResult, NaNVectorBehavior}; pub use add_data::{AddDataBuilder, AddDataMode, AddResult, NaNVectorBehavior};
@@ -146,6 +150,55 @@ pub(crate) fn map_namespace_lance_error(err: lance::Error, table_name: &str) ->
} }
} }
/// Map a `lance::Error::DatasetNotFound` for the table at `uri` into a `lancedb::Error`.
///
/// Lance reports "there is nothing at this location" and "there is a table directory
/// here but nothing loadable inside it" with the same error. Only the first is a
/// `TableNotFound`: a `<name>.lance` directory left behind by an interrupted drop and
/// re-create is still reported by `Connection::table_names`, so callers need to be able
/// to tell "never existed" from "exists but is broken".
///
/// See <https://github.com/lancedb/lancedb/issues/3127>.
async fn map_dataset_not_found(
uri: &str,
name: &str,
params: ReadParams,
err: lance::Error,
) -> Error {
let name = name.to_string();
let source = Box::new(err);
if table_dir_exists(uri, params).await.unwrap_or(false) {
Error::TableCorrupted { name, source }
} else {
Error::TableNotFound { name, source }
}
}
/// Whether a table directory is present at `uri`, even though no dataset could be
/// loaded from it.
///
/// This looks for a `<name>.lance` entry in the parent directory, which is exactly what
/// `ListingDatabase::table_names` lists, so the two APIs agree on whether a table is
/// present. Probing `uri` itself would not work: object stores have no empty
/// directories to probe, and on a local filesystem the interesting case is precisely an
/// empty directory.
async fn table_dir_exists(uri: &str, params: ReadParams) -> Result<bool> {
let (object_store, path, _) = DatasetBuilder::from_uri(uri)
.with_read_params(params)
.build_object_store()
.await?;
// Only `*.lance` entries are ever reported as tables, so nothing else can produce
// the list-then-open mismatch this guards against.
if path.extension() != Some(LANCE_FILE_EXTENSION) {
return Ok(false);
}
let (Some(parent), Some(dir_name)) = (path.parent(), path.filename()) else {
return Ok(false);
};
let entries = object_store.read_dir(parent).await?;
Ok(entries.iter().any(|entry| entry.as_str() == dir_name))
}
/// Defines the type of column /// Defines the type of column
#[derive(Debug, Clone, Serialize, Deserialize)] #[derive(Debug, Clone, Serialize, Deserialize)]
pub enum ColumnKind { pub enum ColumnKind {
@@ -563,6 +616,9 @@ pub trait BaseTable: std::fmt::Display + std::fmt::Debug + Send + Sync {
async fn update(&self, update: UpdateBuilder) -> Result<UpdateResult>; async fn update(&self, update: UpdateBuilder) -> Result<UpdateResult>;
/// Create an index on the provided column(s). /// Create an index on the provided column(s).
async fn create_index(&self, index: IndexBuilder) -> Result<()>; async fn create_index(&self, index: IndexBuilder) -> Result<()>;
/// Starts index creation, returning a handle to the resulting job.
async fn create_index_async(&self, index: IndexBuilder) -> Result<Job>;
/// List the indices on the table. /// List the indices on the table.
async fn list_indices(&self) -> Result<Vec<IndexConfig>>; async fn list_indices(&self) -> Result<Vec<IndexConfig>>;
/// Drop an index from the table. /// Drop an index from the table.
@@ -1566,12 +1622,8 @@ impl Table {
} }
/// Add new columns to the table, providing values to fill in. /// Add new columns to the table, providing values to fill in.
pub async fn add_columns( pub fn add_columns(&self) -> AddColumnsBuilder {
&self, AddColumnsBuilder::new(self.inner.clone())
transforms: NewColumnTransform,
read_columns: Option<Vec<String>>,
) -> Result<AddColumnsResult> {
self.inner.add_columns(transforms, read_columns).await
} }
/// Change a column's name or nullability. /// Change a column's name or nullability.
@@ -2240,6 +2292,8 @@ impl NativeTable {
None => false, None => false,
}; };
// Kept so that a `DatasetNotFound` can be re-checked against storage below.
let recovery_params = params.clone();
let mut builder = DatasetBuilder::from_uri(uri).with_read_params(params); let mut builder = DatasetBuilder::from_uri(uri).with_read_params(params);
// Set up commit handler when managed_versioning is enabled // Set up commit handler when managed_versioning is enabled
@@ -2255,13 +2309,13 @@ impl NativeTable {
builder = builder.with_commit_handler(commit_handler); builder = builder.with_commit_handler(commit_handler);
} }
let dataset = builder.load().await.map_err(|e| match e { let dataset = match builder.load().await {
lance::Error::DatasetNotFound { .. } => Error::TableNotFound { Ok(dataset) => dataset,
name: name.to_string(), Err(e @ lance::Error::DatasetNotFound { .. }) => {
source: Box::new(e), return Err(map_dataset_not_found(uri, name, recovery_params, e).await);
}, }
e => e.into(), Err(e) => return Err(e.into()),
})?; };
let dataset = DatasetConsistencyWrapper::new_latest(dataset, read_consistency_interval); let dataset = DatasetConsistencyWrapper::new_latest(dataset, read_consistency_interval);
let id = Self::build_id(&namespace, name); let id = Self::build_id(&namespace, name);
@@ -3009,29 +3063,18 @@ impl BaseTable for NativeTable {
} }
async fn create_index(&self, opts: IndexBuilder) -> Result<()> { async fn create_index(&self, opts: IndexBuilder) -> Result<()> {
if opts.columns.len() != 1 { let prepared = self.prepare_index(&opts).await?;
return Err(Error::Schema { self.build_index(opts, prepared).await
message: "Multi-column (composite) indices are not yet supported".to_string(), }
});
}
self.dataset.ensure_mutable()?;
let mut dataset = (*self.dataset.get().await?).clone();
let (column, field) = Self::resolve_index_field(dataset.schema(), &opts.columns[0])?;
let lance_idx_params = self.make_index_params(&field, opts.index.clone()).await?; async fn create_index_async(&self, opts: IndexBuilder) -> Result<Job> {
let index_type = self.get_index_type_for_field(&field, &opts.index); // Prepare before spawning so bad input is reported by this call rather
let columns = [column.as_str()]; // than only by the job.
let mut builder = dataset let prepared = self.prepare_index(&opts).await?;
.create_index_builder(&columns, index_type, lance_idx_params.as_ref()) let table = self.clone();
.train(opts.train) Ok(Job::spawned(tokio::spawn(async move {
.replace(opts.replace); table.build_index(opts, prepared).await
})))
if let Some(name) = opts.name {
builder = builder.name(name);
}
builder.await?;
self.dataset.update(dataset);
Ok(())
} }
async fn drop_index(&self, index_name: &str) -> Result<()> { async fn drop_index(&self, index_name: &str) -> Result<()> {
@@ -3584,6 +3627,103 @@ mod tests {
assert!(matches!(table.unwrap_err(), Error::TableNotFound { .. })); assert!(matches!(table.unwrap_err(), Error::TableNotFound { .. }));
} }
#[tokio::test]
async fn test_open_not_found_missing_lance_dir() {
let tmp_dir = tempdir().unwrap();
let dataset_path = tmp_dir.path().join("test.lance");
let err = NativeTable::open(dataset_path.to_str().unwrap())
.await
.unwrap_err();
assert!(
matches!(&err, Error::TableNotFound { name, .. } if name == "test"),
"got {err:?}"
);
}
/// Write a table and then break it, leaving the `<name>.lance` directory in place.
///
/// `remove_all` reproduces an interrupted drop + re-create (the directory is left
/// empty); otherwise only the manifests are removed, leaving the data files behind.
async fn write_then_corrupt_table(dir: &std::path::Path, remove_all: bool) -> String {
let dataset_path = dir.join("test.lance");
let uri = dataset_path.to_str().unwrap().to_string();
let batch = make_test_batches();
let reader = RecordBatchIterator::new(vec![Ok(batch.clone())], batch.schema());
Dataset::write(reader, &uri, None).await.unwrap();
if remove_all {
for entry in std::fs::read_dir(&dataset_path).unwrap() {
let entry = entry.unwrap();
if entry.file_type().unwrap().is_dir() {
std::fs::remove_dir_all(entry.path()).unwrap();
} else {
std::fs::remove_file(entry.path()).unwrap();
}
}
assert_eq!(std::fs::read_dir(&dataset_path).unwrap().count(), 0);
} else {
let versions = dataset_path.join("_versions");
assert!(versions.is_dir(), "expected manifests under {versions:?}");
std::fs::remove_dir_all(&versions).unwrap();
assert!(std::fs::read_dir(&dataset_path).unwrap().count() > 0);
}
uri
}
#[tokio::test]
async fn test_open_corrupt_empty_dir() {
let tmp_dir = tempdir().unwrap();
let uri = write_then_corrupt_table(tmp_dir.path(), true).await;
let err = NativeTable::open(&uri).await.unwrap_err();
assert!(
matches!(&err, Error::TableCorrupted { name, .. } if name == "test"),
"got {err:?}"
);
}
#[tokio::test]
async fn test_open_corrupt_missing_manifest() {
let tmp_dir = tempdir().unwrap();
let uri = write_then_corrupt_table(tmp_dir.path(), false).await;
let err = NativeTable::open(&uri).await.unwrap_err();
assert!(
matches!(&err, Error::TableCorrupted { name, .. } if name == "test"),
"got {err:?}"
);
}
/// A table listed by `table_names()` must not be reported as missing by
/// `open_table()`. See <https://github.com/lancedb/lancedb/issues/3127>.
#[tokio::test]
async fn test_open_table_corrupt_is_still_listed() {
let tmp_dir = tempdir().unwrap();
let db = connect(tmp_dir.path().to_str().unwrap())
.execute()
.await
.unwrap();
write_then_corrupt_table(tmp_dir.path(), true).await;
assert_eq!(
db.table_names().execute().await.unwrap(),
vec!["test".to_string()]
);
let err = db.open_table("test").execute().await.unwrap_err();
assert!(
matches!(&err, Error::TableCorrupted { name, .. } if name == "test"),
"got {err:?}"
);
assert!(
err.to_string().contains("exists but could not be loaded"),
"got {err}"
);
}
#[test] #[test]
#[cfg(not(windows))] #[cfg(not(windows))]
fn test_object_store_path() { fn test_object_store_path() {
@@ -4591,7 +4731,7 @@ mod tests {
.set_lsm_write_spec(LsmWriteSpec::bucket("id", bad)) .set_lsm_write_spec(LsmWriteSpec::bucket("id", bad))
.await .await
.expect_err("should reject"); .expect_err("should reject");
assert!(matches!(err, Error::Lance { .. }), "got {:?}", err); assert!(matches!(err, Error::InvalidInput { .. }), "got {:?}", err);
} }
// Happy path: install spec; verify MemWAL details record it. // Happy path: install spec; verify MemWAL details record it.
+161
View File
@@ -0,0 +1,161 @@
// SPDX-License-Identifier: Apache-2.0
// SPDX-FileCopyrightText: Copyright The LanceDB Authors
//! Builder for adding columns to a table.
use std::sync::Arc;
use lance::dataset::NewColumnTransform;
use super::BaseTable;
use super::schema_evolution::AddColumnsResult;
use crate::{Error, Result};
/// Adds columns to a table. See [`Table::add_columns`](super::Table::add_columns).
pub struct AddColumnsBuilder {
parent: Arc<dyn BaseTable>,
transform: Option<NewColumnTransform>,
read_columns: Option<Vec<String>>,
}
impl std::fmt::Debug for AddColumnsBuilder {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.debug_struct("AddColumnsBuilder")
.field("parent", &self.parent)
.field("has_transform", &self.transform.is_some())
.field("read_columns", &self.read_columns)
.finish()
}
}
impl AddColumnsBuilder {
pub(crate) fn new(parent: Arc<dyn BaseTable>) -> Self {
Self {
parent,
transform: None,
read_columns: None,
}
}
/// Set how the new columns' values are produced. Required.
pub fn transform(mut self, transform: NewColumnTransform) -> Self {
self.transform = Some(transform);
self
}
/// Limit which existing columns a [`NewColumnTransform::BatchUDF`] mapper
/// receives. Every other transform determines what it reads, so setting
/// this alongside one is an error rather than a silent no-op.
pub fn read_columns(mut self, columns: impl IntoIterator<Item = impl Into<String>>) -> Self {
self.read_columns = Some(columns.into_iter().map(Into::into).collect());
self
}
/// Add the columns.
pub async fn execute(self) -> Result<AddColumnsResult> {
let Self {
parent,
transform,
read_columns,
} = self;
let Some(transform) = transform else {
return Err(Error::InvalidInput {
message: "add_columns requires a transform".into(),
});
};
if read_columns.is_some() && !matches!(transform, NewColumnTransform::BatchUDF(_)) {
return Err(Error::InvalidInput {
message: "read_columns applies only to a BatchUDF transform; \
every other transform determines what it reads"
.into(),
});
}
parent.add_columns(transform, read_columns).await
}
}
#[cfg(test)]
mod tests {
use std::sync::Arc;
use arrow_array::{Int32Array, RecordBatch, record_batch};
use arrow_schema::{DataType, Field, Schema};
use lance::dataset::{BatchUDF, NewColumnTransform};
use crate::Table;
use crate::connect;
async fn table_with_two_columns(name: &str) -> Table {
let conn = connect("memory://").execute().await.unwrap();
let batch = record_batch!(("x", Int32, [1, 2, 3]), ("y", Int32, [10, 20, 30])).unwrap();
conn.create_table(name, batch).execute().await.unwrap()
}
#[tokio::test]
async fn test_requires_a_transform() {
let table = table_with_two_columns("no_transform").await;
let err = table.add_columns().execute().await.unwrap_err();
assert!(
err.to_string().contains("requires a transform"),
"got: {err}"
);
}
#[tokio::test]
async fn test_read_columns_with_sql_expressions_is_rejected() {
let table = table_with_two_columns("read_cols_sql").await;
let err = table
.add_columns()
.transform(NewColumnTransform::SqlExpressions(vec![(
"doubled".into(),
"x * 2".into(),
)]))
.read_columns(["x"])
.execute()
.await
.unwrap_err();
assert!(err.to_string().contains("BatchUDF"), "got: {err}");
let schema = table.schema().await.unwrap();
assert!(
schema.field_with_name("doubled").is_err(),
"a rejected call must not commit"
);
}
#[tokio::test]
async fn test_read_columns_limits_what_a_batch_udf_sees() {
let table = table_with_two_columns("read_cols_udf").await;
let output_schema = Arc::new(Schema::new(vec![Field::new("sum", DataType::Int32, true)]));
let mapper_schema = output_schema.clone();
let udf = BatchUDF {
mapper: Box::new(move |batch: &RecordBatch| {
assert!(batch.column_by_name("x").is_some());
assert!(batch.column_by_name("y").is_none(), "y was not requested");
let x = batch["x"].as_any().downcast_ref::<Int32Array>().unwrap();
let doubled: Int32Array = x.iter().map(|v| v.map(|v| v * 2)).collect();
Ok(RecordBatch::try_new(
mapper_schema.clone(),
vec![Arc::new(doubled)],
)?)
}),
output_schema,
result_checkpoint: None,
};
table
.add_columns()
.transform(NewColumnTransform::BatchUDF(udf))
.read_columns(["x"])
.execute()
.await
.unwrap();
let schema = table.schema().await.unwrap();
assert!(schema.field_with_name("sum").is_ok());
}
}
+9 -5
View File
@@ -576,10 +576,12 @@ mod tests {
// Add a new physical column AFTER the embedding column. // Add a new physical column AFTER the embedding column.
table table
.add_columns( .add_columns()
NewColumnTransform::SqlExpressions(vec![("score".into(), "42.0".into())]), .transform(NewColumnTransform::SqlExpressions(vec![(
None, "score".into(),
) "42.0".into(),
)]))
.execute()
.await .await
.unwrap(); .unwrap();
@@ -683,7 +685,9 @@ mod tests {
true, true,
)])); )]));
table table
.add_columns(NewColumnTransform::AllNulls(nested_schema), None) .add_columns()
.transform(NewColumnTransform::AllNulls(nested_schema))
.execute()
.await .await
.unwrap(); .unwrap();
+246
View File
@@ -21,6 +21,9 @@ use lance_index::vector::pq::PQBuildParams;
use lance_index::vector::sq::builder::SQBuildParams; use lance_index::vector::sq::builder::SQBuildParams;
use crate::error::{Error, Result}; use crate::error::{Error, Result};
/// Resolved column, index parameters and index type for one build.
pub(super) type PreparedIndex = (String, Box<dyn lance::index::IndexParams>, IndexType);
use crate::index::Index; use crate::index::Index;
use crate::index::vector::{VectorIndex, suggested_num_sub_vectors}; use crate::index::vector::{VectorIndex, suggested_num_sub_vectors};
use crate::utils::{ use crate::utils::{
@@ -105,6 +108,47 @@ impl NativeTable {
} }
} }
/// Resolves the target column and index parameters, erroring on input the
/// build would reject.
pub(super) async fn prepare_index(
&self,
opts: &crate::index::IndexBuilder,
) -> Result<PreparedIndex> {
if opts.columns.len() != 1 {
return Err(Error::Schema {
message: "Multi-column (composite) indices are not yet supported".to_string(),
});
}
self.dataset.ensure_mutable()?;
let dataset = self.dataset.get().await?;
let (column, field) = Self::resolve_index_field(dataset.schema(), &opts.columns[0])?;
let params = self.make_index_params(&field, opts.index.clone()).await?;
let index_type = self.get_index_type_for_field(&field, &opts.index);
Ok((column, params, index_type))
}
/// Builds a prepared index and publishes the new dataset version.
pub(super) async fn build_index(
&self,
opts: crate::index::IndexBuilder,
prepared: PreparedIndex,
) -> Result<()> {
let (column, lance_idx_params, index_type) = prepared;
let mut dataset = (*self.dataset.get().await?).clone();
let columns = [column.as_str()];
let mut builder = dataset
.create_index_builder(&columns, index_type, lance_idx_params.as_ref())
.train(opts.train)
.replace(opts.replace);
if let Some(name) = opts.name {
builder = builder.name(name);
}
builder.await?;
self.dataset.update(dataset);
Ok(())
}
pub(super) fn resolve_index_field( pub(super) fn resolve_index_field(
schema: &lance_core::datatypes::Schema, schema: &lance_core::datatypes::Schema,
column: &str, column: &str,
@@ -475,6 +519,208 @@ mod tests {
assert_eq!(table.list_indices().await.unwrap().len(), 0); assert_eq!(table.list_indices().await.unwrap().len(), 0);
} }
#[tokio::test]
async fn test_execute_async_job_waits_for_local_build() {
let tmp_dir = tempdir().unwrap();
let conn = connect(tmp_dir.path().to_str().unwrap())
.execute()
.await
.unwrap();
let batch = record_batch!(("id", Int32, (0..512).collect::<Vec<_>>())).unwrap();
let table = conn.create_table("t", batch).execute().await.unwrap();
let job = table
.create_index(&["id"], Index::BTree(BTreeIndexBuilder::default()))
.execute_async()
.await
.unwrap();
// Local jobs run in this process and have no server id.
assert_eq!(job.id(), None);
// The build runs as a task, so the index need not exist yet; it must
// once the job resolves.
job.wait().await.unwrap();
assert_eq!(table.list_indices().await.unwrap().len(), 1);
// Cancelling a finished job is a no-op.
job.cancel().await.unwrap();
}
/// Concurrent waiters, and a wait issued after the job settled, all
/// succeed once the build does.
#[tokio::test]
async fn test_execute_async_job_reports_success_to_every_waiter() {
let tmp_dir = tempdir().unwrap();
let conn = connect(tmp_dir.path().to_str().unwrap())
.execute()
.await
.unwrap();
let batch = record_batch!(("id", Int32, (0..512).collect::<Vec<_>>())).unwrap();
let table = conn.create_table("t", batch).execute().await.unwrap();
let job = Arc::new(
table
.create_index(&["id"], Index::BTree(BTreeIndexBuilder::default()))
.execute_async()
.await
.unwrap(),
);
let waiters = (0..4)
.map(|_| {
let job = job.clone();
tokio::spawn(async move { job.wait().await })
})
.collect::<Vec<_>>();
for waiter in waiters {
waiter.await.unwrap().unwrap();
}
// A wait after the job settled still reports the same outcome.
job.wait().await.unwrap();
assert_eq!(table.list_indices().await.unwrap().len(), 1);
}
/// Every waiter sees a failure, not just the first: a waiter that missed
/// the outcome would be told the job succeeded.
#[tokio::test]
async fn test_execute_async_job_reports_failure_to_every_waiter() {
let tmp_dir = tempdir().unwrap();
let conn = connect(tmp_dir.path().to_str().unwrap())
.execute()
.await
.unwrap();
let batch = record_batch!(("id", Int32, (0..512).collect::<Vec<_>>())).unwrap();
let table = conn.create_table("t", batch).execute().await.unwrap();
table
.create_index(&["id"], Index::BTree(BTreeIndexBuilder::default()))
.execute()
.await
.unwrap();
// Rebuilding the same index without replace fails once the build
// starts, so the failure reaches the job rather than execute_async.
let job = Arc::new(
table
.create_index(&["id"], Index::BTree(BTreeIndexBuilder::default()))
.replace(false)
.execute_async()
.await
.unwrap(),
);
let waiters = (0..3)
.map(|_| {
let job = job.clone();
tokio::spawn(async move { job.wait().await })
})
.collect::<Vec<_>>();
for waiter in waiters {
waiter
.await
.unwrap()
.expect_err("every waiter must see the failure");
}
job.wait().await.expect_err("a later wait still fails");
}
/// A local failure keeps the error it failed with, so a caller can match on
/// the original variant rather than parse a message.
#[tokio::test]
async fn test_execute_async_failure_keeps_the_source_error() {
let tmp_dir = tempdir().unwrap();
let conn = connect(tmp_dir.path().to_str().unwrap())
.execute()
.await
.unwrap();
let batch = record_batch!(("id", Int32, (0..512).collect::<Vec<_>>())).unwrap();
let table = conn.create_table("t", batch).execute().await.unwrap();
table
.create_index(&["id"], Index::BTree(BTreeIndexBuilder::default()))
.execute()
.await
.unwrap();
let job = table
.create_index(&["id"], Index::BTree(BTreeIndexBuilder::default()))
.replace(false)
.execute_async()
.await
.unwrap();
let crate::Error::JobFailed { failure, .. } = job.wait().await.unwrap_err() else {
panic!("a failed job reports JobFailed");
};
let source = failure.source.expect("a local failure carries its error");
assert_eq!(
failure.message.as_deref(),
Some(source.to_string()).as_deref()
);
// Nothing local can report these, so they must be absent rather than invented.
assert!(failure.phase.is_none());
assert!(failure.retryable.is_none());
}
/// Every waiter sees the cancellation, including ones that were already
/// waiting when the cancel landed.
#[tokio::test]
async fn test_execute_async_job_reports_cancellation_to_every_waiter() {
let tmp_dir = tempdir().unwrap();
let conn = connect(tmp_dir.path().to_str().unwrap())
.execute()
.await
.unwrap();
let batch = record_batch!(("id", Int32, (0..512).collect::<Vec<_>>())).unwrap();
let table = conn.create_table("t", batch).execute().await.unwrap();
let job = Arc::new(
table
.create_index(&["id"], Index::BTree(BTreeIndexBuilder::default()))
.execute_async()
.await
.unwrap(),
);
// Cancel before yielding, so the build cannot have started and the
// outcome is always the cancellation.
job.cancel().await.unwrap();
let waiters = (0..2)
.map(|_| {
let job = job.clone();
tokio::spawn(async move { job.wait().await })
})
.collect::<Vec<_>>();
for waiter in waiters {
match waiter.await.unwrap() {
Err(crate::Error::JobCancelled { .. }) => {}
other => panic!("expected the cancellation, got {other:?}"),
}
}
match job.wait().await {
Err(crate::Error::JobCancelled { .. }) => {}
other => panic!("expected the cancellation, got {other:?}"),
}
}
#[tokio::test]
async fn test_execute_async_job_cancel_stops_local_build() {
let tmp_dir = tempdir().unwrap();
let conn = connect(tmp_dir.path().to_str().unwrap())
.execute()
.await
.unwrap();
let batch = record_batch!(("id", Int32, (0..512).collect::<Vec<_>>())).unwrap();
let table = conn.create_table("t", batch).execute().await.unwrap();
let job = table
.create_index(&["id"], Index::BTree(BTreeIndexBuilder::default()))
.execute_async()
.await
.unwrap();
job.cancel().await.unwrap();
match job.wait().await {
Err(crate::Error::JobCancelled { .. }) => {}
// The build may finish before the abort lands.
Ok(()) => {}
other => panic!("unexpected job outcome: {other:?}"),
}
}
#[tokio::test] #[tokio::test]
async fn test_ivf_pq_uses_default_partition_size_for_num_partitions() { async fn test_ivf_pq_uses_default_partition_size_for_num_partitions() {
use crate::index::vector::IvfPqIndexBuilder; use crate::index::vector::IvfPqIndexBuilder;
+24 -19
View File
@@ -193,10 +193,12 @@ mod tests {
// Add a computed column // Add a computed column
let result = table let result = table
.add_columns( .add_columns()
NewColumnTransform::SqlExpressions(vec![("doubled".into(), "id * 2".into())]), .transform(NewColumnTransform::SqlExpressions(vec![(
None, "doubled".into(),
) "id * 2".into(),
)]))
.execute()
.await .await
.unwrap(); .unwrap();
@@ -251,13 +253,12 @@ mod tests {
// Add multiple columns at once // Add multiple columns at once
table table
.add_columns( .add_columns()
NewColumnTransform::SqlExpressions(vec![ .transform(NewColumnTransform::SqlExpressions(vec![
("y".into(), "x + 1".into()), ("y".into(), "x + 1".into()),
("z".into(), "x * x".into()), ("z".into(), "x * x".into()),
]), ]))
None, .execute()
)
.await .await
.unwrap(); .unwrap();
@@ -283,10 +284,12 @@ mod tests {
// Add a column with a constant value // Add a column with a constant value
table table
.add_columns( .add_columns()
NewColumnTransform::SqlExpressions(vec![("constant".into(), "42".into())]), .transform(NewColumnTransform::SqlExpressions(vec![(
None, "constant".into(),
) "42".into(),
)]))
.execute()
.await .await
.unwrap(); .unwrap();
@@ -659,10 +662,12 @@ mod tests {
// Add column increments version // Add column increments version
let add_result = table let add_result = table
.add_columns( .add_columns()
NewColumnTransform::SqlExpressions(vec![("c".into(), "a + b".into())]), .transform(NewColumnTransform::SqlExpressions(vec![(
None, "c".into(),
) "a + b".into(),
)]))
.execute()
.await .await
.unwrap(); .unwrap();
assert!(add_result.version > v1); assert!(add_result.version > v1);
+259 -7
View File
@@ -9,14 +9,17 @@ use arrow_array::{
}; };
use arrow_schema::{DataType, Field, Fields, Schema}; use arrow_schema::{DataType, Field, Fields, Schema};
use futures::TryStreamExt; use futures::TryStreamExt;
use lance::Dataset;
use lance_encoding::version::LanceFileVersion; use lance_encoding::version::LanceFileVersion;
use lancedb::{ use lancedb::{
Connection, Error, Result, Table, Connection, Error, Result, Table,
blob::{BlobRangeRequest, blob}, blob::{BlobRangeRequest, blob},
connect, connect_namespace, connect, connect_namespace,
database::listing::OPT_NEW_TABLE_ENABLE_STABLE_ROW_IDS, database::listing::{
ListingDatabaseOptions, NewTableConfig, OPT_NEW_TABLE_ENABLE_STABLE_ROW_IDS,
},
query::{ExecutableQuery, QueryBase}, query::{ExecutableQuery, QueryBase},
table::{AddDataMode, CompactionOptions, OptimizeAction}, table::{AddDataMode, CompactionOptions, OptimizeAction, OptimizeStats},
}; };
use tempfile::tempdir; use tempfile::tempdir;
@@ -647,7 +650,7 @@ async fn fetch_blob_ranges_validates_requests() -> Result<()> {
.await .await
.unwrap_err(); .unwrap_err();
assert!(matches!(&err, Error::InvalidInput { .. }), "got {err:?}"); assert!(matches!(&err, Error::InvalidInput { .. }), "got {err:?}");
assert!(err.to_string().contains("row ids")); assert!(err.to_string().contains("row IDs"));
Ok(()) Ok(())
} }
@@ -680,7 +683,7 @@ async fn fetch_blobs_out_of_range_id_errors_without_panic() -> Result<()> {
let table = create_inline_blob_table(&db, "t", &[1], &[Some(b"x".as_slice())]).await?; let table = create_inline_blob_table(&db, "t", &[1], &[Some(b"x".as_slice())]).await?;
let err = table.fetch_blobs("image", &[u64::MAX]).await.unwrap_err(); let err = table.fetch_blobs("image", &[u64::MAX]).await.unwrap_err();
assert!(err.to_string().contains("row ids")); assert!(err.to_string().contains("row IDs"));
Ok(()) Ok(())
} }
@@ -694,11 +697,11 @@ async fn fetch_blob_apis_reject_mixed_valid_and_missing_row_ids() -> Result<()>
let err = table.fetch_blobs("image", &row_ids).await.unwrap_err(); let err = table.fetch_blobs("image", &row_ids).await.unwrap_err();
assert!(matches!(&err, Error::InvalidInput { .. }), "got {err:?}"); assert!(matches!(&err, Error::InvalidInput { .. }), "got {err:?}");
assert!(err.to_string().contains("row ids")); assert!(err.to_string().contains("row IDs"));
let err = table.fetch_blob_files("image", &row_ids).await.unwrap_err(); let err = table.fetch_blob_files("image", &row_ids).await.unwrap_err();
assert!(matches!(&err, Error::InvalidInput { .. }), "got {err:?}"); assert!(matches!(&err, Error::InvalidInput { .. }), "got {err:?}");
assert!(err.to_string().contains("row ids")); assert!(err.to_string().contains("row IDs"));
let requests = row_ids.map(|row_id| BlobRangeRequest::new(row_id, 0, 1)); let requests = row_ids.map(|row_id| BlobRangeRequest::new(row_id, 0, 1));
let err = table let err = table
@@ -706,7 +709,7 @@ async fn fetch_blob_apis_reject_mixed_valid_and_missing_row_ids() -> Result<()>
.await .await
.unwrap_err(); .unwrap_err();
assert!(matches!(&err, Error::InvalidInput { .. }), "got {err:?}"); assert!(matches!(&err, Error::InvalidInput { .. }), "got {err:?}");
assert!(err.to_string().contains("row ids")); assert!(err.to_string().contains("row IDs"));
Ok(()) Ok(())
} }
@@ -1075,3 +1078,252 @@ async fn fetch_blob_files_aligns_across_fragments_with_nulls_and_dups() -> Resul
} }
Ok(()) Ok(())
} }
/// Rows exercising the null/empty interleavings from
/// <https://github.com/lancedb/lancedb/issues/3744>: a payload, a null, a valid
/// empty value, then payloads whose descriptors a fragment rewrite used to zero.
fn null_empty_input_batch() -> RecordBatch {
let owned = [
Some(dedicated_blob_bytes(1)),
None,
Some(Vec::new()),
Some(dedicated_blob_bytes(4)),
Some(dedicated_blob_bytes(5)),
Some(dedicated_blob_bytes(6)),
];
let payloads: Vec<Option<&[u8]>> = owned.iter().map(|payload| payload.as_deref()).collect();
binary_input_batch(&[1, 2, 3, 4, 5, 6], &payloads)
}
/// One `(id, Some((payload length, first byte)))` per live row, or `(id, None)`
/// for a null blob. Comparing lengths and first bytes keeps failure output
/// readable where comparing whole payloads would not.
type BlobSummary = Vec<(i64, Option<(usize, Option<u8>)>)>;
/// The rows [`null_empty_input_batch`] leaves behind after `id IN (1, 4)` is
/// deleted: a null, a valid empty value, and the two payloads that follow them.
fn expected_null_empty_survivors() -> BlobSummary {
vec![
(2, None),
(3, Some((0, None))),
(5, Some((DEDICATED_BLOB_LEN, Some(5)))),
(6, Some((DEDICATED_BLOB_LEN, Some(6)))),
]
}
/// `optimize()` only rewrites a fragment when lance's compaction planner selects
/// it — here because the delete pushes the fragment past
/// `materialize_deletions_threshold` (0.1 by default; these tests delete 2 of 6
/// rows). Without this check, a planner or threshold change upstream would leave
/// both regression tests green while no rewrite happened at all.
fn assert_compacted(stats: &OptimizeStats) {
let metrics = stats
.compaction
.as_ref()
.expect("OptimizeAction::All runs compaction");
assert!(
metrics.fragments_removed >= 1,
"optimize() rewrote no fragment, so this test proves nothing: {metrics:?}"
);
}
fn summarize(rows: &[(i64, Option<Vec<u8>>)]) -> BlobSummary {
rows.iter()
.map(|(id, payload)| {
(
*id,
payload
.as_ref()
.map(|bytes| (bytes.len(), bytes.first().copied())),
)
})
.collect()
}
async fn sorted_id_rowid(table: &Table) -> Result<Vec<(i64, u64)>> {
let mut pairs = collect_id_rowid(table).await?;
pairs.sort_by_key(|(id, _)| *id);
Ok(pairs)
}
/// `{position, size}` descriptors of a legacy v1 blob column, keyed by `id`.
async fn v1_blob_descriptors(table: &Table) -> Result<Vec<(i64, Option<(u64, u64)>)>> {
let batches = table
.query()
.execute()
.await?
.try_collect::<Vec<_>>()
.await?;
let batch = arrow_select::concat::concat_batches(&batches[0].schema(), &batches).unwrap();
let ids = batch
.column_by_name("id")
.unwrap()
.as_any()
.downcast_ref::<Int64Array>()
.unwrap();
let descriptors = batch
.column_by_name("image")
.unwrap()
.as_any()
.downcast_ref::<StructArray>()
.expect("v1 blob column reads back as a descriptor struct");
let position = descriptors
.column_by_name("position")
.unwrap()
.as_any()
.downcast_ref::<UInt64Array>()
.unwrap();
let size = descriptors
.column_by_name("size")
.unwrap()
.as_any()
.downcast_ref::<UInt64Array>()
.unwrap();
let mut rows: Vec<(i64, Option<(u64, u64)>)> = (0..batch.num_rows())
.map(|row| {
let descriptor =
(!descriptors.is_null(row)).then(|| (position.value(row), size.value(row)));
(ids.value(row), descriptor)
})
.collect();
rows.sort_by_key(|(id, _)| *id);
Ok(rows)
}
/// Payload bytes of every live row of a legacy v1 blob column, keyed by `id`.
/// [`Table::fetch_blobs`] rejects v1 columns, so read them through lance.
async fn v1_blob_payloads(dataset_uri: &str, table: &Table) -> Result<Vec<(i64, Option<Vec<u8>>)>> {
let pairs = sorted_id_rowid(table).await?;
let row_ids: Vec<u64> = pairs.iter().map(|(_, row_id)| *row_id).collect();
let dataset = Arc::new(Dataset::open(dataset_uri).await?);
let files = dataset.take_blobs(&row_ids, "image").await?;
assert_eq!(
files.len(),
pairs.len(),
"take_blobs returned {} handles for {} live rows",
files.len(),
pairs.len()
);
let mut rows = Vec::with_capacity(pairs.len());
for ((id, _), file) in pairs.iter().zip(files) {
let payload = match file {
Some(file) => Some(file.read().await?.to_vec()),
None => None,
};
rows.push((*id, payload));
}
Ok(rows)
}
/// Length and first byte of every live blob v2 value, keyed by `id`.
async fn blob_v2_values(table: &Table) -> Result<BlobSummary> {
let pairs = sorted_id_rowid(table).await?;
let row_ids: Vec<u64> = pairs.iter().map(|(_, row_id)| *row_id).collect();
let bytes = table.fetch_blobs("image", &row_ids).await?;
Ok(pairs
.iter()
.enumerate()
.map(|(slot, (id, _))| {
let value = (!bytes.is_null(slot))
.then(|| (bytes.value(slot).len(), bytes.value(slot).first().copied()));
(*id, value)
})
.collect())
}
/// Regression test for [#3744]: on storage 2.0 (legacy v1 descriptors),
/// compaction rewrote every payload following a null or empty value in the same
/// fragment as `{position: 0, size: 0}`, so the payload bytes read back as `b""`
/// and the new fragment no longer referenced them at all.
///
/// [#3744]: https://github.com/lancedb/lancedb/issues/3744
#[tokio::test]
async fn optimize_preserves_v1_blob_payloads_with_null_and_empty() -> Result<()> {
let tmp = tempdir().unwrap();
let db_uri = tmp.path().to_str().unwrap().to_string();
let db = connect(&db_uri)
.database_options(&ListingDatabaseOptions {
new_table_config: NewTableConfig {
data_storage_version: Some(LanceFileVersion::V2_0),
..Default::default()
},
..Default::default()
})
.execute()
.await?;
let legacy = Field::new("image", DataType::LargeBinary, true).with_metadata(
std::collections::HashMap::from([("lance-encoding:blob".to_string(), "true".to_string())]),
);
let schema = Arc::new(Schema::new(vec![
Field::new("id", DataType::Int64, false),
legacy,
]));
let table = db.create_empty_table("t", schema).execute().await?;
table.add(null_empty_input_batch()).execute().await?;
assert_eq!(
storage_format_version(&table).await,
LanceFileVersion::V2_0.resolve(),
"v1 blob descriptors only exist below storage 2.2"
);
let dataset_uri = table.uri().await?;
// Any rewrite triggers it; deleting rows is the shape from the issue.
table.delete("id IN (1, 4)").await?;
let descriptors_before = v1_blob_descriptors(&table).await?;
let before = v1_blob_payloads(&dataset_uri, &table).await?;
assert_eq!(
summarize(&before),
expected_null_empty_survivors(),
"test setup no longer produces the null/empty/payload mix"
);
let stats = table.optimize(OptimizeAction::All).await?;
assert_compacted(&stats);
let descriptors_after = v1_blob_descriptors(&table).await?;
let after = v1_blob_payloads(&dataset_uri, &table).await?;
assert_eq!(
summarize(&after),
summarize(&before),
"optimize() lost blob payloads; descriptors before={descriptors_before:?} after={descriptors_after:?}"
);
assert!(after == before, "optimize() changed blob payload bytes");
Ok(())
}
/// Regression test for the blob v2 half of [#3744]: compaction rewrote a valid
/// empty value as null, destroying the null-vs-empty distinction.
///
/// [#3744]: https://github.com/lancedb/lancedb/issues/3744
#[tokio::test]
async fn optimize_preserves_blob_v2_null_and_empty_distinction() -> Result<()> {
let tmp = tempdir().unwrap();
let db = connect(tmp.path().to_str().unwrap()).execute().await?;
let table = db
.create_empty_table("t", blob_table_schema())
.execute()
.await?;
table.add(null_empty_input_batch()).execute().await?;
assert!(
storage_format_version(&table).await >= LanceFileVersion::V2_2,
"blob v2 columns require storage >= 2.2"
);
table.delete("id IN (1, 4)").await?;
let before = blob_v2_values(&table).await?;
assert_eq!(
before,
expected_null_empty_survivors(),
"test setup no longer produces the null/empty/payload mix"
);
let stats = table.optimize(OptimizeAction::All).await?;
assert_compacted(&stats);
assert_eq!(
blob_v2_values(&table).await?,
before,
"optimize() changed blob v2 values"
);
Ok(())
}