Compare commits

..

10 Commits

Author SHA1 Message Date
Wyatt Alt c463ca1503 docs: state what a computed column promises
The declaration API shipped without a runnable example, and none of the three
binding docs said when values appear, what happens to them when an input
changes, which schema operations a declaration blocks, or that the feature is
local-only. Those are the questions a caller has to answer before using it.

Adds Rust doctests on both entry points and the same semantics to the Python
and TypeScript parameter docs, plus a worked Python example. Also adds the
abstract refresh_column that both concrete Python tables already implemented,
so the surface is declared in one place and the cross-references resolve.
2026-08-12 21:01:11 -07:00
Wyatt Alt 79e9dffd06 fix: make computed columns explicitly local-only
Both operations were advertised on remote tables and neither could work. A
declaration reaches the wire as AllNulls, which RemoteTable::add_columns does
not accept, and refresh_column fell through to the BaseTable default; the
TypeScript wrapper reached the same surface. Callers got errors that named
neither the feature nor the reason.

The remote protocol has no representation for a stored expression, and what
one should look like is not settled -- the server persists a declaration under
a different vocabulary. So this states the boundary rather than guessing at a
wire format: an explicit AllNulls arm, a default that names computed columns,
and NotImplementedError raised in Python before the round trip.
2026-08-12 20:00:18 -07:00
Wyatt Alt c3efc320a6 fix(rust): refuse schema changes that invalidate a computed column
A declaration records the columns its expression reads, but nothing consulted
them: renaming an input left an expression naming a column that no longer
exists, and the failure surfaced at refresh time as a plan error rather than
at the operation that caused it. Dropping or retyping an input did the same.

alter_columns and drop_columns now reject a change to a column some
declaration reads. Nullability is not part of what an expression resolves
against, so it stays allowed. A declaration does not read itself and so
travels with its own binding, and paths compare at their root, since a change
to `metadata.age` invalidates an expression reading `metadata` just as surely.

Binding to field ids instead would leave the expression text naming the old
column, so it would need rewriting stored SQL on every rename. Refusing the
operation is what a generated column does elsewhere.
2026-08-12 20:00:18 -07:00
Wyatt Alt 8d6dea6313 fix(rust): fill computed columns row by row
Refresh selected rows with `{column} IS NULL` and rewrote them through an
UPDATE, which made the output value double as the record of whether the row
had been computed. Two consequences: an expression yielding null re-selected
the same rows on every run and reported them as filled forever, and the
target name was interpolated into SQL unquoted, so a column named
`double value` could be declared and never refreshed.

Filling is now per fragment. Each fragment that could hold an unfilled row
has the expression evaluated over its physical rows and the result written as
a standalone column file, published together in one DataReplacement. A row
that already holds a value keeps it -- the computed and current values are
merged on the is-null mask -- and a row counts as filled only when it gains a
value, so a fragment where nothing would change is never staged and a null
expression settles after one pass.

Which fragments are worth looking at comes from the manifest first: one whose
data files do not carry the field cannot hold a filled row. A fragment that
does carry it is still asked, because a row rewrite -- an update, or a
compaction folding an unfilled fragment into a filled one -- leaves nulls
behind a covering file. That case is the reason coverage alone is not the
marker; the test for it fails against a coverage-only implementation.

Names now reach the evaluator through a projection alias or a backtick-quoted
identifier, lance's dialect having no other way to spell one -- a
double-quoted name parses as a string literal.
2026-08-12 20:00:18 -07:00
Wyatt Alt 5a2d3f39e2 chore: update lance dependency to v11.0.0-beta.7
Filling a computed column needs FileFragment::write_column, which lands in
this release.
2026-08-12 20:00:18 -07:00
Wyatt Alt 219f41339d refactor(rust): tag a computed column's definition by kind
ComputedColumn carried a bare expression string, which asserts that every
computed column is a SQL expression. That holds for the only kind there is,
but it is the wrong shape for the next one: a column defined by a registered
function cannot be typed by parsing its definition, so its type and its inputs
have a different provenance than a SQL column's. A single string has nowhere
to say which it is.

The definition is now ComputedColumnKind, non-exhaustive so another kind is
additive, and the metadata carries a matching computed_column.kind tag beside
the payload. Tagging the persisted form is the point -- the Rust type stays
cheap to change and field metadata does not, and a second kind distinguished
only by which keys happen to be present would leave every reader sniffing the
shape.

An unrecognized kind reads back as Unrecognized rather than as absent. A newer
version's declaration is a computed column this one cannot evaluate, not a
plain column: reported as absent it would be redeclarable over and would fail
refresh as "not a computed column".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 18:48:48 -07:00
Wyatt Alt 5982eebbb3 feat(nodejs): expose computed columns and refreshColumn
addComputedColumns declares columns defined by a SQL expression, and
refreshColumn fills a computed column's unfilled rows:

    await table.addComputedColumns([{ name: "doubled", valueSql: "x * 2" }]);
    await table.refreshColumn("doubled");
2026-08-12 18:48:48 -07:00
Wyatt Alt 5a1c839382 feat(python): expose computed columns and refresh_column
add_columns gains a computed= mapping of column name to SQL expression, and
refresh_column fills a computed column's unfilled rows. Both are available on
the sync and async tables:

    table.add_columns(computed={"doubled": "x * 2"})
    table.refresh_column("doubled")

transforms and computed are mutually exclusive, since they commit through
different paths and could half-apply.
2026-08-12 18:48:48 -07:00
Wyatt Alt e40e073a9d feat(rust): fill computed columns with refresh_column
Declaring a computed column stores its expression but computes nothing, so
until now the column stayed null with no way to fill it. refresh_column
evaluates the expression over the rows that still hold no value and commits the
results:

    table.refresh_column("doubled").await?

Rows without a value are the ones to fill, which also makes the operation
idempotent and resumable after a failure: refreshing again picks up whatever
did not land. A row whose expression evaluates to null is indistinguishable
from an unfilled one and is recomputed, which costs work but cannot change the
result.

Values written after a refresh are reachable by the next one, which is the case
that matters -- an expression column populated once and then appended to would
otherwise read null for every later row forever.
2026-08-12 18:48:48 -07:00
Wyatt Alt a47c22b26e feat(rust): declare computed columns through add_columns
A computed column is a column defined by a SQL expression rather than by
values supplied at write time, so it is added through add_columns like any
other:

    table.add_columns().computed("doubled", "x * 2").execute().await?

Declaring does not compute. The column is committed carrying its expression in
field metadata but no data, which makes declaration cost the same on a large
table as on an empty one and leaves a single code path that ever produces
values. A later refresh fills it.

The expression is the whole definition: both the result type and the input
columns are derived from it with lance-datafusion's planner, so a caller writes
neither, and the two can never disagree the way a hand-declared input list can.
Everything statically knowable is rejected at declare time rather than deferred:
an expression that does not parse, one referencing a column that does not
exist, a name already in use, and the same name declared twice in one call.

The binding lives in three field metadata keys: virtual_column marks the
column, virtual_column.expression holds it, and virtual_column.inputs holds the
parsed inputs as a JSON array. computed_columns() reads them back off a schema,
the way a SQL catalog reports a generation expression as another column of its
information schema, so introspection needs no round trip.

A transform and computed columns cannot be combined in one call, since they
commit through different transforms and could half-apply.

Only functions the query engine already knows can be named; resolving a
user-defined one needs a planner aware of the function registry, which does not
exist yet.
2026-08-12 18:48:48 -07:00
63 changed files with 514 additions and 3911 deletions
+1 -1
View File
@@ -1,5 +1,5 @@
[tool.bumpversion]
current_version = "0.38.0-beta.0"
current_version = "0.37.1-beta.1"
parse = """(?x)
(?P<major>0|[1-9]\\d*)\\.
(?P<minor>0|[1-9]\\d*)\\.
@@ -4,14 +4,14 @@ on:
workflow_call:
inputs:
tag:
description: "Tag name from Lance (e.g. `v7.2.0-beta.1`). If omitted, the newest release is resolved automatically — stable releases are preferred over pre-releases — and the run is skipped if it is not newer than the version currently pinned in Cargo.toml."
description: "Tag name from Lance. If omitted, the skill will use the latest Lance release that needs an update."
required: false
default: ""
type: string
workflow_dispatch:
inputs:
tag:
description: "Tag name from Lance (e.g. `v7.2.0-beta.1`). Leave empty to resolve the newest release automatically — stable releases are preferred over pre-releases — and skip the run if it is not newer than the version currently pinned in Cargo.toml."
description: "Tag name from Lance. Leave empty to use the latest Lance release that needs an update."
required: false
default: ""
type: string
-10
View File
@@ -69,16 +69,6 @@ jobs:
uses: actions/setup-python@v6
with:
python-version: "3.10"
- name: Add swap for Arm fat LTO
if: matrix.config.platform == 'aarch64'
shell: bash
run: |
swap_file="$RUNNER_TEMP/lancedb-swap"
sudo fallocate --length 16G "$swap_file"
sudo chmod 600 "$swap_file"
sudo mkswap "$swap_file"
sudo swapon "$swap_file"
free -h
- uses: ./.github/workflows/build_linux_wheel
with:
python-minor-version: 10
Generated
+61 -48
View File
@@ -3455,8 +3455,8 @@ checksum = "42703706b716c37f96a77aea830392ad231f44c9e9a67872fa5548707e11b11c"
[[package]]
name = "fsst"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arrow-array",
"rand 0.9.5",
@@ -4815,8 +4815,8 @@ checksum = "e037a2e1d8d5fdbd49b16a4ea09d5d6401c1f29eca5ff29d03d3824dba16256a"
[[package]]
name = "lance"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arc-swap",
"arrow",
@@ -4832,6 +4832,7 @@ dependencies = [
"async-recursion",
"async-trait",
"async_cell",
"aws-credential-types",
"aws-sdk-dynamodb",
"byteorder",
"bytes",
@@ -4847,6 +4848,7 @@ dependencies = [
"either",
"fst",
"futures",
"half",
"humantime",
"itertools 0.14.0",
"lance-arrow",
@@ -4888,8 +4890,8 @@ dependencies = [
[[package]]
name = "lance-arrow"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arrow-array",
"arrow-buffer",
@@ -4911,7 +4913,7 @@ dependencies = [
[[package]]
name = "lance-arrow-scalar"
version = "58.0.0"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arrow-array",
"arrow-buffer",
@@ -4925,7 +4927,7 @@ dependencies = [
[[package]]
name = "lance-arrow-stats"
version = "58.0.0"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arrow-array",
"arrow-schema",
@@ -4934,8 +4936,8 @@ dependencies = [
[[package]]
name = "lance-bitpacking"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arrayref",
"crunchy",
@@ -4945,8 +4947,8 @@ dependencies = [
[[package]]
name = "lance-core"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arrow-array",
"arrow-buffer",
@@ -4954,10 +4956,12 @@ dependencies = [
"arrow-schema",
"async-trait",
"blake3",
"byteorder",
"bytes",
"datafusion-common",
"datafusion-sql",
"futures",
"itertools 0.14.0",
"lance-arrow",
"lance-derive",
"libc",
@@ -4975,6 +4979,7 @@ dependencies = [
"snafu 0.9.0",
"tempfile",
"tokio",
"tokio-stream",
"tokio-util",
"tracing",
"twox-hash",
@@ -4983,8 +4988,8 @@ dependencies = [
[[package]]
name = "lance-datafusion"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arrow",
"arrow-array",
@@ -5003,6 +5008,7 @@ dependencies = [
"jsonb",
"lance-arrow",
"lance-core",
"lance-datagen",
"log",
"pin-project",
"prost",
@@ -5013,8 +5019,8 @@ dependencies = [
[[package]]
name = "lance-datagen"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arrow",
"arrow-array",
@@ -5031,8 +5037,8 @@ dependencies = [
[[package]]
name = "lance-derive"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"proc-macro2",
"quote",
@@ -5041,8 +5047,8 @@ dependencies = [
[[package]]
name = "lance-encoding"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arrow-arith",
"arrow-array",
@@ -5067,6 +5073,7 @@ dependencies = [
"num-traits",
"prost",
"prost-build",
"rand 0.9.5",
"tokio",
"tracing",
"xxhash-rust",
@@ -5075,8 +5082,8 @@ dependencies = [
[[package]]
name = "lance-file"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arrow-arith",
"arrow-array",
@@ -5107,8 +5114,8 @@ dependencies = [
[[package]]
name = "lance-index"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arc-swap",
"arrow",
@@ -5123,6 +5130,7 @@ dependencies = [
"async-trait",
"bitvec",
"bytes",
"chrono",
"crossbeam-queue",
"datafusion",
"datafusion-common",
@@ -5140,6 +5148,7 @@ dependencies = [
"lance-bitpacking",
"lance-core",
"lance-datafusion",
"lance-datagen",
"lance-encoding",
"lance-file",
"lance-index-core",
@@ -5168,12 +5177,13 @@ dependencies = [
"tempfile",
"tokio",
"tracing",
"uuid",
]
[[package]]
name = "lance-index-core"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arrow-array",
"arrow-schema",
@@ -5195,8 +5205,8 @@ dependencies = [
[[package]]
name = "lance-io"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arrow",
"arrow-array",
@@ -5210,6 +5220,7 @@ dependencies = [
"futures",
"http 1.5.0",
"io-uring",
"lance-arrow",
"lance-core",
"lance-namespace",
"log",
@@ -5227,28 +5238,29 @@ dependencies = [
"tokio",
"tracing",
"url",
"uuid",
]
[[package]]
name = "lance-linalg"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arrow-array",
"arrow-buffer",
"arrow-schema",
"cc",
"half",
"lance-arrow",
"lance-core",
"num-traits",
"rand 0.9.5",
"rayon",
]
[[package]]
name = "lance-namespace"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arrow",
"async-trait",
@@ -5260,8 +5272,8 @@ dependencies = [
[[package]]
name = "lance-namespace-impls"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arrow",
"arrow-ipc",
@@ -5300,9 +5312,9 @@ dependencies = [
[[package]]
name = "lance-namespace-reqwest-client"
version = "0.11.0"
version = "0.8.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0a030196da1c994b63a96a4f0bf5b0cfa459fe6dadc9e962320246ca328da22a"
checksum = "ba3f0a235e3ed5f8805205649ccc7d7d0f3df23ce1294242c9265ad488d7f19d"
dependencies = [
"reqwest 0.12.28",
"serde",
@@ -5314,13 +5326,14 @@ dependencies = [
[[package]]
name = "lance-select"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arrow-array",
"arrow-buffer",
"arrow-schema",
"byteorder",
"bytes",
"itertools 0.14.0",
"lance-core",
"roaring",
@@ -5329,8 +5342,8 @@ dependencies = [
[[package]]
name = "lance-table"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arrow",
"arrow-array",
@@ -5370,8 +5383,8 @@ dependencies = [
[[package]]
name = "lance-testing"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"arrow-array",
"arrow-schema",
@@ -5384,8 +5397,8 @@ dependencies = [
[[package]]
name = "lance-tokenizer"
version = "11.0.0-beta.13"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.13#ee41152ceb9a78e5df4d2456fdbdb98542eb2059"
version = "11.0.0-beta.7"
source = "git+https://github.com/lance-format/lance.git?tag=v11.0.0-beta.7#e581c49338bc83baf1ea50c5e235bd702f3fbeea"
dependencies = [
"frostem",
"icu_segmenter",
@@ -5398,7 +5411,7 @@ dependencies = [
[[package]]
name = "lancedb"
version = "0.38.0-beta.0"
version = "0.37.1-beta.1"
dependencies = [
"ahash",
"anyhow",
@@ -5486,7 +5499,7 @@ dependencies = [
[[package]]
name = "lancedb-nodejs"
version = "0.38.0-beta.0"
version = "0.37.1-beta.1"
dependencies = [
"arrow-array",
"arrow-buffer",
@@ -5511,7 +5524,7 @@ dependencies = [
[[package]]
name = "lancedb-python"
version = "0.38.0-beta.0"
version = "0.37.1-beta.1"
dependencies = [
"arrow",
"async-trait",
+14 -14
View File
@@ -13,20 +13,20 @@ categories = ["database-implementations"]
rust-version = "1.91.0"
[workspace.dependencies]
lance = { "version" = "=11.0.0-beta.13", default-features = false, "tag" = "v11.0.0-beta.13", "git" = "https://github.com/lance-format/lance.git" }
lance-core = { "version" = "=11.0.0-beta.13", "tag" = "v11.0.0-beta.13", "git" = "https://github.com/lance-format/lance.git" }
lance-datagen = { "version" = "=11.0.0-beta.13", "tag" = "v11.0.0-beta.13", "git" = "https://github.com/lance-format/lance.git" }
lance-file = { "version" = "=11.0.0-beta.13", "tag" = "v11.0.0-beta.13", "git" = "https://github.com/lance-format/lance.git" }
lance-io = { "version" = "=11.0.0-beta.13", default-features = false, "tag" = "v11.0.0-beta.13", "git" = "https://github.com/lance-format/lance.git" }
lance-index = { "version" = "=11.0.0-beta.13", "tag" = "v11.0.0-beta.13", "git" = "https://github.com/lance-format/lance.git" }
lance-linalg = { "version" = "=11.0.0-beta.13", "tag" = "v11.0.0-beta.13", "git" = "https://github.com/lance-format/lance.git" }
lance-namespace = { "version" = "=11.0.0-beta.13", "tag" = "v11.0.0-beta.13", "git" = "https://github.com/lance-format/lance.git" }
lance-namespace-impls = { "version" = "=11.0.0-beta.13", default-features = false, "tag" = "v11.0.0-beta.13", "git" = "https://github.com/lance-format/lance.git" }
lance-table = { "version" = "=11.0.0-beta.13", "tag" = "v11.0.0-beta.13", "git" = "https://github.com/lance-format/lance.git" }
lance-testing = { "version" = "=11.0.0-beta.13", "tag" = "v11.0.0-beta.13", "git" = "https://github.com/lance-format/lance.git" }
lance-datafusion = { "version" = "=11.0.0-beta.13", "tag" = "v11.0.0-beta.13", "git" = "https://github.com/lance-format/lance.git" }
lance-encoding = { "version" = "=11.0.0-beta.13", "tag" = "v11.0.0-beta.13", "git" = "https://github.com/lance-format/lance.git" }
lance-arrow = { "version" = "=11.0.0-beta.13", "tag" = "v11.0.0-beta.13", "git" = "https://github.com/lance-format/lance.git" }
lance = { "version" = "=11.0.0-beta.7", default-features = false, "tag" = "v11.0.0-beta.7", "git" = "https://github.com/lance-format/lance.git" }
lance-core = { "version" = "=11.0.0-beta.7", "tag" = "v11.0.0-beta.7", "git" = "https://github.com/lance-format/lance.git" }
lance-datagen = { "version" = "=11.0.0-beta.7", "tag" = "v11.0.0-beta.7", "git" = "https://github.com/lance-format/lance.git" }
lance-file = { "version" = "=11.0.0-beta.7", "tag" = "v11.0.0-beta.7", "git" = "https://github.com/lance-format/lance.git" }
lance-io = { "version" = "=11.0.0-beta.7", default-features = false, "tag" = "v11.0.0-beta.7", "git" = "https://github.com/lance-format/lance.git" }
lance-index = { "version" = "=11.0.0-beta.7", "tag" = "v11.0.0-beta.7", "git" = "https://github.com/lance-format/lance.git" }
lance-linalg = { "version" = "=11.0.0-beta.7", "tag" = "v11.0.0-beta.7", "git" = "https://github.com/lance-format/lance.git" }
lance-namespace = { "version" = "=11.0.0-beta.7", "tag" = "v11.0.0-beta.7", "git" = "https://github.com/lance-format/lance.git" }
lance-namespace-impls = { "version" = "=11.0.0-beta.7", default-features = false, "tag" = "v11.0.0-beta.7", "git" = "https://github.com/lance-format/lance.git" }
lance-table = { "version" = "=11.0.0-beta.7", "tag" = "v11.0.0-beta.7", "git" = "https://github.com/lance-format/lance.git" }
lance-testing = { "version" = "=11.0.0-beta.7", "tag" = "v11.0.0-beta.7", "git" = "https://github.com/lance-format/lance.git" }
lance-datafusion = { "version" = "=11.0.0-beta.7", "tag" = "v11.0.0-beta.7", "git" = "https://github.com/lance-format/lance.git" }
lance-encoding = { "version" = "=11.0.0-beta.7", "tag" = "v11.0.0-beta.7", "git" = "https://github.com/lance-format/lance.git" }
lance-arrow = { "version" = "=11.0.0-beta.7", "tag" = "v11.0.0-beta.7", "git" = "https://github.com/lance-format/lance.git" }
ahash = "0.8"
# Note that this one does not include pyarrow
arrow = { version = "58.0.0", optional = false }
+1 -1
View File
@@ -14,7 +14,7 @@ Add the following dependency to your `pom.xml`:
<dependency>
<groupId>com.lancedb</groupId>
<artifactId>lancedb-core</artifactId>
<version>0.38.0-beta.0</version>
<version>0.37.1-beta.1</version>
</dependency>
```
-23
View File
@@ -386,29 +386,6 @@ Drop an existing table.
***
### dropTableAsync()
```ts
abstract dropTableAsync(name, namespacePath?): Promise<Job>
```
Start dropping a table and return its cleanup job.
The table may become unavailable before its data files are removed. Wait
on the returned job to know when cleanup has finished.
#### Parameters
* **name**: `string`
* **namespacePath?**: `string`[]
#### Returns
`Promise`&lt;[`Job`](Job.md)&gt;
***
### getJob()
```ts
+3 -39
View File
@@ -79,9 +79,8 @@ input leaves the value computed at fill time; recomputing means dropping
the column and declaring it again. While a declaration reads a column,
that column cannot be renamed, retyped or dropped.
On LanceDB Cloud and Enterprise the expression is planned by the
server, and the refresh runs as a server job -- see
[Table#refreshColumnAsync](Table.md#refreshcolumnasync).
Computed columns are local-only: LanceDB Cloud and Enterprise reject a
declaration.
#### Parameters
@@ -755,8 +754,7 @@ Fill the rows of a computed column that hold no value yet.
Rows appended since the last refresh are filled by the next one; rows
already filled are left as they are, so the call is idempotent and does
not observe a mutated input. Local tables only: a remote refresh runs
as a server job, through [Table#refreshColumnAsync](Table.md#refreshcolumnasync).
not observe a mutated input. Local tables only.
#### Parameters
@@ -772,40 +770,6 @@ number of rows filled and the new version number of the table.
***
### refreshColumnAsync()
```ts
abstract refreshColumnAsync(column): Promise<Job>
```
Like [Table#refreshColumn](Table.md#refreshcolumn), but returns a handle to the refresh
job instead of blocking until it completes.
The job may already be complete when returned; callers must not assume
the column is filled until [Job.wait](Job.md#wait) resolves. Invalid input --
an unknown column, or one that is not computed -- rejects here rather
than failing the job. On local tables the job runs in-process; on
LanceDB Cloud and Enterprise it is the server's backfill job.
#### Parameters
* **column**: `string`
The name of the computed column to fill.
#### Returns
`Promise`&lt;[`Job`](Job.md)&gt;
#### Example
```ts
const job = await table.refreshColumnAsync("doubled");
await job.wait();
console.log(await job.status()); // "finished"
```
***
### restore()
```ts
+1 -1
View File
@@ -8,7 +8,7 @@
<parent>
<groupId>com.lancedb</groupId>
<artifactId>lancedb-parent</artifactId>
<version>0.38.0-beta.0</version>
<version>0.37.1-beta.1</version>
<relativePath>../pom.xml</relativePath>
</parent>
+2 -2
View File
@@ -6,7 +6,7 @@
<groupId>com.lancedb</groupId>
<artifactId>lancedb-parent</artifactId>
<version>0.38.0-beta.0</version>
<version>0.37.1-beta.1</version>
<packaging>pom</packaging>
<name>${project.artifactId}</name>
<description>LanceDB Java SDK Parent POM</description>
@@ -28,7 +28,7 @@
<properties>
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
<arrow.version>15.0.0</arrow.version>
<lance-core.version>11.0.0-beta.13</lance-core.version>
<lance-core.version>11.0.0-beta.6</lance-core.version>
<spotless.skip>false</spotless.skip>
<spotless.version>2.30.0</spotless.version>
<spotless.java.googlejavaformat.version>1.7</spotless.java.googlejavaformat.version>
+1 -1
View File
@@ -1,7 +1,7 @@
[package]
name = "lancedb-nodejs"
edition.workspace = true
version = "0.38.0-beta.0"
version = "0.37.1-beta.1"
publish = false
license.workspace = true
description.workspace = true
-10
View File
@@ -89,16 +89,6 @@ describe("given a connection", () => {
await db.createTable("test4", [{ id: 1 }, { id: 2 }]);
});
it("should return a completed job when dropping a local table", async () => {
await db.createTable("async-drop", [{ id: 1 }]);
const job = await db.dropTableAsync("async-drop");
expect(job.id).toBeNull();
await expect(job.status()).resolves.toBe("finished");
await job.wait();
await expect(db.tableNames()).resolves.toEqual([]);
});
it("should fail if creating table twice, unless overwrite is true", async () => {
let tbl = await db.createTable("test", [{ id: 1 }, { id: 2 }]);
await expect(tbl.countRows()).resolves.toBe(2);
-45
View File
@@ -1001,49 +1001,4 @@ describe("remote connection jobs surface", () => {
},
);
});
it("addBases posts the bases array", async () => {
const postedBodies: unknown[] = [];
await withMockDatabase(
(req, res) => {
const path = req.url ?? "";
if (path.endsWith("/describe/")) {
res.writeHead(200, { "Content-Type": "application/json" }).end(
JSON.stringify({
name: "photos",
version: 1,
schema: { fields: [] },
}),
);
return;
}
if (path.endsWith("/bases/")) {
const chunks: Buffer[] = [];
req.on("data", (chunk) => chunks.push(chunk));
req.on("end", () => {
postedBodies.push(JSON.parse(Buffer.concat(chunks).toString()));
res
.writeHead(200, { "Content-Type": "application/json" })
.end(JSON.stringify({ version: 2 }));
});
return;
}
res.writeHead(404).end();
},
async (db) => {
const table = await db.openTable("photos");
await table.addBases({ path: "s3://bucket/media/" });
},
);
expect(postedBodies).toEqual([
{
bases: [
{
path: "s3://bucket/media/",
isDatasetRoot: false,
},
],
},
]);
});
});
-42
View File
@@ -4,7 +4,6 @@
import * as fs from "fs";
import * as path from "path";
import * as tmp from "tmp";
import { pathToFileURL } from "url";
import * as arrow15 from "apache-arrow-15";
import * as arrow16 from "apache-arrow-16";
@@ -3366,28 +3365,6 @@ describe("computed columns", () => {
expect(rows.map((r) => r.doubled).sort()).toEqual([2, 4]);
});
it("returns a job handle from refreshColumnAsync", async () => {
const db = await connect(tmpDir.name);
const table = await db.createTable("computed_job", [{ x: 1 }, { x: 2 }]);
await table.addColumns({
computed: [{ name: "doubled", valueSql: "x * 2" }],
});
const job = await table.refreshColumnAsync("doubled");
expect(job.id).toBeNull();
await job.wait();
expect(await job.status()).toBe("finished");
const rows = await table.query().toArray();
expect(rows.map((r) => r.doubled).sort()).toEqual([2, 4]);
// Bad input rejects at the call, not through the job.
await expect(table.refreshColumnAsync("x")).rejects.toThrow(
"not a computed column",
);
});
it("fills rows added since the last refresh", async () => {
const db = await connect(tmpDir.name);
const table = await db.createTable("computed_append", [{ x: 1 }]);
@@ -3405,22 +3382,3 @@ describe("computed columns", () => {
expect(rows.map((r) => r.doubled).sort()).toEqual([10, 2]);
});
});
describe("table bases", () => {
let tmpDir: tmp.DirResult;
beforeEach(() => {
tmpDir = tmp.dirSync({ unsafeCleanup: true });
});
afterEach(() => tmpDir.removeCallback());
it("addBases accepts a file uri", async () => {
const conn = await connect(tmpDir.name);
const table = await conn.createEmptyTable(
"photos",
new arrow.Schema([new arrow.Field("id", new arrow.Int64(), false)]),
);
const media = path.join(tmpDir.name, "media");
fs.mkdirSync(media);
await table.addBases(pathToFileURL(media).toString());
});
});
-12
View File
@@ -327,14 +327,6 @@ export abstract class Connection {
*/
abstract dropTable(name: string, namespacePath?: string[]): Promise<void>;
/**
* Start dropping a table and return its cleanup job.
*
* The table may become unavailable before its data files are removed. Wait
* on the returned job to know when cleanup has finished.
*/
abstract dropTableAsync(name: string, namespacePath?: string[]): Promise<Job>;
/**
* Drop all tables in the database.
* @param {string[]} namespacePath The namespace path to drop tables from (defaults to root namespace).
@@ -713,10 +705,6 @@ export class LocalConnection extends Connection {
return this.inner.dropTable(name, namespacePath ?? []);
}
async dropTableAsync(name: string, namespacePath?: string[]): Promise<Job> {
return this.inner.dropTableAsync(name, namespacePath ?? []);
}
async dropAllTables(namespacePath?: string[]): Promise<void> {
return this.inner.dropAllTables(namespacePath ?? []);
}
-1
View File
@@ -130,7 +130,6 @@ export {
export {
Table,
TableBase,
Branches,
BranchColumnSummary,
BranchColumnChange,
+3 -77
View File
@@ -78,25 +78,6 @@ export interface WriteProgress {
done: boolean;
}
/**
* An extra storage prefix registered on a table.
*
* `path` is an object-store URI. `name` is an optional alias. `isDatasetRoot`
* is true when `path` points to a Lance dataset root. When false, `path`
* points directly to the directory containing the referenced files.
*/
export interface TableBase {
/** Object store URI such as `s3://bucket/media/`. */
path: string;
/** Optional alias. */
name?: string;
/**
* True when `path` is a Lance dataset root. When false, `path` is the
* directory containing the referenced files.
*/
isDatasetRoot?: boolean;
}
/**
* Options for adding data to a table.
*/
@@ -556,9 +537,8 @@ export abstract class Table {
* the column and declaring it again. While a declaration reads a column,
* that column cannot be renamed, retyped or dropped.
*
* On LanceDB Cloud and Enterprise the expression is planned by the
* server, and the refresh runs as a server job -- see
* {@link Table#refreshColumnAsync}.
* Computed columns are local-only: LanceDB Cloud and Enterprise reject a
* declaration.
* @param {AddColumnsSql[] | Field | Field[] | Schema} newColumnTransforms Either:
* - An array of objects with column names and SQL expressions to calculate values
* - A single Arrow Field defining one column with its data type (column will be initialized with null values)
@@ -582,47 +562,18 @@ export abstract class Table {
| { computed: AddColumnsSql[] },
): Promise<AddColumnsResult>;
/**
* Register additional storage bases for this table.
*
* A URI string is a non-root base with no alias.
*/
abstract addBases(
bases: string | TableBase | Array<string | TableBase>,
): Promise<void>;
/**
* Fill the rows of a computed column that hold no value yet.
*
* Rows appended since the last refresh are filled by the next one; rows
* already filled are left as they are, so the call is idempotent and does
* not observe a mutated input. Local tables only: a remote refresh runs
* as a server job, through {@link Table#refreshColumnAsync}.
* not observe a mutated input. Local tables only.
* @param {string} column The name of the computed column to fill.
* @returns {Promise<RefreshColumnResult>} A promise that resolves to the
* number of rows filled and the new version number of the table.
*/
abstract refreshColumn(column: string): Promise<RefreshColumnResult>;
/**
* Like {@link Table#refreshColumn}, but returns a handle to the refresh
* job instead of blocking until it completes.
*
* The job may already be complete when returned; callers must not assume
* the column is filled until {@link Job.wait} resolves. Invalid input --
* an unknown column, or one that is not computed -- rejects here rather
* than failing the job. On local tables the job runs in-process; on
* LanceDB Cloud and Enterprise it is the server's backfill job.
* @param {string} column The name of the computed column to fill.
* @example
* ```ts
* const job = await table.refreshColumnAsync("doubled");
* await job.wait();
* console.log(await job.status()); // "finished"
* ```
*/
abstract refreshColumnAsync(column: string): Promise<Job>;
/**
* Alter the name or nullability of columns.
* @param {ColumnAlteration[]} columnAlterations One or more alterations to
@@ -1224,20 +1175,10 @@ export class LocalTable extends Table {
throw new Error("Invalid input type for addColumns");
}
async addBases(
bases: string | TableBase | Array<string | TableBase>,
): Promise<void> {
await this.inner.addBases(normalizeBases(bases));
}
async refreshColumn(column: string): Promise<RefreshColumnResult> {
return await this.inner.refreshColumn(column);
}
async refreshColumnAsync(column: string): Promise<Job> {
return await this.inner.refreshColumnAsync(column);
}
async alterColumns(
columnAlterations: ColumnAlteration[],
): Promise<AlterColumnsResult> {
@@ -1430,21 +1371,6 @@ export class LocalTable extends Table {
}
}
function normalizeBases(
bases: string | TableBase | Array<string | TableBase>,
): TableBase[] {
const baseInputs = Array.isArray(bases) ? bases : [bases];
return baseInputs.map((base) =>
typeof base === "string"
? { path: base, isDatasetRoot: false }
: {
path: base.path,
name: base.name,
isDatasetRoot: base.isDatasetRoot ?? false,
},
);
}
/**
* A definition of a column alteration. The alteration changes the column at
* `path` to have the new name `name`, to be nullable if `nullable` is true,
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "@lancedb/lancedb-darwin-arm64",
"version": "0.38.0-beta.0",
"version": "0.37.1-beta.1",
"os": ["darwin"],
"cpu": ["arm64"],
"main": "lancedb.darwin-arm64.node",
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "@lancedb/lancedb-linux-arm64-gnu",
"version": "0.38.0-beta.0",
"version": "0.37.1-beta.1",
"os": ["linux"],
"cpu": ["arm64"],
"main": "lancedb.linux-arm64-gnu.node",
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "@lancedb/lancedb-linux-arm64-musl",
"version": "0.38.0-beta.0",
"version": "0.37.1-beta.1",
"os": ["linux"],
"cpu": ["arm64"],
"main": "lancedb.linux-arm64-musl.node",
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "@lancedb/lancedb-linux-x64-gnu",
"version": "0.38.0-beta.0",
"version": "0.37.1-beta.1",
"os": ["linux"],
"cpu": ["x64"],
"main": "lancedb.linux-x64-gnu.node",
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "@lancedb/lancedb-linux-x64-musl",
"version": "0.38.0-beta.0",
"version": "0.37.1-beta.1",
"os": ["linux"],
"cpu": ["x64"],
"main": "lancedb.linux-x64-musl.node",
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "@lancedb/lancedb-win32-arm64-msvc",
"version": "0.38.0-beta.0",
"version": "0.37.1-beta.1",
"os": [
"win32"
],
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "@lancedb/lancedb-win32-x64-msvc",
"version": "0.38.0-beta.0",
"version": "0.37.1-beta.1",
"os": ["win32"],
"cpu": ["x64"],
"main": "lancedb.win32-x64-msvc.node",
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "@lancedb/lancedb",
"version": "0.38.0-beta.0",
"version": "0.37.1-beta.1",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "@lancedb/lancedb",
"version": "0.38.0-beta.0",
"version": "0.37.1-beta.1",
"cpu": [
"x64",
"arm64"
+1 -1
View File
@@ -11,7 +11,7 @@
"ann"
],
"private": false,
"version": "0.38.0-beta.0",
"version": "0.37.1-beta.1",
"main": "dist/index.js",
"exports": {
".": "./dist/index.js",
-16
View File
@@ -334,22 +334,6 @@ impl Connection {
.default_error()
}
/// Start dropping a table and return its cleanup job.
#[napi(catch_unwind)]
pub async fn drop_table_async(
&self,
name: String,
namespace_path: Option<Vec<String>>,
) -> napi::Result<crate::job::Job> {
let ns = namespace_path.unwrap_or_default();
let job = self
.get_inner()?
.drop_table_async(&name, &ns)
.await
.default_error()?;
Ok(crate::job::Job::new(job))
}
#[napi(catch_unwind)]
pub async fn drop_all_tables(&self, namespace_path: Option<Vec<String>>) -> napi::Result<()> {
let ns = namespace_path.unwrap_or_default();
-35
View File
@@ -10,7 +10,6 @@ use lancedb::table::{
AddDataMode, ColumnAlteration as LanceColumnAlteration, Duration,
FieldMetadataUpdate as LanceFieldMetadataUpdate, FtsToken as LanceDbFtsToken,
NewColumnTransform, OptimizeAction, OptimizeOptions, Ref, Table as LanceDbTable,
TableBase as LanceTableBase,
};
use napi::bindgen_prelude::*;
use napi::threadsafe_function::{ThreadsafeFunction, ThreadsafeFunctionCallMode};
@@ -372,16 +371,6 @@ impl Table {
Ok(res.into())
}
#[napi(catch_unwind)]
pub async fn refresh_column_async(&self, column: String) -> napi::Result<crate::job::Job> {
let job = self
.inner_ref()?
.refresh_column_async(column)
.await
.default_error()?;
Ok(crate::job::Job::new(job))
}
#[napi(catch_unwind)]
pub async fn add_columns_with_schema(
&self,
@@ -447,18 +436,6 @@ impl Table {
Ok(res.into())
}
#[napi(catch_unwind)]
pub async fn add_bases(&self, bases: Vec<TableBase>) -> napi::Result<()> {
self.inner_ref()?
.add_bases(bases.into_iter().map(|base| LanceTableBase {
path: base.path,
name: base.name,
is_dataset_root: base.is_dataset_root,
}))
.await
.default_error()
}
#[napi(catch_unwind)]
pub async fn drop_columns(&self, columns: Vec<String>) -> napi::Result<DropColumnsResult> {
let col_refs = columns.iter().map(String::as_str).collect::<Vec<_>>();
@@ -713,18 +690,6 @@ impl Table {
}
}
#[napi(object)]
/// An extra storage prefix registered on a table.
pub struct TableBase {
/// Object store URI such as `s3://bucket/media/`.
pub path: String,
/// Optional alias.
pub name: Option<String>,
/// True when `path` is a Lance dataset root. When false, `path` is the
/// directory containing the referenced files.
pub is_dataset_root: bool,
}
#[napi(object)]
/// A description of an index currently configured on a column
pub struct IndexConfig {
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "lancedb-python"
version = "0.38.0-beta.0"
version = "0.37.1-beta.1"
publish = false
edition.workspace = true
description = "Python bindings for LanceDB"
+1 -2
View File
@@ -21,7 +21,7 @@ from .remote.db import RemoteDBConnection
from .expr import Expr, col, lit, func
from .schema import blob, vector, BlobType
from .job import AsyncJob, Job
from .table import AsyncTable, Table, TableBase
from .table import AsyncTable, Table
from .types import BaseTokenizerType
from ._lancedb import Session
from .namespace import (
@@ -521,6 +521,5 @@ __all__ = [
"RemoteDBConnection",
"Session",
"Table",
"TableBase",
"__version__",
]
-5
View File
@@ -198,9 +198,6 @@ class Connection(object):
async def drop_table(
self, name: str, namespace_path: Optional[List[str]] = None
) -> None: ...
async def drop_table_async(
self, name: str, namespace_path: Optional[List[str]] = None
) -> Job: ...
async def drop_all_tables(
self, namespace_path: Optional[List[str]] = None
) -> None: ...
@@ -342,7 +339,6 @@ class Table:
self, columns: list[tuple[str, str]]
) -> AddColumnsResult: ...
async def refresh_column(self, column: str) -> RefreshColumnResult: ...
async def refresh_column_async(self, column: str) -> Job: ...
async def add_columns_with_schema(self, schema: pa.Schema) -> AddColumnsResult: ...
async def alter_columns(
self, columns: list[dict[str, Any]]
@@ -377,7 +373,6 @@ class Table:
def take_offsets(self, offsets: list[int]) -> TakeQuery: ...
def take_row_ids(self, row_ids: list[int]) -> TakeQuery: ...
async def blob_columns(self) -> list[str]: ...
async def add_bases(self, bases: list[Any]) -> None: ...
async def fetch_blobs(
self, column: str, row_ids: list[int]
) -> pa.LargeBinaryArray: ...
-37
View File
@@ -524,12 +524,6 @@ class DBConnection(EnforceOverrides):
namespace_path = []
raise NotImplementedError
def drop_table_async(
self, name: str, namespace_path: Optional[List[str]] = None
) -> Job:
"""Start dropping a table and return its cleanup job."""
raise NotImplementedError
def rename_table(
self,
cur_name: str,
@@ -1192,20 +1186,6 @@ class LanceDBConnection(DBConnection):
)
)
@override
def drop_table_async(
self, name: str, namespace_path: Optional[List[str]] = None
) -> Job:
"""Start dropping a table and return its cleanup job.
The table may become unavailable before its data files are removed.
Call :meth:`Job.wait` to wait for cleanup to finish.
"""
if namespace_path is None:
namespace_path = []
job = LOOP.run(self._conn.drop_table_async(name, namespace_path=namespace_path))
return Job(job if isinstance(job, AsyncJob) else AsyncJob(job))
@override
def drop_all_tables(self, namespace_path: Optional[List[str]] = None):
if namespace_path is None:
@@ -1983,23 +1963,6 @@ class AsyncConnection(object):
if f"Table '{name}' was not found" not in str(e):
raise e
async def drop_table_async(
self,
name: str,
*,
namespace_path: Optional[List[str]] = None,
) -> AsyncJob:
"""Start dropping a table and return its cleanup job.
The table may become unavailable before its data files are removed.
Await :meth:`AsyncJob.wait` to wait for cleanup to finish.
"""
if namespace_path is None:
namespace_path = []
return AsyncJob(
await self._inner.drop_table_async(name, namespace_path=namespace_path)
)
async def drop_all_tables(self, namespace_path: Optional[List[str]] = None):
"""Drop all tables from the database.
-21
View File
@@ -49,7 +49,6 @@ from lancedb._lancedb import (
)
from lancedb.background_loop import LOOP
from lancedb.db import AsyncConnection, DBConnection
from lancedb.job import AsyncJob, Job
from lance_namespace import (
LanceNamespace,
connect as namespace_connect,
@@ -625,18 +624,6 @@ class LanceNamespaceDBConnection(DBConnection):
namespace_path = []
LOOP.run(self._inner.drop_table(name, namespace_path=namespace_path))
@override
def drop_table_async(
self, name: str, namespace_path: Optional[List[str]] = None
) -> Job:
"""Start dropping a table and return its cleanup job."""
if namespace_path is None:
namespace_path = []
job = LOOP.run(
self._inner.drop_table_async(name, namespace_path=namespace_path)
)
return Job(job if isinstance(job, AsyncJob) else AsyncJob(job))
@override
def rename_table(
self,
@@ -1147,14 +1134,6 @@ class AsyncLanceNamespaceDBConnection:
namespace_path = []
await self._inner.drop_table(name, namespace_path=namespace_path)
async def drop_table_async(
self, name: str, namespace_path: Optional[List[str]] = None
) -> AsyncJob:
"""Start dropping a table and return its cleanup job."""
if namespace_path is None:
namespace_path = []
return await self._inner.drop_table_async(name, namespace_path=namespace_path)
async def rename_table(
self,
cur_name: str,
+1 -11
View File
@@ -23,7 +23,7 @@ import pyarrow as pa
from ..common import DATA
from ..db import DBConnection, LOOP
from ..job import AsyncJob, Job
from ..job import Job
if TYPE_CHECKING:
from .._lancedb import JobDescription, JobInfo
@@ -663,16 +663,6 @@ class RemoteDBConnection(DBConnection):
namespace_path = []
LOOP.run(self._conn.drop_table(name, namespace_path=namespace_path))
@override
def drop_table_async(
self, name: str, namespace_path: Optional[List[str]] = None
) -> Job:
"""Start dropping a table and return its cleanup job."""
if namespace_path is None:
namespace_path = []
job = LOOP.run(self._conn.drop_table_async(name, namespace_path=namespace_path))
return Job(job if isinstance(job, AsyncJob) else AsyncJob(job))
@override
def rename_table(
self,
+7 -13
View File
@@ -50,7 +50,7 @@ from lancedb.index import (
)
from lancedb.job import Job
from lancedb.remote.db import LOOP
from lancedb.table import IndexConfigType, KNOWN_METRICS, TableBase
from lancedb.table import IndexConfigType, KNOWN_METRICS
import pyarrow as pa
from lancedb.common import DATA, VEC, VECTOR_COLUMN_NAME
@@ -964,13 +964,14 @@ class RemoteTable(Table):
*,
computed: Dict[str, str] | None = None,
) -> AddColumnsResult:
return LOOP.run(self._table.add_columns(transforms, computed=computed))
if computed:
raise NotImplementedError(
"computed columns are supported only on local tables"
)
return LOOP.run(self._table.add_columns(transforms))
def refresh_column(self, column: str):
return LOOP.run(self._table.refresh_column(column))
def refresh_column_async(self, column: str) -> Job:
return Job(LOOP.run(self._table.refresh_column_async(column)))
raise NotImplementedError("computed columns are supported only on local tables")
def alter_columns(
self, *alterations: Iterable[Dict[str, str]]
@@ -1082,13 +1083,6 @@ class RemoteTable(Table):
def blob_columns(self) -> list[str]:
return LOOP.run(self._table.blob_columns())
def add_bases(
self,
bases: Union[str, TableBase, Iterable[Union[str, TableBase]]],
) -> None:
"""Register additional storage bases for this table."""
LOOP.run(self._table.add_bases(bases))
def fetch_blobs(
self, column: str, row_ids: Union[list[int], pa.Table]
) -> pa.LargeBinaryArray:
+7 -140
View File
@@ -19,7 +19,6 @@ from typing import (
Iterable,
List,
Literal,
Mapping,
Optional,
Sequence,
Tuple,
@@ -711,21 +710,6 @@ def _normalize_progress(progress):
return progress, False
@dataclass
class TableBase:
"""An extra storage prefix registered on a table.
``path`` is an object-store URI. ``name`` is an optional alias.
``is_dataset_root`` is true when ``path`` points to a Lance dataset
root. When false, ``path`` points directly to the directory containing
the referenced files.
"""
path: str
name: Optional[str] = None
is_dataset_root: bool = False
class Table(ABC):
"""
A Table is a collection of Records in a LanceDB Database.
@@ -1584,18 +1568,6 @@ class Table(ABC):
def blob_columns(self) -> list[str]:
"""Names of the blob v2 columns declared on this table."""
def add_bases(
self,
bases: Union[str, TableBase, Iterable[Union[str, TableBase]]],
) -> None:
"""Register additional storage bases for this table.
A URI string is a non-root base with no alias::
table.add_bases("s3://bucket/media/")
"""
raise NotImplementedError
@abstractmethod
def fetch_blobs(
self, column: str, row_ids: Union[list[int], pa.Table]
@@ -1982,10 +1954,8 @@ class Table(ABC):
dropping the column and declaring it again. While a declaration
reads a column, that column cannot be renamed, retyped or dropped.
On LanceDB Cloud and Enterprise the expression is planned by the
server, and the refresh runs as a server job -- see
[`refresh_column_async`][lancedb.table.Table.refresh_column_async].
Cannot be combined with ``transforms``.
Local tables only; LanceDB Cloud and Enterprise raise
``NotImplementedError``. Cannot be combined with ``transforms``.
Returns
-------
@@ -2017,8 +1987,8 @@ class Table(ABC):
by the next one; rows already filled are left as they are, so the call
is idempotent and does not observe a mutated input.
Local tables only: a remote refresh runs as a server job, through
[`refresh_column_async`][lancedb.table.Table.refresh_column_async].
Local tables only; LanceDB Cloud and Enterprise raise
``NotImplementedError``.
Parameters
----------
@@ -2032,31 +2002,6 @@ class Table(ABC):
version: the new version number of the table.
"""
@abstractmethod
def refresh_column_async(self, column: str) -> Job:
"""
Like :meth:`refresh_column`, but returns a handle to the refresh job
instead of blocking until it completes.
The job may already be complete when returned; callers must not assume
the column is filled until :meth:`Job.wait` returns. Invalid input --
an unknown column, or one that is not computed -- raises here rather
than failing the job. On local tables the job runs in-process; on
LanceDB Cloud and Enterprise it is the server's backfill job.
Examples
--------
>>> import lancedb
>>> db = lancedb.connect("./.lancedb")
>>> table = db.create_table("computed_job_demo", [{"x": 1}, {"x": 2}])
>>> table.add_columns(computed={"doubled": "x * 2"})
AddColumnsResult(version=2)
>>> job = table.refresh_column_async("doubled")
>>> job.wait()
>>> job.status()
'finished'
"""
@abstractmethod
def alter_columns(self, *alterations: Iterable[Dict[str, str]]):
"""
@@ -2442,12 +2387,6 @@ class LanceTable(Table):
def blob_columns(self) -> list[str]:
return LOOP.run(self._table.blob_columns())
def add_bases(
self,
bases: Union[str, TableBase, Iterable[Union[str, TableBase]]],
) -> None:
LOOP.run(self._table.add_bases(bases))
def fetch_blobs(
self, column: str, row_ids: Union[list[int], pa.Table]
) -> pa.LargeBinaryArray:
@@ -4081,13 +4020,6 @@ class LanceTable(Table):
[`AsyncTable.refresh_column`][lancedb.AsyncTable.refresh_column]."""
return LOOP.run(self._table.refresh_column(column))
def refresh_column_async(self, column: str) -> Job:
"""Fill a computed column's unfilled rows, returning a handle to the
refresh job. See
[`Table.refresh_column_async`][lancedb.table.Table.refresh_column_async].
"""
return Job(LOOP.run(self._table.refresh_column_async(column)))
def alter_columns(
self, *alterations: Iterable[Dict[str, str]]
) -> AlterColumnsResult:
@@ -6035,8 +5967,7 @@ class AsyncTable:
declaration reads a column, that column cannot be renamed, retyped
or dropped.
On LanceDB Cloud and Enterprise the expression is planned by
the server. Cannot be combined with ``transforms``.
Local tables only. Cannot be combined with ``transforms``.
Returns
-------
@@ -6072,8 +6003,8 @@ class AsyncTable:
by the next one; rows already filled are left as they are, so the call
is idempotent and does not observe a mutated input.
Local tables only: a remote refresh runs as a server job, through
[`refresh_column_async`][lancedb.table.Table.refresh_column_async].
Local tables only; LanceDB Cloud and Enterprise raise
``NotImplementedError``.
Parameters
----------
@@ -6087,34 +6018,6 @@ class AsyncTable:
"""
return await self._inner.refresh_column(column)
async def refresh_column_async(self, column: str) -> AsyncJob:
"""
Like :meth:`refresh_column`, but returns a handle to the refresh job
instead of blocking until it completes.
The job may already be complete when returned; callers must not assume
the column is filled until :meth:`AsyncJob.wait` resolves. Invalid
input -- an unknown column, or one that is not computed -- raises here
rather than failing the job. On local tables the job runs
in-process; on LanceDB Cloud and Enterprise it is the server's
backfill job.
Examples
--------
>>> import asyncio
>>> import lancedb
>>> async def refresh_in_background():
... db = await lancedb.connect_async("./.lancedb")
... table = await db.create_table("computed_job_async_demo", [{"x": 1}])
... await table.add_columns(computed={"doubled": "x * 2"})
... job = await table.refresh_column_async("doubled")
... await job.wait()
... return await job.status()
>>> asyncio.run(refresh_in_background())
'finished'
"""
return AsyncJob(await self._inner.refresh_column_async(column))
async def alter_columns(
self, *alterations: Iterable[dict[str, Any]]
) -> AlterColumnsResult:
@@ -6300,18 +6203,6 @@ class AsyncTable:
async def blob_columns(self) -> list[str]:
return await self._inner.blob_columns()
async def add_bases(
self,
bases: Union[str, TableBase, Iterable[Union[str, TableBase]]],
) -> None:
"""Register additional storage bases for this table.
A URI string is a non-root base with no alias::
await table.add_bases("s3://bucket/media/")
"""
await self._inner.add_bases(_normalize_bases(bases))
async def fetch_blobs(
self, column: str, row_ids: Union[list[int], pa.Table]
) -> pa.LargeBinaryArray:
@@ -6530,30 +6421,6 @@ class AsyncTable:
await self._inner.replace_field_metadata(field_name, new_metadata)
def _normalize_bases(
base_inputs: Union[str, TableBase, Iterable[Union[str, TableBase]]],
) -> list[TableBase]:
if isinstance(base_inputs, (str, TableBase)):
items: Iterable[Union[str, TableBase]] = [base_inputs]
elif isinstance(base_inputs, Mapping):
raise TypeError(
"Expected a URI string, TableBase, or an iterable of those values"
)
else:
items = base_inputs
normalized_bases: list[TableBase] = []
for base in items:
if isinstance(base, str):
normalized_bases.append(TableBase(path=base))
elif isinstance(base, TableBase):
normalized_bases.append(base)
else:
raise TypeError(
f"Expected a URI string or TableBase, got {type(base).__name__}"
)
return normalized_bases
@dataclass
class IndexStatistics:
"""
-72
View File
@@ -1,72 +0,0 @@
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright The LanceDB Authors
import pyarrow as pa
import pytest
import lancedb
def test_add_bases_accepts_named_and_dataset_root(tmp_path):
media = tmp_path / "media"
parent = tmp_path / "parent"
media.mkdir()
parent.mkdir()
db = lancedb.connect(tmp_path / "db")
schema = pa.schema([pa.field("id", pa.int64())])
table = db.create_table("photos", schema=schema)
table.add_bases(
[
lancedb.TableBase(path=media.as_uri(), name="media", is_dataset_root=False),
lancedb.TableBase(
path=parent.as_uri(), name="parent", is_dataset_root=True
),
]
)
def test_add_bases_accepts_two_unnamed_paths(tmp_path):
media = tmp_path / "media"
other = tmp_path / "other"
media.mkdir()
other.mkdir()
db = lancedb.connect(tmp_path / "db")
schema = pa.schema([pa.field("id", pa.int64())])
table = db.create_table("photos", schema=schema)
table.add_bases([media.as_uri(), other.as_uri()])
def test_add_bases_rejects_dict_input(tmp_path):
db = lancedb.connect(tmp_path / "db")
schema = pa.schema([pa.field("id", pa.int64())])
table = db.create_table("photos", schema=schema)
with pytest.raises(TypeError, match="TableBase"):
table.add_bases({"path": "s3://bucket/media/"})
@pytest.mark.asyncio
async def test_async_add_bases_accepts_file_uri(tmp_path):
media = tmp_path / "media"
media.mkdir()
db = await lancedb.connect_async(tmp_path / "db")
schema = pa.schema([pa.field("id", pa.int64())])
table = await db.create_table("photos", schema=schema)
await table.add_bases(media.as_uri())
def test_memory_add_bases_accepts_file_uri(tmp_path):
media = tmp_path / "media"
media.mkdir()
db = lancedb.connect("memory:///")
schema = pa.schema([pa.field("id", pa.int64())])
table = db.create_table("photos", schema=schema)
table.add_bases(media.as_uri())
def test_namespace_add_bases_accepts_file_uri(tmp_path):
media = tmp_path / "media"
media.mkdir()
db = lancedb.connect_namespace("dir", {"root": str(tmp_path / "ns")})
schema = pa.schema([pa.field("id", pa.int64())])
table = db.create_table("photos", schema=schema)
table.add_bases(media.as_uri())
+3 -16
View File
@@ -755,7 +755,8 @@ def test_delete_table(tmp_db: lancedb.DBConnection):
assert tmp_db.table_names() == []
def test_drop_table_async(tmp_db: lancedb.DBConnection):
@pytest.mark.asyncio
async def test_delete_table_async(tmp_db: lancedb.DBConnection):
data = pd.DataFrame(
{
"vector": [[3.1, 4.1], [5.9, 26.5]],
@@ -771,10 +772,7 @@ def test_drop_table_async(tmp_db: lancedb.DBConnection):
assert tmp_db.table_names() == ["test"]
job = tmp_db.drop_table_async("test")
assert job.id is None
assert job.status() == "finished"
job.wait()
tmp_db.drop_table("test")
assert tmp_db.table_names() == []
tmp_db.create_table("test", data=data)
@@ -783,17 +781,6 @@ def test_drop_table_async(tmp_db: lancedb.DBConnection):
tmp_db.drop_table("does_not_exist", ignore_missing=True)
@pytest.mark.asyncio
async def test_drop_table_async_connection(tmp_db_async: lancedb.AsyncConnection):
await tmp_db_async.create_table("test", data=pa.table({"id": [1, 2]}))
job = await tmp_db_async.drop_table_async("test")
assert job.id is None
assert await job.status() == "finished"
await job.wait()
assert await tmp_db_async.table_names() == []
def test_drop_database(tmp_db: lancedb.DBConnection):
data = pd.DataFrame(
{
-33
View File
@@ -2306,36 +2306,3 @@ def test_remote_connection_jobs_surface():
assert job.status() == "failed"
with pytest.raises(JobFailedError, match="worker died"):
job.wait(timeout=timedelta(seconds=5))
def test_remote_add_bases_posts_the_bases_array():
captured_body = {}
def handler(request):
if request.path == "/v1/table/test/describe/":
request.send_response(200)
request.send_header("Content-Type", "application/json")
request.end_headers()
request.wfile.write(json.dumps(BLOB_DESCRIBE_RESPONSE).encode())
elif request.path == "/v1/table/test/bases/":
content_len = int(request.headers.get("Content-Length", 0))
captured_body.update(json.loads(request.rfile.read(content_len)))
request.send_response(200)
request.send_header("Content-Type", "application/json")
request.end_headers()
request.wfile.write(b'{"version": 2}')
else:
request.send_response(404)
request.end_headers()
with mock_lancedb_connection(handler) as db:
table = db.open_table("test")
table.add_bases(lancedb.TableBase(path="s3://bucket/media/"))
assert captured_body["bases"] == [
{
"path": "s3://bucket/media/",
"isDatasetRoot": False,
}
]
-28
View File
@@ -3888,31 +3888,3 @@ async def test_computed_column_async(tmp_path):
await table.refresh_column("tripled")
assert (await table.to_arrow())["tripled"].to_pylist() == [9]
def test_refresh_column_async_returns_job(tmp_path):
db = lancedb.connect(tmp_path)
table = db.create_table("computed_job", [{"x": 1}, {"x": 2}])
table.add_columns(computed={"doubled": "x * 2"})
job = table.refresh_column_async("doubled")
assert job.id is None # in-process jobs have no server id
job.wait()
assert job.status() == "finished"
assert sorted(table.to_arrow()["doubled"].to_pylist()) == [2, 4]
# Bad input raises at the call, not through the job.
with pytest.raises(Exception, match="not a computed column"):
table.refresh_column_async("x")
@pytest.mark.asyncio
async def test_refresh_column_async_job_async_table(tmp_path):
db = await lancedb.connect_async(tmp_path)
table = await db.create_table("computed_job_async", [{"x": 3}])
await table.add_columns(computed={"tripled": "x * 3"})
job = await table.refresh_column_async("tripled")
await job.wait()
assert await job.status() == "finished"
assert (await table.to_arrow())["tripled"].to_pylist() == [9]
-17
View File
@@ -346,23 +346,6 @@ impl Connection {
})
}
#[pyo3(signature = (name, namespace_path=None))]
pub fn drop_table_async(
self_: PyRef<'_, Self>,
name: String,
namespace_path: Option<Vec<String>>,
) -> PyResult<Bound<'_, PyAny>> {
let inner = self_.get_inner()?.clone();
let ns_path = namespace_path.unwrap_or_default();
future_into_py(self_.py(), async move {
inner
.drop_table_async(name, &ns_path)
.await
.infer_error()
.map(crate::job::Job::new)
})
}
#[pyo3(signature = (namespace_path=None,))]
pub fn drop_all_tables(
self_: PyRef<'_, Self>,
-38
View File
@@ -22,7 +22,6 @@ use lancedb::index::scalar::FtsIndexBuilder;
use lancedb::table::{
AddDataMode, ColumnAlteration, Duration, FieldMetadataUpdate, FtsToken as LanceDbFtsToken,
NewColumnTransform, OptimizeAction, OptimizeOptions, Ref, Table as LanceDbTable,
TableBase as LanceTableBase,
};
use lancedb::tokenize as lancedb_tokenize;
use pyo3::{
@@ -95,13 +94,6 @@ fn lsm_stats_to_py(py: Python<'_>, stats: &lancedb::table::LsmStats) -> PyResult
Ok(out.unbind())
}
#[derive(FromPyObject)]
pub(crate) struct PyTableBase {
path: String,
name: Option<String>,
is_dataset_root: bool,
}
#[derive(FromPyObject)]
enum PredicateArg {
Expr(PyExpr),
@@ -1246,25 +1238,6 @@ impl Table {
})
}
#[pyo3(signature = (bases))]
pub fn add_bases(
self_: PyRef<'_, Self>,
bases: Vec<PyTableBase>,
) -> PyResult<Bound<'_, PyAny>> {
let inner = self_.inner_ref()?.clone();
let bases: Vec<LanceTableBase> = bases
.into_iter()
.map(|base| LanceTableBase {
path: base.path,
name: base.name,
is_dataset_root: base.is_dataset_root,
})
.collect();
future_into_py(self_.py(), async move {
inner.add_bases(bases).await.infer_error()
})
}
/// Read blob bytes for `row_ids` from blob v2 column `column`.
#[pyo3(signature = (column, row_ids))]
pub fn fetch_blobs(
@@ -1586,17 +1559,6 @@ impl Table {
})
}
pub fn refresh_column_async(
self_: PyRef<'_, Self>,
column: String,
) -> PyResult<Bound<'_, PyAny>> {
let inner = self_.inner_ref()?.clone();
future_into_py(self_.py(), async move {
let job = inner.refresh_column_async(column).await.infer_error()?;
Ok(crate::job::Job::new(job))
})
}
pub fn add_columns_with_schema(
self_: PyRef<'_, Self>,
schema: PyArrowType<Schema>,
+1 -4
View File
@@ -1,6 +1,6 @@
[package]
name = "lancedb"
version = "0.38.0-beta.0"
version = "0.37.1-beta.1"
edition.workspace = true
description = "LanceDB: A serverless, low-latency vector database for AI applications"
license.workspace = true
@@ -188,9 +188,6 @@ required-features = ["bedrock"]
[[example]]
name = "bench_streaming_dataloader"
[[example]]
name = "bench_open_missing_table"
[[example]]
name = "simple"
@@ -1,150 +0,0 @@
// SPDX-License-Identifier: Apache-2.0
// SPDX-FileCopyrightText: Copyright The LanceDB Authors
// Release benchmark for opening a missing table as sibling-table cardinality grows.
//
// The fixture uses real `.lance` directories and marker files. Fixture creation is
// outside the timed section. Defaults intentionally cover 1k, 10k, and 100k siblings
// with 10 warmups and 100 distinct missing-table opens per scale:
//
// ```text
// cargo run --release -p lancedb --example bench_open_missing_table
// ```
//
// `BENCH_SIBLINGS`, `BENCH_WARMUPS`, and `BENCH_TRIALS` override those defaults.
// Reduced settings are useful only as a smoke test. Performance comparisons require
// the same machine, filesystem, fixture sizes, settings, lockfile, and alternating
// baseline/candidate execution order.
use std::time::{Duration, Instant};
use anyhow::{Context, Result, bail};
use lancedb::connection::Connection;
use lancedb::{Error, connect};
use object_store::ObjectStoreExt as _;
use object_store::path::Path;
const MAX_SIBLINGS: usize = 1_000_000;
const MAX_WARMUPS: usize = 10_000;
const MAX_TRIALS: usize = 100_000;
fn env_usize(key: &str, default: usize, max: usize) -> Result<usize> {
let value = match std::env::var(key) {
Ok(value) => value
.parse()
.with_context(|| format!("invalid {key} value: {value}"))?,
Err(std::env::VarError::NotPresent) => default,
Err(error) => return Err(error).with_context(|| format!("reading {key}")),
};
if value == 0 || value > max {
bail!("{key} must be between 1 and {max}");
}
Ok(value)
}
fn sibling_counts() -> Result<Vec<usize>> {
let raw = std::env::var("BENCH_SIBLINGS").unwrap_or_else(|_| "1000,10000,100000".into());
let mut counts = raw
.split(',')
.map(|value| {
value
.trim()
.parse::<usize>()
.with_context(|| format!("invalid BENCH_SIBLINGS value: {value}"))
})
.collect::<Result<Vec<_>>>()?;
counts.sort_unstable();
counts.dedup();
if counts.is_empty() || counts[0] == 0 || counts[counts.len() - 1] > MAX_SIBLINGS {
bail!("BENCH_SIBLINGS values must be between 1 and {MAX_SIBLINGS}");
}
Ok(counts)
}
async fn add_siblings(
store: &object_store::local::LocalFileSystem,
start: usize,
end: usize,
) -> Result<()> {
for index in start..end {
let marker = Path::from(format!("sibling_{index:06}.lance/_marker"));
store
.put(&marker, bytes::Bytes::new().into())
.await
.with_context(|| format!("creating benchmark marker {marker}"))?;
}
Ok(())
}
async fn time_missing_open(db: &Connection, name: &str) -> Result<Duration> {
let started = Instant::now();
let result = db.open_table(name).execute().await;
let elapsed = started.elapsed();
match result {
Err(Error::TableNotFound { .. }) => Ok(elapsed),
Err(error) => bail!("expected TableNotFound for {name}, got {error:?}"),
Ok(_) => bail!("benchmark missing-table name unexpectedly exists: {name}"),
}
}
fn percentile(sorted: &[Duration], percentile: usize) -> Duration {
let rank = (sorted.len() * percentile).div_ceil(100).saturating_sub(1);
sorted[rank]
}
#[tokio::main]
async fn main() -> Result<()> {
let counts = sibling_counts()?;
let warmups = env_usize("BENCH_WARMUPS", 10, MAX_WARMUPS)?;
let trials = env_usize("BENCH_TRIALS", 100, MAX_TRIALS)?;
let fixture = tempfile::tempdir().context("creating benchmark fixture")?;
let database_path = fixture.path();
let fixture_store = object_store::local::LocalFileSystem::new_with_prefix(database_path)
.context("creating benchmark object store")?;
let db = connect(database_path.to_str().context("non-UTF-8 fixture path")?)
.execute()
.await?;
println!(
"config: siblings={counts:?} warmups={warmups} trials={trials} profile={} os={} arch={}",
if cfg!(debug_assertions) {
"debug"
} else {
"release"
},
std::env::consts::OS,
std::env::consts::ARCH,
);
println!("lower is better; fixture setup and teardown are excluded");
println!("| siblings | samples | p50 | p95 | max |");
println!("| ---: | ---: | ---: | ---: | ---: |");
let mut created = 0;
for sibling_count in counts {
add_siblings(&fixture_store, created, sibling_count).await?;
created = sibling_count;
for index in 0..warmups {
let name = format!("__missing_warmup_{sibling_count}_{index}");
let _ = time_missing_open(&db, &name).await?;
}
let mut samples = Vec::with_capacity(trials);
for index in 0..trials {
let name = format!("__missing_trial_{sibling_count}_{index}");
samples.push(time_missing_open(&db, &name).await?);
}
samples.sort_unstable();
println!(
"| {sibling_count} | {} | {:?} | {:?} | {:?} |",
samples.len(),
percentile(&samples, 50),
percentile(&samples, 95),
samples[samples.len() - 1],
);
}
Ok(())
}
+4 -23
View File
@@ -409,11 +409,6 @@ impl Connection {
///
/// The names will be returned in lexicographical order (ascending)
///
/// Listing databases discover physical `*.lance` entries without opening every
/// dataset. The result is a point-in-time discovery snapshot: an entry may still be
/// under creation, may contain only uncommitted storage, or may be concurrently
/// dropped before it is opened.
///
/// The parameters `page_token` and `limit` can be used to paginate the results
pub fn table_names(&self) -> TableNamesBuilder {
TableNamesBuilder::new(self.internal.clone())
@@ -461,9 +456,10 @@ impl Connection {
///
/// # Returns
/// Created [`TableRef`], or [`Error::TableNotFound`] if the table does not exist.
/// On listing databases, a committed Lance manifest is authoritative for table
/// existence. Uncommitted files or a physical `<name>.lance` directory alone do not
/// make a table openable.
/// If the table's storage is present but holds no readable dataset (for example a
/// `<name>.lance` directory left behind by an interrupted drop and re-create, which
/// [`Self::table_names`] still lists) this returns [`Error::TableCorrupted`]
/// instead.
pub fn open_table(&self, name: impl Into<String>) -> OpenTableBuilder {
OpenTableBuilder::new(
self.internal.clone(),
@@ -565,21 +561,6 @@ impl Connection {
.await
}
/// Start dropping a table and return a handle to the cleanup job.
///
/// The table may become unavailable before its physical data is removed.
/// Call [`crate::job::Job::wait`] to wait for cleanup to finish. Local
/// backends may complete the drop before returning the handle.
pub async fn drop_table_async(
&self,
name: impl AsRef<str>,
namespace_path: &[String],
) -> Result<crate::job::Job> {
self.internal
.drop_table_async(name.as_ref(), namespace_path)
.await
}
/// Drop the database
///
/// This is the same as dropping all of the tables
-12
View File
@@ -323,18 +323,6 @@ pub trait Database:
) -> Result<()>;
/// Drop a table in the database
async fn drop_table(&self, name: &str, namespace_path: &[String]) -> Result<()>;
/// Start dropping a table and return a handle to the cleanup job.
///
/// Backends without asynchronous cleanup complete the drop before
/// returning an already-finished job.
async fn drop_table_async(
&self,
name: &str,
namespace_path: &[String],
) -> Result<crate::job::Job> {
self.drop_table(name, namespace_path).await?;
Ok(crate::job::Job::new_done())
}
/// Drop all tables in the database
async fn drop_all_tables(&self, namespace_path: &[String]) -> Result<()>;
fn as_any(&self) -> &dyn std::any::Any;
+2 -116
View File
@@ -1032,7 +1032,6 @@ impl Database for ListingDatabase {
};
Ok(ListTablesResponse {
context: None,
tables: f,
page_token: next_page_token,
})
@@ -1292,21 +1291,16 @@ impl Database for ListingDatabase {
mod tests {
use super::*;
use crate::Table;
use crate::arrow::{SendableRecordBatchStream, SimpleRecordBatchStream};
use crate::connection::ConnectRequest;
use crate::data::scannable::Scannable;
use crate::database::{CreateTableMode, CreateTableRequest};
use crate::query::QueryRequest;
use crate::table::{AnyQuery, WriteOptions};
use arrow_array::{Int32Array, RecordBatch, StringArray};
use arrow_schema::{DataType, Field, Schema, SchemaRef};
use futures::{TryStreamExt, stream::once};
use arrow_schema::{DataType, Field, Schema};
use futures::TryStreamExt;
use std::path::PathBuf;
use std::sync::Arc;
use std::time::Duration;
use tempfile::tempdir;
use tokio::sync::Barrier;
use tokio::time::timeout;
async fn setup_database() -> (tempfile::TempDir, ListingDatabase) {
let tempdir = tempdir().unwrap();
@@ -1330,114 +1324,6 @@ mod tests {
(tempdir, db)
}
struct BarrierScannable {
batch: RecordBatch,
barrier: Arc<Barrier>,
}
impl Scannable for BarrierScannable {
fn schema(&self) -> SchemaRef {
self.batch.schema()
}
fn scan_as_stream(&mut self) -> SendableRecordBatchStream {
let batch = self.batch.clone();
let schema = batch.schema();
let barrier = self.barrier.clone();
Box::pin(SimpleRecordBatchStream {
schema,
stream: once(async move {
barrier.wait().await;
Ok(batch)
}),
})
}
}
fn create_request(name: &str, data: Box<dyn Scannable>) -> CreateTableRequest {
CreateTableRequest {
name: name.to_string(),
namespace_path: vec![],
data,
mode: CreateTableMode::Create,
write_options: Default::default(),
location: None,
namespace_client: None,
}
}
#[tokio::test]
async fn test_create_ignores_uncommitted_storage_without_manifest() {
let (tmp_dir, db) = setup_database().await;
let data_dir = tmp_dir.path().join("test.lance/data");
std::fs::create_dir_all(&data_dir).unwrap();
std::fs::write(data_dir.join("orphan.lance"), b"uncommitted").unwrap();
let schema = Arc::new(Schema::new(vec![Field::new("id", DataType::Int32, false)]));
let batch =
RecordBatch::try_new(schema, vec![Arc::new(Int32Array::from(vec![1]))]).unwrap();
let table = db
.create_table(create_request("test", Box::new(batch)))
.await
.unwrap();
assert_eq!(table.count_rows(None).await.unwrap(), 1);
}
#[tokio::test]
async fn test_concurrent_create_is_arbitrated_by_manifest_commit() {
let uri = format!("memory:///concurrent-create-{}", uuid::Uuid::new_v4());
let db = crate::connect(&uri).execute().await.unwrap();
let store: Arc<dyn object_store::ObjectStore> =
Arc::new(object_store::memory::InMemory::new());
let table_url = url::Url::parse("memory:///database/test.lance").unwrap();
let schema = Arc::new(Schema::new(vec![Field::new("id", DataType::Int32, false)]));
let batch =
RecordBatch::try_new(schema, vec![Arc::new(Int32Array::from(vec![1]))]).unwrap();
let barrier = Arc::new(Barrier::new(2));
#[allow(deprecated)]
let request = |batch, barrier| {
let mut request = create_request("test", Box::new(BarrierScannable { batch, barrier }));
request.write_options = WriteOptions {
lance_write_params: Some(lance::dataset::WriteParams {
store_params: Some(ObjectStoreParams {
object_store: Some((store.clone(), table_url.clone())),
..Default::default()
}),
commit_handler: Some(Arc::new(
lance_table::io::commit::ConditionalPutCommitHandler,
)),
..Default::default()
}),
};
request
};
let left = db
.database()
.create_table(request(batch.clone(), barrier.clone()));
let right = db.database().create_table(request(batch, barrier));
let (left, right) = timeout(Duration::from_secs(30), async { tokio::join!(left, right) })
.await
.expect("concurrent creates deadlocked");
let results = [left, right];
assert_eq!(
results.iter().filter(|result| result.is_ok()).count(),
1,
"expected one successful create, got {results:?}"
);
assert_eq!(
results
.iter()
.filter(|result| matches!(result, Err(Error::TableAlreadyExists { .. })))
.count(),
1,
"expected one manifest conflict, got {results:?}"
);
}
#[tokio::test]
async fn test_listing_database_root_ops_do_not_create_manifest() {
let tempdir = tempdir().unwrap();
+1 -1
View File
@@ -141,7 +141,7 @@ impl SpawnedJob {
Ok(Err(err)) => Outcome::Failed(Arc::new(err)),
Err(err) if err.is_cancelled() => Outcome::Cancelled,
Err(err) => Outcome::Failed(Arc::new(Error::Runtime {
message: format!("job task failed: {err}"),
message: format!("index job task failed: {err}"),
})),
};
let _ = tx.send(Some(outcome));
+1 -1
View File
@@ -214,7 +214,7 @@ use lance_linalg::distance::DistanceType as LanceDistanceType;
/// a built-in pull-based adapter.
#[cfg(feature = "metrics")]
pub use metrics;
pub use table::{FtsToken, Table, TableBase};
pub use table::{FtsToken, Table};
/// Tokenize a full-text search query using an explicit FTS tokenizer configuration.
///
-9
View File
@@ -19,15 +19,6 @@ const ARROW_FILE_CONTENT_TYPE: &str = "application/vnd.apache.arrow.file";
#[cfg(test)]
const JSON_CONTENT_TYPE: &str = "application/json";
fn extract_job_id(body: &str) -> Option<String> {
serde_json::from_str::<serde_json::Value>(body)
.ok()?
.get("job_id")?
.as_str()
.filter(|job_id| !job_id.is_empty())
.map(str::to_string)
}
pub use client::{ClientConfig, HeaderProvider, RetryConfig, TimeoutConfig, TlsConfig};
pub use db::{RemoteDatabaseOptions, RemoteDatabaseOptionsBuilder};
pub use oauth::{OAuthConfig, OAuthFlow, OAuthHeaderProvider};
+8 -103
View File
@@ -9,7 +9,6 @@ use http::StatusCode;
use lance_io::object_store::StorageOptions;
use lance_namespace_impls::{DynamicContextProvider, OperationInfo};
use moka::future::Cache;
use reqwest::Response;
use reqwest::header::CONTENT_TYPE;
use lance_namespace::models::{
@@ -24,17 +23,15 @@ use crate::database::{
JobDescription, JobInfo, OpenTableRequest, ReadConsistency, TableNamesRequest,
};
use crate::error::Result;
use crate::job::Job;
use crate::remote::job::RemoteJob;
use crate::remote::util::stream_as_body;
use crate::table::BaseTable;
use super::ARROW_STREAM_CONTENT_TYPE;
use super::client::{
ClientConfig, HeaderProvider, HttpSend, RequestResultExt, RestfulLanceDbClient, Sender,
};
use super::table::RemoteTable;
use super::util::parse_server_version;
use super::{ARROW_STREAM_CONTENT_TYPE, extract_job_id};
// Request structure for the remote clone table API
#[derive(serde::Serialize)]
@@ -329,22 +326,6 @@ impl RemoteDatabase {
}
}
impl<S: HttpSend> RemoteDatabase<S> {
async fn submit_drop_table(
&self,
name: &str,
namespace_path: &[String],
) -> Result<(String, Response)> {
let identifier = build_table_identifier(name, namespace_path, &self.client.id_delimiter);
let cache_key = build_cache_key(name, namespace_path);
let req = self.client.post(&format!("/v1/table/{}/drop/", identifier));
let (request_id, resp) = self.client.send(req).await?;
let resp = self.client.check_response(&request_id, resp).await?;
self.table_cache.remove(&cache_key).await;
Ok((request_id, resp))
}
}
#[cfg(all(test, feature = "remote"))]
mod test_utils {
use super::*;
@@ -913,28 +894,13 @@ impl<S: HttpSend> Database for RemoteDatabase<S> {
}
async fn drop_table(&self, name: &str, namespace_path: &[String]) -> Result<()> {
self.submit_drop_table(name, namespace_path)
.await
.map(|_| ())
}
async fn drop_table_async(&self, name: &str, namespace_path: &[String]) -> Result<Job> {
let (request_id, response) = self.submit_drop_table(name, namespace_path).await?;
let status = response.status();
let body = response.text().await.err_to_http(request_id.clone())?;
let job_id = extract_job_id(&body);
Ok(match job_id {
Some(job_id) => Job::new(Box::new(RemoteJob::new(self.client.clone(), job_id))),
None if status == StatusCode::ACCEPTED => {
return Err(Error::Http {
source: "asynchronous drop-table response did not contain a valid job_id"
.into(),
request_id,
status_code: Some(status),
});
}
None => Job::new_done(),
})
let identifier = build_table_identifier(name, namespace_path, &self.client.id_delimiter);
let cache_key = build_cache_key(name, namespace_path);
let req = self.client.post(&format!("/v1/table/{}/drop/", identifier));
let (request_id, resp) = self.client.send(req).await?;
self.client.check_response(&request_id, resp).await?;
self.table_cache.remove(&cache_key).await;
Ok(())
}
async fn drop_all_tables(&self, namespace_path: &[String]) -> Result<()> {
@@ -1526,67 +1492,6 @@ mod tests {
// NOTE: the API will return 200 even if the table does not exist. So we shouldn't expect 404.
}
#[tokio::test]
async fn test_drop_table_does_not_read_response_body() {
let conn = Connection::new_with_handler(|_| {
http::Response::builder()
.status(200)
.body(vec![0xff])
.unwrap()
});
conn.drop_table("table1", &[]).await.unwrap();
}
#[tokio::test]
async fn test_drop_table_async_returns_job() {
let conn = Connection::new_with_handler(|request| {
assert_eq!(request.method(), &reqwest::Method::POST);
assert_eq!(request.url().path(), "/v1/table/table1/drop/");
http::Response::builder()
.status(202)
.body(r#"{"job_id":"drop-job-123"}"#)
.unwrap()
});
let job = conn.drop_table_async("table1", &[]).await.unwrap();
assert_eq!(job.id(), Some("drop-job-123"));
}
#[tokio::test]
async fn test_drop_table_async_old_server_returns_done_job() {
let conn = Connection::new_with_handler(|_| {
http::Response::builder().status(200).body("").unwrap()
});
let job = conn.drop_table_async("table1", &[]).await.unwrap();
assert_eq!(job.id(), None);
assert_eq!(job.status().await.unwrap(), "finished");
}
#[tokio::test]
async fn test_drop_table_async_rejects_accepted_response_without_job_id() {
let conn = Connection::new_with_handler(|_| {
http::Response::builder().status(202).body("{}").unwrap()
});
let error = conn.drop_table_async("table1", &[]).await.err().unwrap();
assert!(error.to_string().contains("valid job_id"));
}
#[tokio::test]
async fn test_drop_table_async_rejects_empty_job_id() {
let conn = Connection::new_with_handler(|_| {
http::Response::builder()
.status(202)
.body(r#"{"job_id":""}"#)
.unwrap()
});
let error = conn.drop_table_async("table1", &[]).await.err().unwrap();
assert!(error.to_string().contains("valid job_id"));
}
#[tokio::test]
async fn test_rename_table() {
let conn = Connection::new_with_handler(|request| {
+46 -499
View File
@@ -1,7 +1,6 @@
// SPDX-License-Identifier: Apache-2.0
// SPDX-FileCopyrightText: Copyright The LanceDB Authors
pub mod bases;
pub mod blobs;
pub mod insert;
@@ -9,7 +8,7 @@ use self::insert::{RemoteWriteExec, WriteOp};
use super::client::RequestResultExt;
use super::client::{HttpSend, RestfulLanceDbClient, Sender};
use super::db::ServerVersion;
use super::{ARROW_FILE_CONTENT_TYPE, ARROW_STREAM_CONTENT_TYPE, extract_job_id};
use super::{ARROW_FILE_CONTENT_TYPE, ARROW_STREAM_CONTENT_TYPE};
use crate::blob::BlobFile;
use crate::data::scannable::{PeekedScannable, Scannable, estimate_write_partitions};
use crate::expr::expr_to_sql_string;
@@ -34,9 +33,7 @@ use crate::table::lsm_stats::GetLsmStatsResponse;
use crate::table::merge::MergeFilter;
use crate::table::query::create_multi_vector_plan;
use crate::table::write_progress::FinishOnDrop;
use crate::table::{
AlterColumnsResult, FieldMetadataUpdate, RefreshColumnResult, UpdateFieldMetadataResult,
};
use crate::table::{AlterColumnsResult, FieldMetadataUpdate, UpdateFieldMetadataResult};
use crate::table::{AnyQuery, Filter, Predicate, PreprocessingOutput, TableStatistics};
use crate::utils::background_cache::BackgroundCache;
use crate::utils::{
@@ -143,40 +140,6 @@ impl FreshnessHeaders {
}
}
/// A backfill job whose successful wait establishes a read-freshness
/// baseline on the submitting handle, so a later read cannot be served
/// from a cache older than the completed fill. A handle pinned by checkout
/// at completion keeps its time-travel view instead.
struct FreshnessJob<S: HttpSend> {
inner: RemoteJob<S>,
freshness: Arc<Mutex<FreshnessState>>,
version: Arc<RwLock<Option<u64>>>,
}
#[async_trait]
impl<S: HttpSend> crate::job::JobHandle for FreshnessJob<S> {
fn id(&self) -> Option<&str> {
crate::job::JobHandle::id(&self.inner)
}
async fn status(&self) -> Result<String> {
crate::job::JobHandle::status(&self.inner).await
}
async fn wait(&self) -> Result<()> {
crate::job::JobHandle::wait(&self.inner).await?;
let version = self.version.read().await;
if version.is_none() {
self.freshness.lock().unwrap().checkout_baseline = Some(SystemTime::now());
}
Ok(())
}
async fn cancel(&self) -> Result<()> {
crate::job::JobHandle::cancel(&self.inner).await
}
}
fn compute_min_timestamp(
state: &FreshnessState,
interval: Option<Duration>,
@@ -311,10 +274,10 @@ pub struct RemoteTable<S: HttpSend = Sender> {
identifier: String,
server_version: ServerVersion,
version: Arc<RwLock<Option<u64>>>,
version: RwLock<Option<u64>>,
location: RwLock<Option<String>>,
schema_cache: BackgroundCache<SchemaRef, Error>,
freshness: Arc<Mutex<FreshnessState>>,
freshness: Mutex<FreshnessState>,
/// The branch this handle is scoped to, or `None` for the main branch.
/// Stamped onto every branch-accepting request so reads and writes resolve
/// on the branch's own version chain rather than main's.
@@ -429,7 +392,13 @@ impl<S: HttpSend> RemoteTable<S> {
.text()
.await
.ok()
.and_then(|body| extract_job_id(&body));
.and_then(|body| serde_json::from_str::<serde_json::Value>(&body).ok())
.and_then(|value| {
value
.get("job_id")
.and_then(|id| id.as_str())
.map(str::to_string)
});
if let Some(wait_timeout) = index.wait_timeout {
let index_name = index.name.unwrap_or_else(|| format!("{}_idx", column));
@@ -452,10 +421,10 @@ impl<S: HttpSend> RemoteTable<S> {
namespace,
identifier,
server_version,
version: Arc::new(RwLock::new(None)),
version: RwLock::new(None),
location: RwLock::new(None),
schema_cache: BackgroundCache::new(SCHEMA_CACHE_TTL, SCHEMA_CACHE_REFRESH_WINDOW),
freshness: Arc::new(Mutex::new(FreshnessState::default())),
freshness: Mutex::new(FreshnessState::default()),
branch: None,
}
}
@@ -484,10 +453,10 @@ impl<S: HttpSend> RemoteTable<S> {
namespace: self.namespace.clone(),
identifier: self.identifier.clone(),
server_version: self.server_version.clone(),
version: Arc::new(RwLock::new(None)),
version: RwLock::new(None),
location: RwLock::new(None),
schema_cache: BackgroundCache::new(SCHEMA_CACHE_TTL, SCHEMA_CACHE_REFRESH_WINDOW),
freshness: Arc::new(Mutex::new(FreshnessState::default())),
freshness: Mutex::new(FreshnessState::default()),
branch,
}
}
@@ -1305,10 +1274,10 @@ mod test_utils {
namespace: vec![],
identifier: name,
server_version: version.map(ServerVersion).unwrap_or_default(),
version: Arc::new(RwLock::new(None)),
version: RwLock::new(None),
location: RwLock::new(None),
schema_cache: BackgroundCache::new(SCHEMA_CACHE_TTL, SCHEMA_CACHE_REFRESH_WINDOW),
freshness: Arc::new(Mutex::new(FreshnessState::default())),
freshness: Mutex::new(FreshnessState::default()),
branch: None,
}
}
@@ -1329,10 +1298,10 @@ mod test_utils {
namespace: vec![],
identifier: name,
server_version: ServerVersion::default(),
version: Arc::new(RwLock::new(None)),
version: RwLock::new(None),
location: RwLock::new(None),
schema_cache: BackgroundCache::new(SCHEMA_CACHE_TTL, SCHEMA_CACHE_REFRESH_WINDOW),
freshness: Arc::new(Mutex::new(FreshnessState::default())),
freshness: Mutex::new(FreshnessState::default()),
branch: None,
}
}
@@ -1362,10 +1331,10 @@ mod test_utils {
namespace: vec![],
identifier: name,
server_version: version.map(ServerVersion).unwrap_or_default(),
version: Arc::new(RwLock::new(None)),
version: RwLock::new(None),
location: RwLock::new(None),
schema_cache: BackgroundCache::new(SCHEMA_CACHE_TTL, SCHEMA_CACHE_REFRESH_WINDOW),
freshness: Arc::new(Mutex::new(FreshnessState::default())),
freshness: Mutex::new(FreshnessState::default()),
branch: None,
}
}
@@ -2245,10 +2214,6 @@ impl<S: HttpSend> BaseTable for RemoteTable<S> {
self.blob_columns_impl().await
}
async fn add_bases(&self, bases: &[crate::table::TableBase]) -> Result<()> {
self.add_bases_impl(bases).await
}
async fn fetch_blobs(&self, column: &str, row_ids: &[u64]) -> Result<LargeBinaryArray> {
self.fetch_blobs_impl(column, row_ids).await
}
@@ -2741,6 +2706,13 @@ impl<S: HttpSend> BaseTable for RemoteTable<S> {
Ok(result)
}
// A declaration reaches here as AllNulls, which the remote protocol
// has no representation for.
NewColumnTransform::AllNulls(_) => {
return Err(Error::NotSupported {
message: "computed columns are supported only on local tables".into(),
});
}
_ => {
return Err(Error::NotSupported {
message: "Only SQL expressions are supported for adding columns".into(),
@@ -2749,86 +2721,6 @@ impl<S: HttpSend> BaseTable for RemoteTable<S> {
}
}
async fn add_computed_columns(&self, columns: &[(String, String)]) -> Result<AddColumnsResult> {
self.check_mutable().await?;
// The server plans the declaration: expression validation, type
// inference and the persisted binding all happen there.
let entries = columns
.iter()
.map(
|(name, expression)| lance_namespace::models::AddColumnsEntry {
name: name.clone(),
computed: Some(Some(expression.clone())),
..Default::default()
},
)
.collect::<Vec<_>>();
let mut body = serde_json::json!({ "new_columns": entries });
self.apply_branch_body(&mut body);
let request = self
.client
.post(&format!("/v1/table/{}/add_columns/", self.identifier))
.json(&body);
let (request_id, response) = self.send(request, true).await?;
let response = self.check_table_response(&request_id, response).await?;
let body = response.text().await.err_to_http(request_id.clone())?;
if body.trim().is_empty() {
// Backward compatible with old servers
return Ok(AddColumnsResult { version: 0 });
}
let result: AddColumnsResult = serde_json::from_str(&body).map_err(|e| Error::Http {
source: format!("Failed to parse add_columns response: {}", e).into(),
request_id,
status_code: None,
})?;
self.invalidate_schema_cache();
self.track_write_version(result.version);
Ok(result)
}
async fn refresh_column(&self, _column: &str) -> Result<RefreshColumnResult> {
// The server runs a refresh as a job and does not report a fill
// count, so the blocking form has no honest result to return.
Err(Error::NotSupported {
message: "a remote refresh runs as a server job; use refresh_column_async and \
wait on the returned handle"
.into(),
})
}
async fn refresh_column_async(&self, column: &str) -> Result<Job> {
self.check_mutable().await?;
let mut body = serde_json::json!({ "column": column });
self.apply_branch_body(&mut body);
let request = self
.client
.post(&format!("/v1/table/{}/backfill_column", self.identifier))
.json(&body);
let (request_id, response) = self.send(request, true).await?;
let response = self.check_table_response(&request_id, response).await?;
let body = response.text().await.err_to_http(request_id.clone())?;
#[derive(serde::Deserialize)]
struct BackfillResponse {
job_id: String,
}
let response: BackfillResponse = serde_json::from_str(&body).map_err(|e| Error::Http {
source: format!("Failed to parse backfill_column response: {}", e).into(),
request_id,
status_code: None,
})?;
Ok(Job::new(Box::new(FreshnessJob {
inner: RemoteJob::new(self.client.clone(), response.job_id),
freshness: self.freshness.clone(),
version: self.version.clone(),
})))
}
async fn alter_columns(&self, alterations: &[ColumnAlteration]) -> Result<AlterColumnsResult> {
self.check_mutable().await?;
let body = alterations
@@ -4094,42 +3986,6 @@ mod tests {
.unwrap()
}
#[tokio::test]
async fn test_add_bases_posts_the_bases_array() {
let table = Table::new_with_handler("my_table", |request| {
assert_eq!(request.method(), "POST");
assert_eq!(request.url().path(), "/v1/table/my_table/bases/");
let body: serde_json::Value =
serde_json::from_slice(request.body().unwrap().as_bytes().unwrap()).unwrap();
assert_eq!(
body["bases"],
serde_json::json!([{
"path": "s3://bucket/media/",
"isDatasetRoot": false
}])
);
http::Response::builder()
.status(200)
.body(r#"{"version": 4}"#)
.unwrap()
});
table.add_bases(["s3://bucket/media/"]).await.unwrap();
}
#[tokio::test]
async fn test_add_bases_rejects_empty_response() {
let table = Table::new_with_handler("my_table", |_request| {
http::Response::builder().status(200).body("").unwrap()
});
let err = table.add_bases(["s3://bucket/media/"]).await.unwrap_err();
assert!(
err.to_string()
.contains("invalid response while registering table bases"),
"{err}"
);
}
#[rstest]
#[case(semver::Version::new(0, 1, 0))]
#[case(semver::Version::new(0, 5, 0))]
@@ -6606,346 +6462,37 @@ mod tests {
assert_eq!(result.version, if old_server { 0 } else { 43 });
}
/// A declaration is sent as `{name, computed}` entries for the server to
/// plan; the client never types the expression itself.
/// Computed columns are local-only. Both halves say so here rather than
/// reaching the wire and failing somewhere less legible.
#[tokio::test]
async fn test_add_computed_columns_sends_the_expression() {
let table = Table::new_with_handler("my_table", |request| {
assert_eq!(request.method(), "POST");
assert_eq!(request.url().path(), "/v1/table/my_table/add_columns/");
let body = request.body().unwrap().as_bytes().unwrap();
let value: serde_json::Value = serde_json::from_slice(body).unwrap();
assert_eq!(
value["new_columns"],
serde_json::json!([{"name": "doubled", "computed": "x * 2"}])
);
http::Response::builder()
.status(200)
.body(r#"{"version": 7}"#)
.unwrap()
async fn test_computed_columns_are_refused() {
let table = Table::new_with_handler("my_table", |request| -> http::Response<String> {
panic!("unexpected request: {}", request.url().path())
});
let result = table
let declared = Arc::new(Schema::new(vec![Field::new(
"doubled",
DataType::Int32,
true,
)]));
let err = table
.add_columns()
.computed("doubled", "x * 2")
.transform(NewColumnTransform::AllNulls(declared))
.execute()
.await
.unwrap();
assert_eq!(result.version, 7);
}
/// A remote refresh is a server job: the async form returns its handle,
/// and the blocking form refuses rather than invent a fill count.
#[tokio::test]
async fn test_refresh_column_async_submits_a_backfill_job() {
let table = Table::new_with_handler("my_table", |request| {
assert_eq!(request.method(), "POST");
assert_eq!(request.url().path(), "/v1/table/my_table/backfill_column");
let body = request.body().unwrap().as_bytes().unwrap();
let value: serde_json::Value = serde_json::from_slice(body).unwrap();
assert_eq!(value["column"], "doubled");
http::Response::builder()
.status(202)
.body(r#"{"job_id": "j-42"}"#)
.unwrap()
});
let job = table.refresh_column_async("doubled").await.unwrap();
assert_eq!(job.id(), Some("j-42"));
.unwrap_err();
assert!(
matches!(&err, Error::NotSupported { message } if message.contains("local tables")),
"{err:?}"
);
let err = table.refresh_column("doubled").await.unwrap_err();
assert!(
matches!(&err, Error::NotSupported { message }
if message.contains("refresh_column_async")),
matches!(&err, Error::NotSupported { message } if message.contains("local tables")),
"{err:?}"
);
}
/// The gate's reproducer: after a successful wait, a same-handle read
/// must carry a freshness baseline so a stale server cache cannot serve
/// the pre-backfill snapshot.
#[tokio::test]
async fn test_backfill_wait_establishes_read_freshness() {
let saw_min_timestamp = Arc::new(std::sync::atomic::AtomicBool::new(false));
let saw = saw_min_timestamp.clone();
let table =
Table::new_with_handler("my_table", move |request| match request.url().path() {
"/v1/table/my_table/backfill_column" => http::Response::builder()
.status(202)
.body(r#"{"job_id": "j-7"}"#.to_string())
.unwrap(),
"/v1/jobs/describe" => http::Response::builder()
.status(200)
.body(r#"{"job_id": "j-7", "job_state": "DONE"}"#.to_string())
.unwrap(),
"/v1/table/my_table/count_rows/" => {
saw.store(
request.headers().contains_key("x-lancedb-min-timestamp"),
std::sync::atomic::Ordering::SeqCst,
);
http::Response::builder()
.status(200)
.body("1".to_string())
.unwrap()
}
path => panic!("unexpected request: {path}"),
});
let job = table.refresh_column_async("doubled").await.unwrap();
job.wait().await.unwrap();
table.count_rows(None).await.unwrap();
assert!(
saw_min_timestamp.load(std::sync::atomic::Ordering::SeqCst),
"read after wait carried no freshness baseline"
);
}
/// A checkout after submission wins over the completion fence: the
/// pinned view must not regain a timestamp floor from the job.
#[tokio::test]
async fn test_checkout_after_submit_beats_the_completion_fence() {
let saw_min_timestamp = Arc::new(std::sync::atomic::AtomicBool::new(false));
let saw = saw_min_timestamp.clone();
let table =
Table::new_with_handler("my_table", move |request| match request.url().path() {
"/v1/table/my_table/backfill_column" => http::Response::builder()
.status(202)
.body(r#"{"job_id": "j-8"}"#.to_string())
.unwrap(),
"/v1/jobs/describe" => http::Response::builder()
.status(200)
.body(r#"{"job_id": "j-8", "job_state": "DONE"}"#.to_string())
.unwrap(),
"/v1/table/my_table/describe/" => {
let schema = Schema::new(vec![Field::new("x", DataType::Int32, true)]);
http::Response::builder()
.status(200)
.body(describe_response(&schema))
.unwrap()
}
"/v1/table/my_table/count_rows/" => {
saw.store(
request.headers().contains_key("x-lancedb-min-timestamp"),
std::sync::atomic::Ordering::SeqCst,
);
http::Response::builder()
.status(200)
.body("1".to_string())
.unwrap()
}
path => panic!("unexpected request: {path}"),
});
let job = table.refresh_column_async("doubled").await.unwrap();
table.checkout(3).await.unwrap();
job.wait().await.unwrap();
table.count_rows(None).await.unwrap();
assert!(
!saw_min_timestamp.load(std::sync::atomic::Ordering::SeqCst),
"completion fence overrode an explicit checkout"
);
}
/// Tag checkout resets freshness state wholesale; the fence must not
/// survive it.
#[tokio::test]
async fn test_tag_checkout_after_submit_beats_the_completion_fence() {
let saw_min_timestamp = Arc::new(std::sync::atomic::AtomicBool::new(false));
let saw = saw_min_timestamp.clone();
let table =
Table::new_with_handler("my_table", move |request| match request.url().path() {
"/v1/table/my_table/backfill_column" => http::Response::builder()
.status(202)
.body(r#"{"job_id": "j-9"}"#.to_string())
.unwrap(),
"/v1/jobs/describe" => http::Response::builder()
.status(200)
.body(r#"{"job_id": "j-9", "job_state": "DONE"}"#.to_string())
.unwrap(),
"/v1/table/my_table/tags/version/" => http::Response::builder()
.status(200)
.body(r#"{"version": 5}"#.to_string())
.unwrap(),
"/v1/table/my_table/describe/" => {
let schema = Schema::new(vec![Field::new("x", DataType::Int32, true)]);
http::Response::builder()
.status(200)
.body(describe_response(&schema))
.unwrap()
}
"/v1/table/my_table/count_rows/" => {
saw.store(
request.headers().contains_key("x-lancedb-min-timestamp"),
std::sync::atomic::Ordering::SeqCst,
);
http::Response::builder()
.status(200)
.body("1".to_string())
.unwrap()
}
path => panic!("unexpected request: {path}"),
});
let job = table.refresh_column_async("doubled").await.unwrap();
table.checkout_tag("v1").await.unwrap();
job.wait().await.unwrap();
table.count_rows(None).await.unwrap();
assert!(
!saw_min_timestamp.load(std::sync::atomic::Ordering::SeqCst),
"completion fence overrode a tag checkout"
);
}
/// A checkout landing while the submission request is in flight advances
/// the epoch past the token captured at submit.
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn test_checkout_during_submission_beats_the_completion_fence() {
let saw_min_timestamp = Arc::new(std::sync::atomic::AtomicBool::new(false));
let saw = saw_min_timestamp.clone();
let (release_tx, release_rx) = std::sync::mpsc::channel::<()>();
let release_rx = Arc::new(std::sync::Mutex::new(release_rx));
let (arrived_tx, arrived_rx) = std::sync::mpsc::channel::<()>();
let arrived_tx = Arc::new(std::sync::Mutex::new(arrived_tx));
let table = Table::new_with_handler("my_table", move |request| {
match request.url().path() {
"/v1/table/my_table/backfill_column" => {
// Signal arrival, then hold the response until the
// test's checkout completes.
arrived_tx.lock().unwrap().send(()).unwrap();
release_rx
.lock()
.unwrap()
.recv_timeout(std::time::Duration::from_secs(10))
.unwrap();
http::Response::builder()
.status(202)
.body(r#"{"job_id": "j-10"}"#.to_string())
.unwrap()
}
"/v1/jobs/describe" => http::Response::builder()
.status(200)
.body(r#"{"job_id": "j-10", "job_state": "DONE"}"#.to_string())
.unwrap(),
"/v1/table/my_table/describe/" => {
let schema = Schema::new(vec![Field::new("x", DataType::Int32, true)]);
http::Response::builder()
.status(200)
.body(describe_response(&schema))
.unwrap()
}
"/v1/table/my_table/count_rows/" => {
saw.store(
request.headers().contains_key("x-lancedb-min-timestamp"),
std::sync::atomic::Ordering::SeqCst,
);
http::Response::builder()
.status(200)
.body("1".to_string())
.unwrap()
}
path => panic!("unexpected request: {path}"),
}
});
let submit = tokio::spawn({
let table = table.clone();
async move { table.refresh_column_async("doubled").await }
});
tokio::task::spawn_blocking(move || {
arrived_rx
.recv_timeout(std::time::Duration::from_secs(10))
.unwrap()
})
.await
.unwrap();
table.checkout(7).await.unwrap();
release_tx.send(()).unwrap();
let job = submit.await.unwrap().unwrap();
job.wait().await.unwrap();
table.count_rows(None).await.unwrap();
assert!(
!saw_min_timestamp.load(std::sync::atomic::Ordering::SeqCst),
"completion fence overrode a checkout that landed mid-submission"
);
}
/// checkout_latest keeps the handle on latest, so a completed backfill
/// must still establish its post-fill baseline -- strictly later than the
/// checkout's own, or a pre-fill cache could still serve.
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn test_checkout_latest_during_submission_keeps_the_fence() {
let seen_min_timestamp = Arc::new(std::sync::Mutex::new(None::<String>));
let saw = seen_min_timestamp.clone();
let (release_tx, release_rx) = std::sync::mpsc::channel::<()>();
let release_rx = Arc::new(std::sync::Mutex::new(release_rx));
let (arrived_tx, arrived_rx) = std::sync::mpsc::channel::<()>();
let arrived_tx = Arc::new(std::sync::Mutex::new(arrived_tx));
let table =
Table::new_with_handler("my_table", move |request| match request.url().path() {
"/v1/table/my_table/backfill_column" => {
arrived_tx.lock().unwrap().send(()).unwrap();
release_rx
.lock()
.unwrap()
.recv_timeout(std::time::Duration::from_secs(10))
.unwrap();
http::Response::builder()
.status(202)
.body(r#"{"job_id": "j-11"}"#.to_string())
.unwrap()
}
"/v1/jobs/describe" => http::Response::builder()
.status(200)
.body(r#"{"job_id": "j-11", "job_state": "DONE"}"#.to_string())
.unwrap(),
"/v1/table/my_table/count_rows/" => {
*saw.lock().unwrap() = request
.headers()
.get("x-lancedb-min-timestamp")
.map(|v| v.to_str().unwrap().to_string());
http::Response::builder()
.status(200)
.body("1".to_string())
.unwrap()
}
path => panic!("unexpected request: {path}"),
});
let submit = tokio::spawn({
let table = table.clone();
async move { table.refresh_column_async("doubled").await }
});
tokio::task::spawn_blocking(move || {
arrived_rx
.recv_timeout(std::time::Duration::from_secs(10))
.unwrap()
})
.await
.unwrap();
table.checkout_latest().await.unwrap();
let after_checkout = SystemTime::now();
// Real separation between the checkout baseline and completion.
tokio::time::sleep(std::time::Duration::from_millis(50)).await;
release_tx.send(()).unwrap();
let job = submit.await.unwrap().unwrap();
job.wait().await.unwrap();
table.count_rows(None).await.unwrap();
let header = seen_min_timestamp
.lock()
.unwrap()
.clone()
.expect("no baseline");
let sent: SystemTime = chrono::DateTime::parse_from_rfc3339(&header)
.unwrap()
.into();
assert!(
sent > after_checkout,
"baseline {header} did not advance past the checkout"
);
}
#[tokio::test]
async fn test_prewarm_index() {
let table = Table::new_with_handler("my_table", |request| {
-42
View File
@@ -1,42 +0,0 @@
// SPDX-License-Identifier: Apache-2.0
// SPDX-FileCopyrightText: Copyright The LanceDB Authors
//! Cloud HTTP for registering extra table storage bases.
use serde::Deserialize;
use crate::Error;
use crate::error::Result;
use crate::remote::client::{HttpSend, RequestResultExt};
use super::RemoteTable;
#[derive(Debug, Deserialize)]
struct AddBasesResponse {
version: u64,
}
impl<S: HttpSend> RemoteTable<S> {
pub(super) async fn add_bases_impl(&self, bases: &[crate::table::TableBase]) -> Result<()> {
self.check_mutable().await?;
let mut body = serde_json::json!({ "bases": bases });
self.apply_branch_body(&mut body);
let request = self
.client
.post(&format!("/v1/table/{}/bases/", self.identifier))
.json(&body);
let (request_id, response) = self.send(request, true).await?;
let response = self.check_table_response(&request_id, response).await?;
let body = response.text().await.err_to_http(request_id.clone())?;
let parsed: AddBasesResponse = serde_json::from_str(&body).map_err(|e| Error::Http {
source: format!(
"The server returned an invalid response while registering table bases: {e}"
)
.into(),
request_id,
status_code: None,
})?;
self.track_write_version(parsed.version);
Ok(())
}
}
+106 -387
View File
@@ -34,7 +34,7 @@ use lance_index::scalar::inverted::query::collect_query_tokens;
use lance_namespace::LanceNamespace;
use lance_namespace::error::NamespaceError;
use lance_namespace::models::DescribeTableRequest;
use lance_table::format::{BasePath, Manifest};
use lance_table::format::Manifest;
use lance_table::io::commit::CommitHandler;
use lance_table::io::commit::ManifestNamingScheme;
use lance_table::io::commit::external_manifest::ExternalManifestCommitHandler;
@@ -50,6 +50,7 @@ use crate::DistanceType;
use crate::blob::BlobRangeRequest;
use crate::data::scannable::{PeekedScannable, Scannable, estimate_write_partitions};
use crate::database::Database;
use crate::database::listing::LANCE_FILE_EXTENSION;
use crate::database::read_freshness::TableFreshness;
use crate::embeddings::{EmbeddingDefinition, EmbeddingRegistry, MemoryRegistry};
use crate::error::{Error, Result};
@@ -157,6 +158,55 @@ pub(crate) fn map_namespace_lance_error(err: lance::Error, table_name: &str) ->
}
}
/// Map a `lance::Error::DatasetNotFound` for the table at `uri` into a `lancedb::Error`.
///
/// Lance reports "there is nothing at this location" and "there is a table directory
/// here but nothing loadable inside it" with the same error. Only the first is a
/// `TableNotFound`: a `<name>.lance` directory left behind by an interrupted drop and
/// re-create is still reported by `Connection::table_names`, so callers need to be able
/// to tell "never existed" from "exists but is broken".
///
/// See <https://github.com/lancedb/lancedb/issues/3127>.
async fn map_dataset_not_found(
uri: &str,
name: &str,
params: ReadParams,
err: lance::Error,
) -> Error {
let name = name.to_string();
let source = Box::new(err);
if table_dir_exists(uri, params).await.unwrap_or(false) {
Error::TableCorrupted { name, source }
} else {
Error::TableNotFound { name, source }
}
}
/// Whether a table directory is present at `uri`, even though no dataset could be
/// loaded from it.
///
/// This looks for a `<name>.lance` entry in the parent directory, which is exactly what
/// `ListingDatabase::table_names` lists, so the two APIs agree on whether a table is
/// present. Probing `uri` itself would not work: object stores have no empty
/// directories to probe, and on a local filesystem the interesting case is precisely an
/// empty directory.
async fn table_dir_exists(uri: &str, params: ReadParams) -> Result<bool> {
let (object_store, path, _) = DatasetBuilder::from_uri(uri)
.with_read_params(params)
.build_object_store()
.await?;
// Only `*.lance` entries are ever reported as tables, so nothing else can produce
// the list-then-open mismatch this guards against.
if path.extension() != Some(LANCE_FILE_EXTENSION) {
return Ok(false);
}
let (Some(parent), Some(dir_name)) = (path.parent(), path.filename()) else {
return Ok(false);
};
let entries = object_store.read_dir(parent).await?;
Ok(entries.iter().any(|entry| entry.as_str() == dir_name))
}
/// Defines the type of column
#[derive(Debug, Clone, Serialize, Deserialize)]
pub enum ColumnKind {
@@ -643,15 +693,6 @@ pub trait BaseTable: std::fmt::Display + std::fmt::Debug + Send + Sync {
message: "set_lsm_write_spec is not supported on this table type".into(),
})
}
/// Switch this table to required index catch-up, one way.
///
/// The default implementation returns `NotSupported`. Implementations
/// that support the MemWAL LSM write path must override this.
async fn require_mem_wal_index_catchup(&self) -> Result<()> {
Err(Error::NotSupported {
message: "require_mem_wal_index_catchup is not supported on this table type".into(),
})
}
/// Remove the [`LsmWriteSpec`] from this table.
///
/// This is a no-op if no spec is currently set.
@@ -711,12 +752,6 @@ pub trait BaseTable: std::fmt::Display + std::fmt::Debug + Send + Sync {
message: "blob_columns is not supported on this table type".into(),
})
}
/// Register additional storage bases for this table.
async fn add_bases(&self, _bases: &[TableBase]) -> Result<()> {
Err(Error::NotSupported {
message: "Registering table bases is not supported for this table type.".into(),
})
}
/// Materialize blob bytes for the given row ids. See [`Table::fetch_blobs`].
async fn fetch_blobs(&self, _column: &str, _row_ids: &[u64]) -> Result<LargeBinaryArray> {
Err(Error::NotSupported {
@@ -753,19 +788,6 @@ pub trait BaseTable: std::fmt::Display + std::fmt::Debug + Send + Sync {
transforms: NewColumnTransform,
read_columns: Option<Vec<String>>,
) -> Result<AddColumnsResult>;
/// Declare computed columns, each defined by a SQL expression.
///
/// Where the declaration is planned depends on the backend: a local table
/// validates and types the expression itself, a remote one sends the text
/// for the server to plan.
async fn add_computed_columns(
&self,
_columns: &[(String, String)],
) -> Result<AddColumnsResult> {
Err(Error::NotSupported {
message: "computed columns are not supported on this table type".into(),
})
}
/// Fill a computed column's unfilled rows.
///
/// The default returns `NotSupported`; Lance-backed tables override it.
@@ -774,13 +796,6 @@ pub trait BaseTable: std::fmt::Display + std::fmt::Debug + Send + Sync {
message: "computed columns are supported only on local tables".into(),
})
}
/// Fill a computed column's unfilled rows, returning a [`Job`] tracking
/// the operation.
async fn refresh_column_async(&self, _column: &str) -> Result<Job> {
Err(Error::NotSupported {
message: "computed columns are supported only on local tables".into(),
})
}
/// Alter columns in the table.
async fn alter_columns(&self, alterations: &[ColumnAlteration]) -> Result<AlterColumnsResult>;
/// Drop columns from the table.
@@ -895,54 +910,6 @@ pub trait BaseTable: std::fmt::Display + std::fmt::Debug + Send + Sync {
}
}
/// An extra storage prefix registered on a table.
///
/// `path` is an object-store URI. `name` is an optional alias. `is_dataset_root`
/// is true when `path` points to a Lance dataset root. When false, `path`
/// points directly to the directory containing the referenced files.
#[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize)]
#[serde(rename_all = "camelCase")]
pub struct TableBase {
/// Object store URI such as `s3://bucket/media/`.
pub path: String,
/// Optional alias.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub name: Option<String>,
/// True when `path` is a Lance dataset root. When false, `path` is the
/// directory containing the referenced files.
#[serde(default)]
pub is_dataset_root: bool,
}
impl TableBase {
/// A non-root base with no alias.
pub fn new(path: impl Into<String>) -> Self {
Self {
path: path.into(),
name: None,
is_dataset_root: false,
}
}
}
impl From<&str> for TableBase {
fn from(path: &str) -> Self {
Self::new(path)
}
}
impl From<&String> for TableBase {
fn from(path: &String) -> Self {
Self::new(path.as_str())
}
}
impl From<String> for TableBase {
fn from(path: String) -> Self {
Self::new(path)
}
}
/// A Table is a collection of strong typed Rows.
///
/// The type of the each row is defined in Apache Arrow [Schema].
@@ -1180,25 +1147,6 @@ impl Table {
self.inner.blob_columns().await
}
/// Register additional storage bases for this table.
///
/// A URI string is a non-root base with no alias.
///
/// ```
/// # use lancedb::Table;
/// # async fn register(table: &Table) -> Result<(), Box<dyn std::error::Error>> {
/// table.add_bases(["s3://bucket/media/"]).await?;
/// # Ok(())
/// # }
/// ```
pub async fn add_bases(
&self,
bases: impl IntoIterator<Item = impl Into<TableBase>>,
) -> Result<()> {
let bases: Vec<TableBase> = bases.into_iter().map(Into::into).collect();
self.inner.add_bases(&bases).await
}
/// Materialize blob bytes for the given row ids.
///
/// Output matches `row_ids` in length and order. Null blobs are null;
@@ -1749,8 +1697,7 @@ impl Table {
/// filled are left as they are, so the call is idempotent and does not
/// observe a mutated input.
///
/// Local tables only: a remote refresh runs as a server job, through
/// [`Table::refresh_column_async`].
/// Local tables only.
///
/// ```
/// # use lancedb::Table;
@@ -1764,29 +1711,6 @@ impl Table {
self.inner.refresh_column(column.as_ref()).await
}
/// Like [`Table::refresh_column`], but returns a [`Job`] tracking the
/// operation instead of blocking until it completes.
///
/// The job may already be complete when returned, and callers must not
/// assume the column is filled until [`Job::wait`] returns. Invalid input
/// -- an unknown column, or one that is not computed -- is reported by
/// this call rather than by the job. On local tables the job runs as an
/// in-process task; on LanceDB Cloud and Enterprise it is the server's
/// backfill job.
///
/// ```
/// # use lancedb::Table;
/// # async fn refresh_in_background(table: &Table) -> Result<(), Box<dyn std::error::Error>> {
/// let job = table.refresh_column_async("doubled").await?;
/// println!("refresh running: {:?}", job.status().await?);
/// job.wait().await?;
/// # Ok(())
/// # }
/// ```
pub async fn refresh_column_async(&self, column: impl AsRef<str>) -> Result<Job> {
self.inner.refresh_column_async(column.as_ref()).await
}
/// Change a column's name or nullability.
pub async fn alter_columns(
&self,
@@ -1856,20 +1780,6 @@ impl Table {
self.inner.set_lsm_write_spec(spec).await
}
/// Switch this table to required index catch-up, one way.
///
/// Separate from [`Self::set_lsm_write_spec`] on purpose: a table carrying
/// the bit retains its SSTables until an index records that it holds the
/// compacted rows, so turn it on only once something can repair coverage.
/// A writer that already holds the dataset can call the equivalent on
/// `DatasetMemWalExt` instead; this is the table-level entry point.
///
/// Errors if no spec is set, or if the table already records SSTable
/// compaction progress from before this protocol.
pub async fn require_mem_wal_index_catchup(&self) -> Result<()> {
self.inner.require_mem_wal_index_catchup().await
}
/// Remove the [`LsmWriteSpec`] from this table, reverting to the standard
/// `merge_insert` write path.
///
@@ -2547,6 +2457,8 @@ impl NativeTable {
None => false,
};
// Kept so that a `DatasetNotFound` can be re-checked against storage below.
let recovery_params = params.clone();
let mut builder = DatasetBuilder::from_uri(uri).with_read_params(params);
// Set up commit handler when managed_versioning is enabled
@@ -2565,12 +2477,7 @@ impl NativeTable {
let dataset = match builder.load().await {
Ok(dataset) => dataset,
Err(e @ lance::Error::DatasetNotFound { .. }) => {
// The manifest load is the existence check. A physical prefix may be
// from a concurrent or abandoned create, so it cannot refine this error.
return Err(Error::TableNotFound {
name: name.to_string(),
source: Box::new(e),
});
return Err(map_dataset_not_found(uri, name, recovery_params, e).await);
}
Err(e) => return Err(e.into()),
};
@@ -2782,7 +2689,6 @@ impl NativeTable {
namespace_client: Option<Arc<dyn LanceNamespace>>,
pushdown_operations: HashSet<NamespaceClientPushdownOperation>,
) -> Result<Self> {
computed_columns::ensure_no_foreign_declarations(batches.arrow_schema().fields())?;
// Default params uses format v1.
let params = params.unwrap_or(WriteParams {
..Default::default()
@@ -3231,13 +3137,6 @@ impl BaseTable for NativeTable {
let ds = self.dataset.get().await?;
let table_schema = Schema::from(&ds.schema().clone());
computed_columns::ensure_not_written(
&table_schema,
add.data.schema().fields().iter().map(|f| f.name().as_str()),
)?;
if matches!(add.mode, AddDataMode::Overwrite) {
computed_columns::ensure_no_foreign_declarations(add.data.schema().fields())?;
}
let num_partitions = if let Some(parallelism) = add.write_parallelism {
parallelism
@@ -3398,11 +3297,6 @@ impl BaseTable for NativeTable {
params: MergeInsertBuilder,
new_data: Box<dyn RecordBatchReader + Send>,
) -> Result<MergeResult> {
let source_schema = arrow_array::RecordBatchReader::schema(&new_data);
computed_columns::ensure_not_written(
&Schema::from(self.dataset.get().await?.schema()),
source_schema.fields().iter().map(|f| f.name().as_str()),
)?;
let result = merge::execute_merge_insert(self, params, new_data).await?;
self.bump_freshness();
Ok(result)
@@ -3416,10 +3310,6 @@ impl BaseTable for NativeTable {
merge::lsm::set_lsm_write_spec(self, spec).await
}
async fn require_mem_wal_index_catchup(&self) -> Result<()> {
merge::lsm::require_mem_wal_index_catchup(self).await
}
async fn unset_lsm_write_spec(&self) -> Result<()> {
merge::lsm::unset_lsm_write_spec(self).await
}
@@ -3437,25 +3327,6 @@ impl BaseTable for NativeTable {
Ok(crate::blob::blob_column_names(schema.as_ref()))
}
async fn add_bases(&self, bases: &[TableBase]) -> Result<()> {
self.dataset.ensure_mutable()?;
let dataset = self.dataset.get().await?;
let new_bases = bases
.iter()
.map(|base| {
BasePath::new(
0,
base.path.clone(),
base.name.clone(),
base.is_dataset_root,
)
})
.collect();
let dataset = dataset.add_bases(new_bases, None).await?;
self.dataset.update(dataset);
Ok(())
}
async fn fetch_blobs(&self, column: &str, row_ids: &[u64]) -> Result<LargeBinaryArray> {
let dataset = self.dataset.get().await?;
crate::blob::take_blobs_aligned(&dataset, column, row_ids).await
@@ -3507,22 +3378,12 @@ impl BaseTable for NativeTable {
Ok(result)
}
async fn add_computed_columns(&self, columns: &[(String, String)]) -> Result<AddColumnsResult> {
let result = schema_evolution::execute_declare(self, columns).await?;
self.bump_freshness();
Ok(result)
}
async fn refresh_column(&self, column: &str) -> Result<RefreshColumnResult> {
let result = refresh::execute_refresh_column(self, column).await?;
self.bump_freshness();
Ok(result)
}
async fn refresh_column_async(&self, column: &str) -> Result<Job> {
refresh::execute_refresh_column_async(self, column).await
}
async fn alter_columns(&self, alterations: &[ColumnAlteration]) -> Result<AlterColumnsResult> {
let result = schema_evolution::execute_alter_columns(self, alterations).await?;
self.bump_freshness();
@@ -3890,7 +3751,7 @@ pub struct FragmentSummaryStats {
#[allow(deprecated)]
mod tests {
use std::sync::Arc;
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
use std::sync::atomic::{AtomicBool, Ordering};
use std::time::Duration;
use arrow_array::{
@@ -3972,50 +3833,73 @@ mod tests {
);
}
#[tokio::test]
async fn test_open_not_found_when_empty_directory_exists() {
let tmp_dir = tempdir().unwrap();
let dataset_path = tmp_dir.path().join("test.lance");
std::fs::create_dir(&dataset_path).unwrap();
/// Write a table and then break it, leaving the `<name>.lance` directory in place.
///
/// `remove_all` reproduces an interrupted drop + re-create (the directory is left
/// empty); otherwise only the manifests are removed, leaving the data files behind.
async fn write_then_corrupt_table(dir: &std::path::Path, remove_all: bool) -> String {
let dataset_path = dir.join("test.lance");
let uri = dataset_path.to_str().unwrap().to_string();
let err = NativeTable::open(dataset_path.to_str().unwrap())
.await
.unwrap_err();
let batch = make_test_batches();
let reader = RecordBatchIterator::new(vec![Ok(batch.clone())], batch.schema());
Dataset::write(reader, &uri, None).await.unwrap();
if remove_all {
for entry in std::fs::read_dir(&dataset_path).unwrap() {
let entry = entry.unwrap();
if entry.file_type().unwrap().is_dir() {
std::fs::remove_dir_all(entry.path()).unwrap();
} else {
std::fs::remove_file(entry.path()).unwrap();
}
}
assert_eq!(std::fs::read_dir(&dataset_path).unwrap().count(), 0);
} else {
let versions = dataset_path.join("_versions");
assert!(versions.is_dir(), "expected manifests under {versions:?}");
std::fs::remove_dir_all(&versions).unwrap();
assert!(std::fs::read_dir(&dataset_path).unwrap().count() > 0);
}
uri
}
#[tokio::test]
async fn test_open_corrupt_empty_dir() {
let tmp_dir = tempdir().unwrap();
let uri = write_then_corrupt_table(tmp_dir.path(), true).await;
let err = NativeTable::open(&uri).await.unwrap_err();
assert!(
matches!(&err, Error::TableNotFound { name, .. } if name == "test"),
matches!(&err, Error::TableCorrupted { name, .. } if name == "test"),
"got {err:?}"
);
}
#[tokio::test]
async fn test_open_not_found_when_only_uncommitted_storage_exists() {
async fn test_open_corrupt_missing_manifest() {
let tmp_dir = tempdir().unwrap();
let dataset_path = tmp_dir.path().join("test.lance");
let data_dir = dataset_path.join("data");
std::fs::create_dir_all(&data_dir).unwrap();
std::fs::write(data_dir.join("orphan.lance"), b"uncommitted").unwrap();
let uri = write_then_corrupt_table(tmp_dir.path(), false).await;
let err = NativeTable::open(dataset_path.to_str().unwrap())
.await
.unwrap_err();
let err = NativeTable::open(&uri).await.unwrap_err();
assert!(
matches!(&err, Error::TableNotFound { name, .. } if name == "test"),
matches!(&err, Error::TableCorrupted { name, .. } if name == "test"),
"got {err:?}"
);
}
/// Listing databases discover physical `*.lance` entries. That snapshot is not an
/// authoritative table-existence check: only a committed manifest makes a table
/// openable, and the entry could also be concurrently created or dropped.
/// A table listed by `table_names()` must not be reported as missing by
/// `open_table()`. See <https://github.com/lancedb/lancedb/issues/3127>.
#[tokio::test]
async fn test_table_names_may_include_uncommitted_storage() {
async fn test_open_table_corrupt_is_still_listed() {
let tmp_dir = tempdir().unwrap();
let db = connect(tmp_dir.path().to_str().unwrap())
.execute()
.await
.unwrap();
std::fs::create_dir(tmp_dir.path().join("test.lance")).unwrap();
write_then_corrupt_table(tmp_dir.path(), true).await;
assert_eq!(
db.table_names().execute().await.unwrap(),
@@ -4023,177 +3907,12 @@ mod tests {
);
let err = db.open_table("test").execute().await.unwrap_err();
assert!(
matches!(&err, Error::TableNotFound { name, .. } if name == "test"),
"physical storage without a committed manifest is not a table: {err:?}"
);
}
#[derive(Debug)]
struct ParentListGuardStore {
inner: Arc<dyn object_store::ObjectStore>,
parent: object_store::path::Path,
parent_list_calls: Arc<AtomicUsize>,
}
impl std::fmt::Display for ParentListGuardStore {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.write_str("ParentListGuardStore")
}
}
#[async_trait::async_trait]
#[deny(clippy::missing_trait_methods)]
impl object_store::ObjectStore for ParentListGuardStore {
async fn put_opts(
&self,
location: &object_store::path::Path,
payload: object_store::PutPayload,
opts: object_store::PutOptions,
) -> object_store::Result<object_store::PutResult> {
self.inner.put_opts(location, payload, opts).await
}
async fn put_multipart_opts(
&self,
location: &object_store::path::Path,
opts: object_store::PutMultipartOptions,
) -> object_store::Result<Box<dyn object_store::MultipartUpload>> {
self.inner.put_multipart_opts(location, opts).await
}
async fn get_opts(
&self,
location: &object_store::path::Path,
options: object_store::GetOptions,
) -> object_store::Result<object_store::GetResult> {
self.inner.get_opts(location, options).await
}
async fn get_ranges(
&self,
location: &object_store::path::Path,
ranges: &[std::ops::Range<u64>],
) -> object_store::Result<Vec<bytes::Bytes>> {
self.inner.get_ranges(location, ranges).await
}
fn delete_stream(
&self,
locations: futures::stream::BoxStream<
'static,
object_store::Result<object_store::path::Path>,
>,
) -> futures::stream::BoxStream<'static, object_store::Result<object_store::path::Path>>
{
self.inner.delete_stream(locations)
}
fn list(
&self,
prefix: Option<&object_store::path::Path>,
) -> futures::stream::BoxStream<'static, object_store::Result<object_store::ObjectMeta>>
{
if prefix == Some(&self.parent) {
self.parent_list_calls.fetch_add(1, Ordering::Relaxed);
}
self.inner.list(prefix)
}
fn list_with_offset(
&self,
prefix: Option<&object_store::path::Path>,
offset: &object_store::path::Path,
) -> futures::stream::BoxStream<'static, object_store::Result<object_store::ObjectMeta>>
{
if prefix == Some(&self.parent) {
self.parent_list_calls.fetch_add(1, Ordering::Relaxed);
}
self.inner.list_with_offset(prefix, offset)
}
async fn list_with_delimiter(
&self,
prefix: Option<&object_store::path::Path>,
) -> object_store::Result<object_store::ListResult> {
if prefix == Some(&self.parent) {
self.parent_list_calls.fetch_add(1, Ordering::Relaxed);
}
self.inner.list_with_delimiter(prefix).await
}
async fn copy_opts(
&self,
from: &object_store::path::Path,
to: &object_store::path::Path,
options: object_store::CopyOptions,
) -> object_store::Result<()> {
self.inner.copy_opts(from, to, options).await
}
async fn rename_opts(
&self,
from: &object_store::path::Path,
to: &object_store::path::Path,
options: object_store::RenameOptions,
) -> object_store::Result<()> {
self.inner.rename_opts(from, to, options).await
}
}
#[derive(Debug)]
struct ParentListGuardWrapper {
parent_list_calls: Arc<AtomicUsize>,
}
impl WrappingObjectStore for ParentListGuardWrapper {
fn wrap(
&self,
_store_prefix: &str,
inner: Arc<dyn object_store::ObjectStore>,
) -> Arc<dyn object_store::ObjectStore> {
Arc::new(ParentListGuardStore {
inner,
parent: object_store::path::Path::from("database"),
parent_list_calls: self.parent_list_calls.clone(),
})
}
}
#[tokio::test]
async fn test_open_missing_never_lists_database_parent() {
let parent_list_calls = Arc::new(AtomicUsize::new(0));
let params = ReadParams {
store_options: Some(ObjectStoreParams {
object_store_wrapper: Some(Arc::new(ParentListGuardWrapper {
parent_list_calls: parent_list_calls.clone(),
})),
..Default::default()
}),
..Default::default()
};
let err = NativeTable::open_with_params(
"memory:///database/missing.lance",
"missing",
Vec::new(),
None,
Some(params),
None,
None,
HashSet::new(),
None,
)
.await
.unwrap_err();
assert!(
matches!(&err, Error::TableNotFound { name, .. } if name == "missing"),
matches!(&err, Error::TableCorrupted { name, .. } if name == "test"),
"got {err:?}"
);
assert_eq!(
parent_list_calls.load(Ordering::Relaxed),
0,
"opening one missing table must not enumerate sibling tables"
assert!(
err.to_string().contains("exists but could not be loaded"),
"got {err}"
);
}
+5 -4
View File
@@ -8,6 +8,7 @@ use std::sync::Arc;
use lance::dataset::NewColumnTransform;
use super::BaseTable;
use super::computed_columns;
use super::schema_evolution::AddColumnsResult;
use crate::{Error, Result};
@@ -61,9 +62,8 @@ impl AddColumnsBuilder {
/// column and declaring it again. An input cannot be renamed, retyped or
/// dropped while a declaration reads it, since the expression names it.
///
/// On LanceDB Cloud and Enterprise the expression is planned by the
/// server, and the refresh runs as a server job -- see
/// [`Table::refresh_column_async`](super::Table::refresh_column_async).
/// Local tables only: LanceDB Cloud and Enterprise reject a declaration
/// with `NotSupported`.
///
/// ```
/// # use lancedb::Table;
@@ -130,7 +130,8 @@ impl AddColumnsBuilder {
.into(),
});
}
parent.add_computed_columns(&computed).await
let transform = computed_columns::declare(parent.schema().await?, &computed)?;
parent.add_columns(transform, None).await
}
}
}
+76 -719
View File
@@ -21,9 +21,7 @@
use std::collections::HashMap;
use std::sync::Arc;
use arrow_schema::{DataType, Field as ArrowField, Schema as ArrowSchema, SchemaRef};
use datafusion_common::tree_node::TreeNode;
use datafusion_physical_plan::PhysicalExpr;
use arrow_schema::{Field as ArrowField, Schema as ArrowSchema, SchemaRef};
use lance::dataset::NewColumnTransform;
use lance_datafusion::planner::Planner;
@@ -150,39 +148,20 @@ pub fn computed_columns(schema: &ArrowSchema) -> Vec<ComputedColumn> {
///
/// Paths are compared at their root: a declaration reading `metadata` is
/// invalidated by a change to `metadata.age` just as surely.
pub(crate) fn ensure_not_an_input(schema: &SchemaRef, paths: &[&str]) -> Result<()> {
pub(crate) fn ensure_not_an_input(schema: &ArrowSchema, paths: &[&str]) -> Result<()> {
let root = |path: &str| path.split('.').next().unwrap_or(path).to_string();
for declaration in computed_columns(schema) {
// The expression, not stored inputs, is the source of truth; an
// expression that no longer parses proves nothing, so refuse.
let inputs = match &declaration.kind {
ComputedColumnKind::Sql { expression } => Planner::new(schema.clone())
.parse_expr(expression)
.map(|parsed| Planner::column_names_in_expr(&parsed))
.map_err(|e| Error::InvalidInput {
message: format!(
"computed column '{}' has an unevaluable expression ({e}); drop it \
before changing the schema",
declaration.name
),
})?,
_ => declaration.inputs.clone(),
};
for path in paths {
// Exact target only: the binding travels with the whole column,
// not with a nested field the expression still shapes.
if declaration.name == *path {
// A declaration does not read itself, so it is free to be dropped
// or renamed along with its binding.
if declaration.name == root(path) {
continue;
}
if declaration.name == root(path) {
return Err(Error::InvalidInput {
message: format!(
"'{}' is part of computed column '{}'; drop the column and declare \
it again",
path, declaration.name
),
});
}
if inputs.iter().any(|input| root(input) == root(path)) {
if declaration
.inputs
.iter()
.any(|input| root(input) == root(path))
{
return Err(Error::InvalidInput {
message: format!(
"column '{}' is read by computed column '{}'; drop that column first",
@@ -195,217 +174,6 @@ pub(crate) fn ensure_not_an_input(schema: &SchemaRef, paths: &[&str]) -> Result<
Ok(())
}
/// Reject a write that supplies values for a computed column directly:
/// only refresh materializes one, and refresh never revisits a filled row.
pub(crate) fn ensure_not_written<'a>(
schema: &ArrowSchema,
written: impl IntoIterator<Item = &'a str>,
) -> Result<()> {
let declared: Vec<String> = computed_columns(schema)
.into_iter()
.map(|declaration| declaration.name)
.collect();
for name in written {
if declared.iter().any(|declared| declared == root(name)) {
return Err(Error::InvalidInput {
message: format!(
"column '{}' is computed; its values come from refresh and cannot be \
written directly",
root(name)
),
});
}
}
Ok(())
}
/// Reject a batch holding values for a computed column. Null slots are the
/// declared state, so planner-padded placeholders pass.
pub(crate) fn ensure_batch_writes_no_computed_values(
declared: &[String],
batch: &arrow_array::RecordBatch,
) -> Result<()> {
for name in declared {
if let Some(column) = batch.column_by_name(name)
&& column.null_count() != column.len()
{
return Err(Error::InvalidInput {
message: format!(
"column '{name}' is computed; its values come from refresh and cannot \
be written directly"
),
});
}
}
Ok(())
}
/// Reject fields carrying declaration metadata that did not come through
/// [`plan`]. One authority for creation, overwrite and raw transforms.
pub(crate) fn ensure_no_foreign_declarations<'a>(
fields: impl IntoIterator<Item = &'a Arc<ArrowField>>,
) -> Result<()> {
for field in fields {
if field.metadata().keys().any(|k| is_declaration_key(k)) {
return Err(Error::InvalidInput {
message: format!(
"field '{}' carries computed-column metadata; declare computed columns \
with add_columns().computed()",
field.name()
),
});
}
}
Ok(())
}
/// True for field-metadata keys that belong to a computed-column declaration.
///
/// A declaration is immutable through metadata edits: it is validated as a
/// whole at declare time, and rewriting any piece of it -- the flag, the
/// kind, the expression, the inputs -- would bypass that validation or move
/// a binding out from under a refresh. Drop the column and declare it again.
pub(crate) fn is_declaration_key(key: &str) -> bool {
key == COMPUTED_COLUMN_META_KEY || key.starts_with("computed_column.")
}
/// Reject retyping a computed column itself.
///
/// A cast keeps the stored expression while changing the type it must yield
/// -- and lance's cast rewrites the field without its metadata, so the
/// declaration silently stops being one. Dropping and redeclaring is the
/// coherent way to change a computed column's type.
pub(crate) fn ensure_not_retyped(schema: &ArrowSchema, paths: &[&str]) -> Result<()> {
for declaration in computed_columns(schema) {
for path in paths {
if declaration.name == root(path) {
return Err(Error::InvalidInput {
message: format!(
"column '{}' is computed; drop it and declare it again to change \
its type",
declaration.name
),
});
}
}
}
Ok(())
}
/// The top-level column a possibly nested input path reads.
pub(crate) fn root(path: &str) -> &str {
path.split('.').next().unwrap_or(path)
}
/// A declaration's expression bound to a schema, ready to evaluate.
pub(crate) struct BoundExpression {
/// The columns the expression names, as written; nested inputs keep
/// their dotted path.
pub inputs: Vec<String>,
/// The top-level columns evaluation reads, in [`Self::read_schema`]
/// order. A nested input appears through its root.
pub roots: Vec<String>,
/// The projected schema evaluation runs against.
pub read_schema: SchemaRef,
/// The compiled expression.
pub physical: Arc<dyn PhysicalExpr>,
/// The type the expression yields.
pub data_type: DataType,
}
/// Parse, resolve and compile `expression` against `schema`.
///
/// Inputs come from the expression as written, before optimization: the
/// simplifier can fold a referenced column out entirely (`true OR x > 0`),
/// and the guard protecting the stored SQL has to see every column the text
/// names, not just the ones the simplified form still reads.
pub(crate) fn bind(schema: SchemaRef, column: &str, expression: &str) -> Result<BoundExpression> {
let invalid = |message: String| Error::InvalidExpression {
column: column.to_string(),
message,
};
let planner = Planner::new(schema.clone());
let parsed = planner
.parse_expr(expression)
.map_err(|e| invalid(e.to_string()))?;
// A declaration is evaluated more than once -- staging and writing are
// separate passes, and a refresh years later replays the same text -- so
// a function that can answer differently each time has no coherent value
// to declare.
let mut volatile = None;
parsed
.apply(|expr| {
use datafusion_common::tree_node::TreeNodeRecursion;
if let datafusion_expr::Expr::ScalarFunction(function) = expr
&& function.func.signature().volatility != datafusion_expr::Volatility::Immutable
{
volatile = Some(function.func.name().to_string());
return Ok(TreeNodeRecursion::Stop);
}
Ok(TreeNodeRecursion::Continue)
})
.map_err(|e| invalid(e.to_string()))?;
if let Some(function) = volatile {
return Err(invalid(format!(
"'{function}' is not deterministic; a computed column's expression must \
yield the same value every time it is evaluated"
)));
}
let mut inputs = Planner::column_names_in_expr(&parsed);
inputs.sort();
inputs.dedup();
// A nested input is recorded by its path but read through its root
// column; Schema::index_of resolves top-level names only. Resolved here
// rather than left to the planner so an unknown column names itself in
// the error instead of surfacing as a plan failure.
let mut indices = Vec::with_capacity(inputs.len());
for input in &inputs {
let index = schema
.index_of(root(input))
.map_err(|_| invalid(format!("unknown column '{input}'")))?;
if !indices.contains(&index) {
indices.push(index);
}
}
indices.sort_unstable();
// Physical expressions address columns by position, so the planner that
// compiles the expression has to be built on the projected schema
// evaluation will actually read.
let read_schema = Arc::new(
schema
.project(&indices)
.map_err(|e| invalid(e.to_string()))?,
);
let roots = read_schema
.fields()
.iter()
.map(|field| field.name().clone())
.collect();
let optimized = planner
.optimize_expr(parsed)
.map_err(|e| invalid(e.to_string()))?;
let physical = Planner::new(read_schema.clone())
.create_physical_expr(&optimized)
.map_err(|e| invalid(e.to_string()))?;
let data_type = physical
.data_type(read_schema.as_ref())
.map_err(|e| invalid(e.to_string()))?;
Ok(BoundExpression {
inputs,
roots,
read_schema,
physical,
data_type,
})
}
/// Resolve `(name, expression)` pairs against `schema` into fields carrying
/// their bindings.
///
@@ -420,6 +188,7 @@ pub(crate) fn plan(schema: SchemaRef, columns: &[(String, String)]) -> Result<Ve
});
}
let planner = Planner::new(schema.clone());
let mut fields = Vec::with_capacity(columns.len());
let mut declared: Vec<&str> = Vec::with_capacity(columns.len());
@@ -428,13 +197,62 @@ pub(crate) fn plan(schema: SchemaRef, columns: &[(String, String)]) -> Result<Ve
return Err(Error::ColumnAlreadyExists { name: name.clone() });
}
let bound = bind(schema.clone(), name, expression)?;
let expr = planner
.parse_expr(expression)
.and_then(|expr| planner.optimize_expr(expr))
.map_err(|e| Error::InvalidExpression {
column: name.clone(),
message: e.to_string(),
})?;
let mut inputs = Planner::column_names_in_expr(&expr);
inputs.sort();
inputs.dedup();
// Resolved here rather than left to the planner so an unknown column
// names itself in the error instead of surfacing as a plan failure.
let mut indices = Vec::with_capacity(inputs.len());
for input in &inputs {
let index = schema
.index_of(input)
.map_err(|_| Error::InvalidExpression {
column: name.clone(),
message: format!("unknown column '{input}'"),
})?;
indices.push(index);
}
// Physical expressions address columns by position, so the planner
// that types the expression has to be built on the projected schema
// the refresh will actually read.
let read_schema =
Arc::new(
schema
.project(&indices)
.map_err(|e| Error::InvalidExpression {
column: name.clone(),
message: e.to_string(),
})?,
);
let physical = Planner::new(read_schema.clone())
.create_physical_expr(&expr)
.map_err(|e| Error::InvalidExpression {
column: name.clone(),
message: e.to_string(),
})?;
let data_type =
physical
.data_type(read_schema.as_ref())
.map_err(|e| Error::InvalidExpression {
column: name.clone(),
message: e.to_string(),
})?;
// Declared columns start entirely null, so nullability is a property
// of the declaration rather than of what the expression yields.
fields.push(
ArrowField::new(name, bound.data_type, true)
.with_metadata(computed_column_metadata(expression, &bound.inputs)),
ArrowField::new(name, data_type, true)
.with_metadata(computed_column_metadata(expression, &inputs)),
);
declared.push(name);
}
@@ -460,22 +278,25 @@ pub(crate) fn declare(
}
/// Commit a declaration of a kind this version does not produce, the way a
/// newer lancedb would leave one behind. Bypasses admission, which exists to
/// stop exactly this through the public API.
/// newer lancedb would leave one behind. Shared with the refresh tests, which
/// need the same column to check that refresh refuses it.
#[cfg(test)]
pub(super) async fn add_foreign_kind(table: &crate::Table, name: &str, kind: &str) {
use arrow_schema::DataType;
let field = ArrowField::new(name, DataType::Int32, true).with_metadata(HashMap::from([
(COMPUTED_COLUMN_META_KEY.to_string(), "true".to_string()),
(KIND_META_KEY.to_string(), kind.to_string()),
(INPUTS_META_KEY.to_string(), r#"["x"]"#.to_string()),
]));
super::schema_evolution::commit_add_columns(
table.as_native().unwrap(),
NewColumnTransform::AllNulls(Arc::new(ArrowSchema::new(vec![field]))),
None,
)
.await
.unwrap();
table
.add_columns()
.transform(NewColumnTransform::AllNulls(Arc::new(ArrowSchema::new(
vec![field],
))))
.execute()
.await
.unwrap();
}
#[cfg(test)]
@@ -870,470 +691,6 @@ mod tests {
.unwrap();
}
/// The gate's reproducer: a volatile function evaluates differently in
/// the counting and writing passes, so the declared value is incoherent.
/// Refused at declare time.
#[tokio::test]
async fn test_a_volatile_expression_is_refused() {
let table = table_with_ints("volatile_expr").await;
let err = add_computed(&table, &[("maybe".into(), "random() < 0.5".into())])
.await
.unwrap_err();
assert!(
matches!(&err, Error::InvalidExpression { message, .. }
if message.contains("random") && message.contains("deterministic")),
"{err:?}"
);
}
/// The gate's reproducer: the simplifier folds `true OR x > 0` to a
/// constant, but the stored SQL still names `x`, so the recorded inputs
/// must too -- otherwise dropping `x` is allowed and refresh breaks.
#[tokio::test]
async fn test_inputs_survive_expression_optimization() {
let table = table_with_ints("optimized_inputs").await;
add_computed(&table, &[("flag".into(), "true OR x > 0".into())])
.await
.unwrap();
assert_eq!(declared(&table).await[0].inputs, vec!["x".to_string()]);
let err = table.drop_columns(&["x"]).await.unwrap_err();
assert!(
matches!(&err, Error::InvalidInput { message } if message.contains("flag")),
"{err:?}"
);
}
/// The gate's reproducer: casting a computed column rewrites the field
/// without its metadata, silently destroying the declaration.
#[tokio::test]
async fn test_retyping_the_computed_column_is_refused() {
use arrow_schema::DataType as ArrowDataType;
let table = table_with_ints("retype_computed").await;
add_computed(&table, &[("doubled".into(), "x * 2".into())])
.await
.unwrap();
let err = table
.alter_columns(&[ColumnAlteration::new("doubled".into()).cast_to(ArrowDataType::Int64)])
.await
.unwrap_err();
assert!(
matches!(&err, Error::InvalidInput { message } if message.contains("computed")),
"{err:?}"
);
// The declaration survives the refused change.
table.refresh_column("doubled").await.unwrap();
}
/// A declaration cannot be edited, fabricated or erased through field
/// metadata: it is validated as a whole at declare time.
#[tokio::test]
async fn test_declaration_metadata_is_immutable() {
use crate::table::FieldMetadataUpdate;
let table = table_with_ints("metadata_tamper").await;
add_computed(&table, &[("doubled".into(), "x * 2".into())])
.await
.unwrap();
// Moving the binding.
let err = table
.update_field_metadata(&[
FieldMetadataUpdate::new("doubled").set(EXPRESSION_META_KEY, "x * 3")
])
.await
.unwrap_err();
assert!(matches!(err, Error::InvalidInput { .. }), "{err:?}");
// Fabricating a declaration on a plain column.
let err = table
.update_field_metadata(&[FieldMetadataUpdate::new("x")
.set(COMPUTED_COLUMN_META_KEY, "true")
.set(KIND_META_KEY, SQL_KIND)
.set(EXPRESSION_META_KEY, "x")])
.await
.unwrap_err();
assert!(matches!(err, Error::InvalidInput { .. }), "{err:?}");
// Erasing the declaration wholesale.
let err = table
.update_field_metadata(&[FieldMetadataUpdate::new("doubled")
.set("note", "hi")
.replace()])
.await
.unwrap_err();
assert!(matches!(err, Error::InvalidInput { .. }), "{err:?}");
// Ordinary metadata on a computed column still merges.
table
.update_field_metadata(&[FieldMetadataUpdate::new("doubled").set("note", "hi")])
.await
.unwrap();
table.refresh_column("doubled").await.unwrap();
}
/// The gate's reproducer: only refresh materializes a declared column;
/// a direct write would store an arbitrary durable value.
#[tokio::test]
async fn test_a_computed_column_cannot_be_written_directly() {
let table = table_with_ints("direct_write").await;
add_computed(&table, &[("doubled".into(), "x * 2".into())])
.await
.unwrap();
let batch = record_batch!(("x", Int32, [4]), ("doubled", Int32, [999])).unwrap();
let err = table.add(batch.clone()).execute().await.unwrap_err();
assert!(
matches!(&err, Error::InvalidInput { message } if message.contains("refresh")),
"{err:?}"
);
let err = table
.update()
.column("doubled", "999")
.execute()
.await
.unwrap_err();
assert!(matches!(err, Error::InvalidInput { .. }));
let mut merge = table.merge_insert(&["x"]);
merge
.when_matched_update_all(None)
.when_not_matched_insert_all();
let err = merge
.execute(Box::new(arrow_array::RecordBatchIterator::new(
vec![Ok(batch.clone())],
batch.schema(),
)))
.await
.unwrap_err();
assert!(matches!(err, Error::InvalidInput { .. }));
// The append that omits the column still works.
let plain = record_batch!(("x", Int32, [4])).unwrap();
table.add(plain).execute().await.unwrap();
}
/// The gate's reproducer: the reciprocal of the declare-under-spec check.
#[tokio::test]
async fn test_installing_an_lsm_spec_over_computed_columns_is_refused() {
use crate::table::LsmWriteSpec;
let tmp_dir = tempfile::tempdir().unwrap();
let conn = connect(tmp_dir.path().to_str().unwrap())
.execute()
.await
.unwrap();
let schema = Arc::new(arrow_schema::Schema::new(vec![arrow_schema::Field::new(
"x",
DataType::Int32,
false,
)]));
let batch = arrow_array::RecordBatch::try_new(
schema,
vec![Arc::new(arrow_array::Int32Array::from(vec![1, 2])) as _],
)
.unwrap();
let table = conn
.create_table("lsm_after", batch)
.execute()
.await
.unwrap();
table.set_unenforced_primary_key(["x"]).await.unwrap();
add_computed(&table, &[("doubled".into(), "x * 2".into())])
.await
.unwrap();
let err = table
.set_lsm_write_spec(LsmWriteSpec::unsharded())
.await
.unwrap_err();
assert!(
matches!(&err, Error::NotSupported { message } if message.contains("computed")),
"{err:?}"
);
assert!(table.get_lsm_write_spec().await.unwrap().is_none());
}
/// The gate's reproducer: declaration metadata is admitted only through
/// the validated declare path, never smuggled through a raw transform.
#[tokio::test]
async fn test_forged_declaration_metadata_is_rejected() {
let table = table_with_ints("forged_metadata").await;
let field =
ArrowField::new("doubled", DataType::Int32, true).with_metadata(HashMap::from([
(COMPUTED_COLUMN_META_KEY.to_string(), "true".to_string()),
(KIND_META_KEY.to_string(), SQL_KIND.to_string()),
(EXPRESSION_META_KEY.to_string(), "x * 2".to_string()),
(INPUTS_META_KEY.to_string(), "[]".to_string()),
]));
let err = table
.add_columns()
.transform(NewColumnTransform::AllNulls(Arc::new(ArrowSchema::new(
vec![field],
))))
.execute()
.await
.unwrap_err();
assert!(
matches!(&err, Error::InvalidInput { message } if message.contains("computed()")),
"{err:?}"
);
assert!(declared(&table).await.is_empty());
}
/// The gate's reproducer: SQL INSERT is a write path too.
#[tokio::test]
async fn test_sql_insert_cannot_write_a_computed_column() {
use datafusion::prelude::SessionContext;
let table = table_with_ints("sql_insert").await;
add_computed(&table, &[("doubled".into(), "x * 2".into())])
.await
.unwrap();
let ctx = SessionContext::new();
let provider =
crate::table::datafusion::BaseTableAdapter::try_new(table.base_table().clone())
.await
.unwrap();
ctx.register_table("t", Arc::new(provider)).unwrap();
let result = async {
ctx.sql("INSERT INTO t (x, doubled) VALUES (4, 999)")
.await?
.collect()
.await
}
.await;
let err = result.unwrap_err().to_string();
assert!(err.contains("refresh"), "{err}");
}
/// The gate's reproducer: an overwrite must not smuggle in a filled
/// declaration.
#[tokio::test]
async fn test_overwrite_cannot_inject_a_declaration() {
use crate::table::AddDataMode;
let table = table_with_ints("overwrite_inject").await;
let field =
ArrowField::new("doubled", DataType::Int32, true).with_metadata(HashMap::from([
(COMPUTED_COLUMN_META_KEY.to_string(), "true".to_string()),
(KIND_META_KEY.to_string(), SQL_KIND.to_string()),
(EXPRESSION_META_KEY.to_string(), "x * 2".to_string()),
]));
let schema = Arc::new(ArrowSchema::new(vec![
ArrowField::new("x", DataType::Int32, true),
field,
]));
let batch = arrow_array::RecordBatch::try_new(
schema,
vec![
Arc::new(arrow_array::Int32Array::from(vec![1])) as _,
Arc::new(arrow_array::Int32Array::from(vec![999])) as _,
],
)
.unwrap();
let err = table
.add(batch)
.mode(AddDataMode::Overwrite)
.execute()
.await
.unwrap_err();
assert!(
matches!(&err, Error::InvalidInput { message } if message.contains("declare")),
"{err:?}"
);
}
#[tokio::test]
async fn test_create_table_cannot_inject_a_declaration() {
let conn = connect("memory://").execute().await.unwrap();
let field =
ArrowField::new("doubled", DataType::Int32, true).with_metadata(HashMap::from([
(COMPUTED_COLUMN_META_KEY.to_string(), "true".to_string()),
(KIND_META_KEY.to_string(), SQL_KIND.to_string()),
(EXPRESSION_META_KEY.to_string(), "x * 2".to_string()),
]));
let schema = Arc::new(ArrowSchema::new(vec![
ArrowField::new("x", DataType::Int32, true),
field,
]));
let batch = arrow_array::RecordBatch::try_new(
schema.clone(),
vec![
Arc::new(arrow_array::Int32Array::from(vec![1])) as _,
Arc::new(arrow_array::Int32Array::from(vec![999])) as _,
],
)
.unwrap();
let err = conn
.create_table("forged_create", batch)
.execute()
.await
.unwrap_err();
assert!(
matches!(&err, Error::InvalidInput { message } if message.contains("computed()")),
"{err:?}"
);
}
#[tokio::test]
async fn test_sql_insert_omitting_computed_is_allowed() {
use datafusion::prelude::SessionContext;
let table = table_with_ints("sql_insert_omitted").await;
add_computed(&table, &[("doubled".into(), "x * 2".into())])
.await
.unwrap();
let ctx = SessionContext::new();
let provider =
crate::table::datafusion::BaseTableAdapter::try_new(table.base_table().clone())
.await
.unwrap();
ctx.register_table("t", Arc::new(provider)).unwrap();
ctx.sql("INSERT INTO t (x) VALUES (4)")
.await
.unwrap()
.collect()
.await
.unwrap();
table.checkout_latest().await.unwrap();
assert_eq!(table.count_rows(None).await.unwrap(), 4);
}
#[tokio::test]
async fn test_a_nested_computed_field_cannot_be_renamed() {
let table = table_with_ints("computed_struct_rename").await;
add_computed(&table, &[("payload".into(), "named_struct('a', x)".into())])
.await
.unwrap();
let err = table
.alter_columns(&[ColumnAlteration::new("payload.a".into()).rename("b".into())])
.await
.unwrap_err();
assert!(
matches!(&err, Error::InvalidInput { message } if message.contains("payload")),
"{err:?}"
);
}
/// Stale handles must not commit the computed/LSM state in either order.
#[tokio::test]
async fn test_stale_handles_cannot_mix_computed_and_lsm() {
use crate::table::LsmWriteSpec;
let tmp_dir = tempfile::tempdir().unwrap();
let uri = tmp_dir.path().to_str().unwrap();
let schema = Arc::new(ArrowSchema::new(vec![ArrowField::new(
"x",
DataType::Int32,
false,
)]));
let batch = arrow_array::RecordBatch::try_new(
schema.clone(),
vec![Arc::new(arrow_array::Int32Array::from(vec![1])) as _],
)
.unwrap();
let conn = connect(uri).execute().await.unwrap();
let table = conn.create_table("mix", batch).execute().await.unwrap();
table.set_unenforced_primary_key(["x"]).await.unwrap();
let stale = conn.open_table("mix").execute().await.unwrap();
// Declare on one handle; the stale handle must not install a spec.
add_computed(&table, &[("doubled".into(), "x * 2".into())])
.await
.unwrap();
let err = stale
.set_lsm_write_spec(LsmWriteSpec::unsharded())
.await
.unwrap_err();
assert!(matches!(err, Error::NotSupported { .. }), "install won");
// Reverse order on fresh tables.
let batch = arrow_array::RecordBatch::try_new(
schema,
vec![Arc::new(arrow_array::Int32Array::from(vec![1])) as _],
)
.unwrap();
let table = conn.create_table("mix2", batch).execute().await.unwrap();
table.set_unenforced_primary_key(["x"]).await.unwrap();
let stale = conn.open_table("mix2").execute().await.unwrap();
table
.set_lsm_write_spec(LsmWriteSpec::unsharded())
.await
.unwrap();
let err = add_computed(&stale, &[("doubled".into(), "x * 2".into())])
.await
.unwrap_err();
assert!(matches!(err, Error::NotSupported { .. }), "declare won");
}
/// The gate's reproducer: after catch-up activation, an LSM write, and
/// unset, retained SSTable rows survive without a live spec. The catch-up
/// flag is the durable marker; declaration refuses on it.
#[tokio::test]
async fn test_unset_with_retained_lsm_rows_cannot_admit_a_declaration() {
use crate::table::LsmWriteSpec;
use arrow_array::{Int64Array, RecordBatchIterator};
let tmp_dir = tempfile::tempdir().unwrap();
let conn = connect(tmp_dir.path().to_str().unwrap())
.execute()
.await
.unwrap();
let schema = Arc::new(ArrowSchema::new(vec![
ArrowField::new("id", DataType::Int64, false),
ArrowField::new("value", DataType::Int64, false),
]));
let batch = arrow_array::RecordBatch::try_new(
schema.clone(),
vec![
Arc::new(Int64Array::from(vec![1, 2])) as _,
Arc::new(Int64Array::from(vec![10, 20])) as _,
],
)
.unwrap();
let table = conn
.create_table("t", batch.clone())
.execute()
.await
.unwrap();
table.set_unenforced_primary_key(["id"]).await.unwrap();
table
.set_lsm_write_spec(LsmWriteSpec::unsharded())
.await
.unwrap();
table.require_mem_wal_index_catchup().await.unwrap();
let mut merge = table.merge_insert(&["id"]);
merge
.when_matched_update_all(None)
.when_not_matched_insert_all()
.use_lsm(true);
merge
.execute(Box::new(RecordBatchIterator::new(vec![Ok(batch)], schema)))
.await
.unwrap();
table.unset_lsm_write_spec().await.unwrap();
let err = add_computed(&table, &[("doubled".into(), "value * 2".into())])
.await
.unwrap_err();
assert!(
matches!(&err, Error::NotSupported { message } if message.contains("LSM")),
"{err:?}"
);
}
/// A declaration does not read itself, so it travels with its binding.
#[tokio::test]
async fn test_dropping_the_computed_column_is_allowed() {
+3 -14
View File
@@ -17,7 +17,7 @@ use datafusion_physical_plan::stream::RecordBatchStreamAdapter;
use datafusion_physical_plan::{
DisplayAs, DisplayFormatType, ExecutionPlan, ExecutionPlanProperties, PlanProperties,
};
use futures::StreamExt;
use futures::TryStreamExt;
use lance::Dataset;
use lance::dataset::transaction::{Operation, Transaction};
use lance::dataset::{CommitBuilder, InsertBuilder, WriteParams, WriteProgressFn};
@@ -194,23 +194,12 @@ impl ExecutionPlan for InsertExec {
let output_bytes = MetricBuilder::new(&self.metrics).output_bytes(partition);
let input_schema = input_stream.schema();
let declared: Vec<String> = crate::table::computed_columns::computed_columns(
&arrow_schema::Schema::from(self.dataset.schema()),
)
.into_iter()
.map(|declaration| declaration.name)
.collect();
let input_stream: SendableRecordBatchStream =
Box::pin(InstrumentedRecordBatchStreamAdapter::new(
input_schema,
input_stream.map(move |batch| {
let batch = batch?;
crate::table::computed_columns::ensure_batch_writes_no_computed_values(
&declared, &batch,
)
.map_err(|e| datafusion::error::DataFusionError::External(Box::new(e)))?;
input_stream.map_ok(move |batch| {
output_bytes.add(batch.get_array_memory_size());
Ok(batch)
batch
}),
partition,
&self.metrics,
-39
View File
@@ -94,16 +94,7 @@ pub(crate) async fn set_lsm_write_spec(table: &NativeTable, spec: LsmWriteSpec)
.await?
};
table.checkout_latest().await?;
let mut dataset = (*table.dataset.get().await?).clone();
let schema = arrow_schema::Schema::from(dataset.schema());
if !crate::table::computed_columns::computed_columns(&schema).is_empty() {
return Err(Error::NotSupported {
message: "an LSM write spec cannot be installed on a table with computed \
columns: rows in un-compacted tiers are invisible to refresh"
.into(),
});
}
let mut builder = dataset.initialize_mem_wal();
let writer_config_defaults = match spec {
LsmWriteSpec::Bucket {
@@ -192,36 +183,6 @@ fn index_name_list(indices: &[IndexConfig]) -> String {
format!("[{}]", names.join(", "))
}
// =============================================================================
// require_mem_wal_index_catchup
// =============================================================================
/// Switch this table to required index catch-up, one way.
///
/// Deliberately **not** part of installing the write spec. Until something can
/// actually repair coverage, a table carrying the bit reports every index as
/// not known to hold the compacted rows, so its SSTables are retained
/// indefinitely -- and the WAL pod trims on the legacy rule meanwhile, leaving
/// readers pointed at files that are gone. Turn this on only once remote
/// maintenance owns the merge and the repair for the table.
///
/// Lance refuses the activation if the table already records SSTable
/// compaction progress: those numbers predate this protocol and cannot be
/// validated, so such a table must be drained rather than activated.
#[allow(clippy::redundant_pub_crate)]
pub(crate) async fn require_mem_wal_index_catchup(table: &NativeTable) -> Result<()> {
table.dataset.ensure_mutable()?;
let mut dataset = (*table.dataset.get().await?).clone();
if dataset.mem_wal_index_details().await?.is_none() {
return Err(Error::InvalidInput {
message: "require_mem_wal_index_catchup: no LSM write spec is set on this table".into(),
});
}
dataset.require_mem_wal_index_catchup().await?;
table.dataset.update(dataset);
Ok(())
}
// =============================================================================
// unset_lsm_write_spec
// =============================================================================
+52 -202
View File
@@ -36,7 +36,6 @@ use lance::dataset::mem_wal::{
DatasetMemWalExt, LsmScanner, ShardManifestStore, ShardSnapshot, ShardWriterConfig,
};
use lance_index::mem_wal::{MemWalIndexDetails, ShardManifest};
use lance_table::feature_flags::FLAG_MEM_WAL_INDEX_CATCHUP;
use uuid::Uuid;
use super::NativeTable;
@@ -85,8 +84,9 @@ pub(super) async fn create_lsm_plan(
let pk_columns = pk_columns(&ds_ref)?;
// The base index an indexed arm relies on may lag compaction; resolve it so the
// snapshot retains SSTables the index has not yet caught up to.
let arm_indexes = arm_maintained_index_names(&ds_ref, &query, &details).await?;
let (snapshots, in_memory) = build_read_context(table, &ds_ref, &details, &arm_indexes).await?;
let arm_index = arm_maintained_index_name(&ds_ref, &query, &details).await?;
let (snapshots, in_memory) =
build_read_context(table, &ds_ref, &details, arm_index.as_deref()).await?;
let limit = query.base.limit;
let offset = query.base.offset;
@@ -232,44 +232,28 @@ fn pk_columns(dataset: &Dataset) -> Result<Vec<String>> {
Ok(pk)
}
/// Per-shard SSTable exclusion watermark: the generation at or below which
/// SSTables are safe to drop for this query.
///
/// A generation is droppable only once it is compacted into the base table AND
/// covered by the catch-up of every index the query relies on, so the watermark
/// is the minimum across `index_names`. Gating on fewer than all of them would
/// drop SSTables holding rows an uncounted index has not yet indexed, and that
/// arm would silently return fewer rows.
///
/// See [`arm_maintained_index_names`] for which indexes are collected today: a
/// vector search with a scalar prefilter is not yet among them.
///
/// An empty `index_names` (a plain scan) uses the compaction watermark alone.
/// First occurrence per shard mirrors Lance's `compacted_generation_for_shard`.
/// Per-shard SSTable exclusion watermark: the generation at or below which SSTables
/// are safe to drop for this arm. A generation is droppable only once it is
/// compacted into the base table AND covered by `index_name`'s catch-up (for an
/// indexed arm); a plain scan (`index_name == None`) uses the compaction watermark
/// alone. Capping at the index catch-up keeps rows the base index has not yet
/// indexed visible through their SSTable. First occurrence per shard mirrors Lance's
/// `compacted_generation_for_shard`.
fn exclusion_watermarks(
details: &MemWalIndexDetails,
index_names: &[String],
catchup_required: bool,
index_name: Option<&str>,
) -> HashMap<Uuid, u64> {
let mut exclude: HashMap<Uuid, u64> = HashMap::new();
for entry in &details.compacted_sstables {
let mut watermark = entry.generation;
for name in index_names {
match details
if let Some(name) = index_name
&& let Some(caught_up) = details
.index_catchup
.iter()
.find(|icp| icp.index_name == *name)
.find(|icp| icp.index_name == name)
.and_then(|icp| icp.caught_up_generation_for_shard(&entry.shard_id))
{
Some(caught_up) => watermark = watermark.min(caught_up),
// No entry. On a table that requires catch-up this means the
// index is *not* known to hold these rows, and the base arm is
// index-only -- so every generation stays readable from its
// SSTable. Without the bit the field is not maintained at all,
// and absence carries no information.
None if catchup_required => watermark = 0,
None => {}
}
{
watermark = watermark.min(caught_up);
}
exclude.entry(entry.shard_id).or_insert(watermark);
}
@@ -283,26 +267,13 @@ fn exclusion_watermarks(
/// with a live cached `ShardWriter` (this session's in-flight writes) the
/// writer's authoritative in-memory manifest and memtables override the
/// on-disk view so a read sees data not yet flushed.
/// Whether this table reads a missing `index_catchup` entry as "not caught up".
///
/// Both words must be set. A reader honouring the bit while a writer does not
/// would retain SSTables the writer had already trimmed, and the reverse would
/// serve rows from files the writer still expects to be excluded -- so a
/// half-set manifest is treated as legacy, which is the conservative side.
fn requires_index_catchup(dataset: &Dataset) -> bool {
let manifest = dataset.manifest();
manifest.reader_feature_flags & FLAG_MEM_WAL_INDEX_CATCHUP != 0
&& manifest.writer_feature_flags & FLAG_MEM_WAL_INDEX_CATCHUP != 0
}
async fn build_read_context(
table: &NativeTable,
dataset: &Dataset,
details: &MemWalIndexDetails,
index_names: &[String],
index_name: Option<&str>,
) -> Result<(Vec<ShardSnapshot>, HashMap<Uuid, InMemoryMemTables>)> {
let catchup_required = requires_index_catchup(dataset);
let exclude = exclusion_watermarks(details, index_names, catchup_required);
let exclude = exclusion_watermarks(details, index_name);
let shard_ids = dataset.list_mem_wal_latest_shard_ids().await?;
// Use the dataset's own object store (not `ObjectStore::from_uri`, which
@@ -516,33 +487,19 @@ async fn index_maintained(
}))
}
/// Every maintained base index this query relies on, used to gate SSTable
/// exclusion by index catch-up.
///
/// Returns a list because the watermark must be the lowest across every index a
/// query relies on. Today it never holds more than one: `reject_unsupported`
/// refuses hybrid search, so the vector and full-text arms are mutually
/// exclusive.
///
/// The case that is genuinely multi-index -- a vector search with a scalar or
/// bitmap prefilter -- is **not collected yet**. Identifying those needs the
/// planner's chosen indexes, not the columns the filter names, and no Lance API
/// exposes them. Until it does, such a query is gated on its vector index alone.
///
/// Empty for a plain scan, or when no maintained index covers the searched
/// column.
async fn arm_maintained_index_names(
/// The maintained base index the query's arm relies on (vector index for ANN, FTS
/// index for full-text), used to gate SSTable compaction exclusion by index catch-up.
/// `None` for a plain scan or when no maintained index covers the searched column.
async fn arm_maintained_index_name(
dataset: &Dataset,
query: &VectorQueryRequest,
details: &MemWalIndexDetails,
) -> Result<Vec<String>> {
) -> Result<Option<String>> {
use lance::index::DatasetIndexExt;
// Each arm's searched column, the index-detail type it relies on, and a
// Resolve the arm's searched column, the index-detail type it relies on, and a
// label for diagnostics — catch-up is taken from the vector/FTS index
// specifically, not a BTree on the same column.
let mut arms: Vec<(String, &str, &str)> = Vec::new();
if !query.query_vector.is_empty() {
let (column, type_url_suffix, arm) = if !query.query_vector.is_empty() {
let arrow_schema = ArrowSchema::from(dataset.schema());
let column = match &query.column {
Some(column) => column.clone(),
@@ -551,43 +508,31 @@ async fn arm_maintained_index_names(
default_vector_column(&arrow_schema, dim)?
}
};
arms.push((column, "VectorIndexDetails", "vector"));
}
if let Some(fts) = &query.base.full_text_search
&& let Some(column) = fts.columns().into_iter().next()
{
arms.push((column, "InvertedIndexDetails", "full-text"));
}
if arms.is_empty() {
return Ok(Vec::new());
}
let indices = dataset.load_indices().await?;
let mut names = Vec::with_capacity(arms.len());
for (column, type_url_suffix, arm) in arms {
let Some(field) = dataset.schema().field(&column) else {
continue;
};
let segment_names: Vec<String> = indices
.iter()
.filter(|idx| {
idx.fields.contains(&field.id)
&& idx
.index_details
.as_ref()
.is_some_and(|d| d.type_url.ends_with(type_url_suffix))
})
.map(|idx| idx.name.clone())
.collect();
if let Some(name) =
resolve_single_index(segment_names, &details.maintained_indexes, arm, &column)?
{
names.push(name);
(column, "VectorIndexDetails", "vector")
} else if let Some(fts) = &query.base.full_text_search {
match fts.columns().into_iter().next() {
Some(column) => (column, "InvertedIndexDetails", "full-text"),
None => return Ok(None),
}
}
names.sort();
names.dedup();
Ok(names)
} else {
return Ok(None);
};
let Some(field) = dataset.schema().field(&column) else {
return Ok(None);
};
let indices = dataset.load_indices().await?;
let segment_names: Vec<String> = indices
.iter()
.filter(|idx| {
idx.fields.contains(&field.id)
&& idx
.index_details
.as_ref()
.is_some_and(|d| d.type_url.ends_with(type_url_suffix))
})
.map(|idx| idx.name.clone())
.collect();
resolve_single_index(segment_names, &details.maintained_indexes, arm, &column)
}
/// Resolve the single logical index from the names of its matching physical
@@ -789,117 +734,22 @@ mod tests {
};
// Plain scan: drop every compacted generation (through 5).
assert_eq!(
exclusion_watermarks(&details, &[], false).get(&shard),
Some(&5)
);
assert_eq!(exclusion_watermarks(&details, None).get(&shard), Some(&5));
// FTS arm with a lagging index: exclusion is capped at the index catch-up
// (2), so SSTable generations 3..=5 are retained until the index covers
// them — otherwise those documents would silently vanish from FTS results.
assert_eq!(
exclusion_watermarks(&details, &["fts_idx".to_string()], false).get(&shard),
exclusion_watermarks(&details, Some("fts_idx")).get(&shard),
Some(&2)
);
// A caught-up index — or one untracked in index_catchup — falls back to the
// compaction watermark.
assert_eq!(
exclusion_watermarks(&details, &["caught_up_idx".to_string()], false).get(&shard),
exclusion_watermarks(&details, Some("caught_up_idx")).get(&shard),
Some(&5)
);
// The same missing entry, once the table requires catch-up: absence now
// means "not known to hold these rows", so nothing may be excluded and
// every generation stays readable from its SSTable. This is the whole
// point of the protocol -- an indexed query against a table whose index
// has not caught up must not silently lose rows.
assert_eq!(
exclusion_watermarks(&details, &["untracked_idx".to_string()], true).get(&shard),
Some(&0)
);
// A tracked index is unaffected by the mode: the recorded position is
// information either way, and it still caps the exclusion.
assert_eq!(
exclusion_watermarks(&details, &["fts_idx".to_string()], true).get(&shard),
Some(&2)
);
// One missing entry is enough to hold everything back, even alongside an
// index that has caught up.
let mixed = vec!["fts_idx".to_string(), "untracked_idx".to_string()];
assert_eq!(
exclusion_watermarks(&details, &mixed, true).get(&shard),
Some(&0)
);
}
/// A hybrid search reads a vector and a full-text index, and either may lag.
/// Retaining to the lower of the two is what keeps both arms complete;
/// gating on one alone would drop SSTables the other has not indexed.
#[test]
fn exclusion_watermark_takes_the_minimum_across_every_index_used() {
let shard = Uuid::from_u128(1);
let details = MemWalIndexDetails {
compacted_sstables: vec![CompactedSsTable::new(shard, 9)],
index_catchup: vec![
IndexCatchupProgress::new(
"vec_idx".to_string(),
vec![CompactedSsTable::new(shard, 7)],
),
IndexCatchupProgress::new(
"fts_idx".to_string(),
vec![CompactedSsTable::new(shard, 4)],
),
],
maintained_indexes: vec!["vec_idx".to_string(), "fts_idx".to_string()],
..Default::default()
};
// Each index alone stops at its own catch-up.
assert_eq!(
exclusion_watermarks(&details, &["vec_idx".to_string()], false).get(&shard),
Some(&7)
);
assert_eq!(
exclusion_watermarks(&details, &["fts_idx".to_string()], false).get(&shard),
Some(&4)
);
// Used together, the lower one governs regardless of order.
let both = ["vec_idx".to_string(), "fts_idx".to_string()];
assert_eq!(
exclusion_watermarks(&details, &both, false).get(&shard),
Some(&4)
);
let reversed = ["fts_idx".to_string(), "vec_idx".to_string()];
assert_eq!(
exclusion_watermarks(&details, &reversed, false).get(&shard),
Some(&4)
);
}
/// An index with no catch-up entry contributes no cap today, so a lagging
/// sibling must still govern rather than being widened by the untracked one.
#[test]
fn an_untracked_index_does_not_widen_a_lagging_sibling() {
let shard = Uuid::from_u128(1);
let details = MemWalIndexDetails {
compacted_sstables: vec![CompactedSsTable::new(shard, 9)],
index_catchup: vec![IndexCatchupProgress::new(
"fts_idx".to_string(),
vec![CompactedSsTable::new(shard, 4)],
)],
maintained_indexes: vec!["fts_idx".to_string(), "untracked_idx".to_string()],
..Default::default()
};
let both = ["fts_idx".to_string(), "untracked_idx".to_string()];
assert_eq!(
exclusion_watermarks(&details, &both, false).get(&shard),
Some(&4)
);
}
#[test]
+86 -517
View File
@@ -7,24 +7,16 @@
//! therefore idempotent and does not observe input mutation -- once a row is
//! filled, changing what the expression reads leaves the stored result alone.
//!
//! Two passes per fragment. The first scans only the unfilled live rows and
//! evaluates the expression over them, which yields the exact fill count and
//! decides whether the fragment is staged at all -- a fragment where nothing
//! would change stages nothing, which is what lets an expression yielding
//! null settle instead of restaging forever. The second streams the
//! fragment's physical rows into `write_column` a batch at a time, so peak
//! memory is bounded by a scan batch. The expression is evaluated by this
//! module, never through a projection alias, and only over rows being
//! filled: every other row -- deleted, or already holding a value -- has its
//! inputs masked to null first, so a poison value in a row nobody is filling
//! cannot fail the refresh.
//! Convergence comes from staging nothing when nothing would change, so an
//! expression yielding null settles after one pass rather than re-selecting
//! the same rows forever. Fragments that already cover the column and hold no
//! nulls are skipped without evaluating it at all.
use std::sync::Arc;
use arrow_array::{ArrayRef, BooleanArray, RecordBatch, RecordBatchOptions};
use arrow_array::RecordBatch;
use arrow_schema::Schema as ArrowSchema;
use datafusion_expr::ColumnarValue;
use futures::{Stream, StreamExt, TryStreamExt};
use futures::{TryStreamExt, stream};
use lance::Dataset;
use lance::dataset::WriteDestination;
use lance::dataset::fragment::FileFragment;
@@ -33,11 +25,14 @@ use lance_core::ROW_ID;
use lance_core::datatypes::Schema as LanceSchema;
use serde::{Deserialize, Serialize};
use super::computed_columns::{BoundExpression, ComputedColumnKind, computed_column_from_field};
use super::{BaseTable, NativeTable};
use crate::job::Job;
use super::NativeTable;
use super::computed_columns::{ComputedColumnKind, computed_column_from_field};
use crate::{Error, Result};
/// Alias the expression is projected under, so its result and the column's
/// current values can be read side by side.
const COMPUTED_ALIAS: &str = "__lancedb_computed";
/// The result of refreshing a computed column.
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize, Default)]
pub struct RefreshColumnResult {
@@ -55,12 +50,9 @@ pub(crate) async fn execute_refresh_column(
column: &str,
) -> Result<RefreshColumnResult> {
table.dataset.ensure_mutable()?;
ensure_no_lsm_write_spec(table).await?;
let dataset = table.dataset.get().await?;
let expression = declared_expression(&dataset, column)?;
let schema = Arc::new(ArrowSchema::from(dataset.schema()));
let bound = Arc::new(super::computed_columns::bind(schema, column, &expression)?);
let field = dataset
.schema()
.field(column)
@@ -76,14 +68,18 @@ pub(crate) async fn execute_refresh_column(
let mut rows_filled = 0u64;
let mut replacements = Vec::new();
for fragment in dataset.get_fragments() {
let gained = count_fragment_gains(&dataset, &fragment, &bound, column).await?;
if gained == 0 {
for fragment in fragments_to_consider(&dataset, column, field.id).await? {
let Some((filled, values)) =
fill_fragment(&dataset, &fragment, column, &expression).await?
else {
continue;
}
rows_filled += gained;
let values = fill_stream(&dataset, &fragment, bound.clone(), column).await?;
replacements.push(fragment.write_column(values, &column_schema).await?);
};
rows_filled += filled;
replacements.push(
fragment
.write_column(stream::iter(values.into_iter().map(Ok)), &column_schema)
.await?,
);
}
if replacements.is_empty() {
@@ -94,16 +90,13 @@ pub(crate) async fn execute_refresh_column(
}
let read_version = dataset.version().version;
// The dataset's own session, so registrations and caches survive the
// commit being installed on the handle.
let session = dataset.session();
let new_dataset = Dataset::commit(
WriteDestination::Dataset(dataset.clone()),
Operation::DataReplacement { replacements },
Some(read_version),
None,
None,
session,
Arc::new(Default::default()),
false,
)
.await?;
@@ -116,45 +109,6 @@ pub(crate) async fn execute_refresh_column(
})
}
/// Run the refresh as a [`Job`] in this process.
pub(crate) async fn execute_refresh_column_async(table: &NativeTable, column: &str) -> Result<Job> {
// Validate before spawning so bad input is reported by this call rather
// than only by the job.
table.dataset.ensure_mutable()?;
ensure_no_lsm_write_spec(table).await?;
let dataset = table.dataset.get().await?;
declared_expression(&dataset, column)?;
drop(dataset);
let table = table.clone();
let column = column.to_string();
Ok(Job::spawned(tokio::spawn(async move {
execute_refresh_column(&table, &column).await?;
table.bump_freshness();
Ok(())
})))
}
/// Refuse to refresh under an LSM write spec.
///
/// Refresh enumerates base fragments, and a write spec keeps visible rows in
/// un-compacted MemWAL tiers it cannot reach -- success would silently omit
/// readable rows.
async fn ensure_no_lsm_write_spec(table: &NativeTable) -> Result<()> {
// The catch-up flag outlives unset and marks retained SSTable rows.
let catchup = table.dataset.get().await?.manifest().reader_feature_flags
& lance_table::feature_flags::FLAG_MEM_WAL_INDEX_CATCHUP
!= 0;
if catchup || table.get_lsm_write_spec().await?.is_some() {
return Err(Error::NotSupported {
message: "refresh_column is not supported on a table with an LSM write \
spec: rows in un-compacted tiers are invisible to refresh"
.into(),
});
}
Ok(())
}
/// The SQL expression `column` is declared with.
fn declared_expression(dataset: &Dataset, column: &str) -> Result<String> {
let schema = ArrowSchema::from(dataset.schema());
@@ -186,97 +140,53 @@ fn quote_identifier(name: &str) -> String {
format!("`{}`", name.replace('`', "``"))
}
/// Assemble the batch evaluation runs against: the bound roots, in read-schema
/// order. Built by name so scan-side column order never matters.
fn evaluation_batch(
batch: &RecordBatch,
bound: &BoundExpression,
mask_out: Option<&BooleanArray>,
) -> lance_core::Result<RecordBatch> {
let mut columns = Vec::with_capacity(bound.roots.len());
for name in &bound.roots {
let column = batch.column_by_name(name).ok_or_else(|| {
lance_core::Error::invalid_input(format!(
"refreshing a computed column read no {name} column"
))
})?;
// Rows outside the mask must not reach the expression: a value in a
// deleted or already-filled row can be one it would choke on.
columns.push(match mask_out {
Some(mask) => arrow::compute::nullif(column, mask)?,
None => column.clone(),
});
}
Ok(RecordBatch::try_new_with_options(
bound.read_schema.clone(),
columns,
&RecordBatchOptions::new().with_row_count(Some(batch.num_rows())),
)?)
}
/// Evaluate the expression over `batch`, materializing a constant result to
/// the batch's length.
fn evaluate(bound: &BoundExpression, batch: &RecordBatch) -> lance_core::Result<ArrayRef> {
let value = bound
.physical
.evaluate(batch)
.map_err(lance_core::Error::from)?;
match value {
ColumnarValue::Array(array) => Ok(array),
scalar => scalar
.into_array(batch.num_rows())
.map_err(lance_core::Error::from),
}
}
/// How many rows of one fragment would gain a value.
/// Fragments that could hold a row needing a value.
///
/// Scans only the unfilled live rows -- deleted rows never reach the
/// expression here, the filter having already excluded them -- and counts the
/// non-null results. Exact, so it is both the staging decision and the
/// fragment's contribution to `rows_filled`.
async fn count_fragment_gains(
/// A fragment whose data files do not carry the field cannot hold one that
/// does. One that carries it is asked, since a row rewrite -- an update, or a
/// compaction folding an unfilled fragment into a filled one -- can leave
/// nulls behind a covering file.
async fn fragments_to_consider(
dataset: &Dataset,
column: &str,
field_id: i32,
) -> Result<Vec<FileFragment>> {
let unfilled = format!("{} IS NULL", quote_identifier(column));
let mut considered = Vec::new();
for fragment in dataset.get_fragments() {
let covered = fragment
.metadata()
.files
.iter()
.any(|file| file.fields.contains(&field_id));
if !covered || fragment.count_rows(Some(unfilled.clone())).await? > 0 {
considered.push(fragment);
}
}
Ok(considered)
}
/// Compute one fragment's column, keeping every value it already holds.
///
/// `Ok(None)` when no live row gained a value, which is what keeps a refresh
/// from restaging a fragment whose expression yields null. Deleted rows are
/// carried through so the values line up positionally with the fragment's data
/// files; they are never read back, but the column file has to cover them.
async fn fill_fragment(
dataset: &Dataset,
fragment: &FileFragment,
bound: &BoundExpression,
column: &str,
) -> Result<u64> {
let mut scanner = dataset.scan();
scanner
.with_fragments(vec![fragment.metadata().clone()])
.with_row_id()
.filter(&format!("{} IS NULL", quote_identifier(column)))?
.project(&bound.roots)?;
let mut gained = 0u64;
let mut batches = scanner.try_into_stream().await?;
while let Some(batch) = batches.try_next().await? {
let evaluated = evaluate(bound, &evaluation_batch(&batch, bound, None)?)?;
gained += (batch.num_rows() - evaluated.null_count()) as u64;
}
Ok(gained)
}
/// Stream one fragment's column in physical order, filling the unfilled live
/// rows and keeping every other value.
///
/// Deleted rows are carried through so the values line up positionally with
/// the fragment's data files; they are never read back, but the column file
/// has to cover them.
async fn fill_stream(
dataset: &Dataset,
fragment: &FileFragment,
bound: Arc<BoundExpression>,
column: &str,
) -> Result<impl Stream<Item = lance_core::Result<RecordBatch>> + Send + use<>> {
let mut projection: Vec<String> = bound.roots.clone();
projection.push(column.to_string());
expression: &str,
) -> Result<Option<(u64, Vec<RecordBatch>)>> {
let mut scanner = dataset.scan();
scanner
.with_fragments(vec![fragment.metadata().clone()])
.with_row_id()
.include_deleted_rows()
.project(&projection)?;
.project_with_transform(&[
(column, quote_identifier(column).as_str()),
(COMPUTED_ALIAS, expression),
])?;
let projected = Arc::new(ArrowSchema::new(vec![
ArrowSchema::from(dataset.schema())
@@ -287,39 +197,41 @@ async fn fill_stream(
.clone(),
]));
let column = column.to_string();
let batches = scanner.try_into_stream().await?;
Ok(batches.map(move |batch| {
let batch = batch?;
let missing = |name: &str| {
lance_core::Error::invalid_input(format!(
"refreshing a computed column read no {name} column"
))
};
let missing = |name: &str| Error::Runtime {
message: format!("refreshing {column} produced no {name} column"),
};
let mut filled = 0u64;
let mut values = Vec::new();
let mut batches = scanner.try_into_stream().await?;
while let Some(batch) = batches.try_next().await? {
let existing = batch
.column_by_name(&column)
.ok_or_else(|| missing(&column))?;
.column_by_name(column)
.ok_or_else(|| missing(column))?;
let computed = batch
.column_by_name(COMPUTED_ALIAS)
.ok_or_else(|| missing("expression"))?;
let row_ids = batch
.column_by_name(ROW_ID)
.ok_or_else(|| missing(ROW_ID))?;
// Only an unfilled live row gains a value; a deleted row has a null
// row id and keeps its (null) slot.
// A row is filled only if it gains a value: an expression yielding null
// leaves it as unfilled as it was, which is what lets a refresh settle.
// A deleted row has a null row id; its value is written but not counted.
let unfilled = arrow::compute::is_null(existing.as_ref())?;
let live = arrow::compute::is_not_null(row_ids.as_ref())?;
let fill = arrow::compute::and(&unfilled, &live)?;
let keep = arrow::compute::not(&fill)?;
filled += (0..unfilled.len())
.filter(|i| unfilled.value(*i) && row_ids.is_valid(*i) && computed.is_valid(*i))
.count() as u64;
let computed = evaluate(&bound, &evaluation_batch(&batch, &bound, Some(&keep))?)?;
let merged = arrow_select::zip::zip(&fill, &computed, existing)?;
Ok(RecordBatch::try_new(projected.clone(), vec![merged])?)
}))
let merged = arrow_select::zip::zip(&unfilled, computed, existing)?;
values.push(RecordBatch::try_new(projected.clone(), vec![merged])?);
}
Ok((filled > 0).then_some((filled, values)))
}
#[cfg(test)]
mod tests {
use std::sync::Arc;
use arrow_array::{Int32Array, record_batch};
use futures::TryStreamExt;
@@ -584,89 +496,6 @@ mod tests {
);
}
/// A fragment spanning several scan batches exercises the streamed fill:
/// the probe buffers only until the first gained value and the rest flows
/// through write_column a batch at a time.
#[tokio::test]
async fn test_refresh_streams_a_multi_batch_fragment() {
let values: Vec<i32> = (0..20_000).collect();
let table = table_with("refresh_multi_batch", values.clone()).await;
declare_doubled(&table).await.unwrap();
let result = table.refresh_column("doubled").await.unwrap();
assert_eq!(result.rows_filled, 20_000);
let read_back = read(&table, "doubled").await;
assert_eq!(read_back.len(), 20_000);
let mut expected: Vec<Option<i32>> = values.iter().map(|v| Some(v * 2)).collect();
expected.sort();
assert_eq!(read_back, expected);
}
/// The gate's reproducer: the commit must reuse the configured session,
/// or registrations and caches vanish from the handle after a refresh.
#[tokio::test]
async fn test_refresh_preserves_the_configured_session() {
let session = Arc::new(lance::session::Session::default());
let conn = crate::connect("memory://")
.session(session.clone())
.execute()
.await
.unwrap();
let batch = record_batch!(("x", Int32, [1, 2])).unwrap();
let table = conn
.create_table("session_kept", batch)
.execute()
.await
.unwrap();
declare_doubled(&table).await.unwrap();
table.refresh_column("doubled").await.unwrap();
let dataset = table.as_native().unwrap().dataset.get().await.unwrap();
assert!(Arc::ptr_eq(&dataset.session(), &session));
}
/// The async form's job settles with the fill visible, like
/// create_index's execute_async.
#[tokio::test]
async fn test_refresh_async_job_waits_for_the_fill() {
let table = table_with("refresh_async", vec![1, 2, 3]).await;
declare_doubled(&table).await.unwrap();
let job = table.refresh_column_async("doubled").await.unwrap();
assert!(job.id().is_none(), "in-process jobs have no server id");
job.wait().await.unwrap();
assert_eq!(job.status().await.unwrap(), "finished");
assert_eq!(
read(&table, "doubled").await,
vec![Some(2), Some(4), Some(6)]
);
}
/// Bad input is reported by the call, not by the job.
#[tokio::test]
async fn test_refresh_async_rejects_bad_input_before_spawning() {
let table = table_with("refresh_async_bad", vec![1, 2, 3]).await;
let err = table.refresh_column_async("x").await.unwrap_err();
assert!(matches!(err, Error::NotAComputedColumn { name } if name == "x"));
let err = table.refresh_column_async("nope").await.unwrap_err();
assert!(matches!(err, Error::ColumnNotFound { name } if name == "nope"));
}
#[tokio::test]
async fn test_refresh_async_job_reports_success_to_every_waiter() {
let table = table_with("refresh_async_waiters", vec![1, 2]).await;
declare_doubled(&table).await.unwrap();
let job = table.refresh_column_async("doubled").await.unwrap();
job.wait().await.unwrap();
// A second wait after completion observes the same outcome.
job.wait().await.unwrap();
assert_eq!(job.status().await.unwrap(), "finished");
}
#[tokio::test]
async fn test_refresh_rejects_a_plain_column() {
let table = table_with("refresh_plain", vec![1, 2, 3]).await;
@@ -681,266 +510,6 @@ mod tests {
assert!(matches!(err, Error::ColumnNotFound { name } if name == "nope"));
}
/// The gate's reproducer: a poison value in a deleted row must not
/// abort filling the live rows, since nobody can read it.
#[tokio::test]
async fn test_a_deleted_rows_value_is_never_evaluated() {
let table = table_with("refresh_deleted_poison", vec![1, 0]).await;
table
.add_columns()
.computed("quotient", "10 / x")
.execute()
.await
.unwrap();
table.delete("x = 0").await.unwrap();
let result = table.refresh_column("quotient").await.unwrap();
assert_eq!(result.rows_filled, 1);
assert_eq!(read(&table, "quotient").await, vec![Some(10)]);
}
/// The gate's reproducer: an already-filled row's value must not be
/// re-evaluated either -- its input may have mutated into one the
/// expression chokes on.
#[tokio::test]
async fn test_a_filled_rows_value_is_never_evaluated() {
let table = table_with("refresh_filled_poison", vec![1, 2]).await;
table
.add_columns()
.computed("quotient", "10 / x")
.execute()
.await
.unwrap();
table.refresh_column("quotient").await.unwrap();
table
.update()
.column("x", "0")
.only_if("x = 1")
.execute()
.await
.unwrap();
append(&table, vec![5]).await;
let result = table.refresh_column("quotient").await.unwrap();
assert_eq!(result.rows_filled, 1);
assert_eq!(
read(&table, "quotient").await,
vec![Some(2), Some(5), Some(10)]
);
}
/// The gate's reproducer: the old internal projection alias is an
/// ordinary column name; a computed column may use it.
#[tokio::test]
async fn test_refresh_a_column_named_like_the_old_alias() {
let table = table_with("refresh_alias_name", vec![1, 2]).await;
table
.add_columns()
.computed("__lancedb_computed", "x * 2")
.execute()
.await
.unwrap();
let result = table.refresh_column("__lancedb_computed").await.unwrap();
assert_eq!(result.rows_filled, 2);
assert_eq!(
read(&table, "__lancedb_computed").await,
vec![Some(2), Some(4)]
);
}
/// The gate's reproducer: a late-gain fragment (filled, then one null row
/// compacted onto the end) fills without the old probe's buffering, which
/// this pins behaviorally; the memory bound is structural -- the fill
/// stream retains no batches at all.
#[tokio::test]
async fn test_refresh_fills_a_late_gain_fragment() {
let values: Vec<i32> = (0..20_000).collect();
let table = table_with("refresh_late_gain", values).await;
declare_doubled(&table).await.unwrap();
table.refresh_column("doubled").await.unwrap();
append(&table, vec![2_000_000]).await;
table
.optimize(crate::table::OptimizeAction::Compact {
options: crate::table::CompactionOptions::default(),
remap_options: None,
})
.await
.unwrap();
let result = table.refresh_column("doubled").await.unwrap();
assert_eq!(result.rows_filled, 1);
let read_back = read(&table, "doubled").await;
assert_eq!(read_back.len(), 20_001);
assert_eq!(read_back.last().unwrap(), &Some(4_000_000));
}
/// The gate's reproducer: a nested input declares, refreshes, and guards
/// its root against invalidating schema changes.
#[tokio::test]
async fn test_a_nested_input_declares_and_refreshes() {
use arrow_array::{Int32Array, StructArray};
use arrow_schema::{DataType, Field, Fields};
let conn = connect("memory://").execute().await.unwrap();
let age = Arc::new(Int32Array::from(vec![30, 40]));
let fields = Fields::from(vec![Field::new("age", DataType::Int32, true)]);
let metadata = StructArray::new(fields.clone(), vec![age as _], None);
let schema = Arc::new(arrow_schema::Schema::new(vec![Field::new(
"metadata",
DataType::Struct(fields),
true,
)]));
let batch =
arrow_array::RecordBatch::try_new(schema, vec![Arc::new(metadata) as _]).unwrap();
let table = conn
.create_table("refresh_nested", batch)
.execute()
.await
.unwrap();
table
.add_columns()
.computed("next_age", "metadata.age + 1")
.execute()
.await
.unwrap();
let declaration =
&crate::table::computed_columns(table.schema().await.unwrap().as_ref())[0];
assert_eq!(declaration.inputs, vec!["metadata.age".to_string()]);
let result = table.refresh_column("next_age").await.unwrap();
assert_eq!(result.rows_filled, 2);
assert_eq!(read(&table, "next_age").await, vec![Some(31), Some(41)]);
// The dotted input guards its root.
let err = table.drop_columns(&["metadata"]).await.unwrap_err();
assert!(
matches!(&err, Error::InvalidInput { message } if message.contains("next_age")),
"{err:?}"
);
// Masking a struct input for a deleted row goes through the same
// nullif path as a primitive; a nested input plus deletions must not
// be the combination that breaks it.
table.delete("next_age = 31").await.unwrap();
append_struct_row(&table, 50).await;
let result = table.refresh_column("next_age").await.unwrap();
assert_eq!(result.rows_filled, 1);
assert_eq!(read(&table, "next_age").await, vec![Some(41), Some(51)]);
}
/// Append one `metadata: {age}` row to the nested-input table.
async fn append_struct_row(table: &Table, age: i32) {
use arrow_array::{Int32Array, StructArray};
use arrow_schema::{DataType, Field, Fields};
let ages = Arc::new(Int32Array::from(vec![age]));
let fields = Fields::from(vec![Field::new("age", DataType::Int32, true)]);
let metadata = StructArray::new(fields.clone(), vec![ages as _], None);
let schema = Arc::new(arrow_schema::Schema::new(vec![Field::new(
"metadata",
DataType::Struct(fields),
true,
)]));
let batch =
arrow_array::RecordBatch::try_new(schema, vec![Arc::new(metadata) as _]).unwrap();
table.add(batch).execute().await.unwrap();
}
/// Both orders of declare+spec are refused at the source (see the
/// schema_evolution tests); refresh's own check covers a dataset another
/// writer left in that state.
#[tokio::test]
async fn test_refresh_refuses_a_foreign_lsm_state() {
use crate::table::LsmWriteSpec;
let tmp_dir = tempfile::tempdir().unwrap();
let conn = connect(tmp_dir.path().to_str().unwrap())
.execute()
.await
.unwrap();
let schema = Arc::new(arrow_schema::Schema::new(vec![arrow_schema::Field::new(
"x",
arrow_schema::DataType::Int32,
false,
)]));
let batch =
arrow_array::RecordBatch::try_new(schema, vec![Arc::new(Int32Array::from(vec![1]))])
.unwrap();
let table = conn.create_table("lsm", batch).execute().await.unwrap();
table.set_unenforced_primary_key(["x"]).await.unwrap();
table
.set_lsm_write_spec(LsmWriteSpec::unsharded())
.await
.unwrap();
super::super::computed_columns::add_foreign_kind(&table, "doubled", "sql").await;
let err = table.refresh_column("doubled").await.unwrap_err();
assert!(
matches!(&err, Error::NotSupported { message } if message.contains("LSM")),
"{err:?}"
);
let err = table.refresh_column_async("doubled").await.unwrap_err();
assert!(matches!(err, Error::NotSupported { .. }));
}
/// After catch-up activation and unset, no spec remains but the catch-up
/// flag still marks retained SSTable rows; refresh refuses on the flag.
#[tokio::test]
async fn test_refresh_refuses_retained_catchup_state() {
use crate::table::LsmWriteSpec;
let tmp_dir = tempfile::tempdir().unwrap();
let conn = connect(tmp_dir.path().to_str().unwrap())
.execute()
.await
.unwrap();
let schema = Arc::new(arrow_schema::Schema::new(vec![arrow_schema::Field::new(
"x",
arrow_schema::DataType::Int32,
false,
)]));
let batch = arrow_array::RecordBatch::try_new(
schema.clone(),
vec![Arc::new(Int32Array::from(vec![1]))],
)
.unwrap();
let table = conn
.create_table("catchup", batch.clone())
.execute()
.await
.unwrap();
table.set_unenforced_primary_key(["x"]).await.unwrap();
table
.set_lsm_write_spec(LsmWriteSpec::unsharded())
.await
.unwrap();
table.require_mem_wal_index_catchup().await.unwrap();
let mut merge = table.merge_insert(&["x"]);
merge
.when_matched_update_all(None)
.when_not_matched_insert_all()
.use_lsm(true);
merge
.execute(Box::new(arrow_array::RecordBatchIterator::new(
vec![Ok(batch)],
schema,
)))
.await
.unwrap();
table.unset_lsm_write_spec().await.unwrap();
super::super::computed_columns::add_foreign_kind(&table, "doubled", "sql").await;
let err = table.refresh_column("doubled").await.unwrap_err();
assert!(
matches!(&err, Error::NotSupported { message } if message.contains("LSM")),
"{err:?}"
);
}
/// A declaration of a kind this version cannot evaluate is refused by
/// name, rather than mistaken for a plain column or fed to the SQL path.
#[tokio::test]
+4 -94
View File
@@ -13,9 +13,9 @@ use lance::dataset::{ColumnAlteration, NewColumnTransform};
use serde::{Deserialize, Serialize};
use std::collections::HashMap;
use super::NativeTable;
use super::computed_columns;
use super::{BaseTable, NativeTable};
use crate::{Error, Result};
use crate::Result;
/// The result of an add columns operation.
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize, Default)]
@@ -100,48 +100,6 @@ pub(crate) async fn execute_add_columns(
table: &NativeTable,
transforms: NewColumnTransform,
read_columns: Option<Vec<String>>,
) -> Result<AddColumnsResult> {
// Declarations are admitted only through [`execute_declare`].
match &transforms {
NewColumnTransform::AllNulls(schema) => {
computed_columns::ensure_no_foreign_declarations(schema.fields())?
}
NewColumnTransform::BatchUDF(udf) => {
computed_columns::ensure_no_foreign_declarations(udf.output_schema.fields())?
}
_ => {}
}
commit_add_columns(table, transforms, read_columns).await
}
/// Declare validated computed columns. The only admission path for
/// declaration metadata.
pub(crate) async fn execute_declare(
table: &NativeTable,
columns: &[(String, String)],
) -> Result<AddColumnsResult> {
// An LSM write spec keeps visible rows in tiers refresh cannot reach;
// checked against latest committed state, not this handle's snapshot.
// The catch-up flag outlives unset and marks retained SSTable rows.
table.checkout_latest().await?;
let catchup = table.dataset.get().await?.manifest().reader_feature_flags
& lance_table::feature_flags::FLAG_MEM_WAL_INDEX_CATCHUP
!= 0;
if catchup || table.get_lsm_write_spec().await?.is_some() {
return Err(Error::NotSupported {
message: "computed columns are not supported on a table with an LSM write \
spec: rows in un-compacted tiers are invisible to refresh"
.into(),
});
}
let transform = computed_columns::declare(table.schema().await?, columns)?;
commit_add_columns(table, transform, None).await
}
pub(crate) async fn commit_add_columns(
table: &NativeTable,
transforms: NewColumnTransform,
read_columns: Option<Vec<String>>,
) -> Result<AddColumnsResult> {
table.dataset.ensure_mutable()?;
let mut dataset = (*table.dataset.get().await?).clone();
@@ -162,19 +120,12 @@ pub(crate) async fn execute_alter_columns(
let mut dataset = (*table.dataset.get().await?).clone();
// Nullability is not part of what an expression resolves against, so only
// a rename or a retype can invalidate a binding.
let schema = std::sync::Arc::new(ArrowSchema::from(dataset.schema()));
let rebinding = alterations
.iter()
.filter(|alteration| alteration.rename.is_some() || alteration.data_type.is_some())
.map(|alteration| alteration.path.as_str())
.collect::<Vec<_>>();
computed_columns::ensure_not_an_input(&schema, &rebinding)?;
let retyped = alterations
.iter()
.filter(|alteration| alteration.data_type.is_some())
.map(|alteration| alteration.path.as_str())
.collect::<Vec<_>>();
computed_columns::ensure_not_retyped(schema.as_ref(), &retyped)?;
computed_columns::ensure_not_an_input(&ArrowSchema::from(dataset.schema()), &rebinding)?;
dataset.alter_columns(alterations).await?;
let version = dataset.version().version;
table.dataset.update(dataset);
@@ -190,10 +141,7 @@ pub(crate) async fn execute_drop_columns(
) -> Result<DropColumnsResult> {
table.dataset.ensure_mutable()?;
let mut dataset = (*table.dataset.get().await?).clone();
computed_columns::ensure_not_an_input(
&std::sync::Arc::new(ArrowSchema::from(dataset.schema())),
columns,
)?;
computed_columns::ensure_not_an_input(&ArrowSchema::from(dataset.schema()), columns)?;
dataset.drop_columns(columns).await?;
let version = dataset.version().version;
table.dataset.update(dataset);
@@ -210,44 +158,6 @@ pub(crate) async fn execute_update_field_metadata(
table.dataset.ensure_mutable()?;
let mut dataset = (*table.dataset.get().await?).clone();
// A declaration is validated as a whole at declare time; editing its keys
// here would bypass that, fabricate one on a plain column, or move a
// binding out from under a refresh. A replace on a declared column would
// silently erase it.
let schema = ArrowSchema::from(dataset.schema());
let declared: Vec<String> = computed_columns::computed_columns(&schema)
.into_iter()
.map(|declaration| declaration.name)
.collect();
for update in updates {
if update
.metadata
.keys()
.any(|key| computed_columns::is_declaration_key(key))
{
return Err(Error::InvalidInput {
message: format!(
"metadata keys of a computed-column declaration cannot be edited \
(path '{}'); drop the column and declare it again",
update.path
),
});
}
if update.replace
&& declared
.iter()
.any(|name| name == computed_columns::root(&update.path))
{
return Err(Error::InvalidInput {
message: format!(
"replacing all metadata of computed column '{}' would erase its \
declaration; drop the column and declare it again",
update.path
),
});
}
}
let mut builder = dataset.update_field_metadata();
for update in updates {
let entries = update.metadata.iter().map(|(k, v)| (k.clone(), v.clone()));
-4
View File
@@ -82,10 +82,6 @@ pub(crate) async fn execute_update(
// 1. Snapshot the current dataset
let dataset = table.dataset.get().await?;
super::computed_columns::ensure_not_written(
&arrow_schema::Schema::from(dataset.schema()),
update.columns.iter().map(|(name, _)| name.as_str()),
)?;
// 2. Initialize the Lance Core builder
let mut builder = LanceUpdateBuilder::new(dataset);
-147
View File
@@ -1,147 +0,0 @@
// SPDX-License-Identifier: Apache-2.0
// SPDX-FileCopyrightText: Copyright The LanceDB Authors
use std::sync::Arc;
use arrow_array::{Int64Array, RecordBatch};
use arrow_schema::{DataType, Field, Schema};
use lance::dataset::{WriteMode, WriteParams};
use lancedb::{Result, TableBase, connect, connect_namespace, table::WriteOptions};
use tempfile::tempdir;
use url::Url;
fn empty_schema() -> Arc<Schema> {
Arc::new(Schema::new(vec![Field::new("id", DataType::Int64, false)]))
}
fn file_uri(path: &std::path::Path) -> String {
Url::from_file_path(path)
.unwrap_or_else(|_| panic!("not an absolute path: {}", path.display()))
.to_string()
}
#[tokio::test]
async fn test_add_bases_accepts_named_and_dataset_root_entries() -> Result<()> {
let tmp = tempdir().unwrap();
let db = connect(tmp.path().join("db").to_str().unwrap())
.execute()
.await?;
let table = db.create_empty_table("t", empty_schema()).execute().await?;
let media = tmp.path().join("media");
let parent = tmp.path().join("parent");
std::fs::create_dir_all(&media).unwrap();
std::fs::create_dir_all(&parent).unwrap();
table
.add_bases([
TableBase {
path: file_uri(&media),
name: Some("media".into()),
is_dataset_root: false,
},
TableBase {
path: file_uri(&parent),
name: Some("parent".into()),
is_dataset_root: true,
},
])
.await
}
#[tokio::test]
async fn test_add_bases_accepts_two_unnamed_paths() -> Result<()> {
let tmp = tempdir().unwrap();
let db = connect(tmp.path().join("db").to_str().unwrap())
.execute()
.await?;
let table = db.create_empty_table("t", empty_schema()).execute().await?;
let media = tmp.path().join("media");
let other = tmp.path().join("other");
std::fs::create_dir_all(&media).unwrap();
std::fs::create_dir_all(&other).unwrap();
table
.add_bases([&file_uri(&media), &file_uri(&other)])
.await
}
#[tokio::test]
async fn test_add_bases_write_and_read_through_registered_base() -> Result<()> {
let tmp = tempdir().unwrap();
let db = connect(tmp.path().join("db").to_str().unwrap())
.execute()
.await?;
let table = db.create_empty_table("t", empty_schema()).execute().await?;
let media = tmp.path().join("media");
std::fs::create_dir_all(&media).unwrap();
let media_uri = file_uri(&media);
table.add_bases([&media_uri]).await?;
let batch = RecordBatch::try_new(
empty_schema(),
vec![Arc::new(Int64Array::from(vec![1, 2, 3]))],
)
.unwrap();
table
.add(batch)
.write_options(WriteOptions {
lance_write_params: Some(WriteParams {
mode: WriteMode::Append,
target_base_names_or_paths: Some(vec![media_uri.clone()]),
..Default::default()
}),
})
.execute()
.await?;
assert_eq!(table.count_rows(None).await?, 3);
let dataset = table.dataset().unwrap().get().await?;
let registered = dataset
.manifest()
.base_paths
.values()
.find(|base| base.path == media_uri)
.expect("registered base");
assert_ne!(registered.id, 0);
assert!(registered.name.is_none());
assert!(
dataset.get_fragments().iter().any(|fragment| {
fragment
.metadata()
.files
.iter()
.any(|file| file.base_id == Some(registered.id))
}),
"written fragment should reference the registered base"
);
assert!(
std::fs::read_dir(&media)
.unwrap()
.filter_map(|entry| entry.ok())
.any(|entry| entry.path().extension().is_some_and(|ext| ext == "lance")),
"data file should land under the registered base"
);
Ok(())
}
#[tokio::test]
async fn test_memory_add_bases_accepts_a_file_uri() -> Result<()> {
let tmp = tempdir().unwrap();
let db = connect("memory://").execute().await?;
let table = db.create_empty_table("t", empty_schema()).execute().await?;
let media = tmp.path().join("media");
std::fs::create_dir_all(&media).unwrap();
table.add_bases([file_uri(&media)]).await
}
#[tokio::test]
async fn test_namespace_add_bases_accepts_a_file_uri() -> Result<()> {
let tmp = tempdir().unwrap();
let mut properties = std::collections::HashMap::new();
properties.insert("root".to_string(), tmp.path().to_str().unwrap().to_string());
let db = connect_namespace("dir", properties).execute().await?;
let table = db.create_empty_table("t", empty_schema()).execute().await?;
let media = tmp.path().join("media");
std::fs::create_dir_all(&media).unwrap();
table.add_bases([file_uri(&media)]).await
}