Dropping a view unbinds its name and leaves the definition dataset to a
server-side cleanup job, so the two are separate events a caller may
want to wait on.
`drop_view_async` returns that job — the same shape
`drop_materialized_view_async` and `drop_function_async` already use: a
`202` carries the job id, a `200` (nothing was bound to the name) yields
an already-finished job with no id, and any other success status is an
error rather than a silent no-op.
## `drop_view` waits
`drop_view` now awaits the job before returning, so a caller who does
not want to think about cleanup gets the stronger guarantee: when it
returns, the definition really is deleted.
That is deliberately **different** from `drop_materialized_view` and
`drop_function`, which return as soon as the name is unbound and
document that content may still be deleting. The view API is the newer
one, and waiting is the semantic worth having; the other two are left
alone rather than changing behaviour already released.
## Surfaces
`Database` trait, the remote client, `Connection`, and the Python and
Node bindings — matching where `drop_materialized_view_async` is already
exposed.
Four client tests cover the accepted case reporting its job id, the
nothing-bound case reporting a finished job, a `202` without a usable
`job_id`, and an unexpected success status.
Update the Rust workspace Lance dependencies and Java lance-core from
v13.0.0-beta.8 to
[v13.0.0-beta.12](https://github.com/lance-format/lance/releases/tag/v13.0.0-beta.12),
refreshing Cargo.lock.
No compatibility fixes were required; validation passed with `cargo
clippy --quiet --workspace --tests --all-features -- -D warnings`,
`cargo fmt --all --quiet`, and `git diff --check`.
---------
Co-authored-by: Lu Qiu <luqiujob@gmail.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
## Cause
`RemoteDatabase` cached and returned the same `Arc<RemoteTable>` for
every open of a table on one connection. Checkout mutates version and
freshness state on that shared object, so pinning one handle also pinned
the others.
## Fix
Cache the table's server version rather than its mutable handle. Each
`open_table` now creates its own `RemoteTable` with independent
checkout, schema cache, and freshness state.
## Validation
- Added a cache-enabled regression test that failed before the fix:
pinning one handle changed another handle's reported version. It now
verifies a write through the independent handle succeeds.
- `cargo test --quiet --features remote -p lancedb --lib
remote::db::tests` (90 passed)
- `cargo check --quiet --features remote --tests --examples`
- `cargo clippy --quiet --features remote --tests --examples`
- `cargo fmt --all`
Fixes#4296
<!-- lance-gatekeeper-fix:v1 agent=38593bfbce0279286da9850fe61050e0
generation=1 -->
Co-authored-by: Gatefixer <313497061+lancedb-gatefixer[bot]@users.noreply.github.com>
## Cause
`io.RawIOBase.__del__` checks `closed` and calls `close()` during GC.
Both PyO3 methods entered `block_on`, which panics when GC runs on a
LanceDB Tokio worker.
## Fix
Track Python blob handle closure atomically so `closed` is synchronous.
`close()` marks the handle closed immediately and, when called on a
Tokio thread, schedules the underlying async cleanup without blocking
that thread. Reads and seeks after close fail immediately; ordinary
callers still wait for cleanup.
## Validation
- Rebuilt the Python extension with `uv run --extra tests --extra dev
maturin develop --extras tests,dev`.
- Ran the blob test file and remote blob handle test: 69 passed, 1
skipped.
- Ran `cargo fmt --all`, `ruff format .`, and `ruff check .`.
Fixes#4283
<!-- lance-gatekeeper-fix:v1 agent=15a7bd545e63dea4b69faa0429f0a684
generation=1 -->
Co-authored-by: Gatefixer <313497061+lancedb-gatefixer[bot]@users.noreply.github.com>
## Cause
Schema inference treated unrecognized objects as nested structs. Objects
without enumerable fields produced no field paths, so schema matching
dropped their columns before Arrow or blob conversion. Blob instances
with own fields were mistaken for struct fields.
## Fix
Walk only nonempty plain records as structs. Reject unsupported object
values in scalar fields, nested structs, and list elements with the
field path and row number. Preserve the object types already passed
through to Arrow. Support for converting these new input types is
tracked in #4270.
Nested class instances with enumerable fields previously became structs;
this change rejects them and is labeled `breaking-change`.
## Validation
- `pnpm build`
- `pnpm lint`
- `pnpm test __test__/arrow.test.ts __test__/blob.test.ts
__test__/table.test.ts --runInBand --silent` (626 passed)
- `pnpm run docs`
Fixes#4269
<!-- lance-gatekeeper-fix:v1 agent=426011291c209da027689ae88a0b4da7
generation=1 -->
---------
Co-authored-by: Gatefixer <313497061+lancedb-gatefixer[bot]@users.noreply.github.com>
Since #4241 restored fat LTO / 1 codegen unit for the Windows wheel, the
`windows` job in PyPI Publish has been cancelled at its 90-minute
timeout on every run, including the v0.40.0-beta.6 and v0.40.0-beta.7
tags, so neither release was published. In those runs dependency
compilation finishes after ~26 minutes and the `lancedb` crate compile +
fat-LTO step was still running after 63 minutes.
Move the job to the org's `windows-2025-8x-x64` larger runner and raise
the timeout to 150 minutes. Fat LTO stays, since the thin-LTO wheel
exceeds PyPI's 100 MiB limit.
The final fat-LTO step is largely single-threaded, so the gain from 8
cores is mainly in the parallel dependency phase plus faster/larger
hardware; the PR run of this workflow is the first measurement of fat
LTO on this runner.
Stacked on #4253.
Remote Python Functions could not share helper code, take typed
parameters, or keep per-process state: helpers became `from <module>
import ...` lines a worker cannot resolve, every parameter was an `env=`
string, and state had to be stashed on imported modules. MMLB's OpenAI
preset shows the cost — 24 env vars, a 302-line callable, and a rate
limiter duplicated between the scalar and batch variants, one of which
never constructed it (ENT-2516).
**Class Functions.** `@udf` accepts a class. Each remote instance runs
`__init__` once, calls `__call__` for every row or batch it processes
(row or batch mode is inferred from annotations as before), and calls
`close()` once if defined. Calling the definition locally constructs the
class, so unit tests stay ordinary.
**Initialization.** The annotated `__init__` parameters become the
Function's initialization fields. A binding passes their values beside
its column inputs: `fn(text=col("body"), model="small",
dimensions=512)`. Values are constants of that binding; another column
can bind the same Function version with other values. Types are limited
to booleans, integers, floats, strings, and lists or structs of them,
which have one unambiguous JSON and SQL spelling. A parameter with a
default may be omitted; a null value then takes the default.
Initialization and input names must be disjoint because both are keyword
arguments of one call. Secrets stay on `EnvVarSecret`.
**Helper code.** `code=[package, ...]` ships top-level modules or
packages as Python source in a `python_bundle` artifact (a canonical
JSON object of path → source), imported normally on the worker. Changing
a helper changes the artifact digest and so the Function version; the
environment from `pip`/`conda` is reused. Functions and classes defined
in `__main__` (notebook cells, scripts) are packaged by source,
recursively and dependencies first. An import of a module that lives in
a local source tree, is not shipped with `code=`, and names no declared
package is now rejected at registration instead of failing on the
worker. Closures, lambdas, and nested definitions are still rejected,
with the reason.
A Function without these features packages byte-identically to before
(`python_callable`, same digest, no `initialization` on the wire). Class
sources are located through a method's code object, because
`inspect.getsource` cannot find a class defined in a notebook cell or
doctest.
The packaging contract is documented on `udf`. Sophon changes that build
and execute these artifacts are in lancedb/sophon; they were verified
end to end on a Linux local server (registration → REST, SQL, and
materialized-view bindings with different initialization → refresh →
query, one construction per instance).
A grouped view's refresh runs its whole aggregate in one process, so a
view grouped by ivf_partition(col) over a large table is bounded by a
single worker however many workers a deployment has. The IVF index
already holds each partition's row ids, so the aggregate splits along
the index without a shuffle: every partition's groups are computed from
its own rows alone.
plan_grouped_refresh names the units (one per index partition plus one
for the rows the index cannot place) and the source version they read;
write_grouped_unit computes one unit's groups from the index's row ids,
plus the rows of fragments the index has not covered, assigned in place,
and writes them as uncommitted fragments; commit_grouped_refresh
replaces the view's rows with every unit's fragments in one Update.
The split has to publish what the single-pass refresh would. A unit
applies the view's predicate in its own query, because neither the
indexed take nor the fragment scan filters the way a lance scan does. An
index that does not say which fragments it covers has unknown coverage,
not empty, so it yields no plan at all and the caller refreshes in one
pass rather than reading those rows twice. Each result names its unit,
its plan and the view incarnation it was computed for -- two views of
one shape reach the same counters, and fragments written into one
dataset are not publishable into another -- and the commit publishes the
plan's units exactly once each or nothing. The commit lands on the
planned generation or is refused: lance rebases this Update over a
concurrent append rather than rejecting it, so the version it actually
landed on is checked, as the single-pass rebuild already does.
What holds every unit to one index is the source version the plan pins:
indices live in the source manifest, so a rebuild lands in a version the
units never read. Within that version a segment's postings can still
outlive its ownership -- a column rewrite attaches a new file and takes
the fragment out of the segment's bitmap without dropping its rows from
the posting lists -- so a unit keeps a segment's rows only while it
holds their fragment, and reads the rest from the scan.
The split is the grouping only where the index assigns by its own
centroids. lance also builds an index from precomputed partitions, and
records nowhere that it did, so a posting list can hold a row that
ivf_partition puts elsewhere -- that row's group would then be
aggregated in its own unit as well and published twice, since
concatenated fragments cannot merge two halves of a group. The plan
samples each partition and yields no units when they disagree, every
unit proves the rows it took before grouping them, and the commit
refuses a unit that wrote more than the single group its key allows.
A Function instance is created once with an initialization row
(`create(initialization, context)` in the Function Format), but nothing
in the client model could say what that row holds or where its values
come from, so every source-built Function ran with an empty row and
users pushed configuration through `env=` strings instead (ENT-2516).
This adds initialization to the canonical wire model without changing a
Function's identity:
- `FunctionSignature.initialization` lists the fields of the row, in
order. It is omitted when empty, so existing signatures and
FunctionVersion hashes are unchanged.
- `FunctionApplication.initialization` carries a binding's values as
JSON. The service validates them against the Function's fields and
re-encodes them as Arrow, so this JSON is transport only and never
hashed; that is why floats are accepted here while column-input literals
keep the Slice 1 domain.
- `FunctionBinding.initialization` is the validated one-row Arrow IPC
stream (base64) that every instance of the binding is created with.
Values belong to the binding, so one Function version can back several
columns with different values.
- The binding-shape check that guards schema mutations accepts the new
field; without it a table holding an initialized binding refused further
declarations.
Rust and Python share new canonical goldens for an initialized
application and binding.
The authoring side (class `@udf`, initialization arguments, shipped
helper code) is the stacked PR on top of this one.
A view is a named query a database stores and plans on every read. It
holds no rows, which is the whole difference from a materialized view.
## API
| Verb | Route |
| --- | --- |
| `create_view(name, query, namespace_path)` | `POST
/v1/view/{id}/create` |
| `describe_view(name, namespace_path)` | `POST /v1/view/{id}/describe`
|
| `drop_view(name, namespace_path)` | `POST /v1/view/{id}/drop` |
| `list_views(namespace_path)` | `GET /v1/namespace/{id}/view/list` |
On `Connection` and the `Database` trait, with the remote client, Python
(sync and async) and Node bindings. Local databases return
`NotSupported`: the server side is Sophon's, where a view is an object
of the database manifest.
`ViewDescription` carries the defining query, the database *and
namespace path* unqualified names in it resolve against, and the schema
the query resolved to. `create_view` returns one, so a caller has the
schema without a second call.
Both defaults travel with the view because it outlives the session that
declared it: the server re-plans the stored query on every read, so a
reader resolving an unqualified name against its own defaults would read
a different table. `default_namespace_path` crosses the wire as
`default_namespace`, a path like `namespace`, absent for the root.
There is no replace: a name already taken is an error, and changing a
view is a drop followed by a create, each authorized against what it
actually touches.
Querying a view stays SQL's job. There are no rows behind a view, so
there is no `open_view` returning a `Table`.
## Summary
Closes: #4060
- Convert Python dict/list objects to JSON strings during ingestion.
- Support JSON fields nested inside structs and lists.
- Cover `add` and `merge_insert` JSON ingestion paths.
- Add regression coverage for nested struct and list JSON fields.
## Testing
- `git diff --check`
- Python compilation passed.
- Focused pytest was attempted but could not complete because the native
extension build stalled during `uv` bootstrap.
Grouping near neighbours together, so an all-pairs comparison runs per
bucket instead of over the whole table, needs the partition an IVF index
assigns each vector.
`ivf_partition(column)` returns that partition, assigned by the
centroids and distance type of the IVF index on the column. It is bound
when the view is declared and again at every refresh; an index retrain
commits a new source version, so the next refresh regroups. A column
without an IVF index, or with two, is refused.
Stacked on #4222.
Dropping a Function can leave its content to a server-side cleanup job,
so the name drop and the content deletion become separate events a
caller may want to wait on.
`drop_function_async` returns the unbind result alongside a `Job` for
the cleanup, the same shape `drop_table_async` and
`drop_materialized_view_async` already use: a `202` carries the job id,
and a `200` — an inline deletion, or a name that was not bound — yields
an already-finished job with no id. A `202` without a usable job id is
rejected rather than silently reported as finished.
`drop_function` keeps its `bool` result and now delegates, so nothing
changes for callers that do not care when the content goes.
Available on `Connection` in Rust and on both the sync and asyncio
Python connections.
A view could only map source rows one to one, or one to many through a
FROM-position item, so nothing could see all the rows sharing a key.
A view query now takes `GROUP BY expr, ...` with aggregate projections,
planned and executed by DataFusion over the lance scan. A group spans
fragments, so a grouped view is recomputed in full whenever its source
changes; each row's provenance is its group's smallest source row id.
Such a query is stored as definition format 2, so a reader that predates
grouping reports the view as unrefreshable instead of failing to parse
it.
Stacked on #4190.
The production `remote` client talks to an existing namespace over HTTP.
It needs `lance-namespace-impls/rest`, not the REST *server* adapter.
`RestAdapter` is only used by `cfg(test)` integration tests that stand
up an in-process server. Enabling `rest-adapter` on the production
`remote` feature pulled that server stack, including Axum 0.7, into
default Python wheels.
This keeps `rest` on `remote` and moves `rest-adapter` to a
`lance-namespace-impls` dev-dependency so those tests still compile and
run. Cargo resolver=2 does not leak the extra feature into
`lancedb-python`.
## Measurement
Paired `maturin build --release --strip --target aarch64-apple-darwin
--features fp16kernels` wheels. Source was `3be29228` plus this
Cargo.toml change, which is this PR's tree (`878b2ae5` on `3be29228`).
Same toolchain, profile, and packaging flags; only `rest-adapter` moved.
| Artifact | Before | After | Delta |
| --- | ---: | ---: | ---: |
| Compressed wheel | 64,685,664 | 63,824,281 | −861,383 (−1.33%) |
| `_lancedb.abi3.so` uncompressed | 148,255,440 | 146,275,936 |
−1,979,504 (−1.34%) |
This is macOS arm64, not Windows. It does not resolve the Windows wheel
upload limit. Tonic still pulls Axum 0.8; the change only removes the
adapter's Axum 0.7 stack from the production graph.
The Windows wheel for 0.39.0 exceeded PyPI's 100 MiB file-size limit,
preventing the original release upload. Restore the repository's release
profile (`fat` LTO, 1 codegen unit) by removing the Windows-only
`thin`/16 overrides introduced in #3716. Keep `rust-lld` as the linker.
This recovers wheel-size headroom at the cost of the longer fat-LTO
build. The already-published 0.39.0 Windows wheel was recovered
separately by recompressing the original artifact; this change addresses
the build configuration for future releases.
Fixes#4239.
### Windows size comparison
Built the same v0.38.0 source
(`8c68e0c619f2b1febe92e51a29d72c968d24a5c9`) twice on one AWS
`m7i.4xlarge` Windows Server 2025 machine, changing only the
LTO/codegen-unit overrides:
| Artifact | thin LTO / 16 units | fat LTO / 1 unit |
| --- | ---: | ---: |
| Installable Windows wheel | 104,129,447 bytes (99.31 MiB) | 73,875,399
bytes (70.45 MiB) |
| Uncompressed native module | 309,969,408 bytes (295.61 MiB) |
206,609,408 bytes (197.04 MiB) |
The thin/16 configuration increased wheel size by **40.95%** and
native-module size by **50.03%** relative to fat/1. The compressed
native module accounts for all but 3 bytes of the wheel increase.
Both builds used Rust 1.97.0, maturin 1.12.4, Python 3.13.5, MSVC
14.44.35207, Windows SDK 10.0.26100.0, rust-lld, static CRT, default
features, and `maturin build --release --strip --locked --verbose`. They
ran sequentially with separate empty target directories and identical
locked third-party dependencies. The tag's stale workspace package
versions were normalized once before both builds. This controlled pair
used VS2022 Build Tools; the historical GitHub runner used VS2026.
Both wheels passed ZIP/RECORD integrity checks and installed
successfully. Native smoke checks covered import, database creation, row
count, nearest-vector query, and reopening the database. This
establishes the combined configuration effect; it does not isolate LTO
mode from codegen-unit count or establish a runtime-performance
difference.
The figures above are the v0.38.0 reproduction, **not measurements of
this PR head**. The existing PyPI Publish pull-request workflow rebuilds
the current revision without publishing; its result is pending.
Workflow validation passed with `actionlint -shellcheck=`. Full
actionlint reports the same three pre-existing ShellCheck diagnostics in
the unchanged repository-selection step as on `main`.
Support adding column with pyarrow schema in remote client. Previously
this raised an error. This achieves parity with local client.
Server-side implementation is already complete.
`RemoteTable::add_columns` matched only `SqlExpressions` and refused
everything else, so `add_columns(pa.Field | List[pa.Field] | pa.Schema)`
reached a remote table as `NotSupported`. Every layer above was already
in place: the Python and Node surfaces accept a schema, both bindings
build `NewColumnTransform::AllNulls` from it, and the Function-binding
guard already reads that variant's column names. Only the arm that turns
it into a request was missing.
Send the schema as an Arrow IPC schema message under
`application/vnd.apache.arrow.stream`, which is what the server takes.
It is also the only encoding that round-trips a field whole: the JSON
type representations either hide decimal precision in a length field or
drop a timestamp's unit and timezone. With no JSON envelope the branch
rides the query string, as it does for the other binary-bodied
endpoints.
The response handling is now shared by both arms rather than living
inside the SQL one, so the new path gets the same schema-cache
invalidation, write-version tracking, and old-server empty-body
fallback.
Also widen the sync `RemoteTable.add_columns` annotation to match
`Table.add_columns`. It delegates to the async table, so a schema
already worked at runtime; only the signature and its docs disagreed.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Dropping a column a Function binding writes was refused outright, which
left a bad declaration unrecoverable: the binding is immutable, there is
no rebind, and so the column could never be filled again. A drop that
names every output of a binding now retires it instead.
A binding lives in three places -- the envelope in schema metadata, a
`computed_column.*` marker on each output field, and the columns
themselves -- and readers cross-check the first two, so removing one and
leaving the others is a table that refuses every write. The retirement
therefore commits in two steps, each landing a state that stands on its
own: one `UpdateConfig` rewrites the envelope and clears the markers
together, leaving the outputs as ordinary columns holding their last
values, and the drop follows. Interrupted between them, the columns are
still there to be dropped again. Folding the metadata into the drop's
own `Project` was the obvious alternative and does not work: the
transaction proto records `Project` as fields alone, so the edit would
survive only in the writer's memory.
Naming one output of a multi-output binding is refused and names the
missing siblings, since one refresh writes them in one commit. A
multi-output binding's hidden `__function_assignment_*` column goes with
it. Inputs stay protected while a surviving binding reads them, and drop
in the request that retires the last one.
## Summary
- Remove the `CLAUDE.md` agent instruction files at the repo root,
`python/`, and `nodejs/`.
- Claude now reads `AGENTS.md`, and these files were only symlinks to
the existing `AGENTS.md` copies.
## Test plan
- [x] Confirm `AGENTS.md` remains at the repo root, `python/`, and
`nodejs/`
- [x] Confirm no remaining in-repo references to `CLAUDE.md`
Made with [Cursor](https://cursor.com)
Co-authored-by: Cursor <cursoragent@cursor.com>
A view definition was a structured record under a `kind` tag, one kind
per query shape, and a Function returning `list<struct>` was about to
add a third. That names shapes instead of describing a relation.
A materialized view is now a relation defined by a query, stored as one
canonical SQL string under a format number:
```sql
SELECT columns FROM [ns.]table [, function(args) AS alias | , UNNEST(column) AS alias]
[WHERE predicate] [LIMIT n]
```
Any other clause is refused at parse time. Older readers report the view
as unrefreshable, the pre-format layouts still read, and a legacy view
is rewritten on its next refresh that commits, rebuilt only where its
raw text meant something else under lance's parser.
A Function in FROM position yields one row per element it returns, as a
table function does in any dialect. The server stages its list output in
a hidden table and records that binding beside the query, which stays as
the user wrote it; refresh scans the staging and unnests the column, the
same operator as UNNEST over a list column the table already holds. Row
ids repeat per element, so eviction and incremental append are
unchanged. A local database refuses a Function in FROM position, since
it has no executor.
## What
Decode the path component of `file://` image URIs before passing it to
Pillow.
## Why
`Path.as_uri()` percent-encodes characters such as spaces. Passing
`parsed.path` directly to Pillow therefore tries to open a literal `%20`
path and fails.
## Testing
- Added a regression test that opens an image whose local filename
contains a space.
- Verified the focused URI conversion behavior against the changed
method.
- Ruff check, formatting check, and `compileall` on both changed files.
Co-authored-by: Xuanwo <github@xuanwo.io>
Remote materialized-view deletion currently rejects HTTP 200 even when
the server has completed cleanup synchronously. Treat 200 as an
already-finished Job with no ID, matching table deletion; retain the
cleanup Job for 202, require its ID, and invalidate the table cache in
both cases.
Add regressions for synchronous completion and malformed or unexpected
responses, alongside the existing asynchronous Job coverage. No local
builds or tests were run.
Export `MetadataEraserExec` and its constructor from
`lancedb::table::datafusion`. Engines that serialise a physical plan
containing a LanceDB scan have to rebuild the operator outside this
crate, which a private type makes impossible.
A table with one registered Function binding refused every add_columns,
alter_columns, drop_columns and field-metadata update, whatever column
they named. The hazard is narrower: a binding stores the exact Arrow
fields of its inputs, outputs and assignment column, so editing one of
those strands it and the table stops accepting rows. Any other column
was never at risk.
The guard now compares the columns a request names, a rename's target
included, against the set every binding depends on, and refuses only on
an intersection. Local and remote tables apply the same rule. The
blanket guard stays on update and merge insert, which cannot say what
they touch.
A refresh running under a skip policy records each row it skipped, with
the failing input and the error, but the client could not read that
store: the server exposes it over SQL and, since recently, a REST route.
A user who hit per-row failures still had to open a SQL session.
`Table::function_errors` calls the route. The listing is table-addressed
with optional job and column filters, the same addressing the SQL
surface uses, so the two cannot disagree about what a table's errors
are. The two non-record signals come back as their own fields rather
than as rows: capped-fragment summaries, and whether the listing stopped
at its limit. Local tables refuse rather than answer with an empty list.
Python and Node expose the same call, with the same optional filters.
Adds `TypeSafeReranker`, which reranks vector, FTS, and hybrid results
with the [TypeSafe System One
API](https://docs.typesafe.ai/introduction).
Each result is scored independently: TypeSafe reads `{"query",
"document"}` and answers a yes/no (noul) question, and the probability
of yes becomes `_relevance_score`. Unlike listwise LLM rerankers, the
score is an absolute probability, so it is comparable across queries and
can be thresholded. The question's `instructions` and `true`/`false`
`criteria` are configurable, since domain-specific criteria are what
make this kind of scoring work well ([TypeSafe's re-ranking
cookbook](https://docs.typesafe.ai/cookbooks/rerank_typesafe)).
The API takes one state per request, so the reranker sends one request
per result on a thread pool bounded by `max_concurrency`. It
deliberately does not use the background event loop: rerankers are
called synchronously from inside the async query APIs, where `LOOP.run`
would deadlock.
TypeSafe scores for the same pair vary slightly between calls, so
results with close scores can swap places when a search is repeated. The
shared reranker test helper now takes `deterministic=False` for this
case: it still checks result sizes and descending scores, but not that
two identical searches return the same order.
The SDK is imported only when the client is created and questions are
sent as plain dicts, so the new tests run in CI with a fake client and
without `typesafe-sdk` installed. The live-API test is skipped without
`TYPESAFE_API_KEY`. Ranking quality has not been compared with other API
rerankers.
Adds the client half of database-scoped named Secrets: a Secret is a
name and
an opaque value stored by the service, and a Function binds one to the
environment variable its library already reads. Secrets are addressed by
a
namespace path plus a name.
The UDF body is unchanged and stays portable — it reads `OPENAI_API_KEY`
the
way it always did, and the binding is what puts a value there:
```python
db.create_secret("openai-prod", os.environ["OPENAI_API_KEY"])
function = db.create_function(
analyze_caption,
secrets=[
EnvVarSecret(secret_name="openai-prod", env_variable="OPENAI_API_KEY")
],
)
function.secret_bindings # the Secret's name, never its value
```
- `create_secret` / `alter_secret` / `list_secrets` / `describe_secret`
/
`drop_secret` on sync, async and remote connections, with the pyo3
binding
and the Rust client behind them. Each takes `namespace_path`
keyword-only,
defaulting to the root.
- **There is no read API, by construction rather than by policy** — no
code
path returns a stored credential, and `describe_secret` answers with
metadata only.
- `EnvVarSecret` is a pure local constructor: it contacts no server, so
it
cannot fail on a Secret that does not exist. It exists so that a bare
string
in that position — which would be a credential — is a `TypeError` rather
than a plausible-looking mistake that reads identically in a diff.
- `create_function(..., secrets=[...])` carries the bindings as
`secret_bindings`: a list of `SecretBinding` tagged by `kind`, so a
later
delivery mode is a variant rather than a sibling field. The value never
travels — it is resolved by the service when the Function runs, which is
what
lets a rotation reach columns already pinned to an older
FunctionVersion.
- A binding names its Secret as a `SecretReference` of `{name,
namespace_path}`
rather than one joined string, so no delimiter has to be excluded from
every
name and segment forever, and `ClientConfig.id_delimiter` cannot
contradict
an identity built on a fixed separator.
- A root namespace is omitted from the request body rather than sent
empty, so
a root request is byte-identical to one from a client that predates
namespaces. Tests pin it.
This is the client surface the design's §4 describes; the service side
lives in
sophon.
**Previously split across two PRs.** Namespace addressing was #4151,
stacked on
this one; it is folded in here so the Secret identity contract — name,
namespace path, and the binding that carries both — is reviewable as one
piece
rather than as a shape introduced and then replaced.
## Identifier safety, merged from #4189
**#4189 is merged into this branch**, so the client half of Secrets and
the
guards on the identity it puts in the URL are one PR. What it added:
- Components are checked where the identifier is built, before a request
is
constructed. `create_secret("../jobs", value)` no longer resolves to
`/v1/jobs/create` and delivers a credential-bearing body to a route with
none
of this one's body suppression.
- Each component is percent-encoded and joined by the delimiter, so
nothing
inside a component can end the path segment or add one.
- A component may not be empty, a relative segment (`.`, `..`, and their
`%2e`
spellings), or the delimiter itself — the three ways a component erases
a
boundary the split has to recover. `["prod", ""]` joined to `prod$`,
which
reads back as `["prod"]`.
- `$` is the only accepted `id_delimiter`, refused at client
construction.
`ClientConfig.id_delimiter` remains, since the identifier grammar comes
from
the Lance REST catalog standard, but a value that would produce
identifiers no
service splits the caller's way is now an error where it was written.
- One `build_object_identifier` and one character set serve tables,
namespaces,
Secrets, Functions and materialized views.
Components are checked for *addressability*, not a character set: the
name's
own grammar stays each object's own, so a catalog database keeps the `/`
that
`RemoteCatalog::validate_name` allows.
## Known shortcoming
`secret_bindings` is omitted from a registration body when empty, so a
client
that binds nothing sends what a client without bindings sends. When a
client
does bind a Secret and the service does not know the field, the field is
ignored: registration succeeds, the returned version carries no
bindings, and
the Function fails at execution with the variable unset, far from the
call that
asked for it.
`ServerVersion` is how this codebase refuses a feature the service is
too old
for, and it gates five features already. It does not gate this one: it
is held
per table, and registering a Function is a database-level call. Noted at
the
field in `remote/db.rs`; wiring the gate is follow-up work.
**Tests:** lancedb lib 1340 passed, `first_class_function_slice1` 9,
`first_class_function_slice2` 3, plus Python tests across both slices.
Rebased onto `main` after #4176 (OCI Function identity), #4191 (`.`/`..`
table
names) and #4195 (remote catalogs).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_01UfmeJ533rQDnPBkMtjerV6
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
ZoneMap indexes were added to the Rust API in #4199, but LanceDB reused
the BTree type validation when creating them. That made the public
builder reject some types that Lance ZoneMap can support, including
`LargeUtf8`, `Binary`, and `LargeBinary`. This PR gives ZoneMap its own
validation helper so it can accept the broader scalar set while keeping
the rest of the create-index path unchanged.
Lance supports ZoneMap scalar indexes, but the LanceDB Rust API did not
expose a first-class way to request one through `Table::create_index`.
Users had builders for the other scalar index families, while ZoneMap
was missing from the public `Index` model and remote create-index
serialization. This PR adds ZoneMap as a supported scalar index option
in LanceDB.
This was accomplished with the following changes:
- Added `ZoneMapIndexBuilder` in `rust/lancedb/src/index/scalar.rs`.
- Added `Index::ZoneMap` and `IndexType::ZoneMap`, including
display/from-string aliases for `ZONEMAP` and `ZONE_MAP`.
- Mapped local index creation to
`ScalarIndexParams::for_builtin(BuiltinIndexType::ZoneMap)` and Lance
`IndexType::ZoneMap` in `rust/lancedb/src/table/create_index.rs`.
- Serialized remote create-index requests as `index_type: "ZONEMAP"` in
`rust/lancedb/src/remote/table.rs`.
- Added coverage for both local ZoneMap index creation and remote
request serialization.
Example:
```rust
table
.create_index(&["my_column"], Index::ZoneMap(Default::default()))
.execute()
.await?;
```
### Testing
Added `test_create_zonemap_index` for local index creation and extended
the remote request body test matrix for `ZONEMAP`.
Add a `Catalog` trait and `RemoteCatalog` for managing databases,
exposed through Rust, synchronous/asynchronous Python, and TypeScript. A
remote catalog represents the server's root namespace, and each database
is one child namespace. Create/connect return ordinary LanceDB
connections, so existing table APIs work unchanged.
## Rust API
`Catalog` is an object-safe async trait with `create_database`,
`connect_database`, `list_databases`, and `drop_database`. Backend
create/connect methods return `Arc<dyn Database>`; the public
`CatalogConnection` wraps them as `Connection` values and shares its
embedding registry with those connections. `RemoteCatalog` implements
the trait; `connect_catalog` is the convenience builder, available with
the `remote` feature.
```rust
use lancedb::catalog::{
CreateDatabaseRequest, DropDatabaseRequest, ListDatabasesRequest,
};
let catalog = lancedb::connect_catalog("https://my-server.example")
.api_key("my-api-key")
.execute()
.await?;
let db = catalog.create_database(
CreateDatabaseRequest::new("analytics").exist_ok(true),
).await?;
let connected = catalog.connect_database("analytics").await?;
let page = catalog.list_databases(
ListDatabasesRequest::default().limit(20),
).await?;
// page.databases: Vec<String>; page.page_token: Option<String>
catalog.drop_database(
DropDatabaseRequest::new("analytics").ignore_missing(true),
).await?;
```
Create/drop also accept a plain name for default behavior, e.g.
`catalog.create_database("analytics").await?`. Existing names fail
creation unless `exist_ok` is enabled; missing names fail drop unless
`ignore_missing` is enabled. Drop always requires an empty database.
## Python API
```python
import lancedb
catalog = lancedb.connect_catalog(
"https://my-server.example", api_key="my-api-key"
)
db = catalog.create_database("analytics", exist_ok=True)
connected = catalog.connect_database("analytics")
page = catalog.list_databases(limit=20)
# page.databases: list[str]; page.page_token: Optional[str]
if page.page_token is not None:
next_page = catalog.list_databases(limit=20, page_token=page.page_token)
catalog.drop_database("analytics", ignore_missing=True)
```
`connect_catalog` returns `Catalog`; create/connect return the existing
`DBConnection` API. The async equivalent is `catalog = await
lancedb.connect_catalog_async(...)`, returning `AsyncCatalog`; await
each of the same four methods, with create/connect returning
`AsyncConnection`.
## TypeScript API
```typescript
import { connectCatalog } from "@lancedb/lancedb";
const catalog = await connectCatalog("https://my-server.example", {
apiKey: "my-api-key",
});
const db = await catalog.createDatabase("analytics", { existOk: true });
const connected = await catalog.connectDatabase("analytics");
const page = await catalog.listDatabases({ limit: 20 });
// page.databases: string[]; page.pageToken?: string
if (page.pageToken !== undefined) {
const nextPage = await catalog.listDatabases({
limit: 20, pageToken: page.pageToken,
});
}
await catalog.dropDatabase("analytics", { ignoreMissing: true });
```
Create/connect return the existing `Connection` API. All four methods
are asynchronous.
## REST mapping
All paths below are relative to the catalog endpoint. `{name}` is the
logical database name encoded as one URL path component. The default
namespace delimiter is `$`, so the root identifier is encoded as `%24`.
| Catalog operation | Existing REST route | Request |
| --- | --- | --- |
| `create_database(name)` | `POST /v1/namespace/{name}/create` |
`{"mode":"Create"}`; `exist_ok=true` sends `{"mode":"ExistOk"}` |
| `connect_database(name)` | `POST /v1/namespace/{name}/describe` |
`{}`; verifies existence before returning a scoped connection |
| `list_databases(...)` | `GET /v1/namespace/%24/list` | Optional
`limit` and `page_token` query parameters |
| `drop_database(name)` | `POST /v1/namespace/{name}/drop` |
`{"mode":"Fail","behavior":"Restrict"}`; `ignore_missing=true` changes
mode to `"Skip"` |
For example, database `team/search` uses
`/v1/namespace/team%2Fsearch/create`. A paginated root listing can use
`/v1/namespace/%24/list?limit=20&page_token=a%2Fb`. The list response
retains the existing namespace wire shape,
`{"namespaces":["analytics"],"page_token":"next"}`; the SDK exposes
`namespaces` as `databases` and preserves the opaque continuation token.
An absent or empty token ends pagination. Page limits must be between 1
and 2147483647. Create/drop accept a namespace JSON response or HTTP
204.
Catalog management requests omit both `x-lancedb-database` and
`x-lancedb-database-prefix`, including values supplied through static or
dynamic headers. Returned database connections set `x-lancedb-database`
to the exact logical name and keep independent scope. API keys, OAuth or
dynamic authentication, client settings, table read consistency
settings, and an optional SQL endpoint override carry over to those
connections. OAuth cannot be combined with an API key or a custom header
provider.
For SQL through an HTTPS catalog, configure the existing SQL endpoint
contract with Rust
`.sql_host_override("grpc+tls://sql.example.com:10026")` or Python
`sql_host_override="grpc+tls://sql.example.com:10026"`. TypeScript
catalog options expose the same setting as `sqlHostOverride`. It is
inherited by created/connected databases, retained by Python connection
serialization, and initialized lazily when SQL is executed.
Create HTTP 409 maps to `DatabaseAlreadyExists`; connect/drop HTTP 404
maps to `DatabaseNotFound`, except that `ignore_missing` suppresses a
missing-database drop. Other server errors propagate. The server
enforces restricted deletion; the client never requests cascading
deletion.
Database names preserve literal slashes as part of one name. They must
be nonempty ASCII, with no control characters, surrounding whitespace,
or configured namespace delimiter, and cannot be `.` or `..`. Endpoints
must be HTTP(S) URLs without embedded credentials, query parameters, or
fragments.
## Scope
This PR adds the client API and reuses existing namespace endpoints.
Local catalogs, `__catalog` storage, location generation/sanitization,
and `__manifest` lifecycle support remain deferred; the Lance dependency
is unchanged.
The PR also runs macOS Node tests serially to avoid existing
resource-contention timeouts reproduced across recent main runs.
---------
Co-authored-by: Xuanwo <github@xuanwo.io>
Passing an invalid table name to `open_table` or `create_table` panics
instead of returning an error:
thread '...' panicked at rust/lancedb/src/database/listing.rs:1155:62:
called `Result::unwrap()` on an `Err` value: InvalidTableName { name:
"my table", ... }
Both call sites build the table URI with
`request.location.clone().unwrap_or_else(||
self.table_uri(&request.name).unwrap())`, and `table_uri` is the
function that validates the name — so every rejected name (empty,
spaces, slashes, non-ASCII) hits the inner `unwrap`.
`Error::InvalidTableName` clearly is the intended contract here: the
variant exists for exactly this, and the Python binding maps it to
`ValueError`.
Replaced the closure with a `match` that propagates the validation
error; behavior with an explicit `location` is unchanged (the name is
not validated on that path, as before). Added tests asserting
`InvalidTableName` for `create_table` and `open_table` over a set of
rejected names — both panic without the fix. Full `cargo test -p lancedb
--lib --features remote`: 1214 passed; clippy/fmt clean; `cargo check
--workspace --all-targets` clean.
Co-authored-by: Xuanwo <github@xuanwo.io>
Stacked on #4173 (`jack/restore-oidc-flows`, base branch mirrored to
this repo so the diff shows only this change); context from review:
https://github.com/lancedb/lancedb/pull/4173#issuecomment-5674048100.
Rebase to `main` once #4173 and #4179 merge.
## What moved to the `oauth2` crate (5.0, no default features)
- Authorization URL generation and CSRF state (`authorize_url`,
`CsrfToken`)
- PKCE S256 challenge/verifier generation and code exchange
- Client-credentials, authorization-code, refresh-token, and device-code
grant request construction
- Device authorization request and the device token polling loop
(`authorization_pending`, `slow_down` +5s, expiry deadline, denial,
network backoff capped at 10s)
- Standard success/error response parsing (`RequestTokenError`)
- Token-endpoint client authentication and standards-compliant parameter
encoding (RFC 6749 2.3.1 Basic encoding)
LanceDB keeps ownership of OIDC discovery (compared `openidconnect`: no
measurable win for our 3-field metadata + strict validation, at real
dependency cost), HTTPS-or-loopback endpoint enforcement, the loopback
callback server, browser/stderr prompts, token caching and refresh
orchestration, and the dedicated hardened Azure IMDS source, which is
unchanged.
## Client authentication methods
New `ClientAuthMethod` enum (`none` | `client_secret_basic` |
`client_secret_post`), exposed in Rust, Python, and Node. Unset resolves
to `client_secret_basic` when a secret is present (RFC 6749 2.3.1
recommendation and the normal Okta confidential-app default, so a
default Okta app works without weakening its configuration) and to
`none` for public clients (PKCE/device). Explicit `none` with a secret,
or basic/post without one, is rejected. The method applies to client
credentials, code exchange, refresh, and device requests. Deliberate
behavior change: confidential clients previously always sent the secret
in the POST body; they now default to Basic (Keycloak accepts both).
No `audience`/`resource` parameters were added: the supported target is
an Okta custom authorization server with the API audience configured
server-side, so client-provided audience parameters are unnecessary;
`add_extra_param` support exists if a concrete provider contract ever
needs them.
## Device polling behavior changes (deliberate, tested)
- The first token poll now happens immediately rather than after one
interval (RFC 8628 allows both).
- Transient failures (HTTP 429, 5xx, `temporarily_unavailable`, network
errors) now retry with exponential backoff capped at 10s instead of
retrying at the fixed interval; polling never spins faster than once per
second even if a server reports a zero interval.
## Security and compatibility
- Issuer and discovered endpoints (and device verification URIs) still
require HTTPS except explicit loopback HTTP, enforced before any crate
URL type is built
- Token HTTP client keeps the hardened redirect policy that refuses
insecure redirect targets; regression test added
- Errors never embed raw response bodies (avoids leaking tokens through
parse failures); all credential types stay redacted in Debug
- Transient conditions (429/5xx/`temporarily_unavailable`) remain
retryable in device polling and hard errors elsewhere; refresh keeps
rotation and reauthentication semantics
- Existing public APIs stay source-compatible except the added
`OAuthConfig.client_auth_method` field
## Tests
Rust: client-auth methods across code
exchange/refresh/client-credentials/device (none/basic/post),
auth-method resolution and validation, transient device retries,
denial/expiry, redirect rejection, malformed-response leak check, PKCE
URL assertions, redaction. Python and Node: enum values, conversion,
unknown-method errors, config defaults.
Manual Okta validation recipe (no automated Okta credentials): create a
custom authorization server with an API audience, one confidential web
app (Basic) for authorization-code, one native app (PKCE, no secret),
one native device app; point `issuer_url` at the custom server, set
`client_auth_method` only for the POST-required case; verify token
acquisition, refresh after expiry, and `x-lancedb-credential-type: oidc`
against a LanceDB deployment. Never commit tenant URLs or secrets.
Co-authored-by: Xuanwo <github@xuanwo.io>
Update the Rust workspace and Java Lance dependencies to
[v13.0.0-beta.3](https://github.com/lance-format/lance/releases/tag/v13.0.0-beta.3),
which includes the nullable-list encoding fix in
lance-format/lance#9268. Also resolve existing strict Clippy diagnostics
in crate-private OAuth modules and the Python token-cache conversion,
without changing effective visibility.
Validated the rebuilt Python SDK with nullable and nested lists:
Match/Phrase searches preserve document coordinates before and after
appending data and optimizing the table.
OAuth tests set `LANCEDB_OAUTH_BROWSER` inside Jest's sandbox, which
does not update the process environment read by Rust. As a result, the
native login flow launches a real browser; the [failing macOS main
job](https://github.com/lancedb/lancedb/actions/runs/35067074667/job/104699667840)
ends by terminating an orphaned Safari process.
Set the no-op browser helper while loading Jest configuration, before
creating sandboxes and workers, so both local and CI tests inherit it.
Check the inherited process environment before OAuth login to catch
regressions without opening a browser. Windows uses a no-op command
fixture. Test deadlines and worker counts are unchanged.
A negative control restoring the sandbox-only assignment fails in the
new pre-login check. The full macOS test suite passes with the standard
test launcher; the Windows helper has not been executed locally. The
[first hosted macOS
run](https://github.com/lancedb/lancedb/actions/runs/35071525843/job/104713928139)
passes all 843 tests (5 skipped), with no Safari process in the job log.
Further normal runs are needed to establish sustained stability.