Commit Graph
3034 Commits
Author SHA1 Message Date
Lance Release df7bfcdb65 Bump version: 0.40.0-beta.7 → 0.40.0-beta.8 2026-09-23 17:05:35 +00:00
Xuanwo a9e350c158 ci(python): build Windows release wheel on 8-core runner (#4264)
Since #4241 restored fat LTO / 1 codegen unit for the Windows wheel, the
`windows` job in PyPI Publish has been cancelled at its 90-minute
timeout on every run, including the v0.40.0-beta.6 and v0.40.0-beta.7
tags, so neither release was published. In those runs dependency
compilation finishes after ~26 minutes and the `lancedb` crate compile +
fat-LTO step was still running after 63 minutes.

Move the job to the org's `windows-2025-8x-x64` larger runner and raise
the timeout to 150 minutes. Fat LTO stays, since the thin-LTO wheel
exceeds PyPI's 100 MiB limit.

The final fat-LTO step is largely single-threaded, so the gain from 8
cores is mainly in the parallel dependency phase plus faster/larger
hardware; the PR run of this workflow is the first measurement of fat
LTO on this runner.
2026-09-24 01:04:18 +08:00
Xuanwo b944055b63 feat(python): class Functions with initialization and shipped helper code (#4254)
Stacked on #4253.

Remote Python Functions could not share helper code, take typed
parameters, or keep per-process state: helpers became `from <module>
import ...` lines a worker cannot resolve, every parameter was an `env=`
string, and state had to be stashed on imported modules. MMLB's OpenAI
preset shows the cost — 24 env vars, a 302-line callable, and a rate
limiter duplicated between the scalar and batch variants, one of which
never constructed it (ENT-2516).

**Class Functions.** `@udf` accepts a class. Each remote instance runs
`__init__` once, calls `__call__` for every row or batch it processes
(row or batch mode is inferred from annotations as before), and calls
`close()` once if defined. Calling the definition locally constructs the
class, so unit tests stay ordinary.

**Initialization.** The annotated `__init__` parameters become the
Function's initialization fields. A binding passes their values beside
its column inputs: `fn(text=col("body"), model="small",
dimensions=512)`. Values are constants of that binding; another column
can bind the same Function version with other values. Types are limited
to booleans, integers, floats, strings, and lists or structs of them,
which have one unambiguous JSON and SQL spelling. A parameter with a
default may be omitted; a null value then takes the default.
Initialization and input names must be disjoint because both are keyword
arguments of one call. Secrets stay on `EnvVarSecret`.

**Helper code.** `code=[package, ...]` ships top-level modules or
packages as Python source in a `python_bundle` artifact (a canonical
JSON object of path → source), imported normally on the worker. Changing
a helper changes the artifact digest and so the Function version; the
environment from `pip`/`conda` is reused. Functions and classes defined
in `__main__` (notebook cells, scripts) are packaged by source,
recursively and dependencies first. An import of a module that lives in
a local source tree, is not shipped with `code=`, and names no declared
package is now rejected at registration instead of failing on the
worker. Closures, lambdas, and nested definitions are still rejected,
with the reason.

A Function without these features packages byte-identically to before
(`python_callable`, same digest, no `initialization` on the wire). Class
sources are located through a method's code object, because
`inspect.getsource` cannot find a class defined in a notebook cell or
doctest.

The packaging contract is documented on `udf`. Sophon changes that build
and execute these artifacts are in lancedb/sophon; they were verified
end to end on a Linux local server (registration → REST, SQL, and
materialized-view bindings with different initialization → refresh →
query, one construction per instance).
2026-09-24 00:56:52 +08:00
Wyatt Alt 5538bdb9ff feat: refresh an ivf-grouped materialized view in units (#4232)
A grouped view's refresh runs its whole aggregate in one process, so a
view grouped by ivf_partition(col) over a large table is bounded by a
single worker however many workers a deployment has. The IVF index
already holds each partition's row ids, so the aggregate splits along
the index without a shuffle: every partition's groups are computed from
its own rows alone.

plan_grouped_refresh names the units (one per index partition plus one
for the rows the index cannot place) and the source version they read;
write_grouped_unit computes one unit's groups from the index's row ids,
plus the rows of fragments the index has not covered, assigned in place,
and writes them as uncommitted fragments; commit_grouped_refresh
replaces the view's rows with every unit's fragments in one Update.

The split has to publish what the single-pass refresh would. A unit
applies the view's predicate in its own query, because neither the
indexed take nor the fragment scan filters the way a lance scan does. An
index that does not say which fragments it covers has unknown coverage,
not empty, so it yields no plan at all and the caller refreshes in one
pass rather than reading those rows twice. Each result names its unit,
its plan and the view incarnation it was computed for -- two views of
one shape reach the same counters, and fragments written into one
dataset are not publishable into another -- and the commit publishes the
plan's units exactly once each or nothing. The commit lands on the
planned generation or is refused: lance rebases this Update over a
concurrent append rather than rejecting it, so the version it actually
landed on is checked, as the single-pass rebuild already does.

What holds every unit to one index is the source version the plan pins:
indices live in the source manifest, so a rebuild lands in a version the
units never read. Within that version a segment's postings can still
outlive its ownership -- a column rewrite attaches a new file and takes
the fragment out of the segment's bitmap without dropping its rows from
the posting lists -- so a unit keeps a segment's rows only while it
holds their fragment, and reads the rest from the scan.

The split is the grouping only where the index assigns by its own
centroids. lance also builds an index from precomputed partitions, and
records nowhere that it did, so a posting list can hold a row that
ivf_partition puts elsewhere -- that row's group would then be
aggregated in its own unit as well and published twice, since
concatenated fragments cannot merge two halves of a group. The plan
samples each partition and yields no units when they disagree, every
unit proves the rows it took before grouping them, and the commit
refuses a unit that wrote more than the single group its key allows.
2026-09-23 09:13:17 -07:00
Xuanwo 3728d41f02 feat: carry Function initialization in the canonical model (#4253)
A Function instance is created once with an initialization row
(`create(initialization, context)` in the Function Format), but nothing
in the client model could say what that row holds or where its values
come from, so every source-built Function ran with an empty row and
users pushed configuration through `env=` strings instead (ENT-2516).

This adds initialization to the canonical wire model without changing a
Function's identity:

- `FunctionSignature.initialization` lists the fields of the row, in
order. It is omitted when empty, so existing signatures and
FunctionVersion hashes are unchanged.
- `FunctionApplication.initialization` carries a binding's values as
JSON. The service validates them against the Function's fields and
re-encodes them as Arrow, so this JSON is transport only and never
hashed; that is why floats are accepted here while column-input literals
keep the Slice 1 domain.
- `FunctionBinding.initialization` is the validated one-row Arrow IPC
stream (base64) that every instance of the binding is created with.
Values belong to the binding, so one Function version can back several
columns with different values.
- The binding-shape check that guards schema mutations accepts the new
field; without it a table holding an initialized binding refused further
declarations.

Rust and Python share new canonical goldens for an initialized
application and binding.

The authoring side (class `@udf`, initialization arguments, shipped
helper code) is the stacked PR on top of this one.
2026-09-23 23:31:49 +08:00
Lance Release f359c5c8c5 Bump version: 0.40.0-beta.6 → 0.40.0-beta.7 2026-09-23 07:28:56 +00:00
Jack Ye 8541d6d9df feat: add view CRUD APIs (#4236)
A view is a named query a database stores and plans on every read. It
holds no rows, which is the whole difference from a materialized view.

## API

| Verb | Route |
| --- | --- |
| `create_view(name, query, namespace_path)` | `POST
/v1/view/{id}/create` |
| `describe_view(name, namespace_path)` | `POST /v1/view/{id}/describe`
|
| `drop_view(name, namespace_path)` | `POST /v1/view/{id}/drop` |
| `list_views(namespace_path)` | `GET /v1/namespace/{id}/view/list` |

On `Connection` and the `Database` trait, with the remote client, Python
(sync and async) and Node bindings. Local databases return
`NotSupported`: the server side is Sophon's, where a view is an object
of the database manifest.

`ViewDescription` carries the defining query, the database *and
namespace path* unqualified names in it resolve against, and the schema
the query resolved to. `create_view` returns one, so a caller has the
schema without a second call.

Both defaults travel with the view because it outlives the session that
declared it: the server re-plans the stored query on every read, so a
reader resolving an unqualified name against its own defaults would read
a different table. `default_namespace_path` crosses the wire as
`default_namespace`, a path like `namespace`, absent for the root.

There is no replace: a name already taken is an error, and changing a
view is a drop followed by a create, each authorized against what it
actually touches.

Querying a view stays SQL's job. There are no rows behind a view, so
there is no `open_view` returning a `Table`.
2026-09-23 15:09:40 +08:00
Rudra Prasad Bhuyan dd2539c2ca fix(python): convert objects for JSON fields (#4211)
## Summary

Closes: #4060 

  - Convert Python dict/list objects to JSON strings during ingestion.
  - Support JSON fields nested inside structs and  lists.
  - Cover `add` and `merge_insert` JSON ingestion paths.
  - Add regression coverage for nested struct and list JSON fields.

  ## Testing

  - `git diff --check`
  - Python compilation passed.
- Focused pytest was attempted but could not complete because the native
extension build stalled during `uv` bootstrap.
2026-09-22 14:06:11 -07:00
Lance Release a6a6617ee2 Bump version: 0.40.0-beta.5 → 0.40.0-beta.6 2026-09-22 18:00:27 +00:00
Wyatt Alt e020e13744 feat: group a materialized view by the IVF partition of a vector column (#4223)
Grouping near neighbours together, so an all-pairs comparison runs per
bucket instead of over the whole table, needs the partition an IVF index
assigns each vector.

`ivf_partition(column)` returns that partition, assigned by the
centroids and distance type of the IVF index on the column. It is bound
when the view is declared and again at every refresh; an index retrain
commits a new source version, so the next refresh regroups. A column
without an IVF index, or with two, is refused.


Stacked on #4222.
2026-09-22 10:16:20 -07:00
Jack Ye 814de30c5a feat: return a cleanup job from drop_function (#4237)
Dropping a Function can leave its content to a server-side cleanup job,
so the name drop and the content deletion become separate events a
caller may want to wait on.

`drop_function_async` returns the unbind result alongside a `Job` for
the cleanup, the same shape `drop_table_async` and
`drop_materialized_view_async` already use: a `202` carries the job id,
and a `200` — an inline deletion, or a name that was not bound — yields
an already-finished job with no id. A `202` without a usable job id is
rejected rather than silently reported as finished.

`drop_function` keeps its `bool` result and now delegates, so nothing
changes for callers that do not care when the content goes.

Available on `Connection` in Rust and on both the sync and asyncio
Python connections.
2026-09-22 09:54:23 -07:00
Wyatt Alt 254688df96 feat: group a materialized view by GROUP BY (#4222)
A view could only map source rows one to one, or one to many through a
FROM-position item, so nothing could see all the rows sharing a key.

A view query now takes `GROUP BY expr, ...` with aggregate projections,
planned and executed by DataFusion over the lance scan. A group spans
fragments, so a grouped view is recomputed in full whenever its source
changes; each row's provenance is its group's smallest source row id.
Such a query is stored as definition format 2, so a reader that predates
grouping reports the view as unrefreshable instead of failing to parse
it.


Stacked on #4190.
2026-09-22 08:03:27 -07:00
Xuanwo 518d7ff1dd fix: keep REST server adapter out of production dependencies (#4240)
The production `remote` client talks to an existing namespace over HTTP.
It needs `lance-namespace-impls/rest`, not the REST *server* adapter.
`RestAdapter` is only used by `cfg(test)` integration tests that stand
up an in-process server. Enabling `rest-adapter` on the production
`remote` feature pulled that server stack, including Axum 0.7, into
default Python wheels.

This keeps `rest` on `remote` and moves `rest-adapter` to a
`lance-namespace-impls` dev-dependency so those tests still compile and
run. Cargo resolver=2 does not leak the extra feature into
`lancedb-python`.

## Measurement

Paired `maturin build --release --strip --target aarch64-apple-darwin
--features fp16kernels` wheels. Source was `3be29228` plus this
Cargo.toml change, which is this PR's tree (`878b2ae5` on `3be29228`).
Same toolchain, profile, and packaging flags; only `rest-adapter` moved.

| Artifact | Before | After | Delta |
| --- | ---: | ---: | ---: |
| Compressed wheel | 64,685,664 | 63,824,281 | −861,383 (−1.33%) |
| `_lancedb.abi3.so` uncompressed | 148,255,440 | 146,275,936 |
−1,979,504 (−1.34%) |

This is macOS arm64, not Windows. It does not resolve the Windows wheel
upload limit. Tonic still pulls Axum 0.8; the change only removes the
adapter's Axum 0.7 stack from the production graph.
2026-09-22 15:13:51 +08:00
Xuanwo ff757b74fd fix(python): reduce Windows release wheel size (#4241)
The Windows wheel for 0.39.0 exceeded PyPI's 100 MiB file-size limit,
preventing the original release upload. Restore the repository's release
profile (`fat` LTO, 1 codegen unit) by removing the Windows-only
`thin`/16 overrides introduced in #3716. Keep `rust-lld` as the linker.

This recovers wheel-size headroom at the cost of the longer fat-LTO
build. The already-published 0.39.0 Windows wheel was recovered
separately by recompressing the original artifact; this change addresses
the build configuration for future releases.

Fixes #4239.

### Windows size comparison

Built the same v0.38.0 source
(`8c68e0c619f2b1febe92e51a29d72c968d24a5c9`) twice on one AWS
`m7i.4xlarge` Windows Server 2025 machine, changing only the
LTO/codegen-unit overrides:

| Artifact | thin LTO / 16 units | fat LTO / 1 unit |
| --- | ---: | ---: |
| Installable Windows wheel | 104,129,447 bytes (99.31 MiB) | 73,875,399
bytes (70.45 MiB) |
| Uncompressed native module | 309,969,408 bytes (295.61 MiB) |
206,609,408 bytes (197.04 MiB) |

The thin/16 configuration increased wheel size by **40.95%** and
native-module size by **50.03%** relative to fat/1. The compressed
native module accounts for all but 3 bytes of the wheel increase.

Both builds used Rust 1.97.0, maturin 1.12.4, Python 3.13.5, MSVC
14.44.35207, Windows SDK 10.0.26100.0, rust-lld, static CRT, default
features, and `maturin build --release --strip --locked --verbose`. They
ran sequentially with separate empty target directories and identical
locked third-party dependencies. The tag's stale workspace package
versions were normalized once before both builds. This controlled pair
used VS2022 Build Tools; the historical GitHub runner used VS2026.

Both wheels passed ZIP/RECORD integrity checks and installed
successfully. Native smoke checks covered import, database creation, row
count, nearest-vector query, and reopening the database. This
establishes the combined configuration effect; it does not isolate LTO
mode from codegen-unit count or establish a runtime-performance
difference.

The figures above are the v0.38.0 reproduction, **not measurements of
this PR head**. The existing PyPI Publish pull-request workflow rebuilds
the current revision without publishing; its result is pending.

Workflow validation passed with `actionlint -shellcheck=`. Full
actionlint reports the same three pre-existing ShellCheck diagnostics in
the unchanged repository-selection step as on `main`.
2026-09-22 14:59:52 +08:00
Ryan GreenandClaude Opus 5 bd94fc3572 feat(remote): support add_columns with schema in remote client (#4244)
Support adding column with pyarrow schema in remote client. Previously
this raised an error. This achieves parity with local client.
Server-side implementation is already complete.

`RemoteTable::add_columns` matched only `SqlExpressions` and refused
everything else, so `add_columns(pa.Field | List[pa.Field] | pa.Schema)`
reached a remote table as `NotSupported`. Every layer above was already
in place: the Python and Node surfaces accept a schema, both bindings
build `NewColumnTransform::AllNulls` from it, and the Function-binding
guard already reads that variant's column names. Only the arm that turns
it into a request was missing.

Send the schema as an Arrow IPC schema message under
`application/vnd.apache.arrow.stream`, which is what the server takes.
It is also the only encoding that round-trips a field whole: the JSON
type representations either hide decimal precision in a length field or
drop a timestamp's unit and timezone. With no JSON envelope the branch
rides the query string, as it does for the other binary-bodied
endpoints.

The response handling is now shared by both arms rather than living
inside the SQL one, so the new path gets the same schema-cache
invalidation, write-version tracking, and old-server empty-body
fallback.

Also widen the sync `RemoteTable.add_columns` annotation to match
`Table.add_columns`. It delegates to the async table, so a schema
already worked at runtime; only the signature and its docs disagreed.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 15:37:24 -07:00
Lance Release a8aa19f513 Bump version: 0.40.0-beta.4 → 0.40.0-beta.5 2026-09-21 18:53:12 +00:00
Wyatt Alt df2e785df7 feat: retire a Function binding when its output columns are dropped (#4231)
Dropping a column a Function binding writes was refused outright, which
left a bad declaration unrecoverable: the binding is immutable, there is
no rebind, and so the column could never be filled again. A drop that
names every output of a binding now retires it instead.

A binding lives in three places -- the envelope in schema metadata, a
`computed_column.*` marker on each output field, and the columns
themselves -- and readers cross-check the first two, so removing one and
leaving the others is a table that refuses every write. The retirement
therefore commits in two steps, each landing a state that stands on its
own: one `UpdateConfig` rewrites the envelope and clears the markers
together, leaving the outputs as ordinary columns holding their last
values, and the drop follows. Interrupted between them, the columns are
still there to be dropped again. Folding the metadata into the drop's
own `Project` was the obvious alternative and does not work: the
transaction proto records `Project` as fields alone, so the edit would
survive only in the writer's memory.

Naming one output of a multi-output binding is refused and names the
missing siblings, since one refresh writes them in one commit. A
multi-output binding's hidden `__function_assignment_*` column goes with
it. Inputs stay protected while a surviving binding reads them, and drop
in the request that retires the last one.
2026-09-21 11:51:44 -07:00
LanceDB Robot 7e171f3abb chore: update lance dependency to v13.0.0-beta.8 (#4242)
Updates the Rust workspace Lance dependencies and Java lance-core to
[v13.0.0-beta.8](https://github.com/lance-format/lance/releases/tag/v13.0.0-beta.8),
refreshing Cargo.lock including the required goosefs-sdk 0.2.2 update.
No source compatibility changes were required.

Validation passed: `cargo clippy --quiet --workspace --tests
--all-features -- -D warnings`, `cargo fmt --all --quiet`, and `git diff
--check`.
2026-09-21 11:08:11 -07:00
XuanwoandCursor 3be29228e4 chore: remove CLAUDE.md now that Claude reads AGENTS.md (#4238)
## Summary
- Remove the `CLAUDE.md` agent instruction files at the repo root,
`python/`, and `nodejs/`.
- Claude now reads `AGENTS.md`, and these files were only symlinks to
the existing `AGENTS.md` copies.

## Test plan
- [x] Confirm `AGENTS.md` remains at the repo root, `python/`, and
`nodejs/`
- [x] Confirm no remaining in-repo references to `CLAUDE.md`


Made with [Cursor](https://cursor.com)

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-21 20:33:50 +08:00
Lance Release 73346c0a2c Bump version: 0.40.0-beta.3 → 0.40.0-beta.4 2026-09-20 08:28:58 +00:00
LanceDB Robot 54efd4949c chore: update lance dependency to v13.0.0-beta.7 (#4234)
Updates the Rust workspace Lance dependencies, Cargo lockfile, and Java
lance-core from v13.0.0-beta.6 to
[v13.0.0-beta.7](https://github.com/lance-format/lance/releases/tag/v13.0.0-beta.7);
no compatibility fixes are required.

Validation: `cargo clippy --quiet --workspace --tests --all-features --
-D warnings`, `cargo fmt --all --quiet`, and `git diff --check` passed.
2026-09-20 01:27:13 -07:00
LanceDB Robot df5709efd8 chore: update lance dependency to v13.0.0-beta.6 (#4224)
Updates the Rust workspace Lance dependencies, Cargo lockfile, and Java
lance-core from 13.0.0-beta.4 to
[v13.0.0-beta.6](https://github.com/lance-format/lance/releases/tag/v13.0.0-beta.6).
Fixes the redundant visibility qualifier on the internal identifier
delimiter constant reported by Clippy.

Validation: `cargo clippy --quiet --workspace --tests --all-features --
-D warnings`, `cargo fmt --all --quiet`, and `git diff --check`.
2026-09-18 15:00:33 -05:00
Wyatt Alt 01ee01dbc8 feat: define a materialized view by its query, with Functions in FROM position (#4190)
A view definition was a structured record under a `kind` tag, one kind
per query shape, and a Function returning `list<struct>` was about to
add a third. That names shapes instead of describing a relation.

A materialized view is now a relation defined by a query, stored as one
canonical SQL string under a format number:

```sql
SELECT columns FROM [ns.]table [, function(args) AS alias | , UNNEST(column) AS alias]
[WHERE predicate] [LIMIT n]
```

Any other clause is refused at parse time. Older readers report the view
as unrefreshable, the pre-format layouts still read, and a legacy view
is rewritten on its next refresh that commits, rebuilt only where its
raw text meant something else under lance's parser.

A Function in FROM position yields one row per element it returns, as a
table function does in any dialect. The server stages its list output in
a hidden table and records that binding beside the query, which stays as
the user wrote it; refresh scans the staging and unnests the column, the
same operator as UNNEST over a list column the table already holds. Row
ids repeat per element, so eviction and incremental append are
unchanged. A local database refuses a Function in FROM position, since
it has no executor.
2026-09-18 12:16:25 -07:00
Hongzhu YiandXuanwo 7955c50929 fix(python): decode file URIs in OpenCLIP (#4131)
## What

Decode the path component of `file://` image URIs before passing it to
Pillow.

## Why

`Path.as_uri()` percent-encodes characters such as spaces. Passing
`parsed.path` directly to Pillow therefore tries to open a literal `%20`
path and fails.

## Testing

- Added a regression test that opens an image whose local filename
contains a space.
- Verified the focused URI conversion behavior against the changed
method.
- Ruff check, formatting check, and `compileall` on both changed files.

Co-authored-by: Xuanwo <github@xuanwo.io>
2026-09-18 09:18:39 -07:00
Lance Release 63cd121225 Bump version: 0.40.0-beta.2 → 0.40.0-beta.3 2026-09-17 20:06:31 +00:00
Jack Ye 97de3ccd98 fix: accept synchronous remote materialized-view drops (#4212)
Remote materialized-view deletion currently rejects HTTP 200 even when
the server has completed cleanup synchronously. Treat 200 as an
already-finished Job with no ID, matching table deletion; retain the
cleanup Job for 202, require its ID, and invalidate the table cache in
both cases.

Add regressions for synchronous completion and malformed or unexpected
responses, alongside the existing asynchronous Job coverage. No local
builds or tests were run.
2026-09-17 13:05:00 -07:00
Vivek a0c5612aee feat(rust): make MetadataEraserExec public (#4213)
Export `MetadataEraserExec` and its constructor from
`lancedb::table::datafusion`. Engines that serialise a physical plan
containing a LanceDB scan have to rebuild the operator outside this
crate, which a private type makes impossible.
2026-09-17 10:59:41 -07:00
Lance Release 89ad06c782 Bump version: 0.40.0-beta.1 → 0.40.0-beta.2 2026-09-17 12:52:27 +00:00
Wyatt Alt a91b29efc1 fix: scope the Function binding guard on schema evolution to the bound columns (#4204)
A table with one registered Function binding refused every add_columns,
alter_columns, drop_columns and field-metadata update, whatever column
they named. The hazard is narrower: a binding stores the exact Arrow
fields of its inputs, outputs and assignment column, so editing one of
those strands it and the table stops accepting rows. Any other column
was never at risk.

The guard now compares the columns a request names, a rename's target
included, against the set every binding depends on, and refuses only on
an intersection. Local and remote tables apply the same rule. The
blanket guard stays on update and merge insert, which cannot say what
they touch.
2026-09-17 05:51:12 -07:00
Wyatt Alt 5231d37f8d feat: list a table's per-row Function errors from the client (#4208)
A refresh running under a skip policy records each row it skipped, with
the failing input and the error, but the client could not read that
store: the server exposes it over SQL and, since recently, a REST route.
A user who hit per-row failures still had to open a SQL session.

`Table::function_errors` calls the route. The listing is table-addressed
with optional job and column filters, the same addressing the SQL
surface uses, so the two cannot disagree about what a table's errors
are. The two non-record signals come back as their own fields rather
than as rows: capped-fragment summaries, and whether the listing stopped
at its limit. Local tables refuse rather than answer with an empty list.

Python and Node expose the same call, with the same optional filters.
2026-09-17 05:12:50 -07:00
Xuanwo c72931b30f feat(python): add TypeSafe reranker (#4209)
Adds `TypeSafeReranker`, which reranks vector, FTS, and hybrid results
with the [TypeSafe System One
API](https://docs.typesafe.ai/introduction).

Each result is scored independently: TypeSafe reads `{"query",
"document"}` and answers a yes/no (noul) question, and the probability
of yes becomes `_relevance_score`. Unlike listwise LLM rerankers, the
score is an absolute probability, so it is comparable across queries and
can be thresholded. The question's `instructions` and `true`/`false`
`criteria` are configurable, since domain-specific criteria are what
make this kind of scoring work well ([TypeSafe's re-ranking
cookbook](https://docs.typesafe.ai/cookbooks/rerank_typesafe)).

The API takes one state per request, so the reranker sends one request
per result on a thread pool bounded by `max_concurrency`. It
deliberately does not use the background event loop: rerankers are
called synchronously from inside the async query APIs, where `LOOP.run`
would deadlock.

TypeSafe scores for the same pair vary slightly between calls, so
results with close scores can swap places when a search is repeated. The
shared reranker test helper now takes `deterministic=False` for this
case: it still checks result sizes and descending scores, but not that
two identical searches return the same order.

The SDK is imported only when the client is created and questions are
sent as plain dicts, so the new tests run in CI with a fake client and
without `typesafe-sdk` installed. The live-API test is skipped without
`TYPESAFE_API_KEY`. Ranking quality has not been compared with other API
rerankers.
2026-09-17 15:52:34 +08:00
Jonathan HsiehandClaude Opus 5 60a1b4c219 feat(secrets): named Secrets, bindings, and namespace addressing (#4150)
Adds the client half of database-scoped named Secrets: a Secret is a
name and
an opaque value stored by the service, and a Function binds one to the
environment variable its library already reads. Secrets are addressed by
a
namespace path plus a name.

The UDF body is unchanged and stays portable — it reads `OPENAI_API_KEY`
the
way it always did, and the binding is what puts a value there:

```python
db.create_secret("openai-prod", os.environ["OPENAI_API_KEY"])
function = db.create_function(
    analyze_caption,
    secrets=[
        EnvVarSecret(secret_name="openai-prod", env_variable="OPENAI_API_KEY")
    ],
)
function.secret_bindings   # the Secret's name, never its value
```

- `create_secret` / `alter_secret` / `list_secrets` / `describe_secret`
/
`drop_secret` on sync, async and remote connections, with the pyo3
binding
and the Rust client behind them. Each takes `namespace_path`
keyword-only,
  defaulting to the root.
- **There is no read API, by construction rather than by policy** — no
code
  path returns a stored credential, and `describe_secret` answers with
  metadata only.
- `EnvVarSecret` is a pure local constructor: it contacts no server, so
it
cannot fail on a Secret that does not exist. It exists so that a bare
string
in that position — which would be a credential — is a `TypeError` rather
  than a plausible-looking mistake that reads identically in a diff.
- `create_function(..., secrets=[...])` carries the bindings as
`secret_bindings`: a list of `SecretBinding` tagged by `kind`, so a
later
delivery mode is a variant rather than a sibling field. The value never
travels — it is resolved by the service when the Function runs, which is
what
lets a rotation reach columns already pinned to an older
FunctionVersion.
- A binding names its Secret as a `SecretReference` of `{name,
namespace_path}`
rather than one joined string, so no delimiter has to be excluded from
every
name and segment forever, and `ClientConfig.id_delimiter` cannot
contradict
  an identity built on a fixed separator.
- A root namespace is omitted from the request body rather than sent
empty, so
  a root request is byte-identical to one from a client that predates
  namespaces. Tests pin it.

This is the client surface the design's §4 describes; the service side
lives in
sophon.

**Previously split across two PRs.** Namespace addressing was #4151,
stacked on
this one; it is folded in here so the Secret identity contract — name,
namespace path, and the binding that carries both — is reviewable as one
piece
rather than as a shape introduced and then replaced.

## Identifier safety, merged from #4189

**#4189 is merged into this branch**, so the client half of Secrets and
the
guards on the identity it puts in the URL are one PR. What it added:

- Components are checked where the identifier is built, before a request
is
  constructed. `create_secret("../jobs", value)` no longer resolves to
`/v1/jobs/create` and delivers a credential-bearing body to a route with
none
  of this one's body suppression.
- Each component is percent-encoded and joined by the delimiter, so
nothing
  inside a component can end the path segment or add one.
- A component may not be empty, a relative segment (`.`, `..`, and their
`%2e`
spellings), or the delimiter itself — the three ways a component erases
a
boundary the split has to recover. `["prod", ""]` joined to `prod$`,
which
  reads back as `["prod"]`.
- `$` is the only accepted `id_delimiter`, refused at client
construction.
`ClientConfig.id_delimiter` remains, since the identifier grammar comes
from
the Lance REST catalog standard, but a value that would produce
identifiers no
  service splits the caller's way is now an error where it was written.
- One `build_object_identifier` and one character set serve tables,
namespaces,
  Secrets, Functions and materialized views.

Components are checked for *addressability*, not a character set: the
name's
own grammar stays each object's own, so a catalog database keeps the `/`
that
`RemoteCatalog::validate_name` allows.

## Known shortcoming

`secret_bindings` is omitted from a registration body when empty, so a
client
that binds nothing sends what a client without bindings sends. When a
client
does bind a Secret and the service does not know the field, the field is
ignored: registration succeeds, the returned version carries no
bindings, and
the Function fails at execution with the variable unset, far from the
call that
asked for it.

`ServerVersion` is how this codebase refuses a feature the service is
too old
for, and it gates five features already. It does not gate this one: it
is held
per table, and registering a Function is a database-level call. Noted at
the
field in `remote/db.rs`; wiring the gate is follow-up work.

**Tests:** lancedb lib 1340 passed, `first_class_function_slice1` 9,
`first_class_function_slice2` 3, plus Python tests across both slices.
Rebased onto `main` after #4176 (OCI Function identity), #4191 (`.`/`..`
table
names) and #4195 (remote catalogs).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_01UfmeJ533rQDnPBkMtjerV6

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-16 17:00:23 -07:00
LanceDB Robot f8d73b3447 chore: update lance dependency to v13.0.0-beta.4 (#4207)
Updates the Rust workspace Lance dependencies, Cargo lockfile, and Java
lance-core from v13.0.0-beta.3 to
[v13.0.0-beta.4](https://github.com/lance-format/lance/releases/tag/v13.0.0-beta.4);
no compatibility fixes were required.
Validation passed: `cargo clippy --quiet --workspace --tests
--all-features -- -D warnings`, `cargo fmt --all --quiet`, and `git diff
--check`.
2026-09-16 18:11:15 -05:00
Colin Patrick McCabe 04150b3b82 feat: index builders for NGram, BloomFilter, RTree (#4205)
Add index builders in the lancedb library for NGram, BloomFilter, and
RTree indexes.
2026-09-16 15:03:06 -07:00
Lance Release 3309c71a6e Bump version: 0.40.0-beta.0 → 0.40.0-beta.1 2026-09-16 20:15:36 +00:00
Bruno Ramirez 32871e97a7 fix: widen zonemap index type support (#4206)
ZoneMap indexes were added to the Rust API in #4199, but LanceDB reused
the BTree type validation when creating them. That made the public
builder reject some types that Lance ZoneMap can support, including
`LargeUtf8`, `Binary`, and `LargeBinary`. This PR gives ZoneMap its own
validation helper so it can accept the broader scalar set while keeping
the rest of the create-index path unchanged.
2026-09-16 13:10:19 -07:00
Lance Release 86835da5db Bump version: 0.39.0-beta.10 → 0.40.0-beta.0 2026-09-16 19:00:30 +00:00
Bruno Ramirez a4f66afc69 feat(rust): add zonemap index builder (#4199)
Lance supports ZoneMap scalar indexes, but the LanceDB Rust API did not
expose a first-class way to request one through `Table::create_index`.
Users had builders for the other scalar index families, while ZoneMap
was missing from the public `Index` model and remote create-index
serialization. This PR adds ZoneMap as a supported scalar index option
in LanceDB.

This was accomplished with the following changes:
- Added `ZoneMapIndexBuilder` in `rust/lancedb/src/index/scalar.rs`.
- Added `Index::ZoneMap` and `IndexType::ZoneMap`, including
display/from-string aliases for `ZONEMAP` and `ZONE_MAP`.
- Mapped local index creation to
`ScalarIndexParams::for_builtin(BuiltinIndexType::ZoneMap)` and Lance
`IndexType::ZoneMap` in `rust/lancedb/src/table/create_index.rs`.
- Serialized remote create-index requests as `index_type: "ZONEMAP"` in
`rust/lancedb/src/remote/table.rs`.
- Added coverage for both local ZoneMap index creation and remote
request serialization.

Example:

```rust
table
    .create_index(&["my_column"], Index::ZoneMap(Default::default()))
    .execute()
    .await?;
```

### Testing

Added `test_create_zonemap_index` for local index creation and extended
the remote request body test matrix for `ZONEMAP`.
2026-09-16 11:54:22 -07:00
Jack YeandXuanwo 6a07f88980 feat: add remote catalogs and Python and TypeScript bindings (#4195)
Add a `Catalog` trait and `RemoteCatalog` for managing databases,
exposed through Rust, synchronous/asynchronous Python, and TypeScript. A
remote catalog represents the server's root namespace, and each database
is one child namespace. Create/connect return ordinary LanceDB
connections, so existing table APIs work unchanged.

## Rust API

`Catalog` is an object-safe async trait with `create_database`,
`connect_database`, `list_databases`, and `drop_database`. Backend
create/connect methods return `Arc<dyn Database>`; the public
`CatalogConnection` wraps them as `Connection` values and shares its
embedding registry with those connections. `RemoteCatalog` implements
the trait; `connect_catalog` is the convenience builder, available with
the `remote` feature.

```rust
use lancedb::catalog::{
    CreateDatabaseRequest, DropDatabaseRequest, ListDatabasesRequest,
};

let catalog = lancedb::connect_catalog("https://my-server.example")
    .api_key("my-api-key")
    .execute()
    .await?;

let db = catalog.create_database(
    CreateDatabaseRequest::new("analytics").exist_ok(true),
).await?;
let connected = catalog.connect_database("analytics").await?;
let page = catalog.list_databases(
    ListDatabasesRequest::default().limit(20),
).await?;
// page.databases: Vec<String>; page.page_token: Option<String>
catalog.drop_database(
    DropDatabaseRequest::new("analytics").ignore_missing(true),
).await?;
```

Create/drop also accept a plain name for default behavior, e.g.
`catalog.create_database("analytics").await?`. Existing names fail
creation unless `exist_ok` is enabled; missing names fail drop unless
`ignore_missing` is enabled. Drop always requires an empty database.

## Python API

```python
import lancedb

catalog = lancedb.connect_catalog(
    "https://my-server.example", api_key="my-api-key"
)
db = catalog.create_database("analytics", exist_ok=True)
connected = catalog.connect_database("analytics")
page = catalog.list_databases(limit=20)
# page.databases: list[str]; page.page_token: Optional[str]
if page.page_token is not None:
    next_page = catalog.list_databases(limit=20, page_token=page.page_token)
catalog.drop_database("analytics", ignore_missing=True)
```

`connect_catalog` returns `Catalog`; create/connect return the existing
`DBConnection` API. The async equivalent is `catalog = await
lancedb.connect_catalog_async(...)`, returning `AsyncCatalog`; await
each of the same four methods, with create/connect returning
`AsyncConnection`.

## TypeScript API

```typescript
import { connectCatalog } from "@lancedb/lancedb";

const catalog = await connectCatalog("https://my-server.example", {
  apiKey: "my-api-key",
});
const db = await catalog.createDatabase("analytics", { existOk: true });
const connected = await catalog.connectDatabase("analytics");
const page = await catalog.listDatabases({ limit: 20 });
// page.databases: string[]; page.pageToken?: string
if (page.pageToken !== undefined) {
  const nextPage = await catalog.listDatabases({
    limit: 20, pageToken: page.pageToken,
  });
}
await catalog.dropDatabase("analytics", { ignoreMissing: true });
```

Create/connect return the existing `Connection` API. All four methods
are asynchronous.

## REST mapping

All paths below are relative to the catalog endpoint. `{name}` is the
logical database name encoded as one URL path component. The default
namespace delimiter is `$`, so the root identifier is encoded as `%24`.

| Catalog operation | Existing REST route | Request |
| --- | --- | --- |
| `create_database(name)` | `POST /v1/namespace/{name}/create` |
`{"mode":"Create"}`; `exist_ok=true` sends `{"mode":"ExistOk"}` |
| `connect_database(name)` | `POST /v1/namespace/{name}/describe` |
`{}`; verifies existence before returning a scoped connection |
| `list_databases(...)` | `GET /v1/namespace/%24/list` | Optional
`limit` and `page_token` query parameters |
| `drop_database(name)` | `POST /v1/namespace/{name}/drop` |
`{"mode":"Fail","behavior":"Restrict"}`; `ignore_missing=true` changes
mode to `"Skip"` |

For example, database `team/search` uses
`/v1/namespace/team%2Fsearch/create`. A paginated root listing can use
`/v1/namespace/%24/list?limit=20&page_token=a%2Fb`. The list response
retains the existing namespace wire shape,
`{"namespaces":["analytics"],"page_token":"next"}`; the SDK exposes
`namespaces` as `databases` and preserves the opaque continuation token.
An absent or empty token ends pagination. Page limits must be between 1
and 2147483647. Create/drop accept a namespace JSON response or HTTP
204.

Catalog management requests omit both `x-lancedb-database` and
`x-lancedb-database-prefix`, including values supplied through static or
dynamic headers. Returned database connections set `x-lancedb-database`
to the exact logical name and keep independent scope. API keys, OAuth or
dynamic authentication, client settings, table read consistency
settings, and an optional SQL endpoint override carry over to those
connections. OAuth cannot be combined with an API key or a custom header
provider.

For SQL through an HTTPS catalog, configure the existing SQL endpoint
contract with Rust
`.sql_host_override("grpc+tls://sql.example.com:10026")` or Python
`sql_host_override="grpc+tls://sql.example.com:10026"`. TypeScript
catalog options expose the same setting as `sqlHostOverride`. It is
inherited by created/connected databases, retained by Python connection
serialization, and initialized lazily when SQL is executed.

Create HTTP 409 maps to `DatabaseAlreadyExists`; connect/drop HTTP 404
maps to `DatabaseNotFound`, except that `ignore_missing` suppresses a
missing-database drop. Other server errors propagate. The server
enforces restricted deletion; the client never requests cascading
deletion.

Database names preserve literal slashes as part of one name. They must
be nonempty ASCII, with no control characters, surrounding whitespace,
or configured namespace delimiter, and cannot be `.` or `..`. Endpoints
must be HTTP(S) URLs without embedded credentials, query parameters, or
fragments.

## Scope

This PR adds the client API and reuses existing namespace endpoints.
Local catalogs, `__catalog` storage, location generation/sanitization,
and `__manifest` lifecycle support remain deferred; the Lance dependency
is unchanged.

The PR also runs macOS Node tests serially to avoid existing
resource-contention timeouts reproduced across recent main runs.

---------

Co-authored-by: Xuanwo <github@xuanwo.io>
2026-09-17 01:51:10 +08:00
Joaquin HuiandXuanwo 99ed25f753 fix: return InvalidTableName instead of panicking in open_table/create_table (#4192)
Passing an invalid table name to `open_table` or `create_table` panics
instead of returning an error:

thread '...' panicked at rust/lancedb/src/database/listing.rs:1155:62:
called `Result::unwrap()` on an `Err` value: InvalidTableName { name:
"my table", ... }

Both call sites build the table URI with
`request.location.clone().unwrap_or_else(||
self.table_uri(&request.name).unwrap())`, and `table_uri` is the
function that validates the name — so every rejected name (empty,
spaces, slashes, non-ASCII) hits the inner `unwrap`.
`Error::InvalidTableName` clearly is the intended contract here: the
variant exists for exactly this, and the Python binding maps it to
`ValueError`.

Replaced the closure with a `match` that propagates the validation
error; behavior with an explicit `location` is unchanged (the name is
not validated on that path, as before). Added tests asserting
`InvalidTableName` for `create_table` and `open_table` over a set of
rejected names — both panic without the fix. Full `cargo test -p lancedb
--lib --features remote`: 1214 passed; clippy/fmt clean; `cargo check
--workspace --all-targets` clean.

Co-authored-by: Xuanwo <github@xuanwo.io>
2026-09-16 23:31:06 +08:00
Jack YeandXuanwo f3ef21b8ca feat: use oauth2 crate for OAuth with configurable client auth (#4181)
Stacked on #4173 (`jack/restore-oidc-flows`, base branch mirrored to
this repo so the diff shows only this change); context from review:
https://github.com/lancedb/lancedb/pull/4173#issuecomment-5674048100.
Rebase to `main` once #4173 and #4179 merge.

## What moved to the `oauth2` crate (5.0, no default features)

- Authorization URL generation and CSRF state (`authorize_url`,
`CsrfToken`)
- PKCE S256 challenge/verifier generation and code exchange
- Client-credentials, authorization-code, refresh-token, and device-code
grant request construction
- Device authorization request and the device token polling loop
(`authorization_pending`, `slow_down` +5s, expiry deadline, denial,
network backoff capped at 10s)
- Standard success/error response parsing (`RequestTokenError`)
- Token-endpoint client authentication and standards-compliant parameter
encoding (RFC 6749 2.3.1 Basic encoding)

LanceDB keeps ownership of OIDC discovery (compared `openidconnect`: no
measurable win for our 3-field metadata + strict validation, at real
dependency cost), HTTPS-or-loopback endpoint enforcement, the loopback
callback server, browser/stderr prompts, token caching and refresh
orchestration, and the dedicated hardened Azure IMDS source, which is
unchanged.

## Client authentication methods

New `ClientAuthMethod` enum (`none` | `client_secret_basic` |
`client_secret_post`), exposed in Rust, Python, and Node. Unset resolves
to `client_secret_basic` when a secret is present (RFC 6749 2.3.1
recommendation and the normal Okta confidential-app default, so a
default Okta app works without weakening its configuration) and to
`none` for public clients (PKCE/device). Explicit `none` with a secret,
or basic/post without one, is rejected. The method applies to client
credentials, code exchange, refresh, and device requests. Deliberate
behavior change: confidential clients previously always sent the secret
in the POST body; they now default to Basic (Keycloak accepts both).

No `audience`/`resource` parameters were added: the supported target is
an Okta custom authorization server with the API audience configured
server-side, so client-provided audience parameters are unnecessary;
`add_extra_param` support exists if a concrete provider contract ever
needs them.

## Device polling behavior changes (deliberate, tested)

- The first token poll now happens immediately rather than after one
interval (RFC 8628 allows both).
- Transient failures (HTTP 429, 5xx, `temporarily_unavailable`, network
errors) now retry with exponential backoff capped at 10s instead of
retrying at the fixed interval; polling never spins faster than once per
second even if a server reports a zero interval.

## Security and compatibility

- Issuer and discovered endpoints (and device verification URIs) still
require HTTPS except explicit loopback HTTP, enforced before any crate
URL type is built
- Token HTTP client keeps the hardened redirect policy that refuses
insecure redirect targets; regression test added
- Errors never embed raw response bodies (avoids leaking tokens through
parse failures); all credential types stay redacted in Debug
- Transient conditions (429/5xx/`temporarily_unavailable`) remain
retryable in device polling and hard errors elsewhere; refresh keeps
rotation and reauthentication semantics
- Existing public APIs stay source-compatible except the added
`OAuthConfig.client_auth_method` field

## Tests

Rust: client-auth methods across code
exchange/refresh/client-credentials/device (none/basic/post),
auth-method resolution and validation, transient device retries,
denial/expiry, redirect rejection, malformed-response leak check, PKCE
URL assertions, redaction. Python and Node: enum values, conversion,
unknown-method errors, config defaults.

Manual Okta validation recipe (no automated Okta credentials): create a
custom authorization server with an API audience, one confidential web
app (Basic) for authorization-code, one native app (PKCE, no secret),
one native device app; point `issuer_url` at the custom server, set
`client_auth_method` only for the POST-required case; verify token
acquisition, refresh after expiry, and `x-lancedb-credential-type: oidc`
against a LanceDB deployment. Never commit tenant URLs or secrets.

Co-authored-by: Xuanwo <github@xuanwo.io>
2026-09-16 20:11:08 +08:00
Lance Release ed100ccc31 Bump version: 0.39.0-beta.9 → 0.39.0-beta.10 2026-09-16 09:26:22 +00:00
Xuanwo 161f81276d chore: update lance dependency to v13.0.0-beta.3 (#4197)
Update the Rust workspace and Java Lance dependencies to
[v13.0.0-beta.3](https://github.com/lance-format/lance/releases/tag/v13.0.0-beta.3),
which includes the nullable-list encoding fix in
lance-format/lance#9268. Also resolve existing strict Clippy diagnostics
in crate-private OAuth modules and the Python token-cache conversion,
without changing effective visibility.

Validated the rebuilt Python SDK with nullable and nested lists:
Match/Phrase searches preserve document coordinates before and after
appending data and optimizing the table.
2026-09-16 17:25:02 +08:00
Xuanwo 7991e27f35 fix(node): prevent OAuth tests from launching browsers (#4196)
OAuth tests set `LANCEDB_OAUTH_BROWSER` inside Jest's sandbox, which
does not update the process environment read by Rust. As a result, the
native login flow launches a real browser; the [failing macOS main
job](https://github.com/lancedb/lancedb/actions/runs/35067074667/job/104699667840)
ends by terminating an orphaned Safari process.

Set the no-op browser helper while loading Jest configuration, before
creating sandboxes and workers, so both local and CI tests inherit it.
Check the inherited process environment before OAuth login to catch
regressions without opening a browser. Windows uses a no-op command
fixture. Test deadlines and worker counts are unchanged.

A negative control restoring the sandbox-only assignment fails in the
new pre-login check. The full macOS test suite passes with the standard
test launcher; the Windows helper has not been executed locally. The
[first hosted macOS
run](https://github.com/lancedb/lancedb/actions/runs/35071525843/job/104713928139)
passes all 843 tests (5 skipped), with no Safari process in the job log.
Further normal runs are needed to establish sustained stability.
2026-09-16 16:32:34 +08:00
Xuanwo 34ab278a5b fix: expose public FTS paths in remote index listings (#4194)
Remote `list_indices()` exposes physical FTS paths such as
`docs.item.content`, while index creation, queries, and native table
listings use `docs.content`. Normalize FTS columns to the public path
after parsing the server response, using the same field-ID-based
conversion as native tables.

Keep the physical paths on the wire: existing clients need them to
resolve the Arrow schema. Cover both legacy responses that fetch index
statistics and enriched responses that already include the index type.
2026-09-16 15:08:56 +08:00
3b37ea2c7a feat: represent registered Functions by OCI image identity (#4176)
Function versions identify independent Function objects and their
numeric revisions. Rust and Python expose the object ID, location,
canonical decimal version, metadata, and availability separately from
the OCI image digest. Computed-column applications carry the complete
object reference, so existing bindings retain their identity after a
name is removed and reused.

Source authoring keeps the existing
create_function/create_function_async, job wait, and column-binding
APIs. The server coordinates baking followed by registration; users do
not have to manage manifest digests to create a Function. The previous
stored Function representation is intentionally unsupported. Existing
contract tests and shared wire fixtures are migrated to the new model.

This SDK change accompanies the final integration layer
https://github.com/lancedb/sophon/pull/7887 in the Sophon stack:
https://github.com/lancedb/sophon/pull/7885 →
https://github.com/lancedb/sophon/pull/7886 →
https://github.com/lancedb/sophon/pull/7887. A metadata-only revision
can retain the same executable image; Function version numbers must not
be used as image cache keys.

---------

Co-authored-by: lancedb automation <robot@lancedb.com>
Co-authored-by: Yang Cen <bubble-cal@outlook.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-16 12:38:59 +08:00
Shiduo LiandXuanwo f7579d2aae fix(python): correct pylance test extra pin (#4051)
Closes #4045

## What changed
- replace the unpublished pylance 9.0.0rc1 test-extra pin with the
published 9.0.0 release
- restore dependency resolution for editable installs using the tests
extra

## Validation
- downloaded pylance==9.0.0 from PyPI with pip --no-deps
- parsed python/pyproject.toml with tomllib and verified the tests extra
- git diff --check

---------

Co-authored-by: Xuanwo <github@xuanwo.io>
2026-09-15 19:47:13 -07:00
mikemikimike be3215a6ae feat(nodejs): add JSON field helper (#4082)
## Issue

Fixes #4063

## Background

The Node.js SDK currently requires callers to know the Arrow extension
metadata needed to represent JSON fields. This makes a common LanceDB
schema type unnecessarily verbose and easy to get wrong.

## Changes

- Add `makeJsonField(name, nullable = true)` to create a UTF-8 Arrow
field with the `arrow.json` extension metadata.
- Re-export the helper from the public Node.js entry point.
- Add coverage for the default nullable behavior, explicit non-nullable
fields, and the extension metadata.
- Add the generated TypeDoc function page and public globals entry,
including a usage example.

## Implementation

The helper uses the existing Apache Arrow `Field` type and sets
`ARROW:extension:name` to `arrow.json`, matching the metadata convention
already used by LanceDB.

## Compatibility

This is an additive Node.js API. Existing schema construction and Arrow
behavior are unchanged.

## Verification

- `pnpm test -- arrow.test.ts --runInBand` — 236 tests passed.
- `pnpm exec biome ci lancedb/arrow.ts lancedb/index.ts
__test__/arrow.test.ts` — passed.
- `git diff --check` — passed.

## Not run / known limitations

- `pnpm build` and `pnpm run docs` were attempted after expanding the
checkout. Both are blocked locally by the native binding build/type
declarations: Cargo did not complete, and TypeDoc reported the missing
generated `nodejs/lancedb/native` module. The docs files were generated
from the updated TypeScript comments; full build and docs validation are
left to CI.
2026-09-15 18:01:42 -07:00
Colin Patrick McCabe 3a1d3be256 feat(oidc): support resource and audience (#4193)
Support configuring resource and audience for OAuth authorization, token
exchange, and refresh requests.
2026-09-15 16:35:37 -07:00
Colin Patrick McCabe ba693ae43d fix: do not allow . and .. as table names (#4191)
Do not allow . and .. as table names. They are incompatible with the
local filesystem, and confusing in cases where they are supported.
2026-09-15 15:49:57 -07:00