Stacked on #4253. Remote Python Functions could not share helper code, take typed parameters, or keep per-process state: helpers became `from <module> import ...` lines a worker cannot resolve, every parameter was an `env=` string, and state had to be stashed on imported modules. MMLB's OpenAI preset shows the cost — 24 env vars, a 302-line callable, and a rate limiter duplicated between the scalar and batch variants, one of which never constructed it (ENT-2516). **Class Functions.** `@udf` accepts a class. Each remote instance runs `__init__` once, calls `__call__` for every row or batch it processes (row or batch mode is inferred from annotations as before), and calls `close()` once if defined. Calling the definition locally constructs the class, so unit tests stay ordinary. **Initialization.** The annotated `__init__` parameters become the Function's initialization fields. A binding passes their values beside its column inputs: `fn(text=col("body"), model="small", dimensions=512)`. Values are constants of that binding; another column can bind the same Function version with other values. Types are limited to booleans, integers, floats, strings, and lists or structs of them, which have one unambiguous JSON and SQL spelling. A parameter with a default may be omitted; a null value then takes the default. Initialization and input names must be disjoint because both are keyword arguments of one call. Secrets stay on `EnvVarSecret`. **Helper code.** `code=[package, ...]` ships top-level modules or packages as Python source in a `python_bundle` artifact (a canonical JSON object of path → source), imported normally on the worker. Changing a helper changes the artifact digest and so the Function version; the environment from `pip`/`conda` is reused. Functions and classes defined in `__main__` (notebook cells, scripts) are packaged by source, recursively and dependencies first. An import of a module that lives in a local source tree, is not shipped with `code=`, and names no declared package is now rejected at registration instead of failing on the worker. Closures, lambdas, and nested definitions are still rejected, with the reason. A Function without these features packages byte-identically to before (`python_callable`, same digest, no `initialization` on the wire). Class sources are located through a method's code object, because `inspect.getsource` cannot find a class defined in a notebook cell or doctest. The packaging contract is documented on `udf`. Sophon changes that build and execute these artifacts are in lancedb/sophon; they were verified end to end on a Linux local server (registration → REST, SQL, and materialized-view bindings with different initialization → refresh → query, one construction per instance).
LanceDB Python SDK
A Python library for LanceDB.
Installation
pip install lancedb
Pre-Haswell x86_64 hosts: lancedb-compat
The default lancedb wheel targets x86-64-haswell (AVX2 + FMA + F16C) for full performance on modern hardware. Pre-Haswell hosts — Intel Sandy Bridge / Ivy Bridge / Westmere; AMD Bulldozer / Piledriver / Steamroller — don't have AVX2 and crash with Illegal instruction at import lancedb.
For those hosts, install the lancedb-compat package instead:
pip install lancedb-compat
Same Python API (import lancedb works as usual). The compat wheel is compiled at the x86-64-v2 baseline (Nehalem-class) and uses runtime SIMD dispatch in the embedded lance crate to pick the right kernel tier (scalar / AVX / AVX+FMA / AVX2+FMA / AVX-512) at load time, so it still goes fast on modern hardware while running cleanly on the pre-Haswell silicon. Use lance.simd_info() from Python to verify which tier was selected.
lancedb and lancedb-compat install to the same lancedb/ namespace and conflict at install time. Pick one. To switch, pip uninstall lancedb first, then pip install lancedb-compat (or vice-versa).
If you need a custom baseline (or lancedb-compat isn't yet published for your platform), build from source with the override:
RUSTFLAGS="-C target-cpu=x86-64-v2" maturin build --release
pip install ./target/wheels/lancedb-*.whl
Preview Releases
Stable releases are created about every 2 weeks. For the latest features and bug fixes, you can install the preview release. These releases receive the same level of testing as stable releases, but are not guaranteed to be available for more than 6 months after they are released. Once your application is stable, we recommend switching to stable releases.
pip install --pre --extra-index-url https://pypi.fury.io/lancedb/ lancedb
Threading in CPU-limited containers
LanceDB uses separate pools for compute work and storage I/O. On a container with two visible CPUs, current releases intentionally use one compute worker by default; no manual configuration is needed. If every query logs an I/O core reservation warning on a two-CPU container, upgrade from LanceDB 0.21.1 or earlier.
The two commonly tuned environment variables control different resources:
LANCE_CPU_THREADSoverrides the number of compute workers. One worker is the appropriate setting for a two-CPU container when an explicit override is needed.LANCE_IO_THREADScontrols concurrent storage operations, not reserved CPU cores. Its default can be greater than the number of CPUs because I/O workers spend much of their time waiting for storage.
Keep the defaults unless measurements show that the workload benefits from an override. See the Lance threading model for the current defaults and tuning guidance.
Usage
Basic Example
import lancedb
db = lancedb.connect('<PATH_TO_LANCEDB_DATASET>')
table = db.open_table('my_table')
results = table.search([0.1, 0.3]).limit(20).to_list()
print(results)
Development
See CONTRIBUTING.md for information on how to contribute to LanceDB.