A Function instance is created once with an initialization row (`create(initialization, context)` in the Function Format), but nothing in the client model could say what that row holds or where its values come from, so every source-built Function ran with an empty row and users pushed configuration through `env=` strings instead (ENT-2516). This adds initialization to the canonical wire model without changing a Function's identity: - `FunctionSignature.initialization` lists the fields of the row, in order. It is omitted when empty, so existing signatures and FunctionVersion hashes are unchanged. - `FunctionApplication.initialization` carries a binding's values as JSON. The service validates them against the Function's fields and re-encodes them as Arrow, so this JSON is transport only and never hashed; that is why floats are accepted here while column-input literals keep the Slice 1 domain. - `FunctionBinding.initialization` is the validated one-row Arrow IPC stream (base64) that every instance of the binding is created with. Values belong to the binding, so one Function version can back several columns with different values. - The binding-shape check that guards schema mutations accepts the new field; without it a table holding an initialized binding refused further declarations. Rust and Python share new canonical goldens for an initialized application and binding. The authoring side (class `@udf`, initialization arguments, shipped helper code) is the stacked PR on top of this one.
The Multimodal AI Lakehouse
How to Install ✦ Detailed Documentation ✦ Tutorials and Recipes ✦ Contributors
The ultimate multimodal data platform for AI/ML applications.
LanceDB is designed for fast, scalable, and production-ready vector search. It is built on top of the Lance columnar format. You can store, index, and search over petabytes of multimodal data and vectors with ease. LanceDB is a central location where developers can build, train and analyze their AI workloads.
Demo: Multimodal Search by Keyword, Vector or with SQL
Star LanceDB to get updates!
Key Features:
- Fast Vector Search: Search billions of vectors in milliseconds with state-of-the-art indexing.
- Comprehensive Search: Support for vector similarity search, full-text search and SQL.
- Multimodal Support: Store, query and filter vectors, metadata and multimodal data (text, images, videos, point clouds, and more).
- Advanced Features: Zero-copy, automatic versioning, manage versions of your data without needing extra infrastructure. GPU support in building vector index.
Products:
- Open Source & Local: 100% open source, runs locally or in your cloud. No vendor lock-in.
- Cloud and Enterprise: Production-scale vector search with no servers to manage. Complete data sovereignty and security.
Ecosystem:
- Columnar Storage: Built on the Lance columnar format for efficient storage and analytics.
- Seamless Integration: Python, Node.js, Rust, and REST APIs for easy integration. Native Python and Javascript/Typescript support.
- Rich Ecosystem: Integrations with LangChain 🦜️🔗, LlamaIndex 🦙, Apache-Arrow, Pandas, Polars, DuckDB and more on the way.
How to Install:
Follow the Quickstart doc to set up LanceDB locally.
API & SDK: We also support Python, Typescript and Rust SDKs
| Interface | Documentation |
|---|---|
| Python SDK | https://lancedb.github.io/lancedb/python/python/ |
| Typescript SDK | https://lancedb.github.io/lancedb/js/globals/ |
| Rust SDK | https://docs.rs/lancedb/latest/lancedb/index.html |
| REST API | https://docs.lancedb.com/api-reference/rest |
Join Us and Contribute
We welcome contributions from everyone! Whether you're a developer, researcher, or just someone who wants to help out.
If you have any suggestions or feature requests, please feel free to open an issue on GitHub or discuss it on our Discord server.
Check out the GitHub Issues if you would like to work on the features that are planned for the future. If you have any suggestions or feature requests, please feel free to open an issue on GitHub.
