A grouped view's refresh runs its whole aggregate in one process, so a view grouped by ivf_partition(col) over a large table is bounded by a single worker however many workers a deployment has. The IVF index already holds each partition's row ids, so the aggregate splits along the index without a shuffle: every partition's groups are computed from its own rows alone. plan_grouped_refresh names the units (one per index partition plus one for the rows the index cannot place) and the source version they read; write_grouped_unit computes one unit's groups from the index's row ids, plus the rows of fragments the index has not covered, assigned in place, and writes them as uncommitted fragments; commit_grouped_refresh replaces the view's rows with every unit's fragments in one Update. The split has to publish what the single-pass refresh would. A unit applies the view's predicate in its own query, because neither the indexed take nor the fragment scan filters the way a lance scan does. An index that does not say which fragments it covers has unknown coverage, not empty, so it yields no plan at all and the caller refreshes in one pass rather than reading those rows twice. Each result names its unit, its plan and the view incarnation it was computed for -- two views of one shape reach the same counters, and fragments written into one dataset are not publishable into another -- and the commit publishes the plan's units exactly once each or nothing. The commit lands on the planned generation or is refused: lance rebases this Update over a concurrent append rather than rejecting it, so the version it actually landed on is checked, as the single-pass rebuild already does. What holds every unit to one index is the source version the plan pins: indices live in the source manifest, so a rebuild lands in a version the units never read. Within that version a segment's postings can still outlive its ownership -- a column rewrite attaches a new file and takes the fragment out of the segment's bitmap without dropping its rows from the posting lists -- so a unit keeps a segment's rows only while it holds their fragment, and reads the rest from the scan. The split is the grouping only where the index assigns by its own centroids. lance also builds an index from precomputed partitions, and records nowhere that it did, so a posting list can hold a row that ivf_partition puts elsewhere -- that row's group would then be aggregated in its own unit as well and published twice, since concatenated fragments cannot merge two halves of a group. The plan samples each partition and yields no units when they disagree, every unit proves the rows it took before grouping them, and the commit refuses a unit that wrote more than the single group its key allows.
The Multimodal AI Lakehouse
How to Install ✦ Detailed Documentation ✦ Tutorials and Recipes ✦ Contributors
The ultimate multimodal data platform for AI/ML applications.
LanceDB is designed for fast, scalable, and production-ready vector search. It is built on top of the Lance columnar format. You can store, index, and search over petabytes of multimodal data and vectors with ease. LanceDB is a central location where developers can build, train and analyze their AI workloads.
Demo: Multimodal Search by Keyword, Vector or with SQL
Star LanceDB to get updates!
Key Features:
- Fast Vector Search: Search billions of vectors in milliseconds with state-of-the-art indexing.
- Comprehensive Search: Support for vector similarity search, full-text search and SQL.
- Multimodal Support: Store, query and filter vectors, metadata and multimodal data (text, images, videos, point clouds, and more).
- Advanced Features: Zero-copy, automatic versioning, manage versions of your data without needing extra infrastructure. GPU support in building vector index.
Products:
- Open Source & Local: 100% open source, runs locally or in your cloud. No vendor lock-in.
- Cloud and Enterprise: Production-scale vector search with no servers to manage. Complete data sovereignty and security.
Ecosystem:
- Columnar Storage: Built on the Lance columnar format for efficient storage and analytics.
- Seamless Integration: Python, Node.js, Rust, and REST APIs for easy integration. Native Python and Javascript/Typescript support.
- Rich Ecosystem: Integrations with LangChain 🦜️🔗, LlamaIndex 🦙, Apache-Arrow, Pandas, Polars, DuckDB and more on the way.
How to Install:
Follow the Quickstart doc to set up LanceDB locally.
API & SDK: We also support Python, Typescript and Rust SDKs
| Interface | Documentation |
|---|---|
| Python SDK | https://lancedb.github.io/lancedb/python/python/ |
| Typescript SDK | https://lancedb.github.io/lancedb/js/globals/ |
| Rust SDK | https://docs.rs/lancedb/latest/lancedb/index.html |
| REST API | https://docs.lancedb.com/api-reference/rest |
Join Us and Contribute
We welcome contributions from everyone! Whether you're a developer, researcher, or just someone who wants to help out.
If you have any suggestions or feature requests, please feel free to open an issue on GitHub or discuss it on our Discord server.
Check out the GitHub Issues if you would like to work on the features that are planned for the future. If you have any suggestions or feature requests, please feel free to open an issue on GitHub.
