Wyatt Alt 5538bdb9ff feat: refresh an ivf-grouped materialized view in units (#4232)
A grouped view's refresh runs its whole aggregate in one process, so a
view grouped by ivf_partition(col) over a large table is bounded by a
single worker however many workers a deployment has. The IVF index
already holds each partition's row ids, so the aggregate splits along
the index without a shuffle: every partition's groups are computed from
its own rows alone.

plan_grouped_refresh names the units (one per index partition plus one
for the rows the index cannot place) and the source version they read;
write_grouped_unit computes one unit's groups from the index's row ids,
plus the rows of fragments the index has not covered, assigned in place,
and writes them as uncommitted fragments; commit_grouped_refresh
replaces the view's rows with every unit's fragments in one Update.

The split has to publish what the single-pass refresh would. A unit
applies the view's predicate in its own query, because neither the
indexed take nor the fragment scan filters the way a lance scan does. An
index that does not say which fragments it covers has unknown coverage,
not empty, so it yields no plan at all and the caller refreshes in one
pass rather than reading those rows twice. Each result names its unit,
its plan and the view incarnation it was computed for -- two views of
one shape reach the same counters, and fragments written into one
dataset are not publishable into another -- and the commit publishes the
plan's units exactly once each or nothing. The commit lands on the
planned generation or is refused: lance rebases this Update over a
concurrent append rather than rejecting it, so the version it actually
landed on is checked, as the single-pass rebuild already does.

What holds every unit to one index is the source version the plan pins:
indices live in the source manifest, so a rebuild lands in a version the
units never read. Within that version a segment's postings can still
outlive its ownership -- a column rewrite attaches a new file and takes
the fragment out of the segment's bitmap without dropping its rows from
the posting lists -- so a unit keeps a segment's rows only while it
holds their fragment, and reads the rest from the scan.

The split is the grouping only where the index assigns by its own
centroids. lance also builds an index from precomputed partitions, and
records nowhere that it did, so a posting list can hold a row that
ivf_partition puts elsewhere -- that row's group would then be
aggregated in its own unit as well and published twice, since
concatenated fragments cannot merge two halves of a group. The plan
samples each partition and yields no units when they disagree, every
unit proves the rows it took before grouping them, and the commit
refuses a unit that wrote more than the single group its key allows.
2026-09-23 09:13:17 -07:00
2026-09-09 15:33:04 +08:00
2023-03-17 18:15:19 -07:00
2025-03-10 09:01:23 -07:00

LanceDB Cloud Public Beta

LanceDB Website Blog Discord Twitter LinkedIn

LanceDB

The Multimodal AI Lakehouse

How to Install ✦ Detailed Documentation ✦ Tutorials and Recipes ✦ Contributors

The ultimate multimodal data platform for AI/ML applications.

LanceDB is designed for fast, scalable, and production-ready vector search. It is built on top of the Lance columnar format. You can store, index, and search over petabytes of multimodal data and vectors with ease. LanceDB is a central location where developers can build, train and analyze their AI workloads.


Demo: Multimodal Search by Keyword, Vector or with SQL

LanceDB Multimodal Search

Star LanceDB to get updates!

⭐ Click here ⭐ to see how fast we're growing!

Key Features:

  • Fast Vector Search: Search billions of vectors in milliseconds with state-of-the-art indexing.
  • Comprehensive Search: Support for vector similarity search, full-text search and SQL.
  • Multimodal Support: Store, query and filter vectors, metadata and multimodal data (text, images, videos, point clouds, and more).
  • Advanced Features: Zero-copy, automatic versioning, manage versions of your data without needing extra infrastructure. GPU support in building vector index.

Products:

  • Open Source & Local: 100% open source, runs locally or in your cloud. No vendor lock-in.
  • Cloud and Enterprise: Production-scale vector search with no servers to manage. Complete data sovereignty and security.

Ecosystem:

  • Columnar Storage: Built on the Lance columnar format for efficient storage and analytics.
  • Seamless Integration: Python, Node.js, Rust, and REST APIs for easy integration. Native Python and Javascript/Typescript support.
  • Rich Ecosystem: Integrations with LangChain 🦜️🔗, LlamaIndex 🦙, Apache-Arrow, Pandas, Polars, DuckDB and more on the way.

How to Install:

Follow the Quickstart doc to set up LanceDB locally.

API & SDK: We also support Python, Typescript and Rust SDKs

Interface Documentation
Python SDK https://lancedb.github.io/lancedb/python/python/
Typescript SDK https://lancedb.github.io/lancedb/js/globals/
Rust SDK https://docs.rs/lancedb/latest/lancedb/index.html
REST API https://docs.lancedb.com/api-reference/rest

Join Us and Contribute

We welcome contributions from everyone! Whether you're a developer, researcher, or just someone who wants to help out.

If you have any suggestions or feature requests, please feel free to open an issue on GitHub or discuss it on our Discord server.

Check out the GitHub Issues if you would like to work on the features that are planned for the future. If you have any suggestions or feature requests, please feel free to open an issue on GitHub.

Contributors

Stay in Touch With Us


Website Blog Discord Twitter LinkedIn

S
Description
Developer-friendly OSS embedded retrieval library for multimodal AI. Search More; Manage Less.
Readme Apache-2.0
103 MiB
Languages
Rust 44%
Python 24.9%
HTML 22.9%
TypeScript 7.3%
Java 0.7%
Other 0.1%