Files
lancedb/docs/src/js/classes/AutoQuery.md
T
lancedb-gatefixer[bot] 35b5d015ac fix(node): preserve embedding registration in server bundles (#3806)
## Summary

- lazily initialize built-in OpenAI and Hugging Face providers when
consumers call the public embedding registry API
- choose automatic vector versus FTS search from embedding metadata on a
fresh pinned table revision for every execution
- expose automatic string searches as an `AutoQuery` with only
operations common to both native query families
- keep the registry shared and built-in registration safe across
duplicated module graphs

## Root cause

Nitro treats dependency modules as side-effect-free and removes the bare
OpenAI provider import from its generated route. Registration therefore
never runs, so `getRegistry().get("openai")` remains undefined even when
the registry itself is shared globally. Bundlers may also duplicate the
provider and registry module graphs.

The public embedding entry point now initializes built-in providers only
when `getRegistry()` is explicitly called, keeping initialization on a
live path that Nitro retains. Each terminal automatic-search execution
pins the exact table revision visible at dispatch, reads embedding
metadata and computes an embedding from that snapshot, replays the
builder operations, and constructs and executes the selected native
query against the same snapshot. Pinned native snapshots execute locally
when namespace pushdown cannot carry their revision, while remote
snapshots are seeded directly from one version-and-schema response. The
public `AutoQuery` builder exposes only the operations shared by FTS and
vector search, so runtime class narrowing cannot expose invalid
vector-only methods. Repeated built-in registration replaces stale
constructors from duplicated module graphs while public `register()`
retains its duplicate-alias error.

## Validation

- `cargo fmt --all`
- `cargo check --quiet --features remote --tests --examples`
- `cargo clippy --quiet --features remote --tests --examples`
- `pnpm build`
- `pnpm lint`
- `pnpm run docs`
- `pnpm test --runInBand` (783 passed, 5 skipped)
- serial examples suite with a local OpenAI mock (11 passed), including
`sentence-transformers.test.ts`
- packaged Nitro 2.13.4 server route using the reported imports returned
`{"registered":true}`
- fresh-process FTS fixture initialized both public built-ins and
confirmed automatic string search still returned the indexed row
- schema-consistency regressions cover read-consistency refresh,
checkout, checkoutLatest, restore, runtime class narrowing, concurrent
overwrite during embedding computation, and reused automatic-search
builders
- focused regressions confirm pinned native snapshots bypass unversioned
namespace pushdown and remote snapshots use one describe request

Fixes #2429

<!-- lance-gatekeeper-fix:v1 agent=2adf0f21b8bfb634606ed8897a849e30
generation=1 -->

---------

Co-authored-by: Gatefixer <313497061+lancedb-gatefixer[bot]@users.noreply.github.com>
Co-authored-by: Xuanwo <github@xuanwo.io>
2026-08-26 04:17:32 +08:00

11 KiB

@lancedb/lancedbDocs


@lancedb/lancedb / AutoQuery

Class: AutoQuery

A builder for automatic string searches.

Automatic search determines whether to use full-text or vector search from the table revision selected for each execution. This builder exposes the common operations supported by both query families.

Extends

  • StandardQueryBase<NativeQuery | NativeVectorQuery>

Properties

inner

protected inner: Query | VectorQuery | Promise<Query | VectorQuery>;

Inherited from

StandardQueryBase.inner

Methods

analyzePlan()

analyzePlan(distributedMetrics?): Promise<string>

Executes the query and returns the physical query plan annotated with runtime metrics.

This is useful for debugging and performance analysis, as it shows how the query was executed and includes metrics such as elapsed time, rows processed, and I/O statistics.

Parameters

  • distributedMetrics?: AnalyzePlanDistributedMetrics How distributed worker metrics are displayed for remote query plans. Defaults to "aggregate".

Returns

Promise<string>

A query execution plan with runtime metrics for each step.

Example

import * as lancedb from "@lancedb/lancedb"

const db = await lancedb.connect("./.lancedb");
const table = await db.createTable("my_table", [
  { vector: [1.1, 0.9], id: "1" },
]);

const plan = await table.query().nearestTo([0.5, 0.2]).analyzePlan();

Example output (with runtime metrics inlined):
AnalyzeExec verbose=true, metrics=[]
 ProjectionExec: expr=[id@3 as id, vector@0 as vector, _distance@2 as _distance], metrics=[output_rows=1, elapsed_compute=3.292µs]
  Take: columns="vector, _rowid, _distance, (id)", metrics=[output_rows=1, elapsed_compute=66.001µs, batches_processed=1, bytes_read=8, iops=1, requests=1]
   CoalesceBatchesExec: target_batch_size=1024, metrics=[output_rows=1, elapsed_compute=3.333µs]
    GlobalLimitExec: skip=0, fetch=10, metrics=[output_rows=1, elapsed_compute=167ns]
     FilterExec: _distance@2 IS NOT NULL, metrics=[output_rows=1, elapsed_compute=8.542µs]
      SortExec: TopK(fetch=10), expr=[_distance@2 ASC NULLS LAST], metrics=[output_rows=1, elapsed_compute=63.25µs, row_replacements=1]
       KNNVectorDistance: metric=l2, metrics=[output_rows=1, elapsed_compute=114.333µs, output_batches=1]
        LanceScan: uri=/path/to/data, projection=[vector], row_id=true, row_addr=false, ordered=false, metrics=[output_rows=1, elapsed_compute=103.626µs, bytes_read=549, iops=2, requests=2]

Inherited from

StandardQueryBase.analyzePlan


execute()

protected execute(options?): AsyncGenerator<RecordBatch<any>, void, unknown>

Execute the query and return the results as an

Parameters

Returns

AsyncGenerator<RecordBatch<any>, void, unknown>

See

  • AsyncIterator of
  • RecordBatch.

By default, LanceDb will use many threads to calculate results and, when the result set is large, multiple batches will be processed at one time. This readahead is limited however and backpressure will be applied if this stream is consumed slowly (this constrains the maximum memory used by a single query)

Inherited from

StandardQueryBase.execute


explainPlan()

explainPlan(verbose): Promise<string>

Generates an explanation of the query execution plan.

Parameters

  • verbose: boolean = false If true, provides a more detailed explanation. Defaults to false.

Returns

Promise<string>

A Promise that resolves to a string containing the query execution plan explanation.

Example

import * as lancedb from "@lancedb/lancedb"
const db = await lancedb.connect("./.lancedb");
const table = await db.createTable("my_table", [
  { vector: [1.1, 0.9], id: "1" },
]);
const plan = await table.query().nearestTo([0.5, 0.2]).explainPlan();

Inherited from

StandardQueryBase.explainPlan


fastSearch()

fastSearch(): this

Skip searching un-indexed data. This can make search faster, but will miss any data that is not yet indexed.

Use Table#optimize to index all un-indexed data.

Returns

this

Inherited from

StandardQueryBase.fastSearch


filter()

filter(predicate): this

A filter statement to be applied to this query.

Parameters

  • predicate: string

Returns

this

See

where

Deprecated

Use where instead

Inherited from

StandardQueryBase.filter


fullTextSearch()

fullTextSearch(query, options?): this

Parameters

Returns

this

Inherited from

StandardQueryBase.fullTextSearch


limit()

limit(limit): this

Set the maximum number of results to return.

By default, a plain search has no limit. If this method is not called then every valid row from the table will be returned.

Parameters

  • limit: number

Returns

this

Inherited from

StandardQueryBase.limit


offset()

offset(offset): this

Set the number of rows to skip before returning results.

This is useful for pagination.

Parameters

  • offset: number

Returns

this

Inherited from

StandardQueryBase.offset


orderBy()

orderBy(ordering): this

Sort the results by the specified column(s).

Parameters

Returns

this

This query builder.

Inherited from

StandardQueryBase.orderBy


outputSchema()

outputSchema(): Promise<Schema<any>>

Returns the schema of the output that will be returned by this query.

This can be used to inspect the types and names of the columns that will be returned by the query before executing it.

Returns

Promise<Schema<any>>

An Arrow Schema describing the output columns.

Inherited from

StandardQueryBase.outputSchema


select()

select(columns): this

Return only the specified columns.

By default a query will return all columns from the table. However, this can have a very significant impact on latency. LanceDb stores data in a columnar fashion. This means we can finely tune our I/O to select exactly the columns we need.

As a best practice you should always limit queries to the columns that you need. If you pass in an array of column names then only those columns will be returned.

You can also use this method to create new "dynamic" columns based on your existing columns. For example, you may not care about "a" or "b" but instead simply want "a + b". This is often seen in the SELECT clause of an SQL query (e.g. SELECT a+b FROM my_table).

To create dynamic columns you can pass in a Map<string, string>. A column will be returned for each entry in the map. The key provides the name of the column. The value is an SQL string used to specify how the column is calculated.

For example, an SQL query might state SELECT a + b AS combined, c. The equivalent input to this method would be:

Parameters

  • columns: string | string[] | Record<string, string> | Map<string, string>

Returns

this

Example

new Map([["combined", "a + b"], ["c", "c"]])

Columns will always be returned in the order given, even if that order is different than
the order used when adding the data.

Note that you can pass in a `Record<string, string>` (e.g. an object literal). This method
uses `Object.entries` which should preserve the insertion order of the object.  However,
object insertion order is easy to get wrong and `Map` is more foolproof.

Inherited from

StandardQueryBase.select


toArray()

toArray(options?): Promise<any[]>

Collect the results as an array of objects.

Parameters

Returns

Promise<any[]>

Inherited from

StandardQueryBase.toArray


toArrow()

toArrow(options?): Promise<Table<any>>

Collect the results as an Arrow

Parameters

Returns

Promise<Table<any>>

See

ArrowTable.

Inherited from

StandardQueryBase.toArrow


useLsm()

useLsm(enable): this

Control MemWAL read routing for this query.

By default (unset), when the table carries a MemWAL write spec (see Table#setLsmWriteSpec), reads are routed through the LSM scanner so they also return data written via the mergeInsert LSM path that has not yet been compacted into the base table (the active/frozen in-memory memtables and the flushed generations), deduplicated by primary key; a table without a spec reads the base table.

Parameters

  • enable: boolean true forces the LSM scanner and errors if the table has no MemWAL write spec. false bypasses the MemWAL and reads the base table only, even when a spec is present. Note: the LSM scanner does not support every query shape (e.g. reranking, hybrid search, orderBy). On a MemWAL table those shapes error unless useLsm(false) is set, because a base-only read would silently exclude un-compacted MemWAL data.

Returns

this

Inherited from

StandardQueryBase.useLsm


where()

where(predicate): this

A filter statement to be applied to this query.

The filter should be supplied as an SQL query string. For example:

Parameters

  • predicate: string

Returns

this

Example

x > 10
y > 0 AND y < 100
x > 5 OR y = 'test'

Filtering performance can often be improved by creating a scalar index
on the filter column(s).

Calling this multiple times combines the filters with a logical AND rather
than replacing the previous filter.

Inherited from

StandardQueryBase.where


withRowId()

withRowId(): this

Whether to return the row id in the results.

This column can be used to match results between different queries. For example, to match results from a full text search and a vector search in order to perform hybrid search.

Returns

this

Inherited from

StandardQueryBase.withRowId