Files
lancedb/docs/src/js/classes/Table.md
T
Dan Rammer f1c4967eeb feat: bring the MemWAL LSM surface to parity across the SDKs (#3962)
## Why

Four of the eight LSM methods are **remote-only in the core**. `impl
BaseTable for NativeTable` implements only
`set`/`unset`/`get_lsm_write_spec` and `close_lsm_writers`; `flush_lsm`,
`compact_lsm` and `get_lsm_stats` fall through to trait defaults
returning `NotSupported` (`rust/lancedb/src/table.rs:679,687,696`), and
`checkpoint_lsm` is built on all three.

That explains the state of the bindings: Node had bound the four that
work against a local table and stopped, so a Cloud user could install an
LSM write spec but had no way to observe fresh-tier state or drive a
checkpoint. Java had none of it at all.

| SDK | set/unset/get spec | closeWriters | flush | compact | getStats |
checkpoint |
|---|---|---|---|---|---|---|
| Rust core |  |  |  |  |  |  |
| Python |  |  |  |  |  |  |
| Node *(before)* |  |  | — | — | — | — |
| **Node (after)** |  |  | **new** | **new** | **new** | **new** |
| Java *(before)* | — | — | — | — | — | — |
| **Java (after)** | **new** | n/a | **new** | **new** | **new** |
**new** |

Go and C are separate repos and are out of scope here. `closeLsmWriters`
drains cached in-process shard writers, so it has no meaning for Java,
which is a pure REST client.

## Node

Adds napi bindings for `flushLsm`, `compactLsm`, `checkpointLsm` and
`getLsmStats`, plus typed `LsmStats` / `BucketStats` / `GenerationStats`
/ `MemtableStats` objects — typed rather than a JSON blob, matching the
existing `LsmWriteSpec` object in the same file, with `u64` cast to
`i64` per that file's convention.

Because these four are remote-only, the new tests assert each binding
reaches the core and surfaces `NotSupported` against a local table. That
covers the wiring; behavior against a real endpoint stays covered by the
mocked-endpoint tests in `rust/lancedb/src/remote/table.rs`.

## Python

No new methods. All eight are on `LanceTable`, `AsyncTable` and
`RemoteTable` — the last four landed on the sync `RemoteTable` in #3961,
which is merged into this branch.

What was missing here was reachability. `LsmWriteSpec` was importable
only from the private `lancedb._lancedb`, appearing in `table.py` solely
under `if TYPE_CHECKING:`, and `docs/src/python/python.md` had no
mention of it, which per the repo's docs guidance means it rendered
nowhere in the API reference. It is now `lancedb.LsmWriteSpec`, in
`__all__`, and documented.

## Java

Java reaches LanceDB purely over REST through the generated Lance
Namespace client, and these routes are not in that spec, so they are
issued through a small dedicated client rather than added to the spec.
That call is revisitable — LSM is one of four unspecified route families
alongside `multipart_write`, `page_cache/prewarm` and
`branches/diff|merge`. If those are ever regularized into the spec as a
group, `LanceDbTableLsm` is one file that gets deleted.

`LsmWriteSpec` here is deliberately **not**
`org.lance.memwal.InitializeMemWalParams`. That type defaults to
maintaining *no* indexes where a spec here defaults to maintaining
*every* index, and it cannot express the `null` that asks the server to
resolve the set:

| Value | On the wire | Meaning |
|---|---|---|
| unset (null) | `null` | Server resolves **every** maintainable index |
| `Collections.emptyList()` | `[]` | Maintain **none** |
| `Arrays.asList("id_idx")` | `["id_idx"]` | Exactly those |

A dedicated test pins null and `[]` as distinct on the wire, since
collapsing them is the failure mode that motivated a LanceDB-owned type.

`checkpointLsm` is ported from `rust/lancedb/src/table/checkpoint.rs`
with its constants and status semantics intact: 429/503 retried in place
against an 8-budget, 421 restarting from flush against a 3-budget, 5s
poll, and a target watermark fixed after the seal so it terminates under
write load.

`getLsmStats` returns typed `LsmStats` / `BucketStats` /
`GenerationStats` / `MemtableStats`, mirroring the Rust structs in
`rust/lancedb/src/table/lsm_stats.rs` and the objects Node exposes.
Decoding is strict — see below.

## Review feedback

Both gatekeeper findings were real. Each was reproduced against the
scripted test server first, and each fix ships with the reproducer as a
regression test.

**The transport was doubling every checkpoint retry budget.**
`HttpClients.createDefault()` installs Apache's default response retry
strategy, whose retryable-status list is exactly 429 and 503 — the two
statuses `isRetryable` owns. A 429 held against `flush_lsm` issued
**18** wire requests where the loop intends 9, and `compact_lsm` was
retried in place despite the loop being built to fall through to a fresh
stats poll instead. Timing confirmed the mechanism: that run took 25.4s
≈ 16.3s of the loop's own backoff plus 9 × the transport's 1s retry
interval.

Automatic retries are now disabled, so the checkpoint loop is the sole
owner of the 421/429/503 transitions. A side effect worth noting:
`testCheckpointRetriesRetryableStatusInPlace` was passing on a
transport-absorbed 429 and never reaching `issue()`'s retry branch at
all. It now exercises the real path.

**Stats decoding failed open.** `getLsmStats` read the response with
Jackson's `path()`, which yields a missing node that iterates as an
empty array — making "malformed" indistinguishable from "no buckets",
which is indistinguishable from "drained". Four separate payloads made
`checkpointLsm()` report convergence for a checkpoint that never ran:

| Response | Before | Now |
|---|---|---|
| `{"lsm_stats": null}` or absent key | disabled ✓ | disabled ✓ |
| `{"lsm_stats": {}}` | **reported success** | `IllegalStateException` |
| empty response body | **reported success** | `IllegalStateException` |
| bucket missing required fields | **reported success** |
`IllegalStateException` |

The empty-body row is the one to weight: a proxy 200 with no body is a
realistic production event, and it silently reported a checkpoint that
never happened.

Decoding is now strict and fails closed, matching the serde contract on
the Rust side exactly. One deliberate deviation from the review comment,
which asked that *only* explicit JSON `null` count as disabled: Rust has
`#[serde(default)]` on `lsm_stats`, so an **absent key** decodes to
`None` there too. Java now matches that. It is an absent-or-malformed
**`buckets`** that fails closed, which is the case the comment was
actually protecting.

## Testing

- Java: **33 passing** (8 existing + 25 LSM) against a scripted
`com.sun.net.httpserver.HttpServer` — no new test dependency. Wire
assertions mirror `rust/lancedb/src/remote/table.rs:6581-6748`;
checkpoint tests cover convergence, not piling onto a latched bucket,
421 restart-from-flush, 429 retry-in-place, terminal-status propagation,
reissue exhaustion, the exact wire-request count against the retry
budget, and five malformed stats payloads.
- Node: **19 LSM tests passing**; `cargo check`, `npm run build`, `npm
run tsc`, `npm run lint`, `npm run docs` all clean.
- Python: `ruff format --check` and `ruff check` clean.
- Java formatting: `./mvnw -pl lancedb-core spotless:apply` and
`spotless:check` both clean under a JDK 11 toolchain.

## Note: spotless needs a pre-16 JDK

`./mvnw spotless:apply` fails on JDK 16+ with
`JCTree$JCImport.getQualifiedIdentifier()` — google-java-format 1.7,
pinned at `java/pom.xml:34`, predates JDK 16's compiler API change.
**This is pre-existing** and reproduces on a pristine `main` checkout.

It is not a blocker, just a toolchain requirement. Spotless was run
against these sources under JDK 11 and both `spotless:apply` and
`spotless:check` pass on the whole module:

```shell
JAVA_HOME=/path/to/jdk11 ./mvnw -pl lancedb-core spotless:apply
```

Bumping the plugin so it works on modern JDKs is still worth doing, but
separately from this PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 11:44:46 -05:00

31 KiB

@lancedb/lancedbDocs


@lancedb/lancedb / Table

Class: abstract Table

A Table is a collection of Records in a LanceDB Database.

A Table object is expected to be long lived and reused for multiple operations. Table objects will cache a certain amount of index data in memory. This cache will be freed when the Table is garbage collected. To eagerly free the cache you can call the close method. Once the Table is closed, it cannot be used for any further operations.

Tables are created using the methods Connection#createTable and Connection#createEmptyTable. Existing tables are opened using Connection#openTable.

Closing a table is optional. It not closed, it will be closed when it is garbage collected.

Accessors

name

get abstract name(): string

Returns the name of the table

Returns

string

Methods

add()

abstract add(data, options?): Promise<AddResult>

Insert records into this Table.

Parameters

Returns

Promise<AddResult>

A promise that resolves to an object containing the new version number of the table


addColumns()

abstract addColumns(newColumnTransforms): Promise<AddColumnsResult>

Add new columns with defined values.

The { computed } form stores the expression rather than evaluating it now: the column is committed with no values, and rows get them from Table#refreshColumn. Declaring one therefore costs the same on a large table as on an empty one.

A refresh does not revisit rows it has already filled, so mutating an input leaves the value computed at fill time; recomputing means dropping the column and declaring it again. While a declaration reads a column, that column cannot be renamed, retyped or dropped.

On LanceDB Cloud and Enterprise the expression is planned by the server, and the refresh runs as a server job -- see Table#refreshColumnAsync.

Parameters

  • newColumnTransforms: | Field<any> | Field<any>[] | Schema<any> | AddColumnsSql[] | object Either:
    • An array of objects with column names and SQL expressions to calculate values
    • A single Arrow Field defining one column with its data type (column will be initialized with null values)
    • An array of Arrow Fields defining columns with their data types (columns will be initialized with null values)
    • An Arrow Schema defining columns with their data types (columns will be initialized with null values)
    • { computed }, declaring columns defined by a SQL expression whose type and inputs are derived from it

Returns

Promise<AddColumnsResult>

A promise that resolves to an object containing the new version number of the table after adding the columns.

Example

await table.addColumns({ computed: [{ name: "doubled", valueSql: "x * 2" }] });
const { rowsFilled } = await table.refreshColumn("doubled");

alterColumns()

abstract alterColumns(columnAlterations): Promise<AlterColumnsResult>

Alter the name or nullability of columns.

Parameters

  • columnAlterations: ColumnAlteration[] One or more alterations to apply to columns.

Returns

Promise<AlterColumnsResult>

A promise that resolves to an object containing the new version number of the table after altering the columns.


branches()

abstract branches(): Promise<Branches>

Get the branch manager for this table.

Branches are isolated, writable lines of history forked from another branch (or version). Writes on a branch do not affect main.

Returns

Promise<Branches>


checkout()

abstract checkout(version): Promise<void>

Checks out a specific version of the table This is an in-place operation.

This allows viewing previous versions of the table. If you wish to keep writing to the dataset starting from an old version, then use the restore function.

Calling this method will set the table into time-travel mode. If you wish to return to standard mode, call checkoutLatest.

Parameters

  • version: string | number The version to checkout, could be version number or tag

Returns

Promise<void>

Example

import * as lancedb from "@lancedb/lancedb"
const db = await lancedb.connect("./.lancedb");
const table = await db.createTable("my_table", [
  { vector: [1.1, 0.9], type: "vector" },
]);

console.log(await table.version()); // 1
console.log(table.display());
await table.add([{ vector: [0.5, 0.2], type: "vector" }]);
await table.checkout(1);
console.log(await table.version()); // 2

checkoutLatest()

abstract checkoutLatest(): Promise<void>

Checkout the latest version of the table. This is an in-place operation.

The table will be set back into standard mode, and will track the latest version of the table.

Returns

Promise<void>


checkpointLsm()

abstract checkpointLsm(): Promise<void>

Converge this table's LSM write path into its base table.

Seals once, then triggers compaction and polls until the L0 that existed at the start is gone. The target set is fixed at the start, so generations created during the checkpoint are ignored — that is what lets it terminate under write load, and what makes it best-effort: it converges the fresh tier as of some instant. Idempotent, abandonable at any point, and safe to run on a cadence.

There is no liveness bound — the compactor pool is shared across tables, so a checkpoint queued behind unrelated work looks exactly like one that is merging. The caller owns the deadline.

Returns

Promise<void>

Example

const before = await table.getLsmStats();
await table.checkpointLsm();
const after = await table.getLsmStats();

close()

abstract close(): void

Close the table, releasing any underlying resources.

It is safe to call this method multiple times.

Any attempt to use the table after it is closed will result in an error.

Returns

void


closeLsmWriters()

abstract closeLsmWriters(): Promise<void>

Drain and close any cached MemWAL shard writers held for this table.

When an LsmWriteSpec is installed, mergeInsert opens MemWAL shard writers and caches them for reuse across calls. This closes them, flushing pending data; writers reopen lazily on the next mergeInsert. It is a no-op when no writers are cached.

Returns

Promise<void>


compactLsm()

abstract compactLsm(): Promise<void>

Trigger a background L0 → base compaction pass per bucket.

Returns once the passes are dispatched, not once they finish — watch Table#getLsmStats for progress, or use Table#checkpointLsm to wait for convergence.

Returns

Promise<void>


countRows()

abstract countRows(filter?): Promise<number>

Count the total number of rows in the dataset.

Parameters

  • filter?: string

Returns

Promise<number>


createIndex()

abstract createIndex(column, options?): Promise<void>

Create an index to speed up queries.

Indices can be created on vector columns or scalar columns. Indices on vector columns will speed up vector searches. Indices on scalar columns will speed up filtering (in both vector and non-vector searches)

We currently don't support custom named indexes. The index name will always be ${column}_idx.

Parameters

Returns

Promise<void>

Examples

// If the column has a vector (fixed size list) data type then
// an IvfPq vector index will be created.
const table = await conn.openTable("my_table");
await table.createIndex("vector");
// For advanced control over vector index creation you can specify
// the index type and options.
const table = await conn.openTable("my_table");
await table.createIndex("vector", {
  config: lancedb.Index.ivfPq({
    numPartitions: 128,
    numSubVectors: 16,
  }),
});
// Or create a Scalar index
await table.createIndex("my_float_col");

createIndexAsync()

abstract createIndexAsync(column, options?): Promise<Job>

Create an index, returning a handle to the indexing job.

The job may already be complete when returned; callers must not assume the index exists until Job.wait resolves.

Parameters

Returns

Promise<Job>


currentBranch()

abstract currentBranch(): null | string

The branch this table handle is scoped to, or null for the main branch.

A handle returned by Branches.create or Branches.checkout reports the branch it targets; a handle opened normally reports null.

Returns

null | string


delete()

abstract delete(predicate): Promise<DeleteResult>

Delete the rows that satisfy the predicate.

Parameters

  • predicate: string

Returns

Promise<DeleteResult>

A promise that resolves to an object containing the new version number of the table


display()

abstract display(): string

Return a brief description of the table

Returns

string


dropColumns()

abstract dropColumns(columnNames): Promise<DropColumnsResult>

Drop one or more columns from the dataset

This is a metadata-only operation and does not remove the data from the underlying storage. In order to remove the data, you must subsequently call compact_files to rewrite the data without the removed columns and then call cleanup_files to remove the old files.

Parameters

  • columnNames: string[] The names of the columns to drop. These can be nested column references (e.g. "a.b.c") or top-level column names (e.g. "a").

Returns

Promise<DropColumnsResult>

A promise that resolves to an object containing the new version number of the table after dropping the columns.


dropIndex()

abstract dropIndex(name): Promise<void>

Drop an index from the table.

Parameters

  • name: string The name of the index. This does not delete the index from disk, it just removes it from the table. To delete the index, run Table#optimize after dropping the index. Use Table.listIndices to find the names of the indices.

Returns

Promise<void>


flushLsm()

abstract flushLsm(): Promise<void>

Seal every bucket's active memtable into a new L0 generation.

Returns once the seal is committed. Sealing an empty memtable is a no-op, so this is safe to call repeatedly.

Returns

Promise<void>


getLsmStats()

abstract getLsmStats(includeGenerationRows?): Promise<undefined | LsmStats>

Read live per-bucket LSM state.

Answers "how far behind is my fresh tier", "which bucket is hot", and "why is my fresh-tier vector search brute-force". Mutates no table state.

Resolves to undefined only when the LSM write path is not enabled.

Parameters

  • includeGenerationRows?: boolean Also count rows per L0 generation. Off by default because each count opens an uncached Lance dataset.

Returns

Promise<undefined | LsmStats>


getLsmWriteSpec()

abstract getLsmWriteSpec(): Promise<undefined | LsmWriteSpec>

Read the LsmWriteSpec currently installed on this table.

Resolves to undefined when the MemWAL LSM write path is not enabled (no spec has been set, or it was removed with Table#unsetLsmWriteSpec). The returned spec mirrors what was passed to Table#setLsmWriteSpec, except that maintainedIndexes always reports the concrete list resolved when the spec was set — undefined never round-trips.

Returns

Promise<undefined | LsmWriteSpec>


indexStats()

abstract indexStats(name): Promise<undefined | IndexStatistics>

List all the stats of a specified index

Parameters

  • name: string The name of the index.

Returns

Promise<undefined | IndexStatistics>

The stats of the index. If the index does not exist, it will return undefined

Use Table.listIndices to find the names of the indices.


initialStorageOptions()

abstract initialStorageOptions(): Promise<undefined | null | Record<string, string>>

Get the initial storage options that were passed in when opening this table.

For dynamically refreshed options (e.g., credential vending), use Table.latestStorageOptions.

Warning: This is an internal API and the return value is subject to change.

Returns

Promise<undefined | null | Record<string, string>>

The storage options, or undefined if no storage options were configured.


isOpen()

abstract isOpen(): boolean

Return true if the table has not been closed

Returns

boolean


latestStorageOptions()

abstract latestStorageOptions(): Promise<undefined | null | Record<string, string>>

Get the latest storage options, refreshing from provider if configured.

This method is useful for credential vending scenarios where storage options may be refreshed dynamically. If no dynamic provider is configured, this returns the initial static options.

Warning: This is an internal API and the return value is subject to change.

Returns

Promise<undefined | null | Record<string, string>>

The storage options, or undefined if no storage options were configured.


listIndices()

abstract listIndices(): Promise<IndexConfig[]>

List all indices that have been created with Table.createIndex

Returns

Promise<IndexConfig[]>


listVersions()

abstract listVersions(): Promise<Version[]>

List all the versions of the table

Returns

Promise<Version[]>


mergeInsert()

abstract mergeInsert(on): MergeInsertBuilder

Parameters

  • on: string | string[]

Returns

MergeInsertBuilder


optimize()

abstract optimize(options?): Promise<OptimizeStats>

Optimize the on-disk data and indices for better performance.

Modeled after VACUUM in PostgreSQL.

Optimization covers three operations:

  • Compaction: Merges small files into larger ones
  • Prune: Removes old versions of the dataset
  • Index: Optimizes the indices, adding new data to existing indices

The frequency an application should call optimize is based on the frequency of data modifications. If data is frequently added, deleted, or updated then optimize should be run frequently. A good rule of thumb is to run optimize if you have added or modified 100,000 or more records or run more than 20 data modification operations.

Parameters

Returns

Promise<OptimizeStats>


prewarmData()

abstract prewarmData(columns?): Promise<void>

Prewarm one or more columns of data in the table.

Parameters

  • columns?: string[] The columns to prewarm. If undefined, all columns are prewarmed. This will load the column data into the page cache so that future queries that read those columns avoid the initial cold-start latency. This call initiates prewarming and returns once the request is accepted; the warming itself may continue in the background. Calling it on already-prewarmed columns is a no-op on the server. Prewarming is generally useful for columns used in filters or projections. Large columns (e.g. high-dimensional vectors or binary data) may not be practical to prewarm. This feature is currently only supported on remote tables.

Returns

Promise<void>


prewarmIndex()

abstract prewarmIndex(name): Promise<void>

Prewarm an index in the table.

Parameters

  • name: string The name of the index. This will load the index into memory. This may reduce the cold-start time for future queries. If the index does not fit in the cache then this call may be wasteful.

Returns

Promise<void>


query()

abstract query(): Query

Create a Query Builder.

Queries allow you to search your existing data. By default the query will return all the data in the table in no particular order. The builder returned by this method can be used to control the query using filtering, vector similarity, sorting, and more.

Note: By default, all columns are returned. For best performance, you should only fetch the columns you need.

When appropriate, various indices and statistics based pruning will be used to accelerate the query.

Returns

Query

A builder that can be used to parameterize the query

Examples

// SQL-style filtering
//
// This query will return up to 1000 rows whose value in the `id` column
// is greater than 5. LanceDb supports a broad set of filtering functions.
for await (const batch of table
  .query()
  .where("id > 1")
  .select(["id"])
  .limit(20)) {
  console.log(batch);
}
// Vector Similarity Search
//
// This example will find the 10 rows whose value in the "vector" column are
// closest to the query vector [1.0, 2.0, 3.0].  If an index has been created
// on the "vector" column then this will perform an ANN search.
//
// The `refineFactor` and `nprobes` methods are used to control the recall /
// latency tradeoff of the search.
for await (const batch of table
  .query()
  .where("id > 1")
  .select(["id"])
  .limit(20)) {
  console.log(batch);
}
// Scan the full dataset
//
// This query will return everything in the table in no particular order.
for await (const batch of table.query()) {
  console.log(batch);
}

refreshColumn()

abstract refreshColumn(column): Promise<RefreshColumnResult>

Fill the rows of a computed column that hold no value yet.

Rows appended since the last refresh are filled by the next one; rows already filled are left as they are, so the call is idempotent and does not observe a mutated input. Local tables only: a remote refresh runs as a server job, through Table#refreshColumnAsync.

Parameters

  • column: string The name of the computed column to fill.

Returns

Promise<RefreshColumnResult>

A promise that resolves to the number of rows filled and the new version number of the table.


refreshColumnAsync()

abstract refreshColumnAsync(column): Promise<Job>

Like Table#refreshColumn, but returns a handle to the refresh job instead of blocking until it completes.

The job may already be complete when returned; callers must not assume the column is filled until Job.wait resolves. Invalid input -- an unknown column, or one that is not computed -- rejects here rather than failing the job. On local tables the job runs in-process; on LanceDB Cloud and Enterprise it is the server's backfill job.

Parameters

  • column: string The name of the computed column to fill.

Returns

Promise<Job>

Example

const job = await table.refreshColumnAsync("doubled");
await job.wait();
console.log(await job.status()); // "finished"

restore()

abstract restore(): Promise<void>

Restore the table to the currently checked out version

This operation will fail if checkout has not been called previously

This operation will overwrite the latest version of the table with a previous version. Any changes made since the checked out version will no longer be visible.

Once the operation concludes the table will no longer be in a checked out state and the read_consistency_interval, if any, will apply.

Returns

Promise<void>


schema()

abstract schema(): Promise<Schema<any>>

Get the schema of the table.

Returns

Promise<Schema<any>>


abstract search(
   query,
   queryType?,
   ftsColumns?): Query | VectorQuery

Create a search query to find the nearest neighbors of the given query

Parameters

  • query: string | IntoVector | MultiVector | FullTextQuery the query, a vector or string

  • queryType?: string the type of the query, "vector", "fts", or "auto"

  • ftsColumns?: string | string[] the columns to search in for full text search for now, only one column can be searched at a time. when "auto" is used, if the query is a string and an embedding function is defined, it will be treated as a vector query if the query is a string and no embedding function is defined, it will be treated as a full text search query

Returns

Query | VectorQuery


setLsmWriteSpec()

abstract setLsmWriteSpec(spec): Promise<void>

Install an LsmWriteSpec on this table, selecting Lance's MemWAL LSM-style write path for future mergeInsert calls.

LsmWriteSpec chooses one of three sharding strategies via specType:

  • "bucket" — hash-bucket writes by the single-column unenforced primary key (column and numBuckets required).
  • "identity" — shard by the raw value of a scalar column.
  • "unsharded" — route every write to a single shard.

All variants require the table to have an unenforced primary key (Table#setUnenforcedPrimaryKey); bucket sharding additionally requires it to be the single column being bucketed.

Omitting maintainedIndexes maintains every index on the table, resolved here, failing if one cannot be maintained — name them to install anyway. Naming them pins an exact set, and a still-building index is rejected rather than quietly omitted.

Parameters

Returns

Promise<void>

Example

await table.setUnenforcedPrimaryKey("id");
await table.setLsmWriteSpec({
  specType: "bucket",
  column: "id",
  numBuckets: 16,
  maintainedIndexes: ["id_idx"],
});

setUnenforcedPrimaryKey()

abstract setUnenforcedPrimaryKey(columns): Promise<void>

Set the unenforced primary key for this table to a single column.

"Unenforced" means LanceDB does not check uniqueness on writes; the column is recorded in the schema as the primary key for use by features such as merge_insert. Only single-column primary keys are supported, and the key cannot be changed once set.

Parameters

  • columns: string | string[] The primary key column. A one-element array is also accepted; passing more than one column is rejected.

Returns

Promise<void>


stats()

abstract stats(): Promise<TableStatistics>

Returns table and fragment statistics

Returns

Promise<TableStatistics>

The table and fragment statistics


tags()

abstract tags(): Promise<Tags>

Get a tags manager for this table.

Tags allow you to label specific versions of a table with a human-readable name. The returned tags manager can be used to list, create, update, or delete tags.

Returns

Promise<Tags>

A tags manager for this table

Example

const tagsManager = await table.tags();
await tagsManager.create("v1", 1);
const tags = await tagsManager.list();
console.log(tags); // { "v1": { version: 1, manifestSize: ... } }

takeOffsets()

abstract takeOffsets(offsets): TakeQuery

Create a query that returns a subset of the rows in the table.

Parameters

  • offsets: number[] The offsets of the rows to return.

Returns

TakeQuery

A builder that can be used to parameterize the query.


takeRowIds()

abstract takeRowIds(rowIds): TakeQuery

Create a query that returns a subset of the rows in the table.

Parameters

  • rowIds: readonly (number | bigint)[] The row ids of the rows to return. Row ids returned by withRowId() are bigint, so bigint[] is supported. For convenience / backwards compatibility, number[] is also accepted (for small row ids that fit in a safe integer).

Returns

TakeQuery

A builder that can be used to parameterize the query.


toArrow()

abstract toArrow(): Promise<Table<any>>

Return the table as an arrow table

Returns

Promise<Table<any>>


tokenize()

abstract tokenize(query, options): Promise<FtsToken[]>

Tokenize a full-text search query using the tokenizer configured on an FTS index.

Specify exactly one of column or indexName.

Model-backed tokenizers such as jieba/* and lindera/* are rebuilt in the client process from index metadata. For remote tables, this means the same tokenizer model files must also exist locally.

Parameters

Returns

Promise<FtsToken[]>


unsetLsmWriteSpec()

abstract unsetLsmWriteSpec(): Promise<void>

Remove the LsmWriteSpec from this table, reverting to the standard mergeInsert write path.

Errors if no spec is currently set.

Returns

Promise<void>


update()

update(opts)

abstract update(opts): Promise<UpdateResult>

Update existing records in the Table

Parameters
Returns

Promise<UpdateResult>

A promise that resolves to an object containing the number of rows updated and the new version number

Example
table.update({where:"x = 2", values:{"vector": [10, 10]}})

update(opts)

abstract update(opts): Promise<UpdateResult>

Update existing records in the Table

Parameters
Returns

Promise<UpdateResult>

A promise that resolves to an object containing the number of rows updated and the new version number

Example
table.update({where:"x = 2", valuesSql:{"x": "x + 1"}})

update(updates, options)

abstract update(updates, options?): Promise<UpdateResult>

Update existing records in the Table

An update operation can be used to adjust existing values. Use the returned builder to specify which columns to update. The new value can be a literal value (e.g. replacing nulls with some default value) or an expression applied to the old value (e.g. incrementing a value)

An optional condition can be specified (e.g. "only update if the old value is 0")

Note: if your condition is something like "some_id_column == 7" and you are updating many rows (with different ids) then you will get better performance with a single [merge_insert] call instead of repeatedly calilng this method.

Parameters
  • updates: Record<string, string> | Map<string, string> the columns to update

  • options?: Partial<UpdateOptions> additional options to control the update behavior

Returns

Promise<UpdateResult>

A promise that resolves to an object containing the number of rows updated and the new version number

Keys in the map should specify the name of the column to update. Values in the map provide the new value of the column. These can be SQL literal strings (e.g. "7" or "'foo'") or they can be expressions based on the row being updated (e.g. "my_col + 1")


updateFieldMetadata()

abstract updateFieldMetadata(updates): Promise<UpdateFieldMetadataResult>

Update per-field (column) metadata.

Parameters

  • updates: FieldMetadataUpdate[] One or more per-field updates. Each update's metadata is merged into the field's existing metadata by default; a value of null deletes that key, and replace: true swaps the whole map.

Returns

Promise<UpdateFieldMetadataResult>

resolves to the new table version.


vectorSearch()

abstract vectorSearch(vector): VectorQuery

Search the table with a given query vector.

This is a convenience method for preparing a vector query and is the same thing as calling nearestTo on the builder returned by query.

Parameters

Returns

VectorQuery

See

Query#nearestTo for more details.


version()

abstract version(): Promise<number>

Retrieve the version of the table

Returns

Promise<number>


waitForIndex()

abstract waitForIndex(indexNames, timeoutSeconds): Promise<void>

Waits for asynchronous indexing to complete on the table.

Parameters

  • indexNames: string[] The name of the indices to wait for

  • timeoutSeconds: number The number of seconds to wait before timing out This will raise an error if the indices are not created and fully indexed within the timeout.

Returns

Promise<void>