mirror of
https://github.com/lancedb/lancedb.git
synced 2026-09-06 21:39:03 +00:00
A 1M-row column refresh over 200 fragments produced no visible result,
and the client could only ever say `"running"`. Everything needed to
diagnose it already existed server-side — the job registry records a
`claim`/`claim_complete` pair per fragment carrying `rows_processed` —
but none of it was reachable.
## Before
Four ways to ask about a job, none of which told you much.
```python
job = table.refresh_column_async("embedding")
job.status() # "running". That was the entire debug surface.
db.get_job(job_id) # state, and a spec. No result, no progress.
db.job_history(job_id) # raw record batches, no limit, no filter
db.job(job_id) # a handle that knew nothing
```
## After
Open a job the way you open a table; the handle answers everything.
```python
job = db.open_job(job_id) # raises JobNotFoundError if there is no such job
```
```python
>>> print(job)
Job(
id='job-1',
state='failed',
job_type='refresh_column',
creation_ms=1757000000000,
spec={
"column": "embedding",
"num_workers": 4
},
failure=JobFailureInfo(phase='execute', message='worker died', retryable=True),
)
```
Individual fields are there too — `job.state`, `job.job_type`,
`job.creation_ms`, `job.spec`, `job.result`, `job.failure` — and
`job.result` carries `rows_assigned` / `rows_failed` as soon as the job
succeeds, with no `wait()` required.
Per-fragment progress *while it is still running*:
```python
done = job.events(filter="state = 'claim_complete'", limit=10_000)
done.column("rows_processed").to_pylist() # [5000, 5000, ...]
```
The handle an async action returns is the same object, one `refresh()`
away:
```python
job = table.refresh_column_async("embedding")
job.refresh()
job.state, job.result
```
TypeScript is the same experience, down to `console.log`:
```ts
const job = await db.openJob(jobId); // rejects if there is no such job
console.log(job); // same multi-line layout
job.state; job.jobType; job.spec; job.result; job.failure;
const done = await job.events({ filter: "state = 'claim_complete'", limit: 10_000 });
```
## Why each piece matters
- **A result without waiting.** `rows_assigned` / `rows_failed` used to
live only on the terminal result, so a job that never terminated
reported nothing at all.
- **`limit`.** The server caps event rows at 1000 and truncates without
saying so, which silently hid most of a 200-fragment job's history.
- **`filter`.** `claim_complete` rows carry per-claim `rows_processed` —
the only progress signal that exists mid-flight.
- **Events outlive the worker.** They live in the job registry, not in
pod logs that vanish with the pod.
- **One place to ask.** `open_job` replaces `describe_job`,
`query_job_events` and `job`, so a question about a job has one answer
instead of one per calling location.
- **A missing job is an error, not a `None`.** The common case is a job
id copied out of a log, where absence is the surprise worth raising —
and it matches `open_table`.
- **Printing is the debug surface.** Every field on its own line, JSON
payloads keeping their structure. An unrefreshed handle stays on one
line, because there is nothing to lay out.
- **In-process jobs say so.** A local refresh reports `state` and leaves
the rest null rather than inventing fields it has no record for.
`list_jobs` and `cancel_job` stay as they were: one lists, the other is
a one-shot action that should not need a describe first.
## Breaking
All shipped in 0.38.0. No deprecated aliases.
| Was | Now |
| --- | --- |
| `Connection.get_job` → `describe_job` | `Connection.open_job` returns
a populated `Job`, or raises |
| `Connection.job_history` → `query_job_events` | `job.events(...)` |
| `Connection.job` | `Connection.open_job` |
| Python events → `List[pa.RecordBatch]` | `pa.Table` |
| `JobDescription.spec_json` / `.result_json` | internal; use `job.spec`
/ `job.result` |
Node's `Job` is now a TypeScript class wrapping the native handle, so it
returns an Arrow table and parsed values like Python does. New
`Error::JobNotFound` / `JobNotFoundError`; the three job exceptions are
now in the Python API reference.
183 lines
5.8 KiB
Rust
183 lines
5.8 KiB
Rust
// SPDX-License-Identifier: Apache-2.0
|
|
// SPDX-FileCopyrightText: Copyright The LanceDB Authors
|
|
|
|
use std::sync::Arc;
|
|
|
|
use arrow_array::RecordBatch;
|
|
use lancedb::job::JobEventsRequest;
|
|
use napi::bindgen_prelude::Buffer;
|
|
use napi_derive::napi;
|
|
|
|
use crate::error::NapiErrorExt;
|
|
|
|
/// A handle to an operation that may still be running.
|
|
#[napi]
|
|
pub struct Job {
|
|
inner: Arc<lancedb::Job>,
|
|
}
|
|
|
|
impl Job {
|
|
pub(crate) fn new<T>(inner: lancedb::Job<T>) -> Self
|
|
where
|
|
T: Clone + Send + Sync + 'static,
|
|
{
|
|
Self {
|
|
inner: Arc::new(inner.map(|_| ())),
|
|
}
|
|
}
|
|
}
|
|
|
|
#[napi]
|
|
impl Job {
|
|
/// Identifies the operation on the server that is running it. Operations
|
|
/// that run in this process have no server id. The value is opaque.
|
|
#[napi(getter)]
|
|
pub fn id(&self) -> Option<String> {
|
|
self.inner.id().map(str::to_string)
|
|
}
|
|
|
|
/// The operation's current lifecycle state: "running", "finished",
|
|
/// "failed", or "cancelled".
|
|
///
|
|
/// A point snapshot; unlike {@link Job.wait} it does not block or reject
|
|
/// on a terminal failure state. States a newer server reports that this
|
|
/// client version does not know pass through as-is.
|
|
#[napi(catch_unwind)]
|
|
pub async fn status(&self) -> napi::Result<String> {
|
|
self.inner.status().await.default_error()
|
|
}
|
|
|
|
/// Wait until the operation reaches a terminal state.
|
|
#[napi(catch_unwind)]
|
|
pub async fn wait(&self) -> napi::Result<()> {
|
|
self.inner.wait().await.default_error()
|
|
}
|
|
|
|
/// Request cancellation. Cancelling a finished operation is a no-op.
|
|
#[napi(catch_unwind)]
|
|
pub async fn cancel(&self) -> napi::Result<()> {
|
|
self.inner.cancel().await.default_error()
|
|
}
|
|
|
|
/// Ask the backend for this job's current state, and for a server-side job
|
|
/// its full record, then cache it for the getters below.
|
|
///
|
|
/// They are all null until this runs, because submitting an operation
|
|
/// returns only a job id. {@link Job.status} fetches the whole record too;
|
|
/// {@link Job.wait} records only the terminal state it establishes.
|
|
#[napi(catch_unwind)]
|
|
pub async fn refresh(&self) -> napi::Result<()> {
|
|
self.inner.refresh().await.default_error()
|
|
}
|
|
|
|
/// The last observed lifecycle state, without contacting the backend.
|
|
#[napi(getter)]
|
|
pub fn state(&self) -> Option<String> {
|
|
self.inner.state()
|
|
}
|
|
|
|
/// The job's type, as the server names it. Null for an in-process job,
|
|
/// which has no server-side record.
|
|
#[napi(getter)]
|
|
pub fn job_type(&self) -> Option<String> {
|
|
self.inner.job_type()
|
|
}
|
|
|
|
/// When the job was created, in milliseconds since the epoch.
|
|
#[napi(getter)]
|
|
pub fn creation_ms(&self) -> Option<i64> {
|
|
self.inner.creation_ms()
|
|
}
|
|
|
|
/// The job-type-specific specification as a JSON string, when present.
|
|
#[napi(getter)]
|
|
pub fn spec_json(&self) -> Option<String> {
|
|
self.inner.spec().map(|spec| spec.to_string())
|
|
}
|
|
|
|
/// The job-type-specific terminal result as a JSON string. Null until the
|
|
/// job succeeds, so a job that never terminates reports its progress
|
|
/// through {@link Job.events} instead.
|
|
#[napi(getter)]
|
|
pub fn result_json(&self) -> Option<String> {
|
|
self.inner.result().map(|result| result.to_string())
|
|
}
|
|
|
|
/// Why the job failed, when it failed and the server reports a reason.
|
|
#[napi(getter)]
|
|
pub fn failure(&self) -> Option<JobFailureInfo> {
|
|
self.inner.failure().map(|failure| JobFailureInfo {
|
|
phase: failure.phase,
|
|
message: failure.message,
|
|
retryable: failure.retryable,
|
|
})
|
|
}
|
|
|
|
/// This job's recorded lifecycle events, as an Arrow IPC stream buffer.
|
|
/// The TypeScript wrapper turns it into an Arrow table.
|
|
#[napi(catch_unwind)]
|
|
pub async fn events(&self, limit: Option<u32>, filter: Option<String>) -> napi::Result<Buffer> {
|
|
let batches = self
|
|
.inner
|
|
.events(JobEventsRequest { limit, filter })
|
|
.await
|
|
.default_error()?;
|
|
batches_to_ipc_buffer(&batches)
|
|
}
|
|
}
|
|
|
|
/// Serialise Arrow batches as a single IPC stream for the TypeScript layer.
|
|
pub(crate) fn batches_to_ipc_buffer(batches: &[RecordBatch]) -> napi::Result<Buffer> {
|
|
let Some(first) = batches.first() else {
|
|
return Ok(Buffer::from(Vec::<u8>::new()));
|
|
};
|
|
let mut out = Vec::new();
|
|
let mut writer = arrow_ipc::writer::StreamWriter::try_new(&mut out, &first.schema())
|
|
.map_err(|e| napi::Error::from_reason(e.to_string()))?;
|
|
for batch in batches {
|
|
writer
|
|
.write(batch)
|
|
.map_err(|e| napi::Error::from_reason(e.to_string()))?;
|
|
}
|
|
writer
|
|
.finish()
|
|
.map_err(|e| napi::Error::from_reason(e.to_string()))?;
|
|
drop(writer);
|
|
Ok(Buffer::from(out))
|
|
}
|
|
|
|
/// A row from `Connection.listJobs`: one server-side job.
|
|
#[napi(object)]
|
|
pub struct JobInfo {
|
|
/// The job id -- what `Connection.openJob` and `Connection.cancelJob`
|
|
/// accept.
|
|
pub job_id: String,
|
|
/// The table the job runs against, without URI or namespace.
|
|
pub table: String,
|
|
pub job_type: String,
|
|
/// Lifecycle state: "running", "finished", "failed", or "cancelled".
|
|
pub state: String,
|
|
/// When the job was created, in milliseconds since the epoch.
|
|
pub created_at_millis: i64,
|
|
}
|
|
|
|
impl From<lancedb::database::JobInfo> for JobInfo {
|
|
fn from(info: lancedb::database::JobInfo) -> Self {
|
|
Self {
|
|
job_id: info.job_id,
|
|
table: info.table,
|
|
job_type: info.job_type,
|
|
state: info.state,
|
|
created_at_millis: info.created_at_millis,
|
|
}
|
|
}
|
|
}
|
|
|
|
/// The server's account of why a job failed.
|
|
#[napi(object)]
|
|
pub struct JobFailureInfo {
|
|
pub phase: Option<String>,
|
|
pub message: Option<String>,
|
|
pub retryable: Option<bool>,
|
|
}
|