Files
lancedb/docs/src/js/classes/Job.md
T
Jack Ye 21f11b4463 feat!: replace get_job/job_history with describe_job/query_job_events (#4130)
A 1M-row column refresh over 200 fragments produced no visible result,
and the client could only ever say `"running"`. Everything needed to
diagnose it already existed server-side — the job registry records a
`claim`/`claim_complete` pair per fragment carrying `rows_processed` —
but none of it was reachable.

## Before

Four ways to ask about a job, none of which told you much.

```python
job = table.refresh_column_async("embedding")
job.status()               # "running". That was the entire debug surface.
db.get_job(job_id)         # state, and a spec. No result, no progress.
db.job_history(job_id)     # raw record batches, no limit, no filter
db.job(job_id)             # a handle that knew nothing
```

## After

Open a job the way you open a table; the handle answers everything.

```python
job = db.open_job(job_id)      # raises JobNotFoundError if there is no such job
```

```python
>>> print(job)
Job(
    id='job-1',
    state='failed',
    job_type='refresh_column',
    creation_ms=1757000000000,
    spec={
        "column": "embedding",
        "num_workers": 4
    },
    failure=JobFailureInfo(phase='execute', message='worker died', retryable=True),
)
```

Individual fields are there too — `job.state`, `job.job_type`,
`job.creation_ms`, `job.spec`, `job.result`, `job.failure` — and
`job.result` carries `rows_assigned` / `rows_failed` as soon as the job
succeeds, with no `wait()` required.

Per-fragment progress *while it is still running*:

```python
done = job.events(filter="state = 'claim_complete'", limit=10_000)
done.column("rows_processed").to_pylist()      # [5000, 5000, ...]
```

The handle an async action returns is the same object, one `refresh()`
away:

```python
job = table.refresh_column_async("embedding")
job.refresh()
job.state, job.result
```

TypeScript is the same experience, down to `console.log`:

```ts
const job = await db.openJob(jobId);   // rejects if there is no such job
console.log(job);                      // same multi-line layout
job.state; job.jobType; job.spec; job.result; job.failure;
const done = await job.events({ filter: "state = 'claim_complete'", limit: 10_000 });
```

## Why each piece matters

- **A result without waiting.** `rows_assigned` / `rows_failed` used to
live only on the terminal result, so a job that never terminated
reported nothing at all.
- **`limit`.** The server caps event rows at 1000 and truncates without
saying so, which silently hid most of a 200-fragment job's history.
- **`filter`.** `claim_complete` rows carry per-claim `rows_processed` —
the only progress signal that exists mid-flight.
- **Events outlive the worker.** They live in the job registry, not in
pod logs that vanish with the pod.
- **One place to ask.** `open_job` replaces `describe_job`,
`query_job_events` and `job`, so a question about a job has one answer
instead of one per calling location.
- **A missing job is an error, not a `None`.** The common case is a job
id copied out of a log, where absence is the surprise worth raising —
and it matches `open_table`.
- **Printing is the debug surface.** Every field on its own line, JSON
payloads keeping their structure. An unrefreshed handle stays on one
line, because there is nothing to lay out.
- **In-process jobs say so.** A local refresh reports `state` and leaves
the rest null rather than inventing fields it has no record for.

`list_jobs` and `cancel_job` stay as they were: one lists, the other is
a one-shot action that should not need a describe first.

## Breaking

All shipped in 0.38.0. No deprecated aliases.

| Was | Now |
| --- | --- |
| `Connection.get_job` → `describe_job` | `Connection.open_job` returns
a populated `Job`, or raises |
| `Connection.job_history` → `query_job_events` | `job.events(...)` |
| `Connection.job` | `Connection.open_job` |
| Python events → `List[pa.RecordBatch]` | `pa.Table` |
| `JobDescription.spec_json` / `.result_json` | internal; use `job.spec`
/ `job.result` |

Node's `Job` is now a TypeScript class wrapping the native handle, so it
returns an Arrow table and parsed values like Python does. New
`Error::JobNotFound` / `JobNotFoundError`; the three job exceptions are
now in the Python API reference.
2026-09-05 16:24:23 -07:00

3.9 KiB

@lancedb/lancedbDocs


@lancedb/lancedb / Job

Class: Job

A handle to an operation that may still be running.

The operation may already be complete when the handle is created.

The detail getters read what the handle last observed. Submitting an operation returns only a job id, so populating them eagerly would cost an extra round trip on every call:

  • Job.refresh and Job.status fetch the whole record.
  • Job.wait records the terminal state it establishes, but not the rest of the record.
  • Everything is null until one of those runs.

Accessors

creationMs

get creationMs(): null | number

When the job was created, in milliseconds since the epoch.

Returns

null | number


failure

get failure(): null | JobFailureInfo

Why the job failed, when it failed and the server reports a reason.

Returns

null | JobFailureInfo


id

get id(): null | string

Identifies the operation on the server that is running it.

Operations that run in this process have no server id. The value is opaque: parsing it or storing it to resume the job later is not supported.

Returns

null | string


jobType

get jobType(): null | string

The job's type, as the server names it. Null for an in-process job, which has no server-side record.

Returns

null | string


result

get result(): any

The job-type-specific terminal result. Null until the job succeeds, so a job that never terminates reports its progress through Job.events instead.

Returns

any


spec

get spec(): any

The job-type-specific specification it was submitted with.

Returns

any


state

get state(): null | string

The last observed lifecycle state, without contacting the backend.

Returns

null | string

Methods

cancel()

cancel(): Promise<void>

Request cancellation. Cancelling a finished operation is a no-op.

Returns

Promise<void>


events()

events(options?): Promise<Table<any>>

This job's recorded lifecycle events.

Where the getters above report a terminal result only once the job reaches one, events are written as the job runs and outlive the workers that produced them. A distributed job records a claim/claim_complete pair per unit of work, each carrying rows_processed, so a job that never finishes still accounts for what it did.

The server caps results at 1000 rows by default and 10,000 at most, and truncates without saying so, so pass limit for a job that emits an event per fragment. filter is a SQL-like expression over the state, updated_by, emitted_from, emitted_by, and claim_entity columns.

Parameters

Returns

Promise<Table<any>>


refresh()

refresh(): Promise<void>

Ask the backend for this job's current state, and for a server-side job its full record, then cache it for the getters above.

Returns

Promise<void>


status()

status(): Promise<string>

The operation's current lifecycle state: "running", "finished", "failed", or "cancelled".

A point snapshot; unlike Job.wait it does not block or reject on a terminal failure state. Also refreshes the getters above.

Returns

Promise<string>


toString()

toString(): string

Every field the handle currently knows, one per line, with the JSON payloads indented -- a refresh job's spec and result are the point of printing it.

Returns

string


wait()

wait(): Promise<void>

Wait until the operation reaches a terminal state.

Returns

Promise<void>