A 1M-row column refresh over 200 fragments produced no visible result,
and the client could only ever say `"running"`. Everything needed to
diagnose it already existed server-side — the job registry records a
`claim`/`claim_complete` pair per fragment carrying `rows_processed` —
but none of it was reachable.
## Before
Four ways to ask about a job, none of which told you much.
```python
job = table.refresh_column_async("embedding")
job.status() # "running". That was the entire debug surface.
db.get_job(job_id) # state, and a spec. No result, no progress.
db.job_history(job_id) # raw record batches, no limit, no filter
db.job(job_id) # a handle that knew nothing
```
## After
Open a job the way you open a table; the handle answers everything.
```python
job = db.open_job(job_id) # raises JobNotFoundError if there is no such job
```
```python
>>> print(job)
Job(
id='job-1',
state='failed',
job_type='refresh_column',
creation_ms=1757000000000,
spec={
"column": "embedding",
"num_workers": 4
},
failure=JobFailureInfo(phase='execute', message='worker died', retryable=True),
)
```
Individual fields are there too — `job.state`, `job.job_type`,
`job.creation_ms`, `job.spec`, `job.result`, `job.failure` — and
`job.result` carries `rows_assigned` / `rows_failed` as soon as the job
succeeds, with no `wait()` required.
Per-fragment progress *while it is still running*:
```python
done = job.events(filter="state = 'claim_complete'", limit=10_000)
done.column("rows_processed").to_pylist() # [5000, 5000, ...]
```
The handle an async action returns is the same object, one `refresh()`
away:
```python
job = table.refresh_column_async("embedding")
job.refresh()
job.state, job.result
```
TypeScript is the same experience, down to `console.log`:
```ts
const job = await db.openJob(jobId); // rejects if there is no such job
console.log(job); // same multi-line layout
job.state; job.jobType; job.spec; job.result; job.failure;
const done = await job.events({ filter: "state = 'claim_complete'", limit: 10_000 });
```
## Why each piece matters
- **A result without waiting.** `rows_assigned` / `rows_failed` used to
live only on the terminal result, so a job that never terminated
reported nothing at all.
- **`limit`.** The server caps event rows at 1000 and truncates without
saying so, which silently hid most of a 200-fragment job's history.
- **`filter`.** `claim_complete` rows carry per-claim `rows_processed` —
the only progress signal that exists mid-flight.
- **Events outlive the worker.** They live in the job registry, not in
pod logs that vanish with the pod.
- **One place to ask.** `open_job` replaces `describe_job`,
`query_job_events` and `job`, so a question about a job has one answer
instead of one per calling location.
- **A missing job is an error, not a `None`.** The common case is a job
id copied out of a log, where absence is the surprise worth raising —
and it matches `open_table`.
- **Printing is the debug surface.** Every field on its own line, JSON
payloads keeping their structure. An unrefreshed handle stays on one
line, because there is nothing to lay out.
- **In-process jobs say so.** A local refresh reports `state` and leaves
the rest null rather than inventing fields it has no record for.
`list_jobs` and `cancel_job` stay as they were: one lists, the other is
a one-shot action that should not need a describe first.
## Breaking
All shipped in 0.38.0. No deprecated aliases.
| Was | Now |
| --- | --- |
| `Connection.get_job` → `describe_job` | `Connection.open_job` returns
a populated `Job`, or raises |
| `Connection.job_history` → `query_job_events` | `job.events(...)` |
| `Connection.job` | `Connection.open_job` |
| Python events → `List[pa.RecordBatch]` | `pa.Table` |
| `JobDescription.spec_json` / `.result_json` | internal; use `job.spec`
/ `job.result` |
Node's `Job` is now a TypeScript class wrapping the native handle, so it
returns an Arrow table and parsed values like Python does. New
`Error::JobNotFound` / `JobNotFoundError`; the three job exceptions are
now in the Python API reference.
3.9 KiB
@lancedb/lancedb • Docs
@lancedb/lancedb / Job
Class: Job
A handle to an operation that may still be running.
The operation may already be complete when the handle is created.
The detail getters read what the handle last observed. Submitting an operation returns only a job id, so populating them eagerly would cost an extra round trip on every call:
- Job.refresh and Job.status fetch the whole record.
- Job.wait records the terminal state it establishes, but not the rest of the record.
- Everything is null until one of those runs.
Accessors
creationMs
get creationMs(): null | number
When the job was created, in milliseconds since the epoch.
Returns
null | number
failure
get failure(): null | JobFailureInfo
Why the job failed, when it failed and the server reports a reason.
Returns
null | JobFailureInfo
id
get id(): null | string
Identifies the operation on the server that is running it.
Operations that run in this process have no server id. The value is opaque: parsing it or storing it to resume the job later is not supported.
Returns
null | string
jobType
get jobType(): null | string
The job's type, as the server names it. Null for an in-process job, which has no server-side record.
Returns
null | string
result
get result(): any
The job-type-specific terminal result. Null until the job succeeds, so a job that never terminates reports its progress through Job.events instead.
Returns
any
spec
get spec(): any
The job-type-specific specification it was submitted with.
Returns
any
state
get state(): null | string
The last observed lifecycle state, without contacting the backend.
Returns
null | string
Methods
cancel()
cancel(): Promise<void>
Request cancellation. Cancelling a finished operation is a no-op.
Returns
Promise<void>
events()
events(options?): Promise<Table<any>>
This job's recorded lifecycle events.
Where the getters above report a terminal result only once the job reaches
one, events are written as the job runs and outlive the workers that
produced them. A distributed job records a claim/claim_complete pair
per unit of work, each carrying rows_processed, so a job that never
finishes still accounts for what it did.
The server caps results at 1000 rows by default and 10,000 at most, and
truncates without saying so, so pass limit for a job that emits an event
per fragment. filter is a SQL-like expression over the state,
updated_by, emitted_from, emitted_by, and claim_entity columns.
Parameters
- options?:
JobEventsOptions
Returns
Promise<Table<any>>
refresh()
refresh(): Promise<void>
Ask the backend for this job's current state, and for a server-side job its full record, then cache it for the getters above.
Returns
Promise<void>
status()
status(): Promise<string>
The operation's current lifecycle state: "running", "finished", "failed", or "cancelled".
A point snapshot; unlike Job.wait it does not block or reject on a terminal failure state. Also refreshes the getters above.
Returns
Promise<string>
toString()
toString(): string
Every field the handle currently knows, one per line, with the JSON payloads indented -- a refresh job's spec and result are the point of printing it.
Returns
string
wait()
wait(): Promise<void>
Wait until the operation reaches a terminal state.
Returns
Promise<void>