* fix(promql): apply offset to subquery evaluation window
`prom_subquery_expr_to_plan` destructured `SubqueryExpr` without reading
`offset`, so `<subquery>[range:step] offset <d>` planned exactly the same
window as the un-offset form and silently returned data for the wrong time
range. The plain vector/matrix-selector paths already threaded the offset
through `selector_to_series_normalize_plan` and `RangeManipulate`.
Shift the inner evaluation window back by the offset and pass the offset to
the subquery's `RangeManipulate`, which maps the inner samples forward onto
the evaluation timeline before bucketing them into ranges. This matches
Prometheus, whose `evaluator.subqueryTimeRange` evaluates the inner
expression over `(start - offset - range, end - offset]` and whose
`evalSubquery` then hands the samples to the outer range-vector function as
a `MatrixSelector` that still carries the subquery offset. An offset on the
inner selector composes additively, as `subqueryTimes` documents.
`RangeManipulate`'s protobuf message has no offset field and recovers it on
decode from an immediately underlying `SeriesNormalize`. Since
`RangeManipulate` is commutative in `dist_plan` and can be pushed below a
`MergeScan`, insert that carrier node so the offset survives distributed
planning instead of decoding as zero.
Known divergence, unchanged by this commit: Prometheus anchors subquery step
points on absolute epoch multiples of the step, while GreptimeDB anchors them
on the evaluation start. The two agree whenever the offset is a multiple of
the subquery step; the added sqlness cases stay within that range.
Closes #9330
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Aarav <aaravsjadav@gmail.com>
* docs(promql): correct the subquery-offset rationale and pin the histogram shift
Follow-up to e132c267, which stated two things inaccurately and left one
path untested. No behaviour change beyond a conditional clone.
The `SeriesNormalize` carrier was described as existing "so the offset
survives distributed planning", implying standalone is unaffected. That is
wrong. `local_offset` is read on the substrait decode path
(`range_manipulate.rs`, `instant_manipulate.rs`), and GreptimeDB routes
PromQL plans through `MergeScan`/substrait in standalone too -- the
standalone results `promql/encode_substrait.result` and `precisions.result`
both show `MergeScan [is_placeholder=false, remote_input=[...]]`. Deleting
the carrier therefore empties subquery-with-offset results in standalone as
well, while `count_over_time_subquery_with_offset` keeps passing because it
asserts the pre-serialization plan. Say so at the node, so the next reader
does not remove it believing standalone is safe.
The carrier is also not semantically inert: `SeriesNormalizeStream::normalize`
shifts native histogram `start_timestamp` payloads by `offset`. On the
subquery path that runs on a computed inner result, on top of any offset the
inner selector already applied, and nothing pinned it. It is consistent --
`RangeManipulate` moves the millisecond time index forward by the same
`offset_ms` -- so the distance from a histogram's start timestamp to the
sample carrying it, which reset/rate detection reads, is invariant. Add
`subquery_offset_shifts_histogram_start_and_time_index_together`, which
executes the subquery's node stack over native histogram samples for a
zero and a non-zero inner offset. It cannot be a sqlness case: native
histograms are a struct column with no SQL type or literal
(`sql_data_type_to_concrete_data_type` rejects structs) and the sqlness
runner speaks only MySQL/Postgres, so such rows only arrive over
gRPC/remote-write v2.
The divergence note blamed sub-step offsets. Measured, that condition is too
narrow: an unaligned evaluation timestamp alone diverges, with no offset at
all -- `sum_over_time(fine[20s:10s])` at t=57 samples 47s and 57s here
against Prometheus's 40s and 50s. The real condition is that
`(start - offset)` is not a multiple of the subquery step, and it is
pre-existing: the committed `tql eval (359, 359, '1s')
sum_over_time(metric_total[60s:10s])` case is already an instance of it
(`359 - 60 + 10 = 309`). Reword the comment and the sqlness header
accordingly. A sub-step offset stays accepted: `foo[20s:10s] offset 5s` is
valid PromQL, erroring on it would be a regression and would not close the
gap anyway. This repo has no PromQL-compatibility documentation page, so
there is nowhere else to record it.
Also stop cloning the series key columns when there is no offset, and stop
the `offset 0s` sqlness comment implying its parser error is evidence about
this fix -- the same message comes from a `u32` overflow in
`promql-parser`'s shared duration check (`offset 9999999999d`).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Aarav <aaravsjadav@gmail.com>
* refactor(promql): trim subquery-offset comments and tests per review
Shorten the planner and sqlness comments to the semantic contract, build
`SeriesDivide` with a clone instead of an `Option`, and drop the
normalize.rs histogram test and the `offset 0s` parser case.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bvab9gpXYjKyLN8Mo6JWWL
Signed-off-by: Aarav <aaravsjadav@gmail.com>
* fix(promql): record the subquery offset as the range fold offset
The subquery's `RangeManipulate` now shifts the payload timestamps by the
offset, so `predict_linear` over an offset subquery must recover the
evaluation time with that offset instead of 0.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bvab9gpXYjKyLN8Mo6JWWL
Signed-off-by: Aarav <aaravsjadav@gmail.com>
---------
Signed-off-by: Aarav <aaravsjadav@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sqlness Test
Sqlness manual
Case file
Sqlness has two types of file:
.sql: test input, SQL only.result: expected test output, SQL and its results
.result is the output (execution result) file. If you see .result files is changed,
it means this test gets a different result and indicates it fails. You should
check change logs to solve the problem.
You only need to write test SQL in .sql file, and run the test.
Case organization
The root dir of input cases is tests/cases. It contains several subdirectories stand for different test
modes. E.g., standalone/ contains all the tests to run under greptimedb standalone start mode.
Under the first level of subdirectory (e.g. the cases/standalone), you can organize your cases as you like.
Sqlness walks through every file recursively and runs them.
Kafka WAL
Sqlness supports Kafka WAL. You can either provide a Kafka cluster or let sqlness to start one for you.
To run test with kafka, you need to pass the option -w kafka. If no other options are provided, sqlness will use conf/kafka-cluster.yml to start a Kafka cluster. This requires docker and docker-compose commands in your environment.
Otherwise, you can additionally pass the your existing kafka environment to sqlness with -k option. E.g.:
cargo sqlness bare -w kafka -k localhost:9092
In this case, sqlness will not start its own kafka cluster and the one you provided instead.
Run the test
Unlike other tests, this harness is in a binary target form. You can run it with:
cargo sqlness bare
It automatically finishes the following procedures: compile GreptimeDB, start it, grab tests and feed it to
the server, then collect and compare the results. You only need to check if the .result files are changed.
If not, congratulations, the test is passed 🥳!