mirror of
https://github.com/GreptimeTeam/greptimedb.git
synced 2026-10-03 02:25:35 +00:00
* fix(promql): apply offset to subquery evaluation window
`prom_subquery_expr_to_plan` destructured `SubqueryExpr` without reading
`offset`, so `<subquery>[range:step] offset <d>` planned exactly the same
window as the un-offset form and silently returned data for the wrong time
range. The plain vector/matrix-selector paths already threaded the offset
through `selector_to_series_normalize_plan` and `RangeManipulate`.
Shift the inner evaluation window back by the offset and pass the offset to
the subquery's `RangeManipulate`, which maps the inner samples forward onto
the evaluation timeline before bucketing them into ranges. This matches
Prometheus, whose `evaluator.subqueryTimeRange` evaluates the inner
expression over `(start - offset - range, end - offset]` and whose
`evalSubquery` then hands the samples to the outer range-vector function as
a `MatrixSelector` that still carries the subquery offset. An offset on the
inner selector composes additively, as `subqueryTimes` documents.
`RangeManipulate`'s protobuf message has no offset field and recovers it on
decode from an immediately underlying `SeriesNormalize`. Since
`RangeManipulate` is commutative in `dist_plan` and can be pushed below a
`MergeScan`, insert that carrier node so the offset survives distributed
planning instead of decoding as zero.
Known divergence, unchanged by this commit: Prometheus anchors subquery step
points on absolute epoch multiples of the step, while GreptimeDB anchors them
on the evaluation start. The two agree whenever the offset is a multiple of
the subquery step; the added sqlness cases stay within that range.
Closes #9330
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Aarav <aaravsjadav@gmail.com>
* docs(promql): correct the subquery-offset rationale and pin the histogram shift
Follow-up to e132c267, which stated two things inaccurately and left one
path untested. No behaviour change beyond a conditional clone.
The `SeriesNormalize` carrier was described as existing "so the offset
survives distributed planning", implying standalone is unaffected. That is
wrong. `local_offset` is read on the substrait decode path
(`range_manipulate.rs`, `instant_manipulate.rs`), and GreptimeDB routes
PromQL plans through `MergeScan`/substrait in standalone too -- the
standalone results `promql/encode_substrait.result` and `precisions.result`
both show `MergeScan [is_placeholder=false, remote_input=[...]]`. Deleting
the carrier therefore empties subquery-with-offset results in standalone as
well, while `count_over_time_subquery_with_offset` keeps passing because it
asserts the pre-serialization plan. Say so at the node, so the next reader
does not remove it believing standalone is safe.
The carrier is also not semantically inert: `SeriesNormalizeStream::normalize`
shifts native histogram `start_timestamp` payloads by `offset`. On the
subquery path that runs on a computed inner result, on top of any offset the
inner selector already applied, and nothing pinned it. It is consistent --
`RangeManipulate` moves the millisecond time index forward by the same
`offset_ms` -- so the distance from a histogram's start timestamp to the
sample carrying it, which reset/rate detection reads, is invariant. Add
`subquery_offset_shifts_histogram_start_and_time_index_together`, which
executes the subquery's node stack over native histogram samples for a
zero and a non-zero inner offset. It cannot be a sqlness case: native
histograms are a struct column with no SQL type or literal
(`sql_data_type_to_concrete_data_type` rejects structs) and the sqlness
runner speaks only MySQL/Postgres, so such rows only arrive over
gRPC/remote-write v2.
The divergence note blamed sub-step offsets. Measured, that condition is too
narrow: an unaligned evaluation timestamp alone diverges, with no offset at
all -- `sum_over_time(fine[20s:10s])` at t=57 samples 47s and 57s here
against Prometheus's 40s and 50s. The real condition is that
`(start - offset)` is not a multiple of the subquery step, and it is
pre-existing: the committed `tql eval (359, 359, '1s')
sum_over_time(metric_total[60s:10s])` case is already an instance of it
(`359 - 60 + 10 = 309`). Reword the comment and the sqlness header
accordingly. A sub-step offset stays accepted: `foo[20s:10s] offset 5s` is
valid PromQL, erroring on it would be a regression and would not close the
gap anyway. This repo has no PromQL-compatibility documentation page, so
there is nowhere else to record it.
Also stop cloning the series key columns when there is no offset, and stop
the `offset 0s` sqlness comment implying its parser error is evidence about
this fix -- the same message comes from a `u32` overflow in
`promql-parser`'s shared duration check (`offset 9999999999d`).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Aarav <aaravsjadav@gmail.com>
* refactor(promql): trim subquery-offset comments and tests per review
Shorten the planner and sqlness comments to the semantic contract, build
`SeriesDivide` with a clone instead of an `Option`, and drop the
normalize.rs histogram test and the `offset 0s` parser case.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bvab9gpXYjKyLN8Mo6JWWL
Signed-off-by: Aarav <aaravsjadav@gmail.com>
* fix(promql): record the subquery offset as the range fold offset
The subquery's `RangeManipulate` now shifts the payload timestamps by the
offset, so `predict_linear` over an offset subquery must recover the
evaluation time with that offset instead of 0.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bvab9gpXYjKyLN8Mo6JWWL
Signed-off-by: Aarav <aaravsjadav@gmail.com>
---------
Signed-off-by: Aarav <aaravsjadav@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>