mirror of
https://github.com/GreptimeTeam/greptimedb.git
synced 2026-10-06 03:52:28 +00:00
* fix(promql): derive vector matching result labels and reject ambiguous matchings A vector-vector binary operation projected one operand's whole tag set and inner-joined without any cardinality check, so `on()`/`ignoring()` did not reduce the result labels, `group_left`/`group_right` changed nothing, and a non-unique match group produced a cross product that PromQL cannot represent. Result labels now follow Prometheus `resultMetric`: `on(...)` keeps the matching labels, `ignoring(...)` drops them, and a group modifier keeps the many side's labels plus the `group_x(...)` labels taken from the one side. A label the one side does not carry is deleted from the result. The reduced label set no longer identifies the operand series, so `__tsid` is dropped from the context on this path. Cardinality is enforced with a `count(1) OVER (PARTITION BY match keys, ts)` window and a scalar UDF that fails the query on a repeated group: on the one side before the join, and on the result labels after it, matching where Prometheus raises each of its three errors. Series are unique by their whole tag set, so the window is only planted when the match keys drop a tag; plain arithmetic and `on(<all tags>)` plan exactly as before. Closes #9207, closes #9208, closes #9209. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * perf(promql): keep __tsid when the result labels are an operand's whole tag set Deriving the result labels dropped `__tsid` from the context unconditionally, so an enclosing operation fell back to joining on the tag columns even where the column still identified the result series. Keep it when every result label comes from one operand and covers that operand's whole tag set: no other operand value reaches the labels, and the matching gives each of its rows a single partner, so its `__tsid` is still one per result series. That is the common `on(<all tags>)` and bare `group_left` shape; a matching that actually drops a tag still clears it. The column is re-qualified as the result's own, which is how the enclosing expression and the context look it up. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix(promql): keep the match group count column unambiguous The cardinality check aliased its row count to a fixed `__promql_match_group_count`. An operand carrying a label of that name made the window output two fields with the same name, and planning failed with "Schema contains qualified field name collide_right.__promql_match_group_count and unqualified field name __promql_match_group_count which would be ambiguous". Pick a name the operand does not already have, the way the `or` operator allocates its match key columns. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test(promql): cover a match group spread over several regions The cardinality check runs above the merge of the region scans, so it counts a match group globally. Nothing pinned that: every table in these cases holds a single region, and a check evaluated per region would pass them all. Partition the operand on a column outside the match keys, which puts the two series of one match group in different regions, and assert both the pre-join and the post-join check still reject it. The case runs in the distributed environment too, where the regions sit on different datanodes. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix(promql): group an outer aggregate on the labels the operand kept `by`/`without` planning topped up a missing grouping column by walking down to the table scan and re-projecting it. That is right for a column the scan pruned for efficiency, but the labels a matching modifier deletes are also absent from the operand's output, and they were restored the same way: sum without(host) (a / on(host) b) `on(host)` leaves the operand with `host` alone, so the sum covers everything and Prometheus answers `{} 10`. Instead `device` came back from the scan under `a` and split the result into `{device="d1"} 5` and `{device="d2"} 5`. Same for `sum by(device)` of that operand, which has no `device` to group on at all. Take the grouping labels from the operand's own label set rather than from the row keys of the scan beneath it. A label pruned from the plan is still in that set and still gets restored; a label the operand dropped is not. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * refactor(promql): drop the unreachable aggregation tag top-up `by`/`without` planning could restore a grouping column that the plan no longer carried by rewriting the scan underneath it. Once the grouping labels come from the operand's own label set, there is nothing left for it to restore: a scan projects every label of `ctx.tag_columns` (`scan_tag_columns` only ever adds matcher columns to that set), so a label in the set is always in the schema. Stubbing the rewriter to a no-op passed the whole sqlness suite, in both the standalone and the distributed environment. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test(promql): drop cases that guarded the removed tag top-up Three plain selector aggregates were there to show that restoring a pruned grouping column still worked. With the restore gone they only repeat what the aggregate cases already cover. Also fix a comment that still said the metric engine scan prunes tag columns: it projects every label of the operand. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * docs(promql): note the duplicate a propagated matcher hides A matcher copied onto the one-side operand removes groups without a partner before the cardinality check sees them, so a duplicate in such a group is not reported. Prometheus checks every group of the one side and fails the query. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> --------- Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
222 lines
6.7 KiB
SQL
222 lines
6.7 KiB
SQL
CREATE TABLE counter_metric (
|
|
host STRING NULL,
|
|
device STRING NULL,
|
|
ts TIMESTAMP(3) TIME INDEX,
|
|
greptime_value DOUBLE,
|
|
PRIMARY KEY(host, device)
|
|
);
|
|
|
|
CREATE TABLE gauge_metric (
|
|
host STRING NULL,
|
|
device STRING NULL,
|
|
ts TIMESTAMP(3) TIME INDEX,
|
|
greptime_value DOUBLE,
|
|
PRIMARY KEY(host, device)
|
|
);
|
|
|
|
INSERT INTO counter_metric VALUES
|
|
('host1', 'eth0', 0, 10), ('host1', 'eth1', 0, 20), ('host2', 'eth0', 0, 30);
|
|
|
|
INSERT INTO gauge_metric VALUES
|
|
('host1', 'eth0', 0, 2), ('host1', 'eth1', 0, 4), ('host2', 'eth0', 0, 5);
|
|
|
|
-- Matching on the whole tag set keeps it.
|
|
-- SQLNESS SORT_RESULT 3 1
|
|
TQL EVAL (0, 0, '5s') counter_metric / gauge_metric;
|
|
|
|
-- `on(...)` reduces the result to the matching labels.
|
|
-- SQLNESS SORT_RESULT 3 1
|
|
TQL EVAL (0, 0, '5s') counter_metric{device="eth0"} / on(host) gauge_metric{device="eth0"};
|
|
|
|
-- `ignoring(...)` drops the ignored labels, here from operands that don't even share them.
|
|
-- SQLNESS SORT_RESULT 3 1
|
|
TQL EVAL (0, 0, '5s') counter_metric{device="eth0"} / ignoring(device) gauge_metric{device="eth1"};
|
|
|
|
-- `on()` matches everything against everything, so both sides must hold a single series.
|
|
-- SQLNESS SORT_RESULT 3 1
|
|
TQL EVAL (0, 0, '5s') counter_metric{host="host1",device="eth0"} / on() gauge_metric{host="host1",device="eth0"};
|
|
|
|
-- A filtering comparison keeps the left sample values and the derived labels.
|
|
-- SQLNESS SORT_RESULT 3 1
|
|
TQL EVAL (0, 0, '5s') counter_metric{device="eth0"} > on(host) gauge_metric{device="eth0"};
|
|
|
|
-- SQLNESS SORT_RESULT 3 1
|
|
TQL EVAL (0, 0, '5s') counter_metric{device="eth0"} > bool on(host) gauge_metric{device="eth0"};
|
|
|
|
-- `host1` carries two devices on both sides: the match group repeats on the one side.
|
|
TQL EVAL (0, 0, '5s') counter_metric / on(host) gauge_metric;
|
|
|
|
-- The one side is unique here, the many side is not, and no group modifier makes it explicit.
|
|
TQL EVAL (0, 0, '5s') counter_metric / on(host) gauge_metric{device="eth0"};
|
|
|
|
-- `group_left` takes the result labels from the many side.
|
|
-- SQLNESS SORT_RESULT 3 1
|
|
TQL EVAL (0, 0, '5s') counter_metric / on(host) group_left gauge_metric{device="eth0"};
|
|
|
|
-- `group_right` swaps the sides.
|
|
-- SQLNESS SORT_RESULT 3 1
|
|
TQL EVAL (0, 0, '5s') counter_metric{device="eth0"} / on(host) group_right gauge_metric;
|
|
|
|
-- `group_right` makes the left operand the one side, which `host1` breaks.
|
|
TQL EVAL (0, 0, '5s') counter_metric / on(host) group_right gauge_metric{device="eth0"};
|
|
|
|
-- `group_left(device)` copies `device` from the one side, so both `host1` rows end up with the
|
|
-- same labels.
|
|
TQL EVAL (0, 0, '5s') counter_metric / on(host) group_left(device) gauge_metric{device="eth0"};
|
|
|
|
-- An included label the one side doesn't carry is dropped from the result.
|
|
-- SQLNESS SORT_RESULT 3 1
|
|
TQL EVAL (0, 0, '5s') counter_metric{device="eth0"} / on(host) group_left(device) sum by(host)(gauge_metric);
|
|
|
|
DROP TABLE counter_metric;
|
|
|
|
DROP TABLE gauge_metric;
|
|
|
|
-- A label can carry the name the cardinality check generates for its row count.
|
|
CREATE TABLE collide_left (
|
|
host STRING NULL,
|
|
__promql_match_group_count STRING NULL,
|
|
ts TIMESTAMP(3) TIME INDEX,
|
|
greptime_value DOUBLE,
|
|
PRIMARY KEY(host, __promql_match_group_count)
|
|
);
|
|
|
|
CREATE TABLE collide_right (
|
|
host STRING NULL,
|
|
__promql_match_group_count STRING NULL,
|
|
ts TIMESTAMP(3) TIME INDEX,
|
|
greptime_value DOUBLE,
|
|
PRIMARY KEY(host, __promql_match_group_count)
|
|
);
|
|
|
|
INSERT INTO collide_left VALUES ('host1', 'a', 0, 10);
|
|
|
|
INSERT INTO collide_right VALUES ('host1', 'b', 0, 2);
|
|
|
|
TQL EVAL (0, 0, '5s') collide_left / on(host) collide_right;
|
|
|
|
DROP TABLE collide_left;
|
|
|
|
DROP TABLE collide_right;
|
|
|
|
-- Partitioning on a column outside the match keys puts one match group in several regions. The
|
|
-- cardinality check counts the merged data, so it must still see the duplicate.
|
|
CREATE TABLE spread_left (
|
|
host STRING NULL,
|
|
device STRING NULL,
|
|
ts TIMESTAMP(3) TIME INDEX,
|
|
greptime_value DOUBLE,
|
|
PRIMARY KEY(host, device)
|
|
)
|
|
PARTITION ON COLUMNS (device) (
|
|
device < 'eth1',
|
|
device >= 'eth1'
|
|
);
|
|
|
|
CREATE TABLE spread_right (
|
|
host STRING NULL,
|
|
ts TIMESTAMP(3) TIME INDEX,
|
|
greptime_value DOUBLE,
|
|
PRIMARY KEY(host)
|
|
);
|
|
|
|
INSERT INTO spread_left VALUES ('host1', 'eth0', 0, 10), ('host1', 'eth1', 0, 20);
|
|
|
|
INSERT INTO spread_right VALUES ('host1', 0, 2);
|
|
|
|
TQL EVAL (0, 0, '5s') spread_left / on(host) spread_right;
|
|
|
|
TQL EVAL (0, 0, '5s') spread_right / on(host) group_left spread_left;
|
|
|
|
DROP TABLE spread_left;
|
|
|
|
DROP TABLE spread_right;
|
|
|
|
-- An outer aggregate groups on the labels the operand has left, not on the ones the scan
|
|
-- underneath could still produce.
|
|
CREATE TABLE outer_agg_a (
|
|
host STRING NULL,
|
|
device STRING NULL,
|
|
ts TIMESTAMP(3) TIME INDEX,
|
|
greptime_value DOUBLE,
|
|
PRIMARY KEY(host, device)
|
|
);
|
|
|
|
CREATE TABLE outer_agg_b (
|
|
host STRING NULL,
|
|
ts TIMESTAMP(3) TIME INDEX,
|
|
greptime_value DOUBLE,
|
|
PRIMARY KEY(host)
|
|
);
|
|
|
|
INSERT INTO outer_agg_a VALUES ('h1', 'd1', 0, 10), ('h2', 'd2', 0, 20);
|
|
|
|
INSERT INTO outer_agg_b VALUES ('h1', 0, 2), ('h2', 0, 4);
|
|
|
|
-- `on(host)` leaves only `host`, so `without(host)` groups everything together.
|
|
-- SQLNESS SORT_RESULT 3 1
|
|
TQL EVAL (0, 0, '5s') sum without(host) (outer_agg_a / on(host) outer_agg_b);
|
|
|
|
-- `device` is not a label of the operand any more, so it groups everything together too.
|
|
-- SQLNESS SORT_RESULT 3 1
|
|
TQL EVAL (0, 0, '5s') sum by(device) (outer_agg_a / on(host) outer_agg_b);
|
|
|
|
-- SQLNESS SORT_RESULT 3 1
|
|
TQL EVAL (0, 0, '5s') sum by(host) (outer_agg_a / on(host) outer_agg_b);
|
|
|
|
DROP TABLE outer_agg_a;
|
|
|
|
DROP TABLE outer_agg_b;
|
|
|
|
-- Same over the metric engine, whose operands carry `__tsid`.
|
|
CREATE TABLE outer_agg_physical (
|
|
ts TIMESTAMP(3) TIME INDEX,
|
|
greptime_value DOUBLE,
|
|
) ENGINE = metric WITH ("physical_metric_table" = "");
|
|
|
|
CREATE TABLE outer_agg_metric_a (
|
|
host STRING NULL,
|
|
device STRING NULL,
|
|
ts TIMESTAMP(3) NOT NULL,
|
|
greptime_value DOUBLE NULL,
|
|
TIME INDEX (ts),
|
|
PRIMARY KEY(host, device),
|
|
)
|
|
ENGINE = metric
|
|
WITH(
|
|
on_physical_table = 'outer_agg_physical'
|
|
);
|
|
|
|
CREATE TABLE outer_agg_metric_b (
|
|
host STRING NULL,
|
|
ts TIMESTAMP(3) NOT NULL,
|
|
greptime_value DOUBLE NULL,
|
|
TIME INDEX (ts),
|
|
PRIMARY KEY(host),
|
|
)
|
|
ENGINE = metric
|
|
WITH(
|
|
on_physical_table = 'outer_agg_physical'
|
|
);
|
|
|
|
INSERT INTO outer_agg_metric_a (host, device, ts, greptime_value) VALUES
|
|
('h1', 'd1', 0, 10), ('h2', 'd2', 0, 20);
|
|
|
|
INSERT INTO outer_agg_metric_b (host, ts, greptime_value) VALUES
|
|
('h1', 0, 2), ('h2', 0, 4);
|
|
|
|
-- SQLNESS SORT_RESULT 3 1
|
|
TQL EVAL (0, 0, '5s') sum without(host) (outer_agg_metric_a / on(host) outer_agg_metric_b);
|
|
|
|
-- SQLNESS SORT_RESULT 3 1
|
|
TQL EVAL (0, 0, '5s') sum by(device) (outer_agg_metric_a / on(host) outer_agg_metric_b);
|
|
|
|
-- SQLNESS SORT_RESULT 3 1
|
|
TQL EVAL (0, 0, '5s') sum by(host) (outer_agg_metric_a / on(host) outer_agg_metric_b);
|
|
|
|
DROP TABLE outer_agg_metric_a;
|
|
|
|
DROP TABLE outer_agg_metric_b;
|
|
|
|
DROP TABLE outer_agg_physical;
|