Commit Graph
14 Commits
Author SHA1 Message Date
discord9andDennis Zhuang 528ceb7733 perf(promql): reuse sliding min and max candidates (#9099)
* perf(promql): reuse sliding min and max candidates

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(promql): simplify extrema benchmark parameters

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(promql): record baseline sliding extrema SQL results

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* perf(promql): rescan windows that barely overlap

Reusing candidates loses to a plain scan when consecutive windows overlap
little: the deque bookkeeping then costs more than the rescan it replaces.
A local Criterion run on 4096 samples at width 240 / step 240 measured
10.79 -> 22.14 us for min and 12.83 -> 19.85 us for max.

Pick the evaluator once per batch from the first two windows. RangeManipulate
emits one window length and one step per batch, so that sample decides for all
of them, and both evaluators return identical bits, so a wrong pick costs time
only. Batches that do not qualify fold each window on its own.

Move the incremental state into SlidingExtrema so tests can drive it directly:
the exhaustive four-sample differential test cannot reach it through a UDF
call, because such a batch never qualifies for reuse.

Signed-off-by: Dennis Zhuang <xzhuang@greptime.com>

* fix(promql): select the extrema evaluator from batch averages

Reading the window shape off the first two windows misreads the batch.
RangeManipulate starts a series at max(query start, first aligned sample),
so a series that begins inside the query range gets a first window covering
roughly one step, and a window covering no sample at all is emitted as
(0, 0). Either one closed the gate for the whole batch, including the
one-hour window at a 15s step that candidate reuse was written for.

Compare the batch averages instead: at least 32 samples per window, and a
step advancing at most a quarter of that. Uniform batches select exactly as
before, so the thresholds keep the meaning they were measured with.

The 32-sample rule had also moved most of the benchmark and query-regression
shapes onto the rescan, including the case built to measure reset and
rebuild. Widen those windows to 40 samples, add a step at the selection
boundary, and add an end-to-end case with 40-sample windows advancing 5.

Signed-off-by: Dennis Zhuang <xzhuang@greptime.com>

* fix(promql): ignore empty windows when measuring batch advance

The advance was read from the first and last window offsets, but a window
covering no sample is emitted as (0, 0). A query whose last evaluation lands
exactly one window past the last sample ends on such a window, and its zero
offset made a batch of disjoint windows look like one that never moved, which
selected the evaluator built for overlap. Results stayed correct; the cost was
deque bookkeeping on the shape the scan fallback exists for.

Take the offset span over the windows that cover a sample. Empty windows stay
in the window count, where they only make both conditions stricter.

Signed-off-by: Dennis Zhuang <xzhuang@greptime.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
Signed-off-by: Dennis Zhuang <xzhuang@greptime.com>
Co-authored-by: Dennis Zhuang <xzhuang@greptime.com>
2026-09-17 03:18:21 +00:00
discord9 ba0f7acd93 feat(mito2): add opt-in byte-stream-split encoding for float SST fields (#9069)
* feat(mito2): add opt-in byte stream split encoding

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(mito2): correct float encoding checks

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(compat): cover float SST encoding

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(mito2): compile float encoding tests

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(mito2): release parquet test writer

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(mito2): register float test primary key

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(mito2): verify BSS write lifecycles

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(metric-engine): verify BSS physical SST

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(mito2): verify bulk BSS lifecycle

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(mito2): compile bulk BSS test

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(mito2): narrow bulk encoding constructors

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(compat): accept generated float upgrade output

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(compat): accept generated float downgrade output

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(mito2): narrow bulk encoding builder

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(perf): add default versus BSS storage comparison

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(perf): align BSS reader benchmarks with prior study

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(perf): parse current read benchmark averages

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(perf): retain default float encoding in direct SST fixtures

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(perf): isolate BSS user SSTs and benchmark every file

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* docs(perf): record measured BSS storage and reader tradeoffs

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* docs(perf): expose warm scan variability and evidence limits

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* docs(perf): clarify BSS baseline and storage measurement scope

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(perf): model bounded mixed integer and fractional metric series

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* docs(perf): report bounded mixed BSS measurements and query regressions

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* docs(perf): qualify timings affected by concurrent host builds

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(perf): include float BSS comparison in default regression cases

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(perf): omit unsupported float encoding option from baseline setup

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-14 12:05:42 +00:00
discord9 555485c40e perf(promql): avoid per-window allocations in simple range functions (#9104)
* perf(promql): avoid per-window allocations in simple range functions

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* perf(promql): avoid copying smoothing window values

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(perf): cover smoothing copy removal across window layouts

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(promql): skip null samples in simple range functions

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-11 11:59:45 +00:00
discord9 97648525cf perf(mito2): skip proven all-match prefilters (#9066)
* perf(mito2): skip proven all-match prefilters

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(perf): add manual all-match prefilter reproduction

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(mito2): cover all-match prefilter execution paths

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(mito2): match prefilter fixture to sparse SST schema

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore(mito2): address all-match prefilter lint findings

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(perf): cover all-match prefilters in default regressions

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-11 09:35:57 +00:00
discord9 7cf84892d2 perf(table): filter decoded rows with dynamic predicates (#9004)
* perf(table): filter decoded rows with dynamic predicates

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(table): preserve unknown dynamic filter rows

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(table): use null guards for dynamic filters

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(table): preserve null inputs during dynamic pruning

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(perf): cover frontend join dynamic filter transfer

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(table): reset dynamic filters and scanner pruning state

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-09 12:43:10 +00:00
discord9 da5cb1a190 perf(promql): push down last row for instant queries (#9034)
* perf(promql): push down last row for instant queries

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: guard instant last row correctness

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: update instant query explain results

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: apply last row after source deduplication

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: scope post-merge last row selection

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(perf): cover instant PromQL last row

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(perf): sort generated SST rows before writing

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(promql): cover instant last row selection in sqlness

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(promql): avoid last row hints for lossy timestamp casts

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(promql): preserve stale marker semantics across flushes

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(promql): decode dictionary labels in stale regression

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(promql): avoid reserved column name in stale fixture

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(promql): exercise LastRow hints and filtered results

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor: keep after-merge mode in LastRow selector

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: reject instant LastRow across residual filters

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: expect after-merge selector in instant vector guards

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* docs: explain instant LastRow filter eligibility

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: restrict instant LastRow to safe selector nodes

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor: show LastRow merge mode directly in diagnostics

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: refresh LastRow display in explain expectations

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-09 12:38:52 +00:00
shuiyisong b86da3d35f feat: support raw OTLP delta metrics (#8970)
* feat: support raw OTLP delta metrics

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: fmt

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* test(promql): update sqlness results for normalized label matching

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: derive temporality label from default column prefix

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* test(promql): add analyze coverage for delta temporality

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix(promql): scope label alignment to temporality marker

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: handle count-only histograms and vector broadcasts

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: exclude temporality marker from entity descriptions

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: use a fixed label for OTLP aggregation temporality

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix(promql): preserve mixed-range semantics for raw delta

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

---------

Signed-off-by: shuiyisong <xixing.sys@gmail.com>
2026-09-03 09:28:11 +00:00
discord9 3493d2d0fb perf(servers)!: speed up Prometheus JSON response building with ryu and per-series entry reuse (#8815)
* fix(perf): align direct-SST CREATE TABLE with baked index metadata

The offline fixture generator (query_perf_fixture::direct_sst::
build_region_metadata) bakes greptime:inverted_index /
greptime:skipping_index field metadata into the region manifest for
tag/field columns, but create_table_sql emitted a bare CREATE TABLE
without those declarations. MergeScan's remote-schema validation then
failed on any tag/field projection (HTTP 500 'advertised remote stream
schema field mismatch'), breaking direct_readable_sst perf cases.

CREATE TABLE now declares the matching SKIPPING INDEX WITH
(granularity='1') / INVERTED INDEX column options. A round-trip test
proves the emitted SQL is parser-valid and yields the exact catalog
metadata.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* perf(servers): speed up Prometheus JSON response building with ryu and per-series entry reuse

PrometheusJsonResponse::record_batches_to_data spends ~47% of its CPU
in f64::to_string() per sample and ~32% in IndexMap::entry() per row
(60s profile of concurrent query_range workloads, ~800k series).

- Replace f64::to_string() with ryu::Buffer::format_finite for finite
  values (shortest round-trip, 3-5x faster); NaN/+Inf/-Inf keep the
  previous std formatting so wire output is unchanged.
- Remember the previous row's label vector and entry index; query output
  is clustered by series, so consecutive rows reuse the same IndexMap
  entry via get_index_mut instead of rebuilding and hashing the label
  vector (worst case adds one Vec comparison per series transition).

Also adds a query-regression case (prom_json_response) that measures the
real Prometheus HTTP range API path (/v1/prometheus/api/v1/query_range),
which is the only frontend path that builds the Prometheus JSON response
(TQL ANALYZE formats the SQL JSON shape instead), plus a prom_http query
kind in the regression runner.

Perf (aligned base d90cca4b75, 256 series x 481 points):
- prom_range_2h (JSON response path): 29.31ms -> 21.46ms (-26.8%)
- tql_range_2h_control (non-JSON path): +1.87% (noise)

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(servers): keep Prometheus wire format for integral floats with ryu

ryu::Buffer::format_finite prints integral values as "1.0", but the
Prometheus JSON wire format (matching std f64::to_string) expects "1".
Strip the trailing ".0" that ryu only emits for integral values; extreme
values keep ryu scientific notation, and NaN/Inf keep std output. Adds
wire-format tests covering 1.0, 0.0, -0.0, 1.5, 0.1, 1e21, 1e30, 1e-7,
f64::MAX, NaN, ±Inf.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(servers): address Prometheus response review feedback

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* perf(servers)!: use ryu for Prometheus sample values

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(cmd): skip Prometheus execution time extraction

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-08-25 05:16:50 +00:00
discord9 a7590f8174 perf(promql): avoid repeated scans in sliding range evaluation (#8646)
* perf(promql): use two pointers for sliding range boundaries

Replace the stale cursor heuristic in RangeManipulateStream::calculate_range
with monotonic left/right cursors. The old path rescanned each evaluation
window (O(E x samples-per-window)) and could lose valid samples after sparse
gaps or trailing empty windows. The two pointers keep strict monotonic
progress, reducing boundary generation to O(N + E) while preserving
(curr-range, curr] semantics, start/end shortening, and empty-window output.

Controlled release benchmarks (fixed CPU, ABBA):
- Public RangeManipulate wall time: ~28% faster at 1m/15s, ~66% at 5m/15s,
  ~96% at 1h/15s.
- Warmed distributed TQL ANALYZE 1h queries: ~17-21% faster end to end;
  shorter windows stayed within run-order noise.

Signed-off-by: discord9 <discord9@163.com>

* perf(promql): specialize changes/resets with adaptive edge counting

The generic range_fn macro slices, downcasts, and rescans every overlapping
window for changes() and resets(). Replace the macro path for these two
functions with hand-written UDF wrappers backed by a shared private
edge-count kernel: direct raw-offset scans when requested edges are few,
otherwise one global u64 edge prefix so each window is answered by a prefix
difference.

Behavior is preserved bit-for-bit, including raw null-buffer values, NaN
semantics, signed zero, infinities, empty/singleton windows, independent
timestamp/value offsets, arbitrary window layouts, and exact DataFusion
error messages. The shared proc macro, planner, serializer, and other range
functions are untouched.

Controlled release benchmarks (fixed CPU, ABBA):
- Dense sliding windows (k=4/20/240): 91.7-95.6% less public UDF wall time.
- Low-coverage fallback (N=4096, 8 windows): 73.9-74.4% faster.
- Warmed distributed TQL ANALYZE 5m/1h changes/resets: 12.1-19.7% client
  and 12.0-20.9% server latency improvement; controls stayed within drift.

Signed-off-by: discord9 <discord9@163.com>

* ci(query-regression): include PromQL range boundary case in defaults

An audit of historical query-regression runs found zero range-query
coverage: all 208 PromQL ANALYZE samples were bare selectors, so range
evaluation could regress without CI noticing. Wire the
promql_range_boundary case (introduced in #8646) into DEFAULT_CASES so
label-triggered runs measure the range path. The case is cheap: a ~0.3s
synthetic fixture and about a minute of query execution per base/candidate
pass.

Signed-off-by: discord9 <discord9@163.com>

* chore(promql): address sliding range review nits

Move test-only imports into their test modules and remove the unused
pre-specialization changes and resets helpers.

Signed-off-by: discord9 <discord9@163.com>

* style(promql): apply pinned rustfmt

Signed-off-by: discord9 <discord9@163.com>

* test(promql): cover sparse range results

Share the changes and resets test scaffolding while keeping their behavior
oracles independent. Add an end-to-end sqlness regression for sparse samples,
empty intermediate windows, and a valid trailing sample.

Signed-off-by: discord9 <discord9@163.com>

---------

Signed-off-by: discord9 <discord9@163.com>
2026-08-03 06:53:21 +00:00
shuiyisong f5f5d468bb ci: add OTLP trace ingestion regression testing (#8631)
* chore: update CI config

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* chore: add CI

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* chore: update CI config

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* ci: report otelgen runner diagnostics

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* chore: add script to draw result diagram

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

---------

Signed-off-by: shuiyisong <xixing.sys@gmail.com>
2026-07-26 15:08:27 +00:00
discord9 56addd0623 fix: stream remote analyze metrics while pending (#8405)
* fix: stream remote analyze metrics while pending

Signed-off-by: discord9 <discord9@163.com>

* test: verify flight metrics preserve pending batch

Signed-off-by: discord9 <discord9@163.com>

* fix: preserve direct SST perf queries in plans

Signed-off-by: discord9 <discord9@163.com>

* fix: bind flight metrics capability to query

Signed-off-by: discord9 <discord9@163.com>

---------

Signed-off-by: discord9 <discord9@163.com>
2026-07-21 04:17:50 +00:00
discord9 02283b6ba0 test(perf): add remote write storage inspection (#8444)
* feat: add remote write value distributions

Signed-off-by: discord9 <discord9@163.com>

* test: extend remote write value distributions

Signed-off-by: discord9 <discord9@163.com>

* test: inspect remote write parquet storage

Signed-off-by: discord9 <discord9@163.com>

* test: normalize remote write perf fixtures

Signed-off-by: discord9 <discord9@163.com>

* test: add heavy remote write perf case

Signed-off-by: discord9 <discord9@163.com>

* test: use head greptime for read bench

Signed-off-by: discord9 <discord9@163.com>

* test: keep heavy remote write case local

Signed-off-by: discord9 <discord9@163.com>

* test: tune remote write perf smoke case

Signed-off-by: discord9 <discord9@163.com>

* test: fix query fixture import style

Signed-off-by: discord9 <discord9@163.com>

* test: cover remote write value distributions

Signed-off-by: discord9 <discord9@163.com>

* test: add integer counter perf case

Signed-off-by: discord9 <discord9@163.com>

---------

Signed-off-by: discord9 <discord9@163.com>
2026-07-10 12:27:27 +00:00
discord9 12f83828b1 feat: add Prom remote-write query regression scenario (#8413)
* feat: add Prom remote-write query regression scenario

Signed-off-by: discord9 <discord9@163.com>

* test: add high-cardinality remote-write query case

Signed-off-by: discord9 <discord9@163.com>

* feat: chunk remote-write query regression loads

Signed-off-by: discord9 <discord9@163.com>

* test: use multi-day remote-write regression case

Signed-off-by: discord9 <discord9@163.com>

* fix: address query regression review comments

Signed-off-by: discord9 <discord9@163.com>

* ci: allow large query regression comments

Signed-off-by: discord9 <discord9@163.com>

---------

Signed-off-by: discord9 <discord9@163.com>
2026-07-07 07:58:56 +00:00
discord9 c44f8da646 feat: add query regression perf harness (#8406)
* feat: add query regression perf harness

Signed-off-by: discord9 <discord9@163.com>

* feat: extend query regression cases

Signed-off-by: discord9 <discord9@163.com>

* ci: harden query regression workflows

Signed-off-by: discord9 <discord9@163.com>

* fix: address query regression review comments

Signed-off-by: discord9 <discord9@163.com>

* ci: limit query regression PR triggers

Signed-off-by: discord9 <discord9@163.com>

* ci: run full query regression case set

Signed-off-by: discord9 <discord9@163.com>

* refactor: model query regression scenarios

Signed-off-by: discord9 <discord9@163.com>

* fix: avoid unenforced query regression thresholds

Signed-off-by: discord9 <discord9@163.com>

---------

Signed-off-by: discord9 <discord9@163.com>
2026-07-03 09:09:01 +00:00