Files
greptimedb/tests/perf/fixture-format.md
T
discord9 3493d2d0fb perf(servers)!: speed up Prometheus JSON response building with ryu and per-series entry reuse (#8815)
* fix(perf): align direct-SST CREATE TABLE with baked index metadata

The offline fixture generator (query_perf_fixture::direct_sst::
build_region_metadata) bakes greptime:inverted_index /
greptime:skipping_index field metadata into the region manifest for
tag/field columns, but create_table_sql emitted a bare CREATE TABLE
without those declarations. MergeScan's remote-schema validation then
failed on any tag/field projection (HTTP 500 'advertised remote stream
schema field mismatch'), breaking direct_readable_sst perf cases.

CREATE TABLE now declares the matching SKIPPING INDEX WITH
(granularity='1') / INVERTED INDEX column options. A round-trip test
proves the emitted SQL is parser-valid and yields the exact catalog
metadata.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* perf(servers): speed up Prometheus JSON response building with ryu and per-series entry reuse

PrometheusJsonResponse::record_batches_to_data spends ~47% of its CPU
in f64::to_string() per sample and ~32% in IndexMap::entry() per row
(60s profile of concurrent query_range workloads, ~800k series).

- Replace f64::to_string() with ryu::Buffer::format_finite for finite
  values (shortest round-trip, 3-5x faster); NaN/+Inf/-Inf keep the
  previous std formatting so wire output is unchanged.
- Remember the previous row's label vector and entry index; query output
  is clustered by series, so consecutive rows reuse the same IndexMap
  entry via get_index_mut instead of rebuilding and hashing the label
  vector (worst case adds one Vec comparison per series transition).

Also adds a query-regression case (prom_json_response) that measures the
real Prometheus HTTP range API path (/v1/prometheus/api/v1/query_range),
which is the only frontend path that builds the Prometheus JSON response
(TQL ANALYZE formats the SQL JSON shape instead), plus a prom_http query
kind in the regression runner.

Perf (aligned base d90cca4b75, 256 series x 481 points):
- prom_range_2h (JSON response path): 29.31ms -> 21.46ms (-26.8%)
- tql_range_2h_control (non-JSON path): +1.87% (noise)

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(servers): keep Prometheus wire format for integral floats with ryu

ryu::Buffer::format_finite prints integral values as "1.0", but the
Prometheus JSON wire format (matching std f64::to_string) expects "1".
Strip the trailing ".0" that ryu only emits for integral values; extreme
values keep ryu scientific notation, and NaN/Inf keep std output. Adds
wire-format tests covering 1.0, 0.0, -0.0, 1.5, 0.1, 1e21, 1e30, 1e-7,
f64::MAX, NaN, ±Inf.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(servers): address Prometheus response review feedback

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* perf(servers)!: use ryu for Prometheus sample values

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(cmd): skip Prometheus execution time extraction

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-08-25 05:16:50 +00:00

142 lines
4.8 KiB
Markdown

# Direct-SST fixture format
The phase-1 generator should be a reusable fixture generator, not a collection
of issue-specific data loaders.
## Inputs
Each case describes:
```toml
[case]
name = "example_metric"
description = "example direct-readable-SST regression"
[scenario]
kind = "direct_readable_sst"
seed = 12345
[[scenario.tables]]
database = "public"
name = "example_metric"
engine = "mito"
append_mode = true
sst_format = "flat"
primary_key = ["host", "instance"]
time_index = "ts"
[[scenario.tables.columns]]
name = "host"
type = "STRING"
semantic = "tag"
distribution = { kind = "cardinality", values = 100, prefix = "host" }
[[scenario.tables.columns]]
name = "value"
type = "DOUBLE"
semantic = "field"
distribution = { kind = "deterministic_wave", min = 0.0, max = 100.0 }
[[scenario.tables.columns]]
name = "ts"
type = "TIMESTAMP(9)"
semantic = "timestamp"
[scenario.layout]
regions = 1
sst_count = 1024
rows_per_sst = 4096
row_group_size = 512
time_range_layout = "non_overlapping_per_sst"
series_layout = "round_robin"
[[scenario.queries]]
name = "count_all"
kind = "sql"
query = "SELECT count(*) FROM example_metric"
warmup = 0
iterations = 1
```
### Query kinds
`kind = "prom_http"` runs a Prometheus range query by POSTing form fields to
`/v1/prometheus/api/v1/query_range`. `query`, `start`, `end`, and `step` are
sent as form fields; `start` and `end` define the range, and `step` defines its
evaluation interval. `start`, `end`, and `step` are optional in the case model
and default to empty form values when omitted. The table database is sent as the
`db` URL query parameter, rather than as a form field.
The measurement starts before the HTTP request is sent and ends after the full
response body is read and parsed as JSON. It therefore includes request send,
server-side Prometheus JSON response building, and response-body read and JSON
parsing; it is not server-side-only latency.
`series_layout = "round_robin"` advances the timestamp once per generated row and
cycles series labels across rows. `series_layout = "timestamp_major"` writes all
series for one timestamp before advancing to the next timestamp; use it for
Prometheus-like high-cardinality scrape fixtures where short query windows should
still contain many raw samples. `timestamp_major` requires `rows_per_sst` to be
divisible by `series_count`.
`[scenario]` is required. Other scenario variants are intentionally unsupported
for now, but `scenario.kind` leaves room for future `write_then_query` and
`cache_warm_query` configuration.
The generator should use these declarations to produce:
- object-store SST files written through the real Mito SST writer
- manifest checkpoint and `_last_checkpoint`
- fixture summary with file IDs, row counts, time ranges, and generated schema
## Why this is generic
The same fixture format should support different query-regression families:
- PromQL/TQL time-index pushdown
- SQL predicate pruning
- projection and row-group pruning
- series scan behavior
- joins or aggregation over controlled layouts
Issue-specific behavior belongs in case configs and thresholds, not in the
generator implementation.
## Direct generation vs realism
Direct SST generation is phase-1 because it is fast and reproducible:
- SST count, file time ranges, row groups, and label distributions are fixed by
the case config.
- The same fixture can be used for base and candidate builds.
- Large pruning-sensitive datasets can be created without spending CI time on
ingestion, memtable flush, or compaction.
It is less realistic than ingestion-path data because it bypasses writes,
memtables, flush scheduling, and compaction. That tradeoff is intentional for
PR-level query regression. A later nightly/release suite can add ingestion-based
cases for end-to-end realism.
Multi-table cases are supported by generating one fixture directory per table.
Each table is still limited to one region, and the runner passes `--table`, the
discovered `--region-id`, and the discovered `--table-dir` for each table before
materializing all generated region subtrees into the same datanode data home.
This supports JOIN regression cases without changing existing single-table case
files.
Multi-table cases must use unique table names and unique `(database, name)`
pairs. The runner derives each fixture subdirectory from table index, database,
and table name, sanitizing path-unsafe characters to avoid collisions and unsafe
paths.
The preferred compatibility path is:
1. create an empty table with the target build to seed catalog/table metadata;
2. stop the process;
3. generate readable SSTs and replacement manifest checkpoints offline using the
seeded region metadata;
4. restart and query the fixture.
Fully synthetic metadata is useful for generator smoke tests, but seeded metadata
is safer for end-to-end query performance cases.