* perf: add flight coalesce regression case with high-cardinality aggregations
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* perf: add flight coalesce aggregations bench case
Add a direct_readable_sst case exercising grouped aggregation over
coalesced batches: 16 hosts x 4096 instances, 32 SSTs of 32768 rows,
timestamp-major series layout, three SQL queries (aggregation, topk,
count_by_host) each with a 10% max candidate latency regression
threshold.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
---------
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* perf(promql): reuse sliding min and max candidates
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* test(promql): simplify extrema benchmark parameters
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* test(promql): record baseline sliding extrema SQL results
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* perf(promql): rescan windows that barely overlap
Reusing candidates loses to a plain scan when consecutive windows overlap
little: the deque bookkeeping then costs more than the rescan it replaces.
A local Criterion run on 4096 samples at width 240 / step 240 measured
10.79 -> 22.14 us for min and 12.83 -> 19.85 us for max.
Pick the evaluator once per batch from the first two windows. RangeManipulate
emits one window length and one step per batch, so that sample decides for all
of them, and both evaluators return identical bits, so a wrong pick costs time
only. Batches that do not qualify fold each window on its own.
Move the incremental state into SlidingExtrema so tests can drive it directly:
the exhaustive four-sample differential test cannot reach it through a UDF
call, because such a batch never qualifies for reuse.
Signed-off-by: Dennis Zhuang <xzhuang@greptime.com>
* fix(promql): select the extrema evaluator from batch averages
Reading the window shape off the first two windows misreads the batch.
RangeManipulate starts a series at max(query start, first aligned sample),
so a series that begins inside the query range gets a first window covering
roughly one step, and a window covering no sample at all is emitted as
(0, 0). Either one closed the gate for the whole batch, including the
one-hour window at a 15s step that candidate reuse was written for.
Compare the batch averages instead: at least 32 samples per window, and a
step advancing at most a quarter of that. Uniform batches select exactly as
before, so the thresholds keep the meaning they were measured with.
The 32-sample rule had also moved most of the benchmark and query-regression
shapes onto the rescan, including the case built to measure reset and
rebuild. Widen those windows to 40 samples, add a step at the selection
boundary, and add an end-to-end case with 40-sample windows advancing 5.
Signed-off-by: Dennis Zhuang <xzhuang@greptime.com>
* fix(promql): ignore empty windows when measuring batch advance
The advance was read from the first and last window offsets, but a window
covering no sample is emitted as (0, 0). A query whose last evaluation lands
exactly one window past the last sample ends on such a window, and its zero
offset made a batch of disjoint windows look like one that never moved, which
selected the evaluator built for overlap. Results stayed correct; the cost was
deque bookkeeping on the shape the scan fallback exists for.
Take the offset span over the windows that cover a sample. Empty windows stay
in the window count, where they only make both conditions stricter.
Signed-off-by: Dennis Zhuang <xzhuang@greptime.com>
---------
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
Signed-off-by: Dennis Zhuang <xzhuang@greptime.com>
Co-authored-by: Dennis Zhuang <xzhuang@greptime.com>
* Implement `/query-regression` command handling and admission workflow
- Add `query-regression-slash.py` script for processing `/query-regression` commands in PR comments, validating case arguments, and checking permissions.
- Update `checks.yml` to include tests for the new slash command functionality.
- Modify `query-regression-comment.yml` to trigger on the new `Query Regression Command` workflow.
- Create `query-regression-slash.yml` to handle the dispatched command, validate allowlist and permissions, and initiate the regression workflow.
- Enhance `query-regression.yml` to support additional inputs for PR admission and SHA verification.
- Introduce `slash-command-dispatch.yml` to parse and dispatch commands from PR comments.
- Document the new command admission process in `AGENTS.md` and `README.md`.
- Add unit tests in `test_query_regression_slash.py` to cover command parsing and admission logic.
* refactor: enhance query-regression command handling with comment validation and identity checks
* feat: implement admission identity handling for query regression workflows
* refactor: update PR admission logic in query regression workflow
* refactor: update token usage in slash command dispatch and README for clarity
* test: add cases for handling re-run failed jobs and stale runner artifacts
* refactor: improve repository metadata handling in query regression scripts
* chore: enable overwrite for artifact uploads to handle re-run failed jobs
* chore: enable overwrite for query regression admission uploads
* feat: enhance query-regression admission with HMAC signing and verification
- Introduced HMAC signing for admission markers in query-regression workflows to ensure integrity and authenticity.
- Updated `query-regression-comment.test.cjs` to include tests for signing and verifying admission markers.
- Modified `query-regression-slash.py` to handle admission marker signing and verification, including checks for dispatch sender and head SHA consistency.
- Enhanced workflows to securely manage admission markers and HMAC secrets, ensuring they are not exposed to untrusted contexts.
- Improved documentation to clarify the admission process and the role of HMAC in securing the workflow.
* test: add case to find newly posted marker among newer comments
* test: add case to verify multiline output handling in write_outputs function
* feat(ci): add aliyun ecs ephemeral runner path for query regression
Signed-off-by: paomian <xpaomian@gmail.com>
* fix: improve condition for query-regression job execution in workflow
* feat: update Docker installation to use official repository and add GPG key handling
* Refactor query regression runner setup and configuration
- Removed deprecated PersistentVolumeClaim for build cache.
- Introduced a new bootstrap script for setting up the ECS runner host.
- Deleted obsolete Helm values files for runner configuration.
- Updated the Aliyun ECS runner provisioning script to reflect new cache paths.
- Modified GitHub workflows to use the new Aliyun ECS runner setup.
- Adjusted documentation to clarify the new runner lifecycle and provisioning process.
* fix: enhance runner service management during bootstrap process
* fix: update alibabacloud_tea_openapi dependency version in metadata
* feat: enhance ECS runner scripts with region_id and resource_group_id support
* fix: move containerd content store to data root for improved storage management
* feat: rename query-regression runner to ephemeral-github runner and update related scripts
* fix: update sentinel polling method to use serial console output for improved reliability
* fix: add environment variable checks for Alibaba Cloud access keys in ECS client
* fix: improve error handling in GitHub API requests for better diagnostics
* fix: improve cache disk detection logic for Aliyun ECS instances
* fix: enhance cache disk waiting logic with detailed output and error handling
* fix: update dependency version for alibabacloud_tea_openapi in teardown script
* fix: enhance cache disk waiting logic for better compatibility and clarity
* fix: enhance console output handling and add incremental logging during instance provisioning
* fix: add PATH environment variable for runner jobs in service and provision script
* fix: add machine telemetry sampling and logging during query regression jobs
* fix: update query regression documentation and provision script for cache disk handling
* fix: update SCCACHE_CACHE_SIZE validation to 10G for improved caching efficiency
* fix: remove outdated cache size checks and cleanup logic for fresh system disk runs
* fix: enhance instance deletion logic with region handling and console output export
* fix: add swap file setup and OOM handling for ECS runner to improve stability
* fix: update OOM handling and service restart logic for ECS runner to enhance stability
* fix: increase system disk size to 100 GiB for cold double nightly builds to prevent ENOSPC errors
* fix: increase system disk size to 150 GiB for ECS runner to prevent ENOSPC errors
* fix: add keep_instance option to preserve ECS instance for post-mortem debugging
* fix: disable unattended upgrades to prevent job cancellations during library updates
* fix: reduce system disk size to 40 GiB for ECS runner to prevent ENOSPC errors
* feat: Refactor Aliyun ECS runner provisioning and introduce nightly regression comparison
- Update `aliyun-ecs-runner-provision.py` to remove cache disk handling, simplifying the provisioning process.
- Introduce `query-regression-nightly-refs.py` to resolve and compare SHAs from successful nightly builds.
- Create `query-regression-nightly.yml` workflow to trigger nightly comparisons based on successful builds.
- Enhance `query-regression.yml` to include a `test-tooling` job for validating Python scripts before provisioning.
- Update tests for the new nightly reference selection logic and refactor existing tests to align with the new caching strategy.
- Modify documentation to reflect changes in caching and nightly comparison workflows.
* fix: enhance runner image tool verification with detailed checks
* fix: improve error handling in runner image tool verification
* fix: update tool versions in ECS image and workflow for consistency
* fix: correct typo in error message for unparseable ECS creation time
* fix: update README and workflow files for query regression tests and image hygiene
---------
Signed-off-by: paomian <xpaomian@gmail.com>
* fix(perf): align direct-SST CREATE TABLE with baked index metadata
The offline fixture generator (query_perf_fixture::direct_sst::
build_region_metadata) bakes greptime:inverted_index /
greptime:skipping_index field metadata into the region manifest for
tag/field columns, but create_table_sql emitted a bare CREATE TABLE
without those declarations. MergeScan's remote-schema validation then
failed on any tag/field projection (HTTP 500 'advertised remote stream
schema field mismatch'), breaking direct_readable_sst perf cases.
CREATE TABLE now declares the matching SKIPPING INDEX WITH
(granularity='1') / INVERTED INDEX column options. A round-trip test
proves the emitted SQL is parser-valid and yields the exact catalog
metadata.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* perf(servers): speed up Prometheus JSON response building with ryu and per-series entry reuse
PrometheusJsonResponse::record_batches_to_data spends ~47% of its CPU
in f64::to_string() per sample and ~32% in IndexMap::entry() per row
(60s profile of concurrent query_range workloads, ~800k series).
- Replace f64::to_string() with ryu::Buffer::format_finite for finite
values (shortest round-trip, 3-5x faster); NaN/+Inf/-Inf keep the
previous std formatting so wire output is unchanged.
- Remember the previous row's label vector and entry index; query output
is clustered by series, so consecutive rows reuse the same IndexMap
entry via get_index_mut instead of rebuilding and hashing the label
vector (worst case adds one Vec comparison per series transition).
Also adds a query-regression case (prom_json_response) that measures the
real Prometheus HTTP range API path (/v1/prometheus/api/v1/query_range),
which is the only frontend path that builds the Prometheus JSON response
(TQL ANALYZE formats the SQL JSON shape instead), plus a prom_http query
kind in the regression runner.
Perf (aligned base d90cca4b75, 256 series x 481 points):
- prom_range_2h (JSON response path): 29.31ms -> 21.46ms (-26.8%)
- tql_range_2h_control (non-JSON path): +1.87% (noise)
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* fix(servers): keep Prometheus wire format for integral floats with ryu
ryu::Buffer::format_finite prints integral values as "1.0", but the
Prometheus JSON wire format (matching std f64::to_string) expects "1".
Strip the trailing ".0" that ryu only emits for integral values; extreme
values keep ryu scientific notation, and NaN/Inf keep std output. Adds
wire-format tests covering 1.0, 0.0, -0.0, 1.5, 0.1, 1e21, 1e30, 1e-7,
f64::MAX, NaN, ±Inf.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* fix(servers): address Prometheus response review feedback
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* perf(servers)!: use ryu for Prometheus sample values
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* fix(cmd): skip Prometheus execution time extraction
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
---------
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* refactor: port query regression runner to Rust
Signed-off-by: discord9 <discord9@163.com>
* ci: remove optional OTLP report plotter
Signed-off-by: discord9 <discord9@163.com>
* refactor: split query regression runner into modules
Signed-off-by: discord9 <discord9@163.com>
* style: use crate-qualified imports in query regression runner
Signed-off-by: discord9 <discord9@163.com>
* refactor: simplify query regression runner internals
Signed-off-by: discord9 <discord9@163.com>
* feat: abstract inspect-footer storage access behind object store destination
Add an optional --destination <TOML> to inspect-footer (and
--base-destination/--candidate-destination to finalize-remote) so the
storage inspection reads DB data files through the opendal-backed
object_store abstraction instead of bare std::fs. Local paths keep
working unchanged via the --root shortcut (File backend); remote
backends (S3/GCS/...) are described by a DestinationConfig TOML
reusing the object-store crate's ObjectStoreConfig serde shape.
- inspect_footer: list via ObjectStore::list + ObjectMeta filtering
(parquet keys, non-zero size, metadata/ segment), read footers
async via ParquetObjectReader + ParquetMetaDataReader with known
file size (no extra HEAD); output JSON schema unchanged
- finalize-remote: --base-data-home/--candidate-data-home become
optional, mutually exclusive with the new --*-destination args
- cmd deps: add object_store_opendal + datafusion_object_store
- tests: fs-backend list+footer integration tests (metadata filtering,
destination TOML mode, root/destination exclusivity)
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* style: drop needless borrow in inspect footer test
Fix clippy::needless_borrows_for_generic_args in the inspect-footer test
(fs::create_dir_all(table.join("metadata"))). Missed by the earlier
focused clippy run because it only covered --bin targets.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
---------
Signed-off-by: discord9 <discord9@163.com>
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* perf(promql): use two pointers for sliding range boundaries
Replace the stale cursor heuristic in RangeManipulateStream::calculate_range
with monotonic left/right cursors. The old path rescanned each evaluation
window (O(E x samples-per-window)) and could lose valid samples after sparse
gaps or trailing empty windows. The two pointers keep strict monotonic
progress, reducing boundary generation to O(N + E) while preserving
(curr-range, curr] semantics, start/end shortening, and empty-window output.
Controlled release benchmarks (fixed CPU, ABBA):
- Public RangeManipulate wall time: ~28% faster at 1m/15s, ~66% at 5m/15s,
~96% at 1h/15s.
- Warmed distributed TQL ANALYZE 1h queries: ~17-21% faster end to end;
shorter windows stayed within run-order noise.
Signed-off-by: discord9 <discord9@163.com>
* perf(promql): specialize changes/resets with adaptive edge counting
The generic range_fn macro slices, downcasts, and rescans every overlapping
window for changes() and resets(). Replace the macro path for these two
functions with hand-written UDF wrappers backed by a shared private
edge-count kernel: direct raw-offset scans when requested edges are few,
otherwise one global u64 edge prefix so each window is answered by a prefix
difference.
Behavior is preserved bit-for-bit, including raw null-buffer values, NaN
semantics, signed zero, infinities, empty/singleton windows, independent
timestamp/value offsets, arbitrary window layouts, and exact DataFusion
error messages. The shared proc macro, planner, serializer, and other range
functions are untouched.
Controlled release benchmarks (fixed CPU, ABBA):
- Dense sliding windows (k=4/20/240): 91.7-95.6% less public UDF wall time.
- Low-coverage fallback (N=4096, 8 windows): 73.9-74.4% faster.
- Warmed distributed TQL ANALYZE 5m/1h changes/resets: 12.1-19.7% client
and 12.0-20.9% server latency improvement; controls stayed within drift.
Signed-off-by: discord9 <discord9@163.com>
* ci(query-regression): include PromQL range boundary case in defaults
An audit of historical query-regression runs found zero range-query
coverage: all 208 PromQL ANALYZE samples were bare selectors, so range
evaluation could regress without CI noticing. Wire the
promql_range_boundary case (introduced in #8646) into DEFAULT_CASES so
label-triggered runs measure the range path. The case is cheap: a ~0.3s
synthetic fixture and about a minute of query execution per base/candidate
pass.
Signed-off-by: discord9 <discord9@163.com>
* chore(promql): address sliding range review nits
Move test-only imports into their test modules and remove the unused
pre-specialization changes and resets helpers.
Signed-off-by: discord9 <discord9@163.com>
* style(promql): apply pinned rustfmt
Signed-off-by: discord9 <discord9@163.com>
* test(promql): cover sparse range results
Share the changes and resets test scaffolding while keeping their behavior
oracles independent. Add an end-to-end sqlness regression for sparse samples,
empty intermediate windows, and a valid trailing sample.
Signed-off-by: discord9 <discord9@163.com>
---------
Signed-off-by: discord9 <discord9@163.com>