Files
greptimedb/tests/perf/AGENTS.md
T
localhost 35ea88a4ef feat(ci): run query regression on ephemeral Aliyun ECS runners (#8937)
* feat(ci): add aliyun ecs ephemeral runner path for query regression

Signed-off-by: paomian <xpaomian@gmail.com>

* fix: improve condition for query-regression job execution in workflow

* feat: update Docker installation to use official repository and add GPG key handling

* Refactor query regression runner setup and configuration

- Removed deprecated PersistentVolumeClaim for build cache.
- Introduced a new bootstrap script for setting up the ECS runner host.
- Deleted obsolete Helm values files for runner configuration.
- Updated the Aliyun ECS runner provisioning script to reflect new cache paths.
- Modified GitHub workflows to use the new Aliyun ECS runner setup.
- Adjusted documentation to clarify the new runner lifecycle and provisioning process.

* fix: enhance runner service management during bootstrap process

* fix: update alibabacloud_tea_openapi dependency version in metadata

* feat: enhance ECS runner scripts with region_id and resource_group_id support

* fix: move containerd content store to data root for improved storage management

* feat: rename query-regression runner to ephemeral-github runner and update related scripts

* fix: update sentinel polling method to use serial console output for improved reliability

* fix: add environment variable checks for Alibaba Cloud access keys in ECS client

* fix: improve error handling in GitHub API requests for better diagnostics

* fix: improve cache disk detection logic for Aliyun ECS instances

* fix: enhance cache disk waiting logic with detailed output and error handling

* fix: update dependency version for alibabacloud_tea_openapi in teardown script

* fix: enhance cache disk waiting logic for better compatibility and clarity

* fix: enhance console output handling and add incremental logging during instance provisioning

* fix: add PATH environment variable for runner jobs in service and provision script

* fix: add machine telemetry sampling and logging during query regression jobs

* fix: update query regression documentation and provision script for cache disk handling

* fix: update SCCACHE_CACHE_SIZE validation to 10G for improved caching efficiency

* fix: remove outdated cache size checks and cleanup logic for fresh system disk runs

* fix: enhance instance deletion logic with region handling and console output export

* fix: add swap file setup and OOM handling for ECS runner to improve stability

* fix: update OOM handling and service restart logic for ECS runner to enhance stability

* fix: increase system disk size to 100 GiB for cold double nightly builds to prevent ENOSPC errors

* fix: increase system disk size to 150 GiB for ECS runner to prevent ENOSPC errors

* fix: add keep_instance option to preserve ECS instance for post-mortem debugging

* fix: disable unattended upgrades to prevent job cancellations during library updates

* fix: reduce system disk size to 40 GiB for ECS runner to prevent ENOSPC errors

* feat: Refactor Aliyun ECS runner provisioning and introduce nightly regression comparison

- Update `aliyun-ecs-runner-provision.py` to remove cache disk handling, simplifying the provisioning process.
- Introduce `query-regression-nightly-refs.py` to resolve and compare SHAs from successful nightly builds.
- Create `query-regression-nightly.yml` workflow to trigger nightly comparisons based on successful builds.
- Enhance `query-regression.yml` to include a `test-tooling` job for validating Python scripts before provisioning.
- Update tests for the new nightly reference selection logic and refactor existing tests to align with the new caching strategy.
- Modify documentation to reflect changes in caching and nightly comparison workflows.

* fix: enhance runner image tool verification with detailed checks

* fix: improve error handling in runner image tool verification

* fix: update tool versions in ECS image and workflow for consistency

* fix: correct typo in error message for unparseable ECS creation time

* fix: update README and workflow files for query regression tests and image hygiene

---------

Signed-off-by: paomian <xpaomian@gmail.com>
2026-08-26 12:11:14 +00:00

2.9 KiB

Agent Guidelines for Query Performance Tests

  • Keep GitHub Actions YAML thin. Put non-trivial control flow, case expansion, report generation, and metadata writing in scripts under .github/scripts/; workflow steps should mostly invoke those scripts.
  • Runner lifecycle: the default path provisions one ephemeral Aliyun ECS instance per run via .github/scripts/aliyun-ecs-runner-provision.py and always releases it via aliyun-ecs-runner-teardown.py; a scheduled janitor workflow sweeps leftovers. Build caches live on that instance's system disk and are discarded with the VM. Runs do not share a workflow concurrency group. The ECS custom image is built from the runner Dockerfile by .github/runner-scale-sets/query-regression/ecs-image/build-ecs-image.py; keep the Dockerfile the single source of the tool contract. Dispatching with any other runner value treats it as a literal self-hosted runner label (see ecs-image/bootstrap-runner-host.sh for preparing such a host).
  • Query regression PR runs should build base/candidate binaries once, then run the default case set. Do not hard-code a single case such as promql_pushdown_7913 into the workflow path.
  • Scheduled nightly comparison lives in query-regression-nightly.yml: it waits for a successful Nightly Build, then calls query-regression.yml with the previous vs current nightly SHAs. Keep SHA selection in .github/scripts/query-regression-nightly-refs.py.
  • The case DSL is not required to keep compatibility inside this PR. When the DSL changes, update TOML cases, the outer lifecycle script, Rust helpers, and docs together.
  • [case] is report metadata only. [scenario] is the executable regression configuration and must include kind, data layout, tables, queries, and thresholds. Rust owns case schema, defaults, validation, and normalized plan output through query_perf_fixture plan. .github/scripts/query-regression-run.py owns process lifecycle; query_regression_runner consumes normalized plans, frontend endpoints, direct-SST materialization requests, and OTLP target/finalize requests.
  • Keep the direct-SST generator generic. Issue-specific behavior belongs in case files and thresholds, not in Rust generator logic.
  • Before pushing perf harness changes, run at least:
    • the Python tests in the test-tooling job of .github/workflows/query-regression.yml (ubuntu-latest, not the ECS runner). The Checks workflow runs the same tests on ordinary PRs so they are not gated on the query-regression / heavy-regression labels.
    • cargo fmt --all -- --check
    • cargo build -p cmd --bin query_perf_fixture --features dev-tools
    • cargo build -p cmd --bin query_regression_runner --features dev-tools
    • exercise the outer lifecycle script and Rust fixture generator against all built-in cases when the DSL or workflow case selection changes.