mirror of
https://github.com/GreptimeTeam/greptimedb.git
synced 2026-09-06 05:28:57 +00:00
* feat(ci): add aliyun ecs ephemeral runner path for query regression Signed-off-by: paomian <xpaomian@gmail.com> * fix: improve condition for query-regression job execution in workflow * feat: update Docker installation to use official repository and add GPG key handling * Refactor query regression runner setup and configuration - Removed deprecated PersistentVolumeClaim for build cache. - Introduced a new bootstrap script for setting up the ECS runner host. - Deleted obsolete Helm values files for runner configuration. - Updated the Aliyun ECS runner provisioning script to reflect new cache paths. - Modified GitHub workflows to use the new Aliyun ECS runner setup. - Adjusted documentation to clarify the new runner lifecycle and provisioning process. * fix: enhance runner service management during bootstrap process * fix: update alibabacloud_tea_openapi dependency version in metadata * feat: enhance ECS runner scripts with region_id and resource_group_id support * fix: move containerd content store to data root for improved storage management * feat: rename query-regression runner to ephemeral-github runner and update related scripts * fix: update sentinel polling method to use serial console output for improved reliability * fix: add environment variable checks for Alibaba Cloud access keys in ECS client * fix: improve error handling in GitHub API requests for better diagnostics * fix: improve cache disk detection logic for Aliyun ECS instances * fix: enhance cache disk waiting logic with detailed output and error handling * fix: update dependency version for alibabacloud_tea_openapi in teardown script * fix: enhance cache disk waiting logic for better compatibility and clarity * fix: enhance console output handling and add incremental logging during instance provisioning * fix: add PATH environment variable for runner jobs in service and provision script * fix: add machine telemetry sampling and logging during query regression jobs * fix: update query regression documentation and provision script for cache disk handling * fix: update SCCACHE_CACHE_SIZE validation to 10G for improved caching efficiency * fix: remove outdated cache size checks and cleanup logic for fresh system disk runs * fix: enhance instance deletion logic with region handling and console output export * fix: add swap file setup and OOM handling for ECS runner to improve stability * fix: update OOM handling and service restart logic for ECS runner to enhance stability * fix: increase system disk size to 100 GiB for cold double nightly builds to prevent ENOSPC errors * fix: increase system disk size to 150 GiB for ECS runner to prevent ENOSPC errors * fix: add keep_instance option to preserve ECS instance for post-mortem debugging * fix: disable unattended upgrades to prevent job cancellations during library updates * fix: reduce system disk size to 40 GiB for ECS runner to prevent ENOSPC errors * feat: Refactor Aliyun ECS runner provisioning and introduce nightly regression comparison - Update `aliyun-ecs-runner-provision.py` to remove cache disk handling, simplifying the provisioning process. - Introduce `query-regression-nightly-refs.py` to resolve and compare SHAs from successful nightly builds. - Create `query-regression-nightly.yml` workflow to trigger nightly comparisons based on successful builds. - Enhance `query-regression.yml` to include a `test-tooling` job for validating Python scripts before provisioning. - Update tests for the new nightly reference selection logic and refactor existing tests to align with the new caching strategy. - Modify documentation to reflect changes in caching and nightly comparison workflows. * fix: enhance runner image tool verification with detailed checks * fix: improve error handling in runner image tool verification * fix: update tool versions in ECS image and workflow for consistency * fix: correct typo in error message for unparseable ECS creation time * fix: update README and workflow files for query regression tests and image hygiene --------- Signed-off-by: paomian <xpaomian@gmail.com>
2.9 KiB
2.9 KiB
Agent Guidelines for Query Performance Tests
- Keep GitHub Actions YAML thin. Put non-trivial control flow, case expansion,
report generation, and metadata writing in scripts under
.github/scripts/; workflow steps should mostly invoke those scripts. - Runner lifecycle: the default path provisions one ephemeral Aliyun ECS
instance per run via
.github/scripts/aliyun-ecs-runner-provision.pyand always releases it viaaliyun-ecs-runner-teardown.py; a scheduled janitor workflow sweeps leftovers. Build caches live on that instance's system disk and are discarded with the VM. Runs do not share a workflow concurrency group. The ECS custom image is built from the runner Dockerfile by.github/runner-scale-sets/query-regression/ecs-image/build-ecs-image.py; keep the Dockerfile the single source of the tool contract. Dispatching with any otherrunnervalue treats it as a literal self-hosted runner label (seeecs-image/bootstrap-runner-host.shfor preparing such a host). - Query regression PR runs should build base/candidate binaries once, then run
the default case set. Do not hard-code a single case such as
promql_pushdown_7913into the workflow path. - Scheduled nightly comparison lives in
query-regression-nightly.yml: it waits for a successful Nightly Build, then callsquery-regression.ymlwith the previous vs current nightly SHAs. Keep SHA selection in.github/scripts/query-regression-nightly-refs.py. - The case DSL is not required to keep compatibility inside this PR. When the DSL changes, update TOML cases, the outer lifecycle script, Rust helpers, and docs together.
[case]is report metadata only.[scenario]is the executable regression configuration and must includekind, data layout, tables, queries, and thresholds. Rust owns case schema, defaults, validation, and normalized plan output throughquery_perf_fixture plan..github/scripts/query-regression-run.pyowns process lifecycle;query_regression_runnerconsumes normalized plans, frontend endpoints, direct-SST materialization requests, and OTLP target/finalize requests.- Keep the direct-SST generator generic. Issue-specific behavior belongs in case files and thresholds, not in Rust generator logic.
- Before pushing perf harness changes, run at least:
- the Python tests in the
test-toolingjob of.github/workflows/query-regression.yml(ubuntu-latest, not the ECS runner). The Checks workflow runs the same tests on ordinary PRs so they are not gated on thequery-regression/heavy-regressionlabels. cargo fmt --all -- --checkcargo build -p cmd --bin query_perf_fixture --features dev-toolscargo build -p cmd --bin query_regression_runner --features dev-tools- exercise the outer lifecycle script and Rust fixture generator against all built-in cases when the DSL or workflow case selection changes.
- the Python tests in the