mirror of
https://github.com/GreptimeTeam/greptimedb.git
synced 2026-10-04 19:15:34 +00:00
75bd8e9ce6fa554c2da3fa7d72dc63c8b50e99ce
15
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
43eaea7a9a |
fix(ci): teach check-builder-rust-version.sh to handle stable channels (#9369)
* fix(ci): teach check-builder-rust-version.sh to handle stable channels
The script extracted a YYYY-MM-DD date from rust-toolchain.toml to
compare against the rustc build date inside the dev-builder image —
a nightly-era design. With channel = "1.96.1" there is no date in
the file, so every release build failed with 'Error: No rust toolchain
version found in rust-toolchain.toml'.
Extract the channel token instead and branch on it:
- stable channel (X.Y[.Z]): require the builder image's rustc to
exactly match the pinned version
- nightly-YYYY-MM-DD: keep the legacy date-difference check
Verified against a mocked docker for all four paths (stable match /
mismatch, nightly fresh / stale).
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* chore(toolchain): finish stable-migration cleanup in docs and query-regression pin
- README/AGENTS: the toolchain is now stable Rust pinned by
rust-toolchain.toml, not nightly
- query-regression: align the benchmark toolchain pin with the
workspace (nightly-2026-03-21 = 1.96.0-nightly -> stable 1.96.1),
including the exact-version assertions (cargo 356927216, rustc
31fca3adb, both 2026-06-26) and the runner image default
The query-regression runner image must be rebuilt and
QUERY_REGRESSION_ECS_IMAGE_ID bumped together with these pins
(per .github/runner-scale-sets/query-regression/README.md) before
the next regression run.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci(query-regression): derive the Rust toolchain pin from rust-toolchain.toml
Replace the hard-coded RUSTUP_TOOLCHAIN value and the hard-coded
version strings in the runner Verify assertions with a pin resolved
from rust-toolchain.toml:
- the always-running test-tooling job exports the channel parsed from
rust-toolchain.toml as a job output
- query-regression sets RUSTUP_TOOLCHAIN from that output
- the Verify step escapes the pin into the cargo/rustc/active-toolchain
regexes at runtime; the exact commit hash and date are asserted
generically since a stable version identifies the release
Removing the redundant require_eq (workflow yaml vs runner env) since
both now flow from the single source of truth. When rust-toolchain.toml
is bumped, the run fails with a clear signal until the runner image is
rebuilt with the new toolchain, keeping the existing lockstep contract.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci(query-regression): derive the runner image toolchain from rust-toolchain.toml
Remove the hard-coded 'ARG RUST_TOOLCHAIN=1.96.1' from the
query-regression runner Dockerfile. The pin is now parsed from a
COPY'd rust-toolchain.toml at build time (the bootstrap script builds
with the repo root as context, so the file is in the build context):
- rustup-init installs the parsed channel as the default toolchain
- the baked ENV RUSTUP_TOOLCHAIN is dropped: the rustup default makes
bare cargo/rustc resolve correctly without it, and the workflow
supplies RUSTUP_TOOLCHAIN explicitly at run time
- the build-time self-verification asserts the active toolchain
against the same parsed pin
With this, rust-toolchain.toml is the single source of truth for the
benchmark toolchain end to end: the image bakes whatever the toml says
at build time and the workflow asserts against the toml at run time.
A toolchain bump now only requires rebuilding the image.
The changed mechanics were verified natively with the real rustup-init
1.29.0 and the real 1.96.1 toolchain (registry pulls are unavailable
in this sandbox): parsing, default-toolchain installation without
RUSTUP_TOOLCHAIN env, bare cargo/rustc resolution, and the
active-toolchain assertion for both '(default)' and '(overridden)'
output forms.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci(query-regression): rebuild the runner image automatically on toolchain changes
Mirror the dev-builder automation for the query-regression ECS runner
image: a new rebuild-query-regression-runner-image.yaml workflow runs
whenever rust-toolchain.toml or the query-regression runner directory
changes on main (or via manual dispatch). It drives the existing
build-ecs-image.py ops tool, then completes the documented lockstep
updates in order: bump RUNNER_IMAGE_EPOCH in query-regression.yml and
push the commit to main, and only then point the
QUERY_REGRESSION_ECS_IMAGE_ID repo variable at the new image, so the
next regression run picks up image and epoch together.
Also fix build-ecs-image.py to stage rust-toolchain.toml into the
temporary docker build context: the AMI path builds the embedded
Dockerfile from an empty /tmp/image-context, which would break on the
Dockerfile's COPY of rust-toolchain.toml introduced earlier. The
user-data now base64-stages the toml next to the Dockerfile before
docker build.
Verified: py_compile, render_user_data round-trip (mkdir -> stage ->
docker build ordering), and the RUNNER_IMAGE_EPOCH bump sed against
the real workflow file. Requires a new ALIYUN_ECS_BASE_IMAGE_ID repo
variable (Ubuntu 24.04 public image id in the region); all other
secrets/vars are shared with the provisioning job.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci: fold the query-regression runner rebuild into release-dev-builder-images.yaml
Merge the standalone rebuild workflow into the existing builder-image
release workflow, as one entry point for all builder artifacts:
- push paths extended with .github/runner-scale-sets/query-regression/**
- new 'release_query_regression_runner_image' dispatch input
- a 'changes' job diffs the pushed range (github.event.before..sha,
with an everything-changed fallback for dispatch or unknown bases)
so each expensive rebuild only fires for its own paths:
rust-toolchain.toml gates both, docker/dev-builder/** gates the
dev-builder images, the query-regression runner directory gates the
ECS image rebuild
- the rebuild job itself is unchanged from the standalone workflow
(build-ecs-image.py, then RUNNER_IMAGE_EPOCH commit to main, then
the QUERY_REGRESSION_ECS_IMAGE_ID variable update)
The dev-builder jobs, their ECR/CN/tag-update dependents, and the
runner rebuild now share one workflow; the changes filter preserves
the previous on-push behavior for the dev-builder images while the
runner rebuild keeps its own trigger. The epoch-bump commit only
touches query-regression.yml, which is outside the trigger paths, so
no re-trigger loop.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix(ci): auto-resolve the ECS base image for the runner rebuild
The automated rebuild failed with 'Missing required configuration:
--base-image-id' because the ALIYUN_ECS_BASE_IMAGE_ID repo variable
does not exist yet (it was flagged as a one-time setup item).
Remove the setup dependency instead: build-ecs-image.py now defaults
--base-image-id to the latest public Ubuntu 24.04 x86_64 system image
in the region (DescribeImages with image_owner_alias=system), so no
manual variable is required. The runner Dockerfile pins every tool
version itself, so base-image drift is low-risk; --base-image-id or
the ALIYUN_ECS_BASE_IMAGE_ID variable still pin a specific base image
deterministically, and the workflow only passes the flag when the
variable is set.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* docs(query-regression): clarify what an image rebuild requires
A routine runner-image rebuild needs no manual file updates: the
rebuild job updates QUERY_REGRESSION_ECS_IMAGE_ID and
RUNNER_IMAGE_EPOCH; the toolchain derives from rust-toolchain.toml;
uv, sccache, otelgen, rustup, and the runner base are pinned by
digest/sha/commit in the Dockerfile. Only an apt package revision
bump (mold, protoc, python3) between rebuilds requires bumping the
corresponding Verify pins, and that failure is loud with the observed
version.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix(ci): correct SDK field names in the base-image resolver
DescribeImagesRequest takes 'ostype' (not 'os_type') and the image
items expose 'osname'/'osname_en' (not 'os_name') in the pinned
alibabacloud_ecs20140526 SDK range, so the auto-resolution added in
|
||
|
|
4ecec69bee |
ci: add manual agent observability benchmarks on Aliyun ECS (#9179)
* ci: add manual agent observability benchmarks on Aliyun ECS Signed-off-by: WenyXu <wenymedia@gmail.com> * ci: configure observability ECS budgets and reuse runner actions Signed-off-by: WenyXu <wenymedia@gmail.com> * ci: default observability runners to ecs.c9i.2xlarge Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(ci): prepare observability Docker access during ECS bootstrap Signed-off-by: WenyXu <wenymedia@gmail.com> * docs: remove standalone observability CI guide Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
4d65e8984a |
chore(ci): Implement /query-regression command handling and admission workflow (#8975)
* Implement `/query-regression` command handling and admission workflow - Add `query-regression-slash.py` script for processing `/query-regression` commands in PR comments, validating case arguments, and checking permissions. - Update `checks.yml` to include tests for the new slash command functionality. - Modify `query-regression-comment.yml` to trigger on the new `Query Regression Command` workflow. - Create `query-regression-slash.yml` to handle the dispatched command, validate allowlist and permissions, and initiate the regression workflow. - Enhance `query-regression.yml` to support additional inputs for PR admission and SHA verification. - Introduce `slash-command-dispatch.yml` to parse and dispatch commands from PR comments. - Document the new command admission process in `AGENTS.md` and `README.md`. - Add unit tests in `test_query_regression_slash.py` to cover command parsing and admission logic. * refactor: enhance query-regression command handling with comment validation and identity checks * feat: implement admission identity handling for query regression workflows * refactor: update PR admission logic in query regression workflow * refactor: update token usage in slash command dispatch and README for clarity * test: add cases for handling re-run failed jobs and stale runner artifacts * refactor: improve repository metadata handling in query regression scripts * chore: enable overwrite for artifact uploads to handle re-run failed jobs * chore: enable overwrite for query regression admission uploads * feat: enhance query-regression admission with HMAC signing and verification - Introduced HMAC signing for admission markers in query-regression workflows to ensure integrity and authenticity. - Updated `query-regression-comment.test.cjs` to include tests for signing and verifying admission markers. - Modified `query-regression-slash.py` to handle admission marker signing and verification, including checks for dispatch sender and head SHA consistency. - Enhanced workflows to securely manage admission markers and HMAC secrets, ensuring they are not exposed to untrusted contexts. - Improved documentation to clarify the admission process and the role of HMAC in securing the workflow. * test: add case to find newly posted marker among newer comments * test: add case to verify multiline output handling in write_outputs function |
||
|
|
0de0c01283 | fix: add disk usage logging to GitHub step summary in query regression workflow (#9005) | ||
|
|
35ea88a4ef |
feat(ci): run query regression on ephemeral Aliyun ECS runners (#8937)
* feat(ci): add aliyun ecs ephemeral runner path for query regression Signed-off-by: paomian <xpaomian@gmail.com> * fix: improve condition for query-regression job execution in workflow * feat: update Docker installation to use official repository and add GPG key handling * Refactor query regression runner setup and configuration - Removed deprecated PersistentVolumeClaim for build cache. - Introduced a new bootstrap script for setting up the ECS runner host. - Deleted obsolete Helm values files for runner configuration. - Updated the Aliyun ECS runner provisioning script to reflect new cache paths. - Modified GitHub workflows to use the new Aliyun ECS runner setup. - Adjusted documentation to clarify the new runner lifecycle and provisioning process. * fix: enhance runner service management during bootstrap process * fix: update alibabacloud_tea_openapi dependency version in metadata * feat: enhance ECS runner scripts with region_id and resource_group_id support * fix: move containerd content store to data root for improved storage management * feat: rename query-regression runner to ephemeral-github runner and update related scripts * fix: update sentinel polling method to use serial console output for improved reliability * fix: add environment variable checks for Alibaba Cloud access keys in ECS client * fix: improve error handling in GitHub API requests for better diagnostics * fix: improve cache disk detection logic for Aliyun ECS instances * fix: enhance cache disk waiting logic with detailed output and error handling * fix: update dependency version for alibabacloud_tea_openapi in teardown script * fix: enhance cache disk waiting logic for better compatibility and clarity * fix: enhance console output handling and add incremental logging during instance provisioning * fix: add PATH environment variable for runner jobs in service and provision script * fix: add machine telemetry sampling and logging during query regression jobs * fix: update query regression documentation and provision script for cache disk handling * fix: update SCCACHE_CACHE_SIZE validation to 10G for improved caching efficiency * fix: remove outdated cache size checks and cleanup logic for fresh system disk runs * fix: enhance instance deletion logic with region handling and console output export * fix: add swap file setup and OOM handling for ECS runner to improve stability * fix: update OOM handling and service restart logic for ECS runner to enhance stability * fix: increase system disk size to 100 GiB for cold double nightly builds to prevent ENOSPC errors * fix: increase system disk size to 150 GiB for ECS runner to prevent ENOSPC errors * fix: add keep_instance option to preserve ECS instance for post-mortem debugging * fix: disable unattended upgrades to prevent job cancellations during library updates * fix: reduce system disk size to 40 GiB for ECS runner to prevent ENOSPC errors * feat: Refactor Aliyun ECS runner provisioning and introduce nightly regression comparison - Update `aliyun-ecs-runner-provision.py` to remove cache disk handling, simplifying the provisioning process. - Introduce `query-regression-nightly-refs.py` to resolve and compare SHAs from successful nightly builds. - Create `query-regression-nightly.yml` workflow to trigger nightly comparisons based on successful builds. - Enhance `query-regression.yml` to include a `test-tooling` job for validating Python scripts before provisioning. - Update tests for the new nightly reference selection logic and refactor existing tests to align with the new caching strategy. - Modify documentation to reflect changes in caching and nightly comparison workflows. * fix: enhance runner image tool verification with detailed checks * fix: improve error handling in runner image tool verification * fix: update tool versions in ECS image and workflow for consistency * fix: correct typo in error message for unparseable ECS creation time * fix: update README and workflow files for query regression tests and image hygiene --------- Signed-off-by: paomian <xpaomian@gmail.com> |
||
|
|
1693b2727c |
refactor: port query regression runner to Rust (#8651)
* refactor: port query regression runner to Rust Signed-off-by: discord9 <discord9@163.com> * ci: remove optional OTLP report plotter Signed-off-by: discord9 <discord9@163.com> * refactor: split query regression runner into modules Signed-off-by: discord9 <discord9@163.com> * style: use crate-qualified imports in query regression runner Signed-off-by: discord9 <discord9@163.com> * refactor: simplify query regression runner internals Signed-off-by: discord9 <discord9@163.com> * feat: abstract inspect-footer storage access behind object store destination Add an optional --destination <TOML> to inspect-footer (and --base-destination/--candidate-destination to finalize-remote) so the storage inspection reads DB data files through the opendal-backed object_store abstraction instead of bare std::fs. Local paths keep working unchanged via the --root shortcut (File backend); remote backends (S3/GCS/...) are described by a DestinationConfig TOML reusing the object-store crate's ObjectStoreConfig serde shape. - inspect_footer: list via ObjectStore::list + ObjectMeta filtering (parquet keys, non-zero size, metadata/ segment), read footers async via ParquetObjectReader + ParquetMetaDataReader with known file size (no extra HEAD); output JSON schema unchanged - finalize-remote: --base-data-home/--candidate-data-home become optional, mutually exclusive with the new --*-destination args - cmd deps: add object_store_opendal + datafusion_object_store - tests: fs-backend list+footer integration tests (metadata filtering, destination TOML mode, root/destination exclusivity) Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * style: drop needless borrow in inspect footer test Fix clippy::needless_borrows_for_generic_args in the inspect-footer test (fs::create_dir_all(table.join("metadata"))). Missed by the earlier focused clippy run because it only covered --bin targets. Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> --------- Signed-off-by: discord9 <discord9@163.com> Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> |
||
|
|
457a3f317e |
ci: add heavy query regression label (#8619)
Signed-off-by: discord9 <discord9@163.com> |
||
|
|
f5f5d468bb |
ci: add OTLP trace ingestion regression testing (#8631)
* chore: update CI config Signed-off-by: shuiyisong <xixing.sys@gmail.com> * chore: add CI Signed-off-by: shuiyisong <xixing.sys@gmail.com> * chore: update CI config Signed-off-by: shuiyisong <xixing.sys@gmail.com> * ci: report otelgen runner diagnostics Signed-off-by: shuiyisong <xixing.sys@gmail.com> * chore: add script to draw result diagram Signed-off-by: shuiyisong <xixing.sys@gmail.com> --------- Signed-off-by: shuiyisong <xixing.sys@gmail.com> |
||
|
|
a636dcf6b2 |
ci: gate releases on compat and query regression (#8469)
Signed-off-by: discord9 <discord9@163.com> |
||
|
|
a3816c889e |
fix(ci): harden query regression runner (#8534)
Signed-off-by: discord9 <discord9@163.com> |
||
|
|
16217ff567 |
ci: add persistent query regression cache (#8474)
* ci: add persistent query regression cache Signed-off-by: discord9 <discord9@163.com> * ci: add sccache to query regression runner Signed-off-by: discord9 <discord9@163.com> * ci: pin query regression label revision Signed-off-by: discord9 <discord9@163.com> * ci: isolate query regression toolchain state Signed-off-by: discord9 <discord9@163.com> --------- Signed-off-by: discord9 <discord9@163.com> |
||
|
|
02283b6ba0 |
test(perf): add remote write storage inspection (#8444)
* feat: add remote write value distributions Signed-off-by: discord9 <discord9@163.com> * test: extend remote write value distributions Signed-off-by: discord9 <discord9@163.com> * test: inspect remote write parquet storage Signed-off-by: discord9 <discord9@163.com> * test: normalize remote write perf fixtures Signed-off-by: discord9 <discord9@163.com> * test: add heavy remote write perf case Signed-off-by: discord9 <discord9@163.com> * test: use head greptime for read bench Signed-off-by: discord9 <discord9@163.com> * test: keep heavy remote write case local Signed-off-by: discord9 <discord9@163.com> * test: tune remote write perf smoke case Signed-off-by: discord9 <discord9@163.com> * test: fix query fixture import style Signed-off-by: discord9 <discord9@163.com> * test: cover remote write value distributions Signed-off-by: discord9 <discord9@163.com> * test: add integer counter perf case Signed-off-by: discord9 <discord9@163.com> --------- Signed-off-by: discord9 <discord9@163.com> |
||
|
|
12f83828b1 |
feat: add Prom remote-write query regression scenario (#8413)
* feat: add Prom remote-write query regression scenario Signed-off-by: discord9 <discord9@163.com> * test: add high-cardinality remote-write query case Signed-off-by: discord9 <discord9@163.com> * feat: chunk remote-write query regression loads Signed-off-by: discord9 <discord9@163.com> * test: use multi-day remote-write regression case Signed-off-by: discord9 <discord9@163.com> * fix: address query regression review comments Signed-off-by: discord9 <discord9@163.com> * ci: allow large query regression comments Signed-off-by: discord9 <discord9@163.com> --------- Signed-off-by: discord9 <discord9@163.com> |
||
|
|
d8f692f4c5 |
ci: run query regression on self-hosted runners (#8423)
* ci: run query regression on self-hosted runners Signed-off-by: discord9 <discord9@163.com> * ci: use dedicated perf regression runner labels Signed-off-by: discord9 <discord9@163.com> * ci: keep query regression runner scale set minimal Signed-off-by: discord9 <discord9@163.com> * ci: use custom query regression runner image Signed-off-by: discord9 <discord9@163.com> * ci: host query regression runner image in acr Signed-off-by: discord9 <discord9@163.com> * ci: harden query regression runner workflow Signed-off-by: discord9 <discord9@163.com> * ci: avoid runner uid assumptions in values Signed-off-by: discord9 <discord9@163.com> * ci: fix query regression comments for fork prs Signed-off-by: discord9 <discord9@163.com> * ci: address query regression review comments Signed-off-by: discord9 <discord9@163.com> --------- Signed-off-by: discord9 <discord9@163.com> |
||
|
|
c44f8da646 |
feat: add query regression perf harness (#8406)
* feat: add query regression perf harness Signed-off-by: discord9 <discord9@163.com> * feat: extend query regression cases Signed-off-by: discord9 <discord9@163.com> * ci: harden query regression workflows Signed-off-by: discord9 <discord9@163.com> * fix: address query regression review comments Signed-off-by: discord9 <discord9@163.com> * ci: limit query regression PR triggers Signed-off-by: discord9 <discord9@163.com> * ci: run full query regression case set Signed-off-by: discord9 <discord9@163.com> * refactor: model query regression scenarios Signed-off-by: discord9 <discord9@163.com> * fix: avoid unenforced query regression thresholds Signed-off-by: discord9 <discord9@163.com> --------- Signed-off-by: discord9 <discord9@163.com> |