* fix(ci): teach check-builder-rust-version.sh to handle stable channels
The script extracted a YYYY-MM-DD date from rust-toolchain.toml to
compare against the rustc build date inside the dev-builder image —
a nightly-era design. With channel = "1.96.1" there is no date in
the file, so every release build failed with 'Error: No rust toolchain
version found in rust-toolchain.toml'.
Extract the channel token instead and branch on it:
- stable channel (X.Y[.Z]): require the builder image's rustc to
exactly match the pinned version
- nightly-YYYY-MM-DD: keep the legacy date-difference check
Verified against a mocked docker for all four paths (stable match /
mismatch, nightly fresh / stale).
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* chore(toolchain): finish stable-migration cleanup in docs and query-regression pin
- README/AGENTS: the toolchain is now stable Rust pinned by
rust-toolchain.toml, not nightly
- query-regression: align the benchmark toolchain pin with the
workspace (nightly-2026-03-21 = 1.96.0-nightly -> stable 1.96.1),
including the exact-version assertions (cargo 356927216, rustc
31fca3adb, both 2026-06-26) and the runner image default
The query-regression runner image must be rebuilt and
QUERY_REGRESSION_ECS_IMAGE_ID bumped together with these pins
(per .github/runner-scale-sets/query-regression/README.md) before
the next regression run.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci(query-regression): derive the Rust toolchain pin from rust-toolchain.toml
Replace the hard-coded RUSTUP_TOOLCHAIN value and the hard-coded
version strings in the runner Verify assertions with a pin resolved
from rust-toolchain.toml:
- the always-running test-tooling job exports the channel parsed from
rust-toolchain.toml as a job output
- query-regression sets RUSTUP_TOOLCHAIN from that output
- the Verify step escapes the pin into the cargo/rustc/active-toolchain
regexes at runtime; the exact commit hash and date are asserted
generically since a stable version identifies the release
Removing the redundant require_eq (workflow yaml vs runner env) since
both now flow from the single source of truth. When rust-toolchain.toml
is bumped, the run fails with a clear signal until the runner image is
rebuilt with the new toolchain, keeping the existing lockstep contract.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci(query-regression): derive the runner image toolchain from rust-toolchain.toml
Remove the hard-coded 'ARG RUST_TOOLCHAIN=1.96.1' from the
query-regression runner Dockerfile. The pin is now parsed from a
COPY'd rust-toolchain.toml at build time (the bootstrap script builds
with the repo root as context, so the file is in the build context):
- rustup-init installs the parsed channel as the default toolchain
- the baked ENV RUSTUP_TOOLCHAIN is dropped: the rustup default makes
bare cargo/rustc resolve correctly without it, and the workflow
supplies RUSTUP_TOOLCHAIN explicitly at run time
- the build-time self-verification asserts the active toolchain
against the same parsed pin
With this, rust-toolchain.toml is the single source of truth for the
benchmark toolchain end to end: the image bakes whatever the toml says
at build time and the workflow asserts against the toml at run time.
A toolchain bump now only requires rebuilding the image.
The changed mechanics were verified natively with the real rustup-init
1.29.0 and the real 1.96.1 toolchain (registry pulls are unavailable
in this sandbox): parsing, default-toolchain installation without
RUSTUP_TOOLCHAIN env, bare cargo/rustc resolution, and the
active-toolchain assertion for both '(default)' and '(overridden)'
output forms.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci(query-regression): rebuild the runner image automatically on toolchain changes
Mirror the dev-builder automation for the query-regression ECS runner
image: a new rebuild-query-regression-runner-image.yaml workflow runs
whenever rust-toolchain.toml or the query-regression runner directory
changes on main (or via manual dispatch). It drives the existing
build-ecs-image.py ops tool, then completes the documented lockstep
updates in order: bump RUNNER_IMAGE_EPOCH in query-regression.yml and
push the commit to main, and only then point the
QUERY_REGRESSION_ECS_IMAGE_ID repo variable at the new image, so the
next regression run picks up image and epoch together.
Also fix build-ecs-image.py to stage rust-toolchain.toml into the
temporary docker build context: the AMI path builds the embedded
Dockerfile from an empty /tmp/image-context, which would break on the
Dockerfile's COPY of rust-toolchain.toml introduced earlier. The
user-data now base64-stages the toml next to the Dockerfile before
docker build.
Verified: py_compile, render_user_data round-trip (mkdir -> stage ->
docker build ordering), and the RUNNER_IMAGE_EPOCH bump sed against
the real workflow file. Requires a new ALIYUN_ECS_BASE_IMAGE_ID repo
variable (Ubuntu 24.04 public image id in the region); all other
secrets/vars are shared with the provisioning job.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci: fold the query-regression runner rebuild into release-dev-builder-images.yaml
Merge the standalone rebuild workflow into the existing builder-image
release workflow, as one entry point for all builder artifacts:
- push paths extended with .github/runner-scale-sets/query-regression/**
- new 'release_query_regression_runner_image' dispatch input
- a 'changes' job diffs the pushed range (github.event.before..sha,
with an everything-changed fallback for dispatch or unknown bases)
so each expensive rebuild only fires for its own paths:
rust-toolchain.toml gates both, docker/dev-builder/** gates the
dev-builder images, the query-regression runner directory gates the
ECS image rebuild
- the rebuild job itself is unchanged from the standalone workflow
(build-ecs-image.py, then RUNNER_IMAGE_EPOCH commit to main, then
the QUERY_REGRESSION_ECS_IMAGE_ID variable update)
The dev-builder jobs, their ECR/CN/tag-update dependents, and the
runner rebuild now share one workflow; the changes filter preserves
the previous on-push behavior for the dev-builder images while the
runner rebuild keeps its own trigger. The epoch-bump commit only
touches query-regression.yml, which is outside the trigger paths, so
no re-trigger loop.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix(ci): auto-resolve the ECS base image for the runner rebuild
The automated rebuild failed with 'Missing required configuration:
--base-image-id' because the ALIYUN_ECS_BASE_IMAGE_ID repo variable
does not exist yet (it was flagged as a one-time setup item).
Remove the setup dependency instead: build-ecs-image.py now defaults
--base-image-id to the latest public Ubuntu 24.04 x86_64 system image
in the region (DescribeImages with image_owner_alias=system), so no
manual variable is required. The runner Dockerfile pins every tool
version itself, so base-image drift is low-risk; --base-image-id or
the ALIYUN_ECS_BASE_IMAGE_ID variable still pin a specific base image
deterministically, and the workflow only passes the flag when the
variable is set.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* docs(query-regression): clarify what an image rebuild requires
A routine runner-image rebuild needs no manual file updates: the
rebuild job updates QUERY_REGRESSION_ECS_IMAGE_ID and
RUNNER_IMAGE_EPOCH; the toolchain derives from rust-toolchain.toml;
uv, sccache, otelgen, rustup, and the runner base are pinned by
digest/sha/commit in the Dockerfile. Only an apt package revision
bump (mold, protoc, python3) between rebuilds requires bumping the
corresponding Verify pins, and that failure is loud with the observed
version.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix(ci): correct SDK field names in the base-image resolver
DescribeImagesRequest takes 'ostype' (not 'os_type') and the image
items expose 'osname'/'osname_en' (not 'os_name') in the pinned
alibabacloud_ecs20140526 SDK range, so the auto-resolution added in
5a9fd2c769 crashed with a TypeError before describing anything.
Fix the request fields, move the architecture filter server-side, and
paginate (page_size=100 until a short page) instead of relying on a
single default-sized response. Match Ubuntu 24.04 on the localized
osname or the English osname_en.
Verified against the real SDK models (uv run --with
'alibabacloud_ecs20140526>=4.1.0,<6'): a two-page fake client picks
the newest Ubuntu 24.04 via osname_en and rejects 22.04/Windows
decoys.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix(ci): correct the repo-root path in build-ecs-image.py
ASSETS_DIR.parent.parent lands on runner-scale-sets, not the repo
root -- the toml lookup failed with FileNotFoundError. The root is
four levels above ecs-image; express it as an explicit REPO_ROOT
constant (ASSETS_DIR.parents[3]).
Verified every path main() reads against the real checkout layout
(Dockerfile, rust-toolchain.toml, start-runner.sh, the systemd unit,
plus REPO_ROOT sanity against Cargo.toml/.git), re-checked the
user-data toml staging round-trip, and re-ran the base-image
resolver regression test.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
Query regression self-hosted runners
The Query Regression workflow runs on Aliyun ECS ephemeral runners
(aliyun-ecs, the default and only automated path): a provision job on
ubuntu-latest creates one pay-as-you-go ECS instance per run, and the
instance registers itself as an ephemeral GitHub runner with a per-run
label. A teardown job (if: always()) deletes the instance;
query-regression-janitor.yml sweeps tagged leftovers older than 4 hours
daily.
Build caches live on the instance's system disk, so every run compiles cold; only the within-run reuse (base build warms the candidate build through the shared target dir and sccache) applies. There is no retained data disk.
Dispatching with any other runner value uses it as a literal self-hosted
runner label, which is how a manually prepared host (see
ecs-image/bootstrap-runner-host.sh) runs the workflow. For PR comment
admission, set the repository variable QUERY_REGRESSION_PR_RUNNER to such
a label to redirect those runs away from ECS.
The office ARC scale set perf-regression-8-cores that previously ran this
workflow is retired; see git history for its values files and pause/deploy
procedures. Tearing down the office cluster (helm release, runner namespace,
and the query-regression-build-cache PVC) is a manual operator action
outside this repository. This directory keeps its historical
runner-scale-sets name for path stability.
Nightly vs previous nightly
query-regression-nightly.yml runs after a successful GreptimeDB Nightly Build (workflow_run). It resolves that run's head_sha as the candidate
and the previous successful nightly (same branch, typically Friday when
Monday's nightly fires) as the base, then calls query-regression.yml with
those immutable SHAs. Both binaries are still compiled on the ECS runner;
this is not an artifact-download path. workflow_dispatch can pass explicit
base_ref / candidate_ref or a nightly run id. If there is no previous
nightly, or both nightlies built the same commit, the comparison is skipped.
Aliyun ECS path
Configuration lives in repository variables/secrets:
| Kind | Name | Purpose |
|---|---|---|
| secret | ALICLOUD_ECS_ACCESS_KEY_ID / ALICLOUD_ECS_ACCESS_KEY_SECRET |
RAM user scoped to ECS RunInstances/DeleteInstances/Describe*/CreateImage/RunCommand. Used only by provision/teardown jobs on ubuntu-latest; never reaches the ECS instance. |
| secret | GH_PERSONAL_ACCESS_TOKEN |
Creates the short-lived runner registration token (shared with the jsonbench EC2 path). Slash-command-dispatch uses github.token with contents: write for same-repo repository_dispatch. |
| secret | QUERY_REGRESSION_ADMISSION_HMAC |
HMAC-SHA256 key for the hidden admission marker comment. Injected only into the ubuntu-latest parse job and the sticky-comment workflow. Never reference it in query-regression.yml (the ECS job would inherit it). Empty fails closed: admit refuses, comments are skipped. |
| vars | ALIYUN_ECS_REGION_ID / ALIYUN_ECS_VSWITCH_ID / ALIYUN_ECS_SECURITY_GROUP_ID |
Network placement. The security group should allow egress only; no inbound rules are needed. The vSwitch pins the zone. |
| vars | ALIYUN_ECS_INSTANCE_TYPE |
Dedicated (non-burstable, non-shared) instance family. Prefer 32 GiB (e.g. ecs.g8i.2xlarge); ecs.c9i.2xlarge is 8c16g and nightly thin-LTO of greptime peaks above that. Both base and candidate clusters run on the same machine, so noisy neighbors break thresholds. |
| vars | QUERY_REGRESSION_ECS_IMAGE_ID |
Custom image built by ecs-image/build-ecs-image.py. |
| vars | QUERY_REGRESSION_COMMENT_ALLOWLIST |
Comma/whitespace-separated GitHub logins allowed to comment /query-regression on a PR. Each login must also have repository admin permission. Empty denies all comment commands. |
The system disk is 80 GiB, providing capacity headroom for the image, a 16 GiB
swapfile, the checkout, and cold build caches (target dir, cargo registry,
sccache). ENOSPC stops the runner itself from writing logs, which GitHub reports
as The operation was canceled with no telemetry, indistinguishable from a
platform-side cancellation.
cloud-init masks systemd-oomd, disables unattended-upgrades /
apt-daily-upgrade, creates /swapfile, and sets OOMPolicy=continue
on the runner unit. Ubuntu 24.04 defaults to DefaultOOMPolicy=stop,
which SIGTERM-s the whole unit when rustc is OOM-killed and GitHub
reports The operation was canceled with no telemetry. Unattended
upgrades can do the same via systemctl restart of the runner after a
library update; GitHub then records UserCancelled even though the job
was still valid. Swap is a safety net for 16 GiB types, not a substitute
for 32 GiB; linking on swap is slow.
When a run dies with an unexplained The operation was canceled (no
"Canceled by" banner, healthy machine), re-dispatch with keep_instance
checked: the teardown job is skipped and the instance survives for
post-mortem inspection. The security group is egress-only, so inspect via
Cloud Assistant (RunCommand) or VNC: the runner's _diag logs under the
runner home record reconnects and worker crashes, journalctl -u ephemeral-github-runner.service mirrors the console stream, dmesg -T
shows kernel OOM kills, and machine-telemetry.log in the job workspace
has the 30 s sampler history. The janitor sweep still deletes the instance
after its TTL, so finish the inspection within that window.
Trust model on the ECS path: the instance receives only the one-hour runner
registration token via user data and holds no cloud credentials; the Aliyun
AK/SK exist only in the control-plane jobs. Instance tags
(managed-by=query-regression-ci, query-regression-run-id) feed the janitor
sweep.
Building and updating the ECS image
The runner Dockerfile in the parent directory stays the single source of the
tool contract. The image is rebuilt automatically by the
rebuild-query-regression-runner-image job in
.github/workflows/release-dev-builder-images.yaml whenever
rust-toolchain.toml or anything under this directory changes on main (or via
manual dispatch): it runs the ops tool below, then bumps
RUNNER_IMAGE_EPOCH in query-regression.yml and points the
QUERY_REGRESSION_ECS_IMAGE_ID repo variable at the new image, so the next
regression run picks up image and epoch together.
The manual fallback (also what the workflow runs):
ALIBABA_CLOUD_ACCESS_KEY_ID=... ALIBABA_CLOUD_ACCESS_KEY_SECRET=... \
uv run .github/runner-scale-sets/query-regression/ecs-image/build-ecs-image.py \
--region-id <region> --vswitch-id <vsw-...> --security-group-id <sg-...> \
[--base-image-id <ubuntu-24.04-image-id>]
--base-image-id is optional: the script defaults to the latest public
Ubuntu 24.04 image in the region (the Dockerfile pins every tool version
itself, so base drift is low-risk); pass it — or set the
ALIYUN_ECS_BASE_IMAGE_ID repo variable consumed by the automated job —
to pin a specific base image.
The script boots a temporary builder instance, docker builds the runner
image, materializes /opt/rustup, /opt/cargo, /usr/local/bin tools, and
/home/runner (actions-runner) onto the host, installs the ephemeral-runner
systemd unit from ecs-image/, and snapshots a custom image. It prints the
image id; set it as QUERY_REGRESSION_ECS_IMAGE_ID, and bump
RUNNER_IMAGE_EPOCH in query-regression.yml at the same time so the target
cache invalidates. Builder sentinel polling requires the Cloud Assistant
agent, which Aliyun public Ubuntu images include.
Trust admission for PR runs
An allowlisted repository admin commenting /query-regression on the PR is
trust admission for that exact PR revision. /query-regression runs the
nine routine default cases; /query-regression heavy runs only the
high-cardinality prom_remote_write_7913 remote-write case.
slash-command-dispatch.yml (issue_comment on the default branch) parses
the command and repository_dispatches; query-regression-slash.yml admits
the revision so secrets work for fork PRs. Dispatch must come from
github-actions[bot]. The handler requires the current PR head to equal the
head SHA snapshotted in that dispatch (comment time, not handler start). If
the head moved while queued, admission denies; comment /query-regression
again after reviewing the new revision. It then snapshots merge, head, and
base SHAs at admission. A queued job fetches that immutable event merge SHA
directly, verifies it is a two-parent merge whose parents include the
snapshotted head exactly once, and uses its other parent as the actual base
build revision. The ubuntu-latest admission job posts a hidden HMAC-signed
marker comment on the admitted PR and uploads query-regression-admission as
a lookup hint. The sticky-comment workflow verifies that marker (and that the
runner artifact matches it) before posting. Candidate code on ECS shares the
run and can overwrite artifacts, but it cannot forge the marker: the reusable
workflow's GITHUB_TOKEN is contents: read only, and
QUERY_REGRESSION_ADMISSION_HMAC is never referenced there so it never
reaches ECS. The snapshotted event base is retained for audit only, so a
difference from the merge's non-head
parent is not a failure. The job never
follows a newer mutable PR merge ref. An unavailable event merge, or one that
does not contain exactly one snapshotted head parent, fails closed. An
already-admitted run is not retargeted by a later push; cancel it if it is no
longer wanted.
The commenter must be in QUERY_REGRESSION_COMMENT_ALLOWLIST and have
repository admin permission.
Admission does not relax runner hardening or GitHub permissions. Keep the ECS instance free of cloud credentials and long-lived tokens, keep the security group egress-only, keep GitHub tokens least-privilege, and review workflow changes before admission.
Runner image and workflow tools
The runner Dockerfile builds otelgen from
WenyXu/otelgen commit
863a3f395d062c7322cc1de08a38774b7fdaa6c8
so trace cases do not download or compile tools during a benchmark run. It is
built only as a toolchain factory: build-ecs-image.py and
bootstrap-runner-host.sh both materialize its contents onto a host, and the
benchmark itself runs host-native.
Before builds, the workflow asserts the runner UID/GID (1001 in the ECS image;
overridable via QUERY_REGRESSION_RUNNER_UID/QUERY_REGRESSION_RUNNER_GID)
and exact tool versions: libprotoc 3.21.12, uv 0.11.26, mold 2.40.4,
Python 3.14.4, sccache 0.16.0, otelgen commit
863a3f395d062c7322cc1de08a38774b7fdaa6c8, root-owned rustup 1.29.0, and
the image-baked Rust toolchain matching rust-toolchain.toml (the image
parses the pin from the toml at build time, and the workflow asserts it
dynamically at run time — there is no separately pinned toolchain version).
mold and python3
come from apt at image-build time (not Ubuntu 24.04's default 3.12); if the
Ubuntu archive ships a newer package revision between rebuilds, the Verify
step fails with the observed version — bump those pins in query-regression.yml
when that happens. Everything else (toolchain, uv, sccache, otelgen, rustup,
the runner base) is pinned by digest/sha/commit in the Dockerfile, and the
image id + RUNNER_IMAGE_EPOCH updates are automated by the rebuild job, so
a routine rebuild needs no manual file updates. Rustup, Cargo, and Rustc
must resolve from /opt/cargo/bin; the runner cannot write /opt/rustup or
/opt/cargo/bin. Protobuf well-known includes, including
google/protobuf/any.proto and google/protobuf/empty.proto, are an image
contract and must compile with protoc.
actions-rust-lang/setup-rust-toolchain@v1 is intentionally removed. The
workflow sets its warning-denying mold RUSTFLAGS directly, disables
automatic Rustup installation, and performs no runtime toolchain downloads.
The workflow no longer uses GitHub rust-cache, setup-protoc, setup-uv,
or runtime Rust setup: the image establishes immutable executable state and
each instance's system disk holds only that run's Cargo data. Do not
reintroduce those actions unless the image contract changes.
Capacity
Each run provisions its own ECS instance, so overlapping dispatches proceed
in parallel. There is no workflow concurrency group. All Cargo state
(CARGO_HOME including registry/git, CARGO_TARGET_DIR, sccache, cache
metadata) lives on the 80 GiB system disk and is discarded with the VM.
RUSTUP_HOME=/opt/rustup and /opt/cargo/bin are image-owned. The runner
sets RUSTC_WRAPPER=/usr/local/bin/sccache,
SCCACHE_DIR=/home/runner/.cache/sccache, SCCACHE_CACHE_SIZE=10G, and
CARGO_INCREMENTAL=0. sccache uses its local disk backend and self-evicts
at 10G. Base and candidate builds share the target on that disk; Cargo
fingerprints invalidate source and dependency changes.
The workflow reports du / df in telemetry. It does not try to reclaim
space across runs: a cold 80 GiB disk that fills up fails the build.
Future optional phases
The current phase uses a materialized runner toolchain with sccache 0.16.0 and
no shared cache service. Optional follow-ups are an image additionally seeded
with cargo fetch --locked results, or an internal read/write sccache
backend or Cargo/Git mirror. Evaluate them only if cold compile time on the
system disk becomes the bottleneck.