mirror of
https://github.com/GreptimeTeam/greptimedb.git
synced 2026-09-05 13:08:58 +00:00
feat(ci): run query regression on ephemeral Aliyun ECS runners (#8937)
* feat(ci): add aliyun ecs ephemeral runner path for query regression Signed-off-by: paomian <xpaomian@gmail.com> * fix: improve condition for query-regression job execution in workflow * feat: update Docker installation to use official repository and add GPG key handling * Refactor query regression runner setup and configuration - Removed deprecated PersistentVolumeClaim for build cache. - Introduced a new bootstrap script for setting up the ECS runner host. - Deleted obsolete Helm values files for runner configuration. - Updated the Aliyun ECS runner provisioning script to reflect new cache paths. - Modified GitHub workflows to use the new Aliyun ECS runner setup. - Adjusted documentation to clarify the new runner lifecycle and provisioning process. * fix: enhance runner service management during bootstrap process * fix: update alibabacloud_tea_openapi dependency version in metadata * feat: enhance ECS runner scripts with region_id and resource_group_id support * fix: move containerd content store to data root for improved storage management * feat: rename query-regression runner to ephemeral-github runner and update related scripts * fix: update sentinel polling method to use serial console output for improved reliability * fix: add environment variable checks for Alibaba Cloud access keys in ECS client * fix: improve error handling in GitHub API requests for better diagnostics * fix: improve cache disk detection logic for Aliyun ECS instances * fix: enhance cache disk waiting logic with detailed output and error handling * fix: update dependency version for alibabacloud_tea_openapi in teardown script * fix: enhance cache disk waiting logic for better compatibility and clarity * fix: enhance console output handling and add incremental logging during instance provisioning * fix: add PATH environment variable for runner jobs in service and provision script * fix: add machine telemetry sampling and logging during query regression jobs * fix: update query regression documentation and provision script for cache disk handling * fix: update SCCACHE_CACHE_SIZE validation to 10G for improved caching efficiency * fix: remove outdated cache size checks and cleanup logic for fresh system disk runs * fix: enhance instance deletion logic with region handling and console output export * fix: add swap file setup and OOM handling for ECS runner to improve stability * fix: update OOM handling and service restart logic for ECS runner to enhance stability * fix: increase system disk size to 100 GiB for cold double nightly builds to prevent ENOSPC errors * fix: increase system disk size to 150 GiB for ECS runner to prevent ENOSPC errors * fix: add keep_instance option to preserve ECS instance for post-mortem debugging * fix: disable unattended upgrades to prevent job cancellations during library updates * fix: reduce system disk size to 40 GiB for ECS runner to prevent ENOSPC errors * feat: Refactor Aliyun ECS runner provisioning and introduce nightly regression comparison - Update `aliyun-ecs-runner-provision.py` to remove cache disk handling, simplifying the provisioning process. - Introduce `query-regression-nightly-refs.py` to resolve and compare SHAs from successful nightly builds. - Create `query-regression-nightly.yml` workflow to trigger nightly comparisons based on successful builds. - Enhance `query-regression.yml` to include a `test-tooling` job for validating Python scripts before provisioning. - Update tests for the new nightly reference selection logic and refactor existing tests to align with the new caching strategy. - Modify documentation to reflect changes in caching and nightly comparison workflows. * fix: enhance runner image tool verification with detailed checks * fix: improve error handling in runner image tool verification * fix: update tool versions in ECS image and workflow for consistency * fix: correct typo in error message for unparseable ECS creation time * fix: update README and workflow files for query regression tests and image hygiene --------- Signed-off-by: paomian <xpaomian@gmail.com>
This commit is contained in:
@@ -1,321 +1,184 @@
|
||||
# Query regression self-hosted runners
|
||||
|
||||
The `Query Regression` workflow uses the dedicated ARC runner scale set
|
||||
`perf-regression-8-cores`. ARC runner Pods run in the target Kubernetes cluster
|
||||
and connect outbound to GitHub. The live scale set is currently **paused**:
|
||||
`minRunners=0`, `maxRunners=0`, and no runner Pods. Do not resume it without
|
||||
explicit approval.
|
||||
The `Query Regression` workflow runs on **Aliyun ECS ephemeral runners**
|
||||
(`aliyun-ecs`, the default and only automated path): a `provision` job on
|
||||
`ubuntu-latest` creates one pay-as-you-go ECS instance per run, and the
|
||||
instance registers itself as an ephemeral GitHub runner with a per-run
|
||||
label. A `teardown` job (`if: always()`) deletes the instance;
|
||||
`query-regression-janitor.yml` sweeps tagged leftovers older than 4 hours
|
||||
daily.
|
||||
|
||||
## Prerequisites and trust admission
|
||||
Build caches live on the instance's system disk, so **every run compiles
|
||||
cold**; only the within-run reuse (base build warms the candidate build
|
||||
through the shared target dir and sccache) applies. There is no retained
|
||||
data disk.
|
||||
|
||||
Install the ARC scale set controller if it is not already installed:
|
||||
Dispatching with any other `runner` value uses it as a literal self-hosted
|
||||
runner label, which is how a manually prepared host (see
|
||||
`ecs-image/bootstrap-runner-host.sh`) runs the workflow. For PR labels, set
|
||||
the repository variable `QUERY_REGRESSION_PR_RUNNER` to such a label to
|
||||
redirect PR runs away from ECS.
|
||||
|
||||
The office ARC scale set `perf-regression-8-cores` that previously ran this
|
||||
workflow is retired; see git history for its values files and pause/deploy
|
||||
procedures. Tearing down the office cluster (helm release, runner namespace,
|
||||
and the `query-regression-build-cache` PVC) is a manual operator action
|
||||
outside this repository. This directory keeps its historical
|
||||
`runner-scale-sets` name for path stability.
|
||||
|
||||
## Nightly vs previous nightly
|
||||
|
||||
`query-regression-nightly.yml` runs after a successful `GreptimeDB Nightly
|
||||
Build` (`workflow_run`). It resolves that run's `head_sha` as the candidate
|
||||
and the previous successful nightly (same branch, typically Friday when
|
||||
Monday's nightly fires) as the base, then calls `query-regression.yml` with
|
||||
those immutable SHAs. Both binaries are still compiled on the ECS runner;
|
||||
this is not an artifact-download path. `workflow_dispatch` can pass explicit
|
||||
`base_ref` / `candidate_ref` or a nightly run id. If there is no previous
|
||||
nightly, or both nightlies built the same commit, the comparison is skipped.
|
||||
|
||||
## Aliyun ECS path
|
||||
|
||||
Configuration lives in repository variables/secrets:
|
||||
|
||||
| Kind | Name | Purpose |
|
||||
| --- | --- | --- |
|
||||
| secret | `ALICLOUD_ECS_ACCESS_KEY_ID` / `ALICLOUD_ECS_ACCESS_KEY_SECRET` | RAM user scoped to ECS RunInstances/DeleteInstances/Describe*/CreateImage/RunCommand. Used only by provision/teardown jobs on `ubuntu-latest`; never reaches the ECS instance. |
|
||||
| secret | `GH_PERSONAL_ACCESS_TOKEN` | Creates the short-lived runner registration token (shared with the jsonbench EC2 path). |
|
||||
| vars | `ALIYUN_ECS_REGION_ID` / `ALIYUN_ECS_VSWITCH_ID` / `ALIYUN_ECS_SECURITY_GROUP_ID` | Network placement. The security group should allow egress only; no inbound rules are needed. The vSwitch pins the zone. |
|
||||
| vars | `ALIYUN_ECS_INSTANCE_TYPE` | Dedicated (non-burstable, non-shared) instance family. Prefer 32 GiB (e.g. `ecs.g8i.2xlarge`); `ecs.c9i.2xlarge` is 8c16g and nightly thin-LTO of greptime peaks above that. Both base and candidate clusters run on the same machine, so noisy neighbors break thresholds. |
|
||||
| vars | `QUERY_REGRESSION_ECS_IMAGE_ID` | Custom image built by `ecs-image/build-ecs-image.py`. |
|
||||
|
||||
The system disk is 40 GiB, which covers the image, a 16 GiB swapfile, the
|
||||
checkout, and cold build caches (target dir, cargo registry, sccache). ENOSPC
|
||||
stops the runner itself from writing logs, which GitHub reports as `The
|
||||
operation was canceled` with no telemetry, indistinguishable from a
|
||||
platform-side cancellation.
|
||||
cloud-init masks `systemd-oomd`, disables `unattended-upgrades` /
|
||||
`apt-daily-upgrade`, creates `/swapfile`, and sets `OOMPolicy=continue`
|
||||
on the runner unit. Ubuntu 24.04 defaults to `DefaultOOMPolicy=stop`,
|
||||
which SIGTERM-s the whole unit when rustc is OOM-killed and GitHub
|
||||
reports `The operation was canceled` with no telemetry. Unattended
|
||||
upgrades can do the same via `systemctl restart` of the runner after a
|
||||
library update; GitHub then records `UserCancelled` even though the job
|
||||
was still valid. Swap is a safety net for 16 GiB types, not a substitute
|
||||
for 32 GiB; linking on swap is slow.
|
||||
|
||||
When a run dies with an unexplained `The operation was canceled` (no
|
||||
"Canceled by" banner, healthy machine), re-dispatch with `keep_instance`
|
||||
checked: the teardown job is skipped and the instance survives for
|
||||
post-mortem inspection. The security group is egress-only, so inspect via
|
||||
Cloud Assistant (`RunCommand`) or VNC: the runner's `_diag` logs under the
|
||||
runner home record reconnects and worker crashes, `journalctl -u
|
||||
ephemeral-github-runner.service` mirrors the console stream, `dmesg -T`
|
||||
shows kernel OOM kills, and `machine-telemetry.log` in the job workspace
|
||||
has the 30 s sampler history. The janitor sweep still deletes the instance
|
||||
after its TTL, so finish the inspection within that window.
|
||||
|
||||
Trust model on the ECS path: the instance receives only the one-hour runner
|
||||
registration token via user data and holds no cloud credentials; the Aliyun
|
||||
AK/SK exist only in the control-plane jobs. Instance tags
|
||||
(`managed-by=query-regression-ci`, `query-regression-run-id`) feed the janitor
|
||||
sweep.
|
||||
|
||||
### Building and updating the ECS image
|
||||
|
||||
The runner `Dockerfile` in the parent directory stays the single source of the
|
||||
tool contract. Build a new ECS image from it:
|
||||
|
||||
```bash
|
||||
helm upgrade --install arc \
|
||||
oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set-controller \
|
||||
--namespace arc-systems \
|
||||
--create-namespace \
|
||||
--version 0.14.2
|
||||
ALIBABA_CLOUD_ACCESS_KEY_ID=... ALIBABA_CLOUD_ACCESS_KEY_SECRET=... \
|
||||
uv run .github/runner-scale-sets/query-regression/ecs-image/build-ecs-image.py \
|
||||
--region-id <region> --vswitch-id <vsw-...> --security-group-id <sg-...> \
|
||||
--base-image-id <ubuntu-24.04-image-id>
|
||||
```
|
||||
|
||||
Create the GitHub App secret in the runner namespace. Prefer an App limited to
|
||||
`GreptimeTeam/greptimedb`:
|
||||
The script boots a temporary builder instance, `docker build`s the runner
|
||||
image, materializes `/opt/rustup`, `/opt/cargo`, `/usr/local/bin` tools, and
|
||||
`/home/runner` (actions-runner) onto the host, installs the ephemeral-runner
|
||||
systemd unit from `ecs-image/`, and snapshots a custom image. It prints the
|
||||
image id; set it as `QUERY_REGRESSION_ECS_IMAGE_ID`, and bump
|
||||
`RUNNER_IMAGE_EPOCH` in `query-regression.yml` at the same time so the target
|
||||
cache invalidates. Builder sentinel polling requires the Cloud Assistant
|
||||
agent, which Aliyun public Ubuntu images include.
|
||||
|
||||
```bash
|
||||
kubectl -n arc-runners create secret generic greptimedb-arc-github-app \
|
||||
--from-literal=github_app_id=<app-id> \
|
||||
--from-literal=github_app_installation_id=<installation-id> \
|
||||
--from-file=github_app_private_key=<private-key.pem>
|
||||
```
|
||||
|
||||
The values files here reference that secret by name.
|
||||
## Trust admission for PR runs
|
||||
|
||||
A maintainer applying the `query-regression` or `heavy-regression` label is
|
||||
**trust admission for that exact PR revision**. `query-regression` runs the six
|
||||
routine default cases; `heavy-regression` runs only the high-cardinality
|
||||
`prom_remote_write_7913` remote-write case. The admitted job may use this scale
|
||||
set's dedicated, writable persistent cache. `pull_request: labeled` is the only
|
||||
PR trigger: the label event snapshots its merge, head, and base SHAs. A queued
|
||||
job fetches that immutable event merge SHA directly, verifies it is a two-parent
|
||||
merge whose parents include the snapshotted head exactly once, and uses its
|
||||
other parent as the actual base build revision. The snapshotted event base is
|
||||
retained for audit only, so a difference from the merge's non-head parent is
|
||||
not a failure. The job never follows a newer mutable PR merge ref. An
|
||||
unavailable event merge, or one that does not contain exactly one snapshotted
|
||||
head parent, fails closed. A later PR head change does not retarget an already
|
||||
queued run: it may execute only its previously trusted event revision if that
|
||||
revision remains fetchable. To run the new revision, the maintainer must review
|
||||
it, remove the label, and re-add the desired regression label; cancel the old
|
||||
run if it is no longer wanted. An existing label does not automatically rerun
|
||||
the benchmark.
|
||||
**trust admission for that exact PR revision**. `query-regression` runs the
|
||||
six routine default cases; `heavy-regression` runs only the high-cardinality
|
||||
`prom_remote_write_7913` remote-write case. `pull_request: labeled` is the only PR
|
||||
trigger: the label event snapshots its merge, head, and base SHAs. A queued
|
||||
job fetches that immutable event merge SHA directly, verifies it is a
|
||||
two-parent merge whose parents include the snapshotted head exactly once, and
|
||||
uses its other parent as the actual base build revision. The snapshotted
|
||||
event base is retained for audit only, so a difference from the merge's
|
||||
non-head parent is not a failure. The job never follows a newer mutable PR
|
||||
merge ref. An unavailable event merge, or one that does not contain exactly
|
||||
one snapshotted head parent, fails closed. A later PR head change does not
|
||||
retarget an already queued run: it may execute only its previously trusted
|
||||
event revision if that revision remains fetchable. To run the new revision,
|
||||
the maintainer must review it, remove the label, and re-add the desired
|
||||
regression label; cancel the old run if it is no longer wanted. An existing
|
||||
label does not automatically rerun the benchmark.
|
||||
|
||||
Admission does not relax runner hardening or GitHub permissions. Keep
|
||||
service-account token mounting disabled; do not mount host paths, the Docker
|
||||
socket, kubeconfig, or long-lived credentials. The runner and cache initializer
|
||||
use UID/GID 1001, disallow privilege escalation, drop all capabilities, and use
|
||||
the RuntimeDefault seccomp profile. Keep GitHub tokens least-privilege and
|
||||
review workflow changes before admission. Where the CNI supports it, restrict
|
||||
egress to required GitHub Actions, artifact/cache, Rust/crate/toolchain, DNS,
|
||||
and image-registry endpoints; block unrelated cluster services, private ranges,
|
||||
and metadata endpoints unless a case requires them.
|
||||
|
||||
### Network routing prerequisite
|
||||
|
||||
Required split routing is an **external environment-specific prerequisite**. The
|
||||
responsible network operator must route GitHub Actions, GitHub content,
|
||||
artifact/cache, crates.io, Rust toolchain, and image-registry traffic through
|
||||
the approved path rather than the VPN where applicable. Neither this repository
|
||||
nor Kubernetes configures that route. Verify it with the responsible network
|
||||
operator before any canary.
|
||||
Admission does not relax runner hardening or GitHub permissions. Keep the ECS
|
||||
instance free of cloud credentials and long-lived tokens, keep the security
|
||||
group egress-only, keep GitHub tokens least-privilege, and review workflow
|
||||
changes before admission.
|
||||
|
||||
## Runner image and workflow tools
|
||||
|
||||
Build and push the derived runner image; it preserves the official
|
||||
`/home/runner/run.sh` entrypoint and supplies CI tools needed at runtime. The
|
||||
image builds `otelgen` from
|
||||
The runner `Dockerfile` builds `otelgen` from
|
||||
[`WenyXu/otelgen`](https://github.com/WenyXu/otelgen) commit
|
||||
[`863a3f395d062c7322cc1de08a38774b7fdaa6c8`](https://github.com/WenyXu/otelgen/commit/863a3f395d062c7322cc1de08a38774b7fdaa6c8)
|
||||
so trace cases do not download or compile tools during a benchmark run:
|
||||
so trace cases do not download or compile tools during a benchmark run. It is
|
||||
built only as a toolchain factory: `build-ecs-image.py` and
|
||||
`bootstrap-runner-host.sh` both materialize its contents onto a host, and the
|
||||
benchmark itself runs host-native.
|
||||
|
||||
```bash
|
||||
docker build \
|
||||
--platform linux/amd64 \
|
||||
-f .github/runner-scale-sets/query-regression/Dockerfile \
|
||||
-t greptime-registry.cn-hangzhou.cr.aliyuncs.com/greptime/greptimedb-query-regression-runner:latest \
|
||||
.github/runner-scale-sets/query-regression
|
||||
|
||||
docker push greptime-registry.cn-hangzhou.cr.aliyuncs.com/greptime/greptimedb-query-regression-runner:latest
|
||||
```
|
||||
|
||||
Deploy by digest, not mutable tag, by updating both image references in
|
||||
`values-8-cores.yaml` after a rebuild. Update `RUNNER_IMAGE_DIGEST` and bump
|
||||
`RUNNER_IMAGE_EPOCH` in `query-regression.yml` at the same time. If the registry
|
||||
is private, use a dedicated read-only pull secret only as `imagePullSecrets`;
|
||||
never expose registry credentials to runner containers.
|
||||
Both digest-pinned init and runner containers use `IfNotPresent`: the immutable
|
||||
digest makes a cached image safe and avoids adding a registry dependency to every
|
||||
runner startup.
|
||||
|
||||
The runner optionally imports only the non-sensitive `HTTP_PROXY`,
|
||||
`HTTPS_PROXY`, and `NO_PROXY` variables from the
|
||||
`query-regression-runner-local-env` ConfigMap. Manage that ConfigMap locally in
|
||||
the target namespace; private endpoint configuration must not be committed, and
|
||||
credentials or secrets must never be placed in a ConfigMap.
|
||||
|
||||
Before builds, the workflow asserts UID/GID 1001 and exact image tool versions:
|
||||
`libprotoc 3.21.12`, `uv 0.11.26`, `mold 2.30.0`, `Python 3.12.3`, `sccache
|
||||
0.16.0`, `otelgen` commit `863a3f395d062c7322cc1de08a38774b7fdaa6c8`,
|
||||
root-owned `rustup 1.29.0`, and the image-baked
|
||||
`nightly-2026-03-21` Rust toolchain. Rustup, Cargo, and Rustc must resolve from
|
||||
`/opt/cargo/bin`; the runner cannot write `/opt/rustup` or `/opt/cargo/bin`.
|
||||
Protobuf well-known includes, including `google/protobuf/any.proto` and
|
||||
`google/protobuf/empty.proto`, are an image contract and must compile with
|
||||
`protoc`.
|
||||
Before builds, the workflow asserts the runner UID/GID (1001 in the ECS image;
|
||||
overridable via `QUERY_REGRESSION_RUNNER_UID`/`QUERY_REGRESSION_RUNNER_GID`)
|
||||
and exact tool versions: `libprotoc 3.21.12`, `uv 0.11.26`, `mold 2.40.4`,
|
||||
`Python 3.14.4`, `sccache 0.16.0`, `otelgen` commit
|
||||
`863a3f395d062c7322cc1de08a38774b7fdaa6c8`, root-owned `rustup 1.29.0`, and
|
||||
the image-baked `nightly-2026-03-21` Rust toolchain. `mold` and `python3`
|
||||
come from apt at image-build time (not Ubuntu 24.04's default 3.12); bump
|
||||
the Verify pins together with `QUERY_REGRESSION_ECS_IMAGE_ID` when the
|
||||
image is rebuilt. Rustup, Cargo, and Rustc
|
||||
must resolve from `/opt/cargo/bin`; the runner cannot write `/opt/rustup` or
|
||||
`/opt/cargo/bin`. Protobuf well-known includes, including
|
||||
`google/protobuf/any.proto` and `google/protobuf/empty.proto`, are an image
|
||||
contract and must compile with `protoc`.
|
||||
`actions-rust-lang/setup-rust-toolchain@v1` is intentionally removed. The
|
||||
workflow sets its warning-denying mold `RUSTFLAGS` directly, disables automatic
|
||||
Rustup installation, and performs no runtime toolchain downloads.
|
||||
workflow sets its warning-denying mold `RUSTFLAGS` directly, disables
|
||||
automatic Rustup installation, and performs no runtime toolchain downloads.
|
||||
|
||||
The workflow no longer uses GitHub `rust-cache`, `setup-protoc`, `setup-uv`, or
|
||||
runtime Rust setup: the image establishes immutable executable state and the PVC
|
||||
supplies only reusable Cargo data. Do not reintroduce those actions unless the
|
||||
corresponding cache or image contract changes.
|
||||
The workflow no longer uses GitHub `rust-cache`, `setup-protoc`, `setup-uv`,
|
||||
or runtime Rust setup: the image establishes immutable executable state and
|
||||
each instance's system disk holds only that run's Cargo data. Do not
|
||||
reintroduce those actions unless the image contract changes.
|
||||
|
||||
## Capacity and persistent cache
|
||||
## Capacity
|
||||
|
||||
`values-8-cores.yaml` is normal operation: `minRunners=0`, `maxRunners=1`.
|
||||
`values-paused.yaml` is the mandatory pause overlay: `minRunners=0`,
|
||||
`maxRunners=0`. The job uses group `query-regression-persistent-cache-v1`,
|
||||
`queue: max`, and `cancel-in-progress: false`; admitted jobs queue rather than
|
||||
replacing older pending jobs. During maintenance, cancel admitted queued runs as
|
||||
well as pausing ARC. Runner Pods have `activeDeadlineSeconds=12600`.
|
||||
The runner requests 6 CPU and limits at 8 CPU to preserve `minipc-3`
|
||||
allocatable-capacity scheduling headroom; do not reset the request to 8 CPU
|
||||
without revalidating scheduling capacity.
|
||||
Each run provisions its own ECS instance, so overlapping dispatches proceed
|
||||
in parallel. There is no workflow `concurrency` group. All Cargo state
|
||||
(`CARGO_HOME` including registry/git, `CARGO_TARGET_DIR`, sccache, cache
|
||||
metadata) lives on the 40 GiB system disk and is discarded with the VM.
|
||||
`RUSTUP_HOME=/opt/rustup` and `/opt/cargo/bin` are image-owned. The runner
|
||||
sets `RUSTC_WRAPPER=/usr/local/bin/sccache`,
|
||||
`SCCACHE_DIR=/home/runner/.cache/sccache`, `SCCACHE_CACHE_SIZE=10G`, and
|
||||
`CARGO_INCREMENTAL=0`. sccache uses its local disk backend and self-evicts
|
||||
at 10G. Base and candidate builds share the target on that disk; Cargo
|
||||
fingerprints invalidate source and dependency changes.
|
||||
|
||||
The cache claim `query-regression-build-cache` is a nominal 600Gi `local-path`
|
||||
PVC in `arc-runners`. It is `ReadWriteOnce`; `local-path` uses
|
||||
WaitForFirstConsumer binding and Delete reclaim behavior, produces a
|
||||
node-affine local PV, is non-expandable, and the 600Gi request is not a hard
|
||||
storage quota. The runner's `minipc-3` selector is its only consumer candidate.
|
||||
|
||||
The initializer mounts the PVC root at `/cache`, creates and write-tests these
|
||||
versioned subpaths as non-root UID/GID 1001, and the runner mounts them as:
|
||||
|
||||
| Persistent state | PVC subpath | Runner mount |
|
||||
| --- | --- | --- |
|
||||
| Ephemeral Cargo home | `emptyDir` | `/home/runner/.cargo` |
|
||||
| Cargo registry data | `cargo-registry-v1` | `/home/runner/.cargo/registry` |
|
||||
| Cargo Git data | `cargo-git-v1` | `/home/runner/.cargo/git` |
|
||||
| Cargo target | `query-regression-target-v1` | `/home/runner/query-regression-target` |
|
||||
| Cache metadata | `meta-v1` | `/home/runner/query-regression-cache-meta` |
|
||||
| sccache local disk cache | `sccache-v1` | `/home/runner/.cache/sccache` |
|
||||
| Immutable Rust toolchain | image-owned | `/opt/rustup`, `/opt/cargo/bin` |
|
||||
|
||||
The Pod security context uses UID/GID and `fsGroup` 1001 with
|
||||
`fsGroupChangePolicy: OnRootMismatch`; no privileged `chown` or raw `hostPath`
|
||||
is used. `CARGO_HOME` is a per-Pod `emptyDir`; only its nested `registry` and
|
||||
`git` mounts are persistent. `RUSTUP_HOME=/opt/rustup` and `/opt/cargo/bin` are
|
||||
image-owned immutable paths, while `CARGO_TARGET_DIR`, cache metadata, and
|
||||
`SCCACHE_DIR` are persistent absolute paths. The runner sets
|
||||
`RUSTC_WRAPPER=/usr/local/bin/sccache`,
|
||||
`SCCACHE_DIR=/home/runner/.cache/sccache`, `SCCACHE_CACHE_SIZE=40G`, and
|
||||
`CARGO_INCREMENTAL=0`. sccache uses its local PVC disk backend and self-evicts
|
||||
at 40G; do not add runtime downloads, object storage, or a shared backend.
|
||||
|
||||
The repository's `.cargo/config.toml` remains a trusted per-revision build input.
|
||||
In contrast, `$CARGO_HOME/config*`, credentials, installed bins, and Cargo
|
||||
metadata outside the persistent `registry` and `git` data mounts are ephemeral
|
||||
and cannot survive to another Pod.
|
||||
|
||||
The local disk backend has a one-server constraint. `maxRunners=1` and the
|
||||
unchanged `query-regression-persistent-cache-v1` workflow concurrency group
|
||||
serialize runs; do not increase runner capacity or relax that serialization
|
||||
while this backend is in use. Base and candidate builds share the target; Cargo
|
||||
fingerprints invalidate source and dependency changes. The workflow records the
|
||||
sccache version and relevant environment in the target ABI marker, starts and
|
||||
zeros sccache after cache and toolchain checks, shows initial/base/candidate
|
||||
statistics, and resets statistics between base and candidate builds.
|
||||
|
||||
### Disk preflight and cleanup contract
|
||||
|
||||
Before applying or unpausing, verify the backing filesystem on `minipc-3` has
|
||||
at least 900GiB free. The current local-path provisioner source is
|
||||
`/opt/local-path-provisioner`; measure the filesystem containing it:
|
||||
|
||||
```bash
|
||||
df -PB1G /opt/local-path-provisioner
|
||||
```
|
||||
|
||||
The workflow reports `du`, `df -P`, human-readable free space, and inode
|
||||
availability before builds and in an always-run report. Its cleanup is narrow
|
||||
and non-destructive:
|
||||
|
||||
- warn at target size 400GiB; at 450GiB clear only the complete target root;
|
||||
- warn at Cargo registry-plus-Git data size 60GiB; at 80GiB remove only
|
||||
`registry/src` and `git/checkouts`, then abort if that persistent data remains
|
||||
at least 80GiB;
|
||||
- below 300GiB backing free space, clear the complete target root first,
|
||||
remeasure, then remove only those Cargo extracted trees and checkouts; abort
|
||||
if free space is still below 300GiB;
|
||||
- never automatically remove Cargo registry cache/index, Git database, the
|
||||
image-owned Cargo bin or Rustup toolchain, cache metadata, the self-evicting
|
||||
sccache directory, or the PVC.
|
||||
|
||||
The target clear uses fixed absolute roots and removes all entries, including
|
||||
dotfiles. After migration validation, remove obsolete `cargo-home-v1` and
|
||||
`rustup-home-v1` only in explicit maintenance while ARC is 0/0 and no runner Pod
|
||||
exists; they are not mounted by the current configuration.
|
||||
|
||||
## Deploy and pause safely
|
||||
|
||||
First verify the configured external network route and the disk preflight. Apply the PVC; while
|
||||
the scale set is paused, it is expected to remain `Pending` because
|
||||
WaitForFirstConsumer has no scheduled runner:
|
||||
|
||||
```bash
|
||||
kubectl apply --dry-run=server \
|
||||
-f .github/runner-scale-sets/query-regression/cache-pvc.yaml
|
||||
kubectl apply -f .github/runner-scale-sets/query-regression/cache-pvc.yaml
|
||||
```
|
||||
|
||||
Render normal and paused configurations. Normal values are always first; the
|
||||
pause overlay is always last:
|
||||
|
||||
```bash
|
||||
helm template perf-regression-8-cores \
|
||||
oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set \
|
||||
--namespace arc-runners --version 0.14.2 \
|
||||
--set controllerServiceAccount.name=arc-gha-rs-controller \
|
||||
--set controllerServiceAccount.namespace=arc-systems \
|
||||
-f .github/runner-scale-sets/query-regression/values-8-cores.yaml
|
||||
|
||||
helm template perf-regression-8-cores \
|
||||
oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set \
|
||||
--namespace arc-runners --version 0.14.2 \
|
||||
--set controllerServiceAccount.name=arc-gha-rs-controller \
|
||||
--set controllerServiceAccount.namespace=arc-systems \
|
||||
-f .github/runner-scale-sets/query-regression/values-8-cores.yaml \
|
||||
-f .github/runner-scale-sets/query-regression/values-paused.yaml
|
||||
```
|
||||
|
||||
The **first post-merge Helm deployment must reconcile the release in paused
|
||||
mode**. Keep the pause overlay last:
|
||||
|
||||
```bash
|
||||
# First post-merge deployment and every return to paused mode: 0/0.
|
||||
helm upgrade --install perf-regression-8-cores \
|
||||
oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set \
|
||||
--namespace arc-runners --create-namespace --version 0.14.2 \
|
||||
--reset-values --wait \
|
||||
-f .github/runner-scale-sets/query-regression/values-8-cores.yaml \
|
||||
-f .github/runner-scale-sets/query-regression/values-paused.yaml
|
||||
|
||||
# Expect 0/0 and no runner resources before considering normal mode.
|
||||
kubectl -n arc-runners get autoscalingrunnerset perf-regression-8-cores \
|
||||
-o jsonpath='{.spec.minRunners}{"/"}{.spec.maxRunners}{"\n"}'
|
||||
kubectl -n arc-runners get ephemeralrunners,pods \
|
||||
-l actions.github.com/scale-set-name=perf-regression-8-cores
|
||||
```
|
||||
|
||||
Only after that verification and separate explicit approval, apply normal 0/1
|
||||
operation without the pause overlay:
|
||||
|
||||
```bash
|
||||
helm upgrade --install perf-regression-8-cores \
|
||||
oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set \
|
||||
--namespace arc-runners --create-namespace --version 0.14.2 \
|
||||
--reset-values --wait \
|
||||
-f .github/runner-scale-sets/query-regression/values-8-cores.yaml
|
||||
```
|
||||
|
||||
Do not use bare `helm rollback`, `--atomic`, or `--reuse-values`: a stored
|
||||
revision can restore nonzero runner capacity. Inspect rendered manifests for
|
||||
capacity, the `minipc-3` selector, cache claim and mounts, initializer,
|
||||
security context, and resources. After approved normal mode receives its first
|
||||
canary, the PVC binds to `minipc-3`.
|
||||
|
||||
For local-PV node loss, cache recovery is intentionally disposable: return to
|
||||
0/0, recreate the PVC on a healthy node, cold-fill it, and run a new canary.
|
||||
|
||||
## Canary and rollback
|
||||
|
||||
With explicit approval, run two identical `workflow_dispatch` canaries on
|
||||
`perf-regression-8-cores`, using immutable full base and candidate commit SHAs
|
||||
and `cargo_profile=nightly`. The first is the cold fill; the second verifies warm
|
||||
reuse. Record the workflow's base/candidate build elapsed logs and cache
|
||||
ABI-marker output, initial/base/candidate sccache statistics, and cache report.
|
||||
Confirm the image tool contract (root-owned Rustup/Cargo paths, baked nightly,
|
||||
and non-writable `/opt` roots) and that the ephemeral Cargo home contains only
|
||||
the mounted registry/Git data before Cargo creates per-Pod state.
|
||||
Obtain dependency and tool network byte counters from the configured
|
||||
environment counter source, filtered to `minipc-3` and the dependency/tool
|
||||
destinations.
|
||||
|
||||
Accept the canary only when all of the following hold:
|
||||
|
||||
- exactly one runner Pod runs on `minipc-3`, and both jobs use the same bound PV;
|
||||
- UID/GID 1001 cache mounts are writable; the warm run does not invalidate the
|
||||
target or bulk-redownload crates or toolchains; sccache reports separate base
|
||||
and candidate build statistics without server or cache-path errors; immutable
|
||||
Rustup/Cargo roots remain non-writable and only registry/Git data persists;
|
||||
- warm base build time is at most 50% of cold base build time;
|
||||
- warm dependency/tool network bytes are at most 10% of cold fill bytes;
|
||||
- cache sizes remain below soft watermarks, node free space remains at least
|
||||
300GiB, and the benchmark is correct without TLS EOFs or timeouts;
|
||||
- the configured environment counter source confirms this traffic is outside VPN
|
||||
accounting.
|
||||
|
||||
Immediately return to 0/0 after either canary unless ongoing normal operation
|
||||
has been explicitly approved; return immediately on any traffic, cache, disk,
|
||||
TLS, or correctness failure. To roll back, use the paused Helm upgrade above,
|
||||
or another explicit `helm upgrade --install` with known-good values followed by
|
||||
`values-paused.yaml`, `--reset-values`, and `--wait`. Do not delete the PVC
|
||||
automatically; preserve it for diagnosis unless intentionally discarding cache.
|
||||
The workflow reports `du` / `df` in telemetry. It does not try to reclaim
|
||||
space across runs: a cold 40 GiB disk that fills up fails the build.
|
||||
|
||||
## Future optional phases
|
||||
|
||||
The current phase uses a digest-pinned image with sccache 0.16.0 and no shared
|
||||
cache service. Optional follow-ups are an image additionally seeded with the
|
||||
exact Rust toolchain and `cargo fetch --locked`; or an internal read/write
|
||||
sccache backend or Cargo/Git mirror. Evaluate them only if persistent PVC reuse
|
||||
is insufficient.
|
||||
The current phase uses a materialized runner toolchain with sccache 0.16.0 and
|
||||
no shared cache service. Optional follow-ups are an image additionally seeded
|
||||
with `cargo fetch --locked` results, or an internal read/write sccache
|
||||
backend or Cargo/Git mirror. Evaluate them only if cold compile time on the
|
||||
system disk becomes the bottleneck.
|
||||
|
||||
@@ -1,15 +0,0 @@
|
||||
apiVersion: v1
|
||||
kind: PersistentVolumeClaim
|
||||
metadata:
|
||||
name: query-regression-build-cache
|
||||
namespace: arc-runners
|
||||
labels:
|
||||
app.kubernetes.io/part-of: perf-regression-8-cores
|
||||
spec:
|
||||
volumeMode: Filesystem
|
||||
accessModes:
|
||||
- ReadWriteOnce
|
||||
storageClassName: local-path
|
||||
resources:
|
||||
requests:
|
||||
storage: 600Gi
|
||||
@@ -0,0 +1,195 @@
|
||||
#!/usr/bin/env bash
|
||||
# Copyright 2023 Greptime Team
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# Bootstrap an Ubuntu 24.04 host as a query-regression self-hosted runner.
|
||||
#
|
||||
# This is the host-native equivalent of the ECS image contract: it
|
||||
# installs docker-ce from Docker's official repository, builds the runner
|
||||
# image from the sibling Dockerfile (the single source of the tool contract),
|
||||
# materializes the toolchain onto the host, and optionally registers the
|
||||
# runner and installs its systemd service.
|
||||
#
|
||||
# Large state lives under --data-root (a data disk): docker's data-root, the
|
||||
# runner home, and the rustup/cargo toolchain roots, symlinked back to the
|
||||
# hard-coded contract paths (/home/runner, /opt/rustup, /opt/cargo).
|
||||
#
|
||||
# The script is idempotent: re-running refreshes the toolchain in place, and
|
||||
# the runner registration (.runner/.credentials) and _work survive. A live
|
||||
# runner service is stopped first and restarted at the end.
|
||||
#
|
||||
# Usage:
|
||||
# sudo bash bootstrap-runner-host.sh --repo-dir /path/to/greptimedb
|
||||
# sudo RUNNER_TOKEN=<token> bash bootstrap-runner-host.sh --repo-dir . \
|
||||
# --register --repo-url https://github.com/<owner>/<repo>
|
||||
#
|
||||
# Options:
|
||||
# --repo-dir PATH greptimedb checkout containing the runner Dockerfile (required)
|
||||
# --data-root PATH data disk mount point (default: /data)
|
||||
# --uid / --gid N runner user id (default: 3141; the ECS image contract is 1001)
|
||||
# --register also register the runner and install the systemd service
|
||||
# --repo-url URL repository URL for registration (required with --register)
|
||||
# --runner-name N runner name (default: qreg-host)
|
||||
# --labels L runner labels (default: perf-regression-8-cores)
|
||||
#
|
||||
# With --register, provide a fresh registration token via RUNNER_TOKEN.
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
REPO_DIR=""
|
||||
DATA_ROOT="/data"
|
||||
RUNNER_UID="3141"
|
||||
RUNNER_GID="3141"
|
||||
REGISTER="false"
|
||||
REPO_URL=""
|
||||
RUNNER_NAME="qreg-host"
|
||||
RUNNER_LABELS="perf-regression-8-cores"
|
||||
IMAGE_TAG="qreg-runner:manual"
|
||||
|
||||
while [[ $# -gt 0 ]]; do
|
||||
case "$1" in
|
||||
--repo-dir) REPO_DIR="$2"; shift 2 ;;
|
||||
--data-root) DATA_ROOT="$2"; shift 2 ;;
|
||||
--uid) RUNNER_UID="$2"; shift 2 ;;
|
||||
--gid) RUNNER_GID="$2"; shift 2 ;;
|
||||
--register) REGISTER="true"; shift ;;
|
||||
--repo-url) REPO_URL="$2"; shift 2 ;;
|
||||
--runner-name) RUNNER_NAME="$2"; shift 2 ;;
|
||||
--labels) RUNNER_LABELS="$2"; shift 2 ;;
|
||||
-h | --help) sed -n '17,40p' "$0"; exit 0 ;;
|
||||
*) echo "Unknown argument: $1" >&2; exit 2 ;;
|
||||
esac
|
||||
done
|
||||
|
||||
[[ $EUID -eq 0 ]] || { echo "Run as root, e.g. sudo bash $0 ..." >&2; exit 1; }
|
||||
[[ -n "${REPO_DIR}" ]] || { echo "--repo-dir is required" >&2; exit 2; }
|
||||
DOCKERFILE="${REPO_DIR}/.github/runner-scale-sets/query-regression/Dockerfile"
|
||||
[[ -f "${DOCKERFILE}" ]] || { echo "Dockerfile not found at ${DOCKERFILE}" >&2; exit 1; }
|
||||
if [[ "${REGISTER}" == "true" ]]; then
|
||||
[[ -n "${REPO_URL}" ]] || { echo "--repo-url is required with --register" >&2; exit 2; }
|
||||
[[ -n "${RUNNER_TOKEN:-}" ]] || { echo "RUNNER_TOKEN env is required with --register" >&2; exit 2; }
|
||||
fi
|
||||
|
||||
step() { printf '\n==> %s\n' "$*"; }
|
||||
|
||||
# A re-run refreshes files under the (possibly live) runner service, so stop
|
||||
# it first and restart it at the end (unless --register re-creates it).
|
||||
RUNNER_SERVICE="$(systemctl list-units --type=service --all --no-legend 'actions.runner.*' 2>/dev/null | awk '{print $1}' | head -n 1 || true)"
|
||||
RUNNER_WAS_ACTIVE="false"
|
||||
if [[ -n "${RUNNER_SERVICE}" ]] && systemctl is-active --quiet "${RUNNER_SERVICE}"; then
|
||||
step "Stop running runner service ${RUNNER_SERVICE}"
|
||||
systemctl stop "${RUNNER_SERVICE}"
|
||||
RUNNER_WAS_ACTIVE="true"
|
||||
fi
|
||||
|
||||
step "Install docker-ce from Docker's official repository"
|
||||
for pkg in docker.io docker-doc docker-compose docker-compose-v2 podman-docker containerd runc; do
|
||||
apt-get remove -y "${pkg}" 2>/dev/null || true
|
||||
done
|
||||
apt-get update
|
||||
apt-get install -y ca-certificates curl
|
||||
install -m 0755 -d /etc/apt/keyrings
|
||||
curl -fsSL https://download.docker.com/linux/ubuntu/gpg -o /etc/apt/keyrings/docker.asc
|
||||
chmod a+r /etc/apt/keyrings/docker.asc
|
||||
# shellcheck disable=SC1091
|
||||
. /etc/os-release
|
||||
tee /etc/apt/sources.list.d/docker.sources > /dev/null <<EOF
|
||||
Types: deb
|
||||
URIs: https://download.docker.com/linux/ubuntu
|
||||
Suites: ${UBUNTU_CODENAME:-$VERSION_CODENAME}
|
||||
Components: stable
|
||||
Architectures: $(dpkg --print-architecture)
|
||||
Signed-By: /etc/apt/keyrings/docker.asc
|
||||
EOF
|
||||
apt-get update
|
||||
apt-get install -y docker-ce docker-ce-cli containerd.io docker-buildx-plugin
|
||||
|
||||
step "Move docker data-root to ${DATA_ROOT}/docker"
|
||||
systemctl stop docker containerd
|
||||
mkdir -p "${DATA_ROOT}/docker" "${DATA_ROOT}/containerd"
|
||||
if [[ -f /etc/docker/daemon.json ]] && ! grep -q '"data-root"' /etc/docker/daemon.json; then
|
||||
echo "ERROR: /etc/docker/daemon.json exists without data-root; merge manually." >&2
|
||||
exit 1
|
||||
fi
|
||||
echo "{ \"data-root\": \"${DATA_ROOT}/docker\" }" > /etc/docker/daemon.json
|
||||
# containerd keeps its own content store (where build-time image data lands);
|
||||
# move it off the system disk too.
|
||||
mkdir -p /etc/containerd
|
||||
containerd config default | sed "s|root = \"/var/lib/containerd\"|root = \"${DATA_ROOT}/containerd\"|" > /etc/containerd/config.toml
|
||||
systemctl start containerd docker
|
||||
docker info | grep "Docker Root Dir"
|
||||
|
||||
step "Build the runner image"
|
||||
docker build --platform linux/amd64 -f "${DOCKERFILE}" -t "${IMAGE_TAG}" "${REPO_DIR}"
|
||||
|
||||
step "Materialize the toolchain onto the host"
|
||||
mkdir -p "${DATA_ROOT}/opt" "${DATA_ROOT}/runner" /opt /home
|
||||
container="$(docker create "${IMAGE_TAG}")"
|
||||
trap 'docker rm -f "${container}" >/dev/null 2>&1 || true' EXIT
|
||||
# The `/.` suffix copies directory *contents*, so re-running over an existing
|
||||
# materialization refreshes it in place instead of nesting runner/runner.
|
||||
# Files absent from the image (runner registration, _work) are preserved.
|
||||
docker cp "${container}:/home/runner/." "${DATA_ROOT}/runner"
|
||||
docker cp "${container}:/opt/rustup/." "${DATA_ROOT}/opt/rustup"
|
||||
docker cp "${container}:/opt/cargo/." "${DATA_ROOT}/opt/cargo"
|
||||
for tool in uv uvx otelgen sccache; do
|
||||
docker cp "${container}:/usr/local/bin/${tool}" "/usr/local/bin/${tool}"
|
||||
done
|
||||
docker rm "${container}" > /dev/null
|
||||
trap - EXIT
|
||||
ln -sfn "${DATA_ROOT}/runner" /home/runner
|
||||
ln -sfn "${DATA_ROOT}/opt/rustup" /opt/rustup
|
||||
ln -sfn "${DATA_ROOT}/opt/cargo" /opt/cargo
|
||||
docker image rm "${IMAGE_TAG}" > /dev/null
|
||||
|
||||
step "Install system packages"
|
||||
apt-get install -y --no-install-recommends \
|
||||
build-essential clang cmake git gzip jq libprotobuf-dev libssl-dev mold \
|
||||
openssh-client pkg-config protobuf-compiler python3 sudo tar unzip wget xz-utils zip zstd
|
||||
|
||||
step "Create runner user (${RUNNER_UID}:${RUNNER_GID}) and environment"
|
||||
getent group "${RUNNER_GID}" > /dev/null || groupadd -g "${RUNNER_GID}" runner
|
||||
id -u runner > /dev/null 2>&1 || useradd -u "${RUNNER_UID}" -g "${RUNNER_GID}" -d /home/runner -s /bin/bash runner
|
||||
chown -R "${RUNNER_UID}:${RUNNER_GID}" "${DATA_ROOT}/runner"
|
||||
echo 'PATH=/opt/cargo/bin:/usr/local/bin:/usr/bin:/bin' > "${DATA_ROOT}/runner/.env"
|
||||
chown "${RUNNER_UID}:${RUNNER_GID}" "${DATA_ROOT}/runner/.env"
|
||||
|
||||
if [[ "${REGISTER}" == "true" ]]; then
|
||||
step "Register runner ${RUNNER_NAME} (labels: ${RUNNER_LABELS})"
|
||||
if [[ -n "${RUNNER_SERVICE}" ]]; then
|
||||
(cd /home/runner && ./svc.sh uninstall) || true
|
||||
fi
|
||||
runuser -u runner -- bash -c "cd /home/runner && HOME=/home/runner ./config.sh \
|
||||
--url '${REPO_URL}' --token '${RUNNER_TOKEN}' --name '${RUNNER_NAME}' \
|
||||
--labels '${RUNNER_LABELS}' --unattended --replace --disableupdate"
|
||||
|
||||
step "Install and start the runner service"
|
||||
cd /home/runner
|
||||
./svc.sh install runner
|
||||
./svc.sh start
|
||||
./svc.sh status
|
||||
else
|
||||
step "Skipping registration (pass --register --repo-url ... with RUNNER_TOKEN to enable)"
|
||||
if [[ "${RUNNER_WAS_ACTIVE}" == "true" ]]; then
|
||||
step "Restart runner service ${RUNNER_SERVICE}"
|
||||
systemctl start "${RUNNER_SERVICE}"
|
||||
systemctl --no-pager status "${RUNNER_SERVICE}" || true
|
||||
fi
|
||||
fi
|
||||
|
||||
step "Done"
|
||||
echo "Runner home: /home/runner -> ${DATA_ROOT}/runner"
|
||||
echo "Before each workflow run on this persistent host, clean transient cargo state:"
|
||||
echo " sudo rm -f /home/runner/.cargo/.package-cache"
|
||||
echo " sudo rm -rf /home/runner/.cargo/.global-cache"
|
||||
@@ -0,0 +1,392 @@
|
||||
#!/usr/bin/env python3
|
||||
# Copyright 2023 Greptime Team
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# PEP 723 inline metadata (see .github/scripts/aliyun-ecs-runner-provision.py
|
||||
# for the convention).
|
||||
# /// script
|
||||
# requires-python = ">=3.10"
|
||||
# dependencies = [
|
||||
# "alibabacloud_ecs20140526>=4.1.0,<6",
|
||||
# "alibabacloud_tea_openapi>=0.3.12,<1",
|
||||
# ]
|
||||
# ///
|
||||
|
||||
"""Build the query-regression ECS custom image (manual ops tool).
|
||||
|
||||
Boots a temporary pay-as-you-go ECS instance from a public Ubuntu 24.04 image,
|
||||
builds the existing runner container image (the Dockerfile in the parent
|
||||
directory remains the single source of the tool contract), materializes the
|
||||
tool directories onto the host filesystem so the workflow's "Verify runner
|
||||
image tools" step holds unchanged, installs the ephemeral-runner systemd unit,
|
||||
and snapshots the result as a custom image. The temporary instance is deleted
|
||||
afterwards.
|
||||
|
||||
Sentinel polling reads the instance's serial console output
|
||||
(GetInstanceConsoleOutput) and looks for marker lines the user data writes to
|
||||
/dev/console. This has no in-guest agent dependency.
|
||||
|
||||
Usage:
|
||||
uv run .github/runner-scale-sets/query-regression/ecs-image/build-ecs-image.py \
|
||||
--region-id cn-hangzhou --vswitch-id vsw-... --security-group-id sg-... \
|
||||
--base-image-id ubuntu_24_04_x64_20G_alibase_*.vhd
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import base64
|
||||
import os
|
||||
import time
|
||||
from pathlib import Path
|
||||
|
||||
ASSETS_DIR = Path(__file__).resolve().parent
|
||||
DONE_MARKER = "QREG_IMAGE_BUILD_DONE"
|
||||
FAILED_MARKER = "QREG_IMAGE_BUILD_FAILED"
|
||||
POLL_INTERVAL_SECONDS = 15
|
||||
CONSOLE_POLL_INTERVAL_SECONDS = 30
|
||||
BUILD_TIMEOUT_SECONDS = 60 * 60
|
||||
|
||||
# Same apt package contract as the runner Dockerfile; the base
|
||||
# actions-runner image is Ubuntu 24.04, so an Ubuntu 24.04 host resolves the
|
||||
# same tool versions (protoc 3.21.12, mold 2.40.4, Python 3.14.4).
|
||||
# Docker itself comes from Docker's official repository (docker-ce), not the
|
||||
# distribution-packaged docker.io.
|
||||
APT_PACKAGES = [
|
||||
"build-essential",
|
||||
"ca-certificates",
|
||||
"clang",
|
||||
"cmake",
|
||||
"curl",
|
||||
"git",
|
||||
"gpg",
|
||||
"gzip",
|
||||
"jq",
|
||||
"libprotobuf-dev",
|
||||
"libssl-dev",
|
||||
"mold",
|
||||
"openssh-client",
|
||||
"pkg-config",
|
||||
"protobuf-compiler",
|
||||
"python3",
|
||||
"sudo",
|
||||
"tar",
|
||||
"unzip",
|
||||
"wget",
|
||||
"xz-utils",
|
||||
"zip",
|
||||
"zstd",
|
||||
]
|
||||
|
||||
DOCKER_CE_PACKAGES = "docker-ce docker-ce-cli containerd.io docker-buildx-plugin"
|
||||
|
||||
|
||||
def render_user_data(dockerfile: str, start_runner: str, unit: str) -> str:
|
||||
dockerfile_b64 = base64.b64encode(dockerfile.encode()).decode()
|
||||
start_runner_b64 = base64.b64encode(start_runner.encode()).decode()
|
||||
unit_b64 = base64.b64encode(unit.encode()).decode()
|
||||
packages = " ".join(APT_PACKAGES)
|
||||
return f"""#!/bin/bash
|
||||
set -euo pipefail
|
||||
trap 'echo "{FAILED_MARKER} at line $LINENO" > /dev/console' ERR
|
||||
# Stream the full build log to the serial console so the poller (and anyone
|
||||
# watching GetInstanceConsoleOutput) sees real progress, not a silent login
|
||||
# prompt for the whole build.
|
||||
exec > >(tee -a /dev/console) 2>&1
|
||||
|
||||
apt-get update
|
||||
DEBIAN_FRONTEND=noninteractive apt-get install -y --no-install-recommends {packages}
|
||||
|
||||
# Docker from the official repository, not the distribution-packaged docker.io.
|
||||
install -m 0755 -d /etc/apt/keyrings
|
||||
curl -fsSL https://download.docker.com/linux/ubuntu/gpg | gpg --dearmor -o /etc/apt/keyrings/docker.gpg
|
||||
chmod a+r /etc/apt/keyrings/docker.gpg
|
||||
. /etc/os-release
|
||||
echo "deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.gpg] https://download.docker.com/linux/ubuntu ${{VERSION_CODENAME}} stable" \
|
||||
> /etc/apt/sources.list.d/docker.list
|
||||
apt-get update
|
||||
DEBIAN_FRONTEND=noninteractive apt-get install -y --no-install-recommends {DOCKER_CE_PACKAGES}
|
||||
|
||||
base64 -d > /tmp/Dockerfile <<'EOF'
|
||||
{dockerfile_b64}
|
||||
EOF
|
||||
mkdir -p /tmp/image-context
|
||||
docker build --platform linux/amd64 -f /tmp/Dockerfile -t qreg-runner:local /tmp/image-context
|
||||
|
||||
# Materialize the tool contract onto the host filesystem.
|
||||
container="$(docker create qreg-runner:local)"
|
||||
trap 'docker rm -f "${{container}}" >/dev/null 2>&1 || true' EXIT
|
||||
docker cp "${{container}}:/opt/rustup" /opt/rustup
|
||||
docker cp "${{container}}:/opt/cargo" /opt/cargo
|
||||
for tool in uv uvx otelgen sccache; do
|
||||
docker cp "${{container}}:/usr/local/bin/${{tool}}" "/usr/local/bin/${{tool}}"
|
||||
done
|
||||
docker cp "${{container}}:/home/runner" /home/runner
|
||||
|
||||
# Runner identity mirrors the container image (UID/GID 1001).
|
||||
groupadd --gid 1001 runner 2>/dev/null || true
|
||||
useradd --uid 1001 --gid 1001 --home-dir /home/runner --shell /bin/bash runner 2>/dev/null || true
|
||||
chown -R 1001:1001 /home/runner
|
||||
|
||||
mkdir -p /opt/ephemeral-github-runner
|
||||
base64 -d > /opt/ephemeral-github-runner/start-runner.sh <<'EOF'
|
||||
{start_runner_b64}
|
||||
EOF
|
||||
chmod 0755 /opt/ephemeral-github-runner/start-runner.sh
|
||||
base64 -d > /etc/systemd/system/ephemeral-github-runner.service <<'EOF'
|
||||
{unit_b64}
|
||||
EOF
|
||||
systemctl daemon-reload
|
||||
systemctl enable ephemeral-github-runner.service
|
||||
|
||||
# Keep the image free of the build-time docker state (also shrinks the
|
||||
# snapshot: buildkit cache is several GB).
|
||||
docker rm -f "${{container}}" >/dev/null
|
||||
docker system prune -af >/dev/null
|
||||
trap - EXIT
|
||||
|
||||
echo "{DONE_MARKER}" > /dev/console
|
||||
"""
|
||||
|
||||
|
||||
def make_ecs_client(region_id: str):
|
||||
from alibabacloud_ecs20140526.client import Client as EcsClient
|
||||
from alibabacloud_tea_openapi.models import Config as OpenApiConfig
|
||||
|
||||
return EcsClient(
|
||||
OpenApiConfig(
|
||||
access_key_id=os.environ["ALIBABA_CLOUD_ACCESS_KEY_ID"],
|
||||
access_key_secret=os.environ["ALIBABA_CLOUD_ACCESS_KEY_SECRET"],
|
||||
region_id=region_id,
|
||||
endpoint=f"ecs.{region_id}.aliyuncs.com",
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
def wait_for_instance_status(client, region_id: str, instance_id: str, wanted: str, deadline: float) -> None:
|
||||
from alibabacloud_ecs20140526 import models as ecs_models
|
||||
|
||||
import json
|
||||
|
||||
while time.monotonic() < deadline:
|
||||
response = call_api_with_retry(
|
||||
lambda: client.describe_instances(
|
||||
ecs_models.DescribeInstancesRequest(
|
||||
region_id=region_id, instance_ids=json.dumps([instance_id])
|
||||
)
|
||||
),
|
||||
"DescribeInstances",
|
||||
)
|
||||
instances = response.body.instances.instance
|
||||
if instances and instances[0].status == wanted:
|
||||
return
|
||||
time.sleep(POLL_INTERVAL_SECONDS)
|
||||
raise TimeoutError(f"Instance {instance_id} did not reach status {wanted} in time")
|
||||
|
||||
|
||||
# Aliyun error codes worth retrying: throttling and server-side faults. Client
|
||||
# errors (4xx: permissions, bad parameters) are configuration problems and must
|
||||
# fail fast instead of being retried.
|
||||
TRANSIENT_ERROR_CODES = {"Throttling", "Throttling.User", "InternalError", "ServiceUnavailable"}
|
||||
|
||||
|
||||
def is_transient(error: Exception) -> bool:
|
||||
# UnretryableException comes from the darabonba network layer: connection
|
||||
# reset, read timeout, DNS blip.
|
||||
if type(error).__name__ == "UnretryableException":
|
||||
return True
|
||||
code = getattr(error, "code", "") or ""
|
||||
status = getattr(error, "statusCode", None)
|
||||
return code in TRANSIENT_ERROR_CODES or (isinstance(status, int) and status >= 500)
|
||||
|
||||
|
||||
def call_api_with_retry(fn, description: str, attempts: int = 5):
|
||||
"""Retry transient API/network failures; the build is too long to die on a blip.
|
||||
|
||||
Only call this with idempotent or safely-repeatable requests (reads,
|
||||
StopInstance, CreateImage). Never with RunInstances: if the request
|
||||
succeeded but the response was lost, a retry double-creates instances.
|
||||
"""
|
||||
for attempt in range(1, attempts + 1):
|
||||
try:
|
||||
return fn()
|
||||
except Exception as error: # noqa: BLE001
|
||||
if attempt == attempts or not is_transient(error):
|
||||
raise
|
||||
print(f"{description} failed (attempt {attempt}/{attempts}): {error}", flush=True)
|
||||
time.sleep(POLL_INTERVAL_SECONDS)
|
||||
raise RuntimeError("unreachable: retry loop exited without returning")
|
||||
|
||||
|
||||
def read_console_output(client, region_id: str, instance_id: str) -> str:
|
||||
"""Fetch the instance's serial console output; no in-guest agent needed."""
|
||||
from alibabacloud_ecs20140526 import models as ecs_models
|
||||
|
||||
response = call_api_with_retry(
|
||||
lambda: client.get_instance_console_output(
|
||||
ecs_models.GetInstanceConsoleOutputRequest(region_id=region_id, instance_id=instance_id)
|
||||
),
|
||||
"GetInstanceConsoleOutput",
|
||||
)
|
||||
return base64.b64decode(response.body.console_output or "").decode("utf-8", "replace")
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--region-id", default=os.environ.get("ALIYUN_ECS_REGION_ID"))
|
||||
parser.add_argument("--vswitch-id", default=os.environ.get("ALIYUN_ECS_VSWITCH_ID"))
|
||||
parser.add_argument("--security-group-id", default=os.environ.get("ALIYUN_ECS_SECURITY_GROUP_ID"))
|
||||
parser.add_argument("--base-image-id", default=os.environ.get("ALIYUN_ECS_BASE_IMAGE_ID"))
|
||||
parser.add_argument("--resource-group-id", default=os.environ.get("ALIYUN_ECS_RESOURCE_GROUP_ID"))
|
||||
parser.add_argument("--instance-type", default="ecs.g7.xlarge")
|
||||
parser.add_argument("--image-name", default=None, help="Defaults to a timestamped name.")
|
||||
args = parser.parse_args()
|
||||
|
||||
for name in ("region_id", "vswitch_id", "security_group_id", "base_image_id"):
|
||||
if not getattr(args, name):
|
||||
raise SystemExit(f"Missing required configuration: --{name.replace('_', '-')}")
|
||||
|
||||
from alibabacloud_ecs20140526 import models as ecs_models
|
||||
|
||||
client = make_ecs_client(args.region_id)
|
||||
image_name = args.image_name or time.strftime(
|
||||
"greptimedb-query-regression-runner-%Y%m%d%H%M%S", time.gmtime()
|
||||
)
|
||||
|
||||
user_data = base64.b64encode(
|
||||
render_user_data(
|
||||
(ASSETS_DIR.parent / "Dockerfile").read_text(),
|
||||
(ASSETS_DIR / "start-runner.sh").read_text(),
|
||||
(ASSETS_DIR / "ephemeral-github-runner.service").read_text(),
|
||||
).encode()
|
||||
).decode()
|
||||
|
||||
instance_id = None
|
||||
try:
|
||||
response = client.run_instances(
|
||||
ecs_models.RunInstancesRequest(
|
||||
region_id=args.region_id,
|
||||
image_id=args.base_image_id,
|
||||
resource_group_id=args.resource_group_id,
|
||||
instance_type=args.instance_type,
|
||||
v_switch_id=args.vswitch_id,
|
||||
security_group_id=args.security_group_id,
|
||||
instance_name=f"build-{image_name}",
|
||||
description="Temporary builder for the query-regression ECS image",
|
||||
amount=1,
|
||||
instance_charge_type="PostPaid",
|
||||
spot_strategy="NoSpot",
|
||||
internet_charge_type="PayByTraffic",
|
||||
internet_max_bandwidth_out=100,
|
||||
# Peak usage is ~18G (OS+apt, docker image, and the materialized
|
||||
# toolchain coexist briefly): deliberately tight, and a smaller
|
||||
# disk makes the image snapshot faster. Instances created from
|
||||
# the image get a larger system disk from the provision side.
|
||||
system_disk=ecs_models.RunInstancesRequestSystemDisk(
|
||||
category="cloud_essd", size="20"
|
||||
),
|
||||
user_data=user_data,
|
||||
tag=[
|
||||
ecs_models.RunInstancesRequestTag(key="managed-by", value="query-regression-ci"),
|
||||
ecs_models.RunInstancesRequestTag(key="role", value="image-builder"),
|
||||
],
|
||||
)
|
||||
)
|
||||
instance_id = response.body.instance_id_sets.instance_id_set[0]
|
||||
print(f"Builder instance: {instance_id}", flush=True)
|
||||
wait_for_instance_status(client, args.region_id, instance_id, "Running", time.monotonic() + 10 * 60)
|
||||
|
||||
deadline = time.monotonic() + BUILD_TIMEOUT_SECONDS
|
||||
while time.monotonic() < deadline:
|
||||
console = read_console_output(client, args.region_id, instance_id)
|
||||
if DONE_MARKER in console:
|
||||
break
|
||||
if FAILED_MARKER in console:
|
||||
tail = "\n".join(console.splitlines()[-20:])
|
||||
raise RuntimeError(f"Image build failed on the builder; console tail:\n{tail}")
|
||||
lines = console.splitlines()
|
||||
print(f"Build in progress; last console line: {lines[-1] if lines else '(none yet)'}", flush=True)
|
||||
time.sleep(CONSOLE_POLL_INTERVAL_SECONDS)
|
||||
else:
|
||||
raise TimeoutError("Image build did not finish in time")
|
||||
|
||||
print("Stopping builder before image creation", flush=True)
|
||||
try:
|
||||
call_api_with_retry(
|
||||
lambda: client.stop_instance(
|
||||
ecs_models.StopInstanceRequest(
|
||||
instance_id=instance_id,
|
||||
# The instance is deleted right after the snapshot, but
|
||||
# there is no reason to keep billing vCPU during the stop.
|
||||
stopped_mode="StopCharging",
|
||||
)
|
||||
),
|
||||
"StopInstance",
|
||||
)
|
||||
except Exception as error: # noqa: BLE001
|
||||
# Instance families with local disks do not support StopCharging.
|
||||
print(f"StopCharging unavailable ({error}); stopping with default mode", flush=True)
|
||||
call_api_with_retry(
|
||||
lambda: client.stop_instance(ecs_models.StopInstanceRequest(instance_id=instance_id)),
|
||||
"StopInstance",
|
||||
)
|
||||
wait_for_instance_status(client, args.region_id, instance_id, "Stopped", time.monotonic() + 10 * 60)
|
||||
|
||||
image = call_api_with_retry(
|
||||
lambda: client.create_image(
|
||||
ecs_models.CreateImageRequest(
|
||||
region_id=args.region_id, instance_id=instance_id, image_name=image_name
|
||||
)
|
||||
),
|
||||
"CreateImage",
|
||||
)
|
||||
image_id = image.body.image_id
|
||||
deadline = time.monotonic() + 60 * 60
|
||||
while time.monotonic() < deadline:
|
||||
description = call_api_with_retry(
|
||||
lambda: client.describe_images(
|
||||
ecs_models.DescribeImagesRequest(region_id=args.region_id, image_id=image_id)
|
||||
),
|
||||
"DescribeImages",
|
||||
)
|
||||
images = description.body.images.image
|
||||
if images:
|
||||
status = images[0].status
|
||||
print(f"Image {image_id} status: {status} (progress {images[0].progress})", flush=True)
|
||||
if status == "Available":
|
||||
break
|
||||
time.sleep(POLL_INTERVAL_SECONDS)
|
||||
else:
|
||||
raise TimeoutError(
|
||||
f"Image {image_id} did not become Available in time. The snapshot "
|
||||
f"continues server-side: check its status in the console and reuse it "
|
||||
f"once Available instead of rebuilding."
|
||||
)
|
||||
|
||||
print(f"Custom image ready: {image_id} ({image_name})", flush=True)
|
||||
print(f"Set the repo variable QUERY_REGRESSION_ECS_IMAGE_ID={image_id}", flush=True)
|
||||
return 0
|
||||
finally:
|
||||
if instance_id:
|
||||
try:
|
||||
client.delete_instance(
|
||||
ecs_models.DeleteInstanceRequest(instance_id=instance_id, force=True)
|
||||
)
|
||||
print(f"Deleted builder instance {instance_id}", flush=True)
|
||||
except Exception as error: # noqa: BLE001
|
||||
print(f"Failed to delete builder instance {instance_id}: {error}", flush=True)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,45 @@
|
||||
# Copyright 2023 Greptime Team
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
[Unit]
|
||||
Description=Ephemeral GitHub Actions runner
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
# cloud-final runs the provision user data, which writes
|
||||
# /etc/ephemeral-github-runner.env and the drop-ins below. Ordering after it
|
||||
# makes the runner come up configured on first boot, no restart needed.
|
||||
# (cloud-init's `systemctl restart` in user data remains as the bridge for
|
||||
# images built before this ordering existed.)
|
||||
After=cloud-final.service
|
||||
|
||||
[Service]
|
||||
Type=simple
|
||||
# The toolchain lives in the image at /opt/cargo/bin (the Dockerfile's ENV
|
||||
# PATH does not survive materialization onto the host). Publish it here so
|
||||
# every runner job inherits it. /etc/ephemeral-github-runner.env currently
|
||||
# carries the same PATH line as a bridge for images built before this
|
||||
# directive existed; EnvironmentFile is applied after Environment=, and the
|
||||
# two values are kept identical, so there is no conflict.
|
||||
Environment=PATH=/opt/cargo/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
|
||||
EnvironmentFile=/etc/ephemeral-github-runner.env
|
||||
ExecStart=/opt/ephemeral-github-runner/start-runner.sh
|
||||
Restart=no
|
||||
# If the kernel OOM killer kills a job child (rustc/cargo), keep the runner
|
||||
# service alive so the job is reported as a normal failure with logs instead
|
||||
# of the runner vanishing and GitHub reporting a cancellation. cloud-init's
|
||||
# oom.conf drop-in carries the same directive for images built before this.
|
||||
OOMPolicy=continue
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
@@ -0,0 +1,37 @@
|
||||
#!/bin/bash
|
||||
# Copyright 2023 Greptime Team
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# Configure and start the ephemeral GitHub Actions runner on an ECS instance.
|
||||
# Baked into the custom image at /opt/ephemeral-github-runner/ by
|
||||
# build-ecs-image.py; the per-run environment is written by cloud-init to
|
||||
# /etc/ephemeral-github-runner.env.
|
||||
set -euo pipefail
|
||||
|
||||
# shellcheck disable=SC1091
|
||||
source /etc/ephemeral-github-runner.env
|
||||
|
||||
export HOME=/home/runner
|
||||
cd /home/runner
|
||||
|
||||
if [[ ! -f /home/runner/.runner ]]; then
|
||||
runuser -u runner -- ./config.sh \
|
||||
--url "${REPO_URL}" \
|
||||
--token "${RUNNER_TOKEN}" \
|
||||
--name "${RUNNER_NAME}" \
|
||||
--labels "${RUNNER_LABELS}" \
|
||||
--ephemeral --unattended --replace --disableupdate
|
||||
fi
|
||||
|
||||
exec runuser -u runner -- ./run.sh
|
||||
@@ -1,146 +0,0 @@
|
||||
githubConfigUrl: "https://github.com/GreptimeTeam/greptimedb"
|
||||
githubConfigSecret: greptimedb-arc-github-app
|
||||
|
||||
runnerScaleSetName: "perf-regression-8-cores"
|
||||
minRunners: 0
|
||||
maxRunners: 1
|
||||
|
||||
template:
|
||||
spec:
|
||||
automountServiceAccountToken: false
|
||||
activeDeadlineSeconds: 12600
|
||||
nodeSelector:
|
||||
kubernetes.io/hostname: minipc-3
|
||||
securityContext:
|
||||
runAsNonRoot: true
|
||||
runAsUser: 1001
|
||||
runAsGroup: 1001
|
||||
fsGroup: 1001
|
||||
fsGroupChangePolicy: OnRootMismatch
|
||||
seccompProfile:
|
||||
type: RuntimeDefault
|
||||
volumes:
|
||||
- name: cargo-home
|
||||
emptyDir: {}
|
||||
- name: build-cache
|
||||
persistentVolumeClaim:
|
||||
claimName: query-regression-build-cache
|
||||
initContainers:
|
||||
- name: initialize-build-cache
|
||||
image: greptime-registry.cn-hangzhou.cr.aliyuncs.com/greptime/greptimedb-query-regression-runner@sha256:e713b294e23b7e15184e558866c90025e59930033e72c97650dbc7f1ca022d11
|
||||
imagePullPolicy: IfNotPresent
|
||||
command:
|
||||
- /bin/sh
|
||||
- -ec
|
||||
- |
|
||||
umask 0002
|
||||
for directory in cargo-registry-v1 cargo-git-v1 query-regression-target-v1 meta-v1 sccache-v1; do
|
||||
cache_directory="/cache/${directory}"
|
||||
mkdir -p "${cache_directory}"
|
||||
test -w "${cache_directory}"
|
||||
probe_file="${cache_directory}/.write-probe"
|
||||
: > "${probe_file}"
|
||||
test -f "${probe_file}"
|
||||
rm "${probe_file}"
|
||||
done
|
||||
securityContext:
|
||||
runAsNonRoot: true
|
||||
runAsUser: 1001
|
||||
runAsGroup: 1001
|
||||
allowPrivilegeEscalation: false
|
||||
capabilities:
|
||||
drop:
|
||||
- ALL
|
||||
resources:
|
||||
requests:
|
||||
cpu: 100m
|
||||
memory: 128Mi
|
||||
ephemeral-storage: 1Gi
|
||||
limits:
|
||||
cpu: 500m
|
||||
memory: 512Mi
|
||||
ephemeral-storage: 1Gi
|
||||
volumeMounts:
|
||||
- name: build-cache
|
||||
mountPath: /cache
|
||||
containers:
|
||||
- name: runner
|
||||
image: greptime-registry.cn-hangzhou.cr.aliyuncs.com/greptime/greptimedb-query-regression-runner@sha256:e713b294e23b7e15184e558866c90025e59930033e72c97650dbc7f1ca022d11
|
||||
imagePullPolicy: IfNotPresent
|
||||
command: ["/home/runner/run.sh"]
|
||||
env:
|
||||
- name: CARGO_HOME
|
||||
value: /home/runner/.cargo
|
||||
- name: HTTP_PROXY
|
||||
valueFrom:
|
||||
configMapKeyRef:
|
||||
name: query-regression-runner-local-env
|
||||
key: HTTP_PROXY
|
||||
optional: true
|
||||
- name: HTTPS_PROXY
|
||||
valueFrom:
|
||||
configMapKeyRef:
|
||||
name: query-regression-runner-local-env
|
||||
key: HTTPS_PROXY
|
||||
optional: true
|
||||
- name: NO_PROXY
|
||||
valueFrom:
|
||||
configMapKeyRef:
|
||||
name: query-regression-runner-local-env
|
||||
key: NO_PROXY
|
||||
optional: true
|
||||
- name: UV_CACHE_DIR
|
||||
value: /home/runner/.cargo/uv-cache
|
||||
- name: RUSTUP_HOME
|
||||
value: /opt/rustup
|
||||
- name: RUSTUP_TOOLCHAIN
|
||||
value: nightly-2026-03-21
|
||||
- name: RUSTUP_AUTO_INSTALL
|
||||
value: "0"
|
||||
- name: CARGO_TARGET_DIR
|
||||
value: /home/runner/query-regression-target
|
||||
- name: QUERY_REGRESSION_CACHE_META
|
||||
value: /home/runner/query-regression-cache-meta
|
||||
- name: RUSTC_WRAPPER
|
||||
value: /usr/local/bin/sccache
|
||||
- name: SCCACHE_DIR
|
||||
value: /home/runner/.cache/sccache
|
||||
- name: SCCACHE_CACHE_SIZE
|
||||
value: 40G
|
||||
- name: CARGO_INCREMENTAL
|
||||
value: "0"
|
||||
securityContext:
|
||||
runAsNonRoot: true
|
||||
runAsUser: 1001
|
||||
runAsGroup: 1001
|
||||
allowPrivilegeEscalation: false
|
||||
capabilities:
|
||||
drop:
|
||||
- ALL
|
||||
resources:
|
||||
requests:
|
||||
cpu: "6"
|
||||
memory: 24Gi
|
||||
ephemeral-storage: 80Gi
|
||||
limits:
|
||||
cpu: "8"
|
||||
memory: 24Gi
|
||||
ephemeral-storage: 80Gi
|
||||
volumeMounts:
|
||||
- name: cargo-home
|
||||
mountPath: /home/runner/.cargo
|
||||
- name: build-cache
|
||||
mountPath: /home/runner/.cargo/registry
|
||||
subPath: cargo-registry-v1
|
||||
- name: build-cache
|
||||
mountPath: /home/runner/.cargo/git
|
||||
subPath: cargo-git-v1
|
||||
- name: build-cache
|
||||
mountPath: /home/runner/query-regression-target
|
||||
subPath: query-regression-target-v1
|
||||
- name: build-cache
|
||||
mountPath: /home/runner/query-regression-cache-meta
|
||||
subPath: meta-v1
|
||||
- name: build-cache
|
||||
mountPath: /home/runner/.cache/sccache
|
||||
subPath: sccache-v1
|
||||
@@ -1,2 +0,0 @@
|
||||
minRunners: 0
|
||||
maxRunners: 0
|
||||
@@ -0,0 +1,485 @@
|
||||
#!/usr/bin/env python3
|
||||
# Copyright 2023 Greptime Team
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# This is the first `.github/scripts/` user of PEP 723 inline script metadata:
|
||||
# `uv run` installs the pinned Aliyun SDK before executing. The SDK import stays
|
||||
# lazy (inside `make_ecs_client`) so unit tests can import this module with a
|
||||
# plain stdlib interpreter.
|
||||
# /// script
|
||||
# requires-python = ">=3.10"
|
||||
# dependencies = [
|
||||
# # Pin exactly after the first live run against the Aliyun API.
|
||||
# "alibabacloud_ecs20140526>=4.1.0,<6",
|
||||
# "alibabacloud_tea_openapi>=0.3.12,<1",
|
||||
# ]
|
||||
# ///
|
||||
|
||||
"""Provision an ephemeral Aliyun ECS query-regression runner.
|
||||
|
||||
Creates one pay-as-you-go ECS instance from the query-regression custom image
|
||||
and waits until the instance registers itself as an ephemeral GitHub Actions
|
||||
runner with a per-run label. Build caches live on the instance system disk
|
||||
and disappear with the VM. The instance id is written to job outputs as soon
|
||||
as the VM exists so the teardown job is the single DeleteInstance caller (a
|
||||
timeout here used to delete and then teardown deleted again). On runner
|
||||
online timeout the console is dumped and the job fails; teardown releases
|
||||
the instance.
|
||||
|
||||
The Aliyun credentials are used only in this control-plane job (ubuntu-latest)
|
||||
and are never passed to the ECS instance; the instance receives only a
|
||||
short-lived GitHub runner registration token via user data.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import base64
|
||||
import json
|
||||
import os
|
||||
import time
|
||||
import urllib.error
|
||||
import urllib.request
|
||||
from dataclasses import dataclass
|
||||
|
||||
RUNNER_NAME_PREFIX = "qreg-ecs"
|
||||
RUNNER_LABEL_PREFIX = "query-regression-ecs"
|
||||
MANAGED_BY_TAG_KEY = "managed-by"
|
||||
MANAGED_BY_TAG_VALUE = "query-regression-ci"
|
||||
RUN_TAG_KEY = "query-regression-run-id"
|
||||
|
||||
# Runner cache paths on the instance system disk. They are created empty
|
||||
# every provision and deleted with the VM.
|
||||
CACHE_PATHS = (
|
||||
"/home/runner/.cargo/registry",
|
||||
"/home/runner/.cargo/git",
|
||||
"/home/runner/query-regression-target",
|
||||
"/home/runner/query-regression-cache-meta",
|
||||
"/home/runner/.cache/sccache",
|
||||
)
|
||||
|
||||
RUNNER_ONLINE_TIMEOUT_SECONDS = 10 * 60
|
||||
POLL_INTERVAL_SECONDS = 5
|
||||
|
||||
# Nightly thin-LTO of greptime peaks above ecs.c9i.2xlarge's 16 GiB. A swap
|
||||
# file plus masking systemd-oomd lets the kernel reclaim rustc pages instead
|
||||
# of SIGTERM-ing the runner service cgroup (GitHub then reports the step as
|
||||
# "The operation was canceled"). Sized to match that 16 GiB instance; a
|
||||
# 32 GiB type still benefits as a safety net. Lives on the system disk next
|
||||
# to the cold build caches.
|
||||
SWAP_FILE = "/swapfile"
|
||||
SWAP_SIZE_GIB = 16
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ProvisionConfig:
|
||||
region_id: str
|
||||
vswitch_id: str
|
||||
security_group_id: str
|
||||
image_id: str
|
||||
instance_type: str
|
||||
repo: str
|
||||
run_id: str
|
||||
github_token: str
|
||||
# Optional resource group; required when the RAM grant is scoped to one.
|
||||
resource_group_id: str | None = None
|
||||
# Runner identity inside the image; the workflow asserts the same values.
|
||||
runner_uid: str = "1001"
|
||||
runner_gid: str = "1001"
|
||||
|
||||
|
||||
def runner_name_for_run(run_id: str) -> str:
|
||||
return f"{RUNNER_NAME_PREFIX}-{run_id}"
|
||||
|
||||
|
||||
def runner_label_for_run(run_id: str) -> str:
|
||||
return f"{RUNNER_LABEL_PREFIX}-{run_id}"
|
||||
|
||||
|
||||
def render_user_data(
|
||||
runner_name: str,
|
||||
runner_label: str,
|
||||
runner_token: str,
|
||||
repo: str,
|
||||
runner_uid: str = "1001",
|
||||
runner_gid: str = "1001",
|
||||
) -> str:
|
||||
"""Render the cloud-init shell script for the runner instance."""
|
||||
destinations = " ".join(f'"{dst}"' for dst in CACHE_PATHS)
|
||||
cache_setup = f"""# Caches live on the system disk, are deleted with the instance, and every
|
||||
# run compiles cold. Within-run reuse (base warming candidate via the shared
|
||||
# target dir and sccache) still applies.
|
||||
mkdir -p {destinations}
|
||||
chown -R {runner_uid}:{runner_gid} /home/runner"""
|
||||
swap_setup = f"""# Mask systemd-oomd before enabling swap: Ubuntu 24.04 kills the whole
|
||||
# service cgroup on PSI pressure, and swap thrashing looks like pressure.
|
||||
systemctl disable --now systemd-oomd.socket systemd-oomd.service || true
|
||||
systemctl mask systemd-oomd.socket systemd-oomd.service || true
|
||||
# apt-daily-upgrade / unattended-upgrades can `systemctl restart` services
|
||||
# whose libraries were updated. That SIGTERM-s the runner mid-job; GitHub
|
||||
# records UserCancelled (job still valid, child exit 143). These VMs live
|
||||
# ~1h; security updates belong in the image, not at job time.
|
||||
systemctl disable --now unattended-upgrades.service apt-daily.timer apt-daily-upgrade.timer apt-daily.service apt-daily-upgrade.service || true
|
||||
systemctl mask unattended-upgrades.service apt-daily.timer apt-daily-upgrade.timer apt-daily.service apt-daily-upgrade.service || true
|
||||
systemctl stop unattended-upgrades.service || true
|
||||
killall -9 unattended-upgr || true
|
||||
cat > /etc/apt/apt.conf.d/99disable-auto-updates <<'APTEOF'
|
||||
APT::Periodic::Update-Package-Lists "0";
|
||||
APT::Periodic::Unattended-Upgrade "0";
|
||||
APT::Periodic::Download-Upgradeable-Packages "0";
|
||||
APTEOF
|
||||
if [[ ! -f "{SWAP_FILE}" ]]; then
|
||||
fallocate --length {SWAP_SIZE_GIB}G "{SWAP_FILE}"
|
||||
chmod 600 "{SWAP_FILE}"
|
||||
mkswap "{SWAP_FILE}"
|
||||
fi
|
||||
swapon "{SWAP_FILE}"
|
||||
sysctl --write vm.swappiness=10
|
||||
swapon --show
|
||||
free --human"""
|
||||
return f"""#!/bin/bash
|
||||
set -euo pipefail
|
||||
|
||||
{cache_setup}
|
||||
|
||||
{swap_setup}
|
||||
|
||||
cat > /etc/ephemeral-github-runner.env <<'ENVEOF'
|
||||
RUNNER_NAME={runner_name}
|
||||
RUNNER_LABELS={runner_label}
|
||||
RUNNER_TOKEN={runner_token}
|
||||
REPO_URL=https://github.com/{repo}
|
||||
# The container image baked /opt/cargo/bin into PATH via Dockerfile ENV; on
|
||||
# the host nothing inherits that, so publish it through the unit's
|
||||
# EnvironmentFile: runner job processes inherit the runner's environment.
|
||||
PATH=/opt/cargo/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
|
||||
ENVEOF
|
||||
|
||||
# Stream the runner service's own logs (config output, job lifecycle, crash
|
||||
# messages) to the serial console, next to the kernel's OOM-killer records.
|
||||
# Teardown dumps the console tail before deleting the instance, so the
|
||||
# witness of what killed the runner survives the machine.
|
||||
mkdir -p /etc/systemd/system/ephemeral-github-runner.service.d
|
||||
cat > /etc/systemd/system/ephemeral-github-runner.service.d/console.conf <<'CONFEOF'
|
||||
[Service]
|
||||
StandardOutput=journal+console
|
||||
StandardError=journal+console
|
||||
CONFEOF
|
||||
# systemd's DefaultOOMPolicy=stop SIGTERM-s the whole unit when *any*
|
||||
# process in the cgroup is OOM-killed (typically rustc during thin LTO).
|
||||
# That is the "Session terminated, killing shell..." / job Canceled path:
|
||||
# runuser dies, GitHub deletes the ephemeral registration, always() steps
|
||||
# never run. continue keeps the runner up so cargo can report signal 9.
|
||||
cat > /etc/systemd/system/ephemeral-github-runner.service.d/oom.conf <<'CONFEOF'
|
||||
[Service]
|
||||
OOMPolicy=continue
|
||||
CONFEOF
|
||||
systemctl daemon-reload
|
||||
# restart (not start): the image enables this unit, so it may already
|
||||
# have come up without the drop-ins if it raced cloud-init. restart
|
||||
# applies OOMPolicy=continue to the running cgroup. --no-block matters:
|
||||
# newer image units carry After=cloud-final.service, and this script IS
|
||||
# cloud-final — a blocking restart would deadlock until TimeoutStartSec.
|
||||
systemctl restart --no-block ephemeral-github-runner.service
|
||||
"""
|
||||
|
||||
|
||||
def encode_user_data(script: str) -> str:
|
||||
return base64.b64encode(script.encode("utf-8")).decode("ascii")
|
||||
|
||||
|
||||
def github_api(token: str, method: str, path: str, body: dict | None = None) -> dict:
|
||||
data = json.dumps(body).encode("utf-8") if body is not None else None
|
||||
request = urllib.request.Request(
|
||||
f"https://api.github.com{path}",
|
||||
data=data,
|
||||
method=method,
|
||||
headers={
|
||||
"Accept": "application/vnd.github+json",
|
||||
"Authorization": f"Bearer {token}",
|
||||
"X-GitHub-Api-Version": "2022-11-28",
|
||||
},
|
||||
)
|
||||
try:
|
||||
with urllib.request.urlopen(request, timeout=30) as response:
|
||||
return json.loads(response.read().decode("utf-8"))
|
||||
except urllib.error.HTTPError as error:
|
||||
# GitHub's error body says exactly why (e.g. "Must have admin rights to
|
||||
# Repository" for a PAT without the required scope); surface it instead
|
||||
# of a bare "HTTP Error 403".
|
||||
body = error.read().decode("utf-8", "replace")
|
||||
raise SystemExit(
|
||||
f"GitHub API {method} {path} failed: HTTP {error.code}: {body}\n"
|
||||
"The token comes from the GH_PERSONAL_ACCESS_TOKEN secret; it needs "
|
||||
"'repo' scope (classic PAT) or 'Administration: write' on the "
|
||||
"repository (fine-grained PAT)."
|
||||
) from error
|
||||
|
||||
|
||||
def create_registration_token(github_token: str, repo: str) -> str:
|
||||
response = github_api(
|
||||
github_token,
|
||||
"POST",
|
||||
f"/repos/{repo}/actions/runners/registration-token",
|
||||
body={},
|
||||
)
|
||||
return response["token"]
|
||||
|
||||
|
||||
def find_runner_by_name(github_token: str, repo: str, name: str) -> dict | None:
|
||||
page = 1
|
||||
while True:
|
||||
response = github_api(
|
||||
github_token,
|
||||
"GET",
|
||||
f"/repos/{repo}/actions/runners?per_page=100&page={page}",
|
||||
)
|
||||
runners = response.get("runners", [])
|
||||
for runner in runners:
|
||||
if runner.get("name") == name:
|
||||
return runner
|
||||
if len(runners) < 100:
|
||||
return None
|
||||
page += 1
|
||||
|
||||
|
||||
def make_ecs_client(config: ProvisionConfig):
|
||||
from alibabacloud_ecs20140526.client import Client as EcsClient
|
||||
from alibabacloud_tea_openapi.models import Config as OpenApiConfig
|
||||
|
||||
access_key_id = os.environ.get("ALIBABA_CLOUD_ACCESS_KEY_ID", "")
|
||||
access_key_secret = os.environ.get("ALIBABA_CLOUD_ACCESS_KEY_SECRET", "")
|
||||
if not access_key_id or not access_key_secret:
|
||||
raise SystemExit(
|
||||
"ALIBABA_CLOUD_ACCESS_KEY_ID/SECRET are empty; in CI they come from the "
|
||||
"ALICLOUD_ECS_ACCESS_KEY_ID/SECRET repository secrets (a missing or "
|
||||
"misnamed secret expands to an empty string)."
|
||||
)
|
||||
return EcsClient(
|
||||
OpenApiConfig(
|
||||
access_key_id=access_key_id,
|
||||
access_key_secret=access_key_secret,
|
||||
region_id=config.region_id,
|
||||
endpoint=f"ecs.{config.region_id}.aliyuncs.com",
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
def wait_for_instance_status(client, region_id: str, instance_id: str, wanted: str, deadline: float) -> None:
|
||||
from alibabacloud_ecs20140526 import models as ecs_models
|
||||
|
||||
while time.monotonic() < deadline:
|
||||
response = client.describe_instances(
|
||||
ecs_models.DescribeInstancesRequest(
|
||||
region_id=region_id, instance_ids=json.dumps([instance_id])
|
||||
)
|
||||
)
|
||||
instances = response.body.instances.instance
|
||||
if instances and instances[0].status == wanted:
|
||||
return
|
||||
time.sleep(POLL_INTERVAL_SECONDS)
|
||||
raise TimeoutError(f"Instance {instance_id} did not reach status {wanted} in time")
|
||||
|
||||
|
||||
def fetch_console_output(client, region_id: str, instance_id: str) -> str:
|
||||
from alibabacloud_ecs20140526 import models as ecs_models
|
||||
|
||||
response = client.get_instance_console_output(
|
||||
ecs_models.GetInstanceConsoleOutputRequest(region_id=region_id, instance_id=instance_id)
|
||||
)
|
||||
return base64.b64decode(response.body.console_output or "").decode("utf-8", "replace")
|
||||
|
||||
|
||||
class ConsoleTailer:
|
||||
"""Incrementally prints new serial-console lines so cloud-init progress is
|
||||
visible in the CI log while waiting, not only after a failure dump."""
|
||||
|
||||
def __init__(self, client, region_id: str, instance_id: str) -> None:
|
||||
self.client = client
|
||||
self.region_id = region_id
|
||||
self.instance_id = instance_id
|
||||
self.printed_lines = 0
|
||||
|
||||
def poll(self) -> None:
|
||||
try:
|
||||
lines = fetch_console_output(self.client, self.region_id, self.instance_id).splitlines()
|
||||
except Exception as error: # noqa: BLE001
|
||||
print(f"[console] unable to fetch console output: {error}", flush=True)
|
||||
return
|
||||
# The API returns a bounded tail; if it ever shrinks, restart from the
|
||||
# beginning of the new buffer rather than skipping lines.
|
||||
if len(lines) < self.printed_lines:
|
||||
self.printed_lines = 0
|
||||
for line in lines[self.printed_lines :]:
|
||||
print(f"[console] {line}", flush=True)
|
||||
self.printed_lines = len(lines)
|
||||
|
||||
|
||||
def dump_console_output(client, region_id: str, instance_id: str) -> None:
|
||||
try:
|
||||
output = fetch_console_output(client, region_id, instance_id)
|
||||
except Exception as error: # noqa: BLE001
|
||||
print(f"Unable to fetch console output for {instance_id}: {error}", flush=True)
|
||||
return
|
||||
tail = "\n".join(output.splitlines()[-80:])
|
||||
print(f"::group::Console output tail for {instance_id}\n{tail}\n::endgroup::", flush=True)
|
||||
summary_path = os.environ.get("GITHUB_STEP_SUMMARY")
|
||||
if summary_path:
|
||||
with open(summary_path, "a", encoding="utf-8") as summary:
|
||||
summary.write(f"<details><summary>ECS console output ({instance_id})</summary>\n\n```\n")
|
||||
summary.write(tail)
|
||||
summary.write("\n```\n</details>\n")
|
||||
|
||||
|
||||
def append_github_output(name: str, value: str) -> None:
|
||||
output_path = os.environ.get("GITHUB_OUTPUT")
|
||||
if output_path:
|
||||
with open(output_path, "a", encoding="utf-8") as output:
|
||||
output.write(f"{name}={value}\n")
|
||||
|
||||
|
||||
def run_instance(client, config: ProvisionConfig, user_data: str) -> str:
|
||||
from alibabacloud_ecs20140526 import models as ecs_models
|
||||
|
||||
request = ecs_models.RunInstancesRequest(
|
||||
region_id=config.region_id,
|
||||
image_id=config.image_id,
|
||||
resource_group_id=config.resource_group_id,
|
||||
instance_type=config.instance_type,
|
||||
v_switch_id=config.vswitch_id,
|
||||
security_group_id=config.security_group_id,
|
||||
instance_name=runner_name_for_run(config.run_id),
|
||||
host_name=runner_name_for_run(config.run_id),
|
||||
description=f"Ephemeral query-regression runner for run {config.run_id}",
|
||||
amount=1,
|
||||
instance_charge_type="PostPaid",
|
||||
spot_strategy="NoSpot",
|
||||
internet_charge_type="PayByTraffic",
|
||||
internet_max_bandwidth_out=100,
|
||||
# Holds the image (~12G), the 16G swapfile, the source checkout, the
|
||||
# base+candidate cluster data homes, and cold build caches. Deleted
|
||||
# with the instance.
|
||||
system_disk=ecs_models.RunInstancesRequestSystemDisk(category="cloud_essd", size="40"),
|
||||
user_data=user_data,
|
||||
tag=[
|
||||
ecs_models.RunInstancesRequestTag(key=MANAGED_BY_TAG_KEY, value=MANAGED_BY_TAG_VALUE),
|
||||
ecs_models.RunInstancesRequestTag(key=RUN_TAG_KEY, value=config.run_id),
|
||||
],
|
||||
)
|
||||
response = client.run_instances(request)
|
||||
instance_id = response.body.instance_id_sets.instance_id_set[0]
|
||||
return instance_id
|
||||
|
||||
|
||||
def provision(config: ProvisionConfig) -> int:
|
||||
client = make_ecs_client(config)
|
||||
runner_name = runner_name_for_run(config.run_id)
|
||||
runner_label = runner_label_for_run(config.run_id)
|
||||
|
||||
registration_token = create_registration_token(config.github_token, config.repo)
|
||||
user_data = encode_user_data(
|
||||
render_user_data(
|
||||
runner_name,
|
||||
runner_label,
|
||||
registration_token,
|
||||
config.repo,
|
||||
config.runner_uid,
|
||||
config.runner_gid,
|
||||
)
|
||||
)
|
||||
|
||||
print(f"::group::Run instance {runner_name}", flush=True)
|
||||
instance_id = run_instance(client, config, user_data)
|
||||
print(f"Instance id: {instance_id}", flush=True)
|
||||
# Teardown is the only DeleteInstance caller. Write outputs immediately
|
||||
# so a later status/online failure still releases this VM.
|
||||
append_github_output("label", runner_label)
|
||||
append_github_output("instance_id", instance_id)
|
||||
append_github_output("runner_name", runner_name)
|
||||
wait_for_instance_status(client, config.region_id, instance_id, "Running", time.monotonic() + 5 * 60)
|
||||
print("::endgroup::", flush=True)
|
||||
|
||||
print("::group::Wait for runner online", flush=True)
|
||||
deadline = time.monotonic() + RUNNER_ONLINE_TIMEOUT_SECONDS
|
||||
tailer = ConsoleTailer(client, config.region_id, instance_id)
|
||||
last_console_poll = 0.0
|
||||
while time.monotonic() < deadline:
|
||||
runner = find_runner_by_name(config.github_token, config.repo, runner_name)
|
||||
if runner is not None:
|
||||
status = runner.get("status")
|
||||
print(f"Runner {runner_name} status: {status}", flush=True)
|
||||
if status == "online":
|
||||
print("::endgroup::", flush=True)
|
||||
print(f"Runner {runner_name} is online with label {runner_label}", flush=True)
|
||||
return 0
|
||||
# The serial console lags the guest by a minute or so; 30s polling is
|
||||
# enough to follow cloud-init without tripping API throttling.
|
||||
if time.monotonic() - last_console_poll >= 30:
|
||||
tailer.poll()
|
||||
last_console_poll = time.monotonic()
|
||||
time.sleep(POLL_INTERVAL_SECONDS)
|
||||
|
||||
print("::endgroup::", flush=True)
|
||||
print(
|
||||
f"Runner {runner_name} did not come online in time; "
|
||||
"leaving instance for the teardown job to delete",
|
||||
flush=True,
|
||||
)
|
||||
dump_console_output(client, config.region_id, instance_id)
|
||||
return 1
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--region-id", default=os.environ.get("ALIYUN_ECS_REGION_ID"))
|
||||
parser.add_argument("--vswitch-id", default=os.environ.get("ALIYUN_ECS_VSWITCH_ID"))
|
||||
parser.add_argument("--security-group-id", default=os.environ.get("ALIYUN_ECS_SECURITY_GROUP_ID"))
|
||||
parser.add_argument("--image-id", default=os.environ.get("QUERY_REGRESSION_ECS_IMAGE_ID"))
|
||||
parser.add_argument("--instance-type", default=os.environ.get("ALIYUN_ECS_INSTANCE_TYPE"))
|
||||
parser.add_argument("--repo", default=os.environ.get("GITHUB_REPOSITORY"))
|
||||
parser.add_argument("--run-id", default=os.environ.get("GITHUB_RUN_ID"))
|
||||
parser.add_argument("--github-token", default=os.environ.get("GH_PERSONAL_ACCESS_TOKEN"))
|
||||
parser.add_argument("--runner-uid", default=os.environ.get("QUERY_REGRESSION_RUNNER_UID", "1001"))
|
||||
parser.add_argument("--runner-gid", default=os.environ.get("QUERY_REGRESSION_RUNNER_GID", "1001"))
|
||||
parser.add_argument("--resource-group-id", default=os.environ.get("ALIYUN_ECS_RESOURCE_GROUP_ID"))
|
||||
args = parser.parse_args()
|
||||
|
||||
missing = [
|
||||
name
|
||||
for name, value in vars(args).items()
|
||||
if name not in ("resource_group_id",)
|
||||
and (value is None or (isinstance(value, str) and not value))
|
||||
]
|
||||
if missing:
|
||||
flags = ", ".join(f"--{name.replace('_', '-')}" for name in missing)
|
||||
raise SystemExit(f"Missing required configuration: {flags} (or their env defaults)")
|
||||
|
||||
config = ProvisionConfig(
|
||||
region_id=args.region_id,
|
||||
vswitch_id=args.vswitch_id,
|
||||
security_group_id=args.security_group_id,
|
||||
image_id=args.image_id,
|
||||
instance_type=args.instance_type,
|
||||
repo=args.repo,
|
||||
run_id=args.run_id,
|
||||
github_token=args.github_token,
|
||||
resource_group_id=args.resource_group_id or None,
|
||||
runner_uid=args.runner_uid,
|
||||
runner_gid=args.runner_gid,
|
||||
)
|
||||
return provision(config)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,237 @@
|
||||
#!/usr/bin/env python3
|
||||
# Copyright 2023 Greptime Team
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# PEP 723 inline metadata (see aliyun-ecs-runner-provision.py for the
|
||||
# convention). The SDK import stays lazy so unit tests run on a plain
|
||||
# stdlib interpreter.
|
||||
# /// script
|
||||
# requires-python = ">=3.10"
|
||||
# dependencies = [
|
||||
# "alibabacloud_ecs20140526>=4.1.0,<6",
|
||||
# "alibabacloud_tea_openapi>=0.3.12,<1",
|
||||
# ]
|
||||
# ///
|
||||
|
||||
"""Tear down ephemeral Aliyun ECS query-regression runners.
|
||||
|
||||
Two modes:
|
||||
|
||||
- Targeted (per workflow run): delete one instance by id and deregister its
|
||||
runner by name. Both steps are idempotent and best-effort; a missing
|
||||
instance or runner is not an error.
|
||||
- Sweep (scheduled janitor): delete every instance tagged as managed by
|
||||
query-regression CI whose creation time is older than the given TTL, and
|
||||
deregister the matching runners. This is the safety net for runs whose
|
||||
teardown job never executed.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import importlib.util
|
||||
import os
|
||||
import sys
|
||||
from datetime import datetime, timedelta, timezone
|
||||
from pathlib import Path
|
||||
|
||||
# Reuse the GitHub API client, runner lookup, and tag constants from the
|
||||
# provision script.
|
||||
_PROVISION_SPEC = importlib.util.spec_from_file_location(
|
||||
"aliyun_ecs_runner_provision",
|
||||
Path(__file__).resolve().parent / "aliyun-ecs-runner-provision.py",
|
||||
)
|
||||
assert _PROVISION_SPEC is not None and _PROVISION_SPEC.loader is not None
|
||||
provision = importlib.util.module_from_spec(_PROVISION_SPEC)
|
||||
sys.modules[_PROVISION_SPEC.name] = provision
|
||||
_PROVISION_SPEC.loader.exec_module(provision)
|
||||
|
||||
SWEEP_TTL = timedelta(hours=4)
|
||||
# ECS DescribeInstances creation_time is ISO8601 UTC, e.g. "2026-08-17T02:13Z".
|
||||
CREATION_TIME_FORMATS = ("%Y-%m-%dT%H:%M:%SZ", "%Y-%m-%dT%H:%MZ")
|
||||
|
||||
|
||||
def parse_creation_time(value: str) -> datetime:
|
||||
for fmt in CREATION_TIME_FORMATS:
|
||||
try:
|
||||
return datetime.strptime(value, fmt).replace(tzinfo=timezone.utc)
|
||||
except ValueError:
|
||||
continue
|
||||
raise ValueError(f"Unparsable ECS creation time: {value}")
|
||||
|
||||
|
||||
def expired_instance_names(
|
||||
instances: list[tuple[str, str, str]], now: datetime, ttl: timedelta
|
||||
) -> list[tuple[str, str]]:
|
||||
"""Pick (instance_id, instance_name) pairs whose creation is older than ttl.
|
||||
|
||||
`instances` items are (instance_id, instance_name, creation_time).
|
||||
"""
|
||||
expired = []
|
||||
for instance_id, instance_name, creation_time in instances:
|
||||
age = now - parse_creation_time(creation_time)
|
||||
if age >= ttl:
|
||||
expired.append((instance_id, instance_name))
|
||||
return expired
|
||||
|
||||
|
||||
def make_ecs_client(region_id: str):
|
||||
from alibabacloud_ecs20140526.client import Client as EcsClient
|
||||
from alibabacloud_tea_openapi.models import Config as OpenApiConfig
|
||||
|
||||
access_key_id = os.environ.get("ALIBABA_CLOUD_ACCESS_KEY_ID", "")
|
||||
access_key_secret = os.environ.get("ALIBABA_CLOUD_ACCESS_KEY_SECRET", "")
|
||||
if not access_key_id or not access_key_secret:
|
||||
raise SystemExit(
|
||||
"ALIBABA_CLOUD_ACCESS_KEY_ID/SECRET are empty; in CI they come from the "
|
||||
"ALICLOUD_ECS_ACCESS_KEY_ID/SECRET repository secrets (a missing or "
|
||||
"misnamed secret expands to an empty string)."
|
||||
)
|
||||
return EcsClient(
|
||||
OpenApiConfig(
|
||||
access_key_id=access_key_id,
|
||||
access_key_secret=access_key_secret,
|
||||
region_id=region_id,
|
||||
endpoint=f"ecs.{region_id}.aliyuncs.com",
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
def delete_instance(client, instance_id: str, region_id: str | None = None) -> bool:
|
||||
import time
|
||||
|
||||
from alibabacloud_ecs20140526 import models as ecs_models
|
||||
|
||||
# Export the serial console before destroying the evidence: a cancelled
|
||||
# run still runs this teardown job, and the console tail (cloud-init,
|
||||
# runner service logs, kernel OOM records) is the only witness of what
|
||||
# happened on the machine.
|
||||
if region_id is not None:
|
||||
provision.dump_console_output(client, region_id, instance_id)
|
||||
|
||||
# DeleteInstance rejects instances that are still Initializing (or in
|
||||
# another transitional status). Teardown runs right after a cancelled
|
||||
# run, when the instance may be only seconds old, so retry the transient
|
||||
# status errors for a few minutes instead of leaking the instance.
|
||||
deadline = time.monotonic() + 5 * 60
|
||||
while True:
|
||||
try:
|
||||
client.delete_instance(ecs_models.DeleteInstanceRequest(instance_id=instance_id, force=True))
|
||||
print(f"Deleted instance {instance_id}", flush=True)
|
||||
return True
|
||||
except Exception as error: # noqa: BLE001
|
||||
message = str(error)
|
||||
if "InvalidInstanceId.NotFound" in message:
|
||||
print(f"Instance {instance_id} already gone", flush=True)
|
||||
return True
|
||||
if "IncorrectInstanceStatus" in message and time.monotonic() < deadline:
|
||||
print(f"Instance {instance_id} is in a transitional status; retrying delete", flush=True)
|
||||
time.sleep(15)
|
||||
continue
|
||||
print(f"Failed to delete instance {instance_id}: {error}", flush=True)
|
||||
return False
|
||||
|
||||
|
||||
def deregister_runner(token: str, repo: str, runner_name: str) -> bool:
|
||||
runner = provision.find_runner_by_name(token, repo, runner_name)
|
||||
if runner is None:
|
||||
print(f"Runner {runner_name} is not registered", flush=True)
|
||||
return True
|
||||
try:
|
||||
provision.github_api(token, "DELETE", f"/repos/{repo}/actions/runners/{runner['id']}")
|
||||
print(f"Deregistered runner {runner_name} (id {runner['id']})", flush=True)
|
||||
return True
|
||||
except Exception as error: # noqa: BLE001
|
||||
print(f"Failed to deregister runner {runner_name}: {error}", flush=True)
|
||||
return False
|
||||
|
||||
|
||||
def list_managed_instances(client, region_id: str) -> list[tuple[str, str, str]]:
|
||||
from alibabacloud_ecs20140526 import models as ecs_models
|
||||
|
||||
result: list[tuple[str, str, str]] = []
|
||||
next_token = None
|
||||
while True:
|
||||
request = ecs_models.DescribeInstancesRequest(
|
||||
region_id=region_id,
|
||||
tag=[
|
||||
ecs_models.DescribeInstancesRequestTag(
|
||||
key=provision.MANAGED_BY_TAG_KEY, value=provision.MANAGED_BY_TAG_VALUE
|
||||
)
|
||||
],
|
||||
max_results=100,
|
||||
next_token=next_token,
|
||||
)
|
||||
response = client.describe_instances(request)
|
||||
for instance in response.body.instances.instance:
|
||||
result.append((instance.instance_id, instance.instance_name, instance.creation_time))
|
||||
next_token = response.body.next_token
|
||||
if not next_token:
|
||||
return result
|
||||
|
||||
|
||||
def sweep(client, region_id: str, repo: str, github_token: str, ttl: timedelta) -> int:
|
||||
instances = list_managed_instances(client, region_id)
|
||||
print(f"Found {len(instances)} managed instance(s) in {region_id}", flush=True)
|
||||
expired = expired_instance_names(instances, datetime.now(timezone.utc), ttl)
|
||||
ok = True
|
||||
for instance_id, instance_name in expired:
|
||||
print(f"Instance {instance_id} ({instance_name}) exceeds TTL {ttl}; deleting", flush=True)
|
||||
# Runner names mirror instance names by construction in the provision
|
||||
# script (both are qreg-ecs-<run_id>).
|
||||
ok &= delete_instance(client, instance_id, region_id)
|
||||
ok &= deregister_runner(github_token, repo, instance_name)
|
||||
return 0 if ok else 1
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--region-id", default=os.environ.get("ALIYUN_ECS_REGION_ID"))
|
||||
parser.add_argument("--repo", default=os.environ.get("GITHUB_REPOSITORY"))
|
||||
parser.add_argument("--github-token", default=os.environ.get("GH_PERSONAL_ACCESS_TOKEN"))
|
||||
parser.add_argument("--instance-id", default=os.environ.get("QUERY_REGRESSION_ECS_INSTANCE_ID"))
|
||||
parser.add_argument("--runner-name", default=os.environ.get("QUERY_REGRESSION_ECS_RUNNER_NAME"))
|
||||
parser.add_argument(
|
||||
"--sweep",
|
||||
action="store_true",
|
||||
help="Janitor mode: delete all managed instances older than --sweep-ttl-hours.",
|
||||
)
|
||||
parser.add_argument("--sweep-ttl-hours", type=float, default=SWEEP_TTL.total_seconds() / 3600)
|
||||
args = parser.parse_args()
|
||||
|
||||
for name in ("region_id", "repo", "github_token"):
|
||||
if not getattr(args, name):
|
||||
raise SystemExit(f"Missing required configuration: --{name.replace('_', '-')}")
|
||||
|
||||
client = make_ecs_client(args.region_id)
|
||||
|
||||
if args.sweep:
|
||||
return sweep(
|
||||
client, args.region_id, args.repo, args.github_token, timedelta(hours=args.sweep_ttl_hours)
|
||||
)
|
||||
|
||||
ok = True
|
||||
if args.instance_id:
|
||||
ok &= delete_instance(client, args.instance_id, args.region_id)
|
||||
else:
|
||||
print("No instance id given; skipping instance deletion", flush=True)
|
||||
if args.runner_name:
|
||||
ok &= deregister_runner(args.github_token, args.repo, args.runner_name)
|
||||
else:
|
||||
print("No runner name given; skipping runner deregistration", flush=True)
|
||||
return 0 if ok else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,270 @@
|
||||
#!/usr/bin/env python3
|
||||
# Copyright 2023 Greptime Team
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
"""Resolve consecutive successful Nightly Build SHAs for query-regression.
|
||||
|
||||
Query regression builds both binaries from git; this script only picks the
|
||||
head SHAs of two Nightly Build workflow runs (today vs the previous success).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import sys
|
||||
import urllib.error
|
||||
import urllib.parse
|
||||
import urllib.request
|
||||
from dataclasses import dataclass
|
||||
from typing import Any
|
||||
|
||||
|
||||
NIGHTLY_WORKFLOW = "nightly-build.yml"
|
||||
NIGHTLY_EVENTS = frozenset({"schedule", "workflow_dispatch"})
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class WorkflowRun:
|
||||
id: int
|
||||
head_sha: str
|
||||
head_branch: str
|
||||
html_url: str
|
||||
created_at: str
|
||||
conclusion: str
|
||||
event: str
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class RefPair:
|
||||
base: WorkflowRun | None
|
||||
candidate: WorkflowRun | None
|
||||
skip: bool
|
||||
reason: str
|
||||
|
||||
|
||||
def parse_run(payload: dict[str, Any]) -> WorkflowRun:
|
||||
return WorkflowRun(
|
||||
id=int(payload["id"]),
|
||||
head_sha=str(payload["head_sha"]),
|
||||
head_branch=str(payload.get("head_branch") or ""),
|
||||
html_url=str(payload.get("html_url") or ""),
|
||||
created_at=str(payload.get("created_at") or ""),
|
||||
conclusion=str(payload.get("conclusion") or ""),
|
||||
event=str(payload.get("event") or ""),
|
||||
)
|
||||
|
||||
|
||||
def github_get(token: str, path: str, query: dict[str, str] | None = None) -> dict[str, Any]:
|
||||
encoded = f"?{urllib.parse.urlencode(query)}" if query else ""
|
||||
request = urllib.request.Request(
|
||||
f"https://api.github.com{path}{encoded}",
|
||||
method="GET",
|
||||
headers={
|
||||
"Accept": "application/vnd.github+json",
|
||||
"Authorization": f"Bearer {token}",
|
||||
"X-GitHub-Api-Version": "2022-11-28",
|
||||
},
|
||||
)
|
||||
try:
|
||||
with urllib.request.urlopen(request, timeout=30) as response:
|
||||
return json.loads(response.read().decode("utf-8"))
|
||||
except urllib.error.HTTPError as error:
|
||||
body = error.read().decode("utf-8", "replace")
|
||||
raise SystemExit(
|
||||
f"GitHub API GET {path} failed: HTTP {error.code}: {body}"
|
||||
) from error
|
||||
|
||||
|
||||
def list_successful_nightly_runs(
|
||||
token: str,
|
||||
repo: str,
|
||||
*,
|
||||
branch: str | None = None,
|
||||
per_page: int = 30,
|
||||
) -> list[WorkflowRun]:
|
||||
query = {
|
||||
"status": "success",
|
||||
"per_page": str(per_page),
|
||||
}
|
||||
if branch:
|
||||
query["branch"] = branch
|
||||
payload = github_get(
|
||||
token,
|
||||
f"/repos/{repo}/actions/workflows/{NIGHTLY_WORKFLOW}/runs",
|
||||
query,
|
||||
)
|
||||
runs = [parse_run(item) for item in payload.get("workflow_runs") or []]
|
||||
return [
|
||||
run
|
||||
for run in runs
|
||||
if run.conclusion == "success" and run.event in NIGHTLY_EVENTS and run.head_sha
|
||||
]
|
||||
|
||||
|
||||
def fetch_run(token: str, repo: str, run_id: int) -> WorkflowRun:
|
||||
payload = github_get(token, f"/repos/{repo}/actions/runs/{run_id}")
|
||||
return parse_run(payload)
|
||||
|
||||
|
||||
def select_base_and_candidate(
|
||||
runs: list[WorkflowRun],
|
||||
*,
|
||||
candidate_run_id: int | None = None,
|
||||
candidate: WorkflowRun | None = None,
|
||||
) -> RefPair:
|
||||
"""Pick candidate (newest / requested) and the previous successful run.
|
||||
|
||||
`runs` must be newest-first. Same-SHA consecutive nightlies are skipped
|
||||
(no commits that day).
|
||||
"""
|
||||
if candidate is None:
|
||||
if candidate_run_id is not None:
|
||||
candidate = next((run for run in runs if run.id == candidate_run_id), None)
|
||||
if candidate is None:
|
||||
return RefPair(
|
||||
None,
|
||||
None,
|
||||
True,
|
||||
f"candidate nightly run {candidate_run_id} was not a successful "
|
||||
f"{NIGHTLY_WORKFLOW} run",
|
||||
)
|
||||
elif runs:
|
||||
candidate = runs[0]
|
||||
else:
|
||||
return RefPair(None, None, True, "no successful Nightly Build runs")
|
||||
|
||||
branch = candidate.head_branch
|
||||
older = [
|
||||
run
|
||||
for run in runs
|
||||
if run.id != candidate.id
|
||||
and (not branch or not run.head_branch or run.head_branch == branch)
|
||||
and (not candidate.created_at or run.created_at <= candidate.created_at)
|
||||
]
|
||||
if not older:
|
||||
return RefPair(
|
||||
None,
|
||||
candidate,
|
||||
True,
|
||||
"no previous successful Nightly Build to compare against",
|
||||
)
|
||||
base = older[0]
|
||||
if base.head_sha.lower() == candidate.head_sha.lower():
|
||||
return RefPair(
|
||||
base,
|
||||
candidate,
|
||||
True,
|
||||
f"previous nightly SHA {base.head_sha} matches candidate; nothing to compare",
|
||||
)
|
||||
return RefPair(base, candidate, False, "")
|
||||
|
||||
|
||||
def _override_run(sha: str) -> WorkflowRun:
|
||||
return WorkflowRun(
|
||||
id=0,
|
||||
head_sha=sha,
|
||||
head_branch="",
|
||||
html_url="",
|
||||
created_at="",
|
||||
conclusion="success",
|
||||
event="workflow_dispatch",
|
||||
)
|
||||
|
||||
|
||||
def override_pair(base_ref: str, candidate_ref: str) -> RefPair:
|
||||
return RefPair(_override_run(base_ref), _override_run(candidate_ref), False, "explicit refs")
|
||||
|
||||
|
||||
def write_outputs(pair: RefPair) -> None:
|
||||
values = {
|
||||
"skip": "true" if pair.skip else "false",
|
||||
"reason": pair.reason,
|
||||
"base_sha": pair.base.head_sha if pair.base else "",
|
||||
"candidate_sha": pair.candidate.head_sha if pair.candidate else "",
|
||||
"base_run_url": pair.base.html_url if pair.base else "",
|
||||
"candidate_run_url": pair.candidate.html_url if pair.candidate else "",
|
||||
"base_run_id": str(pair.base.id) if pair.base and pair.base.id else "",
|
||||
"candidate_run_id": str(pair.candidate.id) if pair.candidate and pair.candidate.id else "",
|
||||
}
|
||||
output_path = os.environ.get("GITHUB_OUTPUT")
|
||||
if output_path:
|
||||
with open(output_path, "a", encoding="utf-8") as handle:
|
||||
for key, value in values.items():
|
||||
handle.write(f"{key}={value}\n")
|
||||
for key, value in values.items():
|
||||
print(f"{key}={value}")
|
||||
|
||||
|
||||
def parse_args(argv: list[str] | None = None) -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--repo", default=os.environ.get("GITHUB_REPOSITORY"))
|
||||
parser.add_argument(
|
||||
"--token",
|
||||
default=os.environ.get("GITHUB_TOKEN") or os.environ.get("GH_TOKEN"),
|
||||
)
|
||||
parser.add_argument("--branch", default=os.environ.get("NIGHTLY_BRANCH") or "")
|
||||
parser.add_argument(
|
||||
"--candidate-run-id",
|
||||
default=os.environ.get("CANDIDATE_RUN_ID") or "",
|
||||
)
|
||||
parser.add_argument("--base-ref", default=os.environ.get("BASE_REF") or "")
|
||||
parser.add_argument("--candidate-ref", default=os.environ.get("CANDIDATE_REF") or "")
|
||||
return parser.parse_args(argv)
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
args = parse_args(argv)
|
||||
base_ref = args.base_ref.strip()
|
||||
candidate_ref = args.candidate_ref.strip()
|
||||
if base_ref and candidate_ref:
|
||||
write_outputs(override_pair(base_ref, candidate_ref))
|
||||
return 0
|
||||
if bool(base_ref) != bool(candidate_ref):
|
||||
print("Provide both --base-ref and --candidate-ref, or neither.", file=sys.stderr)
|
||||
return 2
|
||||
if not args.repo or not args.token:
|
||||
print("--repo and --token (or GITHUB_REPOSITORY / GITHUB_TOKEN) are required.", file=sys.stderr)
|
||||
return 2
|
||||
|
||||
candidate_run_id = int(args.candidate_run_id) if str(args.candidate_run_id).strip() else None
|
||||
candidate: WorkflowRun | None = None
|
||||
branch = args.branch.strip() or None
|
||||
if candidate_run_id is not None:
|
||||
candidate = fetch_run(args.token, args.repo, candidate_run_id)
|
||||
if candidate.conclusion != "success":
|
||||
write_outputs(
|
||||
RefPair(
|
||||
None,
|
||||
candidate,
|
||||
True,
|
||||
f"nightly run {candidate_run_id} conclusion is {candidate.conclusion}",
|
||||
)
|
||||
)
|
||||
return 0
|
||||
branch = branch or candidate.head_branch or None
|
||||
|
||||
runs = list_successful_nightly_runs(args.token, args.repo, branch=branch)
|
||||
pair = select_base_and_candidate(
|
||||
runs,
|
||||
candidate_run_id=candidate_run_id if candidate is None else None,
|
||||
candidate=candidate,
|
||||
)
|
||||
write_outputs(pair)
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -80,6 +80,14 @@ jobs:
|
||||
run: python3 scripts/check-enterprise-license-test.py
|
||||
- name: Run fuzz orchestration script tests
|
||||
run: .github/scripts/run-fuzz-targets-test.sh
|
||||
- name: Run query-regression Python tooling tests
|
||||
run: |
|
||||
python3 tests/perf/test_query_regression_runner_compaction_toctou.py
|
||||
python3 tests/perf/test_query_regression_runner_otlp_trace_load.py
|
||||
python3 tests/perf/test_query_regression_summary_otlp.py
|
||||
python3 tests/perf/test_query_regression_case_selection.py
|
||||
python3 tests/perf/test_query_regression_nightly_refs.py
|
||||
python3 tests/perf/test_aliyun_ecs_runner_scripts.py
|
||||
|
||||
check:
|
||||
if: ${{ github.repository == 'GreptimeTeam/greptimedb' }}
|
||||
|
||||
@@ -0,0 +1,36 @@
|
||||
name: Query Regression ECS Janitor
|
||||
|
||||
# Safety net for the aliyun-ecs query-regression path: deletes managed ECS
|
||||
# instances (and their runner registrations) that outlived their workflow run,
|
||||
# so a failed teardown never leaks pay-as-you-go spend.
|
||||
|
||||
on:
|
||||
schedule:
|
||||
- cron: "23 3 * * *"
|
||||
workflow_dispatch:
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
|
||||
jobs:
|
||||
sweep:
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 15
|
||||
steps:
|
||||
- name: Checkout teardown script
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
persist-credentials: false
|
||||
|
||||
- name: Install uv
|
||||
uses: astral-sh/setup-uv@v6
|
||||
|
||||
- name: Sweep expired ECS runners
|
||||
env:
|
||||
ALIBABA_CLOUD_ACCESS_KEY_ID: ${{ secrets.ALICLOUD_ECS_ACCESS_KEY_ID }}
|
||||
ALIBABA_CLOUD_ACCESS_KEY_SECRET: ${{ secrets.ALICLOUD_ECS_ACCESS_KEY_SECRET }}
|
||||
GH_PERSONAL_ACCESS_TOKEN: ${{ secrets.GH_PERSONAL_ACCESS_TOKEN }}
|
||||
ALIYUN_ECS_REGION_ID: ${{ vars.ALIYUN_ECS_REGION_ID }}
|
||||
run: >-
|
||||
uv run .github/scripts/aliyun-ecs-runner-teardown.py
|
||||
--sweep --sweep-ttl-hours 4
|
||||
@@ -0,0 +1,104 @@
|
||||
name: Query Regression Nightly
|
||||
|
||||
# After a successful GreptimeDB Nightly Build, compare that commit against
|
||||
# the previous successful nightly. The reusable Query Regression workflow
|
||||
# still compiles both SHAs; this wrapper only resolves which two SHAs.
|
||||
|
||||
on:
|
||||
workflow_run:
|
||||
workflows:
|
||||
- GreptimeDB Nightly Build
|
||||
types:
|
||||
- completed
|
||||
workflow_dispatch:
|
||||
inputs:
|
||||
base_ref:
|
||||
description: Base ref/SHA (empty = previous successful nightly)
|
||||
required: false
|
||||
type: string
|
||||
default: ""
|
||||
candidate_ref:
|
||||
description: Candidate ref/SHA (empty = latest successful nightly)
|
||||
required: false
|
||||
type: string
|
||||
default: ""
|
||||
candidate_run_id:
|
||||
description: Nightly Build run id to treat as candidate (empty = latest)
|
||||
required: false
|
||||
type: string
|
||||
default: ""
|
||||
case:
|
||||
description: Query perf case path(s), all, or heavy
|
||||
required: false
|
||||
type: string
|
||||
default: all
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
actions: read
|
||||
|
||||
jobs:
|
||||
resolve-refs:
|
||||
name: Resolve previous vs current nightly SHAs
|
||||
if: ${{ github.event_name == 'workflow_dispatch' || github.event.workflow_run.conclusion == 'success' }}
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 5
|
||||
outputs:
|
||||
skip: ${{ steps.refs.outputs.skip }}
|
||||
reason: ${{ steps.refs.outputs.reason }}
|
||||
base_sha: ${{ steps.refs.outputs.base_sha }}
|
||||
candidate_sha: ${{ steps.refs.outputs.candidate_sha }}
|
||||
steps:
|
||||
- name: Checkout ref resolver
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
persist-credentials: false
|
||||
|
||||
- name: Resolve nightly SHAs
|
||||
id: refs
|
||||
env:
|
||||
GITHUB_TOKEN: ${{ github.token }}
|
||||
GITHUB_REPOSITORY: ${{ github.repository }}
|
||||
CANDIDATE_RUN_ID: ${{ github.event_name == 'workflow_run' && github.event.workflow_run.id || inputs.candidate_run_id || '' }}
|
||||
BASE_REF: ${{ github.event_name == 'workflow_dispatch' && inputs.base_ref || '' }}
|
||||
CANDIDATE_REF: ${{ github.event_name == 'workflow_dispatch' && inputs.candidate_ref || '' }}
|
||||
run: python3 .github/scripts/query-regression-nightly-refs.py
|
||||
|
||||
- name: Summarize comparison
|
||||
if: always()
|
||||
env:
|
||||
SKIP: ${{ steps.refs.outputs.skip }}
|
||||
REASON: ${{ steps.refs.outputs.reason }}
|
||||
BASE_SHA: ${{ steps.refs.outputs.base_sha }}
|
||||
CANDIDATE_SHA: ${{ steps.refs.outputs.candidate_sha }}
|
||||
BASE_RUN_URL: ${{ steps.refs.outputs.base_run_url }}
|
||||
CANDIDATE_RUN_URL: ${{ steps.refs.outputs.candidate_run_url }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
{
|
||||
if [[ "${SKIP}" == "true" ]]; then
|
||||
printf 'Skipping query-regression nightly: %s\n' "${REASON}"
|
||||
else
|
||||
printf 'Comparing previous nightly `%s` -> current nightly `%s`\n' \
|
||||
"${BASE_SHA}" "${CANDIDATE_SHA}"
|
||||
if [[ -n "${BASE_RUN_URL}" ]]; then
|
||||
printf -- '- Previous Nightly Build: %s\n' "${BASE_RUN_URL}"
|
||||
fi
|
||||
if [[ -n "${CANDIDATE_RUN_URL}" ]]; then
|
||||
printf -- '- Current Nightly Build: %s\n' "${CANDIDATE_RUN_URL}"
|
||||
fi
|
||||
fi
|
||||
} | tee -a "${GITHUB_STEP_SUMMARY}"
|
||||
|
||||
query-regression:
|
||||
name: Query regression nightly
|
||||
needs: [resolve-refs]
|
||||
if: ${{ needs.resolve-refs.outputs.skip != 'true' && needs.resolve-refs.outputs.base_sha != '' && needs.resolve-refs.outputs.candidate_sha != '' }}
|
||||
uses: ./.github/workflows/query-regression.yml
|
||||
secrets: inherit
|
||||
with:
|
||||
case: ${{ github.event_name == 'workflow_dispatch' && inputs.case || 'all' }}
|
||||
base_ref: ${{ needs.resolve-refs.outputs.base_sha }}
|
||||
candidate_ref: ${{ needs.resolve-refs.outputs.candidate_sha }}
|
||||
cargo_profile: nightly
|
||||
runner: aliyun-ecs
|
||||
@@ -32,10 +32,10 @@ on:
|
||||
type: string
|
||||
default: nightly
|
||||
runner:
|
||||
description: Self-hosted runner label or ARC runner scale set for this query regression run
|
||||
description: Self-hosted runner label; aliyun-ecs provisions a fresh ECS instance per run
|
||||
required: false
|
||||
type: string
|
||||
default: perf-regression-8-cores
|
||||
default: aliyun-ecs
|
||||
workflow_dispatch:
|
||||
inputs:
|
||||
case:
|
||||
@@ -68,12 +68,20 @@ on:
|
||||
- release
|
||||
- dev
|
||||
runner:
|
||||
description: Self-hosted runner label or ARC runner scale set for this query regression run
|
||||
description: >-
|
||||
Self-hosted runner label; aliyun-ecs provisions a fresh ECS instance
|
||||
per run, any other value is used as a literal runner label
|
||||
required: true
|
||||
type: choice
|
||||
default: perf-regression-8-cores
|
||||
options:
|
||||
- perf-regression-8-cores
|
||||
type: string
|
||||
default: aliyun-ecs
|
||||
keep_instance:
|
||||
description: >-
|
||||
Debug: keep the ECS instance after the run (skip teardown) so its
|
||||
runner _diag logs, journal, and telemetry can be inspected; the
|
||||
janitor still sweeps it after the TTL
|
||||
required: false
|
||||
type: boolean
|
||||
default: false
|
||||
pull_request:
|
||||
types: [labeled]
|
||||
|
||||
@@ -81,19 +89,99 @@ permissions:
|
||||
contents: read
|
||||
|
||||
jobs:
|
||||
query-regression:
|
||||
test-tooling:
|
||||
# Stdlib unittests for the workflow Python (case selection, report
|
||||
# helpers, nightly SHA picking, and rendered ECS user-data). They do
|
||||
# not talk to Aliyun or the Actions runner process; running them on
|
||||
# ubuntu-latest fails fast before any ECS spend. Not gated on the
|
||||
# regression labels: those labels boot a VM, these tests should not.
|
||||
# Ordinary PRs also run the same tests from checks.yml, because this
|
||||
# workflow only starts on `labeled` (or dispatch / workflow_call).
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Checkout
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
persist-credentials: false
|
||||
|
||||
- name: Test query regression tooling
|
||||
run: |
|
||||
python3 tests/perf/test_query_regression_runner_compaction_toctou.py
|
||||
python3 tests/perf/test_query_regression_runner_otlp_trace_load.py
|
||||
python3 tests/perf/test_query_regression_summary_otlp.py
|
||||
python3 tests/perf/test_query_regression_case_selection.py
|
||||
python3 tests/perf/test_query_regression_nightly_refs.py
|
||||
python3 tests/perf/test_aliyun_ecs_runner_scripts.py
|
||||
|
||||
provision:
|
||||
# Runs when the aliyun-ecs path is selected: explicitly via the runner
|
||||
# input, or for PR labels by default (a QUERY_REGRESSION_PR_RUNNER
|
||||
# repository variable set to another value redirects PRs to that literal
|
||||
# runner label instead). Uses trusted scripts from the PR base (or the
|
||||
# dispatched ref), never from candidate code. Waits for test-tooling so
|
||||
# a broken user-data template does not still create a VM.
|
||||
needs: [test-tooling]
|
||||
if: >-
|
||||
${{ github.event_name != 'pull_request' ||
|
||||
(github.event_name == 'pull_request' &&
|
||||
!github.event.pull_request.draft &&
|
||||
(github.event.label.name == 'query-regression' ||
|
||||
github.event.label.name == 'heavy-regression')) }}
|
||||
runs-on: ${{ github.event_name != 'pull_request' && inputs.runner || 'perf-regression-8-cores' }}
|
||||
${{ !failure() && !cancelled() &&
|
||||
((github.event_name != 'pull_request' && inputs.runner == 'aliyun-ecs') ||
|
||||
(github.event_name == 'pull_request' &&
|
||||
!github.event.pull_request.draft &&
|
||||
(github.event.label.name == 'query-regression' ||
|
||||
github.event.label.name == 'heavy-regression') &&
|
||||
(vars.QUERY_REGRESSION_PR_RUNNER || 'aliyun-ecs') == 'aliyun-ecs')) }}
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 45
|
||||
outputs:
|
||||
label: ${{ steps.provision.outputs.label }}
|
||||
instance_id: ${{ steps.provision.outputs.instance_id }}
|
||||
runner_name: ${{ steps.provision.outputs.runner_name }}
|
||||
steps:
|
||||
- name: Checkout trusted provisioning scripts
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
repository: ${{ github.repository }}
|
||||
ref: ${{ github.event_name == 'pull_request' && github.event.pull_request.base.sha || github.sha }}
|
||||
persist-credentials: false
|
||||
|
||||
- name: Install uv
|
||||
uses: astral-sh/setup-uv@v6
|
||||
|
||||
- name: Provision ECS runner
|
||||
id: provision
|
||||
env:
|
||||
ALIBABA_CLOUD_ACCESS_KEY_ID: ${{ secrets.ALICLOUD_ECS_ACCESS_KEY_ID }}
|
||||
ALIBABA_CLOUD_ACCESS_KEY_SECRET: ${{ secrets.ALICLOUD_ECS_ACCESS_KEY_SECRET }}
|
||||
GH_PERSONAL_ACCESS_TOKEN: ${{ secrets.GH_PERSONAL_ACCESS_TOKEN }}
|
||||
ALIYUN_ECS_REGION_ID: ${{ vars.ALIYUN_ECS_REGION_ID }}
|
||||
ALIYUN_ECS_VSWITCH_ID: ${{ vars.ALIYUN_ECS_VSWITCH_ID }}
|
||||
ALIYUN_ECS_SECURITY_GROUP_ID: ${{ vars.ALIYUN_ECS_SECURITY_GROUP_ID }}
|
||||
ALIYUN_ECS_INSTANCE_TYPE: ${{ vars.ALIYUN_ECS_INSTANCE_TYPE }}
|
||||
ALIYUN_ECS_RESOURCE_GROUP_ID: ${{ vars.ALIYUN_ECS_RESOURCE_GROUP_ID }}
|
||||
QUERY_REGRESSION_ECS_IMAGE_ID: ${{ vars.QUERY_REGRESSION_ECS_IMAGE_ID }}
|
||||
QUERY_REGRESSION_RUNNER_UID: ${{ vars.QUERY_REGRESSION_RUNNER_UID || '1001' }}
|
||||
QUERY_REGRESSION_RUNNER_GID: ${{ vars.QUERY_REGRESSION_RUNNER_GID || '1001' }}
|
||||
run: >-
|
||||
uv run .github/scripts/aliyun-ecs-runner-provision.py
|
||||
|
||||
query-regression:
|
||||
needs: [provision, test-tooling]
|
||||
# `!failure() && !cancelled()` both suppresses the implicit success() and
|
||||
# lets the job run when provision was intentionally skipped because a
|
||||
# literal runner label was selected; a failed or cancelled provision still
|
||||
# blocks the run because no ECS runner would be waiting.
|
||||
if: >-
|
||||
${{ !failure() && !cancelled() &&
|
||||
(github.event_name != 'pull_request' ||
|
||||
(github.event_name == 'pull_request' &&
|
||||
!github.event.pull_request.draft &&
|
||||
(github.event.label.name == 'query-regression' ||
|
||||
github.event.label.name == 'heavy-regression'))) }}
|
||||
runs-on: >-
|
||||
${{ needs.provision.outputs.label ||
|
||||
(github.event_name != 'pull_request' && inputs.runner ||
|
||||
(vars.QUERY_REGRESSION_PR_RUNNER || 'aliyun-ecs')) }}
|
||||
timeout-minutes: 180
|
||||
concurrency:
|
||||
group: query-regression-persistent-cache-v1
|
||||
queue: max
|
||||
cancel-in-progress: false
|
||||
env:
|
||||
CARGO_PROFILE: ${{ github.event_name == 'pull_request' && 'nightly' || inputs.cargo_profile }}
|
||||
CARGO_HOME: /home/runner/.cargo
|
||||
@@ -105,7 +193,8 @@ jobs:
|
||||
QUERY_REGRESSION_CACHE_META: /home/runner/query-regression-cache-meta
|
||||
RUSTC_WRAPPER: /usr/local/bin/sccache
|
||||
SCCACHE_DIR: /home/runner/.cache/sccache
|
||||
SCCACHE_CACHE_SIZE: 40G
|
||||
# Caps the local sccache on the system disk next to target and cargo.
|
||||
SCCACHE_CACHE_SIZE: 10G
|
||||
CARGO_INCREMENTAL: "0"
|
||||
RUSTFLAGS: -D warnings -C link-arg=-fuse-ld=mold
|
||||
QUERY_REGRESSION_CACHE_EPOCH: "1"
|
||||
@@ -114,7 +203,55 @@ jobs:
|
||||
EVENT_MERGE_SHA: ${{ github.event_name == 'pull_request' && github.sha || '' }}
|
||||
EVENT_HEAD_SHA: ${{ github.event_name == 'pull_request' && github.event.pull_request.head.sha || '' }}
|
||||
EVENT_BASE_SHA: ${{ github.event_name == 'pull_request' && github.event.pull_request.base.sha || '' }}
|
||||
# Runner identity contract. The ECS image uses 1001; a manually prepared
|
||||
# host may override via repo variables when 1001 is already taken.
|
||||
EXPECTED_RUNNER_UID: ${{ vars.QUERY_REGRESSION_RUNNER_UID || '1001' }}
|
||||
EXPECTED_RUNNER_GID: ${{ vars.QUERY_REGRESSION_RUNNER_GID || '1001' }}
|
||||
steps:
|
||||
- name: Report provisioned ECS runner
|
||||
if: ${{ needs.provision.outputs.instance_id != '' }}
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
{
|
||||
printf -- '- ECS instance: `%s` (type `%s`, image `%s`, region `%s`)\n' \
|
||||
"${{ needs.provision.outputs.instance_id }}" \
|
||||
"${{ vars.ALIYUN_ECS_INSTANCE_TYPE }}" \
|
||||
"${{ vars.QUERY_REGRESSION_ECS_IMAGE_ID }}" \
|
||||
"${{ vars.ALIYUN_ECS_REGION_ID }}"
|
||||
} >> "${GITHUB_STEP_SUMMARY}"
|
||||
|
||||
- name: Start machine telemetry sampler
|
||||
shell: bash
|
||||
# The ECS instance is deleted when the run ends, so machine state must
|
||||
# be captured while the job runs. A background sampler appends load,
|
||||
# memory, disk, and top-memory processes to a log every 30s; the
|
||||
# "Dump machine telemetry" step (if: always()) prints and uploads it.
|
||||
# Note: when the job is *cancelled*, even always() steps are killed —
|
||||
# telemetry covers failures, not cancellations.
|
||||
run: |
|
||||
set -euo pipefail
|
||||
log="${GITHUB_WORKSPACE}/machine-telemetry.log"
|
||||
{
|
||||
echo "== baseline $(date -u +%Y-%m-%dT%H:%M:%SZ) =="
|
||||
nproc
|
||||
free -h
|
||||
df -h / /home/runner
|
||||
} >> "${log}"
|
||||
nohup bash -c "
|
||||
while true; do
|
||||
{
|
||||
date -u '+== %Y-%m-%dT%H:%M:%SZ =='
|
||||
uptime
|
||||
free -m
|
||||
df -h / /home/runner
|
||||
ps -eo pid,comm,%mem,%cpu --sort=-%mem | head -6
|
||||
} >> '${log}' 2>&1
|
||||
sleep 30
|
||||
done
|
||||
" >/dev/null 2>&1 &
|
||||
echo "Telemetry sampler started (pid $!), logging to ${log}"
|
||||
|
||||
- name: Checkout base source
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
@@ -263,14 +400,64 @@ jobs:
|
||||
shell: bash
|
||||
run: |
|
||||
set -euo pipefail
|
||||
[[ "$(id -u)" == "1001" ]]
|
||||
[[ "$(id -g)" == "1001" ]]
|
||||
[[ "${UV_CACHE_DIR}" == "/home/runner/.cargo/uv-cache" ]]
|
||||
[[ "$(protoc --version)" == "libprotoc 3.21.12" ]]
|
||||
errors=()
|
||||
record() {
|
||||
printf 'FAIL %s\n' "$*"
|
||||
errors+=("$*")
|
||||
}
|
||||
require_eq() {
|
||||
local name="$1" actual="$2" expected="$3"
|
||||
printf 'check %s: %s\n' "${name}" "${actual}"
|
||||
[[ "${actual}" == "${expected}" ]] || record "${name}: expected '${expected}', got '${actual}'"
|
||||
}
|
||||
require_match() {
|
||||
local name="$1" actual="$2" pattern="$3"
|
||||
printf 'check %s: %s\n' "${name}" "${actual}"
|
||||
[[ "${actual}" =~ $pattern ]] || record "${name}: expected to match ${pattern}, got '${actual}'"
|
||||
}
|
||||
require() {
|
||||
local name="$1"
|
||||
shift
|
||||
printf 'check %s\n' "${name}"
|
||||
"$@" || record "${name}"
|
||||
}
|
||||
capture() {
|
||||
local out
|
||||
if out="$("$@" 2>&1)"; then
|
||||
printf '%s' "${out}"
|
||||
else
|
||||
printf '<failed to run: %s>' "$*"
|
||||
fi
|
||||
}
|
||||
|
||||
printf 'uid=%s gid=%s PATH=%s CARGO_HOME=%s\n' "$(id -u)" "$(id -g)" "${PATH}" "${CARGO_HOME}"
|
||||
ls -la "${CARGO_HOME}" 2>&1 || printf '(CARGO_HOME missing)\n'
|
||||
|
||||
# Image/host hygiene: CARGO_HOME must be empty of config and
|
||||
# install-state before cargo/rustup run. Those tools create
|
||||
# .package-cache / bin on first use, so this cannot come after.
|
||||
cargo_home_dirty=()
|
||||
for entry in config config.toml credentials credentials.toml bin .crates.toml .crates2.json .global-cache .package-cache; do
|
||||
if [[ -e "${CARGO_HOME}/${entry}" ]]; then
|
||||
cargo_home_dirty+=("${entry}")
|
||||
record "CARGO_HOME must not contain ${entry}"
|
||||
fi
|
||||
done
|
||||
if (( ${#cargo_home_dirty[@]} > 0 )); then
|
||||
printf 'CARGO_HOME contents:\n' >&2
|
||||
ls -la "${CARGO_HOME}" >&2 || true
|
||||
fi
|
||||
|
||||
require_eq uid "$(id -u)" "${EXPECTED_RUNNER_UID}"
|
||||
require_eq gid "$(id -g)" "${EXPECTED_RUNNER_GID}"
|
||||
require_eq UV_CACHE_DIR "${UV_CACHE_DIR}" "/home/runner/.cargo/uv-cache"
|
||||
require_eq protoc "$(capture protoc --version)" "libprotoc 3.21.12"
|
||||
temporary_proto_dir="$(mktemp --directory)"
|
||||
trap 'rm -rf "${temporary_proto_dir}"' EXIT
|
||||
test -r /usr/include/google/protobuf/any.proto
|
||||
test -r /usr/include/google/protobuf/empty.proto
|
||||
require "readable /usr/include/google/protobuf/any.proto" \
|
||||
test -r /usr/include/google/protobuf/any.proto
|
||||
require "readable /usr/include/google/protobuf/empty.proto" \
|
||||
test -r /usr/include/google/protobuf/empty.proto
|
||||
printf '%s\n' \
|
||||
'syntax = "proto3";' \
|
||||
'package smoke;' \
|
||||
@@ -278,47 +465,51 @@ jobs:
|
||||
'import "google/protobuf/empty.proto";' \
|
||||
'message Smoke { google.protobuf.Any any = 1; google.protobuf.Empty empty = 2; }' \
|
||||
> "${temporary_proto_dir}/smoke.proto"
|
||||
protoc --proto_path="${temporary_proto_dir}" --proto_path=/usr/include \
|
||||
if protoc --proto_path="${temporary_proto_dir}" --proto_path=/usr/include \
|
||||
--descriptor_set_out="${temporary_proto_dir}/smoke.pb" \
|
||||
"${temporary_proto_dir}/smoke.proto"
|
||||
test -s "${temporary_proto_dir}/smoke.pb"
|
||||
[[ "$(uv --version)" =~ ^uv[[:space:]]0\.11\.26([[:space:]]|$) ]]
|
||||
mold_version="$(mold --version)"
|
||||
[[ "${mold_version}" =~ ^mold[[:space:]]2\.30\.0([[:space:]]|$) ]]
|
||||
[[ "$(python3 --version)" == "Python 3.12.3" ]]
|
||||
"${temporary_proto_dir}/smoke.proto"; then
|
||||
require "protoc smoke descriptor is non-empty" test -s "${temporary_proto_dir}/smoke.pb"
|
||||
else
|
||||
record "protoc smoke compile failed"
|
||||
fi
|
||||
require_match uv "$(capture uv --version)" '^uv[[:space:]]0\.11\.26([[:space:]]|$)'
|
||||
require_match mold "$(capture mold --version)" '^mold[[:space:]]2\.40\.4([[:space:]]|$)'
|
||||
require_eq python3 "$(capture python3 --version)" "Python 3.14.4"
|
||||
otelgen_path="$(command -v otelgen || true)"
|
||||
otelgen_version="<unavailable>"
|
||||
if [[ -n "${otelgen_path}" ]]; then
|
||||
otelgen_version="$("${otelgen_path}" --version 2>&1 || true)"
|
||||
otelgen_version="$(capture "${otelgen_path}" --version)"
|
||||
fi
|
||||
printf 'otelgen path: %s\n' "${otelgen_path:-<missing>}"
|
||||
printf 'otelgen version: %s\n' "${otelgen_version}"
|
||||
[[ "${otelgen_path}" == "/usr/local/bin/otelgen" ]]
|
||||
[[ "${otelgen_version}" == *"863a3f395d062c7322cc1de08a38774b7fdaa6c8"* ]]
|
||||
sccache_version="$(sccache --version)"
|
||||
[[ "${sccache_version}" =~ ^sccache[[:space:]]0\.16\.0([[:space:]]|$) ]]
|
||||
[[ "$(command -v rustup)" == "/opt/cargo/bin/rustup" ]]
|
||||
[[ "$(command -v cargo)" == "/opt/cargo/bin/cargo" ]]
|
||||
[[ "$(command -v rustc)" == "/opt/cargo/bin/rustc" ]]
|
||||
[[ "$(rustup --version)" =~ ^rustup[[:space:]]1\.29\.0([[:space:]]|$) ]]
|
||||
cargo_version="$(cargo --version)"
|
||||
[[ "${cargo_version}" =~ ^cargo[[:space:]]1\.96\.0-nightly[[:space:]]\(cbb9bb8bd[[:space:]][0-9]{4}-[0-9]{2}-[0-9]{2}\)$ ]]
|
||||
rustc_version="$(rustc --version)"
|
||||
[[ "${rustc_version}" =~ ^rustc[[:space:]]1\.96\.0-nightly[[:space:]]\(ac7f9ec7d[[:space:]][0-9]{4}-[0-9]{2}-[0-9]{2}\)$ ]]
|
||||
active_toolchain="$(rustup show active-toolchain)"
|
||||
[[ "${active_toolchain}" =~ ^nightly-2026-03-21-x86_64-unknown-linux-gnu([[:space:]]|$) ]]
|
||||
[[ "${RUSTUP_HOME}" == "/opt/rustup" ]]
|
||||
[[ "${RUSTUP_TOOLCHAIN}" == "nightly-2026-03-21" ]]
|
||||
[[ "${RUSTUP_AUTO_INSTALL}" == "0" ]]
|
||||
test -r /opt/rustup && test -x /opt/rustup
|
||||
test ! -w /opt/rustup
|
||||
test ! -w /opt/cargo/bin
|
||||
for entry in config config.toml credentials credentials.toml bin .crates.toml .crates2.json .global-cache .package-cache; do
|
||||
test ! -e "${CARGO_HOME}/${entry}"
|
||||
done
|
||||
require_eq otelgen_path "${otelgen_path}" "/usr/local/bin/otelgen"
|
||||
printf 'check otelgen_version: %s\n' "${otelgen_version}"
|
||||
[[ "${otelgen_version}" == *"863a3f395d062c7322cc1de08a38774b7fdaa6c8"* ]] \
|
||||
|| record "otelgen_version: expected commit 863a3f395d062c7322cc1de08a38774b7fdaa6c8, got '${otelgen_version}'"
|
||||
require_match sccache "$(capture sccache --version)" '^sccache[[:space:]]0\.16\.0([[:space:]]|$)'
|
||||
require_eq rustup_path "$(command -v rustup || true)" "/opt/cargo/bin/rustup"
|
||||
require_eq cargo_path "$(command -v cargo || true)" "/opt/cargo/bin/cargo"
|
||||
require_eq rustc_path "$(command -v rustc || true)" "/opt/cargo/bin/rustc"
|
||||
require_match rustup "$(capture rustup --version)" '^rustup[[:space:]]1\.29\.0([[:space:]]|$)'
|
||||
require_match cargo "$(capture cargo --version)" \
|
||||
'^cargo[[:space:]]1\.96\.0-nightly[[:space:]]\(cbb9bb8bd[[:space:]][0-9]{4}-[0-9]{2}-[0-9]{2}\)$'
|
||||
require_match rustc "$(capture rustc --version)" \
|
||||
'^rustc[[:space:]]1\.96\.0-nightly[[:space:]]\(ac7f9ec7d[[:space:]][0-9]{4}-[0-9]{2}-[0-9]{2}\)$'
|
||||
require_match active_toolchain "$(capture rustup show active-toolchain)" \
|
||||
'^nightly-2026-03-21-x86_64-unknown-linux-gnu([[:space:]]|$)'
|
||||
require_eq RUSTUP_HOME "${RUSTUP_HOME}" "/opt/rustup"
|
||||
require_eq RUSTUP_TOOLCHAIN "${RUSTUP_TOOLCHAIN}" "nightly-2026-03-21"
|
||||
require_eq RUSTUP_AUTO_INSTALL "${RUSTUP_AUTO_INSTALL}" "0"
|
||||
require "readable /opt/rustup" test -r /opt/rustup
|
||||
require "executable /opt/rustup" test -x /opt/rustup
|
||||
require "runner cannot write /opt/rustup" test ! -w /opt/rustup
|
||||
require "runner cannot write /opt/cargo/bin" test ! -w /opt/cargo/bin
|
||||
mkdir -p "${CARGO_HOME}/registry" "${CARGO_HOME}/git"
|
||||
if (( ${#errors[@]} > 0 )); then
|
||||
printf '\nVerify runner image tools failed (%d checks):\n' "${#errors[@]}" >&2
|
||||
printf ' - %s\n' "${errors[@]}" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
- name: Prepare persistent query regression cache
|
||||
- name: Prepare query regression cache
|
||||
shell: bash
|
||||
working-directory: src
|
||||
run: |
|
||||
@@ -362,24 +553,6 @@ jobs:
|
||||
find "${root}" -mindepth 1 -maxdepth 1 -exec rm -rf -- {} +
|
||||
}
|
||||
|
||||
clear_cargo_extracted_trees() {
|
||||
require_expected_root CARGO_REGISTRY "${CARGO_HOME}/registry" "${EXPECTED_CARGO_REGISTRY}"
|
||||
require_expected_root CARGO_GIT "${CARGO_HOME}/git" "${EXPECTED_CARGO_GIT}"
|
||||
rm -rf -- "${CARGO_HOME}/registry/src" "${CARGO_HOME}/git/checkouts"
|
||||
}
|
||||
|
||||
cargo_size_kib() {
|
||||
du -sk -- "${CARGO_HOME}/registry" "${CARGO_HOME}/git" | awk '{ total += $1 } END { print total }'
|
||||
}
|
||||
|
||||
target_size_kib() {
|
||||
du -sk -- "${CARGO_TARGET_DIR}" | cut -f1
|
||||
}
|
||||
|
||||
free_kib() {
|
||||
df -Pk "${CARGO_TARGET_DIR}" | awk 'NR == 2 { print $4 }'
|
||||
}
|
||||
|
||||
report_cache_usage() {
|
||||
du -sh -- "${CARGO_HOME}" "${CARGO_HOME}/registry" "${CARGO_HOME}/git" \
|
||||
"${RUSTUP_HOME}" "${CARGO_TARGET_DIR}" "${QUERY_REGRESSION_CACHE_META}" "${SCCACHE_DIR}"
|
||||
@@ -400,7 +573,7 @@ jobs:
|
||||
printf 'Refusing unexpected RUSTC_WRAPPER: %s\n' "${RUSTC_WRAPPER}" >&2
|
||||
exit 1
|
||||
}
|
||||
[[ "${SCCACHE_CACHE_SIZE}" == "40G" ]] || {
|
||||
[[ "${SCCACHE_CACHE_SIZE}" == "10G" ]] || {
|
||||
printf 'Refusing unexpected SCCACHE_CACHE_SIZE: %s\n' "${SCCACHE_CACHE_SIZE}" >&2
|
||||
exit 1
|
||||
}
|
||||
@@ -487,46 +660,10 @@ jobs:
|
||||
|
||||
report_cache_usage
|
||||
|
||||
target_size="$(target_size_kib)"
|
||||
if (( target_size >= 400 * 1024 * 1024 )); then
|
||||
printf 'Warning: target cache is at least 400 GiB (%s KiB)\n' "${target_size}" >&2
|
||||
fi
|
||||
if (( target_size >= 450 * 1024 * 1024 )); then
|
||||
printf 'Target cache is at least 450 GiB; clearing complete target root\n' >&2
|
||||
clear_directory "${CARGO_TARGET_DIR}"
|
||||
fi
|
||||
|
||||
cargo_size="$(cargo_size_kib)"
|
||||
if (( cargo_size >= 60 * 1024 * 1024 )); then
|
||||
printf 'Warning: Cargo cache is at least 60 GiB (%s KiB)\n' "${cargo_size}" >&2
|
||||
fi
|
||||
if (( cargo_size >= 80 * 1024 * 1024 )); then
|
||||
printf 'Cargo cache is at least 80 GiB; removing extracted sources and checkouts\n' >&2
|
||||
clear_cargo_extracted_trees
|
||||
cargo_size="$(cargo_size_kib)"
|
||||
if (( cargo_size >= 80 * 1024 * 1024 )); then
|
||||
printf 'Cargo cache remains at least 80 GiB after cleanup (%s KiB)\n' "${cargo_size}" >&2
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
|
||||
free_space="$(free_kib)"
|
||||
if (( free_space < 300 * 1024 * 1024 )); then
|
||||
printf 'Free space is below 300 GiB; clearing complete target root\n' >&2
|
||||
clear_directory "${CARGO_TARGET_DIR}"
|
||||
free_space="$(free_kib)"
|
||||
if (( free_space < 300 * 1024 * 1024 )); then
|
||||
printf 'Free space remains below 300 GiB; removing Cargo extracted sources and checkouts\n' >&2
|
||||
clear_cargo_extracted_trees
|
||||
free_space="$(free_kib)"
|
||||
if (( free_space < 300 * 1024 * 1024 )); then
|
||||
printf 'Free space remains below 300 GiB after cleanup (%s KiB)\n' "${free_space}" >&2
|
||||
exit 1
|
||||
fi
|
||||
fi
|
||||
fi
|
||||
|
||||
report_cache_usage
|
||||
# Every run starts from a fresh system disk, so size/free-space
|
||||
# watermarks from the retired retained-disk era are not restored.
|
||||
# A run that overflows the disk fails the build outright, which the
|
||||
# telemetry step makes diagnosable.
|
||||
|
||||
sccache --start-server
|
||||
sccache --zero-stats
|
||||
@@ -567,14 +704,6 @@ jobs:
|
||||
git reset --hard "${VERIFIED_CANDIDATE_SHA}"
|
||||
git clean -ffdx
|
||||
|
||||
- name: Test query regression tooling
|
||||
working-directory: src
|
||||
run: |
|
||||
uv run --no-project python tests/perf/test_query_regression_runner_compaction_toctou.py
|
||||
uv run --no-project python tests/perf/test_query_regression_runner_otlp_trace_load.py
|
||||
uv run --no-project python tests/perf/test_query_regression_summary_otlp.py
|
||||
uv run --no-project python tests/perf/test_query_regression_case_selection.py
|
||||
|
||||
- name: Build candidate greptime and query regression helpers
|
||||
working-directory: src
|
||||
run: |
|
||||
@@ -643,6 +772,7 @@ jobs:
|
||||
query-regression-work/**/logs/**
|
||||
query-regression-work/**/otelgen/**
|
||||
query-regression-summary.md
|
||||
machine-telemetry.log
|
||||
if-no-files-found: warn
|
||||
retention-days: 7
|
||||
|
||||
@@ -657,7 +787,7 @@ jobs:
|
||||
if-no-files-found: warn
|
||||
retention-days: 7
|
||||
|
||||
- name: Report persistent cache usage
|
||||
- name: Report cache usage
|
||||
if: ${{ always() }}
|
||||
shell: bash
|
||||
run: |
|
||||
@@ -697,6 +827,55 @@ jobs:
|
||||
printf 'sccache is unavailable (report only)\n' >&2
|
||||
fi
|
||||
|
||||
- name: Dump machine telemetry
|
||||
if: ${{ always() }}
|
||||
shell: bash
|
||||
run: |
|
||||
log="${GITHUB_WORKSPACE}/machine-telemetry.log"
|
||||
if [[ -f "${log}" ]]; then
|
||||
echo "::group::Machine telemetry (last 200 lines)"
|
||||
tail -n 200 "${log}"
|
||||
echo "::endgroup::"
|
||||
else
|
||||
echo "No telemetry log found (sampler never started?)"
|
||||
fi
|
||||
echo "::group::dmesg tail (OOM killer records)"
|
||||
sudo dmesg -T 2>/dev/null | tail -n 50 || dmesg -T 2>/dev/null | tail -n 50 || \
|
||||
echo "dmesg unavailable without root"
|
||||
echo "::endgroup::"
|
||||
|
||||
- name: Fail on regression failure
|
||||
if: ${{ steps.run.outputs.status != '0' }}
|
||||
run: exit 1
|
||||
|
||||
teardown:
|
||||
# Releases the dynamically provisioned ECS runner. Runs even when the
|
||||
# benchmark job fails or is cancelled; skipped when a literal runner label
|
||||
# was selected because the provision outputs are empty, or when the
|
||||
# dispatch set keep_instance to preserve the machine for post-mortem
|
||||
# debugging (the janitor sweep still reclaims it after the TTL).
|
||||
if: ${{ always() && needs.provision.outputs.instance_id != '' && !inputs.keep_instance }}
|
||||
needs: [provision, query-regression]
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 15
|
||||
steps:
|
||||
- name: Checkout trusted teardown scripts
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
repository: ${{ github.repository }}
|
||||
ref: ${{ github.event_name == 'pull_request' && github.event.pull_request.base.sha || github.sha }}
|
||||
persist-credentials: false
|
||||
|
||||
- name: Install uv
|
||||
uses: astral-sh/setup-uv@v6
|
||||
|
||||
- name: Teardown ECS runner
|
||||
env:
|
||||
ALIBABA_CLOUD_ACCESS_KEY_ID: ${{ secrets.ALICLOUD_ECS_ACCESS_KEY_ID }}
|
||||
ALIBABA_CLOUD_ACCESS_KEY_SECRET: ${{ secrets.ALICLOUD_ECS_ACCESS_KEY_SECRET }}
|
||||
GH_PERSONAL_ACCESS_TOKEN: ${{ secrets.GH_PERSONAL_ACCESS_TOKEN }}
|
||||
ALIYUN_ECS_REGION_ID: ${{ vars.ALIYUN_ECS_REGION_ID }}
|
||||
QUERY_REGRESSION_ECS_INSTANCE_ID: ${{ needs.provision.outputs.instance_id }}
|
||||
QUERY_REGRESSION_ECS_RUNNER_NAME: ${{ needs.provision.outputs.runner_name }}
|
||||
run: >-
|
||||
uv run .github/scripts/aliyun-ecs-runner-teardown.py
|
||||
|
||||
@@ -510,7 +510,7 @@ jobs:
|
||||
base_ref: ${{ needs.prepare-release-validation.outputs.previous-release-tag }}
|
||||
candidate_ref: ${{ needs.prepare-release-validation.outputs.candidate-ref }}
|
||||
cargo_profile: nightly
|
||||
runner: perf-regression-8-cores
|
||||
runner: aliyun-ecs
|
||||
|
||||
release-images-to-dockerhub:
|
||||
name: Build and push images to DockerHub
|
||||
|
||||
+18
-2
@@ -3,9 +3,23 @@
|
||||
- Keep GitHub Actions YAML thin. Put non-trivial control flow, case expansion,
|
||||
report generation, and metadata writing in scripts under `.github/scripts/`;
|
||||
workflow steps should mostly invoke those scripts.
|
||||
- Runner lifecycle: the default path provisions one ephemeral Aliyun ECS
|
||||
instance per run via `.github/scripts/aliyun-ecs-runner-provision.py` and
|
||||
always releases it via `aliyun-ecs-runner-teardown.py`; a scheduled janitor
|
||||
workflow sweeps leftovers. Build caches live on that instance's system disk
|
||||
and are discarded with the VM. Runs do not share a workflow concurrency
|
||||
group. The ECS custom image is built from the runner Dockerfile by
|
||||
`.github/runner-scale-sets/query-regression/ecs-image/build-ecs-image.py`;
|
||||
keep the Dockerfile the single source of the tool contract. Dispatching with
|
||||
any other `runner` value treats it as a literal self-hosted runner label
|
||||
(see `ecs-image/bootstrap-runner-host.sh` for preparing such a host).
|
||||
- Query regression PR runs should build base/candidate binaries once, then run
|
||||
the default case set. Do not hard-code a single case such as
|
||||
`promql_pushdown_7913` into the workflow path.
|
||||
- Scheduled nightly comparison lives in `query-regression-nightly.yml`: it
|
||||
waits for a successful Nightly Build, then calls `query-regression.yml`
|
||||
with the previous vs current nightly SHAs. Keep SHA selection in
|
||||
`.github/scripts/query-regression-nightly-refs.py`.
|
||||
- The case DSL is not required to keep compatibility inside this PR. When the
|
||||
DSL changes, update TOML cases, the outer lifecycle script, Rust helpers, and
|
||||
docs together.
|
||||
@@ -19,8 +33,10 @@
|
||||
- Keep the direct-SST generator generic. Issue-specific behavior belongs in case
|
||||
files and thresholds, not in Rust generator logic.
|
||||
- Before pushing perf harness changes, run at least:
|
||||
- the Python tests in the `Test query regression tooling` step of
|
||||
`.github/workflows/query-regression.yml`
|
||||
- the Python tests in the `test-tooling` job of
|
||||
`.github/workflows/query-regression.yml` (ubuntu-latest, not the ECS runner).
|
||||
The Checks workflow runs the same tests on ordinary PRs so they are not
|
||||
gated on the `query-regression` / `heavy-regression` labels.
|
||||
- `cargo fmt --all -- --check`
|
||||
- `cargo build -p cmd --bin query_perf_fixture --features dev-tools`
|
||||
- `cargo build -p cmd --bin query_regression_runner --features dev-tools`
|
||||
|
||||
@@ -255,11 +255,12 @@ gh workflow run query-regression.yml \
|
||||
-f candidate_ref=<full-candidate-sha> \
|
||||
-f cargo_profile=nightly \
|
||||
-f http_timeout=300 \
|
||||
-f runner=perf-regression-8-cores
|
||||
-f runner=aliyun-ecs
|
||||
```
|
||||
|
||||
The selected ARC scale set must already be deployed with the runner-image
|
||||
digest built from the current query-regression Dockerfile.
|
||||
The `aliyun-ecs` path provisions a fresh ECS instance per run from the custom
|
||||
image built from the current query-regression Dockerfile; see
|
||||
`.github/runner-scale-sets/query-regression/README.md` for its configuration.
|
||||
|
||||
## Generator contract
|
||||
|
||||
|
||||
@@ -0,0 +1,122 @@
|
||||
#!/usr/bin/env python3
|
||||
# Copyright 2023 Greptime Team
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
"""Coverage for the pure parts of the Aliyun ECS runner provision/teardown scripts."""
|
||||
|
||||
import base64
|
||||
import importlib.util
|
||||
import sys
|
||||
import unittest
|
||||
from datetime import datetime, timedelta, timezone
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
SCRIPTS_DIR = Path(__file__).parents[2] / ".github/scripts"
|
||||
|
||||
|
||||
def load_module(name: str, filename: str):
|
||||
spec = importlib.util.spec_from_file_location(name, SCRIPTS_DIR / filename)
|
||||
assert spec is not None and spec.loader is not None
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
sys.modules[spec.name] = module
|
||||
spec.loader.exec_module(module)
|
||||
return module
|
||||
|
||||
|
||||
provision = load_module("aliyun_ecs_runner_provision_under_test", "aliyun-ecs-runner-provision.py")
|
||||
teardown = load_module("aliyun_ecs_runner_teardown_under_test", "aliyun-ecs-runner-teardown.py")
|
||||
|
||||
|
||||
class ProvisionNamingTest(unittest.TestCase):
|
||||
def test_runner_name_and_label_derive_from_run_id(self) -> None:
|
||||
self.assertEqual(provision.runner_name_for_run("12345"), "qreg-ecs-12345")
|
||||
self.assertEqual(provision.runner_label_for_run("12345"), "query-regression-ecs-12345")
|
||||
|
||||
|
||||
class ProvisionUserDataTest(unittest.TestCase):
|
||||
def render(self) -> str:
|
||||
return provision.render_user_data(
|
||||
runner_name="qreg-ecs-12345",
|
||||
runner_label="query-regression-ecs-12345",
|
||||
runner_token="TOKEN",
|
||||
repo="GreptimeTeam/greptimedb",
|
||||
)
|
||||
|
||||
def test_user_data_creates_cache_paths_on_the_system_disk(self) -> None:
|
||||
script = self.render()
|
||||
self.assertNotIn("DISK_SERIAL", script)
|
||||
self.assertNotIn("mount --bind", script)
|
||||
self.assertNotIn("mkfs.ext4", script)
|
||||
for destination in provision.CACHE_PATHS:
|
||||
self.assertIn(f'"{destination}"', script)
|
||||
|
||||
def test_user_data_wires_runner_registration(self) -> None:
|
||||
script = self.render()
|
||||
self.assertIn("RUNNER_NAME=qreg-ecs-12345", script)
|
||||
self.assertIn("RUNNER_LABELS=query-regression-ecs-12345", script)
|
||||
self.assertIn("RUNNER_TOKEN=TOKEN", script)
|
||||
self.assertIn("REPO_URL=https://github.com/GreptimeTeam/greptimedb", script)
|
||||
self.assertIn("PATH=/opt/cargo/bin:", script)
|
||||
self.assertIn("systemctl restart --no-block ephemeral-github-runner.service", script)
|
||||
|
||||
def test_encode_user_data_round_trips(self) -> None:
|
||||
script = self.render()
|
||||
self.assertEqual(
|
||||
base64.b64decode(provision.encode_user_data(script)).decode("utf-8"), script
|
||||
)
|
||||
|
||||
def test_user_data_enables_swap_and_masks_oomd(self) -> None:
|
||||
script = self.render()
|
||||
self.assertIn("systemctl mask systemd-oomd.socket systemd-oomd.service", script)
|
||||
self.assertIn("systemctl mask unattended-upgrades.service apt-daily.timer apt-daily-upgrade.timer", script)
|
||||
self.assertIn('APT::Periodic::Unattended-Upgrade "0"', script)
|
||||
self.assertIn(f'fallocate --length {provision.SWAP_SIZE_GIB}G "{provision.SWAP_FILE}"', script)
|
||||
self.assertIn(f'swapon "{provision.SWAP_FILE}"', script)
|
||||
self.assertIn("sysctl --write vm.swappiness=10", script)
|
||||
self.assertIn("OOMPolicy=continue", script)
|
||||
self.assertNotIn("OOMScoreAdjust", script)
|
||||
self.assertLess(script.index("swapon"), script.index("systemctl restart --no-block ephemeral-github-runner.service"))
|
||||
|
||||
|
||||
class TeardownExpiryTest(unittest.TestCase):
|
||||
NOW = datetime(2026, 8, 17, 6, 0, tzinfo=timezone.utc)
|
||||
TTL = timedelta(hours=4)
|
||||
|
||||
def test_parse_creation_time_formats(self) -> None:
|
||||
self.assertEqual(
|
||||
teardown.parse_creation_time("2026-08-17T01:02:03Z"),
|
||||
datetime(2026, 8, 17, 1, 2, 3, tzinfo=timezone.utc),
|
||||
)
|
||||
self.assertEqual(
|
||||
teardown.parse_creation_time("2026-08-17T01:02Z"),
|
||||
datetime(2026, 8, 17, 1, 2, tzinfo=timezone.utc),
|
||||
)
|
||||
with self.assertRaises(ValueError):
|
||||
teardown.parse_creation_time("not-a-time")
|
||||
|
||||
def test_expired_instance_names_selects_only_old_instances(self) -> None:
|
||||
instances = [
|
||||
("i-old", "qreg-ecs-1", "2026-08-17T01:00Z"), # 5h old: expired
|
||||
("i-edge", "qreg-ecs-2", "2026-08-17T02:00Z"), # exactly TTL: expired
|
||||
("i-fresh", "qreg-ecs-3", "2026-08-17T05:30Z"), # 30m old: kept
|
||||
]
|
||||
self.assertEqual(
|
||||
teardown.expired_instance_names(instances, self.NOW, self.TTL),
|
||||
[("i-old", "qreg-ecs-1"), ("i-edge", "qreg-ecs-2")],
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,120 @@
|
||||
#!/usr/bin/env python3
|
||||
# Copyright 2023 Greptime Team
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
"""Unit tests for consecutive Nightly Build SHA selection."""
|
||||
|
||||
import importlib.util
|
||||
import sys
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
SCRIPTS_DIR = Path(__file__).parents[2] / ".github/scripts"
|
||||
|
||||
|
||||
def load_module():
|
||||
spec = importlib.util.spec_from_file_location(
|
||||
"query_regression_nightly_refs_under_test",
|
||||
SCRIPTS_DIR / "query-regression-nightly-refs.py",
|
||||
)
|
||||
assert spec is not None and spec.loader is not None
|
||||
module = importlib.util.module_from_spec(spec)
|
||||
sys.modules[spec.name] = module
|
||||
spec.loader.exec_module(module)
|
||||
return module
|
||||
|
||||
|
||||
refs = load_module()
|
||||
|
||||
|
||||
def run(*, run_id: int, sha: str, branch: str = "main", created_at: str = "2026-08-24T00:00:00Z"):
|
||||
return refs.WorkflowRun(
|
||||
id=run_id,
|
||||
head_sha=sha,
|
||||
head_branch=branch,
|
||||
html_url=f"https://github.com/example/run/{run_id}",
|
||||
created_at=created_at,
|
||||
conclusion="success",
|
||||
event="schedule",
|
||||
)
|
||||
|
||||
|
||||
class SelectNightlyRefsTest(unittest.TestCase):
|
||||
def test_picks_newest_and_previous_on_same_branch(self) -> None:
|
||||
pair = refs.select_base_and_candidate(
|
||||
[
|
||||
run(run_id=3, sha="ccc", created_at="2026-08-22T00:00:00Z"),
|
||||
run(run_id=2, sha="bbb", created_at="2026-08-21T00:00:00Z"),
|
||||
run(run_id=1, sha="aaa", created_at="2026-08-20T00:00:00Z"),
|
||||
]
|
||||
)
|
||||
self.assertFalse(pair.skip)
|
||||
assert pair.candidate is not None and pair.base is not None
|
||||
self.assertEqual(pair.candidate.head_sha, "ccc")
|
||||
self.assertEqual(pair.base.head_sha, "bbb")
|
||||
|
||||
def test_candidate_run_id_uses_that_run_and_the_next_older(self) -> None:
|
||||
pair = refs.select_base_and_candidate(
|
||||
[
|
||||
run(run_id=3, sha="ccc", created_at="2026-08-22T00:00:00Z"),
|
||||
run(run_id=2, sha="bbb", created_at="2026-08-21T00:00:00Z"),
|
||||
run(run_id=1, sha="aaa", created_at="2026-08-20T00:00:00Z"),
|
||||
],
|
||||
candidate_run_id=2,
|
||||
)
|
||||
self.assertFalse(pair.skip)
|
||||
assert pair.candidate is not None and pair.base is not None
|
||||
self.assertEqual(pair.candidate.id, 2)
|
||||
self.assertEqual(pair.base.id, 1)
|
||||
|
||||
def test_skips_when_consecutive_nightlies_share_a_sha(self) -> None:
|
||||
pair = refs.select_base_and_candidate(
|
||||
[
|
||||
run(run_id=2, sha="same", created_at="2026-08-22T00:00:00Z"),
|
||||
run(run_id=1, sha="same", created_at="2026-08-21T00:00:00Z"),
|
||||
]
|
||||
)
|
||||
self.assertTrue(pair.skip)
|
||||
self.assertIn("matches candidate", pair.reason)
|
||||
|
||||
def test_skips_other_branches_when_picking_previous(self) -> None:
|
||||
pair = refs.select_base_and_candidate(
|
||||
[
|
||||
run(run_id=3, sha="ccc", branch="main", created_at="2026-08-22T00:00:00Z"),
|
||||
run(run_id=2, sha="other", branch="feat", created_at="2026-08-21T12:00:00Z"),
|
||||
run(run_id=1, sha="aaa", branch="main", created_at="2026-08-21T00:00:00Z"),
|
||||
]
|
||||
)
|
||||
self.assertFalse(pair.skip)
|
||||
assert pair.base is not None
|
||||
self.assertEqual(pair.base.head_sha, "aaa")
|
||||
|
||||
def test_skips_when_only_one_nightly_exists(self) -> None:
|
||||
pair = refs.select_base_and_candidate(
|
||||
[run(run_id=1, sha="aaa")],
|
||||
)
|
||||
self.assertTrue(pair.skip)
|
||||
self.assertIn("no previous", pair.reason)
|
||||
|
||||
def test_explicit_refs_bypass_github(self) -> None:
|
||||
pair = refs.override_pair("base-sha", "cand-sha")
|
||||
self.assertFalse(pair.skip)
|
||||
assert pair.base is not None and pair.candidate is not None
|
||||
self.assertEqual(pair.base.head_sha, "base-sha")
|
||||
self.assertEqual(pair.candidate.head_sha, "cand-sha")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
Reference in New Issue
Block a user