mirror of
https://github.com/GreptimeTeam/greptimedb.git
synced 2026-10-03 02:25:35 +00:00
75bd8e9ce6fa554c2da3fa7d72dc63c8b50e99ce
13
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1d365adb3d |
fix(ci): update the shared Actions runner to v2.337.0 (#9376)
Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
43eaea7a9a |
fix(ci): teach check-builder-rust-version.sh to handle stable channels (#9369)
* fix(ci): teach check-builder-rust-version.sh to handle stable channels
The script extracted a YYYY-MM-DD date from rust-toolchain.toml to
compare against the rustc build date inside the dev-builder image —
a nightly-era design. With channel = "1.96.1" there is no date in
the file, so every release build failed with 'Error: No rust toolchain
version found in rust-toolchain.toml'.
Extract the channel token instead and branch on it:
- stable channel (X.Y[.Z]): require the builder image's rustc to
exactly match the pinned version
- nightly-YYYY-MM-DD: keep the legacy date-difference check
Verified against a mocked docker for all four paths (stable match /
mismatch, nightly fresh / stale).
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* chore(toolchain): finish stable-migration cleanup in docs and query-regression pin
- README/AGENTS: the toolchain is now stable Rust pinned by
rust-toolchain.toml, not nightly
- query-regression: align the benchmark toolchain pin with the
workspace (nightly-2026-03-21 = 1.96.0-nightly -> stable 1.96.1),
including the exact-version assertions (cargo 356927216, rustc
31fca3adb, both 2026-06-26) and the runner image default
The query-regression runner image must be rebuilt and
QUERY_REGRESSION_ECS_IMAGE_ID bumped together with these pins
(per .github/runner-scale-sets/query-regression/README.md) before
the next regression run.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci(query-regression): derive the Rust toolchain pin from rust-toolchain.toml
Replace the hard-coded RUSTUP_TOOLCHAIN value and the hard-coded
version strings in the runner Verify assertions with a pin resolved
from rust-toolchain.toml:
- the always-running test-tooling job exports the channel parsed from
rust-toolchain.toml as a job output
- query-regression sets RUSTUP_TOOLCHAIN from that output
- the Verify step escapes the pin into the cargo/rustc/active-toolchain
regexes at runtime; the exact commit hash and date are asserted
generically since a stable version identifies the release
Removing the redundant require_eq (workflow yaml vs runner env) since
both now flow from the single source of truth. When rust-toolchain.toml
is bumped, the run fails with a clear signal until the runner image is
rebuilt with the new toolchain, keeping the existing lockstep contract.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci(query-regression): derive the runner image toolchain from rust-toolchain.toml
Remove the hard-coded 'ARG RUST_TOOLCHAIN=1.96.1' from the
query-regression runner Dockerfile. The pin is now parsed from a
COPY'd rust-toolchain.toml at build time (the bootstrap script builds
with the repo root as context, so the file is in the build context):
- rustup-init installs the parsed channel as the default toolchain
- the baked ENV RUSTUP_TOOLCHAIN is dropped: the rustup default makes
bare cargo/rustc resolve correctly without it, and the workflow
supplies RUSTUP_TOOLCHAIN explicitly at run time
- the build-time self-verification asserts the active toolchain
against the same parsed pin
With this, rust-toolchain.toml is the single source of truth for the
benchmark toolchain end to end: the image bakes whatever the toml says
at build time and the workflow asserts against the toml at run time.
A toolchain bump now only requires rebuilding the image.
The changed mechanics were verified natively with the real rustup-init
1.29.0 and the real 1.96.1 toolchain (registry pulls are unavailable
in this sandbox): parsing, default-toolchain installation without
RUSTUP_TOOLCHAIN env, bare cargo/rustc resolution, and the
active-toolchain assertion for both '(default)' and '(overridden)'
output forms.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci(query-regression): rebuild the runner image automatically on toolchain changes
Mirror the dev-builder automation for the query-regression ECS runner
image: a new rebuild-query-regression-runner-image.yaml workflow runs
whenever rust-toolchain.toml or the query-regression runner directory
changes on main (or via manual dispatch). It drives the existing
build-ecs-image.py ops tool, then completes the documented lockstep
updates in order: bump RUNNER_IMAGE_EPOCH in query-regression.yml and
push the commit to main, and only then point the
QUERY_REGRESSION_ECS_IMAGE_ID repo variable at the new image, so the
next regression run picks up image and epoch together.
Also fix build-ecs-image.py to stage rust-toolchain.toml into the
temporary docker build context: the AMI path builds the embedded
Dockerfile from an empty /tmp/image-context, which would break on the
Dockerfile's COPY of rust-toolchain.toml introduced earlier. The
user-data now base64-stages the toml next to the Dockerfile before
docker build.
Verified: py_compile, render_user_data round-trip (mkdir -> stage ->
docker build ordering), and the RUNNER_IMAGE_EPOCH bump sed against
the real workflow file. Requires a new ALIYUN_ECS_BASE_IMAGE_ID repo
variable (Ubuntu 24.04 public image id in the region); all other
secrets/vars are shared with the provisioning job.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci: fold the query-regression runner rebuild into release-dev-builder-images.yaml
Merge the standalone rebuild workflow into the existing builder-image
release workflow, as one entry point for all builder artifacts:
- push paths extended with .github/runner-scale-sets/query-regression/**
- new 'release_query_regression_runner_image' dispatch input
- a 'changes' job diffs the pushed range (github.event.before..sha,
with an everything-changed fallback for dispatch or unknown bases)
so each expensive rebuild only fires for its own paths:
rust-toolchain.toml gates both, docker/dev-builder/** gates the
dev-builder images, the query-regression runner directory gates the
ECS image rebuild
- the rebuild job itself is unchanged from the standalone workflow
(build-ecs-image.py, then RUNNER_IMAGE_EPOCH commit to main, then
the QUERY_REGRESSION_ECS_IMAGE_ID variable update)
The dev-builder jobs, their ECR/CN/tag-update dependents, and the
runner rebuild now share one workflow; the changes filter preserves
the previous on-push behavior for the dev-builder images while the
runner rebuild keeps its own trigger. The epoch-bump commit only
touches query-regression.yml, which is outside the trigger paths, so
no re-trigger loop.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix(ci): auto-resolve the ECS base image for the runner rebuild
The automated rebuild failed with 'Missing required configuration:
--base-image-id' because the ALIYUN_ECS_BASE_IMAGE_ID repo variable
does not exist yet (it was flagged as a one-time setup item).
Remove the setup dependency instead: build-ecs-image.py now defaults
--base-image-id to the latest public Ubuntu 24.04 x86_64 system image
in the region (DescribeImages with image_owner_alias=system), so no
manual variable is required. The runner Dockerfile pins every tool
version itself, so base-image drift is low-risk; --base-image-id or
the ALIYUN_ECS_BASE_IMAGE_ID variable still pin a specific base image
deterministically, and the workflow only passes the flag when the
variable is set.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* docs(query-regression): clarify what an image rebuild requires
A routine runner-image rebuild needs no manual file updates: the
rebuild job updates QUERY_REGRESSION_ECS_IMAGE_ID and
RUNNER_IMAGE_EPOCH; the toolchain derives from rust-toolchain.toml;
uv, sccache, otelgen, rustup, and the runner base are pinned by
digest/sha/commit in the Dockerfile. Only an apt package revision
bump (mold, protoc, python3) between rebuilds requires bumping the
corresponding Verify pins, and that failure is loud with the observed
version.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix(ci): correct SDK field names in the base-image resolver
DescribeImagesRequest takes 'ostype' (not 'os_type') and the image
items expose 'osname'/'osname_en' (not 'os_name') in the pinned
alibabacloud_ecs20140526 SDK range, so the auto-resolution added in
|
||
|
|
ba0f7acd93 |
feat(mito2): add opt-in byte-stream-split encoding for float SST fields (#9069)
* feat(mito2): add opt-in byte stream split encoding Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * fix(mito2): correct float encoding checks Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * test(compat): cover float SST encoding Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * fix(mito2): compile float encoding tests Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * fix(mito2): release parquet test writer Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * fix(mito2): register float test primary key Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * test(mito2): verify BSS write lifecycles Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * test(metric-engine): verify BSS physical SST Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * test(mito2): verify bulk BSS lifecycle Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * fix(mito2): compile bulk BSS test Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * refactor(mito2): narrow bulk encoding constructors Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * test(compat): accept generated float upgrade output Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * test(compat): accept generated float downgrade output Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * refactor(mito2): narrow bulk encoding builder Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * test(perf): add default versus BSS storage comparison Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * test(perf): align BSS reader benchmarks with prior study Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * fix(perf): parse current read benchmark averages Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * fix(perf): retain default float encoding in direct SST fixtures Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * test(perf): isolate BSS user SSTs and benchmark every file Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * docs(perf): record measured BSS storage and reader tradeoffs Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * docs(perf): expose warm scan variability and evidence limits Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * docs(perf): clarify BSS baseline and storage measurement scope Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * test(perf): model bounded mixed integer and fractional metric series Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * docs(perf): report bounded mixed BSS measurements and query regressions Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * docs(perf): qualify timings affected by concurrent host builds Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * test(perf): include float BSS comparison in default regression cases Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * fix(perf): omit unsupported float encoding option from baseline setup Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> --------- Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> |
||
|
|
1fce0eb6d8 |
fix(ci): increase query regression ECS disk to 80 GiB (#9123)
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> |
||
|
|
4d65e8984a |
chore(ci): Implement /query-regression command handling and admission workflow (#8975)
* Implement `/query-regression` command handling and admission workflow - Add `query-regression-slash.py` script for processing `/query-regression` commands in PR comments, validating case arguments, and checking permissions. - Update `checks.yml` to include tests for the new slash command functionality. - Modify `query-regression-comment.yml` to trigger on the new `Query Regression Command` workflow. - Create `query-regression-slash.yml` to handle the dispatched command, validate allowlist and permissions, and initiate the regression workflow. - Enhance `query-regression.yml` to support additional inputs for PR admission and SHA verification. - Introduce `slash-command-dispatch.yml` to parse and dispatch commands from PR comments. - Document the new command admission process in `AGENTS.md` and `README.md`. - Add unit tests in `test_query_regression_slash.py` to cover command parsing and admission logic. * refactor: enhance query-regression command handling with comment validation and identity checks * feat: implement admission identity handling for query regression workflows * refactor: update PR admission logic in query regression workflow * refactor: update token usage in slash command dispatch and README for clarity * test: add cases for handling re-run failed jobs and stale runner artifacts * refactor: improve repository metadata handling in query regression scripts * chore: enable overwrite for artifact uploads to handle re-run failed jobs * chore: enable overwrite for query regression admission uploads * feat: enhance query-regression admission with HMAC signing and verification - Introduced HMAC signing for admission markers in query-regression workflows to ensure integrity and authenticity. - Updated `query-regression-comment.test.cjs` to include tests for signing and verifying admission markers. - Modified `query-regression-slash.py` to handle admission marker signing and verification, including checks for dispatch sender and head SHA consistency. - Enhanced workflows to securely manage admission markers and HMAC secrets, ensuring they are not exposed to untrusted contexts. - Improved documentation to clarify the admission process and the role of HMAC in securing the workflow. * test: add case to find newly posted marker among newer comments * test: add case to verify multiline output handling in write_outputs function |
||
|
|
6504af641e | fix: increase system disk size to 50 GiB for ECS instances (#8986) | ||
|
|
35ea88a4ef |
feat(ci): run query regression on ephemeral Aliyun ECS runners (#8937)
* feat(ci): add aliyun ecs ephemeral runner path for query regression Signed-off-by: paomian <xpaomian@gmail.com> * fix: improve condition for query-regression job execution in workflow * feat: update Docker installation to use official repository and add GPG key handling * Refactor query regression runner setup and configuration - Removed deprecated PersistentVolumeClaim for build cache. - Introduced a new bootstrap script for setting up the ECS runner host. - Deleted obsolete Helm values files for runner configuration. - Updated the Aliyun ECS runner provisioning script to reflect new cache paths. - Modified GitHub workflows to use the new Aliyun ECS runner setup. - Adjusted documentation to clarify the new runner lifecycle and provisioning process. * fix: enhance runner service management during bootstrap process * fix: update alibabacloud_tea_openapi dependency version in metadata * feat: enhance ECS runner scripts with region_id and resource_group_id support * fix: move containerd content store to data root for improved storage management * feat: rename query-regression runner to ephemeral-github runner and update related scripts * fix: update sentinel polling method to use serial console output for improved reliability * fix: add environment variable checks for Alibaba Cloud access keys in ECS client * fix: improve error handling in GitHub API requests for better diagnostics * fix: improve cache disk detection logic for Aliyun ECS instances * fix: enhance cache disk waiting logic with detailed output and error handling * fix: update dependency version for alibabacloud_tea_openapi in teardown script * fix: enhance cache disk waiting logic for better compatibility and clarity * fix: enhance console output handling and add incremental logging during instance provisioning * fix: add PATH environment variable for runner jobs in service and provision script * fix: add machine telemetry sampling and logging during query regression jobs * fix: update query regression documentation and provision script for cache disk handling * fix: update SCCACHE_CACHE_SIZE validation to 10G for improved caching efficiency * fix: remove outdated cache size checks and cleanup logic for fresh system disk runs * fix: enhance instance deletion logic with region handling and console output export * fix: add swap file setup and OOM handling for ECS runner to improve stability * fix: update OOM handling and service restart logic for ECS runner to enhance stability * fix: increase system disk size to 100 GiB for cold double nightly builds to prevent ENOSPC errors * fix: increase system disk size to 150 GiB for ECS runner to prevent ENOSPC errors * fix: add keep_instance option to preserve ECS instance for post-mortem debugging * fix: disable unattended upgrades to prevent job cancellations during library updates * fix: reduce system disk size to 40 GiB for ECS runner to prevent ENOSPC errors * feat: Refactor Aliyun ECS runner provisioning and introduce nightly regression comparison - Update `aliyun-ecs-runner-provision.py` to remove cache disk handling, simplifying the provisioning process. - Introduce `query-regression-nightly-refs.py` to resolve and compare SHAs from successful nightly builds. - Create `query-regression-nightly.yml` workflow to trigger nightly comparisons based on successful builds. - Enhance `query-regression.yml` to include a `test-tooling` job for validating Python scripts before provisioning. - Update tests for the new nightly reference selection logic and refactor existing tests to align with the new caching strategy. - Modify documentation to reflect changes in caching and nightly comparison workflows. * fix: enhance runner image tool verification with detailed checks * fix: improve error handling in runner image tool verification * fix: update tool versions in ECS image and workflow for consistency * fix: correct typo in error message for unparseable ECS creation time * fix: update README and workflow files for query regression tests and image hygiene --------- Signed-off-by: paomian <xpaomian@gmail.com> |
||
|
|
a7590f8174 |
perf(promql): avoid repeated scans in sliding range evaluation (#8646)
* perf(promql): use two pointers for sliding range boundaries Replace the stale cursor heuristic in RangeManipulateStream::calculate_range with monotonic left/right cursors. The old path rescanned each evaluation window (O(E x samples-per-window)) and could lose valid samples after sparse gaps or trailing empty windows. The two pointers keep strict monotonic progress, reducing boundary generation to O(N + E) while preserving (curr-range, curr] semantics, start/end shortening, and empty-window output. Controlled release benchmarks (fixed CPU, ABBA): - Public RangeManipulate wall time: ~28% faster at 1m/15s, ~66% at 5m/15s, ~96% at 1h/15s. - Warmed distributed TQL ANALYZE 1h queries: ~17-21% faster end to end; shorter windows stayed within run-order noise. Signed-off-by: discord9 <discord9@163.com> * perf(promql): specialize changes/resets with adaptive edge counting The generic range_fn macro slices, downcasts, and rescans every overlapping window for changes() and resets(). Replace the macro path for these two functions with hand-written UDF wrappers backed by a shared private edge-count kernel: direct raw-offset scans when requested edges are few, otherwise one global u64 edge prefix so each window is answered by a prefix difference. Behavior is preserved bit-for-bit, including raw null-buffer values, NaN semantics, signed zero, infinities, empty/singleton windows, independent timestamp/value offsets, arbitrary window layouts, and exact DataFusion error messages. The shared proc macro, planner, serializer, and other range functions are untouched. Controlled release benchmarks (fixed CPU, ABBA): - Dense sliding windows (k=4/20/240): 91.7-95.6% less public UDF wall time. - Low-coverage fallback (N=4096, 8 windows): 73.9-74.4% faster. - Warmed distributed TQL ANALYZE 5m/1h changes/resets: 12.1-19.7% client and 12.0-20.9% server latency improvement; controls stayed within drift. Signed-off-by: discord9 <discord9@163.com> * ci(query-regression): include PromQL range boundary case in defaults An audit of historical query-regression runs found zero range-query coverage: all 208 PromQL ANALYZE samples were bare selectors, so range evaluation could regress without CI noticing. Wire the promql_range_boundary case (introduced in #8646) into DEFAULT_CASES so label-triggered runs measure the range path. The case is cheap: a ~0.3s synthetic fixture and about a minute of query execution per base/candidate pass. Signed-off-by: discord9 <discord9@163.com> * chore(promql): address sliding range review nits Move test-only imports into their test modules and remove the unused pre-specialization changes and resets helpers. Signed-off-by: discord9 <discord9@163.com> * style(promql): apply pinned rustfmt Signed-off-by: discord9 <discord9@163.com> * test(promql): cover sparse range results Share the changes and resets test scaffolding while keeping their behavior oracles independent. Add an end-to-end sqlness regression for sparse samples, empty intermediate windows, and a valid trailing sample. Signed-off-by: discord9 <discord9@163.com> --------- Signed-off-by: discord9 <discord9@163.com> |
||
|
|
457a3f317e |
ci: add heavy query regression label (#8619)
Signed-off-by: discord9 <discord9@163.com> |
||
|
|
f5f5d468bb |
ci: add OTLP trace ingestion regression testing (#8631)
* chore: update CI config Signed-off-by: shuiyisong <xixing.sys@gmail.com> * chore: add CI Signed-off-by: shuiyisong <xixing.sys@gmail.com> * chore: update CI config Signed-off-by: shuiyisong <xixing.sys@gmail.com> * ci: report otelgen runner diagnostics Signed-off-by: shuiyisong <xixing.sys@gmail.com> * chore: add script to draw result diagram Signed-off-by: shuiyisong <xixing.sys@gmail.com> --------- Signed-off-by: shuiyisong <xixing.sys@gmail.com> |
||
|
|
a3816c889e |
fix(ci): harden query regression runner (#8534)
Signed-off-by: discord9 <discord9@163.com> |
||
|
|
16217ff567 |
ci: add persistent query regression cache (#8474)
* ci: add persistent query regression cache Signed-off-by: discord9 <discord9@163.com> * ci: add sccache to query regression runner Signed-off-by: discord9 <discord9@163.com> * ci: pin query regression label revision Signed-off-by: discord9 <discord9@163.com> * ci: isolate query regression toolchain state Signed-off-by: discord9 <discord9@163.com> --------- Signed-off-by: discord9 <discord9@163.com> |
||
|
|
d8f692f4c5 |
ci: run query regression on self-hosted runners (#8423)
* ci: run query regression on self-hosted runners Signed-off-by: discord9 <discord9@163.com> * ci: use dedicated perf regression runner labels Signed-off-by: discord9 <discord9@163.com> * ci: keep query regression runner scale set minimal Signed-off-by: discord9 <discord9@163.com> * ci: use custom query regression runner image Signed-off-by: discord9 <discord9@163.com> * ci: host query regression runner image in acr Signed-off-by: discord9 <discord9@163.com> * ci: harden query regression runner workflow Signed-off-by: discord9 <discord9@163.com> * ci: avoid runner uid assumptions in values Signed-off-by: discord9 <discord9@163.com> * ci: fix query regression comments for fork prs Signed-off-by: discord9 <discord9@163.com> * ci: address query regression review comments Signed-off-by: discord9 <discord9@163.com> --------- Signed-off-by: discord9 <discord9@163.com> |