* chore(toolchain): switch to stable Rust 1.96.1 and remove all nightly feature gates
Move the workspace from the pinned nightly-2026-03-21 to stable
1.96.1 and drop all 23 '#![feature]' gates across 13 crates,
rewriting the still-unstable API usages with stable equivalents:
- try_blocks: closures / an async block (table, query, common-function, servers)
- duration_constructors: Duration::from_secs(n * 86400) / (n * 60)
- iterator_try_collect: collect::<Result<Vec<_>, _>>()
- box_patterns: as_deref() + matches! chains (sql)
- error_iter: error_chain_root() source-chain walker (common-error);
sources() includes the error itself, so the walker never panics
- int_roundings: div_floor -> div_euclid (equal for positive divisors)
- iter_partition_in_place: stable sort_by_key partition helper (index)
- hash_set_entry: HashSet::insert bool / contains+insert
- trait_alias: lifetime-parameterized dyn FnOnce type aliases (puffin)
- string_from_utf8_lossy_owned: from_utf8_lossy(&v).into_owned()
- never_type: Infallible (common-recordbatch)
- debug_closure_helpers: closure-backed DebugFmt newtype (mito2)
- binary_heap_pop_if: peek().is_some_and() + pop()
- exclusive_wrapper: drop Exclusive; C: Send + Unpin already in bounds
- stmt_expr_attributes: stale gate, no usages
Also fix release-dev-builder-images.yaml, which parsed
rust-toolchain.toml with a date-only regex and would produce empty
image versions with a stable channel; it now extracts the full
channel token. Dev-builder images verified against stable 1.96.1
(image build, default-toolchain behavior, binstall/nextest, riscv64
and android targets, in-image cargo check).
Validated on 1.96.1: cargo check --workspace --all-targets, clippy
--workspace --all-targets --all-features -D warnings, cargo fmt
--check, and nextest on all 13 affected crates (4586 passed).
Part of #9289. Depends on #9298 (fuzz nightly quarantine) merging
first.
Signed-off-by: Ning Sun <sunning@greptime.com>
* chore: update flake checksum
* chore: use wild for linker in flake
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: use relative object keys for Windows filesystem access
Signed-off-by: WenyXu <wenymedia@gmail.com>
* fix: cover Windows path and time limits in full test CI
Signed-off-by: WenyXu <wenymedia@gmail.com>
* fix: use relative keys in metadata snapshot tests
Signed-off-by: WenyXu <wenymedia@gmail.com>
---------
Signed-off-by: WenyXu <wenymedia@gmail.com>
* chore(deps): replace cargo-udeps with cargo-shear for unused dependency checks
cargo-udeps requires a nightly toolchain and its pinned version (0.1.61)
no longer detects unused dependencies against current cargo internals —
unused deps have landed on main undetected (e.g. humantime in
common-frontend since #6689). cargo-shear is a standalone static analyzer
that runs on any toolchain.
- Swap 'make check-udeps' / 'make fix-udeps' recipes to 'cargo shear' /
'cargo shear --fix' and retire scripts/fix-udeps.py
- CI: install cargo-shear in the check-udeps job; drop the build cache
and protoc steps (cargo-shear never compiles)
- Remove ~150 unused dependency declarations found by cargo-shear, move
misplaced deps to the correct sections, drop orphaned
[workspace.dependencies] entries (arrow-cast, rustc-hash)
- Add [package.metadata.cargo-shear] ignored entries with explanations
for dependencies that are structurally required despite no textual
reference: sqlparser (required by sqlparser_derive expansions in
datatypes, common-query), common-error (required by common-macro's
stack_trace_debug expansions in session, tests-fuzz), k8s-openapi
(feature-pinning for the transitive kube dependency in tests-fuzz),
tikv-jemalloc-sys (link-only, enables jemalloc profiling features in
common-mem-prof), protobuf (required by build.rs-generated bindings in
log-store)
- Drop the obsolete [package.metadata.cargo-udeps.ignore] sections
Part of #9289
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix(meta): populate physical metric table column ids (#9286)
* fix(meta): populate physical metric table column ids
Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>
* test(meta): verify physical metric column ids
Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>
---------
Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>
* fix(postgres): return empty responses for comment-only SQL (#9295)
fix(postgres): handle parsed empty queries in both protocols
Signed-off-by: houyuwushang <180804215+houyuwushang@users.noreply.github.com>
* ci: create docs follow-up issue on PR merge instead of on label (#9237)
* ci: create docs follow-up issue on PR merge instead of on label
The docbot workflow previously created a docs-repo issue as soon as the
'docs-required' condition was detected (PR opened/edited with the docs
checkbox ticked), even if the PR was never merged.
Now the workflow also triggers on PR 'closed':
- opened/edited: only manage the docs-required/docs-not-required labels
- closed: create the docs issue only when the PR was actually merged and
carries the docs-required label
This also lets maintainers control issue creation by manually adding or
removing the docs-required label before merging.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: address review comments on docs issue creation timing
- Only touch docs labels when the docs checkbox state actually changed
in an edit. Previously, editing any other part of the PR body while
the checkbox stayed checked removed the docs-required label, silently
dropping the docs follow-up now that issue creation happens at merge.
Unchanged checkbox now leaves labels untouched, which also preserves
manual label overrides.
- Do not trust the closed event's stale label snapshot at merge time:
re-read the live PR via the API and create the docs issue if the
docs-required label is present OR the checkbox is ticked in the
current body.
- Make the workflow concurrency group action-aware so a merge run does
not cancel an in-flight label update from an edit run.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: make docs-required label the single source of truth at merge
The label-OR-checkbox merge condition could not distinguish an
intentional opt-out from an unfinished label update: removing
docs-required while the checkbox stayed checked still produced an
issue, and unchecking the box could still produce one if the merge
read the stale label before the edit run removed it.
At merge time, wait for any pending docbot runs on the PR head SHA to
finish their label updates (bounded to 5 minutes), then decide solely
by the live docs-required label. Adds actions: read permission for
listing workflow runs.
Signed-off-by: Ning Sun <sunning@greptime.com>
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
* perf(promql): push label filters into grouped join inputs (#9280)
* perf(promql): propagate matching filters through grouped joins
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* perf(promql): check matcher safety on the receiving operand
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor(promql): spell out the shapes a filter may cross
`preserves_filter` ended in `_ => true`, which was only sound because
`selector_matchers` independently rejects label rewriting, `count_values`,
subqueries and non-rollup calls on the same operand. Loosening the latter
alone would have silently pushed a matcher below a label rewrite. List the
shapes that carry a scan filter instead and default to `false`.
Cite #9207 for the result labels the grouped cases record: the join
projects the right operand's tag set, so `zone` is missing wherever the
right side aggregates it away.
No behavior change.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* test(promql): assert the new pushdowns reach the scan
The grouped-join unit tests feed tag columns by hand and the SQLness case
only checks results, which are identical whether or not the rewrite fires.
Nothing would have failed if scalar arithmetic, ranking or grouped
matching stopped propagating. Assert through the planner that the matcher
reaches both scans, with a global topk one-side as the counter-example.
Also state that the duplicate-one-side cases record a cross product
Prometheus rejects (#9209), so the baseline is not read as intended
semantics.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(ci): build tests-integration lib with meta-srv/mock (#9299)
* fix(ci): build tests-integration lib with meta-srv/mock
tests-integration's lib code (src/cluster.rs) uses meta_srv::mocks, but
the dependency carrying the mock feature sits in [dev-dependencies].
Builds that only touch the lib, such as the apidoc job's cargo doc
--workspace, resolve meta-srv without mock and fail with E0432.
--all-targets builds unify dev-dependency features, which is why check,
clippy and nextest stayed green.
Move the mock-enabled meta-srv entry back to [dependencies]. The other
testing features moved out in #9072 are not needed by the lib and stay
in [dev-dependencies].
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* test(repartition): split per-case repartition tests
test_repartition_metric ran four format/primary-key-encoding cases in a
single test function, and test_repartition_mito ran two format cases.
Each case builds its own 3-datanode cluster and runs a full repartition
plus GC cycle, so on S3 the metric test took 165-178s against the 180s
nextest terminate-after. Merge queue runs failed on it at random.
Split each case into its own test. Cases were already independent, so
they now run in parallel and each stays far inside the timeout, and a
failure points at one encoding instead of four.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* feat(json2): support altering JSON2 settings (#9029)
* feat(sql): support alter syntax for JSON2 columns
Signed-off-by: fys <fengys1996@gmail.com>
* fix(json2): preserve rows on type hint mismatch during compaction
* refactor(json2): simplify alter settings handling
* fix(json2): preserve coerced values during compaction
* chore: remove unnecessary clone
* chor: reduce memory allocations
* fix: cargo clippy
* chore: update greptime-proto to main branch
* refactor(datatypes): unify string handling with other JSON type hints
* fix: cr
---------
Signed-off-by: fys <fengys1996@gmail.com>
* fix: keep compaction pruning, metadata, and index work on compact runtime (#9304)
* fix: run compaction pruner tasks on compact runtime
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix: keep compaction metadata and index work on compact runtime
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
---------
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* feat: add AI matching, classification, and scoring functions (#9300)
* feat: return matching scores from jev
Replace the experimental three-argument Boolean function with jev(text, prompt) returning a Float64 probability in [0, 1]. Move threshold comparisons into SQL and update tests and migration examples.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* feat: add Jev choice and score functions
Share asynchronous execution across Noul, Choice, and Score. Validate JSON criteria before requests and return typed scalar answers. Add SQL and HTTP mock coverage with usage examples.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor: use generic AI SQL function names
Expose ai_match, ai_choose, and ai_score and move their implementation, tests, and usage guide under generic AI names. Document the current unreleased interface without migration history.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix: share constant AI criteria within each batch
Borrow scalar string arguments and lazily parse constant criteria once per batch. Share the parsed allocation across requests while preserving NULL propagation and batch validation before HTTP calls.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* feat: preserve AI score uncertainty in JSONB results
Return score, confidence, and probabilities in criteria-level order from one evaluation. Validate the distribution and preserve provider precision. Add JSON extraction, uncertainty, and single-request regressions, and document confidence-aware ranking.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* docs: explain reuse of volatile AI evaluations
Document repeated SELECT and WHERE evaluation costs as N + M requests, and show subquery aliases for reusing scalar or structured AI results without additional model calls.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
---------
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* feat: share logical table batching with OTLP metrics (#9288)
* feat: share logical table batching with OTLP metrics
Signed-off-by: WenyXu <wenymedia@gmail.com>
* fix: unify pending rows batch acknowledgement policy
Signed-off-by: WenyXu <wenymedia@gmail.com>
* fix: align logical batcher example configuration expectations
Signed-off-by: WenyXu <wenymedia@gmail.com>
* fix: align batcher worker channel defaults to 65536
Signed-off-by: WenyXu <wenymedia@gmail.com>
---------
Signed-off-by: WenyXu <wenymedia@gmail.com>
* perf(mito2): lazily decode dense primary key columns (#9226)
* perf(mito2): lazily decode dense primary key columns
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* perf(mito2): bypass lazy decoding for full primary keys
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito-codec): preserve prefix decoding errors
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito-codec): align encoded length helper naming
Rename encoded_length to encoded_len and update all callers to match the other length helpers in the module.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito2): clarify conditional dense key decoding
Rename decode_dense_pk to ensure_dense_pk_decoded so callers can see that existing decoded values are preserved and only missing caches are populated.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito-codec): share string framing in row converter
Move encoded_string_len to the parent module so Dense and Sparse use the same framing helper without depending on each other. Preserve its implementation and visibility.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
---------
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* chore: sync lock
* fix: shear and check issues
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>
Signed-off-by: houyuwushang <180804215+houyuwushang@users.noreply.github.com>
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
Signed-off-by: fys <fengys1996@gmail.com>
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
Signed-off-by: WenyXu <wenymedia@gmail.com>
Co-authored-by: Dhruv Vaishnav <dhruvvaishnav687@gmail.com>
Co-authored-by: houyuwushang <180804215+houyuwushang@users.noreply.github.com>
Co-authored-by: dennis zhuang <killme2008@gmail.com>
Co-authored-by: fys <40801205+fengys1996@users.noreply.github.com>
Co-authored-by: Lei, HUANG <6406592+v0y4g3r@users.noreply.github.com>
Co-authored-by: Weny Xu <wenymedia@gmail.com>
* test: cut integration test time and make the storage matrix meaningful
tests-integration is ~85% of workspace test CPU, and 81% of that is the
S3/S3WithCache variants of the HTTP and gRPC suites. Those suites do not
touch the object store: of the 70 matrix HTTP tests only one flushed and
read back an SST, so the matrix was paying real AWS round trips to
re-prove protocol parsing.
- Point the PR CI object-store matrix at the MinIO already started by
tests-integration/fixtures. Three GT_S3_* consumers did not read
GT_S3_ENDPOINT_URL and would have hit real AWS with MinIO credentials;
they now do.
- Add a nightly Linux job against real AWS S3, and pass GT_S3_* into the
release integration-test container. The release previously ran every
remote-backend case as a skip and only exercised the file backend.
- Give each S3WithCache test its own read cache directory. They shared
/tmp/greptimedb_cache, which the datanode wipes on startup, so a
starting test deleted the read cache of a running one.
- Add flush -> read-back assertions to the tests whose columns have a
non-trivial SST representation: JSON/JSON2 columns, native histograms,
metric-engine logical tables, and tables carrying fulltext or skipping
indexes whose puffin files only exist after a flush.
- Move eight tests that create no table out of the storage matrix.
- Make the event recorder flush interval a constructor parameter and
shorten it in the event tests, which otherwise wait a 5s window per DDL
they assert on. It is skipped by serde and never reaches config files.
- Drop duplicates: test_grpc_zstd_compression was a verbatim copy of
test_grpc_message_size_ok and is now rewritten to assert the negotiated
grpc-encoding; test_execute_copy_to_{s3,oss,gcs,azblob} were strict
prefixes of their copy_from siblings; two standalone/distributed event
test pairs shared one assertion body.
- Fix and un-ignore stddev_by_label. stddev_pop merges partial aggregates
in a parallelism-dependent order, so its last digits are unstable; the
test now compares values with a tolerance.
- Rebase the jaeger v1 fixture on the current instant. It carries
ttl=7d with 2025 timestamps, so its rows were only readable as long as
they stayed in the memtable.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* test: address review — wire nightly real-S3 job into check-status, keep the short event interval
The nightly `check-status` job did not depend on the new real-S3 job, so a
failure there would not have reached the status or Slack notification.
In database_ddl_event the short interval was set by a first
`with_event_recorder_options` call and then overwritten by the pre-existing
one, which carries `..Default::default()`. Merged into a single call.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* ci: update cargo fuzz command to use nightly toolchain explicitly
* ci: honor RUSTUP_TOOLCHAIN pin in fuzz orchestration script
An explicit `+toolchain` argument overrides the RUSTUP_TOOLCHAIN env var
in rustup precedence, so the hard-coded `cargo +nightly` in
run-fuzz-targets.sh bypassed the pinned FUZZ_RUST_TOOLCHAIN
(nightly-2026-03-21) configured in the workflow.
- Invoke `cargo +"${RUSTUP_TOOLCHAIN:-nightly}" fuzz run` in the script
so CI uses the pinned toolchain and local runs fall back to the
floating nightly
- Pass RUSTUP_TOOLCHAIN through to all four fuzz-test action invocations,
covering the no-prebuilt-binaries path and making the reproduce command
in the summary print the exact pinned toolchain
- Add test_rustup_toolchain_env_is_honored covering the pinned-env
scenario for both the cargo invocation args and the summary text
Addresses #9298 (review).
Signed-off-by: Ning Sun <sunning@greptime.com>
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci: create docs follow-up issue on PR merge instead of on label
The docbot workflow previously created a docs-repo issue as soon as the
'docs-required' condition was detected (PR opened/edited with the docs
checkbox ticked), even if the PR was never merged.
Now the workflow also triggers on PR 'closed':
- opened/edited: only manage the docs-required/docs-not-required labels
- closed: create the docs issue only when the PR was actually merged and
carries the docs-required label
This also lets maintainers control issue creation by manually adding or
removing the docs-required label before merging.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: address review comments on docs issue creation timing
- Only touch docs labels when the docs checkbox state actually changed
in an edit. Previously, editing any other part of the PR body while
the checkbox stayed checked removed the docs-required label, silently
dropping the docs follow-up now that issue creation happens at merge.
Unchanged checkbox now leaves labels untouched, which also preserves
manual label overrides.
- Do not trust the closed event's stale label snapshot at merge time:
re-read the live PR via the API and create the docs issue if the
docs-required label is present OR the checkbox is ticked in the
current body.
- Make the workflow concurrency group action-aware so a merge run does
not cancel an in-flight label update from an edit run.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: make docs-required label the single source of truth at merge
The label-OR-checkbox merge condition could not distinguish an
intentional opt-out from an unfinished label update: removing
docs-required while the checkbox stayed checked still produced an
issue, and unchecking the box could still produce one if the merge
read the stale label before the edit run removed it.
At merge time, wait for any pending docbot runs on the PR head SHA to
finish their label updates (bounded to 5 minutes), then decide solely
by the live docs-required label. Adds actions: read permission for
listing workflow runs.
Signed-off-by: Ning Sun <sunning@greptime.com>
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
* feat: add experimental Jev SQL filtering
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* feat: gate Jev filtering behind an opt-in Cargo feature
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* ci: verify Jev registration with default features
Run the existing registry regression without the jev feature in both PR tests and merge-queue coverage, alongside the existing feature-enabled test runs.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* docs: clarify Jev concurrency scope and stabilization work
Document the per-expression/batch concurrency bound and track process-wide limiting, rate-limit backoff, and request budgets and metrics as stabilization prerequisites.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* feat: enable renamed ai-functions feature by default
Rename the Jev Cargo feature across the command, query, and function crates and enable it in their defaults. Update CI and documentation, retaining an isolated no-default-features registry check and the runtime API opt-in.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* ci: remove extra AI feature-off checks
Use the regular AI-enabled unit and coverage runs for the default feature configuration. Keep feature-off validation available locally and update the usage guide to match.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor: rename AI feature to ai_functions
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
---------
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
crate-ci/typos renamed its default branch from master to main on
2026-09-18 (0a3d75e) and the master branch is gone, so every workflow
run since then fails at job setup with:
Unable to resolve action `crate-ci/typos@master`, unable to find version `master`
Pin to the latest release tag instead of tracking a branch. This matches
how the other actions in these two workflows are referenced and keeps the
check reproducible.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(ci): check Windows test targets before merge
Signed-off-by: WenyXu <wenymedia@gmail.com>
* fix(ci): use standard Windows runner for checks
Signed-off-by: WenyXu <wenymedia@gmail.com>
* fix(ci): omit dashboard assets from Windows checks
Signed-off-by: WenyXu <wenymedia@gmail.com>
---------
Signed-off-by: WenyXu <wenymedia@gmail.com>
The scheduled (nightly) release version was created by appending
'-nightly-YYYYMMDD' to NEXT_RELEASE_VERSION as-is. When the Cargo.toml
version carries a pre-release extension (e.g. v1.3.0-alpha.1), this
produced invalid tags like 'v1.3.0-alpha.1-nightly-20260907', stacking
'nightly' on top of the 'alpha.1' pre-release.
Strip the pre-release extension first so 'nightly' itself becomes the
only pre-release extension: 'v1.3.0-alpha.1' -> 'v1.3.0-nightly-20260908'.
Stable versions are unaffected; nightly-build ('nightly-YYYYMMDD-sha'),
dev-build, tag push, and manual dispatch paths are unchanged.
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci(backport): label backport PRs with their version name
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci(backport): restore multi-line PR body string
Signed-off-by: Ning Sun <sunning@greptime.com>
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
* Implement `/query-regression` command handling and admission workflow
- Add `query-regression-slash.py` script for processing `/query-regression` commands in PR comments, validating case arguments, and checking permissions.
- Update `checks.yml` to include tests for the new slash command functionality.
- Modify `query-regression-comment.yml` to trigger on the new `Query Regression Command` workflow.
- Create `query-regression-slash.yml` to handle the dispatched command, validate allowlist and permissions, and initiate the regression workflow.
- Enhance `query-regression.yml` to support additional inputs for PR admission and SHA verification.
- Introduce `slash-command-dispatch.yml` to parse and dispatch commands from PR comments.
- Document the new command admission process in `AGENTS.md` and `README.md`.
- Add unit tests in `test_query_regression_slash.py` to cover command parsing and admission logic.
* refactor: enhance query-regression command handling with comment validation and identity checks
* feat: implement admission identity handling for query regression workflows
* refactor: update PR admission logic in query regression workflow
* refactor: update token usage in slash command dispatch and README for clarity
* test: add cases for handling re-run failed jobs and stale runner artifacts
* refactor: improve repository metadata handling in query regression scripts
* chore: enable overwrite for artifact uploads to handle re-run failed jobs
* chore: enable overwrite for query regression admission uploads
* feat: enhance query-regression admission with HMAC signing and verification
- Introduced HMAC signing for admission markers in query-regression workflows to ensure integrity and authenticity.
- Updated `query-regression-comment.test.cjs` to include tests for signing and verifying admission markers.
- Modified `query-regression-slash.py` to handle admission marker signing and verification, including checks for dispatch sender and head SHA consistency.
- Enhanced workflows to securely manage admission markers and HMAC secrets, ensuring they are not exposed to untrusted contexts.
- Improved documentation to clarify the admission process and the role of HMAC in securing the workflow.
* test: add case to find newly posted marker among newer comments
* test: add case to verify multiline output handling in write_outputs function
* ci: add backport workflow to create backport PRs from backport labels
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci: document backport labels in PR template and AGENTS.md
Signed-off-by: Ning Sun <sunning@greptime.com>
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
* feat(runtime): add weighted workload scheduler
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* feat(runtime): switch catio to GreptimeTeam fork with admission-wait metrics
Use the GreptimeTeam/catio fork (pinned c20eafc) which adds
ClassStats::total_admission_wait and ClassStats::admitted, recorded
at each QUEUED -> ADMITTED transition. This exposes the scheduler's
own admission delay (excluding Tokio queueing and poll execution),
enabling admission-wait based fairness gates.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: bump catio to dynamic-config revision
Bump the catio scheduler fork to 9f4b028 which adds
Scheduler::set_weight and Scheduler::set_max_concurrent_polls for
runtime configuration.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* feat(perf): runtime-adjustable workload scheduler parameters
Expose dynamic adjustment of the experimental workload scheduler at
runtime:
- common-runtime: set_workload_scheduler_weights and
set_workload_scheduler_max_concurrent_polls, which forward to the
catio scheduler's set_weight/set_max_concurrent_polls when the
scheduler is enabled and reject zero values.
- servers: /debug/workload_scheduler/weights and
/debug/workload_scheduler/max_concurrent_polls POST handlers, so
operators can rebalance query/write shares or admission concurrency
without restarting the datanode.
Both endpoints return 400 with a clear reason when the scheduler is
disabled or the requested value is invalid.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* feat(perf): add GET /debug/workload_scheduler status endpoint
Returns the current weights (per class), max_concurrent_polls,
active_polls and per-class counters (queued, tasks, wakes, polls,
completed, cancelled, admitted, total_admission_wait) as JSON. When the
scheduler is disabled, returns enabled=false with the other fields
omitted, so operators can distinguish 'disabled' from an error.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: bump catio to time-accounting revision
Bump the catio scheduler fork to 257ba56 which replaces
admission-count accounting with real execution-time accounting
(pass += exec_time / (weight * concurrency)), so CPU share follows the
configured weights regardless of poll length.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: bump catio to lock-free sampling revision
Bump the catio scheduler fork to efdc0a4 which adds an optional
downsampled clock sampling mode (SchedulerBuilder::sample_every_polls,
default off) with a lock-free per-class atomic counter, so the
downsampled path costs one fetch_add per poll instead of a global
mutex.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: pin catio to scheduler PR head
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* feat(runtime): add scheduler bypass control
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: advance catio scheduler fixes
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: pin merged catio scheduler
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: regenerate config docs for workload scheduler
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: pin catio scheduler test fix
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* test(http): satisfy scheduler lifecycle clippy
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* test: add distributed scheduler toggle coverage
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* feat: finalize workload scheduler runtime controls
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: pin merged catio atomic weights
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: preserve unrelated lockfile resolution
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* perf(runtime): downsample scheduler time accounting
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* test(runtime): verify cross-runtime scheduler progress
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* feat(runtime): configure scheduler poll sampling
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* docs(runtime): clarify scheduler activation
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* docs(runtime): explain scheduler use case
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
---------
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
Co-authored-by: Ruihang Xia <waynestxia@gmail.com>
* feat(ci): add aliyun ecs ephemeral runner path for query regression
Signed-off-by: paomian <xpaomian@gmail.com>
* fix: improve condition for query-regression job execution in workflow
* feat: update Docker installation to use official repository and add GPG key handling
* Refactor query regression runner setup and configuration
- Removed deprecated PersistentVolumeClaim for build cache.
- Introduced a new bootstrap script for setting up the ECS runner host.
- Deleted obsolete Helm values files for runner configuration.
- Updated the Aliyun ECS runner provisioning script to reflect new cache paths.
- Modified GitHub workflows to use the new Aliyun ECS runner setup.
- Adjusted documentation to clarify the new runner lifecycle and provisioning process.
* fix: enhance runner service management during bootstrap process
* fix: update alibabacloud_tea_openapi dependency version in metadata
* feat: enhance ECS runner scripts with region_id and resource_group_id support
* fix: move containerd content store to data root for improved storage management
* feat: rename query-regression runner to ephemeral-github runner and update related scripts
* fix: update sentinel polling method to use serial console output for improved reliability
* fix: add environment variable checks for Alibaba Cloud access keys in ECS client
* fix: improve error handling in GitHub API requests for better diagnostics
* fix: improve cache disk detection logic for Aliyun ECS instances
* fix: enhance cache disk waiting logic with detailed output and error handling
* fix: update dependency version for alibabacloud_tea_openapi in teardown script
* fix: enhance cache disk waiting logic for better compatibility and clarity
* fix: enhance console output handling and add incremental logging during instance provisioning
* fix: add PATH environment variable for runner jobs in service and provision script
* fix: add machine telemetry sampling and logging during query regression jobs
* fix: update query regression documentation and provision script for cache disk handling
* fix: update SCCACHE_CACHE_SIZE validation to 10G for improved caching efficiency
* fix: remove outdated cache size checks and cleanup logic for fresh system disk runs
* fix: enhance instance deletion logic with region handling and console output export
* fix: add swap file setup and OOM handling for ECS runner to improve stability
* fix: update OOM handling and service restart logic for ECS runner to enhance stability
* fix: increase system disk size to 100 GiB for cold double nightly builds to prevent ENOSPC errors
* fix: increase system disk size to 150 GiB for ECS runner to prevent ENOSPC errors
* fix: add keep_instance option to preserve ECS instance for post-mortem debugging
* fix: disable unattended upgrades to prevent job cancellations during library updates
* fix: reduce system disk size to 40 GiB for ECS runner to prevent ENOSPC errors
* feat: Refactor Aliyun ECS runner provisioning and introduce nightly regression comparison
- Update `aliyun-ecs-runner-provision.py` to remove cache disk handling, simplifying the provisioning process.
- Introduce `query-regression-nightly-refs.py` to resolve and compare SHAs from successful nightly builds.
- Create `query-regression-nightly.yml` workflow to trigger nightly comparisons based on successful builds.
- Enhance `query-regression.yml` to include a `test-tooling` job for validating Python scripts before provisioning.
- Update tests for the new nightly reference selection logic and refactor existing tests to align with the new caching strategy.
- Modify documentation to reflect changes in caching and nightly comparison workflows.
* fix: enhance runner image tool verification with detailed checks
* fix: improve error handling in runner image tool verification
* fix: update tool versions in ECS image and workflow for consistency
* fix: correct typo in error message for unparseable ECS creation time
* fix: update README and workflow files for query regression tests and image hygiene
---------
Signed-off-by: paomian <xpaomian@gmail.com>
Posting to `/issues/{n}/comments` is authorized against the target object, and
that object is a pull request, so `issues: write` alone is refused with 403 and
the warning comment never lands.
Drafts are no longer counted and no longer warned about. `ready_for_review` is
added to the trigger types so that opening as a draft and flipping it to ready
still goes through the check.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(ci): identify team members by repository permission
`author_association` is computed from what the caller can see, so GITHUB_TOKEN
reports a private organization member as CONTRIBUTOR. Only 5 of GreptimeTeam's
members have public membership, so the open-pull-request check skipped almost
everyone it was written for.
Use the repository permission of the author instead, which is
viewer-independent. On error, apply the limit rather than skipping, so a token
that cannot read permissions cannot silently disable the check again.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(ci): do not log repository permission levels
Job logs are public. Resolving the author's permission is fine; printing the
level is not.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* Potential fix for pull request finding
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
* feat: add riscv64 cross-build support
Add the missing build infrastructure for riscv64gc-unknown-linux-gnu.
The codebase itself already compiles cleanly for riscv64 (verified with
`cargo check --workspace --target riscv64gc-unknown-linux-gnu`): all
architecture-sensitive dependencies support it (tikv-jemalloc-sys,
aws-lc-sys, ring, pprof, simd-json).
- .cargo/config.toml: set riscv64-linux-gnu-gcc as the linker for the
riscv64gc-unknown-linux-gnu target
- rust.yml: add a check-riscv64 CI job that cross-checks the whole
workspace to prevent regressions from future dependency changes
- docker/dev-builder/riscv64/Dockerfile: new cross dev-builder image
with gcc/g++-riscv64-linux-gnu and the riscv64 rust target
- Makefile: add dev-builder-riscv64 and build-riscv64-bin targets
Verified end-to-end: the produced riscv64 binary starts in standalone
mode under qemu and serves SQL (create/insert/select) over HTTP.
Signed-off-by: Lei, HUANG <mrsatangel@gmail.com>
* ci: build riscv64 artifacts in the release workflow
- release.yml: add build-linux-riscv64-artifacts job that cross-compiles
greptime for riscv64gc-unknown-linux-gnu on the amd64 runner with the
dev-builder-riscv64 image, and uploads greptime-linux-riscv64-*
artifacts. The job is wired into the needs of publish-github-release,
release-cn-artifacts and stop-linux-amd64-runner. Integration tests
are skipped since the cross-compiled binary cannot run on the host.
- release-dev-builder-images.yaml + build-dev-builder-images action:
build and push the dev-builder-riscv64 image to DockerHub, and sync
it to ECR and ACR via skopeo like the other dev-builder images.
- Makefile: DEV_BUILDER_RISCV64_IMAGE_TAG now defaults to
DEV_BUILDER_IMAGE_TAG so the existing tag-bump automation keeps the
riscv64 image tag in sync.
Signed-off-by: Lei, HUANG <mrsatangel@gmail.com>
* fix: forward cargo extension in riscv64 build
Pass CARGO_EXTENSION through build-riscv64-bin just like the existing
build-by-dev-builder target, so wrappers such as sccache are preserved
inside the cross-build container.
Signed-off-by: Lei, HUANG <mrsatangel@gmail.com>
* ci: check riscv64 release feature graph
Check all workspace targets with the servers/dashboard feature enabled so
the riscv64 CI job covers the same optional dependency graph used by the
release artifact build.
Signed-off-by: Lei, HUANG <mrsatangel@gmail.com>
* fix: gate riscv64 latest tags to main pushes
Manual dev-builder workflow dispatches now publish only their immutable
version tag. Update DockerHub and ECR latest tags only for the workflow's
main-branch push event, preventing feature-branch builds from replacing
the shared latest image.
Signed-off-by: Lei, HUANG <mrsatangel@gmail.com>
* docs: include riscv64 in release input description
Update the build_linux_artifacts workflow input description to reflect
that it now triggers amd64, arm64, and riscv64 artifact builds.
Signed-off-by: Lei, HUANG <mrsatangel@gmail.com>
* fix: fall back when riscv64 builder is unpublished
Before building a release artifact, pull the pinned RISC-V dev-builder
from ECR. If the image has not been published yet, build the same tag
locally from the current Dockerfile so releases remain unblocked during
the builder-tag update window.
Signed-off-by: Lei, HUANG <mrsatangel@gmail.com>
---------
Signed-off-by: Lei, HUANG <mrsatangel@gmail.com>
* chore(ci): warn when a member has too many open pull requests
Review capacity is the bottleneck. Add a `pull_request_target` workflow that
counts an org member's open pull requests (drafts included) on open/reopen and
posts a warning comment when the count exceeds the limit.
Advisory only for now: nothing is closed and no check fails.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* chore(ci): address review on the pr-open-limit workflow
- Only match marker comments authored by the Actions bot; a marker pasted
by anyone else would otherwise be picked up and fail the edit with 403.
- Validate MAX_OPEN_PRS and fall back to 5 on a non-numeric variable.
- Drop pull-requests write permission; commenting goes through the issues API.
- continue-on-error so a script failure never marks the pull request red.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor: port query regression runner to Rust
Signed-off-by: discord9 <discord9@163.com>
* ci: remove optional OTLP report plotter
Signed-off-by: discord9 <discord9@163.com>
* refactor: split query regression runner into modules
Signed-off-by: discord9 <discord9@163.com>
* style: use crate-qualified imports in query regression runner
Signed-off-by: discord9 <discord9@163.com>
* refactor: simplify query regression runner internals
Signed-off-by: discord9 <discord9@163.com>
* feat: abstract inspect-footer storage access behind object store destination
Add an optional --destination <TOML> to inspect-footer (and
--base-destination/--candidate-destination to finalize-remote) so the
storage inspection reads DB data files through the opendal-backed
object_store abstraction instead of bare std::fs. Local paths keep
working unchanged via the --root shortcut (File backend); remote
backends (S3/GCS/...) are described by a DestinationConfig TOML
reusing the object-store crate's ObjectStoreConfig serde shape.
- inspect_footer: list via ObjectStore::list + ObjectMeta filtering
(parquet keys, non-zero size, metadata/ segment), read footers
async via ParquetObjectReader + ParquetMetaDataReader with known
file size (no extra HEAD); output JSON schema unchanged
- finalize-remote: --base-data-home/--candidate-data-home become
optional, mutually exclusive with the new --*-destination args
- cmd deps: add object_store_opendal + datafusion_object_store
- tests: fs-backend list+footer integration tests (metadata filtering,
destination TOML mode, root/destination exclusivity)
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* style: drop needless borrow in inspect footer test
Fix clippy::needless_borrows_for_generic_args in the inspect-footer test
(fs::create_dir_all(table.join("metadata"))). Missed by the earlier
focused clippy run because it only covered --bin targets.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
---------
Signed-off-by: discord9 <discord9@163.com>
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* ci: automate compatibility version window
Signed-off-by: discord9 <discord9@163.com>
* ci: address compat version window review feedback
Address all four review comments on the compat version window
automation:
- Keep the PR window to the sliding window only (latest patch of the two
newest stable minor lines). Exact =vX.Y.Z anchors from case.toml are no
longer unioned into from_versions; they are validated by the new
--check-anchors mode and exercised by nightly runs via --nightly-window.
- Add --published-only: the window is computed over stable git tags that
have a published, non-draft GitHub release carrying the sqlness compat
artifacts (greptime-linux-amd64 tar.gz and sha256sum), so failed releases
cannot land in the window.
- Run the updater Python tests plus the window/anchor consistency check in
PR and merge-group CI (new compat-updater-check job in integration.yml).
- Regenerate tests/compatibility/ci.toml to the current sliding window
[v1.0.2, v1.1.4].
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
---------
Signed-off-by: discord9 <discord9@163.com>
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: gate soft-drop table behind the enterprise feature
Soft-drop table becomes an enterprise-only feature:
- metasrv rejects gc.experimental_soft_drop.enable=true at startup in
non-enterprise builds, and ddl_soft_drop_enabled is hard-disabled
without the enterprise feature as a second line of defense
- the UNDROP TABLE parser/AST/statement variant, ADMIN purge_table()
registration, and information_schema.recycle_bin registration are
compiled out unless the enterprise feature is enabled
- common-meta procedures, tombstone keys, and DdlTask serde stay
unconditional for persisted-procedure recovery and wire compatibility
- the [gc.experimental_soft_drop] section is removed from the OSS
example config and generated docs (moving to the enterprise repo)
- the soft-drop sqlness cases and their CI job are removed from OSS
(moving to the enterprise repo); affected information_schema .result
files are regenerated
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor: limit unused_variables allow to non-enterprise builds
Addresses review comment: apply the allow via cfg_attr so enterprise
builds still catch accidental unused variables in register_admin_only.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor: include the config key in the soft-drop enterprise gate error
Addresses review comment: name gc.experimental_soft_drop.enable in the
startup validation error so users can locate the setting quickly when
it is set via env vars or layered config.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* test: limit unused_mut allow to non-enterprise builds
Addresses review comment: apply the allow via cfg_attr so enterprise
builds still catch unused mut in the table_ddl_event test setup.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* feat: reject soft-drop DDL submissions in non-enterprise builds
Addresses review comment: clients could bypass the SQL-level gates by
submitting DdlTask::UndropTable or DdlTask::PurgeDroppedTable directly
to the procedure service. Reject fresh submissions at the DdlManager
boundary in non-enterprise builds while keeping the procedure loaders
registered for crash recovery and wire compatibility.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* test: stop --enable-gc from enabling soft drop in the sqlness template
Addresses review comment: the metasrv test template rendered
[gc.experimental_soft_drop] enable = true under the generic --enable-gc
flag, which non-enterprise metasrv now rejects at startup, making the
documented --enable-gc mode unusable in OSS. Keep the flag scoped to
plain GC; enterprise soft-drop coverage moves to the enterprise repo.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix: gate fresh soft-drop procedures
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* test: gate soft-drop fallback coverage
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix: gate soft-drop procedure implementation
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor: gate drop table soft-drop behavior
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor: gate expired soft-drop gc behavior
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* ci: test enterprise table ddl lifecycle
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* chore: mark purge_table as enterprise licensed
The purge_table module is compiled only with the enterprise feature, so
apply the Enterprise License header and register it with both license
header configurations.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* chore: mark recycle_bin as enterprise licensed
The recycle_bin module is compiled only with the enterprise feature, so
apply the Enterprise License header and register it with both license
header configurations.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* chore: mark soft-drop procedure sources as enterprise licensed
The purge and undrop procedure implementations plus the recycle-bin test
module compile only with the enterprise feature. Apply the Enterprise
License header and register them with both license configurations.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
---------
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* chore: check enterprise-gated files are listed in both license configs
A file reachable only through `#[cfg(feature = "enterprise")] mod ...;` is
governed by the GreptimeDB Enterprise License, so it must appear in the
`includes` of licenserc-enterprise.toml and the `excludes` of licenserc.toml.
hawkeye stays silent when it does not: the file keeps its Apache-2.0 header and
passes the default check precisely because it was never excluded from it.
scripts/check-enterprise-license.py walks enterprise-gated `mod` declarations,
resolves them to files (submodules included) and diffs that set against both
configs, also reporting stale entries. It runs in the license job in CI and as
`make check-enterprise-license`.
Documents the split it cannot decide for you — whole enterprise features get
their own file, a gated match arm stays inline — in
.agents/architecture-invariants.md.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix: tighten enterprise license checks
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>