mirror of
https://github.com/GreptimeTeam/greptimedb.git
synced 2026-10-03 02:25:35 +00:00
194bc2fb3c48ea1c55cff8fdd5886daa8b60d8ad
6229
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
194bc2fb3c |
feat(log-store): add the object store WAL durable write path (#9320)
* feat(log-store): add the object store WAL durable write path ObjectStoreLogStore can now write. append_batch admits entries into one open batch and assigns object-sequence-major entry ids at admission. The batch is sealed by size, by the flush interval, or before a region would run past the position range, and sealed batches are uploaded with at most four conditional creates in flight, started in sequence order. Created objects are indexed in sequence order, and an append is acknowledged only once its object is durable and indexed. A transient create failure rolls back the failed batch and every later batch unless a later object is already durable, in which case the store poisons itself with a history-gap error. A conflicting object, an encoding or catalog error, or taking the last representable sequence poisons the store. obsolete now goes through the actor and raises the sequence floor together with the watermark, refusing with a retryable error while the next sequence is not settled. stop drops the open batch and the batches whose create has not started, and lets creates in flight finish. Only the durable acknowledgement mode exists. A testing feature exposes hooks to wait for admissions, seal the open batch, and hold or fail creates. Signed-off-by: jeremyhi <fengjiachun@gmail.com> * test(log-store): synchronize object store WAL write tests with the actor Tests that assert nothing happened round-trip a command through the actor instead of yielding the test task, the conflict test waits for the object it expects, and the obsolete-behind-stop test holds the command channel itself so the unanswered command is deterministic. Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix(log-store): bound admitted WAL appends and never reuse a failed sequence A create that reports an error may still have written its object, so the sequences of failed batches are no longer handed out again: the next batch keeps the sequence after the last sealed one and retries are assigned new ids. A retry batched differently can no longer conflict with that object. As the next sequence never moves back, the sequence floor of obsolete only waits for an open batch that has handed out ids. Appends now arrive on their own bounded channel, which the actor stops reading while MAX_SEALED_BATCHES batches wait to become durable, so a stalled object store holds callers back instead of growing the backlog, while stop and obsolete are still handled. Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix(log-store): reserve room for two seals before admitting a WAL append One admission can seal the open batch before an append that would exhaust its positions and then the append's own batch, so the actor takes an append only while two more sealed batches fit under MAX_SEALED_BATCHES. Signed-off-by: jeremyhi <fengjiachun@gmail.com> * feat(log-store): chain object store WAL objects and recover the latest chain A conditional create can fail with an unknown outcome while its object is stored, or still lands later, and Mito reuses the row sequences of a failed append. Replaying such an object next to a later acknowledged one can let the unacknowledged rows win after a restart. Format version 2 gives every object header its writer's epoch, a link to the object it extends (sequence and writer instance) and a header CRC32, and allows objects without segments. Recovery replays only the chain ending at the complete object with the largest epoch and sequence: a link holds when its predecessor is present with the recorded writer instance, or is missing below every present object. Objects off the chain are orphans that are never replayed but keep their sequences. Each open writes an empty object that starts an epoch above every present object, linked to the recovered tip, before it accepts writes, so a late object of an earlier instance never ends the chain. A start object that meets an object of an earlier epoch moves to the next sequence; one of an equal or later epoch fails the open. Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix(log-store): break the WAL chain at every missing predecessor Accepting a missing predecessor below every present object lets a late object that lands below the chain change which links hold. Nothing collects objects yet, so a missing predecessor now always breaks the link, and recovery fails when objects are present but none completes a chain. Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix(log-store): keep the object store WAL format at version 1 The object store WAL has not been enabled anywhere, so no object in the previous layout exists and the chained header can stay version 1. Signed-off-by: jeremyhi <fengjiachun@gmail.com> * refactor(log-store): link object store WAL objects by epoch Only one store instance writes under an epoch, since every open starts an epoch above every present object and a start object that meets the same or a later epoch fails the open. The epoch therefore identifies the instance, and the random writer instance id is dropped from the header. A link now records the sequence and the epoch of the object it extends, and holds when the predecessor carries that epoch. The header shrinks to 46 bytes, and the store logs its epoch when it opens. Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix(log-store): fail an open whose start object is already present Without a random writer instance, two opens that recover the same objects encode byte-identical start objects, and a conditional create treats the same bytes as its own retry. A start object that is already present therefore fails the open with a retryable error instead of letting both opens claim the epoch; the next open counts the object and starts a later epoch. Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix(log-store): derive the WAL epoch from the claimed start sequence Two opens can recover different listings when a late object lands across a sequence gap between them, pick the same largest epoch plus one, and both create their start objects under different sequences. The epoch of an instance is now one above the sequence its start object claims, so a successful create decides the epoch and no two instances share one. It stays above every epoch recovery listed, and an object that carries an epoch above the next sequence fails the open. Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix(log-store): never run a held WAL create after the store is dropped The create test hook ignored the closed hold channel, so a create parked when the store was dropped could still run. It now returns without creating. Drop the per-admission bookkeeping of issued entry ids, which nothing reads, and move the parked I/O documentation to its helper. Signed-off-by: jeremyhi <fengjiachun@gmail.com> * test(log-store): cover a WAL start object stored with an unknown outcome Add a fault-injection test in which the create of the start object stores the object but reports an error: the open fails without moving to another sequence, and the next open counts the stored object and claims a later epoch. Rename the test helper that writes a whole object from a given header to put_object_with_header, and drop a needless clone. Signed-off-by: jeremyhi <fengjiachun@gmail.com> * test(log-store): make the held-create drop test deterministic Keep the actor running while the store drops the hold sender, so the parked create always completes on the closed channel instead of racing the actor's exit. Drop a comment that restates epoch_of. Signed-off-by: jeremyhi <fengjiachun@gmail.com> --------- Signed-off-by: jeremyhi <fengjiachun@gmail.com> |
||
|
|
3c2aac0a55 |
feat: add manual series index reconciliation (#9323)
* feat(mito): support manual series index reconciliation Signed-off-by: evenyag <realevenyag@gmail.com> * feat(storage): route series index build requests Signed-off-by: evenyag <realevenyag@gmail.com> * feat(admin): add BUILD_SERIES_INDEX Signed-off-by: evenyag <realevenyag@gmail.com> * test: specify compaction type in series index fixtures Signed-off-by: evenyag <realevenyag@gmail.com> * test: correct series index SQL fixtures and error assertions Signed-off-by: evenyag <realevenyag@gmail.com> * test(compat): preserve legacy index rebuild across upgrades Signed-off-by: evenyag <realevenyag@gmail.com> * test(sql): cover series index admin validation Signed-off-by: evenyag <realevenyag@gmail.com> * refactor(mito): bound series index maintenance queue Signed-off-by: evenyag <realevenyag@gmail.com> * docs: remove series index how-to guide Signed-off-by: evenyag <realevenyag@gmail.com> * test: remove index build upgrade compatibility case Signed-off-by: evenyag <realevenyag@gmail.com> * fix: address series index reconciliation review feedback Signed-off-by: evenyag <realevenyag@gmail.com> * fix: bound manual series index reconciliation admission Signed-off-by: evenyag <realevenyag@gmail.com> * chore: update greptime-proto to merged index build options Signed-off-by: evenyag <realevenyag@gmail.com> --------- Signed-off-by: evenyag <realevenyag@gmail.com> |
||
|
|
b547a3e369 |
ci: update dev-builder image tag (#9355)
Signed-off-by: greptimedb-ci <greptimedb-ci@greptime.com> Co-authored-by: greptimedb-ci <greptimedb-ci@greptime.com> |
||
|
|
b5c0e24625 |
feat(otlp): add histogram rejection metrics and compact zero buckets (#9321)
* feat(otlp): add histogram rejection metrics and compact zero buckets Signed-off-by: shuiyisong <xixing.sys@gmail.com> * fix: test isolation Signed-off-by: shuiyisong <xixing.sys@gmail.com> * chore: fix minor naming Signed-off-by: shuiyisong <xixing.sys@gmail.com> --------- Signed-off-by: shuiyisong <xixing.sys@gmail.com> |
||
|
|
357c935804 |
fix: make vector aggregates work with GROUP BY and partial aggregation (#9338)
* fix: make vector aggregates work with GROUP BY and partial aggregation Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test: make partitioned vec_avg case distinguish weighted averages Signed-off-by: Dennis Zhuang <killme2008@gmail.com> --------- Signed-off-by: Dennis Zhuang <killme2008@gmail.com> |
||
|
|
822042c7bb |
feat(log-store): chain object store WAL objects and recover the latest chain (#9334)
* feat(log-store): chain object store WAL objects and recover the latest chain A conditional create can fail with an unknown outcome while its object is stored, or still lands later, and Mito reuses the row sequences of a failed append. Replaying such an object next to a later acknowledged one can let the unacknowledged rows win after a restart. Format version 2 gives every object header its writer's epoch, a link to the object it extends (sequence and writer instance) and a header CRC32, and allows objects without segments. Recovery replays only the chain ending at the complete object with the largest epoch and sequence: a link holds when its predecessor is present with the recorded writer instance, or is missing below every present object. Objects off the chain are orphans that are never replayed but keep their sequences. Each open writes an empty object that starts an epoch above every present object, linked to the recovered tip, before it accepts writes, so a late object of an earlier instance never ends the chain. A start object that meets an object of an earlier epoch moves to the next sequence; one of an equal or later epoch fails the open. Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix(log-store): break the WAL chain at every missing predecessor Accepting a missing predecessor below every present object lets a late object that lands below the chain change which links hold. Nothing collects objects yet, so a missing predecessor now always breaks the link, and recovery fails when objects are present but none completes a chain. Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix(log-store): keep the object store WAL format at version 1 The object store WAL has not been enabled anywhere, so no object in the previous layout exists and the chained header can stay version 1. Signed-off-by: jeremyhi <fengjiachun@gmail.com> * refactor(log-store): link object store WAL objects by epoch Only one store instance writes under an epoch, since every open starts an epoch above every present object and a start object that meets the same or a later epoch fails the open. The epoch therefore identifies the instance, and the random writer instance id is dropped from the header. A link now records the sequence and the epoch of the object it extends, and holds when the predecessor carries that epoch. The header shrinks to 46 bytes, and the store logs its epoch when it opens. Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix(log-store): fail an open whose start object is already present Without a random writer instance, two opens that recover the same objects encode byte-identical start objects, and a conditional create treats the same bytes as its own retry. A start object that is already present therefore fails the open with a retryable error instead of letting both opens claim the epoch; the next open counts the object and starts a later epoch. Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix(log-store): derive the WAL epoch from the claimed start sequence Two opens can recover different listings when a late object lands across a sequence gap between them, pick the same largest epoch plus one, and both create their start objects under different sequences. The epoch of an instance is now one above the sequence its start object claims, so a successful create decides the epoch and no two instances share one. It stays above every epoch recovery listed, and an object that carries an epoch above the next sequence fails the open. Signed-off-by: jeremyhi <fengjiachun@gmail.com> * test(log-store): cover a WAL start object stored with an unknown outcome Add a fault-injection test in which the create of the start object stores the object but reports an error: the open fails without moving to another sequence, and the next open counts the stored object and claims a later epoch. Rename the test helper that writes a whole object from a given header to put_object_with_header, and drop a needless clone. Signed-off-by: jeremyhi <fengjiachun@gmail.com> --------- Signed-off-by: jeremyhi <fengjiachun@gmail.com> |
||
|
|
20619273ee |
chore(toolchain): switch to stable Rust 1.96.1 and remove all nightly feature gates (#9303)
* chore(toolchain): switch to stable Rust 1.96.1 and remove all nightly feature gates Move the workspace from the pinned nightly-2026-03-21 to stable 1.96.1 and drop all 23 '#![feature]' gates across 13 crates, rewriting the still-unstable API usages with stable equivalents: - try_blocks: closures / an async block (table, query, common-function, servers) - duration_constructors: Duration::from_secs(n * 86400) / (n * 60) - iterator_try_collect: collect::<Result<Vec<_>, _>>() - box_patterns: as_deref() + matches! chains (sql) - error_iter: error_chain_root() source-chain walker (common-error); sources() includes the error itself, so the walker never panics - int_roundings: div_floor -> div_euclid (equal for positive divisors) - iter_partition_in_place: stable sort_by_key partition helper (index) - hash_set_entry: HashSet::insert bool / contains+insert - trait_alias: lifetime-parameterized dyn FnOnce type aliases (puffin) - string_from_utf8_lossy_owned: from_utf8_lossy(&v).into_owned() - never_type: Infallible (common-recordbatch) - debug_closure_helpers: closure-backed DebugFmt newtype (mito2) - binary_heap_pop_if: peek().is_some_and() + pop() - exclusive_wrapper: drop Exclusive; C: Send + Unpin already in bounds - stmt_expr_attributes: stale gate, no usages Also fix release-dev-builder-images.yaml, which parsed rust-toolchain.toml with a date-only regex and would produce empty image versions with a stable channel; it now extracts the full channel token. Dev-builder images verified against stable 1.96.1 (image build, default-toolchain behavior, binstall/nextest, riscv64 and android targets, in-image cargo check). Validated on 1.96.1: cargo check --workspace --all-targets, clippy --workspace --all-targets --all-features -D warnings, cargo fmt --check, and nextest on all 13 affected crates (4586 passed). Part of #9289. Depends on #9298 (fuzz nightly quarantine) merging first. Signed-off-by: Ning Sun <sunning@greptime.com> * chore: update flake checksum * chore: use wild for linker in flake --------- Signed-off-by: Ning Sun <sunning@greptime.com> |
||
|
|
abedeb2bea |
fix(query): keep count_values generated label in enclosing expressions (#9223)
* fix(query): keep count_values generated label in enclosing expressions
`count_values("v", m)` projects the generated label as a real output
column, but did not register it in `ctx.tag_columns`. Enclosing
expressions (abs/round/+1/topk/label_replace/vector join) rebuild their
projection from `ctx.tag_columns` and silently drop the label.
Register the generated label in `ctx.tag_columns` after the projection,
and give it the same qualifier as other tag columns so qualified
references resolve. PromQL overwrites an input label with the same name,
so drop any existing tag with that name first to avoid duplicate column
ambiguity.
Fixes https://github.com/GreptimeTeam/greptimedb/issues/9181
Report: .e-agent/greptimedb_promql_compatibility_report_2026-09-16.md P0-1
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* fix(query): use total_cmp in count_values test helper
Silence clippy::needless_borrow on partial_cmp(&right.1); f64 sorting
uses total_cmp, matching the other planner test helpers.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* fix(flow): keep count_values generated label as sink table primary key
FindGroupByFinalName::f_up only renamed a group key when the projection
aliased the key column directly. count_values now projects its generated
label as a unary UDF over the sampled column
(prom_float_to_string(value) AS label) while the aggregate still groups
by the raw column, so the name match failed and the sink table lost the
label from its PRIMARY KEY, demoting it to a DOUBLE value column.
Allow a projection above the aggregate to rename a group key by deriving
its output from a single group-key column (unary scalar function or cast
over that column only). Multi-column expressions, case, aggregates,
windows, subqueries and literals are still rejected so a computed column
cannot be mistaken for the group key.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* fix(query): group count_values by the formatted sample value
The numeric count_values branch grouped by the raw sample column and only
formatted the value into the generated label in the post-aggregate
projection. Two distinct raw values that collapse to one label text
(e.g. BIGINT 9007199254740992 and 9007199254740993, both 9007199254740992
in Float64) were split into two groups, each emitting the same label set
at one timestamp, violating Prometheus' unique-label-set-per-timestamp
invariant.
Return the formatted value expression (prom_float_to_string, with a
CAST to Float64 for non-Float64 inputs) as previous_field_expressions so
the aggregate groups by the same expression that produces the label,
mirroring the existing mixed float/native-histogram precedent.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* fix(query): drop needless borrow in count_values test helper
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* fix(flow): don't replace group key with derived unary expr
FindGroupByFinalName treated any unary expression of a single group
column (scalar fn, Cast, TryCast) as a rename of that group key and
swapped it in as the sink primary key. A derived expression such as
lower(host) is many-to-one, so distinct groups (HOST_A vs host_a) would
collapse to the same primary key and be silently deduplicated.
Narrow the matching to direct name-matched aliases of the actual group
expression only, and drop the derived-unary path (is_unary_expr_of_column)
and the allow_derived flag. count_values already groups by the formatted
expression name, so it is unaffected.
Add a regression asserting host survives and host_lc does not replace it
under both optimizer settings.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
---------
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
|
||
|
|
d59725b04a |
test(sqlness): run environments concurrently and drop redundant restarts and sleeps (#9333)
* test(sqlness): run environments concurrently with external backends The runner forced both environment and instance parallelism to 1 whenever etcd/PG/MySQL or an external Kafka was set up, so the standalone and distributed environments ran one after the other. Only distributed uses the kv backend, and the two environments use different Kafka topic prefixes, so they can run at the same time. Keep one instance per environment, since instances would share the metadata table and Kafka topics. Runs with a runner-managed Kafka, an external server, a test filter, or `-j 1` stay fully serial. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test(sqlness): drop redundant restarts and sleeps - flow_rebuild: drop the R2 block, which repeats R1 from the same state, and the restart in R5, whose result R1 already asserts. - flow_advance_ttl, flow_view, ttl_instant: drop sleeps that no assertion depends on. - flow_basic, flow_null, flow_call_df_func: drop the bytes_log section (covered by flow_insert), a duplicate state_size query, an unchecked insert, and flushes that run with no new data. - alter_table_options, skip_wal: share restarts between independent tables. - session_skip_wal, copy_skip_wal: check every case after a single restart. Tables copied or inserted with skip_wal = true first get a WAL row that is truncated, so the check that truncated WAL entries are not replayed stays. - region_statistics: wait once for the statistics of all three tables. - build_index_table: fold its index_size checks into build_index_table_restart, which runs the same fixture and waits. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test(sqlness): restore session_skip_wal and copy_skip_wal Checking every case after a single restart dropped two state transitions the original cases cover: recovering a region whose memtable only held skipped rows, and writing to the recovered region before restarting again. Restore the original cases. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test(sqlness): restore the one-restart check in flow_auto_sink_table #5112 checked the auto-created sink before a restart and the flow and sink after it. #5987 moved the restart in front of the first SHOW and added a second one. Flow recovery only creates the sink when it is missing, so the second restart recovers from the same persisted state as the first. Restore the original before/after check with one restart. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test(sqlness): keep one instance per environment with external store addresses Instances of the distributed environment share the etcd behind --store-addrs, so run one instance per environment when it is set, same as --setup-etcd. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test(sqlness): restore the false-branch batch in flow_basic and assert it The kept batch (20, 22) is all above the threshold. The second batch (10, 23) is the only input that hits the false branch of the CASE and makes the flow compute again. Restore it and check the result, which the original case never did. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> --------- Signed-off-by: Dennis Zhuang <killme2008@gmail.com> |
||
|
|
affc0a1b1d |
fix: align ordered aggregate state type with the accumulator output (#9340)
* fix: align ordered aggregate state type with the accumulator output Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix: reject mismatched aggregate states and keep hard-ordered aggregates unsplit Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix: keep WITHIN GROUP aggregates splittable and relabel state fields by position Signed-off-by: Dennis Zhuang <killme2008@gmail.com> --------- Signed-off-by: Dennis Zhuang <killme2008@gmail.com> |
||
|
|
def0c2ec5c |
fix: only push down aggregates grouped by the partition columns themselves (#9337)
* fix: only push down aggregates grouped by the partition columns themselves Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix: keep grouping sets on the frontend for partitioned tables Signed-off-by: Dennis Zhuang <killme2008@gmail.com> --------- Signed-off-by: Dennis Zhuang <killme2008@gmail.com> |
||
|
|
b9f991502c |
refactor: remove experimental vector index (#9345)
Signed-off-by: Dennis Zhuang <killme2008@gmail.com> |
||
|
|
cc1be82858 |
fix: merge every state row in geo_path and json_encode_path (#9339)
Signed-off-by: Dennis Zhuang <killme2008@gmail.com> |
||
|
|
c7fa48ef95 |
feat(mito): limit approximate series index disk usage (#9313)
* feat(mito): limit series index disk usage Signed-off-by: evenyag <realevenyag@gmail.com> * refactor(mito): enforce series index quota during reconciliation Signed-off-by: evenyag <realevenyag@gmail.com> * refactor(mito): remove series index disk budget layer Signed-off-by: evenyag <realevenyag@gmail.com> * refactor(mito): simplify series index limit to estimated usage Signed-off-by: evenyag <realevenyag@gmail.com> * refactor(mito): trust index catalogs when loading snapshots Signed-off-by: evenyag <realevenyag@gmail.com> * refactor(mito): minimize series index disk limit changes Signed-off-by: evenyag <realevenyag@gmail.com> * fix(mito2): keep series index cleanup running at capacity Signed-off-by: evenyag <realevenyag@gmail.com> * fix(mito2): avoid no-op index clones and stabilize capacity tests Signed-off-by: evenyag <realevenyag@gmail.com> * fix: update config API expectation and stabilize index build tests Signed-off-by: evenyag <realevenyag@gmail.com> --------- Signed-off-by: evenyag <realevenyag@gmail.com> |
||
|
|
1955955b2d |
fix(ci): grant PR write permission for CI command replies (#9331)
Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
20dde2601f |
feat: enable native histogram ingestion by default (#9301)
* feat: enable native histogram ingestion by default Signed-off-by: shuiyisong <xixing.sys@gmail.com> * chore: add comments Signed-off-by: shuiyisong <xixing.sys@gmail.com> * fix: test Signed-off-by: shuiyisong <xixing.sys@gmail.com> --------- Signed-off-by: shuiyisong <xixing.sys@gmail.com> |
||
|
|
4c98fa5265 |
ci: add optional AWS runners for observability benchmarks (#9322)
* ci: add optional AWS runners for observability benchmarks Signed-off-by: WenyXu <wenymedia@gmail.com> * ci: restore automatic benchmark disk sizing Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
2bc18c483d |
feat: add schema metadata stream to SchemaManager (#9325)
Signed-off-by: shuiyisong <xixing.sys@gmail.com> |
||
|
|
88197f4019 |
fix(promql): derive vector matching result labels and reject ambiguous matchings (#9306)
* fix(promql): derive vector matching result labels and reject ambiguous matchings A vector-vector binary operation projected one operand's whole tag set and inner-joined without any cardinality check, so `on()`/`ignoring()` did not reduce the result labels, `group_left`/`group_right` changed nothing, and a non-unique match group produced a cross product that PromQL cannot represent. Result labels now follow Prometheus `resultMetric`: `on(...)` keeps the matching labels, `ignoring(...)` drops them, and a group modifier keeps the many side's labels plus the `group_x(...)` labels taken from the one side. A label the one side does not carry is deleted from the result. The reduced label set no longer identifies the operand series, so `__tsid` is dropped from the context on this path. Cardinality is enforced with a `count(1) OVER (PARTITION BY match keys, ts)` window and a scalar UDF that fails the query on a repeated group: on the one side before the join, and on the result labels after it, matching where Prometheus raises each of its three errors. Series are unique by their whole tag set, so the window is only planted when the match keys drop a tag; plain arithmetic and `on(<all tags>)` plan exactly as before. Closes #9207, closes #9208, closes #9209. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * perf(promql): keep __tsid when the result labels are an operand's whole tag set Deriving the result labels dropped `__tsid` from the context unconditionally, so an enclosing operation fell back to joining on the tag columns even where the column still identified the result series. Keep it when every result label comes from one operand and covers that operand's whole tag set: no other operand value reaches the labels, and the matching gives each of its rows a single partner, so its `__tsid` is still one per result series. That is the common `on(<all tags>)` and bare `group_left` shape; a matching that actually drops a tag still clears it. The column is re-qualified as the result's own, which is how the enclosing expression and the context look it up. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix(promql): keep the match group count column unambiguous The cardinality check aliased its row count to a fixed `__promql_match_group_count`. An operand carrying a label of that name made the window output two fields with the same name, and planning failed with "Schema contains qualified field name collide_right.__promql_match_group_count and unqualified field name __promql_match_group_count which would be ambiguous". Pick a name the operand does not already have, the way the `or` operator allocates its match key columns. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test(promql): cover a match group spread over several regions The cardinality check runs above the merge of the region scans, so it counts a match group globally. Nothing pinned that: every table in these cases holds a single region, and a check evaluated per region would pass them all. Partition the operand on a column outside the match keys, which puts the two series of one match group in different regions, and assert both the pre-join and the post-join check still reject it. The case runs in the distributed environment too, where the regions sit on different datanodes. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix(promql): group an outer aggregate on the labels the operand kept `by`/`without` planning topped up a missing grouping column by walking down to the table scan and re-projecting it. That is right for a column the scan pruned for efficiency, but the labels a matching modifier deletes are also absent from the operand's output, and they were restored the same way: sum without(host) (a / on(host) b) `on(host)` leaves the operand with `host` alone, so the sum covers everything and Prometheus answers `{} 10`. Instead `device` came back from the scan under `a` and split the result into `{device="d1"} 5` and `{device="d2"} 5`. Same for `sum by(device)` of that operand, which has no `device` to group on at all. Take the grouping labels from the operand's own label set rather than from the row keys of the scan beneath it. A label pruned from the plan is still in that set and still gets restored; a label the operand dropped is not. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * refactor(promql): drop the unreachable aggregation tag top-up `by`/`without` planning could restore a grouping column that the plan no longer carried by rewriting the scan underneath it. Once the grouping labels come from the operand's own label set, there is nothing left for it to restore: a scan projects every label of `ctx.tag_columns` (`scan_tag_columns` only ever adds matcher columns to that set), so a label in the set is always in the schema. Stubbing the rewriter to a no-op passed the whole sqlness suite, in both the standalone and the distributed environment. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test(promql): drop cases that guarded the removed tag top-up Three plain selector aggregates were there to show that restoring a pruned grouping column still worked. With the restore gone they only repeat what the aggregate cases already cover. Also fix a comment that still said the metric engine scan prunes tag columns: it projects every label of the operand. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * docs(promql): note the duplicate a propagated matcher hides A matcher copied onto the one-side operand removes groups without a partner before the cardinality check sees them, so a duplicate in such a group is not reported. Prometheus checks every group of the one side and fails the query. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> --------- Signed-off-by: Dennis Zhuang <killme2008@gmail.com> |
||
|
|
37fed9f12b |
fix(ci): repair draft PR command dispatch (#9271)
* fix(ci): repair draft PR command dispatch Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(ci): dispatch command workflows by branch ref Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(ci): admit PR authors and writers for CI commands Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(ci): preserve dispatch guard dependency semantics Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(ci): configure slash command permissions individually Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(ci): run command workflows at dispatch branch head Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(ci): rerun fork PR checks and report command failures Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
d910183cc4 |
feat(flow): freeze recovery windows and retention bounds for incremental flows (#9312)
* feat(flow): discover recovery windows from exact sequence reads Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * feat(flow): freeze explicit recovery windows and retention bounds Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * fix(flow): validate retained source windows before recovery Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * fix(flow): qualify recovery source tables before serialization Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> --------- Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> |
||
|
|
5a78e07be6 |
fix(json2): restrict JSON2 type hints (#9316)
* fix(sql): restrict JSON2 type hints Signed-off-by: fys <fengys1996@gmail.com> * fix(sql): allow equivalent 64-bit aliases in JSON2 type hints Signed-off-by: fys <fengys1996@gmail.com> * fix(sql): support UInt64 conversion and preserve JSON2 hint aliases Signed-off-by: fys <fengys1996@gmail.com> * refactor(sql): remove redundant JSON2 type hint normalization Signed-off-by: fys <fengys1996@gmail.com> * fix(datatypes): restrict JSON2 hint types in JsonSettings::try_new Signed-off-by: fys <fengys1996@gmail.com> --------- Signed-off-by: fys <fengys1996@gmail.com> |
||
|
|
bc80071409 |
feat: expose region open failure metrics (#9283)
Signed-off-by: evenyag <realevenyag@gmail.com> |
||
|
|
a35f9d5c78 |
fix: address Windows test failures and run full Windows CI (#9305)
* fix: use relative object keys for Windows filesystem access Signed-off-by: WenyXu <wenymedia@gmail.com> * fix: cover Windows path and time limits in full test CI Signed-off-by: WenyXu <wenymedia@gmail.com> * fix: use relative keys in metadata snapshot tests Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
1d8d95c12a |
chore(deps): replace cargo-udeps with cargo-shear for unused dependency checks (#9294)
* chore(deps): replace cargo-udeps with cargo-shear for unused dependency checks cargo-udeps requires a nightly toolchain and its pinned version (0.1.61) no longer detects unused dependencies against current cargo internals — unused deps have landed on main undetected (e.g. humantime in common-frontend since #6689). cargo-shear is a standalone static analyzer that runs on any toolchain. - Swap 'make check-udeps' / 'make fix-udeps' recipes to 'cargo shear' / 'cargo shear --fix' and retire scripts/fix-udeps.py - CI: install cargo-shear in the check-udeps job; drop the build cache and protoc steps (cargo-shear never compiles) - Remove ~150 unused dependency declarations found by cargo-shear, move misplaced deps to the correct sections, drop orphaned [workspace.dependencies] entries (arrow-cast, rustc-hash) - Add [package.metadata.cargo-shear] ignored entries with explanations for dependencies that are structurally required despite no textual reference: sqlparser (required by sqlparser_derive expansions in datatypes, common-query), common-error (required by common-macro's stack_trace_debug expansions in session, tests-fuzz), k8s-openapi (feature-pinning for the transitive kube dependency in tests-fuzz), tikv-jemalloc-sys (link-only, enables jemalloc profiling features in common-mem-prof), protobuf (required by build.rs-generated bindings in log-store) - Drop the obsolete [package.metadata.cargo-udeps.ignore] sections Part of #9289 Signed-off-by: Ning Sun <sunning@greptime.com> * fix(meta): populate physical metric table column ids (#9286) * fix(meta): populate physical metric table column ids Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> * test(meta): verify physical metric column ids Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> --------- Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> * fix(postgres): return empty responses for comment-only SQL (#9295) fix(postgres): handle parsed empty queries in both protocols Signed-off-by: houyuwushang <180804215+houyuwushang@users.noreply.github.com> * ci: create docs follow-up issue on PR merge instead of on label (#9237) * ci: create docs follow-up issue on PR merge instead of on label The docbot workflow previously created a docs-repo issue as soon as the 'docs-required' condition was detected (PR opened/edited with the docs checkbox ticked), even if the PR was never merged. Now the workflow also triggers on PR 'closed': - opened/edited: only manage the docs-required/docs-not-required labels - closed: create the docs issue only when the PR was actually merged and carries the docs-required label This also lets maintainers control issue creation by manually adding or removing the docs-required label before merging. Signed-off-by: Ning Sun <sunning@greptime.com> * fix: address review comments on docs issue creation timing - Only touch docs labels when the docs checkbox state actually changed in an edit. Previously, editing any other part of the PR body while the checkbox stayed checked removed the docs-required label, silently dropping the docs follow-up now that issue creation happens at merge. Unchanged checkbox now leaves labels untouched, which also preserves manual label overrides. - Do not trust the closed event's stale label snapshot at merge time: re-read the live PR via the API and create the docs issue if the docs-required label is present OR the checkbox is ticked in the current body. - Make the workflow concurrency group action-aware so a merge run does not cancel an in-flight label update from an edit run. Signed-off-by: Ning Sun <sunning@greptime.com> * fix: make docs-required label the single source of truth at merge The label-OR-checkbox merge condition could not distinguish an intentional opt-out from an unfinished label update: removing docs-required while the checkbox stayed checked still produced an issue, and unchecking the box could still produce one if the merge read the stale label before the edit run removed it. At merge time, wait for any pending docbot runs on the PR head SHA to finish their label updates (bounded to 5 minutes), then decide solely by the live docs-required label. Adds actions: read permission for listing workflow runs. Signed-off-by: Ning Sun <sunning@greptime.com> --------- Signed-off-by: Ning Sun <sunning@greptime.com> * perf(promql): push label filters into grouped join inputs (#9280) * perf(promql): propagate matching filters through grouped joins Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * perf(promql): check matcher safety on the receiving operand Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * refactor(promql): spell out the shapes a filter may cross `preserves_filter` ended in `_ => true`, which was only sound because `selector_matchers` independently rejects label rewriting, `count_values`, subqueries and non-rollup calls on the same operand. Loosening the latter alone would have silently pushed a matcher below a label rewrite. List the shapes that carry a scan filter instead and default to `false`. Cite #9207 for the result labels the grouped cases record: the join projects the right operand's tag set, so `zone` is missing wherever the right side aggregates it away. No behavior change. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test(promql): assert the new pushdowns reach the scan The grouped-join unit tests feed tag columns by hand and the SQLness case only checks results, which are identical whether or not the rewrite fires. Nothing would have failed if scalar arithmetic, ranking or grouped matching stopped propagating. Assert through the planner that the matcher reaches both scans, with a global topk one-side as the counter-example. Also state that the duplicate-one-side cases record a cross product Prometheus rejects (#9209), so the baseline is not read as intended semantics. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> --------- Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix(ci): build tests-integration lib with meta-srv/mock (#9299) * fix(ci): build tests-integration lib with meta-srv/mock tests-integration's lib code (src/cluster.rs) uses meta_srv::mocks, but the dependency carrying the mock feature sits in [dev-dependencies]. Builds that only touch the lib, such as the apidoc job's cargo doc --workspace, resolve meta-srv without mock and fail with E0432. --all-targets builds unify dev-dependency features, which is why check, clippy and nextest stayed green. Move the mock-enabled meta-srv entry back to [dependencies]. The other testing features moved out in #9072 are not needed by the lib and stay in [dev-dependencies]. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test(repartition): split per-case repartition tests test_repartition_metric ran four format/primary-key-encoding cases in a single test function, and test_repartition_mito ran two format cases. Each case builds its own 3-datanode cluster and runs a full repartition plus GC cycle, so on S3 the metric test took 165-178s against the 180s nextest terminate-after. Merge queue runs failed on it at random. Split each case into its own test. Cases were already independent, so they now run in parallel and each stays far inside the timeout, and a failure points at one encoding instead of four. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> --------- Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * feat(json2): support altering JSON2 settings (#9029) * feat(sql): support alter syntax for JSON2 columns Signed-off-by: fys <fengys1996@gmail.com> * fix(json2): preserve rows on type hint mismatch during compaction * refactor(json2): simplify alter settings handling * fix(json2): preserve coerced values during compaction * chore: remove unnecessary clone * chor: reduce memory allocations * fix: cargo clippy * chore: update greptime-proto to main branch * refactor(datatypes): unify string handling with other JSON type hints * fix: cr --------- Signed-off-by: fys <fengys1996@gmail.com> * fix: keep compaction pruning, metadata, and index work on compact runtime (#9304) * fix: run compaction pruner tasks on compact runtime Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * fix: keep compaction metadata and index work on compact runtime Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> --------- Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * feat: add AI matching, classification, and scoring functions (#9300) * feat: return matching scores from jev Replace the experimental three-argument Boolean function with jev(text, prompt) returning a Float64 probability in [0, 1]. Move threshold comparisons into SQL and update tests and migration examples. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * feat: add Jev choice and score functions Share asynchronous execution across Noul, Choice, and Score. Validate JSON criteria before requests and return typed scalar answers. Add SQL and HTTP mock coverage with usage examples. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * refactor: use generic AI SQL function names Expose ai_match, ai_choose, and ai_score and move their implementation, tests, and usage guide under generic AI names. Document the current unreleased interface without migration history. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * fix: share constant AI criteria within each batch Borrow scalar string arguments and lazily parse constant criteria once per batch. Share the parsed allocation across requests while preserving NULL propagation and batch validation before HTTP calls. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * feat: preserve AI score uncertainty in JSONB results Return score, confidence, and probabilities in criteria-level order from one evaluation. Validate the distribution and preserve provider precision. Add JSON extraction, uncertainty, and single-request regressions, and document confidence-aware ranking. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * docs: explain reuse of volatile AI evaluations Document repeated SELECT and WHERE evaluation costs as N + M requests, and show subquery aliases for reusing scalar or structured AI results without additional model calls. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> --------- Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * feat: share logical table batching with OTLP metrics (#9288) * feat: share logical table batching with OTLP metrics Signed-off-by: WenyXu <wenymedia@gmail.com> * fix: unify pending rows batch acknowledgement policy Signed-off-by: WenyXu <wenymedia@gmail.com> * fix: align logical batcher example configuration expectations Signed-off-by: WenyXu <wenymedia@gmail.com> * fix: align batcher worker channel defaults to 65536 Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> * perf(mito2): lazily decode dense primary key columns (#9226) * perf(mito2): lazily decode dense primary key columns Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * perf(mito2): bypass lazy decoding for full primary keys Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * fix(mito-codec): preserve prefix decoding errors Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * refactor(mito-codec): align encoded length helper naming Rename encoded_length to encoded_len and update all callers to match the other length helpers in the module. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * refactor(mito2): clarify conditional dense key decoding Rename decode_dense_pk to ensure_dense_pk_decoded so callers can see that existing decoded values are preserved and only missing caches are populated. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * refactor(mito-codec): share string framing in row converter Move encoded_string_len to the parent module so Dense and Sparse use the same framing helper without depending on each other. Preserve its implementation and visibility. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> --------- Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * chore: sync lock * fix: shear and check issues --------- Signed-off-by: Ning Sun <sunning@greptime.com> Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> Signed-off-by: houyuwushang <180804215+houyuwushang@users.noreply.github.com> Signed-off-by: Dennis Zhuang <killme2008@gmail.com> Signed-off-by: fys <fengys1996@gmail.com> Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> Signed-off-by: WenyXu <wenymedia@gmail.com> Co-authored-by: Dhruv Vaishnav <dhruvvaishnav687@gmail.com> Co-authored-by: houyuwushang <180804215+houyuwushang@users.noreply.github.com> Co-authored-by: dennis zhuang <killme2008@gmail.com> Co-authored-by: fys <40801205+fengys1996@users.noreply.github.com> Co-authored-by: Lei, HUANG <6406592+v0y4g3r@users.noreply.github.com> Co-authored-by: Weny Xu <wenymedia@gmail.com> |
||
|
|
953d01ac54 |
feat: support pending rows batching for MySQL and PostgreSQL (#9302)
* feat: support pending rows batching for MySQL and PostgreSQL Signed-off-by: WenyXu <wenymedia@gmail.com> * style: group batcher imports before item definitions Signed-off-by: WenyXu <wenymedia@gmail.com> * fix: use 65536 as the default batcher worker channel capacity Signed-off-by: WenyXu <wenymedia@gmail.com> * test: use a distinct custom worker channel capacity Signed-off-by: WenyXu <wenymedia@gmail.com> * test: complete Prom config in worker capacity override case Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
9281965bac |
feat(flow): allow execution hook to rewrite completed plans (#9290)
* refactor(flow): expose incremental aggregate plan analysis Signed-off-by: discord9 <discord9@163.com> * feat(flow): allow execution hook to rewrite completed plans Signed-off-by: discord9 <discord9@163.com> * test(flow): cover the execution hook receiving the completed plan The hook must see the plan that is dispatched after the incremental delta-sink merge, so the test records the plan a collaborator is handed and asserts the merge is already part of it. Signed-off-by: discord9 <discord9@163.com> --------- Signed-off-by: discord9 <discord9@163.com> |
||
|
|
0f118bc5ba |
fix: serialize struct to json in postgres (#9170)
* feat: serialize struct to json in postgres * fix: support view scalars and preserve null structs in scalar-to-value conversion Address PR review: - Utf8View/BinaryView ScalarValues now convert like their non-view forms instead of failing row extraction for struct columns - a null struct scalar converts to Value::Null so a null struct inside a list stays null in the serialized JSON Signed-off-by: Ning Sun <sunning@greptime.com> * fix: return errors instead of panics for unsupported arrow field types Struct-typed query results with arrow field types greptimedb cannot represent (e.g. Decimal256) used to panic during schema conversion and row extraction, dropping the client connection. They now surface as query errors: - ConcreteDataType::try_from builds struct types fallibly via the new StructType::try_from_arrow_fields - Value::try_from(ScalarValue::Struct) uses the same fallible path - new try_value_from_array converts an arrow element to Value with error propagation, used by the postgres struct encoding Signed-off-by: Ning Sun <sunning@greptime.com> --------- Signed-off-by: Ning Sun <sunning@greptime.com> |
||
|
|
f3eb8e6c72 |
test: cut integration test time and make the storage matrix meaningful (#9308)
* test: cut integration test time and make the storage matrix meaningful
tests-integration is ~85% of workspace test CPU, and 81% of that is the
S3/S3WithCache variants of the HTTP and gRPC suites. Those suites do not
touch the object store: of the 70 matrix HTTP tests only one flushed and
read back an SST, so the matrix was paying real AWS round trips to
re-prove protocol parsing.
- Point the PR CI object-store matrix at the MinIO already started by
tests-integration/fixtures. Three GT_S3_* consumers did not read
GT_S3_ENDPOINT_URL and would have hit real AWS with MinIO credentials;
they now do.
- Add a nightly Linux job against real AWS S3, and pass GT_S3_* into the
release integration-test container. The release previously ran every
remote-backend case as a skip and only exercised the file backend.
- Give each S3WithCache test its own read cache directory. They shared
/tmp/greptimedb_cache, which the datanode wipes on startup, so a
starting test deleted the read cache of a running one.
- Add flush -> read-back assertions to the tests whose columns have a
non-trivial SST representation: JSON/JSON2 columns, native histograms,
metric-engine logical tables, and tables carrying fulltext or skipping
indexes whose puffin files only exist after a flush.
- Move eight tests that create no table out of the storage matrix.
- Make the event recorder flush interval a constructor parameter and
shorten it in the event tests, which otherwise wait a 5s window per DDL
they assert on. It is skipped by serde and never reaches config files.
- Drop duplicates: test_grpc_zstd_compression was a verbatim copy of
test_grpc_message_size_ok and is now rewritten to assert the negotiated
grpc-encoding; test_execute_copy_to_{s3,oss,gcs,azblob} were strict
prefixes of their copy_from siblings; two standalone/distributed event
test pairs shared one assertion body.
- Fix and un-ignore stddev_by_label. stddev_pop merges partial aggregates
in a parallelism-dependent order, so its last digits are unstable; the
test now compares values with a tolerance.
- Rebase the jaeger v1 fixture on the current instant. It carries
ttl=7d with 2025 timestamps, so its rows were only readable as long as
they stayed in the memtable.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* test: address review — wire nightly real-S3 job into check-status, keep the short event interval
The nightly `check-status` job did not depend on the new real-S3 job, so a
failure there would not have reached the status or Slack notification.
In database_ddl_event the short interval was set by a first
`with_event_recorder_options` call and then overwritten by the pre-existing
one, which carries `..Default::default()`. Merged into a single call.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
|
||
|
|
edc81c2355 |
ci: update cargo fuzz command to use nightly toolchain explicitly (#9298)
* ci: update cargo fuzz command to use nightly toolchain explicitly
* ci: honor RUSTUP_TOOLCHAIN pin in fuzz orchestration script
An explicit `+toolchain` argument overrides the RUSTUP_TOOLCHAIN env var
in rustup precedence, so the hard-coded `cargo +nightly` in
run-fuzz-targets.sh bypassed the pinned FUZZ_RUST_TOOLCHAIN
(nightly-2026-03-21) configured in the workflow.
- Invoke `cargo +"${RUSTUP_TOOLCHAIN:-nightly}" fuzz run` in the script
so CI uses the pinned toolchain and local runs fall back to the
floating nightly
- Pass RUSTUP_TOOLCHAIN through to all four fuzz-test action invocations,
covering the no-prebuilt-binaries path and making the reproduce command
in the summary print the exact pinned toolchain
- Add test_rustup_toolchain_env_is_honored covering the pinned-env
scenario for both the cargo invocation args and the summary text
Addresses #9298 (review).
Signed-off-by: Ning Sun <sunning@greptime.com>
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
|
||
|
|
e91faa9df8 |
perf(mito2): lazily decode dense primary key columns (#9226)
* perf(mito2): lazily decode dense primary key columns Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * perf(mito2): bypass lazy decoding for full primary keys Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * fix(mito-codec): preserve prefix decoding errors Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * refactor(mito-codec): align encoded length helper naming Rename encoded_length to encoded_len and update all callers to match the other length helpers in the module. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * refactor(mito2): clarify conditional dense key decoding Rename decode_dense_pk to ensure_dense_pk_decoded so callers can see that existing decoded values are preserved and only missing caches are populated. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * refactor(mito-codec): share string framing in row converter Move encoded_string_len to the parent module so Dense and Sparse use the same framing helper without depending on each other. Preserve its implementation and visibility. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> --------- Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> |
||
|
|
723da69b21 |
feat: share logical table batching with OTLP metrics (#9288)
* feat: share logical table batching with OTLP metrics Signed-off-by: WenyXu <wenymedia@gmail.com> * fix: unify pending rows batch acknowledgement policy Signed-off-by: WenyXu <wenymedia@gmail.com> * fix: align logical batcher example configuration expectations Signed-off-by: WenyXu <wenymedia@gmail.com> * fix: align batcher worker channel defaults to 65536 Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
d0f8f4b80c |
feat: add AI matching, classification, and scoring functions (#9300)
* feat: return matching scores from jev Replace the experimental three-argument Boolean function with jev(text, prompt) returning a Float64 probability in [0, 1]. Move threshold comparisons into SQL and update tests and migration examples. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * feat: add Jev choice and score functions Share asynchronous execution across Noul, Choice, and Score. Validate JSON criteria before requests and return typed scalar answers. Add SQL and HTTP mock coverage with usage examples. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * refactor: use generic AI SQL function names Expose ai_match, ai_choose, and ai_score and move their implementation, tests, and usage guide under generic AI names. Document the current unreleased interface without migration history. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * fix: share constant AI criteria within each batch Borrow scalar string arguments and lazily parse constant criteria once per batch. Share the parsed allocation across requests while preserving NULL propagation and batch validation before HTTP calls. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * feat: preserve AI score uncertainty in JSONB results Return score, confidence, and probabilities in criteria-level order from one evaluation. Validate the distribution and preserve provider precision. Add JSON extraction, uncertainty, and single-request regressions, and document confidence-aware ranking. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * docs: explain reuse of volatile AI evaluations Document repeated SELECT and WHERE evaluation costs as N + M requests, and show subquery aliases for reusing scalar or structured AI results without additional model calls. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> --------- Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> |
||
|
|
a9e2a89b7b |
fix: keep compaction pruning, metadata, and index work on compact runtime (#9304)
* fix: run compaction pruner tasks on compact runtime Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * fix: keep compaction metadata and index work on compact runtime Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> --------- Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> |
||
|
|
045441e3cc |
feat(json2): support altering JSON2 settings (#9029)
* feat(sql): support alter syntax for JSON2 columns Signed-off-by: fys <fengys1996@gmail.com> * fix(json2): preserve rows on type hint mismatch during compaction * refactor(json2): simplify alter settings handling * fix(json2): preserve coerced values during compaction * chore: remove unnecessary clone * chor: reduce memory allocations * fix: cargo clippy * chore: update greptime-proto to main branch * refactor(datatypes): unify string handling with other JSON type hints * fix: cr --------- Signed-off-by: fys <fengys1996@gmail.com> |
||
|
|
aa36f74feb |
fix(ci): build tests-integration lib with meta-srv/mock (#9299)
* fix(ci): build tests-integration lib with meta-srv/mock tests-integration's lib code (src/cluster.rs) uses meta_srv::mocks, but the dependency carrying the mock feature sits in [dev-dependencies]. Builds that only touch the lib, such as the apidoc job's cargo doc --workspace, resolve meta-srv without mock and fail with E0432. --all-targets builds unify dev-dependency features, which is why check, clippy and nextest stayed green. Move the mock-enabled meta-srv entry back to [dependencies]. The other testing features moved out in #9072 are not needed by the lib and stay in [dev-dependencies]. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test(repartition): split per-case repartition tests test_repartition_metric ran four format/primary-key-encoding cases in a single test function, and test_repartition_mito ran two format cases. Each case builds its own 3-datanode cluster and runs a full repartition plus GC cycle, so on S3 the metric test took 165-178s against the 180s nextest terminate-after. Merge queue runs failed on it at random. Split each case into its own test. Cases were already independent, so they now run in parallel and each stays far inside the timeout, and a failure points at one encoding instead of four. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> --------- Signed-off-by: Dennis Zhuang <killme2008@gmail.com> |
||
|
|
25bca4609f |
perf(promql): push label filters into grouped join inputs (#9280)
* perf(promql): propagate matching filters through grouped joins Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * perf(promql): check matcher safety on the receiving operand Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * refactor(promql): spell out the shapes a filter may cross `preserves_filter` ended in `_ => true`, which was only sound because `selector_matchers` independently rejects label rewriting, `count_values`, subqueries and non-rollup calls on the same operand. Loosening the latter alone would have silently pushed a matcher below a label rewrite. List the shapes that carry a scan filter instead and default to `false`. Cite #9207 for the result labels the grouped cases record: the join projects the right operand's tag set, so `zone` is missing wherever the right side aggregates it away. No behavior change. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test(promql): assert the new pushdowns reach the scan The grouped-join unit tests feed tag columns by hand and the SQLness case only checks results, which are identical whether or not the rewrite fires. Nothing would have failed if scalar arithmetic, ranking or grouped matching stopped propagating. Assert through the planner that the matcher reaches both scans, with a global topk one-side as the counter-example. Also state that the duplicate-one-side cases record a cross product Prometheus rejects (#9209), so the baseline is not read as intended semantics. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> --------- Signed-off-by: Dennis Zhuang <killme2008@gmail.com> |
||
|
|
2388257c35 |
ci: create docs follow-up issue on PR merge instead of on label (#9237)
* ci: create docs follow-up issue on PR merge instead of on label The docbot workflow previously created a docs-repo issue as soon as the 'docs-required' condition was detected (PR opened/edited with the docs checkbox ticked), even if the PR was never merged. Now the workflow also triggers on PR 'closed': - opened/edited: only manage the docs-required/docs-not-required labels - closed: create the docs issue only when the PR was actually merged and carries the docs-required label This also lets maintainers control issue creation by manually adding or removing the docs-required label before merging. Signed-off-by: Ning Sun <sunning@greptime.com> * fix: address review comments on docs issue creation timing - Only touch docs labels when the docs checkbox state actually changed in an edit. Previously, editing any other part of the PR body while the checkbox stayed checked removed the docs-required label, silently dropping the docs follow-up now that issue creation happens at merge. Unchanged checkbox now leaves labels untouched, which also preserves manual label overrides. - Do not trust the closed event's stale label snapshot at merge time: re-read the live PR via the API and create the docs issue if the docs-required label is present OR the checkbox is ticked in the current body. - Make the workflow concurrency group action-aware so a merge run does not cancel an in-flight label update from an edit run. Signed-off-by: Ning Sun <sunning@greptime.com> * fix: make docs-required label the single source of truth at merge The label-OR-checkbox merge condition could not distinguish an intentional opt-out from an unfinished label update: removing docs-required while the checkbox stayed checked still produced an issue, and unchecking the box could still produce one if the merge read the stale label before the edit run removed it. At merge time, wait for any pending docbot runs on the PR head SHA to finish their label updates (bounded to 5 minutes), then decide solely by the live docs-required label. Adds actions: read permission for listing workflow runs. Signed-off-by: Ning Sun <sunning@greptime.com> --------- Signed-off-by: Ning Sun <sunning@greptime.com> |
||
|
|
614b46e5f5 |
fix(postgres): return empty responses for comment-only SQL (#9295)
fix(postgres): handle parsed empty queries in both protocols Signed-off-by: houyuwushang <180804215+houyuwushang@users.noreply.github.com> |
||
|
|
f1e9a74f00 |
fix(meta): populate physical metric table column ids (#9286)
* fix(meta): populate physical metric table column ids Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> * test(meta): verify physical metric column ids Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> --------- Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> |
||
|
|
9dfe057199 |
test: make export chunk deletion failure deterministic (#9291)
* test: make export chunk deletion failure deterministic Signed-off-by: jeremyhi <fengjiachun@gmail.com> * test: simplify export chunk deletion failure fixture Signed-off-by: jeremyhi <fengjiachun@gmail.com> * docs: guide deterministic storage failure tests Signed-off-by: jeremyhi <fengjiachun@gmail.com> --------- Signed-off-by: jeremyhi <fengjiachun@gmail.com> |
||
|
|
67497f14e5 |
refactor: isolate logical table preparation and reuse Flow notifications (#9212)
refactor: isolate logical table preparation and share Flow notifications Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
e65f10c8ad |
fix(query): stop encoding oversized dynamic filters (#9267)
* fix(query): stop encoding oversized dynamic filters Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test(query): stop bounded encoding partway through an IN list Signed-off-by: Dennis Zhuang <killme2008@gmail.com> --------- Signed-off-by: Dennis Zhuang <killme2008@gmail.com> |
||
|
|
2a63ea2660 |
feat(log-store): add object store WAL reads and obsolete watermarks (#9248)
* feat(log-store): add object store WAL reads and obsolete watermarks Signed-off-by: jeremyhi <fengjiachun@gmail.com> * perf(log-store): reuse decoded WAL payload allocations Signed-off-by: jeremyhi <fengjiachun@gmail.com> * refactor(log-store): clarify WAL region mismatch error Signed-off-by: jeremyhi <fengjiachun@gmail.com> --------- Signed-off-by: jeremyhi <fengjiachun@gmail.com> |
||
|
|
b5199bc59a |
feat: add experimental Jev SQL filtering (#9265)
* feat: add experimental Jev SQL filtering Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * feat: gate Jev filtering behind an opt-in Cargo feature Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * ci: verify Jev registration with default features Run the existing registry regression without the jev feature in both PR tests and merge-queue coverage, alongside the existing feature-enabled test runs. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * docs: clarify Jev concurrency scope and stabilization work Document the per-expression/batch concurrency bound and track process-wide limiting, rate-limit backoff, and request budgets and metrics as stabilization prerequisites. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * feat: enable renamed ai-functions feature by default Rename the Jev Cargo feature across the command, query, and function crates and enable it in their defaults. Update CI and documentation, retaining an isolated no-default-features registry check and the runtime API opt-in. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * ci: remove extra AI feature-off checks Use the regular AI-enabled unit and coverage runs for the default feature configuration. Keep feature-off validation available locally and update the usage guide to match. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * refactor: rename AI feature to ai_functions Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> --------- Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> |
||
|
|
16978cf6c2 |
feat(trace): support Semantic Graph for Trace V2 (follow-up to #9192) (#9278)
* feat(trace): support Semantic Graph for Trace V2 Signed-off-by: luofucong <luofc@foxmail.com> * fix: remove unused annotation context import Signed-off-by: luofucong <luofc@foxmail.com> --------- Signed-off-by: luofucong <luofc@foxmail.com> |
||
|
|
4a980f350f |
docs: make packed snapshots part of the Metric export/import RFC (#9247)
* docs: revise Metric snapshots around packed Parquet objects Signed-off-by: jeremyhi <fengjiachun@gmail.com> * docs: summarize experiment conclusions and bound multipart fallback Signed-off-by: jeremyhi <fengjiachun@gmail.com> * docs: specify packed snapshot persistence contract Signed-off-by: jeremyhi <fengjiachun@gmail.com> --------- Signed-off-by: jeremyhi <fengjiachun@gmail.com> |
||
|
|
19c127d9c3 |
feat: restore packed metric snapshots (#9250)
* test: cover snapshot parquet restore compatibility Signed-off-by: jeremyhi <fengjiachun@gmail.com> * feat: restore packed metric snapshots Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix: stream large packed parquet entries Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix: allow default S3 endpoint in packed restore test Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix: skip unconfigured S3 in packed restore test Signed-off-by: jeremyhi <fengjiachun@gmail.com> * test: use portable file URLs in packed restore fixtures Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix: validate packed snapshot structure before restore Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix: preserve strict manifest decoding and verify fixture Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix: reject packed layout on both database export paths Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix: implement file size in coordinator test storage Signed-off-by: jeremyhi <fengjiachun@gmail.com> * test: use expect_err for rejected export layouts Signed-off-by: jeremyhi <fengjiachun@gmail.com> --------- Signed-off-by: jeremyhi <fengjiachun@gmail.com> |
||
|
|
edbcb8224e |
fix: fail startup on duplicate region engine configs (#9281)
* fix: reject duplicate region engine configurations Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * refactor: reuse canonical engine names in config validation Replace duplicated engine-name literals with the existing common-catalog constants so duplicate-config validation uses the shared engine names. Keep the TOML regression inputs independent to verify the public configuration tags. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * fix: clarify zero-based duplicate engine config indices State explicitly that duplicate region engine configuration indices are zero-based so users can map them to the order of TOML entries. Preserve the existing index values and error classification. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> --------- Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> |
||
|
|
263e229103 |
feat: add experimental Metric export to V2 snapshots (#9233)
* feat: add experimental Metric export to V2 snapshots Signed-off-by: jeremyhi <fengjiachun@gmail.com> * fix: validate the complete Metric export capability response Signed-off-by: jeremyhi <fengjiachun@gmail.com> * test: construct portable file URLs for Metric export fixtures Signed-off-by: jeremyhi <fengjiachun@gmail.com> * refactor: address Metric export review nits Signed-off-by: jeremyhi <fengjiachun@gmail.com> --------- Signed-off-by: jeremyhi <fengjiachun@gmail.com> |