Commit Graph
5169 Commits
Author SHA1 Message Date
Lei, HUANG 5ccfcd4644 feat: support automatic column addition for Flight bulk inserts (#9285)
* fix: auto-add columns when initializing bulk insert streams

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: reject nested columns in bulk schema auto-add

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: reject unknown columns in non-empty bulk inserts

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-09-29 06:47:06 +00:00
DeviousCardiandClaude Opus 5.5 d6474b961c fix(promql): apply offset to subquery evaluation window (#9364)
* fix(promql): apply offset to subquery evaluation window

`prom_subquery_expr_to_plan` destructured `SubqueryExpr` without reading
`offset`, so `<subquery>[range:step] offset <d>` planned exactly the same
window as the un-offset form and silently returned data for the wrong time
range. The plain vector/matrix-selector paths already threaded the offset
through `selector_to_series_normalize_plan` and `RangeManipulate`.

Shift the inner evaluation window back by the offset and pass the offset to
the subquery's `RangeManipulate`, which maps the inner samples forward onto
the evaluation timeline before bucketing them into ranges. This matches
Prometheus, whose `evaluator.subqueryTimeRange` evaluates the inner
expression over `(start - offset - range, end - offset]` and whose
`evalSubquery` then hands the samples to the outer range-vector function as
a `MatrixSelector` that still carries the subquery offset. An offset on the
inner selector composes additively, as `subqueryTimes` documents.

`RangeManipulate`'s protobuf message has no offset field and recovers it on
decode from an immediately underlying `SeriesNormalize`. Since
`RangeManipulate` is commutative in `dist_plan` and can be pushed below a
`MergeScan`, insert that carrier node so the offset survives distributed
planning instead of decoding as zero.

Known divergence, unchanged by this commit: Prometheus anchors subquery step
points on absolute epoch multiples of the step, while GreptimeDB anchors them
on the evaluation start. The two agree whenever the offset is a multiple of
the subquery step; the added sqlness cases stay within that range.

Closes #9330

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Aarav <aaravsjadav@gmail.com>

* docs(promql): correct the subquery-offset rationale and pin the histogram shift

Follow-up to e132c267, which stated two things inaccurately and left one
path untested. No behaviour change beyond a conditional clone.

The `SeriesNormalize` carrier was described as existing "so the offset
survives distributed planning", implying standalone is unaffected. That is
wrong. `local_offset` is read on the substrait decode path
(`range_manipulate.rs`, `instant_manipulate.rs`), and GreptimeDB routes
PromQL plans through `MergeScan`/substrait in standalone too -- the
standalone results `promql/encode_substrait.result` and `precisions.result`
both show `MergeScan [is_placeholder=false, remote_input=[...]]`. Deleting
the carrier therefore empties subquery-with-offset results in standalone as
well, while `count_over_time_subquery_with_offset` keeps passing because it
asserts the pre-serialization plan. Say so at the node, so the next reader
does not remove it believing standalone is safe.

The carrier is also not semantically inert: `SeriesNormalizeStream::normalize`
shifts native histogram `start_timestamp` payloads by `offset`. On the
subquery path that runs on a computed inner result, on top of any offset the
inner selector already applied, and nothing pinned it. It is consistent --
`RangeManipulate` moves the millisecond time index forward by the same
`offset_ms` -- so the distance from a histogram's start timestamp to the
sample carrying it, which reset/rate detection reads, is invariant. Add
`subquery_offset_shifts_histogram_start_and_time_index_together`, which
executes the subquery's node stack over native histogram samples for a
zero and a non-zero inner offset. It cannot be a sqlness case: native
histograms are a struct column with no SQL type or literal
(`sql_data_type_to_concrete_data_type` rejects structs) and the sqlness
runner speaks only MySQL/Postgres, so such rows only arrive over
gRPC/remote-write v2.

The divergence note blamed sub-step offsets. Measured, that condition is too
narrow: an unaligned evaluation timestamp alone diverges, with no offset at
all -- `sum_over_time(fine[20s:10s])` at t=57 samples 47s and 57s here
against Prometheus's 40s and 50s. The real condition is that
`(start - offset)` is not a multiple of the subquery step, and it is
pre-existing: the committed `tql eval (359, 359, '1s')
sum_over_time(metric_total[60s:10s])` case is already an instance of it
(`359 - 60 + 10 = 309`). Reword the comment and the sqlness header
accordingly. A sub-step offset stays accepted: `foo[20s:10s] offset 5s` is
valid PromQL, erroring on it would be a regression and would not close the
gap anyway. This repo has no PromQL-compatibility documentation page, so
there is nowhere else to record it.

Also stop cloning the series key columns when there is no offset, and stop
the `offset 0s` sqlness comment implying its parser error is evidence about
this fix -- the same message comes from a `u32` overflow in
`promql-parser`'s shared duration check (`offset 9999999999d`).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Aarav <aaravsjadav@gmail.com>

* refactor(promql): trim subquery-offset comments and tests per review

Shorten the planner and sqlness comments to the semantic contract, build
`SeriesDivide` with a clone instead of an `Option`, and drop the
normalize.rs histogram test and the `offset 0s` parser case.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bvab9gpXYjKyLN8Mo6JWWL
Signed-off-by: Aarav <aaravsjadav@gmail.com>

* fix(promql): record the subquery offset as the range fold offset

The subquery's `RangeManipulate` now shifts the payload timestamps by the
offset, so `predict_linear` over an offset subquery must recover the
evaluation time with that offset instead of 0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bvab9gpXYjKyLN8Mo6JWWL
Signed-off-by: Aarav <aaravsjadav@gmail.com>

---------

Signed-off-by: Aarav <aaravsjadav@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-29 03:38:43 +00:00
Mohd Quamar Tyagi dd2c1d1aca fix(meta-srv): use NoTls for disabled and Unix socket Postgres KV backends (#9059)
* fix(meta-srv): use NoTls for disabled and Unix socket Postgres KV backends

Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com>

* fix(meta-srv): treat Postgres config as unix socket only when every host is a socket

tokio-postgres dials hostaddr over TCP even when host is a socket path,
and a mixed host list with Require/VerifyFull TLS would otherwise end up
sending plaintext over TCP. is_unix_socket_url now parses via
tokio_postgres::Config and requires no hostaddr and all-Unix hosts.
Adds regression cases for the libpq keyword form, the user@ percent
encoded socket URL, and the mixed host / hostaddr negatives.

Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com>

---------

Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com>
2026-09-29 00:07:49 +00:00
jeremyhi 3c0e2a8d55 feat: export Metric snapshots with packed Parquet objects (#9382)
* feat: export Metric snapshots with packed Parquet objects

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: reject empty packed export time ranges

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test: cover cancellation during packed export I/O

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-09-28 15:22:02 +00:00
dennis zhuang 079ffec116 chore: remove dead code left behind by removed features (#9377)
* chore: remove dead code left behind by removed features

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore: remove unused RouteInfoCorrupted error

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-28 12:35:47 +00:00
dennis zhuang 408536af18 chore: remove leftover flow worker config docs and unused flow metrics (#9375)
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-28 12:34:32 +00:00
Ning Sun fedce5c5ec feat: support non-millisecond time index units in the logical batcher (#9346)
* feat: allow customized time index unit for metric engine table

* test: provide query tests

* refactor: revert unnecessary change

* refactor: share timestamp unit conversions in api helper

Address review feedback on the time index unit changeset:

- Add shared timestamp_unit/timestamp_datatype helpers to api::helper
  (the only crate that sees both proto ColumnDataType and TimeUnit due
  to layering; common-time and datatypes have no greptime-proto dep).
  This removes the ColumnDataType -> TimeUnit match duplicated between
  operator's insert path and the OTLP logs path.
- Collapse the two TimeUnit <-> ValueData matches in
  convert_timestamp_value_data by reusing api::helper::to_grpc_value
  for the construction side.
- Note that convert_rows_time_unit rewrites the schema before the
  values, so an overflow mid-batch leaves the request half-converted;
  harmless because the error aborts the whole insert request.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: align time units per destination table and floor remote-read timestamps

Address review feedback on PR #9236:

- Align each metric insert request to the unit of the table it actually
  targets: an existing logical table keeps its own unit (it may be bound
  to a different physical table than the one selected by the request),
  and only new tables use the selected physical table's unit. The
  previous blanket conversion rewrote valid millisecond samples to the
  selected physical table's unit and the engine rejected them.
  Regression test: writing an existing millisecond logical table and a
  new table in one request that selects a microsecond physical table.
- Remote read now floors narrowing timestamp conversions towards
  negative infinity (div_euclid), consistent with
  Timestamp::convert_to on the ingestion path; arrow's cast truncates
  towards zero and returned -1ms for a stored -1001us. Widening
  (second -> millisecond) keeps the exact arrow cast. Regression test:
  a negative, non-aligned timestamp round-trips as -2ms.

Signed-off-by: Ning Sun <sunning@greptime.com>

* perf: fold time unit alignment into existing table lookups

Address review feedback on PR #9236:

- The per-destination unit alignment no longer runs its own pass of
  table lookups: create_or_alter_tables_on_demand gains an
  align_time_index_unit parameter (metric engine path only) and
  converts each request inside the lookups it already performs —
  existing tables to their own unit, new tables to the selected
  physical table's. Default ingest paths now issue zero additional
  catalog lookups compared to main; the separate alignment pass remains
  only in the opt-in logical batcher pre-gate, next to the eligibility
  check that already looks up the same tables.
- convert_rows_time_unit indexes the time index position directly
  (validate_column_count_match guarantees row widths) instead of
  Optional get_mut; the gate-side alignment validates widths itself.

Signed-off-by: Ning Sun <sunning@greptime.com>

* perf: resolve the batcher time index guard once per write target

All batches of one remote write request share the same write target
(catalog, schema, physical table), so the batcher time index guard now
resolves each distinct target once instead of once per batch.

Signed-off-by: Ning Sun <sunning@greptime.com>

* feat: support non-millisecond time index units in the logical batcher

Make the logical table batcher's bulk encode path unit-aware so physical
metric tables with a non-millisecond time index (e.g. TIMESTAMP(6)) can
use logical batching instead of falling back to the ordinary insert path.

- rows_to_aligned_record_batch builds the time index column in the
  TARGET schema's unit, converting any timestamp encoding via
  Timestamp::convert_to (flooring on narrowing, consistent with the
  ordinary insert path).
- New tables created by the batcher use the selected physical table's
  time index unit (resolved once per submit; a missing physical table
  keeps the millisecond auto-create default).
- columns_taxonomy and the can_batch_metric_rows schema whitelist accept
  any timestamp unit; the prometheus remote write v1/v2 batcher gates
  and the OTLP pre-gate alignment are removed together with
  Inserter::align_metric_row_inserts_time_unit, as the batcher now
  converts internally.

Closes #9342

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: address review comments

* fix: address review issue

* refactor: drop the OTLP pre-gate unit alignment made redundant by the bulk path

The main merge of #9236 (squash) resurrected the OTLP pre-gate
alignment and Inserter::align_metric_row_inserts_time_unit, which this
branch had removed. Drop them again:

- The pre-gate existed because the #9236-era bulk eligibility gate only
  accepted millisecond schemas, so nanosecond-encoded OTLP requests had
  to be converted before the check. This branch makes the bulk path
  unit-aware (the gate accepts all time index units and batch alignment
  converts each request to its destination's unit), so the pre-gate is
  redundant and only added N+1 catalog lookups per batched request —
  the very lookup-count overhead raised in the #9236 review.
- The per-destination unit semantics it implemented remain enforced in
  the two paths that need them: the ordinary insert path
  (create_or_alter_tables_on_demand converts inside its existing table
  lookups) and the batched path (batch alignment resolves each
  destination schema and converts to it).

test_otlp_logical_batcher_alignment (the test the pre-gate originally
fixed) and the mixed-physical-table regression both pass without it.

Signed-off-by: Ning Sun <sunning@greptime.com>

* test: cover OTLP batcher cross-physical fallback and nanosecond physical

Extend the logical batcher integration coverage for the cases
previously guarded by the removed OTLP pre-gate alignment:

- test_otlp_logical_batcher_fallback_for_cross_physical_destination:
  with the batcher enabled, an OTLP request targeting an existing
  logical table bound to another physical table must NOT enter the
  batcher (the bulk eligibility check rejects the destination binding)
  and the ordinary insert path must convert it to the destination's
  unit (60s -> 60_000_000us).
- test_otlp_logical_batcher_non_millisecond_physical_table now covers
  both microsecond and nanosecond physical tables (parameterized),
  asserting batcher submissions and unit-precise stored values.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: validate per-batch physical bindings and avoid intermediate timestamp buffers

Address review feedback on PR #9346:

- accepts_bulk_destinations dedupes on (schema, table, selected physical)
  instead of (schema, table): one request can select different physical
  tables per batch (per-series x_greptime_physical_table labels), and the
  old key let a second selection skip validation and flush rows through
  the wrong physical's regions. Missing tables additionally reject
  conflicting physical selections within the same request. Regression
  test covers an existing destination, a missing destination, and a
  consistent selection (which must still batch).
- The timestamp column builder appends each value directly into the
  target-unit Arrow builder; values already in the target unit (the
  unchanged millisecond fast path) are appended without conversion, so
  the default millisecond physical pays no Timestamp construction or
  intermediate Vec allocation.
- The non-millisecond batching tests assert the submit_build_and_align
  counter, which increments on every batcher submission in both
  acknowledgement modes, so a silent fallback to ordinary insertion
  fails the tests instead of passing on stored values alone.

Signed-off-by: Ning Sun <sunning@greptime.com>

---------

Signed-off-by: Ning Sun <sunning@greptime.com>
2026-09-28 12:27:36 +00:00
discord9 28e415a103 feat(query): implement PromQL @ modifier on vector and matrix selectors (#9224)
* feat(query): implement PromQL @ modifier on vector and matrix selectors

The planner previously dropped the `at` field of selectors (`at: _`),
so `some_metric @ 300` silently returned step-following values instead
of the samples anchored at the fixed timestamp.

Anchoring follows Prometheus's `setOffsetForAtModifier` + `refetch`
semantics: resolve the anchor (`@ <ts>`, `@ start()`, `@ end()`), apply
`offset` to the anchor, rewrite the selector offset to
`eval_start - anchor`, scan only the anchored window, then report the
same window at every evaluation step via a grid-wide-replay
`InstantManipulate`. `@ start()` / `@ end()` resolve against the
statement's evaluation range (`stmt_start` / `stmt_end`), which subquery
planning does not rewrite.

Timestamps before the Unix epoch are accepted; unrepresentable ones are
rejected with `AtModifierTimestampOutOfRange` instead of wrapping.
Selectors without `@` are planned exactly as before.

Report: .e-agent/greptimedb_promql_compatibility_report_2026-09-16.md P0-2
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(promql): center predict_linear on the evaluation timestamp

predict_linear_impl used the window's last sample time as the evaluation
timestamp, and the UDF took only (ts_range, value_range, t) with no channel
for the current instant. After a range selector is folded for @ (or with an
offset) the same window is replayed at every step, so the prediction stayed
constant at the anchor's answer; even a plain window ended before the step
produced a stale value.

Give prom_predict_linear a 4th argument carrying the step's evaluation
instant (ms timestamp), derived in the planner from the row's time index
plus the offset the window was folded with (at_offset for @, offset_ms
otherwise, recorded on PromPlannerContext). The regression is centered on
that instant, matching Prometheus' use of enh.Ts. The mixed
float/native-histogram path forwards the same argument.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(query): tighten @ modifier planner and predict_linear

- range_fold_offset is always a concrete millisecond offset, not an
  Option; the single-step replay helper no longer wraps an infallible
  plan in Result.
- predict_linear's eval timestamp is always cast to Timestamp(ms) at the
  call site, so the UDF drops its dead Int64 branch and the extra
  func_name argument.
- Drop two redundant comparison cases from the at_modifier sqlness test
  and fix the lookback window comment to the half-open (0s, 300s].

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(query): accept parse-time rejection of @ on Windows

`@ 1e16` is 10^19 milliseconds, beyond i64::MAX. A Unix SystemTime
holds it and the planner rejects the anchor it cannot represent, but a
Windows SystemTime tops out near 1.8e12 seconds, so the parser's
checked_add fails first and the same literal is rejected while parsing.
Accept either rejection path so the test passes on both platforms.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(query): avoid replaying label_join over rewritten series keys

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(query): narrow anchored range call promotion

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(query): address @ modifier review feedback

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(query): satisfy super import format check

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* docs(query): address at modifier review follow-ups

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-28 09:38:12 +00:00
jeremyhi 03823a9a01 feat(log-store): add the enqueued acknowledgement mode to the object store WAL (#9358)
* feat(log-store): add the enqueued acknowledgement mode to the object store WAL

Add `ack_mode` (`durable` by default, or `enqueued`) and the backlog
thresholds `max_unpersisted_bytes` and `max_unpersisted_age` to the
object store WAL config, validated by the datanode and the store.

In the `enqueued` mode `append_batch` returns on admission with the
entry ids assigned and the object is created in the background. At a
backlog threshold the next append is held back until an upload
completes. A transient create failure is repeated under the same
sequence with the same bytes; any conflicting object poisons the
store. `stop` uploads the backlog, or returns the error that dropped
it once stop began. `obsolete` clamps the watermark to the durable
entry id, and an id the store handed out needs no sequence floor.

Add `LogStore::wait_durable` with a default that returns at once. The
object store WAL answers it once the region is durable and indexed
through the entry id, and fails it after a backlog was lost.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): poison an enqueued store on permanent create failures

Repeat a failed create in the enqueued mode only when the storage error
is retryable; any other storage error poisons the store. A conflicting
object poisons an enqueued store without reading its epoch, so a failed
header read cannot turn the conflict into a retry.

A durability wait for an id above the highest id the store handed out
now waits for the handed-out ids of the region below it instead of
returning at once.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): poison on permanent create failures after stop begins

A create that fails with a storage error that is not retryable poisons
an enqueued store even after stop began; only a transient failure drops
the backlog without poisoning. Durability waiters whose callers stopped
waiting are pruned before a new waiter is queued. The backlog age test
no longer depends on a follow-up append finishing within the threshold.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): answer a durability wait once no earlier entry is pending

A durability wait now returns once the region holds no entry at or
below the target that is handed out but not durable, instead of waiting
for the largest id handed out to the region. A later object of the
region that is still being created no longer holds back a wait whose
target it does not cover.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test(log-store): order the pending durability wait check after the actor

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* docs(store-api): state that the default wait_durable keeps each store's guarantee

The default `LogStore::wait_durable` returns at once, which keeps each
log store's own acknowledgement guarantee; Raft Engine with
`sync_write = false` acknowledges before its periodic sync, so the
documentation no longer claims that every entry id a caller holds is
durable. The object store WAL configuration test now also serializes
the new options and reads them back.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* docs(log-store): limit the acknowledgement guarantees to the durable mode

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): repeat enqueued creates that a retry layer marks persistent

An object store wrapped in the OpenDAL retry layer reports a temporary
error that outlasted its retries as persistent rather than temporary.
The enqueued mode now repeats a create after any storage error that is
not permanent, so a transient outage behind the retry layer no longer
poisons the store and drops the acknowledged backlog; after stop began
such a failure still drops the backlog without poisoning.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test(log-store): cover a persistent create failure after stop begins

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-09-28 09:17:45 +00:00
75bd8e9ce6 feat: add HDFS object storage backend (#8701)
* feat: add HDFS object storage backend

Signed-off-by: Minghan2005 <cambrianocean@gmail.com>

* fix: make HDFS storage operations durable

Gate the native HDFS backend behind an explicit feature. Publish writes through same-directory temporary files and atomic HDFS Rename2 replacement, and provide streaming copy fallback for COPY_REGION. Add regression coverage for interrupted writes and the region-copy path.

Signed-off-by: Minghan2005 <cambrianocean@gmail.com>

* ci: run HDFS object store tests

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* docs: note HDFS temporary file cleanup follow-up

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* feat: enable HDFS object storage by default

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* docs: remove redundant HDFS build feature notes

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: Minghan2005 <cambrianocean@gmail.com>
Signed-off-by: jeremyhi <fengjiachun@gmail.com>
Co-authored-by: Minghan2005 <cambrianocean@gmail.com>
Co-authored-by: jeremyhi <fengjiachun@gmail.com>
2026-09-28 08:39:48 +00:00
dennis zhuang ffdd6d09a6 perf(index): cut allocations when building bloom and inverted indexes (#9359)
* perf(index): build bloom filters from element hashes without per-token allocation

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* perf(index): speed up inverted index building with hashed buffers and sync pushes

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore(index): use BuildHasher::hash_one for element hashes

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore(index): require callers to act on spill requests

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(index): insert bloom hashes one by one to keep segment set capacity bounded

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(index): add index build and bloom search benchmarks

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(index): keep applier setup out of the bloom search benchmark

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(index): skip empty spills and keep the old inverted sort memory estimate

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(index): stream fulltext token hashes and test spill dispatch

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* docs(index): note non-ASCII tokens still allocate in analyze_text_hashes

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* docs(index): narrow the allocation note to case-insensitive non-ASCII tokens

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-28 08:22:42 +00:00
dennis zhuang bc5ad4b416 perf(mito2): batch index page loads and share cached bloom metadata (#9360)
* perf(mito2): load missing index pages of all ranges in one read

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor(mito2): drop unused IndexCacheMetrics::merge

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* perf(mito2): share cached bloom filter metadata instead of cloning it

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(mito2): only merge adjacent ranges when batch reading index pages

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(mito2): skip empty ranges when batch loading index pages

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* perf(mito2): fetch index ranges concurrently

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(index): return Arc<BloomFilterMeta> from RecordingReader

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-28 07:57:04 +00:00
Ning Sun 55b7e08a1a feat: allow customized time index unit for metric engine table (#9236)
* feat: allow customized time index unit for metric engine table

* test: provide query tests

* refactor: revert unnecessary change

* refactor: share timestamp unit conversions in api helper

Address review feedback on the time index unit changeset:

- Add shared timestamp_unit/timestamp_datatype helpers to api::helper
  (the only crate that sees both proto ColumnDataType and TimeUnit due
  to layering; common-time and datatypes have no greptime-proto dep).
  This removes the ColumnDataType -> TimeUnit match duplicated between
  operator's insert path and the OTLP logs path.
- Collapse the two TimeUnit <-> ValueData matches in
  convert_timestamp_value_data by reusing api::helper::to_grpc_value
  for the construction side.
- Note that convert_rows_time_unit rewrites the schema before the
  values, so an overflow mid-batch leaves the request half-converted;
  harmless because the error aborts the whole insert request.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: align time units per destination table and floor remote-read timestamps

Address review feedback on PR #9236:

- Align each metric insert request to the unit of the table it actually
  targets: an existing logical table keeps its own unit (it may be bound
  to a different physical table than the one selected by the request),
  and only new tables use the selected physical table's unit. The
  previous blanket conversion rewrote valid millisecond samples to the
  selected physical table's unit and the engine rejected them.
  Regression test: writing an existing millisecond logical table and a
  new table in one request that selects a microsecond physical table.
- Remote read now floors narrowing timestamp conversions towards
  negative infinity (div_euclid), consistent with
  Timestamp::convert_to on the ingestion path; arrow's cast truncates
  towards zero and returned -1ms for a stored -1001us. Widening
  (second -> millisecond) keeps the exact arrow cast. Regression test:
  a negative, non-aligned timestamp round-trips as -2ms.

Signed-off-by: Ning Sun <sunning@greptime.com>

* perf: fold time unit alignment into existing table lookups

Address review feedback on PR #9236:

- The per-destination unit alignment no longer runs its own pass of
  table lookups: create_or_alter_tables_on_demand gains an
  align_time_index_unit parameter (metric engine path only) and
  converts each request inside the lookups it already performs —
  existing tables to their own unit, new tables to the selected
  physical table's. Default ingest paths now issue zero additional
  catalog lookups compared to main; the separate alignment pass remains
  only in the opt-in logical batcher pre-gate, next to the eligibility
  check that already looks up the same tables.
- convert_rows_time_unit indexes the time index position directly
  (validate_column_count_match guarantees row widths) instead of
  Optional get_mut; the gate-side alignment validates widths itself.

Signed-off-by: Ning Sun <sunning@greptime.com>

* perf: resolve the batcher time index guard once per write target

All batches of one remote write request share the same write target
(catalog, schema, physical table), so the batcher time index guard now
resolves each distinct target once instead of once per batch.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: address review comments

* fix: address review issue

---------

Signed-off-by: Ning Sun <sunning@greptime.com>
2026-09-28 07:15:36 +00:00
dennis zhuang 0a03e64d8c perf(index): batch bloom filter searches across row groups (#9361)
* perf(index): search bloom filters of all row groups in one read

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* perf(index): bound the filter bytes of each batched bloom search

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* perf(index): count batch filter bytes exactly as the search reads them

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore(index): allow single-range groups in the budget test

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(index): cover filters shared across segments in the budget test

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-28 06:22:15 +00:00
discord9 f956f4da30 refactor(promql): split planner into focused modules (#9354)
* refactor(promql): move planner tests into separate module

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(promql): move function-specific planner methods into module

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(promql): move OR planner method into module

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(promql): move binary island planner into module

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(promql): keep binary result labels in planner root

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(promql): limit planner helper visibility to parent module

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(promql): group set operator planning and localize imports

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(promql): use crate-rooted planner imports

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-28 05:30:33 +00:00
discord9 678aa81cae fix(client): complete transport lane isolation (#9030)
* fix(client): complete transport lane isolation

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(client): deprecate legacy single-manager constructors

Per review: mark the legacy single-manager constructors and helper as
deprecated so callers move to the isolated query/control manager pair.
Tests intentionally exercising the legacy path are annotated with
`#[allow(deprecated)]`.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(client): update deprecated constructor callers for CI

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-28 04:39:20 +00:00
DeviousCardiandClaude Opus 5.5 d5a6a9ac05 fix(auth): follow symlink chains in watch_file_user_provider (#9365)
* fix(auth): follow symlink chains in watch_file_user_provider

The shared file watcher only watched the immediate parent directory of the
configured path, so swapping a directory symlink elsewhere in the chain
(e.g. `ln -sfn secrets.2 secrets`, as sops-nix does) never produced an
event and the new users file was not loaded until restart.

The watcher now resolves each path component by component and also watches
the directory holding every symlink it meets, plus the directory of the
final file. After each relevant event the chain is re-resolved and the
watches are re-armed, so later in-place edits of the new target are seen.
Only the final file's directory is required at startup; a link directory
that cannot be watched (e.g. mode 0711) is logged and retried later.

Events are now filtered to the chain's links and final files, so that
watching e.g. /tmp does not reload on unrelated files. Names are compared
case-insensitively on macOS and Windows. TLS cert reload uses the same
helper.

Close #9310

Signed-off-by: Aarav <aaravsjadav@gmail.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bvab9gpXYjKyLN8Mo6JWWL

* fix(common-config): warn when the file watcher watches the root dir

On macOS /tmp, /var and /etc are symlinks under /, so following a symlink
chain there adds / to the watched directories, and the FSEvents backend
then sees every event on the volume. Log a warning when that happens.
Also shorten the comment on skipped link directories.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bvab9gpXYjKyLN8Mo6JWWL
Signed-off-by: Aarav <aaravsjadav@gmail.com>

---------

Signed-off-by: Aarav <aaravsjadav@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 04:33:46 +00:00
discord9 4ffd93d8dc fix(query): preserve global limits with DataFusion optimizer fixes (#9071)
* fix(query): preserve global fetch through physical-optimizer repartitioning

Aggregate soft-limit pushdown inserts a global CoalescePartitionsExec(fetch)
below an AggregateExec whose input is HashPartitioned/KeyPartitioned. The
DataFusion physical optimizer's EnforceDistribution/EnforceSorting passes then
dropped or rewrote that fetched coalesce, losing the global limit. Root-caused
and fixed in the DataFusion fork (GreptimeTeam/datafusion PR #35); this repo
pins to that fix commit and adds regression coverage.

- pin datafusion to GreptimeTeam/datafusion 48510c7f5 (fetch preservation fix)
- global_limit.rs: adapt to new DF API (input_distribution_requirements,
  child_distribution, replace_children+Recompute); accept HashPartitioned and
  KeyPartitioned as partitioning to restore; add regression unit tests
- sqlness: extend standalone common ssts and refresh distributed ssts_limit
  result for the preserved fetch plan

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor: drop dead HashPartitioned compat arms and test wrapper

- remove HashPartitioned compatibility matches (pinned DF only emits
  KeyPartitioned); drop their #[expect(deprecated)] attributes
- remove single-use agg_with_limit helper whose seed limit is always
  overwritten by the soft-limit transform
- restore Cargo.lock dependency edges unrelated to the DataFusion pin
  bump (cargo update --precise had downgraded 26 unrelated packages)

Addresses oracle ablation review.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore(deps): pin DataFusion to merged fetch preservation fix

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-28 03:35:36 +00:00
jeremyhi bf67246645 perf: bound concurrent Metric export writers (#9296)
* fix: preserve objects after conditional Metric export collisions

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* perf: share Metric export writer and payload budgets

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test: cover bounded Metric export conversion and mixed roundtrips

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: preserve export output after ambiguous close

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test: align HTTP export timeout with managed cancellation

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test: trim redundant Metric export roundtrips

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: bound pending Metric export writer handles

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: protect export files on failed writes

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: preserve ordinary export cleanup and large rows

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: report Metric export collisions as invalid arguments

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: clean up failed conditional export closes

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: support view arrays in managed ordinary export

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test: expect successful cleanup after failed close

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: clean up unsynced overwrite after close failure

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: scope export JSON expansion estimates to JSON columns

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: rechunk ordinary export batches within the write budget

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-09-28 03:10:21 +00:00
dennis zhuang 411c0f173e fix(mito2): stop parallel flat scan tasks once the receiver is dropped (#9368)
* fix(mito2): stop parallel flat scan tasks once the receiver is dropped

spawn_flat_scan_task ignored send errors, so after a query was cancelled the
detached task kept reading its source to EOF. Break out of the loop when the
receiver is gone.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore(mito2): log when a parallel scan task stops on a dropped receiver

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore(mito2): include the send error in the parallel scan stop log

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-28 02:50:35 +00:00
Weny Xu 5ffd01a70a fix: preserve primary key order when syncing columns (#9189)
* fix: preserve primary key order when syncing columns

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: reject invalid SyncColumns metadata

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-09-27 02:15:47 +00:00
jeremyhi 194bc2fb3c feat(log-store): add the object store WAL durable write path (#9320)
* feat(log-store): add the object store WAL durable write path

ObjectStoreLogStore can now write. append_batch admits entries into one
open batch and assigns object-sequence-major entry ids at admission. The
batch is sealed by size, by the flush interval, or before a region would
run past the position range, and sealed batches are uploaded with at most
four conditional creates in flight, started in sequence order. Created
objects are indexed in sequence order, and an append is acknowledged only
once its object is durable and indexed.

A transient create failure rolls back the failed batch and every later
batch unless a later object is already durable, in which case the store
poisons itself with a history-gap error. A conflicting object, an
encoding or catalog error, or taking the last representable sequence
poisons the store.

obsolete now goes through the actor and raises the sequence floor
together with the watermark, refusing with a retryable error while the
next sequence is not settled. stop drops the open batch and the batches
whose create has not started, and lets creates in flight finish.

Only the durable acknowledgement mode exists. A testing feature exposes
hooks to wait for admissions, seal the open batch, and hold or fail
creates.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test(log-store): synchronize object store WAL write tests with the actor

Tests that assert nothing happened round-trip a command through the actor
instead of yielding the test task, the conflict test waits for the object
it expects, and the obsolete-behind-stop test holds the command channel
itself so the unanswered command is deterministic.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): bound admitted WAL appends and never reuse a failed sequence

A create that reports an error may still have written its object, so the
sequences of failed batches are no longer handed out again: the next
batch keeps the sequence after the last sealed one and retries are
assigned new ids. A retry batched differently can no longer conflict
with that object. As the next sequence never moves back, the sequence
floor of obsolete only waits for an open batch that has handed out ids.

Appends now arrive on their own bounded channel, which the actor stops
reading while MAX_SEALED_BATCHES batches wait to become durable, so a
stalled object store holds callers back instead of growing the backlog,
while stop and obsolete are still handled.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): reserve room for two seals before admitting a WAL append

One admission can seal the open batch before an append that would exhaust
its positions and then the append's own batch, so the actor takes an
append only while two more sealed batches fit under MAX_SEALED_BATCHES.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* feat(log-store): chain object store WAL objects and recover the latest chain

A conditional create can fail with an unknown outcome while its object is
stored, or still lands later, and Mito reuses the row sequences of a
failed append. Replaying such an object next to a later acknowledged one
can let the unacknowledged rows win after a restart.

Format version 2 gives every object header its writer's epoch, a link
to the object it extends (sequence and writer instance) and a header
CRC32, and allows objects without segments. Recovery replays only the
chain ending at the complete object with the largest epoch and sequence:
a link holds when its predecessor is present with the recorded writer
instance, or is missing below every present object. Objects off the
chain are orphans that are never replayed but keep their sequences.

Each open writes an empty object that starts an epoch above every
present object, linked to the recovered tip, before it accepts writes,
so a late object of an earlier instance never ends the chain. A start
object that meets an object of an earlier epoch moves to the next
sequence; one of an equal or later epoch fails the open.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): break the WAL chain at every missing predecessor

Accepting a missing predecessor below every present object lets a late
object that lands below the chain change which links hold. Nothing
collects objects yet, so a missing predecessor now always breaks the
link, and recovery fails when objects are present but none completes a
chain.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): keep the object store WAL format at version 1

The object store WAL has not been enabled anywhere, so no object in the
previous layout exists and the chained header can stay version 1.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* refactor(log-store): link object store WAL objects by epoch

Only one store instance writes under an epoch, since every open starts an
epoch above every present object and a start object that meets the same
or a later epoch fails the open. The epoch therefore identifies the
instance, and the random writer instance id is dropped from the header.
A link now records the sequence and the epoch of the object it extends,
and holds when the predecessor carries that epoch. The header shrinks to
46 bytes, and the store logs its epoch when it opens.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): fail an open whose start object is already present

Without a random writer instance, two opens that recover the same objects
encode byte-identical start objects, and a conditional create treats the
same bytes as its own retry. A start object that is already present
therefore fails the open with a retryable error instead of letting both
opens claim the epoch; the next open counts the object and starts a
later epoch.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): derive the WAL epoch from the claimed start sequence

Two opens can recover different listings when a late object lands across
a sequence gap between them, pick the same largest epoch plus one, and
both create their start objects under different sequences. The epoch of
an instance is now one above the sequence its start object claims, so a
successful create decides the epoch and no two instances share one. It
stays above every epoch recovery listed, and an object that carries an
epoch above the next sequence fails the open.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): never run a held WAL create after the store is dropped

The create test hook ignored the closed hold channel, so a create parked
when the store was dropped could still run. It now returns without
creating. Drop the per-admission bookkeeping of issued entry ids, which
nothing reads, and move the parked I/O documentation to its helper.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test(log-store): cover a WAL start object stored with an unknown outcome

Add a fault-injection test in which the create of the start object
stores the object but reports an error: the open fails without moving
to another sequence, and the next open counts the stored object and
claims a later epoch. Rename the test helper that writes a whole object
from a given header to put_object_with_header, and drop a needless
clone.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test(log-store): make the held-create drop test deterministic

Keep the actor running while the store drops the hold sender, so the
parked create always completes on the closed channel instead of racing
the actor's exit. Drop a comment that restates epoch_of.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-09-24 11:07:29 +00:00
Yingwen 3c2aac0a55 feat: add manual series index reconciliation (#9323)
* feat(mito): support manual series index reconciliation

Signed-off-by: evenyag <realevenyag@gmail.com>

* feat(storage): route series index build requests

Signed-off-by: evenyag <realevenyag@gmail.com>

* feat(admin): add BUILD_SERIES_INDEX

Signed-off-by: evenyag <realevenyag@gmail.com>

* test: specify compaction type in series index fixtures

Signed-off-by: evenyag <realevenyag@gmail.com>

* test: correct series index SQL fixtures and error assertions

Signed-off-by: evenyag <realevenyag@gmail.com>

* test(compat): preserve legacy index rebuild across upgrades

Signed-off-by: evenyag <realevenyag@gmail.com>

* test(sql): cover series index admin validation

Signed-off-by: evenyag <realevenyag@gmail.com>

* refactor(mito): bound series index maintenance queue

Signed-off-by: evenyag <realevenyag@gmail.com>

* docs: remove series index how-to guide

Signed-off-by: evenyag <realevenyag@gmail.com>

* test: remove index build upgrade compatibility case

Signed-off-by: evenyag <realevenyag@gmail.com>

* fix: address series index reconciliation review feedback

Signed-off-by: evenyag <realevenyag@gmail.com>

* fix: bound manual series index reconciliation admission

Signed-off-by: evenyag <realevenyag@gmail.com>

* chore: update greptime-proto to merged index build options

Signed-off-by: evenyag <realevenyag@gmail.com>

---------

Signed-off-by: evenyag <realevenyag@gmail.com>
2026-09-24 09:39:21 +00:00
shuiyisong b5c0e24625 feat(otlp): add histogram rejection metrics and compact zero buckets (#9321)
* feat(otlp): add histogram rejection metrics and compact zero buckets

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: test isolation

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* chore: fix minor naming

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

---------

Signed-off-by: shuiyisong <xixing.sys@gmail.com>
2026-09-24 08:45:45 +00:00
dennis zhuang 357c935804 fix: make vector aggregates work with GROUP BY and partial aggregation (#9338)
* fix: make vector aggregates work with GROUP BY and partial aggregation

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test: make partitioned vec_avg case distinguish weighted averages

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-24 08:45:40 +00:00
jeremyhi 822042c7bb feat(log-store): chain object store WAL objects and recover the latest chain (#9334)
* feat(log-store): chain object store WAL objects and recover the latest chain

A conditional create can fail with an unknown outcome while its object is
stored, or still lands later, and Mito reuses the row sequences of a
failed append. Replaying such an object next to a later acknowledged one
can let the unacknowledged rows win after a restart.

Format version 2 gives every object header its writer's epoch, a link
to the object it extends (sequence and writer instance) and a header
CRC32, and allows objects without segments. Recovery replays only the
chain ending at the complete object with the largest epoch and sequence:
a link holds when its predecessor is present with the recorded writer
instance, or is missing below every present object. Objects off the
chain are orphans that are never replayed but keep their sequences.

Each open writes an empty object that starts an epoch above every
present object, linked to the recovered tip, before it accepts writes,
so a late object of an earlier instance never ends the chain. A start
object that meets an object of an earlier epoch moves to the next
sequence; one of an equal or later epoch fails the open.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): break the WAL chain at every missing predecessor

Accepting a missing predecessor below every present object lets a late
object that lands below the chain change which links hold. Nothing
collects objects yet, so a missing predecessor now always breaks the
link, and recovery fails when objects are present but none completes a
chain.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): keep the object store WAL format at version 1

The object store WAL has not been enabled anywhere, so no object in the
previous layout exists and the chained header can stay version 1.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* refactor(log-store): link object store WAL objects by epoch

Only one store instance writes under an epoch, since every open starts an
epoch above every present object and a start object that meets the same
or a later epoch fails the open. The epoch therefore identifies the
instance, and the random writer instance id is dropped from the header.
A link now records the sequence and the epoch of the object it extends,
and holds when the predecessor carries that epoch. The header shrinks to
46 bytes, and the store logs its epoch when it opens.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): fail an open whose start object is already present

Without a random writer instance, two opens that recover the same objects
encode byte-identical start objects, and a conditional create treats the
same bytes as its own retry. A start object that is already present
therefore fails the open with a retryable error instead of letting both
opens claim the epoch; the next open counts the object and starts a
later epoch.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): derive the WAL epoch from the claimed start sequence

Two opens can recover different listings when a late object lands across
a sequence gap between them, pick the same largest epoch plus one, and
both create their start objects under different sequences. The epoch of
an instance is now one above the sequence its start object claims, so a
successful create decides the epoch and no two instances share one. It
stays above every epoch recovery listed, and an object that carries an
epoch above the next sequence fails the open.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test(log-store): cover a WAL start object stored with an unknown outcome

Add a fault-injection test in which the create of the start object
stores the object but reports an error: the open fails without moving
to another sequence, and the next open counts the stored object and
claims a later epoch. Rename the test helper that writes a whole object
from a given header to put_object_with_header, and drop a needless
clone.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-09-24 08:41:56 +00:00
Ning Sun 20619273ee chore(toolchain): switch to stable Rust 1.96.1 and remove all nightly feature gates (#9303)
* chore(toolchain): switch to stable Rust 1.96.1 and remove all nightly feature gates

Move the workspace from the pinned nightly-2026-03-21 to stable
1.96.1 and drop all 23 '#![feature]' gates across 13 crates,
rewriting the still-unstable API usages with stable equivalents:

- try_blocks: closures / an async block (table, query, common-function, servers)
- duration_constructors: Duration::from_secs(n * 86400) / (n * 60)
- iterator_try_collect: collect::<Result<Vec<_>, _>>()
- box_patterns: as_deref() + matches! chains (sql)
- error_iter: error_chain_root() source-chain walker (common-error);
  sources() includes the error itself, so the walker never panics
- int_roundings: div_floor -> div_euclid (equal for positive divisors)
- iter_partition_in_place: stable sort_by_key partition helper (index)
- hash_set_entry: HashSet::insert bool / contains+insert
- trait_alias: lifetime-parameterized dyn FnOnce type aliases (puffin)
- string_from_utf8_lossy_owned: from_utf8_lossy(&v).into_owned()
- never_type: Infallible (common-recordbatch)
- debug_closure_helpers: closure-backed DebugFmt newtype (mito2)
- binary_heap_pop_if: peek().is_some_and() + pop()
- exclusive_wrapper: drop Exclusive; C: Send + Unpin already in bounds
- stmt_expr_attributes: stale gate, no usages

Also fix release-dev-builder-images.yaml, which parsed
rust-toolchain.toml with a date-only regex and would produce empty
image versions with a stable channel; it now extracts the full
channel token. Dev-builder images verified against stable 1.96.1
(image build, default-toolchain behavior, binstall/nextest, riscv64
and android targets, in-image cargo check).

Validated on 1.96.1: cargo check --workspace --all-targets, clippy
--workspace --all-targets --all-features -D warnings, cargo fmt
--check, and nextest on all 13 affected crates (4586 passed).

Part of #9289. Depends on #9298 (fuzz nightly quarantine) merging
first.

Signed-off-by: Ning Sun <sunning@greptime.com>

* chore: update flake checksum

* chore: use wild for linker in flake

---------

Signed-off-by: Ning Sun <sunning@greptime.com>
2026-09-24 08:37:43 +00:00
discord9 abedeb2bea fix(query): keep count_values generated label in enclosing expressions (#9223)
* fix(query): keep count_values generated label in enclosing expressions

`count_values("v", m)` projects the generated label as a real output
column, but did not register it in `ctx.tag_columns`. Enclosing
expressions (abs/round/+1/topk/label_replace/vector join) rebuild their
projection from `ctx.tag_columns` and silently drop the label.

Register the generated label in `ctx.tag_columns` after the projection,
and give it the same qualifier as other tag columns so qualified
references resolve. PromQL overwrites an input label with the same name,
so drop any existing tag with that name first to avoid duplicate column
ambiguity.

Fixes https://github.com/GreptimeTeam/greptimedb/issues/9181
Report: .e-agent/greptimedb_promql_compatibility_report_2026-09-16.md P0-1

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(query): use total_cmp in count_values test helper

Silence clippy::needless_borrow on partial_cmp(&right.1); f64 sorting
uses total_cmp, matching the other planner test helpers.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(flow): keep count_values generated label as sink table primary key

FindGroupByFinalName::f_up only renamed a group key when the projection
aliased the key column directly. count_values now projects its generated
label as a unary UDF over the sampled column
(prom_float_to_string(value) AS label) while the aggregate still groups
by the raw column, so the name match failed and the sink table lost the
label from its PRIMARY KEY, demoting it to a DOUBLE value column.

Allow a projection above the aggregate to rename a group key by deriving
its output from a single group-key column (unary scalar function or cast
over that column only). Multi-column expressions, case, aggregates,
windows, subqueries and literals are still rejected so a computed column
cannot be mistaken for the group key.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(query): group count_values by the formatted sample value

The numeric count_values branch grouped by the raw sample column and only
formatted the value into the generated label in the post-aggregate
projection. Two distinct raw values that collapse to one label text
(e.g. BIGINT 9007199254740992 and 9007199254740993, both 9007199254740992
in Float64) were split into two groups, each emitting the same label set
at one timestamp, violating Prometheus' unique-label-set-per-timestamp
invariant.

Return the formatted value expression (prom_float_to_string, with a
CAST to Float64 for non-Float64 inputs) as previous_field_expressions so
the aggregate groups by the same expression that produces the label,
mirroring the existing mixed float/native-histogram precedent.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(query): drop needless borrow in count_values test helper

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(flow): don't replace group key with derived unary expr

FindGroupByFinalName treated any unary expression of a single group
column (scalar fn, Cast, TryCast) as a rename of that group key and
swapped it in as the sink primary key. A derived expression such as
lower(host) is many-to-one, so distinct groups (HOST_A vs host_a) would
collapse to the same primary key and be silently deduplicated.

Narrow the matching to direct name-matched aliases of the actual group
expression only, and drop the derived-unary path (is_unary_expr_of_column)
and the allow_derived flag. count_values already groups by the formatted
expression name, so it is unaffected.

Add a regression asserting host survives and host_lc does not replace it
under both optimizer settings.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-24 08:18:41 +00:00
dennis zhuang affc0a1b1d fix: align ordered aggregate state type with the accumulator output (#9340)
* fix: align ordered aggregate state type with the accumulator output

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix: reject mismatched aggregate states and keep hard-ordered aggregates unsplit

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix: keep WITHIN GROUP aggregates splittable and relabel state fields by position

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-24 07:02:14 +00:00
dennis zhuang def0c2ec5c fix: only push down aggregates grouped by the partition columns themselves (#9337)
* fix: only push down aggregates grouped by the partition columns themselves

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix: keep grouping sets on the frontend for partitioned tables

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-24 06:49:52 +00:00
dennis zhuang b9f991502c refactor: remove experimental vector index (#9345)
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-24 05:38:54 +00:00
dennis zhuang cc1be82858 fix: merge every state row in geo_path and json_encode_path (#9339)
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-24 02:12:11 +00:00
Yingwen c7fa48ef95 feat(mito): limit approximate series index disk usage (#9313)
* feat(mito): limit series index disk usage

Signed-off-by: evenyag <realevenyag@gmail.com>

* refactor(mito): enforce series index quota during reconciliation

Signed-off-by: evenyag <realevenyag@gmail.com>

* refactor(mito): remove series index disk budget layer

Signed-off-by: evenyag <realevenyag@gmail.com>

* refactor(mito): simplify series index limit to estimated usage

Signed-off-by: evenyag <realevenyag@gmail.com>

* refactor(mito): trust index catalogs when loading snapshots

Signed-off-by: evenyag <realevenyag@gmail.com>

* refactor(mito): minimize series index disk limit changes

Signed-off-by: evenyag <realevenyag@gmail.com>

* fix(mito2): keep series index cleanup running at capacity

Signed-off-by: evenyag <realevenyag@gmail.com>

* fix(mito2): avoid no-op index clones and stabilize capacity tests

Signed-off-by: evenyag <realevenyag@gmail.com>

* fix: update config API expectation and stabilize index build tests

Signed-off-by: evenyag <realevenyag@gmail.com>

---------

Signed-off-by: evenyag <realevenyag@gmail.com>
2026-09-23 14:18:42 +00:00
shuiyisong 20dde2601f feat: enable native histogram ingestion by default (#9301)
* feat: enable native histogram ingestion by default

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* chore: add comments

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: test

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

---------

Signed-off-by: shuiyisong <xixing.sys@gmail.com>
2026-09-23 13:03:51 +00:00
shuiyisong 2bc18c483d feat: add schema metadata stream to SchemaManager (#9325)
Signed-off-by: shuiyisong <xixing.sys@gmail.com>
2026-09-23 12:25:56 +00:00
dennis zhuang 88197f4019 fix(promql): derive vector matching result labels and reject ambiguous matchings (#9306)
* fix(promql): derive vector matching result labels and reject ambiguous matchings

A vector-vector binary operation projected one operand's whole tag set and
inner-joined without any cardinality check, so `on()`/`ignoring()` did not
reduce the result labels, `group_left`/`group_right` changed nothing, and a
non-unique match group produced a cross product that PromQL cannot represent.

Result labels now follow Prometheus `resultMetric`: `on(...)` keeps the
matching labels, `ignoring(...)` drops them, and a group modifier keeps the
many side's labels plus the `group_x(...)` labels taken from the one side.
A label the one side does not carry is deleted from the result. The reduced
label set no longer identifies the operand series, so `__tsid` is dropped
from the context on this path.

Cardinality is enforced with a `count(1) OVER (PARTITION BY match keys, ts)`
window and a scalar UDF that fails the query on a repeated group: on the one
side before the join, and on the result labels after it, matching where
Prometheus raises each of its three errors. Series are unique by their whole
tag set, so the window is only planted when the match keys drop a tag; plain
arithmetic and `on(<all tags>)` plan exactly as before.

Closes #9207, closes #9208, closes #9209.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* perf(promql): keep __tsid when the result labels are an operand's whole tag set

Deriving the result labels dropped `__tsid` from the context unconditionally,
so an enclosing operation fell back to joining on the tag columns even where
the column still identified the result series.

Keep it when every result label comes from one operand and covers that
operand's whole tag set: no other operand value reaches the labels, and the
matching gives each of its rows a single partner, so its `__tsid` is still one
per result series. That is the common `on(<all tags>)` and bare `group_left`
shape; a matching that actually drops a tag still clears it.

The column is re-qualified as the result's own, which is how the enclosing
expression and the context look it up.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(promql): keep the match group count column unambiguous

The cardinality check aliased its row count to a fixed `__promql_match_group_count`.
An operand carrying a label of that name made the window output two fields with
the same name, and planning failed with "Schema contains qualified field name
collide_right.__promql_match_group_count and unqualified field name
__promql_match_group_count which would be ambiguous".

Pick a name the operand does not already have, the way the `or` operator
allocates its match key columns.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(promql): cover a match group spread over several regions

The cardinality check runs above the merge of the region scans, so it counts a
match group globally. Nothing pinned that: every table in these cases holds a
single region, and a check evaluated per region would pass them all.

Partition the operand on a column outside the match keys, which puts the two
series of one match group in different regions, and assert both the pre-join
and the post-join check still reject it. The case runs in the distributed
environment too, where the regions sit on different datanodes.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(promql): group an outer aggregate on the labels the operand kept

`by`/`without` planning topped up a missing grouping column by walking down to
the table scan and re-projecting it. That is right for a column the scan pruned
for efficiency, but the labels a matching modifier deletes are also absent from
the operand's output, and they were restored the same way:

  sum without(host) (a / on(host) b)

`on(host)` leaves the operand with `host` alone, so the sum covers everything
and Prometheus answers `{} 10`. Instead `device` came back from the scan under
`a` and split the result into `{device="d1"} 5` and `{device="d2"} 5`. Same for
`sum by(device)` of that operand, which has no `device` to group on at all.

Take the grouping labels from the operand's own label set rather than from the
row keys of the scan beneath it. A label pruned from the plan is still in that
set and still gets restored; a label the operand dropped is not.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor(promql): drop the unreachable aggregation tag top-up

`by`/`without` planning could restore a grouping column that the plan no longer
carried by rewriting the scan underneath it. Once the grouping labels come from
the operand's own label set, there is nothing left for it to restore: a scan
projects every label of `ctx.tag_columns` (`scan_tag_columns` only ever adds
matcher columns to that set), so a label in the set is always in the schema.

Stubbing the rewriter to a no-op passed the whole sqlness suite, in both the
standalone and the distributed environment.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(promql): drop cases that guarded the removed tag top-up

Three plain selector aggregates were there to show that restoring a pruned
grouping column still worked. With the restore gone they only repeat what the
aggregate cases already cover. Also fix a comment that still said the metric
engine scan prunes tag columns: it projects every label of the operand.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* docs(promql): note the duplicate a propagated matcher hides

A matcher copied onto the one-side operand removes groups without a partner
before the cardinality check sees them, so a duplicate in such a group is not
reported. Prometheus checks every group of the one side and fails the query.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-23 11:08:18 +00:00
discord9 d910183cc4 feat(flow): freeze recovery windows and retention bounds for incremental flows (#9312)
* feat(flow): discover recovery windows from exact sequence reads

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* feat(flow): freeze explicit recovery windows and retention bounds

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(flow): validate retained source windows before recovery

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(flow): qualify recovery source tables before serialization

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-23 09:21:40 +00:00
fys 5a78e07be6 fix(json2): restrict JSON2 type hints (#9316)
* fix(sql): restrict JSON2 type hints

Signed-off-by: fys <fengys1996@gmail.com>

* fix(sql): allow equivalent 64-bit aliases in JSON2 type hints

Signed-off-by: fys <fengys1996@gmail.com>

* fix(sql): support UInt64 conversion and preserve JSON2 hint aliases

Signed-off-by: fys <fengys1996@gmail.com>

* refactor(sql): remove redundant JSON2 type hint normalization

Signed-off-by: fys <fengys1996@gmail.com>

* fix(datatypes): restrict JSON2 hint types in JsonSettings::try_new

Signed-off-by: fys <fengys1996@gmail.com>

---------

Signed-off-by: fys <fengys1996@gmail.com>
2026-09-23 09:19:18 +00:00
Yingwen bc80071409 feat: expose region open failure metrics (#9283)
Signed-off-by: evenyag <realevenyag@gmail.com>
2026-09-23 06:23:52 +00:00
Weny Xu a35f9d5c78 fix: address Windows test failures and run full Windows CI (#9305)
* fix: use relative object keys for Windows filesystem access

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: cover Windows path and time limits in full test CI

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: use relative keys in metadata snapshot tests

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-09-23 05:16:41 +00:00
1d8d95c12a chore(deps): replace cargo-udeps with cargo-shear for unused dependency checks (#9294)
* chore(deps): replace cargo-udeps with cargo-shear for unused dependency checks

cargo-udeps requires a nightly toolchain and its pinned version (0.1.61)
no longer detects unused dependencies against current cargo internals —
unused deps have landed on main undetected (e.g. humantime in
common-frontend since #6689). cargo-shear is a standalone static analyzer
that runs on any toolchain.

- Swap 'make check-udeps' / 'make fix-udeps' recipes to 'cargo shear' /
  'cargo shear --fix' and retire scripts/fix-udeps.py
- CI: install cargo-shear in the check-udeps job; drop the build cache
  and protoc steps (cargo-shear never compiles)
- Remove ~150 unused dependency declarations found by cargo-shear, move
  misplaced deps to the correct sections, drop orphaned
  [workspace.dependencies] entries (arrow-cast, rustc-hash)
- Add [package.metadata.cargo-shear] ignored entries with explanations
  for dependencies that are structurally required despite no textual
  reference: sqlparser (required by sqlparser_derive expansions in
  datatypes, common-query), common-error (required by common-macro's
  stack_trace_debug expansions in session, tests-fuzz), k8s-openapi
  (feature-pinning for the transitive kube dependency in tests-fuzz),
  tikv-jemalloc-sys (link-only, enables jemalloc profiling features in
  common-mem-prof), protobuf (required by build.rs-generated bindings in
  log-store)
- Drop the obsolete [package.metadata.cargo-udeps.ignore] sections

Part of #9289

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix(meta): populate physical metric table column ids (#9286)

* fix(meta): populate physical metric table column ids

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>

* test(meta): verify physical metric column ids

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>

---------

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>

* fix(postgres): return empty responses for comment-only SQL (#9295)

fix(postgres): handle parsed empty queries in both protocols

Signed-off-by: houyuwushang <180804215+houyuwushang@users.noreply.github.com>

* ci: create docs follow-up issue on PR merge instead of on label (#9237)

* ci: create docs follow-up issue on PR merge instead of on label

The docbot workflow previously created a docs-repo issue as soon as the
'docs-required' condition was detected (PR opened/edited with the docs
checkbox ticked), even if the PR was never merged.

Now the workflow also triggers on PR 'closed':
- opened/edited: only manage the docs-required/docs-not-required labels
- closed: create the docs issue only when the PR was actually merged and
  carries the docs-required label

This also lets maintainers control issue creation by manually adding or
removing the docs-required label before merging.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: address review comments on docs issue creation timing

- Only touch docs labels when the docs checkbox state actually changed
  in an edit. Previously, editing any other part of the PR body while
  the checkbox stayed checked removed the docs-required label, silently
  dropping the docs follow-up now that issue creation happens at merge.
  Unchanged checkbox now leaves labels untouched, which also preserves
  manual label overrides.
- Do not trust the closed event's stale label snapshot at merge time:
  re-read the live PR via the API and create the docs issue if the
  docs-required label is present OR the checkbox is ticked in the
  current body.
- Make the workflow concurrency group action-aware so a merge run does
  not cancel an in-flight label update from an edit run.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: make docs-required label the single source of truth at merge

The label-OR-checkbox merge condition could not distinguish an
intentional opt-out from an unfinished label update: removing
docs-required while the checkbox stayed checked still produced an
issue, and unchecking the box could still produce one if the merge
read the stale label before the edit run removed it.

At merge time, wait for any pending docbot runs on the PR head SHA to
finish their label updates (bounded to 5 minutes), then decide solely
by the live docs-required label. Adds actions: read permission for
listing workflow runs.

Signed-off-by: Ning Sun <sunning@greptime.com>

---------

Signed-off-by: Ning Sun <sunning@greptime.com>

* perf(promql): push label filters into grouped join inputs (#9280)

* perf(promql): propagate matching filters through grouped joins

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* perf(promql): check matcher safety on the receiving operand

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor(promql): spell out the shapes a filter may cross

`preserves_filter` ended in `_ => true`, which was only sound because
`selector_matchers` independently rejects label rewriting, `count_values`,
subqueries and non-rollup calls on the same operand. Loosening the latter
alone would have silently pushed a matcher below a label rewrite. List the
shapes that carry a scan filter instead and default to `false`.

Cite #9207 for the result labels the grouped cases record: the join
projects the right operand's tag set, so `zone` is missing wherever the
right side aggregates it away.

No behavior change.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(promql): assert the new pushdowns reach the scan

The grouped-join unit tests feed tag columns by hand and the SQLness case
only checks results, which are identical whether or not the rewrite fires.
Nothing would have failed if scalar arithmetic, ranking or grouped
matching stopped propagating. Assert through the planner that the matcher
reaches both scans, with a global topk one-side as the counter-example.

Also state that the duplicate-one-side cases record a cross product
Prometheus rejects (#9209), so the baseline is not read as intended
semantics.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(ci): build tests-integration lib with meta-srv/mock (#9299)

* fix(ci): build tests-integration lib with meta-srv/mock

tests-integration's lib code (src/cluster.rs) uses meta_srv::mocks, but
the dependency carrying the mock feature sits in [dev-dependencies].
Builds that only touch the lib, such as the apidoc job's cargo doc
--workspace, resolve meta-srv without mock and fail with E0432.
--all-targets builds unify dev-dependency features, which is why check,
clippy and nextest stayed green.

Move the mock-enabled meta-srv entry back to [dependencies]. The other
testing features moved out in #9072 are not needed by the lib and stay
in [dev-dependencies].

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(repartition): split per-case repartition tests

test_repartition_metric ran four format/primary-key-encoding cases in a
single test function, and test_repartition_mito ran two format cases.
Each case builds its own 3-datanode cluster and runs a full repartition
plus GC cycle, so on S3 the metric test took 165-178s against the 180s
nextest terminate-after. Merge queue runs failed on it at random.

Split each case into its own test. Cases were already independent, so
they now run in parallel and each stays far inside the timeout, and a
failure points at one encoding instead of four.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* feat(json2): support altering JSON2 settings (#9029)

* feat(sql): support alter syntax for JSON2 columns

Signed-off-by: fys <fengys1996@gmail.com>

* fix(json2): preserve rows on type hint mismatch during compaction

* refactor(json2): simplify alter settings handling

* fix(json2): preserve coerced values during compaction

* chore: remove unnecessary clone

* chor: reduce memory allocations

* fix: cargo clippy

* chore: update greptime-proto to main branch

* refactor(datatypes): unify string handling with other JSON type hints

* fix: cr

---------

Signed-off-by: fys <fengys1996@gmail.com>

* fix: keep compaction pruning, metadata, and index work on compact runtime (#9304)

* fix: run compaction pruner tasks on compact runtime

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: keep compaction metadata and index work on compact runtime

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* feat: add AI matching, classification, and scoring functions (#9300)

* feat: return matching scores from jev

Replace the experimental three-argument Boolean function with jev(text, prompt) returning a Float64 probability in [0, 1]. Move threshold comparisons into SQL and update tests and migration examples.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* feat: add Jev choice and score functions

Share asynchronous execution across Noul, Choice, and Score. Validate JSON criteria before requests and return typed scalar answers. Add SQL and HTTP mock coverage with usage examples.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor: use generic AI SQL function names

Expose ai_match, ai_choose, and ai_score and move their implementation, tests, and usage guide under generic AI names. Document the current unreleased interface without migration history.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: share constant AI criteria within each batch

Borrow scalar string arguments and lazily parse constant criteria once per batch. Share the parsed allocation across requests while preserving NULL propagation and batch validation before HTTP calls.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* feat: preserve AI score uncertainty in JSONB results

Return score, confidence, and probabilities in criteria-level order from one evaluation. Validate the distribution and preserve provider precision. Add JSON extraction, uncertainty, and single-request regressions, and document confidence-aware ranking.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* docs: explain reuse of volatile AI evaluations

Document repeated SELECT and WHERE evaluation costs as N + M requests, and show subquery aliases for reusing scalar or structured AI results without additional model calls.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* feat: share logical table batching with OTLP metrics (#9288)

* feat: share logical table batching with OTLP metrics

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: unify pending rows batch acknowledgement policy

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: align logical batcher example configuration expectations

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: align batcher worker channel defaults to 65536

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>

* perf(mito2): lazily decode dense primary key columns (#9226)

* perf(mito2): lazily decode dense primary key columns

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* perf(mito2): bypass lazy decoding for full primary keys

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(mito-codec): preserve prefix decoding errors

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor(mito-codec): align encoded length helper naming

Rename encoded_length to encoded_len and update all callers to match the other length helpers in the module.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor(mito2): clarify conditional dense key decoding

Rename decode_dense_pk to ensure_dense_pk_decoded so callers can see that existing decoded values are preserved and only missing caches are populated.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor(mito-codec): share string framing in row converter

Move encoded_string_len to the parent module so Dense and Sparse use the same framing helper without depending on each other. Preserve its implementation and visibility.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* chore: sync lock

* fix: shear and check issues

---------

Signed-off-by: Ning Sun <sunning@greptime.com>
Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>
Signed-off-by: houyuwushang <180804215+houyuwushang@users.noreply.github.com>
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
Signed-off-by: fys <fengys1996@gmail.com>
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
Signed-off-by: WenyXu <wenymedia@gmail.com>
Co-authored-by: Dhruv Vaishnav <dhruvvaishnav687@gmail.com>
Co-authored-by: houyuwushang <180804215+houyuwushang@users.noreply.github.com>
Co-authored-by: dennis zhuang <killme2008@gmail.com>
Co-authored-by: fys <40801205+fengys1996@users.noreply.github.com>
Co-authored-by: Lei, HUANG <6406592+v0y4g3r@users.noreply.github.com>
Co-authored-by: Weny Xu <wenymedia@gmail.com>
2026-09-23 04:56:44 +00:00
Weny Xu 953d01ac54 feat: support pending rows batching for MySQL and PostgreSQL (#9302)
* feat: support pending rows batching for MySQL and PostgreSQL

Signed-off-by: WenyXu <wenymedia@gmail.com>

* style: group batcher imports before item definitions

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: use 65536 as the default batcher worker channel capacity

Signed-off-by: WenyXu <wenymedia@gmail.com>

* test: use a distinct custom worker channel capacity

Signed-off-by: WenyXu <wenymedia@gmail.com>

* test: complete Prom config in worker capacity override case

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-09-23 04:41:13 +00:00
discord9 9281965bac feat(flow): allow execution hook to rewrite completed plans (#9290)
* refactor(flow): expose incremental aggregate plan analysis

Signed-off-by: discord9 <discord9@163.com>

* feat(flow): allow execution hook to rewrite completed plans

Signed-off-by: discord9 <discord9@163.com>

* test(flow): cover the execution hook receiving the completed plan

The hook must see the plan that is dispatched after the incremental
delta-sink merge, so the test records the plan a collaborator is handed and
asserts the merge is already part of it.

Signed-off-by: discord9 <discord9@163.com>

---------

Signed-off-by: discord9 <discord9@163.com>
2026-09-23 04:20:02 +00:00
Ning Sun 0f118bc5ba fix: serialize struct to json in postgres (#9170)
* feat: serialize struct to json in postgres

* fix: support view scalars and preserve null structs in scalar-to-value conversion

Address PR review:
- Utf8View/BinaryView ScalarValues now convert like their non-view forms
  instead of failing row extraction for struct columns
- a null struct scalar converts to Value::Null so a null struct inside a
  list stays null in the serialized JSON

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: return errors instead of panics for unsupported arrow field types

Struct-typed query results with arrow field types greptimedb cannot
represent (e.g. Decimal256) used to panic during schema conversion and
row extraction, dropping the client connection. They now surface as
query errors:

- ConcreteDataType::try_from builds struct types fallibly via the new
  StructType::try_from_arrow_fields
- Value::try_from(ScalarValue::Struct) uses the same fallible path
- new try_value_from_array converts an arrow element to Value with
  error propagation, used by the postgres struct encoding

Signed-off-by: Ning Sun <sunning@greptime.com>

---------

Signed-off-by: Ning Sun <sunning@greptime.com>
2026-09-23 01:55:46 +00:00
dennis zhuang f3eb8e6c72 test: cut integration test time and make the storage matrix meaningful (#9308)
* test: cut integration test time and make the storage matrix meaningful

tests-integration is ~85% of workspace test CPU, and 81% of that is the
S3/S3WithCache variants of the HTTP and gRPC suites. Those suites do not
touch the object store: of the 70 matrix HTTP tests only one flushed and
read back an SST, so the matrix was paying real AWS round trips to
re-prove protocol parsing.

- Point the PR CI object-store matrix at the MinIO already started by
  tests-integration/fixtures. Three GT_S3_* consumers did not read
  GT_S3_ENDPOINT_URL and would have hit real AWS with MinIO credentials;
  they now do.
- Add a nightly Linux job against real AWS S3, and pass GT_S3_* into the
  release integration-test container. The release previously ran every
  remote-backend case as a skip and only exercised the file backend.
- Give each S3WithCache test its own read cache directory. They shared
  /tmp/greptimedb_cache, which the datanode wipes on startup, so a
  starting test deleted the read cache of a running one.
- Add flush -> read-back assertions to the tests whose columns have a
  non-trivial SST representation: JSON/JSON2 columns, native histograms,
  metric-engine logical tables, and tables carrying fulltext or skipping
  indexes whose puffin files only exist after a flush.
- Move eight tests that create no table out of the storage matrix.
- Make the event recorder flush interval a constructor parameter and
  shorten it in the event tests, which otherwise wait a 5s window per DDL
  they assert on. It is skipped by serde and never reaches config files.
- Drop duplicates: test_grpc_zstd_compression was a verbatim copy of
  test_grpc_message_size_ok and is now rewritten to assert the negotiated
  grpc-encoding; test_execute_copy_to_{s3,oss,gcs,azblob} were strict
  prefixes of their copy_from siblings; two standalone/distributed event
  test pairs shared one assertion body.
- Fix and un-ignore stddev_by_label. stddev_pop merges partial aggregates
  in a parallelism-dependent order, so its last digits are unstable; the
  test now compares values with a tolerance.
- Rebase the jaeger v1 fixture on the current instant. It carries
  ttl=7d with 2025 timestamps, so its rows were only readable as long as
  they stayed in the memtable.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test: address review — wire nightly real-S3 job into check-status, keep the short event interval

The nightly `check-status` job did not depend on the new real-S3 job, so a
failure there would not have reached the status or Slack notification.

In database_ddl_event the short interval was set by a first
`with_event_recorder_options` call and then overwritten by the pre-existing
one, which carries `..Default::default()`. Merged into a single call.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-23 01:52:52 +00:00
Lei, HUANG e91faa9df8 perf(mito2): lazily decode dense primary key columns (#9226)
* perf(mito2): lazily decode dense primary key columns

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* perf(mito2): bypass lazy decoding for full primary keys

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(mito-codec): preserve prefix decoding errors

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor(mito-codec): align encoded length helper naming

Rename encoded_length to encoded_len and update all callers to match the other length helpers in the module.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor(mito2): clarify conditional dense key decoding

Rename decode_dense_pk to ensure_dense_pk_decoded so callers can see that existing decoded values are preserved and only missing caches are populated.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor(mito-codec): share string framing in row converter

Move encoded_string_len to the parent module so Dense and Sparse use the same framing helper without depending on each other. Preserve its implementation and visibility.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-09-22 15:52:54 +00:00
Weny Xu 723da69b21 feat: share logical table batching with OTLP metrics (#9288)
* feat: share logical table batching with OTLP metrics

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: unify pending rows batch acknowledgement policy

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: align logical batcher example configuration expectations

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: align batcher worker channel defaults to 65536

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-09-22 14:38:17 +00:00
Lei, HUANG d0f8f4b80c feat: add AI matching, classification, and scoring functions (#9300)
* feat: return matching scores from jev

Replace the experimental three-argument Boolean function with jev(text, prompt) returning a Float64 probability in [0, 1]. Move threshold comparisons into SQL and update tests and migration examples.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* feat: add Jev choice and score functions

Share asynchronous execution across Noul, Choice, and Score. Validate JSON criteria before requests and return typed scalar answers. Add SQL and HTTP mock coverage with usage examples.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor: use generic AI SQL function names

Expose ai_match, ai_choose, and ai_score and move their implementation, tests, and usage guide under generic AI names. Document the current unreleased interface without migration history.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: share constant AI criteria within each batch

Borrow scalar string arguments and lazily parse constant criteria once per batch. Share the parsed allocation across requests while preserving NULL propagation and batch validation before HTTP calls.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* feat: preserve AI score uncertainty in JSONB results

Return score, confidence, and probabilities in criteria-level order from one evaluation. Validate the distribution and preserve provider precision. Add JSON extraction, uncertainty, and single-request regressions, and document confidence-aware ranking.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* docs: explain reuse of volatile AI evaluations

Document repeated SELECT and WHERE evaluation costs as N + M requests, and show subquery aliases for reusing scalar or structured AI results without additional model calls.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-09-22 14:12:24 +00:00
Lei, HUANG a9e2a89b7b fix: keep compaction pruning, metadata, and index work on compact runtime (#9304)
* fix: run compaction pruner tasks on compact runtime

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: keep compaction metadata and index work on compact runtime

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-09-22 13:12:39 +00:00
fys 045441e3cc feat(json2): support altering JSON2 settings (#9029)
* feat(sql): support alter syntax for JSON2 columns

Signed-off-by: fys <fengys1996@gmail.com>

* fix(json2): preserve rows on type hint mismatch during compaction

* refactor(json2): simplify alter settings handling

* fix(json2): preserve coerced values during compaction

* chore: remove unnecessary clone

* chor: reduce memory allocations

* fix: cargo clippy

* chore: update greptime-proto to main branch

* refactor(datatypes): unify string handling with other JSON type hints

* fix: cr

---------

Signed-off-by: fys <fengys1996@gmail.com>
2026-09-22 12:56:26 +00:00