Commit Graph
628 Commits
Author SHA1 Message Date
Yingwen 3c2aac0a55 feat: add manual series index reconciliation (#9323)
* feat(mito): support manual series index reconciliation

Signed-off-by: evenyag <realevenyag@gmail.com>

* feat(storage): route series index build requests

Signed-off-by: evenyag <realevenyag@gmail.com>

* feat(admin): add BUILD_SERIES_INDEX

Signed-off-by: evenyag <realevenyag@gmail.com>

* test: specify compaction type in series index fixtures

Signed-off-by: evenyag <realevenyag@gmail.com>

* test: correct series index SQL fixtures and error assertions

Signed-off-by: evenyag <realevenyag@gmail.com>

* test(compat): preserve legacy index rebuild across upgrades

Signed-off-by: evenyag <realevenyag@gmail.com>

* test(sql): cover series index admin validation

Signed-off-by: evenyag <realevenyag@gmail.com>

* refactor(mito): bound series index maintenance queue

Signed-off-by: evenyag <realevenyag@gmail.com>

* docs: remove series index how-to guide

Signed-off-by: evenyag <realevenyag@gmail.com>

* test: remove index build upgrade compatibility case

Signed-off-by: evenyag <realevenyag@gmail.com>

* fix: address series index reconciliation review feedback

Signed-off-by: evenyag <realevenyag@gmail.com>

* fix: bound manual series index reconciliation admission

Signed-off-by: evenyag <realevenyag@gmail.com>

* chore: update greptime-proto to merged index build options

Signed-off-by: evenyag <realevenyag@gmail.com>

---------

Signed-off-by: evenyag <realevenyag@gmail.com>
2026-09-24 09:39:21 +00:00
1d8d95c12a chore(deps): replace cargo-udeps with cargo-shear for unused dependency checks (#9294)
* chore(deps): replace cargo-udeps with cargo-shear for unused dependency checks

cargo-udeps requires a nightly toolchain and its pinned version (0.1.61)
no longer detects unused dependencies against current cargo internals —
unused deps have landed on main undetected (e.g. humantime in
common-frontend since #6689). cargo-shear is a standalone static analyzer
that runs on any toolchain.

- Swap 'make check-udeps' / 'make fix-udeps' recipes to 'cargo shear' /
  'cargo shear --fix' and retire scripts/fix-udeps.py
- CI: install cargo-shear in the check-udeps job; drop the build cache
  and protoc steps (cargo-shear never compiles)
- Remove ~150 unused dependency declarations found by cargo-shear, move
  misplaced deps to the correct sections, drop orphaned
  [workspace.dependencies] entries (arrow-cast, rustc-hash)
- Add [package.metadata.cargo-shear] ignored entries with explanations
  for dependencies that are structurally required despite no textual
  reference: sqlparser (required by sqlparser_derive expansions in
  datatypes, common-query), common-error (required by common-macro's
  stack_trace_debug expansions in session, tests-fuzz), k8s-openapi
  (feature-pinning for the transitive kube dependency in tests-fuzz),
  tikv-jemalloc-sys (link-only, enables jemalloc profiling features in
  common-mem-prof), protobuf (required by build.rs-generated bindings in
  log-store)
- Drop the obsolete [package.metadata.cargo-udeps.ignore] sections

Part of #9289

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix(meta): populate physical metric table column ids (#9286)

* fix(meta): populate physical metric table column ids

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>

* test(meta): verify physical metric column ids

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>

---------

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>

* fix(postgres): return empty responses for comment-only SQL (#9295)

fix(postgres): handle parsed empty queries in both protocols

Signed-off-by: houyuwushang <180804215+houyuwushang@users.noreply.github.com>

* ci: create docs follow-up issue on PR merge instead of on label (#9237)

* ci: create docs follow-up issue on PR merge instead of on label

The docbot workflow previously created a docs-repo issue as soon as the
'docs-required' condition was detected (PR opened/edited with the docs
checkbox ticked), even if the PR was never merged.

Now the workflow also triggers on PR 'closed':
- opened/edited: only manage the docs-required/docs-not-required labels
- closed: create the docs issue only when the PR was actually merged and
  carries the docs-required label

This also lets maintainers control issue creation by manually adding or
removing the docs-required label before merging.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: address review comments on docs issue creation timing

- Only touch docs labels when the docs checkbox state actually changed
  in an edit. Previously, editing any other part of the PR body while
  the checkbox stayed checked removed the docs-required label, silently
  dropping the docs follow-up now that issue creation happens at merge.
  Unchanged checkbox now leaves labels untouched, which also preserves
  manual label overrides.
- Do not trust the closed event's stale label snapshot at merge time:
  re-read the live PR via the API and create the docs issue if the
  docs-required label is present OR the checkbox is ticked in the
  current body.
- Make the workflow concurrency group action-aware so a merge run does
  not cancel an in-flight label update from an edit run.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: make docs-required label the single source of truth at merge

The label-OR-checkbox merge condition could not distinguish an
intentional opt-out from an unfinished label update: removing
docs-required while the checkbox stayed checked still produced an
issue, and unchecking the box could still produce one if the merge
read the stale label before the edit run removed it.

At merge time, wait for any pending docbot runs on the PR head SHA to
finish their label updates (bounded to 5 minutes), then decide solely
by the live docs-required label. Adds actions: read permission for
listing workflow runs.

Signed-off-by: Ning Sun <sunning@greptime.com>

---------

Signed-off-by: Ning Sun <sunning@greptime.com>

* perf(promql): push label filters into grouped join inputs (#9280)

* perf(promql): propagate matching filters through grouped joins

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* perf(promql): check matcher safety on the receiving operand

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor(promql): spell out the shapes a filter may cross

`preserves_filter` ended in `_ => true`, which was only sound because
`selector_matchers` independently rejects label rewriting, `count_values`,
subqueries and non-rollup calls on the same operand. Loosening the latter
alone would have silently pushed a matcher below a label rewrite. List the
shapes that carry a scan filter instead and default to `false`.

Cite #9207 for the result labels the grouped cases record: the join
projects the right operand's tag set, so `zone` is missing wherever the
right side aggregates it away.

No behavior change.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(promql): assert the new pushdowns reach the scan

The grouped-join unit tests feed tag columns by hand and the SQLness case
only checks results, which are identical whether or not the rewrite fires.
Nothing would have failed if scalar arithmetic, ranking or grouped
matching stopped propagating. Assert through the planner that the matcher
reaches both scans, with a global topk one-side as the counter-example.

Also state that the duplicate-one-side cases record a cross product
Prometheus rejects (#9209), so the baseline is not read as intended
semantics.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(ci): build tests-integration lib with meta-srv/mock (#9299)

* fix(ci): build tests-integration lib with meta-srv/mock

tests-integration's lib code (src/cluster.rs) uses meta_srv::mocks, but
the dependency carrying the mock feature sits in [dev-dependencies].
Builds that only touch the lib, such as the apidoc job's cargo doc
--workspace, resolve meta-srv without mock and fail with E0432.
--all-targets builds unify dev-dependency features, which is why check,
clippy and nextest stayed green.

Move the mock-enabled meta-srv entry back to [dependencies]. The other
testing features moved out in #9072 are not needed by the lib and stay
in [dev-dependencies].

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(repartition): split per-case repartition tests

test_repartition_metric ran four format/primary-key-encoding cases in a
single test function, and test_repartition_mito ran two format cases.
Each case builds its own 3-datanode cluster and runs a full repartition
plus GC cycle, so on S3 the metric test took 165-178s against the 180s
nextest terminate-after. Merge queue runs failed on it at random.

Split each case into its own test. Cases were already independent, so
they now run in parallel and each stays far inside the timeout, and a
failure points at one encoding instead of four.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* feat(json2): support altering JSON2 settings (#9029)

* feat(sql): support alter syntax for JSON2 columns

Signed-off-by: fys <fengys1996@gmail.com>

* fix(json2): preserve rows on type hint mismatch during compaction

* refactor(json2): simplify alter settings handling

* fix(json2): preserve coerced values during compaction

* chore: remove unnecessary clone

* chor: reduce memory allocations

* fix: cargo clippy

* chore: update greptime-proto to main branch

* refactor(datatypes): unify string handling with other JSON type hints

* fix: cr

---------

Signed-off-by: fys <fengys1996@gmail.com>

* fix: keep compaction pruning, metadata, and index work on compact runtime (#9304)

* fix: run compaction pruner tasks on compact runtime

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: keep compaction metadata and index work on compact runtime

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* feat: add AI matching, classification, and scoring functions (#9300)

* feat: return matching scores from jev

Replace the experimental three-argument Boolean function with jev(text, prompt) returning a Float64 probability in [0, 1]. Move threshold comparisons into SQL and update tests and migration examples.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* feat: add Jev choice and score functions

Share asynchronous execution across Noul, Choice, and Score. Validate JSON criteria before requests and return typed scalar answers. Add SQL and HTTP mock coverage with usage examples.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor: use generic AI SQL function names

Expose ai_match, ai_choose, and ai_score and move their implementation, tests, and usage guide under generic AI names. Document the current unreleased interface without migration history.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: share constant AI criteria within each batch

Borrow scalar string arguments and lazily parse constant criteria once per batch. Share the parsed allocation across requests while preserving NULL propagation and batch validation before HTTP calls.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* feat: preserve AI score uncertainty in JSONB results

Return score, confidence, and probabilities in criteria-level order from one evaluation. Validate the distribution and preserve provider precision. Add JSON extraction, uncertainty, and single-request regressions, and document confidence-aware ranking.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* docs: explain reuse of volatile AI evaluations

Document repeated SELECT and WHERE evaluation costs as N + M requests, and show subquery aliases for reusing scalar or structured AI results without additional model calls.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* feat: share logical table batching with OTLP metrics (#9288)

* feat: share logical table batching with OTLP metrics

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: unify pending rows batch acknowledgement policy

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: align logical batcher example configuration expectations

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: align batcher worker channel defaults to 65536

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>

* perf(mito2): lazily decode dense primary key columns (#9226)

* perf(mito2): lazily decode dense primary key columns

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* perf(mito2): bypass lazy decoding for full primary keys

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(mito-codec): preserve prefix decoding errors

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor(mito-codec): align encoded length helper naming

Rename encoded_length to encoded_len and update all callers to match the other length helpers in the module.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor(mito2): clarify conditional dense key decoding

Rename decode_dense_pk to ensure_dense_pk_decoded so callers can see that existing decoded values are preserved and only missing caches are populated.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor(mito-codec): share string framing in row converter

Move encoded_string_len to the parent module so Dense and Sparse use the same framing helper without depending on each other. Preserve its implementation and visibility.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* chore: sync lock

* fix: shear and check issues

---------

Signed-off-by: Ning Sun <sunning@greptime.com>
Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>
Signed-off-by: houyuwushang <180804215+houyuwushang@users.noreply.github.com>
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
Signed-off-by: fys <fengys1996@gmail.com>
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
Signed-off-by: WenyXu <wenymedia@gmail.com>
Co-authored-by: Dhruv Vaishnav <dhruvvaishnav687@gmail.com>
Co-authored-by: houyuwushang <180804215+houyuwushang@users.noreply.github.com>
Co-authored-by: dennis zhuang <killme2008@gmail.com>
Co-authored-by: fys <40801205+fengys1996@users.noreply.github.com>
Co-authored-by: Lei, HUANG <6406592+v0y4g3r@users.noreply.github.com>
Co-authored-by: Weny Xu <wenymedia@gmail.com>
2026-09-23 04:56:44 +00:00
fys 045441e3cc feat(json2): support altering JSON2 settings (#9029)
* feat(sql): support alter syntax for JSON2 columns

Signed-off-by: fys <fengys1996@gmail.com>

* fix(json2): preserve rows on type hint mismatch during compaction

* refactor(json2): simplify alter settings handling

* fix(json2): preserve coerced values during compaction

* chore: remove unnecessary clone

* chor: reduce memory allocations

* fix: cargo clippy

* chore: update greptime-proto to main branch

* refactor(datatypes): unify string handling with other JSON type hints

* fix: cr

---------

Signed-off-by: fys <fengys1996@gmail.com>
2026-09-22 12:56:26 +00:00
Lei, HUANG 6a084d9a4f fix: release completed SST write buffers during compaction (#9243)
* fix: upgrade OpenDAL and honor SST write buffer size

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* docs: track removal of the OpenDAL bridge fork

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(mito2): honor SST write buffers across upload paths

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-09-21 07:47:23 +00:00
shuiyisongandLei, HUANG 66d38e8e1c feat: add database ingestion admission through metering (#9239)
* feat: add `ingest_rows_rate_limit` database option

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* feat: add `InsertLimitInterceptor` hook to `Inserter`

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: attribute insert limit checks to the target table's database

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor: unify write admission through metering

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: enforce write admission for pending row batches

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* feat: distinguish internal requests for ingestion metering

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: admit split ingestion requests once per database

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: OpenTSDB throws error reason

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* chore: use meter crate main rev

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: exclude database ingest rate limit from table options

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* refactor: reserve channel 255 for internal requests

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
Signed-off-by: shuiyisong <xixing.sys@gmail.com>
Co-authored-by: Lei, HUANG <ratuthomm@gmail.com>
2026-09-21 07:27:22 +00:00
discord9 6647dbc437 feat!: upgrade DataFusion fork to 55.1.0 (#9177)
Rebase the GreptimeDB DataFusion fork from official 55.0.0 to 55.1.0.
DataFusion 55.1.0 is a patch release on branch-55 containing eleven
cherry-picked fixes (schema-adaptation struct filters, cast/projection
metadata propagation, nested-nullability aggregation adaptation,
UnnestExec batch_size, RightMark join ordering panic, empty-struct
ScalarValue, and FFI codec fixes). All twenty GreptimeDB fork patches
rebase onto it with no textual or semantic overlap; none of them is
absorbed upstream, so all are retained.

Fork pin moves to discord9/datafusion branch greptimedb-55.1.0,
commit 2aa87d52cdce7006af492330064738f33ed294c1 (55.1.0 +
20 patches).

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-21 04:14:22 +00:00
Ning Sun 055bbbc93d feat: update pgwire to 0.41 (#9025)
* feat: update pgwire to 0.41

* feat: update datafusion-pg-catalog and arrow-pg
2026-09-21 02:35:05 +00:00
LFC 697cc5fa29 fix(json): fix JSONPath panic with jsonb 0.5.6 (follow-up to #9192) (#9228)
* fix(json): upgrade jsonb to fix unterminated JSONPath panic

Signed-off-by: luofucong <luofc@foxmail.com>

* fix(json): align integer extraction with JSON2 conversions

Signed-off-by: luofucong <luofc@foxmail.com>

* style: format jsonb dependency declaration

Signed-off-by: luofucong <luofc@foxmail.com>

* fix(otlp): report concrete unsupported JSONB type names

Signed-off-by: luofucong <luofc@foxmail.com>

---------

Signed-off-by: luofucong <luofc@foxmail.com>
2026-09-18 04:55:02 +00:00
jeremyhi dca01654bc feat: prepare and execute database Metric exports (#9180)
* feat: prepare and authorize captured database exports

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* feat: bound database export jobs and drain failures

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test: validate database export identity and restore equivalence

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: validate database export directory URLs

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: preserve Windows export directory paths

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: reject local database export filename aliases

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* refactor: consolidate database export planning policies

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: restore escaped database export filenames

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test: align database restore assertions with shared policies

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* refactor: clarify database export boundaries and names

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-09-17 12:02:39 +00:00
Weny Xu be15c88e92 feat: batch ordinary table writes across HTTP protocols (#9115)
* feat: integrate table batching across HTTP protocols

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: skip empty prepared writes before batch admission

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: load batching protocols from environment and document frontend wiring

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor: remove experimental prefix from pending rows batcher config

Signed-off-by: WenyXu <wenymedia@gmail.com>

* style: sort frontend test dependencies

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: count batched ingestion once and update config snapshot

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-09-17 09:32:12 +00:00
jeremyhi d45f5d6eaf feat(log-store): add the object store WAL object format (#9200)
* feat(log-store): add the object store WAL object format

Add the byte format of a single object store WAL object: a header with
the GTWALOBJ magic, format version 1, the object sequence and the writer
instance id; one segment per region ordered by region id with entries
ordered by entry id; a footer that records each segment's region id,
entry id range, entry count, byte range and CRC32; and a fixed trailer
with the GTWALTRL magic, the footer location, the footer CRC32 and the
whole-object CRC32.

The module encodes objects deterministically and decodes the header,
trailer, footer and segments separately, with structural checks on
footer ranges and segment tiling. It has no callers yet; the store that
writes and reads objects follows in later changes.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test(log-store): pin the object store WAL format with a byte fixture

Add a fixed version 1 object with two regions and five entries as a hex
literal. The test decodes it and checks the exact header, trailer,
footer entries and records, and checks that encoding the same records,
in either input order, reproduces the fixture byte for byte.

Round-trip tests alone pass when a refactor changes field order,
endianness or checksum coverage in both the encoder and the decoder.
The fixture bytes were derived from the documented layout rather than
from the encoder, so such a change now fails.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): reject an empty footer when decoding a WAL object

The encoder never writes an object without records, but decode_footer
accepted a footer that declares zero segments, and
verify_segment_ranges accepts an empty footer too. Only the test-only
decode_object rejected it, so a checksum-valid empty object would pass
the header, trailer and footer checks that recovery runs.

Reject a zero entry count in decode_footer and drop the now unreachable
check in decode_object. Add a test that builds a checksum-valid object
with an empty footer and checks that decode_footer and decode_object
reject it.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* refactor(log-store): use pub(crate) for the WAL object format API

Other log-store modules use pub(crate) for items shared across module
boundaries. Switch the format module from pub(super) to pub(crate) to
follow that convention. No behavior change.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-09-17 09:14:22 +00:00
Weny Xu e0915a8380 refactor: reuse common batching components in Prom ingestion (#9114)
* refactor: compose Prom ingestion with common batcher components

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor: share bulk insert IPC encoding

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor: centralize batcher worker channel creation

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor: share timestamp extraction for batched writes

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor: preserve sharded batcher worker lookup

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: align batcher worker lookup and configuration validation

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor: validate batcher limits inside fallible constructors

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-09-16 09:21:49 +00:00
discord9andNing Sun 94d7e2c7fc feat!: upgrade DataFusion to 55 (#8555)
* feat!: upgrade DataFusion dependencies to 55

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor: migrate DataFusion 55 APIs

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: preserve table function planning behavior

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: preserve PostgreSQL query compatibility

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: preserve distributed execution plan behavior

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: cover DataFusion 55 behavior regressions

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: update DataFusion 55 SQLness expectations

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: complete DataFusion 55 test API migration

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: address DataFusion 55 CI regressions

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: address remaining DataFusion 55 regressions

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: adapt latest base code to DataFusion 55

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: normalize environment-specific DataFusion 55 plans

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: align final DataFusion 55 expectations

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: isolate DataFusion 55 regression cases

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: preserve empty result schema in timestamp widening

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: preserve JSON source column order

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: use released DataFusion 55 integrations

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: adapt latest execution plan mock to DataFusion 55

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: pin DataFusion recursive schema and date repairs

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(promql): align dictionary temporality match keys

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: retain Greptime DataFusion fork behaviors on version 55

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: restore ordinary function error expectations

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: refresh distributed count compatibility plan

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(query): adapt last-row cast hint to DataFusion 55

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: refresh instant last-row empty results for Arrow 59

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* style: simplify DataFusion expression visitor imports

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: restore sorting and PostgreSQL column-order assertions

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(function): restore primitive numeric coercion signatures

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(function): share geo integer signature types

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: cover timestamp widening overflow boundaries

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: fix decimal coercion regression imports

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(function): preserve scalar count_hash NULL state semantics

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: simplify decimal clamp case type inference

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: retain historical count_hash wrapper result

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: restore timestamp widening equality and IN pruning

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: carry upstream aggregate dynamic filter correctness fix

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: carry upstream null and predicate simplification fixes

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: restore baseline JSON ordering expectations

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: restore histogram JSON ordering expectations

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: refresh empty PromQL range result schemas

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: align native timestamp plan with DF55 decimal display

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: refresh native timestamp SQLness results for DF55

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: regenerate NULL sample empty result headers for DF55

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: use DF55 child replacement API in timestamp regressions

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: expose pushed scan dynamic filters to DF55 producers

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: encode string-backed PostgreSQL OID aliases in binary results

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: verify REGPROC binary and text over PostgreSQL protocol

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: register real PostgreSQL catalogs in server fixtures

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: complete DF55 expression inventories for custom query plans

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: correct RangeSelect expression fixture and column identities

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* ci: wait for Kafka WAL helper deployment rollout

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: update custom storage empty result headers

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: require exact row counts in scan statistics

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: suppress deprecated partition_statistics warning in test

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
Co-authored-by: Ning Sun <sunng@protonmail.com>
2026-09-15 11:42:38 +00:00
Weny Xu 737025760e feat: support request-level WAL skipping for bulk inserts (#9110)
* feat: support request-level WAL skipping for bulk inserts

Signed-off-by: WenyXu <wenymedia@gmail.com>

* test: cover bulk insert WAL skipping across protocols

Signed-off-by: WenyXu <wenymedia@gmail.com>

* test: align WAL snapshot naming with sequence watermarks

Signed-off-by: WenyXu <wenymedia@gmail.com>

* chore: bump proto

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-09-14 09:38:05 +00:00
Lei, HUANG 9619085310 chore: bump tikv-jemalloc-sys patch to jemalloc dev (ff80bf2d) (#9119)
* chore: bump tikv-jemalloc-sys patch to jemalloc dev (ff80bf2d)

Update the [patch.crates-io] rev for tikv-jemalloc-sys from
e1846d8c (5.3.1 + 54f22c83 backport) to ff444d4 (upstream dev HEAD,
161 commits ahead of 5.3.1).

The dev branch includes additional TSD/tcache fixes beyond the
original backport:
- fb5499aa9c: Handle jemalloc calls after TSD teardown
- 61dc1da395: Fix possible tcache corruption on fiber migration
- 1e92317014: Fix thread-exit TSD cleanup

See GreptimeTeam/jemallocator branch bump-jemalloc-dev and
tikv/jemallocator#182.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* chore: point tikv-jemalloc-sys patch at GreptimeTeam/jemallocator main

GreptimeTeam/jemallocator#2 has been merged; reference the merge
commit e254a7ea on main instead of the PR head branch. Jemalloc
source content is unchanged (still upstream dev ff80bf2d).

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-09-11 07:29:43 +00:00
Lei, HUANG ab0b1f5cce fix: bump jemalloc crates to 0.7 and patch tikv-jemalloc-sys with tcache init fix (#9103)
* fix: bump jemalloc crates to 0.7 and patch tikv-jemalloc-sys with tcache init fix

Upgrade tikv-jemallocator / tikv-jemalloc-ctl / tikv-jemalloc-sys from
0.6 to 0.7, which embeds jemalloc 5.3.1 (includes a056c20d 'Handle
tcache init failures gracefully').

On top of that, patch tikv-jemalloc-sys to the GreptimeTeam fork that
adds the remaining upstream fix 54f22c83 'Initialize TSD tcache before
enabling it' (GreptimeTeam/jemalloc#1, GreptimeTeam/jemallocator#1).
Without the ordering fix, a reentrant allocation during TSD bootstrap
(e.g. heap-profiling prof_tdata init / sampled backtrace when prof:true
is active) can observe an enabled-but-uninitialized tcache, corrupting
per-thread tcache metadata and crashing the process in arena_stats_merge,
calloc, or the libgcc unwinder.

The patch is pinned by rev and should be removed once tikv/jemallocator
ships a jemalloc snapshot that includes 54f22c83.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* chore: bump tikv-jemalloc-sys patch rev to merged release-5.3.1

GreptimeTeam/jemalloc#1 has been merged; point the patch at the
jemallocator commit referencing the merge commit on release-5.3.1.
Jemalloc source content is unchanged.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* chore: point tikv-jemalloc-sys patch at GreptimeTeam/jemallocator main

GreptimeTeam/jemallocator#1 has been merged; reference the merge
commit e1846d8c on main instead of the PR head branch. Jemalloc
source content is unchanged.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(servers): bump tikv-jemallocator dev-dependency to 0.7

main added a target.'cfg(not(windows))'.dev-dependencies entry on
tikv-jemallocator 0.6 for servers after this branch diverged. On the
merge ref it pulled tikv-jemalloc-sys 0.6 from crates.io, which
conflicts with the patched 0.7 (links = "jemalloc" may only appear
once in the dependency graph), failing version selection in CI.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-09-10 13:06:43 +00:00
discord9 5d5f0d6d70 fix(query): prevent incomplete aggregate dynamic filtering (#9102)
* fix(query): prevent incomplete aggregate dynamic filtering

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(query): cover mixed expression and column maxima

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(query): pin upstream aggregate regression backport

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(query): assert aggregate filter behavior on datanode scans

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(query): pin merged aggregate dynamic filter fix

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-10 12:52:35 +00:00
Weny Xu fa128adb8e feat: support request-level insert WAL skipping (#9088)
* refactor: add skip_wal fields to internal write requests

Signed-off-by: WenyXu <wenymedia@gmail.com>

* chore(deps): update greptime-proto for insert skip_wal

Signed-off-by: WenyXu <wenymedia@gmail.com>

* feat(mito): support request-level WAL skipping

Signed-off-by: WenyXu <wenymedia@gmail.com>

* feat(metric-engine): handle request-level WAL policies

Signed-off-by: WenyXu <wenymedia@gmail.com>

* feat: propagate insert WAL policy through query context

Signed-off-by: WenyXu <wenymedia@gmail.com>

* feat: support session-level insert WAL policy via SET

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(metric-engine): require uniform WAL policy in batch puts

Signed-off-by: WenyXu <wenymedia@gmail.com>

* test(metric-engine): simplify WAL policy coverage

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor(servers): simplify gRPC hint extraction

Signed-off-by: WenyXu <wenymedia@gmail.com>

* test: flatten Mito and Metric WAL scenario orchestration

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor: separate WAL and memtable-only mutations

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor: carry skip-WAL policy in table insert requests

Signed-off-by: WenyXu <wenymedia@gmail.com>

* test: flatten skip-WAL policy cases

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor: clarify WAL notifier naming

Signed-off-by: WenyXu <wenymedia@gmail.com>

* chore: pin merged skip-WAL proto revision

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-09-10 11:40:51 +00:00
Lei, HUANGandCopilot Autofix powered by AI be482553b0 fix(mito2): harden flat merge and use winner_tree (#9064)
* fix(mito2): reject non-flat batches in flat merge instead of panicking

- SortColumns::new becomes fallible try_new: batches missing the
  flat-format internal columns (time index, __primary_key, __sequence)
  at the fixed trailing positions now yield InvalidRecordBatch instead
  of a downcast panic, completing the generic-schema gate that only
  covered BatchBuilder output assembly. Document the flat-format input
  contract on FlatMergeIterator/FlatMergeReader.
- Clarify why BatchBuilder's schema gate uses >= 3 columns when a real
  flat-format schema always has at least 4.
- Add schema-structure tests: empty primary keys (tables without tags),
  dictionary-encoded string tag columns with per-source dictionaries,
  and graceful rejection of batches without internal columns.

Signed-off-by: Lei, HUANG <mrsatangel@gmail.com>

* test(mito2): scan tables with various schemas through flat merge

Add an engine-level test that writes, flushes and scans regions without
tags (empty primary key) and with multiple string tags (dictionary-encoded
in the flat input schema), so the flat merge reader merges an SST with
the memtable on real schemas instead of hand-built batches.

Signed-off-by: Lei, HUANG <mrsatangel@gmail.com>

* refactor(mito2): use winner_tree dependency

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* Update comments for FlatMergeIterator struct

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

---------

Signed-off-by: Lei, HUANG <mrsatangel@gmail.com>
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-09-08 07:20:18 +00:00
XuanwoandWenyXu 4c12ea1aba chore(deps): bump opendal to 0.58.1 (#8742)
* chore(deps): bump opendal to 0.58.1

Upgrade direct opendal dependency and workspace object_store_opendal pin
from 0.57 to 0.58 (lockfile resolves opendal 0.58.1 / object_store_opendal
0.58.0). Adapt to OpenDAL 0.58 composition API:

- Operator::new returns a finished operator; drop .finish() call sites
- Replace HttpClientLayer / raw::HttpClient with OperationContext +
  HttpTransporter (ReqwestTransport)
- Migrate SecureFsBackend and MockLayer from Access/LayeredAccess to
  Service + Layer::apply_service
- Rewrite SecureFs reader/writer/lister for sync factories and StreamRead
- Use OperatorInfo::capability() instead of removed native_capability()

Signed-off-by: Xuanwo <github@xuanwo.io>
Signed-off-by: WenyXu <wenymedia@gmail.com>

* chore: retrigger CI after udeps runner segfault

Signed-off-by: Xuanwo <github@xuanwo.io>
Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(object-store): restore suffix read simulation for secure fs

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: adapt remaining callers to OpenDAL 0.58

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: Xuanwo <github@xuanwo.io>
Signed-off-by: WenyXu <wenymedia@gmail.com>
Co-authored-by: WenyXu <wenymedia@gmail.com>
2026-09-07 08:16:47 +00:00
Weny Xu f9df4def74 chore(deps): switch rskafka to upstream main (#9047)
Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-09-07 07:14:53 +00:00
Weny Xu a932433d21 fix(wal): bound Kafka requests and extend latency buckets (#9026)
* fix(wal): bound Kafka requests and extend latency buckets

Signed-off-by: WenyXu <wenymedia@gmail.com>

* chore(wal): update rskafka request timeout revision

Signed-off-by: WenyXu <wenymedia@gmail.com>

* docs(config): document Kafka WAL timeouts in MetaSrv

Signed-off-by: WenyXu <wenymedia@gmail.com>

* style: sort common-wal dev dependencies

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-09-04 09:30:02 +00:00
LFC 84bd993131 refactor(json2): concretize JSON2 schemas at merge scan boundaries (#9016)
* refactor(json2): concretize JSON2 schemas at merge scan boundaries

Infer concrete JSON2 output types from remote plans and expose them on MergeScanLogicalPlan before physical planning. Recompute affected local schemas and remove the JSON2-specific rewrite from MergeScanExec.

Add SQLness coverage for whole JSON2 columns in windows and joins.

Signed-off-by: luofucong <luofc@foxmail.com>

* fix ci

Signed-off-by: luofucong <luofc@foxmail.com>

---------

Signed-off-by: luofucong <luofc@foxmail.com>
2026-09-04 02:02:24 +00:00
Weny Xu d664326b1e chore: bump version to 1.3.0-alpha.1 (#9014)
* chore: bump version to 1.3.0-alpha.1

Signed-off-by: WenyXu <wenymedia@gmail.com>

* chore: update Cargo lockfile

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-09-03 08:51:39 +00:00
LFC 15317a131b feat(json2): support JSON2 paths in SQL functions (#9007)
feat(query): support JSON2 paths in SQL functions

Update the DataFusion fork to expose scalar function planning hooks.

Infer JSON2 path output types from scalar, aggregate, and window function signatures, while preserving the default Utf8View behavior for functions that accept arbitrary inputs.

Add unit and sqlness coverage for type conflicts, mixed typed and untyped JSON paths, filters, aggregates, and window functions.

Signed-off-by: luofucong <luofc@foxmail.com>
2026-09-03 04:19:55 +00:00
discord9andRuihang Xia 43c30d1446 feat(runtime): add weighted workload scheduler (#8736)
* feat(runtime): add weighted workload scheduler

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* feat(runtime): switch catio to GreptimeTeam fork with admission-wait metrics

Use the GreptimeTeam/catio fork (pinned c20eafc) which adds
ClassStats::total_admission_wait and ClassStats::admitted, recorded
at each QUEUED -> ADMITTED transition. This exposes the scheduler's
own admission delay (excluding Tokio queueing and poll execution),
enabling admission-wait based fairness gates.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: bump catio to dynamic-config revision

Bump the catio scheduler fork to 9f4b028 which adds
Scheduler::set_weight and Scheduler::set_max_concurrent_polls for
runtime configuration.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* feat(perf): runtime-adjustable workload scheduler parameters

Expose dynamic adjustment of the experimental workload scheduler at
runtime:

- common-runtime: set_workload_scheduler_weights and
  set_workload_scheduler_max_concurrent_polls, which forward to the
  catio scheduler's set_weight/set_max_concurrent_polls when the
  scheduler is enabled and reject zero values.
- servers: /debug/workload_scheduler/weights and
  /debug/workload_scheduler/max_concurrent_polls POST handlers, so
  operators can rebalance query/write shares or admission concurrency
  without restarting the datanode.

Both endpoints return 400 with a clear reason when the scheduler is
disabled or the requested value is invalid.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* feat(perf): add GET /debug/workload_scheduler status endpoint

Returns the current weights (per class), max_concurrent_polls,
active_polls and per-class counters (queued, tasks, wakes, polls,
completed, cancelled, admitted, total_admission_wait) as JSON. When the
scheduler is disabled, returns enabled=false with the other fields
omitted, so operators can distinguish 'disabled' from an error.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: bump catio to time-accounting revision

Bump the catio scheduler fork to 257ba56 which replaces
admission-count accounting with real execution-time accounting
(pass += exec_time / (weight * concurrency)), so CPU share follows the
configured weights regardless of poll length.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: bump catio to lock-free sampling revision

Bump the catio scheduler fork to efdc0a4 which adds an optional
downsampled clock sampling mode (SchedulerBuilder::sample_every_polls,
default off) with a lock-free per-class atomic counter, so the
downsampled path costs one fetch_add per poll instead of a global
mutex.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: pin catio to scheduler PR head

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* feat(runtime): add scheduler bypass control

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: advance catio scheduler fixes

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: pin merged catio scheduler

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: regenerate config docs for workload scheduler

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: pin catio scheduler test fix

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(http): satisfy scheduler lifecycle clippy

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: add distributed scheduler toggle coverage

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* feat: finalize workload scheduler runtime controls

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: pin merged catio atomic weights

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: preserve unrelated lockfile resolution

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* perf(runtime): downsample scheduler time accounting

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(runtime): verify cross-runtime scheduler progress

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* feat(runtime): configure scheduler poll sampling

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* docs(runtime): clarify scheduler activation

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* docs(runtime): explain scheduler use case

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
Co-authored-by: Ruihang Xia <waynestxia@gmail.com>
2026-09-02 07:02:39 +00:00
Ning Sun b31f05eb59 fix: update tokio-postgres and correct explain/fetch cursor output schema (#8955)
* chore(deps): update tokio-postgres

* fix: describing fetch cursor and analyze
2026-08-27 03:37:59 +00:00
discord9 3493d2d0fb perf(servers)!: speed up Prometheus JSON response building with ryu and per-series entry reuse (#8815)
* fix(perf): align direct-SST CREATE TABLE with baked index metadata

The offline fixture generator (query_perf_fixture::direct_sst::
build_region_metadata) bakes greptime:inverted_index /
greptime:skipping_index field metadata into the region manifest for
tag/field columns, but create_table_sql emitted a bare CREATE TABLE
without those declarations. MergeScan's remote-schema validation then
failed on any tag/field projection (HTTP 500 'advertised remote stream
schema field mismatch'), breaking direct_readable_sst perf cases.

CREATE TABLE now declares the matching SKIPPING INDEX WITH
(granularity='1') / INVERTED INDEX column options. A round-trip test
proves the emitted SQL is parser-valid and yields the exact catalog
metadata.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* perf(servers): speed up Prometheus JSON response building with ryu and per-series entry reuse

PrometheusJsonResponse::record_batches_to_data spends ~47% of its CPU
in f64::to_string() per sample and ~32% in IndexMap::entry() per row
(60s profile of concurrent query_range workloads, ~800k series).

- Replace f64::to_string() with ryu::Buffer::format_finite for finite
  values (shortest round-trip, 3-5x faster); NaN/+Inf/-Inf keep the
  previous std formatting so wire output is unchanged.
- Remember the previous row's label vector and entry index; query output
  is clustered by series, so consecutive rows reuse the same IndexMap
  entry via get_index_mut instead of rebuilding and hashing the label
  vector (worst case adds one Vec comparison per series transition).

Also adds a query-regression case (prom_json_response) that measures the
real Prometheus HTTP range API path (/v1/prometheus/api/v1/query_range),
which is the only frontend path that builds the Prometheus JSON response
(TQL ANALYZE formats the SQL JSON shape instead), plus a prom_http query
kind in the regression runner.

Perf (aligned base d90cca4b75, 256 series x 481 points):
- prom_range_2h (JSON response path): 29.31ms -> 21.46ms (-26.8%)
- tql_range_2h_control (non-JSON path): +1.87% (noise)

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(servers): keep Prometheus wire format for integral floats with ryu

ryu::Buffer::format_finite prints integral values as "1.0", but the
Prometheus JSON wire format (matching std f64::to_string) expects "1".
Strip the trailing ".0" that ryu only emits for integral values; extreme
values keep ryu scientific notation, and NaN/Inf keep std output. Adds
wire-format tests covering 1.0, 0.0, -0.0, 1.5, 0.1, 1e21, 1e30, 1e-7,
f64::MAX, NaN, ±Inf.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(servers): address Prometheus response review feedback

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* perf(servers)!: use ryu for Prometheus sample values

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(cmd): skip Prometheus execution time extraction

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-08-25 05:16:50 +00:00
discord9anddiscord9 8d887ddd00 fix(query): restore columnar group-by for dictionary-encoded tags (#8902)
* fix(query): restore columnar group-by for dictionary-encoded tags

Bump the DataFusion fork to be93ffd85 (feat/dict-group-column-53), which
backports apache/datafusion #23187: DictionaryGroupValuesColumn lets
dictionary-encoded group keys use the columnar GroupValuesColumn fast
path (hash distinct dictionary values once per batch, resolve rows by
key index) instead of falling back to row-based GroupValuesRows.

This fixes the TSBS double-groupby regression introduced by #8541
(preserve dictionary-encoded query labels): v1.2.0-beta.1 scan output
changed tag columns to Dictionary(UInt32, Utf8), which DataFusion 53.1.0
did not support in GroupValuesColumn's supported_type allow-list, so
GROUP BY queries silently dropped to the ~60% slower row path
(time_calculating_group_ids +57%, peak_mem +50%, end-to-end +38%).

Adds an end-to-end integration test (dict_groupby_sst) that flushes a
flat-format SST with dictionary-encoded hostname, runs the tsbs-style
double-groupby query, and asserts correct results with no CastExec
inserted before the aggregate.

Signed-off-by: discord9 <discord9@greptime.dev>

* chore(deps): pin merged dictionary group-by support

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <discord9@greptime.dev>
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
Co-authored-by: discord9 <discord9@greptime.dev>
2026-08-21 03:50:30 +00:00
LFC b96ea86a62 refactor(json2): add JSON2 v2 physical layout primitives (#8901)
* refactor(json2): add JSON2 v2 physical layout primitives

Signed-off-by: luofucong <luofc@foxmail.com>

* resolve PR comments

Signed-off-by: luofucong <luofc@foxmail.com>

* fix ci

Signed-off-by: luofucong <luofc@foxmail.com>

---------

Signed-off-by: luofucong <luofc@foxmail.com>
2026-08-19 02:27:06 +00:00
Lei, HUANG a8924bb95c refactor(udaf): replace uddsketch implementation (#8867)
* refactor(function): replace uddsketch implementation

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* bench(function): compare uddsketch batch ingestion

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* perf(function): avoid copying non-null uddsketch batches

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(function): decode legacy uddsketch states

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(function): harden legacy uddsketch validation

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* format: taplo

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* test: add compatibility tests for uddsketch functions

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-08-14 12:20:04 +00:00
jeremyhi 764c93bf43 perf(query): choose bounded CTE as hash join build side (#8807)
* fix(query): choose bounded CTE as hash join build side

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* perf(query): remove join estimate cap

* test(compat): accept repartition in analyze plan

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-08-13 10:51:53 +00:00
Weny Xu 1af4c33524 refactor(event): separate procedure submission context (#8856)
* refactor(event): separate procedure submission context

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(event): map extensions and forward GC context

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(gc): initialize integration test context

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(test): pass procedure context to DDL helpers

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor: simplify procedure submission contexts

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor(event): separate procedure and query contexts

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(event): clarify procedure context propagation

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(event): preserve procedure submission context

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor(event): move DDL context by value

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(test): retain manual GC event context

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor(event): tighten procedure context API

Signed-off-by: WenyXu <wenymedia@gmail.com>

* chore: update greptime-proto

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-08-13 07:29:00 +00:00
dennis zhuang 546625c45a feat: embedded convention pack for the entity graph (prom/k8s, gen_ai naming) (#8854)
* feat: embed the derivation conventions as data and adopt gen_ai entity naming

Move the co-declared edge vocabulary, the agent-edge vocabulary and the
virtual-destination candidates from Rust consts into an embedded
conventions.yaml (include_str!), parsed once behind a LazyLock and
validated against the entity-type grammar and the closed rel_type set; a
broken file propagates as a plan error instead of panicking. The agent
vocabulary entity types follow the GenAI semantic-convention namespace
as written: gen_ai.agent / gen_ai.model / gen_ai.tool.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* feat: drop the tag requirement for entity identity columns

Entity declarations no longer require id columns to be tag/primary-key
columns; only column existence is validated. Trace pipelines flatten the
identifying attributes (span_attributes.gen_ai.agent.id, ...) into field
columns, so the tag rule locked real trace tables out of declaring
entities while buying no correctness — the read-time derivation works on
any column.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* feat: implicit declarations for well-known prometheus info metrics

Tables stamped signal_type=metric + source=prometheus whose name matches
the conventions.yaml whitelist (kube_pod_info, kube_node_info,
kube_pod_owner, target_info) get implicit entity declarations: k8s.pod /
k8s.node / k8s.workload with name-based identity and target_info's
service / service.instance with the remaining tags as the descriptive
snapshot. The existing co-declared vocabulary then derives runs_on and
part_of from the same rows, so no new edge branch is needed. Explicit
declarations of a type always suppress the implicit one, and the metric
engine's physical table is excluded (it aggregates every logical
table's columns).

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test: cover the prometheus conventions in sqlness and compact the graph cases

Add the whitelisted-info-metric scenario (kube_pod_info, kube_pod_owner,
target_info deriving runs_on / part_of, a non-whitelisted metric
contributing nothing), fold the single-table calls, cross-table pairing
and virtual-node cases into one trace scenario (they exercise the same
union-before-join path), merge the two declaring-metric-table cases, and
reuse one rename probe for both reserved names.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix: reject entity id columns without a stable string form

Review follow-ups: the DDL check now validates against the schema and
rejects binary-backed and nested types for identity columns (the
derivation renders ids via CAST to Utf8, so the failure used to surface
only when the graph was scanned); the agent sqlness case keeps its
identity columns as fields to cover the relaxed tag rule end to end;
stale tag-rule comments and a dangling const reference are cleaned up.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix: type-check every entity column role, not only ids

The registry renders scope and descriptive values through the same
CAST-to-string path as ids, so a binary-backed column in any role fails
at scan time; the DDL check is now role-independent (and simpler).

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor: name the code-anchored vocabulary constants

Entity types and edge attributes the derivation code itself anchors on
(service, gen_ai.agent, calls, trace/attribute provenance) become
constants in the conventions module; the rest of the vocabulary stays
YAML-only data. ImplicitEntity is renamed PromImplicitEntity, and the
implicit-declaration path logs each skip of a whitelisted info metric
(wrong stamps, suppressed by an explicit declaration, missing id
column) so a missing graph entity is diagnosable.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor: single-source the graph constants

The graph tables' column names move to common-catalog (the schemas
catalog exposes and the plans operator builds must match column by
column), and the conventions module now carries the complete built-in
vocabulary — entity types, rel_types, provenances and connection types —
with the embedded YAML validated by membership against it, so an edit
drifting outside the vocabulary fails the conventions test instead of
deriving nothing.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix: treat empty identity components as absent

kube-state-metrics emits empty-string labels an entity id must not be
built from: an unscheduled pod's node and an owner-less pod's owner_kind
/ owner_name. Standard Prometheus drops empty labels (they arrive as
NULL and the existing predicate handles them), but other remote-write
agents may keep them, which produced ghost entities with empty ids and
false runs_on / part_of edges. Every identity predicate (registry,
co-declared edges, span endpoints) now requires non-NULL and non-empty
components through one shared helper.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor: tighten the conventions DSL semantics

Rename the co-declaration rule lists to what they are (co_declared_edges
/ trace_co_declared_edges — derivation rules, not a relation
vocabulary), stop overstating the GenAI entity types (Greptime types
derived from GenAI attributes; OTel defines no model/tool entities),
move target_info's descriptive snapshot to service.instance (the
remaining labels are the target's resource attributes, and instances
would write conflicting snapshots onto the logical service), and extend
the descriptor whitelist with the stable KSM sources: container info
metrics (closing the k8s.pod contains k8s.container rule),
kube_service_info (new k8s.service entity type) and the fuller
descriptive label sets.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix: guard entity column types on ALTER as well

ALTER MODIFY COLUMN could change a declared entity column to a type
without a stable string form, deferring the failure to graph scan time;
verify_alter now checks the post-alter schema. Dropping a declared
column stays allowed — the read-time derivation skips the stale
declaration, and semantic options cannot be altered off yet.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* feat: bridge traces and kube-state-metrics on the pod UID

Trace-v1 tables now get implicit declarations from their flattened
resource attributes (otlp_trace_entities in conventions.yaml): the
service identity — replacing the hardcoded fallback — plus
service.instance and k8s.pod, each applied only when its columns exist.
A new co-declared rule derives service.instance runs_on k8s.pod, and
the whitelisted kube-state-metrics pod identity switches from
namespace+pod names to the UID, so the trace-side pod and every KSM
descriptor land on one entity while names stay descriptive. This also
removes pod identity from the multi-cluster same-name collision.

The conventions rejection tests were passing for the wrong reason (a
half-renamed fixture key failed deserialization before reaching any
validation rule); they now assert the specific error each case targets.
Sqlness covers the UID merge across descriptor tables, pod-contains-
container, the k8s.service node, and the empty-uid/empty-node rows
deriving nothing.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test: cover the OTLP-to-graph chain end to end

One real OTLP export must come out of semantic_relationships as the
zero-configuration chain: service calls service, instance part_of
service, instance runs_on pod (bridged by k8s.pod.uid). Resources
without service.instance.id or k8s.pod.uid derive nothing extra.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix: identify k8s.service by UID

Same reasoning as pods: a recreated same-name service must not merge
into the old entity and same-named services across clusters must not
collide; kube_service_info carries a stable uid and nothing joins on the
service's name. Also drop a stale tag-rule mention from the option
validation docs.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore: cut duplicated test coverage and redundant comments

The trace service-fallback test collapsed into the resource-entities
test (same synthesis path since the fallback moved to YAML; only the
invalid-explicit-no-fallback case was distinct), role-duplicate and
subsumed DDL cases are gone, the embedded-conventions test is just the
parse (its assertions were decorative), and the YAML section comments no
longer restate the struct docs.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-08-13 02:31:56 +00:00
discord9 c851430349 chore(deps): bump datafusion to 452cb4b (support Dictionary literals in substrait) (#8839)
Bump the GreptimeTeam/datafusion fork rev from 6d6ae9a to 452cb4b,
which includes fix(substrait): support Dictionary literals in producer.

This fixes flow queries against dictionary-encoded PK string columns
(metric tables) failing with:
  Failed to encode DataFusion plan:
  NotImplemented("Unsupported literal: Dictionary(UInt32, Utf8(...))")

The substrait producer now encodes ScalarValue::Dictionary as its inner
value wrapped in a cast to the dictionary type, so the original SQL
works without CAST workarounds.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-08-11 07:16:06 +00:00
Yingwen 78084a9d44 feat: add admin function to discard unflushed data (#8768)
* feat: add admin function to discard unflushed data

Signed-off-by: evenyag <realevenyag@gmail.com>

* test: cover discarding unflushed data by table

Signed-off-by: evenyag <realevenyag@gmail.com>

* chore: fix license header

Signed-off-by: evenyag <realevenyag@gmail.com>

* fix: reject discarding logical metric table data

Signed-off-by: evenyag <realevenyag@gmail.com>

* refactor: defer table name formatting in error paths

Signed-off-by: evenyag <realevenyag@gmail.com>

* chore(deps): update greptime-proto revision

Signed-off-by: evenyag <realevenyag@gmail.com>

* refactor: rename discard unflushed admin function

Signed-off-by: evenyag <realevenyag@gmail.com>

---------

Signed-off-by: evenyag <realevenyag@gmail.com>
2026-08-10 12:19:27 +00:00
shuiyisongandfys d4af650ec0 perf: reduce cold workspace compile time (#8801)
* refactor: remove datanode and meta-srv dep from frontend

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* refactor: use on-device protoc if possible

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* refactor: remove unused dep

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: CR issue

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: version and docs

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* chore: update logs

Co-authored-by: fys <40801205+fengys1996@users.noreply.github.com>

---------

Signed-off-by: shuiyisong <xixing.sys@gmail.com>
Co-authored-by: fys <40801205+fengys1996@users.noreply.github.com>
2026-08-10 07:18:16 +00:00
Ning Sun 606e5fd6a0 feat: update opentelemetry family to 0.32 series (#8776)
* feat: update openetelemtry family to 0.32 series

* chore: resolve warning

* fix: update tests
2026-08-10 03:42:53 +00:00
Ning Sun 62e70027a4 chore: bump version to 1.3.0 on default branch (#8749)
chore: dump version to 1.3.0 on default branch
2026-08-06 08:36:56 +00:00
jeremyhi 9f724aa5f6 feat: make frontend heartbeat extensible and lifecycle-safe (#8726)
* feat: make frontend heartbeat extensible and lifecycle-safe

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: isolate heartbeat extension response handlers

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: cancel in-flight heartbeat response handling

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test: cover heartbeat wire compatibility

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: clean up failed heartbeat startup

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: address frontend heartbeat review feedback

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-08-05 04:38:54 +00:00
jeremyhi 448f973593 fix: sandbox SQL local filesystem access (#8708)
* fix: sandbox SQL local filesystem access

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: address local file sandbox review findings

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: support Windows local copy paths

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: improve sandbox path errors

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* refactor: simplify local path error context

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* perf: stream secure filesystem listings

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* style: derive local file access default

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: improve local file access errors

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: address local file access review findings

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test: simplify local file access coverage

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: harden sandboxed local file backends

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: reject directory copy targets before creation

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: avoid implicit string clone in file table listing

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-07-31 13:23:15 +00:00
discord9 64c66f0a6b chore(deps): pin timestamp widening preimage fix (#8686)
Signed-off-by: discord9 <discord9@163.com>
2026-07-30 02:29:35 +00:00
Lei, HUANG f524a0b5b4 feat: support time range in manual compaction (#8669)
* feat: support time range in manual compaction

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: reject overflowing compaction range alignment

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: preserve range across compaction continuations

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* perf: use graph traversal for compaction windows

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* chore: bump proto to commit on main

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* docs: explain compaction window dependency closure

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-07-29 13:33:10 +00:00
discord9 25217100b0 chore: pin DataFusion cardinality effect (#8562)
Signed-off-by: discord9 <discord9@163.com>
2026-07-22 12:37:09 +00:00
Lei, HUANG ed8a4990fe feat(flow): handle time_ranges in DirtyWindowRequest (#8582)
* feat(flow): handle time_ranges in DirtyWindowRequest

Bump greptime-proto to include the new `time_ranges` field on
DirtyWindowRequest (GreptimeTeam/greptime-proto#330) and mark the
corresponding aligned time windows as dirty in the batching engine,
in addition to the existing per-timestamp dirty marking.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* style(flow): fix doc comment spacing in align_time_window

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* test(flow): cover time_ranges in handle_mark_dirty_time_window

Verify a valid [start_inclusive, end_exclusive) range is aligned to
time window boundaries and stored with an explicit end, and that empty
or reversed ranges are skipped.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(flow): union merged dirty windows with the larger end

Merging a bounded dirty range with a window contained in it (e.g.
[0s, 15s) with nested [5s, 10s), or an unbounded dirty window inside a
bounded range) previously assigned the contained window's upper bound,
shrinking the merged window and permanently dropping the tail range
from re-computation. Keep max(prev_upper, cur_upper) instead.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(flow): clip bounded dirty ranges at the expire bound

merge_dirty_time_windows dropped every window whose start is before
expire_lower_bound, so a bounded dirty range crossing the bound (e.g.
[0s, 15s) with expire 10s) lost its still-live suffix [10s, 15s). Now
bounded ranges are dropped only when their end is at/before the expire
bound, and crossing ranges are clipped to the bound (which the caller
aligns to the time window boundary). Unbounded windows keep the
existing start-based behavior.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(flow): fall back to full dirty on dirty-window alignment failure

An eval/alignment error previously aborted the per-task dirty-marking
closure, losing every dirty timestamp and range accumulated for that
task, while the RPC still returned Ok so the producer would not retry.
On alignment failure now log a warning and mark the whole task dirty
(set_dirty) instead, so the affected data is conservatively
recomputed.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* style(flow): apply rustfmt to new dirty-window merge tests

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* test(flow): cover time index units in dirty window marking

Document that DirtyWindowRequest timestamps/time_ranges are bare i64s
interpreted in the source table's time index native unit, and add a
test expressing the same [3s, 11s) range in second/millisecond/
microsecond/nanosecond units across four tables, asserting all align
to the same dirty window [0s, 15s).

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* chore: bump greptime-proto to 8127f179

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(flow): return dirty-window alignment errors to callers

Do not acknowledge a DirtyWindowRequest when a time-windowed task cannot
align a timestamp or range. The previous conservative fallback used
set_dirty(), but that marker only represents a single epoch-start window
for time-windowed flows, so it could still lose the affected dirty
range. Propagate task errors through the join loop instead so producers
can retry.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* chore: update proto to commits on main

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-07-21 12:56:01 +00:00
dennis zhuang 67683cef2e feat: support SCRAM auth for Postgres (#8304)
* feat: support SCRAM auth for Postgres

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* feat: add pg_scram_sha256 format to hash-password command

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix: harden Postgres SCRAM auth

- Verify the client-final nonce matches the server-issued nonce, per RFC 5802
  transcript validation, instead of only checking the channel-binding field.
- Replace the per-connection PBKDF2 over a random password for unknown users
  with a deterministic mock verifier keyed by the username and a process-wide
  secret. This avoids a CPU-exhaustion DoS on unknown usernames and removes a
  username-enumeration oracle: the SCRAM server-first salt and iteration count
  are now stable per username and indistinguishable from a real user, with no
  PBKDF2 cost and random keys that never accept a proof.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* style: format PG_SCRAM_MOCK_SECRET declaration

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix: precompute stable SCRAM verifier for plaintext users

Plaintext-backed credentials derived a Postgres SCRAM verifier on the fly
on every connection, using a fresh random salt and running PBKDF2 each
time. That made a known plaintext user distinguishable from stored-hash
and unknown (mock) users through both the unstable server-first salt and
the per-connection timing, enabling username enumeration.

Precompute the SCRAM verifier once at load time (stable salt, default
iteration count) and reuse it, matching the mock verifier handed to
unknown users. Document that non-default iteration counts remain
observable in the SCRAM handshake and weaken enumeration resistance.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix: normalize passwords for Postgres SCRAM

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore: docs

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-07-15 11:49:43 +00:00
discord9 b513dbaf4a fix: reject datanode startup on GC config mismatch (#8509)
* fix: reject datanode gc config mismatch

Signed-off-by: discord9 <discord9@163.com>

* refactor: minimize datanode gc startup check

Signed-off-by: discord9 <discord9@163.com>

* chore: update greptime-proto revision

Signed-off-by: discord9 <discord9@163.com>

* chore: use merged greptime-proto revision

Signed-off-by: discord9 <discord9@163.com>

---------

Signed-off-by: discord9 <discord9@163.com>
2026-07-14 12:48:13 +00:00
Lei, HUANG 56e9158819 feat: prepare soft-drop WAL retirement (#8475)
* fix: flush soft-dropped regions on close

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* feat(mito2): handle flush-on-close race with concurrent in-flight flush

When a region close with `flush_on_close: true` races with an
already-running flush, pass the actual close request (including the
flush_on_close flag) to the DDL handler instead of a default request
so the pending flush is correctly awaited.

Files: `src/mito2/src/worker/handle_close.rs`

Also adds a test verifying that closing with flush-on-close while a
flush is in progress still persists all written data correctly.

Files: `src/mito2/src/engine/close_test.rs`
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* feat: support full WAL retirement

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: complete close request migration

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: finish close request callsites

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor: guard Kafka provider setup behind index collector check

Move Kafka provider initialization and `get_or_insert` inside the
existing `if let Some(collector)` block so these operations are
skipped when no global index collector is configured.

Affected file:
- `src/log-store/src/kafka/log_store.rs`

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* chore: avoid to_vec

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor: replace imperative close-region loop with functional combinators

Transform the region close dispatch in `DropTableExecutor` from mutable
`Vec` and `push` loops to iterator chains with `join_all`, improving
idiomatic Rust style and readability.

- `src/common/meta/src/ddl/drop_table/executor.rs` — rewired datanode
  region-close logic to use `peers.map()` and nested `join_all`, moving
  `node_manager.datanode()` inside the closure to align with the new
  structure

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: decouple Kafka client from WAL checkpoint

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: merge Kafka WAL index checkpoints

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: delegate Kafka WAL retirement to metasrv

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* chore: rebase main and resolve conflicts

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: license header

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* chore: bump proto to commits on main

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor: remove Kafka obsolete-all index changes

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-07-13 12:32:46 +00:00
shuiyisong d667728fde chore!: update promql-parser to v0.10.0, remove holt_winters (#8457)
* chore: update promql parser and fix compile

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: sqlness

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: cr issue

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: cr issue

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

---------

Signed-off-by: shuiyisong <xixing.sys@gmail.com>
2026-07-13 03:27:06 +00:00
jeremyhi f12a1da3de fix: upgrade datafusion fork (#8438)
Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-07-08 06:04:13 +00:00