Commit Graph
1671 Commits
Author SHA1 Message Date
discord9 8af3a04ed7 fix: preserve structured query errors through distributed execution (#9161)
* fix: preserve structured query errors through distributed execution

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: update SQL expectations for preserved query error codes

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-16 06:28:15 +00:00
jeremyhi d7ada1761d feat: export logical tables from Metric physical scans (#9159)
* feat: add physical Metric table exporter

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: drain Metric export writes before cancellation cleanup

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* perf: construct Metric export error context lazily

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* refactor: name the logical table export entry point

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* refactor: share Parquet writer for logical table exports

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: preserve Parquet destinations and cancellation boundaries

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* refactor: clarify logical table export field names

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* refactor: clarify logical table export helper responsibilities

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: validate logical export membership by table ID

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test: simplify logical table export coverage and strengthen assertions

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-09-16 02:48:47 +00:00
discord9andNing Sun 94d7e2c7fc feat!: upgrade DataFusion to 55 (#8555)
* feat!: upgrade DataFusion dependencies to 55

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor: migrate DataFusion 55 APIs

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: preserve table function planning behavior

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: preserve PostgreSQL query compatibility

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: preserve distributed execution plan behavior

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: cover DataFusion 55 behavior regressions

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: update DataFusion 55 SQLness expectations

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: complete DataFusion 55 test API migration

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: address DataFusion 55 CI regressions

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: address remaining DataFusion 55 regressions

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: adapt latest base code to DataFusion 55

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: normalize environment-specific DataFusion 55 plans

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: align final DataFusion 55 expectations

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: isolate DataFusion 55 regression cases

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: preserve empty result schema in timestamp widening

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: preserve JSON source column order

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: use released DataFusion 55 integrations

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: adapt latest execution plan mock to DataFusion 55

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: pin DataFusion recursive schema and date repairs

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(promql): align dictionary temporality match keys

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: retain Greptime DataFusion fork behaviors on version 55

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: restore ordinary function error expectations

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: refresh distributed count compatibility plan

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(query): adapt last-row cast hint to DataFusion 55

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: refresh instant last-row empty results for Arrow 59

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* style: simplify DataFusion expression visitor imports

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: restore sorting and PostgreSQL column-order assertions

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(function): restore primitive numeric coercion signatures

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(function): share geo integer signature types

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: cover timestamp widening overflow boundaries

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: fix decimal coercion regression imports

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(function): preserve scalar count_hash NULL state semantics

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: simplify decimal clamp case type inference

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: retain historical count_hash wrapper result

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: restore timestamp widening equality and IN pruning

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: carry upstream aggregate dynamic filter correctness fix

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: carry upstream null and predicate simplification fixes

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: restore baseline JSON ordering expectations

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: restore histogram JSON ordering expectations

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: refresh empty PromQL range result schemas

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: align native timestamp plan with DF55 decimal display

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: refresh native timestamp SQLness results for DF55

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: regenerate NULL sample empty result headers for DF55

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: use DF55 child replacement API in timestamp regressions

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: expose pushed scan dynamic filters to DF55 producers

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: encode string-backed PostgreSQL OID aliases in binary results

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: verify REGPROC binary and text over PostgreSQL protocol

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: register real PostgreSQL catalogs in server fixtures

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: complete DF55 expression inventories for custom query plans

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: correct RangeSelect expression fixture and column identities

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* ci: wait for Kafka WAL helper deployment rollout

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: update custom storage empty result headers

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: require exact row counts in scan statistics

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: suppress deprecated partition_statistics warning in test

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
Co-authored-by: Ning Sun <sunng@protonmail.com>
2026-09-15 11:42:38 +00:00
dennis zhuang 7cbd20a053 fix(mysql): strip leading comments before the federated statement filter (#9156)
* fix(mysql): strip leading comments before the federated statement filter

JDBC clients prefix every statement with a comment. DataGrip sends
`/* ApplicationName=DataGrip <version> */` in front of each one, and every
pattern in the federated filter is anchored with `^`, so the prefix makes all
of them miss.

Two failures follow. `SET TRANSACTION READ WRITE` reaches the SQL parser and is
rejected. Worse, the DBeaver-specific entries such as
`^(/\* ApplicationName=(.*)SELECT @@(.*))` do match DataGrip's prefix, so
`SELECT @@GLOBAL.event_scheduler` and `SHOW VARIABLES LIKE ...` are absorbed
into a zero-column output, which the MySQL writer sends as an OK packet. The
JDBC driver reports that the statement returned no cursor.

Strip leading whitespace and comments before matching, and drop the
ApplicationName-specific entries the stripping makes redundant. The three that
had no unprefixed counterpart (`SHOW PLUGINS`, `SHOW ENGINES`, `SHOW @@...`)
keep their behaviour as plain patterns.

Also covers gaps found while probing the same path:

- `BEGIN` is absorbed like `START TRANSACTION`/`COMMIT`/`ROLLBACK`.
- `SHOW [GLOBAL|SESSION|LOCAL] VARIABLES|STATUS` parses; the scope is ignored
  because GreptimeDB keeps no global/session split.
- `USER()`, `CURRENT_USER()`, `SYSTEM_USER()` and `SCHEMA()` are registered.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(mysql): keep executable comments, reject multi-statement absorption

Review follow-up on the comment stripping, plus the remaining MySQL
compatibility gaps from #9155.

`/*!...*/` is an executable comment: mysqldump emits its initialization as
`/*!40101 SET NAMES ... */`, and the patterns match those verbatim. Stripping
it left an empty statement, so nothing matched and the original SQL reached the
parser, which rejects `SetNames` and `MultipleAssignments`. Leave executable
comments in place.

Absorbing a request also has to stop at a statement boundary, because every
pattern ends in `(.*)`. `BEGIN; INSERT INTO t VALUES (1)` used to fail on the
unsupported `BEGIN`; once `BEGIN` became absorbable the whole request would
report success and write nothing. A request is now scanned for a second
statement and handed to the query engine if it has one. The scan skips plain
comments and string literals so a `;` inside either is not a boundary, and
treats a `/*!...*/` that is not the request itself as a statement, since it
carries SQL.

`check()` now dispatches on the leading statement keyword, so INSERT, UPDATE,
CREATE and ordinary SELECTs run no regex at all — this replaces the narrower
hand-rolled INSERT shortcut. The statement scan runs only for a request the
patterns already matched, which is always a short one.

New compatibility surface:

- `SELECT CURRENT_USER|SESSION_USER|SYSTEM_USER|USER` without parentheses, and
  `SELECT @var`, are answered here rather than in the parser. Both are anchored
  to the whole statement, so `SELECT user FROM t` still reads the column.
- `information_schema.plugins`, `user_privileges` and `processlist` are
  registered as empty tables, like the other MySQL-shape tables around them.
  Sessions are reported through `information_schema.process_list` and
  `SHOW PROCESSLIST`; `processlist` carries the column shape only.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-15 08:59:30 +00:00
Yingwen 23a1a5ecaa refactor: remove constant vector and replicate operation (#8999)
* refactor: remove constant vector

Signed-off-by: evenyag <realevenyag@gmail.com>

* refactor: remove vector replicate operation

Signed-off-by: evenyag <realevenyag@gmail.com>

* fix: preserve scalar vector types with optional type hints

Signed-off-by: evenyag <realevenyag@gmail.com>

* fix: remove obsolete mutable vector helper and import

Signed-off-by: evenyag <realevenyag@gmail.com>

* fix: preserve typed nulls in struct scalar conversion

Signed-off-by: evenyag <realevenyag@gmail.com>

---------

Signed-off-by: evenyag <realevenyag@gmail.com>
2026-09-14 09:25:33 +00:00
jeremyhi 0059508e29 perf: avoid directory listing for explicit COPY input files (#9126)
* perf: avoid directory listing for explicit COPY input files

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: preserve COPY input symlink filtering

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-09-14 07:52:25 +00:00
discord9 13c69cdaee perf(gc): pack file reference exchange (#9009)
* perf(gc): pack file reference exchange

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(gc): address packed reference review feedback

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(gc): stop without retry when maintenance is enabled

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(meta): avoid logging malformed mailbox payloads

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-14 07:30:39 +00:00
discord9 555485c40e perf(promql): avoid per-window allocations in simple range functions (#9104)
* perf(promql): avoid per-window allocations in simple range functions

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* perf(promql): avoid copying smoothing window values

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(perf): cover smoothing copy removal across window layouts

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(promql): skip null samples in simple range functions

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-11 11:59:45 +00:00
Lei, HUANG 21aed0371f feat(pprof): switch CPU profiler to framehop unwinder (#9125)
* feat(pprof): switch CPU profiler to framehop unwinder

Replace the default libgcc-based unwinder in pprof-rs with the
framehop unwinder, which is designed to be async-signal-safe:

- framehop performs no heap allocation during unwinding
  (MustNotAllocateDuringUnwind)
- It handles prologue/epilogue interruption correctly
- It falls back to frame-pointer unwinding when CFI is unavailable
- It does not depend on libgcc's unwind implementation, which is
  documented as not signal-safe (see tikv/pprof-rs#36)

Bump pprof from 0.14 to 0.15 in all three consumers (common-pprof,
cmd, servers) to unify on a single version. pprof 0.15 also replaces
parking_lot with spin-rs to avoid a potential profiler deadlock (#268).

This addresses the libgcc_s.so.1 #GP crash observed in production
when CPU profiling is active, by eliminating the libgcc unwinder
from the signal handler path entirely.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(pprof): gate framehop-unwinder to supported targets

framehop-unwinder is only available on x86_64/aarch64 Linux/macOS.
Enabling it unconditionally for all Unix targets breaks the build on
riscv64 and other platforms: pprof disables its backtrace-rs fallback
when framehop-unwinder is set, but the framehop module is not compiled
on unsupported targets, leaving no TraceImpl implementation.

Split the pprof dependency: the base target.'cfg(unix)' block carries
the common features (flamegraph, prost-codec, protobuf), and a separate
target block adds framehop-unwinder only on supported targets.
Cargo unions features from both blocks on matching targets, so x86_64
and aarch64 Linux/macOS get the full feature set while other Unix
targets fall back to the default backtrace-rs implementation.

Addresses review comment discussion_r3987313632.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-09-11 09:39:13 +00:00
Weny Xu 1c4d3a1729 fix(ci): check Windows test targets before merge (#9116)
* fix(ci): check Windows test targets before merge

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(ci): use standard Windows runner for checks

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(ci): omit dashboard assets from Windows checks

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-09-11 07:04:41 +00:00
Dhruv Vaishnav 83ff0d8138 feat(meta): record catalog and database reconciliation events (#8896)
* feat(meta): record catalog and database reconciliation events

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>

* test(meta): join catalog and database reconciliation events

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>

---------

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>
2026-09-11 07:00:44 +00:00
Lei, HUANG ab0b1f5cce fix: bump jemalloc crates to 0.7 and patch tikv-jemalloc-sys with tcache init fix (#9103)
* fix: bump jemalloc crates to 0.7 and patch tikv-jemalloc-sys with tcache init fix

Upgrade tikv-jemallocator / tikv-jemalloc-ctl / tikv-jemalloc-sys from
0.6 to 0.7, which embeds jemalloc 5.3.1 (includes a056c20d 'Handle
tcache init failures gracefully').

On top of that, patch tikv-jemalloc-sys to the GreptimeTeam fork that
adds the remaining upstream fix 54f22c83 'Initialize TSD tcache before
enabling it' (GreptimeTeam/jemalloc#1, GreptimeTeam/jemallocator#1).
Without the ordering fix, a reentrant allocation during TSD bootstrap
(e.g. heap-profiling prof_tdata init / sampled backtrace when prof:true
is active) can observe an enabled-but-uninitialized tcache, corrupting
per-thread tcache metadata and crashing the process in arena_stats_merge,
calloc, or the libgcc unwinder.

The patch is pinned by rev and should be removed once tikv/jemallocator
ships a jemalloc snapshot that includes 54f22c83.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* chore: bump tikv-jemalloc-sys patch rev to merged release-5.3.1

GreptimeTeam/jemalloc#1 has been merged; point the patch at the
jemallocator commit referencing the merge commit on release-5.3.1.
Jemalloc source content is unchanged.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* chore: point tikv-jemalloc-sys patch at GreptimeTeam/jemallocator main

GreptimeTeam/jemallocator#1 has been merged; reference the merge
commit e1846d8c on main instead of the PR head branch. Jemalloc
source content is unchanged.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(servers): bump tikv-jemallocator dev-dependency to 0.7

main added a target.'cfg(not(windows))'.dev-dependencies entry on
tikv-jemallocator 0.6 for servers after this branch diverged. On the
merge ref it pulled tikv-jemalloc-sys 0.6 from crates.io, which
conflicts with the patched 0.7 (links = "jemalloc" may only appear
once in the dependency graph), failing version selection in CI.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-09-10 13:06:43 +00:00
Weny Xu adda50e03f feat: add repartition partition count hint (#9080)
Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-09-10 05:04:21 +00:00
Lei, HUANG 38577c8854 refactor: support plugin-backed admin functions (#9083)
* feat: add bulk load SQL extension points

Add `CREATE BULK LOAD` parsing and `BulkLoadHandler` dispatch.

Expose private system table factories through `src/catalog`.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor: support plugin-backed admin functions

Replace the custom bulk-load SQL hooks with access to process `Plugins` through `FunctionState`.

Preserve typed external errors from ADMIN functions in `src/operator/src/error.rs`.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-09-10 03:10:17 +00:00
Lei, HUANG 6fa1023b7f feat(mito2): introduce TWCS active window compaction (#9011)
* feat(mito2): support independent TWCS trigger_file_num for active and inactive windows

Split the single TWCS trigger_file_num into per-window-state thresholds:
the active window keeps the existing trigger (default 4, legacy
compaction.twcs.trigger_file_num stays a compatible alias), while
inactive windows use a new trigger (default 2). Inactive windows
additionally fall back from balanced L0-only/L1-only candidates to a
progress-making unbalanced mixed candidate so historical windows can
converge; the active window retains the row/byte balance guards.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(mito2): bound inactive TWCS window convergence by rewrite budget

Inactive windows that cannot compact within one level previously either
stayed stuck (a threshold-qualified but unbalanced level returned no
candidate without trying any fallback) or fell back to a mixed merge
with no balance checks at all, which could rewrite a huge compacted file
to absorb tiny fresh files.

Inactive windows now converge progressively: threshold-qualified
balanced picks, sub-threshold balanced single-level picks, a mixed merge
whose total rewrite must fit in the output file budget, and finally an
L0-only merge without balance checks. Windows that qualify for none of
these are left uncompacted, bounding write amplification.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* test(mito2): cover TWCS window trigger options in alter_table_options sqlness case

Exercise SET/UNSET of compaction.twcs.active_window.trigger_file_num and
compaction.twcs.inactive_window.trigger_file_num end to end, including
that setting the canonical active key removes the legacy
compaction.twcs.trigger_file_num alias from the table options.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(mito2): derive TWCS active window from the max-sequence file

The active window was determined by the max event-time window among
level-0 files. Between an L0 compaction removing its inputs and the next
flush landing, level 0 is empty, so the active window transiently became
None and every window fell back to the inactive rules - triggering
full-window convergence merges during ongoing ingestion whose outputs
are then superseded by new data.

Flush and compaction outputs both inherit the max input sequence, so the
file with the highest sequence across all levels always tracks the most
recent write. Use its window as the active window, falling back to the
previous L0-based rule when no file carries a sequence (legacy files).

The new helper deliberately computes window keys with the
assign_to_windows convention (truncate to seconds, then align up),
because the result is compared against window keys produced there; the
older ceil-based helper is kept unchanged for the legacy fallback path.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* feat(mito2): add active-window L1 compaction safety trigger

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* test(compat): cover TWCS active window options

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(mito2): align TWCS window trigger validation

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(mito2): resolve database TWCS trigger aliases

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(mito2): prioritize newer compaction windows

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(mito2): repick serial compaction outputs

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* feat(mito2): configure inactive-window L1 trigger

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* feat(mito2): prioritize TWCS compaction candidates

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(meta): preserve TWCS trigger downgrade compatibility

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(sql): distinguish invalid database option values

Separate database option key and value validation so recognized keys report the invalid value and its constraint. Add parser coverage for invalid, valid, and unknown options.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* docs(meta): explain TWCS legacy key compatibility

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor(mito2): clarify active window trigger field

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(options): normalize TWCS trigger aliases

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(mito2): ignore ineligible files for active window

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(mito2): mark explicit TWCS options as overrides

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(mito2): normalize zero compaction output size

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(options): validate database TWCS trigger values

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* style(store-api): collapse TWCS alias conflict condition

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-09-09 13:02:26 +00:00
Weny Xu b46de8c828 feat(telemetry): add log directory size retention (#8997)
* feat(telemetry): add log directory size retention

Signed-off-by: WenyXu <wenymedia@gmail.com>

* test: update config API logging fixture

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(telemetry): recover log retention state

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(telemetry): handle log retention cleanup errors

Signed-off-by: WenyXu <wenymedia@gmail.com>

* test(telemetry): cover log count retention on rotation

Signed-off-by: WenyXu <wenymedia@gmail.com>

* test(telemetry): cover log directory retention

Signed-off-by: WenyXu <wenymedia@gmail.com>

* perf(telemetry): avoid log filename allocation

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-09-08 07:31:10 +00:00
Dhruv Vaishnav 35f5485974 feat(meta): record logical-table reconciliation events (#8941)
* feat(meta): add logical table reconciliation events

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>

* fix(meta): preserve logical reconciliation progress

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>

* fix(meta): preserve logical region retry progress

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>

* fix(meta): simplify logical reconciliation events

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>

* test(meta): assert logical event values

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>

---------

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>
2026-09-07 08:29:26 +00:00
XuanwoandWenyXu 4c12ea1aba chore(deps): bump opendal to 0.58.1 (#8742)
* chore(deps): bump opendal to 0.58.1

Upgrade direct opendal dependency and workspace object_store_opendal pin
from 0.57 to 0.58 (lockfile resolves opendal 0.58.1 / object_store_opendal
0.58.0). Adapt to OpenDAL 0.58 composition API:

- Operator::new returns a finished operator; drop .finish() call sites
- Replace HttpClientLayer / raw::HttpClient with OperationContext +
  HttpTransporter (ReqwestTransport)
- Migrate SecureFsBackend and MockLayer from Access/LayeredAccess to
  Service + Layer::apply_service
- Rewrite SecureFs reader/writer/lister for sync factories and StreamRead
- Use OperatorInfo::capability() instead of removed native_capability()

Signed-off-by: Xuanwo <github@xuanwo.io>
Signed-off-by: WenyXu <wenymedia@gmail.com>

* chore: retrigger CI after udeps runner segfault

Signed-off-by: Xuanwo <github@xuanwo.io>
Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(object-store): restore suffix read simulation for secure fs

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: adapt remaining callers to OpenDAL 0.58

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: Xuanwo <github@xuanwo.io>
Signed-off-by: WenyXu <wenymedia@gmail.com>
Co-authored-by: WenyXu <wenymedia@gmail.com>
2026-09-07 08:16:47 +00:00
dennis zhuang bb9b7e8778 fix: re-scan stream-backed tables in recursive CTEs (#9039)
* fix: re-scan stream-backed tables in recursive CTEs

A recursive CTE re-executes its recursive term on every iteration, but
DfTableProviderAdapter hands StreamScanAdapter a single-use stream built at
planning time. The second iteration failed with "Stream already exhausted"
for every table served through DataSource::get_stream — information_schema,
pg_catalog, the computed entity-graph tables and numbers.

Keep that stream for the first execution and open a new one over the same
scan request for later executions.

Closes #9037

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor: drop redundant binding in stream factory

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-07 07:42:41 +00:00
dennis zhuang fa794fae7a fix: match system schema names case-insensitively (#9040)
* fix: match system schema names case-insensitively

Database names that arrive over a protocol (the MySQL handshake and
COM_INIT_DB, the Postgres startup parameter, the HTTP `db` parameter, the
gRPC dbname header) never reach the SQL parser, which is what lowercases
unquoted identifiers. Since #8062 stopped lowercasing them wholesale,
connecting to `INFORMATION_SCHEMA` in any spelling but the canonical one
fails with "Unknown database" -- including the `USE <db>` that a MySQL
client turns into COM_INIT_DB.

Fold only system schema names to their canonical spelling, so user schema
names keep the case they were created with. `is_reserved_schema_name` uses
the same match, otherwise a quoted `CREATE DATABASE "INFORMATION_SCHEMA"`
creates a schema shadowed by the system one.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor: hoist system schema names into a const

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-05 08:18:12 +00:00
Weny Xu a932433d21 fix(wal): bound Kafka requests and extend latency buckets (#9026)
* fix(wal): bound Kafka requests and extend latency buckets

Signed-off-by: WenyXu <wenymedia@gmail.com>

* chore(wal): update rskafka request timeout revision

Signed-off-by: WenyXu <wenymedia@gmail.com>

* docs(config): document Kafka WAL timeouts in MetaSrv

Signed-off-by: WenyXu <wenymedia@gmail.com>

* style: sort common-wal dev dependencies

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-09-04 09:30:02 +00:00
discord9 ad7b0ace64 feat(flow): support eval schedule offsets (#8878)
* feat(flow): support eval schedule offsets

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(flow): remove redundant schedule assertion

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(flow): trim eval offset compatibility scope

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(flow): trim eval offset edge coverage

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* docs(flow): trim eval offset comment noise

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(flow): address eval offset review feedback

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(compat): cover Flow eval offset persistence

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-04 02:39:59 +00:00
shuiyisong b86da3d35f feat: support raw OTLP delta metrics (#8970)
* feat: support raw OTLP delta metrics

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: fmt

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* test(promql): update sqlness results for normalized label matching

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: derive temporality label from default column prefix

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* test(promql): add analyze coverage for delta temporality

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix(promql): scope label alignment to temporality marker

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: handle count-only histograms and vector broadcasts

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: exclude temporality marker from entity descriptions

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: use a fixed label for OTLP aggregation temporality

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix(promql): preserve mixed-range semantics for raw delta

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

---------

Signed-off-by: shuiyisong <xixing.sys@gmail.com>
2026-09-03 09:28:11 +00:00
discord9 9c135ebcb3 feat!: stabilize streaming analyze metrics (#8966)
* feat: stabilize streaming analyze metrics

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* feat: expose analyze memory usage

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor: simplify analyze stream handling

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix: preserve analyze stream sequence on panic

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: log analyze stream worker panic

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-02 10:28:36 +00:00
discord9andRuihang Xia 43c30d1446 feat(runtime): add weighted workload scheduler (#8736)
* feat(runtime): add weighted workload scheduler

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* feat(runtime): switch catio to GreptimeTeam fork with admission-wait metrics

Use the GreptimeTeam/catio fork (pinned c20eafc) which adds
ClassStats::total_admission_wait and ClassStats::admitted, recorded
at each QUEUED -> ADMITTED transition. This exposes the scheduler's
own admission delay (excluding Tokio queueing and poll execution),
enabling admission-wait based fairness gates.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: bump catio to dynamic-config revision

Bump the catio scheduler fork to 9f4b028 which adds
Scheduler::set_weight and Scheduler::set_max_concurrent_polls for
runtime configuration.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* feat(perf): runtime-adjustable workload scheduler parameters

Expose dynamic adjustment of the experimental workload scheduler at
runtime:

- common-runtime: set_workload_scheduler_weights and
  set_workload_scheduler_max_concurrent_polls, which forward to the
  catio scheduler's set_weight/set_max_concurrent_polls when the
  scheduler is enabled and reject zero values.
- servers: /debug/workload_scheduler/weights and
  /debug/workload_scheduler/max_concurrent_polls POST handlers, so
  operators can rebalance query/write shares or admission concurrency
  without restarting the datanode.

Both endpoints return 400 with a clear reason when the scheduler is
disabled or the requested value is invalid.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* feat(perf): add GET /debug/workload_scheduler status endpoint

Returns the current weights (per class), max_concurrent_polls,
active_polls and per-class counters (queued, tasks, wakes, polls,
completed, cancelled, admitted, total_admission_wait) as JSON. When the
scheduler is disabled, returns enabled=false with the other fields
omitted, so operators can distinguish 'disabled' from an error.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: bump catio to time-accounting revision

Bump the catio scheduler fork to 257ba56 which replaces
admission-count accounting with real execution-time accounting
(pass += exec_time / (weight * concurrency)), so CPU share follows the
configured weights regardless of poll length.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: bump catio to lock-free sampling revision

Bump the catio scheduler fork to efdc0a4 which adds an optional
downsampled clock sampling mode (SchedulerBuilder::sample_every_polls,
default off) with a lock-free per-class atomic counter, so the
downsampled path costs one fetch_add per poll instead of a global
mutex.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: pin catio to scheduler PR head

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* feat(runtime): add scheduler bypass control

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: advance catio scheduler fixes

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: pin merged catio scheduler

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: regenerate config docs for workload scheduler

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: pin catio scheduler test fix

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(http): satisfy scheduler lifecycle clippy

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test: add distributed scheduler toggle coverage

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* feat: finalize workload scheduler runtime controls

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: pin merged catio atomic weights

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore: preserve unrelated lockfile resolution

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* perf(runtime): downsample scheduler time accounting

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(runtime): verify cross-runtime scheduler progress

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* feat(runtime): configure scheduler poll sampling

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* docs(runtime): clarify scheduler activation

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* docs(runtime): explain scheduler use case

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
Co-authored-by: Ruihang Xia <waynestxia@gmail.com>
2026-09-02 07:02:39 +00:00
Ning Sun 6f1dd0cb76 feat: allow widening the time index column's timestamp unit via ALTER TABLE, mito2 table only (#8894)
* feat: allow widening the time index column's timestamp unit via ALTER TABLE ... MODIFY COLUMN

Previously MODIFY COLUMN rejected the time index column outright. Now the
time index unit can be widened (Second -> Milli -> Micro -> Nano), which is
lossless for data that fits the target unit: historical data in old SSTs is
cast to the new unit on read by the existing schema-compat layer, and
compaction rewrites it lazily. Narrowing and non-timestamp targets remain
rejected; tag columns keep being rejected. Widening is rejected if any
SST's time range would overflow the target unit's i64 range (e.g.
millisecond -> nanosecond beyond year 2262), since the cast would silently
null those values.

Read-path correctness for old-unit SSTs (verified by new engine e2e tests
and sqlness WHERE queries):

- row-group min/max pruning: parquet statistics of a timestamp column are
  raw integers in the file's unit; when the region metadata's type differs
  (also the case for altered field columns), stats are now interpreted in
  the file's type and converted to the expected type before pruning.
  Without this, a new-unit predicate silently pruned whole row groups of
  old-unit files (wrong results, rows missing).
- SST-level simple filters are skipped for columns whose file type differs
  from the expected type; the predicate is applied by the query layer's
  residual filter above the region scan. Also drop the stale
  "timestamp columns cannot change type" debug_assert.
- retry idempotency: a same-type ModifyColumnType on the time index
  validates as a no-op and `need_alter` returns false, so a retried alter
  procedure (region already altered before the previous attempt failed)
  converges instead of aborting forever.

- add TimeUnit ordering and ConcreteDataType::is_timestamp_unit_widening_to
- relax ModifyColumnType validation in store-api and table metadata
- tests: unit tests in datatypes/store-api/table; mito2 engine e2e tests
  (flushed SST + new writes + reopen + retry + overflow + predicate scans,
  cross-unit dedup, mixed-unit SSTs, compaction); sqlness cases incl.
  partitioned table and WHERE filters over old-unit data

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: resolve parquet filter issue

* test: provide sqlness tests

* refactor: drop trivial test cases and shorten comments

Review pass over the branch's additions:

- datatypes: keep a representative subset of the widening-matrix asserts
- table: collapse the three single-branch rejection blocks into one loop
- mito2: drop the boundary gt_eq and post-compaction predicate asserts
  (covered by the exact-filter regression test and sqlness); drop the
  engine-level gt_eq/lt_eq casts (full operator matrix stays in the
  cast_timestamp_unit unit tests)
- sqlness: drop a bare full scan already covered by the filter above it
- shorten function doc comments across datatypes/store-api/table/mito2/
  recordbatch to the essential semantics

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: drop physical prefilter for columns whose file type differs

Follow-up to the review feedback on CompatBatch/prune reader/filter
handling for widened time index units.

Between/InList/IsNull predicates are prefiltered by PhysicalFilterContext,
which builds its physical expression against the FILE's schema while the
predicate literals are in the expected (post-alter) unit. Evaluating them
against an old-unit SST raised a cross-unit comparison error (Timestamp(ms)
>= Timestamp(µs)) that failed the whole scan. Physical prefilter predicates
are best-effort pruning hints (the query layer re-applies them above the
scan), so drop the prefilter when the column's file type differs from the
expected type, mirroring the simple-filter strategy.

Verified: Between and a non-rewritten (large) InList on old-unit data no
longer error and filter exactly end-to-end (sqlness), and no matching rows
are lost at the engine level (engine test).

Signed-off-by: Ning Sun <sunning@greptime.com>

* test: add direct unit tests for stats cast and prefilter drop

The two-step stats cast (reinterpret raw Int64 stats in the file's
timestamp type, then rescale to the expected type) and the physical
prefilter drop on file/expected type mismatch were only covered
end-to-end; add localized unit tests so a regression fails at the
exact site:

- stats.rs (previously no tests): RowGroupPruningStats min/max over a
  hand-built RowGroupMetaData — passthrough with no expected metadata,
  passthrough on same type, and rescale (1000ms -> 1_000_000us, not
  1000us) on a widened expected unit
- reader.rs: PhysicalFilterContext::new_opt keeps a Between prefilter
  when file and expected types match and drops it on unit mismatch

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: tolerate mixed time units range cache key coverage check

* docs: flag mixed-unit hazard in the (unwired) series index

The series index stores per-series min/max ts as raw Int64 in the unit
of the region metadata at write time, and the searcher builds its range
predicates from a single per-region metadata. After a time index unit
widen, files of one region would carry mixed units, so a per-file unit
(or an index rebuild on such alters) is required before this index is
wired into scans. Leave notes at both sites.

Signed-off-by: Ning Sun <sunning@greptime.com>

* test: cover mixed-unit compaction for sparse encoding and strict windows

Compaction-path audit follow-up. The compat cast and window math were
already covered for dense regions; add the two remaining e2e scenarios:

- sparse primary key encoding (used by metric-engine physical regions):
  widening then compacting mixed-unit files rewrites the old-unit time
  index correctly through the sparse compaction compat path
- strict-window manual compaction: each window output trims rows with a
  predicate built in the region's new unit against an old-unit file;
  every instant must survive exactly once (no loss, no cross-window
  duplication), rescaled

Also documents the audit finding that Regular ranged (manual)
compaction never trims rows: TwcsPicker sets output_time_range to None
and the request time range only selects candidate windows.

Signed-off-by: Ning Sun <sunning@greptime.com>

* test: cover time index unit change in FlatCompatBatch directly

The compat layer's rescaling of a widened time index was only verified
end-to-end; add direct unit tests for both paths:

- dense: identical units skip compat entirely; a widened unit rescales
  the time index column (1000ms -> 1_000_000us, not reinterpreted) while
  other columns pass through and the output schema matches the expected
  metadata
- compact sparse (the metric-engine compaction path): same rescaling

Signed-off-by: Ning Sun <sunning@greptime.com>

* refactor: move timestamp unit division into common-time

The exact unit division (UnitQuotient + div_mod_units) is time
semantics, not filter logic; move it next to TimeUnit in common-time
with a compact test covering representable/non-representable values,
negative (floor) instants, and quotient overflow. The ScalarValue
helpers stay in filter.rs since common-time has no datafusion
dependency.

Signed-off-by: Ning Sun <sunning@greptime.com>

* test: compat rescales a widened time index and fills an added column together

The realistic multi-alter sequence (widen at T1, add column at T2,
read a T0 SST) exercises cast and default-fill in the same
compute_index_and_fields pass; assert both in one output batch.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: preflight time index widening overflow before any region alters

Address review feedback on the overflow guard:

- preflight: when the frontend operator receives a widening alter on the
  time index, run an existence scan (ts outside the target unit's i64
  range, LIMIT 1, via the query engine so it covers every region of the
  table in both standalone and distributed modes) BEFORE any DDL task is
  submitted. A region that fits can no longer commit the new schema
  while another region rejects the alter with a non-retryable error.
  File and row-group pruning keep the scan cheap when nothing overflows.
  The per-region check in mito2 stays as the final guard for data
  written after the preflight (the remaining race window); without a
  validate-only wire field (region.proto lives in the external
  greptime-proto repo) a fully atomic two-phase validate/commit is out
  of scope here.

- fast path: cast_timestamp_unit returns the filter unchanged when the
  literal is already in the target unit, skipping the div-mod rebuild.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: address review comments

* fix: address auto review comments

* fix: remove time index widening overflow preflight

Overflow needs timestamps beyond the target unit's i64 range (~year
2262 for nanoseconds), which real workloads never write, so the two
existence scans before every widening alter are not worth the cost.
Region validation already rejects the alter when an SST's time range
overflows the target unit; it now logs the rejection (with the
offending file) and returns a deterministic client-facing message.

Signed-off-by: Ning Sun <sunning@greptime.com>

* test: make sqlness test stable

* fix: log instead of rejecting time index widening overflow

Overflowing values cast to NULL on read but do not otherwise affect
reads or writes, so the alter is allowed; the region-level check now
only logs (with the offending file) when an SST's time range exceeds
the target unit's i64 range.

Signed-off-by: Ning Sun <sunning@greptime.com>

---------

Signed-off-by: Ning Sun <sunning@greptime.com>
2026-09-02 04:29:32 +00:00
discord9 7252ceb4bb feat(flow): add generic delta merge for incremental aggregates (#8938)
* feat(function): add internal delta merge aggregates

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(function): cover delta merge aggregates in SQL

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* feat(function): add Welford delta merge aggregate

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-01 12:27:54 +00:00
Lei, HUANG c6b10bfbb9 feat(function): add mergeable stddev_pop state functions (#8972)
* feat(function): add Welford stddev functions

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* test(function): cover merged Welford time windows

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor(function): use stddev_pop SQL names

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(function): make Welford arithmetic partition-stable

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(function): reject invalid singleton Welford states

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* test(function): pin Welford state compatibility

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* refactor(function): remove unreachable variance clamp

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(function): reject DISTINCT Welford aggregates

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* test(compat): bound Welford downgrade targets

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-08-31 14:15:33 +00:00
dennis zhuang 109257505e feat: report what the graph derives and fix two duplicate-node bugs (#8936)
* feat(semantic-graph): let the generic container yield to k8s.container

A container inside a Kubernetes pod reaches the graph twice: as the
k8s.container entity kube-state-metrics describes, identified by
[pod uid, container name], and as the generic container the OTel resource
attributes describe, identified by container.id. One physical container,
two nodes.

Kubernetes is the primary scenario, so k8s.container keeps its identity
and the generic type stands down where it applies. Conventions gain a
row-level condition for that: `suppressed_by` withdraws a declaration on
rows where any of the named columns has a value. The test has to be per
row, not per table — one descriptor table holds both pod rows and
bare-runtime rows.

Every branch that turns a declaration into rows now shares one guard
(`declaration_predicate`), so the condition cannot apply to entities but
not to the edges they carry.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* feat(semantic-graph): report derived entity declarations in table_semantics

information_schema.table_semantics only read table options, so the
declarations the built-in conventions derive — for trace tables, for
whitelisted Prometheus and OTel descriptor metrics — were invisible.
"Why is my table not in the graph?" was answerable only from debug logs,
which is not a self-service path.

A new `entity_declarations` column reports the entities a table actually
contributes: each one's identity, whether it came from an option or from
the conventions, and any row-level condition attached to it. An expected
entity missing from the list is the answer — the table name is not
whitelisted, the source stamp is wrong, an id column is absent.

The row filter widens to match: a table that declares nothing by option
but derives entities by convention now appears, since it is in the graph
and the view has to say so. The provider reaches the derivation through a
new metadata-only trait method, keeping catalog below operator.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(semantic-graph): never yield a container to an entity nothing derives

The generic container withdrew on any row carrying a pod UID, on the
assumption that k8s.container would cover it. Nothing guaranteed that:
k8s.container came only from kube-state-metrics, so an OTLP-only
deployment — or one whose KSM data had expired or fell outside the query
window — lost the container node and its edges entirely instead of
gaining a more specific one.

The rule now names the superseding type rather than a trigger column, and
withdraws only where that type's full identity is on the row. Both OTel
sources declare k8s.container themselves, under the identity
kube-state-metrics gives it ([pod uid, container name]), so the two
sources name one node; `k8s.container.name` joins the descriptor's
projected attributes to carry it.

Resolution runs once every declaration for the table is known, so a
superseding type that ends up undeclared — its columns are gone, or an
explicit declaration of it was skipped — leaves the guard empty and the
generic container standing. A container can change type; it cannot
disappear.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor(semantic-graph): build the declaration JSON by serializing a type

The column was assembled entry by entry into a serde_json::Map, cloning
every value and allocating a String per key. A Serialize struct that
consumes the declaration moves the same data instead, and the field order
now reads type, origin, identity, then description.

The scan path around it was doing the same kind of avoidable work:
declarations were derived before the predicate could discard the table,
the supersession pass cloned every declaration's identity to look up one,
and the option parse claimed the time index for tables that turn out to
declare nothing.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* feat(semantic-graph): keep the structured identity for single-column ids

`entity_id_attrs` was NULL whenever the identity came from one column, so a
consumer holding `host` = `a3f2...` had no way to tell which column produced
it, and no way back to the source table. Most identities are single-column —
host, k8s.pod, k8s.node, service — so the common case was the opaque one.

Build the JSON object unconditionally. Entity equality still reads `entity_id`
alone, so this changes no merging: it records which attributes the id was
assembled from, beside an id that deliberately omits them.

Also corrects the `entity_id` column doc, which still described the `k=v,k=v`
rendering replaced in #8904 by values joined in declared order.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(semantic-graph): keep what a container carried when it yields

Yielding to k8s.container cost the row two things it had as a generic
container. The runtime container id, which kube_pod_container_info keeps
descriptive precisely because it is the handle back to runtime logs and
metrics, went missing entirely: neither an id nor an attribute of any
node. And the edge vocabulary knew only the generic type, so a row with
the more complete labels ended up with fewer connections than one
without — the container layer no longer reached its host.

Both OTel sources now keep the runtime id and name descriptive on
k8s.container, and the vocabulary gains the two edges that mirror the
generic type's. Nothing checks a supersession against the edge
vocabulary, so that requirement is written down where the rule is.

Also: the conventions-failure path now reports the explicit half as its
comment always claimed, entity_declarations reports scope columns, and
identifies() is private again now that only declaration_predicate calls
it. The two RFCs catch up with entity_id_attrs being unconditional, the
view listing convention-derived tables, and supersession existing.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* feat(semantic-graph): report unmatched clients and the longest request

`real_wins` pins a `calls` edge's RED metrics to the observed span pairs:
when a window's edge key holds any pair, the unmatched client spans are
suppressed so one (window, edge) yields one row. That leaves no way to
tell a callee that stopped responding from traffic that stopped arriving —
both show up as a lower request_count.

Add two columns to `semantic_relationships`:

- `unmatched_count` — client spans with no server span, counted outside
  `real_wins` on the same row, so the suppressed population stays visible
  without splitting the edge into two rows. NULL for agent calls, whose
  inner join leaves nothing unmatched, and for declared edges.
- `duration_max` — the longest single request. It goes through
  `real_wins`: a pair is timed by the server span while an unmatched
  client is timed by its own (network wait included), so mixing them would
  make the max describe a different population than duration_sum and
  duration_count. Agent calls compute it from the child spans they already
  aggregate.

The projection contract goes from 16 to 18 columns; every branch projects
both. Explicit column queries are unaffected, `SELECT *` and
ordinal-based readers see the new shape.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(semantic-graph): drop the duplicated source-gate case

The whitelist gate on `source=opentelemetry` is already covered by
`otel_implicit_declarations_are_gated` and by the wrong-source case in
`table_semantics`; here it only paid for another table create, insert and
full graph derivation.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* style(semantic-graph): trim the comments back to what the code cannot say

Several comments restated the code, repeated a rationale already stated at
the type or in the RFC, or explained a test in more words than the test
body.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-08-31 10:39:17 +00:00
Ning Sun d32cd77505 fix: postgres describe for more statements (#8974)
* fix: postgres describe for more statements

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: cover more show statements

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: address review comments

- add missing `clippy::too_many_arguments` allow on
  `query_from_information_schema_dataframe` (CI clippy failure)
- take `&ShowKind` in the information-schema dataframe helper so `kind`
  is no longer cloned at every call site; only the WHERE arm (which needs
  an owned expression for `sql_to_expr`) clones internally
- document why re-applying TQL explain formats never overwrites an
  existing value (per-query context state)

Signed-off-by: Ning Sun <sunning@greptime.com>

* chore: trim comments to essentials

Signed-off-by: Ning Sun <sunning@greptime.com>

---------

Signed-off-by: Ning Sun <sunning@greptime.com>
2026-08-30 13:24:17 +00:00
Weny Xu c01de4afdc fix(operator): whitelist private system table auto create (#8930)
Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-08-28 08:49:49 +00:00
Dhruv Vaishnav 9198462869 feat(meta): record physical table reconciliation events (#8935)
* feat(meta): record physical table reconciliation events

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>

* docs(config): add reconciliation table event

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>

* fix(meta): address reconciliation event review feedback

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>

* fix(meta): keep reconciliation event summary volatile

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>

* refactor(meta): remove unused table state downcasting

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>

---------

Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>
2026-08-28 06:17:55 +00:00
wy471xandNing Sun aaa843104b fix(mysql): interpret prepared statement datetime params in session timezone (#8923)
* fix(mysql): interpret prepared statement datetime params in session timezone

Binary DATETIME parameters of server-side prepared statements were
converted as if UTC, ignoring the session timezone set via SET time_zone.
Convert them with the session timezone and add an integration test
covering prepared inserts and predicates under Asia/Shanghai.

Signed-off-by: wy471x <wy471x@gmail.com>

* refactor: share naive datetime timezone policy via common-time

Address review feedback on the prepared-statement timezone fix:

- Expose Timestamp::from_naive_datetime in common-time so the DST policy
  (gap -> error, ambiguous -> earlier instant) lives in one place, shared
  by the text protocol (Timestamp::from_str) and the MySQL binary protocol.
- Route the MySQL prepared-statement datetime conversion through it.
- Match the target type before converting datetime params so
  PreparedStmtTypeMismatch fails fast without wasted conversion.
- Use the short Timezone import form for consistency with the rest of servers.

Signed-off-by: wy471x <wy471x@gmail.com>

---------

Signed-off-by: wy471x <wy471x@gmail.com>
Co-authored-by: Ning Sun <sunng@protonmail.com>
2026-08-28 02:40:25 +00:00
Ning Sun b31f05eb59 fix: update tokio-postgres and correct explain/fetch cursor output schema (#8955)
* chore(deps): update tokio-postgres

* fix: describing fetch cursor and analyze
2026-08-27 03:37:59 +00:00
dennis zhuang 6d86e6ff06 feat: synthesize OTLP resource descriptor for the semantic entity graph (#8904)
* fix(servers): compose OTLP metrics job from service.namespace/service.name

The OTel Prometheus compatibility spec defines job as
"<service.namespace>/<service.name>" when the namespace is present.
The OTLP metrics path only used the bare service.name, so the job tag
diverged from target_info produced by Prometheus-side exporters for the
same resource. Compose the namespace form, and keep not fabricating a
job when service.name is absent.

Behavior change: resources carrying service.namespace now get
"namespace/name" as their job tag value.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* feat(otlp): synthesize otel_resource_info at OTLP metrics ingestion

Ordinary OTLP metrics scatter filtered resource attributes as tags over
every logical metric table, so metrics-only services contribute nothing
to the semantic entity graph. Each request now also projects its
distinct resources into one info-metric-shaped mito table,
otel_resource_info: a fixed allowlist of identity-relevant attributes
under their raw OTel keys (independent of the label translation
strategy and the promote/ignore headers) plus derived job/instance
compatibility columns, value 1.0, and the newest data-point timestamp.

The descriptor is written after the main insert is committed; a failure
there (conflicting pre-existing table, auto-create disabled) degrades
to an OTLP partial_success warning with rejected_data_points = 0
instead of failing the request and triggering client retries of
already-accepted data. A request writing a metric named
otel_resource_info suppresses synthesis. Legacy mode is unchanged.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* feat(operator): otel info-metric conventions with host/container entities

Whitelist the ingestion-synthesized otel_resource_info descriptor via a
new otel_info_metrics conventions map, gated on source=opentelemetry
(the existing gate hardcoded source=prometheus). Its declarations use
explicit descriptive lists instead of descriptive_rest so identifying
attributes of other entities do not leak into service.instance.

Conventions tightened per the Astronomy Shop findings: host identity is
host.id with host.name descriptive only (host.name is not stable across
SDKs and resource detectors), a generic container entity (new entity
type) is declared only when container.id is present, and trace-v1
tables now synthesize host/container from their flattened resource
attributes too. New co-declared edges: service.instance runs_on
container, container runs_on host.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(otlp): cover the resource descriptor in integration tests

Covers the descriptor's raw-key columns and info-metric options through
the HTTP path, the namespace/name job composition end-to-end, column
names being independent of the translation strategy, the allowlist
excluding unlisted resource attributes, auto-create after a drop, the
metric-name collision suppressing synthesis, and the partial-success
warning (rejected_data_points = 0) when a pre-existing incompatible
table fails the descriptor write while metric data is accepted.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore: cargo fmt

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(frontend): degrade descriptor permission denial to a warning

A table-level permission policy denying otel_resource_info would have
failed the whole OTLP metrics request because the descriptor's
permission check ran before the main insert. The descriptor is derived
enrichment: check its permission in the degrade path so a denial skips
the write and surfaces as the partial-success warning, like any other
descriptor write failure.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(otlp): guard descriptor writes with semantic ownership markers

A pre-existing schema-compatible table named otel_resource_info would
silently receive descriptor rows while its missing semantic stamps kept
it out of the entity graph. The descriptor write now requires the
auto-created table's ownership markers (mito engine + signal_type +
source + metric.type=info + metadata_quality=declared) and otherwise
degrades to the partial-success warning; the entity-graph gate for the
otel whitelist likewise requires metric.type=info, so a user table
stamped with only signal/source no longer picks up implicit
declarations.

Also fold the descriptor write cost into the response and surface the
degrade warning through the otel-arrow BatchStatus status_message.
Integration tests pin the full marker set on auto-create and that an
existing owned descriptor keeps accepting writes without degrading —
a missing marker would otherwise silently stop every descriptor write
after the first request.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* perf(otlp): build descriptor rows without the per-resource BTreeMap

Projecting a resource allocated a BTreeMap and then collected it into the
row key, and every attribute was matched against the allowlist by linear
scan. Collect the tags into a Vec and sort once, and match the allowlist
instead of scanning it. Measured on the conversion path: descriptor work
drops 16-18%, from 10.6% to 8.9% of conversion CPU on the worst shape
(1000 resources with 4 data points each), where the cost tracks resource
count rather than data-point count.

Also trims the comments and tests added with the descriptor to what
carries information.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(otlp): pin the descriptor permission-denial degrade path

A policy denying the descriptor table must not fail the metrics request,
which the fix in 2401b3dd9c does but nothing covered. Verified as a
regression guard by mutation: moving the permission check back before
the main write makes this test fail.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(otlp): keep legacy mode covered after trimming the unit tests

Trimming the descriptor tests dropped the only assertion that legacy
mode skips the job/instance remap and the promote filter. Both alter
the columns of tables already in use, so fold the check into the legacy
conversion test rather than leaving it uncovered.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(semantic-graph): stop encoding column names into composite entity ids

A composite entity id rendered the identifying columns as sorted
`col=value` pairs, so the same identity split into one entity per
signal: a trace table names its columns service_name and
resource_attributes.service.instance.id where a metric table names them
job and instance. One service instance became two nodes with two
parallel edge sets, breaking the walk from a trace to that instance's
metrics.

Render an id as its values in declared order instead, escaping the
separator so components stay distinguishable, which is what single-column
ids already did by keeping only the value. entity_id_attrs still carries
the structured form.

Values alone are not enough for a namespaced service: the metric side
folds service.namespace into job while traces keep the bare name. Add
qualified_by to the conventions so the trace declarations compose the
namespace the same way, per the OTel rule that job is
<service.namespace>/<service.name> or the bare name when the namespace
is empty. A table without the namespace column keeps the unqualified
identity rather than losing the declaration.

Conventions validation now rejects one entity type declared with a
different number of id columns by two sources, which would silently
produce ids that can never match.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* style(otlp): import the parent module by crate path

check-super-imports.py, part of the CI format gate, rejects a
file-level `use super::`.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* feat(otlp): gate the resource descriptor, and fix what review found

Synthesizing greptime_otel_resource_info creates and writes a table the
user never sent, so it is now off unless
otlp.experimental_enable_resource_info says otherwise. With it off the
request costs exactly what it did before the descriptor existed: nothing
is projected, no table is created, no write and no permission check
happen. Tests run with it on. StandaloneOptions carried no otlp field,
so the whole [otlp] section was silently dropped in standalone mode; map
it through, or the new option (and trace_ingest_chunk_size before it)
would do nothing there.

Renamed from otel_resource_info: the greptime_ prefix marks the table as
engine-managed and makes a collision with a user metric unlikely, which
is what the pre-existing-table ownership check and its per-request
catalog lookup were defending against. Both are gone.

A request may carry data for several graph windows, but the descriptor
folded every data-point time into one row at the newest of them, leaving
the earlier windows with metric rows and no entities. Key the rows by
window as well, and take the times from the data points the encoder
actually writes: it drops exponential histograms, and a resource
carrying nothing else was being described as an entity with no
measurements.

Projecting a resource cloned its attributes once per data point. Nest
the windows under the attributes instead, so they are moved once per
resource, and walk the data-point times through a visitor rather than
collecting a Vec per metric.

Also documents what the two maps key and hold, and lifts the projected
attribute names to constants beside KEY_SERVICE_NAME.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(otlp): skip the descriptor's work entirely when it is disabled

The collision scan over the request's output tables ran even with the
feature off. Short-circuit on the option instead, and update the config
snapshot the new [otlp] section changed.

Also drops the doc comment orphaned by the deleted ownership check: it
had attached itself to the trait impl and described a check that no
longer exists.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor(semantic-graph): drop the expect and name the service identity

The CASE is built through Case directly rather than the fallible
when().otherwise() builder, so the non-test path no longer carries an
expect (architecture-invariants $4).

service_identity returned two same-typed Options that both call sites
destructured positionally; a named struct makes a swap fail to compile.

Also records that id-column order is part of the identity, where the
option docs and the conventions authors will read it.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore(semantic-graph): drop comments that narrate the code

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(semantic-graph): cast duration_nano before the trace-table union

Trace tables written before the signed-integer ingest change hold
duration_nano as UInt64 and later ones as Int64. The calls derivation
unions the per-table selects, and the two have no common integer type,
so a deployment holding both shapes could not build the plan. The
cross-table test now spans both.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor(semantic-graph): drop the redundant duration_nano casts

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(otlp): decide exponential histogram acceptance in one place

The resource descriptor mirrored only the experimental gate, so with both
experimental flags on a resource whose only metric is a delta exponential
histogram was described as an entity with no measurements. The encoder's
whole-metric rules move into exponential_histogram_gate, which both call,
and the descriptor takes its timestamps through exponential_histogram_value
so per-point rejections drop out too.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-08-25 13:08:44 +00:00
dennis zhuang 1851f6bf4d fix(query): keep INSERT timestamp conversion out of the source query (#8911)
* fix(query): keep INSERT timestamp conversion out of the source query

Interpreting an INSERT's string timestamps used to work by pushing the
conversion down into the source query, which changed what that query
means. Two consequences:

- Pushing through a UNION's DISTINCT moved the dedup key from the raw
  strings to parsed instants, so rows spelling the same instant
  differently collapsed into one. On an append-only table that is a
  silently dropped row.
- A UNION branch that needed no conversion (a NULL, or an explicit cast)
  made the whole column give up, leaving sibling branches on UTC while
  the rest of the row used the session timezone.

Convert at the assignment instead, by routing its cast through a
timezone-carrying timestamp type and back. Arrow applies the timezone
when a cast target carries one, and stripping it afterwards preserves
the value. The source query is no longer touched, so both cases go away
and the tree-walking rewrite (roughly 160 lines) is deleted.

The rewrite reads source types, so it now runs TypeCoercion first: a
UNION still carries its loose per-branch schema before coercion, and
retargeting a cast whose input later becomes a timestamp would shift the
value rather than reinterpret it.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(query): address review on INSERT assignment rewrite

- Clone the input `Arc` instead of the whole subtree, and only rebuild it
  when a `Values` row actually changes.
- Defer cloning the cast source until the literal-folding path has been
  ruled out.
- Move the UTC check onto `Timezone::is_utc`, replacing a bare string
  compare.
- Cover a prepared `INSERT ... VALUES (?)`: an untyped placeholder types
  as `Null`, so the assignment cast is left for parameter substitution.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-08-25 12:34:53 +00:00
Lei, HUANG 73c4140938 feat(function): expose uddsketch rank (#8929)
* feat(function): expose uddsketch rank

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* test(function): cover uddsketch rank in sqlness

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(function): support legacy uddsketch rank

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* docs(function): document uddsketch rank behavior

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-08-25 11:31:47 +00:00
LFC 932f87f7a8 refactor(json2): support querying v2 storage layout (#8940)
* feat(json2): support querying v2 storage layout

- route missing JSON2 paths to the v2 remainder field
- reconstruct complete values from explicit fields and remainder data
- preserve root JSON2 columns across projections and filters
- support nested JSON values in json_get string results
- add and reorganize JSON2 sqlness coverage

Signed-off-by: luofucong <luofc@foxmail.com>

* resolve PR comments

Signed-off-by: luofucong <luofc@foxmail.com>

---------

Signed-off-by: luofucong <luofc@foxmail.com>
2026-08-25 11:11:06 +00:00
shuiyisong 04614175fe refactor: centralize native histogram encoding in common-query (#8945)
* refactor: centralize native histogram encoding in common-query

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: fmt

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

---------

Signed-off-by: shuiyisong <xixing.sys@gmail.com>
2026-08-25 06:26:09 +00:00
shuiyisong 82444635f5 feat: support quantile and fraction queries on mixed histograms (#8874)
* fix: test

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: remove try_build_float_literal

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: remove NonCommutative

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: NaN

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* chore: add comments

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* test(query): cover frontend-only histogram fold planning

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix(promql): ignore unparseable histogram bucket labels

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: typo

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix: sqlness

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

---------

Signed-off-by: shuiyisong <xixing.sys@gmail.com>
2026-08-24 12:45:35 +00:00
Lei, HUANG 2182dccd9b fix: cap default runtime sizes to a minimum of 2 threads (#8908)
* fix: cap default runtime sizes to a minimum of 2 threads

RuntimeOptions derived its default sizes directly from num_cpus. On
single-core machines every runtime (global, compact, query, ingest)
ended up with one worker thread, which can easily deadlock async code
(e.g. block_on combined with spawn).

Clamp all CPU-derived runtime sizes to at least 2 threads.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: init logging before runtimes so runtime options are logged

The global runtimes were initialized before the global logging
subscriber, so the "Creating runtime ..." info logs that carry the
runtime sizes were silently dropped. Initialize logging first in all
node start paths; common-telemetry has no dependency on
common-runtime, so the reorder is safe.

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-08-19 10:04:50 +00:00
dennis zhuang 09c0b23a23 feat: manage semantic table options via ALTER TABLE SET/UNSET (#8880)
* fix(meta): actually acquire logical table locks in alter-logical-tables procedure

The procedure listed its logical table locks from table_info_values,
which is only filled during Prepare, while procedure lock keys are
fixed at submission — so the logical locks were never acquired. Today
every writer of a logical table's info is serialized by the physical
table lock, which hides the problem; a metadata-only alter procedure
targeting a single logical table would race it.

Resolve the logical table ids at submission, persist them in the
procedure state (serde(default): state dumped by older versions keeps
the previous behavior), lock physical + logical tables, and re-check
the resolved ids against the locked set at Prepare so a table dropped
and recreated after submission cannot be mutated without a lock.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* feat: manage semantic table options via ALTER TABLE SET/UNSET

CREATE TABLE accepts greptime.semantic.* options, but ALTER TABLE SET
routed every option through SetRegionOption, whose closed match
rejects them — tables auto-created by ingestion could never receive
semantic declarations after the fact.

Semantic options are pure metadata markers no region consumes, so
they now take a metadata-only alter, following the repartition-hint
precedent:

- New AlterKind::SetAnnotations/UnsetAnnotations carrying an
  AnnotationFamily (currently only Semantic), so future marker-style
  option families reuse the same machinery. The converter classifies
  a SET/UNSET batch by key prefix and rejects batches that mix
  annotation keys with regular options.
- The procedure reuses the MetadataOnly flow: no region dispatch,
  table-info update plus cache invalidation only.
- Validation lives in the table-meta mutation layer, so it runs at
  frontend verification and again in the procedure's prepare step
  under the table lock: SET is strict (known key, value domain,
  entity columns exist and render as strings); UNSET is lenient
  inside the namespace so stale keys can be cleaned up.
  ModifyColumnTypes re-checks columns referenced by entity
  declarations at the same layer, closing a verify-then-execute race.
- Logical metric tables are supported: an annotation alter submits a
  regular alter-table task locking only the logical table, and the
  DDL manager's physical-route guard admits it.
- create_table_info re-checks semantic value domains for gRPC-built
  expressions that bypass the SQL parser.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor(table): centralize annotation option classification and validation

Address review feedback on the AnnotationFamily abstraction: with only
one variant that every consumer immediately destructured, the
generality was fake. Make it real and exhaustive instead:

- AnnotationFamily gains RepartitionHint: repartition.column.hint is
  the same kind of marker option (pure metadata, no region consumes
  it) and previously had a hand-rolled special case in the converter,
  the metadata-only classifier, and a dedicated AlterKind pair — all
  deleted, one classification API remains. Per-family logical-table
  eligibility (allows_logical_tables) replaces the hard-coded
  Semantic check in the DDL manager guard.
- One validation core in the table crate (check_annotation) serves
  both DDL entry points. CREATE and ALTER previously duplicated the
  rules; each keeps its existing error variants, status codes and
  messages via thin adapters over a typed error (ALTER missing column
  stays 4002 TableColumnNotFound, CREATE stays InvalidArguments).
- The batch classifier returns Result instead of swallowing the
  mixed-batch error: a mixed SET on a logical table now reports the
  actual problem instead of UnexpectedLogicalRouteTable, and the flow
  classifiers propagate instead of guessing. The converter also moves
  its owned payloads instead of cloning them.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(meta): cover logical-table annotation alter routing

The route-guard branch admitting metadata-only annotation alters on
logical tables was only exercised end to end by sqlness. Pin it at the
DDL manager level: a semantic SET on a logical table succeeds, updates
only the logical table's metadata and dispatches nothing to datanodes;
a mixed batch reports its own error instead of the route guard's; the
repartition hint stays rejected on logical routes.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(table): keep entity guard on ADD COLUMN and report missing columns first

Review follow-ups: the old verify_alter loop scanned the post-alter
schema, so it also caught DROP COLUMN followed by re-adding the
declared column with a non-string type — the mutation-layer move only
kept the MODIFY path. Guard add_columns the same way (this also covers
ingestion auto-alter). And run the MODIFY drift check after the
existence lookup, so altering a dropped-but-still-declared column
reports ColumnNotExists (4002) like every other MODIFY on a missing
column.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* style(grpc-expr): drop a test comment restating the classifier doc

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor(table): rename annotation validation helpers per review

check_annotation* validated and normalized; align the names with the
validate_and_normalize_* convention nearby, and spell out
AnnotationContext (Cx is not used in this repo).

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-08-17 09:53:59 +00:00
Lei, HUANG a8924bb95c refactor(udaf): replace uddsketch implementation (#8867)
* refactor(function): replace uddsketch implementation

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* bench(function): compare uddsketch batch ingestion

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* perf(function): avoid copying non-null uddsketch batches

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(function): decode legacy uddsketch states

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix(function): harden legacy uddsketch validation

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* format: taplo

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* test: add compatibility tests for uddsketch functions

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-08-14 12:20:04 +00:00
Weny Xu ed4271af40 feat(procedure): record event actor (#8849)
* feat(event): record procedure actor

Signed-off-by: WenyXu <wenymedia@gmail.com>

* test: handle streamed region migration output

Signed-off-by: WenyXu <wenymedia@gmail.com>

* test: cover procedure actors across SQL protocols

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-08-14 09:36:01 +00:00
dennis zhuang 4dd92c774e feat: add json_object function and use it in the entity-graph derivation (#8870)
* feat: add json_object scalar function

Builds a JSONB object from interleaved (key, value, ...) arguments, like
MySQL's JSON_OBJECT. Values are written into the binary directly, so
JSON-hostile characters (quotes, backslashes, control characters) need no
text-level escaping. Keys must be non-NULL strings; values may be strings,
numbers, booleans, or NULL (JSON null).

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix: build entity-graph JSON objects with json_object

The derivation assembled entity_id_attrs and descriptive by concatenating a
JSON text and parsing it, escaping only backslash and double quote in runtime
values. A label containing a control character (e.g. a newline) produced
unparseable text and failed the whole semantic_entities scan instead of one
attribute. json_object assembles the JSONB binary directly from the value
columns, so no text escaping is involved; NULL-to-'' stays at the call site.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore: trim comments and fold duplicate test coverage

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix: json_object() returns an empty object; narrow values to integers and floats

MySQL's JSON_OBJECT allows an empty pair list, so the signature accepts zero
arguments and the row count falls back to number_rows. Decimals stay rejected
instead of casting to Float64: JSONB numbers (i64/u64/f64) cannot represent
them exactly and a silent precision loss is worse than an explicit cast.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore: document key-to-string conversion and align test naming

Keys follow MySQL JSON_OBJECT: any castable type is converted to string.
Rustdoc and the cast-failure message now say so, with a numeric-key test.
Test names take the module-conventional test_ prefix.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-08-14 08:24:50 +00:00
Weny Xu 1af4c33524 refactor(event): separate procedure submission context (#8856)
* refactor(event): separate procedure submission context

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(event): map extensions and forward GC context

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(gc): initialize integration test context

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(test): pass procedure context to DDL helpers

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor: simplify procedure submission contexts

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor(event): separate procedure and query contexts

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(event): clarify procedure context propagation

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(event): preserve procedure submission context

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor(event): move DDL context by value

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(test): retain manual GC event context

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor(event): tighten procedure context API

Signed-off-by: WenyXu <wenymedia@gmail.com>

* chore: update greptime-proto

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-08-13 07:29:00 +00:00
dennis zhuang 546625c45a feat: embedded convention pack for the entity graph (prom/k8s, gen_ai naming) (#8854)
* feat: embed the derivation conventions as data and adopt gen_ai entity naming

Move the co-declared edge vocabulary, the agent-edge vocabulary and the
virtual-destination candidates from Rust consts into an embedded
conventions.yaml (include_str!), parsed once behind a LazyLock and
validated against the entity-type grammar and the closed rel_type set; a
broken file propagates as a plan error instead of panicking. The agent
vocabulary entity types follow the GenAI semantic-convention namespace
as written: gen_ai.agent / gen_ai.model / gen_ai.tool.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* feat: drop the tag requirement for entity identity columns

Entity declarations no longer require id columns to be tag/primary-key
columns; only column existence is validated. Trace pipelines flatten the
identifying attributes (span_attributes.gen_ai.agent.id, ...) into field
columns, so the tag rule locked real trace tables out of declaring
entities while buying no correctness — the read-time derivation works on
any column.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* feat: implicit declarations for well-known prometheus info metrics

Tables stamped signal_type=metric + source=prometheus whose name matches
the conventions.yaml whitelist (kube_pod_info, kube_node_info,
kube_pod_owner, target_info) get implicit entity declarations: k8s.pod /
k8s.node / k8s.workload with name-based identity and target_info's
service / service.instance with the remaining tags as the descriptive
snapshot. The existing co-declared vocabulary then derives runs_on and
part_of from the same rows, so no new edge branch is needed. Explicit
declarations of a type always suppress the implicit one, and the metric
engine's physical table is excluded (it aggregates every logical
table's columns).

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test: cover the prometheus conventions in sqlness and compact the graph cases

Add the whitelisted-info-metric scenario (kube_pod_info, kube_pod_owner,
target_info deriving runs_on / part_of, a non-whitelisted metric
contributing nothing), fold the single-table calls, cross-table pairing
and virtual-node cases into one trace scenario (they exercise the same
union-before-join path), merge the two declaring-metric-table cases, and
reuse one rename probe for both reserved names.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix: reject entity id columns without a stable string form

Review follow-ups: the DDL check now validates against the schema and
rejects binary-backed and nested types for identity columns (the
derivation renders ids via CAST to Utf8, so the failure used to surface
only when the graph was scanned); the agent sqlness case keeps its
identity columns as fields to cover the relaxed tag rule end to end;
stale tag-rule comments and a dangling const reference are cleaned up.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix: type-check every entity column role, not only ids

The registry renders scope and descriptive values through the same
CAST-to-string path as ids, so a binary-backed column in any role fails
at scan time; the DDL check is now role-independent (and simpler).

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor: name the code-anchored vocabulary constants

Entity types and edge attributes the derivation code itself anchors on
(service, gen_ai.agent, calls, trace/attribute provenance) become
constants in the conventions module; the rest of the vocabulary stays
YAML-only data. ImplicitEntity is renamed PromImplicitEntity, and the
implicit-declaration path logs each skip of a whitelisted info metric
(wrong stamps, suppressed by an explicit declaration, missing id
column) so a missing graph entity is diagnosable.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor: single-source the graph constants

The graph tables' column names move to common-catalog (the schemas
catalog exposes and the plans operator builds must match column by
column), and the conventions module now carries the complete built-in
vocabulary — entity types, rel_types, provenances and connection types —
with the embedded YAML validated by membership against it, so an edit
drifting outside the vocabulary fails the conventions test instead of
deriving nothing.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix: treat empty identity components as absent

kube-state-metrics emits empty-string labels an entity id must not be
built from: an unscheduled pod's node and an owner-less pod's owner_kind
/ owner_name. Standard Prometheus drops empty labels (they arrive as
NULL and the existing predicate handles them), but other remote-write
agents may keep them, which produced ghost entities with empty ids and
false runs_on / part_of edges. Every identity predicate (registry,
co-declared edges, span endpoints) now requires non-NULL and non-empty
components through one shared helper.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor: tighten the conventions DSL semantics

Rename the co-declaration rule lists to what they are (co_declared_edges
/ trace_co_declared_edges — derivation rules, not a relation
vocabulary), stop overstating the GenAI entity types (Greptime types
derived from GenAI attributes; OTel defines no model/tool entities),
move target_info's descriptive snapshot to service.instance (the
remaining labels are the target's resource attributes, and instances
would write conflicting snapshots onto the logical service), and extend
the descriptor whitelist with the stable KSM sources: container info
metrics (closing the k8s.pod contains k8s.container rule),
kube_service_info (new k8s.service entity type) and the fuller
descriptive label sets.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix: guard entity column types on ALTER as well

ALTER MODIFY COLUMN could change a declared entity column to a type
without a stable string form, deferring the failure to graph scan time;
verify_alter now checks the post-alter schema. Dropping a declared
column stays allowed — the read-time derivation skips the stale
declaration, and semantic options cannot be altered off yet.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* feat: bridge traces and kube-state-metrics on the pod UID

Trace-v1 tables now get implicit declarations from their flattened
resource attributes (otlp_trace_entities in conventions.yaml): the
service identity — replacing the hardcoded fallback — plus
service.instance and k8s.pod, each applied only when its columns exist.
A new co-declared rule derives service.instance runs_on k8s.pod, and
the whitelisted kube-state-metrics pod identity switches from
namespace+pod names to the UID, so the trace-side pod and every KSM
descriptor land on one entity while names stay descriptive. This also
removes pod identity from the multi-cluster same-name collision.

The conventions rejection tests were passing for the wrong reason (a
half-renamed fixture key failed deserialization before reaching any
validation rule); they now assert the specific error each case targets.
Sqlness covers the UID merge across descriptor tables, pod-contains-
container, the k8s.service node, and the empty-uid/empty-node rows
deriving nothing.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test: cover the OTLP-to-graph chain end to end

One real OTLP export must come out of semantic_relationships as the
zero-configuration chain: service calls service, instance part_of
service, instance runs_on pod (bridged by k8s.pod.uid). Resources
without service.instance.id or k8s.pod.uid derive nothing extra.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix: identify k8s.service by UID

Same reasoning as pods: a recreated same-name service must not merge
into the old entity and same-named services across clusters must not
collide; kube_service_info carries a stable uid and nothing joins on the
service's name. Also drop a stale tag-rule mention from the option
validation docs.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore: cut duplicated test coverage and redundant comments

The trace service-fallback test collapsed into the resource-entities
test (same synthesis path since the fallback moved to YAML; only the
invalid-explicit-no-fallback case was distinct), role-duplicate and
subsumed DDL cases are gone, the embedded-conventions test is just the
parse (its assertions were decorative), and the YAML section comments no
longer restate the struct docs.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-08-13 02:31:56 +00:00
shuiyisong 6493435bee feat(promql): support native histogram aggregations (#8848)
* feat(promql): support native histogram aggregations

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix(promql): correct mixed native histogram aggregations

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix(promql): format mixed count_values labels consistently

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

* fix(promql): preserve reset hint warnings for incompatible histograms

Signed-off-by: shuiyisong <xixing.sys@gmail.com>

---------

Signed-off-by: shuiyisong <xixing.sys@gmail.com>
2026-08-12 07:50:54 +00:00
Weny Xu 943eee852f feat(event): record admin function executions (#8835)
* feat(event): record admin function executions

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(event): handle admin function recording edge cases

Signed-off-by: WenyXu <wenymedia@gmail.com>

* feat(event): record actor for admin functions

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(event): preserve admin function event values

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix(event): preserve non-finite admin results

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-08-11 12:50:50 +00:00
Weny Xu 72f6cf09bf refactor(procedure): centralize event context handling (#8834)
* refactor(procedure): centralize event context handling

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor(meta): simplify migration trigger reason handling

Signed-off-by: WenyXu <wenymedia@gmail.com>

* refactor(meta): avoid cloning event context

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-08-11 09:05:26 +00:00