mirror of
https://github.com/GreptimeTeam/greptimedb.git
synced 2026-09-10 07:22:17 +00:00
6fb5d2ebad42df98ebcde2bf9db5a4382f76a483
580
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6fb5d2ebad |
feat(mito2): add series index catalog and lifecycle foundation (#9053)
* refactor(mito2): add series index catalog and lifecycle components Signed-off-by: evenyag <realevenyag@gmail.com> * feat(mito2): restore series index catalogs on region open Signed-off-by: evenyag <realevenyag@gmail.com> * refactor(mito2): simplify series index foundation and maintenance Signed-off-by: evenyag <realevenyag@gmail.com> * refactor(mito2): separate series index purge task and simplify tests Signed-off-by: evenyag <realevenyag@gmail.com> * docs: defer experimental series index configuration examples Signed-off-by: evenyag <realevenyag@gmail.com> * feat(mito): make series index maintenance interval configurable Signed-off-by: evenyag <realevenyag@gmail.com> * fix(mito2): correct series index cleanup on close and drop Signed-off-by: evenyag <realevenyag@gmail.com> * test(mito2): revert drop test changes Signed-off-by: evenyag <realevenyag@gmail.com> * test: update config API expectation for series index settings Signed-off-by: evenyag <realevenyag@gmail.com> * refactor: use tokio unbounded channel for series index purger Signed-off-by: evenyag <realevenyag@gmail.com> * chore(mito2): simplify review test scope and clarify index config Signed-off-by: evenyag <realevenyag@gmail.com> --------- Signed-off-by: evenyag <realevenyag@gmail.com> |
||
|
|
b46de8c828 |
feat(telemetry): add log directory size retention (#8997)
* feat(telemetry): add log directory size retention Signed-off-by: WenyXu <wenymedia@gmail.com> * test: update config API logging fixture Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(telemetry): recover log retention state Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(telemetry): handle log retention cleanup errors Signed-off-by: WenyXu <wenymedia@gmail.com> * test(telemetry): cover log count retention on rotation Signed-off-by: WenyXu <wenymedia@gmail.com> * test(telemetry): cover log directory retention Signed-off-by: WenyXu <wenymedia@gmail.com> * perf(telemetry): avoid log filename allocation Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
35f5485974 |
feat(meta): record logical-table reconciliation events (#8941)
* feat(meta): add logical table reconciliation events Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> * fix(meta): preserve logical reconciliation progress Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> * fix(meta): preserve logical region retry progress Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> * fix(meta): simplify logical reconciliation events Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> * test(meta): assert logical event values Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> --------- Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> |
||
|
|
fa794fae7a |
fix: match system schema names case-insensitively (#9040)
* fix: match system schema names case-insensitively Database names that arrive over a protocol (the MySQL handshake and COM_INIT_DB, the Postgres startup parameter, the HTTP `db` parameter, the gRPC dbname header) never reach the SQL parser, which is what lowercases unquoted identifiers. Since #8062 stopped lowercasing them wholesale, connecting to `INFORMATION_SCHEMA` in any spelling but the canonical one fails with "Unknown database" -- including the `USE <db>` that a MySQL client turns into COM_INIT_DB. Fold only system schema names to their canonical spelling, so user schema names keep the case they were created with. `is_reserved_schema_name` uses the same match, otherwise a quoted `CREATE DATABASE "INFORMATION_SCHEMA"` creates a schema shadowed by the system one. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * refactor: hoist system schema names into a const Signed-off-by: Dennis Zhuang <killme2008@gmail.com> --------- Signed-off-by: Dennis Zhuang <killme2008@gmail.com> |
||
|
|
ed1f2d9f4e |
fix(pipeline): coalesce concurrent pipeline cache misses (#9022)
* fix(pipeline): coalesce concurrent pipeline cache misses The pipeline cache reads with a plain `moka::sync::Cache::get` and falls through to a distributed query on a miss, so when the 10s TTL expires every in-flight write request on a frontend issues its own scan of the single-region `greptime_private.pipelines` table. Concurrent scans per expiry scale with write QPS, and every frontend's burst lands on the same datanode. A user running high-throughput ingestion through a pipeline saw that datanode overloaded. Switch to `moka::future::Cache::try_get_with` so concurrent misses on the same key share one loader. This requires a single-key lookup, so cache entries are now keyed by the requested schema rather than the schema the pipeline is stored under; resolving a request to a stored schema stays in the loader, which is the authoritative path and already handles the empty-schema and multi-schema cases. A lookup for a schema not yet cached costs one extra read, now protected from amplification by the coalescing it enables. `remove_cache` previously only walked the compiled-pipeline cache, so an entry populated by `get_pipeline_str` alone (the pipeline read API) survived deletion until it expired. It now walks all three caches. Also make the TTL configurable as `pipeline.cache_ttl`, default unchanged at 10s. The TTL is what propagates a pipeline change to other frontends, so raising it trades staleness for fewer reads. Refs #9021 Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix(pipeline): restore cross-schema semantics broken by the new cache key Keying cache entries by the requested schema dropped two behaviours that the previous stored-schema key provided for free. Creating a new version only wrote the creating request's schema, so another schema on the same frontend kept serving its cached `latest` — an older version — until the entry expired. Since the whole point of making the TTL configurable is to let operators raise it, that window is not bounded by anything useful. Creation now invalidates every schema's `latest` alias for that name before priming the cache, leaving the version-pinned keys alone. The failover cache lost its reach across schemas the same way: a global pipeline (stored under the empty schema) loaded by schema A was cached under `A`, so schema B using it for the first time while the pipeline table was down missed and failed ingestion. The failover cache has no loader and so is not subject to the single-key model of `try_get_with`; it keeps the stored-schema key and the empty-schema-first resolution. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * refactor(pipeline): drop cache priming on create and fold the sweep helpers Priming the cache on create saved one read on a low-frequency operation and cost a concept: entries were written under the creating request's schema while `PipelineContent.schema` said empty, so the two schemas in play disagreed. Invalidating the `latest` aliases is required regardless — that is what makes a new version visible to other schemas — so dropping the priming loses only the saved read, which coalescing now protects anyway. `insert_and_compile` no longer needs the caller's schema. `remove_cache` and the create-time invalidation collapse into one `invalidate(name, version)`; `None` sweeps only the `latest` aliases, which is exactly what creation wants. That leaves `invalidate_by_suffixes` and `cache_keys` with a single caller each, so both are inlined. Drop the `PipelineOptions` humantime test: `load_config_test` loads both example TOMLs, which now carry `cache_ttl = "10s"`, and would fail the same way if the serde attribute were lost. The `toml` dev-dependency goes with it. The two invalidation tests are now checked to be orthogonal: removing the version suffix fails only the delete test, and sweeping just the compiled cache fails both. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix(pipeline): keep failover populated across a create The `latest` sweep on create clears the failover cache along with the loaded ones, and after dropping the priming there was nothing writing it back. An outage between the create and the first read-back left neither `latest` nor the explicit version with anything to fall back on, failing ingestion — worse than before, since the previous version's failover entry was swept too. Creation now goes through `PipelineCache::on_pipeline_created`, which pairs the sweep with a failover write of the new empty-schema definition. The two must happen together, so they live behind one method rather than at the call site. Also commit the Cargo.lock entry for the dropped `toml` dev-dependency, and trim the comments added over the last few commits down to what the code does not already say. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> --------- Signed-off-by: Dennis Zhuang <killme2008@gmail.com> |
||
|
|
b86da3d35f |
feat: support raw OTLP delta metrics (#8970)
* feat: support raw OTLP delta metrics Signed-off-by: shuiyisong <xixing.sys@gmail.com> * fix: fmt Signed-off-by: shuiyisong <xixing.sys@gmail.com> * test(promql): update sqlness results for normalized label matching Signed-off-by: shuiyisong <xixing.sys@gmail.com> * fix: derive temporality label from default column prefix Signed-off-by: shuiyisong <xixing.sys@gmail.com> * test(promql): add analyze coverage for delta temporality Signed-off-by: shuiyisong <xixing.sys@gmail.com> * fix(promql): scope label alignment to temporality marker Signed-off-by: shuiyisong <xixing.sys@gmail.com> * fix: handle count-only histograms and vector broadcasts Signed-off-by: shuiyisong <xixing.sys@gmail.com> * fix: exclude temporality marker from entity descriptions Signed-off-by: shuiyisong <xixing.sys@gmail.com> * fix: use a fixed label for OTLP aggregation temporality Signed-off-by: shuiyisong <xixing.sys@gmail.com> * fix(promql): preserve mixed-range semantics for raw delta Signed-off-by: shuiyisong <xixing.sys@gmail.com> --------- Signed-off-by: shuiyisong <xixing.sys@gmail.com> |
||
|
|
9c135ebcb3 |
feat!: stabilize streaming analyze metrics (#8966)
* feat: stabilize streaming analyze metrics Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * feat: expose analyze memory usage Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * refactor: simplify analyze stream handling Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * fix: preserve analyze stream sequence on panic Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * chore: log analyze stream worker panic Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> --------- Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> |
||
|
|
00d43b29ad |
feat(query): add experimental DataFusion spill-to-disk controls (#8884)
* feat(query): add experimental DataFusion spill-to-disk controls Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * docs(config): regenerate configuration reference Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * test: update config API for spill defaults Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * fix(query): address spill configuration review Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> * fix(query): preserve spill settings with runtime plugins Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> --------- Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> |
||
|
|
d32cd77505 |
fix: postgres describe for more statements (#8974)
* fix: postgres describe for more statements Signed-off-by: Ning Sun <sunning@greptime.com> * fix: cover more show statements Signed-off-by: Ning Sun <sunning@greptime.com> * fix: address review comments - add missing `clippy::too_many_arguments` allow on `query_from_information_schema_dataframe` (CI clippy failure) - take `&ShowKind` in the information-schema dataframe helper so `kind` is no longer cloned at every call site; only the WHERE arm (which needs an owned expression for `sql_to_expr`) clones internally - document why re-applying TQL explain formats never overwrites an existing value (per-query context state) Signed-off-by: Ning Sun <sunning@greptime.com> * chore: trim comments to essentials Signed-off-by: Ning Sun <sunning@greptime.com> --------- Signed-off-by: Ning Sun <sunning@greptime.com> |
||
|
|
c01de4afdc |
fix(operator): whitelist private system table auto create (#8930)
Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
9198462869 |
feat(meta): record physical table reconciliation events (#8935)
* feat(meta): record physical table reconciliation events Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> * docs(config): add reconciliation table event Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> * fix(meta): address reconciliation event review feedback Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> * fix(meta): keep reconciliation event summary volatile Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> * refactor(meta): remove unused table state downcasting Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> --------- Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com> |
||
|
|
ba3c5a939e |
chore(mito2): reduce default auto flush interval (#8971)
Signed-off-by: evenyag <realevenyag@gmail.com> |
||
|
|
aaa843104b |
fix(mysql): interpret prepared statement datetime params in session timezone (#8923)
* fix(mysql): interpret prepared statement datetime params in session timezone Binary DATETIME parameters of server-side prepared statements were converted as if UTC, ignoring the session timezone set via SET time_zone. Convert them with the session timezone and add an integration test covering prepared inserts and predicates under Asia/Shanghai. Signed-off-by: wy471x <wy471x@gmail.com> * refactor: share naive datetime timezone policy via common-time Address review feedback on the prepared-statement timezone fix: - Expose Timestamp::from_naive_datetime in common-time so the DST policy (gap -> error, ambiguous -> earlier instant) lives in one place, shared by the text protocol (Timestamp::from_str) and the MySQL binary protocol. - Route the MySQL prepared-statement datetime conversion through it. - Match the target type before converting datetime params so PreparedStmtTypeMismatch fails fast without wasted conversion. - Use the short Timezone import form for consistency with the rest of servers. Signed-off-by: wy471x <wy471x@gmail.com> --------- Signed-off-by: wy471x <wy471x@gmail.com> Co-authored-by: Ning Sun <sunng@protonmail.com> |
||
|
|
1409e66837 |
refactor(flight): add request builder and defer DoGet execution (#8953)
* refactor: add flight request builder Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor: use flight request builder in flow Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor: defer frontend flight query execution Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(grpc): avoid cloning requests during auth Signed-off-by: WenyXu <wenymedia@gmail.com> * test(grpc): cover flight request timeout Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(client): share Flight message reader Signed-off-by: WenyXu <wenymedia@gmail.com> * feat(flow): add Flight DoGet timeout Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(client): gate Flight DDL helpers for testing Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(client): restore Flight stream error semantics Signed-off-by: WenyXu <wenymedia@gmail.com> * docs(grpc): document Flight stream input constructors Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(flight): preserve deferred stream context Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(client): use Flight stream SNAFU context Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
6d86e6ff06 |
feat: synthesize OTLP resource descriptor for the semantic entity graph (#8904)
* fix(servers): compose OTLP metrics job from service.namespace/service.name
The OTel Prometheus compatibility spec defines job as
"<service.namespace>/<service.name>" when the namespace is present.
The OTLP metrics path only used the bare service.name, so the job tag
diverged from target_info produced by Prometheus-side exporters for the
same resource. Compose the namespace form, and keep not fabricating a
job when service.name is absent.
Behavior change: resources carrying service.namespace now get
"namespace/name" as their job tag value.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* feat(otlp): synthesize otel_resource_info at OTLP metrics ingestion
Ordinary OTLP metrics scatter filtered resource attributes as tags over
every logical metric table, so metrics-only services contribute nothing
to the semantic entity graph. Each request now also projects its
distinct resources into one info-metric-shaped mito table,
otel_resource_info: a fixed allowlist of identity-relevant attributes
under their raw OTel keys (independent of the label translation
strategy and the promote/ignore headers) plus derived job/instance
compatibility columns, value 1.0, and the newest data-point timestamp.
The descriptor is written after the main insert is committed; a failure
there (conflicting pre-existing table, auto-create disabled) degrades
to an OTLP partial_success warning with rejected_data_points = 0
instead of failing the request and triggering client retries of
already-accepted data. A request writing a metric named
otel_resource_info suppresses synthesis. Legacy mode is unchanged.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* feat(operator): otel info-metric conventions with host/container entities
Whitelist the ingestion-synthesized otel_resource_info descriptor via a
new otel_info_metrics conventions map, gated on source=opentelemetry
(the existing gate hardcoded source=prometheus). Its declarations use
explicit descriptive lists instead of descriptive_rest so identifying
attributes of other entities do not leak into service.instance.
Conventions tightened per the Astronomy Shop findings: host identity is
host.id with host.name descriptive only (host.name is not stable across
SDKs and resource detectors), a generic container entity (new entity
type) is declared only when container.id is present, and trace-v1
tables now synthesize host/container from their flattened resource
attributes too. New co-declared edges: service.instance runs_on
container, container runs_on host.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* test(otlp): cover the resource descriptor in integration tests
Covers the descriptor's raw-key columns and info-metric options through
the HTTP path, the namespace/name job composition end-to-end, column
names being independent of the translation strategy, the allowlist
excluding unlisted resource attributes, auto-create after a drop, the
metric-name collision suppressing synthesis, and the partial-success
warning (rejected_data_points = 0) when a pre-existing incompatible
table fails the descriptor write while metric data is accepted.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* chore: cargo fmt
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(frontend): degrade descriptor permission denial to a warning
A table-level permission policy denying otel_resource_info would have
failed the whole OTLP metrics request because the descriptor's
permission check ran before the main insert. The descriptor is derived
enrichment: check its permission in the degrade path so a denial skips
the write and surfaces as the partial-success warning, like any other
descriptor write failure.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(otlp): guard descriptor writes with semantic ownership markers
A pre-existing schema-compatible table named otel_resource_info would
silently receive descriptor rows while its missing semantic stamps kept
it out of the entity graph. The descriptor write now requires the
auto-created table's ownership markers (mito engine + signal_type +
source + metric.type=info + metadata_quality=declared) and otherwise
degrades to the partial-success warning; the entity-graph gate for the
otel whitelist likewise requires metric.type=info, so a user table
stamped with only signal/source no longer picks up implicit
declarations.
Also fold the descriptor write cost into the response and surface the
degrade warning through the otel-arrow BatchStatus status_message.
Integration tests pin the full marker set on auto-create and that an
existing owned descriptor keeps accepting writes without degrading —
a missing marker would otherwise silently stop every descriptor write
after the first request.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* perf(otlp): build descriptor rows without the per-resource BTreeMap
Projecting a resource allocated a BTreeMap and then collected it into the
row key, and every attribute was matched against the allowlist by linear
scan. Collect the tags into a Vec and sort once, and match the allowlist
instead of scanning it. Measured on the conversion path: descriptor work
drops 16-18%, from 10.6% to 8.9% of conversion CPU on the worst shape
(1000 resources with 4 data points each), where the cost tracks resource
count rather than data-point count.
Also trims the comments and tests added with the descriptor to what
carries information.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* test(otlp): pin the descriptor permission-denial degrade path
A policy denying the descriptor table must not fail the metrics request,
which the fix in
|
||
|
|
8a473c5bf0 |
fix(meta): allow manual migration from offline datanodes (#8934)
* fix(meta): allow migration from offline datanodes Signed-off-by: WenyXu <wenymedia@gmail.com> * test: fix offline migration event actor Signed-off-by: WenyXu <wenymedia@gmail.com> * test: read migration routes from metadata Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
3493d2d0fb |
perf(servers)!: speed up Prometheus JSON response building with ryu and per-series entry reuse (#8815)
* fix(perf): align direct-SST CREATE TABLE with baked index metadata
The offline fixture generator (query_perf_fixture::direct_sst::
build_region_metadata) bakes greptime:inverted_index /
greptime:skipping_index field metadata into the region manifest for
tag/field columns, but create_table_sql emitted a bare CREATE TABLE
without those declarations. MergeScan's remote-schema validation then
failed on any tag/field projection (HTTP 500 'advertised remote stream
schema field mismatch'), breaking direct_readable_sst perf cases.
CREATE TABLE now declares the matching SKIPPING INDEX WITH
(granularity='1') / INVERTED INDEX column options. A round-trip test
proves the emitted SQL is parser-valid and yields the exact catalog
metadata.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* perf(servers): speed up Prometheus JSON response building with ryu and per-series entry reuse
PrometheusJsonResponse::record_batches_to_data spends ~47% of its CPU
in f64::to_string() per sample and ~32% in IndexMap::entry() per row
(60s profile of concurrent query_range workloads, ~800k series).
- Replace f64::to_string() with ryu::Buffer::format_finite for finite
values (shortest round-trip, 3-5x faster); NaN/+Inf/-Inf keep the
previous std formatting so wire output is unchanged.
- Remember the previous row's label vector and entry index; query output
is clustered by series, so consecutive rows reuse the same IndexMap
entry via get_index_mut instead of rebuilding and hashing the label
vector (worst case adds one Vec comparison per series transition).
Also adds a query-regression case (prom_json_response) that measures the
real Prometheus HTTP range API path (/v1/prometheus/api/v1/query_range),
which is the only frontend path that builds the Prometheus JSON response
(TQL ANALYZE formats the SQL JSON shape instead), plus a prom_http query
kind in the regression runner.
Perf (aligned base
|
||
|
|
1c5eabcbbf |
feat(otlp): support cumulative exponential histograms (#8900)
* feat: implement exponential histogram Signed-off-by: shuiyisong <xixing.sys@gmail.com> * chore: remove duplicate tests Signed-off-by: shuiyisong <xixing.sys@gmail.com> * fix(otlp): enforce exponential histogram ingestion safety Signed-off-by: shuiyisong <xixing.sys@gmail.com> * chore: update rfc Signed-off-by: shuiyisong <xixing.sys@gmail.com> * fix: test Signed-off-by: shuiyisong <xixing.sys@gmail.com> * fix(otlp): remove protocol-coupled histogram checks Signed-off-by: shuiyisong <xixing.sys@gmail.com> * perf(otlp): reuse native histogram schema across data points Signed-off-by: shuiyisong <xixing.sys@gmail.com> * fix: merge repeated OTLP histogram fragments Signed-off-by: shuiyisong <xixing.sys@gmail.com> * fix(otlp): build rejection messages lazily Signed-off-by: shuiyisong <xixing.sys@gmail.com> * fix: add doc Signed-off-by: shuiyisong <xixing.sys@gmail.com> --------- Signed-off-by: shuiyisong <xixing.sys@gmail.com> |
||
|
|
174577164a |
feat: otlp duration_nano and trace_flag signed integer coercion (#8816)
* feat(servers): add signed→unsigned int coercion to OTLP ingest path Phase 0 of transitioning built-in data models from unsigned to signed integers (#8793): add lossless Int64→UInt64 and Int32→UInt32 coercion arms so existing UInt64/UInt32 columns (e.g. trace `duration_nano`, log `trace_flags`) keep accepting new signed ingest without an ALTER TABLE. The OTLP ingest path already reconciles every incoming column against the existing table schema and treats it as authoritative. With these arms, `choose_trace_reconcile_decision` returns `UseExisting(UInt64/Uint32)` for an existing unsigned column receiving signed data: the table keeps its type byte-for-byte and the request value is coerced. No persisted format is mutated; existing data stays readable as-is. This is the safety net that makes the actual schema flip (Phase 1) safe. Only the signed→unsigned direction is supported — the reverse would be lossy for values above the signed range and is intentionally rejected. Tests cover both new arms plus an end-to-end log test proving an existing UInt64 column coerces an incoming Int64 value while keeping its type. Signed-off-by: Ning Sun <sunning@greptime.com> * feat: add compatibility layer for uint trace/log fields Signed-off-by: Ning Sun <sunning@greptime.com> * fix: jaeger test Signed-off-by: Ning Sun <sunning@greptime.com> * fix: keep trace v0 unsigned, reject negative span durations - Keep the frozen v0 data model on UInt64 duration_nano: the signed ingest compatibility layer only runs on the v1 path, so flipping v0 would break writes into every pre-existing v0 table at mito's schema check. Pin the schema with unit and integration tests. - Reject spans whose end precedes their start (or whose duration does not fit i64) on the v1 path instead of wrapping: new Int64 tables and existing UInt64 tables now fail identically, rather than storing negative durations that break the Jaeger query API. - Extract is_supported_signed_to_unsigned_coercion so the trace and log ingest paths share one supported-pair predicate and cannot drift. Signed-off-by: Ning Sun <sunning@greptime.com> * fix: clamp negative span durations to zero, revert semantic_graph comment - Record duration 0 for spans whose end precedes their start instead of erroring: a malformed span no longer fails the request, and the value written is always a non-negative, in-range i64 so new Int64 tables and existing UInt64 tables (via the checked coercion) behave identically. Durations above i64::MAX saturate rather than wrap. - Revert the doc-comment tweak on the semantic_graph test fixture; the file is untouched by this PR again. Signed-off-by: Ning Sun <sunning@greptime.com> --------- Signed-off-by: Ning Sun <sunning@greptime.com> |
||
|
|
8d887ddd00 |
fix(query): restore columnar group-by for dictionary-encoded tags (#8902)
* fix(query): restore columnar group-by for dictionary-encoded tags Bump the DataFusion fork to be93ffd85 (feat/dict-group-column-53), which backports apache/datafusion #23187: DictionaryGroupValuesColumn lets dictionary-encoded group keys use the columnar GroupValuesColumn fast path (hash distinct dictionary values once per batch, resolve rows by key index) instead of falling back to row-based GroupValuesRows. This fixes the TSBS double-groupby regression introduced by #8541 (preserve dictionary-encoded query labels): v1.2.0-beta.1 scan output changed tag columns to Dictionary(UInt32, Utf8), which DataFusion 53.1.0 did not support in GroupValuesColumn's supported_type allow-list, so GROUP BY queries silently dropped to the ~60% slower row path (time_calculating_group_ids +57%, peak_mem +50%, end-to-end +38%). Adds an end-to-end integration test (dict_groupby_sst) that flushes a flat-format SST with dictionary-encoded hostname, runs the tsbs-style double-groupby query, and asserts correct results with no CastExec inserted before the aggregate. Signed-off-by: discord9 <discord9@greptime.dev> * chore(deps): pin merged dictionary group-by support Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> --------- Signed-off-by: discord9 <discord9@greptime.dev> Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> Co-authored-by: discord9 <discord9@greptime.dev> |
||
|
|
3de4feddd5 |
feat(otlp): report the cause of rejected trace spans (#8897)
* feat(otlp): report the cause of rejected trace spans
When trace-v1 ingestion cannot coerce an attribute value, it falls back to
single-span writes and rejects the bad span. That behavior is correct, but the
OTLP partial-success message only carried `Rejected span <trace_id>:<span_id>
(InvalidArguments)`: the column, the source value, the source type and the
target type were all dropped, so locating the bad attribute required adding a
detailed exporter on the collector side and replaying traffic.
Two places lost the information. `prepare_trace_column_rewrites` built a message
without the failing value, and the span rejection path kept only the status code
from the error.
Coercion errors now name the failing value, e.g.
failed to coerce trace column 'span_attributes.http.response.body.size'
in table 'opentelemetry_traces' from String("") to Int64
and the rejection detail carries that cause. Values are user data, so a string
keeps at most 16 characters, is escaped, and binary payloads report only their
length; the cause itself is bounded at 256 characters. Both truncations cut on a
char boundary.
Failure details now deduplicate: repeats of the same (site, cause) collapse into
one entry with an occurrence count, keyed on the untruncated cause so two
failures that differ past the display limit stay separate. Only four distinct
entries are retained and the rest are counted, which keeps the state bounded no
matter how many distinct bad values a request carries. A fully rejected request
is logged at warn level and a partial success at debug level, since the latter
repeats every export interval; the detail goes out as a Debug field so a newline
in an attribute key cannot forge log lines.
Rejection semantics are unchanged: partial success, same accepted and rejected
counts, same HTTP status mapping.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor(otlp): compare failure keys directly instead of hashing
The dedup identity was a DefaultHasher fingerprint of `(label, key)`, which
bought a fixed 8 bytes per entry at the cost of an import, four lines, and a
collision argument the reader has to make. Entries are capped at four and a
cause runs a couple of hundred characters, so the saving is about a kilobyte
per in-flight request while the column name it avoids retaining is already held
several times over by the request itself.
Compare the strings instead, keeping the untruncated cause as the key so
failures differing past the display limit still stay apart. Labels are metric
label values and always static, so the entry borrows them.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
|
||
|
|
ed4271af40 |
feat(procedure): record event actor (#8849)
* feat(event): record procedure actor Signed-off-by: WenyXu <wenymedia@gmail.com> * test: handle streamed region migration output Signed-off-by: WenyXu <wenymedia@gmail.com> * test: cover procedure actors across SQL protocols Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
76924c2d36 |
feat(mito2): introduce two-phase metric series scans (#8826)
* feat(mito2): add two-phase series scan Signed-off-by: evenyag <realevenyag@gmail.com> * docs: regenerate configuration reference Signed-off-by: evenyag <realevenyag@gmail.com> * test(sqlness): update series scan explain results Signed-off-by: evenyag <realevenyag@gmail.com> * test: update config API expectation Signed-off-by: evenyag <realevenyag@gmail.com> * fix(mito2): bound two-phase series discovery Signed-off-by: evenyag <realevenyag@gmail.com> * fix(mito2): avoid candidate distribution deadlock Signed-off-by: evenyag <realevenyag@gmail.com> * chore(mito2): remove obsolete dead code allowances Signed-off-by: evenyag <realevenyag@gmail.com> * fix(mito2): share series scan memory pool Signed-off-by: evenyag <realevenyag@gmail.com> --------- Signed-off-by: evenyag <realevenyag@gmail.com> |
||
|
|
1af4c33524 |
refactor(event): separate procedure submission context (#8856)
* refactor(event): separate procedure submission context Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(event): map extensions and forward GC context Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(gc): initialize integration test context Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(test): pass procedure context to DDL helpers Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor: simplify procedure submission contexts Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(event): separate procedure and query contexts Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(event): clarify procedure context propagation Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(event): preserve procedure submission context Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(event): move DDL context by value Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(test): retain manual GC event context Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(event): tighten procedure context API Signed-off-by: WenyXu <wenymedia@gmail.com> * chore: update greptime-proto Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
546625c45a |
feat: embedded convention pack for the entity graph (prom/k8s, gen_ai naming) (#8854)
* feat: embed the derivation conventions as data and adopt gen_ai entity naming Move the co-declared edge vocabulary, the agent-edge vocabulary and the virtual-destination candidates from Rust consts into an embedded conventions.yaml (include_str!), parsed once behind a LazyLock and validated against the entity-type grammar and the closed rel_type set; a broken file propagates as a plan error instead of panicking. The agent vocabulary entity types follow the GenAI semantic-convention namespace as written: gen_ai.agent / gen_ai.model / gen_ai.tool. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * feat: drop the tag requirement for entity identity columns Entity declarations no longer require id columns to be tag/primary-key columns; only column existence is validated. Trace pipelines flatten the identifying attributes (span_attributes.gen_ai.agent.id, ...) into field columns, so the tag rule locked real trace tables out of declaring entities while buying no correctness — the read-time derivation works on any column. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * feat: implicit declarations for well-known prometheus info metrics Tables stamped signal_type=metric + source=prometheus whose name matches the conventions.yaml whitelist (kube_pod_info, kube_node_info, kube_pod_owner, target_info) get implicit entity declarations: k8s.pod / k8s.node / k8s.workload with name-based identity and target_info's service / service.instance with the remaining tags as the descriptive snapshot. The existing co-declared vocabulary then derives runs_on and part_of from the same rows, so no new edge branch is needed. Explicit declarations of a type always suppress the implicit one, and the metric engine's physical table is excluded (it aggregates every logical table's columns). Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test: cover the prometheus conventions in sqlness and compact the graph cases Add the whitelisted-info-metric scenario (kube_pod_info, kube_pod_owner, target_info deriving runs_on / part_of, a non-whitelisted metric contributing nothing), fold the single-table calls, cross-table pairing and virtual-node cases into one trace scenario (they exercise the same union-before-join path), merge the two declaring-metric-table cases, and reuse one rename probe for both reserved names. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix: reject entity id columns without a stable string form Review follow-ups: the DDL check now validates against the schema and rejects binary-backed and nested types for identity columns (the derivation renders ids via CAST to Utf8, so the failure used to surface only when the graph was scanned); the agent sqlness case keeps its identity columns as fields to cover the relaxed tag rule end to end; stale tag-rule comments and a dangling const reference are cleaned up. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix: type-check every entity column role, not only ids The registry renders scope and descriptive values through the same CAST-to-string path as ids, so a binary-backed column in any role fails at scan time; the DDL check is now role-independent (and simpler). Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * refactor: name the code-anchored vocabulary constants Entity types and edge attributes the derivation code itself anchors on (service, gen_ai.agent, calls, trace/attribute provenance) become constants in the conventions module; the rest of the vocabulary stays YAML-only data. ImplicitEntity is renamed PromImplicitEntity, and the implicit-declaration path logs each skip of a whitelisted info metric (wrong stamps, suppressed by an explicit declaration, missing id column) so a missing graph entity is diagnosable. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * refactor: single-source the graph constants The graph tables' column names move to common-catalog (the schemas catalog exposes and the plans operator builds must match column by column), and the conventions module now carries the complete built-in vocabulary — entity types, rel_types, provenances and connection types — with the embedded YAML validated by membership against it, so an edit drifting outside the vocabulary fails the conventions test instead of deriving nothing. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix: treat empty identity components as absent kube-state-metrics emits empty-string labels an entity id must not be built from: an unscheduled pod's node and an owner-less pod's owner_kind / owner_name. Standard Prometheus drops empty labels (they arrive as NULL and the existing predicate handles them), but other remote-write agents may keep them, which produced ghost entities with empty ids and false runs_on / part_of edges. Every identity predicate (registry, co-declared edges, span endpoints) now requires non-NULL and non-empty components through one shared helper. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * refactor: tighten the conventions DSL semantics Rename the co-declaration rule lists to what they are (co_declared_edges / trace_co_declared_edges — derivation rules, not a relation vocabulary), stop overstating the GenAI entity types (Greptime types derived from GenAI attributes; OTel defines no model/tool entities), move target_info's descriptive snapshot to service.instance (the remaining labels are the target's resource attributes, and instances would write conflicting snapshots onto the logical service), and extend the descriptor whitelist with the stable KSM sources: container info metrics (closing the k8s.pod contains k8s.container rule), kube_service_info (new k8s.service entity type) and the fuller descriptive label sets. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix: guard entity column types on ALTER as well ALTER MODIFY COLUMN could change a declared entity column to a type without a stable string form, deferring the failure to graph scan time; verify_alter now checks the post-alter schema. Dropping a declared column stays allowed — the read-time derivation skips the stale declaration, and semantic options cannot be altered off yet. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * feat: bridge traces and kube-state-metrics on the pod UID Trace-v1 tables now get implicit declarations from their flattened resource attributes (otlp_trace_entities in conventions.yaml): the service identity — replacing the hardcoded fallback — plus service.instance and k8s.pod, each applied only when its columns exist. A new co-declared rule derives service.instance runs_on k8s.pod, and the whitelisted kube-state-metrics pod identity switches from namespace+pod names to the UID, so the trace-side pod and every KSM descriptor land on one entity while names stay descriptive. This also removes pod identity from the multi-cluster same-name collision. The conventions rejection tests were passing for the wrong reason (a half-renamed fixture key failed deserialization before reaching any validation rule); they now assert the specific error each case targets. Sqlness covers the UID merge across descriptor tables, pod-contains- container, the k8s.service node, and the empty-uid/empty-node rows deriving nothing. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test: cover the OTLP-to-graph chain end to end One real OTLP export must come out of semantic_relationships as the zero-configuration chain: service calls service, instance part_of service, instance runs_on pod (bridged by k8s.pod.uid). Resources without service.instance.id or k8s.pod.uid derive nothing extra. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix: identify k8s.service by UID Same reasoning as pods: a recreated same-name service must not merge into the old entity and same-named services across clusters must not collide; kube_service_info carries a stable uid and nothing joins on the service's name. Also drop a stale tag-rule mention from the option validation docs. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * chore: cut duplicated test coverage and redundant comments The trace service-fallback test collapsed into the resource-entities test (same synthesis path since the fallback moved to YAML; only the invalid-explicit-no-fallback case was distinct), role-duplicate and subsumed DDL cases are gone, the embedded-conventions test is just the parse (its assertions were decorative), and the YAML section comments no longer restate the struct docs. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> --------- Signed-off-by: Dennis Zhuang <killme2008@gmail.com> |
||
|
|
3b3a9032e0 |
feat(servers): expose native histograms over Prometheus HTTP (#8850)
* feat(servers): expose native histograms over Prometheus HTTP Signed-off-by: shuiyisong <xixing.sys@gmail.com> * fix(servers): refine Prometheus HTTP metadata handling Signed-off-by: shuiyisong <xixing.sys@gmail.com> * test(servers): expand Prometheus metadata coverage Signed-off-by: shuiyisong <xixing.sys@gmail.com> * fix(servers): return OpenMetrics units from metadata API Signed-off-by: shuiyisong <xixing.sys@gmail.com> --------- Signed-off-by: shuiyisong <xixing.sys@gmail.com> |
||
|
|
943eee852f |
feat(event): record admin function executions (#8835)
* feat(event): record admin function executions Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(event): handle admin function recording edge cases Signed-off-by: WenyXu <wenymedia@gmail.com> * feat(event): record actor for admin functions Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(event): preserve admin function event values Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(event): preserve non-finite admin results Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
72f6cf09bf |
refactor(procedure): centralize event context handling (#8834)
* refactor(procedure): centralize event context handling Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(meta): simplify migration trigger reason handling Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(meta): avoid cloning event context Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
3510ef7d4c |
feat!: update native histogram unsigned int types (#8824)
* feat(native-histogram): store counts and span lengths as signed integers Native histograms are unreleased, so the on-disk integer payload columns are switched from unsigned to signed types without backward-compat: - count_u64 / zero_count_u64: uint64 -> int64 - positive_span_lengths / negative_span_lengths: list(uint32) -> list(int32) - Span.length (query-time model): u32 -> i32 The Prometheus remote-write v2 source carries these as uint64/uint32, so the unsigned->signed conversion at the ingestion boundary is overflow checked: an integer count >= 2^63 or a span length >= 2^31 is rejected with an explicit error rather than silently wrapping to a negative value. read_spans additionally rejects negative stored lengths to keep the non-negative invariant sound for downstream `as usize` casts. The UDAF accumulator's own observation counter (transient aggregation state, not part of the persisted histogram value) is intentionally left as uint64. Signed-off-by: Ning Sun <sunning@greptime.com> * refactor(native-histogram): rename count/zero_count fields to _i64 Now that the integer payload columns are stored as int64, rename the field constants and persisted names to match: COUNT_U64_FIELD ("count_u64") -> COUNT_I64_FIELD ("count_i64") ZERO_COUNT_U64_FIELD ("zero_count_u64") -> ZERO_COUNT_I64_FIELD ("zero_count_i64") The local builder variables and the docs/JSON snapshot are updated to match. No backward-compat (unreleased feature). Signed-off-by: Ning Sun <sunning@greptime.com> * test(native-histogram): refresh planner plan snapshot for signed types The mixed native-histogram range test embeds the full histogram Struct type in its expected plan string, which still carried the pre-rename unsigned fields. Update the snapshot to match the signed schema: positive/negative_span_lengths: List(UInt32) -> List(Int32) count_u64/zero_count_u64: UInt64 -> count_i64/zero_count_i64: Int64 Signed-off-by: Ning Sun <sunning@greptime.com> --------- Signed-off-by: Ning Sun <sunning@greptime.com> |
||
|
|
32e215cad0 |
fix(event): preserve procedure lifecycle locators (#8787)
* fix(event): preserve procedure lifecycle locators Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(event): preserve dropped table lifecycle locators Signed-off-by: WenyXu <wenymedia@gmail.com> * test(event): cover lifecycle locators Signed-off-by: WenyXu <wenymedia@gmail.com> * test(event): fix lifecycle context expectations Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
4eeb4052d6 |
feat(servers): stamp prometheus remote write v2 metadata as semantic table options (#8797)
* feat(servers): stamp prometheus remote write v2 metadata as semantic table options
Remote write 2.0 carries per-series metadata (type, unit, help) that the
v2 ingest decoded and dropped; tables kept the name-based 'inferred'
quality. Wire it into the semantic layer:
- generalize the OTLP per-table semantic index into a shared, schema-
aware servers::semantic module: v2 lets each series override its
target schema, so the index is keyed {schema -> table -> options} and
the same metric name in two schemas no longer collapses;
- into_write_requests records metric type and unit per written table;
an explicit type upgrades the table's metadata quality to declared,
UNSPECIFIED series keep the request-level inferred stamp, and units
are canonicalised from OpenMetrics words to the UCUM codes the
vocabulary is defined in (unknown units are dropped, help text is not
persisted);
- the consumer folds the index in on both auto-create paths: the
operator row-insert path and the pending-rows batched create, which
bypasses the former.
Table options are stamped at auto-create only; updating existing tables
from later metadata is future work (a metadata registry, see the native
histograms RFC).
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* chore: trim over-commenting in the remote write metadata path
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix: parse the per-table semantic index once per create round
The index was re-parsed from JSON for every table being created — a
first write creating N tables (a fleet's first scrape) paid
O(N x index size). Parsing now happens lazily once per create-planning
round, on both consumers: the operator row-insert auto-create (also
serving OTLP metrics) and the pending-rows batched create.
Also validate every non-zero metadata symbol reference up front, as the
remote write 2.0 spec requires: help_ref was never checked, and
unit_ref escaped checking when the metric type was UNSPECIFIED.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(servers): stamp remote-write v2 unit independently of metric type
OpenMetrics models TYPE and UNIT as independent MetricFamily metadata, and
the Prometheus v2 sender emits UNSPECIFIED-type series that still carry a
unit. The early return on UNSPECIFIED dropped that unit, which is
unrecoverable after table auto-create (units are not stored in rows).
Stamp the mapped unit whenever present; the type and the
metadata_quality=declared upgrade still require an explicit type.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* chore: trim restating comments in the v2 metadata path
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
|
||
|
|
00e45c8602 |
test: fix unknown channel event context expectation (#8832)
Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
1f1c9270a8 |
feat(event): add event context to procedure events (#8734)
* feat(event): add trigger context to procedure events Signed-off-by: WenyXu <wenymedia@gmail.com> * test(event): fix trigger context event contracts Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(event): preserve trigger context origins Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(event): preserve lifecycle trigger contexts Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(meta): gate enterprise trigger imports Signed-off-by: WenyXu <wenymedia@gmail.com> * test(event): assert trigger contexts exactly Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(test): restore migration test literals Signed-off-by: WenyXu <wenymedia@gmail.com> * feat(event): add trigger context to non-table ddl Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(event): simplify trigger context encoding Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(event): preserve auto alter trigger reason Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(event): require explicit trigger context Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(event): derive trigger context at ddl boundary Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(event): propagate DDL trigger reasons Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(meta): resolve alter table rebase conflict Signed-off-by: WenyXu <wenymedia@gmail.com> * test(meta): fix alter table trigger context setup Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(meta): pass trigger context to region migration Signed-off-by: WenyXu <wenymedia@gmail.com> * test(event): fix unknown trigger context assertions Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(meta): centralize trigger context protocol mapping Signed-off-by: WenyXu <wenymedia@gmail.com> * test(event): fix migration trigger context assertion Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(procedure): rename event runtime context Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(event): rename trigger context Signed-off-by: WenyXu <wenymedia@gmail.com> * test(event): fix migration event context assertion Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
57b8239ff8 |
test: rename internal bug numbers in tests to semantic names (#8779)
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com> |
||
|
|
3d12273c84 |
feat: read-time entity relationships graph over telemetry (M0+M1) (#8614)
* feat(table): add entity semantic declarations Define open-ended greptime.semantic.entity.* options, validate entity columns at DDL time, and stamp OTLP trace tables with the service entity declaration. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * feat: add read-time entity relationships graph Add computed semantic graph tables, typed DataFusion derivation plans for entity registry and trace calls edges, and streaming read-time execution. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test: exclude semantic graph tables from table constraints Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * refactor(operator): name the plan-builder source groupings Review feedback: build_registry_plan / build_calls_plan took anonymous (declarations, DataFrame) tuples while the caller already grouped the same fields. Introduce RegistrySource { declarations, scan } and CallsSource { service, scan } next to the builders and flow them through the frontend caller and tests. The frontend-side EntitySource keeps holding a TableRef (the operator builders stay pure over already-built scans), so the named structs live in operator rather than reusing that type. No behavior change. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> --------- Signed-off-by: Dennis Zhuang <killme2008@gmail.com> |
||
|
|
f25836a6ed |
fix: support Utf8View labels in Prometheus response (#8754)
Signed-off-by: evenyag <realevenyag@gmail.com> |
||
|
|
8026064659 |
feat: add health-aware gRPC client routing (#8684)
* feat: add gRPC client health routing Signed-off-by: WenyXu <wenymedia@gmail.com> * fix: harden gRPC client health routing Signed-off-by: WenyXu <wenymedia@gmail.com> * fix: defer gRPC client health checks until first use Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
aa72563783 |
refactor!: move native histogram config and prom_validation_mode to prom_store (#8744)
* chore: adjust the position of experimental_enable_prometheus_native_histogram Signed-off-by: shuiyisong <xixing.sys@gmail.com> * chore: move prom_validation_mode as well Signed-off-by: shuiyisong <xixing.sys@gmail.com> --------- Signed-off-by: shuiyisong <xixing.sys@gmail.com> |
||
|
|
c55f297dec |
chore!: gate soft-drop table behind the enterprise feature (#8747)
* chore: gate soft-drop table behind the enterprise feature Soft-drop table becomes an enterprise-only feature: - metasrv rejects gc.experimental_soft_drop.enable=true at startup in non-enterprise builds, and ddl_soft_drop_enabled is hard-disabled without the enterprise feature as a second line of defense - the UNDROP TABLE parser/AST/statement variant, ADMIN purge_table() registration, and information_schema.recycle_bin registration are compiled out unless the enterprise feature is enabled - common-meta procedures, tombstone keys, and DdlTask serde stay unconditional for persisted-procedure recovery and wire compatibility - the [gc.experimental_soft_drop] section is removed from the OSS example config and generated docs (moving to the enterprise repo) - the soft-drop sqlness cases and their CI job are removed from OSS (moving to the enterprise repo); affected information_schema .result files are regenerated Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * refactor: limit unused_variables allow to non-enterprise builds Addresses review comment: apply the allow via cfg_attr so enterprise builds still catch accidental unused variables in register_admin_only. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * refactor: include the config key in the soft-drop enterprise gate error Addresses review comment: name gc.experimental_soft_drop.enable in the startup validation error so users can locate the setting quickly when it is set via env vars or layered config. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * test: limit unused_mut allow to non-enterprise builds Addresses review comment: apply the allow via cfg_attr so enterprise builds still catch unused mut in the table_ddl_event test setup. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * feat: reject soft-drop DDL submissions in non-enterprise builds Addresses review comment: clients could bypass the SQL-level gates by submitting DdlTask::UndropTable or DdlTask::PurgeDroppedTable directly to the procedure service. Reject fresh submissions at the DdlManager boundary in non-enterprise builds while keeping the procedure loaders registered for crash recovery and wire compatibility. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * test: stop --enable-gc from enabling soft drop in the sqlness template Addresses review comment: the metasrv test template rendered [gc.experimental_soft_drop] enable = true under the generic --enable-gc flag, which non-enterprise metasrv now rejects at startup, making the documented --enable-gc mode unusable in OSS. Keep the flag scoped to plain GC; enterprise soft-drop coverage moves to the enterprise repo. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * fix: gate fresh soft-drop procedures Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * test: gate soft-drop fallback coverage Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * fix: gate soft-drop procedure implementation Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * refactor: gate drop table soft-drop behavior Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * refactor: gate expired soft-drop gc behavior Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * ci: test enterprise table ddl lifecycle Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * chore: mark purge_table as enterprise licensed The purge_table module is compiled only with the enterprise feature, so apply the Enterprise License header and register it with both license header configurations. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * chore: mark recycle_bin as enterprise licensed The recycle_bin module is compiled only with the enterprise feature, so apply the Enterprise License header and register it with both license header configurations. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> * chore: mark soft-drop procedure sources as enterprise licensed The purge and undrop procedure implementations plus the recycle-bin test module compile only with the enterprise feature. Apply the Enterprise License header and register them with both license configurations. Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> --------- Signed-off-by: Lei, HUANG <ratuthomm@gmail.com> |
||
|
|
e58f21ed6d |
feat(logging): add enable_file_logging option to disable file logging (#8721)
Signed-off-by: xhwhis <hi@whis.me> |
||
|
|
1aa35716cb |
refactor(event): store procedure trigger as JSONB (#8700)
* refactor(event): store procedure trigger as JSONB Signed-off-by: WenyXu <wenymedia@gmail.com> * test(event): fix JSONB procedure trigger assertions Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(event): remove procedure trigger display Signed-off-by: WenyXu <wenymedia@gmail.com> * test(event): update batch GC trigger assertions Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
8f11629e34 |
feat(metasrv): add batch GC lifecycle events (#8673)
* feat(metasrv): add batch GC lifecycle events Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(metasrv): reduce batch GC event fanout Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(metasrv): fix batch GC event import Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(metasrv): refine batch GC events Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(metasrv): preserve batch GC reports on failure Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(metasrv): retain batch GC reports on retry Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(metasrv): harden batch GC report merging Signed-off-by: WenyXu <wenymedia@gmail.com> * test: scope repartition SST assertions to target table Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
c3d8c976a2 |
feat(query): add native histogram result plumbing (#8693)
* feat(query): add native histogram result plumbing Signed-off-by: shuiyisong <xixing.sys@gmail.com> * fix: cr issue Signed-off-by: shuiyisong <xixing.sys@gmail.com> --------- Signed-off-by: shuiyisong <xixing.sys@gmail.com> |
||
|
|
979a22b38a |
feat(metasrv): record WAL prune procedure events (#8677)
* feat(metasrv): record WAL prune procedure events Signed-off-by: WenyXu <wenymedia@gmail.com> * feat(metasrv): expand WAL prune procedure events Signed-off-by: WenyXu <wenymedia@gmail.com> * fix: fix toml fmt Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(metasrv): clarify WAL prune event semantics Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
47ca5c362e |
feat: add table DDL procedure events (#8627)
* feat(meta): emit table DDL procedure events Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(meta): honor table DDL event filters Signed-off-by: WenyXu <wenymedia@gmail.com> * test(meta): cover table DDL event filters Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(meta): align table DDL event conventions Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(meta): bound table DDL event payloads Signed-off-by: WenyXu <wenymedia@gmail.com> * test(meta): consolidate table DDL event tests Signed-off-by: WenyXu <wenymedia@gmail.com> * style(meta): use crate visibility in event tests Signed-off-by: WenyXu <wenymedia@gmail.com> * test: stabilize table DDL event assertions Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(meta): exclude repartition from alter table events Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(meta): resolve table event rebase conflicts Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
31f9a9a6fd |
feat(metasrv): add repartition lifecycle events (#8665)
* feat(metasrv): add repartition lifecycle events Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(metasrv): simplify event module names Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(procedure): emit submitted events for child procedures Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(metasrv): flatten repartition event payload Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(metasrv): defer repartition topology rows Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(procedure): avoid events on failed child spawn Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
deb688f572 |
feat: add a dedicated http api server port (#8657)
* feat: add a dedicated http api server port * fix: integration test * refactor: make http-api-port opt-in * refactor: rename attribute to http-api-server * feat: use middleware to check different http server port * refactor: rename config option |
||
|
|
8ca6132b84 |
feat: add events for create and drop view (#8626)
* feat(procedure): add view ddl events Signed-off-by: WenyXu <wenymedia@gmail.com> * test(procedure): satisfy view event clippy Signed-off-by: WenyXu <wenymedia@gmail.com> * feat(meta): add view DDL procedure events Signed-off-by: WenyXu <wenymedia@gmail.com> * test(meta): group view event tests Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(meta): align view DDL events Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(meta): centralize view event schema Signed-off-by: WenyXu <wenymedia@gmail.com> * test(integration): use singular view event module Signed-off-by: WenyXu <wenymedia@gmail.com> * test(integration): align view event assertions Signed-off-by: WenyXu <wenymedia@gmail.com> * test(integration): share DDL event assertions Signed-off-by: WenyXu <wenymedia@gmail.com> * docs: document view event recorder types Signed-off-by: WenyXu <wenymedia@gmail.com> * refactor(meta): align view DDL event conventions Signed-off-by: WenyXu <wenymedia@gmail.com> * test(meta): simplify view event tests Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
775a9af3b8 |
feat: add procedure events for Flow DDL (#8632)
* feat(meta): record Flow DDL procedure events Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(meta): remove query schema from Flow events Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(tests): fix Flow DDL event test Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |
||
|
|
09e1d24365 |
feat: add database DDL procedure events (#8623)
* feat(meta): add database DDL procedure events Signed-off-by: WenyXu <wenymedia@gmail.com> * test(sqlness): disable event recording Signed-off-by: WenyXu <wenymedia@gmail.com> * test: poll database DDL event assertions Signed-off-by: WenyXu <wenymedia@gmail.com> * docs(config): list database DDL event types Signed-off-by: WenyXu <wenymedia@gmail.com> * fix(test): satisfy clippy Signed-off-by: WenyXu <wenymedia@gmail.com> --------- Signed-off-by: WenyXu <wenymedia@gmail.com> |