* feat: add CORS support for gRPC-Web on the frontend gRPC server
A browser-based gRPC-Web client cannot read a cross-origin response. The
request carries a non-simple content type, so the browser sends an
`OPTIONS` preflight first, and it drops the call when the response has no
`Access-Control-Allow-Origin`.
Add `enable_cors` and `cors_allowed_origins` to `[grpc]`, threaded from
`GrpcOptions` to `GrpcServerConfig` the same way `tls` is.
The CORS layer sits outside `tonic_web::GrpcWebLayer`, so it answers the
preflight before the request reaches the gRPC routes. It exposes
`grpc-status`, `grpc-message` and `grpc-status-details-bin`, which a
browser cannot read otherwise. An empty `cors_allowed_origins` allows any
origin.
`enable_cors` defaults to false, unlike `[http]`. `GrpcOptions` is shared
by the public `[grpc]` section and the internal `[internal_grpc]` one, and
serde cannot tell them apart, so a true default would also turn CORS on
for the internal gRPC listeners: the frontend internal gRPC server, and
the datanode and flownode gRPC servers. None of them authenticate callers,
and a browser on the host or in the cluster network can reach all of them.
Set `enable_cors = true` to turn it on.
Tests start a real server on an ephemeral port and cover the preflight, a
custom origin list, and the disabled case.
Update the example TOMLs and regenerate config/config.md.
Signed-off-by: lczllx <2181719471@qq.com>
* fix: drop the redundant OPTIONS from the gRPC CORS allow_methods
Signed-off-by: lczllx <2181719471@qq.com>
* test: assert the allow-headers and allow-methods headers in the gRPC CORS preflight
The preflight answers with `access-control-allow-headers: *` from
`AllowHeaders::any()`, and with `access-control-allow-methods: post` from the
POST-only `allow_methods` list. Pin both, the way the HTTP CORS test pins its
own headers: a missing allow-headers header would let the preflight pass the
origin check and still have the browser block every gRPC-Web call, which a
non-browser client would never notice.
Signed-off-by: lczllx <2181719471@qq.com>
* docs(config): stop the example configs from setting the new gRPC CORS keys
Signed-off-by: lczllx <2181719471@qq.com>
* test(servers): cover the exposed gRPC CORS headers and pin the option plumbing
Signed-off-by: lczllx <2181719471@qq.com>
* docs(config): document the gRPC CORS origin example with `#+`
Signed-off-by: lczllx <2181719471@qq.com>
* fix(frontend): keep CORS off the internal gRPC server
Signed-off-by: lczllx <2181719471@qq.com>
* docs(grpc): document that the internal gRPC server never serves CORS
Signed-off-by: lczllx <2181719471@qq.com>
* Update src/servers/src/grpc.rs
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* chore: trim redundant comments
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor: share CORS origin parsing and simplify gRPC-Web CORS tests
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor: enable gRPC CORS only on the frontend public server
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* docs(config): clarify the gRPC CORS options
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: lczllx <2181719471@qq.com>
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
Co-authored-by: Dennis Zhuang <killme2008@gmail.com>
* fix(promql): align timestamp() and label_join with Prometheus semantics
- timestamp() over a selector reports the selected sample's timestamp, without adding the offset.
- timestamp() over any other expression reports the evaluation time instead of the input value.
- label_join that overwrites an existing label rejects duplicate label sets at runtime.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(promql): align label_replace and label_join edge cases with Prometheus
- label_replace leaves non-matching series unchanged instead of copying the source value into the destination label.
- label_replace may overwrite an existing label; duplicate label sets are rejected at runtime like label_join.
- label_join with no source labels removes the destination label, and an empty source label name is rejected.
- Empty label values produced by these functions are NULL, the same as an absent label.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(promql): join absent labels as empty strings in label_join
concat_ws skips NULL arguments together with their separator, so a label that is absent (NULL) dropped its separator, e.g. label_join(label_join(vector(1), "a", "", "missing"), "b", "-", "a", "a") lost b="-". Read every source label as coalesce(label, '') like label_replace does.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(promql): skip rows without a sample in timestamp()
timestamp() replaced the input fields with the timestamp before the empty-value filter ran, so a row whose fields were all NULL got a timestamp, e.g. timestamp(-m) for a NULL sample. Filter such rows before the projection, for both the selector and the evaluation-time paths.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix: bound decompressed request size and close memory-admission gaps for compressed requests
The request-memory accounting only bounds and charges the encoded bytes on
the wire, but a compressed body can expand far beyond that during
decompression, so tiny requests could allocate disproportionate frontend
memory before any protobuf validation or quota charge.
- Handler-level decompression (Prometheus remote read/write v1+v2, Loki)
now enforces a hard 512 MiB decoded-size cap, checked before any output
buffer is allocated, and charges the decoded bytes to the shared
ServerMemoryLimiter, holding the permits for the lifetime of the
decompressed buffer.
- gRPC requests with transport compression reserve the configured
max_recv_message_size before tonic decompresses, so the decoding phase
is admitted against max_in_flight_write_bytes; the later per-message
charge is skipped to avoid double accounting.
- The HTTP memory-limit middleware keeps its upfront Content-Length charge
but now also accounts the bytes actually streamed beyond it, so chunked
requests and understated Content-Length headers no longer bypass the
aggregate quota.
- Routes that decompress request bodies via RequestDecompressionLayer
(InfluxDB, OTLP, Loki, Splunk, Elasticsearch, pipelines, dashboards)
now charge the decompressed bytes as handlers consume them: the global
middleware marks Content-Encoding requests, and a route-local
accounting layer inside the decompression layer charges the decoded
stream. Plain requests are skipped to avoid double-counting.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: address review comments on memory admission guard lifetimes and 429 mapping
- Retain the gRPC pre-decode reservation for the whole request: the
extensions holding the guard were dropped when the request was consumed
(into_inner), releasing the reservation while per-message charges
stayed skipped. Both the unary and streaming handlers now clone and
hold the reservation marker for the duration of request handling.
- Retain the HTTP body permits across the handler: the AccountedBody
wrapper and the request extensions are dropped once the extractors
finish collecting the body, before the handler is done with the decoded
data. Both accounting middlewares now keep their own accounting handle
alive across next.run(req).await.
- Return a ChargedBuffer from the Loki snappy decompressor so the
reservation outlives the decompressed bytes through the caller's
protobuf decoding, matching the Prometheus path.
- Map mid-stream quota exhaustion to 429 instead of the generic body
error (400): the accounting flags exhaustion and the middlewares
rewrite the extractor rejection, so clients can distinguish
backpressure from malformed input.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: resolve the in-flight body acquisition before trying a new one
A parked acquisition in AccountedBody::charge always belongs to the
currently buffered frame, so it must be resolved before any new
acquisition is tried. Letting the fast-path try_acquire succeed while a
waiter is parked left the waiter alive, and its late completion would
then be credited with a later frame's byte count, under-reserving memory
relative to what was marked charged.
Adds a regression test that reproduces the misattribution: with the
buggy ordering the test delivers a body the quota cannot cover; with the
fix the over-quota frame waits for its own acquisition and times out.
Signed-off-by: Ning Sun <sunning@greptime.com>
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
* feat: allow customized time index unit for metric engine table
* test: provide query tests
* refactor: revert unnecessary change
* refactor: share timestamp unit conversions in api helper
Address review feedback on the time index unit changeset:
- Add shared timestamp_unit/timestamp_datatype helpers to api::helper
(the only crate that sees both proto ColumnDataType and TimeUnit due
to layering; common-time and datatypes have no greptime-proto dep).
This removes the ColumnDataType -> TimeUnit match duplicated between
operator's insert path and the OTLP logs path.
- Collapse the two TimeUnit <-> ValueData matches in
convert_timestamp_value_data by reusing api::helper::to_grpc_value
for the construction side.
- Note that convert_rows_time_unit rewrites the schema before the
values, so an overflow mid-batch leaves the request half-converted;
harmless because the error aborts the whole insert request.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: align time units per destination table and floor remote-read timestamps
Address review feedback on PR #9236:
- Align each metric insert request to the unit of the table it actually
targets: an existing logical table keeps its own unit (it may be bound
to a different physical table than the one selected by the request),
and only new tables use the selected physical table's unit. The
previous blanket conversion rewrote valid millisecond samples to the
selected physical table's unit and the engine rejected them.
Regression test: writing an existing millisecond logical table and a
new table in one request that selects a microsecond physical table.
- Remote read now floors narrowing timestamp conversions towards
negative infinity (div_euclid), consistent with
Timestamp::convert_to on the ingestion path; arrow's cast truncates
towards zero and returned -1ms for a stored -1001us. Widening
(second -> millisecond) keeps the exact arrow cast. Regression test:
a negative, non-aligned timestamp round-trips as -2ms.
Signed-off-by: Ning Sun <sunning@greptime.com>
* perf: fold time unit alignment into existing table lookups
Address review feedback on PR #9236:
- The per-destination unit alignment no longer runs its own pass of
table lookups: create_or_alter_tables_on_demand gains an
align_time_index_unit parameter (metric engine path only) and
converts each request inside the lookups it already performs —
existing tables to their own unit, new tables to the selected
physical table's. Default ingest paths now issue zero additional
catalog lookups compared to main; the separate alignment pass remains
only in the opt-in logical batcher pre-gate, next to the eligibility
check that already looks up the same tables.
- convert_rows_time_unit indexes the time index position directly
(validate_column_count_match guarantees row widths) instead of
Optional get_mut; the gate-side alignment validates widths itself.
Signed-off-by: Ning Sun <sunning@greptime.com>
* perf: resolve the batcher time index guard once per write target
All batches of one remote write request share the same write target
(catalog, schema, physical table), so the batcher time index guard now
resolves each distinct target once instead of once per batch.
Signed-off-by: Ning Sun <sunning@greptime.com>
* feat: support non-millisecond time index units in the logical batcher
Make the logical table batcher's bulk encode path unit-aware so physical
metric tables with a non-millisecond time index (e.g. TIMESTAMP(6)) can
use logical batching instead of falling back to the ordinary insert path.
- rows_to_aligned_record_batch builds the time index column in the
TARGET schema's unit, converting any timestamp encoding via
Timestamp::convert_to (flooring on narrowing, consistent with the
ordinary insert path).
- New tables created by the batcher use the selected physical table's
time index unit (resolved once per submit; a missing physical table
keeps the millisecond auto-create default).
- columns_taxonomy and the can_batch_metric_rows schema whitelist accept
any timestamp unit; the prometheus remote write v1/v2 batcher gates
and the OTLP pre-gate alignment are removed together with
Inserter::align_metric_row_inserts_time_unit, as the batcher now
converts internally.
Closes#9342
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: address review comments
* fix: address review issue
* refactor: drop the OTLP pre-gate unit alignment made redundant by the bulk path
The main merge of #9236 (squash) resurrected the OTLP pre-gate
alignment and Inserter::align_metric_row_inserts_time_unit, which this
branch had removed. Drop them again:
- The pre-gate existed because the #9236-era bulk eligibility gate only
accepted millisecond schemas, so nanosecond-encoded OTLP requests had
to be converted before the check. This branch makes the bulk path
unit-aware (the gate accepts all time index units and batch alignment
converts each request to its destination's unit), so the pre-gate is
redundant and only added N+1 catalog lookups per batched request —
the very lookup-count overhead raised in the #9236 review.
- The per-destination unit semantics it implemented remain enforced in
the two paths that need them: the ordinary insert path
(create_or_alter_tables_on_demand converts inside its existing table
lookups) and the batched path (batch alignment resolves each
destination schema and converts to it).
test_otlp_logical_batcher_alignment (the test the pre-gate originally
fixed) and the mixed-physical-table regression both pass without it.
Signed-off-by: Ning Sun <sunning@greptime.com>
* test: cover OTLP batcher cross-physical fallback and nanosecond physical
Extend the logical batcher integration coverage for the cases
previously guarded by the removed OTLP pre-gate alignment:
- test_otlp_logical_batcher_fallback_for_cross_physical_destination:
with the batcher enabled, an OTLP request targeting an existing
logical table bound to another physical table must NOT enter the
batcher (the bulk eligibility check rejects the destination binding)
and the ordinary insert path must convert it to the destination's
unit (60s -> 60_000_000us).
- test_otlp_logical_batcher_non_millisecond_physical_table now covers
both microsecond and nanosecond physical tables (parameterized),
asserting batcher submissions and unit-precise stored values.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: validate per-batch physical bindings and avoid intermediate timestamp buffers
Address review feedback on PR #9346:
- accepts_bulk_destinations dedupes on (schema, table, selected physical)
instead of (schema, table): one request can select different physical
tables per batch (per-series x_greptime_physical_table labels), and the
old key let a second selection skip validation and flush rows through
the wrong physical's regions. Missing tables additionally reject
conflicting physical selections within the same request. Regression
test covers an existing destination, a missing destination, and a
consistent selection (which must still batch).
- The timestamp column builder appends each value directly into the
target-unit Arrow builder; values already in the target unit (the
unchanged millisecond fast path) are appended without conversion, so
the default millisecond physical pays no Timestamp construction or
intermediate Vec allocation.
- The non-millisecond batching tests assert the submit_build_and_align
counter, which increments on every batcher submission in both
acknowledgement modes, so a silent fallback to ordinary insertion
fails the tests instead of passing on stored values alone.
Signed-off-by: Ning Sun <sunning@greptime.com>
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
* feat: allow customized time index unit for metric engine table
* test: provide query tests
* refactor: revert unnecessary change
* refactor: share timestamp unit conversions in api helper
Address review feedback on the time index unit changeset:
- Add shared timestamp_unit/timestamp_datatype helpers to api::helper
(the only crate that sees both proto ColumnDataType and TimeUnit due
to layering; common-time and datatypes have no greptime-proto dep).
This removes the ColumnDataType -> TimeUnit match duplicated between
operator's insert path and the OTLP logs path.
- Collapse the two TimeUnit <-> ValueData matches in
convert_timestamp_value_data by reusing api::helper::to_grpc_value
for the construction side.
- Note that convert_rows_time_unit rewrites the schema before the
values, so an overflow mid-batch leaves the request half-converted;
harmless because the error aborts the whole insert request.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: align time units per destination table and floor remote-read timestamps
Address review feedback on PR #9236:
- Align each metric insert request to the unit of the table it actually
targets: an existing logical table keeps its own unit (it may be bound
to a different physical table than the one selected by the request),
and only new tables use the selected physical table's unit. The
previous blanket conversion rewrote valid millisecond samples to the
selected physical table's unit and the engine rejected them.
Regression test: writing an existing millisecond logical table and a
new table in one request that selects a microsecond physical table.
- Remote read now floors narrowing timestamp conversions towards
negative infinity (div_euclid), consistent with
Timestamp::convert_to on the ingestion path; arrow's cast truncates
towards zero and returned -1ms for a stored -1001us. Widening
(second -> millisecond) keeps the exact arrow cast. Regression test:
a negative, non-aligned timestamp round-trips as -2ms.
Signed-off-by: Ning Sun <sunning@greptime.com>
* perf: fold time unit alignment into existing table lookups
Address review feedback on PR #9236:
- The per-destination unit alignment no longer runs its own pass of
table lookups: create_or_alter_tables_on_demand gains an
align_time_index_unit parameter (metric engine path only) and
converts each request inside the lookups it already performs —
existing tables to their own unit, new tables to the selected
physical table's. Default ingest paths now issue zero additional
catalog lookups compared to main; the separate alignment pass remains
only in the opt-in logical batcher pre-gate, next to the eligibility
check that already looks up the same tables.
- convert_rows_time_unit indexes the time index position directly
(validate_column_count_match guarantees row widths) instead of
Optional get_mut; the gate-side alignment validates widths itself.
Signed-off-by: Ning Sun <sunning@greptime.com>
* perf: resolve the batcher time index guard once per write target
All batches of one remote write request share the same write target
(catalog, schema, physical table), so the batcher time index guard now
resolves each distinct target once instead of once per batch.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: address review comments
* fix: address review issue
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix(client): complete transport lane isolation
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* refactor(client): deprecate legacy single-manager constructors
Per review: mark the legacy single-manager constructors and helper as
deprecated so callers move to the isolated query/control manager pair.
Tests intentionally exercising the legacy path are annotated with
`#[allow(deprecated)]`.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* fix(client): update deprecated constructor callers for CI
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
---------
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* feat(mito): limit series index disk usage
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito): enforce series index quota during reconciliation
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito): remove series index disk budget layer
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito): simplify series index limit to estimated usage
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito): trust index catalogs when loading snapshots
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito): minimize series index disk limit changes
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix(mito2): keep series index cleanup running at capacity
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix(mito2): avoid no-op index clones and stabilize capacity tests
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix: update config API expectation and stabilize index build tests
Signed-off-by: evenyag <realevenyag@gmail.com>
---------
Signed-off-by: evenyag <realevenyag@gmail.com>
* feat: serialize struct to json in postgres
* fix: support view scalars and preserve null structs in scalar-to-value conversion
Address PR review:
- Utf8View/BinaryView ScalarValues now convert like their non-view forms
instead of failing row extraction for struct columns
- a null struct scalar converts to Value::Null so a null struct inside a
list stays null in the serialized JSON
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: return errors instead of panics for unsupported arrow field types
Struct-typed query results with arrow field types greptimedb cannot
represent (e.g. Decimal256) used to panic during schema conversion and
row extraction, dropping the client connection. They now surface as
query errors:
- ConcreteDataType::try_from builds struct types fallibly via the new
StructType::try_from_arrow_fields
- Value::try_from(ScalarValue::Struct) uses the same fallible path
- new try_value_from_array converts an arrow element to Value with
error propagation, used by the postgres struct encoding
Signed-off-by: Ning Sun <sunning@greptime.com>
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
* test: cut integration test time and make the storage matrix meaningful
tests-integration is ~85% of workspace test CPU, and 81% of that is the
S3/S3WithCache variants of the HTTP and gRPC suites. Those suites do not
touch the object store: of the 70 matrix HTTP tests only one flushed and
read back an SST, so the matrix was paying real AWS round trips to
re-prove protocol parsing.
- Point the PR CI object-store matrix at the MinIO already started by
tests-integration/fixtures. Three GT_S3_* consumers did not read
GT_S3_ENDPOINT_URL and would have hit real AWS with MinIO credentials;
they now do.
- Add a nightly Linux job against real AWS S3, and pass GT_S3_* into the
release integration-test container. The release previously ran every
remote-backend case as a skip and only exercised the file backend.
- Give each S3WithCache test its own read cache directory. They shared
/tmp/greptimedb_cache, which the datanode wipes on startup, so a
starting test deleted the read cache of a running one.
- Add flush -> read-back assertions to the tests whose columns have a
non-trivial SST representation: JSON/JSON2 columns, native histograms,
metric-engine logical tables, and tables carrying fulltext or skipping
indexes whose puffin files only exist after a flush.
- Move eight tests that create no table out of the storage matrix.
- Make the event recorder flush interval a constructor parameter and
shorten it in the event tests, which otherwise wait a 5s window per DDL
they assert on. It is skipped by serde and never reaches config files.
- Drop duplicates: test_grpc_zstd_compression was a verbatim copy of
test_grpc_message_size_ok and is now rewritten to assert the negotiated
grpc-encoding; test_execute_copy_to_{s3,oss,gcs,azblob} were strict
prefixes of their copy_from siblings; two standalone/distributed event
test pairs shared one assertion body.
- Fix and un-ignore stddev_by_label. stddev_pop merges partial aggregates
in a parallelism-dependent order, so its last digits are unstable; the
test now compares values with a tolerance.
- Rebase the jaeger v1 fixture on the current instant. It carries
ttl=7d with 2025 timestamps, so its rows were only readable as long as
they stayed in the memtable.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* test: address review — wire nightly real-S3 job into check-status, keep the short event interval
The nightly `check-status` job did not depend on the new real-S3 job, so a
failure there would not have reached the status or Slack notification.
In database_ddl_event the short interval was set by a first
`with_event_recorder_options` call and then overwritten by the pre-existing
one, which carries `..Default::default()`. Merged into a single call.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(ci): build tests-integration lib with meta-srv/mock
tests-integration's lib code (src/cluster.rs) uses meta_srv::mocks, but
the dependency carrying the mock feature sits in [dev-dependencies].
Builds that only touch the lib, such as the apidoc job's cargo doc
--workspace, resolve meta-srv without mock and fail with E0432.
--all-targets builds unify dev-dependency features, which is why check,
clippy and nextest stayed green.
Move the mock-enabled meta-srv entry back to [dependencies]. The other
testing features moved out in #9072 are not needed by the lib and stay
in [dev-dependencies].
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* test(repartition): split per-case repartition tests
test_repartition_metric ran four format/primary-key-encoding cases in a
single test function, and test_repartition_mito ran two format cases.
Each case builds its own 3-datanode cluster and runs a full repartition
plus GC cycle, so on S3 the metric test took 165-178s against the 180s
nextest terminate-after. Merge queue runs failed on it at random.
Split each case into its own test. Cases were already independent, so
they now run in parallel and each stays far inside the timeout, and a
failure points at one encoding instead of four.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* feat: add config flag to control range index reads
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix: disable range index builds when configured and default to off
Signed-off-by: evenyag <realevenyag@gmail.com>
---------
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix(prometheus): honor label matchers in __name__ values query
`/api/v1/label/__name__/values?match[]={pod="abc"}` dropped every matcher
other than `__name__` and returned all metrics in the schema. No error,
just the wrong list. Grafana's metrics browser sends this request, so
picking a label value there did nothing.
Selectors that only constrain `__name__` keep answering from table
metadata. A selector constraining an ordinary label now goes to the data:
scan each metric engine physical table for distinct `__table_id` in the
time range, map the ids back to metric names, then apply the selector's
own `__name__` matchers.
Only metric engine tables are covered; other engines share no column space
to scan.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor(prometheus): batch-resolve metric names by table id
Building a full table-id-to-name map meant walking every table in the
schema and holding all of them in memory, just to name the handful the
scan returned. Use `tables_by_ids` instead — one batch KV read over the
ids the scan actually produced.
The catalog walk stays, but only to find the physical tables to scan.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(promql): read an absent label as the empty string
A matcher on a label the series does not carry only worked when the table
had no column for it at all. Where the column exists but is NULL on that
row -- the norm for logical metrics sharing a metric engine physical
table, which holds the union of their label columns -- three-valued logic
dropped the row, so `host!="host1"` and `host=""` missed every metric
without a host label.
Coalesce nullable string label columns to "" for matchers that accept the
empty string, rather than only for the OTLP temporality marker. Equality
matchers are untouched; they cannot match NULL either way.
This is the Prometheus compatibility fix#8970 deliberately kept out of
its own scope. The cost is visible in the regex sqlness plan: the
predicate becomes a CASE, so the scan loses its LastRow selector and
grows a FilterExec. Only negative and empty-accepting matchers pay it.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(promql): don't panic on pre-epoch label value bounds
`rewrite_label_values_query` unwrapped `duration_since(UNIX_EPOCH)`, which
returns an error for an instant before the epoch. `start=1969-12-31T23:59:59Z`
parses as valid RFC3339, so the request panicked instead of answering.
Recover the sign from the error branch, and report a value beyond i64
milliseconds as an error rather than wrapping the cast.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor(prometheus): drop applicable_matchers, share the distinct scan
With the planner reading an absent label as empty, the frontend no longer
needs to pre-filter matchers per physical table. Removing that exposed a
second problem: a physical table that never took a column from a logical
table exposes no `__table_id`, and projecting it failed the whole request.
Skip those tables; the only thing that can miss is a metric with no labels.
Also pulls out the plan-build-execute-collect sequence the two label value
scans had in common.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* feat(mito2): reconcile series indexes in background
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix(mito2): clean up series indexes published during region drop
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix(config): align series index examples with upstream enable flag
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix(mito2): run series index tasks on compaction runtime
Signed-off-by: evenyag <realevenyag@gmail.com>
---------
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix: use join_dir for series index config path
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix: configure series indexes with an enable flag
Signed-off-by: evenyag <realevenyag@gmail.com>
* docs: omit experimental series index from example configs
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix: preserve legacy cache cleanup path behavior
Signed-off-by: evenyag <realevenyag@gmail.com>
* test: remove trivial path joining tests
Signed-off-by: evenyag <realevenyag@gmail.com>
* test: isolate worker group WAL directories on Windows
Signed-off-by: evenyag <realevenyag@gmail.com>
---------
Signed-off-by: evenyag <realevenyag@gmail.com>
Building a matrix response allocated one `String` per sample while the
record batches were scanned, then dropped it after the JSON body was
written. Keep the `f64` in `PromSampleValue::Number` instead and format
it with ryu while serializing, so no per-sample string is allocated.
`PromSampleValue::Text` keeps values parsed from a JSON body, so
deserializing and re-serializing a response is unchanged. Vector and
scalar results still expose `String`, since they hold a single sample.
Signed-off-by: Dennis Zhuang <xzhuang@greptime.com>
* refactor(mito2): add series index catalog and lifecycle components
Signed-off-by: evenyag <realevenyag@gmail.com>
* feat(mito2): restore series index catalogs on region open
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito2): simplify series index foundation and maintenance
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito2): separate series index purge task and simplify tests
Signed-off-by: evenyag <realevenyag@gmail.com>
* docs: defer experimental series index configuration examples
Signed-off-by: evenyag <realevenyag@gmail.com>
* feat(mito): make series index maintenance interval configurable
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix(mito2): correct series index cleanup on close and drop
Signed-off-by: evenyag <realevenyag@gmail.com>
* test(mito2): revert drop test changes
Signed-off-by: evenyag <realevenyag@gmail.com>
* test: update config API expectation for series index settings
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor: use tokio unbounded channel for series index purger
Signed-off-by: evenyag <realevenyag@gmail.com>
* chore(mito2): simplify review test scope and clarify index config
Signed-off-by: evenyag <realevenyag@gmail.com>
---------
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix: match system schema names case-insensitively
Database names that arrive over a protocol (the MySQL handshake and
COM_INIT_DB, the Postgres startup parameter, the HTTP `db` parameter, the
gRPC dbname header) never reach the SQL parser, which is what lowercases
unquoted identifiers. Since #8062 stopped lowercasing them wholesale,
connecting to `INFORMATION_SCHEMA` in any spelling but the canonical one
fails with "Unknown database" -- including the `USE <db>` that a MySQL
client turns into COM_INIT_DB.
Fold only system schema names to their canonical spelling, so user schema
names keep the case they were created with. `is_reserved_schema_name` uses
the same match, otherwise a quoted `CREATE DATABASE "INFORMATION_SCHEMA"`
creates a schema shadowed by the system one.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor: hoist system schema names into a const
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(pipeline): coalesce concurrent pipeline cache misses
The pipeline cache reads with a plain `moka::sync::Cache::get` and falls
through to a distributed query on a miss, so when the 10s TTL expires every
in-flight write request on a frontend issues its own scan of the single-region
`greptime_private.pipelines` table. Concurrent scans per expiry scale with
write QPS, and every frontend's burst lands on the same datanode. A user
running high-throughput ingestion through a pipeline saw that datanode
overloaded.
Switch to `moka::future::Cache::try_get_with` so concurrent misses on the same
key share one loader. This requires a single-key lookup, so cache entries are
now keyed by the requested schema rather than the schema the pipeline is stored
under; resolving a request to a stored schema stays in the loader, which is the
authoritative path and already handles the empty-schema and multi-schema cases.
A lookup for a schema not yet cached costs one extra read, now protected from
amplification by the coalescing it enables.
`remove_cache` previously only walked the compiled-pipeline cache, so an entry
populated by `get_pipeline_str` alone (the pipeline read API) survived deletion
until it expired. It now walks all three caches.
Also make the TTL configurable as `pipeline.cache_ttl`, default unchanged at
10s. The TTL is what propagates a pipeline change to other frontends, so
raising it trades staleness for fewer reads.
Refs #9021
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(pipeline): restore cross-schema semantics broken by the new cache key
Keying cache entries by the requested schema dropped two behaviours that the
previous stored-schema key provided for free.
Creating a new version only wrote the creating request's schema, so another
schema on the same frontend kept serving its cached `latest` — an older
version — until the entry expired. Since the whole point of making the TTL
configurable is to let operators raise it, that window is not bounded by
anything useful. Creation now invalidates every schema's `latest` alias for
that name before priming the cache, leaving the version-pinned keys alone.
The failover cache lost its reach across schemas the same way: a global
pipeline (stored under the empty schema) loaded by schema A was cached under
`A`, so schema B using it for the first time while the pipeline table was down
missed and failed ingestion. The failover cache has no loader and so is not
subject to the single-key model of `try_get_with`; it keeps the stored-schema
key and the empty-schema-first resolution.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor(pipeline): drop cache priming on create and fold the sweep helpers
Priming the cache on create saved one read on a low-frequency operation and
cost a concept: entries were written under the creating request's schema while
`PipelineContent.schema` said empty, so the two schemas in play disagreed.
Invalidating the `latest` aliases is required regardless — that is what makes
a new version visible to other schemas — so dropping the priming loses only
the saved read, which coalescing now protects anyway. `insert_and_compile` no
longer needs the caller's schema.
`remove_cache` and the create-time invalidation collapse into one
`invalidate(name, version)`; `None` sweeps only the `latest` aliases, which is
exactly what creation wants. That leaves `invalidate_by_suffixes` and
`cache_keys` with a single caller each, so both are inlined.
Drop the `PipelineOptions` humantime test: `load_config_test` loads both
example TOMLs, which now carry `cache_ttl = "10s"`, and would fail the same
way if the serde attribute were lost. The `toml` dev-dependency goes with it.
The two invalidation tests are now checked to be orthogonal: removing the
version suffix fails only the delete test, and sweeping just the compiled
cache fails both.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(pipeline): keep failover populated across a create
The `latest` sweep on create clears the failover cache along with the loaded
ones, and after dropping the priming there was nothing writing it back. An
outage between the create and the first read-back left neither `latest` nor the
explicit version with anything to fall back on, failing ingestion — worse than
before, since the previous version's failover entry was swept too.
Creation now goes through `PipelineCache::on_pipeline_created`, which pairs the
sweep with a failover write of the new empty-schema definition. The two must
happen together, so they live behind one method rather than at the call site.
Also commit the Cargo.lock entry for the dropped `toml` dev-dependency, and
trim the comments added over the last few commits down to what the code does
not already say.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix: postgres describe for more statements
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: cover more show statements
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: address review comments
- add missing `clippy::too_many_arguments` allow on
`query_from_information_schema_dataframe` (CI clippy failure)
- take `&ShowKind` in the information-schema dataframe helper so `kind`
is no longer cloned at every call site; only the WHERE arm (which needs
an owned expression for `sql_to_expr`) clones internally
- document why re-applying TQL explain formats never overwrites an
existing value (per-query context state)
Signed-off-by: Ning Sun <sunning@greptime.com>
* chore: trim comments to essentials
Signed-off-by: Ning Sun <sunning@greptime.com>
---------
Signed-off-by: Ning Sun <sunning@greptime.com>