* feat(mito): limit series index disk usage
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito): enforce series index quota during reconciliation
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito): remove series index disk budget layer
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito): simplify series index limit to estimated usage
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito): trust index catalogs when loading snapshots
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito): minimize series index disk limit changes
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix(mito2): keep series index cleanup running at capacity
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix(mito2): avoid no-op index clones and stabilize capacity tests
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix: update config API expectation and stabilize index build tests
Signed-off-by: evenyag <realevenyag@gmail.com>
---------
Signed-off-by: evenyag <realevenyag@gmail.com>
* chore(deps): replace cargo-udeps with cargo-shear for unused dependency checks
cargo-udeps requires a nightly toolchain and its pinned version (0.1.61)
no longer detects unused dependencies against current cargo internals —
unused deps have landed on main undetected (e.g. humantime in
common-frontend since #6689). cargo-shear is a standalone static analyzer
that runs on any toolchain.
- Swap 'make check-udeps' / 'make fix-udeps' recipes to 'cargo shear' /
'cargo shear --fix' and retire scripts/fix-udeps.py
- CI: install cargo-shear in the check-udeps job; drop the build cache
and protoc steps (cargo-shear never compiles)
- Remove ~150 unused dependency declarations found by cargo-shear, move
misplaced deps to the correct sections, drop orphaned
[workspace.dependencies] entries (arrow-cast, rustc-hash)
- Add [package.metadata.cargo-shear] ignored entries with explanations
for dependencies that are structurally required despite no textual
reference: sqlparser (required by sqlparser_derive expansions in
datatypes, common-query), common-error (required by common-macro's
stack_trace_debug expansions in session, tests-fuzz), k8s-openapi
(feature-pinning for the transitive kube dependency in tests-fuzz),
tikv-jemalloc-sys (link-only, enables jemalloc profiling features in
common-mem-prof), protobuf (required by build.rs-generated bindings in
log-store)
- Drop the obsolete [package.metadata.cargo-udeps.ignore] sections
Part of #9289
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix(meta): populate physical metric table column ids (#9286)
* fix(meta): populate physical metric table column ids
Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>
* test(meta): verify physical metric column ids
Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>
---------
Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>
* fix(postgres): return empty responses for comment-only SQL (#9295)
fix(postgres): handle parsed empty queries in both protocols
Signed-off-by: houyuwushang <180804215+houyuwushang@users.noreply.github.com>
* ci: create docs follow-up issue on PR merge instead of on label (#9237)
* ci: create docs follow-up issue on PR merge instead of on label
The docbot workflow previously created a docs-repo issue as soon as the
'docs-required' condition was detected (PR opened/edited with the docs
checkbox ticked), even if the PR was never merged.
Now the workflow also triggers on PR 'closed':
- opened/edited: only manage the docs-required/docs-not-required labels
- closed: create the docs issue only when the PR was actually merged and
carries the docs-required label
This also lets maintainers control issue creation by manually adding or
removing the docs-required label before merging.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: address review comments on docs issue creation timing
- Only touch docs labels when the docs checkbox state actually changed
in an edit. Previously, editing any other part of the PR body while
the checkbox stayed checked removed the docs-required label, silently
dropping the docs follow-up now that issue creation happens at merge.
Unchanged checkbox now leaves labels untouched, which also preserves
manual label overrides.
- Do not trust the closed event's stale label snapshot at merge time:
re-read the live PR via the API and create the docs issue if the
docs-required label is present OR the checkbox is ticked in the
current body.
- Make the workflow concurrency group action-aware so a merge run does
not cancel an in-flight label update from an edit run.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: make docs-required label the single source of truth at merge
The label-OR-checkbox merge condition could not distinguish an
intentional opt-out from an unfinished label update: removing
docs-required while the checkbox stayed checked still produced an
issue, and unchecking the box could still produce one if the merge
read the stale label before the edit run removed it.
At merge time, wait for any pending docbot runs on the PR head SHA to
finish their label updates (bounded to 5 minutes), then decide solely
by the live docs-required label. Adds actions: read permission for
listing workflow runs.
Signed-off-by: Ning Sun <sunning@greptime.com>
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
* perf(promql): push label filters into grouped join inputs (#9280)
* perf(promql): propagate matching filters through grouped joins
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* perf(promql): check matcher safety on the receiving operand
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor(promql): spell out the shapes a filter may cross
`preserves_filter` ended in `_ => true`, which was only sound because
`selector_matchers` independently rejects label rewriting, `count_values`,
subqueries and non-rollup calls on the same operand. Loosening the latter
alone would have silently pushed a matcher below a label rewrite. List the
shapes that carry a scan filter instead and default to `false`.
Cite #9207 for the result labels the grouped cases record: the join
projects the right operand's tag set, so `zone` is missing wherever the
right side aggregates it away.
No behavior change.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* test(promql): assert the new pushdowns reach the scan
The grouped-join unit tests feed tag columns by hand and the SQLness case
only checks results, which are identical whether or not the rewrite fires.
Nothing would have failed if scalar arithmetic, ranking or grouped
matching stopped propagating. Assert through the planner that the matcher
reaches both scans, with a global topk one-side as the counter-example.
Also state that the duplicate-one-side cases record a cross product
Prometheus rejects (#9209), so the baseline is not read as intended
semantics.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(ci): build tests-integration lib with meta-srv/mock (#9299)
* fix(ci): build tests-integration lib with meta-srv/mock
tests-integration's lib code (src/cluster.rs) uses meta_srv::mocks, but
the dependency carrying the mock feature sits in [dev-dependencies].
Builds that only touch the lib, such as the apidoc job's cargo doc
--workspace, resolve meta-srv without mock and fail with E0432.
--all-targets builds unify dev-dependency features, which is why check,
clippy and nextest stayed green.
Move the mock-enabled meta-srv entry back to [dependencies]. The other
testing features moved out in #9072 are not needed by the lib and stay
in [dev-dependencies].
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* test(repartition): split per-case repartition tests
test_repartition_metric ran four format/primary-key-encoding cases in a
single test function, and test_repartition_mito ran two format cases.
Each case builds its own 3-datanode cluster and runs a full repartition
plus GC cycle, so on S3 the metric test took 165-178s against the 180s
nextest terminate-after. Merge queue runs failed on it at random.
Split each case into its own test. Cases were already independent, so
they now run in parallel and each stays far inside the timeout, and a
failure points at one encoding instead of four.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* feat(json2): support altering JSON2 settings (#9029)
* feat(sql): support alter syntax for JSON2 columns
Signed-off-by: fys <fengys1996@gmail.com>
* fix(json2): preserve rows on type hint mismatch during compaction
* refactor(json2): simplify alter settings handling
* fix(json2): preserve coerced values during compaction
* chore: remove unnecessary clone
* chor: reduce memory allocations
* fix: cargo clippy
* chore: update greptime-proto to main branch
* refactor(datatypes): unify string handling with other JSON type hints
* fix: cr
---------
Signed-off-by: fys <fengys1996@gmail.com>
* fix: keep compaction pruning, metadata, and index work on compact runtime (#9304)
* fix: run compaction pruner tasks on compact runtime
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix: keep compaction metadata and index work on compact runtime
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
---------
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* feat: add AI matching, classification, and scoring functions (#9300)
* feat: return matching scores from jev
Replace the experimental three-argument Boolean function with jev(text, prompt) returning a Float64 probability in [0, 1]. Move threshold comparisons into SQL and update tests and migration examples.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* feat: add Jev choice and score functions
Share asynchronous execution across Noul, Choice, and Score. Validate JSON criteria before requests and return typed scalar answers. Add SQL and HTTP mock coverage with usage examples.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor: use generic AI SQL function names
Expose ai_match, ai_choose, and ai_score and move their implementation, tests, and usage guide under generic AI names. Document the current unreleased interface without migration history.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix: share constant AI criteria within each batch
Borrow scalar string arguments and lazily parse constant criteria once per batch. Share the parsed allocation across requests while preserving NULL propagation and batch validation before HTTP calls.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* feat: preserve AI score uncertainty in JSONB results
Return score, confidence, and probabilities in criteria-level order from one evaluation. Validate the distribution and preserve provider precision. Add JSON extraction, uncertainty, and single-request regressions, and document confidence-aware ranking.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* docs: explain reuse of volatile AI evaluations
Document repeated SELECT and WHERE evaluation costs as N + M requests, and show subquery aliases for reusing scalar or structured AI results without additional model calls.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
---------
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* feat: share logical table batching with OTLP metrics (#9288)
* feat: share logical table batching with OTLP metrics
Signed-off-by: WenyXu <wenymedia@gmail.com>
* fix: unify pending rows batch acknowledgement policy
Signed-off-by: WenyXu <wenymedia@gmail.com>
* fix: align logical batcher example configuration expectations
Signed-off-by: WenyXu <wenymedia@gmail.com>
* fix: align batcher worker channel defaults to 65536
Signed-off-by: WenyXu <wenymedia@gmail.com>
---------
Signed-off-by: WenyXu <wenymedia@gmail.com>
* perf(mito2): lazily decode dense primary key columns (#9226)
* perf(mito2): lazily decode dense primary key columns
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* perf(mito2): bypass lazy decoding for full primary keys
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito-codec): preserve prefix decoding errors
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito-codec): align encoded length helper naming
Rename encoded_length to encoded_len and update all callers to match the other length helpers in the module.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito2): clarify conditional dense key decoding
Rename decode_dense_pk to ensure_dense_pk_decoded so callers can see that existing decoded values are preserved and only missing caches are populated.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito-codec): share string framing in row converter
Move encoded_string_len to the parent module so Dense and Sparse use the same framing helper without depending on each other. Preserve its implementation and visibility.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
---------
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* chore: sync lock
* fix: shear and check issues
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
Signed-off-by: dhruvxvaishnav <dhruvvaishnav687@gmail.com>
Signed-off-by: houyuwushang <180804215+houyuwushang@users.noreply.github.com>
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
Signed-off-by: fys <fengys1996@gmail.com>
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
Signed-off-by: WenyXu <wenymedia@gmail.com>
Co-authored-by: Dhruv Vaishnav <dhruvvaishnav687@gmail.com>
Co-authored-by: houyuwushang <180804215+houyuwushang@users.noreply.github.com>
Co-authored-by: dennis zhuang <killme2008@gmail.com>
Co-authored-by: fys <40801205+fengys1996@users.noreply.github.com>
Co-authored-by: Lei, HUANG <6406592+v0y4g3r@users.noreply.github.com>
Co-authored-by: Weny Xu <wenymedia@gmail.com>
* feat: serialize struct to json in postgres
* fix: support view scalars and preserve null structs in scalar-to-value conversion
Address PR review:
- Utf8View/BinaryView ScalarValues now convert like their non-view forms
instead of failing row extraction for struct columns
- a null struct scalar converts to Value::Null so a null struct inside a
list stays null in the serialized JSON
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: return errors instead of panics for unsupported arrow field types
Struct-typed query results with arrow field types greptimedb cannot
represent (e.g. Decimal256) used to panic during schema conversion and
row extraction, dropping the client connection. They now surface as
query errors:
- ConcreteDataType::try_from builds struct types fallibly via the new
StructType::try_from_arrow_fields
- Value::try_from(ScalarValue::Struct) uses the same fallible path
- new try_value_from_array converts an arrow element to Value with
error propagation, used by the postgres struct encoding
Signed-off-by: Ning Sun <sunning@greptime.com>
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
* test: cut integration test time and make the storage matrix meaningful
tests-integration is ~85% of workspace test CPU, and 81% of that is the
S3/S3WithCache variants of the HTTP and gRPC suites. Those suites do not
touch the object store: of the 70 matrix HTTP tests only one flushed and
read back an SST, so the matrix was paying real AWS round trips to
re-prove protocol parsing.
- Point the PR CI object-store matrix at the MinIO already started by
tests-integration/fixtures. Three GT_S3_* consumers did not read
GT_S3_ENDPOINT_URL and would have hit real AWS with MinIO credentials;
they now do.
- Add a nightly Linux job against real AWS S3, and pass GT_S3_* into the
release integration-test container. The release previously ran every
remote-backend case as a skip and only exercised the file backend.
- Give each S3WithCache test its own read cache directory. They shared
/tmp/greptimedb_cache, which the datanode wipes on startup, so a
starting test deleted the read cache of a running one.
- Add flush -> read-back assertions to the tests whose columns have a
non-trivial SST representation: JSON/JSON2 columns, native histograms,
metric-engine logical tables, and tables carrying fulltext or skipping
indexes whose puffin files only exist after a flush.
- Move eight tests that create no table out of the storage matrix.
- Make the event recorder flush interval a constructor parameter and
shorten it in the event tests, which otherwise wait a 5s window per DDL
they assert on. It is skipped by serde and never reaches config files.
- Drop duplicates: test_grpc_zstd_compression was a verbatim copy of
test_grpc_message_size_ok and is now rewritten to assert the negotiated
grpc-encoding; test_execute_copy_to_{s3,oss,gcs,azblob} were strict
prefixes of their copy_from siblings; two standalone/distributed event
test pairs shared one assertion body.
- Fix and un-ignore stddev_by_label. stddev_pop merges partial aggregates
in a parallelism-dependent order, so its last digits are unstable; the
test now compares values with a tolerance.
- Rebase the jaeger v1 fixture on the current instant. It carries
ttl=7d with 2025 timestamps, so its rows were only readable as long as
they stayed in the memtable.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* test: address review — wire nightly real-S3 job into check-status, keep the short event interval
The nightly `check-status` job did not depend on the new real-S3 job, so a
failure there would not have reached the status or Slack notification.
In database_ddl_event the short interval was set by a first
`with_event_recorder_options` call and then overwritten by the pre-existing
one, which carries `..Default::default()`. Merged into a single call.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(ci): build tests-integration lib with meta-srv/mock
tests-integration's lib code (src/cluster.rs) uses meta_srv::mocks, but
the dependency carrying the mock feature sits in [dev-dependencies].
Builds that only touch the lib, such as the apidoc job's cargo doc
--workspace, resolve meta-srv without mock and fail with E0432.
--all-targets builds unify dev-dependency features, which is why check,
clippy and nextest stayed green.
Move the mock-enabled meta-srv entry back to [dependencies]. The other
testing features moved out in #9072 are not needed by the lib and stay
in [dev-dependencies].
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* test(repartition): split per-case repartition tests
test_repartition_metric ran four format/primary-key-encoding cases in a
single test function, and test_repartition_mito ran two format cases.
Each case builds its own 3-datanode cluster and runs a full repartition
plus GC cycle, so on S3 the metric test took 165-178s against the 180s
nextest terminate-after. Merge queue runs failed on it at random.
Split each case into its own test. Cases were already independent, so
they now run in parallel and each stays far inside the timeout, and a
failure points at one encoding instead of four.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(mito2): compare primary key ranges across schema versions
Bind FileHandle ranges to the pinned region schema and append cached constant defaults to historical Dense keys. Preserve raw SST statistics, reject inexact bounds, and avoid invalidating views for unrelated metadata changes.
Cover schema evolution, default changes, and tombstone retention through real compaction and reopen regressions.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito2): report invalid primary key ranges
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito2): scope primary key ranges to comparisons
Keep raw PK bounds only in FileHandleInner and move schema-aware mapping and caching into task-local comparison contexts.
Use explicit contexts for compaction overlap checks, window aggregation, and series scans. Preserve pinned-schema isolation, late statistics, and shared file lifecycle state without rebinding every handle.
Cover cache isolation across region owners and adapt range fixtures to real Dense encodings. All 1530 mito2 tests and Clippy for all targets with the testing feature pass.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito2): cache aligned primary key ranges per file
Replace task-local range maps with a single-slot cache in FileHandleInner, keyed by the target schema version. Preserve raw bounds for realignment across snapshots and default changes.
Share schema mappers across comparison paths and use copy-on-write SST lists for metadata updates. Simplify range mapping to accept encoded bounds and assert the same-table contract at the file accessor.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* docs(mito2): clarify primary key mapper schema snapshot
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* test(mito2): align primary key range fixtures with table contract
Remove obsolete cross-table fallback expectations after region validation became a caller contract. Give compaction fixtures matching table identities, including the active-window L1 scenario.
Clarify the mapper precondition and format the simplified alignment call. All 1530 mito2 tests and Clippy for all targets with the testing feature pass.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* test(mito2): enable filesystem GC in release unit tests
Let unit tests use the filesystem-backed object-store GC path regardless of optimization profile. Keep the production release GC selection unchanged.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* perf(mito-codec): skip release value decoding in PK prefix counts
Validate field values only in debug builds and unit tests while keeping boundary, truncation, and trailing-byte checks in every build.
Cover the linked library in debug and release integration tests, and verify that release unit tests still perform value validation.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* test(mito-codec): remove redundant prefix integration tests
Retain the codec unit tests and cross-schema compaction regressions while dropping the standalone build-profile test file and its release-only invalid-value expectation.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* ci: wait for MySQL to accept authenticated TCP queries
Add a healthcheck using the configured test account and database. Docker Compose --wait previously only observed container startup because the fixture image had no healthcheck, allowing metasrv to connect before MySQL initialization completed.
Verify readiness with SELECT 1 over TCP rather than the initialization socket.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
---------
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* feat: add config flag to control range index reads
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix: disable range index builds when configured and default to off
Signed-off-by: evenyag <realevenyag@gmail.com>
---------
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix(prometheus): honor label matchers in __name__ values query
`/api/v1/label/__name__/values?match[]={pod="abc"}` dropped every matcher
other than `__name__` and returned all metrics in the schema. No error,
just the wrong list. Grafana's metrics browser sends this request, so
picking a label value there did nothing.
Selectors that only constrain `__name__` keep answering from table
metadata. A selector constraining an ordinary label now goes to the data:
scan each metric engine physical table for distinct `__table_id` in the
time range, map the ids back to metric names, then apply the selector's
own `__name__` matchers.
Only metric engine tables are covered; other engines share no column space
to scan.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor(prometheus): batch-resolve metric names by table id
Building a full table-id-to-name map meant walking every table in the
schema and holding all of them in memory, just to name the handful the
scan returned. Use `tables_by_ids` instead — one batch KV read over the
ids the scan actually produced.
The catalog walk stays, but only to find the physical tables to scan.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(promql): read an absent label as the empty string
A matcher on a label the series does not carry only worked when the table
had no column for it at all. Where the column exists but is NULL on that
row -- the norm for logical metrics sharing a metric engine physical
table, which holds the union of their label columns -- three-valued logic
dropped the row, so `host!="host1"` and `host=""` missed every metric
without a host label.
Coalesce nullable string label columns to "" for matchers that accept the
empty string, rather than only for the OTLP temporality marker. Equality
matchers are untouched; they cannot match NULL either way.
This is the Prometheus compatibility fix#8970 deliberately kept out of
its own scope. The cost is visible in the regex sqlness plan: the
predicate becomes a CASE, so the scan loses its LastRow selector and
grows a FilterExec. Only negative and empty-accepting matchers pay it.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(promql): don't panic on pre-epoch label value bounds
`rewrite_label_values_query` unwrapped `duration_since(UNIX_EPOCH)`, which
returns an error for an instant before the epoch. `start=1969-12-31T23:59:59Z`
parses as valid RFC3339, so the request panicked instead of answering.
Recover the sign from the error branch, and report a value beyond i64
milliseconds as an error rather than wrapping the cast.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor(prometheus): drop applicable_matchers, share the distinct scan
With the planner reading an absent label as empty, the frontend no longer
needs to pre-filter matchers per physical table. Removing that exposed a
second problem: a physical table that never took a column from a logical
table exposes no `__table_id`, and projecting it failed the whole request.
Skip those tables; the only thing that can miss is a metric with no labels.
Also pulls out the plan-build-execute-collect sequence the two label value
scans had in common.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* feat(mito2): reconcile series indexes in background
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix(mito2): clean up series indexes published during region drop
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix(config): align series index examples with upstream enable flag
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix(mito2): run series index tasks on compaction runtime
Signed-off-by: evenyag <realevenyag@gmail.com>
---------
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix: use join_dir for series index config path
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix: configure series indexes with an enable flag
Signed-off-by: evenyag <realevenyag@gmail.com>
* docs: omit experimental series index from example configs
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix: preserve legacy cache cleanup path behavior
Signed-off-by: evenyag <realevenyag@gmail.com>
* test: remove trivial path joining tests
Signed-off-by: evenyag <realevenyag@gmail.com>
* test: isolate worker group WAL directories on Windows
Signed-off-by: evenyag <realevenyag@gmail.com>
---------
Signed-off-by: evenyag <realevenyag@gmail.com>
Building a matrix response allocated one `String` per sample while the
record batches were scanned, then dropped it after the JSON body was
written. Keep the `f64` in `PromSampleValue::Number` instead and format
it with ryu while serializing, so no per-sample string is allocated.
`PromSampleValue::Text` keeps values parsed from a JSON body, so
deserializing and re-serializing a response is unchanged. Vector and
scalar results still expose `String`, since they hold a single sample.
Signed-off-by: Dennis Zhuang <xzhuang@greptime.com>
* refactor(mito2): add series index catalog and lifecycle components
Signed-off-by: evenyag <realevenyag@gmail.com>
* feat(mito2): restore series index catalogs on region open
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito2): simplify series index foundation and maintenance
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito2): separate series index purge task and simplify tests
Signed-off-by: evenyag <realevenyag@gmail.com>
* docs: defer experimental series index configuration examples
Signed-off-by: evenyag <realevenyag@gmail.com>
* feat(mito): make series index maintenance interval configurable
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix(mito2): correct series index cleanup on close and drop
Signed-off-by: evenyag <realevenyag@gmail.com>
* test(mito2): revert drop test changes
Signed-off-by: evenyag <realevenyag@gmail.com>
* test: update config API expectation for series index settings
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor: use tokio unbounded channel for series index purger
Signed-off-by: evenyag <realevenyag@gmail.com>
* chore(mito2): simplify review test scope and clarify index config
Signed-off-by: evenyag <realevenyag@gmail.com>
---------
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix: match system schema names case-insensitively
Database names that arrive over a protocol (the MySQL handshake and
COM_INIT_DB, the Postgres startup parameter, the HTTP `db` parameter, the
gRPC dbname header) never reach the SQL parser, which is what lowercases
unquoted identifiers. Since #8062 stopped lowercasing them wholesale,
connecting to `INFORMATION_SCHEMA` in any spelling but the canonical one
fails with "Unknown database" -- including the `USE <db>` that a MySQL
client turns into COM_INIT_DB.
Fold only system schema names to their canonical spelling, so user schema
names keep the case they were created with. `is_reserved_schema_name` uses
the same match, otherwise a quoted `CREATE DATABASE "INFORMATION_SCHEMA"`
creates a schema shadowed by the system one.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor: hoist system schema names into a const
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(pipeline): coalesce concurrent pipeline cache misses
The pipeline cache reads with a plain `moka::sync::Cache::get` and falls
through to a distributed query on a miss, so when the 10s TTL expires every
in-flight write request on a frontend issues its own scan of the single-region
`greptime_private.pipelines` table. Concurrent scans per expiry scale with
write QPS, and every frontend's burst lands on the same datanode. A user
running high-throughput ingestion through a pipeline saw that datanode
overloaded.
Switch to `moka::future::Cache::try_get_with` so concurrent misses on the same
key share one loader. This requires a single-key lookup, so cache entries are
now keyed by the requested schema rather than the schema the pipeline is stored
under; resolving a request to a stored schema stays in the loader, which is the
authoritative path and already handles the empty-schema and multi-schema cases.
A lookup for a schema not yet cached costs one extra read, now protected from
amplification by the coalescing it enables.
`remove_cache` previously only walked the compiled-pipeline cache, so an entry
populated by `get_pipeline_str` alone (the pipeline read API) survived deletion
until it expired. It now walks all three caches.
Also make the TTL configurable as `pipeline.cache_ttl`, default unchanged at
10s. The TTL is what propagates a pipeline change to other frontends, so
raising it trades staleness for fewer reads.
Refs #9021
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(pipeline): restore cross-schema semantics broken by the new cache key
Keying cache entries by the requested schema dropped two behaviours that the
previous stored-schema key provided for free.
Creating a new version only wrote the creating request's schema, so another
schema on the same frontend kept serving its cached `latest` — an older
version — until the entry expired. Since the whole point of making the TTL
configurable is to let operators raise it, that window is not bounded by
anything useful. Creation now invalidates every schema's `latest` alias for
that name before priming the cache, leaving the version-pinned keys alone.
The failover cache lost its reach across schemas the same way: a global
pipeline (stored under the empty schema) loaded by schema A was cached under
`A`, so schema B using it for the first time while the pipeline table was down
missed and failed ingestion. The failover cache has no loader and so is not
subject to the single-key model of `try_get_with`; it keeps the stored-schema
key and the empty-schema-first resolution.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor(pipeline): drop cache priming on create and fold the sweep helpers
Priming the cache on create saved one read on a low-frequency operation and
cost a concept: entries were written under the creating request's schema while
`PipelineContent.schema` said empty, so the two schemas in play disagreed.
Invalidating the `latest` aliases is required regardless — that is what makes
a new version visible to other schemas — so dropping the priming loses only
the saved read, which coalescing now protects anyway. `insert_and_compile` no
longer needs the caller's schema.
`remove_cache` and the create-time invalidation collapse into one
`invalidate(name, version)`; `None` sweeps only the `latest` aliases, which is
exactly what creation wants. That leaves `invalidate_by_suffixes` and
`cache_keys` with a single caller each, so both are inlined.
Drop the `PipelineOptions` humantime test: `load_config_test` loads both
example TOMLs, which now carry `cache_ttl = "10s"`, and would fail the same
way if the serde attribute were lost. The `toml` dev-dependency goes with it.
The two invalidation tests are now checked to be orthogonal: removing the
version suffix fails only the delete test, and sweeping just the compiled
cache fails both.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(pipeline): keep failover populated across a create
The `latest` sweep on create clears the failover cache along with the loaded
ones, and after dropping the priming there was nothing writing it back. An
outage between the create and the first read-back left neither `latest` nor the
explicit version with anything to fall back on, failing ingestion — worse than
before, since the previous version's failover entry was swept too.
Creation now goes through `PipelineCache::on_pipeline_created`, which pairs the
sweep with a failover write of the new empty-schema definition. The two must
happen together, so they live behind one method rather than at the call site.
Also commit the Cargo.lock entry for the dropped `toml` dev-dependency, and
trim the comments added over the last few commits down to what the code does
not already say.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>