* feat: record time unit per series index file
The series index stores __series_min_ts/__series_max_ts as raw i64 in
the time index unit at write time, but the searcher built its range
predicates from the region's current unit, so files written before a
time index unit widening would be compared in the wrong unit.
Record the unit in the min/max ts fields' Arrow metadata when writing
and build the per-file time predicates from it when searching, so each
file is interpreted in the unit it was written with. The writer now
also requires a timestamp time index.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: address review comments on series index time units
- split schema validation (validate_index_schema) from unit extraction
(index_time_unit), distinguishing missing vs unsupported unit metadata
in the errors instead of one misleading 'missing a valid metadata'
- encode the recorded unit with an explicit exhaustive match rather than
Debug formatting, so the on-disk encoding is reviewed next to its parser
- extract time_index_unit to drop the unwrap in series_index_schema and
the duplicated timestamp-time-index ensure in validate_metadata
- test one searcher reading files with different recorded units, and the
rejection of missing, unknown and mismatched units
Signed-off-by: Ning Sun <sunning@greptime.com>
* refactor: reuse TimeUnit's Display form for the recorded unit string
common_time::timestamp::TimeUnit already implements Display with the
exact strings the series index records ("Second"/"Millisecond"/
"Microsecond"/"Nanosecond"), so drop the local time_unit_as_str
mapping and use it; the parse side stays local since no FromStr
counterpart exists anywhere yet.
Signed-off-by: Ning Sun <sunning@greptime.com>
* feat: parse TimeUnit from its Display form in common-time
Add FromStr for common_time::timestamp::TimeUnit, accepting exactly the
Display form ("Second"/"Millisecond"/"Microsecond"/"Nanosecond") and
failing with a new UnsupportedTimeUnit error (InvalidArguments). The
series index now records and parses the unit with the common codec,
dropping its local parse_time_unit.
Signed-off-by: Ning Sun <sunning@greptime.com>
* feat: parse TimeUnit case-insensitively
Lowercase the input before matching so "millisecond" and "MILLISECOND"
parse like "Millisecond"; the error still reports the original string.
Signed-off-by: Ning Sun <sunning@greptime.com>
* feat: store series index min/max timestamps as native Timestamp columns
Replace the Int64 min/max columns plus 'time_unit' field metadata with
native Timestamp(unit) columns, so the unit rides on the datatype and
each file is interpreted in the unit it was written with naturally.
- series_index_schema types the columns from the time index unit; the
writer reinterprets the raw i64 series bounds in that type (arrow's
Int64->Timestamp cast reinterprets, it does not rescale)
- the searcher reads the unit from each file's column datatype, builds
Timestamp-typed predicates via datatypes' timestamp_to_scalar_value,
and reinterprets parquet INT64 statistics in the column type so
row-group pruning compares like-typed values
- index files whose min/max columns are not Timestamp (written before
this change) are rejected
- the pruning test now asserts time-range predicates prune row groups,
not just tag predicates
Signed-off-by: Ning Sun <sunning@greptime.com>
* refactor: drop the TimeUnit string codec from common-time
With the unit carried by the Timestamp datatype, the FromStr impl and
UnsupportedTimeUnit error added for the field-metadata approach have no
consumer; remove them.
Signed-off-by: Ning Sun <sunning@greptime.com>
* refactor: fold index file validation into a single pass
The series index format is unreleased and unwired, so no compatibility
classes are needed: validate_index_schema checks all columns and
returns the min/max columns' unit directly, replacing the separate
index_time_unit extraction.
Signed-off-by: Ning Sun <sunning@greptime.com>
* refactor: carry timestamps through SeriesIndexRow
SeriesIndexRow and the aggregation path now hold common_time::Timestamp
instead of raw i64s: timestamp_values interprets the input column in the
writer's unit (rejecting a timestamp array whose unit differs, instead
of silently reinterpreting it), and rows_to_batch builds the native
Timestamp columns directly from the rows' units without an Int64 round
trip through arrow cast.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: require a timestamp time index column in series index input
The writer already refuses non-timestamp time indexes on the metadata
side, and its input batches always carry the region's ts column as a
timestamp array, so accepting plain Int64 columns only left a silent
unit-interpretation hole; reject them instead.
Signed-off-by: Ning Sun <sunning@greptime.com>
* refactor: rescale series index input timestamps into the file unit
The input array's unit is self-describing, so converting with
Timestamp::convert_to cannot mislabel values; a mismatch no longer
needs to be an error. Only a value that overflows the file's unit
fails the write. This also makes the writer ready to aggregate
old-unit batches after a time index widening.
Signed-off-by: Ning Sun <sunning@greptime.com>
* refactor: drop redundant unit checks in series index writer
The alter path flushes memtables before widening the region's time
index unit, so a writer never receives batches in the region's
previous unit. Reject a unit mismatch at the input boundary instead
of rescaling per value, and build the index batch in the writer's
recorded unit instead of re-deriving it from the schema and
re-checking every row.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix: address review comments
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci(backport): label backport PRs with their version name
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci(backport): restore multi-line PR body string
Signed-off-by: Ning Sun <sunning@greptime.com>
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix(prometheus): align batch flush deadline with creation
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* test(servers): sync pending worker submission with explicit ack
Replace the mpsc capacity() polling in the pending rows batcher deadline
test with a test-only WorkerCommand::Ack round trip. The FIFO channel
guarantees the worker has dequeued and processed the submission (and
anchored the flush deadline) before the test advances virtual time,
removing reliance on an implementation detail that can be flaky under
scheduling variance.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* test(servers): await pending worker flush results with bounded timeout
Replace the try_recv() yield-polling loop in the pending rows batcher
deadline test with a direct await bounded by tokio::time::timeout. Under
paused time the timeout auto-advances the clock and fires
deterministically, so the test fails reliably instead of intermittently
missing the result on slower CI or under contention.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
---------
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito2): add series index catalog and lifecycle components
Signed-off-by: evenyag <realevenyag@gmail.com>
* feat(mito2): restore series index catalogs on region open
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito2): simplify series index foundation and maintenance
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito2): separate series index purge task and simplify tests
Signed-off-by: evenyag <realevenyag@gmail.com>
* docs: defer experimental series index configuration examples
Signed-off-by: evenyag <realevenyag@gmail.com>
* feat(mito): make series index maintenance interval configurable
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix(mito2): correct series index cleanup on close and drop
Signed-off-by: evenyag <realevenyag@gmail.com>
* test(mito2): revert drop test changes
Signed-off-by: evenyag <realevenyag@gmail.com>
* test: update config API expectation for series index settings
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor: use tokio unbounded channel for series index purger
Signed-off-by: evenyag <realevenyag@gmail.com>
* chore(mito2): simplify review test scope and clarify index config
Signed-off-by: evenyag <realevenyag@gmail.com>
---------
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix(mito2): reject non-flat batches in flat merge instead of panicking
- SortColumns::new becomes fallible try_new: batches missing the
flat-format internal columns (time index, __primary_key, __sequence)
at the fixed trailing positions now yield InvalidRecordBatch instead
of a downcast panic, completing the generic-schema gate that only
covered BatchBuilder output assembly. Document the flat-format input
contract on FlatMergeIterator/FlatMergeReader.
- Clarify why BatchBuilder's schema gate uses >= 3 columns when a real
flat-format schema always has at least 4.
- Add schema-structure tests: empty primary keys (tables without tags),
dictionary-encoded string tag columns with per-source dictionaries,
and graceful rejection of batches without internal columns.
Signed-off-by: Lei, HUANG <mrsatangel@gmail.com>
* test(mito2): scan tables with various schemas through flat merge
Add an engine-level test that writes, flushes and scans regions without
tags (empty primary key) and with multiple string tags (dictionary-encoded
in the flat input schema), so the flat merge reader merges an SST with
the memtable on real schemas instead of hand-built batches.
Signed-off-by: Lei, HUANG <mrsatangel@gmail.com>
* refactor(mito2): use winner_tree dependency
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* Update comments for FlatMergeIterator struct
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
---------
Signed-off-by: Lei, HUANG <mrsatangel@gmail.com>
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
* fix(cli): sanitize store_addrs in kvbackend build log
Close#7525. The CLI's kvbackend construction log was printing raw
store_addrs which could contain sensitive connection strings (e.g.
PostgreSQL DSNs with passwords).
Changes:
- Add sanitize_store_addrs() helper that reuses
common_meta::kv_backend::util::sanitize_connection_string(),
consistent with MetasrvOptions and StartCommand patterns.
- Replace raw store_addrs in the info! log with sanitized version.
- Add unit tests covering MySQL URLs, PostgreSQL DSNs, etcd addresses,
and empty store_addrs cases.
Signed-off-by: qiang_liu
Signed-off-by: qiang_liu <qiang_liu@trendmicro.com>
Signed-off-by: LiuQhahah <liuqiang9596@gmail.com>
* fix(cli): drop redundant sanitize tests per review
sanitize_connection_string in common_meta already covers MySQL URLs,
PostgreSQL DSNs and credential-free etcd addresses with its own tests.
The added tests only exercised a trivial map+collect wrapper, so remove
them per review nit.
Signed-off-by: LiuQhahah <liuqiang9596@gmail.com>
---------
Signed-off-by: qiang_liu
Signed-off-by: qiang_liu <qiang_liu@trendmicro.com>
Signed-off-by: LiuQhahah <liuqiang9596@gmail.com>
Co-authored-by: dennis zhuang <killme2008@gmail.com>
* feat(mito2): adapt bulk memtable encode bytes threshold to write buffer size
The default encode_bytes_threshold is now max(64MB,
min(global_write_buffer_size / 32, 512MB)) instead of a fixed 64MB, so
it scales with the memtable budget. GREPTIME_BULK_ENCODE_BYTES_THRESHOLD
and the per-region option still override the default.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* test(mito2): clarify binary units in bulk threshold test
The threshold test uses powers of 1024, so label its values as MiB and GiB instead of decimal MB and GB.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito2): resolve bulk encode threshold in memtable provider
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
---------
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* perf: add scanbench partition diagnostics
Port benchmark source and documentation from c5dd5e6356a8903f1dcf493eb46dc318612dc953. Exclude the Mito metrics changes.
Signed-off-by: evenyag <realevenyag@gmail.com>
* feat: support scanbench query suites
Signed-off-by: evenyag <realevenyag@gmail.com>
(cherry picked from commit def9c22069b0d29c611695a68ad1b31110c845cb)
Signed-off-by: evenyag <realevenyag@gmail.com>
* docs: align scanbench port with existing engine metrics
Remove documentation and fixture references to unported Mito metrics. Enable dev-tools in the build example and remove the obsolete force-flat-format option.
Signed-off-by: evenyag <realevenyag@gmail.com>
---------
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix: re-scan stream-backed tables in recursive CTEs
A recursive CTE re-executes its recursive term on every iteration, but
DfTableProviderAdapter hands StreamScanAdapter a single-use stream built at
planning time. The second iteration failed with "Stream already exhausted"
for every table served through DataSource::get_stream — information_schema,
pg_catalog, the computed entity-graph tables and numbers.
Keep that stream for the first execution and open a new one over the same
scan request for later executions.
Closes#9037
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor: drop redundant binding in stream factory
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* Implement `/query-regression` command handling and admission workflow
- Add `query-regression-slash.py` script for processing `/query-regression` commands in PR comments, validating case arguments, and checking permissions.
- Update `checks.yml` to include tests for the new slash command functionality.
- Modify `query-regression-comment.yml` to trigger on the new `Query Regression Command` workflow.
- Create `query-regression-slash.yml` to handle the dispatched command, validate allowlist and permissions, and initiate the regression workflow.
- Enhance `query-regression.yml` to support additional inputs for PR admission and SHA verification.
- Introduce `slash-command-dispatch.yml` to parse and dispatch commands from PR comments.
- Document the new command admission process in `AGENTS.md` and `README.md`.
- Add unit tests in `test_query_regression_slash.py` to cover command parsing and admission logic.
* refactor: enhance query-regression command handling with comment validation and identity checks
* feat: implement admission identity handling for query regression workflows
* refactor: update PR admission logic in query regression workflow
* refactor: update token usage in slash command dispatch and README for clarity
* test: add cases for handling re-run failed jobs and stale runner artifacts
* refactor: improve repository metadata handling in query regression scripts
* chore: enable overwrite for artifact uploads to handle re-run failed jobs
* chore: enable overwrite for query regression admission uploads
* feat: enhance query-regression admission with HMAC signing and verification
- Introduced HMAC signing for admission markers in query-regression workflows to ensure integrity and authenticity.
- Updated `query-regression-comment.test.cjs` to include tests for signing and verifying admission markers.
- Modified `query-regression-slash.py` to handle admission marker signing and verification, including checks for dispatch sender and head SHA consistency.
- Enhanced workflows to securely manage admission markers and HMAC secrets, ensuring they are not exposed to untrusted contexts.
- Improved documentation to clarify the admission process and the role of HMAC in securing the workflow.
* test: add case to find newly posted marker among newer comments
* test: add case to verify multiline output handling in write_outputs function
* perf(mito2): optimize flat merge heap and primary-key interleave
Replace the per-row BinaryHeap pop/push cycle in FlatMerge with an
in-place root mutation plus a single sift-down repair on a custom
RootHeap, keeping the cold heap and direct-batch fast path unchanged.
Fallible or awaiting batch transitions move the hot node out of the
heap first, preserving error and cancellation semantics.
Exploit the globally sorted merge output to build the internal
Dictionary<UInt32, Binary> primary-key column with a one-pass ordered
gather: append a Binary value only when the PK changes and reuse the
current key for adjacent equal PKs, bypassing Arrow dictionary masks,
hash interning and key remapping. Non-PK columns still use Arrow
interleave.
Also cache the current primary-key byte range in RowCursor to avoid
repeated dictionary range decoding during comparisons, and add a
setup-free Criterion benchmark with exact output-row assertions.
32-way/1 row-per-series/40-tag improves 955.79ms -> 562.36ms (-41.2%);
0-tag -39.5%, 64 rows/series -79.8%, 8-way -56.1%, single-iterator
control +0.3%.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* test(mito2): add rows-per-series sweep to flat merge bench
Add 32-way/40-tag shapes for 1, 10, 100, 1000 and 10000 rows per
series, and allow FLAT_MERGE_BENCH_SHAPE to match shape name prefixes
so the whole sweep can run in one invocation.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* test(mito2): add oracle-based correctness tests for RootHeap
Drive RootHeap and a std BinaryHeap oracle with the same seeded op
sequence (push / pop / mutate-root + repair) and assert peek, len,
best_child and the full drain order after every operation. A second
run with a tiny value range makes duplicates dominate, covering the
equal-key branches of sift_up/sift_down and best_child.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* perf(mito2): replace hot heap with a tournament tree in flat merge
Replace the hot RootHeap with a fixed-capacity tournament (winner) tree
over per-node slots: every internal node caches the champion of its
subtree, so advancing the winner only replays the ~log2(k) nodes on its
leaf-to-root path with one compare per level, instead of the heap's
two-compares-per-level sift that also re-compares the same node pairs
on every row.
Two fast paths keep dense shapes at O(1) per row:
- champion retention: after mutating the winner in place, skip the
replay entirely when it still beats the runner-up (its path caches
are unchanged by construction);
- a second-best slot cache, invalidated on any structural change, so
the retention check costs a single compare without walking the tree.
The cold heap, hot/cold overlap window, direct-batch fast path and the
remove-before-fallible-fetch batch transition semantics are unchanged.
Vs the RootHeap version: 1rps/32way/40tag -19.7%, 0tag -34.4%,
8way -15.9%, 64rps -30.5%, sweep 10/100/1000/10000rps -29~32%;
vs the original BinaryHeap baseline the main shape is -52.8%.
The single-iterator control is +8% (+50ns one-time construction
allocation, no merge work).
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito2): support generic schemas in flat merge
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito2): satisfy clippy in flat merge benchmark
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* perf(mito2): cache flat merge primary key index
Compute the internal primary-key column index once when constructing BatchBuilder and reuse it for every output batch. Preserve the column-name gate for generic schemas.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* test(mito2): benchmark high-fan-in flat merges
Add sparse 64, 128, 256, and 512-way merge shapes while keeping the total input fixed at 3.2 million rows. Compared with the merge-base heap implementation, median time improves by 56.0%, 60.8%, 55.6%, and 56.1%, respectively.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
---------
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix: match system schema names case-insensitively
Database names that arrive over a protocol (the MySQL handshake and
COM_INIT_DB, the Postgres startup parameter, the HTTP `db` parameter, the
gRPC dbname header) never reach the SQL parser, which is what lowercases
unquoted identifiers. Since #8062 stopped lowercasing them wholesale,
connecting to `INFORMATION_SCHEMA` in any spelling but the canonical one
fails with "Unknown database" -- including the `USE <db>` that a MySQL
client turns into COM_INIT_DB.
Fold only system schema names to their canonical spelling, so user schema
names keep the case they were created with. `is_reserved_schema_name` uses
the same match, otherwise a quoted `CREATE DATABASE "INFORMATION_SCHEMA"`
creates a schema shadowed by the system one.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor: hoist system schema names into a const
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* ci: add backport workflow to create backport PRs from backport labels
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci: document backport labels in PR template and AGENTS.md
Signed-off-by: Ning Sun <sunning@greptime.com>
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix(pipeline): coalesce concurrent pipeline cache misses
The pipeline cache reads with a plain `moka::sync::Cache::get` and falls
through to a distributed query on a miss, so when the 10s TTL expires every
in-flight write request on a frontend issues its own scan of the single-region
`greptime_private.pipelines` table. Concurrent scans per expiry scale with
write QPS, and every frontend's burst lands on the same datanode. A user
running high-throughput ingestion through a pipeline saw that datanode
overloaded.
Switch to `moka::future::Cache::try_get_with` so concurrent misses on the same
key share one loader. This requires a single-key lookup, so cache entries are
now keyed by the requested schema rather than the schema the pipeline is stored
under; resolving a request to a stored schema stays in the loader, which is the
authoritative path and already handles the empty-schema and multi-schema cases.
A lookup for a schema not yet cached costs one extra read, now protected from
amplification by the coalescing it enables.
`remove_cache` previously only walked the compiled-pipeline cache, so an entry
populated by `get_pipeline_str` alone (the pipeline read API) survived deletion
until it expired. It now walks all three caches.
Also make the TTL configurable as `pipeline.cache_ttl`, default unchanged at
10s. The TTL is what propagates a pipeline change to other frontends, so
raising it trades staleness for fewer reads.
Refs #9021
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(pipeline): restore cross-schema semantics broken by the new cache key
Keying cache entries by the requested schema dropped two behaviours that the
previous stored-schema key provided for free.
Creating a new version only wrote the creating request's schema, so another
schema on the same frontend kept serving its cached `latest` — an older
version — until the entry expired. Since the whole point of making the TTL
configurable is to let operators raise it, that window is not bounded by
anything useful. Creation now invalidates every schema's `latest` alias for
that name before priming the cache, leaving the version-pinned keys alone.
The failover cache lost its reach across schemas the same way: a global
pipeline (stored under the empty schema) loaded by schema A was cached under
`A`, so schema B using it for the first time while the pipeline table was down
missed and failed ingestion. The failover cache has no loader and so is not
subject to the single-key model of `try_get_with`; it keeps the stored-schema
key and the empty-schema-first resolution.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor(pipeline): drop cache priming on create and fold the sweep helpers
Priming the cache on create saved one read on a low-frequency operation and
cost a concept: entries were written under the creating request's schema while
`PipelineContent.schema` said empty, so the two schemas in play disagreed.
Invalidating the `latest` aliases is required regardless — that is what makes
a new version visible to other schemas — so dropping the priming loses only
the saved read, which coalescing now protects anyway. `insert_and_compile` no
longer needs the caller's schema.
`remove_cache` and the create-time invalidation collapse into one
`invalidate(name, version)`; `None` sweeps only the `latest` aliases, which is
exactly what creation wants. That leaves `invalidate_by_suffixes` and
`cache_keys` with a single caller each, so both are inlined.
Drop the `PipelineOptions` humantime test: `load_config_test` loads both
example TOMLs, which now carry `cache_ttl = "10s"`, and would fail the same
way if the serde attribute were lost. The `toml` dev-dependency goes with it.
The two invalidation tests are now checked to be orthogonal: removing the
version suffix fails only the delete test, and sweeping just the compiled
cache fails both.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(pipeline): keep failover populated across a create
The `latest` sweep on create clears the failover cache along with the loaded
ones, and after dropping the priming there was nothing writing it back. An
outage between the create and the first read-back left neither `latest` nor the
explicit version with anything to fall back on, failing ingestion — worse than
before, since the previous version's failover entry was swept too.
Creation now goes through `PipelineCache::on_pipeline_created`, which pairs the
sweep with a failover write of the new empty-schema definition. The two must
happen together, so they live behind one method rather than at the call site.
Also commit the Cargo.lock entry for the dropped `toml` dev-dependency, and
trim the comments added over the last few commits down to what the code does
not already say.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor(json2): concretize JSON2 schemas at merge scan boundaries
Infer concrete JSON2 output types from remote plans and expose them on MergeScanLogicalPlan before physical planning. Recompute affected local schemas and remove the JSON2-specific rewrite from MergeScanExec.
Add SQLness coverage for whole JSON2 columns in windows and joins.
Signed-off-by: luofucong <luofc@foxmail.com>
* fix ci
Signed-off-by: luofucong <luofc@foxmail.com>
---------
Signed-off-by: luofucong <luofc@foxmail.com>
feat(query): support list indexing for JSON2 columns
Extend JSON2 paths through DataFusion field-access planning, including nested list indexes and object fields following an index.
Preserve Variant reads for bracket JSONPath expressions and normalize dot accesses after subscripts to work around the current DataFusion planner limitation.
Add unit and sqlness coverage for nested indexes, type conflicts, missing paths, flushes, and compacted SSTs.
Signed-off-by: luofucong <luofc@foxmail.com>
Extend `WriteCacheUploadStoreWrapper::wrap` with the `OperationType` of the upload so implementations can apply per-operation policies (e.g. throttling compaction uploads but not flush uploads). Flush and compaction paths forward their existing `SstWriteRequest::op_type`; `put_and_upload_sst` is flush-only and index rebuild uploads are reported as compaction uploads.
Files: `src/mito2/src/cache/write_cache.rs`, `src/mito2/src/sst/index.rs`.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* feat(mito2): support exact sequence range reads
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* test(mito2): cover preserve row sequence table alter
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* fix(mito2): clear preserve_row_sequence marker on copy_region_from
copy_region_from copies source FileMeta into the target region, which has an
independent sequence domain. The physical per-row sequences in the copied
file belong to the source region only; trusting them in the target would let
an exact sequence-range request replay source-domain rows as if they were
target sequences. Clear the preserve_row_sequence marker on copied files so
the target fails closed with SequenceRangeUnsupported until the scan provably
cannot intersect the copied rows.
Add a regression test: copying from a preserve-enabled source into a
preserve-enabled target clears the marker, and an exact (2, 7] request on the
target returns SequenceRangeUnsupported instead of replaying source rows.
Fixes#8865
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* style: remove redundant doc comments for exact sequence range options
Approved comment-cleanup-only changes for #8865: drop outdated doc
summaries duplicated on the exact_sequence_range wrapper and the
preserve_row_sequence field, drop pure-restatement doc comments on the
SetRegionOption/UnsetRegionOption PreserveRowSequence variants, and
remove the four structural SQL comments from the alter_preserve_row_sequence
case. No behavior changes; .result regenerated by the sqlness runner.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* fix(mito2): fail closed exact reads on copied files and extension ranges
Address review feedback on #8865:
- copy_region_from: clear the source-domain FileMeta::sequence along with
the preserve_row_sequence marker. An unmarked file retaining a stale
source-domain max sequence could be silently skipped by
files_allow_exact_sequence_range() as 'proven disjoint' in the target's
independent sequence domain, dropping rows on exact (C, H] reads. With
sequence=None the capability check fails closed (SequenceRangeUnsupported)
until the copied rows are provably disjoint.
- Engine/reader: reject exact sequence-range reads whenever a follower
region has an extension range provider attached. Extension streams are
returned without a row-level sequence filter, so exactness cannot be
proven; treat the capability as missing (fail closed) instead of emitting
out-of-range rows. The reader also fails closed as defense in depth.
Tests: extend copy_region_from regression to assert the copied file's
sequence hint is cleared; mito2 suite 1148/1148 passing.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* style(mito2): use doc comments for test function descriptions
Elevate the block comments describing test functions (in scan_test and
copy_region_from_test) to /// doc comments, matching the convention used
elsewhere in the exact sequence range change. No logic change.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* refactor(mito2): extract helpers and trim comment noise in exact sequence reads
PR finalization for #8865 (zero behavior change, full mito2 suite green):
- engine: extract validate_sequence_fences and
sequence_range_unsupported_reason, keeping error variants, check order
and reason strings identical; OSS binds the extension blocker to false.
- handle_copy_region: extract remap_copied_file_meta and
file_descriptors_for_meta; rename file_ids -> source_file_ids and
files_to_copy -> new_file_metas.
- compactor: rename max_input_sequence -> known_max_input_sequence,
document the None semantics (empty input vs unknown sequence).
- Remove restating/outdated comments (ScanInput::sequence_range doc
first line, outdated file-pruning note, options test restatements),
compress verbatim comments while keeping why/invariants/contracts.
Verified: cargo check -p mito2 (+ --features enterprise), cargo fmt,
git diff --check, mito2 suite 1148/1148 passing.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* fix(mito2): reject foreign-region SSTs in exact sequence reads
Reading an SST whose FileMeta.region_id differs from the scanned region
means the region's sequence domain is broken (manifest corruption or a
repartition/copy path that leaked a source-domain file). Treat this as
an explicit RegionSequenceDomainBroken error instead of silently
ignoring the file's sequence or falling back to a full scan: the region
is unusable for exact sequence-range reads until the foreign lineage is
compacted away or repaired.
- files_allow_exact_sequence_range / exact_sequence_range now return
Result and propagate the error through engine fence validation and
scan construction (StatusCode::Internal, distinct from the
fallback-capable SequenceRangeUnsupported).
- Row-level flat-batch sequence filtering rejects foreign-region files
as defense in depth.
- Engine test asserts the broken-domain error rather than
Unsupported/fallback.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* fix(mito2): never trust unmarked SSTs for exact-range disjoint skipping
An unmarked file's FileMeta.sequence may be synthesized by the
region-edit or repartition paths (committed+1 import barrier), not a
physical max of its rows. Treating it as a whole-file disjoint proof
could permanently skip rows that were never incrementally consumed
once the flow checkpoint passes that value.
Exact sequence-range capability now requires every SST in the region to
carry the preserve_row_sequence marker; any unmarked file disables
exactness (fallback), and the (C, H] file-selection skip also only
applies to marked files. Foreign-region files still raise
RegionSequenceDomainBroken as before.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* fix(mito2): scope exact sequence-range capability to the time-selected read set
The exact capability check used to walk the entire SstVersion, so a
single unmarked or foreign-region SST anywhere in the region disabled
exact reads or raised RegionSequenceDomainBroken even when the
request's time range could never touch that file.
Both the engine fence and the scan builder now derive the read set with
shared time-pruning + exact-min/sst-min selection and validate
capability only over the files actually selected: a time-pruned file
cannot contribute a row to (C, H], so it cannot affect exactness. The
existing fail-loud semantics are unchanged for every selected file
(foreign region id -> RegionSequenceDomainBroken; unmarked -> exact
unavailable).
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* fix(mito2): wash untrusted sequences in compaction and restore barrier skipping
Compaction with any non-preserved input now writes a sequence-less
output: the physical __sequence column is zeroed (the flat format
requires the internal columns) and FileMeta.sequence records the
region-local admission barrier committed_sequence + 1 (falling back to
the flushed frontier). preserve_row_sequence stays false.
Exact sequence-range scans interpret an unmarked file's sequence as an
admission barrier: barrier <= C means flow has already consumed the
whole file, so it is skipped at file level; a missing or newer barrier
fails closed. Foreign-region files stay in the selected read set so the
capability fence still raises RegionSequenceDomainBroken.
This closes the recovery loop: after a region repartition, one
time-scoped fallback consumes the migrated rows, then compaction washes
the untrusted per-row sequences away and exact incremental reads resume
via file-level barrier skipping.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore(mito2): drop restating comments in known_max_input_sequence tests
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* test(mito2): trim SQLness result EOF whitespace
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* refactor(mito2): reuse exact scan file selection
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* test(mito2): strengthen sequence scan coverage
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* style(mito2): trim ALTER option comments
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* test(mito2): remove no-op bulk compaction check
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* fix(mito2): preserve trusted row sequences when reading SSTs
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* fix(mito2): preserve target sequence domain for imported SSTs
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* test(mito2): add trailing blank line to SQLness result EOF
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore(mito2): trim exact sequence scan plumbing
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* refactor(mito2): fold exact SST selection checks
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* test(mito2): make legacy compaction rewrite deterministic
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* test(mito2): make PK compaction rewrite deterministic
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
---------
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
feat(query): support JSON2 paths in SQL functions
Update the DataFusion fork to expose scalar function planning hooks.
Infer JSON2 path output types from scalar, aggregate, and window function signatures, while preserving the default Utf8View behavior for functions that accept arbitrary inputs.
Add unit and sqlness coverage for type conflicts, mixed typed and untyped JSON paths, filters, aggregates, and window functions.
Signed-off-by: luofucong <luofc@foxmail.com>
* feat(runtime): add weighted workload scheduler
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* feat(runtime): switch catio to GreptimeTeam fork with admission-wait metrics
Use the GreptimeTeam/catio fork (pinned c20eafc) which adds
ClassStats::total_admission_wait and ClassStats::admitted, recorded
at each QUEUED -> ADMITTED transition. This exposes the scheduler's
own admission delay (excluding Tokio queueing and poll execution),
enabling admission-wait based fairness gates.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: bump catio to dynamic-config revision
Bump the catio scheduler fork to 9f4b028 which adds
Scheduler::set_weight and Scheduler::set_max_concurrent_polls for
runtime configuration.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* feat(perf): runtime-adjustable workload scheduler parameters
Expose dynamic adjustment of the experimental workload scheduler at
runtime:
- common-runtime: set_workload_scheduler_weights and
set_workload_scheduler_max_concurrent_polls, which forward to the
catio scheduler's set_weight/set_max_concurrent_polls when the
scheduler is enabled and reject zero values.
- servers: /debug/workload_scheduler/weights and
/debug/workload_scheduler/max_concurrent_polls POST handlers, so
operators can rebalance query/write shares or admission concurrency
without restarting the datanode.
Both endpoints return 400 with a clear reason when the scheduler is
disabled or the requested value is invalid.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* feat(perf): add GET /debug/workload_scheduler status endpoint
Returns the current weights (per class), max_concurrent_polls,
active_polls and per-class counters (queued, tasks, wakes, polls,
completed, cancelled, admitted, total_admission_wait) as JSON. When the
scheduler is disabled, returns enabled=false with the other fields
omitted, so operators can distinguish 'disabled' from an error.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: bump catio to time-accounting revision
Bump the catio scheduler fork to 257ba56 which replaces
admission-count accounting with real execution-time accounting
(pass += exec_time / (weight * concurrency)), so CPU share follows the
configured weights regardless of poll length.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: bump catio to lock-free sampling revision
Bump the catio scheduler fork to efdc0a4 which adds an optional
downsampled clock sampling mode (SchedulerBuilder::sample_every_polls,
default off) with a lock-free per-class atomic counter, so the
downsampled path costs one fetch_add per poll instead of a global
mutex.
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: pin catio to scheduler PR head
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* feat(runtime): add scheduler bypass control
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: advance catio scheduler fixes
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: pin merged catio scheduler
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: regenerate config docs for workload scheduler
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: pin catio scheduler test fix
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* test(http): satisfy scheduler lifecycle clippy
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* test: add distributed scheduler toggle coverage
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* feat: finalize workload scheduler runtime controls
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: pin merged catio atomic weights
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* chore: preserve unrelated lockfile resolution
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* perf(runtime): downsample scheduler time accounting
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* test(runtime): verify cross-runtime scheduler progress
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* feat(runtime): configure scheduler poll sampling
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* docs(runtime): clarify scheduler activation
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
* docs(runtime): explain scheduler use case
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
---------
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
Co-authored-by: Ruihang Xia <waynestxia@gmail.com>
The k8s.pod runs_on k8s.node rule can only fire from kube_pod_info: it is
the only table declaring both endpoints, and co-declared edges require
both on the same row. A deployment sending only OTLP has no
kube-state-metrics tables, so its node layer is invisible and that rule
has no source at all, even though the resource attributes carry
k8s.node.name.
Declare k8s.node in otlp_trace_entities and in the synthesized resource
descriptor, and project k8s.node.name in the descriptor writer so the
column the declaration needs exists. Identity is the name rather than
k8s.node.uid: kube-state-metrics carries no node UID, so the name is the
only identity both sources can agree on.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>