* feat(semantic-graph): let the generic container yield to k8s.container
A container inside a Kubernetes pod reaches the graph twice: as the
k8s.container entity kube-state-metrics describes, identified by
[pod uid, container name], and as the generic container the OTel resource
attributes describe, identified by container.id. One physical container,
two nodes.
Kubernetes is the primary scenario, so k8s.container keeps its identity
and the generic type stands down where it applies. Conventions gain a
row-level condition for that: `suppressed_by` withdraws a declaration on
rows where any of the named columns has a value. The test has to be per
row, not per table — one descriptor table holds both pod rows and
bare-runtime rows.
Every branch that turns a declaration into rows now shares one guard
(`declaration_predicate`), so the condition cannot apply to entities but
not to the edges they carry.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* feat(semantic-graph): report derived entity declarations in table_semantics
information_schema.table_semantics only read table options, so the
declarations the built-in conventions derive — for trace tables, for
whitelisted Prometheus and OTel descriptor metrics — were invisible.
"Why is my table not in the graph?" was answerable only from debug logs,
which is not a self-service path.
A new `entity_declarations` column reports the entities a table actually
contributes: each one's identity, whether it came from an option or from
the conventions, and any row-level condition attached to it. An expected
entity missing from the list is the answer — the table name is not
whitelisted, the source stamp is wrong, an id column is absent.
The row filter widens to match: a table that declares nothing by option
but derives entities by convention now appears, since it is in the graph
and the view has to say so. The provider reaches the derivation through a
new metadata-only trait method, keeping catalog below operator.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(semantic-graph): never yield a container to an entity nothing derives
The generic container withdrew on any row carrying a pod UID, on the
assumption that k8s.container would cover it. Nothing guaranteed that:
k8s.container came only from kube-state-metrics, so an OTLP-only
deployment — or one whose KSM data had expired or fell outside the query
window — lost the container node and its edges entirely instead of
gaining a more specific one.
The rule now names the superseding type rather than a trigger column, and
withdraws only where that type's full identity is on the row. Both OTel
sources declare k8s.container themselves, under the identity
kube-state-metrics gives it ([pod uid, container name]), so the two
sources name one node; `k8s.container.name` joins the descriptor's
projected attributes to carry it.
Resolution runs once every declaration for the table is known, so a
superseding type that ends up undeclared — its columns are gone, or an
explicit declaration of it was skipped — leaves the guard empty and the
generic container standing. A container can change type; it cannot
disappear.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor(semantic-graph): build the declaration JSON by serializing a type
The column was assembled entry by entry into a serde_json::Map, cloning
every value and allocating a String per key. A Serialize struct that
consumes the declaration moves the same data instead, and the field order
now reads type, origin, identity, then description.
The scan path around it was doing the same kind of avoidable work:
declarations were derived before the predicate could discard the table,
the supersession pass cloned every declaration's identity to look up one,
and the option parse claimed the time index for tables that turn out to
declare nothing.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* feat(semantic-graph): keep the structured identity for single-column ids
`entity_id_attrs` was NULL whenever the identity came from one column, so a
consumer holding `host` = `a3f2...` had no way to tell which column produced
it, and no way back to the source table. Most identities are single-column —
host, k8s.pod, k8s.node, service — so the common case was the opaque one.
Build the JSON object unconditionally. Entity equality still reads `entity_id`
alone, so this changes no merging: it records which attributes the id was
assembled from, beside an id that deliberately omits them.
Also corrects the `entity_id` column doc, which still described the `k=v,k=v`
rendering replaced in #8904 by values joined in declared order.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(semantic-graph): keep what a container carried when it yields
Yielding to k8s.container cost the row two things it had as a generic
container. The runtime container id, which kube_pod_container_info keeps
descriptive precisely because it is the handle back to runtime logs and
metrics, went missing entirely: neither an id nor an attribute of any
node. And the edge vocabulary knew only the generic type, so a row with
the more complete labels ended up with fewer connections than one
without — the container layer no longer reached its host.
Both OTel sources now keep the runtime id and name descriptive on
k8s.container, and the vocabulary gains the two edges that mirror the
generic type's. Nothing checks a supersession against the edge
vocabulary, so that requirement is written down where the rule is.
Also: the conventions-failure path now reports the explicit half as its
comment always claimed, entity_declarations reports scope columns, and
identifies() is private again now that only declaration_predicate calls
it. The two RFCs catch up with entity_id_attrs being unconditional, the
view listing convention-derived tables, and supersession existing.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* feat(semantic-graph): report unmatched clients and the longest request
`real_wins` pins a `calls` edge's RED metrics to the observed span pairs:
when a window's edge key holds any pair, the unmatched client spans are
suppressed so one (window, edge) yields one row. That leaves no way to
tell a callee that stopped responding from traffic that stopped arriving —
both show up as a lower request_count.
Add two columns to `semantic_relationships`:
- `unmatched_count` — client spans with no server span, counted outside
`real_wins` on the same row, so the suppressed population stays visible
without splitting the edge into two rows. NULL for agent calls, whose
inner join leaves nothing unmatched, and for declared edges.
- `duration_max` — the longest single request. It goes through
`real_wins`: a pair is timed by the server span while an unmatched
client is timed by its own (network wait included), so mixing them would
make the max describe a different population than duration_sum and
duration_count. Agent calls compute it from the child spans they already
aggregate.
The projection contract goes from 16 to 18 columns; every branch projects
both. Explicit column queries are unaffected, `SELECT *` and
ordinal-based readers see the new shape.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* test(semantic-graph): drop the duplicated source-gate case
The whitelist gate on `source=opentelemetry` is already covered by
`otel_implicit_declarations_are_gated` and by the wrong-source case in
`table_semantics`; here it only paid for another table create, insert and
full graph derivation.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* style(semantic-graph): trim the comments back to what the code cannot say
Several comments restated the code, repeated a rationale already stated at
the type or in the RFC, or explained a test in more words than the test
body.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(meta): actually acquire logical table locks in alter-logical-tables procedure
The procedure listed its logical table locks from table_info_values,
which is only filled during Prepare, while procedure lock keys are
fixed at submission — so the logical locks were never acquired. Today
every writer of a logical table's info is serialized by the physical
table lock, which hides the problem; a metadata-only alter procedure
targeting a single logical table would race it.
Resolve the logical table ids at submission, persist them in the
procedure state (serde(default): state dumped by older versions keeps
the previous behavior), lock physical + logical tables, and re-check
the resolved ids against the locked set at Prepare so a table dropped
and recreated after submission cannot be mutated without a lock.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* feat: manage semantic table options via ALTER TABLE SET/UNSET
CREATE TABLE accepts greptime.semantic.* options, but ALTER TABLE SET
routed every option through SetRegionOption, whose closed match
rejects them — tables auto-created by ingestion could never receive
semantic declarations after the fact.
Semantic options are pure metadata markers no region consumes, so
they now take a metadata-only alter, following the repartition-hint
precedent:
- New AlterKind::SetAnnotations/UnsetAnnotations carrying an
AnnotationFamily (currently only Semantic), so future marker-style
option families reuse the same machinery. The converter classifies
a SET/UNSET batch by key prefix and rejects batches that mix
annotation keys with regular options.
- The procedure reuses the MetadataOnly flow: no region dispatch,
table-info update plus cache invalidation only.
- Validation lives in the table-meta mutation layer, so it runs at
frontend verification and again in the procedure's prepare step
under the table lock: SET is strict (known key, value domain,
entity columns exist and render as strings); UNSET is lenient
inside the namespace so stale keys can be cleaned up.
ModifyColumnTypes re-checks columns referenced by entity
declarations at the same layer, closing a verify-then-execute race.
- Logical metric tables are supported: an annotation alter submits a
regular alter-table task locking only the logical table, and the
DDL manager's physical-route guard admits it.
- create_table_info re-checks semantic value domains for gRPC-built
expressions that bypass the SQL parser.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor(table): centralize annotation option classification and validation
Address review feedback on the AnnotationFamily abstraction: with only
one variant that every consumer immediately destructured, the
generality was fake. Make it real and exhaustive instead:
- AnnotationFamily gains RepartitionHint: repartition.column.hint is
the same kind of marker option (pure metadata, no region consumes
it) and previously had a hand-rolled special case in the converter,
the metadata-only classifier, and a dedicated AlterKind pair — all
deleted, one classification API remains. Per-family logical-table
eligibility (allows_logical_tables) replaces the hard-coded
Semantic check in the DDL manager guard.
- One validation core in the table crate (check_annotation) serves
both DDL entry points. CREATE and ALTER previously duplicated the
rules; each keeps its existing error variants, status codes and
messages via thin adapters over a typed error (ALTER missing column
stays 4002 TableColumnNotFound, CREATE stays InvalidArguments).
- The batch classifier returns Result instead of swallowing the
mixed-batch error: a mixed SET on a logical table now reports the
actual problem instead of UnexpectedLogicalRouteTable, and the flow
classifiers propagate instead of guessing. The converter also moves
its owned payloads instead of cloning them.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* test(meta): cover logical-table annotation alter routing
The route-guard branch admitting metadata-only annotation alters on
logical tables was only exercised end to end by sqlness. Pin it at the
DDL manager level: a semantic SET on a logical table succeeds, updates
only the logical table's metadata and dispatches nothing to datanodes;
a mixed batch reports its own error instead of the route guard's; the
repartition hint stays rejected on logical routes.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(table): keep entity guard on ADD COLUMN and report missing columns first
Review follow-ups: the old verify_alter loop scanned the post-alter
schema, so it also caught DROP COLUMN followed by re-adding the
declared column with a non-string type — the mutation-layer move only
kept the MODIFY path. Guard add_columns the same way (this also covers
ingestion auto-alter). And run the MODIFY drift check after the
existence lookup, so altering a dropped-but-still-declared column
reports ColumnNotExists (4002) like every other MODIFY on a missing
column.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* style(grpc-expr): drop a test comment restating the classifier doc
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor(table): rename annotation validation helpers per review
check_annotation* validated and normalized; align the names with the
validate_and_normalize_* convention nearby, and spell out
AnnotationContext (Cx is not used in this repo).
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* feat(table): add entity semantic declarations
Define open-ended greptime.semantic.entity.* options, validate entity columns at DDL time, and stamp OTLP trace tables with the service entity declaration.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* feat: add read-time entity relationships graph
Add computed semantic graph tables, typed DataFusion derivation plans for entity registry and trace calls edges, and streaming read-time execution.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* test: exclude semantic graph tables from table constraints
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor(operator): name the plan-builder source groupings
Review feedback: build_registry_plan / build_calls_plan took anonymous
(declarations, DataFrame) tuples while the caller already grouped the same
fields. Introduce RegistrySource { declarations, scan } and CallsSource
{ service, scan } next to the builders and flow them through the frontend
caller and tests. The frontend-side EntitySource keeps holding a TableRef
(the operator builders stay pure over already-built scans), so the named
structs live in operator rather than reusing that type. No behavior change.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* feat: report region read load in heartbeat
Signed-off-by: WenyXu <wenymedia@gmail.com>
* feat: expose region query stats in information schema
Signed-off-by: WenyXu <wenymedia@gmail.com>
* chore: update sqlness result
Signed-off-by: WenyXu <wenymedia@gmail.com>
* fix: record region query stats on stream drop
Signed-off-by: WenyXu <wenymedia@gmail.com>
* fix: keep region query cpu stats in nanoseconds
Signed-off-by: WenyXu <wenymedia@gmail.com>
---------
Signed-off-by: WenyXu <wenymedia@gmail.com>
* feat: table semantic layer information_schema view (Phase 3)
Add `information_schema.table_semantics`, a queryable view over the table
semantic layer. One row per table that carries at least one
`greptime.semantic.*` option: the signal-agnostic keys
(signal_type/source/pipeline/metadata_quality) are promoted to columns and
the remaining signal-specific keys are folded into a `semantic_options`
JSON string. Tables with no semantic key are excluded.
Stacked on Phase 2.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* chore: address PR review on table_semantics
- fold JSON serialization failure into None instead of unwrap/panic
- drop per-row Vec allocation in predicate eval; use a fixed array
- align RFC view name with the shipped `table_semantics`
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* chore: update results
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* feat: switch partition tree to bulk
Signed-off-by: evenyag <realevenyag@gmail.com>
* chore: keep partition tree memtable for migration test
Restore PartitionTreeMemtable construction when memtable.type=partition_tree
is explicit, and move the sparse-encoding bulk override into the default
(no explicit memtable.type) arm so phase 2's memtable.type=bulk wins on
reopen. Rewrite test_reopen_time_series_sparse_memtable_with_bulk to use a
metric-engine-shaped schema and sparse-encoded rows with WriteHint::Sparse,
so the test actually exercises a PartitionTreeMemtable in phase 1 and
verifies WAL replay into the new BulkMemtable on reopen without flushing.
Signed-off-by: evenyag <realevenyag@gmail.com>
* chore: drop partition tree memtable from runtime
Re-apply the unconditional sparse-encoding override in
`MemtableBuilderProvider::builder_for_options` and route the
`MemtableOptions::PartitionTree` arm to `BulkMemtable` with a deprecation
warning. After this change, `PartitionTreeMemtableBuilder` is no longer
reachable from the engine runtime; benchmarks still reference the type.
Remove `test_reopen_time_series_sparse_memtable_with_bulk` and the
`put_sparse_rows` helper added in the previous commit — that test only
existed to validate the PartitionTree -> Bulk reopen migration and is
unnecessary now that the override is in place.
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito2): move timestamp_array_to_i64_slice into read module
Relocate the timestamp_array_to_i64_slice helper from
memtable/partition_tree/data.rs to the read module so that the read
path no longer depends on the partition_tree internals. All call sites
(both inside and outside the partition_tree module) now import from
crate::read.
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito2): use TimeSeriesMemtableBuilder in time_partition tests
The time_partition tests use the memtable builder purely as a generic
backend for the TimePartitions write/scan paths; nothing in them is
specific to the partition-tree memtable. Switch the seven affected
tests to TimeSeriesMemtableBuilder so the tests no longer depend on
PartitionTreeMemtableBuilder.
Signed-off-by: evenyag <realevenyag@gmail.com>
* chore(mito2): delete PartitionTreeMemtable implementation
The runtime already falls back to BulkMemtable for the PartitionTree
variant. Drop the now-unreachable implementation, its metrics, the
partition_tree benchmarks, the metric-engine Unsupported fallback in
bulk_insert.rs, and the test helpers that only existed for the deleted
module.
MemtableOptions::PartitionTree, its parsing, the runtime fallback, the
store-api MEMTABLE_PARTITION_TREE_* constants, and the SQL fixtures
remain so existing region options keep round-tripping.
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor(mito-codec): drop skip_partition_column parameter
PartitionTreeMemtable was the only caller passing
skip_partition_column=true; every other caller passes false. Now that
the partition_tree module is gone, the parameter is uniformly false
and the guard branch is dead. Drop the parameter from the trait method
and both impls, remove the guard and the is_partition_column helper,
and update the four remaining call sites in mito2 plus the bench.
Signed-off-by: evenyag <realevenyag@gmail.com>
* chore(mito2): remove unused MemtableConfig enum
Signed-off-by: evenyag <realevenyag@gmail.com>
* chore: fmt code
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor: remove unused variant
Signed-off-by: evenyag <realevenyag@gmail.com>
* test: update test_config_api
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix: remove unused memtable test helpers
Signed-off-by: evenyag <realevenyag@gmail.com>
* chore: address review comment
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix: support bulk memtable options
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix: sanitize config
Signed-off-by: evenyag <realevenyag@gmail.com>
* feat: remove partition tree options from region options
Move primary_key_encoding to the top level
Signed-off-by: evenyag <realevenyag@gmail.com>
* test: make ssts test datetime replaced text stable
Signed-off-by: evenyag <realevenyag@gmail.com>
* test: update sqlness result
Signed-off-by: evenyag <realevenyag@gmail.com>
* chore: validate_enum_options consider bulk memtable
Signed-off-by: evenyag <realevenyag@gmail.com>
* refactor: pass region id when parsing region options
Replace the `TryFrom<&HashMap>` impl for `RegionOptions` with
`try_from_options(region_id, options_map)` so the legacy partition_tree
fallback can log the affected region. The fallback now also overrides
the SST format to flat in addition to clearing the memtable type.
Signed-off-by: evenyag <realevenyag@gmail.com>
* fix: align sst_format with bulk memtable on parse and open
Signed-off-by: evenyag <realevenyag@gmail.com>
---------
Signed-off-by: evenyag <realevenyag@gmail.com>
* change dep
Signed-off-by: Ruihang Xia <waynestxia@gmail.com>
* feat: adapt to arrow's interval array
* chore: fix compile errors in datatypes crate
* chore: fix api crate compiler errors
* chore: fix compiler errors in common-grpc
* chore: fix common-datasource errors
* chore: fix deprecated code in common-datasource
* fix promql and physical plan related
Signed-off-by: Ruihang Xia <waynestxia@gmail.com>
* wip: upgrading network deps
Signed-off-by: Ruihang Xia <waynestxia@gmail.com>
* block on updating `sqlparser`
* upgrade sqlparser
Signed-off-by: Ruihang Xia <waynestxia@gmail.com>
* adapt new df's trait requirements
Signed-off-by: Ruihang Xia <waynestxia@gmail.com>
* chore: fix compiler errors in mito2
* chore: fix common-function crate errors
* chore: fix catalog errors
* change import path
Signed-off-by: Ruihang Xia <waynestxia@gmail.com>
* chore: fix some errors in query crate
* chore: fix some errors in query crate
* aggr expr and some other tiny fixes
Signed-off-by: Ruihang Xia <waynestxia@gmail.com>
* chore: fix expr related errors in query crate
* chore: fix query serializer and admin command
* chore: fix grpc services
* feat: axum serve
* chore: fix http server
* remove handle_error handler
* refactor timeout layer
* serve axum
* chore: fix flow aggr functions
* chore: fix flow
* feat: fix errors in meta-srv
* boxed()
* use TokioIo
* feat!: Remove script crate and python feature (#5321)
* feat: exclude script crate
* chore: simplify feature
* feat: remove the script crate
* chore: remove python feature and some comments
* chore: fix warning
* chore: fix servers tests compiler errors
* feat: fix tests-integration errors
* chore: fix unused
* test: fix catalog test
* chore: fix compiler errors for crates using common-meta
testing feature is enabled when check with --workspace
* test: use display for logical plan test
* test: implement rewrite for ScanHintRule
* fix: http server build panic
* test: fix mito test
* fix: sql parser type alias error
* test: fix TestClient not listen
* test: some flow tests
* test(flow): more fix
* fix: test_otlp_logs
* test: fix promql test that using deprecated method fun()
* fix: sql type replace supports Int8 ~ Int64, UInt8 ~ UInt64
* test: fix infer schema test case
* test: fix tests related to plan display
* chore: fix last flow test
* test: fix function format related assertion
* test: use larger port range for tests
* fix: test_otlp_traces
* fix: test_otlp_metrics
* fix range query and dist plan
Signed-off-by: Ruihang Xia <waynestxia@gmail.com>
* fix: flow handle distinct use deprecated field
* fix: can't pass Join plan expressions to LogicalPlan::with_new_exprs
* test: fix deserialize test
* test: reduce split key case num
* tests: lower case aggr func name
* test: fix some sqlness tests
* tests: more sqlness fix
* tests: fixed sqlness test
* commit non-bug changes
Signed-off-by: Ruihang Xia <waynestxia@gmail.com>
* fix: make our udf correct
* fix: implement empty methods of ContextProvider for DfContextProviderAdapter
* test: update sqlness test result
* chore: remove unused
* fix: provide alias name for AggregateExprBuilder in range plan
* test: update range query result
* fix: implement missing ContextProvider methods for DfContextProviderAdapter
* test: update timestamps, cte result
* fix: supports empty projection in mito
* test: update comment for cte test
* fix: support projection for numbers
* test: update test cases after projection fix
* fix: fix range select first_value/last_value
* fix: handle CAST and time index conflict
* fix: handle order by correctly in range first_value/last_value
* test: update sqlness result
* test: update view test result
* test: update decimal test
wait for https://github.com/apache/datafusion/pull/14126 to fix this
* feat: remove redundant physical optimization
todo(ruihang): Check if we can remove this.
* test: update sqlness test result
* chore: range select default sort use nulls_first = false
* test: update filter push down test result
* test: comment deciaml test to avoid different panic message
* test: update some distributed test result
* test: update test for distributed count and filter push down
* test: update subqueries test
* fix: SessionState may overwrite our UDFs
* chore: fix compiler errors after merging main
* fix: fix elasticsearch and dashboard router panic
* chore: fix common-functions tests
* chore: update sqlness result
* test: fix id keyword and update sqlness result
* test: fix flow_null test
* fix: enlarge thread size in debug mode to avoid overflow
* chore: fix warnings in common-function
* chore: fix warning in flow
* chore: fix warnings in query crate
* chore: remove unused warnings
* chore: fix deprecated warnings for parquet
* chore: fix deprecated warning in servers crate
* style: fix clippy
* test: enlarge mito cache tttl test ttl time
* chore: fix typo
* style: fmt toml
* refactor: reimplement PartialOrd for RangeSelect
* chore: remove script crate files introduced by merge
* fix: return error if sql option is not kv
* chore: do not use ..default::default()
* chore: per review
* chore: update error message in BuildAdminFunctionArgsSnafu
Co-authored-by: jeremyhi <jiachun_feng@proton.me>
* refactor: typed precision
* update sqlness view case
Signed-off-by: Ruihang Xia <waynestxia@gmail.com>
* chore: flow per review
* chore: add example in comment
* chore: warn if parquet stats of timestamp is not INT64
* style: add a newline before derive to make the comment more clear
* test: update sqlness result
* fix: flow from substrait
* chore: change update_range_context log to debug level
* chore: move axum-extra axum-macros to workspace
---------
Signed-off-by: Ruihang Xia <waynestxia@gmail.com>
Co-authored-by: Ruihang Xia <waynestxia@gmail.com>
Co-authored-by: luofucong <luofc@foxmail.com>
Co-authored-by: discord9 <discord9@163.com>
Co-authored-by: shuiyisong <xixing.sys@gmail.com>
Co-authored-by: jeremyhi <jiachun_feng@proton.me>
* feat: introduce `PrimaryKeyEncoding`
* fix: fix unit tests
* chore: add empty line
* test: add unit tests
* chore: fmt code
* refactor: introduce new codec trait to support various encoding
* fix: fix unit tests
* chore: update sqlness result
* chore: apply suggestions from CR
* chore: apply suggestions from CR
* fix: data_length, index_length, table_rows in tables
* feat: table stats only works for mito engine currently
* fix: tests
* fix: typo
* chore: log error when region_stats fails
* feat: adds index size to region statistics
* feat: adds the number of rows for region statistics
* test: adds sqlness test for region_statistics
* fix: test
* fix: table resolving logic related to pg_catalog
refer to
https://github.com/GreptimeTeam/greptimedb/issues/3560#issuecomment-2287794348
and #4543
* refactor: remove CatalogProtocol type
* fix: sqlness
* fix: forbid create database pg_catalog with mysql client
* refactor: use QueryContext as arguments rather than Channel
* refactor: pass None as default behaviour in information_schema
* test: fix test
* feat: add function 'pg_catalog.pg_table_is_visible'q
* feat: add 'pg_class' and 'pg_namespace', now we can run '\d' and '\dt'!
* refactor: move memory_table::tables to utils::tables
* refactor: move out predicate to system_schema to reuse it
* feat: predicates pushdown
* test: add pg_namespace, pg_class related sqlness test
* fix: typos and license header
* fix: sqlness test
* refactor: use `expect` instead of `unwrap` here
* refactor: remove the `information_schema::utils` mod
* doc: make the comment in pg_get_userbyid more precise
* doc: add TODO and comment in pg_catalog
* fix: typo
* fix: sqlness
* doc: change to comment on PGClassBuilder to TODO
* WIP: pg_catalog
* refactor: move memory_table to crate public level to reuse it in pgcatalog
* refactor: new system_schema mod to manage implementation of information_schema and pg_catalog
* feat: pg_catalog.pg_type
* fix: remove unused code to avoid warning
* test: add pg_catalog sqlness test
* feat: pg_catalog_cache in system_catalog
* fix: integration test
* test: rollback unit test
* refactor: mix pg_catalog table_id with old ones
* fix: add todo information
* tests: rerun sqlness
---------
Co-authored-by: johnsonlee <johnsonlee@localhost.localdomain>
* fix: forbid to change tables in information_schema
* refactor: use unified read-only check function
* test: add more sqlness tests for information_schema
* refactor: move is_readonly_schema to common_catalog
* fix: `region_peers` returns same region_id for multi logical tables
* test: add sqlness test for information_schema.region_peers
* refactor: region_peers sqlness