mirror of
https://github.com/GreptimeTeam/greptimedb.git
synced 2026-09-05 13:08:58 +00:00
* feat(semantic-graph): let the generic container yield to k8s.container A container inside a Kubernetes pod reaches the graph twice: as the k8s.container entity kube-state-metrics describes, identified by [pod uid, container name], and as the generic container the OTel resource attributes describe, identified by container.id. One physical container, two nodes. Kubernetes is the primary scenario, so k8s.container keeps its identity and the generic type stands down where it applies. Conventions gain a row-level condition for that: `suppressed_by` withdraws a declaration on rows where any of the named columns has a value. The test has to be per row, not per table — one descriptor table holds both pod rows and bare-runtime rows. Every branch that turns a declaration into rows now shares one guard (`declaration_predicate`), so the condition cannot apply to entities but not to the edges they carry. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * feat(semantic-graph): report derived entity declarations in table_semantics information_schema.table_semantics only read table options, so the declarations the built-in conventions derive — for trace tables, for whitelisted Prometheus and OTel descriptor metrics — were invisible. "Why is my table not in the graph?" was answerable only from debug logs, which is not a self-service path. A new `entity_declarations` column reports the entities a table actually contributes: each one's identity, whether it came from an option or from the conventions, and any row-level condition attached to it. An expected entity missing from the list is the answer — the table name is not whitelisted, the source stamp is wrong, an id column is absent. The row filter widens to match: a table that declares nothing by option but derives entities by convention now appears, since it is in the graph and the view has to say so. The provider reaches the derivation through a new metadata-only trait method, keeping catalog below operator. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix(semantic-graph): never yield a container to an entity nothing derives The generic container withdrew on any row carrying a pod UID, on the assumption that k8s.container would cover it. Nothing guaranteed that: k8s.container came only from kube-state-metrics, so an OTLP-only deployment — or one whose KSM data had expired or fell outside the query window — lost the container node and its edges entirely instead of gaining a more specific one. The rule now names the superseding type rather than a trigger column, and withdraws only where that type's full identity is on the row. Both OTel sources declare k8s.container themselves, under the identity kube-state-metrics gives it ([pod uid, container name]), so the two sources name one node; `k8s.container.name` joins the descriptor's projected attributes to carry it. Resolution runs once every declaration for the table is known, so a superseding type that ends up undeclared — its columns are gone, or an explicit declaration of it was skipped — leaves the guard empty and the generic container standing. A container can change type; it cannot disappear. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * refactor(semantic-graph): build the declaration JSON by serializing a type The column was assembled entry by entry into a serde_json::Map, cloning every value and allocating a String per key. A Serialize struct that consumes the declaration moves the same data instead, and the field order now reads type, origin, identity, then description. The scan path around it was doing the same kind of avoidable work: declarations were derived before the predicate could discard the table, the supersession pass cloned every declaration's identity to look up one, and the option parse claimed the time index for tables that turn out to declare nothing. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * feat(semantic-graph): keep the structured identity for single-column ids `entity_id_attrs` was NULL whenever the identity came from one column, so a consumer holding `host` = `a3f2...` had no way to tell which column produced it, and no way back to the source table. Most identities are single-column — host, k8s.pod, k8s.node, service — so the common case was the opaque one. Build the JSON object unconditionally. Entity equality still reads `entity_id` alone, so this changes no merging: it records which attributes the id was assembled from, beside an id that deliberately omits them. Also corrects the `entity_id` column doc, which still described the `k=v,k=v` rendering replaced in #8904 by values joined in declared order. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix(semantic-graph): keep what a container carried when it yields Yielding to k8s.container cost the row two things it had as a generic container. The runtime container id, which kube_pod_container_info keeps descriptive precisely because it is the handle back to runtime logs and metrics, went missing entirely: neither an id nor an attribute of any node. And the edge vocabulary knew only the generic type, so a row with the more complete labels ended up with fewer connections than one without — the container layer no longer reached its host. Both OTel sources now keep the runtime id and name descriptive on k8s.container, and the vocabulary gains the two edges that mirror the generic type's. Nothing checks a supersession against the edge vocabulary, so that requirement is written down where the rule is. Also: the conventions-failure path now reports the explicit half as its comment always claimed, entity_declarations reports scope columns, and identifies() is private again now that only declaration_predicate calls it. The two RFCs catch up with entity_id_attrs being unconditional, the view listing convention-derived tables, and supersession existing. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * feat(semantic-graph): report unmatched clients and the longest request `real_wins` pins a `calls` edge's RED metrics to the observed span pairs: when a window's edge key holds any pair, the unmatched client spans are suppressed so one (window, edge) yields one row. That leaves no way to tell a callee that stopped responding from traffic that stopped arriving — both show up as a lower request_count. Add two columns to `semantic_relationships`: - `unmatched_count` — client spans with no server span, counted outside `real_wins` on the same row, so the suppressed population stays visible without splitting the edge into two rows. NULL for agent calls, whose inner join leaves nothing unmatched, and for declared edges. - `duration_max` — the longest single request. It goes through `real_wins`: a pair is timed by the server span while an unmatched client is timed by its own (network wait included), so mixing them would make the max describe a different population than duration_sum and duration_count. Agent calls compute it from the child spans they already aggregate. The projection contract goes from 16 to 18 columns; every branch projects both. Explicit column queries are unaffected, `SELECT *` and ordinal-based readers see the new shape. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test(semantic-graph): drop the duplicated source-gate case The whitelist gate on `source=opentelemetry` is already covered by `otel_implicit_declarations_are_gated` and by the wrong-source case in `table_semantics`; here it only paid for another table create, insert and full graph derivation. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * style(semantic-graph): trim the comments back to what the code cannot say Several comments restated the code, repeated a rationale already stated at the type or in the RFC, or explained a test in more words than the test body. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> --------- Signed-off-by: Dennis Zhuang <killme2008@gmail.com>