* feat(semantic-graph): let the generic container yield to k8s.container A container inside a Kubernetes pod reaches the graph twice: as the k8s.container entity kube-state-metrics describes, identified by [pod uid, container name], and as the generic container the OTel resource attributes describe, identified by container.id. One physical container, two nodes. Kubernetes is the primary scenario, so k8s.container keeps its identity and the generic type stands down where it applies. Conventions gain a row-level condition for that: `suppressed_by` withdraws a declaration on rows where any of the named columns has a value. The test has to be per row, not per table — one descriptor table holds both pod rows and bare-runtime rows. Every branch that turns a declaration into rows now shares one guard (`declaration_predicate`), so the condition cannot apply to entities but not to the edges they carry. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * feat(semantic-graph): report derived entity declarations in table_semantics information_schema.table_semantics only read table options, so the declarations the built-in conventions derive — for trace tables, for whitelisted Prometheus and OTel descriptor metrics — were invisible. "Why is my table not in the graph?" was answerable only from debug logs, which is not a self-service path. A new `entity_declarations` column reports the entities a table actually contributes: each one's identity, whether it came from an option or from the conventions, and any row-level condition attached to it. An expected entity missing from the list is the answer — the table name is not whitelisted, the source stamp is wrong, an id column is absent. The row filter widens to match: a table that declares nothing by option but derives entities by convention now appears, since it is in the graph and the view has to say so. The provider reaches the derivation through a new metadata-only trait method, keeping catalog below operator. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix(semantic-graph): never yield a container to an entity nothing derives The generic container withdrew on any row carrying a pod UID, on the assumption that k8s.container would cover it. Nothing guaranteed that: k8s.container came only from kube-state-metrics, so an OTLP-only deployment — or one whose KSM data had expired or fell outside the query window — lost the container node and its edges entirely instead of gaining a more specific one. The rule now names the superseding type rather than a trigger column, and withdraws only where that type's full identity is on the row. Both OTel sources declare k8s.container themselves, under the identity kube-state-metrics gives it ([pod uid, container name]), so the two sources name one node; `k8s.container.name` joins the descriptor's projected attributes to carry it. Resolution runs once every declaration for the table is known, so a superseding type that ends up undeclared — its columns are gone, or an explicit declaration of it was skipped — leaves the guard empty and the generic container standing. A container can change type; it cannot disappear. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * refactor(semantic-graph): build the declaration JSON by serializing a type The column was assembled entry by entry into a serde_json::Map, cloning every value and allocating a String per key. A Serialize struct that consumes the declaration moves the same data instead, and the field order now reads type, origin, identity, then description. The scan path around it was doing the same kind of avoidable work: declarations were derived before the predicate could discard the table, the supersession pass cloned every declaration's identity to look up one, and the option parse claimed the time index for tables that turn out to declare nothing. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * feat(semantic-graph): keep the structured identity for single-column ids `entity_id_attrs` was NULL whenever the identity came from one column, so a consumer holding `host` = `a3f2...` had no way to tell which column produced it, and no way back to the source table. Most identities are single-column — host, k8s.pod, k8s.node, service — so the common case was the opaque one. Build the JSON object unconditionally. Entity equality still reads `entity_id` alone, so this changes no merging: it records which attributes the id was assembled from, beside an id that deliberately omits them. Also corrects the `entity_id` column doc, which still described the `k=v,k=v` rendering replaced in #8904 by values joined in declared order. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * fix(semantic-graph): keep what a container carried when it yields Yielding to k8s.container cost the row two things it had as a generic container. The runtime container id, which kube_pod_container_info keeps descriptive precisely because it is the handle back to runtime logs and metrics, went missing entirely: neither an id nor an attribute of any node. And the edge vocabulary knew only the generic type, so a row with the more complete labels ended up with fewer connections than one without — the container layer no longer reached its host. Both OTel sources now keep the runtime id and name descriptive on k8s.container, and the vocabulary gains the two edges that mirror the generic type's. Nothing checks a supersession against the edge vocabulary, so that requirement is written down where the rule is. Also: the conventions-failure path now reports the explicit half as its comment always claimed, entity_declarations reports scope columns, and identifies() is private again now that only declaration_predicate calls it. The two RFCs catch up with entity_id_attrs being unconditional, the view listing convention-derived tables, and supersession existing. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * feat(semantic-graph): report unmatched clients and the longest request `real_wins` pins a `calls` edge's RED metrics to the observed span pairs: when a window's edge key holds any pair, the unmatched client spans are suppressed so one (window, edge) yields one row. That leaves no way to tell a callee that stopped responding from traffic that stopped arriving — both show up as a lower request_count. Add two columns to `semantic_relationships`: - `unmatched_count` — client spans with no server span, counted outside `real_wins` on the same row, so the suppressed population stays visible without splitting the edge into two rows. NULL for agent calls, whose inner join leaves nothing unmatched, and for declared edges. - `duration_max` — the longest single request. It goes through `real_wins`: a pair is timed by the server span while an unmatched client is timed by its own (network wait included), so mixing them would make the max describe a different population than duration_sum and duration_count. Agent calls compute it from the child spans they already aggregate. The projection contract goes from 16 to 18 columns; every branch projects both. Explicit column queries are unaffected, `SELECT *` and ordinal-based readers see the new shape. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * test(semantic-graph): drop the duplicated source-gate case The whitelist gate on `source=opentelemetry` is already covered by `otel_implicit_declarations_are_gated` and by the wrong-source case in `table_semantics`; here it only paid for another table create, insert and full graph derivation. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> * style(semantic-graph): trim the comments back to what the code cannot say Several comments restated the code, repeated a rationale already stated at the type or in the RFC, or explained a test in more words than the test body. Signed-off-by: Dennis Zhuang <killme2008@gmail.com> --------- Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
Metrics, logs, and traces.
One engine, on your infrastructure.
A columnar database for metrics, logs, and traces on object storage. Apache-2.0 licensed core.
User Guide | API Docs | Roadmap 2026
stable for production · canary includes pre-releases · nightly is a weekly snapshot of main
- Introduction
- Why You Might Use It
- Overview
- What's Supported
- Compatibility and Migration
- Limitations and Edition Boundary
- Architecture
- Try GreptimeDB
- Getting Started
- Build From Source
- Tools & Extensions
- Project Status
- Community
- License
- Commercial Support
- Contributing
- Acknowledgement
Introduction
GreptimeDB is an open-source observability database. Metrics, logs, and traces run on one columnar engine over object storage and share one table model: tags, timestamp, and fields. When signals carry common identifiers such as service, host, or trace ID, you can correlate them in SQL without moving data between databases.
Ingest through OpenTelemetry, Prometheus Remote Write, Loki Push, or Elasticsearch Bulk. Use SQL across observability data and PromQL for metrics. Migrate ingestion one signal at a time without rebuilding your collectors.
Why You Might Use It
- You run Prometheus plus Loki or Elasticsearch and want one backend instead of three
- You have outgrown Prometheus on cardinality or retention and don't want the Thanos/Mimir operational surface
- You need long retention on object storage without a separate analytics stack
- You want to query telemetry with SQL, not only a domain query language
- You are storing GenAI or agent telemetry (OTel GenAI conventions) alongside infrastructure signals
- You need the same engine and semantics on resource-constrained devices
Learn more in Why GreptimeDB.
Overview
A quick overview of what GreptimeDB ingests, how it connects to other systems, and what its distributed engine lets you do.
What's Supported
| Ingest | OpenTelemetry (OTLP), Prometheus Remote Write, Loki Push, Elasticsearch Bulk, InfluxDB line protocol, gRPC |
| Query | SQL, PromQL, Jaeger-compatible trace queries, MySQL and PostgreSQL wire protocols |
| Storage | S3, GCS, Azure Blob and S3-compatible endpoints as primary storage, with memory and local-disk caches |
| Built in | Retention policies, downsampling, continuous aggregation, explicit table partitioning, and inverted / skipping / fulltext indexes |
Compute and storage are disaggregated: object storage holds the data, while memory and local-disk caches keep recent and frequently queried data close to compute.
Compatibility and Migration
Compatibility is per protocol, and query-side coverage is narrower than ingestion.
| Compatible | Not compatible | |
|---|---|---|
| Prometheus | Remote Write ingestion; PromQL queries | Gaps are listed in PromQL compatibility |
| Loki | Push ingestion; dual-write through Grafana Alloy makes the cutover gradual | LogQL and the rest of the Loki query API |
| Elasticsearch | _bulk ingestion in the open-source core; QueryDSL partially, in Enterprise |
Most other Elasticsearch APIs |
Benchmarks:
Limitations and Edition Boundary
Cluster deployment, object storage, the Flow engine, and every ingestion protocol listed above are in the Apache-2.0 build. Repartitioning, region migration, and index creation are manual operations there.
Read replicas, workload isolation, and automated repartitioning are GreptimeDB Enterprise features, along with enterprise security and governance. The Enterprise overview has the current list, and pricing has the edition comparison.
Architecture
GreptimeDB can run in two modes:
- Standalone — single binary for development and small deployments.
- Distributed — four components, each independently scalable:
- Frontend — protocol entry (OTel, Prometheus, MySQL/PostgreSQL, gRPC, ingestion APIs for Elasticsearch/InfluxDB/Loki) and the distributed query engine. Stateless, scales horizontally.
- Datanode — region engine with WAL, memtable, SST, cache, compaction, and indexes. Persists data to object storage. Elastic.
- Metasrv — metadata, routing, repartitioning, and security. Backed by a pluggable KV layer (etcd or RDS).
- Flownode (optional) — continuous flow computation (streaming and materialized views).
For deeper coverage, see the architecture doc or DeepWiki.
Try GreptimeDB
For AI agents — paste this prompt into your agent:
Read https://docs.greptime.com/SKILL.md and follow the instructions
to deploy, configure, ingest, and query GreptimeDB.
docker run -p 127.0.0.1:4000-4003:4000-4003 \
-v "$(pwd)/greptimedb_data:/greptimedb_data" \
--name greptime --rm \
greptime/greptimedb:latest standalone start \
--http-addr 0.0.0.0:4000 \
--grpc-bind-addr 0.0.0.0:4001 \
--mysql-addr 0.0.0.0:4002 \
--postgres-addr 0.0.0.0:4003
Dashboard: http://localhost:4000/dashboard
Read more in the full Install Guide.
Troubleshooting:
- Cannot connect to the database? Ensure that ports
4000,4001,4002, and4003are not blocked by a firewall or used by other services. - Failed to start? Check the container logs with
docker logs greptimefor further details.
Getting Started
Build From Source
Prerequisites:
- Rust toolchain — nightly, pinned by
rust-toolchain.toml - Protobuf compiler (>= 3.15)
- C/C++ building essentials:
gcc/g++/autoconfand the glibc dev package (libc6-devon Ubuntu,glibc-develon Fedora) - Python toolchain (optional, only for some test scripts)
Build and run:
make # build greptime binary
cargo run -- standalone start # start in standalone mode
Common dev commands:
make fmt # format Rust code
make clippy # lint (fails on warnings)
make test # unit + integration tests (uses cargo-nextest)
make sqlness-test # SQL regression tests
See the Contribution Guidelines for the full developer workflow.
Tools & Extensions
- Kubernetes: GreptimeDB Operator
- Helm Charts: Greptime Helm Charts
- Dashboard: Web UI
- gRPC Ingester: Go, Java, C++, Erlang, Rust, .NET, TypeScript
- Grafana Data Source: GreptimeDB Grafana data source plugin
- Grafana Dashboard: Official Dashboard for monitoring
Project Status
GreptimeDB is generally available, with stable APIs and regular releases. It runs in production at scale — OceanBase Cloud operates 80+ GreptimeDB clusters managing 300 TB of logs, cutting log storage cost by 60%+ after migrating from Grafana Loki. See more in case studies.
Read the v1.0 highlights and 2026 roadmap, or browse the version reference.
If GreptimeDB is useful to you, please star the repo.
Community
We invite you to engage and contribute!
License
GreptimeDB is an open-core project. Its core is licensed under the Apache License 2.0.
A small set of peripheral, enterprise-only features are gated behind the
enterprise Cargo feature (not built by default) and are governed by the
separate GreptimeDB Enterprise License. Source files under
that license carry an explicit Enterprise License header.
Commercial Support
Scaling observability on your infrastructure? GreptimeDB Enterprise adds the operational, security, and support layer for production deployments. Contact us for details.
Contributing
- Read our Contribution Guidelines.
- Explore Internal Concepts and DeepWiki.
- Pick up a good first issue and join the #contributors Slack channel.
Acknowledgement
Special thanks to all contributors! See AUTHOR.md.
- Uses Apache Arrow™ (memory model)
- Apache Parquet™ (file storage)
- Apache DataFusion™ (query engine)
- Apache OpenDAL™ (data access abstraction)
All trademarks, logos, and brand names referenced in this README and in the Overview diagram are the property of their respective owners. Their use is for identification purposes only and does not imply endorsement or affiliation.
