dennis zhuang 8d9f24d1b6 docs: add entity relationships and graph query RFC (#8605)
* docs: add entity relationships and graph query RFC

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* docs: revise entity-graph RFC after review

- Zero-configuration declarations: the Prometheus-on-Kubernetes convention
  pack (job/instance per the OTel-Prometheus compatibility spec, k8s SD
  labels, *_info descriptors, target_info enrichment) alongside the OTLP
  trace auto-stamp, plus Remote Write 2.0 inline metadata.
- Calls endpoints follow the service entity declaration; self-calls
  compare full endpoint ids; no silent identity fallback.
- Strict time-window contract: the source window is never narrower than
  the query's observed_at range; unsafe-to-extract predicates error
  instead of silently defaulting.
- scope removed from the relationship schema (kept on entities as a
  display property); entity row contract restated per projected
  observation; endpoint encoding documented as the v1 storage-level key
  with its known collision limitation.
- Sampling caveats corrected (ratios are representative only under
  unbiased sampling); snapshot relation synthesizes endpoint-only
  vertices; shared attributes provide join keys while co-declaration
  provides relationship semantics.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* docs: align entity-graph RFC contracts and tighten prose

- Metric/trace service unification is promised only when service.namespace
  is empty and job is not relabeled (the compatibility spec renders job as
  <namespace>/<name>); otherwise alignment needs pipeline normalization or
  explicit declarations.
- Entity row contract stated once (per projected observation); the calls
  defining SQL is marked as the single-column simplification of the
  declaration-derived endpoint ids; Remote Write 2.0 metadata intake and
  the Prometheus implicit declarations are listed as M1 work.
- Compress survey/example/reference prose.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* docs: pin down the k8s convention pack rules and edge directions

- kube_pod_owner implicitly declares k8s.workload with id
  (namespace, owner_kind, owner_name) and derives pod part_of workload;
  target_info's non-job/instance labels are implicit service descriptive
  columns — fixed rules, no new declaration syntax.
- One direction for pod placement: k8s.pod runs_on k8s.node (pod added to
  runs_on sources; node->pod removed from contains).
- Drop the remaining 'canonical' wording for the v1 storage-level id.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* docs: make contains/part_of a true inverse pair

part_of covers service.instance->service and k8s.pod->k8s.workload with
contains as its inverse; has_instance is dropped.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* docs: rewrite entity-graph RFC as a design narrative

Restructure for human review: mainline first, cases illustrate the
design instead of specifying it.

- New Architecture and Benefits and Drawbacks sections; the calls
  derivation stays expanded as the flagship example while schema
  enumerations, window-rule listings, and executor edge-case handling
  move out of the document.
- The cross-signal promise is stated honestly: neighbours and their
  source tables are discovered first, their telemetry is the next
  query — one engine, two statements; the worked example shows the
  full declaration -> entity -> edge -> telemetry flow.
- Default materialisation added as the most direct alternative to
  read-time derivation, with its costs.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* docs: final wording pass on the entity-graph RFC

Mark the property-graph DDL as illustrative rather than settled M2
syntax, credit standards as foundations rather than claiming wholesale
alignment, and clean up punctuation-heavy prose.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* docs: keep the RFC at design altitude

Demote the convention-pack rule details and snapshot property-merge
semantics to the implementing changes; correct the single-trace-table
assumption (traces can be routed to multiple tables); record
attribute-key participation in entity equality as an open question.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* docs: correct service graph terminology

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-07-23 06:47:41 +00:00
2023-08-10 08:08:37 +00:00
2023-06-25 11:05:46 +08:00
2023-11-09 10:38:12 +00:00
2023-03-28 19:14:29 +08:00

GreptimeDB Logo

One database for metrics, logs, and traces
replacing Prometheus, Loki, and Elasticsearch

The unified OpenTelemetry backend — with SQL + PromQL on object storage.

Introduction

GreptimeDB is an open-source observability database built for Observability 2.0 — treating metrics, logs, and traces as one unified data model (wide events) instead of three separate pillars.

Use it as the single OpenTelemetry backend — replacing Prometheus, Loki, and Elasticsearch with one database built on object storage. Query with SQL and PromQL, scale without pain, cut costs up to 50×.

Overview

A quick overview of what GreptimeDB ingests, how it connects to other systems, and what its distributed engine lets you do.

GreptimeDB Overview

Features

Feature Description
Observability 2.0 native Logs, metrics, and traces in one engine with SQL + PromQL. Native OpenTelemetry, Prometheus remote write, and Jaeger. Migrate one signal at a time, or use as a single backend.
Elastic compute-storage separation Scale reads independently with horizontal replicas. Serve high-concurrency workloads from dashboards, alerting, and AI agents — without resharding or data migration.
Sub-second on PBEB-scale data Columnar engine with fulltext, inverted, and skipping indexes. Written in Rust. Designed for high-concurrency point queries, not just analytical scans.
50× lower cost Object storage (S3, GCS, Azure Blob) as primary storage, with a tiered cache (memory + local disk) to keep writes and queries fast.

Perfect for:

  • Replacing Prometheus + Loki + Elasticsearch with a single observability backend
  • Scaling past Prometheus — high cardinality, long-term storage, no Thanos/Mimir overhead
  • AI/agent workloads — store GenAI telemetry (OTel GenAI conventions), and serve high-concurrency reads from SRE/developer agents via horizontal read replicas
  • Cutting observability costs with object storage (up to 50× savings on traces, 30% on logs)
  • Edge-to-cloud observability with unified APIs on resource-constrained devices

Why Observability 2.0? Three separate databases for metrics, logs, and traces means three storage layers, three query languages, and three sets of dashboards. GreptimeDB stores all three as timestamped wide events in one columnar engine — JOIN across signals in SQL, run one stack instead of three, and ingest AI agent telemetry the same way. Read more: Observability 2.0 and the Database for It.

Learn more in Why GreptimeDB.

How GreptimeDB Compares

Capability GreptimeDB Prometheus / Thanos / Mimir Grafana Loki Elasticsearch
Data types Metrics, logs, traces Metrics only Logs only Logs, traces
Query language SQL + PromQL PromQL LogQL Query DSL
Storage Native object storage (S3, etc.) Local disk + object storage (Thanos/Mimir) Object storage (chunks) Local disk
Scaling Compute-storage separation, stateless nodes Federation / Thanos / Mimir — multi-component, ops heavy Stateless + object storage Shard-based, ops heavy
Cost efficiency Up to 50× lower storage cost High at scale Moderate High (inverted index overhead)
OpenTelemetry Native (metrics + logs + traces) Partial (metrics only) Partial (logs only) Via instrumentation

Benchmarks:

Architecture

GreptimeDB can run in two modes:

  • Standalone — single binary for development and small deployments.
  • Distributed — four components, each independently scalable:
    • Frontend — protocol entry (OTel, Prometheus, MySQL/PostgreSQL, gRPC, ingestion APIs for Elasticsearch/InfluxDB/Loki) and the distributed query engine. Stateless, scales horizontally.
    • Datanode — region engine with WAL, memtable, SST, cache, compaction, and indexes. Persists data to object storage. Elastic.
    • Metasrv — metadata, routing, repartitioning, autopilot, and security. Backed by a pluggable KV layer (etcd or RDS).
    • Flownode (optional) — continuous flow computation (streaming and materialized views).

For deeper coverage, see the architecture doc or DeepWiki.

GreptimeDB System Overview

Try GreptimeDB

For AI agents — paste this prompt into your agent:

Read https://docs.greptime.com/SKILL.md and follow the instructions
to deploy, configure, ingest, and query GreptimeDB.
docker run -p 127.0.0.1:4000-4003:4000-4003 \
  -v "$(pwd)/greptimedb_data:/greptimedb_data" \
  --name greptime --rm \
  greptime/greptimedb:latest standalone start \
  --http-addr 0.0.0.0:4000 \
  --rpc-bind-addr 0.0.0.0:4001 \
  --mysql-addr 0.0.0.0:4002 \
  --postgres-addr 0.0.0.0:4003

Dashboard: http://localhost:4000/dashboard

Read more in the full Install Guide.

Troubleshooting:

  • Cannot connect to the database? Ensure that ports 4000, 4001, 4002, and 4003 are not blocked by a firewall or used by other services.
  • Failed to start? Check the container logs with docker logs greptime for further details.

Getting Started

Build From Source

Prerequisites:

  • Rust toolchain — nightly, pinned by rust-toolchain.toml
  • Protobuf compiler (>= 3.15)
  • C/C++ building essentials: gcc / g++ / autoconf and the glibc dev package (libc6-dev on Ubuntu, glibc-devel on Fedora)
  • Python toolchain (optional, only for some test scripts)

Build and run:

make                          # build greptime binary
cargo run -- standalone start # start in standalone mode

Common dev commands:

make fmt            # format Rust code
make clippy         # lint (fails on warnings)
make test           # unit + integration tests (uses cargo-nextest)
make sqlness-test   # SQL regression tests

See the Contribution Guidelines for the full developer workflow.

Tools & Extensions

Project Status

GreptimeDB is at v1.0 GA with stable APIs and regular releases. It runs in production at scale — OceanBase Cloud operates 80+ GreptimeDB clusters managing 300 TB of logs, cutting log storage cost by 60% after migrating from Grafana Loki. See more in case studies.

Read the v1.0 highlights and 2026 roadmap, or browse the version reference.

If GreptimeDB is useful to you, please star the repo.

Star History Chart Known Users

Community

We invite you to engage and contribute!

License

GreptimeDB is an open-core project. Its core is licensed under the Apache License 2.0.

A small set of peripheral, enterprise-only features are gated behind the enterprise Cargo feature (not built by default) and are governed by the separate GreptimeDB Enterprise License. Source files under that license carry an explicit Enterprise License header.

Commercial Support

Running GreptimeDB in your organization? We offer enterprise add-ons, services, training, and consulting. Contact us for details.

Contributing

Acknowledgement

Special thanks to all contributors! See AUTHOR.md.


All trademarks, logos, and brand names referenced in this README and in the Overview diagram are the property of their respective owners. Their use is for identification purposes only and does not imply endorsement or affiliation.

S
Description
Languages
Rust 98.8%
Python 0.6%
Shell 0.3%
JavaScript 0.1%