* fix(ci): teach check-builder-rust-version.sh to handle stable channels
The script extracted a YYYY-MM-DD date from rust-toolchain.toml to
compare against the rustc build date inside the dev-builder image —
a nightly-era design. With channel = "1.96.1" there is no date in
the file, so every release build failed with 'Error: No rust toolchain
version found in rust-toolchain.toml'.
Extract the channel token instead and branch on it:
- stable channel (X.Y[.Z]): require the builder image's rustc to
exactly match the pinned version
- nightly-YYYY-MM-DD: keep the legacy date-difference check
Verified against a mocked docker for all four paths (stable match /
mismatch, nightly fresh / stale).
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* chore(toolchain): finish stable-migration cleanup in docs and query-regression pin
- README/AGENTS: the toolchain is now stable Rust pinned by
rust-toolchain.toml, not nightly
- query-regression: align the benchmark toolchain pin with the
workspace (nightly-2026-03-21 = 1.96.0-nightly -> stable 1.96.1),
including the exact-version assertions (cargo 356927216, rustc
31fca3adb, both 2026-06-26) and the runner image default
The query-regression runner image must be rebuilt and
QUERY_REGRESSION_ECS_IMAGE_ID bumped together with these pins
(per .github/runner-scale-sets/query-regression/README.md) before
the next regression run.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci(query-regression): derive the Rust toolchain pin from rust-toolchain.toml
Replace the hard-coded RUSTUP_TOOLCHAIN value and the hard-coded
version strings in the runner Verify assertions with a pin resolved
from rust-toolchain.toml:
- the always-running test-tooling job exports the channel parsed from
rust-toolchain.toml as a job output
- query-regression sets RUSTUP_TOOLCHAIN from that output
- the Verify step escapes the pin into the cargo/rustc/active-toolchain
regexes at runtime; the exact commit hash and date are asserted
generically since a stable version identifies the release
Removing the redundant require_eq (workflow yaml vs runner env) since
both now flow from the single source of truth. When rust-toolchain.toml
is bumped, the run fails with a clear signal until the runner image is
rebuilt with the new toolchain, keeping the existing lockstep contract.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci(query-regression): derive the runner image toolchain from rust-toolchain.toml
Remove the hard-coded 'ARG RUST_TOOLCHAIN=1.96.1' from the
query-regression runner Dockerfile. The pin is now parsed from a
COPY'd rust-toolchain.toml at build time (the bootstrap script builds
with the repo root as context, so the file is in the build context):
- rustup-init installs the parsed channel as the default toolchain
- the baked ENV RUSTUP_TOOLCHAIN is dropped: the rustup default makes
bare cargo/rustc resolve correctly without it, and the workflow
supplies RUSTUP_TOOLCHAIN explicitly at run time
- the build-time self-verification asserts the active toolchain
against the same parsed pin
With this, rust-toolchain.toml is the single source of truth for the
benchmark toolchain end to end: the image bakes whatever the toml says
at build time and the workflow asserts against the toml at run time.
A toolchain bump now only requires rebuilding the image.
The changed mechanics were verified natively with the real rustup-init
1.29.0 and the real 1.96.1 toolchain (registry pulls are unavailable
in this sandbox): parsing, default-toolchain installation without
RUSTUP_TOOLCHAIN env, bare cargo/rustc resolution, and the
active-toolchain assertion for both '(default)' and '(overridden)'
output forms.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci(query-regression): rebuild the runner image automatically on toolchain changes
Mirror the dev-builder automation for the query-regression ECS runner
image: a new rebuild-query-regression-runner-image.yaml workflow runs
whenever rust-toolchain.toml or the query-regression runner directory
changes on main (or via manual dispatch). It drives the existing
build-ecs-image.py ops tool, then completes the documented lockstep
updates in order: bump RUNNER_IMAGE_EPOCH in query-regression.yml and
push the commit to main, and only then point the
QUERY_REGRESSION_ECS_IMAGE_ID repo variable at the new image, so the
next regression run picks up image and epoch together.
Also fix build-ecs-image.py to stage rust-toolchain.toml into the
temporary docker build context: the AMI path builds the embedded
Dockerfile from an empty /tmp/image-context, which would break on the
Dockerfile's COPY of rust-toolchain.toml introduced earlier. The
user-data now base64-stages the toml next to the Dockerfile before
docker build.
Verified: py_compile, render_user_data round-trip (mkdir -> stage ->
docker build ordering), and the RUNNER_IMAGE_EPOCH bump sed against
the real workflow file. Requires a new ALIYUN_ECS_BASE_IMAGE_ID repo
variable (Ubuntu 24.04 public image id in the region); all other
secrets/vars are shared with the provisioning job.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* ci: fold the query-regression runner rebuild into release-dev-builder-images.yaml
Merge the standalone rebuild workflow into the existing builder-image
release workflow, as one entry point for all builder artifacts:
- push paths extended with .github/runner-scale-sets/query-regression/**
- new 'release_query_regression_runner_image' dispatch input
- a 'changes' job diffs the pushed range (github.event.before..sha,
with an everything-changed fallback for dispatch or unknown bases)
so each expensive rebuild only fires for its own paths:
rust-toolchain.toml gates both, docker/dev-builder/** gates the
dev-builder images, the query-regression runner directory gates the
ECS image rebuild
- the rebuild job itself is unchanged from the standalone workflow
(build-ecs-image.py, then RUNNER_IMAGE_EPOCH commit to main, then
the QUERY_REGRESSION_ECS_IMAGE_ID variable update)
The dev-builder jobs, their ECR/CN/tag-update dependents, and the
runner rebuild now share one workflow; the changes filter preserves
the previous on-push behavior for the dev-builder images while the
runner rebuild keeps its own trigger. The epoch-bump commit only
touches query-regression.yml, which is outside the trigger paths, so
no re-trigger loop.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix(ci): auto-resolve the ECS base image for the runner rebuild
The automated rebuild failed with 'Missing required configuration:
--base-image-id' because the ALIYUN_ECS_BASE_IMAGE_ID repo variable
does not exist yet (it was flagged as a one-time setup item).
Remove the setup dependency instead: build-ecs-image.py now defaults
--base-image-id to the latest public Ubuntu 24.04 x86_64 system image
in the region (DescribeImages with image_owner_alias=system), so no
manual variable is required. The runner Dockerfile pins every tool
version itself, so base-image drift is low-risk; --base-image-id or
the ALIYUN_ECS_BASE_IMAGE_ID variable still pin a specific base image
deterministically, and the workflow only passes the flag when the
variable is set.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* docs(query-regression): clarify what an image rebuild requires
A routine runner-image rebuild needs no manual file updates: the
rebuild job updates QUERY_REGRESSION_ECS_IMAGE_ID and
RUNNER_IMAGE_EPOCH; the toolchain derives from rust-toolchain.toml;
uv, sccache, otelgen, rustup, and the runner base are pinned by
digest/sha/commit in the Dockerfile. Only an apt package revision
bump (mold, protoc, python3) between rebuilds requires bumping the
corresponding Verify pins, and that failure is loud with the observed
version.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix(ci): correct SDK field names in the base-image resolver
DescribeImagesRequest takes 'ostype' (not 'os_type') and the image
items expose 'osname'/'osname_en' (not 'os_name') in the pinned
alibabacloud_ecs20140526 SDK range, so the auto-resolution added in
5a9fd2c769 crashed with a TypeError before describing anything.
Fix the request fields, move the architecture filter server-side, and
paginate (page_size=100 until a short page) instead of relying on a
single default-sized response. Match Ubuntu 24.04 on the localized
osname or the English osname_en.
Verified against the real SDK models (uv run --with
'alibabacloud_ecs20140526>=4.1.0,<6'): a two-page fake client picks
the newest Ubuntu 24.04 via osname_en and rejects 22.04/Windows
decoys.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
* fix(ci): correct the repo-root path in build-ecs-image.py
ASSETS_DIR.parent.parent lands on runner-scale-sets, not the repo
root -- the toml lookup failed with FileNotFoundError. The root is
four levels above ecs-image; express it as an explicit REPO_ROOT
constant (ASSETS_DIR.parents[3]).
Verified every path main() reads against the real checkout layout
(Dockerfile, rust-toolchain.toml, start-runner.sh, the systemd unit,
plus REPO_ROOT sanity against Cargo.toml/.git), re-checked the
user-data toml staging round-trip, and re-ran the base-image
resolver regression test.
Part of #9289.
Signed-off-by: Ning Sun <sunning@greptime.com>
---------
Signed-off-by: Ning Sun <sunning@greptime.com>
Metrics, logs, and traces.
One engine, on your infrastructure.
A columnar database for metrics, logs, and traces on object storage. Apache-2.0 licensed core.
User Guide · API Docs · Roadmap 2026 · Slack
stable for production · canary includes pre-releases · nightly is a weekly snapshot of main
Introduction
GreptimeDB is an open-source observability database. Metrics, logs, and traces run on one columnar engine over object storage and share one table model: tags, timestamp, and fields. When signals carry common identifiers such as service, host, or trace ID, you can correlate them in SQL without moving data between databases.
Ingest through OpenTelemetry, Prometheus Remote Write, Loki Push, or Elasticsearch Bulk. Use SQL across observability data and PromQL for metrics. Migrate ingestion one signal at a time without rebuilding your collectors.
One Query Across Signals
OpenTelemetry ingestion writes spans to opentelemetry_traces and log records to
opentelemetry_logs. Both tables carry trace_id, so correlating them is a join:
-- The slowest failed spans in the last hour,
-- with the log lines emitted inside those same traces.
SELECT
t.service_name,
t.span_name,
t.duration_nano / 1000000 AS duration_ms,
l.timestamp AS log_time,
l.severity_text,
l.body
FROM opentelemetry_traces t
JOIN opentelemetry_logs l ON l.trace_id = t.trace_id
WHERE t.timestamp > now() - INTERVAL '1' HOUR
AND t.span_status_code = 'STATUS_CODE_ERROR'
ORDER BY t.duration_nano DESC
LIMIT 20;
Metrics join the same way, on any tag the tables share, such as service,
host, or pod.
Why You Might Use It
- You run Prometheus plus Loki or Elasticsearch and want one backend instead of three
- You have outgrown Prometheus on cardinality or retention and don't want the Thanos/Mimir operational surface
- You are hitting Loki's query performance limits as log volume grows
- You need long retention on object storage without a separate analytics stack
- You want to query telemetry with SQL, not only a domain query language
- You are storing GenAI or agent telemetry (OTel GenAI conventions) alongside infrastructure signals
Learn more in Why GreptimeDB.
What's Supported
| Ingest | OpenTelemetry (OTLP), Prometheus Remote Write, Loki Push, Elasticsearch Bulk, InfluxDB line protocol, gRPC |
| Query | SQL, PromQL, Jaeger-compatible trace queries, MySQL and PostgreSQL wire protocols |
| Storage | S3, GCS, Azure Blob and S3-compatible endpoints as primary storage, with memory and local-disk caches |
| Built in | Retention policies, downsampling, continuous aggregation, explicit table partitioning, and inverted / skipping / fulltext indexes |
Compute and storage are disaggregated: object storage holds the data, while memory and local-disk caches keep recent and frequently queried data close to compute.
Benchmarks
- Agent RCA Bench: LLM agents doing root cause analysis over GreptimeDB versus Prometheus + Loki + Tempo. 40% fewer wrong diagnoses, 48% fewer input tokens, 45% lower cost (write-up)
- GreptimeDB tops JSONBench's billion-record cold run test
- TSBS Benchmark
- More benchmark reports
Compatibility and Migration
Compatibility is per protocol, and query-side coverage is narrower than ingestion.
| Compatible | Not compatible | |
|---|---|---|
| Prometheus | Remote Write ingestion; PromQL queries | Gaps are listed in PromQL compatibility |
| Loki | Push ingestion; dual-write through Grafana Alloy makes the cutover gradual | LogQL and the rest of the Loki query API |
| Elasticsearch | _bulk ingestion in the open-source core; QueryDSL partially, in Enterprise |
Most other Elasticsearch APIs |
Limitations and Edition Boundary
Cluster deployment, object storage, the Flow engine, and every ingestion protocol listed above are in the Apache-2.0 build. Repartitioning, region migration, and index creation are manual operations there.
Read replicas, workload isolation, and automated repartitioning are GreptimeDB Enterprise features, along with enterprise security and governance. The Enterprise overview has the current list, and pricing has the edition comparison.
Architecture
GreptimeDB can run in two modes:
- Standalone — single binary for development and small deployments.
- Distributed — four components, each independently scalable:
- Frontend — protocol entry (OTel, Prometheus, MySQL/PostgreSQL, gRPC, ingestion APIs for Elasticsearch/InfluxDB/Loki) and the distributed query engine. Stateless, scales horizontally.
- Datanode — region engine with WAL, memtable, SST, cache, compaction, and indexes. Persists data to object storage. Elastic.
- Metasrv — metadata, routing, repartitioning, and security. Backed by a pluggable KV layer (etcd or RDS).
- Flownode (optional) — continuous flow computation (streaming and materialized views).
For deeper coverage, see the architecture doc or DeepWiki.
Try GreptimeDB
For AI agents — paste this prompt into your agent:
Read https://docs.greptime.com/SKILL.md and follow the instructions
to deploy, configure, ingest, and query GreptimeDB.
docker run -p 127.0.0.1:4000-4003:4000-4003 \
-v "$(pwd)/greptimedb_data:/greptimedb_data" \
--name greptime --rm \
greptime/greptimedb:latest standalone start \
--http-addr 0.0.0.0:4000 \
--grpc-bind-addr 0.0.0.0:4001 \
--mysql-addr 0.0.0.0:4002 \
--postgres-addr 0.0.0.0:4003
Dashboard: http://localhost:4000/dashboard
Read more in the full Install Guide.
Troubleshooting:
- Cannot connect to the database? Ensure that ports
4000,4001,4002, and4003are not blocked by a firewall or used by other services. - Failed to start? Check the container logs with
docker logs greptimefor further details.
Getting Started
Build From Source
Prerequisites:
- Rust toolchain — stable, pinned by
rust-toolchain.toml - Protobuf compiler (>= 3.15)
- C/C++ building essentials:
gcc/g++/autoconfand the glibc dev package (libc6-devon Ubuntu,glibc-develon Fedora) - Python toolchain (optional, only for some test scripts)
Build and run:
make # build greptime binary
cargo run -- standalone start # start in standalone mode
Common dev commands:
make fmt # format Rust code
make clippy # lint (fails on warnings)
make test # unit + integration tests (uses cargo-nextest)
make sqlness-test # SQL regression tests
See the Contribution Guidelines for the full developer workflow.
Tools & Extensions
- Kubernetes: GreptimeDB Operator
- Helm Charts: Greptime Helm Charts
- Dashboard: Web UI
- gRPC Ingester: Go, Java, C++, Erlang, Rust, .NET, TypeScript
- Grafana Data Source: GreptimeDB Grafana data source plugin
- Grafana Dashboard: Official Dashboard for monitoring
Project Status
GreptimeDB is generally available, with stable APIs and regular releases. It runs in production at scale — OceanBase Cloud operates 80+ GreptimeDB clusters managing 300 TB of logs, cutting log storage cost by 60%+ after migrating from Grafana Loki. See more in case studies.
Release lines and support windows are in the version reference. For where the project is going, read the v1.0 highlights and the 2026 roadmap.
Community
We invite you to engage and contribute!
If GreptimeDB is useful to you, please star the repo.
License
GreptimeDB is an open-core project. Its core is licensed under the Apache License 2.0.
A small set of peripheral, enterprise-only features are gated behind the
enterprise Cargo feature (not built by default) and are governed by the
separate GreptimeDB Enterprise License. Source files under
that license carry an explicit Enterprise License header.
Commercial Support
Scaling observability on your infrastructure? GreptimeDB Enterprise adds the operational, security, and support layer for production deployments. Contact us for details.
Contributing
- Read our Contribution Guidelines.
- Explore Internal Concepts and DeepWiki.
- Pick up a good first issue and join the #contributors Slack channel.
Acknowledgement
Special thanks to all contributors! See AUTHOR.md.
- Uses Apache Arrow™ (memory model)
- Apache Parquet™ (file storage)
- Apache DataFusion™ (query engine)
- Apache OpenDAL™ (data access abstraction)
All trademarks, logos, and brand names referenced in this README and in the Overview diagram are the property of their respective owners. Their use is for identification purposes only and does not imply endorsement or affiliation.
