jeremyhi 194bc2fb3c feat(log-store): add the object store WAL durable write path (#9320)
* feat(log-store): add the object store WAL durable write path

ObjectStoreLogStore can now write. append_batch admits entries into one
open batch and assigns object-sequence-major entry ids at admission. The
batch is sealed by size, by the flush interval, or before a region would
run past the position range, and sealed batches are uploaded with at most
four conditional creates in flight, started in sequence order. Created
objects are indexed in sequence order, and an append is acknowledged only
once its object is durable and indexed.

A transient create failure rolls back the failed batch and every later
batch unless a later object is already durable, in which case the store
poisons itself with a history-gap error. A conflicting object, an
encoding or catalog error, or taking the last representable sequence
poisons the store.

obsolete now goes through the actor and raises the sequence floor
together with the watermark, refusing with a retryable error while the
next sequence is not settled. stop drops the open batch and the batches
whose create has not started, and lets creates in flight finish.

Only the durable acknowledgement mode exists. A testing feature exposes
hooks to wait for admissions, seal the open batch, and hold or fail
creates.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test(log-store): synchronize object store WAL write tests with the actor

Tests that assert nothing happened round-trip a command through the actor
instead of yielding the test task, the conflict test waits for the object
it expects, and the obsolete-behind-stop test holds the command channel
itself so the unanswered command is deterministic.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): bound admitted WAL appends and never reuse a failed sequence

A create that reports an error may still have written its object, so the
sequences of failed batches are no longer handed out again: the next
batch keeps the sequence after the last sealed one and retries are
assigned new ids. A retry batched differently can no longer conflict
with that object. As the next sequence never moves back, the sequence
floor of obsolete only waits for an open batch that has handed out ids.

Appends now arrive on their own bounded channel, which the actor stops
reading while MAX_SEALED_BATCHES batches wait to become durable, so a
stalled object store holds callers back instead of growing the backlog,
while stop and obsolete are still handled.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): reserve room for two seals before admitting a WAL append

One admission can seal the open batch before an append that would exhaust
its positions and then the append's own batch, so the actor takes an
append only while two more sealed batches fit under MAX_SEALED_BATCHES.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* feat(log-store): chain object store WAL objects and recover the latest chain

A conditional create can fail with an unknown outcome while its object is
stored, or still lands later, and Mito reuses the row sequences of a
failed append. Replaying such an object next to a later acknowledged one
can let the unacknowledged rows win after a restart.

Format version 2 gives every object header its writer's epoch, a link
to the object it extends (sequence and writer instance) and a header
CRC32, and allows objects without segments. Recovery replays only the
chain ending at the complete object with the largest epoch and sequence:
a link holds when its predecessor is present with the recorded writer
instance, or is missing below every present object. Objects off the
chain are orphans that are never replayed but keep their sequences.

Each open writes an empty object that starts an epoch above every
present object, linked to the recovered tip, before it accepts writes,
so a late object of an earlier instance never ends the chain. A start
object that meets an object of an earlier epoch moves to the next
sequence; one of an equal or later epoch fails the open.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): break the WAL chain at every missing predecessor

Accepting a missing predecessor below every present object lets a late
object that lands below the chain change which links hold. Nothing
collects objects yet, so a missing predecessor now always breaks the
link, and recovery fails when objects are present but none completes a
chain.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): keep the object store WAL format at version 1

The object store WAL has not been enabled anywhere, so no object in the
previous layout exists and the chained header can stay version 1.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* refactor(log-store): link object store WAL objects by epoch

Only one store instance writes under an epoch, since every open starts an
epoch above every present object and a start object that meets the same
or a later epoch fails the open. The epoch therefore identifies the
instance, and the random writer instance id is dropped from the header.
A link now records the sequence and the epoch of the object it extends,
and holds when the predecessor carries that epoch. The header shrinks to
46 bytes, and the store logs its epoch when it opens.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): fail an open whose start object is already present

Without a random writer instance, two opens that recover the same objects
encode byte-identical start objects, and a conditional create treats the
same bytes as its own retry. A start object that is already present
therefore fails the open with a retryable error instead of letting both
opens claim the epoch; the next open counts the object and starts a
later epoch.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): derive the WAL epoch from the claimed start sequence

Two opens can recover different listings when a late object lands across
a sequence gap between them, pick the same largest epoch plus one, and
both create their start objects under different sequences. The epoch of
an instance is now one above the sequence its start object claims, so a
successful create decides the epoch and no two instances share one. It
stays above every epoch recovery listed, and an object that carries an
epoch above the next sequence fails the open.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): never run a held WAL create after the store is dropped

The create test hook ignored the closed hold channel, so a create parked
when the store was dropped could still run. It now returns without
creating. Drop the per-admission bookkeeping of issued entry ids, which
nothing reads, and move the parked I/O documentation to its helper.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test(log-store): cover a WAL start object stored with an unknown outcome

Add a fault-injection test in which the create of the start object
stores the object but reports an error: the open fails without moving
to another sequence, and the next open counts the stored object and
claims a later epoch. Rename the test helper that writes a whole object
from a given header to put_object_with_header, and drop a needless
clone.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test(log-store): make the held-create drop test deterministic

Keep the actor running while the store drops the hold sender, so the
parked create always completes on the closed channel instead of racing
the actor's exit. Drop a comment that restates epoch_of.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-09-24 11:07:29 +00:00
2023-08-10 08:08:37 +00:00
2023-06-25 11:05:46 +08:00
2023-11-09 10:38:12 +00:00
2023-03-28 19:14:29 +08:00

GreptimeDB Logo

Metrics, logs, and traces.
One engine, on your infrastructure.

A columnar database for metrics, logs, and traces on object storage. Apache-2.0 licensed core.

User Guide  ·  API Docs  ·  Roadmap 2026  ·  Slack

Stable Canary Nightly Docker Pulls License

stable for production  ·  canary includes pre-releases  ·  nightly is a weekly snapshot of main

Introduction

GreptimeDB is an open-source observability database. Metrics, logs, and traces run on one columnar engine over object storage and share one table model: tags, timestamp, and fields. When signals carry common identifiers such as service, host, or trace ID, you can correlate them in SQL without moving data between databases.

Ingest through OpenTelemetry, Prometheus Remote Write, Loki Push, or Elasticsearch Bulk. Use SQL across observability data and PromQL for metrics. Migrate ingestion one signal at a time without rebuilding your collectors.

One Query Across Signals

OpenTelemetry ingestion writes spans to opentelemetry_traces and log records to opentelemetry_logs. Both tables carry trace_id, so correlating them is a join:

-- The slowest failed spans in the last hour,
-- with the log lines emitted inside those same traces.
SELECT
    t.service_name,
    t.span_name,
    t.duration_nano / 1000000 AS duration_ms,
    l.timestamp AS log_time,
    l.severity_text,
    l.body
FROM opentelemetry_traces t
JOIN opentelemetry_logs l ON l.trace_id = t.trace_id
WHERE t.timestamp > now() - INTERVAL '1' HOUR
  AND t.span_status_code = 'STATUS_CODE_ERROR'
ORDER BY t.duration_nano DESC
LIMIT 20;

Metrics join the same way, on any tag the tables share, such as service, host, or pod.

Why You Might Use It

  • You run Prometheus plus Loki or Elasticsearch and want one backend instead of three
  • You have outgrown Prometheus on cardinality or retention and don't want the Thanos/Mimir operational surface
  • You are hitting Loki's query performance limits as log volume grows
  • You need long retention on object storage without a separate analytics stack
  • You want to query telemetry with SQL, not only a domain query language
  • You are storing GenAI or agent telemetry (OTel GenAI conventions) alongside infrastructure signals

Learn more in Why GreptimeDB.

What's Supported

Ingest OpenTelemetry (OTLP), Prometheus Remote Write, Loki Push, Elasticsearch Bulk, InfluxDB line protocol, gRPC
Query SQL, PromQL, Jaeger-compatible trace queries, MySQL and PostgreSQL wire protocols
Storage S3, GCS, Azure Blob and S3-compatible endpoints as primary storage, with memory and local-disk caches
Built in Retention policies, downsampling, continuous aggregation, explicit table partitioning, and inverted / skipping / fulltext indexes

Compute and storage are disaggregated: object storage holds the data, while memory and local-disk caches keep recent and frequently queried data close to compute.

GreptimeDB Overview

Benchmarks

Compatibility and Migration

Compatibility is per protocol, and query-side coverage is narrower than ingestion.

Compatible Not compatible
Prometheus Remote Write ingestion; PromQL queries Gaps are listed in PromQL compatibility
Loki Push ingestion; dual-write through Grafana Alloy makes the cutover gradual LogQL and the rest of the Loki query API
Elasticsearch _bulk ingestion in the open-source core; QueryDSL partially, in Enterprise Most other Elasticsearch APIs

Limitations and Edition Boundary

Cluster deployment, object storage, the Flow engine, and every ingestion protocol listed above are in the Apache-2.0 build. Repartitioning, region migration, and index creation are manual operations there.

Read replicas, workload isolation, and automated repartitioning are GreptimeDB Enterprise features, along with enterprise security and governance. The Enterprise overview has the current list, and pricing has the edition comparison.

Architecture

GreptimeDB can run in two modes:

  • Standalone — single binary for development and small deployments.
  • Distributed — four components, each independently scalable:
    • Frontend — protocol entry (OTel, Prometheus, MySQL/PostgreSQL, gRPC, ingestion APIs for Elasticsearch/InfluxDB/Loki) and the distributed query engine. Stateless, scales horizontally.
    • Datanode — region engine with WAL, memtable, SST, cache, compaction, and indexes. Persists data to object storage. Elastic.
    • Metasrv — metadata, routing, repartitioning, and security. Backed by a pluggable KV layer (etcd or RDS).
    • Flownode (optional) — continuous flow computation (streaming and materialized views).

For deeper coverage, see the architecture doc or DeepWiki.

GreptimeDB System Overview

Try GreptimeDB

For AI agents — paste this prompt into your agent:

Read https://docs.greptime.com/SKILL.md and follow the instructions
to deploy, configure, ingest, and query GreptimeDB.
docker run -p 127.0.0.1:4000-4003:4000-4003 \
  -v "$(pwd)/greptimedb_data:/greptimedb_data" \
  --name greptime --rm \
  greptime/greptimedb:latest standalone start \
  --http-addr 0.0.0.0:4000 \
  --grpc-bind-addr 0.0.0.0:4001 \
  --mysql-addr 0.0.0.0:4002 \
  --postgres-addr 0.0.0.0:4003

Dashboard: http://localhost:4000/dashboard

Read more in the full Install Guide.

Troubleshooting:

  • Cannot connect to the database? Ensure that ports 4000, 4001, 4002, and 4003 are not blocked by a firewall or used by other services.
  • Failed to start? Check the container logs with docker logs greptime for further details.

Getting Started

Build From Source

Prerequisites:

  • Rust toolchain — nightly, pinned by rust-toolchain.toml
  • Protobuf compiler (>= 3.15)
  • C/C++ building essentials: gcc / g++ / autoconf and the glibc dev package (libc6-dev on Ubuntu, glibc-devel on Fedora)
  • Python toolchain (optional, only for some test scripts)

Build and run:

make                          # build greptime binary
cargo run -- standalone start # start in standalone mode

Common dev commands:

make fmt            # format Rust code
make clippy         # lint (fails on warnings)
make test           # unit + integration tests (uses cargo-nextest)
make sqlness-test   # SQL regression tests

See the Contribution Guidelines for the full developer workflow.

Tools & Extensions

Project Status

GreptimeDB is generally available, with stable APIs and regular releases. It runs in production at scale — OceanBase Cloud operates 80+ GreptimeDB clusters managing 300 TB of logs, cutting log storage cost by 60%+ after migrating from Grafana Loki. See more in case studies.

Release lines and support windows are in the version reference. For where the project is going, read the v1.0 highlights and the 2026 roadmap.

Community

We invite you to engage and contribute!

If GreptimeDB is useful to you, please star the repo.

Known Users

License

GreptimeDB is an open-core project. Its core is licensed under the Apache License 2.0.

A small set of peripheral, enterprise-only features are gated behind the enterprise Cargo feature (not built by default) and are governed by the separate GreptimeDB Enterprise License. Source files under that license carry an explicit Enterprise License header.

Commercial Support

Scaling observability on your infrastructure? GreptimeDB Enterprise adds the operational, security, and support layer for production deployments. Contact us for details.

Contributing

Integration CI Codecov

Acknowledgement

Special thanks to all contributors! See AUTHOR.md.


All trademarks, logos, and brand names referenced in this README and in the Overview diagram are the property of their respective owners. Their use is for identification purposes only and does not imply endorsement or affiliation.

S
Description
Open-source, cloud-native, unified observability database for metrics, logs and traces, supporting SQL/PromQL/Streaming.
Readme Apache-2.0
1.2 GiB
Languages
Rust 98.2%
Python 1.1%
Shell 0.3%
JavaScript 0.2%