Ning Sun c65a4e545e feat: make the series index resilient to time index unit widening (#8996)
* feat: record time unit per series index file

The series index stores __series_min_ts/__series_max_ts as raw i64 in
the time index unit at write time, but the searcher built its range
predicates from the region's current unit, so files written before a
time index unit widening would be compared in the wrong unit.

Record the unit in the min/max ts fields' Arrow metadata when writing
and build the per-file time predicates from it when searching, so each
file is interpreted in the unit it was written with. The writer now
also requires a timestamp time index.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: address review comments on series index time units

- split schema validation (validate_index_schema) from unit extraction
  (index_time_unit), distinguishing missing vs unsupported unit metadata
  in the errors instead of one misleading 'missing a valid metadata'
- encode the recorded unit with an explicit exhaustive match rather than
  Debug formatting, so the on-disk encoding is reviewed next to its parser
- extract time_index_unit to drop the unwrap in series_index_schema and
  the duplicated timestamp-time-index ensure in validate_metadata
- test one searcher reading files with different recorded units, and the
  rejection of missing, unknown and mismatched units

Signed-off-by: Ning Sun <sunning@greptime.com>

* refactor: reuse TimeUnit's Display form for the recorded unit string

common_time::timestamp::TimeUnit already implements Display with the
exact strings the series index records ("Second"/"Millisecond"/
"Microsecond"/"Nanosecond"), so drop the local time_unit_as_str
mapping and use it; the parse side stays local since no FromStr
counterpart exists anywhere yet.

Signed-off-by: Ning Sun <sunning@greptime.com>

* feat: parse TimeUnit from its Display form in common-time

Add FromStr for common_time::timestamp::TimeUnit, accepting exactly the
Display form ("Second"/"Millisecond"/"Microsecond"/"Nanosecond") and
failing with a new UnsupportedTimeUnit error (InvalidArguments). The
series index now records and parses the unit with the common codec,
dropping its local parse_time_unit.

Signed-off-by: Ning Sun <sunning@greptime.com>

* feat: parse TimeUnit case-insensitively

Lowercase the input before matching so "millisecond" and "MILLISECOND"
parse like "Millisecond"; the error still reports the original string.

Signed-off-by: Ning Sun <sunning@greptime.com>

* feat: store series index min/max timestamps as native Timestamp columns

Replace the Int64 min/max columns plus 'time_unit' field metadata with
native Timestamp(unit) columns, so the unit rides on the datatype and
each file is interpreted in the unit it was written with naturally.

- series_index_schema types the columns from the time index unit; the
  writer reinterprets the raw i64 series bounds in that type (arrow's
  Int64->Timestamp cast reinterprets, it does not rescale)
- the searcher reads the unit from each file's column datatype, builds
  Timestamp-typed predicates via datatypes' timestamp_to_scalar_value,
  and reinterprets parquet INT64 statistics in the column type so
  row-group pruning compares like-typed values
- index files whose min/max columns are not Timestamp (written before
  this change) are rejected
- the pruning test now asserts time-range predicates prune row groups,
  not just tag predicates

Signed-off-by: Ning Sun <sunning@greptime.com>

* refactor: drop the TimeUnit string codec from common-time

With the unit carried by the Timestamp datatype, the FromStr impl and
UnsupportedTimeUnit error added for the field-metadata approach have no
consumer; remove them.

Signed-off-by: Ning Sun <sunning@greptime.com>

* refactor: fold index file validation into a single pass

The series index format is unreleased and unwired, so no compatibility
classes are needed: validate_index_schema checks all columns and
returns the min/max columns' unit directly, replacing the separate
index_time_unit extraction.

Signed-off-by: Ning Sun <sunning@greptime.com>

* refactor: carry timestamps through SeriesIndexRow

SeriesIndexRow and the aggregation path now hold common_time::Timestamp
instead of raw i64s: timestamp_values interprets the input column in the
writer's unit (rejecting a timestamp array whose unit differs, instead
of silently reinterpreting it), and rows_to_batch builds the native
Timestamp columns directly from the rows' units without an Int64 round
trip through arrow cast.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: require a timestamp time index column in series index input

The writer already refuses non-timestamp time indexes on the metadata
side, and its input batches always carry the region's ts column as a
timestamp array, so accepting plain Int64 columns only left a silent
unit-interpretation hole; reject them instead.

Signed-off-by: Ning Sun <sunning@greptime.com>

* refactor: rescale series index input timestamps into the file unit

The input array's unit is self-describing, so converting with
Timestamp::convert_to cannot mislabel values; a mismatch no longer
needs to be an error. Only a value that overflows the file's unit
fails the write. This also makes the writer ready to aggregate
old-unit batches after a time index widening.

Signed-off-by: Ning Sun <sunning@greptime.com>

* refactor: drop redundant unit checks in series index writer

The alter path flushes memtables before widening the region's time
index unit, so a writer never receives batches in the region's
previous unit. Reject a unit mismatch at the input boundary instead
of rescaling per value, and build the index batch in the writer's
recorded unit instead of re-deriving it from the schema and
re-checking every row.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: address review comments

---------

Signed-off-by: Ning Sun <sunning@greptime.com>
2026-09-09 08:58:00 +00:00
2023-08-10 08:08:37 +00:00
2023-06-25 11:05:46 +08:00
2023-11-09 10:38:12 +00:00
2023-03-28 19:14:29 +08:00

GreptimeDB Logo

Metrics, logs, and traces.
One engine, on your infrastructure.

A columnar database for metrics, logs, and traces on object storage. Apache-2.0 licensed core.

User Guide | API Docs | Roadmap 2026

Stable Canary Nightly

stable for production  ·  canary includes pre-releases  ·  nightly is a weekly snapshot of main

Docker Pulls GitHub Actions Codecov License
Slack Twitter LinkedIn

Introduction

GreptimeDB is an open-source observability database. Metrics, logs, and traces run on one columnar engine over object storage and share one table model: tags, timestamp, and fields. When signals carry common identifiers such as service, host, or trace ID, you can correlate them in SQL without moving data between databases.

Ingest through OpenTelemetry, Prometheus Remote Write, Loki Push, or Elasticsearch Bulk. Use SQL across observability data and PromQL for metrics. Migrate ingestion one signal at a time without rebuilding your collectors.

Why You Might Use It

  • You run Prometheus plus Loki or Elasticsearch and want one backend instead of three
  • You have outgrown Prometheus on cardinality or retention and don't want the Thanos/Mimir operational surface
  • You need long retention on object storage without a separate analytics stack
  • You want to query telemetry with SQL, not only a domain query language
  • You are storing GenAI or agent telemetry (OTel GenAI conventions) alongside infrastructure signals
  • You need the same engine and semantics on resource-constrained devices

Learn more in Why GreptimeDB.

Overview

A quick overview of what GreptimeDB ingests, how it connects to other systems, and what its distributed engine lets you do.

GreptimeDB Overview

What's Supported

Ingest OpenTelemetry (OTLP), Prometheus Remote Write, Loki Push, Elasticsearch Bulk, InfluxDB line protocol, gRPC
Query SQL, PromQL, Jaeger-compatible trace queries, MySQL and PostgreSQL wire protocols
Storage S3, GCS, Azure Blob and S3-compatible endpoints as primary storage, with memory and local-disk caches
Built in Retention policies, downsampling, continuous aggregation, explicit table partitioning, and inverted / skipping / fulltext indexes

Compute and storage are disaggregated: object storage holds the data, while memory and local-disk caches keep recent and frequently queried data close to compute.

Compatibility and Migration

Compatibility is per protocol, and query-side coverage is narrower than ingestion.

Compatible Not compatible
Prometheus Remote Write ingestion; PromQL queries Gaps are listed in PromQL compatibility
Loki Push ingestion; dual-write through Grafana Alloy makes the cutover gradual LogQL and the rest of the Loki query API
Elasticsearch _bulk ingestion in the open-source core; QueryDSL partially, in Enterprise Most other Elasticsearch APIs

Benchmarks:

Limitations and Edition Boundary

Cluster deployment, object storage, the Flow engine, and every ingestion protocol listed above are in the Apache-2.0 build. Repartitioning, region migration, and index creation are manual operations there.

Read replicas, workload isolation, and automated repartitioning are GreptimeDB Enterprise features, along with enterprise security and governance. The Enterprise overview has the current list, and pricing has the edition comparison.

Architecture

GreptimeDB can run in two modes:

  • Standalone — single binary for development and small deployments.
  • Distributed — four components, each independently scalable:
    • Frontend — protocol entry (OTel, Prometheus, MySQL/PostgreSQL, gRPC, ingestion APIs for Elasticsearch/InfluxDB/Loki) and the distributed query engine. Stateless, scales horizontally.
    • Datanode — region engine with WAL, memtable, SST, cache, compaction, and indexes. Persists data to object storage. Elastic.
    • Metasrv — metadata, routing, repartitioning, and security. Backed by a pluggable KV layer (etcd or RDS).
    • Flownode (optional) — continuous flow computation (streaming and materialized views).

For deeper coverage, see the architecture doc or DeepWiki.

GreptimeDB System Overview

Try GreptimeDB

For AI agents — paste this prompt into your agent:

Read https://docs.greptime.com/SKILL.md and follow the instructions
to deploy, configure, ingest, and query GreptimeDB.
docker run -p 127.0.0.1:4000-4003:4000-4003 \
  -v "$(pwd)/greptimedb_data:/greptimedb_data" \
  --name greptime --rm \
  greptime/greptimedb:latest standalone start \
  --http-addr 0.0.0.0:4000 \
  --grpc-bind-addr 0.0.0.0:4001 \
  --mysql-addr 0.0.0.0:4002 \
  --postgres-addr 0.0.0.0:4003

Dashboard: http://localhost:4000/dashboard

Read more in the full Install Guide.

Troubleshooting:

  • Cannot connect to the database? Ensure that ports 4000, 4001, 4002, and 4003 are not blocked by a firewall or used by other services.
  • Failed to start? Check the container logs with docker logs greptime for further details.

Getting Started

Build From Source

Prerequisites:

  • Rust toolchain — nightly, pinned by rust-toolchain.toml
  • Protobuf compiler (>= 3.15)
  • C/C++ building essentials: gcc / g++ / autoconf and the glibc dev package (libc6-dev on Ubuntu, glibc-devel on Fedora)
  • Python toolchain (optional, only for some test scripts)

Build and run:

make                          # build greptime binary
cargo run -- standalone start # start in standalone mode

Common dev commands:

make fmt            # format Rust code
make clippy         # lint (fails on warnings)
make test           # unit + integration tests (uses cargo-nextest)
make sqlness-test   # SQL regression tests

See the Contribution Guidelines for the full developer workflow.

Tools & Extensions

Project Status

GreptimeDB is generally available, with stable APIs and regular releases. It runs in production at scale — OceanBase Cloud operates 80+ GreptimeDB clusters managing 300 TB of logs, cutting log storage cost by 60%+ after migrating from Grafana Loki. See more in case studies.

Read the v1.0 highlights and 2026 roadmap, or browse the version reference.

If GreptimeDB is useful to you, please star the repo.

Known Users

Community

We invite you to engage and contribute!

License

GreptimeDB is an open-core project. Its core is licensed under the Apache License 2.0.

A small set of peripheral, enterprise-only features are gated behind the enterprise Cargo feature (not built by default) and are governed by the separate GreptimeDB Enterprise License. Source files under that license carry an explicit Enterprise License header.

Commercial Support

Scaling observability on your infrastructure? GreptimeDB Enterprise adds the operational, security, and support layer for production deployments. Contact us for details.

Contributing

Acknowledgement

Special thanks to all contributors! See AUTHOR.md.


All trademarks, logos, and brand names referenced in this README and in the Overview diagram are the property of their respective owners. Their use is for identification purposes only and does not imply endorsement or affiliation.

S
Description
Open-source, cloud-native, unified observability database for metrics, logs and traces, supporting SQL/PromQL/Streaming.
Readme Apache-2.0
1.1 GiB
Languages
Rust 98.4%
Python 0.8%
Shell 0.4%
JavaScript 0.2%