feat: add HDFS object storage backend (#8701)

* feat: add HDFS object storage backend

Signed-off-by: Minghan2005 <cambrianocean@gmail.com>

* fix: make HDFS storage operations durable

Gate the native HDFS backend behind an explicit feature. Publish writes through same-directory temporary files and atomic HDFS Rename2 replacement, and provide streaming copy fallback for COPY_REGION. Add regression coverage for interrupted writes and the region-copy path.

Signed-off-by: Minghan2005 <cambrianocean@gmail.com>

* ci: run HDFS object store tests

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* docs: note HDFS temporary file cleanup follow-up

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* feat: enable HDFS object storage by default

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* docs: remove redundant HDFS build feature notes

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: Minghan2005 <cambrianocean@gmail.com>
Signed-off-by: jeremyhi <fengjiachun@gmail.com>
Co-authored-by: Minghan2005 <cambrianocean@gmail.com>
Co-authored-by: jeremyhi <fengjiachun@gmail.com>
This commit is contained in:
Minghan Jiang
2026-09-28 08:39:48 +00:00
committed by GitHub
co-authored by Minghan2005 jeremyhi
parent ffdd6d09a6
commit 75bd8e9ce6
17 changed files with 1090 additions and 47 deletions
+20 -2
View File
@@ -300,6 +300,13 @@ overwrite_entry_start_id = false
# credential = "base64-credential"
# endpoint = "https://storage.googleapis.com"
# Example of using HDFS as the storage with the native Rust client.
# [storage]
# type = "Hdfs"
# root = "/greptimedb"
# name_node = "hdfs://127.0.0.1:9000"
# options = { "dfs.client.block.write.replace-datanode-on-failure.enable" = "true" }
## The query engine options.
[query]
## Parallelism of the query engine.
@@ -348,6 +355,7 @@ data_home = "./greptimedb_data"
## - `Gcs`: the data is stored in the Google Cloud Storage.
## - `Azblob`: the data is stored in the Azure Blob Storage.
## - `Oss`: the data is stored in the Aliyun OSS.
## - `Hdfs`: the data is stored in the Hadoop Distributed File System.
type = "File"
## The S3 bucket name.
@@ -355,11 +363,21 @@ type = "File"
## @toml2docs:none-default
bucket = "greptimedb"
## The S3 data will be stored in the specified prefix, for example, `s3://${bucket}/${root}`.
## **It's only used when the storage type is `S3`, `Oss` and `Azblob`**.
## The directory or object prefix under which data is stored.
## **It's only used when the storage type is `S3`, `Oss`, `Gcs`, `Azblob` and `Hdfs`**.
## @toml2docs:none-default
root = "greptimedb"
## The HDFS NameNode URI, for example, `hdfs://127.0.0.1:9000`.
## **It's only used when the storage type is `Hdfs`**.
## @toml2docs:none-default
name_node = "hdfs://127.0.0.1:9000"
## Additional options passed to the native HDFS client.
## **It's only used when the storage type is `Hdfs`**.
## @toml2docs:none-default
options = {}
## The access key id of the aws account.
## It's **highly recommended** to use AWS IAM roles instead of hardcoding the access key id and secret key.
## **It's only used when the storage type is `S3` and `Oss`**.