feat: add HDFS object storage backend (#8701)

* feat: add HDFS object storage backend

Signed-off-by: Minghan2005 <cambrianocean@gmail.com>

* fix: make HDFS storage operations durable

Gate the native HDFS backend behind an explicit feature. Publish writes through same-directory temporary files and atomic HDFS Rename2 replacement, and provide streaming copy fallback for COPY_REGION. Add regression coverage for interrupted writes and the region-copy path.

Signed-off-by: Minghan2005 <cambrianocean@gmail.com>

* ci: run HDFS object store tests

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* docs: note HDFS temporary file cleanup follow-up

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* feat: enable HDFS object storage by default

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* docs: remove redundant HDFS build feature notes

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: Minghan2005 <cambrianocean@gmail.com>
Signed-off-by: jeremyhi <fengjiachun@gmail.com>
Co-authored-by: Minghan2005 <cambrianocean@gmail.com>
Co-authored-by: jeremyhi <fengjiachun@gmail.com>
This commit is contained in:
Minghan Jiang
2026-09-28 08:39:48 +00:00
committed by GitHub
co-authored by Minghan2005 jeremyhi
parent ffdd6d09a6
commit 75bd8e9ce6
17 changed files with 1090 additions and 47 deletions
+8 -4
View File
@@ -148,9 +148,11 @@
| `storage` | -- | -- | The data storage options. |
| `storage.data_home` | String | `./greptimedb_data` | The working home directory. |
| `storage.copy_root` | String | `./greptimedb_data/copy` | Root directory for standalone SQL access to local files.<br/>Relative SQL paths are resolved below this directory. Absolute paths are accepted only when<br/>they are inside this directory. Defaults to `<data_home>/copy`.<br/>Distributed deployments always reject SQL access to local files.<br/>Upgrade note: COPY commands and existing external tables that reference paths outside this<br/>directory will fail. Move those files below the copy root, set this option to a dedicated<br/>directory containing them, or migrate the files to object storage before upgrading. |
| `storage.type` | String | `File` | The storage type used to store the data.<br/>- `File`: the data is stored in the local file system.<br/>- `S3`: the data is stored in the S3 object storage.<br/>- `Gcs`: the data is stored in the Google Cloud Storage.<br/>- `Azblob`: the data is stored in the Azure Blob Storage.<br/>- `Oss`: the data is stored in the Aliyun OSS. |
| `storage.type` | String | `File` | The storage type used to store the data.<br/>- `File`: the data is stored in the local file system.<br/>- `S3`: the data is stored in the S3 object storage.<br/>- `Gcs`: the data is stored in the Google Cloud Storage.<br/>- `Azblob`: the data is stored in the Azure Blob Storage.<br/>- `Oss`: the data is stored in the Aliyun OSS.<br/>- `Hdfs`: the data is stored in the Hadoop Distributed File System. |
| `storage.bucket` | String | Unset | The S3 bucket name.<br/>**It's only used when the storage type is `S3`, `Oss` and `Gcs`**. |
| `storage.root` | String | Unset | The S3 data will be stored in the specified prefix, for example, `s3://${bucket}/${root}`.<br/>**It's only used when the storage type is `S3`, `Oss` and `Azblob`**. |
| `storage.root` | String | Unset | The directory or object prefix under which data is stored.<br/>**It's only used when the storage type is `S3`, `Oss`, `Gcs`, `Azblob` and `Hdfs`**. |
| `storage.name_node` | String | Unset | The HDFS NameNode URI, for example, `hdfs://127.0.0.1:9000`.<br/>**It's only used when the storage type is `Hdfs`**. |
| `storage.options` | InlineTable | Unset | Additional options passed to the native HDFS client.<br/>**It's only used when the storage type is `Hdfs`**. |
| `storage.access_key_id` | String | Unset | The access key id of the aws account.<br/>It's **highly recommended** to use AWS IAM roles instead of hardcoding the access key id and secret key.<br/>**It's only used when the storage type is `S3` and `Oss`**. |
| `storage.secret_access_key` | String | Unset | The secret access key of the aws account.<br/>It's **highly recommended** to use AWS IAM roles instead of hardcoding the access key id and secret key.<br/>**It's only used when the storage type is `S3`**. |
| `storage.access_key_secret` | String | Unset | The secret access key of the aliyun account.<br/>**It's only used when the storage type is `Oss`**. |
@@ -609,9 +611,11 @@
| `query.experimental_spill_compression` | String | `uncompressed` | Compression for spilled data files: "uncompressed" (default), "lz4_frame", "zstd".<br/>Ignored unless mode is "custom". |
| `storage` | -- | -- | The data storage options. |
| `storage.data_home` | String | `./greptimedb_data` | The working home directory. |
| `storage.type` | String | `File` | The storage type used to store the data.<br/>- `File`: the data is stored in the local file system.<br/>- `S3`: the data is stored in the S3 object storage.<br/>- `Gcs`: the data is stored in the Google Cloud Storage.<br/>- `Azblob`: the data is stored in the Azure Blob Storage.<br/>- `Oss`: the data is stored in the Aliyun OSS. |
| `storage.type` | String | `File` | The storage type used to store the data.<br/>- `File`: the data is stored in the local file system.<br/>- `S3`: the data is stored in the S3 object storage.<br/>- `Gcs`: the data is stored in the Google Cloud Storage.<br/>- `Azblob`: the data is stored in the Azure Blob Storage.<br/>- `Oss`: the data is stored in the Aliyun OSS.<br/>- `Hdfs`: the data is stored in the Hadoop Distributed File System. |
| `storage.bucket` | String | Unset | The S3 bucket name.<br/>**It's only used when the storage type is `S3`, `Oss` and `Gcs`**. |
| `storage.root` | String | Unset | The S3 data will be stored in the specified prefix, for example, `s3://${bucket}/${root}`.<br/>**It's only used when the storage type is `S3`, `Oss` and `Azblob`**. |
| `storage.root` | String | Unset | The directory or object prefix under which data is stored.<br/>**It's only used when the storage type is `S3`, `Oss`, `Gcs`, `Azblob` and `Hdfs`**. |
| `storage.name_node` | String | Unset | The HDFS NameNode URI, for example, `hdfs://127.0.0.1:9000`.<br/>**It's only used when the storage type is `Hdfs`**. |
| `storage.options` | InlineTable | Unset | Additional options passed to the native HDFS client.<br/>**It's only used when the storage type is `Hdfs`**. |
| `storage.access_key_id` | String | Unset | The access key id of the aws account.<br/>It's **highly recommended** to use AWS IAM roles instead of hardcoding the access key id and secret key.<br/>**It's only used when the storage type is `S3` and `Oss`**. |
| `storage.secret_access_key` | String | Unset | The secret access key of the aws account.<br/>It's **highly recommended** to use AWS IAM roles instead of hardcoding the access key id and secret key.<br/>**It's only used when the storage type is `S3`**. |
| `storage.access_key_secret` | String | Unset | The secret access key of the aliyun account.<br/>**It's only used when the storage type is `Oss`**. |
+20 -2
View File
@@ -300,6 +300,13 @@ overwrite_entry_start_id = false
# credential = "base64-credential"
# endpoint = "https://storage.googleapis.com"
# Example of using HDFS as the storage with the native Rust client.
# [storage]
# type = "Hdfs"
# root = "/greptimedb"
# name_node = "hdfs://127.0.0.1:9000"
# options = { "dfs.client.block.write.replace-datanode-on-failure.enable" = "true" }
## The query engine options.
[query]
## Parallelism of the query engine.
@@ -348,6 +355,7 @@ data_home = "./greptimedb_data"
## - `Gcs`: the data is stored in the Google Cloud Storage.
## - `Azblob`: the data is stored in the Azure Blob Storage.
## - `Oss`: the data is stored in the Aliyun OSS.
## - `Hdfs`: the data is stored in the Hadoop Distributed File System.
type = "File"
## The S3 bucket name.
@@ -355,11 +363,21 @@ type = "File"
## @toml2docs:none-default
bucket = "greptimedb"
## The S3 data will be stored in the specified prefix, for example, `s3://${bucket}/${root}`.
## **It's only used when the storage type is `S3`, `Oss` and `Azblob`**.
## The directory or object prefix under which data is stored.
## **It's only used when the storage type is `S3`, `Oss`, `Gcs`, `Azblob` and `Hdfs`**.
## @toml2docs:none-default
root = "greptimedb"
## The HDFS NameNode URI, for example, `hdfs://127.0.0.1:9000`.
## **It's only used when the storage type is `Hdfs`**.
## @toml2docs:none-default
name_node = "hdfs://127.0.0.1:9000"
## Additional options passed to the native HDFS client.
## **It's only used when the storage type is `Hdfs`**.
## @toml2docs:none-default
options = {}
## The access key id of the aws account.
## It's **highly recommended** to use AWS IAM roles instead of hardcoding the access key id and secret key.
## **It's only used when the storage type is `S3` and `Oss`**.
+20 -2
View File
@@ -509,6 +509,13 @@ max_running_procedures = 128
# credential = "base64-credential"
# endpoint = "https://storage.googleapis.com"
# Example of using HDFS as the storage with the native Rust client.
# [storage]
# type = "Hdfs"
# root = "/greptimedb"
# name_node = "hdfs://127.0.0.1:9000"
# options = { "dfs.client.block.write.replace-datanode-on-failure.enable" = "true" }
## The query engine options.
[query]
## Parallelism of the query engine.
@@ -566,6 +573,7 @@ data_home = "./greptimedb_data"
## - `Gcs`: the data is stored in the Google Cloud Storage.
## - `Azblob`: the data is stored in the Azure Blob Storage.
## - `Oss`: the data is stored in the Aliyun OSS.
## - `Hdfs`: the data is stored in the Hadoop Distributed File System.
type = "File"
## The S3 bucket name.
@@ -573,11 +581,21 @@ type = "File"
## @toml2docs:none-default
bucket = "greptimedb"
## The S3 data will be stored in the specified prefix, for example, `s3://${bucket}/${root}`.
## **It's only used when the storage type is `S3`, `Oss` and `Azblob`**.
## The directory or object prefix under which data is stored.
## **It's only used when the storage type is `S3`, `Oss`, `Gcs`, `Azblob` and `Hdfs`**.
## @toml2docs:none-default
root = "greptimedb"
## The HDFS NameNode URI, for example, `hdfs://127.0.0.1:9000`.
## **It's only used when the storage type is `Hdfs`**.
## @toml2docs:none-default
name_node = "hdfs://127.0.0.1:9000"
## Additional options passed to the native HDFS client.
## **It's only used when the storage type is `Hdfs`**.
## @toml2docs:none-default
options = {}
## The access key id of the aws account.
## It's **highly recommended** to use AWS IAM roles instead of hardcoding the access key id and secret key.
## **It's only used when the storage type is `S3` and `Oss`**.