fix: conn timeout&refactor: better err msg (#5974 )

* fix: conn timeout&refactor: better err msg * chore: clippy * chore: make test work * chore: comment * todo: fix null cast * fix: retry conn&udd_calc * chore: comment * chore: apply suggestion --------- Co-authored-by: dennis zhuang <killme2008@gmail.com>
fix: security update (#5982 )
2025-12-25 23:49:58 +00:00 · 2025-04-25 19:12:30 +00:00 · 2025-04-25 18:11:09 +00:00 · 2025-04-25 17:34:49 +00:00 · 2025-04-24 23:31:26 +00:00 · 2025-04-24 08:11:45 +00:00
180 changed files with 21395 additions and 13187 deletions
--- a/.coderabbit.yaml
+++ b/.coderabbit.yaml
@@ -1,15 +0,0 @@
-# yaml-language-server: $schema=https://coderabbit.ai/integrations/schema.v2.json
-language: "en-US"
-early_access: false
-reviews:
-  profile: "chill"
-  request_changes_workflow: false
-  high_level_summary: true
-  poem: true
-  review_status: true
-  collapse_walkthrough: false
-  auto_review:
-    enabled: false
-    drafts: false
-chat:
-  auto_reply: true
--- a/.github/actions/setup-greptimedb-cluster/with-remote-wal.yaml
+++ b/.github/actions/setup-greptimedb-cluster/with-remote-wal.yaml
@@ -7,7 +7,8 @@ meta:
    provider = "kafka"
    broker_endpoints = ["kafka.kafka-cluster.svc.cluster.local:9092"]
    num_topics = 3
-    auto_prune_topic_records = true
+    auto_prune_interval = "30s"
+    trigger_flush_threshold = 100

    [datanode]
    [datanode.client]
@@ -22,6 +23,7 @@ datanode:
    provider = "kafka"
    broker_endpoints = ["kafka.kafka-cluster.svc.cluster.local:9092"]
    linger = "2ms"
+    overwrite_entry_start_id = true
 frontend:
  configData: |-
    [runtime]
--- a/.github/workflows/grafana.yml
+++ b/.github/workflows/grafana.yml
@@ -21,32 +21,6 @@ jobs:
        run: sudo apt-get install -y jq

      # Make the check.sh script executable
-      - name: Make check.sh executable
-        run: chmod +x grafana/check.sh
-
-      # Run the check.sh script
-      - name: Run check.sh
-        run: ./grafana/check.sh
-
-      # Only run summary.sh for pull_request events (not for merge queues or final pushes)
-      - name: Check if this is a pull request
-        id: check-pr
+      - name: Check grafana dashboards
        run: |
-          if [[ "${{ github.event_name }}" == "pull_request" ]]; then
-            echo "is_pull_request=true" >> $GITHUB_OUTPUT
-          else
-            echo "is_pull_request=false" >> $GITHUB_OUTPUT
-          fi
-
-      # Make the summary.sh script executable
-      - name: Make summary.sh executable
-        if: steps.check-pr.outputs.is_pull_request == 'true'
-        run: chmod +x grafana/summary.sh
-
-      # Run the summary.sh script and add its output to the GitHub Job Summary
-      - name: Run summary.sh and add to Job Summary
-        if: steps.check-pr.outputs.is_pull_request == 'true'
-        run: |
-          SUMMARY=$(./grafana/summary.sh)
-          echo "### Summary of Grafana Panels" >> $GITHUB_STEP_SUMMARY
-          echo "$SUMMARY" >> $GITHUB_STEP_SUMMARY
+          make check-dashboards
--- a/Cargo.lock
+++ b/Cargo.lock
--- a/Cargo.toml
+++ b/Cargo.toml
@@ -112,15 +112,15 @@ clap = { version = "4.4", features = ["derive"] }
 config = "0.13.0"
 crossbeam-utils = "0.8"
 dashmap = "6.1"
-datafusion = { git = "https://github.com/waynexia/arrow-datafusion.git", rev = "5bbedc6704162afb03478f56ffb629405a4e1220" }
-datafusion-common = { git = "https://github.com/waynexia/arrow-datafusion.git", rev = "5bbedc6704162afb03478f56ffb629405a4e1220" }
-datafusion-expr = { git = "https://github.com/waynexia/arrow-datafusion.git", rev = "5bbedc6704162afb03478f56ffb629405a4e1220" }
-datafusion-functions = { git = "https://github.com/waynexia/arrow-datafusion.git", rev = "5bbedc6704162afb03478f56ffb629405a4e1220" }
-datafusion-optimizer = { git = "https://github.com/waynexia/arrow-datafusion.git", rev = "5bbedc6704162afb03478f56ffb629405a4e1220" }
-datafusion-physical-expr = { git = "https://github.com/waynexia/arrow-datafusion.git", rev = "5bbedc6704162afb03478f56ffb629405a4e1220" }
-datafusion-physical-plan = { git = "https://github.com/waynexia/arrow-datafusion.git", rev = "5bbedc6704162afb03478f56ffb629405a4e1220" }
-datafusion-sql = { git = "https://github.com/waynexia/arrow-datafusion.git", rev = "5bbedc6704162afb03478f56ffb629405a4e1220" }
-datafusion-substrait = { git = "https://github.com/waynexia/arrow-datafusion.git", rev = "5bbedc6704162afb03478f56ffb629405a4e1220" }
+datafusion = { git = "https://github.com/waynexia/arrow-datafusion.git", rev = "e104c7cf62b11dd5fe41461b82514978234326b4" }
+datafusion-common = { git = "https://github.com/waynexia/arrow-datafusion.git", rev = "e104c7cf62b11dd5fe41461b82514978234326b4" }
+datafusion-expr = { git = "https://github.com/waynexia/arrow-datafusion.git", rev = "e104c7cf62b11dd5fe41461b82514978234326b4" }
+datafusion-functions = { git = "https://github.com/waynexia/arrow-datafusion.git", rev = "e104c7cf62b11dd5fe41461b82514978234326b4" }
+datafusion-optimizer = { git = "https://github.com/waynexia/arrow-datafusion.git", rev = "e104c7cf62b11dd5fe41461b82514978234326b4" }
+datafusion-physical-expr = { git = "https://github.com/waynexia/arrow-datafusion.git", rev = "e104c7cf62b11dd5fe41461b82514978234326b4" }
+datafusion-physical-plan = { git = "https://github.com/waynexia/arrow-datafusion.git", rev = "e104c7cf62b11dd5fe41461b82514978234326b4" }
+datafusion-sql = { git = "https://github.com/waynexia/arrow-datafusion.git", rev = "e104c7cf62b11dd5fe41461b82514978234326b4" }
+datafusion-substrait = { git = "https://github.com/waynexia/arrow-datafusion.git", rev = "e104c7cf62b11dd5fe41461b82514978234326b4" }
 deadpool = "0.12"
 deadpool-postgres = "0.14"
 derive_builder = "0.20"
@@ -129,7 +129,7 @@ etcd-client = "0.14"
 fst = "0.4.7"
 futures = "0.3"
 futures-util = "0.3"
-greptime-proto = { git = "https://github.com/GreptimeTeam/greptime-proto.git", rev = "b6d9cffd43c4e6358805a798f17e03e232994b82" }
+greptime-proto = { git = "https://github.com/GreptimeTeam/greptime-proto.git", rev = "e82b0158cd38d4021edb4e4c0ae77f999051e62f" }
 hex = "0.4"
 http = "1"
 humantime = "2.1"
@@ -191,7 +191,7 @@ simd-json = "0.15"
 similar-asserts = "1.6.0"
 smallvec = { version = "1", features = ["serde"] }
 snafu = "0.8"
-sqlparser = { git = "https://github.com/GreptimeTeam/sqlparser-rs.git", rev = "e98e6b322426a9d397a71efef17075966223c089", features = [
+sqlparser = { git = "https://github.com/GreptimeTeam/sqlparser-rs.git", rev = "0cf6c04490d59435ee965edd2078e8855bd8471e", features = [
    "visitor",
    "serde",
 ] } # branch = "v0.54.x"
@@ -269,6 +269,9 @@ metric-engine = { path = "src/metric-engine" }
 mito2 = { path = "src/mito2" }
 object-store = { path = "src/object-store" }
 operator = { path = "src/operator" }
+otel-arrow-rust = { git = "https://github.com/open-telemetry/otel-arrow", rev = "5d551412d2a12e689cde4d84c14ef29e36784e51", features = [
+    "server",
+] }
 partition = { path = "src/partition" }
 pipeline = { path = "src/pipeline" }
 plugins = { path = "src/plugins" }
--- a/10
+++ b/10
@@ -222,6 +222,16 @@ start-cluster: ## Start the greptimedb cluster with etcd by using docker compose
 stop-cluster: ## Stop the greptimedb cluster that created by docker compose.
 	docker compose -f ./docker/docker-compose/cluster-with-etcd.yaml stop

+##@ Grafana
+
+.PHONY: check-dashboards
+check-dashboards: ## Check the Grafana dashboards.
+	@./grafana/scripts/check.sh
+
+.PHONY: dashboards
+dashboards: ## Generate the Grafana dashboards for standalone mode and intermediate dashboards.
+	@./grafana/scripts/gen-dashboards.sh
+
 ##@ Docs
 config-docs: ## Generate configuration documentation from toml files.
 	docker run --rm \
--- a/README.md
+++ b/README.md
@@ -6,7 +6,7 @@
  </picture>
 </p>

-<h2 align="center">Unified & Cost-Effective Observability  Database for Metrics, Logs, and Events</h2>
+<h2 align="center">Real-Time & Cloud-Native Observability  Database<br/>for metrics, logs, and traces</h2>

 <div align="center">
 <h3 align="center">
--- a/config/config.md
+++ b/config/config.md
@@ -319,6 +319,7 @@
 | `selector` | String | `round_robin` | Datanode selector type.<br/>- `round_robin` (default value)<br/>- `lease_based`<br/>- `load_based`<br/>For details, please see "https://docs.greptime.com/developer-guide/metasrv/selector". |
 | `use_memory_store` | Bool | `false` | Store data in memory. |
 | `enable_region_failover` | Bool | `false` | Whether to enable region failover.<br/>This feature is only available on GreptimeDB running on cluster mode and<br/>- Using Remote WAL<br/>- Using shared storage (e.g., s3). |
+| `allow_region_failover_on_local_wal` | Bool | `false` | Whether to allow region failover on local WAL.<br/>**This option is not recommended to be set to true, because it may lead to data loss during failover.** |
 | `node_max_idle_time` | String | `24hours` | Max allowed idle time before removing node info from metasrv memory. |
 | `enable_telemetry` | Bool | `true` | Whether to enable greptimedb telemetry. Enabled by default. |
 | `runtime` | -- | -- | The runtime options. |
--- a/config/metasrv.example.toml
+++ b/config/metasrv.example.toml
@@ -50,6 +50,10 @@ use_memory_store = false
 ## - Using shared storage (e.g., s3).
 enable_region_failover = false

+## Whether to allow region failover on local WAL.
+## **This option is not recommended to be set to true, because it may lead to data loss during failover.**
+allow_region_failover_on_local_wal = false
+
 ## Max allowed idle time before removing node info from metasrv memory.
 node_max_idle_time = "24hours"

--- a/grafana/README.md
+++ b/grafana/README.md
@@ -1,61 +1,89 @@
-Grafana dashboard for GreptimeDB
--------------------------------
+# Grafana dashboards for GreptimeDB

-GreptimeDB's official Grafana dashboard.
+## Overview

-Status notify: we are still working on this config. It's expected to change frequently in the recent days. Please feel free to submit your feedback and/or contribution to this dashboard 🤗
+This repository maintains the Grafana dashboards for GreptimeDB. It has two types of dashboards:

-If you use Helm [chart](https://github.com/GreptimeTeam/helm-charts) to deploy GreptimeDB cluster, you can enable self-monitoring by setting the following values in your Helm chart:
+- `cluster/dashboard.json`: The Grafana dashboard for the GreptimeDB cluster. Read the [dashboard.md](./dashboards/cluster/dashboard.md) for more details.
+- `standalone/dashboard.json`: The Grafana dashboard for the standalone GreptimeDB instance. **It's generated from the `cluster/dashboard.json` by removing the instance filter through the `make dashboards` command**. Read the [dashboard.md](./dashboards/standalone/dashboard.md) for more details.
+
+As the rapid development of GreptimeDB, the metrics may be changed, and please feel free to submit your feedback and/or contribution to this dashboard 🤗
+
+**NOTE**: 
+
+- The Grafana version should be greater than 9.0.
+
+- If you want to modify the dashboards, you only need to modify the `cluster/dashboard.json` and run the `make dashboards` command to generate the `standalone/dashboard.json` and other related files.
+
+To maintain the dashboards easily, we use the [`dac`](https://github.com/zyy17/dac) tool to generate the intermediate dashboards and markdown documents:
+
+- `cluster/dashboard.yaml`: The intermediate dashboard for the GreptimeDB cluster.
+- `standalone/dashboard.yaml`: The intermediate dashboard for the standalone GreptimeDB instance.
+
+## Data Sources
+
+There are two data sources for the dashboards to fetch the metrics:
+
+- **Prometheus**: Expose the metrics of GreptimeDB.
+- **Information Schema**: It is the MySQL port of the current monitored instance. The `overview` dashboard will use this datasource to show the information schema of the current instance.
+
+## Instance Filters
+
+To deploy the dashboards for multiple scenarios (K8s, bare metal, etc.), we prefer to use the `instance` label when filtering instances.
+
+Additionally, we recommend including the `pod` label in the legend to make it easier to identify each instance, even though this field will be empty in bare metal scenarios.
+
+For example, the following query is recommended:
+
+```promql
+sum(process_resident_memory_bytes{instance=~"$datanode"}) by (instance, pod)
+```
+
+And the legend will be like: `[{{instance}}]-[{{ pod }}]`.
+
+## Deployment
+
+### Helm
+
+If you use the Helm [chart](https://github.com/GreptimeTeam/helm-charts) to deploy a GreptimeDB cluster, you can enable self-monitoring by setting the following values in your Helm chart:

 - `monitoring.enabled=true`: Deploys a standalone GreptimeDB instance dedicated to monitoring the cluster;
 - `grafana.enabled=true`: Deploys Grafana and automatically imports the monitoring dashboard;

-The standalone GreptimeDB instance will collect metrics from your cluster and the dashboard will be available in the Grafana UI. For detailed deployment instructions, please refer to our [Kubernetes deployment guide](https://docs.greptime.com/nightly/user-guide/deployments/deploy-on-kubernetes/getting-started).
+The standalone GreptimeDB instance will collect metrics from your cluster, and the dashboard will be available in the Grafana UI. For detailed deployment instructions, please refer to our [Kubernetes deployment guide](https://docs.greptime.com/nightly/user-guide/deployments/deploy-on-kubernetes/getting-started).

-# How to use
+### Self-host Prometheus and import dashboards manually

-## `greptimedb.json`
+1. **Configure Prometheus to scrape the cluster**

-Open Grafana Dashboard page, choose `New` -> `Import`. And upload `greptimedb.json` file.
+   The following is an example configuration(**Please modify it according to your actual situation**):

-## `greptimedb-cluster.json`
+    ```yml
+    # example config
+    # only to indicate how to assign labels to each target
+    # modify yours accordingly
+    scrape_configs:
+      - job_name: metasrv
+        static_configs:
+        - targets: ['<metasrv-ip>:<port>']

-This cluster dashboard provides a comprehensive view of incoming requests, response statuses, and internal activities such as flush and compaction, with a layered structure from frontend to datanode. Designed with a focus on alert functionality, its primary aim is to highlight any anomalies in metrics, allowing users to quickly pinpoint the cause of errors.
+      - job_name: datanode
+        static_configs:
+        - targets: ['<datanode0-ip>:<port>', '<datanode1-ip>:<port>', '<datanode2-ip>:<port>']

-We use Prometheus to scrape off metrics from nodes in GreptimeDB cluster, Grafana to visualize the diagram. Any compatible stack should work too.
+      - job_name: frontend
+        static_configs:
+        - targets: ['<frontend-ip>:<port>']
+    ```

-__Note__: This dashboard is still in an early stage of development. Any issue or advice on improvement is welcomed.
+2. **Configure the data sources in Grafana**

-### Configuration
+   You need to add two data sources in Grafana:

-Please ensure the following configuration before importing the dashboard into Grafana.
+   - Prometheus: It is the Prometheus instance that scrapes the GreptimeDB metrics.
+   - Information Schema: It is the MySQL port of the current monitored instance. The dashboard will use this datasource to show the information schema of the current instance.

-__1. Prometheus scrape config__
+3. **Import the dashboards based on your deployment scenario**

-Configure Prometheus to scrape the cluster.
-
-```yml
-# example config
-# only to indicate how to assign labels to each target
-# modify yours accordingly
-scrape_configs:
-  - job_name: metasrv
-    static_configs:
-    - targets: ['<metasrv-ip>:<port>']
-
-  - job_name: datanode
-    static_configs:
-    - targets: ['<datanode0-ip>:<port>', '<datanode1-ip>:<port>', '<datanode2-ip>:<port>']
-
-  - job_name: frontend
-    static_configs:
-    - targets: ['<frontend-ip>:<port>']
-```
-
-__2. Grafana config__
-
-Create a Prometheus data source in Grafana before using this dashboard. We use `datasource` as a variable in Grafana dashboard so that multiple environments are supported.
-
-### Usage
-
-Use `datasource` or `instance` on the upper-left corner to filter data from certain node.
+   - **Cluster**: Import the `cluster/dashboard.json` dashboard.
+   - **Standalone**: Import the `standalone/dashboard.json` dashboard.
--- a/grafana/check.sh
+++ b/grafana/check.sh
@@ -1,19 +0,0 @@
-#!/usr/bin/env bash
-
-BASEDIR=$(dirname "$0")
-
-# Use jq to check for panels with empty or missing descriptions
-invalid_panels=$(cat $BASEDIR/greptimedb-cluster.json | jq -r '
-  .panels[]
-  | select((.type == "stats" or .type == "timeseries") and (.description == "" or .description == null))
-')
-
-# Check if any invalid panels were found
-if [[ -n "$invalid_panels" ]]; then
-  echo "Error: The following panels have empty or missing descriptions:"
-  echo "$invalid_panels"
-  exit 1
-else
-  echo "All panels with type 'stats' or 'timeseries' have valid descriptions."
-  exit 0
-fi
--- a/grafana/dashboards/cluster/dashboard.json
+++ b/grafana/dashboards/cluster/dashboard.json
--- a/grafana/dashboards/cluster/dashboard.md
+++ b/grafana/dashboards/cluster/dashboard.md
@@ -0,0 +1,97 @@
+# Overview
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| Uptime | `time() - process_start_time_seconds` | `stat` | The start time of GreptimeDB. | `prometheus` | `s` | `__auto` |
+| Version | `SELECT pkg_version FROM information_schema.build_info` | `stat` | GreptimeDB version. | `mysql` | -- | -- |
+| Total Ingestion Rate | `sum(rate(greptime_table_operator_ingest_rows[$__rate_interval]))` | `stat` | Total ingestion rate. | `prometheus` | `rowsps` | `__auto` |
+| Total Storage Size | `select SUM(disk_size) from information_schema.region_statistics;` | `stat` | Total number of data file size. | `mysql` | `decbytes` | -- |
+| Total Rows | `select SUM(region_rows) from information_schema.region_statistics;` | `stat` | Total number of data rows in the cluster. Calculated by sum of rows from each region. | `mysql` | `sishort` | -- |
+| Deployment | `SELECT count(*) as datanode FROM information_schema.cluster_info WHERE peer_type = 'DATANODE';`<br/>`SELECT count(*) as frontend FROM information_schema.cluster_info WHERE peer_type = 'FRONTEND';`<br/>`SELECT count(*) as metasrv FROM information_schema.cluster_info WHERE peer_type = 'METASRV';`<br/>`SELECT count(*) as flownode FROM information_schema.cluster_info WHERE peer_type = 'FLOWNODE';` | `stat` | The deployment topology of GreptimeDB. | `mysql` | -- | -- |
+| Database Resources | `SELECT COUNT(*) as databases FROM information_schema.schemata WHERE schema_name NOT IN ('greptime_private', 'information_schema')`<br/>`SELECT COUNT(*) as tables FROM information_schema.tables WHERE table_schema != 'information_schema'`<br/>`SELECT COUNT(region_id) as regions FROM information_schema.region_peers`<br/>`SELECT COUNT(*) as flows FROM information_schema.flows` | `stat` | The number of the key resources in GreptimeDB. | `mysql` | -- | -- |
+| Data Size | `SELECT SUM(memtable_size) * 0.42825 as WAL FROM information_schema.region_statistics;`<br/>`SELECT SUM(index_size) as index FROM information_schema.region_statistics;`<br/>`SELECT SUM(manifest_size) as manifest FROM information_schema.region_statistics;` | `stat` | The data size of wal/index/manifest in the GreptimeDB. | `mysql` | `decbytes` | -- |
+# Ingestion
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| Total Ingestion Rate | `sum(rate(greptime_table_operator_ingest_rows{instance=~"$frontend"}[$__rate_interval]))` | `timeseries` | Total ingestion rate.<br/><br/>Here we listed 3 primary protocols:<br/><br/>- Prometheus remote write<br/>- Greptime's gRPC API (when using our ingest SDK)<br/>- Log ingestion http API<br/> | `prometheus` | `rowsps` | `ingestion` |
+| Ingestion Rate by Type | `sum(rate(greptime_servers_http_logs_ingestion_counter[$__rate_interval]))`<br/>`sum(rate(greptime_servers_prometheus_remote_write_samples[$__rate_interval]))` | `timeseries` | Total ingestion rate.<br/><br/>Here we listed 3 primary protocols:<br/><br/>- Prometheus remote write<br/>- Greptime's gRPC API (when using our ingest SDK)<br/>- Log ingestion http API<br/> | `prometheus` | `rowsps` | `http-logs` |
+# Queries
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| Total Query Rate | `sum (rate(greptime_servers_mysql_query_elapsed_count{instance=~"$frontend"}[$__rate_interval]))`<br/>`sum (rate(greptime_servers_postgres_query_elapsed_count{instance=~"$frontend"}[$__rate_interval]))`<br/>`sum (rate(greptime_servers_http_promql_elapsed_counte{instance=~"$frontend"}[$__rate_interval]))` | `timeseries` | Total rate of query API calls by protocol. This metric is collected from frontends.<br/><br/>Here we listed 3 main protocols:<br/>- MySQL<br/>- Postgres<br/>- Prometheus API<br/><br/>Note that there are some other minor query APIs like /sql are not included | `prometheus` | `reqps` | `mysql` |
+# Resources
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| Datanode Memory per Instance | `sum(process_resident_memory_bytes{instance=~"$datanode"}) by (instance, pod)` | `timeseries` | Current memory usage by instance | `prometheus` | `decbytes` | `[{{instance}}]-[{{ pod }}]` |
+| Datanode CPU Usage per Instance | `sum(rate(process_cpu_seconds_total{instance=~"$datanode"}[$__rate_interval]) * 1000) by (instance, pod)` | `timeseries` | Current cpu usage by instance | `prometheus` | `none` | `[{{ instance }}]-[{{ pod }}]` |
+| Frontend Memory per Instance | `sum(process_resident_memory_bytes{instance=~"$frontend"}) by (instance, pod)` | `timeseries` | Current memory usage by instance | `prometheus` | `decbytes` | `[{{ instance }}]-[{{ pod }}]` |
+| Frontend CPU Usage per Instance | `sum(rate(process_cpu_seconds_total{instance=~"$frontend"}[$__rate_interval]) * 1000) by (instance, pod)` | `timeseries` | Current cpu usage by instance | `prometheus` | `none` | `[{{ instance }}]-[{{ pod }}]-cpu` |
+| Metasrv Memory per Instance | `sum(process_resident_memory_bytes{instance=~"$metasrv"}) by (instance, pod)` | `timeseries` | Current memory usage by instance | `prometheus` | `decbytes` | `[{{ instance }}]-[{{ pod }}]-resident` |
+| Metasrv CPU Usage per Instance | `sum(rate(process_cpu_seconds_total{instance=~"$metasrv"}[$__rate_interval]) * 1000) by (instance, pod)` | `timeseries` | Current cpu usage by instance | `prometheus` | `none` | `[{{ instance }}]-[{{ pod }}]` |
+| Flownode Memory per Instance | `sum(process_resident_memory_bytes{instance=~"$flownode"}) by (instance, pod)` | `timeseries` | Current memory usage by instance | `prometheus` | `decbytes` | `[{{ instance }}]-[{{ pod }}]` |
+| Flownode CPU Usage per Instance | `sum(rate(process_cpu_seconds_total{instance=~"$flownode"}[$__rate_interval]) * 1000) by (instance, pod)` | `timeseries` | Current cpu usage by instance | `prometheus` | `none` | `[{{ instance }}]-[{{ pod }}]` |
+# Frontend Requests
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| HTTP QPS per Instance | `sum by(instance, pod, path, method, code) (rate(greptime_servers_http_requests_elapsed_count{instance=~"$frontend",path!~"/health\|/metrics"}[$__rate_interval]))` | `timeseries` | HTTP QPS per Instance. | `prometheus` | `reqps` | `[{{instance}}]-[{{pod}}]-[{{path}}]-[{{method}}]-[{{code}}]` |
+| HTTP P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, path, method, code) (rate(greptime_servers_http_requests_elapsed_bucket{instance=~"$frontend",path!~"/health\|/metrics"}[$__rate_interval])))` | `timeseries` | HTTP P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{path}}]-[{{method}}]-[{{code}}]-p99` |
+| gRPC QPS per Instance | `sum by(instance, pod, path, code) (rate(greptime_servers_grpc_requests_elapsed_count{instance=~"$frontend"}[$__rate_interval]))` | `timeseries` | gRPC QPS per Instance. | `prometheus` | `reqps` | `[{{instance}}]-[{{pod}}]-[{{path}}]-[{{code}}]` |
+| gRPC P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, path, code) (rate(greptime_servers_grpc_requests_elapsed_bucket{instance=~"$frontend"}[$__rate_interval])))` | `timeseries` | gRPC P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{path}}]-[{{method}}]-[{{code}}]-p99` |
+| MySQL QPS per Instance | `sum by(pod, instance)(rate(greptime_servers_mysql_query_elapsed_count{instance=~"$frontend"}[$__rate_interval]))` | `timeseries` | MySQL QPS per Instance. | `prometheus` | `reqps` | `[{{instance}}]-[{{pod}}]` |
+| MySQL P99 per Instance | `histogram_quantile(0.99, sum by(pod, instance, le) (rate(greptime_servers_mysql_query_elapsed_bucket{instance=~"$frontend"}[$__rate_interval])))` | `timeseries` | MySQL P99 per Instance. | `prometheus` | `s` | `[{{ instance }}]-[{{ pod }}]-p99` |
+| PostgreSQL QPS per Instance | `sum by(pod, instance)(rate(greptime_servers_postgres_query_elapsed_count{instance=~"$frontend"}[$__rate_interval]))` | `timeseries` | PostgreSQL QPS per Instance. | `prometheus` | `reqps` | `[{{instance}}]-[{{pod}}]` |
+| PostgreSQL P99 per Instance | `histogram_quantile(0.99, sum by(pod,instance,le) (rate(greptime_servers_postgres_query_elapsed_bucket{instance=~"$frontend"}[$__rate_interval])))` | `timeseries` | PostgreSQL P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-p99` |
+# Frontend to Datanode
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| Ingest Rows per Instance | `sum by(instance, pod)(rate(greptime_table_operator_ingest_rows{instance=~"$frontend"}[$__rate_interval]))` | `timeseries` | Ingestion rate by row as in each frontend | `prometheus` | `rowsps` | `[{{instance}}]-[{{pod}}]` |
+| Region Call QPS per Instance | `sum by(instance, pod, request_type) (rate(greptime_grpc_region_request_count{instance=~"$frontend"}[$__rate_interval]))` | `timeseries` | Region Call QPS per Instance. | `prometheus` | `ops` | `[{{instance}}]-[{{pod}}]-[{{request_type}}]` |
+| Region Call P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, request_type) (rate(greptime_grpc_region_request_bucket{instance=~"$frontend"}[$__rate_interval])))` | `timeseries` | Region Call P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{request_type}}]` |
+# Mito Engine
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| Request OPS per Instance | `sum by(instance, pod, type) (rate(greptime_mito_handle_request_elapsed_count{instance=~"$datanode"}[$__rate_interval]))` | `timeseries` | Request QPS per Instance. | `prometheus` | `ops` | `[{{instance}}]-[{{pod}}]-[{{type}}]` |
+| Request P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, type) (rate(greptime_mito_handle_request_elapsed_bucket{instance=~"$datanode"}[$__rate_interval])))` | `timeseries` | Request P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{type}}]` |
+| Write Buffer per Instance | `greptime_mito_write_buffer_bytes{instance=~"$datanode"}` | `timeseries` | Write Buffer per Instance. | `prometheus` | `decbytes` | `[{{instance}}]-[{{pod}}]` |
+| Write Rows per Instance | `sum by (instance, pod) (rate(greptime_mito_write_rows_total{instance=~"$datanode"}[$__rate_interval]))` | `timeseries` | Ingestion size by row counts. | `prometheus` | `rowsps` | `[{{instance}}]-[{{pod}}]` |
+| Flush OPS per Instance | `sum by(instance, pod, reason) (rate(greptime_mito_flush_requests_total{instance=~"$datanode"}[$__rate_interval]))` | `timeseries` | Flush QPS per Instance. | `prometheus` | `ops` | `[{{instance}}]-[{{pod}}]-[{{reason}}]` |
+| Write Stall per Instance | `sum by(instance, pod) (greptime_mito_write_stall_total{instance=~"$datanode"})` | `timeseries` | Write Stall per Instance. | `prometheus` | -- | `[{{instance}}]-[{{pod}}]` |
+| Read Stage OPS per Instance | `sum by(instance, pod) (rate(greptime_mito_read_stage_elapsed_count{instance=~"$datanode", stage="total"}[$__rate_interval]))` | `timeseries` | Read Stage OPS per Instance. | `prometheus` | `ops` | `[{{instance}}]-[{{pod}}]` |
+| Read Stage P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, stage) (rate(greptime_mito_read_stage_elapsed_bucket{instance=~"$datanode"}[$__rate_interval])))` | `timeseries` | Read Stage P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{stage}}]` |
+| Write Stage P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, stage) (rate(greptime_mito_write_stage_elapsed_bucket{instance=~"$datanode"}[$__rate_interval])))` | `timeseries` | Write Stage P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{stage}}]` |
+| Compaction OPS per Instance | `sum by(instance, pod) (rate(greptime_mito_compaction_total_elapsed_count{instance=~"$datanode"}[$__rate_interval]))` | `timeseries` | Compaction OPS per Instance. | `prometheus` | `ops` | `[{{ instance }}]-[{{pod}}]` |
+| Compaction P99 per Instance by Stage | `histogram_quantile(0.99, sum by(instance, pod, le, stage) (rate(greptime_mito_compaction_stage_elapsed_bucket{instance=~"$datanode"}[$__rate_interval])))` | `timeseries` | Compaction latency by stage | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{stage}}]-p99` |
+| Compaction P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le,stage) (rate(greptime_mito_compaction_total_elapsed_bucket{instance=~"$datanode"}[$__rate_interval])))` | `timeseries` | Compaction P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{stage}}]-compaction` |
+| WAL write size | `histogram_quantile(0.95, sum by(le,instance, pod) (rate(raft_engine_write_size_bucket[$__rate_interval])))`<br/>`histogram_quantile(0.99, sum by(le,instance,pod) (rate(raft_engine_write_size_bucket[$__rate_interval])))`<br/>`sum by (instance, pod)(rate(raft_engine_write_size_sum[$__rate_interval]))` | `timeseries` | Write-ahead logs write size as bytes. This chart includes stats of p95 and p99 size by instance, total WAL write rate. | `prometheus` | `bytes` | `[{{instance}}]-[{{pod}}]-req-size-p95` |
+| Cached Bytes per Instance | `greptime_mito_cache_bytes{instance=~"$datanode"}` | `timeseries` | Cached Bytes per Instance. | `prometheus` | `decbytes` | `[{{instance}}]-[{{pod}}]-[{{type}}]` |
+| Inflight Compaction | `greptime_mito_inflight_compaction_count` | `timeseries` | Ongoing compaction task count | `prometheus` | `none` | `[{{instance}}]-[{{pod}}]` |
+| WAL sync duration seconds | `histogram_quantile(0.99, sum by(le, type, node, instance, pod) (rate(raft_engine_sync_log_duration_seconds_bucket[$__rate_interval])))` | `timeseries` | Raft engine (local disk) log store sync latency, p99 | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-p99` |
+| Log Store op duration seconds | `histogram_quantile(0.99, sum by(le,logstore,optype,instance, pod) (rate(greptime_logstore_op_elapsed_bucket[$__rate_interval])))` | `timeseries` | Write-ahead log operations latency at p99 | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{logstore}}]-[{{optype}}]-p99` |
+| Inflight Flush | `greptime_mito_inflight_flush_count` | `timeseries` | Ongoing flush task count | `prometheus` | `none` | `[{{instance}}]-[{{pod}}]` |
+# OpenDAL
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| QPS per Instance | `sum by(instance, pod, scheme, operation) (rate(opendal_operation_duration_seconds_count{instance=~"$datanode"}[$__rate_interval]))` | `timeseries` | QPS per Instance. | `prometheus` | `ops` | `[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]` |
+| Read QPS per Instance | `sum by(instance, pod, scheme) (rate(opendal_operation_duration_seconds_count{instance=~"$datanode", operation="read"}[$__rate_interval]))` | `timeseries` | Read QPS per Instance. | `prometheus` | `ops` | `[{{instance}}]-[{{pod}}]-[{{scheme}}]` |
+| Read P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, scheme) (rate(opendal_operation_duration_seconds_bucket{instance=~"$datanode",operation="read"}[$__rate_interval])))` | `timeseries` | Read P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-{{scheme}}` |
+| Write QPS per Instance | `sum by(instance, pod, scheme) (rate(opendal_operation_duration_seconds_count{instance=~"$datanode", operation="write"}[$__rate_interval]))` | `timeseries` | Write QPS per Instance. | `prometheus` | `ops` | `[{{instance}}]-[{{pod}}]-{{scheme}}` |
+| Write P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, scheme) (rate(opendal_operation_duration_seconds_bucket{instance=~"$datanode", operation="write"}[$__rate_interval])))` | `timeseries` | Write P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{scheme}}]` |
+| List QPS per Instance | `sum by(instance, pod, scheme) (rate(opendal_operation_duration_seconds_count{instance=~"$datanode", operation="list"}[$__rate_interval]))` | `timeseries` | List QPS per Instance. | `prometheus` | `ops` | `[{{instance}}]-[{{pod}}]-[{{scheme}}]` |
+| List P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, scheme) (rate(opendal_operation_duration_seconds_bucket{instance=~"$datanode", operation="list"}[$__rate_interval])))` | `timeseries` | List P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{scheme}}]` |
+| Other Requests per Instance | `sum by(instance, pod, scheme, operation) (rate(opendal_operation_duration_seconds_count{instance=~"$datanode",operation!~"read\|write\|list\|stat"}[$__rate_interval]))` | `timeseries` | Other Requests per Instance. | `prometheus` | `ops` | `[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]` |
+| Other Request P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, scheme, operation) (rate(opendal_operation_duration_seconds_bucket{instance=~"$datanode", operation!~"read\|write\|list"}[$__rate_interval])))` | `timeseries` | Other Request P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]` |
+| Opendal traffic | `sum by(instance, pod, scheme, operation) (rate(opendal_operation_bytes_sum{instance=~"$datanode"}[$__rate_interval]))` | `timeseries` | Total traffic as in bytes by instance and operation | `prometheus` | `decbytes` | `[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]` |
+| OpenDAL errors per Instance | `sum by(instance, pod, scheme, operation, error) (rate(opendal_operation_errors_total{instance=~"$datanode", error!="NotFound"}[$__rate_interval]))` | `timeseries` | OpenDAL error counts per Instance. | `prometheus` | -- | `[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]-[{{error}}]` |
+# Metasrv
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| Region migration datanode | `greptime_meta_region_migration_stat{datanode_type="src"}`<br/>`greptime_meta_region_migration_stat{datanode_type="desc"}` | `state-timeline` | Counter of region migration by source and destination | `prometheus` | `none` | `from-datanode-{{datanode_id}}` |
+| Region migration error | `greptime_meta_region_migration_error` | `timeseries` | Counter of region migration error | `prometheus` | `none` | `__auto` |
+| Datanode load | `greptime_datanode_load` | `timeseries` | Gauge of load information of each datanode, collected via heartbeat between datanode and metasrv. This information is for metasrv to schedule workloads. | `prometheus` | `none` | `__auto` |
+# Flownode
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| Flow Ingest / Output Rate | `sum by(instance, pod, direction) (rate(greptime_flow_processed_rows[$__rate_interval]))` | `timeseries` | Flow Ingest / Output Rate. | `prometheus` | -- | `[{{pod}}]-[{{instance}}]-[{{direction}}]` |
+| Flow Ingest Latency | `histogram_quantile(0.95, sum(rate(greptime_flow_insert_elapsed_bucket[$__rate_interval])) by (le, instance, pod))`<br/>`histogram_quantile(0.99, sum(rate(greptime_flow_insert_elapsed_bucket[$__rate_interval])) by (le, instance, pod))` | `timeseries` | Flow Ingest Latency. | `prometheus` | -- | `[{{instance}}]-[{{pod}}]-p95` |
+| Flow Operation Latency | `histogram_quantile(0.95, sum(rate(greptime_flow_processing_time_bucket[$__rate_interval])) by (le,instance,pod,type))`<br/>`histogram_quantile(0.99, sum(rate(greptime_flow_processing_time_bucket[$__rate_interval])) by (le,instance,pod,type))` | `timeseries` | Flow Operation Latency. | `prometheus` | -- | `[{{instance}}]-[{{pod}}]-[{{type}}]-p95` |
+| Flow Buffer Size per Instance | `greptime_flow_input_buf_size` | `timeseries` | Flow Buffer Size per Instance. | `prometheus` | -- | `[{{instance}}]-[{{pod}]` |
+| Flow Processing Error per Instance | `sum by(instance,pod,code) (rate(greptime_flow_errors[$__rate_interval]))` | `timeseries` | Flow Processing Error per Instance. | `prometheus` | -- | `[{{instance}}]-[{{pod}}]-[{{code}}]` |
--- a/grafana/dashboards/cluster/dashboard.yaml
+++ b/grafana/dashboards/cluster/dashboard.yaml
@@ -0,0 +1,769 @@
+groups:
+    - title: Overview
+      panels:
+        - title: Uptime
+          type: stat
+          description: The start time of GreptimeDB.
+          unit: s
+          queries:
+            - expr: time() - process_start_time_seconds
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: __auto
+        - title: Version
+          type: stat
+          description: GreptimeDB version.
+          queries:
+            - expr: SELECT pkg_version FROM information_schema.build_info
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+        - title: Total Ingestion Rate
+          type: stat
+          description: Total ingestion rate.
+          unit: rowsps
+          queries:
+            - expr: sum(rate(greptime_table_operator_ingest_rows[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: __auto
+        - title: Total Storage Size
+          type: stat
+          description: Total number of data file size.
+          unit: decbytes
+          queries:
+            - expr: select SUM(disk_size) from information_schema.region_statistics;
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+        - title: Total Rows
+          type: stat
+          description: Total number of data rows in the cluster. Calculated by sum of rows from each region.
+          unit: sishort
+          queries:
+            - expr: select SUM(region_rows) from information_schema.region_statistics;
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+        - title: Deployment
+          type: stat
+          description: The deployment topology of GreptimeDB.
+          queries:
+            - expr: SELECT count(*) as datanode FROM information_schema.cluster_info WHERE peer_type = 'DATANODE';
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+            - expr: SELECT count(*) as frontend FROM information_schema.cluster_info WHERE peer_type = 'FRONTEND';
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+            - expr: SELECT count(*) as metasrv FROM information_schema.cluster_info WHERE peer_type = 'METASRV';
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+            - expr: SELECT count(*) as flownode FROM information_schema.cluster_info WHERE peer_type = 'FLOWNODE';
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+        - title: Database Resources
+          type: stat
+          description: The number of the key resources in GreptimeDB.
+          queries:
+            - expr: SELECT COUNT(*) as databases FROM information_schema.schemata WHERE schema_name NOT IN ('greptime_private', 'information_schema')
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+            - expr: SELECT COUNT(*) as tables FROM information_schema.tables WHERE table_schema != 'information_schema'
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+            - expr: SELECT COUNT(region_id) as regions FROM information_schema.region_peers
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+            - expr: SELECT COUNT(*) as flows FROM information_schema.flows
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+        - title: Data Size
+          type: stat
+          description: The data size of wal/index/manifest in the GreptimeDB.
+          unit: decbytes
+          queries:
+            - expr: SELECT SUM(memtable_size) * 0.42825 as WAL FROM information_schema.region_statistics;
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+            - expr: SELECT SUM(index_size) as index FROM information_schema.region_statistics;
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+            - expr: SELECT SUM(manifest_size) as manifest FROM information_schema.region_statistics;
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+    - title: Ingestion
+      panels:
+        - title: Total Ingestion Rate
+          type: timeseries
+          description: |
+            Total ingestion rate.
+
+            Here we listed 3 primary protocols:
+
+            - Prometheus remote write
+            - Greptime's gRPC API (when using our ingest SDK)
+            - Log ingestion http API
+          unit: rowsps
+          queries:
+            - expr: sum(rate(greptime_table_operator_ingest_rows{instance=~"$frontend"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: ingestion
+        - title: Ingestion Rate by Type
+          type: timeseries
+          description: |
+            Total ingestion rate.
+
+            Here we listed 3 primary protocols:
+
+            - Prometheus remote write
+            - Greptime's gRPC API (when using our ingest SDK)
+            - Log ingestion http API
+          unit: rowsps
+          queries:
+            - expr: sum(rate(greptime_servers_http_logs_ingestion_counter[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: http-logs
+            - expr: sum(rate(greptime_servers_prometheus_remote_write_samples[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: prometheus-remote-write
+    - title: Queries
+      panels:
+        - title: Total Query Rate
+          type: timeseries
+          description: |-
+            Total rate of query API calls by protocol. This metric is collected from frontends.
+
+            Here we listed 3 main protocols:
+            - MySQL
+            - Postgres
+            - Prometheus API
+
+            Note that there are some other minor query APIs like /sql are not included
+          unit: reqps
+          queries:
+            - expr: sum (rate(greptime_servers_mysql_query_elapsed_count{instance=~"$frontend"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: mysql
+            - expr: sum (rate(greptime_servers_postgres_query_elapsed_count{instance=~"$frontend"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: pg
+            - expr: sum (rate(greptime_servers_http_promql_elapsed_counte{instance=~"$frontend"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: promql
+    - title: Resources
+      panels:
+        - title: Datanode Memory per Instance
+          type: timeseries
+          description: Current memory usage by instance
+          unit: decbytes
+          queries:
+            - expr: sum(process_resident_memory_bytes{instance=~"$datanode"}) by (instance, pod)
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{ pod }}]'
+        - title: Datanode CPU Usage per Instance
+          type: timeseries
+          description: Current cpu usage by instance
+          unit: none
+          queries:
+            - expr: sum(rate(process_cpu_seconds_total{instance=~"$datanode"}[$__rate_interval]) * 1000) by (instance, pod)
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{ instance }}]-[{{ pod }}]'
+        - title: Frontend Memory per Instance
+          type: timeseries
+          description: Current memory usage by instance
+          unit: decbytes
+          queries:
+            - expr: sum(process_resident_memory_bytes{instance=~"$frontend"}) by (instance, pod)
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{ instance }}]-[{{ pod }}]'
+        - title: Frontend CPU Usage per Instance
+          type: timeseries
+          description: Current cpu usage by instance
+          unit: none
+          queries:
+            - expr: sum(rate(process_cpu_seconds_total{instance=~"$frontend"}[$__rate_interval]) * 1000) by (instance, pod)
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{ instance }}]-[{{ pod }}]-cpu'
+        - title: Metasrv Memory per Instance
+          type: timeseries
+          description: Current memory usage by instance
+          unit: decbytes
+          queries:
+            - expr: sum(process_resident_memory_bytes{instance=~"$metasrv"}) by (instance, pod)
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{ instance }}]-[{{ pod }}]-resident'
+        - title: Metasrv CPU Usage per Instance
+          type: timeseries
+          description: Current cpu usage by instance
+          unit: none
+          queries:
+            - expr: sum(rate(process_cpu_seconds_total{instance=~"$metasrv"}[$__rate_interval]) * 1000) by (instance, pod)
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{ instance }}]-[{{ pod }}]'
+        - title: Flownode Memory per Instance
+          type: timeseries
+          description: Current memory usage by instance
+          unit: decbytes
+          queries:
+            - expr: sum(process_resident_memory_bytes{instance=~"$flownode"}) by (instance, pod)
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{ instance }}]-[{{ pod }}]'
+        - title: Flownode CPU Usage per Instance
+          type: timeseries
+          description: Current cpu usage by instance
+          unit: none
+          queries:
+            - expr: sum(rate(process_cpu_seconds_total{instance=~"$flownode"}[$__rate_interval]) * 1000) by (instance, pod)
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{ instance }}]-[{{ pod }}]'
+    - title: Frontend Requests
+      panels:
+        - title: HTTP QPS per Instance
+          type: timeseries
+          description: HTTP QPS per Instance.
+          unit: reqps
+          queries:
+            - expr: sum by(instance, pod, path, method, code) (rate(greptime_servers_http_requests_elapsed_count{instance=~"$frontend",path!~"/health|/metrics"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{path}}]-[{{method}}]-[{{code}}]'
+        - title: HTTP P99 per Instance
+          type: timeseries
+          description: HTTP P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, path, method, code) (rate(greptime_servers_http_requests_elapsed_bucket{instance=~"$frontend",path!~"/health|/metrics"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{path}}]-[{{method}}]-[{{code}}]-p99'
+        - title: gRPC QPS per Instance
+          type: timeseries
+          description: gRPC QPS per Instance.
+          unit: reqps
+          queries:
+            - expr: sum by(instance, pod, path, code) (rate(greptime_servers_grpc_requests_elapsed_count{instance=~"$frontend"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{path}}]-[{{code}}]'
+        - title: gRPC P99 per Instance
+          type: timeseries
+          description: gRPC P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, path, code) (rate(greptime_servers_grpc_requests_elapsed_bucket{instance=~"$frontend"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{path}}]-[{{method}}]-[{{code}}]-p99'
+        - title: MySQL QPS per Instance
+          type: timeseries
+          description: MySQL QPS per Instance.
+          unit: reqps
+          queries:
+            - expr: sum by(pod, instance)(rate(greptime_servers_mysql_query_elapsed_count{instance=~"$frontend"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]'
+        - title: MySQL P99 per Instance
+          type: timeseries
+          description: MySQL P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(pod, instance, le) (rate(greptime_servers_mysql_query_elapsed_bucket{instance=~"$frontend"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{ instance }}]-[{{ pod }}]-p99'
+        - title: PostgreSQL QPS per Instance
+          type: timeseries
+          description: PostgreSQL QPS per Instance.
+          unit: reqps
+          queries:
+            - expr: sum by(pod, instance)(rate(greptime_servers_postgres_query_elapsed_count{instance=~"$frontend"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]'
+        - title: PostgreSQL P99 per Instance
+          type: timeseries
+          description: PostgreSQL P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(pod,instance,le) (rate(greptime_servers_postgres_query_elapsed_bucket{instance=~"$frontend"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-p99'
+    - title: Frontend to Datanode
+      panels:
+        - title: Ingest Rows per Instance
+          type: timeseries
+          description: Ingestion rate by row as in each frontend
+          unit: rowsps
+          queries:
+            - expr: sum by(instance, pod)(rate(greptime_table_operator_ingest_rows{instance=~"$frontend"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]'
+        - title: Region Call QPS per Instance
+          type: timeseries
+          description: Region Call QPS per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod, request_type) (rate(greptime_grpc_region_request_count{instance=~"$frontend"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{request_type}}]'
+        - title: Region Call P99 per Instance
+          type: timeseries
+          description: Region Call P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, request_type) (rate(greptime_grpc_region_request_bucket{instance=~"$frontend"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{request_type}}]'
+    - title: Mito Engine
+      panels:
+        - title: Request OPS per Instance
+          type: timeseries
+          description: Request QPS per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod, type) (rate(greptime_mito_handle_request_elapsed_count{instance=~"$datanode"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{type}}]'
+        - title: Request P99 per Instance
+          type: timeseries
+          description: Request P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, type) (rate(greptime_mito_handle_request_elapsed_bucket{instance=~"$datanode"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{type}}]'
+        - title: Write Buffer per Instance
+          type: timeseries
+          description: Write Buffer per Instance.
+          unit: decbytes
+          queries:
+            - expr: greptime_mito_write_buffer_bytes{instance=~"$datanode"}
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]'
+        - title: Write Rows per Instance
+          type: timeseries
+          description: Ingestion size by row counts.
+          unit: rowsps
+          queries:
+            - expr: sum by (instance, pod) (rate(greptime_mito_write_rows_total{instance=~"$datanode"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]'
+        - title: Flush OPS per Instance
+          type: timeseries
+          description: Flush QPS per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod, reason) (rate(greptime_mito_flush_requests_total{instance=~"$datanode"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{reason}}]'
+        - title: Write Stall per Instance
+          type: timeseries
+          description: Write Stall per Instance.
+          queries:
+            - expr: sum by(instance, pod) (greptime_mito_write_stall_total{instance=~"$datanode"})
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]'
+        - title: Read Stage OPS per Instance
+          type: timeseries
+          description: Read Stage OPS per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod) (rate(greptime_mito_read_stage_elapsed_count{instance=~"$datanode", stage="total"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]'
+        - title: Read Stage P99 per Instance
+          type: timeseries
+          description: Read Stage P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, stage) (rate(greptime_mito_read_stage_elapsed_bucket{instance=~"$datanode"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{stage}}]'
+        - title: Write Stage P99 per Instance
+          type: timeseries
+          description: Write Stage P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, stage) (rate(greptime_mito_write_stage_elapsed_bucket{instance=~"$datanode"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{stage}}]'
+        - title: Compaction OPS per Instance
+          type: timeseries
+          description: Compaction OPS per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod) (rate(greptime_mito_compaction_total_elapsed_count{instance=~"$datanode"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{ instance }}]-[{{pod}}]'
+        - title: Compaction P99 per Instance by Stage
+          type: timeseries
+          description: Compaction latency by stage
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, stage) (rate(greptime_mito_compaction_stage_elapsed_bucket{instance=~"$datanode"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{stage}}]-p99'
+        - title: Compaction P99 per Instance
+          type: timeseries
+          description: Compaction P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le,stage) (rate(greptime_mito_compaction_total_elapsed_bucket{instance=~"$datanode"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{stage}}]-compaction'
+        - title: WAL write size
+          type: timeseries
+          description: Write-ahead logs write size as bytes. This chart includes stats of p95 and p99 size by instance, total WAL write rate.
+          unit: bytes
+          queries:
+            - expr: histogram_quantile(0.95, sum by(le,instance, pod) (rate(raft_engine_write_size_bucket[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-req-size-p95'
+            - expr: histogram_quantile(0.99, sum by(le,instance,pod) (rate(raft_engine_write_size_bucket[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-req-size-p99'
+            - expr: sum by (instance, pod)(rate(raft_engine_write_size_sum[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-throughput'
+        - title: Cached Bytes per Instance
+          type: timeseries
+          description: Cached Bytes per Instance.
+          unit: decbytes
+          queries:
+            - expr: greptime_mito_cache_bytes{instance=~"$datanode"}
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{type}}]'
+        - title: Inflight Compaction
+          type: timeseries
+          description: Ongoing compaction task count
+          unit: none
+          queries:
+            - expr: greptime_mito_inflight_compaction_count
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]'
+        - title: WAL sync duration seconds
+          type: timeseries
+          description: Raft engine (local disk) log store sync latency, p99
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(le, type, node, instance, pod) (rate(raft_engine_sync_log_duration_seconds_bucket[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-p99'
+        - title: Log Store op duration seconds
+          type: timeseries
+          description: Write-ahead log operations latency at p99
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(le,logstore,optype,instance, pod) (rate(greptime_logstore_op_elapsed_bucket[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{logstore}}]-[{{optype}}]-p99'
+        - title: Inflight Flush
+          type: timeseries
+          description: Ongoing flush task count
+          unit: none
+          queries:
+            - expr: greptime_mito_inflight_flush_count
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]'
+    - title: OpenDAL
+      panels:
+        - title: QPS per Instance
+          type: timeseries
+          description: QPS per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod, scheme, operation) (rate(opendal_operation_duration_seconds_count{instance=~"$datanode"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]'
+        - title: Read QPS per Instance
+          type: timeseries
+          description: Read QPS per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod, scheme) (rate(opendal_operation_duration_seconds_count{instance=~"$datanode", operation="read"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{scheme}}]'
+        - title: Read P99 per Instance
+          type: timeseries
+          description: Read P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, scheme) (rate(opendal_operation_duration_seconds_bucket{instance=~"$datanode",operation="read"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-{{scheme}}'
+        - title: Write QPS per Instance
+          type: timeseries
+          description: Write QPS per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod, scheme) (rate(opendal_operation_duration_seconds_count{instance=~"$datanode", operation="write"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-{{scheme}}'
+        - title: Write P99 per Instance
+          type: timeseries
+          description: Write P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, scheme) (rate(opendal_operation_duration_seconds_bucket{instance=~"$datanode", operation="write"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{scheme}}]'
+        - title: List QPS per Instance
+          type: timeseries
+          description: List QPS per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod, scheme) (rate(opendal_operation_duration_seconds_count{instance=~"$datanode", operation="list"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{scheme}}]'
+        - title: List P99 per Instance
+          type: timeseries
+          description: List P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, scheme) (rate(opendal_operation_duration_seconds_bucket{instance=~"$datanode", operation="list"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{scheme}}]'
+        - title: Other Requests per Instance
+          type: timeseries
+          description: Other Requests per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod, scheme, operation) (rate(opendal_operation_duration_seconds_count{instance=~"$datanode",operation!~"read|write|list|stat"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]'
+        - title: Other Request P99 per Instance
+          type: timeseries
+          description: Other Request P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, scheme, operation) (rate(opendal_operation_duration_seconds_bucket{instance=~"$datanode", operation!~"read|write|list"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]'
+        - title: Opendal traffic
+          type: timeseries
+          description: Total traffic as in bytes by instance and operation
+          unit: decbytes
+          queries:
+            - expr: sum by(instance, pod, scheme, operation) (rate(opendal_operation_bytes_sum{instance=~"$datanode"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]'
+        - title: OpenDAL errors per Instance
+          type: timeseries
+          description: OpenDAL error counts per Instance.
+          queries:
+            - expr: sum by(instance, pod, scheme, operation, error) (rate(opendal_operation_errors_total{instance=~"$datanode", error!="NotFound"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]-[{{error}}]'
+    - title: Metasrv
+      panels:
+        - title: Region migration datanode
+          type: state-timeline
+          description: Counter of region migration by source and destination
+          unit: none
+          queries:
+            - expr: greptime_meta_region_migration_stat{datanode_type="src"}
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: from-datanode-{{datanode_id}}
+            - expr: greptime_meta_region_migration_stat{datanode_type="desc"}
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: to-datanode-{{datanode_id}}
+        - title: Region migration error
+          type: timeseries
+          description: Counter of region migration error
+          unit: none
+          queries:
+            - expr: greptime_meta_region_migration_error
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: __auto
+        - title: Datanode load
+          type: timeseries
+          description: Gauge of load information of each datanode, collected via heartbeat between datanode and metasrv. This information is for metasrv to schedule workloads.
+          unit: none
+          queries:
+            - expr: greptime_datanode_load
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: __auto
+    - title: Flownode
+      panels:
+        - title: Flow Ingest / Output Rate
+          type: timeseries
+          description: Flow Ingest / Output Rate.
+          queries:
+            - expr: sum by(instance, pod, direction) (rate(greptime_flow_processed_rows[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{pod}}]-[{{instance}}]-[{{direction}}]'
+        - title: Flow Ingest Latency
+          type: timeseries
+          description: Flow Ingest Latency.
+          queries:
+            - expr: histogram_quantile(0.95, sum(rate(greptime_flow_insert_elapsed_bucket[$__rate_interval])) by (le, instance, pod))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-p95'
+            - expr: histogram_quantile(0.99, sum(rate(greptime_flow_insert_elapsed_bucket[$__rate_interval])) by (le, instance, pod))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-p99'
+        - title: Flow Operation Latency
+          type: timeseries
+          description: Flow Operation Latency.
+          queries:
+            - expr: histogram_quantile(0.95, sum(rate(greptime_flow_processing_time_bucket[$__rate_interval])) by (le,instance,pod,type))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{type}}]-p95'
+            - expr: histogram_quantile(0.99, sum(rate(greptime_flow_processing_time_bucket[$__rate_interval])) by (le,instance,pod,type))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{type}}]-p99'
+        - title: Flow Buffer Size per Instance
+          type: timeseries
+          description: Flow Buffer Size per Instance.
+          queries:
+            - expr: greptime_flow_input_buf_size
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}]'
+        - title: Flow Processing Error per Instance
+          type: timeseries
+          description: Flow Processing Error per Instance.
+          queries:
+            - expr: sum by(instance,pod,code) (rate(greptime_flow_errors[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{code}}]'
--- a/grafana/dashboards/standalone/dashboard.json
+++ b/grafana/dashboards/standalone/dashboard.json
--- a/grafana/dashboards/standalone/dashboard.md
+++ b/grafana/dashboards/standalone/dashboard.md
@@ -0,0 +1,97 @@
+# Overview
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| Uptime | `time() - process_start_time_seconds` | `stat` | The start time of GreptimeDB. | `prometheus` | `s` | `__auto` |
+| Version | `SELECT pkg_version FROM information_schema.build_info` | `stat` | GreptimeDB version. | `mysql` | -- | -- |
+| Total Ingestion Rate | `sum(rate(greptime_table_operator_ingest_rows[$__rate_interval]))` | `stat` | Total ingestion rate. | `prometheus` | `rowsps` | `__auto` |
+| Total Storage Size | `select SUM(disk_size) from information_schema.region_statistics;` | `stat` | Total number of data file size. | `mysql` | `decbytes` | -- |
+| Total Rows | `select SUM(region_rows) from information_schema.region_statistics;` | `stat` | Total number of data rows in the cluster. Calculated by sum of rows from each region. | `mysql` | `sishort` | -- |
+| Deployment | `SELECT count(*) as datanode FROM information_schema.cluster_info WHERE peer_type = 'DATANODE';`<br/>`SELECT count(*) as frontend FROM information_schema.cluster_info WHERE peer_type = 'FRONTEND';`<br/>`SELECT count(*) as metasrv FROM information_schema.cluster_info WHERE peer_type = 'METASRV';`<br/>`SELECT count(*) as flownode FROM information_schema.cluster_info WHERE peer_type = 'FLOWNODE';` | `stat` | The deployment topology of GreptimeDB. | `mysql` | -- | -- |
+| Database Resources | `SELECT COUNT(*) as databases FROM information_schema.schemata WHERE schema_name NOT IN ('greptime_private', 'information_schema')`<br/>`SELECT COUNT(*) as tables FROM information_schema.tables WHERE table_schema != 'information_schema'`<br/>`SELECT COUNT(region_id) as regions FROM information_schema.region_peers`<br/>`SELECT COUNT(*) as flows FROM information_schema.flows` | `stat` | The number of the key resources in GreptimeDB. | `mysql` | -- | -- |
+| Data Size | `SELECT SUM(memtable_size) * 0.42825 as WAL FROM information_schema.region_statistics;`<br/>`SELECT SUM(index_size) as index FROM information_schema.region_statistics;`<br/>`SELECT SUM(manifest_size) as manifest FROM information_schema.region_statistics;` | `stat` | The data size of wal/index/manifest in the GreptimeDB. | `mysql` | `decbytes` | -- |
+# Ingestion
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| Total Ingestion Rate | `sum(rate(greptime_table_operator_ingest_rows{}[$__rate_interval]))` | `timeseries` | Total ingestion rate.<br/><br/>Here we listed 3 primary protocols:<br/><br/>- Prometheus remote write<br/>- Greptime's gRPC API (when using our ingest SDK)<br/>- Log ingestion http API<br/> | `prometheus` | `rowsps` | `ingestion` |
+| Ingestion Rate by Type | `sum(rate(greptime_servers_http_logs_ingestion_counter[$__rate_interval]))`<br/>`sum(rate(greptime_servers_prometheus_remote_write_samples[$__rate_interval]))` | `timeseries` | Total ingestion rate.<br/><br/>Here we listed 3 primary protocols:<br/><br/>- Prometheus remote write<br/>- Greptime's gRPC API (when using our ingest SDK)<br/>- Log ingestion http API<br/> | `prometheus` | `rowsps` | `http-logs` |
+# Queries
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| Total Query Rate | `sum (rate(greptime_servers_mysql_query_elapsed_count{}[$__rate_interval]))`<br/>`sum (rate(greptime_servers_postgres_query_elapsed_count{}[$__rate_interval]))`<br/>`sum (rate(greptime_servers_http_promql_elapsed_counte{}[$__rate_interval]))` | `timeseries` | Total rate of query API calls by protocol. This metric is collected from frontends.<br/><br/>Here we listed 3 main protocols:<br/>- MySQL<br/>- Postgres<br/>- Prometheus API<br/><br/>Note that there are some other minor query APIs like /sql are not included | `prometheus` | `reqps` | `mysql` |
+# Resources
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| Datanode Memory per Instance | `sum(process_resident_memory_bytes{}) by (instance, pod)` | `timeseries` | Current memory usage by instance | `prometheus` | `decbytes` | `[{{instance}}]-[{{ pod }}]` |
+| Datanode CPU Usage per Instance | `sum(rate(process_cpu_seconds_total{}[$__rate_interval]) * 1000) by (instance, pod)` | `timeseries` | Current cpu usage by instance | `prometheus` | `none` | `[{{ instance }}]-[{{ pod }}]` |
+| Frontend Memory per Instance | `sum(process_resident_memory_bytes{}) by (instance, pod)` | `timeseries` | Current memory usage by instance | `prometheus` | `decbytes` | `[{{ instance }}]-[{{ pod }}]` |
+| Frontend CPU Usage per Instance | `sum(rate(process_cpu_seconds_total{}[$__rate_interval]) * 1000) by (instance, pod)` | `timeseries` | Current cpu usage by instance | `prometheus` | `none` | `[{{ instance }}]-[{{ pod }}]-cpu` |
+| Metasrv Memory per Instance | `sum(process_resident_memory_bytes{}) by (instance, pod)` | `timeseries` | Current memory usage by instance | `prometheus` | `decbytes` | `[{{ instance }}]-[{{ pod }}]-resident` |
+| Metasrv CPU Usage per Instance | `sum(rate(process_cpu_seconds_total{}[$__rate_interval]) * 1000) by (instance, pod)` | `timeseries` | Current cpu usage by instance | `prometheus` | `none` | `[{{ instance }}]-[{{ pod }}]` |
+| Flownode Memory per Instance | `sum(process_resident_memory_bytes{}) by (instance, pod)` | `timeseries` | Current memory usage by instance | `prometheus` | `decbytes` | `[{{ instance }}]-[{{ pod }}]` |
+| Flownode CPU Usage per Instance | `sum(rate(process_cpu_seconds_total{}[$__rate_interval]) * 1000) by (instance, pod)` | `timeseries` | Current cpu usage by instance | `prometheus` | `none` | `[{{ instance }}]-[{{ pod }}]` |
+# Frontend Requests
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| HTTP QPS per Instance | `sum by(instance, pod, path, method, code) (rate(greptime_servers_http_requests_elapsed_count{path!~"/health\|/metrics"}[$__rate_interval]))` | `timeseries` | HTTP QPS per Instance. | `prometheus` | `reqps` | `[{{instance}}]-[{{pod}}]-[{{path}}]-[{{method}}]-[{{code}}]` |
+| HTTP P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, path, method, code) (rate(greptime_servers_http_requests_elapsed_bucket{path!~"/health\|/metrics"}[$__rate_interval])))` | `timeseries` | HTTP P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{path}}]-[{{method}}]-[{{code}}]-p99` |
+| gRPC QPS per Instance | `sum by(instance, pod, path, code) (rate(greptime_servers_grpc_requests_elapsed_count{}[$__rate_interval]))` | `timeseries` | gRPC QPS per Instance. | `prometheus` | `reqps` | `[{{instance}}]-[{{pod}}]-[{{path}}]-[{{code}}]` |
+| gRPC P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, path, code) (rate(greptime_servers_grpc_requests_elapsed_bucket{}[$__rate_interval])))` | `timeseries` | gRPC P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{path}}]-[{{method}}]-[{{code}}]-p99` |
+| MySQL QPS per Instance | `sum by(pod, instance)(rate(greptime_servers_mysql_query_elapsed_count{}[$__rate_interval]))` | `timeseries` | MySQL QPS per Instance. | `prometheus` | `reqps` | `[{{instance}}]-[{{pod}}]` |
+| MySQL P99 per Instance | `histogram_quantile(0.99, sum by(pod, instance, le) (rate(greptime_servers_mysql_query_elapsed_bucket{}[$__rate_interval])))` | `timeseries` | MySQL P99 per Instance. | `prometheus` | `s` | `[{{ instance }}]-[{{ pod }}]-p99` |
+| PostgreSQL QPS per Instance | `sum by(pod, instance)(rate(greptime_servers_postgres_query_elapsed_count{}[$__rate_interval]))` | `timeseries` | PostgreSQL QPS per Instance. | `prometheus` | `reqps` | `[{{instance}}]-[{{pod}}]` |
+| PostgreSQL P99 per Instance | `histogram_quantile(0.99, sum by(pod,instance,le) (rate(greptime_servers_postgres_query_elapsed_bucket{}[$__rate_interval])))` | `timeseries` | PostgreSQL P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-p99` |
+# Frontend to Datanode
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| Ingest Rows per Instance | `sum by(instance, pod)(rate(greptime_table_operator_ingest_rows{}[$__rate_interval]))` | `timeseries` | Ingestion rate by row as in each frontend | `prometheus` | `rowsps` | `[{{instance}}]-[{{pod}}]` |
+| Region Call QPS per Instance | `sum by(instance, pod, request_type) (rate(greptime_grpc_region_request_count{}[$__rate_interval]))` | `timeseries` | Region Call QPS per Instance. | `prometheus` | `ops` | `[{{instance}}]-[{{pod}}]-[{{request_type}}]` |
+| Region Call P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, request_type) (rate(greptime_grpc_region_request_bucket{}[$__rate_interval])))` | `timeseries` | Region Call P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{request_type}}]` |
+# Mito Engine
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| Request OPS per Instance | `sum by(instance, pod, type) (rate(greptime_mito_handle_request_elapsed_count{}[$__rate_interval]))` | `timeseries` | Request QPS per Instance. | `prometheus` | `ops` | `[{{instance}}]-[{{pod}}]-[{{type}}]` |
+| Request P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, type) (rate(greptime_mito_handle_request_elapsed_bucket{}[$__rate_interval])))` | `timeseries` | Request P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{type}}]` |
+| Write Buffer per Instance | `greptime_mito_write_buffer_bytes{}` | `timeseries` | Write Buffer per Instance. | `prometheus` | `decbytes` | `[{{instance}}]-[{{pod}}]` |
+| Write Rows per Instance | `sum by (instance, pod) (rate(greptime_mito_write_rows_total{}[$__rate_interval]))` | `timeseries` | Ingestion size by row counts. | `prometheus` | `rowsps` | `[{{instance}}]-[{{pod}}]` |
+| Flush OPS per Instance | `sum by(instance, pod, reason) (rate(greptime_mito_flush_requests_total{}[$__rate_interval]))` | `timeseries` | Flush QPS per Instance. | `prometheus` | `ops` | `[{{instance}}]-[{{pod}}]-[{{reason}}]` |
+| Write Stall per Instance | `sum by(instance, pod) (greptime_mito_write_stall_total{})` | `timeseries` | Write Stall per Instance. | `prometheus` | -- | `[{{instance}}]-[{{pod}}]` |
+| Read Stage OPS per Instance | `sum by(instance, pod) (rate(greptime_mito_read_stage_elapsed_count{ stage="total"}[$__rate_interval]))` | `timeseries` | Read Stage OPS per Instance. | `prometheus` | `ops` | `[{{instance}}]-[{{pod}}]` |
+| Read Stage P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, stage) (rate(greptime_mito_read_stage_elapsed_bucket{}[$__rate_interval])))` | `timeseries` | Read Stage P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{stage}}]` |
+| Write Stage P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, stage) (rate(greptime_mito_write_stage_elapsed_bucket{}[$__rate_interval])))` | `timeseries` | Write Stage P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{stage}}]` |
+| Compaction OPS per Instance | `sum by(instance, pod) (rate(greptime_mito_compaction_total_elapsed_count{}[$__rate_interval]))` | `timeseries` | Compaction OPS per Instance. | `prometheus` | `ops` | `[{{ instance }}]-[{{pod}}]` |
+| Compaction P99 per Instance by Stage | `histogram_quantile(0.99, sum by(instance, pod, le, stage) (rate(greptime_mito_compaction_stage_elapsed_bucket{}[$__rate_interval])))` | `timeseries` | Compaction latency by stage | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{stage}}]-p99` |
+| Compaction P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le,stage) (rate(greptime_mito_compaction_total_elapsed_bucket{}[$__rate_interval])))` | `timeseries` | Compaction P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{stage}}]-compaction` |
+| WAL write size | `histogram_quantile(0.95, sum by(le,instance, pod) (rate(raft_engine_write_size_bucket[$__rate_interval])))`<br/>`histogram_quantile(0.99, sum by(le,instance,pod) (rate(raft_engine_write_size_bucket[$__rate_interval])))`<br/>`sum by (instance, pod)(rate(raft_engine_write_size_sum[$__rate_interval]))` | `timeseries` | Write-ahead logs write size as bytes. This chart includes stats of p95 and p99 size by instance, total WAL write rate. | `prometheus` | `bytes` | `[{{instance}}]-[{{pod}}]-req-size-p95` |
+| Cached Bytes per Instance | `greptime_mito_cache_bytes{}` | `timeseries` | Cached Bytes per Instance. | `prometheus` | `decbytes` | `[{{instance}}]-[{{pod}}]-[{{type}}]` |
+| Inflight Compaction | `greptime_mito_inflight_compaction_count` | `timeseries` | Ongoing compaction task count | `prometheus` | `none` | `[{{instance}}]-[{{pod}}]` |
+| WAL sync duration seconds | `histogram_quantile(0.99, sum by(le, type, node, instance, pod) (rate(raft_engine_sync_log_duration_seconds_bucket[$__rate_interval])))` | `timeseries` | Raft engine (local disk) log store sync latency, p99 | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-p99` |
+| Log Store op duration seconds | `histogram_quantile(0.99, sum by(le,logstore,optype,instance, pod) (rate(greptime_logstore_op_elapsed_bucket[$__rate_interval])))` | `timeseries` | Write-ahead log operations latency at p99 | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{logstore}}]-[{{optype}}]-p99` |
+| Inflight Flush | `greptime_mito_inflight_flush_count` | `timeseries` | Ongoing flush task count | `prometheus` | `none` | `[{{instance}}]-[{{pod}}]` |
+# OpenDAL
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| QPS per Instance | `sum by(instance, pod, scheme, operation) (rate(opendal_operation_duration_seconds_count{}[$__rate_interval]))` | `timeseries` | QPS per Instance. | `prometheus` | `ops` | `[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]` |
+| Read QPS per Instance | `sum by(instance, pod, scheme) (rate(opendal_operation_duration_seconds_count{ operation="read"}[$__rate_interval]))` | `timeseries` | Read QPS per Instance. | `prometheus` | `ops` | `[{{instance}}]-[{{pod}}]-[{{scheme}}]` |
+| Read P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, scheme) (rate(opendal_operation_duration_seconds_bucket{operation="read"}[$__rate_interval])))` | `timeseries` | Read P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-{{scheme}}` |
+| Write QPS per Instance | `sum by(instance, pod, scheme) (rate(opendal_operation_duration_seconds_count{ operation="write"}[$__rate_interval]))` | `timeseries` | Write QPS per Instance. | `prometheus` | `ops` | `[{{instance}}]-[{{pod}}]-{{scheme}}` |
+| Write P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, scheme) (rate(opendal_operation_duration_seconds_bucket{ operation="write"}[$__rate_interval])))` | `timeseries` | Write P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{scheme}}]` |
+| List QPS per Instance | `sum by(instance, pod, scheme) (rate(opendal_operation_duration_seconds_count{ operation="list"}[$__rate_interval]))` | `timeseries` | List QPS per Instance. | `prometheus` | `ops` | `[{{instance}}]-[{{pod}}]-[{{scheme}}]` |
+| List P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, scheme) (rate(opendal_operation_duration_seconds_bucket{ operation="list"}[$__rate_interval])))` | `timeseries` | List P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{scheme}}]` |
+| Other Requests per Instance | `sum by(instance, pod, scheme, operation) (rate(opendal_operation_duration_seconds_count{operation!~"read\|write\|list\|stat"}[$__rate_interval]))` | `timeseries` | Other Requests per Instance. | `prometheus` | `ops` | `[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]` |
+| Other Request P99 per Instance | `histogram_quantile(0.99, sum by(instance, pod, le, scheme, operation) (rate(opendal_operation_duration_seconds_bucket{ operation!~"read\|write\|list"}[$__rate_interval])))` | `timeseries` | Other Request P99 per Instance. | `prometheus` | `s` | `[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]` |
+| Opendal traffic | `sum by(instance, pod, scheme, operation) (rate(opendal_operation_bytes_sum{}[$__rate_interval]))` | `timeseries` | Total traffic as in bytes by instance and operation | `prometheus` | `decbytes` | `[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]` |
+| OpenDAL errors per Instance | `sum by(instance, pod, scheme, operation, error) (rate(opendal_operation_errors_total{ error!="NotFound"}[$__rate_interval]))` | `timeseries` | OpenDAL error counts per Instance. | `prometheus` | -- | `[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]-[{{error}}]` |
+# Metasrv
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| Region migration datanode | `greptime_meta_region_migration_stat{datanode_type="src"}`<br/>`greptime_meta_region_migration_stat{datanode_type="desc"}` | `state-timeline` | Counter of region migration by source and destination | `prometheus` | `none` | `from-datanode-{{datanode_id}}` |
+| Region migration error | `greptime_meta_region_migration_error` | `timeseries` | Counter of region migration error | `prometheus` | `none` | `__auto` |
+| Datanode load | `greptime_datanode_load` | `timeseries` | Gauge of load information of each datanode, collected via heartbeat between datanode and metasrv. This information is for metasrv to schedule workloads. | `prometheus` | `none` | `__auto` |
+# Flownode
+| Title | Query | Type | Description | Datasource | Unit | Legend Format |
+| --- | --- | --- | --- | --- | --- | --- |
+| Flow Ingest / Output Rate | `sum by(instance, pod, direction) (rate(greptime_flow_processed_rows[$__rate_interval]))` | `timeseries` | Flow Ingest / Output Rate. | `prometheus` | -- | `[{{pod}}]-[{{instance}}]-[{{direction}}]` |
+| Flow Ingest Latency | `histogram_quantile(0.95, sum(rate(greptime_flow_insert_elapsed_bucket[$__rate_interval])) by (le, instance, pod))`<br/>`histogram_quantile(0.99, sum(rate(greptime_flow_insert_elapsed_bucket[$__rate_interval])) by (le, instance, pod))` | `timeseries` | Flow Ingest Latency. | `prometheus` | -- | `[{{instance}}]-[{{pod}}]-p95` |
+| Flow Operation Latency | `histogram_quantile(0.95, sum(rate(greptime_flow_processing_time_bucket[$__rate_interval])) by (le,instance,pod,type))`<br/>`histogram_quantile(0.99, sum(rate(greptime_flow_processing_time_bucket[$__rate_interval])) by (le,instance,pod,type))` | `timeseries` | Flow Operation Latency. | `prometheus` | -- | `[{{instance}}]-[{{pod}}]-[{{type}}]-p95` |
+| Flow Buffer Size per Instance | `greptime_flow_input_buf_size` | `timeseries` | Flow Buffer Size per Instance. | `prometheus` | -- | `[{{instance}}]-[{{pod}]` |
+| Flow Processing Error per Instance | `sum by(instance,pod,code) (rate(greptime_flow_errors[$__rate_interval]))` | `timeseries` | Flow Processing Error per Instance. | `prometheus` | -- | `[{{instance}}]-[{{pod}}]-[{{code}}]` |
--- a/grafana/dashboards/standalone/dashboard.yaml
+++ b/grafana/dashboards/standalone/dashboard.yaml
@@ -0,0 +1,769 @@
+groups:
+    - title: Overview
+      panels:
+        - title: Uptime
+          type: stat
+          description: The start time of GreptimeDB.
+          unit: s
+          queries:
+            - expr: time() - process_start_time_seconds
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: __auto
+        - title: Version
+          type: stat
+          description: GreptimeDB version.
+          queries:
+            - expr: SELECT pkg_version FROM information_schema.build_info
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+        - title: Total Ingestion Rate
+          type: stat
+          description: Total ingestion rate.
+          unit: rowsps
+          queries:
+            - expr: sum(rate(greptime_table_operator_ingest_rows[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: __auto
+        - title: Total Storage Size
+          type: stat
+          description: Total number of data file size.
+          unit: decbytes
+          queries:
+            - expr: select SUM(disk_size) from information_schema.region_statistics;
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+        - title: Total Rows
+          type: stat
+          description: Total number of data rows in the cluster. Calculated by sum of rows from each region.
+          unit: sishort
+          queries:
+            - expr: select SUM(region_rows) from information_schema.region_statistics;
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+        - title: Deployment
+          type: stat
+          description: The deployment topology of GreptimeDB.
+          queries:
+            - expr: SELECT count(*) as datanode FROM information_schema.cluster_info WHERE peer_type = 'DATANODE';
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+            - expr: SELECT count(*) as frontend FROM information_schema.cluster_info WHERE peer_type = 'FRONTEND';
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+            - expr: SELECT count(*) as metasrv FROM information_schema.cluster_info WHERE peer_type = 'METASRV';
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+            - expr: SELECT count(*) as flownode FROM information_schema.cluster_info WHERE peer_type = 'FLOWNODE';
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+        - title: Database Resources
+          type: stat
+          description: The number of the key resources in GreptimeDB.
+          queries:
+            - expr: SELECT COUNT(*) as databases FROM information_schema.schemata WHERE schema_name NOT IN ('greptime_private', 'information_schema')
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+            - expr: SELECT COUNT(*) as tables FROM information_schema.tables WHERE table_schema != 'information_schema'
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+            - expr: SELECT COUNT(region_id) as regions FROM information_schema.region_peers
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+            - expr: SELECT COUNT(*) as flows FROM information_schema.flows
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+        - title: Data Size
+          type: stat
+          description: The data size of wal/index/manifest in the GreptimeDB.
+          unit: decbytes
+          queries:
+            - expr: SELECT SUM(memtable_size) * 0.42825 as WAL FROM information_schema.region_statistics;
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+            - expr: SELECT SUM(index_size) as index FROM information_schema.region_statistics;
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+            - expr: SELECT SUM(manifest_size) as manifest FROM information_schema.region_statistics;
+              datasource:
+                type: mysql
+                uid: ${information_schema}
+    - title: Ingestion
+      panels:
+        - title: Total Ingestion Rate
+          type: timeseries
+          description: |
+            Total ingestion rate.
+
+            Here we listed 3 primary protocols:
+
+            - Prometheus remote write
+            - Greptime's gRPC API (when using our ingest SDK)
+            - Log ingestion http API
+          unit: rowsps
+          queries:
+            - expr: sum(rate(greptime_table_operator_ingest_rows{}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: ingestion
+        - title: Ingestion Rate by Type
+          type: timeseries
+          description: |
+            Total ingestion rate.
+
+            Here we listed 3 primary protocols:
+
+            - Prometheus remote write
+            - Greptime's gRPC API (when using our ingest SDK)
+            - Log ingestion http API
+          unit: rowsps
+          queries:
+            - expr: sum(rate(greptime_servers_http_logs_ingestion_counter[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: http-logs
+            - expr: sum(rate(greptime_servers_prometheus_remote_write_samples[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: prometheus-remote-write
+    - title: Queries
+      panels:
+        - title: Total Query Rate
+          type: timeseries
+          description: |-
+            Total rate of query API calls by protocol. This metric is collected from frontends.
+
+            Here we listed 3 main protocols:
+            - MySQL
+            - Postgres
+            - Prometheus API
+
+            Note that there are some other minor query APIs like /sql are not included
+          unit: reqps
+          queries:
+            - expr: sum (rate(greptime_servers_mysql_query_elapsed_count{}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: mysql
+            - expr: sum (rate(greptime_servers_postgres_query_elapsed_count{}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: pg
+            - expr: sum (rate(greptime_servers_http_promql_elapsed_counte{}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: promql
+    - title: Resources
+      panels:
+        - title: Datanode Memory per Instance
+          type: timeseries
+          description: Current memory usage by instance
+          unit: decbytes
+          queries:
+            - expr: sum(process_resident_memory_bytes{}) by (instance, pod)
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{ pod }}]'
+        - title: Datanode CPU Usage per Instance
+          type: timeseries
+          description: Current cpu usage by instance
+          unit: none
+          queries:
+            - expr: sum(rate(process_cpu_seconds_total{}[$__rate_interval]) * 1000) by (instance, pod)
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{ instance }}]-[{{ pod }}]'
+        - title: Frontend Memory per Instance
+          type: timeseries
+          description: Current memory usage by instance
+          unit: decbytes
+          queries:
+            - expr: sum(process_resident_memory_bytes{}) by (instance, pod)
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{ instance }}]-[{{ pod }}]'
+        - title: Frontend CPU Usage per Instance
+          type: timeseries
+          description: Current cpu usage by instance
+          unit: none
+          queries:
+            - expr: sum(rate(process_cpu_seconds_total{}[$__rate_interval]) * 1000) by (instance, pod)
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{ instance }}]-[{{ pod }}]-cpu'
+        - title: Metasrv Memory per Instance
+          type: timeseries
+          description: Current memory usage by instance
+          unit: decbytes
+          queries:
+            - expr: sum(process_resident_memory_bytes{}) by (instance, pod)
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{ instance }}]-[{{ pod }}]-resident'
+        - title: Metasrv CPU Usage per Instance
+          type: timeseries
+          description: Current cpu usage by instance
+          unit: none
+          queries:
+            - expr: sum(rate(process_cpu_seconds_total{}[$__rate_interval]) * 1000) by (instance, pod)
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{ instance }}]-[{{ pod }}]'
+        - title: Flownode Memory per Instance
+          type: timeseries
+          description: Current memory usage by instance
+          unit: decbytes
+          queries:
+            - expr: sum(process_resident_memory_bytes{}) by (instance, pod)
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{ instance }}]-[{{ pod }}]'
+        - title: Flownode CPU Usage per Instance
+          type: timeseries
+          description: Current cpu usage by instance
+          unit: none
+          queries:
+            - expr: sum(rate(process_cpu_seconds_total{}[$__rate_interval]) * 1000) by (instance, pod)
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{ instance }}]-[{{ pod }}]'
+    - title: Frontend Requests
+      panels:
+        - title: HTTP QPS per Instance
+          type: timeseries
+          description: HTTP QPS per Instance.
+          unit: reqps
+          queries:
+            - expr: sum by(instance, pod, path, method, code) (rate(greptime_servers_http_requests_elapsed_count{path!~"/health|/metrics"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{path}}]-[{{method}}]-[{{code}}]'
+        - title: HTTP P99 per Instance
+          type: timeseries
+          description: HTTP P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, path, method, code) (rate(greptime_servers_http_requests_elapsed_bucket{path!~"/health|/metrics"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{path}}]-[{{method}}]-[{{code}}]-p99'
+        - title: gRPC QPS per Instance
+          type: timeseries
+          description: gRPC QPS per Instance.
+          unit: reqps
+          queries:
+            - expr: sum by(instance, pod, path, code) (rate(greptime_servers_grpc_requests_elapsed_count{}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{path}}]-[{{code}}]'
+        - title: gRPC P99 per Instance
+          type: timeseries
+          description: gRPC P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, path, code) (rate(greptime_servers_grpc_requests_elapsed_bucket{}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{path}}]-[{{method}}]-[{{code}}]-p99'
+        - title: MySQL QPS per Instance
+          type: timeseries
+          description: MySQL QPS per Instance.
+          unit: reqps
+          queries:
+            - expr: sum by(pod, instance)(rate(greptime_servers_mysql_query_elapsed_count{}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]'
+        - title: MySQL P99 per Instance
+          type: timeseries
+          description: MySQL P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(pod, instance, le) (rate(greptime_servers_mysql_query_elapsed_bucket{}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{ instance }}]-[{{ pod }}]-p99'
+        - title: PostgreSQL QPS per Instance
+          type: timeseries
+          description: PostgreSQL QPS per Instance.
+          unit: reqps
+          queries:
+            - expr: sum by(pod, instance)(rate(greptime_servers_postgres_query_elapsed_count{}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]'
+        - title: PostgreSQL P99 per Instance
+          type: timeseries
+          description: PostgreSQL P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(pod,instance,le) (rate(greptime_servers_postgres_query_elapsed_bucket{}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-p99'
+    - title: Frontend to Datanode
+      panels:
+        - title: Ingest Rows per Instance
+          type: timeseries
+          description: Ingestion rate by row as in each frontend
+          unit: rowsps
+          queries:
+            - expr: sum by(instance, pod)(rate(greptime_table_operator_ingest_rows{}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]'
+        - title: Region Call QPS per Instance
+          type: timeseries
+          description: Region Call QPS per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod, request_type) (rate(greptime_grpc_region_request_count{}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{request_type}}]'
+        - title: Region Call P99 per Instance
+          type: timeseries
+          description: Region Call P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, request_type) (rate(greptime_grpc_region_request_bucket{}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{request_type}}]'
+    - title: Mito Engine
+      panels:
+        - title: Request OPS per Instance
+          type: timeseries
+          description: Request QPS per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod, type) (rate(greptime_mito_handle_request_elapsed_count{}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{type}}]'
+        - title: Request P99 per Instance
+          type: timeseries
+          description: Request P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, type) (rate(greptime_mito_handle_request_elapsed_bucket{}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{type}}]'
+        - title: Write Buffer per Instance
+          type: timeseries
+          description: Write Buffer per Instance.
+          unit: decbytes
+          queries:
+            - expr: greptime_mito_write_buffer_bytes{}
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]'
+        - title: Write Rows per Instance
+          type: timeseries
+          description: Ingestion size by row counts.
+          unit: rowsps
+          queries:
+            - expr: sum by (instance, pod) (rate(greptime_mito_write_rows_total{}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]'
+        - title: Flush OPS per Instance
+          type: timeseries
+          description: Flush QPS per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod, reason) (rate(greptime_mito_flush_requests_total{}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{reason}}]'
+        - title: Write Stall per Instance
+          type: timeseries
+          description: Write Stall per Instance.
+          queries:
+            - expr: sum by(instance, pod) (greptime_mito_write_stall_total{})
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]'
+        - title: Read Stage OPS per Instance
+          type: timeseries
+          description: Read Stage OPS per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod) (rate(greptime_mito_read_stage_elapsed_count{ stage="total"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]'
+        - title: Read Stage P99 per Instance
+          type: timeseries
+          description: Read Stage P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, stage) (rate(greptime_mito_read_stage_elapsed_bucket{}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{stage}}]'
+        - title: Write Stage P99 per Instance
+          type: timeseries
+          description: Write Stage P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, stage) (rate(greptime_mito_write_stage_elapsed_bucket{}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{stage}}]'
+        - title: Compaction OPS per Instance
+          type: timeseries
+          description: Compaction OPS per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod) (rate(greptime_mito_compaction_total_elapsed_count{}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{ instance }}]-[{{pod}}]'
+        - title: Compaction P99 per Instance by Stage
+          type: timeseries
+          description: Compaction latency by stage
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, stage) (rate(greptime_mito_compaction_stage_elapsed_bucket{}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{stage}}]-p99'
+        - title: Compaction P99 per Instance
+          type: timeseries
+          description: Compaction P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le,stage) (rate(greptime_mito_compaction_total_elapsed_bucket{}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{stage}}]-compaction'
+        - title: WAL write size
+          type: timeseries
+          description: Write-ahead logs write size as bytes. This chart includes stats of p95 and p99 size by instance, total WAL write rate.
+          unit: bytes
+          queries:
+            - expr: histogram_quantile(0.95, sum by(le,instance, pod) (rate(raft_engine_write_size_bucket[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-req-size-p95'
+            - expr: histogram_quantile(0.99, sum by(le,instance,pod) (rate(raft_engine_write_size_bucket[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-req-size-p99'
+            - expr: sum by (instance, pod)(rate(raft_engine_write_size_sum[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-throughput'
+        - title: Cached Bytes per Instance
+          type: timeseries
+          description: Cached Bytes per Instance.
+          unit: decbytes
+          queries:
+            - expr: greptime_mito_cache_bytes{}
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{type}}]'
+        - title: Inflight Compaction
+          type: timeseries
+          description: Ongoing compaction task count
+          unit: none
+          queries:
+            - expr: greptime_mito_inflight_compaction_count
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]'
+        - title: WAL sync duration seconds
+          type: timeseries
+          description: Raft engine (local disk) log store sync latency, p99
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(le, type, node, instance, pod) (rate(raft_engine_sync_log_duration_seconds_bucket[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-p99'
+        - title: Log Store op duration seconds
+          type: timeseries
+          description: Write-ahead log operations latency at p99
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(le,logstore,optype,instance, pod) (rate(greptime_logstore_op_elapsed_bucket[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{logstore}}]-[{{optype}}]-p99'
+        - title: Inflight Flush
+          type: timeseries
+          description: Ongoing flush task count
+          unit: none
+          queries:
+            - expr: greptime_mito_inflight_flush_count
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]'
+    - title: OpenDAL
+      panels:
+        - title: QPS per Instance
+          type: timeseries
+          description: QPS per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod, scheme, operation) (rate(opendal_operation_duration_seconds_count{}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]'
+        - title: Read QPS per Instance
+          type: timeseries
+          description: Read QPS per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod, scheme) (rate(opendal_operation_duration_seconds_count{ operation="read"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{scheme}}]'
+        - title: Read P99 per Instance
+          type: timeseries
+          description: Read P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, scheme) (rate(opendal_operation_duration_seconds_bucket{operation="read"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-{{scheme}}'
+        - title: Write QPS per Instance
+          type: timeseries
+          description: Write QPS per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod, scheme) (rate(opendal_operation_duration_seconds_count{ operation="write"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-{{scheme}}'
+        - title: Write P99 per Instance
+          type: timeseries
+          description: Write P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, scheme) (rate(opendal_operation_duration_seconds_bucket{ operation="write"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{scheme}}]'
+        - title: List QPS per Instance
+          type: timeseries
+          description: List QPS per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod, scheme) (rate(opendal_operation_duration_seconds_count{ operation="list"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{scheme}}]'
+        - title: List P99 per Instance
+          type: timeseries
+          description: List P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, scheme) (rate(opendal_operation_duration_seconds_bucket{ operation="list"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{scheme}}]'
+        - title: Other Requests per Instance
+          type: timeseries
+          description: Other Requests per Instance.
+          unit: ops
+          queries:
+            - expr: sum by(instance, pod, scheme, operation) (rate(opendal_operation_duration_seconds_count{operation!~"read|write|list|stat"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]'
+        - title: Other Request P99 per Instance
+          type: timeseries
+          description: Other Request P99 per Instance.
+          unit: s
+          queries:
+            - expr: histogram_quantile(0.99, sum by(instance, pod, le, scheme, operation) (rate(opendal_operation_duration_seconds_bucket{ operation!~"read|write|list"}[$__rate_interval])))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]'
+        - title: Opendal traffic
+          type: timeseries
+          description: Total traffic as in bytes by instance and operation
+          unit: decbytes
+          queries:
+            - expr: sum by(instance, pod, scheme, operation) (rate(opendal_operation_bytes_sum{}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]'
+        - title: OpenDAL errors per Instance
+          type: timeseries
+          description: OpenDAL error counts per Instance.
+          queries:
+            - expr: sum by(instance, pod, scheme, operation, error) (rate(opendal_operation_errors_total{ error!="NotFound"}[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{scheme}}]-[{{operation}}]-[{{error}}]'
+    - title: Metasrv
+      panels:
+        - title: Region migration datanode
+          type: state-timeline
+          description: Counter of region migration by source and destination
+          unit: none
+          queries:
+            - expr: greptime_meta_region_migration_stat{datanode_type="src"}
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: from-datanode-{{datanode_id}}
+            - expr: greptime_meta_region_migration_stat{datanode_type="desc"}
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: to-datanode-{{datanode_id}}
+        - title: Region migration error
+          type: timeseries
+          description: Counter of region migration error
+          unit: none
+          queries:
+            - expr: greptime_meta_region_migration_error
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: __auto
+        - title: Datanode load
+          type: timeseries
+          description: Gauge of load information of each datanode, collected via heartbeat between datanode and metasrv. This information is for metasrv to schedule workloads.
+          unit: none
+          queries:
+            - expr: greptime_datanode_load
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: __auto
+    - title: Flownode
+      panels:
+        - title: Flow Ingest / Output Rate
+          type: timeseries
+          description: Flow Ingest / Output Rate.
+          queries:
+            - expr: sum by(instance, pod, direction) (rate(greptime_flow_processed_rows[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{pod}}]-[{{instance}}]-[{{direction}}]'
+        - title: Flow Ingest Latency
+          type: timeseries
+          description: Flow Ingest Latency.
+          queries:
+            - expr: histogram_quantile(0.95, sum(rate(greptime_flow_insert_elapsed_bucket[$__rate_interval])) by (le, instance, pod))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-p95'
+            - expr: histogram_quantile(0.99, sum(rate(greptime_flow_insert_elapsed_bucket[$__rate_interval])) by (le, instance, pod))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-p99'
+        - title: Flow Operation Latency
+          type: timeseries
+          description: Flow Operation Latency.
+          queries:
+            - expr: histogram_quantile(0.95, sum(rate(greptime_flow_processing_time_bucket[$__rate_interval])) by (le,instance,pod,type))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{type}}]-p95'
+            - expr: histogram_quantile(0.99, sum(rate(greptime_flow_processing_time_bucket[$__rate_interval])) by (le,instance,pod,type))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{type}}]-p99'
+        - title: Flow Buffer Size per Instance
+          type: timeseries
+          description: Flow Buffer Size per Instance.
+          queries:
+            - expr: greptime_flow_input_buf_size
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}]'
+        - title: Flow Processing Error per Instance
+          type: timeseries
+          description: Flow Processing Error per Instance.
+          queries:
+            - expr: sum by(instance,pod,code) (rate(greptime_flow_errors[$__rate_interval]))
+              datasource:
+                type: prometheus
+                uid: ${metrics}
+              legendFormat: '[{{instance}}]-[{{pod}}]-[{{code}}]'
--- a/grafana/greptimedb-cluster.json
+++ b/grafana/greptimedb-cluster.json
--- a/grafana/greptimedb.json
+++ b/grafana/greptimedb.json
--- a/grafana/scripts/check.sh
+++ b/grafana/scripts/check.sh
@@ -0,0 +1,54 @@
+#!/usr/bin/env bash
+
+DASHBOARD_DIR=${1:-grafana/dashboards}
+
+check_dashboard_description() {
+  for dashboard in $(find $DASHBOARD_DIR -name "*.json"); do
+    echo "Checking $dashboard description"
+
+    # Use jq to check for panels with empty or missing descriptions
+    invalid_panels=$(cat $dashboard | jq -r '
+      .panels[]
+    | select((.type == "stats" or .type == "timeseries") and (.description == "" or .description == null))')
+
+    # Check if any invalid panels were found
+    if [[ -n "$invalid_panels" ]]; then
+      echo "Error: The following panels have empty or missing descriptions:"
+      echo "$invalid_panels"
+      exit 1
+    else
+      echo "All panels with type 'stats' or 'timeseries' have valid descriptions."
+    fi
+  done
+}
+
+check_dashboards_generation() {
+  ./grafana/scripts/gen-dashboards.sh
+
+  if [[ -n "$(git diff --name-only grafana/dashboards)" ]]; then
+    echo "Error: The dashboards are not generated correctly. You should execute the `make dashboards` command."
+    exit 1
+  fi
+}
+
+check_datasource() {
+  for dashboard in $(find $DASHBOARD_DIR -name "*.json"); do
+    echo "Checking $dashboard datasource"
+    jq -r '.panels[] | select(.type != "row") | .targets[] | [.datasource.type, .datasource.uid] | @tsv' $dashboard | while read -r type uid; do
+    # if the datasource is prometheus, check if the uid is ${metrics}
+    if [[ "$type" == "prometheus" && "$uid" != "\${metrics}" ]]; then
+      echo "Error: The datasource uid of $dashboard is not valid. It should be \${metrics}, got $uid"
+      exit 1
+    fi
+    # if the datasource is mysql, check if the uid is ${information_schema}
+    if [[ "$type" == "mysql" && "$uid" != "\${information_schema}" ]]; then
+      echo "Error: The datasource uid of $dashboard is not valid. It should be \${information_schema}, got $uid"
+      exit 1
+    fi
+    done
+  done
+}
+
+check_dashboards_generation
+check_dashboard_description
+check_datasource
--- a/grafana/scripts/gen-dashboards.sh
+++ b/grafana/scripts/gen-dashboards.sh
@@ -0,0 +1,25 @@
+#! /usr/bin/env bash
+
+CLUSTER_DASHBOARD_DIR=${1:-grafana/dashboards/cluster}
+STANDALONE_DASHBOARD_DIR=${2:-grafana/dashboards/standalone}
+DAC_IMAGE=ghcr.io/zyy17/dac:20250423-522bd35
+
+remove_instance_filters() {
+  # Remove the instance filters for the standalone dashboards.
+  sed 's/instance=~\\"$datanode\\",//; s/instance=~\\"$datanode\\"//; s/instance=~\\"$frontend\\",//; s/instance=~\\"$frontend\\"//; s/instance=~\\"$metasrv\\",//; s/instance=~\\"$metasrv\\"//; s/instance=~\\"$flownode\\",//; s/instance=~\\"$flownode\\"//;' $CLUSTER_DASHBOARD_DIR/dashboard.json > $STANDALONE_DASHBOARD_DIR/dashboard.json
+}
+
+generate_intermediate_dashboards_and_docs() {
+  docker run -v ${PWD}:/greptimedb --rm ${DAC_IMAGE} \
+    -i /greptimedb/$CLUSTER_DASHBOARD_DIR/dashboard.json \
+    -o /greptimedb/$CLUSTER_DASHBOARD_DIR/dashboard.yaml \
+    -m /greptimedb/$CLUSTER_DASHBOARD_DIR/dashboard.md
+
+  docker run -v ${PWD}:/greptimedb --rm ${DAC_IMAGE} \
+    -i /greptimedb/$STANDALONE_DASHBOARD_DIR/dashboard.json \
+    -o /greptimedb/$STANDALONE_DASHBOARD_DIR/dashboard.yaml \
+    -m /greptimedb/$STANDALONE_DASHBOARD_DIR/dashboard.md
+}
+
+remove_instance_filters
+generate_intermediate_dashboards_and_docs
--- a/grafana/summary.sh
+++ b/grafana/summary.sh
@@ -1,11 +0,0 @@
-#!/usr/bin/env bash
-
-BASEDIR=$(dirname "$0")
-echo '| Title | Description | Expressions |
-|---|---|---|'
-
-cat $BASEDIR/greptimedb-cluster.json | jq -r '
-  .panels |
-  map(select(.type == "stat" or .type == "timeseries")) |
-  .[] | "| \(.title) | \(.description | gsub("\n"; "<br>")) | \(.targets | map(.expr // .rawSql | "`\(.|gsub("\n"; "<br>"))`")  | join("<br>")) |"
-'
--- a/src/api/src/helper.rs
+++ b/src/api/src/helper.rs
@@ -514,6 +514,7 @@ fn query_request_type(request: &QueryRequest) -> &'static str {
        Some(Query::Sql(_)) => "query.sql",
        Some(Query::LogicalPlan(_)) => "query.logical_plan",
        Some(Query::PromRangeQuery(_)) => "query.prom_range",
+        Some(Query::InsertIntoPlan(_)) => "query.insert_into_plan",
        None => "query.empty",
    }
 }
--- a/src/catalog/src/table_source.rs
+++ b/src/catalog/src/table_source.rs
@@ -27,7 +27,7 @@ use session::context::QueryContextRef;
 use snafu::{ensure, OptionExt, ResultExt};
 use table::metadata::TableType;
 use table::table::adapter::DfTableProviderAdapter;
-mod dummy_catalog;
+pub mod dummy_catalog;
 use dummy_catalog::DummyCatalogList;
 use table::TableRef;

--- a/src/cmd/src/flownode.rs
+++ b/src/cmd/src/flownode.rs
@@ -345,7 +345,7 @@ impl StartCommand {
        let client = Arc::new(NodeClients::new(channel_config));

        let invoker = FrontendInvoker::build_from(
-            flownode.flow_worker_manager().clone(),
+            flownode.flow_engine().streaming_engine(),
            catalog_manager.clone(),
            cached_meta_backend.clone(),
            layered_cache_registry.clone(),
@@ -355,7 +355,9 @@ impl StartCommand {
        .await
        .context(StartFlownodeSnafu)?;
        flownode
-            .flow_worker_manager()
+            .flow_engine()
+            .streaming_engine()
+            // TODO(discord9): refactor and avoid circular reference
            .set_frontend_invoker(invoker)
            .await;

--- a/src/cmd/src/metasrv.rs
+++ b/src/cmd/src/metasrv.rs
@@ -132,7 +132,7 @@ impl SubCommand {
 }

 #[derive(Debug, Default, Parser)]
-struct StartCommand {
+pub struct StartCommand {
    /// The address to bind the gRPC server.
    #[clap(long, alias = "bind-addr")]
    rpc_bind_addr: Option<String>,
@@ -172,7 +172,7 @@ struct StartCommand {
 }

 impl StartCommand {
-    fn load_options(&self, global_options: &GlobalOptions) -> Result<MetasrvOptions> {
+    pub fn load_options(&self, global_options: &GlobalOptions) -> Result<MetasrvOptions> {
        let mut opts = MetasrvOptions::load_layered_options(
            self.config_file.as_deref(),
            self.env_prefix.as_ref(),
@@ -261,7 +261,7 @@ impl StartCommand {
        Ok(())
    }

-    async fn build(&self, opts: MetasrvOptions) -> Result<Instance> {
+    pub async fn build(&self, opts: MetasrvOptions) -> Result<Instance> {
        common_runtime::init_global_runtimes(&opts.runtime);

        let guard = common_telemetry::init_global_logging(
--- a/src/cmd/src/standalone.rs
+++ b/src/cmd/src/standalone.rs
@@ -56,8 +56,8 @@ use datanode::datanode::{Datanode, DatanodeBuilder};
 use datanode::region_server::RegionServer;
 use file_engine::config::EngineConfig as FileEngineConfig;
 use flow::{
-    FlowConfig, FlowWorkerManager, FlownodeBuilder, FlownodeInstance, FlownodeOptions,
-    FrontendClient, FrontendInvoker,
+    FlowConfig, FlownodeBuilder, FlownodeInstance, FlownodeOptions, FrontendClient,
+    FrontendInvoker, GrpcQueryHandlerWithBoxedError, StreamingEngine,
 };
 use frontend::frontend::{Frontend, FrontendOptions};
 use frontend::instance::builder::FrontendBuilder;
@@ -524,17 +524,17 @@ impl StartCommand {
            ..Default::default()
        };

-        // TODO(discord9): for standalone not use grpc, but just somehow get a handler to frontend grpc client without
+        // for standalone not use grpc, but get a handler to frontend grpc client without
        // actually make a connection
-        let fe_server_addr = fe_opts.grpc.bind_addr.clone();
-        let frontend_client = FrontendClient::from_static_grpc_addr(fe_server_addr);
+        let (frontend_client, frontend_instance_handler) =
+            FrontendClient::from_empty_grpc_handler();
        let flow_builder = FlownodeBuilder::new(
            flownode_options,
            plugins.clone(),
            table_metadata_manager.clone(),
            catalog_manager.clone(),
            flow_metadata_manager.clone(),
-            Arc::new(frontend_client),
+            Arc::new(frontend_client.clone()),
        );
        let flownode = flow_builder
            .build()
@@ -544,15 +544,15 @@ impl StartCommand {

        // set the ref to query for the local flow state
        {
-            let flow_worker_manager = flownode.flow_worker_manager();
+            let flow_streaming_engine = flownode.flow_engine().streaming_engine();
            information_extension
-                .set_flow_worker_manager(flow_worker_manager.clone())
+                .set_flow_streaming_engine(flow_streaming_engine)
                .await;
        }

        let node_manager = Arc::new(StandaloneDatanodeManager {
            region_server: datanode.region_server(),
-            flow_server: flownode.flow_worker_manager(),
+            flow_server: flownode.flow_engine(),
        });

        let table_id_sequence = Arc::new(
@@ -606,10 +606,19 @@ impl StartCommand {
        .context(error::StartFrontendSnafu)?;
        let fe_instance = Arc::new(fe_instance);

-        let flow_worker_manager = flownode.flow_worker_manager();
+        // set the frontend client for flownode
+        let grpc_handler = fe_instance.clone() as Arc<dyn GrpcQueryHandlerWithBoxedError>;
+        let weak_grpc_handler = Arc::downgrade(&grpc_handler);
+        frontend_instance_handler
+            .lock()
+            .unwrap()
+            .replace(weak_grpc_handler);
+
+        // set the frontend invoker for flownode
+        let flow_streaming_engine = flownode.flow_engine().streaming_engine();
        // flow server need to be able to use frontend to write insert requests back
        let invoker = FrontendInvoker::build_from(
-            flow_worker_manager.clone(),
+            flow_streaming_engine.clone(),
            catalog_manager.clone(),
            kv_backend.clone(),
            layered_cache_registry.clone(),
@@ -618,7 +627,7 @@ impl StartCommand {
        )
        .await
        .context(error::StartFlownodeSnafu)?;
-        flow_worker_manager.set_frontend_invoker(invoker).await;
+        flow_streaming_engine.set_frontend_invoker(invoker).await;

        let export_metrics_task = ExportMetricsTask::try_new(&opts.export_metrics, Some(&plugins))
            .context(error::ServersSnafu)?;
@@ -694,7 +703,7 @@ pub struct StandaloneInformationExtension {
    region_server: RegionServer,
    procedure_manager: ProcedureManagerRef,
    start_time_ms: u64,
-    flow_worker_manager: RwLock<Option<Arc<FlowWorkerManager>>>,
+    flow_streaming_engine: RwLock<Option<Arc<StreamingEngine>>>,
 }

 impl StandaloneInformationExtension {
@@ -703,14 +712,14 @@ impl StandaloneInformationExtension {
            region_server,
            procedure_manager,
            start_time_ms: common_time::util::current_time_millis() as u64,
-            flow_worker_manager: RwLock::new(None),
+            flow_streaming_engine: RwLock::new(None),
        }
    }

-    /// Set the flow worker manager for the standalone instance.
-    pub async fn set_flow_worker_manager(&self, flow_worker_manager: Arc<FlowWorkerManager>) {
-        let mut guard = self.flow_worker_manager.write().await;
-        *guard = Some(flow_worker_manager);
+    /// Set the flow streaming engine for the standalone instance.
+    pub async fn set_flow_streaming_engine(&self, flow_streaming_engine: Arc<StreamingEngine>) {
+        let mut guard = self.flow_streaming_engine.write().await;
+        *guard = Some(flow_streaming_engine);
    }
 }

@@ -789,7 +798,7 @@ impl InformationExtension for StandaloneInformationExtension {

    async fn flow_stats(&self) -> std::result::Result<Option<FlowStat>, Self::Error> {
        Ok(Some(
-            self.flow_worker_manager
+            self.flow_streaming_engine
                .read()
                .await
                .as_ref()
--- a/src/cmd/tests/load_config_test.rs
+++ b/src/cmd/tests/load_config_test.rs
@@ -74,6 +74,7 @@ fn test_load_datanode_example_config() {
                RegionEngineConfig::File(FileEngineConfig {}),
                RegionEngineConfig::Metric(MetricEngineConfig {
                    experimental_sparse_primary_key_encoding: false,
+                    flush_metadata_region_interval: Duration::from_secs(30),
                }),
            ],
            logging: LoggingOptions {
@@ -216,6 +217,7 @@ fn test_load_standalone_example_config() {
                RegionEngineConfig::File(FileEngineConfig {}),
                RegionEngineConfig::Metric(MetricEngineConfig {
                    experimental_sparse_primary_key_encoding: false,
+                    flush_metadata_region_interval: Duration::from_secs(30),
                }),
            ],
            storage: StorageConfig {
--- a/src/common/base/src/plugins.rs
+++ b/src/common/base/src/plugins.rs
@@ -31,7 +31,8 @@ impl Plugins {
    }

    pub fn insert<T: 'static + Send + Sync>(&self, value: T) {
-        let _ = self.write().insert(value);
+        let last = self.write().insert(value);
+        assert!(last.is_none(), "each type of plugins must be one and only");
    }

    pub fn get<T: 'static + Send + Sync + Clone>(&self) -> Option<T> {
@@ -137,4 +138,12 @@ mod tests {
        assert_eq!(plugins.len(), 2);
        assert!(!plugins.is_empty());
    }
+
+    #[test]
+    #[should_panic(expected = "each type of plugins must be one and only")]
+    fn test_plugin_uniqueness() {
+        let plugins = Plugins::new();
+        plugins.insert(1i32);
+        plugins.insert(2i32);
+    }
 }
--- a/src/common/function/src/scalars/uddsketch_calc.rs
+++ b/src/common/function/src/scalars/uddsketch_calc.rs
@@ -115,6 +115,13 @@ impl Function for UddSketchCalcFunction {
                }
            };

+            // Check if the sketch is empty, if so, return null
+            // This is important to avoid panics when calling estimate_quantile on an empty sketch
+            // In practice, this will happen if input is all null
+            if sketch.bucket_iter().count() == 0 {
+                builder.push_null();
+                continue;
+            }
            // Compute the estimated quantile from the sketch
            let result = sketch.estimate_quantile(perc);
            builder.push(Some(result));
--- a/src/common/meta/src/cache/flow/table_flownode.rs
+++ b/src/common/meta/src/cache/flow/table_flownode.rs
@@ -187,6 +187,7 @@ mod tests {
                    },
                    flownode_ids: BTreeMap::from([(0, 1), (1, 2), (2, 3)]),
                    catalog_name: DEFAULT_CATALOG_NAME.to_string(),
+                    query_context: None,
                    flow_name: "my_flow".to_string(),
                    raw_sql: "sql".to_string(),
                    expire_after: Some(300),
--- a/src/common/meta/src/ddl/create_flow.rs
+++ b/src/common/meta/src/ddl/create_flow.rs
@@ -38,7 +38,7 @@ use table::metadata::TableId;
 use crate::cache_invalidator::Context;
 use crate::ddl::utils::{add_peer_context_if_needed, handle_retry_error};
 use crate::ddl::DdlContext;
-use crate::error::{self, Result};
+use crate::error::{self, Result, UnexpectedSnafu};
 use crate::instruction::{CacheIdent, CreateFlow};
 use crate::key::flow::flow_info::FlowInfoValue;
 use crate::key::flow::flow_route::FlowRouteValue;
@@ -171,7 +171,7 @@ impl CreateFlowProcedure {
        }
        self.data.state = CreateFlowState::CreateFlows;
        // determine flow type
-        self.data.flow_type = Some(determine_flow_type(&self.data.task));
+        self.data.flow_type = Some(get_flow_type_from_options(&self.data.task)?);

        Ok(Status::executing(true))
    }
@@ -196,8 +196,8 @@ impl CreateFlowProcedure {
            });
        }
        info!(
-            "Creating flow({:?}) on flownodes with peers={:?}",
-            self.data.flow_id, self.data.peers
+            "Creating flow({:?}, type={:?}) on flownodes with peers={:?}",
+            self.data.flow_id, self.data.flow_type, self.data.peers
        );
        join_all(create_flow)
            .await
@@ -306,8 +306,20 @@ impl Procedure for CreateFlowProcedure {
    }
 }

-pub fn determine_flow_type(_flow_task: &CreateFlowTask) -> FlowType {
-    FlowType::Batching
+pub fn get_flow_type_from_options(flow_task: &CreateFlowTask) -> Result<FlowType> {
+    let flow_type = flow_task
+        .flow_options
+        .get(FlowType::FLOW_TYPE_KEY)
+        .map(|s| s.as_str());
+    match flow_type {
+        Some(FlowType::BATCHING) => Ok(FlowType::Batching),
+        Some(FlowType::STREAMING) => Ok(FlowType::Streaming),
+        Some(unknown) => UnexpectedSnafu {
+            err_msg: format!("Unknown flow type: {}", unknown),
+        }
+        .fail(),
+        None => Ok(FlowType::Batching),
+    }
 }

 /// The state of [CreateFlowProcedure].
@@ -437,6 +449,7 @@ impl From<&CreateFlowData> for (FlowInfoValue, Vec<(FlowPartitionId, FlowRouteVa
            sink_table_name,
            flownode_ids,
            catalog_name,
+            query_context: Some(value.query_context.clone()),
            flow_name,
            raw_sql: sql,
            expire_after,
--- a/src/common/meta/src/ddl/tests/create_flow.rs
+++ b/src/common/meta/src/ddl/tests/create_flow.rs
@@ -46,7 +46,7 @@ pub(crate) fn test_create_flow_task(
        create_if_not_exists,
        expire_after: Some(300),
        comment: "".to_string(),
-        sql: "raw_sql".to_string(),
+        sql: "select 1".to_string(),
        flow_options: Default::default(),
    }
 }
--- a/src/common/meta/src/error.rs
+++ b/src/common/meta/src/error.rs
@@ -401,6 +401,13 @@ pub enum Error {
        location: Location,
    },

+    #[snafu(display("Invalid flow request body: {:?}", body))]
+    InvalidFlowRequestBody {
+        body: Box<Option<api::v1::flow::flow_request::Body>>,
+        #[snafu(implicit)]
+        location: Location,
+    },
+
    #[snafu(display("Failed to get kv cache, err: {}", err_msg))]
    GetKvCache { err_msg: String },

@@ -783,6 +790,14 @@ pub enum Error {
        #[snafu(source)]
        source: common_procedure::error::Error,
    },
+
+    #[snafu(display("Failed to parse timezone"))]
+    InvalidTimeZone {
+        #[snafu(implicit)]
+        location: Location,
+        #[snafu(source)]
+        error: common_time::error::Error,
+    },
 }

 pub type Result<T> = std::result::Result<T, Error>;
@@ -853,7 +868,9 @@ impl ErrorExt for Error {
            | TlsConfig { .. }
            | InvalidSetDatabaseOption { .. }
            | InvalidUnsetDatabaseOption { .. }
-            | InvalidTopicNamePrefix { .. } => StatusCode::InvalidArguments,
+            | InvalidTopicNamePrefix { .. }
+            | InvalidTimeZone { .. } => StatusCode::InvalidArguments,
+            InvalidFlowRequestBody { .. } => StatusCode::InvalidArguments,

            FlowNotFound { .. } => StatusCode::FlowNotFound,
            FlowRouteNotFound { .. } => StatusCode::Unexpected,
--- a/src/common/meta/src/key/flow.rs
+++ b/src/common/meta/src/key/flow.rs
@@ -452,6 +452,7 @@ mod tests {
        };
        FlowInfoValue {
            catalog_name: catalog_name.to_string(),
+            query_context: None,
            flow_name: flow_name.to_string(),
            source_table_ids,
            sink_table_name,
@@ -625,6 +626,7 @@ mod tests {
        };
        let flow_value = FlowInfoValue {
            catalog_name: "greptime".to_string(),
+            query_context: None,
            flow_name: "flow".to_string(),
            source_table_ids: vec![1024, 1025, 1026],
            sink_table_name: another_sink_table_name,
@@ -864,6 +866,7 @@ mod tests {
        };
        let flow_value = FlowInfoValue {
            catalog_name: "greptime".to_string(),
+            query_context: None,
            flow_name: "flow".to_string(),
            source_table_ids: vec![1024, 1025, 1026],
            sink_table_name: another_sink_table_name,
--- a/src/common/meta/src/key/flow/flow_info.rs
+++ b/src/common/meta/src/key/flow/flow_info.rs
@@ -121,6 +121,13 @@ pub struct FlowInfoValue {
    pub(crate) flownode_ids: BTreeMap<FlowPartitionId, FlownodeId>,
    /// The catalog name.
    pub(crate) catalog_name: String,
+    /// The query context used when create flow.
+    /// Although flow doesn't belong to any schema, this query_context is needed to remember
+    /// the query context when `create_flow` is executed
+    /// for recovering flow using the same sql&query_context after db restart.
+    /// if none, should use default query context
+    #[serde(default)]
+    pub(crate) query_context: Option<crate::rpc::ddl::QueryContext>,
    /// The flow name.
    pub(crate) flow_name: String,
    /// The raw sql.
@@ -155,6 +162,10 @@ impl FlowInfoValue {
        &self.catalog_name
    }

+    pub fn query_context(&self) -> &Option<crate::rpc::ddl::QueryContext> {
+        &self.query_context
+    }
+
    pub fn flow_name(&self) -> &String {
        &self.flow_name
    }
--- a/src/common/meta/src/region_registry.rs
+++ b/src/common/meta/src/region_registry.rs
@@ -113,8 +113,10 @@ impl LeaderRegionManifestInfo {
    pub fn prunable_entry_id(&self) -> u64 {
        match self {
            LeaderRegionManifestInfo::Mito {
-                flushed_entry_id, ..
-            } => *flushed_entry_id,
+                flushed_entry_id,
+                topic_latest_entry_id,
+                ..
+            } => (*flushed_entry_id).max(*topic_latest_entry_id),
            LeaderRegionManifestInfo::Metric {
                data_flushed_entry_id,
                data_topic_latest_entry_id,
--- a/src/common/meta/src/rpc/ddl.rs
+++ b/src/common/meta/src/rpc/ddl.rs
@@ -35,17 +35,20 @@ use api::v1::{
 };
 use base64::engine::general_purpose;
 use base64::Engine as _;
-use common_time::DatabaseTimeToLive;
+use common_time::{DatabaseTimeToLive, Timezone};
 use prost::Message;
 use serde::{Deserialize, Serialize};
 use serde_with::{serde_as, DefaultOnNull};
-use session::context::QueryContextRef;
+use session::context::{QueryContextBuilder, QueryContextRef};
 use snafu::{OptionExt, ResultExt};
 use table::metadata::{RawTableInfo, TableId};
 use table::table_name::TableName;
 use table::table_reference::TableReference;

-use crate::error::{self, InvalidSetDatabaseOptionSnafu, InvalidUnsetDatabaseOptionSnafu, Result};
+use crate::error::{
+    self, InvalidSetDatabaseOptionSnafu, InvalidTimeZoneSnafu, InvalidUnsetDatabaseOptionSnafu,
+    Result,
+};
 use crate::key::FlowId;

 /// DDL tasks
@@ -1202,7 +1205,7 @@ impl From<DropFlowTask> for PbDropFlowTask {
    }
 }

-#[derive(Debug, Clone, Serialize, Deserialize)]
+#[derive(Debug, Clone, Serialize, Deserialize, PartialEq)]
 pub struct QueryContext {
    current_catalog: String,
    current_schema: String,
@@ -1223,6 +1226,19 @@ impl From<QueryContextRef> for QueryContext {
    }
 }

+impl TryFrom<QueryContext> for session::context::QueryContext {
+    type Error = error::Error;
+    fn try_from(value: QueryContext) -> std::result::Result<Self, Self::Error> {
+        Ok(QueryContextBuilder::default()
+            .current_catalog(value.current_catalog)
+            .current_schema(value.current_schema)
+            .timezone(Timezone::from_tz_string(&value.timezone).context(InvalidTimeZoneSnafu)?)
+            .extensions(value.extensions)
+            .channel((value.channel as u32).into())
+            .build())
+    }
+}
+
 impl From<QueryContext> for PbQueryContext {
    fn from(
        QueryContext {
--- a/src/common/query/src/logical_plan.rs
+++ b/src/common/query/src/logical_plan.rs
@@ -18,16 +18,19 @@ mod udaf;

 use std::sync::Arc;

+use api::v1::TableName;
 use datafusion::catalog::CatalogProviderList;
 use datafusion::error::Result as DatafusionResult;
 use datafusion::logical_expr::{LogicalPlan, LogicalPlanBuilder};
-use datafusion_common::Column;
-use datafusion_expr::col;
+use datafusion_common::{Column, TableReference};
+use datafusion_expr::dml::InsertOp;
+use datafusion_expr::{col, DmlStatement, WriteOp};
 pub use expr::{build_filter_from_timestamp, build_same_type_ts_filter};
+use snafu::ResultExt;

 pub use self::accumulator::{Accumulator, AggregateFunctionCreator, AggregateFunctionCreatorRef};
 pub use self::udaf::AggregateFunction;
-use crate::error::Result;
+use crate::error::{GeneralDataFusionSnafu, Result};
 use crate::logical_plan::accumulator::*;
 use crate::signature::{Signature, Volatility};

@@ -79,6 +82,74 @@ pub fn rename_logical_plan_columns(
    LogicalPlanBuilder::from(plan).project(projection)?.build()
 }

+/// Convert a insert into logical plan to an (table_name, logical_plan)
+/// where table_name is the name of the table to insert into.
+/// logical_plan is the plan to be executed.
+///
+/// if input logical plan is not `insert into table_name <input>`, return None
+///
+/// Returned TableName will use provided catalog and schema if not specified in the logical plan,
+/// if table scan in logical plan have full table name, will **NOT** override it.
+pub fn breakup_insert_plan(
+    plan: &LogicalPlan,
+    default_catalog: &str,
+    default_schema: &str,
+) -> Option<(TableName, Arc<LogicalPlan>)> {
+    if let LogicalPlan::Dml(dml) = plan {
+        if dml.op != WriteOp::Insert(InsertOp::Append) {
+            return None;
+        }
+        let table_name = &dml.table_name;
+        let table_name = match table_name {
+            TableReference::Bare { table } => TableName {
+                catalog_name: default_catalog.to_string(),
+                schema_name: default_schema.to_string(),
+                table_name: table.to_string(),
+            },
+            TableReference::Partial { schema, table } => TableName {
+                catalog_name: default_catalog.to_string(),
+                schema_name: schema.to_string(),
+                table_name: table.to_string(),
+            },
+            TableReference::Full {
+                catalog,
+                schema,
+                table,
+            } => TableName {
+                catalog_name: catalog.to_string(),
+                schema_name: schema.to_string(),
+                table_name: table.to_string(),
+            },
+        };
+        let logical_plan = dml.input.clone();
+        Some((table_name, logical_plan))
+    } else {
+        None
+    }
+}
+
+/// create a `insert into table_name <input>` logical plan
+pub fn add_insert_to_logical_plan(
+    table_name: TableName,
+    table_schema: datafusion_common::DFSchemaRef,
+    input: LogicalPlan,
+) -> Result<LogicalPlan> {
+    let table_name = TableReference::Full {
+        catalog: table_name.catalog_name.into(),
+        schema: table_name.schema_name.into(),
+        table: table_name.table_name.into(),
+    };
+
+    let plan = LogicalPlan::Dml(DmlStatement::new(
+        table_name,
+        table_schema,
+        WriteOp::Insert(InsertOp::Append),
+        Arc::new(input),
+    ));
+    let plan = plan.recompute_schema().context(GeneralDataFusionSnafu)?;
+    Ok(plan)
+}
+
 /// The datafusion `[LogicalPlan]` decoder.
 #[async_trait::async_trait]
 pub trait SubstraitPlanDecoder {
--- a/src/common/wal/src/config/kafka/common.rs
+++ b/src/common/wal/src/config/kafka/common.rs
@@ -30,10 +30,10 @@ pub const DEFAULT_BACKOFF_CONFIG: BackoffConfig = BackoffConfig {
    deadline: Some(Duration::from_secs(120)),
 };

-/// Default interval for active WAL pruning.
-pub const DEFAULT_ACTIVE_PRUNE_INTERVAL: Duration = Duration::ZERO;
-/// Default limit for concurrent active pruning tasks.
-pub const DEFAULT_ACTIVE_PRUNE_TASK_LIMIT: usize = 10;
+/// Default interval for auto WAL pruning.
+pub const DEFAULT_AUTO_PRUNE_INTERVAL: Duration = Duration::ZERO;
+/// Default limit for concurrent auto pruning tasks.
+pub const DEFAULT_AUTO_PRUNE_PARALLELISM: usize = 10;
 /// Default interval for sending flush request to regions when pruning remote WAL.
 pub const DEFAULT_TRIGGER_FLUSH_THRESHOLD: u64 = 0;

--- a/src/common/wal/src/config/kafka/datanode.rs
+++ b/src/common/wal/src/config/kafka/datanode.rs
@@ -18,8 +18,8 @@ use common_base::readable_size::ReadableSize;
 use serde::{Deserialize, Serialize};

 use crate::config::kafka::common::{
-    KafkaConnectionConfig, KafkaTopicConfig, DEFAULT_ACTIVE_PRUNE_INTERVAL,
-    DEFAULT_ACTIVE_PRUNE_TASK_LIMIT, DEFAULT_TRIGGER_FLUSH_THRESHOLD,
+    KafkaConnectionConfig, KafkaTopicConfig, DEFAULT_AUTO_PRUNE_INTERVAL,
+    DEFAULT_AUTO_PRUNE_PARALLELISM, DEFAULT_TRIGGER_FLUSH_THRESHOLD,
 };

 /// Kafka wal configurations for datanode.
@@ -47,9 +47,8 @@ pub struct DatanodeKafkaConfig {
    pub dump_index_interval: Duration,
    /// Ignore missing entries during read WAL.
    pub overwrite_entry_start_id: bool,
-    // Active WAL pruning.
-    pub auto_prune_topic_records: bool,
    // Interval of WAL pruning.
+    #[serde(with = "humantime_serde")]
    pub auto_prune_interval: Duration,
    // Threshold for sending flush request when pruning remote WAL.
    // `None` stands for never sending flush request.
@@ -70,10 +69,9 @@ impl Default for DatanodeKafkaConfig {
            create_index: true,
            dump_index_interval: Duration::from_secs(60),
            overwrite_entry_start_id: false,
-            auto_prune_topic_records: false,
-            auto_prune_interval: DEFAULT_ACTIVE_PRUNE_INTERVAL,
+            auto_prune_interval: DEFAULT_AUTO_PRUNE_INTERVAL,
            trigger_flush_threshold: DEFAULT_TRIGGER_FLUSH_THRESHOLD,
-            auto_prune_parallelism: DEFAULT_ACTIVE_PRUNE_TASK_LIMIT,
+            auto_prune_parallelism: DEFAULT_AUTO_PRUNE_PARALLELISM,
        }
    }
 }
--- a/src/common/wal/src/config/kafka/metasrv.rs
+++ b/src/common/wal/src/config/kafka/metasrv.rs
@@ -17,8 +17,8 @@ use std::time::Duration;
 use serde::{Deserialize, Serialize};

 use crate::config::kafka::common::{
-    KafkaConnectionConfig, KafkaTopicConfig, DEFAULT_ACTIVE_PRUNE_INTERVAL,
-    DEFAULT_ACTIVE_PRUNE_TASK_LIMIT, DEFAULT_TRIGGER_FLUSH_THRESHOLD,
+    KafkaConnectionConfig, KafkaTopicConfig, DEFAULT_AUTO_PRUNE_INTERVAL,
+    DEFAULT_AUTO_PRUNE_PARALLELISM, DEFAULT_TRIGGER_FLUSH_THRESHOLD,
 };

 /// Kafka wal configurations for metasrv.
@@ -34,6 +34,7 @@ pub struct MetasrvKafkaConfig {
    // Automatically create topics for WAL.
    pub auto_create_topics: bool,
    // Interval of WAL pruning.
+    #[serde(with = "humantime_serde")]
    pub auto_prune_interval: Duration,
    // Threshold for sending flush request when pruning remote WAL.
    // `None` stands for never sending flush request.
@@ -48,9 +49,9 @@ impl Default for MetasrvKafkaConfig {
            connection: Default::default(),
            kafka_topic: Default::default(),
            auto_create_topics: true,
-            auto_prune_interval: DEFAULT_ACTIVE_PRUNE_INTERVAL,
+            auto_prune_interval: DEFAULT_AUTO_PRUNE_INTERVAL,
            trigger_flush_threshold: DEFAULT_TRIGGER_FLUSH_THRESHOLD,
-            auto_prune_parallelism: DEFAULT_ACTIVE_PRUNE_TASK_LIMIT,
+            auto_prune_parallelism: DEFAULT_AUTO_PRUNE_PARALLELISM,
        }
    }
 }
--- a/src/datanode/src/datanode.rs
+++ b/src/datanode/src/datanode.rs
@@ -57,9 +57,9 @@ use tokio::sync::Notify;

 use crate::config::{DatanodeOptions, RegionEngineConfig, StorageConfig};
 use crate::error::{
-    self, BuildMitoEngineSnafu, CreateDirSnafu, GetMetadataSnafu, MissingCacheSnafu,
-    MissingKvBackendSnafu, MissingNodeIdSnafu, OpenLogStoreSnafu, Result, ShutdownInstanceSnafu,
-    ShutdownServerSnafu, StartServerSnafu,
+    self, BuildMetricEngineSnafu, BuildMitoEngineSnafu, CreateDirSnafu, GetMetadataSnafu,
+    MissingCacheSnafu, MissingKvBackendSnafu, MissingNodeIdSnafu, OpenLogStoreSnafu, Result,
+    ShutdownInstanceSnafu, ShutdownServerSnafu, StartServerSnafu,
 };
 use crate::event_listener::{
    new_region_server_event_channel, NoopRegionServerEventListener, RegionServerEventListenerRef,
@@ -416,10 +416,11 @@ impl DatanodeBuilder {
                    )
                    .await?;

-                    let metric_engine = MetricEngine::new(
+                    let metric_engine = MetricEngine::try_new(
                        mito_engine.clone(),
                        metric_engine_config.take().unwrap_or_default(),
-                    );
+                    )
+                    .context(BuildMetricEngineSnafu)?;
                    engines.push(Arc::new(mito_engine) as _);
                    engines.push(Arc::new(metric_engine) as _);
                }
--- a/src/datanode/src/error.rs
+++ b/src/datanode/src/error.rs
@@ -336,6 +336,13 @@ pub enum Error {
        location: Location,
    },

+    #[snafu(display("Failed to build metric engine"))]
+    BuildMetricEngine {
+        source: metric_engine::error::Error,
+        #[snafu(implicit)]
+        location: Location,
+    },
+
    #[snafu(display("Failed to serialize options to TOML"))]
    TomlFormat {
        #[snafu(implicit)]
@@ -452,6 +459,7 @@ impl ErrorExt for Error {

            FindLogicalRegions { source, .. } => source.status_code(),
            BuildMitoEngine { source, .. } => source.status_code(),
+            BuildMetricEngine { source, .. } => source.status_code(),
            ConcurrentQueryLimiterClosed { .. } | ConcurrentQueryLimiterTimeout { .. } => {
                StatusCode::RegionBusy
            }
--- a/src/flow/src/adapter.rs
+++ b/src/flow/src/adapter.rs
@@ -58,7 +58,7 @@ use crate::metrics::{METRIC_FLOW_INSERT_ELAPSED, METRIC_FLOW_ROWS, METRIC_FLOW_R
 use crate::repr::{self, DiffRow, RelationDesc, Row, BATCH_SIZE};
 use crate::{CreateFlowArgs, FlowId, TableName};

-mod flownode_impl;
+pub(crate) mod flownode_impl;
 mod parse_expr;
 pub(crate) mod refill;
 mod stat;
@@ -135,12 +135,13 @@ impl Configurable for FlownodeOptions {
 }

 /// Arc-ed FlowNodeManager, cheaper to clone
-pub type FlowWorkerManagerRef = Arc<FlowWorkerManager>;
+pub type FlowStreamingEngineRef = Arc<StreamingEngine>;

 /// FlowNodeManager manages the state of all tasks in the flow node, which should be run on the same thread
 ///
 /// The choice of timestamp is just using current system timestamp for now
-pub struct FlowWorkerManager {
+///
+pub struct StreamingEngine {
    /// The handler to the worker that will run the dataflow
    /// which is `!Send` so a handle is used
    pub worker_handles: Vec<WorkerHandle>,
@@ -158,7 +159,8 @@ pub struct FlowWorkerManager {
    flow_err_collectors: RwLock<BTreeMap<FlowId, ErrCollector>>,
    src_send_buf_lens: RwLock<BTreeMap<TableId, watch::Receiver<usize>>>,
    tick_manager: FlowTickManager,
-    node_id: Option<u32>,
+    /// This node id is only available in distributed mode, on standalone mode this is guaranteed to be `None`
+    pub node_id: Option<u32>,
    /// Lock for flushing, will be `read` by `handle_inserts` and `write` by `flush_flow`
    ///
    /// So that a series of event like `inserts -> flush` can be handled correctly
@@ -168,7 +170,7 @@ pub struct FlowWorkerManager {
 }

 /// Building FlownodeManager
-impl FlowWorkerManager {
+impl StreamingEngine {
    /// set frontend invoker
    pub async fn set_frontend_invoker(&self, frontend: FrontendInvoker) {
        *self.frontend_invoker.write().await = Some(frontend);
@@ -187,7 +189,7 @@ impl FlowWorkerManager {
        let node_context = FlownodeContext::new(Box::new(srv_map.clone()) as _);
        let tick_manager = FlowTickManager::new();
        let worker_handles = Vec::new();
-        FlowWorkerManager {
+        StreamingEngine {
            worker_handles,
            worker_selector: Mutex::new(0),
            query_engine,
@@ -263,7 +265,7 @@ pub fn batches_to_rows_req(batches: Vec<Batch>) -> Result<Vec<DiffRequest>, Erro
 }

 /// This impl block contains methods to send writeback requests to frontend
-impl FlowWorkerManager {
+impl StreamingEngine {
    /// Return the number of requests it made
    pub async fn send_writeback_requests(&self) -> Result<usize, Error> {
        let all_reqs = self.generate_writeback_request().await?;
@@ -534,7 +536,7 @@ impl FlowWorkerManager {
 }

 /// Flow Runtime related methods
-impl FlowWorkerManager {
+impl StreamingEngine {
    /// Start state report handler, which will receive a sender from HeartbeatTask to send state size report back
    ///
    /// if heartbeat task is shutdown, this future will exit too
@@ -659,7 +661,7 @@ impl FlowWorkerManager {
        }
        // flow is now shutdown, drop frontend_invoker early so a ref cycle(in standalone mode) can be prevent:
        // FlowWorkerManager.frontend_invoker -> FrontendInvoker.inserter
-        // -> Inserter.node_manager -> NodeManager.flownode -> Flownode.flow_worker_manager.frontend_invoker
+        // -> Inserter.node_manager -> NodeManager.flownode -> Flownode.flow_streaming_engine.frontend_invoker
        self.frontend_invoker.write().await.take();
    }

@@ -728,7 +730,7 @@ impl FlowWorkerManager {
 }

 /// Create&Remove flow
-impl FlowWorkerManager {
+impl StreamingEngine {
    /// remove a flow by it's id
    pub async fn remove_flow_inner(&self, flow_id: FlowId) -> Result<(), Error> {
        for handle in self.worker_handles.iter() {
@@ -746,7 +748,6 @@ impl FlowWorkerManager {
    /// steps to create task:
    /// 1. parse query into typed plan(and optional parse expire_after expr)
    /// 2. render source/sink with output table id and used input table id
-    #[allow(clippy::too_many_arguments)]
    pub async fn create_flow_inner(&self, args: CreateFlowArgs) -> Result<Option<FlowId>, Error> {
        let CreateFlowArgs {
            flow_id,
--- a/src/flow/src/adapter/flownode_impl.rs
+++ b/src/flow/src/adapter/flownode_impl.rs
@@ -20,35 +20,394 @@ use api::v1::flow::{
    flow_request, CreateRequest, DropRequest, FlowRequest, FlowResponse, FlushFlow,
 };
 use api::v1::region::InsertRequests;
+use catalog::CatalogManager;
 use common_error::ext::BoxedError;
 use common_meta::ddl::create_flow::FlowType;
-use common_meta::error::{Result as MetaResult, UnexpectedSnafu};
+use common_meta::error::Result as MetaResult;
+use common_meta::key::flow::FlowMetadataManager;
 use common_runtime::JoinHandle;
-use common_telemetry::{trace, warn};
+use common_telemetry::{error, info, trace, warn};
 use datatypes::value::Value;
+use futures::TryStreamExt;
 use itertools::Itertools;
-use snafu::{IntoError, OptionExt, ResultExt};
+use session::context::QueryContextBuilder;
+use snafu::{ensure, IntoError, OptionExt, ResultExt};
 use store_api::storage::{RegionId, TableId};
+use tokio::sync::{Mutex, RwLock};

-use crate::adapter::{CreateFlowArgs, FlowWorkerManager};
+use crate::adapter::{CreateFlowArgs, StreamingEngine};
 use crate::batching_mode::engine::BatchingEngine;
 use crate::engine::FlowEngine;
-use crate::error::{CreateFlowSnafu, FlowNotFoundSnafu, InsertIntoFlowSnafu, InternalSnafu};
+use crate::error::{
+    CreateFlowSnafu, ExternalSnafu, FlowNotFoundSnafu, IllegalCheckTaskStateSnafu,
+    InsertIntoFlowSnafu, InternalSnafu, JoinTaskSnafu, ListFlowsSnafu, SyncCheckTaskSnafu,
+    UnexpectedSnafu,
+};
 use crate::metrics::METRIC_FLOW_TASK_COUNT;
 use crate::repr::{self, DiffRow};
 use crate::{Error, FlowId};

+/// Ref to [`FlowDualEngine`]
+pub type FlowDualEngineRef = Arc<FlowDualEngine>;
+
 /// Manage both streaming and batching mode engine
 ///
 /// including create/drop/flush flow
 /// and redirect insert requests to the appropriate engine
 pub struct FlowDualEngine {
-    streaming_engine: Arc<FlowWorkerManager>,
+    streaming_engine: Arc<StreamingEngine>,
    batching_engine: Arc<BatchingEngine>,
    /// helper struct for faster query flow by table id or vice versa
-    src_table2flow: std::sync::RwLock<SrcTableToFlow>,
+    src_table2flow: RwLock<SrcTableToFlow>,
+    flow_metadata_manager: Arc<FlowMetadataManager>,
+    catalog_manager: Arc<dyn CatalogManager>,
+    check_task: tokio::sync::Mutex<Option<ConsistentCheckTask>>,
 }

+impl FlowDualEngine {
+    pub fn new(
+        streaming_engine: Arc<StreamingEngine>,
+        batching_engine: Arc<BatchingEngine>,
+        flow_metadata_manager: Arc<FlowMetadataManager>,
+        catalog_manager: Arc<dyn CatalogManager>,
+    ) -> Self {
+        Self {
+            streaming_engine,
+            batching_engine,
+            src_table2flow: RwLock::new(SrcTableToFlow::default()),
+            flow_metadata_manager,
+            catalog_manager,
+            check_task: Mutex::new(None),
+        }
+    }
+
+    pub fn streaming_engine(&self) -> Arc<StreamingEngine> {
+        self.streaming_engine.clone()
+    }
+
+    pub fn batching_engine(&self) -> Arc<BatchingEngine> {
+        self.batching_engine.clone()
+    }
+
+    /// Try to sync with check task, this is only used in drop flow&flush flow, so a flow id is required
+    ///
+    /// the need to sync is to make sure flush flow actually get called
+    async fn try_sync_with_check_task(
+        &self,
+        flow_id: FlowId,
+        allow_drop: bool,
+    ) -> Result<(), Error> {
+        // this function rarely get called so adding some log is helpful
+        info!("Try to sync with check task for flow {}", flow_id);
+        let mut retry = 0;
+        let max_retry = 10;
+        // keep trying to trigger consistent check
+        while retry < max_retry {
+            if let Some(task) = self.check_task.lock().await.as_ref() {
+                task.trigger(false, allow_drop).await?;
+                break;
+            }
+            retry += 1;
+            tokio::time::sleep(std::time::Duration::from_millis(500)).await;
+        }
+
+        if retry == max_retry {
+            error!(
+                "Can't sync with check task for flow {} with allow_drop={}",
+                flow_id, allow_drop
+            );
+            return SyncCheckTaskSnafu {
+                flow_id,
+                allow_drop,
+            }
+            .fail();
+        }
+        info!("Successfully sync with check task for flow {}", flow_id);
+
+        Ok(())
+    }
+
+    /// Spawn a task to consistently check if all flow tasks in metasrv is created on flownode,
+    /// so on startup, this will create all missing flow tasks, and constantly check at a interval
+    async fn check_flow_consistent(
+        &self,
+        allow_create: bool,
+        allow_drop: bool,
+    ) -> Result<(), Error> {
+        // use nodeid to determine if this is standalone/distributed mode, and retrieve all flows in this node(in distributed mode)/or all flows(in standalone mode)
+        let nodeid = self.streaming_engine.node_id;
+        let should_exists: Vec<_> = if let Some(nodeid) = nodeid {
+            // nodeid is available, so we only need to check flows on this node
+            // which also means we are in distributed mode
+            let to_be_recover = self
+                .flow_metadata_manager
+                .flownode_flow_manager()
+                .flows(nodeid.into())
+                .try_collect::<Vec<_>>()
+                .await
+                .context(ListFlowsSnafu {
+                    id: Some(nodeid.into()),
+                })?;
+            to_be_recover.into_iter().map(|(id, _)| id).collect()
+        } else {
+            // nodeid is not available, so we need to check all flows
+            // which also means we are in standalone mode
+            let all_catalogs = self
+                .catalog_manager
+                .catalog_names()
+                .await
+                .map_err(BoxedError::new)
+                .context(ExternalSnafu)?;
+            let mut all_flow_ids = vec![];
+            for catalog in all_catalogs {
+                let flows = self
+                    .flow_metadata_manager
+                    .flow_name_manager()
+                    .flow_names(&catalog)
+                    .await
+                    .try_collect::<Vec<_>>()
+                    .await
+                    .map_err(BoxedError::new)
+                    .context(ExternalSnafu)?;
+
+                all_flow_ids.extend(flows.into_iter().map(|(_, id)| id.flow_id()));
+            }
+            all_flow_ids
+        };
+        let should_exists = should_exists
+            .into_iter()
+            .map(|i| i as FlowId)
+            .collect::<HashSet<_>>();
+        let actual_exists = self.list_flows().await?.into_iter().collect::<HashSet<_>>();
+        let to_be_created = should_exists
+            .iter()
+            .filter(|id| !actual_exists.contains(id))
+            .collect::<Vec<_>>();
+        let to_be_dropped = actual_exists
+            .iter()
+            .filter(|id| !should_exists.contains(id))
+            .collect::<Vec<_>>();
+
+        if !to_be_created.is_empty() {
+            if allow_create {
+                info!(
+                    "Recovering {} flows: {:?}",
+                    to_be_created.len(),
+                    to_be_created
+                );
+                let mut errors = vec![];
+                for flow_id in to_be_created {
+                    let flow_id = *flow_id;
+                    let info = self
+                        .flow_metadata_manager
+                        .flow_info_manager()
+                        .get(flow_id as u32)
+                        .await
+                        .map_err(BoxedError::new)
+                        .context(ExternalSnafu)?
+                        .context(FlowNotFoundSnafu { id: flow_id })?;
+
+                    let sink_table_name = [
+                        info.sink_table_name().catalog_name.clone(),
+                        info.sink_table_name().schema_name.clone(),
+                        info.sink_table_name().table_name.clone(),
+                    ];
+                    let args = CreateFlowArgs {
+                        flow_id,
+                        sink_table_name,
+                        source_table_ids: info.source_table_ids().to_vec(),
+                        // because recover should only happen on restart the `create_if_not_exists` and `or_replace` can be arbitrary value(since flow doesn't exist)
+                        // but for the sake of consistency and to make sure recover of flow actually happen, we set both to true
+                        // (which is also fine since checks for not allow both to be true is on metasrv and we already pass that)
+                        create_if_not_exists: true,
+                        or_replace: true,
+                        expire_after: info.expire_after(),
+                        comment: Some(info.comment().clone()),
+                        sql: info.raw_sql().clone(),
+                        flow_options: info.options().clone(),
+                        query_ctx: info
+                            .query_context()
+                            .clone()
+                            .map(|ctx| {
+                                ctx.try_into()
+                                    .map_err(BoxedError::new)
+                                    .context(ExternalSnafu)
+                            })
+                            .transpose()?
+                            // or use default QueryContext with catalog_name from info
+                            // to keep compatibility with old version
+                            .or_else(|| {
+                                Some(
+                                    QueryContextBuilder::default()
+                                        .current_catalog(info.catalog_name().to_string())
+                                        .build(),
+                                )
+                            }),
+                    };
+                    if let Err(err) = self
+                        .create_flow(args)
+                        .await
+                        .map_err(BoxedError::new)
+                        .with_context(|_| CreateFlowSnafu {
+                            sql: info.raw_sql().clone(),
+                        })
+                    {
+                        errors.push((flow_id, err));
+                    }
+                }
+                for (flow_id, err) in errors {
+                    warn!("Failed to recreate flow {}, err={:#?}", flow_id, err);
+                }
+            } else {
+                warn!(
+                    "Flownode {:?} found flows not exist in flownode, flow_ids={:?}",
+                    nodeid, to_be_created
+                );
+            }
+        }
+        if !to_be_dropped.is_empty() {
+            if allow_drop {
+                info!("Dropping flows: {:?}", to_be_dropped);
+                let mut errors = vec![];
+                for flow_id in to_be_dropped {
+                    let flow_id = *flow_id;
+                    if let Err(err) = self.remove_flow(flow_id).await {
+                        errors.push((flow_id, err));
+                    }
+                }
+                for (flow_id, err) in errors {
+                    warn!("Failed to drop flow {}, err={:#?}", flow_id, err);
+                }
+            } else {
+                warn!(
+                    "Flownode {:?} found flows not exist in flownode, flow_ids={:?}",
+                    nodeid, to_be_dropped
+                );
+            }
+        }
+        Ok(())
+    }
+
+    // TODO(discord9): consider sync this with heartbeat(might become necessary in the future)
+    pub async fn start_flow_consistent_check_task(self: &Arc<Self>) -> Result<(), Error> {
+        let mut check_task = self.check_task.lock().await;
+        ensure!(
+            check_task.is_none(),
+            IllegalCheckTaskStateSnafu {
+                reason: "Flow consistent check task already exists",
+            }
+        );
+        let task = ConsistentCheckTask::start_check_task(self).await?;
+        *check_task = Some(task);
+        Ok(())
+    }
+
+    pub async fn stop_flow_consistent_check_task(&self) -> Result<(), Error> {
+        info!("Stopping flow consistent check task");
+        let mut check_task = self.check_task.lock().await;
+
+        ensure!(
+            check_task.is_some(),
+            IllegalCheckTaskStateSnafu {
+                reason: "Flow consistent check task does not exist",
+            }
+        );
+
+        check_task.take().unwrap().stop().await?;
+        info!("Stopped flow consistent check task");
+        Ok(())
+    }
+
+    /// TODO(discord9): also add a `exists` api using flow metadata manager's `exists` method
+    async fn flow_exist_in_metadata(&self, flow_id: FlowId) -> Result<bool, Error> {
+        self.flow_metadata_manager
+            .flow_info_manager()
+            .get(flow_id as u32)
+            .await
+            .map_err(BoxedError::new)
+            .context(ExternalSnafu)
+            .map(|info| info.is_some())
+    }
+}
+
+struct ConsistentCheckTask {
+    handle: JoinHandle<()>,
+    shutdown_tx: tokio::sync::mpsc::Sender<()>,
+    trigger_tx: tokio::sync::mpsc::Sender<(bool, bool, tokio::sync::oneshot::Sender<()>)>,
+}
+
+impl ConsistentCheckTask {
+    async fn start_check_task(engine: &Arc<FlowDualEngine>) -> Result<Self, Error> {
+        // first do recover flows
+        engine.check_flow_consistent(true, false).await?;
+
+        let inner = engine.clone();
+        let (tx, mut rx) = tokio::sync::mpsc::channel(1);
+        let (trigger_tx, mut trigger_rx) =
+            tokio::sync::mpsc::channel::<(bool, bool, tokio::sync::oneshot::Sender<()>)>(10);
+        let handle = common_runtime::spawn_global(async move {
+            let (mut allow_create, mut allow_drop) = (false, false);
+            let mut ret_signal: Option<tokio::sync::oneshot::Sender<()>> = None;
+            loop {
+                if let Err(err) = inner.check_flow_consistent(allow_create, allow_drop).await {
+                    error!(err; "Failed to check flow consistent");
+                }
+                if let Some(done) = ret_signal.take() {
+                    let _ = done.send(());
+                }
+                tokio::select! {
+                    _ = rx.recv() => break,
+                    incoming = trigger_rx.recv() => if let Some(incoming) = incoming {
+                        (allow_create, allow_drop) = (incoming.0, incoming.1);
+                        ret_signal = Some(incoming.2);
+                    },
+                    _ = tokio::time::sleep(std::time::Duration::from_secs(10)) => {
+                        (allow_create, allow_drop) = (false, false);
+                    },
+                }
+            }
+        });
+        Ok(ConsistentCheckTask {
+            handle,
+            shutdown_tx: tx,
+            trigger_tx,
+        })
+    }
+
+    async fn trigger(&self, allow_create: bool, allow_drop: bool) -> Result<(), Error> {
+        let (tx, rx) = tokio::sync::oneshot::channel();
+        self.trigger_tx
+            .send((allow_create, allow_drop, tx))
+            .await
+            .map_err(|_| {
+                IllegalCheckTaskStateSnafu {
+                    reason: "Failed to send trigger signal",
+                }
+                .build()
+            })?;
+        rx.await.map_err(|_| {
+            IllegalCheckTaskStateSnafu {
+                reason: "Failed to receive trigger signal",
+            }
+            .build()
+        })?;
+        Ok(())
+    }
+
+    async fn stop(self) -> Result<(), Error> {
+        self.shutdown_tx.send(()).await.map_err(|_| {
+            IllegalCheckTaskStateSnafu {
+                reason: "Failed to send shutdown signal",
+            }
+            .build()
+        })?;
+        // abort so no need to wait
+        self.handle.abort();
+        Ok(())
+    }
+}
+
+#[derive(Default)]
 struct SrcTableToFlow {
    /// mapping of table ids to flow ids for streaming mode
    stream: HashMap<TableId, HashSet<FlowId>>,
@@ -138,35 +497,49 @@ impl FlowEngine for FlowDualEngine {

        self.src_table2flow
            .write()
-            .unwrap()
+            .await
            .add_flow(flow_id, flow_type, src_table_ids);

        Ok(res)
    }

    async fn remove_flow(&self, flow_id: FlowId) -> Result<(), Error> {
-        let flow_type = self.src_table2flow.read().unwrap().get_flow_type(flow_id);
+        let flow_type = self.src_table2flow.read().await.get_flow_type(flow_id);
+
        match flow_type {
            Some(FlowType::Batching) => self.batching_engine.remove_flow(flow_id).await,
            Some(FlowType::Streaming) => self.streaming_engine.remove_flow(flow_id).await,
-            None => FlowNotFoundSnafu { id: flow_id }.fail(),
+            None => {
+                // this can happen if flownode just restart, and is stilling creating the flow
+                // since now that this flow should dropped, we need to trigger the consistent check and allow drop
+                // this rely on drop flow ddl delete metadata first, see src/common/meta/src/ddl/drop_flow.rs
+                warn!(
+                    "Flow {} is not exist in the underlying engine, but exist in metadata",
+                    flow_id
+                );
+                self.try_sync_with_check_task(flow_id, true).await?;
+
+                Ok(())
+            }
        }?;
        // remove mapping
-        self.src_table2flow.write().unwrap().remove_flow(flow_id);
+        self.src_table2flow.write().await.remove_flow(flow_id);
        Ok(())
    }

    async fn flush_flow(&self, flow_id: FlowId) -> Result<usize, Error> {
-        let flow_type = self.src_table2flow.read().unwrap().get_flow_type(flow_id);
+        // sync with check task
+        self.try_sync_with_check_task(flow_id, false).await?;
+        let flow_type = self.src_table2flow.read().await.get_flow_type(flow_id);
        match flow_type {
            Some(FlowType::Batching) => self.batching_engine.flush_flow(flow_id).await,
            Some(FlowType::Streaming) => self.streaming_engine.flush_flow(flow_id).await,
-            None => FlowNotFoundSnafu { id: flow_id }.fail(),
+            None => Ok(0),
        }
    }

    async fn flow_exist(&self, flow_id: FlowId) -> Result<bool, Error> {
-        let flow_type = self.src_table2flow.read().unwrap().get_flow_type(flow_id);
+        let flow_type = self.src_table2flow.read().await.get_flow_type(flow_id);
        // not using `flow_type.is_some()` to make sure the flow is actually exist in the underlying engine
        match flow_type {
            Some(FlowType::Batching) => self.batching_engine.flow_exist(flow_id).await,
@@ -175,6 +548,13 @@ impl FlowEngine for FlowDualEngine {
        }
    }

+    async fn list_flows(&self) -> Result<impl IntoIterator<Item = FlowId>, Error> {
+        let stream_flows = self.streaming_engine.list_flows().await?;
+        let batch_flows = self.batching_engine.list_flows().await?;
+
+        Ok(stream_flows.into_iter().chain(batch_flows))
+    }
+
    async fn handle_flow_inserts(
        &self,
        request: api::v1::region::InsertRequests,
@@ -184,7 +564,7 @@ impl FlowEngine for FlowDualEngine {
        let mut to_batch_engine = request.requests;

        {
-            let src_table2flow = self.src_table2flow.read().unwrap();
+            let src_table2flow = self.src_table2flow.read().await;
            to_batch_engine.retain(|req| {
                let region_id = RegionId::from(req.region_id);
                let table_id = region_id.table_id();
@@ -221,12 +601,7 @@ impl FlowEngine for FlowDualEngine {
                requests: to_batch_engine,
            })
            .await?;
-        stream_handler.await.map_err(|e| {
-            crate::error::UnexpectedSnafu {
-                reason: format!("JoinError when handle inserts for flow stream engine: {e:?}"),
-            }
-            .build()
-        })??;
+        stream_handler.await.context(JoinTaskSnafu)??;

        Ok(())
    }
@@ -307,14 +682,7 @@ impl common_meta::node_manager::Flownode for FlowDualEngine {
                    ..Default::default()
                })
            }
-            None => UnexpectedSnafu {
-                err_msg: "Missing request body",
-            }
-            .fail(),
-            _ => UnexpectedSnafu {
-                err_msg: "Invalid request body.",
-            }
-            .fail(),
+            other => common_meta::error::InvalidFlowRequestBodySnafu { body: other }.fail(),
        }
    }

@@ -339,7 +707,7 @@ fn to_meta_err(
 }

 #[async_trait::async_trait]
-impl common_meta::node_manager::Flownode for FlowWorkerManager {
+impl common_meta::node_manager::Flownode for StreamingEngine {
    async fn handle(&self, request: FlowRequest) -> MetaResult<FlowResponse> {
        let query_ctx = request
            .header
@@ -413,14 +781,7 @@ impl common_meta::node_manager::Flownode for FlowWorkerManager {
                    ..Default::default()
                })
            }
-            None => UnexpectedSnafu {
-                err_msg: "Missing request body",
-            }
-            .fail(),
-            _ => UnexpectedSnafu {
-                err_msg: "Invalid request body.",
-            }
-            .fail(),
+            other => common_meta::error::InvalidFlowRequestBodySnafu { body: other }.fail(),
        }
    }

@@ -432,7 +793,7 @@ impl common_meta::node_manager::Flownode for FlowWorkerManager {
    }
 }

-impl FlowEngine for FlowWorkerManager {
+impl FlowEngine for StreamingEngine {
    async fn create_flow(&self, args: CreateFlowArgs) -> Result<Option<FlowId>, Error> {
        self.create_flow_inner(args).await
    }
@@ -449,6 +810,16 @@ impl FlowEngine for FlowWorkerManager {
        self.flow_exist_inner(flow_id).await
    }

+    async fn list_flows(&self) -> Result<impl IntoIterator<Item = FlowId>, Error> {
+        Ok(self
+            .flow_err_collectors
+            .read()
+            .await
+            .keys()
+            .cloned()
+            .collect::<Vec<_>>())
+    }
+
    async fn handle_flow_inserts(
        &self,
        request: api::v1::region::InsertRequests,
@@ -474,7 +845,7 @@ impl FetchFromRow {
    }
 }

-impl FlowWorkerManager {
+impl StreamingEngine {
    async fn handle_inserts_inner(
        &self,
        request: InsertRequests,
@@ -552,7 +923,7 @@ impl FlowWorkerManager {
                            .copied()
                            .map(FetchFromRow::Idx)
                            .or_else(|| col_default_val.clone().map(FetchFromRow::Default))
-                            .with_context(|| crate::error::UnexpectedSnafu {
+                            .with_context(|| UnexpectedSnafu {
                                reason: format!(
                                    "Column not found: {}, default_value: {:?}",
                                    col_name, col_default_val
--- a/src/flow/src/adapter/refill.rs
+++ b/src/flow/src/adapter/refill.rs
@@ -31,7 +31,7 @@ use snafu::{ensure, OptionExt, ResultExt};
 use table::metadata::TableId;

 use crate::adapter::table_source::ManagedTableSource;
-use crate::adapter::{FlowId, FlowWorkerManager, FlowWorkerManagerRef};
+use crate::adapter::{FlowId, FlowStreamingEngineRef, StreamingEngine};
 use crate::error::{FlowNotFoundSnafu, JoinTaskSnafu, UnexpectedSnafu};
 use crate::expr::error::ExternalSnafu;
 use crate::expr::utils::find_plan_time_window_expr_lower_bound;
@@ -39,10 +39,10 @@ use crate::repr::RelationDesc;
 use crate::server::get_all_flow_ids;
 use crate::{Error, FrontendInvoker};

-impl FlowWorkerManager {
+impl StreamingEngine {
    /// Create and start refill flow tasks in background
    pub async fn create_and_start_refill_flow_tasks(
-        self: &FlowWorkerManagerRef,
+        self: &FlowStreamingEngineRef,
        flow_metadata_manager: &FlowMetadataManagerRef,
        catalog_manager: &CatalogManagerRef,
    ) -> Result<(), Error> {
@@ -130,7 +130,7 @@ impl FlowWorkerManager {

    /// Starting to refill flows, if any error occurs, will rebuild the flow and retry
    pub(crate) async fn starting_refill_flows(
-        self: &FlowWorkerManagerRef,
+        self: &FlowStreamingEngineRef,
        tasks: Vec<RefillTask>,
    ) -> Result<(), Error> {
        // TODO(discord9): add a back pressure mechanism
@@ -266,7 +266,7 @@ impl TaskState<()> {
    fn start_running(
        &mut self,
        task_data: &TaskData,
-        manager: FlowWorkerManagerRef,
+        manager: FlowStreamingEngineRef,
        mut output_stream: SendableRecordBatchStream,
    ) -> Result<(), Error> {
        let data = (*task_data).clone();
@@ -383,7 +383,7 @@ impl RefillTask {
    /// Start running the task in background, non-blocking
    pub async fn start_running(
        &mut self,
-        manager: FlowWorkerManagerRef,
+        manager: FlowStreamingEngineRef,
        invoker: &FrontendInvoker,
    ) -> Result<(), Error> {
        let TaskState::Prepared { sql } = &mut self.state else {
--- a/src/flow/src/adapter/stat.rs
+++ b/src/flow/src/adapter/stat.rs
@@ -16,9 +16,9 @@ use std::collections::BTreeMap;

 use common_meta::key::flow::flow_state::FlowStat;

-use crate::FlowWorkerManager;
+use crate::StreamingEngine;

-impl FlowWorkerManager {
+impl StreamingEngine {
    pub async fn gen_state_report(&self) -> FlowStat {
        let mut full_report = BTreeMap::new();
        let mut last_exec_time_map = BTreeMap::new();
--- a/src/flow/src/adapter/util.rs
+++ b/src/flow/src/adapter/util.rs
@@ -33,8 +33,8 @@ use crate::adapter::table_source::TableDesc;
 use crate::adapter::{TableName, WorkerHandle, AUTO_CREATED_PLACEHOLDER_TS_COL};
 use crate::error::{Error, ExternalSnafu, UnexpectedSnafu};
 use crate::repr::{ColumnType, RelationDesc, RelationType};
-use crate::FlowWorkerManager;
-impl FlowWorkerManager {
+use crate::StreamingEngine;
+impl StreamingEngine {
    /// Get a worker handle for creating flow, using round robin to select a worker
    pub(crate) async fn get_worker_handle_for_create_flow(&self) -> &WorkerHandle {
        let use_idx = {
--- a/src/flow/src/batching_mode.rs
+++ b/src/flow/src/batching_mode.rs
@@ -32,3 +32,9 @@ pub const SLOW_QUERY_THRESHOLD: Duration = Duration::from_secs(60);

 /// The minimum duration between two queries execution by batching mode task
 const MIN_REFRESH_DURATION: Duration = Duration::new(5, 0);
+
+/// Grpc connection timeout
+const GRPC_CONN_TIMEOUT: Duration = Duration::from_secs(5);
+
+/// Grpc max retry number
+const GRPC_MAX_RETRIES: u32 = 3;
--- a/src/flow/src/batching_mode/engine.rs
+++ b/src/flow/src/batching_mode/engine.rs
@@ -17,14 +17,16 @@
 use std::collections::{BTreeMap, HashMap};
 use std::sync::Arc;

+use catalog::CatalogManagerRef;
 use common_error::ext::BoxedError;
 use common_meta::ddl::create_flow::FlowType;
 use common_meta::key::flow::FlowMetadataManagerRef;
-use common_meta::key::table_info::TableInfoManager;
+use common_meta::key::table_info::{TableInfoManager, TableInfoValue};
 use common_meta::key::TableMetadataManagerRef;
 use common_runtime::JoinHandle;
-use common_telemetry::info;
 use common_telemetry::tracing::warn;
+use common_telemetry::{debug, info};
+use common_time::TimeToLive;
 use query::QueryEngineRef;
 use snafu::{ensure, OptionExt, ResultExt};
 use store_api::storage::RegionId;
@@ -36,7 +38,9 @@ use crate::batching_mode::task::BatchingTask;
 use crate::batching_mode::time_window::{find_time_window_expr, TimeWindowExpr};
 use crate::batching_mode::utils::sql_to_df_plan;
 use crate::engine::FlowEngine;
-use crate::error::{ExternalSnafu, FlowAlreadyExistSnafu, TableNotFoundMetaSnafu, UnexpectedSnafu};
+use crate::error::{
+    ExternalSnafu, FlowAlreadyExistSnafu, TableNotFoundMetaSnafu, UnexpectedSnafu, UnsupportedSnafu,
+};
 use crate::{CreateFlowArgs, Error, FlowId, TableName};

 /// Batching mode Engine, responsible for driving all the batching mode tasks
@@ -48,6 +52,7 @@ pub struct BatchingEngine {
    frontend_client: Arc<FrontendClient>,
    flow_metadata_manager: FlowMetadataManagerRef,
    table_meta: TableMetadataManagerRef,
+    catalog_manager: CatalogManagerRef,
    query_engine: QueryEngineRef,
 }

@@ -57,6 +62,7 @@ impl BatchingEngine {
        query_engine: QueryEngineRef,
        flow_metadata_manager: FlowMetadataManagerRef,
        table_meta: TableMetadataManagerRef,
+        catalog_manager: CatalogManagerRef,
    ) -> Self {
        Self {
            tasks: Default::default(),
@@ -64,6 +70,7 @@ impl BatchingEngine {
            frontend_client,
            flow_metadata_manager,
            table_meta,
+            catalog_manager,
            query_engine,
        }
    }
@@ -179,6 +186,16 @@ async fn get_table_name(
    table_info: &TableInfoManager,
    table_id: &TableId,
 ) -> Result<TableName, Error> {
+    get_table_info(table_info, table_id).await.map(|info| {
+        let name = info.table_name();
+        [name.catalog_name, name.schema_name, name.table_name]
+    })
+}
+
+async fn get_table_info(
+    table_info: &TableInfoManager,
+    table_id: &TableId,
+) -> Result<TableInfoValue, Error> {
    table_info
        .get(*table_id)
        .await
@@ -187,8 +204,7 @@ async fn get_table_name(
        .with_context(|| UnexpectedSnafu {
            reason: format!("Table id = {:?}, couldn't found table name", table_id),
        })
-        .map(|name| name.table_name())
-        .map(|name| [name.catalog_name, name.schema_name, name.table_name])
+        .map(|info| info.into_inner())
 }

 impl BatchingEngine {
@@ -248,7 +264,20 @@ impl BatchingEngine {
        let query_ctx = Arc::new(query_ctx);
        let mut source_table_names = Vec::with_capacity(2);
        for src_id in source_table_ids {
+            // also check table option to see if ttl!=instant
            let table_name = get_table_name(self.table_meta.table_info_manager(), &src_id).await?;
+            let table_info = get_table_info(self.table_meta.table_info_manager(), &src_id).await?;
+            ensure!(
+                table_info.table_info.meta.options.ttl != Some(TimeToLive::Instant),
+                UnsupportedSnafu {
+                    reason: format!(
+                        "Source table `{}`(id={}) has instant TTL, Instant TTL is not supported under batching mode. Consider using a TTL longer than flush interval",
+                        table_name.join("."),
+                        src_id
+                    ),
+                }
+            );
+
            source_table_names.push(table_name);
        }

@@ -273,7 +302,14 @@ impl BatchingEngine {
            })
            .transpose()?;

-        info!("Flow id={}, found time window expr={:?}", flow_id, phy_expr);
+        info!(
+            "Flow id={}, found time window expr={}",
+            flow_id,
+            phy_expr
+                .as_ref()
+                .map(|phy_expr| phy_expr.to_string())
+                .unwrap_or("None".to_string())
+        );

        let task = BatchingTask::new(
            flow_id,
@@ -284,7 +320,7 @@ impl BatchingEngine {
            sink_table_name,
            source_table_names,
            query_ctx,
-            self.table_meta.clone(),
+            self.catalog_manager.clone(),
            rx,
        );

@@ -295,10 +331,11 @@ impl BatchingEngine {
        // check execute once first to detect any error early
        task.check_execute(&engine, &frontend).await?;

-        // TODO(discord9): also save handle & use time wheel or what for better
-        let _handle = common_runtime::spawn_global(async move {
+        // TODO(discord9): use time wheel or what for better
+        let handle = common_runtime::spawn_global(async move {
            task_inner.start_executing_loop(engine, frontend).await;
        });
+        task.state.write().unwrap().task_handle = Some(handle);

        // only replace here not earlier because we want the old one intact if something went wrong before this line
        let replaced_old_task_opt = self.tasks.write().await.insert(flow_id, task);
@@ -326,15 +363,23 @@ impl BatchingEngine {
    }

    pub async fn flush_flow_inner(&self, flow_id: FlowId) -> Result<usize, Error> {
+        debug!("Try flush flow {flow_id}");
        let task = self.tasks.read().await.get(&flow_id).cloned();
        let task = task.with_context(|| UnexpectedSnafu {
            reason: format!("Can't found task for flow {flow_id}"),
        })?;

+        task.mark_all_windows_as_dirty()?;
+
        let res = task
            .gen_exec_once(&self.query_engine, &self.frontend_client)
            .await?;
+
        let affected_rows = res.map(|(r, _)| r).unwrap_or_default() as usize;
+        debug!(
+            "Successfully flush flow {flow_id}, affected rows={}",
+            affected_rows
+        );
        Ok(affected_rows)
    }

@@ -357,6 +402,9 @@ impl FlowEngine for BatchingEngine {
    async fn flow_exist(&self, flow_id: FlowId) -> Result<bool, Error> {
        Ok(self.flow_exist_inner(flow_id).await)
    }
+    async fn list_flows(&self) -> Result<impl IntoIterator<Item = FlowId>, Error> {
+        Ok(self.tasks.read().await.keys().cloned().collect::<Vec<_>>())
+    }
    async fn handle_flow_inserts(
        &self,
        request: api::v1::region::InsertRequests,
--- a/src/flow/src/batching_mode/frontend_client.rs
+++ b/src/flow/src/batching_mode/frontend_client.rs
@@ -14,44 +14,109 @@

 //! Frontend client to run flow as batching task which is time-window-aware normal query triggered every tick set by user

-use std::sync::Arc;
+use std::sync::{Arc, Weak};

-use client::{Client, Database, DEFAULT_CATALOG_NAME, DEFAULT_SCHEMA_NAME};
-use common_error::ext::BoxedError;
+use api::v1::greptime_request::Request;
+use api::v1::CreateTableExpr;
+use client::{Client, Database};
+use common_error::ext::{BoxedError, ErrorExt};
 use common_grpc::channel_manager::{ChannelConfig, ChannelManager};
 use common_meta::cluster::{NodeInfo, NodeInfoKey, Role};
 use common_meta::peer::Peer;
 use common_meta::rpc::store::RangeRequest;
+use common_query::Output;
+use common_telemetry::warn;
 use meta_client::client::MetaClient;
-use snafu::ResultExt;
+use servers::query_handler::grpc::GrpcQueryHandler;
+use session::context::{QueryContextBuilder, QueryContextRef};
+use snafu::{OptionExt, ResultExt};

-use crate::batching_mode::DEFAULT_BATCHING_ENGINE_QUERY_TIMEOUT;
-use crate::error::{ExternalSnafu, UnexpectedSnafu};
+use crate::batching_mode::{
+    DEFAULT_BATCHING_ENGINE_QUERY_TIMEOUT, GRPC_CONN_TIMEOUT, GRPC_MAX_RETRIES,
+};
+use crate::error::{ExternalSnafu, InvalidRequestSnafu, UnexpectedSnafu};
 use crate::Error;

-fn default_channel_mgr() -> ChannelManager {
-    let cfg = ChannelConfig::new().timeout(DEFAULT_BATCHING_ENGINE_QUERY_TIMEOUT);
-    ChannelManager::with_config(cfg)
+/// Just like [`GrpcQueryHandler`] but use BoxedError
+///
+/// basically just a specialized `GrpcQueryHandler<Error=BoxedError>`
+///
+/// this is only useful for flownode to
+/// invoke frontend Instance in standalone mode
+#[async_trait::async_trait]
+pub trait GrpcQueryHandlerWithBoxedError: Send + Sync + 'static {
+    async fn do_query(
+        &self,
+        query: Request,
+        ctx: QueryContextRef,
+    ) -> std::result::Result<Output, BoxedError>;
 }

-fn client_from_urls(addrs: Vec<String>) -> Client {
-    Client::with_manager_and_urls(default_channel_mgr(), addrs)
+/// auto impl
+#[async_trait::async_trait]
+impl<
+        E: ErrorExt + Send + Sync + 'static,
+        T: GrpcQueryHandler<Error = E> + Send + Sync + 'static,
+    > GrpcQueryHandlerWithBoxedError for T
+{
+    async fn do_query(
+        &self,
+        query: Request,
+        ctx: QueryContextRef,
+    ) -> std::result::Result<Output, BoxedError> {
+        self.do_query(query, ctx).await.map_err(BoxedError::new)
+    }
 }

+type HandlerMutable = Arc<std::sync::Mutex<Option<Weak<dyn GrpcQueryHandlerWithBoxedError>>>>;
+
 /// A simple frontend client able to execute sql using grpc protocol
-#[derive(Debug)]
+///
+/// This is for computation-heavy query which need to offload computation to frontend, lifting the load from flownode
+#[derive(Debug, Clone)]
 pub enum FrontendClient {
    Distributed {
        meta_client: Arc<MetaClient>,
+        chnl_mgr: ChannelManager,
    },
    Standalone {
        /// for the sake of simplicity still use grpc even in standalone mode
        /// notice the client here should all be lazy, so that can wait after frontend is booted then make conn
-        /// TODO(discord9): not use grpc under standalone mode
-        database_client: DatabaseWithPeer,
+        database_client: HandlerMutable,
    },
 }

+impl FrontendClient {
+    /// Create a new empty frontend client, with a `HandlerMutable` to set the grpc handler later
+    pub fn from_empty_grpc_handler() -> (Self, HandlerMutable) {
+        let handler = Arc::new(std::sync::Mutex::new(None));
+        (
+            Self::Standalone {
+                database_client: handler.clone(),
+            },
+            handler,
+        )
+    }
+
+    pub fn from_meta_client(meta_client: Arc<MetaClient>) -> Self {
+        Self::Distributed {
+            meta_client,
+            chnl_mgr: {
+                let cfg = ChannelConfig::new()
+                    .connect_timeout(GRPC_CONN_TIMEOUT)
+                    .timeout(DEFAULT_BATCHING_ENGINE_QUERY_TIMEOUT);
+                ChannelManager::with_config(cfg)
+            },
+        }
+    }
+
+    pub fn from_grpc_handler(grpc_handler: Weak<dyn GrpcQueryHandlerWithBoxedError>) -> Self {
+        Self::Standalone {
+            database_client: Arc::new(std::sync::Mutex::new(Some(grpc_handler))),
+        }
+    }
+}
+
 #[derive(Debug, Clone)]
 pub struct DatabaseWithPeer {
    pub database: Database,
@@ -64,25 +129,6 @@ impl DatabaseWithPeer {
    }
 }

-impl FrontendClient {
-    pub fn from_meta_client(meta_client: Arc<MetaClient>) -> Self {
-        Self::Distributed { meta_client }
-    }
-
-    pub fn from_static_grpc_addr(addr: String) -> Self {
-        let peer = Peer {
-            id: 0,
-            addr: addr.clone(),
-        };
-
-        let client = client_from_urls(vec![addr]);
-        let database = Database::new(DEFAULT_CATALOG_NAME, DEFAULT_SCHEMA_NAME, client);
-        Self::Standalone {
-            database_client: DatabaseWithPeer::new(database, peer),
-        }
-    }
-}
-
 impl FrontendClient {
    async fn scan_for_frontend(&self) -> Result<Vec<(NodeInfoKey, NodeInfo)>, Error> {
        let Self::Distributed { meta_client, .. } = self else {
@@ -115,10 +161,21 @@ impl FrontendClient {
    }

    /// Get the database with max `last_activity_ts`
-    async fn get_last_active_frontend(&self) -> Result<DatabaseWithPeer, Error> {
-        if let Self::Standalone { database_client } = self {
-            return Ok(database_client.clone());
-        }
+    async fn get_last_active_frontend(
+        &self,
+        catalog: &str,
+        schema: &str,
+    ) -> Result<DatabaseWithPeer, Error> {
+        let Self::Distributed {
+            meta_client: _,
+            chnl_mgr,
+        } = self
+        else {
+            return UnexpectedSnafu {
+                reason: "Expect distributed mode",
+            }
+            .fail();
+        };

        let frontends = self.scan_for_frontend().await?;
        let mut peer = None;
@@ -133,16 +190,139 @@ impl FrontendClient {
            }
            .fail()?
        };
-        let client = client_from_urls(vec![peer.addr.clone()]);
-        let database = Database::new(DEFAULT_CATALOG_NAME, DEFAULT_SCHEMA_NAME, client);
+        let client = Client::with_manager_and_urls(chnl_mgr.clone(), vec![peer.addr.clone()]);
+        let database = Database::new(catalog, schema, client);
        Ok(DatabaseWithPeer::new(database, peer))
    }

-    /// Get a database client, and possibly update it before returning.
-    pub async fn get_database_client(&self) -> Result<DatabaseWithPeer, Error> {
+    pub async fn create(
+        &self,
+        create: CreateTableExpr,
+        catalog: &str,
+        schema: &str,
+    ) -> Result<u32, Error> {
+        self.handle(
+            Request::Ddl(api::v1::DdlRequest {
+                expr: Some(api::v1::ddl_request::Expr::CreateTable(create)),
+            }),
+            catalog,
+            schema,
+            &mut None,
+        )
+        .await
+    }
+
+    /// Handle a request to frontend
+    pub(crate) async fn handle(
+        &self,
+        req: api::v1::greptime_request::Request,
+        catalog: &str,
+        schema: &str,
+        peer_desc: &mut Option<PeerDesc>,
+    ) -> Result<u32, Error> {
        match self {
-            Self::Standalone { database_client } => Ok(database_client.clone()),
-            Self::Distributed { meta_client: _ } => self.get_last_active_frontend().await,
+            FrontendClient::Distributed { .. } => {
+                let db = self.get_last_active_frontend(catalog, schema).await?;
+
+                *peer_desc = Some(PeerDesc::Dist {
+                    peer: db.peer.clone(),
+                });
+
+                let mut retry = 0;
+
+                loop {
+                    let ret = db.database.handle(req.clone()).await.with_context(|_| {
+                        InvalidRequestSnafu {
+                            context: format!("Failed to handle request: {:?}", req),
+                        }
+                    });
+                    if let Err(err) = ret {
+                        if retry < GRPC_MAX_RETRIES {
+                            retry += 1;
+                            warn!(
+                                "Failed to send request to grpc handle at Peer={:?}, retry = {}, error = {:?}",
+                                db.peer, retry, err
+                            );
+                            continue;
+                        } else {
+                            common_telemetry::error!(
+                                "Failed to send request to grpc handle at Peer={:?} after {} retries, error = {:?}",
+                                db.peer, retry, err
+                            );
+                            return Err(err);
+                        }
+                    }
+                    return ret;
+                }
+            }
+            FrontendClient::Standalone { database_client } => {
+                let ctx = QueryContextBuilder::default()
+                    .current_catalog(catalog.to_string())
+                    .current_schema(schema.to_string())
+                    .build();
+                let ctx = Arc::new(ctx);
+                {
+                    let database_client = {
+                        database_client
+                            .lock()
+                            .map_err(|e| {
+                                UnexpectedSnafu {
+                                    reason: format!("Failed to lock database client: {e}"),
+                                }
+                                .build()
+                            })?
+                            .as_ref()
+                            .context(UnexpectedSnafu {
+                                reason: "Standalone's frontend instance is not set",
+                            })?
+                            .upgrade()
+                            .context(UnexpectedSnafu {
+                                reason: "Failed to upgrade database client",
+                            })?
+                    };
+                    let resp: common_query::Output = database_client
+                        .do_query(req.clone(), ctx)
+                        .await
+                        .map_err(BoxedError::new)
+                        .context(ExternalSnafu)?;
+                    match resp.data {
+                        common_query::OutputData::AffectedRows(rows) => {
+                            Ok(rows.try_into().map_err(|_| {
+                                UnexpectedSnafu {
+                                    reason: format!("Failed to convert rows to u32: {}", rows),
+                                }
+                                .build()
+                            })?)
+                        }
+                        _ => UnexpectedSnafu {
+                            reason: "Unexpected output data",
+                        }
+                        .fail(),
+                    }
+                }
+            }
+        }
+    }
+}
+
+/// Describe a peer of frontend
+#[derive(Debug, Default)]
+pub(crate) enum PeerDesc {
+    /// Distributed mode's frontend peer address
+    Dist {
+        /// frontend peer address
+        peer: Peer,
+    },
+    /// Standalone mode
+    #[default]
+    Standalone,
+}
+
+impl std::fmt::Display for PeerDesc {
+    fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
+        match self {
+            PeerDesc::Dist { peer } => write!(f, "{}", peer.addr),
+            PeerDesc::Standalone => write!(f, "standalone"),
        }
    }
 }
--- a/src/flow/src/batching_mode/state.rs
+++ b/src/flow/src/batching_mode/state.rs
@@ -22,13 +22,14 @@ use common_telemetry::tracing::warn;
 use common_time::Timestamp;
 use datatypes::value::Value;
 use session::context::QueryContextRef;
-use snafu::ResultExt;
+use snafu::{OptionExt, ResultExt};
 use tokio::sync::oneshot;
 use tokio::time::Instant;

 use crate::batching_mode::task::BatchingTask;
+use crate::batching_mode::time_window::TimeWindowExpr;
 use crate::batching_mode::MIN_REFRESH_DURATION;
-use crate::error::{DatatypesSnafu, InternalSnafu, TimeSnafu};
+use crate::error::{DatatypesSnafu, InternalSnafu, TimeSnafu, UnexpectedSnafu};
 use crate::{Error, FlowId};

 /// The state of the [`BatchingTask`].
@@ -46,6 +47,8 @@ pub struct TaskState {
    exec_state: ExecState,
    /// Shutdown receiver
    pub(crate) shutdown_rx: oneshot::Receiver<()>,
+    /// Task handle
+    pub(crate) task_handle: Option<tokio::task::JoinHandle<()>>,
 }
 impl TaskState {
    pub fn new(query_ctx: QueryContextRef, shutdown_rx: oneshot::Receiver<()>) -> Self {
@@ -56,6 +59,7 @@ impl TaskState {
            dirty_time_windows: Default::default(),
            exec_state: ExecState::Idle,
            shutdown_rx,
+            task_handle: None,
        }
    }

@@ -70,7 +74,11 @@ impl TaskState {
    /// wait for at least `last_query_duration`, at most `max_timeout` to start next query
    ///
    /// if have more dirty time window, exec next query immediately
-    pub fn get_next_start_query_time(&self, max_timeout: Option<Duration>) -> Instant {
+    pub fn get_next_start_query_time(
+        &self,
+        flow_id: FlowId,
+        max_timeout: Option<Duration>,
+    ) -> Instant {
        let next_duration = max_timeout
            .unwrap_or(self.last_query_duration)
            .min(self.last_query_duration);
@@ -80,6 +88,12 @@ impl TaskState {
        if self.dirty_time_windows.windows.is_empty() {
            self.last_update_time + next_duration
        } else {
+            debug!(
+                "Flow id = {}, still have {} dirty time window({:?}), execute immediately",
+                flow_id,
+                self.dirty_time_windows.windows.len(),
+                self.dirty_time_windows.windows
+            );
            Instant::now()
        }
    }
@@ -115,6 +129,15 @@ impl DirtyTimeWindows {
        }
    }

+    pub fn add_window(&mut self, start: Timestamp, end: Option<Timestamp>) {
+        self.windows.insert(start, end);
+    }
+
+    /// Clean all dirty time windows, useful when can't found time window expr
+    pub fn clean(&mut self) {
+        self.windows.clear();
+    }
+
    /// Generate all filter expressions consuming all time windows
    pub fn gen_filter_exprs(
        &mut self,
@@ -177,6 +200,18 @@ impl DirtyTimeWindows {

        let mut expr_lst = vec![];
        for (start, end) in first_nth.into_iter() {
+            // align using time window exprs
+            let (start, end) = if let Some(ctx) = task_ctx {
+                let Some(time_window_expr) = &ctx.config.time_window_expr else {
+                    UnexpectedSnafu {
+                        reason: "time_window_expr is not set",
+                    }
+                    .fail()?
+                };
+                self.align_time_window(start, end, time_window_expr)?
+            } else {
+                (start, end)
+            };
            debug!(
                "Time window start: {:?}, end: {:?}",
                start.to_iso8601_string(),
@@ -199,6 +234,30 @@ impl DirtyTimeWindows {
        Ok(expr)
    }

+    fn align_time_window(
+        &self,
+        start: Timestamp,
+        end: Option<Timestamp>,
+        time_window_expr: &TimeWindowExpr,
+    ) -> Result<(Timestamp, Option<Timestamp>), Error> {
+        let align_start = time_window_expr.eval(start)?.0.context(UnexpectedSnafu {
+            reason: format!(
+                "Failed to align start time {:?} with time window expr {:?}",
+                start, time_window_expr
+            ),
+        })?;
+        let align_end = end
+            .and_then(|end| {
+                time_window_expr
+                    .eval(end)
+                    // if after aligned, end is the same, then use end(because it's already aligned) else use aligned end
+                    .map(|r| if r.0 == Some(end) { r.0 } else { r.1 })
+                    .transpose()
+            })
+            .transpose()?;
+        Ok((align_start, align_end))
+    }
+
    /// Merge time windows that overlaps or get too close
    pub fn merge_dirty_time_windows(
        &mut self,
@@ -287,8 +346,12 @@ enum ExecState {
 #[cfg(test)]
 mod test {
    use pretty_assertions::assert_eq;
+    use session::context::QueryContext;

    use super::*;
+    use crate::batching_mode::time_window::find_time_window_expr;
+    use crate::batching_mode::utils::sql_to_df_plan;
+    use crate::test_utils::create_test_query_engine;

    #[test]
    fn test_merge_dirty_time_windows() {
@@ -404,4 +467,59 @@ mod test {
            assert_eq!(expected_filter_expr, to_sql.as_deref());
        }
    }
+
+    #[tokio::test]
+    async fn test_align_time_window() {
+        type TimeWindow = (Timestamp, Option<Timestamp>);
+        struct TestCase {
+            sql: String,
+            aligns: Vec<(TimeWindow, TimeWindow)>,
+        }
+        let testcases: Vec<TestCase> = vec![TestCase{
+            sql: "SELECT date_bin(INTERVAL '5 second', ts) AS time_window FROM numbers_with_ts GROUP BY time_window;".to_string(),
+            aligns: vec![
+                ((Timestamp::new_second(3), None), (Timestamp::new_second(0), None)),
+                ((Timestamp::new_second(8), None), (Timestamp::new_second(5), None)),
+                ((Timestamp::new_second(8), Some(Timestamp::new_second(10))), (Timestamp::new_second(5), Some(Timestamp::new_second(10)))),
+                ((Timestamp::new_second(8), Some(Timestamp::new_second(9))), (Timestamp::new_second(5), Some(Timestamp::new_second(10)))),
+            ],
+        }];
+
+        let query_engine = create_test_query_engine();
+        let ctx = QueryContext::arc();
+        for TestCase { sql, aligns } in testcases {
+            let plan = sql_to_df_plan(ctx.clone(), query_engine.clone(), &sql, true)
+                .await
+                .unwrap();
+
+            let (column_name, time_window_expr, _, df_schema) = find_time_window_expr(
+                &plan,
+                query_engine.engine_state().catalog_manager().clone(),
+                ctx.clone(),
+            )
+            .await
+            .unwrap();
+
+            let time_window_expr = time_window_expr
+                .map(|expr| {
+                    TimeWindowExpr::from_expr(
+                        &expr,
+                        &column_name,
+                        &df_schema,
+                        &query_engine.engine_state().session_state(),
+                    )
+                })
+                .transpose()
+                .unwrap()
+                .unwrap();
+
+            let dirty = DirtyTimeWindows::default();
+            for (before_align, expected_after_align) in aligns {
+                let after_align = dirty
+                    .align_time_window(before_align.0, before_align.1, &time_window_expr)
+                    .unwrap();
+                assert_eq!(expected_after_align, after_align);
+            }
+        }
+    }
 }
--- a/src/flow/src/batching_mode/task.rs
+++ b/src/flow/src/batching_mode/task.rs
@@ -12,33 +12,32 @@
 // See the License for the specific language governing permissions and
 // limitations under the License.

-use std::collections::HashSet;
+use std::collections::{BTreeSet, HashSet};
 use std::ops::Deref;
 use std::sync::{Arc, RwLock};
 use std::time::{Duration, SystemTime, UNIX_EPOCH};

 use api::v1::CreateTableExpr;
 use arrow_schema::Fields;
+use catalog::CatalogManagerRef;
 use common_error::ext::BoxedError;
-use common_meta::key::table_name::TableNameKey;
-use common_meta::key::TableMetadataManagerRef;
+use common_query::logical_plan::breakup_insert_plan;
 use common_telemetry::tracing::warn;
 use common_telemetry::{debug, info};
 use common_time::Timestamp;
+use datafusion::optimizer::analyzer::count_wildcard_rule::CountWildcardRule;
+use datafusion::optimizer::AnalyzerRule;
 use datafusion::sql::unparser::expr_to_sql;
-use datafusion_common::tree_node::TreeNode;
+use datafusion_common::tree_node::{Transformed, TreeNode};
 use datafusion_expr::{DmlStatement, LogicalPlan, WriteOp};
 use datatypes::prelude::ConcreteDataType;
-use datatypes::schema::constraint::NOW_FN;
-use datatypes::schema::{ColumnDefaultConstraint, ColumnSchema};
-use datatypes::value::Value;
+use datatypes::schema::{ColumnSchema, Schema};
 use operator::expr_helper::column_schemas_to_defs;
 use query::query_engine::DefaultSerializer;
 use query::QueryEngineRef;
 use session::context::QueryContextRef;
-use snafu::{OptionExt, ResultExt};
+use snafu::{ensure, OptionExt, ResultExt};
 use substrait::{DFLogicalSubstraitConvertor, SubstraitPlan};
-use table::metadata::RawTableMeta;
 use tokio::sync::oneshot;
 use tokio::sync::oneshot::error::TryRecvError;
 use tokio::time::Instant;
@@ -48,14 +47,16 @@ use crate::batching_mode::frontend_client::FrontendClient;
 use crate::batching_mode::state::TaskState;
 use crate::batching_mode::time_window::TimeWindowExpr;
 use crate::batching_mode::utils::{
-    sql_to_df_plan, AddAutoColumnRewriter, AddFilterRewriter, FindGroupByFinalName,
+    get_table_info_df_schema, sql_to_df_plan, AddAutoColumnRewriter, AddFilterRewriter,
+    FindGroupByFinalName,
 };
 use crate::batching_mode::{
    DEFAULT_BATCHING_ENGINE_QUERY_TIMEOUT, MIN_REFRESH_DURATION, SLOW_QUERY_THRESHOLD,
 };
+use crate::df_optimizer::apply_df_optimizer;
 use crate::error::{
-    ConvertColumnSchemaSnafu, DatafusionSnafu, DatatypesSnafu, ExternalSnafu, InvalidRequestSnafu,
-    SubstraitEncodeLogicalPlanSnafu, TableNotFoundMetaSnafu, TableNotFoundSnafu, UnexpectedSnafu,
+    ConvertColumnSchemaSnafu, DatafusionSnafu, ExternalSnafu, InvalidQuerySnafu,
+    SubstraitEncodeLogicalPlanSnafu, UnexpectedSnafu,
 };
 use crate::metrics::{
    METRIC_FLOW_BATCHING_ENGINE_QUERY_TIME, METRIC_FLOW_BATCHING_ENGINE_SLOW_QUERY,
@@ -73,7 +74,7 @@ pub struct TaskConfig {
    pub expire_after: Option<i64>,
    sink_table_name: [String; 3],
    pub source_table_names: HashSet<[String; 3]>,
-    table_meta: TableMetadataManagerRef,
+    catalog_manager: CatalogManagerRef,
 }

 #[derive(Clone)]
@@ -93,7 +94,7 @@ impl BatchingTask {
        sink_table_name: [String; 3],
        source_table_names: Vec<[String; 3]>,
        query_ctx: QueryContextRef,
-        table_meta: TableMetadataManagerRef,
+        catalog_manager: CatalogManagerRef,
        shutdown_rx: oneshot::Receiver<()>,
    ) -> Self {
        Self {
@@ -105,12 +106,42 @@ impl BatchingTask {
                expire_after,
                sink_table_name,
                source_table_names: source_table_names.into_iter().collect(),
-                table_meta,
+                catalog_manager,
            }),
            state: Arc::new(RwLock::new(TaskState::new(query_ctx, shutdown_rx))),
        }
    }

+    /// mark time window range (now - expire_after, now) as dirty (or (0, now) if expire_after not set)
+    ///
+    /// useful for flush_flow to flush dirty time windows range
+    pub fn mark_all_windows_as_dirty(&self) -> Result<(), Error> {
+        let now = SystemTime::now();
+        let now = Timestamp::new_second(
+            now.duration_since(UNIX_EPOCH)
+                .expect("Time went backwards")
+                .as_secs() as _,
+        );
+        let lower_bound = self
+            .config
+            .expire_after
+            .map(|e| now.sub_duration(Duration::from_secs(e as _)))
+            .transpose()
+            .map_err(BoxedError::new)
+            .context(ExternalSnafu)?
+            .unwrap_or(Timestamp::new_second(0));
+        debug!(
+            "Flow {} mark range ({:?}, {:?}) as dirty",
+            self.config.flow_id, lower_bound, now
+        );
+        self.state
+            .write()
+            .unwrap()
+            .dirty_time_windows
+            .add_window(lower_bound, Some(now));
+        Ok(())
+    }
+
    /// Test execute, for check syntax or such
    pub async fn check_execute(
        &self,
@@ -148,13 +179,8 @@ impl BatchingTask {

    async fn is_table_exist(&self, table_name: &[String; 3]) -> Result<bool, Error> {
        self.config
-            .table_meta
-            .table_name_manager()
-            .exists(TableNameKey {
-                catalog: &table_name[0],
-                schema: &table_name[1],
-                table: &table_name[2],
-            })
+            .catalog_manager
+            .table_exists(&table_name[0], &table_name[1], &table_name[2], None)
            .await
            .map_err(BoxedError::new)
            .context(ExternalSnafu)
@@ -166,8 +192,10 @@ impl BatchingTask {
        frontend_client: &Arc<FrontendClient>,
    ) -> Result<Option<(u32, Duration)>, Error> {
        if let Some(new_query) = self.gen_insert_plan(engine).await? {
+            debug!("Generate new query: {:#?}", new_query);
            self.execute_logical_plan(frontend_client, &new_query).await
        } else {
+            debug!("Generate no query");
            Ok(None)
        }
    }
@@ -176,67 +204,35 @@ impl BatchingTask {
        &self,
        engine: &QueryEngineRef,
    ) -> Result<Option<LogicalPlan>, Error> {
-        let full_table_name = self.config.sink_table_name.clone().join(".");
-
-        let table_id = self
-            .config
-            .table_meta
-            .table_name_manager()
-            .get(common_meta::key::table_name::TableNameKey::new(
-                &self.config.sink_table_name[0],
-                &self.config.sink_table_name[1],
-                &self.config.sink_table_name[2],
-            ))
-            .await
-            .with_context(|_| TableNotFoundMetaSnafu {
-                msg: full_table_name.clone(),
-            })?
-            .map(|t| t.table_id())
-            .with_context(|| TableNotFoundSnafu {
-                name: full_table_name.clone(),
-            })?;
-
-        let table = self
-            .config
-            .table_meta
-            .table_info_manager()
-            .get(table_id)
-            .await
-            .with_context(|_| TableNotFoundMetaSnafu {
-                msg: full_table_name.clone(),
-            })?
-            .with_context(|| TableNotFoundSnafu {
-                name: full_table_name.clone(),
-            })?
-            .into_inner();
-
-        let schema: datatypes::schema::Schema = table
-            .table_info
-            .meta
-            .schema
-            .clone()
-            .try_into()
-            .with_context(|_| DatatypesSnafu {
-                extra: format!(
-                    "Failed to convert schema from raw schema, raw_schema={:?}",
-                    table.table_info.meta.schema
-                ),
-            })?;
-
-        let df_schema = Arc::new(schema.arrow_schema().clone().try_into().with_context(|_| {
-            DatafusionSnafu {
-                context: format!(
-                    "Failed to convert arrow schema to datafusion schema, arrow_schema={:?}",
-                    schema.arrow_schema()
-                ),
-            }
-        })?);
+        let (table, df_schema) = get_table_info_df_schema(
+            self.config.catalog_manager.clone(),
+            self.config.sink_table_name.clone(),
+        )
+        .await?;

        let new_query = self
-            .gen_query_with_time_window(engine.clone(), &table.table_info.meta)
+            .gen_query_with_time_window(engine.clone(), &table.meta.schema)
            .await?;

        let insert_into = if let Some((new_query, _column_cnt)) = new_query {
+            // first check if all columns in input query exists in sink table
+            // since insert into ref to names in record batch generate by given query
+            let table_columns = df_schema
+                .columns()
+                .into_iter()
+                .map(|c| c.name)
+                .collect::<BTreeSet<_>>();
+            for column in new_query.schema().columns() {
+                ensure!(
+                    table_columns.contains(column.name()),
+                    InvalidQuerySnafu {
+                        reason: format!(
+                            "Column {} not found in sink table with columns {:?}",
+                            column, table_columns
+                        ),
+                    }
+                );
+            }
            // update_at& time index placeholder (if exists) should have default value
            LogicalPlan::Dml(DmlStatement::new(
                datafusion_common::TableReference::Full {
@@ -251,6 +247,9 @@ impl BatchingTask {
        } else {
            return Ok(None);
        };
+        let insert_into = insert_into.recompute_schema().context(DatafusionSnafu {
+            context: "Failed to recompute schema",
+        })?;
        Ok(Some(insert_into))
    }

@@ -259,14 +258,11 @@ impl BatchingTask {
        frontend_client: &Arc<FrontendClient>,
        expr: CreateTableExpr,
    ) -> Result<(), Error> {
-        let db_client = frontend_client.get_database_client().await?;
-        db_client
-            .database
-            .create(expr.clone())
-            .await
-            .with_context(|_| InvalidRequestSnafu {
-                context: format!("Failed to create table with expr: {:?}", expr),
-            })?;
+        let catalog = &self.config.sink_table_name[0];
+        let schema = &self.config.sink_table_name[1];
+        frontend_client
+            .create(expr.clone(), catalog, schema)
+            .await?;
        Ok(())
    }

@@ -277,27 +273,78 @@ impl BatchingTask {
    ) -> Result<Option<(u32, Duration)>, Error> {
        let instant = Instant::now();
        let flow_id = self.config.flow_id;
-        let db_client = frontend_client.get_database_client().await?;
-        let peer_addr = db_client.peer.addr;
+
        debug!(
-            "Executing flow {flow_id}(expire_after={:?} secs) on {:?} with query {}",
-            self.config.expire_after, peer_addr, &plan
+            "Executing flow {flow_id}(expire_after={:?} secs) with query {}",
+            self.config.expire_after, &plan
        );

-        let timer = METRIC_FLOW_BATCHING_ENGINE_QUERY_TIME
-            .with_label_values(&[flow_id.to_string().as_str()])
-            .start_timer();
+        let catalog = &self.config.sink_table_name[0];
+        let schema = &self.config.sink_table_name[1];

-        let message = DFLogicalSubstraitConvertor {}
-            .encode(plan, DefaultSerializer)
-            .context(SubstraitEncodeLogicalPlanSnafu)?;
+        // fix all table ref by make it fully qualified, i.e. "table_name" => "catalog_name.schema_name.table_name"
+        let fixed_plan = plan
+            .clone()
+            .transform_down_with_subqueries(|p| {
+                if let LogicalPlan::TableScan(mut table_scan) = p {
+                    let resolved = table_scan.table_name.resolve(catalog, schema);
+                    table_scan.table_name = resolved.into();
+                    Ok(Transformed::yes(LogicalPlan::TableScan(table_scan)))
+                } else {
+                    Ok(Transformed::no(p))
+                }
+            })
+            .with_context(|_| DatafusionSnafu {
+                context: format!("Failed to fix table ref in logical plan, plan={:?}", plan),
+            })?
+            .data;

-        let req = api::v1::greptime_request::Request::Query(api::v1::QueryRequest {
-            query: Some(api::v1::query_request::Query::LogicalPlan(message.to_vec())),
-        });
+        let expanded_plan = CountWildcardRule::new()
+            .analyze(fixed_plan.clone(), &Default::default())
+            .with_context(|_| DatafusionSnafu {
+                context: format!(
+                    "Failed to expand wildcard in logical plan, plan={:?}",
+                    fixed_plan
+                ),
+            })?;

-        let res = db_client.database.handle(req).await;
-        drop(timer);
+        let plan = expanded_plan;
+        let mut peer_desc = None;
+
+        let res = {
+            let _timer = METRIC_FLOW_BATCHING_ENGINE_QUERY_TIME
+                .with_label_values(&[flow_id.to_string().as_str()])
+                .start_timer();
+
+            // hack and special handling the insert logical plan
+            let req = if let Some((insert_to, insert_plan)) =
+                breakup_insert_plan(&plan, catalog, schema)
+            {
+                let message = DFLogicalSubstraitConvertor {}
+                    .encode(&insert_plan, DefaultSerializer)
+                    .context(SubstraitEncodeLogicalPlanSnafu)?;
+                api::v1::greptime_request::Request::Query(api::v1::QueryRequest {
+                    query: Some(api::v1::query_request::Query::InsertIntoPlan(
+                        api::v1::InsertIntoPlan {
+                            table_name: Some(insert_to),
+                            logical_plan: message.to_vec(),
+                        },
+                    )),
+                })
+            } else {
+                let message = DFLogicalSubstraitConvertor {}
+                    .encode(&plan, DefaultSerializer)
+                    .context(SubstraitEncodeLogicalPlanSnafu)?;
+
+                api::v1::greptime_request::Request::Query(api::v1::QueryRequest {
+                    query: Some(api::v1::query_request::Query::LogicalPlan(message.to_vec())),
+                })
+            };
+
+            frontend_client
+                .handle(req, catalog, schema, &mut peer_desc)
+                .await
+        };

        let elapsed = instant.elapsed();
        if let Ok(affected_rows) = &res {
@@ -307,19 +354,23 @@ impl BatchingTask {
            );
        } else if let Err(err) = &res {
            warn!(
-                "Failed to execute Flow {flow_id} on frontend {}, result: {err:?}, elapsed: {:?} with query: {}",
-                peer_addr, elapsed, &plan
+                "Failed to execute Flow {flow_id} on frontend {:?}, result: {err:?}, elapsed: {:?} with query: {}",
+                peer_desc, elapsed, &plan
            );
        }

        // record slow query
        if elapsed >= SLOW_QUERY_THRESHOLD {
            warn!(
-                "Flow {flow_id} on frontend {} executed for {:?} before complete, query: {}",
-                peer_addr, elapsed, &plan
+                "Flow {flow_id} on frontend {:?} executed for {:?} before complete, query: {}",
+                peer_desc, elapsed, &plan
            );
            METRIC_FLOW_BATCHING_ENGINE_SLOW_QUERY
-                .with_label_values(&[flow_id.to_string().as_str(), &plan.to_string(), &peer_addr])
+                .with_label_values(&[
+                    flow_id.to_string().as_str(),
+                    &plan.to_string(),
+                    &peer_desc.unwrap_or_default().to_string(),
+                ])
                .observe(elapsed.as_secs_f64());
        }

@@ -328,12 +379,7 @@ impl BatchingTask {
            .unwrap()
            .after_query_exec(elapsed, res.is_ok());

-        let res = res.context(InvalidRequestSnafu {
-            context: format!(
-                "Failed to execute query for flow={}: \'{}\'",
-                self.config.flow_id, &plan
-            ),
-        })?;
+        let res = res?;

        Ok(Some((res, elapsed)))
    }
@@ -372,7 +418,10 @@ impl BatchingTask {
                            }
                            Err(TryRecvError::Empty) => (),
                        }
-                        state.get_next_start_query_time(Some(DEFAULT_BATCHING_ENGINE_QUERY_TIMEOUT))
+                        state.get_next_start_query_time(
+                            self.config.flow_id,
+                            Some(DEFAULT_BATCHING_ENGINE_QUERY_TIMEOUT),
+                        )
                    };
                    tokio::time::sleep_until(sleep_until).await;
                }
@@ -386,14 +435,18 @@ impl BatchingTask {
                    continue;
                }
                // TODO(discord9): this error should have better place to go, but for now just print error, also more context is needed
-                Err(err) => match new_query {
-                    Some(query) => {
-                        common_telemetry::error!(err; "Failed to execute query for flow={} with query: {query}", self.config.flow_id)
+                Err(err) => {
+                    match new_query {
+                        Some(query) => {
+                            common_telemetry::error!(err; "Failed to execute query for flow={} with query: {query}", self.config.flow_id)
+                        }
+                        None => {
+                            common_telemetry::error!(err; "Failed to generate query for flow={}", self.config.flow_id)
+                        }
                    }
-                    None => {
-                        common_telemetry::error!(err; "Failed to generate query for flow={}", self.config.flow_id)
-                    }
-                },
+                    // also sleep for a little while before try again to prevent flooding logs
+                    tokio::time::sleep(MIN_REFRESH_DURATION).await;
+                }
            }
        }
    }
@@ -418,7 +471,7 @@ impl BatchingTask {
    async fn gen_query_with_time_window(
        &self,
        engine: QueryEngineRef,
-        sink_table_meta: &RawTableMeta,
+        sink_table_schema: &Arc<Schema>,
    ) -> Result<Option<(LogicalPlan, usize)>, Error> {
        let query_ctx = self.state.read().unwrap().query_ctx.clone();
        let start = SystemTime::now();
@@ -477,9 +530,11 @@ impl BatchingTask {
                        debug!(
                            "Flow id = {:?}, can't get window size: precise_lower_bound={expire_time_window_bound:?}, using the same query", self.config.flow_id
                        );
+                        // clean dirty time window too, this could be from create flow's check_execute
+                        self.state.write().unwrap().dirty_time_windows.clean();

                        let mut add_auto_column =
-                            AddAutoColumnRewriter::new(sink_table_meta.schema.clone());
+                            AddAutoColumnRewriter::new(sink_table_schema.clone());
                        let plan = self
                            .config
                            .plan
@@ -487,7 +542,10 @@ impl BatchingTask {
                            .clone()
                            .rewrite(&mut add_auto_column)
                            .with_context(|_| DatafusionSnafu {
-                                context: format!("Failed to rewrite plan {:?}", self.config.plan),
+                                context: format!(
+                                    "Failed to rewrite plan:\n {}\n",
+                                    self.config.plan
+                                ),
                            })?
                            .data;
                        let schema_len = plan.schema().fields().len();
@@ -515,18 +573,23 @@ impl BatchingTask {
                return Ok(None);
            };

+            // TODO(discord9): add auto column or not? This might break compatibility for auto created sink table before this, but that's ok right?
+
            let mut add_filter = AddFilterRewriter::new(expr);
-            let mut add_auto_column = AddAutoColumnRewriter::new(sink_table_meta.schema.clone());
-            // make a not optimized plan for clearer unparse
+            let mut add_auto_column = AddAutoColumnRewriter::new(sink_table_schema.clone());
+
            let plan = sql_to_df_plan(query_ctx.clone(), engine.clone(), &self.config.query, false)
                .await?;
-            plan.clone()
+            let rewrite = plan
+                .clone()
                .rewrite(&mut add_filter)
                .and_then(|p| p.data.rewrite(&mut add_auto_column))
                .with_context(|_| DatafusionSnafu {
-                    context: format!("Failed to rewrite plan {plan:?}"),
+                    context: format!("Failed to rewrite plan:\n {}\n", plan),
                })?
-                .data
+                .data;
+            // only apply optimize after complex rewrite is done
+            apply_df_optimizer(rewrite).await?
        };

        Ok(Some((new_plan, schema_len)))
@@ -534,7 +597,7 @@ impl BatchingTask {
 }

 // auto created table have a auto added column `update_at`, and optional have a `AUTO_CREATED_PLACEHOLDER_TS_COL` column for time index placeholder if no timestamp column is specified
-// TODO(discord9): unit test
+// TODO(discord9): for now no default value is set for auto added column for compatibility reason with streaming mode, but this might change in favor of simpler code?
 fn create_table_with_expr(
    plan: &LogicalPlan,
    sink_table_name: &[String; 3],
@@ -558,11 +621,7 @@ fn create_table_with_expr(
        AUTO_CREATED_UPDATE_AT_TS_COL,
        ConcreteDataType::timestamp_millisecond_datatype(),
        true,
-    )
-    .with_default_constraint(Some(ColumnDefaultConstraint::Function(NOW_FN.to_string())))
-    .context(DatatypesSnafu {
-        extra: "Failed to build column `update_at TimestampMillisecond default now()`",
-    })?;
+    );
    column_schemas.push(update_at_schema);

    let time_index = if let Some(time_index) = first_time_stamp {
@@ -574,16 +633,7 @@ fn create_table_with_expr(
                ConcreteDataType::timestamp_millisecond_datatype(),
                false,
            )
-            .with_time_index(true)
-            .with_default_constraint(Some(ColumnDefaultConstraint::Value(Value::Timestamp(
-                Timestamp::new_millisecond(0),
-            ))))
-            .context(DatatypesSnafu {
-                extra: format!(
-                    "Failed to build column `{} TimestampMillisecond TIME INDEX default 0`",
-                    AUTO_CREATED_PLACEHOLDER_TS_COL
-                ),
-            })?,
+            .with_time_index(true),
        );
        AUTO_CREATED_PLACEHOLDER_TS_COL.to_string()
    };
@@ -675,20 +725,14 @@ mod test {
            AUTO_CREATED_UPDATE_AT_TS_COL,
            ConcreteDataType::timestamp_millisecond_datatype(),
            true,
-        )
-        .with_default_constraint(Some(ColumnDefaultConstraint::Function(NOW_FN.to_string())))
-        .unwrap();
+        );

        let ts_placeholder_schema = ColumnSchema::new(
            AUTO_CREATED_PLACEHOLDER_TS_COL,
            ConcreteDataType::timestamp_millisecond_datatype(),
            false,
        )
-        .with_time_index(true)
-        .with_default_constraint(Some(ColumnDefaultConstraint::Value(Value::Timestamp(
-            Timestamp::new_millisecond(0),
-        ))))
-        .unwrap();
+        .with_time_index(true);

        let testcases = vec![
            TestCase {
--- a/src/flow/src/batching_mode/time_window.rs
+++ b/src/flow/src/batching_mode/time_window.rs
@@ -72,6 +72,17 @@ pub struct TimeWindowExpr {
    df_schema: DFSchema,
 }

+impl std::fmt::Display for TimeWindowExpr {
+    fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
+        f.debug_struct("TimeWindowExpr")
+            .field("phy_expr", &self.phy_expr.to_string())
+            .field("column_name", &self.column_name)
+            .field("logical_expr", &self.logical_expr.to_string())
+            .field("df_schema", &self.df_schema)
+            .finish()
+    }
+}
+
 impl TimeWindowExpr {
    pub fn from_expr(
        expr: &Expr,
@@ -256,7 +267,7 @@ fn columnar_to_ts_vector(columnar: &ColumnarValue) -> Result<Vec<Option<Timestam
    Ok(val)
 }

-/// Return (the column name of time index column, the time window expr, the expected time unit of time index column, the expr's schema for evaluating the time window)
+/// Return (`the column name of time index column`, `the time window expr`, `the expected time unit of time index column`, `the expr's schema for evaluating the time window`)
 ///
 /// The time window expr is expected to have one input column with Timestamp type, and also return Timestamp type, the time window expr is expected
 /// to be monotonic increasing and appears in the innermost GROUP BY clause
@@ -693,6 +704,28 @@ mod test {
                ),
                "SELECT arrow_cast(date_bin(INTERVAL '1 MINS', numbers_with_ts.ts), 'Timestamp(Second, None)') AS time_window FROM numbers_with_ts WHERE ((ts >= CAST('2025-02-24 10:48:00' AS TIMESTAMP)) AND (ts <= CAST('2025-02-24 10:49:00' AS TIMESTAMP))) GROUP BY arrow_cast(date_bin(INTERVAL '1 MINS', numbers_with_ts.ts), 'Timestamp(Second, None)')"
            ),
+            // complex time window index with where
+            (
+                "SELECT arrow_cast(date_bin(INTERVAL '1 MINS', numbers_with_ts.ts), 'Timestamp(Second, None)') AS time_window FROM numbers_with_ts WHERE number in (2, 3, 4) GROUP BY time_window;",
+                Timestamp::new(1740394109, TimeUnit::Second),
+                (
+                    "ts".to_string(),
+                    Some(Timestamp::new(1740394080, TimeUnit::Second)),
+                    Some(Timestamp::new(1740394140, TimeUnit::Second)),
+                ),
+                "SELECT arrow_cast(date_bin(INTERVAL '1 MINS', numbers_with_ts.ts), 'Timestamp(Second, None)') AS time_window FROM numbers_with_ts WHERE numbers_with_ts.number IN (2, 3, 4) AND ((ts >= CAST('2025-02-24 10:48:00' AS TIMESTAMP)) AND (ts <= CAST('2025-02-24 10:49:00' AS TIMESTAMP))) GROUP BY arrow_cast(date_bin(INTERVAL '1 MINS', numbers_with_ts.ts), 'Timestamp(Second, None)')"
+            ),
+            // complex time window index with between and
+            (
+                "SELECT arrow_cast(date_bin(INTERVAL '1 MINS', numbers_with_ts.ts), 'Timestamp(Second, None)') AS time_window FROM numbers_with_ts WHERE number BETWEEN 2 AND 4 GROUP BY time_window;",
+                Timestamp::new(1740394109, TimeUnit::Second),
+                (
+                    "ts".to_string(),
+                    Some(Timestamp::new(1740394080, TimeUnit::Second)),
+                    Some(Timestamp::new(1740394140, TimeUnit::Second)),
+                ),
+                "SELECT arrow_cast(date_bin(INTERVAL '1 MINS', numbers_with_ts.ts), 'Timestamp(Second, None)') AS time_window FROM numbers_with_ts WHERE (numbers_with_ts.number BETWEEN 2 AND 4) AND ((ts >= CAST('2025-02-24 10:48:00' AS TIMESTAMP)) AND (ts <= CAST('2025-02-24 10:49:00' AS TIMESTAMP))) GROUP BY arrow_cast(date_bin(INTERVAL '1 MINS', numbers_with_ts.ts), 'Timestamp(Second, None)')"
+            ),
            // no time index
            (
                "SELECT date_bin('5 minutes', ts) FROM numbers_with_ts;",
--- a/src/flow/src/batching_mode/utils.rs
+++ b/src/flow/src/batching_mode/utils.rs
@@ -14,29 +14,63 @@

 //! some utils for helping with batching mode

-use std::collections::HashSet;
+use std::collections::{BTreeSet, HashSet};
 use std::sync::Arc;

+use catalog::CatalogManagerRef;
 use common_error::ext::BoxedError;
-use common_telemetry::{debug, info};
+use common_telemetry::debug;
 use datafusion::error::Result as DfResult;
 use datafusion::logical_expr::Expr;
 use datafusion::sql::unparser::Unparser;
 use datafusion_common::tree_node::{
    Transformed, TreeNodeRecursion, TreeNodeRewriter, TreeNodeVisitor,
 };
-use datafusion_common::DataFusionError;
-use datafusion_expr::{Distinct, LogicalPlan};
-use datatypes::schema::RawSchema;
+use datafusion_common::{DFSchema, DataFusionError, ScalarValue};
+use datafusion_expr::{Distinct, LogicalPlan, Projection};
+use datatypes::schema::SchemaRef;
 use query::parser::QueryLanguageParser;
 use query::QueryEngineRef;
 use session::context::QueryContextRef;
-use snafu::ResultExt;
+use snafu::{OptionExt, ResultExt};
+use table::metadata::TableInfo;

 use crate::adapter::AUTO_CREATED_PLACEHOLDER_TS_COL;
 use crate::df_optimizer::apply_df_optimizer;
-use crate::error::{DatafusionSnafu, ExternalSnafu};
-use crate::Error;
+use crate::error::{DatafusionSnafu, ExternalSnafu, TableNotFoundSnafu};
+use crate::{Error, TableName};
+
+pub async fn get_table_info_df_schema(
+    catalog_mr: CatalogManagerRef,
+    table_name: TableName,
+) -> Result<(Arc<TableInfo>, Arc<DFSchema>), Error> {
+    let full_table_name = table_name.clone().join(".");
+    let table = catalog_mr
+        .table(&table_name[0], &table_name[1], &table_name[2], None)
+        .await
+        .map_err(BoxedError::new)
+        .context(ExternalSnafu)?
+        .context(TableNotFoundSnafu {
+            name: &full_table_name,
+        })?;
+    let table_info = table.table_info().clone();
+
+    let schema = table_info.meta.schema.clone();
+
+    let df_schema: Arc<DFSchema> = Arc::new(
+        schema
+            .arrow_schema()
+            .clone()
+            .try_into()
+            .with_context(|_| DatafusionSnafu {
+                context: format!(
+                    "Failed to convert arrow schema to datafusion schema, arrow_schema={:?}",
+                    schema.arrow_schema()
+                ),
+            })?,
+    );
+    Ok((table_info, df_schema))
+}

 /// Convert sql to datafusion logical plan
 pub async fn sql_to_df_plan(
@@ -164,14 +198,16 @@ impl TreeNodeVisitor<'_> for FindGroupByFinalName {
 /// (which doesn't necessary need to have exact name just need to be a extra timestamp column)
 /// and `__ts_placeholder`(this column need to have exact this name and be a timestamp)
 /// with values like `now()` and `0`
+///
+/// it also give existing columns alias to column in sink table if needed
 #[derive(Debug)]
 pub struct AddAutoColumnRewriter {
-    pub schema: RawSchema,
+    pub schema: SchemaRef,
    pub is_rewritten: bool,
 }

 impl AddAutoColumnRewriter {
-    pub fn new(schema: RawSchema) -> Self {
+    pub fn new(schema: SchemaRef) -> Self {
        Self {
            schema,
            is_rewritten: false,
@@ -181,37 +217,97 @@ impl AddAutoColumnRewriter {

 impl TreeNodeRewriter for AddAutoColumnRewriter {
    type Node = LogicalPlan;
-    fn f_down(&mut self, node: Self::Node) -> DfResult<Transformed<Self::Node>> {
+    fn f_down(&mut self, mut node: Self::Node) -> DfResult<Transformed<Self::Node>> {
        if self.is_rewritten {
            return Ok(Transformed::no(node));
        }

-        // if is distinct all, go one level down
-        if let LogicalPlan::Distinct(Distinct::All(_)) = node {
-            return Ok(Transformed::no(node));
+        // if is distinct all, wrap it in a projection
+        if let LogicalPlan::Distinct(Distinct::All(_)) = &node {
+            let mut exprs = vec![];
+
+            for field in node.schema().fields().iter() {
+                exprs.push(Expr::Column(datafusion::common::Column::new_unqualified(
+                    field.name(),
+                )));
+            }
+
+            let projection =
+                LogicalPlan::Projection(Projection::try_new(exprs, Arc::new(node.clone()))?);
+
+            node = projection;
+        }
+        // handle table_scan by wrap it in a projection
+        else if let LogicalPlan::TableScan(table_scan) = node {
+            let mut exprs = vec![];
+
+            for field in table_scan.projected_schema.fields().iter() {
+                exprs.push(Expr::Column(datafusion::common::Column::new(
+                    Some(table_scan.table_name.clone()),
+                    field.name(),
+                )));
+            }
+
+            let projection = LogicalPlan::Projection(Projection::try_new(
+                exprs,
+                Arc::new(LogicalPlan::TableScan(table_scan)),
+            )?);
+
+            node = projection;
        }

-        // FIXME(discord9): just read plan.expr and do stuffs
-        let mut exprs = node.expressions();
+        // only do rewrite if found the outermost projection
+        let mut exprs = if let LogicalPlan::Projection(project) = &node {
+            project.expr.clone()
+        } else {
+            return Ok(Transformed::no(node));
+        };
+
+        let all_names = self
+            .schema
+            .column_schemas()
+            .iter()
+            .map(|c| c.name.clone())
+            .collect::<BTreeSet<_>>();
+        // first match by position
+        for (idx, expr) in exprs.iter_mut().enumerate() {
+            if !all_names.contains(&expr.qualified_name().1) {
+                if let Some(col_name) = self
+                    .schema
+                    .column_schemas()
+                    .get(idx)
+                    .map(|c| c.name.clone())
+                {
+                    // if the data type mismatched, later check_execute will error out
+                    // hence no need to check it here, beside, optimize pass might be able to cast it
+                    // so checking here is not necessary
+                    *expr = expr.clone().alias(col_name);
+                }
+            }
+        }

        // add columns if have different column count
        let query_col_cnt = exprs.len();
-        let table_col_cnt = self.schema.column_schemas.len();
-        info!("query_col_cnt={query_col_cnt}, table_col_cnt={table_col_cnt}");
+        let table_col_cnt = self.schema.column_schemas().len();
+        debug!("query_col_cnt={query_col_cnt}, table_col_cnt={table_col_cnt}");
+
+        let placeholder_ts_expr =
+            datafusion::logical_expr::lit(ScalarValue::TimestampMillisecond(Some(0), None))
+                .alias(AUTO_CREATED_PLACEHOLDER_TS_COL);
+
        if query_col_cnt == table_col_cnt {
-            self.is_rewritten = true;
-            return Ok(Transformed::no(node));
+            // still need to add alias, see below
        } else if query_col_cnt + 1 == table_col_cnt {
-            let last_col_schema = self.schema.column_schemas.last().unwrap();
+            let last_col_schema = self.schema.column_schemas().last().unwrap();

            // if time index column is auto created add it
            if last_col_schema.name == AUTO_CREATED_PLACEHOLDER_TS_COL
-                && self.schema.timestamp_index == Some(table_col_cnt - 1)
+                && self.schema.timestamp_index() == Some(table_col_cnt - 1)
            {
-                exprs.push(datafusion::logical_expr::lit(0));
+                exprs.push(placeholder_ts_expr);
            } else if last_col_schema.data_type.is_timestamp() {
                // is the update at column
-                exprs.push(datafusion::prelude::now());
+                exprs.push(datafusion::prelude::now().alias(&last_col_schema.name));
            } else {
                // helpful error message
                return Err(DataFusionError::Plan(format!(
@@ -221,11 +317,11 @@ impl TreeNodeRewriter for AddAutoColumnRewriter {
                )));
            }
        } else if query_col_cnt + 2 == table_col_cnt {
-            let mut col_iter = self.schema.column_schemas.iter().rev();
+            let mut col_iter = self.schema.column_schemas().iter().rev();
            let last_col_schema = col_iter.next().unwrap();
            let second_last_col_schema = col_iter.next().unwrap();
            if second_last_col_schema.data_type.is_timestamp() {
-                exprs.push(datafusion::prelude::now());
+                exprs.push(datafusion::prelude::now().alias(&second_last_col_schema.name));
            } else {
                return Err(DataFusionError::Plan(format!(
                    "Expect the second last column in the table to be timestamp column, found column {} with type {:?}",
@@ -235,9 +331,9 @@ impl TreeNodeRewriter for AddAutoColumnRewriter {
            }

            if last_col_schema.name == AUTO_CREATED_PLACEHOLDER_TS_COL
-                && self.schema.timestamp_index == Some(table_col_cnt - 1)
+                && self.schema.timestamp_index() == Some(table_col_cnt - 1)
            {
-                exprs.push(datafusion::logical_expr::lit(0));
+                exprs.push(placeholder_ts_expr);
            } else {
                return Err(DataFusionError::Plan(format!(
                    "Expect timestamp column {}, found {:?}",
@@ -247,7 +343,7 @@ impl TreeNodeRewriter for AddAutoColumnRewriter {
        } else {
            return Err(DataFusionError::Plan(format!(
                    "Expect table have 0,1 or 2 columns more than query columns, found {} query columns {:?}, {} table columns {:?}",
-                    query_col_cnt, node.expressions(), table_col_cnt, self.schema.column_schemas
+                    query_col_cnt, exprs, table_col_cnt, self.schema.column_schemas()
                )));
        }

@@ -255,9 +351,12 @@ impl TreeNodeRewriter for AddAutoColumnRewriter {
        let new_plan = node.with_new_exprs(exprs, node.inputs().into_iter().cloned().collect())?;
        Ok(Transformed::yes(new_plan))
    }
-}

-// TODO(discord9): a method to found out the precise time window
+    /// We might add new columns, so we need to recompute the schema
+    fn f_up(&mut self, node: Self::Node) -> DfResult<Transformed<Self::Node>> {
+        node.recompute_schema().map(Transformed::yes)
+    }
+}

 /// Find out the `Filter` Node corresponding to innermost(deepest) `WHERE` and add a new filter expr to it
 #[derive(Debug)]
@@ -301,11 +400,15 @@ impl TreeNodeRewriter for AddFilterRewriter {

 #[cfg(test)]
 mod test {
+    use std::sync::Arc;
+
    use datafusion_common::tree_node::TreeNode as _;
    use datatypes::prelude::ConcreteDataType;
-    use datatypes::schema::ColumnSchema;
+    use datatypes::schema::{ColumnSchema, Schema};
    use pretty_assertions::assert_eq;
+    use query::query_engine::DefaultSerializer;
    use session::context::QueryContext;
+    use substrait::{DFLogicalSubstraitConvertor, SubstraitPlan};

    use super::*;
    use crate::test_utils::create_test_query_engine;
@@ -386,7 +489,7 @@ mod test {
            // add update_at
            (
                "SELECT number FROM numbers_with_ts",
-                Ok("SELECT numbers_with_ts.number, now() FROM numbers_with_ts"),
+                Ok("SELECT numbers_with_ts.number, now() AS ts FROM numbers_with_ts"),
                vec![
                    ColumnSchema::new("number", ConcreteDataType::int32_datatype(), true),
                    ColumnSchema::new(
@@ -400,7 +503,7 @@ mod test {
            // add ts placeholder
            (
                "SELECT number FROM numbers_with_ts",
-                Ok("SELECT numbers_with_ts.number, 0 FROM numbers_with_ts"),
+                Ok("SELECT numbers_with_ts.number, CAST('1970-01-01 00:00:00' AS TIMESTAMP) AS __ts_placeholder FROM numbers_with_ts"),
                vec![
                    ColumnSchema::new("number", ConcreteDataType::int32_datatype(), true),
                    ColumnSchema::new(
@@ -428,7 +531,7 @@ mod test {
            // add update_at and ts placeholder
            (
                "SELECT number FROM numbers_with_ts",
-                Ok("SELECT numbers_with_ts.number, now(), 0 FROM numbers_with_ts"),
+                Ok("SELECT numbers_with_ts.number, now() AS update_at, CAST('1970-01-01 00:00:00' AS TIMESTAMP) AS __ts_placeholder FROM numbers_with_ts"),
                vec![
                    ColumnSchema::new("number", ConcreteDataType::int32_datatype(), true),
                    ColumnSchema::new(
@@ -447,7 +550,7 @@ mod test {
            // add ts placeholder
            (
                "SELECT number, ts FROM numbers_with_ts",
-                Ok("SELECT numbers_with_ts.number, numbers_with_ts.ts, 0 FROM numbers_with_ts"),
+                Ok("SELECT numbers_with_ts.number, numbers_with_ts.ts AS update_at, CAST('1970-01-01 00:00:00' AS TIMESTAMP) AS __ts_placeholder FROM numbers_with_ts"),
                vec![
                    ColumnSchema::new("number", ConcreteDataType::int32_datatype(), true),
                    ColumnSchema::new(
@@ -466,7 +569,7 @@ mod test {
            // add update_at after time index column
            (
                "SELECT number, ts FROM numbers_with_ts",
-                Ok("SELECT numbers_with_ts.number, numbers_with_ts.ts, now() FROM numbers_with_ts"),
+                Ok("SELECT numbers_with_ts.number, numbers_with_ts.ts, now() AS update_atat FROM numbers_with_ts"),
                vec![
                    ColumnSchema::new("number", ConcreteDataType::int32_datatype(), true),
                    ColumnSchema::new(
@@ -528,8 +631,8 @@ mod test {
        let query_engine = create_test_query_engine();
        let ctx = QueryContext::arc();
        for (before, after, column_schemas) in testcases {
-            let raw_schema = RawSchema::new(column_schemas);
-            let mut add_auto_column_rewriter = AddAutoColumnRewriter::new(raw_schema);
+            let schema = Arc::new(Schema::new(column_schemas));
+            let mut add_auto_column_rewriter = AddAutoColumnRewriter::new(schema);

            let plan = sql_to_df_plan(ctx.clone(), query_engine.clone(), before, false)
                .await
@@ -600,4 +703,18 @@ mod test {
            );
        }
    }
+
+    #[tokio::test]
+    async fn test_null_cast() {
+        let query_engine = create_test_query_engine();
+        let ctx = QueryContext::arc();
+        let sql = "SELECT NULL::DOUBLE FROM numbers_with_ts";
+        let plan = sql_to_df_plan(ctx, query_engine.clone(), sql, false)
+            .await
+            .unwrap();
+
+        let _sub_plan = DFLogicalSubstraitConvertor {}
+            .encode(&plan, DefaultSerializer)
+            .unwrap();
+    }
 }
--- a/src/flow/src/df_optimizer.rs
+++ b/src/flow/src/df_optimizer.rs
@@ -25,7 +25,6 @@ use datafusion::config::ConfigOptions;
 use datafusion::error::DataFusionError;
 use datafusion::functions_aggregate::count::count_udaf;
 use datafusion::functions_aggregate::sum::sum_udaf;
-use datafusion::optimizer::analyzer::count_wildcard_rule::CountWildcardRule;
 use datafusion::optimizer::analyzer::type_coercion::TypeCoercion;
 use datafusion::optimizer::common_subexpr_eliminate::CommonSubexprEliminate;
 use datafusion::optimizer::optimize_projections::OptimizeProjections;
@@ -42,6 +41,7 @@ use datafusion_expr::{
    BinaryExpr, ColumnarValue, Expr, Operator, Projection, ScalarFunctionArgs, ScalarUDFImpl,
    Signature, TypeSignature, Volatility,
 };
+use query::optimizer::count_wildcard::CountWildcardToTimeIndexRule;
 use query::parser::QueryLanguageParser;
 use query::query_engine::DefaultSerializer;
 use query::QueryEngine;
@@ -61,9 +61,9 @@ pub async fn apply_df_optimizer(
 ) -> Result<datafusion_expr::LogicalPlan, Error> {
    let cfg = ConfigOptions::new();
    let analyzer = Analyzer::with_rules(vec![
-        Arc::new(CountWildcardRule::new()),
-        Arc::new(AvgExpandRule::new()),
-        Arc::new(TumbleExpandRule::new()),
+        Arc::new(CountWildcardToTimeIndexRule),
+        Arc::new(AvgExpandRule),
+        Arc::new(TumbleExpandRule),
        Arc::new(CheckGroupByRule::new()),
        Arc::new(TypeCoercion::new()),
    ]);
@@ -128,13 +128,7 @@ pub async fn sql_to_flow_plan(
 }

 #[derive(Debug)]
-struct AvgExpandRule {}
-
-impl AvgExpandRule {
-    pub fn new() -> Self {
-        Self {}
-    }
-}
+struct AvgExpandRule;

 impl AnalyzerRule for AvgExpandRule {
    fn analyze(
@@ -331,13 +325,7 @@ impl TreeNodeRewriter for ExpandAvgRewriter<'_> {

 /// expand tumble in aggr expr to tumble_start and tumble_end with column name like `window_start`
 #[derive(Debug)]
-struct TumbleExpandRule {}
-
-impl TumbleExpandRule {
-    pub fn new() -> Self {
-        Self {}
-    }
-}
+struct TumbleExpandRule;

 impl AnalyzerRule for TumbleExpandRule {
    fn analyze(
--- a/src/flow/src/engine.rs
+++ b/src/flow/src/engine.rs
@@ -49,6 +49,8 @@ pub trait FlowEngine {
    async fn flush_flow(&self, flow_id: FlowId) -> Result<usize, Error>;
    /// Check if the flow exists
    async fn flow_exist(&self, flow_id: FlowId) -> Result<bool, Error>;
+    /// List all flows
+    async fn list_flows(&self) -> Result<impl IntoIterator<Item = FlowId>, Error>;
    /// Handle the insert requests for the flow
    async fn handle_flow_inserts(
        &self,
--- a/src/flow/src/error.rs
+++ b/src/flow/src/error.rs
@@ -149,6 +149,13 @@ pub enum Error {
        location: Location,
    },

+    #[snafu(display("Unsupported: {reason}"))]
+    Unsupported {
+        reason: String,
+        #[snafu(implicit)]
+        location: Location,
+    },
+
    #[snafu(display("Unsupported temporal filter: {reason}"))]
    UnsupportedTemporalFilter {
        reason: String,
@@ -189,6 +196,25 @@ pub enum Error {
        location: Location,
    },

+    #[snafu(display("Illegal check task state: {reason}"))]
+    IllegalCheckTaskState {
+        reason: String,
+        #[snafu(implicit)]
+        location: Location,
+    },
+
+    #[snafu(display(
+        "Failed to sync with check task for flow {} with allow_drop={}",
+        flow_id,
+        allow_drop
+    ))]
+    SyncCheckTask {
+        flow_id: FlowId,
+        allow_drop: bool,
+        #[snafu(implicit)]
+        location: Location,
+    },
+
    #[snafu(display("Failed to start server"))]
    StartServer {
        #[snafu(implicit)]
@@ -280,10 +306,12 @@ impl ErrorExt for Error {
            Self::CreateFlow { .. } | Self::Arrow { .. } | Self::Time { .. } => {
                StatusCode::EngineExecuteQuery
            }
-            Self::Unexpected { .. } => StatusCode::Unexpected,
-            Self::NotImplemented { .. } | Self::UnsupportedTemporalFilter { .. } => {
-                StatusCode::Unsupported
-            }
+            Self::Unexpected { .. }
+            | Self::SyncCheckTask { .. }
+            | Self::IllegalCheckTaskState { .. } => StatusCode::Unexpected,
+            Self::NotImplemented { .. }
+            | Self::UnsupportedTemporalFilter { .. }
+            | Self::Unsupported { .. } => StatusCode::Unsupported,
            Self::External { source, .. } => source.status_code(),
            Self::Internal { .. } | Self::CacheRequired { .. } => StatusCode::Internal,
            Self::StartServer { source, .. } | Self::ShutdownServer { source, .. } => {
--- a/src/flow/src/lib.rs
+++ b/src/flow/src/lib.rs
@@ -43,8 +43,8 @@ mod utils;
 #[cfg(test)]
 mod test_utils;

-pub use adapter::{FlowConfig, FlowWorkerManager, FlowWorkerManagerRef, FlownodeOptions};
-pub use batching_mode::frontend_client::FrontendClient;
+pub use adapter::{FlowConfig, FlowStreamingEngineRef, FlownodeOptions, StreamingEngine};
+pub use batching_mode::frontend_client::{FrontendClient, GrpcQueryHandlerWithBoxedError};
 pub(crate) use engine::{CreateFlowArgs, FlowId, TableName};
 pub use error::{Error, Result};
 pub use server::{
--- a/src/flow/src/server.rs
+++ b/src/flow/src/server.rs
@@ -29,6 +29,7 @@ use common_meta::key::TableMetadataManagerRef;
 use common_meta::kv_backend::KvBackendRef;
 use common_meta::node_manager::{Flownode, NodeManagerRef};
 use common_query::Output;
+use common_runtime::JoinHandle;
 use common_telemetry::tracing::info;
 use futures::{FutureExt, TryStreamExt};
 use greptime_proto::v1::flow::{flow_server, FlowRequest, FlowResponse, InsertRequests};
@@ -50,7 +51,10 @@ use tonic::codec::CompressionEncoding;
 use tonic::transport::server::TcpIncoming;
 use tonic::{Request, Response, Status};

-use crate::adapter::{create_worker, FlowWorkerManagerRef};
+use crate::adapter::flownode_impl::{FlowDualEngine, FlowDualEngineRef};
+use crate::adapter::{create_worker, FlowStreamingEngineRef};
+use crate::batching_mode::engine::BatchingEngine;
+use crate::engine::FlowEngine;
 use crate::error::{
    to_status_with_last_err, CacheRequiredSnafu, CreateFlowSnafu, ExternalSnafu, FlowNotFoundSnafu,
    ListFlowsSnafu, ParseAddrSnafu, ShutdownServerSnafu, StartServerSnafu, UnexpectedSnafu,
@@ -59,19 +63,20 @@ use crate::heartbeat::HeartbeatTask;
 use crate::metrics::{METRIC_FLOW_PROCESSING_TIME, METRIC_FLOW_ROWS};
 use crate::transform::register_function_to_query_engine;
 use crate::utils::{SizeReportSender, StateReportHandler};
-use crate::{CreateFlowArgs, Error, FlowWorkerManager, FlownodeOptions, FrontendClient};
+use crate::{CreateFlowArgs, Error, FlownodeOptions, FrontendClient, StreamingEngine};

 pub const FLOW_NODE_SERVER_NAME: &str = "FLOW_NODE_SERVER";
 /// wrapping flow node manager to avoid orphan rule with Arc<...>
 #[derive(Clone)]
 pub struct FlowService {
-    /// TODO(discord9): replace with dual engine
-    pub manager: FlowWorkerManagerRef,
+    pub dual_engine: FlowDualEngineRef,
 }

 impl FlowService {
-    pub fn new(manager: FlowWorkerManagerRef) -> Self {
-        Self { manager }
+    pub fn new(manager: FlowDualEngineRef) -> Self {
+        Self {
+            dual_engine: manager,
+        }
    }
 }

@@ -86,7 +91,7 @@ impl flow_server::Flow for FlowService {
            .start_timer();

        let request = request.into_inner();
-        self.manager
+        self.dual_engine
            .handle(request)
            .await
            .map_err(|err| {
@@ -126,7 +131,7 @@ impl flow_server::Flow for FlowService {
            .with_label_values(&["in"])
            .inc_by(row_count as u64);

-        self.manager
+        self.dual_engine
            .handle_inserts(request)
            .await
            .map(Response::new)
@@ -139,11 +144,16 @@ pub struct FlownodeServer {
    inner: Arc<FlownodeServerInner>,
 }

+/// FlownodeServerInner is the inner state of FlownodeServer,
+/// this struct mostly useful for construct/start and stop the
+/// flow node server
 struct FlownodeServerInner {
    /// worker shutdown signal, not to be confused with server_shutdown_tx
    worker_shutdown_tx: Mutex<broadcast::Sender<()>>,
    /// server shutdown signal for shutdown grpc server
    server_shutdown_tx: Mutex<broadcast::Sender<()>>,
+    /// streaming task handler
+    streaming_task_handler: Mutex<Option<JoinHandle<()>>>,
    flow_service: FlowService,
 }

@@ -156,16 +166,28 @@ impl FlownodeServer {
                flow_service,
                worker_shutdown_tx: Mutex::new(tx),
                server_shutdown_tx: Mutex::new(server_tx),
+                streaming_task_handler: Mutex::new(None),
            }),
        }
    }

    /// Start the background task for streaming computation.
    async fn start_workers(&self) -> Result<(), Error> {
-        let manager_ref = self.inner.flow_service.manager.clone();
-        let _handle = manager_ref
-            .clone()
+        let manager_ref = self.inner.flow_service.dual_engine.clone();
+        let handle = manager_ref
+            .streaming_engine()
            .run_background(Some(self.inner.worker_shutdown_tx.lock().await.subscribe()));
+        self.inner
+            .streaming_task_handler
+            .lock()
+            .await
+            .replace(handle);
+
+        self.inner
+            .flow_service
+            .dual_engine
+            .start_flow_consistent_check_task()
+            .await?;

        Ok(())
    }
@@ -176,6 +198,11 @@ impl FlownodeServer {
        if tx.send(()).is_err() {
            info!("Receiver dropped, the flow node server has already shutdown");
        }
+        self.inner
+            .flow_service
+            .dual_engine
+            .stop_flow_consistent_check_task()
+            .await?;
        Ok(())
    }
 }
@@ -272,8 +299,8 @@ impl FlownodeInstance {
        &self.flownode_server
    }

-    pub fn flow_worker_manager(&self) -> FlowWorkerManagerRef {
-        self.flownode_server.inner.flow_service.manager.clone()
+    pub fn flow_engine(&self) -> FlowDualEngineRef {
+        self.flownode_server.inner.flow_service.dual_engine.clone()
    }

    pub fn setup_services(&mut self, services: ServerHandlers) {
@@ -342,12 +369,21 @@ impl FlownodeBuilder {
            self.build_manager(query_engine_factory.query_engine())
                .await?,
        );
+        let batching = Arc::new(BatchingEngine::new(
+            self.frontend_client.clone(),
+            query_engine_factory.query_engine(),
+            self.flow_metadata_manager.clone(),
+            self.table_meta.clone(),
+            self.catalog_manager.clone(),
+        ));
+        let dual = FlowDualEngine::new(
+            manager.clone(),
+            batching,
+            self.flow_metadata_manager.clone(),
+            self.catalog_manager.clone(),
+        );

-        if let Err(err) = self.recover_flows(&manager).await {
-            common_telemetry::error!(err; "Failed to recover flows");
-        }
-
-        let server = FlownodeServer::new(FlowService::new(manager.clone()));
+        let server = FlownodeServer::new(FlowService::new(Arc::new(dual)));

        let heartbeat_task = self.heartbeat_task;

@@ -364,7 +400,7 @@ impl FlownodeBuilder {
    /// or recover all existing flow tasks if in standalone mode(nodeid is None)
    ///
    /// TODO(discord9): persistent flow tasks with internal state
-    async fn recover_flows(&self, manager: &FlowWorkerManagerRef) -> Result<usize, Error> {
+    async fn recover_flows(&self, manager: &FlowDualEngine) -> Result<usize, Error> {
        let nodeid = self.opts.node_id;
        let to_be_recovered: Vec<_> = if let Some(nodeid) = nodeid {
            let to_be_recover = self
@@ -401,6 +437,7 @@ impl FlownodeBuilder {
        let cnt = to_be_recovered.len();

        // TODO(discord9): recover in parallel
+        info!("Recovering {} flows: {:?}", cnt, to_be_recovered);
        for flow_id in to_be_recovered {
            let info = self
                .flow_metadata_manager
@@ -416,6 +453,7 @@ impl FlownodeBuilder {
                info.sink_table_name().schema_name.clone(),
                info.sink_table_name().table_name.clone(),
            ];
+
            let args = CreateFlowArgs {
                flow_id: flow_id as _,
                sink_table_name,
@@ -429,14 +467,27 @@ impl FlownodeBuilder {
                comment: Some(info.comment().clone()),
                sql: info.raw_sql().clone(),
                flow_options: info.options().clone(),
-                query_ctx: Some(
-                    QueryContextBuilder::default()
-                        .current_catalog(info.catalog_name().clone())
-                        .build(),
-                ),
+                query_ctx: info
+                    .query_context()
+                    .clone()
+                    .map(|ctx| {
+                        ctx.try_into()
+                            .map_err(BoxedError::new)
+                            .context(ExternalSnafu)
+                    })
+                    .transpose()?
+                    // or use default QueryContext with catalog_name from info
+                    // to keep compatibility with old version
+                    .or_else(|| {
+                        Some(
+                            QueryContextBuilder::default()
+                                .current_catalog(info.catalog_name().to_string())
+                                .build(),
+                        )
+                    }),
            };
            manager
-                .create_flow_inner(args)
+                .create_flow(args)
                .await
                .map_err(BoxedError::new)
                .with_context(|_| CreateFlowSnafu {
@@ -452,7 +503,7 @@ impl FlownodeBuilder {
    async fn build_manager(
        &mut self,
        query_engine: Arc<dyn QueryEngine>,
-    ) -> Result<FlowWorkerManager, Error> {
+    ) -> Result<StreamingEngine, Error> {
        let table_meta = self.table_meta.clone();

        register_function_to_query_engine(&query_engine);
@@ -461,7 +512,7 @@ impl FlownodeBuilder {

        let node_id = self.opts.node_id.map(|id| id as u32);

-        let mut man = FlowWorkerManager::new(node_id, query_engine, table_meta);
+        let mut man = StreamingEngine::new(node_id, query_engine, table_meta);
        for worker_id in 0..num_workers {
            let (tx, rx) = oneshot::channel();

@@ -543,6 +594,10 @@ impl<'a> FlownodeServiceBuilder<'a> {
    }
 }

+/// Basically a tiny frontend that communicates with datanode, different from [`FrontendClient`] which
+/// connect to a real frontend instead, this is used for flow's streaming engine. And is for simple query.
+///
+/// For heavy query use [`FrontendClient`] which offload computation to frontend, lifting the load from flownode
 #[derive(Clone)]
 pub struct FrontendInvoker {
    inserter: Arc<Inserter>,
@@ -564,7 +619,7 @@ impl FrontendInvoker {
    }

    pub async fn build_from(
-        flow_worker_manager: FlowWorkerManagerRef,
+        flow_streaming_engine: FlowStreamingEngineRef,
        catalog_manager: CatalogManagerRef,
        kv_backend: KvBackendRef,
        layered_cache_registry: LayeredCacheRegistryRef,
@@ -599,7 +654,7 @@ impl FrontendInvoker {
            node_manager.clone(),
        ));

-        let query_engine = flow_worker_manager.query_engine.clone();
+        let query_engine = flow_streaming_engine.query_engine.clone();

        let statement_executor = Arc::new(StatementExecutor::new(
            catalog_manager.clone(),
--- a/src/frontend/Cargo.toml
+++ b/src/frontend/Cargo.toml
@@ -15,6 +15,7 @@ api.workspace = true
 arc-swap = "1.0"
 async-trait.workspace = true
 auth.workspace = true
+bytes.workspace = true
 cache.workspace = true
 catalog.workspace = true
 client.workspace = true
@@ -39,6 +40,7 @@ datafusion.workspace = true
 datafusion-expr.workspace = true
 datanode.workspace = true
 datatypes.workspace = true
+futures.workspace = true
 humantime-serde.workspace = true
 lazy_static.workspace = true
 log-query.workspace = true
@@ -47,6 +49,7 @@ meta-client.workspace = true
 num_cpus.workspace = true
 opentelemetry-proto.workspace = true
 operator.workspace = true
+otel-arrow-rust.workspace = true
 partition.workspace = true
 pipeline.workspace = true
 prometheus.workspace = true
--- a/src/frontend/src/error.rs
+++ b/src/frontend/src/error.rs
@@ -19,6 +19,8 @@ use common_error::define_into_tonic_status;
 use common_error::ext::{BoxedError, ErrorExt};
 use common_error::status_code::StatusCode;
 use common_macro::stack_trace_debug;
+use common_query::error::datafusion_status_code;
+use datafusion::error::DataFusionError;
 use session::ReadPreference;
 use snafu::{Location, Snafu};
 use store_api::storage::RegionId;
@@ -345,7 +347,15 @@ pub enum Error {
    SubstraitDecodeLogicalPlan {
        #[snafu(implicit)]
        location: Location,
-        source: substrait::error::Error,
+        source: common_query::error::Error,
+    },
+
+    #[snafu(display("DataFusionError"))]
+    DataFusion {
+        #[snafu(source)]
+        error: DataFusionError,
+        #[snafu(implicit)]
+        location: Location,
    },
 }

@@ -423,6 +433,8 @@ impl ErrorExt for Error {
            Error::TableOperation { source, .. } => source.status_code(),

            Error::InFlightWriteBytesExceeded { .. } => StatusCode::RateLimited,
+
+            Error::DataFusion { error, .. } => datafusion_status_code::<Self>(error, None),
        }
    }

--- a/src/frontend/src/instance.rs
+++ b/src/frontend/src/instance.rs
@@ -278,7 +278,7 @@ impl SqlQueryHandler for Instance {
        // plan should be prepared before exec
        // we'll do check there
        self.query_engine
-            .execute(plan, query_ctx)
+            .execute(plan.clone(), query_ctx)
            .await
            .context(ExecLogicalPlanSnafu)
    }
--- a/src/frontend/src/instance/grpc.rs
+++ b/src/frontend/src/instance/grpc.rs
@@ -12,29 +12,33 @@
 // See the License for the specific language governing permissions and
 // limitations under the License.

+use std::sync::Arc;
+
 use api::v1::ddl_request::{Expr as DdlExpr, Expr};
 use api::v1::greptime_request::Request;
 use api::v1::query_request::Query;
-use api::v1::{DeleteRequests, DropFlowExpr, InsertRequests, RowDeleteRequests, RowInsertRequests};
+use api::v1::{
+    DeleteRequests, DropFlowExpr, InsertIntoPlan, InsertRequests, RowDeleteRequests,
+    RowInsertRequests,
+};
 use async_trait::async_trait;
 use auth::{PermissionChecker, PermissionCheckerRef, PermissionReq};
 use common_base::AffectedRows;
+use common_query::logical_plan::add_insert_to_logical_plan;
 use common_query::Output;
 use common_telemetry::tracing::{self};
-use datafusion::execution::SessionStateBuilder;
 use query::parser::PromQuery;
 use servers::interceptor::{GrpcQueryInterceptor, GrpcQueryInterceptorRef};
 use servers::query_handler::grpc::{GrpcQueryHandler, RawRecordBatch};
 use servers::query_handler::sql::SqlQueryHandler;
 use session::context::QueryContextRef;
 use snafu::{ensure, OptionExt, ResultExt};
-use substrait::{DFLogicalSubstraitConvertor, SubstraitPlan};
 use table::table_name::TableName;

 use crate::error::{
-    CatalogSnafu, Error, InFlightWriteBytesExceededSnafu, IncompleteGrpcRequestSnafu,
-    NotSupportedSnafu, PermissionSnafu, Result, SubstraitDecodeLogicalPlanSnafu,
-    TableNotFoundSnafu, TableOperationSnafu,
+    CatalogSnafu, DataFusionSnafu, Error, InFlightWriteBytesExceededSnafu,
+    IncompleteGrpcRequestSnafu, NotSupportedSnafu, PermissionSnafu, PlanStatementSnafu, Result,
+    SubstraitDecodeLogicalPlanSnafu, TableNotFoundSnafu, TableOperationSnafu,
 };
 use crate::instance::{attach_timer, Instance};
 use crate::metrics::{
@@ -91,14 +95,31 @@ impl GrpcQueryHandler for Instance {
                    Query::LogicalPlan(plan) => {
                        // this path is useful internally when flownode needs to execute a logical plan through gRPC interface
                        let timer = GRPC_HANDLE_PLAN_ELAPSED.start_timer();
-                        let plan = DFLogicalSubstraitConvertor {}
-                            .decode(&*plan, SessionStateBuilder::default().build())
+
+                        // use dummy catalog to provide table
+                        let plan_decoder = self
+                            .query_engine()
+                            .engine_context(ctx.clone())
+                            .new_plan_decoder()
+                            .context(PlanStatementSnafu)?;
+
+                        let dummy_catalog_list =
+                            Arc::new(catalog::table_source::dummy_catalog::DummyCatalogList::new(
+                                self.catalog_manager().clone(),
+                            ));
+
+                        let logical_plan = plan_decoder
+                            .decode(bytes::Bytes::from(plan), dummy_catalog_list, true)
                            .await
                            .context(SubstraitDecodeLogicalPlanSnafu)?;
-                        let output = SqlQueryHandler::do_exec_plan(self, plan, ctx.clone()).await?;
+                        let output =
+                            SqlQueryHandler::do_exec_plan(self, logical_plan, ctx.clone()).await?;

                        attach_timer(output, timer)
                    }
+                    Query::InsertIntoPlan(insert) => {
+                        self.handle_insert_plan(insert, ctx.clone()).await?
+                    }
                    Query::PromRangeQuery(promql) => {
                        let timer = GRPC_HANDLE_PROMQL_ELAPSED.start_timer();
                        let prom_query = PromQuery {
@@ -284,6 +305,91 @@ fn fill_catalog_and_schema_from_context(ddl_expr: &mut DdlExpr, ctx: &QueryConte
 }

 impl Instance {
+    async fn handle_insert_plan(
+        &self,
+        insert: InsertIntoPlan,
+        ctx: QueryContextRef,
+    ) -> Result<Output> {
+        let timer = GRPC_HANDLE_PLAN_ELAPSED.start_timer();
+        let table_name = insert.table_name.context(IncompleteGrpcRequestSnafu {
+            err_msg: "'table_name' is absent in InsertIntoPlan",
+        })?;
+
+        // use dummy catalog to provide table
+        let plan_decoder = self
+            .query_engine()
+            .engine_context(ctx.clone())
+            .new_plan_decoder()
+            .context(PlanStatementSnafu)?;
+
+        let dummy_catalog_list =
+            Arc::new(catalog::table_source::dummy_catalog::DummyCatalogList::new(
+                self.catalog_manager().clone(),
+            ));
+
+        // no optimize yet since we still need to add stuff
+        let logical_plan = plan_decoder
+            .decode(
+                bytes::Bytes::from(insert.logical_plan),
+                dummy_catalog_list,
+                false,
+            )
+            .await
+            .context(SubstraitDecodeLogicalPlanSnafu)?;
+
+        let table = self
+            .catalog_manager()
+            .table(
+                &table_name.catalog_name,
+                &table_name.schema_name,
+                &table_name.table_name,
+                None,
+            )
+            .await
+            .context(CatalogSnafu)?
+            .with_context(|| TableNotFoundSnafu {
+                table_name: [
+                    table_name.catalog_name.clone(),
+                    table_name.schema_name.clone(),
+                    table_name.table_name.clone(),
+                ]
+                .join("."),
+            })?;
+
+        let table_info = table.table_info();
+
+        let df_schema = Arc::new(
+            table_info
+                .meta
+                .schema
+                .arrow_schema()
+                .clone()
+                .try_into()
+                .context(DataFusionSnafu)?,
+        );
+
+        let insert_into = add_insert_to_logical_plan(table_name, df_schema, logical_plan)
+            .context(SubstraitDecodeLogicalPlanSnafu)?;
+
+        let engine_ctx = self.query_engine().engine_context(ctx.clone());
+        let state = engine_ctx.state();
+        // Analyze the plan
+        let analyzed_plan = state
+            .analyzer()
+            .execute_and_check(insert_into, state.config_options(), |_, _| {})
+            .context(common_query::error::GeneralDataFusionSnafu)
+            .context(SubstraitDecodeLogicalPlanSnafu)?;
+
+        // Optimize the plan
+        let optimized_plan = state
+            .optimize(&analyzed_plan)
+            .context(common_query::error::GeneralDataFusionSnafu)
+            .context(SubstraitDecodeLogicalPlanSnafu)?;
+
+        let output = SqlQueryHandler::do_exec_plan(self, optimized_plan, ctx.clone()).await?;
+
+        Ok(attach_timer(output, timer))
+    }
    #[tracing::instrument(skip_all)]
    pub async fn handle_inserts(
        &self,
--- a/src/frontend/src/server.rs
+++ b/src/frontend/src/server.rs
@@ -27,6 +27,7 @@ use servers::http::{HttpServer, HttpServerBuilder};
 use servers::interceptor::LogIngestInterceptorRef;
 use servers::metrics_handler::MetricsHandler;
 use servers::mysql::server::{MysqlServer, MysqlSpawnConfig, MysqlSpawnRef};
+use servers::otel_arrow::OtelArrowServiceHandler;
 use servers::postgres::PostgresServer;
 use servers::query_handler::grpc::ServerGrpcQueryHandlerAdapter;
 use servers::query_handler::sql::ServerSqlQueryHandlerAdapter;
@@ -162,6 +163,7 @@ where
        let grpc_server = builder
            .database_handler(greptime_request_handler.clone())
            .prometheus_handler(self.instance.clone(), user_provider.clone())
+            .otel_arrow_handler(OtelArrowServiceHandler(self.instance.clone()))
            .flight_handler(Arc::new(greptime_request_handler))
            .build();
        Ok(grpc_server)
--- a/src/index/src/fulltext_index/tokenizer.rs
+++ b/src/index/src/fulltext_index/tokenizer.rs
@@ -46,7 +46,11 @@ pub struct ChineseTokenizer;

 impl Tokenizer for ChineseTokenizer {
    fn tokenize<'a>(&self, text: &'a str) -> Vec<&'a str> {
-        JIEBA.cut(text, false)
+        if text.is_ascii() {
+            EnglishTokenizer {}.tokenize(text)
+        } else {
+            JIEBA.cut(text, false)
+        }
    }
 }

--- a/src/meta-srv/src/bootstrap.rs
+++ b/src/meta-srv/src/bootstrap.rs
@@ -66,10 +66,12 @@ use crate::election::postgres::PgElection;
 #[cfg(any(feature = "pg_kvbackend", feature = "mysql_kvbackend"))]
 use crate::election::CANDIDATE_LEASE_SECS;
 use crate::metasrv::builder::MetasrvBuilder;
-use crate::metasrv::{BackendImpl, Metasrv, MetasrvOptions, SelectorRef};
+use crate::metasrv::{BackendImpl, Metasrv, MetasrvOptions, SelectTarget, SelectorRef};
+use crate::node_excluder::NodeExcluderRef;
 use crate::selector::lease_based::LeaseBasedSelector;
 use crate::selector::load_based::LoadBasedSelector;
 use crate::selector::round_robin::RoundRobinSelector;
+use crate::selector::weight_compute::RegionNumsBasedWeightCompute;
 use crate::selector::SelectorType;
 use crate::service::admin;
 use crate::{error, Result};
@@ -294,10 +296,31 @@ pub async fn metasrv_builder(

    let in_memory = Arc::new(MemoryKvBackend::new()) as ResettableKvBackendRef;

-    let selector = match opts.selector {
-        SelectorType::LoadBased => Arc::new(LoadBasedSelector::default()) as SelectorRef,
-        SelectorType::LeaseBased => Arc::new(LeaseBasedSelector) as SelectorRef,
-        SelectorType::RoundRobin => Arc::new(RoundRobinSelector::default()) as SelectorRef,
+    let node_excluder = plugins
+        .get::<NodeExcluderRef>()
+        .unwrap_or_else(|| Arc::new(Vec::new()) as NodeExcluderRef);
+    let selector = if let Some(selector) = plugins.get::<SelectorRef>() {
+        info!("Using selector from plugins");
+        selector
+    } else {
+        let selector = match opts.selector {
+            SelectorType::LoadBased => Arc::new(LoadBasedSelector::new(
+                RegionNumsBasedWeightCompute,
+                node_excluder,
+            )) as SelectorRef,
+            SelectorType::LeaseBased => {
+                Arc::new(LeaseBasedSelector::new(node_excluder)) as SelectorRef
+            }
+            SelectorType::RoundRobin => Arc::new(RoundRobinSelector::new(
+                SelectTarget::Datanode,
+                node_excluder,
+            )) as SelectorRef,
+        };
+        info!(
+            "Using selector from options, selector type: {}",
+            opts.selector.as_ref()
+        );
+        selector
    };

    Ok(MetasrvBuilder::new()
--- a/src/meta-srv/src/error.rs
+++ b/src/meta-srv/src/error.rs
@@ -336,6 +336,13 @@ pub enum Error {
        location: Location,
    },

+    #[snafu(display("Region's leader peer changed: {}", msg))]
+    LeaderPeerChanged {
+        msg: String,
+        #[snafu(implicit)]
+        location: Location,
+    },
+
    #[snafu(display("Invalid arguments: {}", err_msg))]
    InvalidArguments {
        err_msg: String,
@@ -914,7 +921,8 @@ impl ErrorExt for Error {
            | Error::ProcedureNotFound { .. }
            | Error::TooManyPartitions { .. }
            | Error::TomlFormat { .. }
-            | Error::HandlerNotFound { .. } => StatusCode::InvalidArguments,
+            | Error::HandlerNotFound { .. }
+            | Error::LeaderPeerChanged { .. } => StatusCode::InvalidArguments,
            Error::LeaseKeyFromUtf8 { .. }
            | Error::LeaseValueFromUtf8 { .. }
            | Error::InvalidRegionKeyFromUtf8 { .. }
--- a/src/meta-srv/src/flow_meta_alloc.rs
+++ b/src/meta-srv/src/flow_meta_alloc.rs
@@ -12,6 +12,8 @@
 // See the License for the specific language governing permissions and
 // limitations under the License.

+use std::collections::HashSet;
+
 use common_error::ext::BoxedError;
 use common_meta::ddl::flow_meta::PartitionPeerAllocator;
 use common_meta::peer::Peer;
@@ -40,6 +42,7 @@ impl PartitionPeerAllocator for FlowPeerAllocator {
                SelectorOptions {
                    min_required_items: partitions,
                    allow_duplication: true,
+                    exclude_peer_ids: HashSet::new(),
                },
            )
            .await
--- a/src/meta-srv/src/lib.rs
+++ b/src/meta-srv/src/lib.rs
@@ -31,6 +31,7 @@ pub mod metasrv;
 pub mod metrics;
 #[cfg(feature = "mock")]
 pub mod mocks;
+pub mod node_excluder;
 pub mod procedure;
 pub mod pubsub;
 pub mod region;
--- a/src/meta-srv/src/metasrv.rs
+++ b/src/meta-srv/src/metasrv.rs
@@ -111,6 +111,11 @@ pub struct MetasrvOptions {
    pub use_memory_store: bool,
    /// Whether to enable region failover.
    pub enable_region_failover: bool,
+    /// Whether to allow region failover on local WAL.
+    ///
+    /// If it's true, the region failover will be allowed even if the local WAL is used.
+    /// Note that this option is not recommended to be set to true, because it may lead to data loss during failover.
+    pub allow_region_failover_on_local_wal: bool,
    /// The HTTP server options.
    pub http: HttpOptions,
    /// The logging options.
@@ -173,6 +178,7 @@ impl Default for MetasrvOptions {
            selector: SelectorType::default(),
            use_memory_store: false,
            enable_region_failover: false,
+            allow_region_failover_on_local_wal: false,
            http: HttpOptions::default(),
            logging: LoggingOptions {
                dir: format!("{METASRV_HOME}/logs"),
--- a/src/meta-srv/src/metasrv/builder.rs
+++ b/src/meta-srv/src/metasrv/builder.rs
@@ -40,7 +40,8 @@ use common_meta::state_store::KvStateStore;
 use common_meta::wal_options_allocator::{build_kafka_client, build_wal_options_allocator};
 use common_procedure::local::{LocalManager, ManagerConfig};
 use common_procedure::ProcedureManagerRef;
-use snafu::ResultExt;
+use common_telemetry::warn;
+use snafu::{ensure, ResultExt};

 use crate::cache_invalidator::MetasrvCacheInvalidator;
 use crate::cluster::{MetaPeerClientBuilder, MetaPeerClientRef};
@@ -190,7 +191,7 @@ impl MetasrvBuilder {

        let meta_peer_client = meta_peer_client
            .unwrap_or_else(|| build_default_meta_peer_client(&election, &in_memory));
-        let selector = selector.unwrap_or_else(|| Arc::new(LeaseBasedSelector));
+        let selector = selector.unwrap_or_else(|| Arc::new(LeaseBasedSelector::default()));
        let pushers = Pushers::default();
        let mailbox = build_mailbox(&kv_backend, &pushers);
        let procedure_manager = build_procedure_manager(&options, &kv_backend);
@@ -234,13 +235,17 @@ impl MetasrvBuilder {
            ))
        });

+        let flow_selector = Arc::new(RoundRobinSelector::new(
+            SelectTarget::Flownode,
+            Arc::new(Vec::new()),
+        )) as SelectorRef;
+
        let flow_metadata_allocator = {
            // for now flownode just use round-robin selector
-            let flow_selector = RoundRobinSelector::new(SelectTarget::Flownode);
            let flow_selector_ctx = selector_ctx.clone();
            let peer_allocator = Arc::new(FlowPeerAllocator::new(
                flow_selector_ctx,
-                Arc::new(flow_selector),
+                flow_selector.clone(),
            ));
            let seq = Arc::new(
                SequenceBuilder::new(FLOW_ID_SEQ, kv_backend.clone())
@@ -272,18 +277,25 @@ impl MetasrvBuilder {
            },
        ));
        let peer_lookup_service = Arc::new(MetaPeerLookupService::new(meta_peer_client.clone()));
+
        if !is_remote_wal && options.enable_region_failover {
-            return error::UnexpectedSnafu {
-                violated: "Region failover is not supported in the local WAL implementation!",
+            ensure!(
+                options.allow_region_failover_on_local_wal,
+                error::UnexpectedSnafu {
+                    violated: "Region failover is not supported in the local WAL implementation! 
+                    If you want to enable region failover for local WAL, please set `allow_region_failover_on_local_wal` to true.",
+                }
+            );
+            if options.allow_region_failover_on_local_wal {
+                warn!("Region failover is force enabled in the local WAL implementation! This may lead to data loss during failover!");
            }
-            .fail();
        }

        let (tx, rx) = RegionSupervisor::channel();
        let (region_failure_detector_controller, region_supervisor_ticker): (
            RegionFailureDetectorControllerRef,
            Option<std::sync::Arc<RegionSupervisorTicker>>,
-        ) = if options.enable_region_failover && is_remote_wal {
+        ) = if options.enable_region_failover {
            (
                Arc::new(RegionFailureDetectorControl::new(tx.clone())) as _,
                Some(Arc::new(RegionSupervisorTicker::new(
@@ -309,7 +321,7 @@ impl MetasrvBuilder {
        ));
        region_migration_manager.try_start()?;

-        let region_failover_handler = if options.enable_region_failover && is_remote_wal {
+        let region_failover_handler = if options.enable_region_failover {
            let region_supervisor = RegionSupervisor::new(
                rx,
                options.failure_detector,
@@ -420,7 +432,7 @@ impl MetasrvBuilder {
            meta_peer_client: meta_peer_client.clone(),
            selector,
            // TODO(jeremy): We do not allow configuring the flow selector.
-            flow_selector: Arc::new(RoundRobinSelector::new(SelectTarget::Flownode)),
+            flow_selector,
            handler_group: RwLock::new(None),
            handler_group_builder: Mutex::new(Some(handler_group_builder)),
            election,
--- a/src/meta-srv/src/metrics.rs
+++ b/src/meta-srv/src/metrics.rs
@@ -62,11 +62,22 @@ lazy_static! {
        register_int_counter!("greptime_meta_region_migration_fail", "meta region migration fail").unwrap();
    // The heartbeat stat memory size histogram.
    pub static ref METRIC_META_HEARTBEAT_STAT_MEMORY_SIZE: Histogram =
-        register_histogram!("greptime_meta_heartbeat_stat_memory_size", "meta heartbeat stat memory size").unwrap();
+        register_histogram!("greptime_meta_heartbeat_stat_memory_size", "meta heartbeat stat memory size", vec![
+            100.0, 500.0, 1000.0, 1500.0, 2000.0, 3000.0, 5000.0, 10000.0, 20000.0
+        ]).unwrap();
    // The heartbeat rate counter.
    pub static ref METRIC_META_HEARTBEAT_RATE: IntCounter =
        register_int_counter!("greptime_meta_heartbeat_rate", "meta heartbeat arrival rate").unwrap();
    /// The remote WAL prune execute counter.
    pub static ref METRIC_META_REMOTE_WAL_PRUNE_EXECUTE: IntCounterVec =
        register_int_counter_vec!("greptime_meta_remote_wal_prune_execute", "meta remote wal prune execute", &["topic_name"]).unwrap();
+    /// The migration stage elapsed histogram.
+    pub static ref METRIC_META_REGION_MIGRATION_STAGE_ELAPSED: HistogramVec = register_histogram_vec!(
+        "greptime_meta_region_migration_stage_elapsed",
+        "meta region migration stage elapsed",
+        &["stage"],
+        // 0.01 ~ 1000
+        exponential_buckets(0.01, 10.0, 7).unwrap(),
+    )
+    .unwrap();
 }
--- a/src/meta-srv/src/node_excluder.rs
+++ b/src/meta-srv/src/node_excluder.rs
@@ -0,0 +1,32 @@
+// Copyright 2023 Greptime Team
+//
+// Licensed under the Apache License, Version 2.0 (the "License");
+// you may not use this file except in compliance with the License.
+// You may obtain a copy of the License at
+//
+//     http://www.apache.org/licenses/LICENSE-2.0
+//
+// Unless required by applicable law or agreed to in writing, software
+// distributed under the License is distributed on an "AS IS" BASIS,
+// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+// See the License for the specific language governing permissions and
+// limitations under the License.
+
+use std::sync::Arc;
+
+use common_meta::DatanodeId;
+
+pub type NodeExcluderRef = Arc<dyn NodeExcluder>;
+
+/// [NodeExcluder] is used to help decide whether some nodes should be excluded (out of consideration)
+/// in certain situations. For example, in some node selectors.
+pub trait NodeExcluder: Send + Sync {
+    /// Returns the excluded datanode ids.
+    fn excluded_datanode_ids(&self) -> &Vec<DatanodeId>;
+}
+
+impl NodeExcluder for Vec<DatanodeId> {
+    fn excluded_datanode_ids(&self) -> &Vec<DatanodeId> {
+        self
+    }
+}
--- a/src/meta-srv/src/procedure/region_migration.rs
+++ b/src/meta-srv/src/procedure/region_migration.rs
@@ -25,7 +25,7 @@ pub(crate) mod update_metadata;
 pub(crate) mod upgrade_candidate_region;

 use std::any::Any;
-use std::fmt::Debug;
+use std::fmt::{Debug, Display};
 use std::time::Duration;

 use common_error::ext::BoxedError;
@@ -43,7 +43,7 @@ use common_procedure::error::{
    Error as ProcedureError, FromJsonSnafu, Result as ProcedureResult, ToJsonSnafu,
 };
 use common_procedure::{Context as ProcedureContext, LockKey, Procedure, Status, StringKey};
-use common_telemetry::info;
+use common_telemetry::{error, info};
 use manager::RegionMigrationProcedureGuard;
 pub use manager::{
    RegionMigrationManagerRef, RegionMigrationProcedureTask, RegionMigrationProcedureTracker,
@@ -55,9 +55,15 @@ use tokio::time::Instant;

 use self::migration_start::RegionMigrationStart;
 use crate::error::{self, Result};
-use crate::metrics::{METRIC_META_REGION_MIGRATION_ERROR, METRIC_META_REGION_MIGRATION_EXECUTE};
+use crate::metrics::{
+    METRIC_META_REGION_MIGRATION_ERROR, METRIC_META_REGION_MIGRATION_EXECUTE,
+    METRIC_META_REGION_MIGRATION_STAGE_ELAPSED,
+};
 use crate::service::mailbox::MailboxRef;

+/// The default timeout for region migration.
+pub const DEFAULT_REGION_MIGRATION_TIMEOUT: Duration = Duration::from_secs(120);
+
 /// It's shared in each step and available even after recovering.
 ///
 /// It will only be updated/stored after the Red node has succeeded.
@@ -100,6 +106,82 @@ impl PersistentContext {
    }
 }

+/// Metrics of region migration.
+#[derive(Debug, Clone, Default)]
+pub struct Metrics {
+    /// Elapsed time of downgrading region and upgrading region.
+    operations_elapsed: Duration,
+    /// Elapsed time of downgrading leader region.
+    downgrade_leader_region_elapsed: Duration,
+    /// Elapsed time of open candidate region.
+    open_candidate_region_elapsed: Duration,
+    /// Elapsed time of upgrade candidate region.
+    upgrade_candidate_region_elapsed: Duration,
+}
+
+impl Display for Metrics {
+    fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
+        write!(
+            f,
+            "operations_elapsed: {:?}, downgrade_leader_region_elapsed: {:?}, open_candidate_region_elapsed: {:?}, upgrade_candidate_region_elapsed: {:?}",
+            self.operations_elapsed,
+            self.downgrade_leader_region_elapsed,
+            self.open_candidate_region_elapsed,
+            self.upgrade_candidate_region_elapsed
+        )
+    }
+}
+
+impl Metrics {
+    /// Updates the elapsed time of downgrading region and upgrading region.
+    pub fn update_operations_elapsed(&mut self, elapsed: Duration) {
+        self.operations_elapsed += elapsed;
+    }
+
+    /// Updates the elapsed time of downgrading leader region.
+    pub fn update_downgrade_leader_region_elapsed(&mut self, elapsed: Duration) {
+        self.downgrade_leader_region_elapsed += elapsed;
+    }
+
+    /// Updates the elapsed time of open candidate region.
+    pub fn update_open_candidate_region_elapsed(&mut self, elapsed: Duration) {
+        self.open_candidate_region_elapsed += elapsed;
+    }
+
+    /// Updates the elapsed time of upgrade candidate region.
+    pub fn update_upgrade_candidate_region_elapsed(&mut self, elapsed: Duration) {
+        self.upgrade_candidate_region_elapsed += elapsed;
+    }
+}
+
+impl Drop for Metrics {
+    fn drop(&mut self) {
+        if !self.operations_elapsed.is_zero() {
+            METRIC_META_REGION_MIGRATION_STAGE_ELAPSED
+                .with_label_values(&["operations"])
+                .observe(self.operations_elapsed.as_secs_f64());
+        }
+
+        if !self.downgrade_leader_region_elapsed.is_zero() {
+            METRIC_META_REGION_MIGRATION_STAGE_ELAPSED
+                .with_label_values(&["downgrade_leader_region"])
+                .observe(self.downgrade_leader_region_elapsed.as_secs_f64());
+        }
+
+        if !self.open_candidate_region_elapsed.is_zero() {
+            METRIC_META_REGION_MIGRATION_STAGE_ELAPSED
+                .with_label_values(&["open_candidate_region"])
+                .observe(self.open_candidate_region_elapsed.as_secs_f64());
+        }
+
+        if !self.upgrade_candidate_region_elapsed.is_zero() {
+            METRIC_META_REGION_MIGRATION_STAGE_ELAPSED
+                .with_label_values(&["upgrade_candidate_region"])
+                .observe(self.upgrade_candidate_region_elapsed.as_secs_f64());
+        }
+    }
+}
+
 /// It's shared in each step and available in executing (including retrying).
 ///
 /// It will be dropped if the procedure runner crashes.
@@ -129,8 +211,8 @@ pub struct VolatileContext {
    leader_region_last_entry_id: Option<u64>,
    /// The last_entry_id of leader metadata region (Only used for metric engine).
    leader_region_metadata_last_entry_id: Option<u64>,
-    /// Elapsed time of downgrading region and upgrading region.
-    operations_elapsed: Duration,
+    /// Metrics of region migration.
+    metrics: Metrics,
 }

 impl VolatileContext {
@@ -228,12 +310,35 @@ impl Context {
    pub fn next_operation_timeout(&self) -> Option<Duration> {
        self.persistent_ctx
            .timeout
-            .checked_sub(self.volatile_ctx.operations_elapsed)
+            .checked_sub(self.volatile_ctx.metrics.operations_elapsed)
    }

    /// Updates operations elapsed.
    pub fn update_operations_elapsed(&mut self, instant: Instant) {
-        self.volatile_ctx.operations_elapsed += instant.elapsed();
+        self.volatile_ctx
+            .metrics
+            .update_operations_elapsed(instant.elapsed());
+    }
+
+    /// Updates the elapsed time of downgrading leader region.
+    pub fn update_downgrade_leader_region_elapsed(&mut self, instant: Instant) {
+        self.volatile_ctx
+            .metrics
+            .update_downgrade_leader_region_elapsed(instant.elapsed());
+    }
+
+    /// Updates the elapsed time of open candidate region.
+    pub fn update_open_candidate_region_elapsed(&mut self, instant: Instant) {
+        self.volatile_ctx
+            .metrics
+            .update_open_candidate_region_elapsed(instant.elapsed());
+    }
+
+    /// Updates the elapsed time of upgrade candidate region.
+    pub fn update_upgrade_candidate_region_elapsed(&mut self, instant: Instant) {
+        self.volatile_ctx
+            .metrics
+            .update_upgrade_candidate_region_elapsed(instant.elapsed());
    }

    /// Returns address of meta server.
@@ -547,6 +652,14 @@ impl Procedure for RegionMigrationProcedure {
                    .inc();
                ProcedureError::retry_later(e)
            } else {
+                error!(
+                    e;
+                    "Region migration procedure failed, region_id: {}, from_peer: {}, to_peer: {}, {}",
+                    self.context.region_id(),
+                    self.context.persistent_ctx.from_peer,
+                    self.context.persistent_ctx.to_peer,
+                    self.context.volatile_ctx.metrics,
+                );
                METRIC_META_REGION_MIGRATION_ERROR
                    .with_label_values(&[name, "external"])
                    .inc();
--- a/src/meta-srv/src/procedure/region_migration/close_downgraded_region.rs
+++ b/src/meta-srv/src/procedure/region_migration/close_downgraded_region.rs
@@ -46,7 +46,13 @@ impl State for CloseDowngradedRegion {
            let region_id = ctx.region_id();
            warn!(err; "Failed to close downgraded leader region: {region_id} on datanode {:?}", downgrade_leader_datanode);
        }
-
+        info!(
+            "Region migration is finished: region_id: {}, from_peer: {}, to_peer: {}, {}",
+            ctx.region_id(),
+            ctx.persistent_ctx.from_peer,
+            ctx.persistent_ctx.to_peer,
+            ctx.volatile_ctx.metrics,
+        );
        Ok((Box::new(RegionMigrationEnd), Status::done()))
    }

--- a/src/meta-srv/src/procedure/region_migration/downgrade_leader_region.rs
+++ b/src/meta-srv/src/procedure/region_migration/downgrade_leader_region.rs
@@ -54,6 +54,7 @@ impl Default for DowngradeLeaderRegion {
 #[typetag::serde]
 impl State for DowngradeLeaderRegion {
    async fn next(&mut self, ctx: &mut Context) -> Result<(Box<dyn State>, Status)> {
+        let now = Instant::now();
        // Ensures the `leader_region_lease_deadline` must exist after recovering.
        ctx.volatile_ctx
            .set_leader_region_lease_deadline(Duration::from_secs(REGION_LEASE_SECS));
@@ -77,6 +78,7 @@ impl State for DowngradeLeaderRegion {
                }
            }
        }
+        ctx.update_downgrade_leader_region_elapsed(now);

        Ok((
            Box::new(UpgradeCandidateRegion::default()),
@@ -348,7 +350,8 @@ mod tests {
        let env = TestingEnv::new();
        let mut ctx = env.context_factory().new_context(persistent_context);
        prepare_table_metadata(&ctx, HashMap::default()).await;
-        ctx.volatile_ctx.operations_elapsed = ctx.persistent_ctx.timeout + Duration::from_secs(1);
+        ctx.volatile_ctx.metrics.operations_elapsed =
+            ctx.persistent_ctx.timeout + Duration::from_secs(1);

        let err = state.downgrade_region(&mut ctx).await.unwrap_err();

@@ -591,7 +594,8 @@ mod tests {
        let mut ctx = env.context_factory().new_context(persistent_context);
        let mailbox_ctx = env.mailbox_context();
        let mailbox = mailbox_ctx.mailbox().clone();
-        ctx.volatile_ctx.operations_elapsed = ctx.persistent_ctx.timeout + Duration::from_secs(1);
+        ctx.volatile_ctx.metrics.operations_elapsed =
+            ctx.persistent_ctx.timeout + Duration::from_secs(1);

        let (tx, rx) = tokio::sync::mpsc::channel(1);
        mailbox_ctx
--- a/src/meta-srv/src/procedure/region_migration/manager.rs
+++ b/src/meta-srv/src/procedure/region_migration/manager.rs
@@ -267,8 +267,8 @@ impl RegionMigrationManager {

        ensure!(
            leader_peer.id == task.from_peer.id,
-            error::InvalidArgumentsSnafu {
-                err_msg: format!(
+            error::LeaderPeerChangedSnafu {
+                msg: format!(
                    "Region's leader peer({}) is not the `from_peer`({}), region: {}",
                    leader_peer.id, task.from_peer.id, task.region_id
                ),
@@ -507,8 +507,8 @@ mod test {
            .await;

        let err = manager.submit_procedure(task).await.unwrap_err();
-        assert_matches!(err, error::Error::InvalidArguments { .. });
-        assert_eq!(err.to_string(), "Invalid arguments: Region's leader peer(3) is not the `from_peer`(1), region: 4398046511105(1024, 1)");
+        assert_matches!(err, error::Error::LeaderPeerChanged { .. });
+        assert_eq!(err.to_string(), "Region's leader peer changed: Region's leader peer(3) is not the `from_peer`(1), region: 4398046511105(1024, 1)");
    }

    #[tokio::test]
--- a/src/meta-srv/src/procedure/region_migration/migration_abort.rs
+++ b/src/meta-srv/src/procedure/region_migration/migration_abort.rs
@@ -15,6 +15,7 @@
 use std::any::Any;

 use common_procedure::Status;
+use common_telemetry::warn;
 use serde::{Deserialize, Serialize};

 use crate::error::{self, Result};
@@ -37,7 +38,15 @@ impl RegionMigrationAbort {
 #[async_trait::async_trait]
 #[typetag::serde]
 impl State for RegionMigrationAbort {
-    async fn next(&mut self, _: &mut Context) -> Result<(Box<dyn State>, Status)> {
+    async fn next(&mut self, ctx: &mut Context) -> Result<(Box<dyn State>, Status)> {
+        warn!(
+            "Region migration is aborted: {}, region_id: {}, from_peer: {}, to_peer: {}, {}",
+            self.reason,
+            ctx.region_id(),
+            ctx.persistent_ctx.from_peer,
+            ctx.persistent_ctx.to_peer,
+            ctx.volatile_ctx.metrics,
+        );
        error::MigrationAbortSnafu {
            reason: &self.reason,
        }
--- a/src/meta-srv/src/procedure/region_migration/open_candidate_region.rs
+++ b/src/meta-srv/src/procedure/region_migration/open_candidate_region.rs
@@ -13,7 +13,7 @@
 // limitations under the License.

 use std::any::Any;
-use std::time::{Duration, Instant};
+use std::time::Duration;

 use api::v1::meta::MailboxMessage;
 use common_meta::distributed_time_constants::REGION_LEASE_SECS;
@@ -24,6 +24,7 @@ use common_procedure::Status;
 use common_telemetry::info;
 use serde::{Deserialize, Serialize};
 use snafu::{OptionExt, ResultExt};
+use tokio::time::Instant;

 use crate::error::{self, Result};
 use crate::handler::HeartbeatMailbox;
@@ -42,7 +43,9 @@ pub struct OpenCandidateRegion;
 impl State for OpenCandidateRegion {
    async fn next(&mut self, ctx: &mut Context) -> Result<(Box<dyn State>, Status)> {
        let instruction = self.build_open_region_instruction(ctx).await?;
+        let now = Instant::now();
        self.open_candidate_region(ctx, instruction).await?;
+        ctx.update_open_candidate_region_elapsed(now);

        Ok((
            Box::new(UpdateMetadata::Downgrade),
--- a/src/meta-srv/src/procedure/region_migration/upgrade_candidate_region.rs
+++ b/src/meta-srv/src/procedure/region_migration/upgrade_candidate_region.rs
@@ -54,9 +54,12 @@ impl Default for UpgradeCandidateRegion {
 #[typetag::serde]
 impl State for UpgradeCandidateRegion {
    async fn next(&mut self, ctx: &mut Context) -> Result<(Box<dyn State>, Status)> {
+        let now = Instant::now();
        if self.upgrade_region_with_retry(ctx).await {
+            ctx.update_upgrade_candidate_region_elapsed(now);
            Ok((Box::new(UpdateMetadata::Upgrade), Status::executing(false)))
        } else {
+            ctx.update_upgrade_candidate_region_elapsed(now);
            Ok((Box::new(UpdateMetadata::Rollback), Status::executing(false)))
        }
    }
@@ -288,7 +291,8 @@ mod tests {
        let persistent_context = new_persistent_context();
        let env = TestingEnv::new();
        let mut ctx = env.context_factory().new_context(persistent_context);
-        ctx.volatile_ctx.operations_elapsed = ctx.persistent_ctx.timeout + Duration::from_secs(1);
+        ctx.volatile_ctx.metrics.operations_elapsed =
+            ctx.persistent_ctx.timeout + Duration::from_secs(1);

        let err = state.upgrade_region(&ctx).await.unwrap_err();

@@ -558,7 +562,8 @@ mod tests {
        let mut ctx = env.context_factory().new_context(persistent_context);
        let mailbox_ctx = env.mailbox_context();
        let mailbox = mailbox_ctx.mailbox().clone();
-        ctx.volatile_ctx.operations_elapsed = ctx.persistent_ctx.timeout + Duration::from_secs(1);
+        ctx.volatile_ctx.metrics.operations_elapsed =
+            ctx.persistent_ctx.timeout + Duration::from_secs(1);

        let (tx, rx) = tokio::sync::mpsc::channel(1);
        mailbox_ctx
--- a/src/meta-srv/src/procedure/wal_prune.rs
+++ b/src/meta-srv/src/procedure/wal_prune.rs
@@ -335,22 +335,21 @@ impl WalPruneProcedure {
            })?;
        partition_client
            .delete_records(
-                (self.data.prunable_entry_id + 1) as i64,
+                // notice here no "+1" is needed because the offset arg is exclusive, and it's defensive programming just in case somewhere else have a off by one error, see https://kafka.apache.org/36/javadoc/org/apache/kafka/clients/consumer/KafkaConsumer.html#endOffsets(java.util.Collection) which we use to get the end offset from high watermark
+                self.data.prunable_entry_id as i64,
                DELETE_RECORDS_TIMEOUT.as_millis() as i32,
            )
            .await
            .context(DeleteRecordsSnafu {
                topic: &self.data.topic,
                partition: DEFAULT_PARTITION,
-                offset: (self.data.prunable_entry_id + 1),
+                offset: self.data.prunable_entry_id,
            })
            .map_err(BoxedError::new)
            .with_context(|_| error::RetryLaterWithSourceSnafu {
                reason: format!(
                    "Failed to delete records for topic: {}, partition: {}, offset: {}",
-                    self.data.topic,
-                    DEFAULT_PARTITION,
-                    self.data.prunable_entry_id + 1
+                    self.data.topic, DEFAULT_PARTITION, self.data.prunable_entry_id
                ),
            })?;
        info!(
@@ -605,19 +604,19 @@ mod tests {
                // Step 3: Test `on_prune`.
                let status = procedure.on_prune().await.unwrap();
                assert_matches!(status, Status::Done { output: None });
-                // Check if the entry ids after `prunable_entry_id` still exist.
-                check_entry_id_existence(
-                    procedure.context.client.clone(),
-                    &topic_name,
-                    procedure.data.prunable_entry_id as i64 + 1,
-                    true,
-                )
-                .await;
-                // Check if the entry s before `prunable_entry_id` are deleted.
+                // Check if the entry ids after(include) `prunable_entry_id` still exist.
                check_entry_id_existence(
                    procedure.context.client.clone(),
                    &topic_name,
                    procedure.data.prunable_entry_id as i64,
+                    true,
+                )
+                .await;
+                // Check if the entry ids before `prunable_entry_id` are deleted.
+                check_entry_id_existence(
+                    procedure.context.client.clone(),
+                    &topic_name,
+                    procedure.data.prunable_entry_id as i64 - 1,
                    false,
                )
                .await;
--- a/src/meta-srv/src/region/supervisor.rs
+++ b/src/meta-srv/src/region/supervisor.rs
@@ -12,6 +12,7 @@
 // See the License for the specific language governing permissions and
 // limitations under the License.

+use std::collections::{HashMap, HashSet};
 use std::fmt::Debug;
 use std::sync::{Arc, Mutex};
 use std::time::Duration;
@@ -24,9 +25,9 @@ use common_meta::leadership_notifier::LeadershipChangeListener;
 use common_meta::peer::PeerLookupServiceRef;
 use common_meta::DatanodeId;
 use common_runtime::JoinHandle;
-use common_telemetry::{error, info, warn};
+use common_telemetry::{debug, error, info, warn};
 use common_time::util::current_time_millis;
-use error::Error::{MigrationRunning, TableRouteNotFound};
+use error::Error::{LeaderPeerChanged, MigrationRunning, TableRouteNotFound};
 use snafu::{OptionExt, ResultExt};
 use store_api::storage::RegionId;
 use tokio::sync::mpsc::{Receiver, Sender};
@@ -36,7 +37,9 @@ use crate::error::{self, Result};
 use crate::failure_detector::PhiAccrualFailureDetectorOptions;
 use crate::metasrv::{SelectorContext, SelectorRef};
 use crate::procedure::region_migration::manager::RegionMigrationManagerRef;
-use crate::procedure::region_migration::RegionMigrationProcedureTask;
+use crate::procedure::region_migration::{
+    RegionMigrationProcedureTask, DEFAULT_REGION_MIGRATION_TIMEOUT,
+};
 use crate::region::failure_detector::RegionFailureDetector;
 use crate::selector::SelectorOptions;

@@ -205,6 +208,8 @@ pub const DEFAULT_TICK_INTERVAL: Duration = Duration::from_secs(1);
 pub struct RegionSupervisor {
    /// Used to detect the failure of regions.
    failure_detector: RegionFailureDetector,
+    /// Tracks the number of failovers for each region.
+    failover_counts: HashMap<DetectingRegion, u32>,
    /// Receives [Event]s.
    receiver: Receiver<Event>,
    /// The context of [`SelectorRef`]
@@ -290,6 +295,7 @@ impl RegionSupervisor {
    ) -> Self {
        Self {
            failure_detector: RegionFailureDetector::new(options),
+            failover_counts: HashMap::new(),
            receiver: event_receiver,
            selector_context,
            selector,
@@ -333,13 +339,14 @@ impl RegionSupervisor {
        }
    }

-    async fn deregister_failure_detectors(&self, detecting_regions: Vec<DetectingRegion>) {
+    async fn deregister_failure_detectors(&mut self, detecting_regions: Vec<DetectingRegion>) {
        for region in detecting_regions {
-            self.failure_detector.remove(&region)
+            self.failure_detector.remove(&region);
+            self.failover_counts.remove(&region);
        }
    }

-    async fn handle_region_failures(&self, mut regions: Vec<(DatanodeId, RegionId)>) {
+    async fn handle_region_failures(&mut self, mut regions: Vec<(DatanodeId, RegionId)>) {
        if regions.is_empty() {
            return;
        }
@@ -362,16 +369,15 @@ impl RegionSupervisor {
            .collect::<Vec<_>>();

        for (datanode_id, region_id) in migrating_regions {
-            self.failure_detector.remove(&(datanode_id, region_id));
+            debug!(
+                "Removed region failover for region: {region_id}, datanode: {datanode_id} because it's migrating"
+            );
        }

        warn!("Detects region failures: {:?}", regions);
        for (datanode_id, region_id) in regions {
-            match self.do_failover(datanode_id, region_id).await {
-                Ok(_) => self.failure_detector.remove(&(datanode_id, region_id)),
-                Err(err) => {
-                    error!(err; "Failed to execute region failover for region: {region_id}, datanode: {datanode_id}");
-                }
+            if let Err(err) = self.do_failover(datanode_id, region_id).await {
+                error!(err; "Failed to execute region failover for region: {region_id}, datanode: {datanode_id}");
            }
        }
    }
@@ -383,7 +389,12 @@ impl RegionSupervisor {
            .context(error::MaintenanceModeManagerSnafu)
    }

-    async fn do_failover(&self, datanode_id: DatanodeId, region_id: RegionId) -> Result<()> {
+    async fn do_failover(&mut self, datanode_id: DatanodeId, region_id: RegionId) -> Result<()> {
+        let count = *self
+            .failover_counts
+            .entry((datanode_id, region_id))
+            .and_modify(|count| *count += 1)
+            .or_insert(1);
        let from_peer = self
            .peer_lookup
            .datanode(datanode_id)
@@ -401,6 +412,7 @@ impl RegionSupervisor {
                SelectorOptions {
                    min_required_items: 1,
                    allow_duplication: false,
+                    exclude_peer_ids: HashSet::from([from_peer.id]),
                },
            )
            .await?;
@@ -411,17 +423,44 @@ impl RegionSupervisor {
            );
            return Ok(());
        }
+        info!(
+            "Failover for region: {region_id}, from_peer: {from_peer}, to_peer: {to_peer}, tries: {count}"
+        );
        let task = RegionMigrationProcedureTask {
            region_id,
            from_peer,
            to_peer,
-            timeout: Duration::from_secs(60),
+            timeout: DEFAULT_REGION_MIGRATION_TIMEOUT * count,
        };

        if let Err(err) = self.region_migration_manager.submit_procedure(task).await {
            return match err {
                // Returns Ok if it's running or table is dropped.
-                MigrationRunning { .. } | TableRouteNotFound { .. } => Ok(()),
+                MigrationRunning { .. } => {
+                    info!(
+                        "Another region migration is running, skip failover for region: {}, datanode: {}",
+                        region_id, datanode_id
+                    );
+                    Ok(())
+                }
+                TableRouteNotFound { .. } => {
+                    self.deregister_failure_detectors(vec![(datanode_id, region_id)])
+                        .await;
+                    info!(
+                        "Table route is not found, the table is dropped, removed failover detector for region: {}, datanode: {}",
+                        region_id, datanode_id
+                    );
+                    Ok(())
+                }
+                LeaderPeerChanged { .. } => {
+                    self.deregister_failure_detectors(vec![(datanode_id, region_id)])
+                        .await;
+                    info!(
+                        "Region's leader peer changed, removed failover detector for region: {}, datanode: {}",
+                        region_id, datanode_id
+                    );
+                    Ok(())
+                }
                err => Err(err),
            };
        };
--- a/src/meta-srv/src/selector.rs
+++ b/src/meta-srv/src/selector.rs
@@ -12,15 +12,18 @@
 // See the License for the specific language governing permissions and
 // limitations under the License.

-mod common;
+pub mod common;
 pub mod lease_based;
 pub mod load_based;
 pub mod round_robin;
 #[cfg(test)]
 pub(crate) mod test_utils;
-mod weight_compute;
-mod weighted_choose;
+pub mod weight_compute;
+pub mod weighted_choose;
+use std::collections::HashSet;
+
 use serde::{Deserialize, Serialize};
+use strum::AsRefStr;

 use crate::error;
 use crate::error::Result;
@@ -39,6 +42,8 @@ pub struct SelectorOptions {
    pub min_required_items: usize,
    /// Whether duplicates are allowed in the selected result, default false.
    pub allow_duplication: bool,
+    /// The peers to exclude from the selection.
+    pub exclude_peer_ids: HashSet<u64>,
 }

 impl Default for SelectorOptions {
@@ -46,12 +51,13 @@ impl Default for SelectorOptions {
        Self {
            min_required_items: 1,
            allow_duplication: false,
+            exclude_peer_ids: HashSet::new(),
        }
    }
 }

 /// [`SelectorType`] refers to the load balancer used when creating tables.
-#[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize, Default)]
+#[derive(Clone, Debug, PartialEq, Eq, Serialize, Deserialize, Default, AsRefStr)]
 #[serde(try_from = "String")]
 pub enum SelectorType {
    /// The current load balancing is based on the number of regions on each datanode node;
--- a/src/meta-srv/src/selector/common.rs
+++ b/src/meta-srv/src/selector/common.rs
@@ -12,15 +12,25 @@
 // See the License for the specific language governing permissions and
 // limitations under the License.

+use std::collections::HashSet;
+
 use common_meta::peer::Peer;
 use snafu::ensure;

 use crate::error;
 use crate::error::Result;
 use crate::metasrv::SelectTarget;
-use crate::selector::weighted_choose::WeightedChoose;
+use crate::selector::weighted_choose::{WeightedChoose, WeightedItem};
 use crate::selector::SelectorOptions;

+/// Filter out the excluded peers from the `weight_array`.
+pub fn filter_out_excluded_peers(
+    weight_array: &mut Vec<WeightedItem<Peer>>,
+    exclude_peer_ids: &HashSet<u64>,
+) {
+    weight_array.retain(|peer| !exclude_peer_ids.contains(&peer.item.id));
+}
+
 /// According to the `opts`, choose peers from the `weight_array` through `weighted_choose`.
 pub fn choose_items<W>(opts: &SelectorOptions, weighted_choose: &mut W) -> Result<Vec<Peer>>
 where
@@ -80,7 +90,7 @@ mod tests {

    use common_meta::peer::Peer;

-    use crate::selector::common::choose_items;
+    use crate::selector::common::{choose_items, filter_out_excluded_peers};
    use crate::selector::weighted_choose::{RandomWeightedChoose, WeightedItem};
    use crate::selector::SelectorOptions;

@@ -92,35 +102,35 @@ mod tests {
                    id: 1,
                    addr: "127.0.0.1:3001".to_string(),
                },
-                weight: 1,
+                weight: 1.0,
            },
            WeightedItem {
                item: Peer {
                    id: 2,
                    addr: "127.0.0.1:3001".to_string(),
                },
-                weight: 1,
+                weight: 1.0,
            },
            WeightedItem {
                item: Peer {
                    id: 3,
                    addr: "127.0.0.1:3001".to_string(),
                },
-                weight: 1,
+                weight: 1.0,
            },
            WeightedItem {
                item: Peer {
                    id: 4,
                    addr: "127.0.0.1:3001".to_string(),
                },
-                weight: 1,
+                weight: 1.0,
            },
            WeightedItem {
                item: Peer {
                    id: 5,
                    addr: "127.0.0.1:3001".to_string(),
                },
-                weight: 1,
+                weight: 1.0,
            },
        ];

@@ -128,6 +138,7 @@ mod tests {
            let opts = SelectorOptions {
                min_required_items: i,
                allow_duplication: false,
+                exclude_peer_ids: HashSet::new(),
            };

            let selected_peers: HashSet<_> =
@@ -142,6 +153,7 @@ mod tests {
        let opts = SelectorOptions {
            min_required_items: 6,
            allow_duplication: false,
+            exclude_peer_ids: HashSet::new(),
        };

        let selected_result =
@@ -152,6 +164,7 @@ mod tests {
            let opts = SelectorOptions {
                min_required_items: i,
                allow_duplication: true,
+                exclude_peer_ids: HashSet::new(),
            };

            let selected_peers =
@@ -160,4 +173,30 @@ mod tests {
            assert_eq!(i, selected_peers.len());
        }
    }
+
+    #[test]
+    fn test_filter_out_excluded_peers() {
+        let mut weight_array = vec![
+            WeightedItem {
+                item: Peer {
+                    id: 1,
+                    addr: "127.0.0.1:3001".to_string(),
+                },
+                weight: 1.0,
+            },
+            WeightedItem {
+                item: Peer {
+                    id: 2,
+                    addr: "127.0.0.1:3002".to_string(),
+                },
+                weight: 1.0,
+            },
+        ];
+
+        let exclude_peer_ids = HashSet::from([1]);
+        filter_out_excluded_peers(&mut weight_array, &exclude_peer_ids);
+
+        assert_eq!(weight_array.len(), 1);
+        assert_eq!(weight_array[0].item.id, 2);
+    }
 }
--- a/src/meta-srv/src/selector/lease_based.rs
+++ b/src/meta-srv/src/selector/lease_based.rs
@@ -12,17 +12,37 @@
 // See the License for the specific language governing permissions and
 // limitations under the License.

+use std::collections::HashSet;
+use std::sync::Arc;
+
 use common_meta::peer::Peer;

 use crate::error::Result;
 use crate::lease;
 use crate::metasrv::SelectorContext;
-use crate::selector::common::choose_items;
+use crate::node_excluder::NodeExcluderRef;
+use crate::selector::common::{choose_items, filter_out_excluded_peers};
 use crate::selector::weighted_choose::{RandomWeightedChoose, WeightedItem};
 use crate::selector::{Selector, SelectorOptions};

 /// Select all alive datanodes based using a random weighted choose.
-pub struct LeaseBasedSelector;
+pub struct LeaseBasedSelector {
+    node_excluder: NodeExcluderRef,
+}
+
+impl LeaseBasedSelector {
+    pub fn new(node_excluder: NodeExcluderRef) -> Self {
+        Self { node_excluder }
+    }
+}
+
+impl Default for LeaseBasedSelector {
+    fn default() -> Self {
+        Self {
+            node_excluder: Arc::new(Vec::new()),
+        }
+    }
+}

 #[async_trait::async_trait]
 impl Selector for LeaseBasedSelector {
@@ -35,18 +55,26 @@ impl Selector for LeaseBasedSelector {
            lease::alive_datanodes(&ctx.meta_peer_client, ctx.datanode_lease_secs).await?;

        // 2. compute weight array, but the weight of each item is the same.
-        let weight_array = lease_kvs
+        let mut weight_array = lease_kvs
            .into_iter()
            .map(|(k, v)| WeightedItem {
                item: Peer {
                    id: k.node_id,
                    addr: v.node_addr.clone(),
                },
-                weight: 1,
+                weight: 1.0,
            })
            .collect();

        // 3. choose peers by weight_array.
+        let mut exclude_peer_ids = self
+            .node_excluder
+            .excluded_datanode_ids()
+            .iter()
+            .cloned()
+            .collect::<HashSet<_>>();
+        exclude_peer_ids.extend(opts.exclude_peer_ids.iter());
+        filter_out_excluded_peers(&mut weight_array, &exclude_peer_ids);
        let mut weighted_choose = RandomWeightedChoose::new(weight_array);
        let selected = choose_items(&opts, &mut weighted_choose)?;

--- a/src/meta-srv/src/selector/load_based.rs
+++ b/src/meta-srv/src/selector/load_based.rs
@@ -12,7 +12,8 @@
 // See the License for the specific language governing permissions and
 // limitations under the License.

-use std::collections::HashMap;
+use std::collections::{HashMap, HashSet};
+use std::sync::Arc;

 use common_meta::datanode::{DatanodeStatKey, DatanodeStatValue};
 use common_meta::key::TableMetadataManager;
@@ -26,18 +27,23 @@ use crate::error::{self, Result};
 use crate::key::{DatanodeLeaseKey, LeaseValue};
 use crate::lease;
 use crate::metasrv::SelectorContext;
-use crate::selector::common::choose_items;
+use crate::node_excluder::NodeExcluderRef;
+use crate::selector::common::{choose_items, filter_out_excluded_peers};
 use crate::selector::weight_compute::{RegionNumsBasedWeightCompute, WeightCompute};
 use crate::selector::weighted_choose::RandomWeightedChoose;
 use crate::selector::{Selector, SelectorOptions};

 pub struct LoadBasedSelector<C> {
    weight_compute: C,
+    node_excluder: NodeExcluderRef,
 }

 impl<C> LoadBasedSelector<C> {
-    pub fn new(weight_compute: C) -> Self {
-        Self { weight_compute }
+    pub fn new(weight_compute: C, node_excluder: NodeExcluderRef) -> Self {
+        Self {
+            weight_compute,
+            node_excluder,
+        }
    }
 }

@@ -45,6 +51,7 @@ impl Default for LoadBasedSelector<RegionNumsBasedWeightCompute> {
    fn default() -> Self {
        Self {
            weight_compute: RegionNumsBasedWeightCompute,
+            node_excluder: Arc::new(Vec::new()),
        }
    }
 }
@@ -85,9 +92,17 @@ where
        };

        // 4. compute weight array.
-        let weight_array = self.weight_compute.compute(&stat_kvs);
+        let mut weight_array = self.weight_compute.compute(&stat_kvs);

        // 5. choose peers by weight_array.
+        let mut exclude_peer_ids = self
+            .node_excluder
+            .excluded_datanode_ids()
+            .iter()
+            .cloned()
+            .collect::<HashSet<_>>();
+        exclude_peer_ids.extend(opts.exclude_peer_ids.iter());
+        filter_out_excluded_peers(&mut weight_array, &exclude_peer_ids);
        let mut weighted_choose = RandomWeightedChoose::new(weight_array);
        let selected = choose_items(&opts, &mut weighted_choose)?;

--- a/src/meta-srv/src/selector/round_robin.rs
+++ b/src/meta-srv/src/selector/round_robin.rs
@@ -12,7 +12,9 @@
 // See the License for the specific language governing permissions and
 // limitations under the License.

+use std::collections::HashSet;
 use std::sync::atomic::AtomicUsize;
+use std::sync::Arc;

 use common_meta::peer::Peer;
 use snafu::ensure;
@@ -20,6 +22,7 @@ use snafu::ensure;
 use crate::error::{NoEnoughAvailableNodeSnafu, Result};
 use crate::lease;
 use crate::metasrv::{SelectTarget, SelectorContext};
+use crate::node_excluder::NodeExcluderRef;
 use crate::selector::{Selector, SelectorOptions};

 /// Round-robin selector that returns the next peer in the list in sequence.
@@ -32,6 +35,7 @@ use crate::selector::{Selector, SelectorOptions};
 pub struct RoundRobinSelector {
    select_target: SelectTarget,
    counter: AtomicUsize,
+    node_excluder: NodeExcluderRef,
 }

 impl Default for RoundRobinSelector {
@@ -39,32 +43,38 @@ impl Default for RoundRobinSelector {
        Self {
            select_target: SelectTarget::Datanode,
            counter: AtomicUsize::new(0),
+            node_excluder: Arc::new(Vec::new()),
        }
    }
 }

 impl RoundRobinSelector {
-    pub fn new(select_target: SelectTarget) -> Self {
+    pub fn new(select_target: SelectTarget, node_excluder: NodeExcluderRef) -> Self {
        Self {
            select_target,
+            node_excluder,
            ..Default::default()
        }
    }

-    async fn get_peers(
-        &self,
-        min_required_items: usize,
-        ctx: &SelectorContext,
-    ) -> Result<Vec<Peer>> {
+    async fn get_peers(&self, opts: &SelectorOptions, ctx: &SelectorContext) -> Result<Vec<Peer>> {
        let mut peers = match self.select_target {
            SelectTarget::Datanode => {
                // 1. get alive datanodes.
                let lease_kvs =
                    lease::alive_datanodes(&ctx.meta_peer_client, ctx.datanode_lease_secs).await?;

+                let mut exclude_peer_ids = self
+                    .node_excluder
+                    .excluded_datanode_ids()
+                    .iter()
+                    .cloned()
+                    .collect::<HashSet<_>>();
+                exclude_peer_ids.extend(opts.exclude_peer_ids.iter());
                // 2. map into peers
                lease_kvs
                    .into_iter()
+                    .filter(|(k, _)| !exclude_peer_ids.contains(&k.node_id))
                    .map(|(k, v)| Peer::new(k.node_id, v.node_addr))
                    .collect::<Vec<_>>()
            }
@@ -84,8 +94,8 @@ impl RoundRobinSelector {
        ensure!(
            !peers.is_empty(),
            NoEnoughAvailableNodeSnafu {
-                required: min_required_items,
-                available: 0usize,
+                required: opts.min_required_items,
+                available: peers.len(),
                select_target: self.select_target
            }
        );
@@ -103,7 +113,7 @@ impl Selector for RoundRobinSelector {
    type Output = Vec<Peer>;

    async fn select(&self, ctx: &Self::Context, opts: SelectorOptions) -> Result<Vec<Peer>> {
-        let peers = self.get_peers(opts.min_required_items, ctx).await?;
+        let peers = self.get_peers(&opts, ctx).await?;
        // choose peers
        let mut selected = Vec::with_capacity(opts.min_required_items);
        for _ in 0..opts.min_required_items {
@@ -120,6 +130,8 @@ impl Selector for RoundRobinSelector {

 #[cfg(test)]
 mod test {
+    use std::collections::HashSet;
+
    use super::*;
    use crate::test_util::{create_selector_context, put_datanodes};

@@ -149,6 +161,7 @@ mod test {
                SelectorOptions {
                    min_required_items: 4,
                    allow_duplication: true,
+                    exclude_peer_ids: HashSet::new(),
                },
            )
            .await
@@ -165,6 +178,7 @@ mod test {
                SelectorOptions {
                    min_required_items: 2,
                    allow_duplication: true,
+                    exclude_peer_ids: HashSet::new(),
                },
            )
            .await
@@ -172,4 +186,42 @@ mod test {
        assert_eq!(peers.len(), 2);
        assert_eq!(peers, vec![peer2.clone(), peer3.clone()]);
    }
+
+    #[tokio::test]
+    async fn test_round_robin_selector_with_exclude_peer_ids() {
+        let selector = RoundRobinSelector::new(SelectTarget::Datanode, Arc::new(vec![5]));
+        let ctx = create_selector_context();
+        // add three nodes
+        let peer1 = Peer {
+            id: 2,
+            addr: "node1".to_string(),
+        };
+        let peer2 = Peer {
+            id: 5,
+            addr: "node2".to_string(),
+        };
+        let peer3 = Peer {
+            id: 8,
+            addr: "node3".to_string(),
+        };
+        put_datanodes(
+            &ctx.meta_peer_client,
+            vec![peer1.clone(), peer2.clone(), peer3.clone()],
+        )
+        .await;
+
+        let peers = selector
+            .select(
+                &ctx,
+                SelectorOptions {
+                    min_required_items: 1,
+                    allow_duplication: true,
+                    exclude_peer_ids: HashSet::from([2]),
+                },
+            )
+            .await
+            .unwrap();
+        assert_eq!(peers.len(), 1);
+        assert_eq!(peers, vec![peer3.clone()]);
+    }
 }
--- a/src/meta-srv/src/selector/weight_compute.rs
+++ b/src/meta-srv/src/selector/weight_compute.rs
@@ -84,7 +84,7 @@ impl WeightCompute for RegionNumsBasedWeightCompute {
            .zip(region_nums)
            .map(|(peer, region_num)| WeightedItem {
                item: peer,
-                weight: (max_weight - region_num + base_weight) as usize,
+                weight: (max_weight - region_num + base_weight) as f64,
            })
            .collect()
    }
@@ -148,7 +148,7 @@ mod tests {
            2,
        );
        for weight in weight_array.iter() {
-            assert_eq!(*expected.get(&weight.item).unwrap(), weight.weight,);
+            assert_eq!(*expected.get(&weight.item).unwrap(), weight.weight as usize);
        }

        let mut expected = HashMap::new();
--- a/src/meta-srv/src/selector/weighted_choose.rs
+++ b/src/meta-srv/src/selector/weighted_choose.rs
@@ -42,10 +42,10 @@ pub trait WeightedChoose<Item>: Send + Sync {
 }

 /// The struct represents a weighted item.
-#[derive(Debug, Clone, PartialEq, Eq)]
+#[derive(Debug, Clone, PartialEq)]
 pub struct WeightedItem<Item> {
    pub item: Item,
-    pub weight: usize,
+    pub weight: f64,
 }

 /// A implementation of weighted balance: random weighted choose.
@@ -87,7 +87,7 @@ where
        // unwrap safety: whether weighted_index is none has been checked before.
        let item = self
            .items
-            .choose_weighted(&mut rng(), |item| item.weight as f64)
+            .choose_weighted(&mut rng(), |item| item.weight)
            .context(error::ChooseItemsSnafu)?
            .item
            .clone();
@@ -95,11 +95,11 @@ where
    }

    fn choose_multiple(&mut self, amount: usize) -> Result<Vec<Item>> {
-        let amount = amount.min(self.items.iter().filter(|item| item.weight > 0).count());
+        let amount = amount.min(self.items.iter().filter(|item| item.weight > 0.0).count());

        Ok(self
            .items
-            .choose_multiple_weighted(&mut rng(), amount, |item| item.weight as f64)
+            .choose_multiple_weighted(&mut rng(), amount, |item| item.weight)
            .context(error::ChooseItemsSnafu)?
            .cloned()
            .map(|item| item.item)
@@ -120,9 +120,12 @@ mod tests {
        let mut choose = RandomWeightedChoose::new(vec![
            WeightedItem {
                item: 1,
-                weight: 100,
+                weight: 100.0,
+            },
+            WeightedItem {
+                item: 2,
+                weight: 0.0,
            },
-            WeightedItem { item: 2, weight: 0 },
        ]);

        for _ in 0..100 {
--- a/src/meta-srv/src/table_meta_alloc.rs
+++ b/src/meta-srv/src/table_meta_alloc.rs
@@ -12,6 +12,8 @@
 // See the License for the specific language governing permissions and
 // limitations under the License.

+use std::collections::HashSet;
+
 use async_trait::async_trait;
 use common_error::ext::BoxedError;
 use common_meta::ddl::table_meta::PeerAllocator;
@@ -51,6 +53,7 @@ impl MetasrvPeerAllocator {
                SelectorOptions {
                    min_required_items: regions,
                    allow_duplication: true,
+                    exclude_peer_ids: HashSet::new(),
                },
            )
            .await?;
--- a/src/metric-engine/Cargo.toml
+++ b/src/metric-engine/Cargo.toml
@@ -18,11 +18,13 @@ common-error.workspace = true
 common-macro.workspace = true
 common-query.workspace = true
 common-recordbatch.workspace = true
+common-runtime.workspace = true
 common-telemetry.workspace = true
 common-time.workspace = true
 datafusion.workspace = true
 datatypes.workspace = true
 futures-util.workspace = true
+humantime-serde.workspace = true
 itertools.workspace = true
 lazy_static = "1.4"
 mito2.workspace = true
--- a/src/metric-engine/src/config.rs
+++ b/src/metric-engine/src/config.rs
@@ -12,9 +12,49 @@
 // See the License for the specific language governing permissions and
 // limitations under the License.

+use std::time::Duration;
+
+use common_telemetry::warn;
 use serde::{Deserialize, Serialize};

-#[derive(Debug, Clone, Default, Serialize, Deserialize, PartialEq, Eq)]
+/// The default flush interval of the metadata region.  
+pub(crate) const DEFAULT_FLUSH_METADATA_REGION_INTERVAL: Duration = Duration::from_secs(30);
+
+/// Configuration for the metric engine.
+#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq)]
 pub struct EngineConfig {
+    /// Experimental feature to use sparse primary key encoding.
    pub experimental_sparse_primary_key_encoding: bool,
+    /// The flush interval of the metadata region.
+    #[serde(
+        with = "humantime_serde",
+        default = "EngineConfig::default_flush_metadata_region_interval"
+    )]
+    pub flush_metadata_region_interval: Duration,
+}
+
+impl Default for EngineConfig {
+    fn default() -> Self {
+        Self {
+            flush_metadata_region_interval: DEFAULT_FLUSH_METADATA_REGION_INTERVAL,
+            experimental_sparse_primary_key_encoding: false,
+        }
+    }
+}
+
+impl EngineConfig {
+    fn default_flush_metadata_region_interval() -> Duration {
+        DEFAULT_FLUSH_METADATA_REGION_INTERVAL
+    }
+
+    /// Sanitizes the configuration.
+    pub fn sanitize(&mut self) {
+        if self.flush_metadata_region_interval.is_zero() {
+            warn!(
+                "Flush metadata region interval is zero, override with default value: {:?}. Disable metadata region flush is forbidden.",
+                DEFAULT_FLUSH_METADATA_REGION_INTERVAL
+            );
+            self.flush_metadata_region_interval = DEFAULT_FLUSH_METADATA_REGION_INTERVAL;
+        }
+    }
 }
--- a/src/metric-engine/src/engine.rs
+++ b/src/metric-engine/src/engine.rs
@@ -34,9 +34,11 @@ use api::region::RegionResponse;
 use async_trait::async_trait;
 use common_error::ext::{BoxedError, ErrorExt};
 use common_error::status_code::StatusCode;
+use common_runtime::RepeatedTask;
 use mito2::engine::MitoEngine;
 pub(crate) use options::IndexOptions;
 use snafu::ResultExt;
+pub(crate) use state::MetricEngineState;
 use store_api::metadata::RegionMetadataRef;
 use store_api::metric_engine_consts::METRIC_ENGINE_NAME;
 use store_api::region_engine::{
@@ -47,11 +49,11 @@ use store_api::region_engine::{
 use store_api::region_request::{BatchRegionDdlRequest, RegionRequest};
 use store_api::storage::{RegionId, ScanRequest, SequenceNumber};

-use self::state::MetricEngineState;
 use crate::config::EngineConfig;
 use crate::data_region::DataRegion;
-use crate::error::{self, Result, UnsupportedRegionRequestSnafu};
+use crate::error::{self, Error, Result, StartRepeatedTaskSnafu, UnsupportedRegionRequestSnafu};
 use crate::metadata_region::MetadataRegion;
+use crate::repeated_task::FlushMetadataRegionTask;
 use crate::row_modifier::RowModifier;
 use crate::utils::{self, get_region_statistic};

@@ -359,19 +361,32 @@ impl RegionEngine for MetricEngine {
 }

 impl MetricEngine {
-    pub fn new(mito: MitoEngine, config: EngineConfig) -> Self {
+    pub fn try_new(mito: MitoEngine, mut config: EngineConfig) -> Result<Self> {
        let metadata_region = MetadataRegion::new(mito.clone());
        let data_region = DataRegion::new(mito.clone());
-        Self {
-            inner: Arc::new(MetricEngineInner {
-                mito,
-                metadata_region,
-                data_region,
-                state: RwLock::default(),
-                config,
-                row_modifier: RowModifier::new(),
-            }),
-        }
+        let state = Arc::new(RwLock::default());
+        config.sanitize();
+        let flush_interval = config.flush_metadata_region_interval;
+        let inner = Arc::new(MetricEngineInner {
+            mito: mito.clone(),
+            metadata_region,
+            data_region,
+            state: state.clone(),
+            config,
+            row_modifier: RowModifier::new(),
+            flush_task: RepeatedTask::new(
+                flush_interval,
+                Box::new(FlushMetadataRegionTask {
+                    state: state.clone(),
+                    mito: mito.clone(),
+                }),
+            ),
+        });
+        inner
+            .flush_task
+            .start(common_runtime::global_runtime())
+            .context(StartRepeatedTaskSnafu { name: "flush_task" })?;
+        Ok(Self { inner })
    }

    pub fn mito(&self) -> MitoEngine {
@@ -426,15 +441,21 @@ impl MetricEngine {
    ) -> Result<common_recordbatch::SendableRecordBatchStream, BoxedError> {
        self.inner.scan_to_stream(region_id, request).await
    }
+
+    /// Returns the configuration of the engine.
+    pub fn config(&self) -> &EngineConfig {
+        &self.inner.config
+    }
 }

 struct MetricEngineInner {
    mito: MitoEngine,
    metadata_region: MetadataRegion,
    data_region: DataRegion,
-    state: RwLock<MetricEngineState>,
+    state: Arc<RwLock<MetricEngineState>>,
    config: EngineConfig,
    row_modifier: RowModifier,
+    flush_task: RepeatedTask<Error>,
 }

 #[cfg(test)]
--- a/src/metric-engine/src/engine/create.rs
+++ b/src/metric-engine/src/engine/create.rs
@@ -737,7 +737,7 @@ mod test {

        // set up
        let env = TestEnv::new().await;
-        let engine = MetricEngine::new(env.mito(), EngineConfig::default());
+        let engine = MetricEngine::try_new(env.mito(), EngineConfig::default()).unwrap();
        let engine_inner = engine.inner;

        // check create data region request
--- a/src/metric-engine/src/error.rs
+++ b/src/metric-engine/src/error.rs
@@ -282,6 +282,14 @@ pub enum Error {
        #[snafu(implicit)]
        location: Location,
    },
+
+    #[snafu(display("Failed to start repeated task: {}", name))]
+    StartRepeatedTask {
+        name: String,
+        source: common_runtime::error::Error,
+        #[snafu(implicit)]
+        location: Location,
+    },
 }

 pub type Result<T, E = Error> = std::result::Result<T, E>;
@@ -335,6 +343,8 @@ impl ErrorExt for Error {

            CollectRecordBatchStream { source, .. } => source.status_code(),

+            StartRepeatedTask { source, .. } => source.status_code(),
+
            MetricManifestInfo { .. } => StatusCode::Internal,
        }
    }
--- a/src/metric-engine/src/lib.rs
+++ b/src/metric-engine/src/lib.rs
@@ -59,6 +59,7 @@ pub mod engine;
 pub mod error;
 mod metadata_region;
 mod metrics;
+mod repeated_task;
 pub mod row_modifier;
 #[cfg(test)]
 mod test_util;
--- a/src/metric-engine/src/repeated_task.rs
+++ b/src/metric-engine/src/repeated_task.rs
@@ -0,0 +1,167 @@
+// Copyright 2023 Greptime Team
+//
+// Licensed under the Apache License, Version 2.0 (the "License");
+// you may not use this file except in compliance with the License.
+// You may obtain a copy of the License at
+//
+//     http://www.apache.org/licenses/LICENSE-2.0
+//
+// Unless required by applicable law or agreed to in writing, software
+// distributed under the License is distributed on an "AS IS" BASIS,
+// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+// See the License for the specific language governing permissions and
+// limitations under the License.
+
+use std::sync::{Arc, RwLock};
+use std::time::Instant;
+
+use common_runtime::TaskFunction;
+use common_telemetry::{debug, error};
+use mito2::engine::MitoEngine;
+use store_api::region_engine::{RegionEngine, RegionRole};
+use store_api::region_request::{RegionFlushRequest, RegionRequest};
+
+use crate::engine::MetricEngineState;
+use crate::error::{Error, Result};
+use crate::utils;
+
+/// Task to flush metadata regions.
+///
+/// This task is used to send flush requests to the metadata regions
+/// periodically.
+pub(crate) struct FlushMetadataRegionTask {
+    pub(crate) state: Arc<RwLock<MetricEngineState>>,
+    pub(crate) mito: MitoEngine,
+}
+
+#[async_trait::async_trait]
+impl TaskFunction<Error> for FlushMetadataRegionTask {
+    fn name(&self) -> &str {
+        "FlushMetadataRegionTask"
+    }
+
+    async fn call(&mut self) -> Result<()> {
+        let region_ids = {
+            let state = self.state.read().unwrap();
+            state
+                .physical_region_states()
+                .keys()
+                .cloned()
+                .collect::<Vec<_>>()
+        };
+
+        let num_region = region_ids.len();
+        let now = Instant::now();
+        for region_id in region_ids {
+            let Some(role) = self.mito.role(region_id) else {
+                continue;
+            };
+            if role == RegionRole::Follower {
+                continue;
+            }
+            let metadata_region_id = utils::to_metadata_region_id(region_id);
+            if let Err(e) = self
+                .mito
+                .handle_request(
+                    metadata_region_id,
+                    RegionRequest::Flush(RegionFlushRequest {
+                        row_group_size: None,
+                    }),
+                )
+                .await
+            {
+                error!(e; "Failed to flush metadata region {}", metadata_region_id);
+            }
+        }
+        debug!(
+            "Flushed {} metadata regions, elapsed: {:?}",
+            num_region,
+            now.elapsed()
+        );
+
+        Ok(())
+    }
+}
+
+#[cfg(test)]
+mod tests {
+    use std::assert_matches::assert_matches;
+    use std::time::Duration;
+
+    use store_api::region_engine::{RegionEngine, RegionManifestInfo};
+
+    use crate::config::{EngineConfig, DEFAULT_FLUSH_METADATA_REGION_INTERVAL};
+    use crate::test_util::TestEnv;
+
+    #[tokio::test]
+    async fn test_flush_metadata_region_task() {
+        let env = TestEnv::with_prefix_and_config(
+            "test_flush_metadata_region_task",
+            EngineConfig {
+                flush_metadata_region_interval: Duration::from_millis(100),
+                ..Default::default()
+            },
+        )
+        .await;
+        env.init_metric_region().await;
+        let engine = env.metric();
+        // Wait for flush task run
+        tokio::time::sleep(Duration::from_millis(200)).await;
+        let physical_region_id = env.default_physical_region_id();
+        let stat = engine.region_statistic(physical_region_id).unwrap();
+
+        assert_matches!(
+            stat.manifest,
+            RegionManifestInfo::Metric {
+                metadata_manifest_version: 1,
+                metadata_flushed_entry_id: 1,
+                ..
+            }
+        )
+    }
+
+    #[tokio::test]
+    async fn test_flush_metadata_region_task_with_long_interval() {
+        let env = TestEnv::with_prefix_and_config(
+            "test_flush_metadata_region_task_with_long_interval",
+            EngineConfig {
+                flush_metadata_region_interval: Duration::from_secs(60),
+                ..Default::default()
+            },
+        )
+        .await;
+        env.init_metric_region().await;
+        let engine = env.metric();
+        // Wait for flush task run, should not flush metadata region
+        tokio::time::sleep(Duration::from_millis(200)).await;
+        let physical_region_id = env.default_physical_region_id();
+        let stat = engine.region_statistic(physical_region_id).unwrap();
+
+        assert_matches!(
+            stat.manifest,
+            RegionManifestInfo::Metric {
+                metadata_manifest_version: 0,
+                metadata_flushed_entry_id: 0,
+                ..
+            }
+        )
+    }
+
+    #[tokio::test]
+    async fn test_flush_metadata_region_sanitize() {
+        let env = TestEnv::with_prefix_and_config(
+            "test_flush_metadata_region_sanitize",
+            EngineConfig {
+                flush_metadata_region_interval: Duration::from_secs(0),
+                ..Default::default()
+            },
+        )
+        .await;
+        let metric = env.metric();
+        let config = metric.config();
+        assert_eq!(
+            config.flush_metadata_region_interval,
+            DEFAULT_FLUSH_METADATA_REGION_INTERVAL
+        );
+    }
+}
--- a/src/metric-engine/src/test_util.rs
+++ b/src/metric-engine/src/test_util.rs
@@ -54,9 +54,14 @@ impl TestEnv {

    /// Returns a new env with specific `prefix` for test.
    pub async fn with_prefix(prefix: &str) -> Self {
+        Self::with_prefix_and_config(prefix, EngineConfig::default()).await
+    }
+
+    /// Returns a new env with specific `prefix` and `config` for test.
+    pub async fn with_prefix_and_config(prefix: &str, config: EngineConfig) -> Self {
        let mut mito_env = MitoTestEnv::with_prefix(prefix);
        let mito = mito_env.create_engine(MitoConfig::default()).await;
-        let metric = MetricEngine::new(mito.clone(), EngineConfig::default());
+        let metric = MetricEngine::try_new(mito.clone(), config).unwrap();
        Self {
            mito_env,
            mito,
@@ -84,7 +89,7 @@ impl TestEnv {
            .mito_env
            .create_follower_engine(MitoConfig::default())
            .await;
-        let metric = MetricEngine::new(mito.clone(), EngineConfig::default());
+        let metric = MetricEngine::try_new(mito.clone(), EngineConfig::default()).unwrap();

        let region_id = self.default_physical_region_id();
        debug!("opening default physical region: {region_id}");
--- a/Show More
+++ b/Show More
Author	SHA1	Message	Date
discord9	66e2242e46	fix: conn timeout&refactor: better err msg (#5974 ) * fix: conn timeout&refactor: better err msg * chore: clippy * chore: make test work * chore: comment * todo: fix null cast * fix: retry conn&udd_calc * chore: comment * chore: apply suggestion --------- Co-authored-by: dennis zhuang <killme2008@gmail.com>	2025-04-25 19:12:30 +00:00
Ning Sun	489b16ae30	fix: security update (#5982 )	2025-04-25 18:11:09 +00:00
dennis zhuang	85d564b0fb	fix: upgrade sqlparse and validate align in range query (#5958 ) * fix: upgrade sqlparse and validate align in range query * update sqlparser to the merged commit Signed-off-by: Ruihang Xia <waynestxia@gmail.com> --------- Signed-off-by: Ruihang Xia <waynestxia@gmail.com> Co-authored-by: Ruihang Xia <waynestxia@gmail.com> Co-authored-by: Zhenchi <zhongzc_arch@outlook.com>	2025-04-25 17:34:49 +00:00
Zhenchi	d5026f3491	perf: optimize fulltext zh tokenizer for ascii-only text (#5975 ) Signed-off-by: Zhenchi <zhongzc_arch@outlook.com>	2025-04-24 23:31:26 +00:00
Weny Xu	e30753fc31	feat: allow forced region failover for local WAL (#5972 ) * feat: allow forced region failover for local WAL * chore: upgrade config.md * chore: apply suggestions from CR	2025-04-24 08:11:45 +00:00
Ruihang Xia	b476584f56	feat: remove hyper parameter from promql functions (#5955 ) * quantile udaf Signed-off-by: Ruihang Xia <waynestxia@gmail.com> * extrapolate rate Signed-off-by: Ruihang Xia <waynestxia@gmail.com> * predict_linear, round, holt_winters, quantile_overtime Signed-off-by: Ruihang Xia <waynestxia@gmail.com> * fix clippy Signed-off-by: Ruihang Xia <waynestxia@gmail.com> * fix quantile function Signed-off-by: Ruihang Xia <waynestxia@gmail.com> --------- Signed-off-by: Ruihang Xia <waynestxia@gmail.com>	2025-04-24 07:17:10 +00:00
Weny Xu	ff3a46b1d0	feat: improve observability of region migration procedure (#5967 ) * feat: improve observability of region migration procedure * chore: apply suggestions from CR * chore: observe non-zero value	2025-04-24 04:00:14 +00:00
Weny Xu	a533ac2555	feat: enhance selector with node exclusion support (#5966 )	2025-04-24 02:27:27 +00:00
dennis zhuang	cc5629b4a1	chore: remove coderabbit (#5969 )	2025-04-24 02:15:44 +00:00
Weny Xu	f3d000f6ec	feat: track region failover attempts and adjust timeout (#5952 )	2025-04-23 18:19:18 +00:00
discord9	9557b76224	fix: try prune one less (#5965 ) * try prune one less * test: also not add one * ci: use longer fuzz time * revert fuzz time&per review * chore: no ( * docs: add explain to offset used in delete records * test: fix test_procedure_execution	2025-04-23 16:57:54 +00:00
discord9	a0900f5b90	feat(flow): use batching mode&fix sqlness (#5903 ) * feat: use flow batching engine broken: try using logical plan fix: use dummy catalog for logical plan fix: insert plan exec&sqlness grpc addr feat: use frontend instance in flownode in standalone feat: flow type in metasrv&fix: flush flow out of sync& column name alias tests: sqlness update tests: sqlness flow rebuild udpate chore: per review refactor: keep chnl mgr refactor: use catalog mgr for get table tests: use valid sql fix: add more check refactor: put flow type determine to frontend * chore: update proto * chore: update proto to main branch * fix: add locks for create/drop flow&docs: update docs * feat: flush_flow flush all ranges now * test: add align time window test * docs: explain `nodeid` use in check task * refactor: AddAutoColumnRewriter check for Projection * refactor: per review * fix: query without time window also clean dirty time window * chore: better logging * chore: add comments per review * refactor: per review * chore: per review * chore: per review rename args * refactor: per review partially * chore: update docs * chore: use better error variant * chore: better error variant * refactor: rename FlowWorkerManager to FlowStreamingEngine * rename again * refactor: per review * chore: rebase after #5963 merged * refactor: rename all flow_worker_manager occurs * docs: rm resolved TODO	2025-04-23 15:12:16 +00:00
Yingwen	45a05fb08c	docs: fix some units and adds the opendal errors panel (#5962 ) * docs: fixes units in the dashboard * docs: add opendal errors panel * docs: opendal traffic use decbytes * docs: update readme --------- Co-authored-by: zyy17 <zyylsxm@gmail.com>	2025-04-23 13:31:29 +00:00
LFC	71db79c8d6	feat: node excluder (#5964 ) * feat: node excluder * fix ci * fix ci	2025-04-23 10:48:46 +00:00
discord9	79ed7bbc44	fix: store flow query ctx on creation (#5963 ) * fix: store flow schema on creation * chore: update sqlness * refactor: save the entire query context to flow info * chore: sqlness update * chore: rm pub * fix: keep old version compatibility	2025-04-23 09:59:09 +00:00
zyy17	02e9a66d7a	chore: update dac tools image and docs (#5961 )	2025-04-23 05:00:37 +00:00
Weny Xu	55cadcd2c0	feat: introduce flush metadata region task for metric engine (#5951 ) * feat: introduce flush metadata region task for metric engine * docs: generate config.md * chore: add header * test: fix unit test * fix: fix unit tests * chore: apply suggestions from CR * chore: remove docs * fix: fix unit tests	2025-04-23 04:51:22 +00:00
fys	8c4796734a	chore: remove unused attribute (#5960 )	2025-04-23 03:17:13 +00:00
Yuhan Wang	919956999b	fix: use max in flushed entry id and topic latest entry id (#5946 )	2025-04-22 23:48:32 +00:00
ZonaHe	7e5f6cbeae	feat: update dashboard to v0.9.0 (#5948 ) Co-authored-by: ZonaHex <ZonaHex@users.noreply.github.com>	2025-04-22 11:35:33 +00:00
shuiyisong	5c07f0dec7	refactor: `run_pipeline` parameters (#5954 ) * refactor: simplify run_pipeline params * refactor: remove unnecessory function wrap	2025-04-22 11:34:19 +00:00
discord9	9fb0487e67	fix: parse flow expire after interval (#5953 ) * fix: parse flow expire after interval * fix: correct 30.44&comments	2025-04-22 08:44:04 +00:00
discord9	6e407ae4b9	test: use random seed for window sort fuzz test (#5950 ) tests: use random seed for window sort fuzz test	2025-04-22 08:14:27 +00:00
Ning Sun	bcefc6b83f	feat: add format support for promql http api (not prometheus) (#5939 ) * feat: add format support for promql http api (not prometheus) * test: add csv format test	2025-04-22 08:10:35 +00:00
Weny Xu	0f77135ef9	feat: add `exclude_peer_ids` to `SelectorOptions` (#5949 ) * feat: add `exclude_peer_ids` to `SelectorOptions` * chore: apply suggestions from CR * fix: clippy	2025-04-22 07:49:22 +00:00
Weny Xu	0a4594c9e2	fix: remove obsolete failover detectors after region leader change (#5944 ) * fix: remove obsolete failover detectors after region leader change * chore: apply suggestions from CR * fix: fix unit tests * fix: fix unit test * fix: failover logic	2025-04-22 06:15:47 +00:00
LFC	d9437c6da7	chore: assert plugin uniqueness (#5947 )	2025-04-22 06:04:06 +00:00
zyy17	35f4fa3c3e	refactor: unify all dashboards and use `dac` tool to generate intermediate dashboards (#5933 ) * refactor: split cluster metrics into multiple dashboards * chore: merge multiple dashboards into one dashboard * refactor: add 'dac' tool to generate a intermediate dashboards * refactor: generate markdown docs for dashboards	2025-04-22 06:03:01 +00:00
jeremyhi	60e4607b64	chore: better buckets for heartbeat stat size histogram (#5945 ) chore: better buckets for METRIC_META_HEARTBEAT_STAT_MEMORY_SIZE	2025-04-21 16:12:27 +00:00
shuiyisong	3b8c6d5ce3	chore: use `once_cell` to avoid parse everytime in pipeline exec (#5943 ) * chore: use once_cell to avoid parse everytime * chore: remove pub on options	2025-04-21 12:55:48 +00:00
Weny Xu	7a8e1bc3f9	feat: support building `metasrv` with selector from plugins (#5942 ) * chore: expose selector * feat: use f64 * chore: expose selector::common * feat: build metasrv with selector from plugins	2025-04-21 10:59:24 +00:00
Yuhan Wang	ee07b9bfa8	test: update configs to enable auto wal prune (#5938 ) * test: update configs to enable auto wal prune * fix: add humantime_serde * fix: enable overwrite_entry_start_id * fix: not in metasrv * chore: update default value name * Apply suggestions from code review Co-authored-by: jeremyhi <jiachun_feng@proton.me> * fix: kafka use overwrite_entry_start_id --------- Co-authored-by: jeremyhi <jiachun_feng@proton.me>	2025-04-21 07:57:43 +00:00
Lei, HUANG	90ffaa8a62	feat: implement otel-arrow protocol for GreptimeDB (#5840 ) * [wip]: implement arrow service * add service * feat/otel-arrow: ### Add OpenTelemetry Arrow Support - `Cargo.toml`, `Cargo.lock`: Updated `otel-arrow-rust` dependency to use a local path and added `arrow-ipc` as a dependency. - `src/servers/src/grpc.rs`, `src/servers/src/grpc/builder.rs`: Integrated `ArrowMetricsServiceServer` with gRPC server, including support for custom header interception and message compression. - `src/servers/src/otel_arrow.rs`: Implemented `OtelArrowServiceHandler` for handling OpenTelemetry Arrow metrics and added `HeaderInterceptor` for custom header handling. * feat/otel-arrow: Add error handling for OpenTelemetry Arrow requests - `src/error.rs`: Introduced a new error variant `HandleOtelArrowRequest` to handle failures in processing OpenTelemetry Arrow requests. - `src/otel_arrow.rs`: Implemented error handling for receiving and consuming batches from the OpenTelemetry Arrow client. Added logging for errors and updated the response status accordingly. * feat/otel-arrow: Remove `otel_arrow` Module from gRPC Server - Deleted the `otel_arrow` module from the gRPC server implementation. - Removed the `otel_arrow` module import from `grpc.rs`. - Deleted the `otel_arrow.rs` file, which contained the `OtelArrowServer` struct and its implementation. * feat/otel-arrow: ## Remove `Arc` Implementations for Protocol and Pipeline Handlers - Removed `Arc` Implementations: Deleted `Arc` implementations for `OpenTelemetryProtocolHandler` and `PipelineHandler` traits in `query_handler.rs`. This change simplifies the code by removing redundant async trait implementations for `Arc<T>`. - File Affected: `src/servers/src/query_handler.rs` * feat/otel-arrow: Improve error handling and metadata processing in `otel_arrow.rs` - Updated error handling by ignoring the result of `sender.send` to prevent panic on failure. - Enhanced metadata processing in `HeaderInterceptor` by using `Ok` to safely handle `grpc-encoding` entry retrieval. * fix dependency * feat/otel-arrow: - Update Dependencies: - Moved `otel-arrow-rust` dependency in `Cargo.toml`. - Adjusted workspace dependencies in `src/frontend/Cargo.toml`. - Error Handling: - Removed `MissingQueryContext` error variant from `src/servers/src/error.rs`. * fix: toml format * remove useless code * chore: resolve conflicts	2025-04-21 07:24:23 +00:00
Yingwen	56f319a707	fix: filter doesn't consider default values after schema change (#5912 ) * test: sqlness test case * feat: use correct default while pruning row groups * fix: consider default in SimpleFilterContext * test: update sqlness test * test: add order by	2025-04-21 06:32:26 +00:00
shuiyisong	9df493988b	fix: wrong error msg in pipeline (#5937 )	2025-04-21 04:05:46 +00:00
dennis zhuang	ad1b77ab04	feat: update readme (#5936 ) * fix: title * chore: format * chore: format * chore: format	2025-04-21 02:44:44 +00:00