6277 Commits
Author SHA1 Message Date
Yingwen 3a4c24a863 fix: temporarily disable v2 series scan by default (#9436)
* fix: disable v2 series scan by default temporarily

Signed-off-by: evenyag <realevenyag@gmail.com>

* test: explicitly enable v2 series scan in sqlness configs

Signed-off-by: evenyag <realevenyag@gmail.com>

* test: regenerate sqlness results for legacy series scan default

Signed-off-by: evenyag <realevenyag@gmail.com>

---------

Signed-off-by: evenyag <realevenyag@gmail.com>
2026-10-02 12:09:16 +00:00
Ning Sunandevenyag a9ce6e2d01 ci: make query regression non-blocking for scheduled nightly releases (#9434)
* ci: make query regression non-blocking for scheduled nightly releases

Scheduled nightly releases no longer gate publishing on the query
regression release test. The test still runs (validation policy stays
'all' for automatic releases) and uploads its report, but any result -
failure, timeout, cancellation, or skip - no longer blocks image
publishing or the GitHub release for schedule-triggered runs.

Since reusable-workflow caller jobs cannot use continue-on-error, the
gate is relaxed at the consumers instead: the if conditions of
release-images-to-dockerhub and publish-github-release now also pass
when github.event_name == 'schedule' regardless of the query-regression
result. All other gates (runner allocation, artifact builds,
prepare-release-validation) are unchanged.

Tag-push and manual-dispatch releases remain fully blocking unless a
skip policy is explicitly chosen via the release_validation input.

Signed-off-by: Ning Sun <sunning@greptime.com>

* ci: preserve downstream nightly release jobs after regression failures

Signed-off-by: evenyag <realevenyag@gmail.com>

---------

Signed-off-by: Ning Sun <sunning@greptime.com>
Signed-off-by: evenyag <realevenyag@gmail.com>
Co-authored-by: evenyag <realevenyag@gmail.com>
2026-10-02 10:19:50 +00:00
Weny Xu acebd6b35b fix(catalog): skip backend batch_get on full cache hits (#9431)
fix(catalog): return cached batch values without an empty backend call

Count backend batch_get invocations in regression tests and cover empty input.

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-10-02 09:50:08 +00:00
Ning Sun 3e86422ef8 chore: bump version and start new 1.4 release cycle (#9403)
chore: start new 1.4 release cycle
2026-10-02 03:45:42 +00:00
Yingwen 0d502e1fc1 ci: upgrade EC2 runner action to restore runner registration (#9430)
ci: update EC2 runner action for minimum runner version enforcement

Signed-off-by: evenyag <realevenyag@gmail.com>
2026-10-01 07:04:10 +00:00
Dragon RoarandDennis Zhuang d7f5331876 feat: add CORS support for gRPC-Web on the frontend gRPC server (#9374)
* feat: add CORS support for gRPC-Web on the frontend gRPC server

A browser-based gRPC-Web client cannot read a cross-origin response. The
request carries a non-simple content type, so the browser sends an
`OPTIONS` preflight first, and it drops the call when the response has no
`Access-Control-Allow-Origin`.

Add `enable_cors` and `cors_allowed_origins` to `[grpc]`, threaded from
`GrpcOptions` to `GrpcServerConfig` the same way `tls` is.

The CORS layer sits outside `tonic_web::GrpcWebLayer`, so it answers the
preflight before the request reaches the gRPC routes. It exposes
`grpc-status`, `grpc-message` and `grpc-status-details-bin`, which a
browser cannot read otherwise. An empty `cors_allowed_origins` allows any
origin.

`enable_cors` defaults to false, unlike `[http]`. `GrpcOptions` is shared
by the public `[grpc]` section and the internal `[internal_grpc]` one, and
serde cannot tell them apart, so a true default would also turn CORS on
for the internal gRPC listeners: the frontend internal gRPC server, and
the datanode and flownode gRPC servers. None of them authenticate callers,
and a browser on the host or in the cluster network can reach all of them.
Set `enable_cors = true` to turn it on.

Tests start a real server on an ephemeral port and cover the preflight, a
custom origin list, and the disabled case.

Update the example TOMLs and regenerate config/config.md.

Signed-off-by: lczllx <2181719471@qq.com>

* fix: drop the redundant OPTIONS from the gRPC CORS allow_methods

Signed-off-by: lczllx <2181719471@qq.com>

* test: assert the allow-headers and allow-methods headers in the gRPC CORS preflight

The preflight answers with `access-control-allow-headers: *` from
`AllowHeaders::any()`, and with `access-control-allow-methods: post` from the
POST-only `allow_methods` list. Pin both, the way the HTTP CORS test pins its
own headers: a missing allow-headers header would let the preflight pass the
origin check and still have the browser block every gRPC-Web call, which a
non-browser client would never notice.

Signed-off-by: lczllx <2181719471@qq.com>

* docs(config): stop the example configs from setting the new gRPC CORS keys

Signed-off-by: lczllx <2181719471@qq.com>

* test(servers): cover the exposed gRPC CORS headers and pin the option plumbing

Signed-off-by: lczllx <2181719471@qq.com>

* docs(config): document the gRPC CORS origin example with `#+`

Signed-off-by: lczllx <2181719471@qq.com>

* fix(frontend): keep CORS off the internal gRPC server

Signed-off-by: lczllx <2181719471@qq.com>

* docs(grpc): document that the internal gRPC server never serves CORS

Signed-off-by: lczllx <2181719471@qq.com>

* Update src/servers/src/grpc.rs

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore: trim redundant comments

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor: share CORS origin parsing and simplify gRPC-Web CORS tests

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor: enable gRPC CORS only on the frontend public server

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* docs(config): clarify the gRPC CORS options

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: lczllx <2181719471@qq.com>
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
Co-authored-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-30 10:40:17 +00:00
dennis zhuang a314ac4281 fix(mito2): truncate parquet column index min/max for SST writes (#9420)
* fix(mito2): truncate parquet column index min/max for SST writes

The SST writer and the bulk part encoder disabled column index truncation
together with statistics truncation in #6977. Mito never reads the column
index, but every page of a large string or binary column still stored its
full min and max values, uncompressed. Truncate the column index to the
parquet default of 64 bytes and keep chunk statistics untruncated, since
primary key pruning decodes them.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* docs(mito2): note that the column index must not be used to decode primary keys

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-30 08:40:27 +00:00
jeremyhi 78eefbe4c4 perf: batch schema export requests (#9408)
* perf: batch schema export requests

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: propagate legacy schema export failures

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: keep credentials out of SQL response errors

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test: use a typo-safe dotted catalog name

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* refactor(cli): address schema export review feedback

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-09-30 06:23:33 +00:00
jeremyhi 8553db3f27 feat(mito): wait for WAL durability before publishing a manifest watermark (#9410)
* feat(mito): wait for WAL durability before publishing a manifest watermark

A log store that acknowledges appends before their entries are durable can
hand out entry ids that a crash loses. Before a flush, a full truncate or a
discard of unflushed data records an entry id in the manifest, Mito now waits
on `LogStore::wait_durable` for that id. The flush waits after its SSTs are
written and observes its cancellation, so a drop or truncate queued behind it
is not held up by the upload.

Add engine tests of the object store WAL across restarts and of the barrier,
and a log store testing hook that stores an object and then reports its create
as failed.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test(mito): synchronize the WAL barrier tests on store signals

The barrier tests waited with fixed sleeps, which do not establish that a
flush or truncate has reached its durability wait, and the lost backlog test
could stop the store before the seal was handled. The object store WAL gets
testing hooks that observe parked creates and durability waits, and one that
ends the actor as a crash would; the tests wait on them, crash the store and
the engine before a reopen until nothing holds the old store, and cover a
truncate or discard of unflushed data that completes once its entry is
durable.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test(mito): rename the recovery state and the crash test in the WAL recovery tests

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test(mito): drop a comment that restates the empty recovery state

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-09-30 05:20:59 +00:00
dennis zhuang 78a7b93292 fix(promql): align timestamp(), label_join and label_replace with Prometheus semantics (#9385)
* fix(promql): align timestamp() and label_join with Prometheus semantics

- timestamp() over a selector reports the selected sample's timestamp, without adding the offset.
- timestamp() over any other expression reports the evaluation time instead of the input value.
- label_join that overwrites an existing label rejects duplicate label sets at runtime.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(promql): align label_replace and label_join edge cases with Prometheus

- label_replace leaves non-matching series unchanged instead of copying the source value into the destination label.
- label_replace may overwrite an existing label; duplicate label sets are rejected at runtime like label_join.
- label_join with no source labels removes the destination label, and an empty source label name is rejected.
- Empty label values produced by these functions are NULL, the same as an absent label.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(promql): join absent labels as empty strings in label_join

concat_ws skips NULL arguments together with their separator, so a label that is absent (NULL) dropped its separator, e.g. label_join(label_join(vector(1), "a", "", "missing"), "b", "-", "a", "a") lost b="-". Read every source label as coalesce(label, '') like label_replace does.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(promql): skip rows without a sample in timestamp()

timestamp() replaced the input fields with the timestamp before the empty-value filter ran, so a row whose fields were all NULL got a timestamp, e.g. timestamp(-m) for a NULL sample. Filter such rows before the projection, for both the selector and the evaluation-time paths.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-30 01:42:11 +00:00
jeremyhi a99cb56d71 feat(datanode): wire the object store WAL into the datanode and standalone (#9386)
* feat(datanode): wire the object store WAL into the datanode and standalone

A datanode configured with `provider = "experimental_object_store"` now
builds an `ObjectStoreLogStore` instead of failing with the placeholder
"not supported yet" error. The builder resolves `storage_provider`
against the object store manager (empty selects the default store, an
unknown name is rejected as an invalid `storage_provider`), opens the
store under the node id and the standalone generation, stops it if a
later build step fails, and hands it to `Datanode`, whose shutdown stops
it after the region server. Datanode shutdown is now best-effort: every
step runs and the first error is returned. A datanode with a metasrv
client still rejects the provider.

The standalone bootstrap builds `WalProvider::ObjectStore` with the same
derived node prefix before the metasrv WAL conversion, so the WAL
options it allocates match the prefix the store runs under.

`log-store` exports `ObjectStoreLogStore` and its testing hooks. Mito's
test utilities can build an engine on the object store WAL over an
in-memory object store, and new engine tests cover create, reopen,
replay isolation between regions, batch open, drop and offline cleanup,
the latest entry id after replay, and the rejection of object store WAL
options on a Raft Engine log store.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(datanode): stop every region engine before the object store WAL

The region server returned at the first engine whose stop failed, so the
engines after it in map order, Mito included, were never stopped before
the datanode stopped the object store WAL. It now attempts every engine,
returns the first error and logs the later ones.

`ObjectStoreLogStore` is also exported from the `log-store` crate root,
and its callers import it from there.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-09-29 13:17:42 +00:00
Ning Sun 3131cdbcb6 fix: bound decompressed request size and close memory-admission gaps for compressed requests (#9264)
* fix: bound decompressed request size and close memory-admission gaps for compressed requests

The request-memory accounting only bounds and charges the encoded bytes on
the wire, but a compressed body can expand far beyond that during
decompression, so tiny requests could allocate disproportionate frontend
memory before any protobuf validation or quota charge.

- Handler-level decompression (Prometheus remote read/write v1+v2, Loki)
  now enforces a hard 512 MiB decoded-size cap, checked before any output
  buffer is allocated, and charges the decoded bytes to the shared
  ServerMemoryLimiter, holding the permits for the lifetime of the
  decompressed buffer.
- gRPC requests with transport compression reserve the configured
  max_recv_message_size before tonic decompresses, so the decoding phase
  is admitted against max_in_flight_write_bytes; the later per-message
  charge is skipped to avoid double accounting.
- The HTTP memory-limit middleware keeps its upfront Content-Length charge
  but now also accounts the bytes actually streamed beyond it, so chunked
  requests and understated Content-Length headers no longer bypass the
  aggregate quota.
- Routes that decompress request bodies via RequestDecompressionLayer
  (InfluxDB, OTLP, Loki, Splunk, Elasticsearch, pipelines, dashboards)
  now charge the decompressed bytes as handlers consume them: the global
  middleware marks Content-Encoding requests, and a route-local
  accounting layer inside the decompression layer charges the decoded
  stream. Plain requests are skipped to avoid double-counting.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: address review comments on memory admission guard lifetimes and 429 mapping

- Retain the gRPC pre-decode reservation for the whole request: the
  extensions holding the guard were dropped when the request was consumed
  (into_inner), releasing the reservation while per-message charges
  stayed skipped. Both the unary and streaming handlers now clone and
  hold the reservation marker for the duration of request handling.
- Retain the HTTP body permits across the handler: the AccountedBody
  wrapper and the request extensions are dropped once the extractors
  finish collecting the body, before the handler is done with the decoded
  data. Both accounting middlewares now keep their own accounting handle
  alive across next.run(req).await.
- Return a ChargedBuffer from the Loki snappy decompressor so the
  reservation outlives the decompressed bytes through the caller's
  protobuf decoding, matching the Prometheus path.
- Map mid-stream quota exhaustion to 429 instead of the generic body
  error (400): the accounting flags exhaustion and the middlewares
  rewrite the extractor rejection, so clients can distinguish
  backpressure from malformed input.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: resolve the in-flight body acquisition before trying a new one

A parked acquisition in AccountedBody::charge always belongs to the
currently buffered frame, so it must be resolved before any new
acquisition is tried. Letting the fast-path try_acquire succeed while a
waiter is parked left the waiter alive, and its late completion would
then be credited with a later frame's byte count, under-reserving memory
relative to what was marked charged.

Adds a regression test that reproduces the misattribution: with the
buggy ordering the test delivers a body the quota cannot cover; with the
fix the over-quota frame waits for its own acquisition and times out.

Signed-off-by: Ning Sun <sunning@greptime.com>

---------

Signed-off-by: Ning Sun <sunning@greptime.com>
2026-09-29 10:07:18 +00:00
dennis zhuang 218000e21b fix(promql): resolve dotted column names as unqualified columns (#9391)
* fix(promql): resolve dotted column names as unqualified columns

col(), From<&str>/From<String> for Column and string join keys go through
Column::from_qualified_name, which splits `service.name` into relation
`service` and column `name` and lowercases unquoted identifiers. Build
PromQL column references with Column::from_name / ident() instead.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(servers): resolve remote read matcher labels as unqualified columns

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(promql): cover same-name columns differing in case and without()

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-29 08:13:28 +00:00
discord9 abf1396c28 fix(client): yield Flight batches and affected rows without waiting for next message (#8918)
* fix(client): avoid Flight metrics lookahead stalls

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(client): extract trailing Flight metrics task

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(client): synchronize trailing metrics and bound cancellation

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-29 08:05:50 +00:00
Ning Sun a7f99b3425 ci: retry failed nightly release on following weekdays (#9394)
Previously the nightly release was scheduled only on Mondays, so a
failed run left users without nightly builds for a whole week.

Now the schedule triggers every weekday at 00:00 UTC, but the release
only proceeds when the latest published nightly release is older than
NIGHTLY_RELEASE_MAX_AGE_DAYS (5) days. The check runs in
allocate-runners before any EC2 runner is allocated, and gates all
schedule-driven jobs (including the Slack notification, so a skipped
nightly is silent). Tag pushes and manual dispatches are unaffected.

Signed-off-by: Ning Sun <sunning@greptime.com>
2026-09-29 07:48:29 +00:00
Lei, HUANG 5ccfcd4644 feat: support automatic column addition for Flight bulk inserts (#9285)
* fix: auto-add columns when initializing bulk insert streams

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: reject nested columns in bulk schema auto-add

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

* fix: reject unknown columns in non-empty bulk inserts

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>

---------

Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
2026-09-29 06:47:06 +00:00
DeviousCardiandClaude Opus 5.5 d6474b961c fix(promql): apply offset to subquery evaluation window (#9364)
* fix(promql): apply offset to subquery evaluation window

`prom_subquery_expr_to_plan` destructured `SubqueryExpr` without reading
`offset`, so `<subquery>[range:step] offset <d>` planned exactly the same
window as the un-offset form and silently returned data for the wrong time
range. The plain vector/matrix-selector paths already threaded the offset
through `selector_to_series_normalize_plan` and `RangeManipulate`.

Shift the inner evaluation window back by the offset and pass the offset to
the subquery's `RangeManipulate`, which maps the inner samples forward onto
the evaluation timeline before bucketing them into ranges. This matches
Prometheus, whose `evaluator.subqueryTimeRange` evaluates the inner
expression over `(start - offset - range, end - offset]` and whose
`evalSubquery` then hands the samples to the outer range-vector function as
a `MatrixSelector` that still carries the subquery offset. An offset on the
inner selector composes additively, as `subqueryTimes` documents.

`RangeManipulate`'s protobuf message has no offset field and recovers it on
decode from an immediately underlying `SeriesNormalize`. Since
`RangeManipulate` is commutative in `dist_plan` and can be pushed below a
`MergeScan`, insert that carrier node so the offset survives distributed
planning instead of decoding as zero.

Known divergence, unchanged by this commit: Prometheus anchors subquery step
points on absolute epoch multiples of the step, while GreptimeDB anchors them
on the evaluation start. The two agree whenever the offset is a multiple of
the subquery step; the added sqlness cases stay within that range.

Closes #9330

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Aarav <aaravsjadav@gmail.com>

* docs(promql): correct the subquery-offset rationale and pin the histogram shift

Follow-up to e132c267, which stated two things inaccurately and left one
path untested. No behaviour change beyond a conditional clone.

The `SeriesNormalize` carrier was described as existing "so the offset
survives distributed planning", implying standalone is unaffected. That is
wrong. `local_offset` is read on the substrait decode path
(`range_manipulate.rs`, `instant_manipulate.rs`), and GreptimeDB routes
PromQL plans through `MergeScan`/substrait in standalone too -- the
standalone results `promql/encode_substrait.result` and `precisions.result`
both show `MergeScan [is_placeholder=false, remote_input=[...]]`. Deleting
the carrier therefore empties subquery-with-offset results in standalone as
well, while `count_over_time_subquery_with_offset` keeps passing because it
asserts the pre-serialization plan. Say so at the node, so the next reader
does not remove it believing standalone is safe.

The carrier is also not semantically inert: `SeriesNormalizeStream::normalize`
shifts native histogram `start_timestamp` payloads by `offset`. On the
subquery path that runs on a computed inner result, on top of any offset the
inner selector already applied, and nothing pinned it. It is consistent --
`RangeManipulate` moves the millisecond time index forward by the same
`offset_ms` -- so the distance from a histogram's start timestamp to the
sample carrying it, which reset/rate detection reads, is invariant. Add
`subquery_offset_shifts_histogram_start_and_time_index_together`, which
executes the subquery's node stack over native histogram samples for a
zero and a non-zero inner offset. It cannot be a sqlness case: native
histograms are a struct column with no SQL type or literal
(`sql_data_type_to_concrete_data_type` rejects structs) and the sqlness
runner speaks only MySQL/Postgres, so such rows only arrive over
gRPC/remote-write v2.

The divergence note blamed sub-step offsets. Measured, that condition is too
narrow: an unaligned evaluation timestamp alone diverges, with no offset at
all -- `sum_over_time(fine[20s:10s])` at t=57 samples 47s and 57s here
against Prometheus's 40s and 50s. The real condition is that
`(start - offset)` is not a multiple of the subquery step, and it is
pre-existing: the committed `tql eval (359, 359, '1s')
sum_over_time(metric_total[60s:10s])` case is already an instance of it
(`359 - 60 + 10 = 309`). Reword the comment and the sqlness header
accordingly. A sub-step offset stays accepted: `foo[20s:10s] offset 5s` is
valid PromQL, erroring on it would be a regression and would not close the
gap anyway. This repo has no PromQL-compatibility documentation page, so
there is nowhere else to record it.

Also stop cloning the series key columns when there is no offset, and stop
the `offset 0s` sqlness comment implying its parser error is evidence about
this fix -- the same message comes from a `u32` overflow in
`promql-parser`'s shared duration check (`offset 9999999999d`).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Aarav <aaravsjadav@gmail.com>

* refactor(promql): trim subquery-offset comments and tests per review

Shorten the planner and sqlness comments to the semantic contract, build
`SeriesDivide` with a clone instead of an `Option`, and drop the
normalize.rs histogram test and the `offset 0s` parser case.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bvab9gpXYjKyLN8Mo6JWWL
Signed-off-by: Aarav <aaravsjadav@gmail.com>

* fix(promql): record the subquery offset as the range fold offset

The subquery's `RangeManipulate` now shifts the payload timestamps by the
offset, so `predict_linear` over an offset subquery must recover the
evaluation time with that offset instead of 0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bvab9gpXYjKyLN8Mo6JWWL
Signed-off-by: Aarav <aaravsjadav@gmail.com>

---------

Signed-off-by: Aarav <aaravsjadav@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-29 03:38:43 +00:00
LFCandgreptimedb-ci 886e02f0dc ci(query-regression): bump RUNNER_IMAGE_EPOCH to 7 (#9395)
ci(query-regression): bump RUNNER_IMAGE_EPOCH to 7 for image m-0xihbfzm5xpbxolybo6i

Signed-off-by: greptimedb-ci <greptimedb-ci@greptime.com>
Co-authored-by: greptimedb-ci <greptimedb-ci@greptime.com>
2026-09-29 03:32:36 +00:00
Weny Xu 133296f298 chore: bump version to 1.3.0-beta.1 (#9392)
Signed-off-by: WenyXu <wenymedia@gmail.com>
v1.3.0-beta.1
2026-09-29 03:22:10 +00:00
Ning Sun f8cc842468 chore: add mold back to nix flake (#9388) 2026-09-29 02:37:42 +00:00
Ning Sun 28f01d2ffe revert(ci): pin the query-regression runner toolchain to nightly-2026-03-21 (#9389)
* revert(ci): pin the query-regression runner toolchain to nightly-2026-03-21

The query-regression benchmark compiles both the candidate and the
BASE checkout (the previous nightly build). Base refs can predate the
stable-toolchain migration and still use #![feature] gates, so the
runner toolchain must stay a nightly that can build historical
revisions; deriving it from the workspace rust-toolchain.toml (now
stable 1.96.1) breaks base builds.

Revert the toolchain-toml coupling introduced in #9369 and keep the
parts that were correct:

- query-regression.yml: RUSTUP_TOOLCHAIN hard-pinned to
  nightly-2026-03-21 again (with a comment explaining why), Verify
  assertions back to the exact nightly versions, and the
  test-tooling pin-derivation machinery removed
- runner Dockerfile: ARG RUST_TOOLCHAIN=nightly-2026-03-21 + baked
  ENV restored; the COPY rust-toolchain.toml parsing removed
- build-ecs-image.py: the toml staging in the builder user-data
  removed; base-image auto-resolution kept but retargeted to Ubuntu
  26.04 to match the Verify tool pins (python3 3.14 etc.)
- the rebuild job keeps the epoch-bump lockstep, now as a PR from a
  timestamped ci/ branch mirroring update-dev-builder-version.sh
  (direct pushes to main fail GH006 under branch protection)

Complements #9376 (actions-runner bump), which requires a rebuilt
image with the non-deprecated runner.

Part of #9289.

Signed-off-by: Ning Sun <sunning@greptime.com>

* chore: skip qreg for rust-toolchain change

---------

Signed-off-by: Ning Sun <sunning@greptime.com>
2026-09-29 02:21:27 +00:00
dennis zhuang 0686f72dc6 chore: add cargo min-publish-age of 7 days (#9384)
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
v1.3.0-alpha.1-3c0e2a8d5-20260928-1790639248
2026-09-29 00:37:34 +00:00
Mohd Quamar Tyagi dd2c1d1aca fix(meta-srv): use NoTls for disabled and Unix socket Postgres KV backends (#9059)
* fix(meta-srv): use NoTls for disabled and Unix socket Postgres KV backends

Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com>

* fix(meta-srv): treat Postgres config as unix socket only when every host is a socket

tokio-postgres dials hostaddr over TCP even when host is a socket path,
and a mixed host list with Require/VerifyFull TLS would otherwise end up
sending plaintext over TCP. is_unix_socket_url now parses via
tokio_postgres::Config and requires no hostaddr and all-Unix hosts.
Adds regression cases for the libpq keyword form, the user@ percent
encoded socket URL, and the mixed host / hostaddr negatives.

Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com>

---------

Signed-off-by: Tyagiquamar <mohdquamartyagi@gmail.com>
2026-09-29 00:07:49 +00:00
jeremyhi 3c0e2a8d55 feat: export Metric snapshots with packed Parquet objects (#9382)
* feat: export Metric snapshots with packed Parquet objects

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: reject empty packed export time ranges

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test: cover cancellation during packed export I/O

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-09-28 15:22:02 +00:00
Ning Sun e2b7e5e0a8 chore: disable unstable rustfmt features (#9379)
* chore: disable unstable rustfmt features

* typo fix
2026-09-28 12:40:35 +00:00
dennis zhuang 079ffec116 chore: remove dead code left behind by removed features (#9377)
* chore: remove dead code left behind by removed features

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore: remove unused RouteInfoCorrupted error

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-28 12:35:47 +00:00
dennis zhuang 180221240a test: remove unused legacy compatibility test suites (#9378)
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-28 12:34:54 +00:00
dennis zhuang 408536af18 chore: remove leftover flow worker config docs and unused flow metrics (#9375)
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-28 12:34:32 +00:00
Ning Sun fedce5c5ec feat: support non-millisecond time index units in the logical batcher (#9346)
* feat: allow customized time index unit for metric engine table

* test: provide query tests

* refactor: revert unnecessary change

* refactor: share timestamp unit conversions in api helper

Address review feedback on the time index unit changeset:

- Add shared timestamp_unit/timestamp_datatype helpers to api::helper
  (the only crate that sees both proto ColumnDataType and TimeUnit due
  to layering; common-time and datatypes have no greptime-proto dep).
  This removes the ColumnDataType -> TimeUnit match duplicated between
  operator's insert path and the OTLP logs path.
- Collapse the two TimeUnit <-> ValueData matches in
  convert_timestamp_value_data by reusing api::helper::to_grpc_value
  for the construction side.
- Note that convert_rows_time_unit rewrites the schema before the
  values, so an overflow mid-batch leaves the request half-converted;
  harmless because the error aborts the whole insert request.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: align time units per destination table and floor remote-read timestamps

Address review feedback on PR #9236:

- Align each metric insert request to the unit of the table it actually
  targets: an existing logical table keeps its own unit (it may be bound
  to a different physical table than the one selected by the request),
  and only new tables use the selected physical table's unit. The
  previous blanket conversion rewrote valid millisecond samples to the
  selected physical table's unit and the engine rejected them.
  Regression test: writing an existing millisecond logical table and a
  new table in one request that selects a microsecond physical table.
- Remote read now floors narrowing timestamp conversions towards
  negative infinity (div_euclid), consistent with
  Timestamp::convert_to on the ingestion path; arrow's cast truncates
  towards zero and returned -1ms for a stored -1001us. Widening
  (second -> millisecond) keeps the exact arrow cast. Regression test:
  a negative, non-aligned timestamp round-trips as -2ms.

Signed-off-by: Ning Sun <sunning@greptime.com>

* perf: fold time unit alignment into existing table lookups

Address review feedback on PR #9236:

- The per-destination unit alignment no longer runs its own pass of
  table lookups: create_or_alter_tables_on_demand gains an
  align_time_index_unit parameter (metric engine path only) and
  converts each request inside the lookups it already performs —
  existing tables to their own unit, new tables to the selected
  physical table's. Default ingest paths now issue zero additional
  catalog lookups compared to main; the separate alignment pass remains
  only in the opt-in logical batcher pre-gate, next to the eligibility
  check that already looks up the same tables.
- convert_rows_time_unit indexes the time index position directly
  (validate_column_count_match guarantees row widths) instead of
  Optional get_mut; the gate-side alignment validates widths itself.

Signed-off-by: Ning Sun <sunning@greptime.com>

* perf: resolve the batcher time index guard once per write target

All batches of one remote write request share the same write target
(catalog, schema, physical table), so the batcher time index guard now
resolves each distinct target once instead of once per batch.

Signed-off-by: Ning Sun <sunning@greptime.com>

* feat: support non-millisecond time index units in the logical batcher

Make the logical table batcher's bulk encode path unit-aware so physical
metric tables with a non-millisecond time index (e.g. TIMESTAMP(6)) can
use logical batching instead of falling back to the ordinary insert path.

- rows_to_aligned_record_batch builds the time index column in the
  TARGET schema's unit, converting any timestamp encoding via
  Timestamp::convert_to (flooring on narrowing, consistent with the
  ordinary insert path).
- New tables created by the batcher use the selected physical table's
  time index unit (resolved once per submit; a missing physical table
  keeps the millisecond auto-create default).
- columns_taxonomy and the can_batch_metric_rows schema whitelist accept
  any timestamp unit; the prometheus remote write v1/v2 batcher gates
  and the OTLP pre-gate alignment are removed together with
  Inserter::align_metric_row_inserts_time_unit, as the batcher now
  converts internally.

Closes #9342

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: address review comments

* fix: address review issue

* refactor: drop the OTLP pre-gate unit alignment made redundant by the bulk path

The main merge of #9236 (squash) resurrected the OTLP pre-gate
alignment and Inserter::align_metric_row_inserts_time_unit, which this
branch had removed. Drop them again:

- The pre-gate existed because the #9236-era bulk eligibility gate only
  accepted millisecond schemas, so nanosecond-encoded OTLP requests had
  to be converted before the check. This branch makes the bulk path
  unit-aware (the gate accepts all time index units and batch alignment
  converts each request to its destination's unit), so the pre-gate is
  redundant and only added N+1 catalog lookups per batched request —
  the very lookup-count overhead raised in the #9236 review.
- The per-destination unit semantics it implemented remain enforced in
  the two paths that need them: the ordinary insert path
  (create_or_alter_tables_on_demand converts inside its existing table
  lookups) and the batched path (batch alignment resolves each
  destination schema and converts to it).

test_otlp_logical_batcher_alignment (the test the pre-gate originally
fixed) and the mixed-physical-table regression both pass without it.

Signed-off-by: Ning Sun <sunning@greptime.com>

* test: cover OTLP batcher cross-physical fallback and nanosecond physical

Extend the logical batcher integration coverage for the cases
previously guarded by the removed OTLP pre-gate alignment:

- test_otlp_logical_batcher_fallback_for_cross_physical_destination:
  with the batcher enabled, an OTLP request targeting an existing
  logical table bound to another physical table must NOT enter the
  batcher (the bulk eligibility check rejects the destination binding)
  and the ordinary insert path must convert it to the destination's
  unit (60s -> 60_000_000us).
- test_otlp_logical_batcher_non_millisecond_physical_table now covers
  both microsecond and nanosecond physical tables (parameterized),
  asserting batcher submissions and unit-precise stored values.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: validate per-batch physical bindings and avoid intermediate timestamp buffers

Address review feedback on PR #9346:

- accepts_bulk_destinations dedupes on (schema, table, selected physical)
  instead of (schema, table): one request can select different physical
  tables per batch (per-series x_greptime_physical_table labels), and the
  old key let a second selection skip validation and flush rows through
  the wrong physical's regions. Missing tables additionally reject
  conflicting physical selections within the same request. Regression
  test covers an existing destination, a missing destination, and a
  consistent selection (which must still batch).
- The timestamp column builder appends each value directly into the
  target-unit Arrow builder; values already in the target unit (the
  unchanged millisecond fast path) are appended without conversion, so
  the default millisecond physical pays no Timestamp construction or
  intermediate Vec allocation.
- The non-millisecond batching tests assert the submit_build_and_align
  counter, which increments on every batcher submission in both
  acknowledgement modes, so a silent fallback to ordinary insertion
  fails the tests instead of passing on stored values alone.

Signed-off-by: Ning Sun <sunning@greptime.com>

---------

Signed-off-by: Ning Sun <sunning@greptime.com>
2026-09-28 12:27:36 +00:00
discord9 28e415a103 feat(query): implement PromQL @ modifier on vector and matrix selectors (#9224)
* feat(query): implement PromQL @ modifier on vector and matrix selectors

The planner previously dropped the `at` field of selectors (`at: _`),
so `some_metric @ 300` silently returned step-following values instead
of the samples anchored at the fixed timestamp.

Anchoring follows Prometheus's `setOffsetForAtModifier` + `refetch`
semantics: resolve the anchor (`@ <ts>`, `@ start()`, `@ end()`), apply
`offset` to the anchor, rewrite the selector offset to
`eval_start - anchor`, scan only the anchored window, then report the
same window at every evaluation step via a grid-wide-replay
`InstantManipulate`. `@ start()` / `@ end()` resolve against the
statement's evaluation range (`stmt_start` / `stmt_end`), which subquery
planning does not rewrite.

Timestamps before the Unix epoch are accepted; unrepresentable ones are
rejected with `AtModifierTimestampOutOfRange` instead of wrapping.
Selectors without `@` are planned exactly as before.

Report: .e-agent/greptimedb_promql_compatibility_report_2026-09-16.md P0-2
Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(promql): center predict_linear on the evaluation timestamp

predict_linear_impl used the window's last sample time as the evaluation
timestamp, and the UDF took only (ts_range, value_range, t) with no channel
for the current instant. After a range selector is folded for @ (or with an
offset) the same window is replayed at every step, so the prediction stayed
constant at the anchor's answer; even a plain window ended before the step
produced a stale value.

Give prom_predict_linear a 4th argument carrying the step's evaluation
instant (ms timestamp), derived in the planner from the row's time index
plus the offset the window was folded with (at_offset for @, offset_ms
otherwise, recorded on PromPlannerContext). The regression is centered on
that instant, matching Prometheus' use of enh.Ts. The mixed
float/native-histogram path forwards the same argument.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(query): tighten @ modifier planner and predict_linear

- range_fold_offset is always a concrete millisecond offset, not an
  Option; the single-step replay helper no longer wraps an infallible
  plan in Result.
- predict_linear's eval timestamp is always cast to Timestamp(ms) at the
  call site, so the UDF drops its dead Int64 branch and the extra
  func_name argument.
- Drop two redundant comparison cases from the at_modifier sqlness test
  and fix the lookback window comment to the half-open (0s, 300s].

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* test(query): accept parse-time rejection of @ on Windows

`@ 1e16` is 10^19 milliseconds, beyond i64::MAX. A Unix SystemTime
holds it and the planner rejects the anchor it cannot represent, but a
Windows SystemTime tops out near 1.8e12 seconds, so the parser's
checked_add fails first and the same literal is rejected while parsing.
Accept either rejection path so the test passes on both platforms.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(query): avoid replaying label_join over rewritten series keys

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(query): narrow anchored range call promotion

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(query): address @ modifier review feedback

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(query): satisfy super import format check

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* docs(query): address at modifier review follow-ups

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-28 09:38:12 +00:00
jeremyhi 03823a9a01 feat(log-store): add the enqueued acknowledgement mode to the object store WAL (#9358)
* feat(log-store): add the enqueued acknowledgement mode to the object store WAL

Add `ack_mode` (`durable` by default, or `enqueued`) and the backlog
thresholds `max_unpersisted_bytes` and `max_unpersisted_age` to the
object store WAL config, validated by the datanode and the store.

In the `enqueued` mode `append_batch` returns on admission with the
entry ids assigned and the object is created in the background. At a
backlog threshold the next append is held back until an upload
completes. A transient create failure is repeated under the same
sequence with the same bytes; any conflicting object poisons the
store. `stop` uploads the backlog, or returns the error that dropped
it once stop began. `obsolete` clamps the watermark to the durable
entry id, and an id the store handed out needs no sequence floor.

Add `LogStore::wait_durable` with a default that returns at once. The
object store WAL answers it once the region is durable and indexed
through the entry id, and fails it after a backlog was lost.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): poison an enqueued store on permanent create failures

Repeat a failed create in the enqueued mode only when the storage error
is retryable; any other storage error poisons the store. A conflicting
object poisons an enqueued store without reading its epoch, so a failed
header read cannot turn the conflict into a retry.

A durability wait for an id above the highest id the store handed out
now waits for the handed-out ids of the region below it instead of
returning at once.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): poison on permanent create failures after stop begins

A create that fails with a storage error that is not retryable poisons
an enqueued store even after stop began; only a transient failure drops
the backlog without poisoning. Durability waiters whose callers stopped
waiting are pruned before a new waiter is queued. The backlog age test
no longer depends on a follow-up append finishing within the threshold.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): answer a durability wait once no earlier entry is pending

A durability wait now returns once the region holds no entry at or
below the target that is handed out but not durable, instead of waiting
for the largest id handed out to the region. A later object of the
region that is still being created no longer holds back a wait whose
target it does not cover.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test(log-store): order the pending durability wait check after the actor

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* docs(store-api): state that the default wait_durable keeps each store's guarantee

The default `LogStore::wait_durable` returns at once, which keeps each
log store's own acknowledgement guarantee; Raft Engine with
`sync_write = false` acknowledges before its periodic sync, so the
documentation no longer claims that every entry id a caller holds is
durable. The object store WAL configuration test now also serializes
the new options and reads them back.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* docs(log-store): limit the acknowledgement guarantees to the durable mode

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): repeat enqueued creates that a retry layer marks persistent

An object store wrapped in the OpenDAL retry layer reports a temporary
error that outlasted its retries as persistent rather than temporary.
The enqueued mode now repeats a create after any storage error that is
not permanent, so a transient outage behind the retry layer no longer
poisons the store and drops the acknowledged backlog; after stop began
such a failure still drops the backlog without poisoning.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test(log-store): cover a persistent create failure after stop begins

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-09-28 09:17:45 +00:00
LFCandgreptimedb-ci a310ca2bcf ci(query-regression): bump RUNNER_IMAGE_EPOCH to 6 (#9381)
ci(query-regression): bump RUNNER_IMAGE_EPOCH to 6 for image m-0xidkbbavm3suxqd0y1m

Co-authored-by: greptimedb-ci <greptimedb-ci@users.noreply.github.com>
2026-09-28 09:16:23 +00:00
75bd8e9ce6 feat: add HDFS object storage backend (#8701)
* feat: add HDFS object storage backend

Signed-off-by: Minghan2005 <cambrianocean@gmail.com>

* fix: make HDFS storage operations durable

Gate the native HDFS backend behind an explicit feature. Publish writes through same-directory temporary files and atomic HDFS Rename2 replacement, and provide streaming copy fallback for COPY_REGION. Add regression coverage for interrupted writes and the region-copy path.

Signed-off-by: Minghan2005 <cambrianocean@gmail.com>

* ci: run HDFS object store tests

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* docs: note HDFS temporary file cleanup follow-up

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* feat: enable HDFS object storage by default

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* docs: remove redundant HDFS build feature notes

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: Minghan2005 <cambrianocean@gmail.com>
Signed-off-by: jeremyhi <fengjiachun@gmail.com>
Co-authored-by: Minghan2005 <cambrianocean@gmail.com>
Co-authored-by: jeremyhi <fengjiachun@gmail.com>
2026-09-28 08:39:48 +00:00
dennis zhuang ffdd6d09a6 perf(index): cut allocations when building bloom and inverted indexes (#9359)
* perf(index): build bloom filters from element hashes without per-token allocation

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* perf(index): speed up inverted index building with hashed buffers and sync pushes

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore(index): use BuildHasher::hash_one for element hashes

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore(index): require callers to act on spill requests

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(index): insert bloom hashes one by one to keep segment set capacity bounded

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(index): add index build and bloom search benchmarks

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(index): keep applier setup out of the bloom search benchmark

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(index): skip empty spills and keep the old inverted sort memory estimate

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(index): stream fulltext token hashes and test spill dispatch

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* docs(index): note non-ASCII tokens still allocate in analyze_text_hashes

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* docs(index): narrow the allocation note to case-insensitive non-ASCII tokens

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-28 08:22:42 +00:00
dennis zhuang bc5ad4b416 perf(mito2): batch index page loads and share cached bloom metadata (#9360)
* perf(mito2): load missing index pages of all ranges in one read

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* refactor(mito2): drop unused IndexCacheMetrics::merge

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* perf(mito2): share cached bloom filter metadata instead of cloning it

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(mito2): only merge adjacent ranges when batch reading index pages

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* fix(mito2): skip empty ranges when batch loading index pages

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* perf(mito2): fetch index ranges concurrently

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(index): return Arc<BloomFilterMeta> from RecordingReader

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-28 07:57:04 +00:00
Weny Xu 1d365adb3d fix(ci): update the shared Actions runner to v2.337.0 (#9376)
Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-09-28 07:54:37 +00:00
Ning Sun 55b7e08a1a feat: allow customized time index unit for metric engine table (#9236)
* feat: allow customized time index unit for metric engine table

* test: provide query tests

* refactor: revert unnecessary change

* refactor: share timestamp unit conversions in api helper

Address review feedback on the time index unit changeset:

- Add shared timestamp_unit/timestamp_datatype helpers to api::helper
  (the only crate that sees both proto ColumnDataType and TimeUnit due
  to layering; common-time and datatypes have no greptime-proto dep).
  This removes the ColumnDataType -> TimeUnit match duplicated between
  operator's insert path and the OTLP logs path.
- Collapse the two TimeUnit <-> ValueData matches in
  convert_timestamp_value_data by reusing api::helper::to_grpc_value
  for the construction side.
- Note that convert_rows_time_unit rewrites the schema before the
  values, so an overflow mid-batch leaves the request half-converted;
  harmless because the error aborts the whole insert request.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: align time units per destination table and floor remote-read timestamps

Address review feedback on PR #9236:

- Align each metric insert request to the unit of the table it actually
  targets: an existing logical table keeps its own unit (it may be bound
  to a different physical table than the one selected by the request),
  and only new tables use the selected physical table's unit. The
  previous blanket conversion rewrote valid millisecond samples to the
  selected physical table's unit and the engine rejected them.
  Regression test: writing an existing millisecond logical table and a
  new table in one request that selects a microsecond physical table.
- Remote read now floors narrowing timestamp conversions towards
  negative infinity (div_euclid), consistent with
  Timestamp::convert_to on the ingestion path; arrow's cast truncates
  towards zero and returned -1ms for a stored -1001us. Widening
  (second -> millisecond) keeps the exact arrow cast. Regression test:
  a negative, non-aligned timestamp round-trips as -2ms.

Signed-off-by: Ning Sun <sunning@greptime.com>

* perf: fold time unit alignment into existing table lookups

Address review feedback on PR #9236:

- The per-destination unit alignment no longer runs its own pass of
  table lookups: create_or_alter_tables_on_demand gains an
  align_time_index_unit parameter (metric engine path only) and
  converts each request inside the lookups it already performs —
  existing tables to their own unit, new tables to the selected
  physical table's. Default ingest paths now issue zero additional
  catalog lookups compared to main; the separate alignment pass remains
  only in the opt-in logical batcher pre-gate, next to the eligibility
  check that already looks up the same tables.
- convert_rows_time_unit indexes the time index position directly
  (validate_column_count_match guarantees row widths) instead of
  Optional get_mut; the gate-side alignment validates widths itself.

Signed-off-by: Ning Sun <sunning@greptime.com>

* perf: resolve the batcher time index guard once per write target

All batches of one remote write request share the same write target
(catalog, schema, physical table), so the batcher time index guard now
resolves each distinct target once instead of once per batch.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix: address review comments

* fix: address review issue

---------

Signed-off-by: Ning Sun <sunning@greptime.com>
2026-09-28 07:15:36 +00:00
dennis zhuang 0a03e64d8c perf(index): batch bloom filter searches across row groups (#9361)
* perf(index): search bloom filters of all row groups in one read

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* perf(index): bound the filter bytes of each batched bloom search

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* perf(index): count batch filter bytes exactly as the search reads them

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore(index): allow single-range groups in the budget test

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* test(index): cover filters shared across segments in the budget test

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-28 06:22:15 +00:00
Ning Sun 43eaea7a9a fix(ci): teach check-builder-rust-version.sh to handle stable channels (#9369)
* fix(ci): teach check-builder-rust-version.sh to handle stable channels

The script extracted a YYYY-MM-DD date from rust-toolchain.toml to
compare against the rustc build date inside the dev-builder image —
a nightly-era design. With channel = "1.96.1" there is no date in
the file, so every release build failed with 'Error: No rust toolchain
version found in rust-toolchain.toml'.

Extract the channel token instead and branch on it:
- stable channel (X.Y[.Z]): require the builder image's rustc to
  exactly match the pinned version
- nightly-YYYY-MM-DD: keep the legacy date-difference check

Verified against a mocked docker for all four paths (stable match /
mismatch, nightly fresh / stale).

Part of #9289.

Signed-off-by: Ning Sun <sunning@greptime.com>

* chore(toolchain): finish stable-migration cleanup in docs and query-regression pin

- README/AGENTS: the toolchain is now stable Rust pinned by
  rust-toolchain.toml, not nightly
- query-regression: align the benchmark toolchain pin with the
  workspace (nightly-2026-03-21 = 1.96.0-nightly -> stable 1.96.1),
  including the exact-version assertions (cargo 356927216, rustc
  31fca3adb, both 2026-06-26) and the runner image default

The query-regression runner image must be rebuilt and
QUERY_REGRESSION_ECS_IMAGE_ID bumped together with these pins
(per .github/runner-scale-sets/query-regression/README.md) before
the next regression run.

Part of #9289.

Signed-off-by: Ning Sun <sunning@greptime.com>

* ci(query-regression): derive the Rust toolchain pin from rust-toolchain.toml

Replace the hard-coded RUSTUP_TOOLCHAIN value and the hard-coded
version strings in the runner Verify assertions with a pin resolved
from rust-toolchain.toml:

- the always-running test-tooling job exports the channel parsed from
  rust-toolchain.toml as a job output
- query-regression sets RUSTUP_TOOLCHAIN from that output
- the Verify step escapes the pin into the cargo/rustc/active-toolchain
  regexes at runtime; the exact commit hash and date are asserted
  generically since a stable version identifies the release

Removing the redundant require_eq (workflow yaml vs runner env) since
both now flow from the single source of truth. When rust-toolchain.toml
is bumped, the run fails with a clear signal until the runner image is
rebuilt with the new toolchain, keeping the existing lockstep contract.

Part of #9289.

Signed-off-by: Ning Sun <sunning@greptime.com>

* ci(query-regression): derive the runner image toolchain from rust-toolchain.toml

Remove the hard-coded 'ARG RUST_TOOLCHAIN=1.96.1' from the
query-regression runner Dockerfile. The pin is now parsed from a
COPY'd rust-toolchain.toml at build time (the bootstrap script builds
with the repo root as context, so the file is in the build context):

- rustup-init installs the parsed channel as the default toolchain
- the baked ENV RUSTUP_TOOLCHAIN is dropped: the rustup default makes
  bare cargo/rustc resolve correctly without it, and the workflow
  supplies RUSTUP_TOOLCHAIN explicitly at run time
- the build-time self-verification asserts the active toolchain
  against the same parsed pin

With this, rust-toolchain.toml is the single source of truth for the
benchmark toolchain end to end: the image bakes whatever the toml says
at build time and the workflow asserts against the toml at run time.
A toolchain bump now only requires rebuilding the image.

The changed mechanics were verified natively with the real rustup-init
1.29.0 and the real 1.96.1 toolchain (registry pulls are unavailable
in this sandbox): parsing, default-toolchain installation without
RUSTUP_TOOLCHAIN env, bare cargo/rustc resolution, and the
active-toolchain assertion for both '(default)' and '(overridden)'
output forms.

Part of #9289.

Signed-off-by: Ning Sun <sunning@greptime.com>

* ci(query-regression): rebuild the runner image automatically on toolchain changes

Mirror the dev-builder automation for the query-regression ECS runner
image: a new rebuild-query-regression-runner-image.yaml workflow runs
whenever rust-toolchain.toml or the query-regression runner directory
changes on main (or via manual dispatch). It drives the existing
build-ecs-image.py ops tool, then completes the documented lockstep
updates in order: bump RUNNER_IMAGE_EPOCH in query-regression.yml and
push the commit to main, and only then point the
QUERY_REGRESSION_ECS_IMAGE_ID repo variable at the new image, so the
next regression run picks up image and epoch together.

Also fix build-ecs-image.py to stage rust-toolchain.toml into the
temporary docker build context: the AMI path builds the embedded
Dockerfile from an empty /tmp/image-context, which would break on the
Dockerfile's COPY of rust-toolchain.toml introduced earlier. The
user-data now base64-stages the toml next to the Dockerfile before
docker build.

Verified: py_compile, render_user_data round-trip (mkdir -> stage ->
docker build ordering), and the RUNNER_IMAGE_EPOCH bump sed against
the real workflow file. Requires a new ALIYUN_ECS_BASE_IMAGE_ID repo
variable (Ubuntu 24.04 public image id in the region); all other
secrets/vars are shared with the provisioning job.

Part of #9289.

Signed-off-by: Ning Sun <sunning@greptime.com>

* ci: fold the query-regression runner rebuild into release-dev-builder-images.yaml

Merge the standalone rebuild workflow into the existing builder-image
release workflow, as one entry point for all builder artifacts:

- push paths extended with .github/runner-scale-sets/query-regression/**
- new 'release_query_regression_runner_image' dispatch input
- a 'changes' job diffs the pushed range (github.event.before..sha,
  with an everything-changed fallback for dispatch or unknown bases)
  so each expensive rebuild only fires for its own paths:
  rust-toolchain.toml gates both, docker/dev-builder/** gates the
  dev-builder images, the query-regression runner directory gates the
  ECS image rebuild
- the rebuild job itself is unchanged from the standalone workflow
  (build-ecs-image.py, then RUNNER_IMAGE_EPOCH commit to main, then
  the QUERY_REGRESSION_ECS_IMAGE_ID variable update)

The dev-builder jobs, their ECR/CN/tag-update dependents, and the
runner rebuild now share one workflow; the changes filter preserves
the previous on-push behavior for the dev-builder images while the
runner rebuild keeps its own trigger. The epoch-bump commit only
touches query-regression.yml, which is outside the trigger paths, so
no re-trigger loop.

Part of #9289.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix(ci): auto-resolve the ECS base image for the runner rebuild

The automated rebuild failed with 'Missing required configuration:
--base-image-id' because the ALIYUN_ECS_BASE_IMAGE_ID repo variable
does not exist yet (it was flagged as a one-time setup item).

Remove the setup dependency instead: build-ecs-image.py now defaults
--base-image-id to the latest public Ubuntu 24.04 x86_64 system image
in the region (DescribeImages with image_owner_alias=system), so no
manual variable is required. The runner Dockerfile pins every tool
version itself, so base-image drift is low-risk; --base-image-id or
the ALIYUN_ECS_BASE_IMAGE_ID variable still pin a specific base image
deterministically, and the workflow only passes the flag when the
variable is set.

Part of #9289.

Signed-off-by: Ning Sun <sunning@greptime.com>

* docs(query-regression): clarify what an image rebuild requires

A routine runner-image rebuild needs no manual file updates: the
rebuild job updates QUERY_REGRESSION_ECS_IMAGE_ID and
RUNNER_IMAGE_EPOCH; the toolchain derives from rust-toolchain.toml;
uv, sccache, otelgen, rustup, and the runner base are pinned by
digest/sha/commit in the Dockerfile. Only an apt package revision
bump (mold, protoc, python3) between rebuilds requires bumping the
corresponding Verify pins, and that failure is loud with the observed
version.

Part of #9289.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix(ci): correct SDK field names in the base-image resolver

DescribeImagesRequest takes 'ostype' (not 'os_type') and the image
items expose 'osname'/'osname_en' (not 'os_name') in the pinned
alibabacloud_ecs20140526 SDK range, so the auto-resolution added in
5a9fd2c769 crashed with a TypeError before describing anything.

Fix the request fields, move the architecture filter server-side, and
paginate (page_size=100 until a short page) instead of relying on a
single default-sized response. Match Ubuntu 24.04 on the localized
osname or the English osname_en.

Verified against the real SDK models (uv run --with
'alibabacloud_ecs20140526>=4.1.0,<6'): a two-page fake client picks
the newest Ubuntu 24.04 via osname_en and rejects 22.04/Windows
decoys.

Part of #9289.

Signed-off-by: Ning Sun <sunning@greptime.com>

* fix(ci): correct the repo-root path in build-ecs-image.py

ASSETS_DIR.parent.parent lands on runner-scale-sets, not the repo
root -- the toml lookup failed with FileNotFoundError. The root is
four levels above ecs-image; express it as an explicit REPO_ROOT
constant (ASSETS_DIR.parents[3]).

Verified every path main() reads against the real checkout layout
(Dockerfile, rust-toolchain.toml, start-runner.sh, the systemd unit,
plus REPO_ROOT sanity against Cargo.toml/.git), re-checked the
user-data toml staging round-trip, and re-ran the base-image
resolver regression test.

Part of #9289.

Signed-off-by: Ning Sun <sunning@greptime.com>

---------

Signed-off-by: Ning Sun <sunning@greptime.com>
2026-09-28 06:08:14 +00:00
discord9 f956f4da30 refactor(promql): split planner into focused modules (#9354)
* refactor(promql): move planner tests into separate module

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(promql): move function-specific planner methods into module

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(promql): move OR planner method into module

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(promql): move binary island planner into module

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(promql): keep binary result labels in planner root

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(promql): limit planner helper visibility to parent module

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(promql): group set operator planning and localize imports

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(promql): use crate-rooted planner imports

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-28 05:30:33 +00:00
Weny Xu 825402dc70 ci: add manually triggered tracesbench workflow (#9372)
Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-09-28 05:07:16 +00:00
discord9 678aa81cae fix(client): complete transport lane isolation (#9030)
* fix(client): complete transport lane isolation

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor(client): deprecate legacy single-manager constructors

Per review: mark the legacy single-manager constructors and helper as
deprecated so callers move to the isolated query/control manager pair.
Tests intentionally exercising the legacy path are annotated with
`#[allow(deprecated)]`.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* fix(client): update deprecated constructor callers for CI

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-28 04:39:20 +00:00
DeviousCardiandClaude Opus 5.5 d5a6a9ac05 fix(auth): follow symlink chains in watch_file_user_provider (#9365)
* fix(auth): follow symlink chains in watch_file_user_provider

The shared file watcher only watched the immediate parent directory of the
configured path, so swapping a directory symlink elsewhere in the chain
(e.g. `ln -sfn secrets.2 secrets`, as sops-nix does) never produced an
event and the new users file was not loaded until restart.

The watcher now resolves each path component by component and also watches
the directory holding every symlink it meets, plus the directory of the
final file. After each relevant event the chain is re-resolved and the
watches are re-armed, so later in-place edits of the new target are seen.
Only the final file's directory is required at startup; a link directory
that cannot be watched (e.g. mode 0711) is logged and retried later.

Events are now filtered to the chain's links and final files, so that
watching e.g. /tmp does not reload on unrelated files. Names are compared
case-insensitively on macOS and Windows. TLS cert reload uses the same
helper.

Close #9310

Signed-off-by: Aarav <aaravsjadav@gmail.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bvab9gpXYjKyLN8Mo6JWWL

* fix(common-config): warn when the file watcher watches the root dir

On macOS /tmp, /var and /etc are symlinks under /, so following a symlink
chain there adds / to the watched directories, and the FSEvents backend
then sees every event on the volume. Log a warning when that happens.
Also shorten the comment on skipped link directories.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bvab9gpXYjKyLN8Mo6JWWL
Signed-off-by: Aarav <aaravsjadav@gmail.com>

---------

Signed-off-by: Aarav <aaravsjadav@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-28 04:33:46 +00:00
discord9 4ffd93d8dc fix(query): preserve global limits with DataFusion optimizer fixes (#9071)
* fix(query): preserve global fetch through physical-optimizer repartitioning

Aggregate soft-limit pushdown inserts a global CoalescePartitionsExec(fetch)
below an AggregateExec whose input is HashPartitioned/KeyPartitioned. The
DataFusion physical optimizer's EnforceDistribution/EnforceSorting passes then
dropped or rewrote that fetched coalesce, losing the global limit. Root-caused
and fixed in the DataFusion fork (GreptimeTeam/datafusion PR #35); this repo
pins to that fix commit and adds regression coverage.

- pin datafusion to GreptimeTeam/datafusion 48510c7f5 (fetch preservation fix)
- global_limit.rs: adapt to new DF API (input_distribution_requirements,
  child_distribution, replace_children+Recompute); accept HashPartitioned and
  KeyPartitioned as partitioning to restore; add regression unit tests
- sqlness: extend standalone common ssts and refresh distributed ssts_limit
  result for the preserved fetch plan

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* refactor: drop dead HashPartitioned compat arms and test wrapper

- remove HashPartitioned compatibility matches (pinned DF only emits
  KeyPartitioned); drop their #[expect(deprecated)] attributes
- remove single-use agg_with_limit helper whose seed limit is always
  overwritten by the soft-limit transform
- restore Cargo.lock dependency edges unrelated to the DataFusion pin
  bump (cargo update --precise had downgraded 26 unrelated packages)

Addresses oracle ablation review.

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

* chore(deps): pin DataFusion to merged fetch preservation fix

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>

---------

Signed-off-by: discord9 <55937128+discord9@users.noreply.github.com>
2026-09-28 03:35:36 +00:00
jeremyhi bf67246645 perf: bound concurrent Metric export writers (#9296)
* fix: preserve objects after conditional Metric export collisions

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* perf: share Metric export writer and payload budgets

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test: cover bounded Metric export conversion and mixed roundtrips

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: preserve export output after ambiguous close

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test: align HTTP export timeout with managed cancellation

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test: trim redundant Metric export roundtrips

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: bound pending Metric export writer handles

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: protect export files on failed writes

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: preserve ordinary export cleanup and large rows

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: report Metric export collisions as invalid arguments

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: clean up failed conditional export closes

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: support view arrays in managed ordinary export

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test: expect successful cleanup after failed close

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: clean up unsynced overwrite after close failure

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: scope export JSON expansion estimates to JSON columns

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix: rechunk ordinary export batches within the write budget

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-09-28 03:10:21 +00:00
dennis zhuang 411c0f173e fix(mito2): stop parallel flat scan tasks once the receiver is dropped (#9368)
* fix(mito2): stop parallel flat scan tasks once the receiver is dropped

spawn_flat_scan_task ignored send errors, so after a query was cancelled the
detached task kept reading its source to EOF. Break out of the loop when the
receiver is gone.

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore(mito2): log when a parallel scan task stops on a dropped receiver

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

* chore(mito2): include the send error in the parallel scan stop log

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>

---------

Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-28 02:50:35 +00:00
dennis zhuang 2b596fd52d ci: deploy MinIO chart with Silo images (#9362)
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
2026-09-27 02:36:11 +00:00
Weny Xu 5ffd01a70a fix: preserve primary key order when syncing columns (#9189)
* fix: preserve primary key order when syncing columns

Signed-off-by: WenyXu <wenymedia@gmail.com>

* fix: reject invalid SyncColumns metadata

Signed-off-by: WenyXu <wenymedia@gmail.com>

---------

Signed-off-by: WenyXu <wenymedia@gmail.com>
2026-09-27 02:15:47 +00:00
jeremyhi 194bc2fb3c feat(log-store): add the object store WAL durable write path (#9320)
* feat(log-store): add the object store WAL durable write path

ObjectStoreLogStore can now write. append_batch admits entries into one
open batch and assigns object-sequence-major entry ids at admission. The
batch is sealed by size, by the flush interval, or before a region would
run past the position range, and sealed batches are uploaded with at most
four conditional creates in flight, started in sequence order. Created
objects are indexed in sequence order, and an append is acknowledged only
once its object is durable and indexed.

A transient create failure rolls back the failed batch and every later
batch unless a later object is already durable, in which case the store
poisons itself with a history-gap error. A conflicting object, an
encoding or catalog error, or taking the last representable sequence
poisons the store.

obsolete now goes through the actor and raises the sequence floor
together with the watermark, refusing with a retryable error while the
next sequence is not settled. stop drops the open batch and the batches
whose create has not started, and lets creates in flight finish.

Only the durable acknowledgement mode exists. A testing feature exposes
hooks to wait for admissions, seal the open batch, and hold or fail
creates.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test(log-store): synchronize object store WAL write tests with the actor

Tests that assert nothing happened round-trip a command through the actor
instead of yielding the test task, the conflict test waits for the object
it expects, and the obsolete-behind-stop test holds the command channel
itself so the unanswered command is deterministic.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): bound admitted WAL appends and never reuse a failed sequence

A create that reports an error may still have written its object, so the
sequences of failed batches are no longer handed out again: the next
batch keeps the sequence after the last sealed one and retries are
assigned new ids. A retry batched differently can no longer conflict
with that object. As the next sequence never moves back, the sequence
floor of obsolete only waits for an open batch that has handed out ids.

Appends now arrive on their own bounded channel, which the actor stops
reading while MAX_SEALED_BATCHES batches wait to become durable, so a
stalled object store holds callers back instead of growing the backlog,
while stop and obsolete are still handled.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): reserve room for two seals before admitting a WAL append

One admission can seal the open batch before an append that would exhaust
its positions and then the append's own batch, so the actor takes an
append only while two more sealed batches fit under MAX_SEALED_BATCHES.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* feat(log-store): chain object store WAL objects and recover the latest chain

A conditional create can fail with an unknown outcome while its object is
stored, or still lands later, and Mito reuses the row sequences of a
failed append. Replaying such an object next to a later acknowledged one
can let the unacknowledged rows win after a restart.

Format version 2 gives every object header its writer's epoch, a link
to the object it extends (sequence and writer instance) and a header
CRC32, and allows objects without segments. Recovery replays only the
chain ending at the complete object with the largest epoch and sequence:
a link holds when its predecessor is present with the recorded writer
instance, or is missing below every present object. Objects off the
chain are orphans that are never replayed but keep their sequences.

Each open writes an empty object that starts an epoch above every
present object, linked to the recovered tip, before it accepts writes,
so a late object of an earlier instance never ends the chain. A start
object that meets an object of an earlier epoch moves to the next
sequence; one of an equal or later epoch fails the open.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): break the WAL chain at every missing predecessor

Accepting a missing predecessor below every present object lets a late
object that lands below the chain change which links hold. Nothing
collects objects yet, so a missing predecessor now always breaks the
link, and recovery fails when objects are present but none completes a
chain.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): keep the object store WAL format at version 1

The object store WAL has not been enabled anywhere, so no object in the
previous layout exists and the chained header can stay version 1.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* refactor(log-store): link object store WAL objects by epoch

Only one store instance writes under an epoch, since every open starts an
epoch above every present object and a start object that meets the same
or a later epoch fails the open. The epoch therefore identifies the
instance, and the random writer instance id is dropped from the header.
A link now records the sequence and the epoch of the object it extends,
and holds when the predecessor carries that epoch. The header shrinks to
46 bytes, and the store logs its epoch when it opens.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): fail an open whose start object is already present

Without a random writer instance, two opens that recover the same objects
encode byte-identical start objects, and a conditional create treats the
same bytes as its own retry. A start object that is already present
therefore fails the open with a retryable error instead of letting both
opens claim the epoch; the next open counts the object and starts a
later epoch.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): derive the WAL epoch from the claimed start sequence

Two opens can recover different listings when a late object lands across
a sequence gap between them, pick the same largest epoch plus one, and
both create their start objects under different sequences. The epoch of
an instance is now one above the sequence its start object claims, so a
successful create decides the epoch and no two instances share one. It
stays above every epoch recovery listed, and an object that carries an
epoch above the next sequence fails the open.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* fix(log-store): never run a held WAL create after the store is dropped

The create test hook ignored the closed hold channel, so a create parked
when the store was dropped could still run. It now returns without
creating. Drop the per-admission bookkeeping of issued entry ids, which
nothing reads, and move the parked I/O documentation to its helper.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test(log-store): cover a WAL start object stored with an unknown outcome

Add a fault-injection test in which the create of the start object
stores the object but reports an error: the open fails without moving
to another sequence, and the next open counts the stored object and
claims a later epoch. Rename the test helper that writes a whole object
from a given header to put_object_with_header, and drop a needless
clone.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

* test(log-store): make the held-create drop test deterministic

Keep the actor running while the store drops the hold sender, so the
parked create always completes on the closed channel instead of racing
the actor's exit. Drop a comment that restates epoch_of.

Signed-off-by: jeremyhi <fengjiachun@gmail.com>

---------

Signed-off-by: jeremyhi <fengjiachun@gmail.com>
2026-09-24 11:07:29 +00:00
Yingwen 3c2aac0a55 feat: add manual series index reconciliation (#9323)
* feat(mito): support manual series index reconciliation

Signed-off-by: evenyag <realevenyag@gmail.com>

* feat(storage): route series index build requests

Signed-off-by: evenyag <realevenyag@gmail.com>

* feat(admin): add BUILD_SERIES_INDEX

Signed-off-by: evenyag <realevenyag@gmail.com>

* test: specify compaction type in series index fixtures

Signed-off-by: evenyag <realevenyag@gmail.com>

* test: correct series index SQL fixtures and error assertions

Signed-off-by: evenyag <realevenyag@gmail.com>

* test(compat): preserve legacy index rebuild across upgrades

Signed-off-by: evenyag <realevenyag@gmail.com>

* test(sql): cover series index admin validation

Signed-off-by: evenyag <realevenyag@gmail.com>

* refactor(mito): bound series index maintenance queue

Signed-off-by: evenyag <realevenyag@gmail.com>

* docs: remove series index how-to guide

Signed-off-by: evenyag <realevenyag@gmail.com>

* test: remove index build upgrade compatibility case

Signed-off-by: evenyag <realevenyag@gmail.com>

* fix: address series index reconciliation review feedback

Signed-off-by: evenyag <realevenyag@gmail.com>

* fix: bound manual series index reconciliation admission

Signed-off-by: evenyag <realevenyag@gmail.com>

* chore: update greptime-proto to merged index build options

Signed-off-by: evenyag <realevenyag@gmail.com>

---------

Signed-off-by: evenyag <realevenyag@gmail.com>
2026-09-24 09:39:21 +00:00