test_writer_abort_is_unsupported_without_atomic_write asserted the file
content immediately after abort() returned Unsupported. SecureFsWriter
writes through tokio::fs::File, whose write_all() only enqueues a blocking
write task (tokio's poll_write returns Ready before the write completes),
so the data may not be visible yet when the test reads the file. Drop the
race-prone content assertion and only verify the Unsupported contract.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(auth): warn when credential load disables Postgres SCRAM or drops a line
Static and watch user providers degraded silently in two ways:
- A single non-SCRAM verifier (mysql_native_password, or a legacy
pbkdf2_sha256 hash that predates SCRAM) disables Postgres SCRAM for
every user and falls back to cleartext, with no signal to the operator.
- A malformed credential line (commonly a plaintext password containing
'=', which splits into more than two parts) was dropped without a trace.
Emit a warning at each credential load for both cases so operators don't
unknowingly serve cleartext passwords over Postgres or lose a user. This
is logging only; authentication behavior is unchanged. The SCRAM check
never logs secrets, and the malformed-line warning logs the line number
and file, never the line content.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(auth): warn on credential file read error before truncating
A read error from lines() (I/O failure or invalid UTF-8) ends the
iterator via map_while, silently dropping every remaining credential.
Warn with the line number and file before truncating, matching the
malformed-line handling, so the drop is observable.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(mito2): fence async index builds by schema generation
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(mito2): retry stale index builds after schema changes
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor(mito2): split compaction module into scheduler/reader submodules
Extract the compaction scheduler lifecycle (scheduler, status, phases,
execution, SST reservations, pending requests) and its tests out of
compaction.rs into compaction/scheduler.rs and compaction/scheduler_test.rs.
Split the remaining helpers by responsibility:
- estimate_compaction_bytes/refresh_picker_output move to scheduler.rs,
the only call site
- get_expired_ssts moves to picker.rs, shared by the TWCS and window
pickers
- CompactionSstReaderBuilder/time_range_to_predicate/ts_to_lit move to
the new compaction/reader.rs
The root compaction.rs keeps the shared output types and
find_dynamic_options, and re-exports the moved types so existing call
paths stay unchanged. Pure code motion, no behavior change.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito2): separate compaction scheduling from execution details
Turn compaction/scheduler.rs into a directory module to make the
scheduling flow easier to review:
- scheduler.rs keeps the pure scheduling core: the CompactionScheduler
state machine, scheduling entry points, termination chaining, DDL
coordination and region lifecycle events
- scheduler/planning.rs holds the execution-facing parts: background
planning dispatch, picker invocation, plan acceptance, remote/local
submission and memory estimation
- scheduler/state.rs holds the per-region lifecycle types:
CompactionStatus, ActiveCompaction, CompactionPhase, CompactingFiles,
LocalCompactionState, CompactionExecution and PendingCompaction
Child modules keep access to the scheduler's private methods, so the
split is pure code motion with minimal visibility changes (pub(super)
only where the parent module or tests reach into child items).
No behavior change.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito2): use absolute scheduler imports
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* docs(mito2): document compaction scheduler modules
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
---------
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito2): make async index publication conditional
Check the captured SST generation and commit index metadata under the same manifest write lock. Revalidate the committed metadata before applying it to the in-memory version, and clean exact-version artifacts when either publication stage becomes stale.
Add deterministic compaction and overlapping-index tests covering reopen consistency, duplicate rows, cache cleanup, and both file purgers.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* refactor(mito2): centralize manifest update finalization
Share the locked update, lock release, follower check, and hook firing path between regular manifest updates and conditional index publication.
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(mito2): retain index build leases across reopen
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(mito2): avoid retiring scheduler on index failure
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(mito2): handle cross-region index publication
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(mito2): use physical region for index paths
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(cmd): update noop index builder
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(mito2): avoid reusing published index versions
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* chore(mito2): log untracked index build stops
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* fix(mito2): set compaction time range in index test
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
---------
Signed-off-by: Dennis Zhuang <killme2008@gmail.com>
* feat: expose MitoRegion::all_manifest_files for metadata rebuild
Adds a public read-only accessor that returns all live SST file metas
and the current manifest version from the region manifest. Used by the
downstream project admin path to enumerate the authoritative live file
set without going through the worker loop.
* fix(mito2): merge staging manifest files in all_manifest_files
The original implementation only read the normal manifest
(manifest_ctx.manifest()) and skipped staging_manifest(). While the
region is in staging mode (e.g. region copy/migration), the authoritative
live file set lives in the staging manifest, so callers would silently
miss those files.
Now matches the semantics of manifest_sst_entries() (~L771), which
explicitly merges manifest().files with staging_manifest().files via a
HashMap collect (dedup by FileId). The returned manifest version is the
staging version when a staging manifest is present, otherwise the normal
manifest version.
Also removed downstream-specific references from the rustdoc comments.
* feat: add a dedicated http api server port
* fix: integration test
* refactor: make http-api-port opt-in
* refactor: rename attribute to http-api-server
* feat: use middleware to check different http server port
* refactor: rename config option
* refactor(mito2): run compaction picking in background with plan tracking
Move the compaction picker out of the region worker's critical path by
dispatching planning to a background task and reporting the result back
via CompactionPickFinished. CompactionStatus now tracks an explicit
picking phase keyed by a monotonic plan id, so stale planning results
are rejected and duplicate regular triggers coalesce while picking.
Before submitting a prepared compaction, the picker output is refreshed
against the current SST version (file handles are re-resolved and
conflicts roll back reservations), ensuring the plan still matches live
state. CompactionExecution identifies the running task by
(plan id, kind, version control) so finish/cancel/fail notifications
from outdated executions are ignored.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito2): notify pick finished even when compaction planning panics
The worker only leaves the picking phase after receiving the
CompactionPickFinished notification. Previously the planning task was
spawned fire-and-forget: if it panicked before sending the notification,
the region would be stuck in the picking phase forever, blocking all
future compactions and pending DDLs (e.g. entering staging) of the
region.
Wrap the planning future with catch_unwind so a panic is converted into
a CompactionPlanningResult::Error and the notification is always sent,
letting the worker run the normal error cleanup path.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito2): dec inflight compaction gauge after re-entrancy guard
DefaultNotifier::notify decremented INFLIGHT_COMPACTION_COUNT before the
re-entrancy guard, so a duplicate notify (which should never happen, but
the guard exists to defend against it) would decrement the gauge an
extra time and let it drift negative. Move the decrement after the
guard, matching the local compaction path's guard-then-account order.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito2): remove idle compaction status to prevent zombie state
When a compaction finished within min_compaction_interval with no
pending requests, on_compaction_finished left an idle status
(phase = None) behind. The worker then skipped chaining the next
compaction due to the interval gate, and the leftover status made
schedule_compaction swallow all future triggers of the region: regular
waiters were queued but never woken, and manual requests stayed pending
forever. The region stopped compacting until close/drop/truncate.
Add CompactionScheduler::remove_idle_status and call it from
handle_compaction_finished when the interval has not elapsed and no
chained planning is scheduled. The chain-until-no-plan semantics for
compactions that outlast the interval is preserved.
Also drops an unused import left by the previous commit.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito2): make compaction scheduling methods synchronous
schedule_compaction, handle_pending_compaction_request and
schedule_next_compaction no longer await anything after compaction
planning became fire-and-forget. Drop the async signature to make the
no-suspension-point invariant explicit: these methods always run to
completion on the worker loop without reentrancy.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito2): report stale compaction execution instead of region closed
When a compaction finishes but its execution no longer matches the
current one, the region may have been reopened or truncated, or the
compaction was superseded. Reporting RegionClosed to waiters is
misleading; introduce a neutral StaleCompactionExecution error (same
Cancelled status code) for this case.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito2): serialize truncate with compaction
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito2): avoid panic-based compaction status lookups
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito2): keep in-flight compaction plan when scheduling next
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito2): notify cancelled compaction pending ddl
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* typo: rename prefence to pre_fence to bypass typo check
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* test(mito2): trim redundant compaction tests
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* test(mito2): move compaction tests to dedicated file
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix: typo and format
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* Revert "test(mito2): move compaction tests to dedicated file"
This reverts commit e202f1f5
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* chore: revert test movement
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito2): remove redundant compaction status lookups
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix: typo
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* test(mito2): remove duplicate compaction test file
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* test(mito2): prune redundant compaction tests
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito2): simplify compaction plan identity
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* docs(mito2): design pending regular state simplification
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* docs(mito2): plan pending regular state simplification
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito2): simplify pending regular compaction state
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* remove: plan files
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito2): clarify compaction completion handling
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito2): inline compaction phase execution lookup
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito2): warn instead of panic on pending DDL for non-compacting region
add_ddl_request_to_pending unwrapped the region status and panicked when
the region was not compacting. Log a warning and skip the request instead,
and inline the now-trivial CompactionStatus::queue_ddl helper.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito2): dispatch pending DDLs before chaining regular compaction
A DDL queued behind a TooLateToCancel compaction (commit started or
remote execution) was deferred behind a whole extra plan/execution
cycle when a regular trigger had been retained during picking. Dispatch
the pending DDLs as soon as the current task finishes instead: satisfy
the retained regular waiters with the just-finished compaction, remove
the region status, and return the DDLs immediately.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito2): demote stale compaction failure logs to debug
Check region presence and execution staleness before logging, so a
superseded execution's terminal failure no longer emits a misleading
error log.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito2): atomically queue compaction DDLs
Combine compaction cancellation and dependent DDL enqueueing under one
status borrow. Return the typed request unchanged when no compaction is
running, avoiding both an unreachable warning branch and silent DDL
loss if the invariant changes.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* docs(mito2): explain why picker output handles are re-resolved
Addresses review question on refresh_picker_output: picking runs in
background on a possibly-stale version snapshot, so handles must be
re-resolved against the current version at accept time to detect
removed files and to read/reserve the up-to-date handle.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito2): keep compaction gate in test module
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* style(mito2): format compaction DDL helper calls
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito2): group active compaction state
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito2): fence compaction triggers behind pending DDL
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* refactor(mito2): simplify pending DDL collection in compaction scheduler
Replace the take-and-restore dance of the active compaction state with
an up-front busy check before handling pending compaction requests,
then take the active state once to drain DDL waiters.
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* docs: clarify that pending_request only carries manual StrictWindow compaction in production
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
* fix(mito2): arm DDL gate before cancellation
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>
---------
Signed-off-by: Lei, HUANG <ratuthomm@gmail.com>