neon/storage_controller at 342607473ad4afcee7e199de8ee2c133c23b73b9 - neon

rust/neon

mirror of https://github.com/neondatabase/neon.git synced 2026-01-08 14:02:55 +00:00

Files

John Spray 52dee408dc storage controller: improve safety of shard splits coinciding with controller restarts (#11412 )

## Problem

The graceful leadership transfer process involves calling step_down on
the old controller, but this was not waiting for shard splits to
complete, and the new controller could therefore end up trying to abort
a shard split while it was still going on.

We mitigated this already in #11256 by avoiding the case where shard
split completion would update the database incorrectly, but this was a
fragile fix because it assumes that is the only problematic part of the
split running concurrently.

Precursors:
- #11290 
- #11256

Closes: #11254 

## Summary of changes

- Hold the reconciler gate from shard splits, so that step_down will
wait for them. Splits should always be fairly prompt, so it is okay to
wait here.
- Defense in depth: if step_down times out (hardcoded 10 second limit),
then fully terminate the controller process rather than letting it
continue running, potentially doing split-brainy things. This makes
sense because the new controller will always declare itself leader
unilaterally if step_down fails, so leaving an old controller running is
not beneficial.
- Tests: extend
`test_storage_controller_leadership_transfer_during_split` to separately
exercise the case of a split holding up step_down, and the case where
the overall timeout on step_down is hit and the controller terminates.

2025-04-10 16:55:37 +00:00

client

storcon + safekeeper + scrubber: propagate root CA certs everywhere (#11418 )

2025-04-04 06:30:48 +00:00

migrations

storcon: timetime table, creation and deletion (#11058 )

2025-03-11 02:31:22 +00:00

src

storage controller: improve safety of shard splits coinciding with controller restarts (#11412 )

2025-04-10 16:55:37 +00:00

Cargo.toml

storcon: add https API (#11239 )

2025-03-20 08:22:02 +00:00