rust/neon

mirror of https://github.com/neondatabase/neon.git synced 2025-12-28 00:23:00 +00:00

Files

Alex Chi Z. 5c57e8a11b feat(pageserver): rework reldirv2 rollout (#12576 )

## Problem

LKB-197, #9516 

To make sure the migration path is smooth.

The previous plan is to store new relations in new keyspace and old ones
in old keyspace until it gets dropped. This makes the migration path
hard as we can't validate v2 writes and can't rollback. This patch gives
us a more smooth migration path:

- The first time we enable reldirv2 for a tenant, we copy over
everything in the old keyspace to the new one. This might create a short
spike of latency for the create relation operation, but it's oneoff.
- After that, we have identical v1/v2 keyspace and read/write both of
them. We validate reads every time we list the reldirs.
- If we are in `migrating` mode, use v1 as source of truth and log a
warning for failed v2 operations. If we are in `migrated` mode, use v2
as source of truth and error when writes fail.
- One compatibility test uses dataset from the time where we enabled
reldirv2 (of the original rollout plan), which only has relations
written to the v2 keyspace instead of the v1 keyspace. We had to adjust
it accordingly.
- Add `migrated_at` in index_part to indicate the LSN where we did the
initialize.

TODOs:

- Test if relv1 can be read below the migrated_at LSN.
- Move the initialization process to L0 compaction instead of doing it
on the write path.
- Disable relcache in the relv2 test case so that all code path gets
fully tested.

## Summary of changes

- New behavior of reldirv2 migration flags as described above.

---------

Signed-off-by: Alex Chi Z <chi@neon.tech>

2025-07-23 16:12:46 +00:00

large_synthetic_oltp

Increase tenant size for large tenant oltp workload (#12260 )

2025-06-18 12:40:25 +00:00

many_relations

Run pgbench on 10 GB scale factor on database with n relations (e.g. 10k) (#10172 )

2024-12-19 10:25:44 +00:00

pageserver

pageserver: Introduce config to enable/disable eviction task (#12496 )

2025-07-08 21:14:04 +00:00

pgvector

Enable all pyupgrade checks in ruff

2024-10-08 14:32:26 -05:00

tpc-h

Nightly Benchmarks: add TPC-H benchmark (#2978 )

2022-12-08 15:32:49 +00:00

__init__.py

Enable all pyupgrade checks in ruff

2024-10-08 14:32:26 -05:00

out_dir_to_csv.py

benchmarking: extend test_page_service_batching.py to cover concurrent IO + batching under random reads (#10466 )

2025-05-15 17:48:13 +00:00

README.md

benchmarking: extend test_page_service_batching.py to cover concurrent IO + batching under random reads (#10466 )

2025-05-15 17:48:13 +00:00

test_branch_creation.py

introduce new runners: unit-perf and use them for benchmark jobs (#11409 )

2025-04-15 08:21:44 +00:00

test_branching.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_bulk_insert.py

test_bulk_insert: fix typing for PgVersion (#9854 )

2024-11-22 16:13:53 +00:00

test_bulk_tenant_create.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_bulk_update.py

use a prod-like shared_buffers size for some perf unit tests (#11373 )

2025-04-02 10:43:05 +00:00

test_compaction.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_compare_pg_stats.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_compute_ctl_api.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_compute_startup.py

tests: use endpoint http wrapper to get auth (#11628 )

2025-04-17 15:03:23 +00:00

test_copy.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_cumulative_statistics_persistence.py

delete orphan left over projects (#11826 )

2025-05-05 14:30:13 +00:00

test_dup_key.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_gc_feedback.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_gist_build.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_hot_page.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_hot_table.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_ingest_insert_bulk.py

introduce new runners: unit-perf and use them for benchmark jobs (#11409 )

2025-04-15 08:21:44 +00:00

test_ingest_logical_message.py

use a prod-like shared_buffers size for some perf unit tests (#11373 )

2025-04-02 10:43:05 +00:00

test_latency.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_layer_map.py

Record more timings in test_layer_map (#10670 )

2025-02-05 17:00:26 +00:00

test_lfc_prewarm.py

LFC prewarm perftest: increase timeout for initialization job (#12594 )

2025-07-14 17:37:47 +00:00

test_logical_replication.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_parallel_copy_to.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_parallel_copy.py

use a prod-like shared_buffers size for some perf unit tests (#11373 )

2025-04-02 10:43:05 +00:00

test_perf_ingest_using_pgcopydb.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_perf_many_relations.py

feat(pageserver): rework reldirv2 rollout (#12576 )

2025-07-23 16:12:46 +00:00

test_perf_olap.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_perf_oltp_large_tenant.py

Increase tenant size for large tenant oltp workload (#12260 )

2025-06-18 12:40:25 +00:00

test_perf_pgbench.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_perf_pgvector_queries.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_physical_replication.py

delete orphan left over projects (#11826 )

2025-05-05 14:30:13 +00:00

test_random_writes.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_seqscans.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_sharded_ingest.py

pageserver: remove handling of vanilla protocol (#12126 )

2025-06-05 11:43:04 +00:00

test_sharding_autosplit.py

storcon: validate intent state before applying optimization (#12593 )

2025-07-16 14:37:40 +00:00

test_storage_controller_scale.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_wal_backpressure.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

test_write_amplification.py

ruff: enable TC — flake8-type-checking (#11368 )

2025-03-30 18:58:33 +00:00

README.md

Running locally

First make a release build. The -s flag silences a lot of output, and makes it easier to see if you have compile errors without scrolling up. BUILD_TYPE=release CARGO_BUILD_FLAGS="--features=testing" make -s -j8

You may also need to run ./scripts/pysync.

Then run the tests DEFAULT_PG_VERSION=17 NEON_BIN=./target/release poetry run pytest test_runner/performance

Some handy pytest flags for local development:

-x tells pytest to stop on first error
-s shows test output
-k selects a test to run
--timeout=0 disables our default timeout of 300s (see setup.cfg)
--preserve-database-files to skip cleanup
--out-dir to produce a JSON with the recorded test metrics. There is a post-processing tool at test_runner/performance/out_dir_to_csv.py.

What performance tests do we have and how we run them

Performance tests are built using the same infrastructure as our usual python integration tests. There are some extra fixtures that help to collect performance metrics, and to run tests against both vanilla PostgreSQL and Neon for comparison.

Tests that are run against local installation

Most of the performance tests run against a local installation. This is not very representative of a production environment. Firstly, Postgres, safekeeper(s) and the pageserver have to share CPU and I/O resources, which can add noise to the results. Secondly, network overhead is eliminated.

In the CI, the performance tests are run in the same environment as the other integration tests. We don't have control over the host that the CI runs on, so the environment may vary widely from one run to another, which makes the results across different runs noisy to compare.

Remote tests

There are a few tests that marked with pytest.mark.remote_cluster. These tests do not set up a local environment, and instead require a libpq connection string to connect to. So they can be run on any Postgres compatible database. Currently, the CI runs these tests on our staging and captest environments daily. Those are not an isolated environments, so there can be noise in the results due to activity of other clusters.

Noise

All tests run only once. Usually to obtain more consistent performance numbers, a test should be repeated multiple times and the results be aggregated, for example by taking min, max, avg, or median.

Results collection

Local test results for main branch, and results of daily performance tests, are stored in a neon project deployed in production environment. There is a Grafana dashboard that visualizes the results. Here is the dashboard. The main problem with it is the unavailability to point at particular commit, though the data for that is available in the database. Needs some tweaking from someone who knows Grafana tricks.

There is also an inconsistency in test naming. Test name should be the same across platforms, and results can be differentiated by the platform field. But currently, platform is sometimes included in test name because of the way how parametrization works in pytest. I.e. there is a platform switch in the dashboard with neon-local-ci and neon-staging variants. I.e. some tests under neon-local-ci value for a platform switch are displayed as Test test_runner/performance/test_bulk_insert.py::test_bulk_insert[vanilla] and Test test_runner/performance/test_bulk_insert.py::test_bulk_insert[neon] which is highly confusing.