rust/neon - neon - Gitea: Git with a cup of tea

rust/neon

mirror of https://github.com/neondatabase/neon.git synced 2026-01-06 13:02:55 +00:00

Author	SHA1	Message	Date
Egor Suvorov	eaff0cd568	Check python for the whole repository and improve docs (#813 )	2021-11-09 22:23:29 +03:00
Egor Suvorov	587935ebed	Add Safekeeper metrics tests (#746 ) * zenith_fixtures.py: add SafekeeperHttpClient.get_metrics() * Ensure that `collect_lsn` and `flush_lsn`'s reported values look reasonable in `test_many_timelines`	2021-11-09 22:18:59 +03:00
Dmitry Rodionov	07dddfed28	Use more robust way to persist safekeeper control file. Now safekeeper control file updated in a following way: 1. Write data to temp file 2. Fsync the temporary file (if sync option is specified) 3. Rename temporary file to actual control file 4. Fsync containing directory (if sync option is specified) 5. Fsync file after rename (if sync option is specified). Note that action 5 is not mentioned anywhere as required but it is done in postgres this way (see durable_rename). Also because of the rename machinery switch to use dedicated lock file to prevent running several safekeepers concurrently on the same data cleanup fsync control file after rename to match postgres behaviour	2021-11-09 17:51:46 +03:00
Arseny Sher	229dc7704f	Bump vendor/postgres.	2021-11-08 17:32:13 +03:00
Dmitry Rodionov	067f2ac814	fix perf repo branch name	2021-11-08 13:27:23 +03:00
Dmitry Rodionov	865870a8e5	Follow up staging benchmarking * change zenith-perf-data checkout ref to be main * set cluster id through secrets so there is no code changes required when we wipe out clusters on staging * display full pgbench output on error	2021-11-05 14:07:11 +03:00
Arthur Petukhovsky	d19263aec8	Adjust timeouts for test_restarts_under_load (#830 ) * Adjust timeouts for test_restarts_under_load * Add test timeout for test_restarts_under_load	2021-11-04 19:58:40 +03:00
Heikki Linnakangas	6d742719a1	Fix infinite loop in looking up predecessor layer Commit `960c7d69a8` changed the LSN returned in the Continue case in InMemoryLayer::get_page_reconstruct_data(), but neglected to make the same change in DeltaLayer. Also add an escape hatch to the loop in materialize_page() to avoid getting stuck in an infinite loop, if a bug like this reoccurs.	2021-11-04 16:07:12 +02:00
Dmitry Rodionov	c75bc9b8b0	Change benchmark plugin layout so pytest loads it properly when running all tests (not necessary performance ones) resolves #837	2021-11-04 16:33:31 +03:00
Egor Suvorov	33007cc0bb	Safekeeper's `START_REPLICATION` handler: remove `stop_point`, do not handle `start_point == 0` (#777 )	2021-11-04 14:50:33 +03:00
Dmitry Rodionov	987833e0b9	Propagate git SHA to zenith binaries Git commit sha is displayed when --version flag is used and is written to logs during service startup. Uses git_version crate when git is available, and GIT_VERSION environment variable otherwise which is the case for docker builds.	2021-11-04 14:22:29 +03:00
Kirill Bulatov	f36acf00de	Reduce "relish" word usages in remote storage	2021-11-04 12:53:42 +02:00
Kirill Bulatov	956fc3dec9	Tidy up and make consistent the remote storate API	2021-11-04 12:53:42 +02:00
Heikki Linnakangas	b38e841f2d	Use poll() in communication with WAL redo process. The tokio futures added some overhead, so switch to plain non-blocking I/O with poll(). In a simple pgbench test on my laptop (select-only queries, scale-factor 1 `pgbench -P1 -T50 -S`), this gives about 10% improvement, from about 4300 TPS to 4800 TPS.	2021-11-04 10:39:04 +02:00
Heikki Linnakangas	3a0111c75e	Refactor functions for constructing WAL redo messages. Instead of building a separate Vec<u8> to hold each message, serialize all the messages to one big Vec<u8>. This eliminates some Vec allocation and memcpy() overhead. The downside is that if there are a lot of records to replay, we have to serialize them all into one big chunk of memory. That shouldn't be a problem in practice. If you need to replay millions of records to reconstruct a page, we should've materialized a new image of that page earlier already.	2021-11-04 10:39:00 +02:00
Heikki Linnakangas	086a02ab92	Add performance test for simple seq scans. Fixes https://github.com/zenithdb/zenith/issues/831	2021-11-04 10:36:45 +02:00
Heikki Linnakangas	7ed39655dc	Bump vendor/postgres	2021-11-04 10:35:50 +02:00
Dmitry Rodionov	c6172dae47	implement performance tests against our staging environment tests are based on self-hosted runner which is physically close to our staging deployment in aws, currently tests consist of various configurations of pgbenchi runs. Also these changes rework benchmark fixture by removing globals and allowing to collect reports with desired metrics and dump them to json for further analysis. This is also applicable to usual performance tests which use local zenith binaries.	2021-11-04 02:15:46 +03:00
Heikki Linnakangas	4ba783d0af	Remove a couple of unused functions. We might want to have custom serialize/deserialize functions for WALRecords and PageVersions for performance reasons, see github issue 832. But that would probably look a bit different from this, and currently these functions are just dead.	2021-11-03 19:10:23 +02:00
Patrick Insinger	0457fe81a9	pageserver - make `PageVersion` an enum	2021-11-03 09:28:49 -07:00
Heikki Linnakangas	fb524dd973	Put a global limit on memory used by in-memory layers. Adds simple global tracking of memory used by the in-memory layers. It's very approximate, it doesn't take into account allocator, memory fragmentation or many other things, but it's a good first step. After storing a WAL record in the repository, the WAL receiver checks if the global memory usage. If it's above a configurable threshold (hard coded at 128 MB at the moment), it evicts a layer. The victim layer is chosen by GClock algorithm, similar to that used in the Postgres buffer cache. This stops the page server from using an unbounded amount of memory. It's pretty crude, the eviction and materializing and writing a layer to disk happens now in the WAL receiver thread. It would be nice to move that to a background thread, and it would be nice to have a smarter policy on when to materialize a new image layer and when to just write out a delta layer, and it would be nice to have more accurate accounting of memory. But this should fix the most pressing OOM issues, and is a step in the right direction. Co-authored-by: Patrick Insinger <patrickinsinger@gmail.com>	2021-11-02 15:49:39 +02:00
Heikki Linnakangas	8c6d2664c0	Support removing arbitrary open layers, not just the oldest one	2021-11-02 15:43:16 +02:00
Patrick Insinger	cdbbd15eb9	pageserver - add InMemoryLayer global map (#817 )	2021-11-01 12:20:24 -07:00
anastasia	85f8bf97f5	Name walkeeper threads to make debugging more convenient	2021-11-01 19:09:57 +03:00
anastasia	83ed930bc2	WIP. Launch and shutdown tenant threads together with walreceiver. TODO: now walreceiver only disconnects if safekeeper was shut down. Implemnt proper walreceiver disconnection.	2021-11-01 18:04:00 +03:00
anastasia	071e30cc53	Expose TENANT_THREADS_COUNT metric to observe number of currently active checkpointer and GC threads	2021-11-01 18:04:00 +03:00
Kirill Bulatov	e6ef27637b	Better API to handle timeline metadata properly	2021-10-29 23:51:40 +03:00
Patrick Insinger	b532470792	Set SO_REUSEADDR for all TCP listeners	2021-10-29 12:45:26 -07:00
Heikki Linnakangas	e0d7ecf91c	Refactor 'zenith' CLI subcommand handling Also fixes 'zenith safekeeper restart -m immediate'. The stop-mode was previously ignored.	2021-10-29 19:01:01 +03:00
Kirill Bulatov	edba2e9744	Use a proper extension for the readme file	2021-10-28 18:55:14 +03:00
Egor Suvorov	7e552b645f	Add disk write/sync metrics to Safekeeper (#745 )	2021-10-28 18:38:36 +03:00
anastasia	ea5900f155	Refactoring of checkpointer and GC. Move them to a separate tenant_threads module to detangle thread management from LayeredRepository implementation.	2021-10-27 20:50:26 +03:00
anastasia	28ab40c8b7	fix init_repo() call in register_relish_download()	2021-10-27 20:50:26 +03:00
Alexey Kondratov	d423142623	Proxy: wait for kick on .pgpass connection (zenithdb/console#227 )	2021-10-27 20:24:23 +03:00
Dmitry Rodionov	1c0e85f9a0	review cleanups	2021-10-27 13:30:34 +03:00
Dmitry Rodionov	5bc09074ea	add a flag to avoid non incremental size calculation in pageserver http api This calculation is not that heavy but it is needed only in tests, and in case the number of tenants/timelines is high the calculation can take noticeable time. Resolves https://github.com/zenithdb/zenith/issues/804	2021-10-27 13:30:34 +03:00
Heikki Linnakangas	1fac4a3c91	Fix a few messages. Pointed out by Egor in https://github.com/zenithdb/zenith/pull/788, but I accidentally pushed that before fixing these.	2021-10-27 10:58:21 +03:00
Heikki Linnakangas	1bc917324d	Use -m immediate for 'immediate' shutdown	2021-10-27 10:49:38 +03:00
Heikki Linnakangas	af429fb401	Improve 'zenith' CLI utility for safekeepers and a config file. The 'zenith' CLI utility can now be used to launch safekeepers. By default, one safekeeper is configured. There are new 'safekeeper start/stop' subcommands to manage the safekeepers. Each safekeeper is given a name that can be used to identify the safekeeper to start/stop with the 'zenith start/stop' commands. The safekeeper data is stored in '.zenith/safekeepers/<name>'. The 'zenith start' command now starts the pageserver and also all safekeepers. 'zenith stop' stops pageserver, all safekeepers, and all postgres nodes. Introduce new 'zenith pageserver start/stop' subcommands for starting/stopping just the page server. The biggest change here is to the 'zenith init' command. This adds a new 'zenith init --config=<path to toml file>' option. It takes a toml config file that describes the environment. In the config file, you can specify options for the pageserver, like the pg and http ports, and authentication. For each safekeeper, you can define a name and the pg and http ports. If you don't use the --config option, you get a default configuration with a pageserver and one safekeeper. Note that that's different from the previous default of no safekeepers. Any fields that are omitted in the configuration file are filled with defaults. You can also specify the initial tenant ID in the config file. A couple of sample config files are added in the control_plane/ directory. The --pageserver-pg-port, --pageserver-http-port, and --pageserver-auth options to 'zenith init' are removed. Use a config file instead. Finally, change the python test fixtures to use the new 'zenith' commands and the config file to describe the environment.	2021-10-27 10:49:38 +03:00
Heikki Linnakangas	710fe02d0b	Return success on 'zenith stop' if the page server is already stopped.	2021-10-27 01:10:24 +03:00
Heikki Linnakangas	de87aad990	Remove a few unused functions	2021-10-27 01:10:24 +03:00
Heikki Linnakangas	41d48719e1	In python tests, skip ports that are already in use. We've seen some failures with "Address already in use" errors in the tests. It's not clear why, perhaps some server processes are not cleaned up properly after test, or maybe the socket is still in TIME_WAIT state. In any case, let's make the tests more robust by checking that the port is free, before trying to use it.	2021-10-27 00:46:24 +03:00
Kirill Bulatov	d88377f9f0	Remove log from zenith_utils	2021-10-26 23:24:11 +03:00
Kirill Bulatov	ecd577c934	Simplify tracing declarations	2021-10-26 23:24:11 +03:00
anastasia	f43f8401ee	Don't wait for wal-redo process for non-relational records replay	2021-10-26 19:30:28 +03:00
Arseny Sher	1877bbc7cb	bump vendor/postgres to fix reconnection busy loop	2021-10-26 15:43:19 +03:00
Heikki Linnakangas	a064ebb64c	Cope with missing 'tenantid' in '.zenith/config' file. We generate the initial tenantid and store it in the file, so it shouldn't be missing. But let's cope with it. (This comes handy with the bigger changes I'm working on at https://github.com/zenithdb/zenith/pull/788)	2021-10-25 21:24:11 +03:00
Heikki Linnakangas	4726870e8d	Remove obsolete comment. We store the pageserver port in the .zenith/config file.	2021-10-25 21:16:58 +03:00
Heikki Linnakangas	3bbc106c70	Prefer long CLI option name for clarity.	2021-10-25 21:16:58 +03:00
Heikki Linnakangas	66eb081876	Improve comment on 'base_dir'	2021-10-25 21:16:58 +03:00

1 2 3 4 5 ...

1036 Commits