rust/neon - neon - Gitea: Git with a cup of tea

rust/neon

mirror of https://github.com/neondatabase/neon.git synced 2026-05-31 12:00:42 +00:00

Author	SHA1	Message	Date
anastasia	81e94d1897	Add LSN and Backpressure descriptions to glossary.md	2022-01-24 12:52:30 +03:00
Konstantin Knizhnik	7bc1274a03	Fix comparison with disk_consistent_lsn in newer_image_layer_exists (#1167 )	2022-01-24 12:19:18 +03:00
Dmitry Rodionov	5f5a11525c	Switch our python package management solution to poetry. Mainly because it has better support for installing the packages from different python versions. It also has better dependency resolver than Pipenv. And supports modern standard for python dependency management. This includes usage of pyproject.toml for project specific configuration instead of per tool conf files. See following links for details: https://pip.pypa.io/en/stable/reference/build-system/pyproject-toml/ https://www.python.org/dev/peps/pep-0518/	2022-01-24 11:33:47 +03:00
Konstantin Knizhnik	e209764877	Do not delete layers beyand cutoff LSN (#1128 ) * Do not delete layers beyand cutoff LSN * Update pageserver/src/layered_repository/layer_map.rs Co-authored-by: Heikki Linnakangas <heikki.linnakangas@iki.fi> Co-authored-by: Heikki Linnakangas <heikki.linnakangas@iki.fi>	2022-01-24 10:42:40 +03:00
Kirill Bulatov	65290b2e96	Ensure every submodule compiles on its own	2022-01-21 17:34:15 +03:00
Dmitry Ivanov	127df96635	[proxy] Make `NUM_BYTES_PROXIED_COUNTER` more precise	2022-01-21 17:31:19 +03:00
Kirill Bulatov	924d8d489a	Allow enabling S3 mock in all existing tests with an env var	2022-01-20 18:42:47 +02:00
Dmitry Rodionov	026eb64a83	Use python lib to mock s3	2022-01-20 18:42:47 +02:00
Kirill Bulatov	45124856b1	Better S3 remote storage logging	2022-01-20 18:42:47 +02:00
Kirill Bulatov	38c6f6ce16	Allow specifying custom endpoint in s3	2022-01-20 18:42:47 +02:00
Heikki Linnakangas	caa62eff2a	Fix description of proxy --auth-endpoint option.	2022-01-20 14:50:27 +03:00
Dmitry Ivanov	d3542c34f1	Refactoring: use anyhow::Context's methods where possible	2022-01-19 16:33:48 +03:00
Kirill Bulatov	7fb62fc849	Fix macos compilation	2022-01-18 23:01:04 +02:00
Andrey Taranik	9d6ae06663	monitoring turn on for proxy (#1146 )	2022-01-18 19:23:53 +03:00
Alexey Kondratov	06c28174c2	Integrate compute_tools into zenith workspace and improve logging (zenithdb/console#487 )	2022-01-18 18:47:31 +03:00
bojanserafimov	8af1b43074	proxy: Add new metrics (#1132 )	2022-01-14 19:12:43 -05:00
Heikki Linnakangas	17b7caddcb	Update vendor/postgres: silence excessive logging from walproposer.	2022-01-14 20:51:02 +02:00
Heikki Linnakangas	dab30c27b6	Refactor thread management and shutdown This introduces a new module to handle thread creation and shutdown. All page server threads are now registered in a global hash map, and there's a function to request individual threads to shut down gracefully. Thread shutdown request is signalled to the thread with a flag, as well as a Future that can be used to wake up async operations if shutdown is requested. Use that facility to have the libpq listener thread respond to pageserver shutdown, based on Kirill's earlier prototype (https://github.com/zenithdb/zenith/pull/1088). That addresses https://github.com/zenithdb/zenith/issues/1036, previously the libpq listener thread would not exit until one more connection arrives. This also eliminates a resource leak in the accept() loop. Previously, we added the JoinHanlde of each new thread to a vector but old handles for threads that had already exited were never removed.	2022-01-14 18:36:10 +02:00
Heikki Linnakangas	bad1dd9759	Don't panic if spawning a new WAL receiver thread fails. The panic would kill the page service thread. That's not too bad, but still let's try to handle it more gracefully.	2022-01-14 18:02:34 +02:00
Heikki Linnakangas	d29836d0d5	Don't panic if spawning a thread to handle a connection fails. Log the error and continue. Hopefully it's a transient failure. This might have been happening in staging earlier, when the safekeeper had a problem where it opened connections very frequently to issue "callmemaybe" commands. If you launch too many threads too fast, you might run out of file descriptors or something. It's not totally clear what happened, but with commit, at least the page server will continue to run and accept new connections, if a transient error happens.	2022-01-14 18:02:30 +02:00
Heikki Linnakangas	adb0b3dada	Include backtrace in error messages in the log. 'anyhow' crate can include a backtrace in all errors, when the 'backtrace' feature is enabled. Enable it, and change the places that used '{:#}' or '{}' to '{:?}', so that the backtrace is printed.	2022-01-14 10:10:17 +02:00
bojanserafimov	5e0f39cc9e	Add proxy metrics (#1093 )	2022-01-13 20:34:30 -05:00
Arthur Petukhovsky	0a34a592d5	Bump vendor/postgres (#1120 )	2022-01-13 20:28:37 +03:00
Heikki Linnakangas	19aaa91f6d	Timeline IDs are not globally unique, fix some code that assumed that. A timeline ID is only guaranteed to be unique for a particular tenant, so you need to use tenant ID + timeline ID as the key, rather than just timeline ID. The safekeeper currently makes the same assumption, and we should fix that too, but this commit just addresses this one case in the page server. In the passing, reorder some function arguments to be more consistent.	2022-01-13 18:45:30 +02:00
Konstantin Knizhnik	404aab9373	Use mutex to prevent concurrent checkpoints (#1115 ) * Use mutex to prevent concurrent checkpoints * Fix comment	2022-01-13 17:48:24 +03:00
Konstantin Knizhnik	bc6db2c10e	Implement IO metrics in VirtualFile (#1112 ) * Implement IO metrics in VirtualFile * Do not group virtual file close statistics by tenantid/timelineid * Add comments concenring close metrics	2022-01-13 17:36:53 +03:00
Heikki Linnakangas	772d853dcf	Fix race condition leading to panic in walkeeper. The walkeeper launch two threads for each connection, and uses a guard object to remove entry from 'replicas' array, when finishes. But only the background thread held onto the guard object, so if the background thread finished before the other thread, the array entry would be removed prematurely, which lead to panic in the check_stop_streaming() call. Fixes https://github.com/zenithdb/zenith/issues/1103	2022-01-13 11:21:11 +02:00
Arseny Sher	ab4d272149	Add safekeeper --dump-control-file option. Hexalize zids there for better output; since Serde doesn't support several formats for one struct, on-disk representation is changed as well, make upgrade.rs cope with it.	2022-01-12 19:47:24 +03:00
Konstantin Knizhnik	f70a5cad61	Fix releasing of timelines lock (#1100 ) refer #1087	2022-01-12 15:05:08 +03:00
anastasia	7aba299dbd	Use safekeeper in test_branch_behind (#1068 ) to avoid a subtle race condition. Without safekeeper, walreceiver reconnection can stuck, because of IO deadlock between walsender auth and regular backend.	2022-01-12 14:38:04 +03:00
Kirill Bulatov	4b3b19f444	Support prefixes when working with s3 buckets	2022-01-11 15:44:50 +02:00
Kirill Bulatov	8ab4c8a050	Code review fixes	2022-01-11 15:44:23 +02:00
Kirill Bulatov	7c4a653230	Propagate Zenith CLI's RUST_LOG env var to subprocesses	2022-01-11 15:44:23 +02:00
Kirill Bulatov	a3cd8f0e6d	Add the remote storage test	2022-01-11 15:44:23 +02:00
Kirill Bulatov	65c851a451	Test pageserver's timeline http methods z	2022-01-11 15:44:23 +02:00
Kirill Bulatov	23cf2fa984	Properly shutdown storage sync loop	2022-01-11 15:44:23 +02:00
Kirill Bulatov	ce8d6ae958	Allow using remote storage in tests	2022-01-11 15:44:23 +02:00
Kirill Bulatov	384b2a91fa	Pass generic pageserver params through zenith cli	2022-01-11 15:44:23 +02:00
Arseny Sher	233c4811db	Fix default safekeeper http port.	2022-01-11 10:13:27 +03:00
Konstantin Knizhnik	2fd4c390cb	Do not hold timelines lock during GC (#1089 ) * Do not hold timelines lock during GC refer #1087 * Add gc_cs mutex for preveting creation of new timelines during GC * Make clippy happy * Use Mutex<()> instead of Mutex<i32> for GC critical section	2022-01-10 14:41:15 +03:00
bojanserafimov	5b9391b51d	Support "query cancel" in proxy (#1052 )	2022-01-05 17:27:12 -05:00
Arthur Petukhovsky	5a6405848d	Bump vendor/postgres (#1086 )	2022-01-05 14:27:51 +03:00
Patrick Insinger	191d9d2b74	par_fsync - use VirtualFile	2022-01-04 20:40:57 -08:00
Patrick Insinger	24c8dab86f	pageserver - parallelize checkpoint fsyncs	2022-01-04 20:40:57 -08:00
Heikki Linnakangas	55a4cf64a1	Refactor WAL record handling. Introduce the concept of a "ZenithWalRecord", which can be a Postgres WAL record that is replayed with the Postgres WAL redo process, or a built-in type that is handled entirely by pageserver code. Replace the special code to replay Postgres XACT commit/abort records with new Zenith WAL records. A separate zenith WAL record is created for each modified CLOG page. This allows removing the 'main_data_offset' field from stored PostgreSQL WAL records, which saves some memory and some disk space in delta layers. Introduce zenith WAL records for updating bits in the visibility map. Previously, when e.g. a heap insert cleared the VM bit, we duplicated the heap insert WAL record for the affected VM page. That was very wasteful. The heap WAL record could be massive, containing a full page image in the worst case. This addresses github issue #941.	2022-01-04 11:26:37 +02:00
Heikki Linnakangas	722667f189	Add test case for performance issue #941 . The first COPY generates about 230 MB of write I/O, but the second COPY, after deleting most of the rows and vacuuming the rows away, generates 370 MB of writes. Both COPYs insert the same amount of data, so they should generate roughly the same amount of I/O. This commit doesn't try to fix the issue, just adds a test case to demonstrate it. Add a new 'checkpoint' command to the pageserver API. Previously, we've used 'do_gc' for that, but many tests, including this new one, really only want to perform a checkpoint and don't care about GC. For now, I only used the command in the new test, though, and didn't convert any existing tests to use it.	2022-01-04 11:26:37 +02:00
Arseny Sher	25a515b968	Don't call immediately on resume in callmemaybe. It creates busy loop if pageserver <-> safekeeper connection fails after it was established (e.g. currently due to 'segment checkpoint not found' error on pageserver). Also wake up callmemaybe thread regularly once in recall_period regardless of channel activity.	2022-01-03 20:44:36 +03:00
Konstantin Knizhnik	1c47fbae81	Do not write image layers during enforced checkpoint (#1057 ) * Do not write image layers during enforced checkpoint refer #1056 * Add Flush option to CheckpointConfig refer #1057	2022-01-01 19:08:09 +03:00
Alexey Kondratov	8f0cd7fb9f	[compute_tools] Switch cluster_id in spec to string (zenithdb/console#72 )	2021-12-29 16:35:29 +03:00
Dmitry Rodionov	c910132d4b	Fix wal receiver shutdown This patch allows to shutdown wal receiver when there are no messages and wal receiver is blocked inside tokio-postgres. In this case it cannot check the shutdown flag. This patch switches to use async interface of tokio-postgres directly without sync wrappers. It opens the possibility to use tokio::select! between the phsycal_stream.next() and a shutdown channel readiness to interrupt replication process. Also this allows to shutdown only particular wal receiver without using global shutdown_requested flag.	2021-12-29 14:42:29 +03:00

1 2 3 4 5 ...

1193 Commits