rust/neon - neon - Gitea: Git with a cup of tea

rust/neon

mirror of https://github.com/neondatabase/neon.git synced 2026-01-08 05:52:55 +00:00

Author	SHA1	Message	Date
Konstantin Knizhnik	7bc1274a03	Fix comparison with disk_consistent_lsn in newer_image_layer_exists (#1167 )	2022-01-24 12:19:18 +03:00
Konstantin Knizhnik	e209764877	Do not delete layers beyand cutoff LSN (#1128 ) * Do not delete layers beyand cutoff LSN * Update pageserver/src/layered_repository/layer_map.rs Co-authored-by: Heikki Linnakangas <heikki.linnakangas@iki.fi> Co-authored-by: Heikki Linnakangas <heikki.linnakangas@iki.fi>	2022-01-24 10:42:40 +03:00
Dmitry Rodionov	026eb64a83	Use python lib to mock s3	2022-01-20 18:42:47 +02:00
Kirill Bulatov	45124856b1	Better S3 remote storage logging	2022-01-20 18:42:47 +02:00
Kirill Bulatov	38c6f6ce16	Allow specifying custom endpoint in s3	2022-01-20 18:42:47 +02:00
Dmitry Ivanov	d3542c34f1	Refactoring: use anyhow::Context's methods where possible	2022-01-19 16:33:48 +03:00
Heikki Linnakangas	dab30c27b6	Refactor thread management and shutdown This introduces a new module to handle thread creation and shutdown. All page server threads are now registered in a global hash map, and there's a function to request individual threads to shut down gracefully. Thread shutdown request is signalled to the thread with a flag, as well as a Future that can be used to wake up async operations if shutdown is requested. Use that facility to have the libpq listener thread respond to pageserver shutdown, based on Kirill's earlier prototype (https://github.com/zenithdb/zenith/pull/1088). That addresses https://github.com/zenithdb/zenith/issues/1036, previously the libpq listener thread would not exit until one more connection arrives. This also eliminates a resource leak in the accept() loop. Previously, we added the JoinHanlde of each new thread to a vector but old handles for threads that had already exited were never removed.	2022-01-14 18:36:10 +02:00
Heikki Linnakangas	bad1dd9759	Don't panic if spawning a new WAL receiver thread fails. The panic would kill the page service thread. That's not too bad, but still let's try to handle it more gracefully.	2022-01-14 18:02:34 +02:00
Heikki Linnakangas	d29836d0d5	Don't panic if spawning a thread to handle a connection fails. Log the error and continue. Hopefully it's a transient failure. This might have been happening in staging earlier, when the safekeeper had a problem where it opened connections very frequently to issue "callmemaybe" commands. If you launch too many threads too fast, you might run out of file descriptors or something. It's not totally clear what happened, but with commit, at least the page server will continue to run and accept new connections, if a transient error happens.	2022-01-14 18:02:30 +02:00
Heikki Linnakangas	adb0b3dada	Include backtrace in error messages in the log. 'anyhow' crate can include a backtrace in all errors, when the 'backtrace' feature is enabled. Enable it, and change the places that used '{:#}' or '{}' to '{:?}', so that the backtrace is printed.	2022-01-14 10:10:17 +02:00
Heikki Linnakangas	19aaa91f6d	Timeline IDs are not globally unique, fix some code that assumed that. A timeline ID is only guaranteed to be unique for a particular tenant, so you need to use tenant ID + timeline ID as the key, rather than just timeline ID. The safekeeper currently makes the same assumption, and we should fix that too, but this commit just addresses this one case in the page server. In the passing, reorder some function arguments to be more consistent.	2022-01-13 18:45:30 +02:00
Konstantin Knizhnik	404aab9373	Use mutex to prevent concurrent checkpoints (#1115 ) * Use mutex to prevent concurrent checkpoints * Fix comment	2022-01-13 17:48:24 +03:00
Konstantin Knizhnik	bc6db2c10e	Implement IO metrics in VirtualFile (#1112 ) * Implement IO metrics in VirtualFile * Do not group virtual file close statistics by tenantid/timelineid * Add comments concenring close metrics	2022-01-13 17:36:53 +03:00
Konstantin Knizhnik	f70a5cad61	Fix releasing of timelines lock (#1100 ) refer #1087	2022-01-12 15:05:08 +03:00
Kirill Bulatov	4b3b19f444	Support prefixes when working with s3 buckets	2022-01-11 15:44:50 +02:00
Kirill Bulatov	8ab4c8a050	Code review fixes	2022-01-11 15:44:23 +02:00
Kirill Bulatov	65c851a451	Test pageserver's timeline http methods z	2022-01-11 15:44:23 +02:00
Kirill Bulatov	23cf2fa984	Properly shutdown storage sync loop	2022-01-11 15:44:23 +02:00
Kirill Bulatov	384b2a91fa	Pass generic pageserver params through zenith cli	2022-01-11 15:44:23 +02:00
Konstantin Knizhnik	2fd4c390cb	Do not hold timelines lock during GC (#1089 ) * Do not hold timelines lock during GC refer #1087 * Add gc_cs mutex for preveting creation of new timelines during GC * Make clippy happy * Use Mutex<()> instead of Mutex<i32> for GC critical section	2022-01-10 14:41:15 +03:00
Patrick Insinger	191d9d2b74	par_fsync - use VirtualFile	2022-01-04 20:40:57 -08:00
Patrick Insinger	24c8dab86f	pageserver - parallelize checkpoint fsyncs	2022-01-04 20:40:57 -08:00
Heikki Linnakangas	55a4cf64a1	Refactor WAL record handling. Introduce the concept of a "ZenithWalRecord", which can be a Postgres WAL record that is replayed with the Postgres WAL redo process, or a built-in type that is handled entirely by pageserver code. Replace the special code to replay Postgres XACT commit/abort records with new Zenith WAL records. A separate zenith WAL record is created for each modified CLOG page. This allows removing the 'main_data_offset' field from stored PostgreSQL WAL records, which saves some memory and some disk space in delta layers. Introduce zenith WAL records for updating bits in the visibility map. Previously, when e.g. a heap insert cleared the VM bit, we duplicated the heap insert WAL record for the affected VM page. That was very wasteful. The heap WAL record could be massive, containing a full page image in the worst case. This addresses github issue #941.	2022-01-04 11:26:37 +02:00
Heikki Linnakangas	722667f189	Add test case for performance issue #941 . The first COPY generates about 230 MB of write I/O, but the second COPY, after deleting most of the rows and vacuuming the rows away, generates 370 MB of writes. Both COPYs insert the same amount of data, so they should generate roughly the same amount of I/O. This commit doesn't try to fix the issue, just adds a test case to demonstrate it. Add a new 'checkpoint' command to the pageserver API. Previously, we've used 'do_gc' for that, but many tests, including this new one, really only want to perform a checkpoint and don't care about GC. For now, I only used the command in the new test, though, and didn't convert any existing tests to use it.	2022-01-04 11:26:37 +02:00
Konstantin Knizhnik	1c47fbae81	Do not write image layers during enforced checkpoint (#1057 ) * Do not write image layers during enforced checkpoint refer #1056 * Add Flush option to CheckpointConfig refer #1057	2022-01-01 19:08:09 +03:00
Dmitry Rodionov	c910132d4b	Fix wal receiver shutdown This patch allows to shutdown wal receiver when there are no messages and wal receiver is blocked inside tokio-postgres. In this case it cannot check the shutdown flag. This patch switches to use async interface of tokio-postgres directly without sync wrappers. It opens the possibility to use tokio::select! between the phsycal_stream.next() and a shutdown channel readiness to interrupt replication process. Also this allows to shutdown only particular wal receiver without using global shutdown_requested flag.	2021-12-29 14:42:29 +03:00
Kirill Bulatov	f0afd08667	Fix zenith init defaults	2021-12-28 00:21:48 +02:00
Kirill Bulatov	b494ac1ea0	Remove redundant pageserver cli params	2021-12-27 18:38:54 +02:00
Arseny Sher	a163650a99	Refactor Postgres command parsing in safekeeper. Do it separately with SafekeeperPostgresCommand enum as a result. Since query is always C string, switch postgres_backend process_query argument from Bytes to &str. Make passing ztli/ztenant id in safekeeper connection string optional; this is needed for upcoming intra-safekeeper heartbeat cmd which is not bound to any timeline.	2021-12-24 15:48:13 +03:00
anastasia	980f5f8440	Propagate remote_consistent_lsn to safekeepers. Change meaning of lsns in HOT_STANDBY_FEEDBACK: flush_lsn = disk_consistent_lsn, apply_lsn = remote_consistent_lsn Update compute node backpressure configuration respectively. Update compute node configuration: set 'synchronous_commit=remote_write' in setup without safekeepers. This way compute node doesn't have to wait for data checkpoint on pageserver. This doesn't guarantee data durability, but we only use this setup for tests, so it's fine.	2021-12-24 15:32:54 +03:00
bojanserafimov	b807570f46	Use parking_lot::Mutex instead of std::Mutex in walreceiver (#1045 )	2021-12-23 14:25:44 -05:00
Kirill Bulatov	114a757d1c	Use generic config parameters in pageserver cli Co-authored-by: Heikki Linnakangas <heikki.linnakangas@iki.fi>	2021-12-23 18:58:28 +02:00
Heikki Linnakangas	fdd987c3ad	Refactor the way Image- and DeltaLayers are created Introduce builder objects, DeltaLayerWriter and ImageLayerWriter. This gives more flexibility, as the DeltaLayer::create and ImageLayer::create functions don't need to know about the details of the format of where the page versions are coming from. This allows us to change the format used in InMemoryLayer more easily, without having to modify Delta- and ImageLayer code. Also refactor the code in InMemoryLayer::write_to_disk for clarity.	2021-12-23 00:33:16 +02:00
Heikki Linnakangas	da62407fce	Change the meaning of 'blknum' argument in Layer trait Previously, the 'blknum' argument of various Layer functions was the block number within the overall relation. That was pretty confusing, because an individual layer only holds data from a one segment of the relation. Furthermore, the 'put_truncation' function already dealt with per-segment size, not overall relation size, adding to the confusion. Change the meaning of the 'blknum' argument to mean the block number within the segment, not the overall relation.	2021-12-22 16:55:37 +02:00
Heikki Linnakangas	1cc181ca32	Fix WAL redo of commit records with subtransactions. If a commit record contains XIDs that are stored on different CLOG pages, we duplicate the commit record for each affected CLOG page. In the redo routine, we must only apply the parts of the record that apply to the CLOG page being restored. We got that right in the loop that handles the sub-XIDs, but incorrectly always set the bit that corresponds to the main XID.	2021-12-21 23:08:01 +02:00
Heikki Linnakangas	bcf80eaa95	Fix multixacts members WAL redo. The logic to compute the page number was broken, and as a result, only the first page of multixact members was updated correctly. All the rest were left as zeros. Improve test_multixact.py to generate more multixacts, to cover this case. Also fix the check that the restored PG data directory matches the original one. Previously, the test compared the 'pg_new' cluster, which is a bit silly because the test restored the 'pg_new' cluster only a few lines earlier, so if the multixact WAL redo is somehow broken, the comparison will just compare two broken data directories and report success. Change it to compare the original datadir, the one where the multixacts were originally created, with a restored image of the same.	2021-12-21 17:50:06 +02:00
Konstantin Knizhnik	68aa9d2715	Set utf8 encoding in initdb (#993 ) refer #992	2021-12-21 15:43:34 +03:00
Konstantin Knizhnik	76777f5812	Add utility for dumping/editing metadata file (#1031 )	2021-12-21 15:43:15 +03:00
anastasia	90e5b6f983	Don't try to reconnect failed walreceiver. If necessary, wal service will send new callmemaybe request	2021-12-20 12:34:28 +03:00
Heikki Linnakangas	75cbaafb96	Remove old ephemeral files on pageserver restart. The ephemeral files are not usable after restart, so just delete them. Before this, you got "unrecognized filename in timeline dir" warnings about them, as Konstantin noted at: https://github.com/zenithdb/zenith/issues/906#issuecomment-995530870. While we're at it, refactor away the list_files() function, moving the logic fully into the caller. Seems more straightforward.	2021-12-17 00:00:02 +02:00
Heikki Linnakangas	d8a367dd32	Remove dead code, fix typos.	2021-12-15 19:58:03 +02:00
Kirill Bulatov	ca60561a01	Propagate disk consistent lsn in timeline sync statuses	2021-12-15 15:13:21 +02:00
Heikki Linnakangas	7f78e80c51	Refactor WAL ingestion code. Rename save_decoded_record() to ingest_record(), and move the responsibility for decoding the record into ingest_record(). Also move the responsibility of updating the CheckPoint relish to ingest_record(). Put it in a new WalIngest struct, to help with tracking that.	2021-12-14 20:24:03 +02:00
Heikki Linnakangas	f8f88154d5	Split restore_local_repo.rs into two files, with more descriptive names.	2021-12-14 20:24:03 +02:00
Kirill Bulatov	5cff7d1de9	Use proper download order	2021-12-14 15:32:22 +02:00
Heikki Linnakangas	e0d41ac6a3	Move constants related to metadata file to metadata.rs. They're not used anywhere else, so seems like a better place.	2021-12-13 16:57:16 +02:00
Heikki Linnakangas	72ef59c378	Fix small typos in comments, add a comment. The introducing paragraph README could use some more love, but let's at least fix the typos.	2021-12-13 13:44:08 +02:00
Kirill Bulatov	673c297949	Download timelines on demand	2021-12-10 17:23:35 +02:00
Kirill Bulatov	e61732ca7c	Compress checkpoint files before streaming into S3	2021-12-10 17:23:35 +02:00
Heikki Linnakangas	cb4a8396fb	Use rustls rather than native-tls in all dependencies. We depends on rustls in postgres_backend anyway, so might as well use it for all TLS stuff. Seems better to depend on only one library both from a security point of view, and because fewer dependencies means less code to compile. With this commit, we no longer depend on OpenSSL.	2021-12-10 15:14:27 +02:00

1 2 3 4 5 ...

575 Commits