rust/neon - neon - Gitea: Git with a cup of tea

rust/neon

mirror of https://github.com/neondatabase/neon.git synced 2025-12-25 23:29:59 +00:00

Author	SHA1	Message	Date
Bojan Serafimov	9d6b78861d	WIP	2022-01-11 12:06:32 -05:00
bojanserafimov	5b9391b51d	Support "query cancel" in proxy (#1052 )	2022-01-05 17:27:12 -05:00
Arthur Petukhovsky	5a6405848d	Bump vendor/postgres (#1086 )	2022-01-05 14:27:51 +03:00
Patrick Insinger	191d9d2b74	par_fsync - use VirtualFile	2022-01-04 20:40:57 -08:00
Patrick Insinger	24c8dab86f	pageserver - parallelize checkpoint fsyncs	2022-01-04 20:40:57 -08:00
Heikki Linnakangas	55a4cf64a1	Refactor WAL record handling. Introduce the concept of a "ZenithWalRecord", which can be a Postgres WAL record that is replayed with the Postgres WAL redo process, or a built-in type that is handled entirely by pageserver code. Replace the special code to replay Postgres XACT commit/abort records with new Zenith WAL records. A separate zenith WAL record is created for each modified CLOG page. This allows removing the 'main_data_offset' field from stored PostgreSQL WAL records, which saves some memory and some disk space in delta layers. Introduce zenith WAL records for updating bits in the visibility map. Previously, when e.g. a heap insert cleared the VM bit, we duplicated the heap insert WAL record for the affected VM page. That was very wasteful. The heap WAL record could be massive, containing a full page image in the worst case. This addresses github issue #941.	2022-01-04 11:26:37 +02:00
Heikki Linnakangas	722667f189	Add test case for performance issue #941 . The first COPY generates about 230 MB of write I/O, but the second COPY, after deleting most of the rows and vacuuming the rows away, generates 370 MB of writes. Both COPYs insert the same amount of data, so they should generate roughly the same amount of I/O. This commit doesn't try to fix the issue, just adds a test case to demonstrate it. Add a new 'checkpoint' command to the pageserver API. Previously, we've used 'do_gc' for that, but many tests, including this new one, really only want to perform a checkpoint and don't care about GC. For now, I only used the command in the new test, though, and didn't convert any existing tests to use it.	2022-01-04 11:26:37 +02:00
Arseny Sher	25a515b968	Don't call immediately on resume in callmemaybe. It creates busy loop if pageserver <-> safekeeper connection fails after it was established (e.g. currently due to 'segment checkpoint not found' error on pageserver). Also wake up callmemaybe thread regularly once in recall_period regardless of channel activity.	2022-01-03 20:44:36 +03:00
Konstantin Knizhnik	1c47fbae81	Do not write image layers during enforced checkpoint (#1057 ) * Do not write image layers during enforced checkpoint refer #1056 * Add Flush option to CheckpointConfig refer #1057	2022-01-01 19:08:09 +03:00
Alexey Kondratov	8f0cd7fb9f	[compute_tools] Switch cluster_id in spec to string (zenithdb/console#72 )	2021-12-29 16:35:29 +03:00
Dmitry Rodionov	c910132d4b	Fix wal receiver shutdown This patch allows to shutdown wal receiver when there are no messages and wal receiver is blocked inside tokio-postgres. In this case it cannot check the shutdown flag. This patch switches to use async interface of tokio-postgres directly without sync wrappers. It opens the possibility to use tokio::select! between the phsycal_stream.next() and a shutdown channel readiness to interrupt replication process. Also this allows to shutdown only particular wal receiver without using global shutdown_requested flag.	2021-12-29 14:42:29 +03:00
Arthur Petukhovsky	70778058d9	Add test for safekeeper setup without pageserver (#1000 )	2021-12-29 12:58:27 +03:00
nikitashamgunov	a379b45257	Update README.md	2021-12-28 14:26:42 -08:00
bojanserafimov	24eca8d58b	Parse cancel message in pq_proto (#1060 )	2021-12-28 16:43:44 -05:00
Bojan Serafimov	1e3ddd43bc	Add struct for key data	2021-12-28 22:40:22 +03:00
Bojan Serafimov	989371493b	Add BeMessage::BackendKeyData variant	2021-12-28 22:40:22 +03:00
Alexey Kondratov	f64074c609	Move compute_tools from console repo (zenithdb/console#383 ) Currently it's included with minimal changes and lives aside of the main workspace. Later we may re-use and combine common parts with zenith control_plane. This change is mostly needed to unify cloud deployment pipeline: 1.1. build compute-tools image 1.2. build compute-node image based on the freshly built compute-tools 2. build zenith image So we can roll new compute image and new storage required by it to operate properly. Also it becomes easier to test console against some specific version of compute-node/-tools.	2021-12-28 20:17:29 +03:00
anastasia	eba897ffe7	send CallmeEvent::Unsubscribe request only when pageserver is caught up with safekeeper and it's time to stop streaming	2021-12-28 17:50:48 +03:00
anastasia	5ef2b1baf7	Add new test illustrating issue with sync-safekeepers. If safekeepers sync fast enough, callmemaybe thread may never make a call before receiving Unsubscribe request. This leads to the situation, when pageserver lacks data that exists on safekeepers.	2021-12-28 17:50:48 +03:00
Kirill Bulatov	f0afd08667	Fix zenith init defaults	2021-12-28 00:21:48 +02:00
Kirill Bulatov	b494ac1ea0	Remove redundant pageserver cli params	2021-12-27 18:38:54 +02:00
Arseny Sher	a163650a99	Refactor Postgres command parsing in safekeeper. Do it separately with SafekeeperPostgresCommand enum as a result. Since query is always C string, switch postgres_backend process_query argument from Bytes to &str. Make passing ztli/ztenant id in safekeeper connection string optional; this is needed for upcoming intra-safekeeper heartbeat cmd which is not bound to any timeline.	2021-12-24 15:48:13 +03:00
anastasia	980f5f8440	Propagate remote_consistent_lsn to safekeepers. Change meaning of lsns in HOT_STANDBY_FEEDBACK: flush_lsn = disk_consistent_lsn, apply_lsn = remote_consistent_lsn Update compute node backpressure configuration respectively. Update compute node configuration: set 'synchronous_commit=remote_write' in setup without safekeepers. This way compute node doesn't have to wait for data checkpoint on pageserver. This doesn't guarantee data durability, but we only use this setup for tests, so it's fine.	2021-12-24 15:32:54 +03:00
Kirill Bulatov	42647f606e	Use correct pageserver CLI parameters in docker entrypoint	2021-12-24 03:41:45 +02:00
bojanserafimov	b807570f46	Use parking_lot::Mutex instead of std::Mutex in walreceiver (#1045 )	2021-12-23 14:25:44 -05:00
Kirill Bulatov	114a757d1c	Use generic config parameters in pageserver cli Co-authored-by: Heikki Linnakangas <heikki.linnakangas@iki.fi>	2021-12-23 18:58:28 +02:00
Andrey Taranik	9854ded56b	Feature/proxy deploy (#1046 ) * zenith proxy deployment * proxy deploy ci fix * ci cleanup or zenith proxy deploy	2021-12-23 15:53:28 +03:00
Heikki Linnakangas	fdd987c3ad	Refactor the way Image- and DeltaLayers are created Introduce builder objects, DeltaLayerWriter and ImageLayerWriter. This gives more flexibility, as the DeltaLayer::create and ImageLayer::create functions don't need to know about the details of the format of where the page versions are coming from. This allows us to change the format used in InMemoryLayer more easily, without having to modify Delta- and ImageLayer code. Also refactor the code in InMemoryLayer::write_to_disk for clarity.	2021-12-23 00:33:16 +02:00
Heikki Linnakangas	da62407fce	Change the meaning of 'blknum' argument in Layer trait Previously, the 'blknum' argument of various Layer functions was the block number within the overall relation. That was pretty confusing, because an individual layer only holds data from a one segment of the relation. Furthermore, the 'put_truncation' function already dealt with per-segment size, not overall relation size, adding to the confusion. Change the meaning of the 'blknum' argument to mean the block number within the segment, not the overall relation.	2021-12-22 16:55:37 +02:00
Heikki Linnakangas	1cc181ca32	Fix WAL redo of commit records with subtransactions. If a commit record contains XIDs that are stored on different CLOG pages, we duplicate the commit record for each affected CLOG page. In the redo routine, we must only apply the parts of the record that apply to the CLOG page being restored. We got that right in the loop that handles the sub-XIDs, but incorrectly always set the bit that corresponds to the main XID.	2021-12-21 23:08:01 +02:00
Heikki Linnakangas	927587cec8	Fix comments in tests	2021-12-21 22:38:33 +02:00
Heikki Linnakangas	bcf80eaa95	Fix multixacts members WAL redo. The logic to compute the page number was broken, and as a result, only the first page of multixact members was updated correctly. All the rest were left as zeros. Improve test_multixact.py to generate more multixacts, to cover this case. Also fix the check that the restored PG data directory matches the original one. Previously, the test compared the 'pg_new' cluster, which is a bit silly because the test restored the 'pg_new' cluster only a few lines earlier, so if the multixact WAL redo is somehow broken, the comparison will just compare two broken data directories and report success. Change it to compare the original datadir, the one where the multixacts were originally created, with a restored image of the same.	2021-12-21 17:50:06 +02:00
Arthur Petukhovsky	f56db3da68	Bump vendor/postgres (#996 )	2021-12-21 16:53:08 +03:00
Konstantin Knizhnik	68aa9d2715	Set utf8 encoding in initdb (#993 ) refer #992	2021-12-21 15:43:34 +03:00
Konstantin Knizhnik	76777f5812	Add utility for dumping/editing metadata file (#1031 )	2021-12-21 15:43:15 +03:00
Arseny Sher	56312522f9	Make safekeeper namings more consistent with reality. s/send_wal.rs/handler.rs s/SendWalHandler/SafekeeperPostgresHandler s/replication.rs/send_wal.rs	2021-12-21 13:24:23 +03:00
Dmitry Rodionov	2d9d0658e8	adjust benchmarking script for go console	2021-12-20 13:54:10 +03:00
anastasia	3b61f364f7	Stop WAL streaming threads, when compute node is shut down. WAL stream uses the 2 connections: 1. Compute node (walproposer) -> Safekeeper (ReceiveWalConn module) When compute node is shut down, safekeeper needs to stop the respective receiving thread. Prior to this PR it didn't work because PostgresBackend haven't handled disconnection properly. 2. Safekeeper (ReplicationConn module) -> pageserver (walreceiver thread) When incoming WAL stream is gone, safekeeper can stop streaming WAL and cancel connection as soon as replica is caught up. Note that the WAL can be streamed to multiple replicas simultaneously, only disconnect ones that are caught up to the last_recieved_lsn.	2021-12-20 12:34:28 +03:00
anastasia	90e5b6f983	Don't try to reconnect failed walreceiver. If necessary, wal service will send new callmemaybe request	2021-12-20 12:34:28 +03:00
Heikki Linnakangas	75cbaafb96	Remove old ephemeral files on pageserver restart. The ephemeral files are not usable after restart, so just delete them. Before this, you got "unrecognized filename in timeline dir" warnings about them, as Konstantin noted at: https://github.com/zenithdb/zenith/issues/906#issuecomment-995530870. While we're at it, refactor away the list_files() function, moving the logic fully into the caller. Seems more straightforward.	2021-12-17 00:00:02 +02:00
Andrey Taranik	5d5c2738a6	staging deployment flow fix (#1029 )	2021-12-16 22:54:01 +03:00
Andrey Taranik	cbe155ff48	storage CI flow for staging environment (#1003 ) * storage CI flow for staging environment * prevent deploy version older than already deployed	2021-12-16 17:05:20 +03:00
Kirill Bulatov	29143b018e	Disable rustc incremental compilation to avoid ICEs	2021-12-15 21:57:34 +03:00
Heikki Linnakangas	d8a367dd32	Remove dead code, fix typos.	2021-12-15 19:58:03 +02:00
Kirill Bulatov	ca60561a01	Propagate disk consistent lsn in timeline sync statuses	2021-12-15 15:13:21 +02:00
Andrey Taranik	86a409a174	cleanup circleci config after test	2021-12-15 16:08:31 +03:00
Andrey Taranik	66242f0d0e	tag docker image by commit sha and add docker build for compute	2021-12-15 16:08:31 +03:00
Heikki Linnakangas	7f78e80c51	Refactor WAL ingestion code. Rename save_decoded_record() to ingest_record(), and move the responsibility for decoding the record into ingest_record(). Also move the responsibility of updating the CheckPoint relish to ingest_record(). Put it in a new WalIngest struct, to help with tracking that.	2021-12-14 20:24:03 +02:00
Heikki Linnakangas	f8f88154d5	Split restore_local_repo.rs into two files, with more descriptive names.	2021-12-14 20:24:03 +02:00
Kirill Bulatov	5cff7d1de9	Use proper download order	2021-12-14 15:32:22 +02:00

1 2 3 4 5 ...

1154 Commits