MatrixSynapse

Commit Graph

Author	SHA1	Message	Date
Erik Johnston	8de3703d21	Make event persisters periodically announce position over replication. (#8499 ) Currently background proccesses stream the events stream use the "minimum persisted position" (i.e. `get_current_token()`) rather than the vector clock style tokens. This is broadly fine as it doesn't matter if the background processes lag a small amount. However, in extreme cases (i.e. SyTests) where we only write to one event persister the background processes will never make progress. This PR changes it so that the `MultiWriterIDGenerator` keeps the current position of a given instance as up to date as possible (i.e using the latest token it sees if its not in the process of persisting anything), and then periodically announces that over replication. This then allows the "minimum persisted position" to advance, albeit with a small lag.	2020-10-12 15:51:41 +01:00
Erik Johnston	6c5d5e507e	Add unit test for event persister sharding (#8433 )	2020-10-02 09:57:12 +01:00
Erik Johnston	04cc249b43	Add experimental support for sharding event persister. Again. (#8294 ) This is not ready for production yet. Caveats: 1. We should write some tests... 2. The stream token that we use for events can get stalled at the minimum position of all writers. This means that new events may not be processed and e.g. sent down sync streams if a writer isn't writing or is slow.	2020-09-14 10:16:41 +01:00
Brendan Abolivier	9f8abdcc38	Revert "Add experimental support for sharding event persister. (#8170 )" (#8242 ) * Revert "Add experimental support for sharding event persister. (#8170)" This reverts commit `82c1ee1c22`. * Changelog	2020-09-04 10:19:42 +01:00
Erik Johnston	82c1ee1c22	Add experimental support for sharding event persister. (#8170 ) This is not ready for production yet. Caveats: 1. We should write some tests... 2. The stream token that we use for events can get stalled at the minimum position of all writers. This means that new events may not be processed and e.g. sent down sync streams if a writer isn't writing or is slow.	2020-09-02 15:48:37 +01:00
Richard van der Hoff	f57b99af22	Handle replication commands synchronously where possible (#7876 ) Most of the stuff we do for replication commands can be done synchronously. There's no point spinning up background processes if we're not going to need them.	2020-07-27 18:54:43 +01:00
Richard van der Hoff	931b026844	Remove an unused prometheus metric (#7878 )	2020-07-22 00:40:55 +01:00
Richard van der Hoff	e5300063ed	Optimise queueing of inbound replication commands (#7861 ) When we get behind on replication, we tend to stack up background processes behind a linearizer. Bg processes are heavy (particularly with respect to prometheus metrics) and linearizers aren't terribly efficient once the queue gets long either. A better approach is to maintain a queue of requests to be processed, and nominate a single process to work its way through the queue. Fixes: #7444	2020-07-16 15:49:37 +01:00
Erik Johnston	f2e38ca867	Allow moving typing off master (#7869 )	2020-07-16 15:12:54 +01:00
Erik Johnston	f299441cc6	Add ability to shard the federation sender (#7798 )	2020-07-10 18:26:36 +01:00
Will Hunt	62b1ce8539	isort 5 compatibility (#7786 ) The CI appears to use the latest version of isort, which is a problem when isort gets a major version bump. Rather than try to pin the version, I've done the necessary to make isort5 happy with synapse.	2020-07-05 16:32:02 +01:00
Patrick Cloke	7d2532be36	Discard RDATA from already seen positions. (#7648 )	2020-06-15 08:44:54 -04:00
Erik Johnston	9bac5d62b3	Ensure ReplicationStreamer is always started when replication enabled. (#7579 ) Fixes #7566.	2020-05-27 11:44:19 +01:00
Erik Johnston	e5c67d04db	Add option to move event persistence off master (#7517 )	2020-05-22 16:11:35 +01:00
Erik Johnston	7ee24c5674	Have all instances correctly respond to REPLICATE command. (#7475 ) Before all streams were only written to from master, so only master needed to respond to `REPLICATE` commands. Before all instances wrote to the cache invalidation stream, but didn't respond to `REPLICATE`. This was a bug, which could lead to missed rows from cache invalidation stream if an instance is restarted, however all the caches would be empty in that case so it wasn't a problem.	2020-05-13 10:27:02 +01:00
Erik Johnston	8ca79613e6	Fix Redis reconnection logic (#7482 ) Proactively send out `POSITION` commands (as if we had just received a `REPLICATE`) when we connect to Redis. This is important as other instances won't notice we've connected to issue a `REPLICATE` command (unlike for direct TCP connections). This is only currently an issue if master process reconnects without restarting (if it restarts then it won't have written anything and so other instances probably won't have missed anything).	2020-05-13 09:57:15 +01:00
Andrew Morgan	5cf758cdd6	Merge branch 'release-v1.13.0' into develop * release-v1.13.0: Don't UPGRADE database rows RST indenting Put rollback instructions in upgrade notes Fix changelog typo Oh yeah, RST Absolute URL it is then Fix upgrade notes link Provide summary of upgrade issues in changelog. Fix ) Move next version notes from changelog to upgrade notes Changelog fixes 1.13.0rc1 Documentation on setting up redis (#7446) Rework UI Auth session validation for registration (#7455) Fix errors from malformed log line (#7454) Drop support for redis.dbid (#7450)	2020-05-11 16:46:33 +01:00
Richard van der Hoff	da9b2db3af	Drop support for redis.dbid (#7450 ) Since we only use pubsub, the dbid is irrelevant.	2020-05-07 16:46:15 +01:00
Erik Johnston	d7983b63a6	Support any process writing to cache invalidation stream. (#7436 )	2020-05-07 13:51:08 +01:00
Richard van der Hoff	a8c17da245	Merge branch 'release-v1.13.0' into rav/fix_dropped_messages	2020-05-05 23:01:12 +01:00
Richard van der Hoff	1242267316	Merge branch 'release-v1.13.0' into rav/fix_dropped_messages	2020-05-05 22:38:44 +01:00
Richard van der Hoff	7f7eedbebb	Wait for a POSITION on the right connection before accepting RDATA ... otherwise we can believe we're up to date when we're not.	2020-05-05 22:38:16 +01:00
Brendan Abolivier	5b8023dc7f	Move logs about discarded RDATA to debug (#7421 )	2020-05-05 21:07:33 +02:00
Richard van der Hoff	d78265af0c	Wait to subscribe before sending REPLICATE	2020-05-05 19:31:37 +01:00
Erik Johnston	0e719f2398	Thread through instance name to replication client. (#7369 ) For in memory streams when fetching updates on workers we need to query the source of the stream, which currently is hard coded to be master. This PR threads through the source instance we received via `POSITION` through to the update function in each stream, which can then be passed to the replication client for in memory streams.	2020-05-01 17:19:56 +01:00
Erik Johnston	3085cde577	Use `stream.current_token()` and remove `stream_positions()` (#7172 ) We move the processing of typing and federation replication traffic into their handlers so that `Stream.current_token()` points to a valid token. This allows us to remove `get_streams_to_replicate()` and `stream_positions()`.	2020-05-01 15:21:35 +01:00
Erik Johnston	37f6823f5b	Add instance name to RDATA/POSITION commands (#7364 ) This is primarily for allowing us to send those commands from workers, but for now simply allows us to ignore echoed RDATA/POSITION commands that we sent (we get echoes of sent commands when using redis). Currently we log a WARNING on the master process every time we receive an echoed RDATA.	2020-04-29 16:23:08 +01:00
Erik Johnston	3eab76ad43	Don't relay REMOTE_SERVER_UP cmds to same conn. (#7352 ) For direct TCP connections we need the master to relay REMOTE_SERVER_UP commands to the other connections so that all instances get notified about it. The old implementation just relayed to all connections, assuming that sending back to the original sender of the command was safe. This is not true for redis, where commands sent get echoed back to the sender, which was causing master to effectively infinite loop sending and then re-receiving REMOTE_SERVER_UP commands that it sent. The fix is to ensure that we only relay to other connections and not to the connection we received the notification from. Fixes #7334.	2020-04-29 14:10:59 +01:00
Richard van der Hoff	c2e1a2110f	Fix limit logic for EventsStream (#7358 ) * Factor out functions for injecting events into database I want to add some more flexibility to the tools for injecting events into the database, and I don't want to clutter up HomeserverTestCase with them, so let's factor them out to a new file. * Rework TestReplicationDataHandler This wasn't very easy to work with: the mock wrapping was largely superfluous, and it's useful to be able to inspect the received rows, and clear out the received list. * Fix AssertionErrors being thrown by EventsStream Part of the problem was that there was an off-by-one error in the assertion, but also the limit logic was too simple. Fix it all up and add some tests.	2020-04-29 12:30:36 +01:00
Richard van der Hoff	71a1abb8a1	Stop the master relaying USER_SYNC for other workers (#7318 ) Long story short: if we're handling presence on the current worker, we shouldn't be sending USER_SYNC commands over replication. In an attempt to figure out what is going on here, I ended up refactoring some bits of the presencehandler code, so the first 4 commits here are non-functional refactors to move this code slightly closer to sanity. (There's still plenty to do here :/). Suggest reviewing individual commits. Fixes (I hope) #7257.	2020-04-22 22:39:04 +01:00
Erik Johnston	51f7eaf908	Add ability to run replication protocol over redis. (#7040 ) This is configured via the `redis` config options.	2020-04-22 13:07:41 +01:00
Richard van der Hoff	0f8f02bc39	On catchup, process each row with its own stream id (#7286 ) Other parts of the code (such as the StreamChangeCache) assume that there will not be multiple changes with the same stream id. This code was introduced in #7024, and I hope this fixes #7206.	2020-04-20 11:43:29 +01:00
Richard van der Hoff	6a519a0ca0	Remove vestigal references to SYNC replication command We've ripped pretty much all of this out: let's remove the remains.	2020-04-07 17:40:07 +01:00
Erik Johnston	ce72355d7f	Fix race in replication (#7226 ) Fixes a race between handling `POSITION` and `RDATA` commands. We do this by simply linearizing handling of them.	2020-04-07 11:01:04 +01:00
Erik Johnston	82498ee901	Move server command handling out of TCP protocol (#7187 ) This completes the merging of server and client command processing.	2020-04-07 10:51:07 +01:00
Erik Johnston	5016b162fc	Move client command handling out of TCP protocol (#7185 ) The aim here is to move the command handling out of the TCP protocol classes and to also merge the client and server command handling (so that we can reuse them for redis protocol). This PR simply moves the client paths to the new `ReplicationCommandHandler`, a future PR will move the server paths too.	2020-04-06 09:58:42 +01:00

36 Commits (37eaf9c27271197320d4bedcceaf58d746935e53)