fix: failure in crash recovery in suggestion rematching - #1309
Merged
Merged
Conversation
Due to how `pgpubsub` works, a listener would lock all notifications in the recovery phase on startup. If it crashes, all notifications are put back to the queue, and only picked up again on next startup. This would result in old notifications never getting processed. In our case, rematching only needs to get done against the latest evaluation. Therefore we can skip recovery entirely, since a new evaluation should appear every day anyway. Then we only need garbage-collect the dangling notifications from crashed listeners once in a while.
Collaborator
Author
|
Tested this locally against a live database, pretty confident it works. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This one was really hard to figure out, because the logs weren't very informative and that particular process wasn't leaving a lot of traces in the database...
Due to how
pgpubsubworks, a listener would lock all notifications in the recovery phase on startup.If it crashes, all notifications are put back to the queue, and only picked up again on next startup.
This would result in old notifications never getting processed.
In our case, rematching only needs to get done against the latest evaluation.
Therefore we can skip recovery entirely, since a new evaluation should appear every day anyway.
Then we only need garbage-collect the dangling notifications from crashed listeners once in a while.
Closes #618 (finally)