Hi Team, today’s PostgreSQL incident call was quite intense, with several questions from the application team due to a similar Sev 4 impact in May. As no PostgreSQL SRE was available at that time, I joined the call, reviewed the issue and shared the resolution within around 13 minutes.
The alert had not been generated, which added to the concern. I have shared the possible causes, relevant checks and observations over email, with Ranjit, Khozema and the relevant SREs included.
Following the discussion, a handful of SREs from other regions are now helping to analyse the condition further and explore how it can be included as part of the alerting. This also highlights the need for broader PostgreSQL SRE ownership so that incidents and follow-up actions do not continue to depend mainly on the existing SMEs.





0 comments:
Post a Comment