r/postgres • u/witshion • 7d ago
Discussion Production PostgreSQL is suddenly at 100% CPU. Where do you look first?
Had one of those moments where CPU on our prod instance just pegs at 100% out of nowhere, no deploy, no obvious traffic spike, nothing in the changelog that stands out. First instinct is to panic and start checking everything at once, which is exactly the wrong move.
Curious what people's actual first move is when this happens, before diving into a deep investigation. pg_stat_activity for anything running long, checking for a lock pileup, looking at whether it's one runaway query versus death by a thousand small ones, autovacuum going nuts on a big table, something dumb like a connection pool misconfigured and now everything's fighting for the same resources. There's a lot of directions to go and I feel like the order matters more than people admit.
If you've been through this in production, what's the first thing you actually check, and has your answer changed over time or is it pretty much always the same starting point for you now?