Autovacuum Wraparound Emergency: What to Do When the Cluster is at Risk
Wraparound emergencies are preventable, but once warnings start you need a calm runbook: identify old XIDs, unblock vacuum, freeze priority tables, and protect availability.
Notes for the problems that show up after launch: bad plans, awkward migrations, index debt, vacuum pressure, replica lag, and the small decisions that keep PostgreSQL, MySQL, MariaDB, and SQL Server easier to operate.
Find the queries dragging your Postgres down, read their plans, and fix them in order of impact.
Read the guide →Which indexes to add, which to drop, and how to spot redundant index debt before it costs you writes.
Read the guide →How autovacuum works, when it falls behind, and how to keep bloat from eating your storage and your buffer cache.
Read the guide →Sizing pools, PgBouncer modes, and how to stop idle-in-transaction sessions from exhausting max_connections.
Read the guide →Track lag, WAL throughput, and replication slots before they turn into stale reads or a full disk.
Read the guide →
Backups, restores, PITR, archive gaps, and recovery drills for PostgreSQL.
AWS RDS, Azure Database for PostgreSQL, Google Cloud SQL, AlloyDB, and managed PostgreSQL monitoring notes.
Connection storms, pool saturation, PgBouncer behavior, and max_connections pressure.
CPU, memory, disk, temp files, checkpoints, buffers, and workload pressure.
Index design, index debt, reindexing, constraint indexes, and evidence-driven DDL.
Lock waits, deadlocks, blocking chains, idle transactions, and transaction hygiene.
Galera clustering, InnoDB and Aria internals, replication, optimizer behavior, and MariaDB operations notes.
InnoDB internals, replication, locking, query performance, and MySQL operations notes.
HNSW, IVFFlat, recall, embedding search, and production RAG performance notes.
Replica lag, WAL growth, failover readiness, hot standby behavior, and replication slots.
Query plans, EXPLAIN ANALYZE, planner regressions, pagination, joins, and statistics.
Wait statistics, tempdb, Query Store, Always On availability groups, and SQL Server operations notes.
Autovacuum, dead tuples, bloat, wraparound, MVCC, freeze age, and visibility.
Wraparound emergencies are preventable, but once warnings start you need a calm runbook: identify old XIDs, unblock vacuum, freeze priority tables, and protect availability.
Plan regressions are painful because the SQL did not change. The work is proving the plan changed, finding the estimate or stats shift, and restoring the safe path.
psql is more capable than people use it for. A handful of meta-commands and shortcuts make it the most productive shell for database work.
pg_stat_statements is cumulative since cluster start or last reset. If you reset it, you lose the data you would have used to debug the next incident.
A useful Postgres health check is not a wall of green checks. It is a short path from symptom to evidence: sessions, locks, slow SQL, vacuum, replication, and WAL.
auto_explain is for the slow plan you cannot reproduce later. It captures the execution plan when the bad thing actually happens.
pgbench measures Postgres throughput under a synthetic workload. It tells you something useful, but only if you understand what its numbers mean.
Production migration review is mostly lock review. The SQL can be correct and still dangerous if it rewrites a table, validates too much, or blocks writes.
JSONB columns are great. JSONB indexing has a steeper learning curve than the docs admit. Here is what works in production.