An analytics engineer hooked up a Debezium CDC pipeline on staging, pointed it to production for a test, and shut down their laptop. Four hours later, 1.3TB of WAL filled the disk. Here is how replication slots retain WAL and…
Database Topic Archive
Replication and WAL Articles
Replica lag, WAL growth, failover readiness, hot standby behavior, and replication slots.
Migrated to 80,000 IOPS NVMe drives and still getting periodic latency spikes? Here is why PostgreSQL checkpoints cause latency sawteeth and how to tune max_wal_size, completion targets, and Linux dirty page writeback.
Our first planned switchover after adding logical replication broke the Debezium pipeline: the slot lived on the old primary, and the new primary knew nothing about it. PostgreSQL 17's failover slots sync the slot to the standby. The drill that…
The reporting subscriber was twenty-six hours behind, but the publisher swore replication was fine — and it was, because it was happily streaming WAL nobody could apply. One duplicate key had killed the apply worker on Tuesday, and the only…
A 40-minute analytics query on a PostgreSQL standby started dying with 'canceling statement due to conflict with recovery' every night at 02:00. MySQL replicas never cancel your query; they just fall behind. Here is the trade-off both engines make and…
One Debezium-style connector on a busy primary cost us a third of a core, twenty percent more WAL, and — when the consumer stalled over a long weekend — 92 GB of retained WAL. The full cost sheet, and the…
Both MySQL and PostgreSQL replicas serve stale data by default. What differs is the toolbox: GTID waits and group replication consistency levels versus LSN tokens and remote_apply.
full_page_writes protects you from torn pages but can dominate WAL volume. How checkpoints set the FPI rate, what wal_compression with lz4 or zstd buys you, and how to measure WAL with pg_stat_wal.
One word took our nightly staging load from 40 minutes to 9: UNLOGGED. Two months later a failover emptied the table and nobody on call could explain why. The semantics are absolute — if someone tells you.
A decommissioned standby left its replication slot behind, and over one quiet weekend the slot pinned 214 GB of WAL until the primary ran out of disk and PANIC'd. Here is the mechanism, the monitoring queries, and the circuit breaker…