MonPG Engineering avatar MonPG Engineering Engineering Team Replication and WAL 3 min read

PostgreSQL Replication Slots: How an Inactive Slot Filled a 2TB Disk in Four Hours

An analytics engineer hooked up a Debezium CDC pipeline on staging, pointed it to production for a test, and shut down their laptop. Four hours later, 1.3TB of WAL filled the disk. Here is how replication slots retain WAL and how to safeguard against disk exhaustion.

Replication and WAL

At 3:45 AM, an emergency high-priority pager woke me up: CRITICAL: PostgreSQL Primary Disk Usage > 95%. Four hours earlier, the primary database disk had been sitting comfortably at 32% of its 2TB capacity.

In four hours, over 1.3 terabytes of Write-Ahead Logs had accumulated in pg_wal. When we checked our standby replicas, replication lag was under 200 milliseconds. We checked max_wal_size—it was configured at 32GB. There were no long-running transactions holding back vacuum.

The culprit was an abandoned logical replication slot. An analytics engineer had tested an Apache Kafka Debezium connector from their local machine, connected it to production, and closed their laptop lid. The replication slot remained active in PostgreSQL, silently forbidding the engine from recycling a single byte of WAL.

How replication slots manage WAL retention

Normally, PostgreSQL recycles old WAL files during checkpoints once they are no longer needed for crash recovery or standby replicas. But when you create a physical or logical replication slot, PostgreSQL makes a strict guarantee to the subscriber: it will never delete any WAL segment that the subscriber has not yet acknowledged.

The slot tracks this progress using restart_lsn. As long as the subscriber is consuming changes and sending back acknowledgment flushes, restart_lsn advances, and PostgreSQL safely recycles older WAL segments.

However, if the subscriber crashes, disconnects, or slows down, restart_lsn freezes. From that moment forward, every single transaction executed on the primary produces WAL files that must be retained on disk indefinitely.

Monitoring replication slot lag in bytes

Do not rely solely on replication lag in seconds. A subscriber might report 0 seconds of lag simply because it is completely disconnected. You must monitor replication slot lag in physical bytes:

-- Checking replication slot status and byte lag from current WAL position
SELECT 
  slot_name,
  plugin,
  slot_type,
  active,
  pg_size_pretty(pg_wal_lsn_diff(pg_current_wal_lsn(), restart_lsn)) AS retained_wal_bytes,
  restart_lsn
FROM pg_replication_slots;

If active is false, the subscriber is disconnected. If retained_wal_bytes starts climbing into gigabytes, you are on a direct path to an out-of-disk emergency.

The safety valve: max_slot_wal_keep_size

In PostgreSQL 13, the core team introduced the ultimate defense against disk exhaustion: max_slot_wal_keep_size.

This parameter sets the maximum amount of WAL files that replication slots are allowed to retain. If a replication slot lags behind by more than this limit, PostgreSQL automatically marks the slot as invalid (wal_status = ‘lost’) and resumes recycling WAL segments to protect disk space.

The disconnected subscriber will fail and need to be re-initialized from a snapshot, but your primary database will remain online and healthy.

-- Setting a hard ceiling on WAL retention for replication slots
ALTER SYSTEM SET max_slot_wal_keep_size = '64GB';
SELECT pg_reload_conf();

Emergency recovery when disk hits 100%

If PostgreSQL halts because the disk is completely full, you cannot easily log in and run standard queries. Never run rm pg_wal/* manually; deleting active WAL files corrupts the database state.

Identify the offending inactive slot by name and drop it via psql:

-- Dropping an abandoned slot to immediately free WAL disk space
SELECT pg_drop_replication_slot('abandoned_debezium_slot');
CHECKPOINT;

The practical standard

High-level database architecture is not about drawn boxes on an infrastructure diagram. It is about how the engine manages shared resources under concurrency — memory, latches, write-ahead logs, and lock tables. When things break at 2 AM, the fix is rarely adding another replica or throwing more CPU at the host. The fix is understanding the underlying resource bottleneck, measuring the exact wait event, and applying the architectural constraint that makes the system predictable.