PostgreSQL 17 Failover Slots: Keeping Logical Subscribers Alive Through a Switchover
Our first planned switchover after adding logical replication broke the Debezium pipeline: the slot lived on the old primary, and the new primary knew nothing about it. PostgreSQL 17's failover slots sync the slot to the standby. The drill that used to cost a nine-hour resnapshot now costs four minutes.
The switchover itself was textbook: replica caught up, primary demoted, replica promoted, applications reconnected, total write downtime under ninety seconds. The pager went off forty minutes later, when the warehouse team noticed the CDC feed had gone silent. Debezium’s slot, debezium_orders, was a logical replication slot, and logical replication slots live on the exact instance that created them — the promoted replica had never heard of it. The connector sat there polling a slot that did not exist, and the fix was to create a new slot on the new primary and resnapshot the tables: nine hours of catch-up for a pipeline that was supposed to survive failover as its main selling point. We accepted the pain once, upgraded that cluster to PostgreSQL 17 over the winter, and rebuilt the setup around failover slots — slots the primary actively synchronizes to its standbys so a promotion brings the subscription state along with it.
What problem do failover slots actually solve?
Physical replication slots are a streaming-replication concern and follow the standby naturally; logical slots are a server-local catalog object, and before PostgreSQL 17 they simply did not exist anywhere except where they were created. Any failover or switchover meant every downstream logical consumer — Debezium, a logical subscriber, a homegrown pg_logical_slot_get_changes poller — either re-baselined from a snapshot or silently lost changes, and we have written before about what unmanaged slots do to disk on the primary in the retained-WAL disk-full notes. Failover slots change the contract: create the slot with failover = true, and the primary’s slot synchronization machinery keeps a matching slot current on each physical standby that opts in, advancing its confirmed_flush_lsn as the primary’s slot advances. When the standby is promoted, the slot is already there, already at a position that reflects what the subscriber had confirmed, and the consumer reconnects and resumes instead of restarting. The nine-hour resnapshot in our drill shrank to about four minutes of connector reconnection and offset validation, and the warehouse team stopped being paged for our database maintenance.
How do you set the machinery up?
The moving parts are three: the slot itself, the standby’s sync process, and a set of preconditions that are easy to miss. The slot is created on the primary with the failover flag, either at creation time with pg_create_logical_replication_slot(name, plugin, temporary, twophase, failover) or through SQL. The standby needs primary_slot_name pointing at a physical slot on the primary, hot_standby_feedback on so its own catalog state is protected, and — the piece that was new in 17 — a slotsync worker that wakes up periodically and reconciles the standby’s slots against the primary’s, which you can also run on demand:
-- On the primary: create the logical slot with failover support
SELECT *
FROM pg_create_logical_replication_slot(
'debezium_orders', 'pgoutput',
temporary := false, twophase := false, failover := true
);
-- On the standby: force a sync now instead of waiting for the worker
SELECT pg_sync_replication_slots();
-- Confirm the standby's copy exists and tracks the primary
SELECT slot_name, failover, synced, confirmed_flush_lsn
FROM pg_replication_slots;
The preconditions that bit us, in order: the wal_level must be logical everywhere in the chain, which for us meant an extra restart of the standby because it had been physical-only since provisioning; the standby must connect using a replication role that can read the primary’s slots, which interacted with the permission tightening from the default privileges notes; and pg_sync_replication_slots fails, by design, if the promotion would leave the synced slot ahead of confirmed data — the function is conservative about consistency, which is exactly what you want and exactly what will confuse you the first time it refuses.
What happens to the subscriber during the promotion?
Less than it used to, but not nothing, and the honest drill report matters here. The promoted standby has the slot at the position it had most recently synced, which can lag the primary’s slot by the sync interval — so the subscriber may see a small window of already-consumed changes re-delivered, which any at-least-once consumer must tolerate anyway and which Debezium handled via its normal offset reconciliation. What the subscriber does not do anymore is lose its place entirely. One sharp edge deserves its own sentence: synchronized_standby_slots, the companion setting that can make the primary wait until standbys have flushed the WAL a failover slot needs, closes the remaining gap but turns slot sync into a synchronous-availability dependency — we left it off, accepting seconds of re-delivery over a new way for a sick standby to stall primary commits. That trade-off is the same family as the durability levels in the synchronous commit notes: every guarantee you add is an availability you spend.
What did the drills look like after the rebuild?
We ran the switchover drill monthly for a quarter before trusting it, and the failure modes were all in the seam between sync and promotion rather than in the feature itself. Twice the standby’s synced slot had stalled because the slotsync worker had died quietly after a network partition and nobody was watching synced = false; once, a second logical slot someone had created by hand without failover = true confused a new on-call into thinking the sync had partially failed. Both produced the same runbook additions: monitor the lag between the primary’s confirmed_flush_lsn and the standby’s synced copy, alert on synced flipping false, and audit the slot list after any hand-created change. The promotion drill now ends with the warehouse feed catching up from a four-minute backlog instead of a nine-hour snapshot, and the difference between those two numbers is the entire argument for the feature.
Watching slot health with MonPG
Failover slots add a second slot timeline to care about — the standby’s synced copy — and the failure mode that matters is drift between the two plus the classic slot-retention disk pressure if any consumer stalls, the same pressure the replication monitoring notes track on the physical side. MonPG monitors PostgreSQL in production today, and its PostgreSQL monitoring puts replication slot age, retained WAL, and replica lag on one view, so a stalled slotsync worker shows up as a growing gap on a graph instead of a silent warehouse feed discovered by the analytics team at 09:00. Failover you have not drilled is a hypothesis; failover slots you do not monitor are the same hypothesis with better marketing.