MonPG Engineering avatar MonPG Engineering Engineering Team 7 min read

pg_rewind: Rejoining a Failed PostgreSQL Primary Without a Full Rebuild

The failover took ninety seconds. Bringing the old primary back as a standby was the expensive part — rsync estimated four hours for 1.1 TB. pg_rewind walked the WAL history to the divergence point and copied only the blocks that had actually changed: six minutes, 3.8 GB transferred.

The failover itself was the easy part. The primary’s host lost its storage at 02:11, the promote happened at 02:12, applications were writing against the new primary by 02:13, and by 03:00 the failed host was back with a healthy-looking PostgreSQL data directory on it. That healthy-looking directory was the problem: it contained a cluster that had accepted writes on the old timeline for a few seconds after the new primary had already been promoted, and a diverged timeline cannot simply be pointed at the new primary as a standby. The standard answer is to rebuild it — pg_basebackup or rsync, four hours for our 1.1 TB cluster, plus four hours of running without a standby. The better answer, which we had never actually rehearsed, was pg_rewind: it found the point of divergence in the WAL history and copied only the blocks that differed afterward — 3.8 GB, six minutes, standby attached before the incident review started. This is what the tool does, what it demands in advance, and where it cannot help.

Why can’t the old primary just rejoin as a standby?

Because PostgreSQL replication follows WAL timelines, and a timeline is a one-way street. When the replica was promoted it created a new timeline and began writing WAL that the old primary knows nothing about; meanwhile the old primary, before it noticed it was dead, wrote a handful of its own WAL records on the old timeline — a few committed transactions in our case, made by an application server whose retry logic outlived the failover by about four seconds. Those seconds are enough. The data files now disagree with the new timeline not just in WAL position but in content: pages modified by those zombie transactions diverge from the pages the new primary produced. A standby works by replaying its primary’s WAL onto data files that are byte-consistent with it, and the old primary’s files no longer qualify. Simply starting it with a primary_conninfo does not merge the two histories — it either errors on the timeline mismatch or, worse configuration by worse configuration, replays new WAL onto subtly wrong pages. The divergence has to be resolved physically, block by block, and the only choice is whether you copy everything or copy what changed.

How does pg_rewind decide what to copy?

By reading WAL from the point of divergence forward, on the old cluster, and treating every block a WAL record touched as suspect. Concretely: pg_rewind connects to the running new primary (or reads a source data directory), walks the timeline history to find the last common checkpoint the two clusters share, then scans the old cluster’s WAL from that point to its end, collecting every block reference. Those blocks — and only those blocks — are copied from the source, after which the target’s file map is reconciled and the configuration adjusted so the cluster starts as a standby of the source. The elegance is that the copy set is driven by write activity, not by comparison: there is no full-file diff, no checksum walk of 1.1 TB. Our four zombie seconds produced a change list of about 3.8 GB because checkpoint activity had dirtied shared buffers across hot tables; on a quieter primary the same incident rewinds in tens of megabytes. The one hard prerequisite hides in that WAL scan: to know which blocks a record touched, pg_rewind needs block references in WAL, which come from either full_page_writes with the standard logging or, critically for clusters that changed defaults, wal_log_hints = on — or data checksums enabled, which imply them. Our cluster had checksums from day one for the reasons in the checksums notes, and we discovered, reading the docs at 03:00, that this was the decision that made pg_rewind available at all. The knob cannot be enabled without a restart, so it belongs in the provisioning template, not the incident runbook.

What does the procedure look like end to end?

Short, scripted, and unforgiving about order. The old cluster must be shut down cleanly — pg_rewind needs a consistent shutdown state; a crashed target has to be started in single-user-ish recovery mode and stopped cleanly first, or pg_rewind runs its –restore-target-wal path against an archive — then the tool runs against the source, then the node rejoins:

-- On the old primary's host (as the postgres OS user):
--   pg_ctl stop -m fast            -- must end in a clean shutdown
--
--   pg_rewind --target-pgdata=/var/lib/postgresql/data --     --source-server="host=new-primary port=5432 user=rewind" --     --progress
--
-- Then write standby.signal + primary_conninfo and start.
-- Verify the rejoin from the new primary:
SELECT client_addr, state, sync_state,
       pg_wal_lsn_diff(pg_current_wal_lsn(), replay_lsn) AS replay_gap
FROM pg_stat_replication;

Two operational details earned their own runbook lines. The rewind source connection needs a role that can read the timeline history — we use a dedicated rewind role provisioned in advance, because creating credentials during an incident is how typos get into pg_authid. And after the rewind, the rejoined node’s first startup replays WAL from the divergence point forward, so its catch-up time depends on how long it was down — a node rewound after six hours of absence still has six hours of WAL to replay, the same replay-time physics the read replica lag notes measure from the steady-state side. Rewind is fast; catch-up is honest.

When can pg_rewind not save you?

The list is short but each entry is fatal to the plan, and knowing it in advance is what makes the tool a strategy instead of a hope. If the old cluster’s WAL no longer reaches back to the point of divergence — recycled by a long gap between failover and rewind, the same retention physics as the slot retention notes viewed from the other end — pg_rewind cannot build its change list and you rebuild from base backup. If the old data directory took actual damage in whatever killed it — our 02:11 storage event could easily have been that story — rewinding copies a healthy structure onto rotten blocks and you find out later, so a suspect disk means rebuild, not rewind. If the failover tooling used pg_promote with the old primary still writable for minutes rather than seconds, the change set grows toward the size of a full copy and the time advantage evaporates. And pg_rewind never substitutes for fixing why the failover happened — it restores redundancy, not root cause. Our incident review ended with wal_log_hints confirmed on every provisioned cluster, a monthly rewind drill against staging, and a standing rule: rebuilds are for doubt, rewinds are for the well-understood divergence we drill.

Watching replication topology with MonPG

The whole incident was a topology problem: one node diverged, one node promoted, redundancy gone until the pair was whole again. MonPG monitors PostgreSQL in production today, and its PostgreSQL monitoring tracks replication state and replay lag per standby alongside the WAL generation that determines both rewind scope and catch-up time, so the hour after a failover shows the rejoined node’s replay gap closing on a graph instead of someone polling pg_stat_replication by hand. Failover makes you a one-node cluster; the rewind is how fast you stop being one.

Related documentation