MariaDB Warm Restarts: Buffer Pool Dump and Load, and the 40-Minute Hangover
The kernel patch reboot took ninety seconds. The hangover took forty minutes: p99 latency eight times normal while InnoDB re-read its working set from disk, one cold page at a time. The fix had been compiled into the server the whole time.
The maintenance itself was textbook: kernel security patch, clean shutdown, reboot, MariaDB back up and accepting connections ninety seconds after going down. Then the hangover started. For the next forty minutes, p99 latency ran at eight times normal, disk reads pinned near the ceiling of what the volume could do, and the application team watched error budgets burn while the database — up, healthy, replicating — served every request the slow way. The explanation is one sentence: the server restarted with a 64 GB buffer pool containing nothing, and InnoDB re-fetched its working set from storage one 16 KB page at a time, in the order traffic happened to ask for it. The fix is two configuration lines that had been compiled into the server the entire time, dumping the buffer pool’s page map at shutdown and reloading it in the background at startup. Our restart now has a two-minute hangover instead of a forty-minute one, and the difference cost us nothing but an afternoon of testing.
Buffer pool dump and load is the most disproportionately valuable feature in MariaDB relative to how rarely I see it configured deliberately. These are the notes from making warm restarts the default on our fleet.
What does the dump and load machinery actually do?
The dump does not copy the buffer pool — 64 GB of pages would take far too long — it records a map: the tablespace and page number of the most recently used pages, a few bytes each, written to a small file named ib_buffer_pool in the datadir by default. At startup, InnoDB reads that map and fetches those pages back from the tablespaces in the background, restoring the cache’s recent-working-set shape while the server is already up and serving. Two variables switch the behavior on at the boundaries, and two more let you drive it by hand:
-- the two switches that make restarts warm
SELECT @@innodb_buffer_pool_dump_at_shutdown,
@@innodb_buffer_pool_load_at_startup,
@@innodb_buffer_pool_dump_pct,
@@innodb_buffer_pool_filename;
-- drive it by hand: snapshot now, reload now
SET GLOBAL innodb_buffer_pool_dump_now = ON;
SET GLOBAL innodb_buffer_pool_load_now = ON;
-- watch a load progress (10.5+)
SHOW STATUS LIKE 'Innodb_buffer_pool_load_status';
Three details separate a working setup from a cargo-culted one. First, innodb_buffer_pool_dump_pct defaults to 25 — the dump records only the hottest quarter of the pool, which is almost always right, because the coldest three quarters are churn that is cheap to re-read and expensive to reload; raising it toward 50 makes sense only on write-heavy servers whose hot set genuinely sprawls. Second, the load is asynchronous by design: the server accepts connections immediately and warms in the background, so your readiness probe will say healthy while the cache is still cold — plan for degraded-but-functional, not for a gate. Third, the dump file is only as fresh as the last dump: the at-shutdown switch handles planned restarts, but a crash leaves you with whatever the last scheduled dump captured, which is why we also trigger innodb_buffer_pool_dump_now from a cron job every few hours on the largest pools.
How do you make it work in practice?
The operational checklist is short and every item on it was learned the hard way. Verify the dump actually happens before you need it: after enabling dump_at_shutdown, do a test restart and confirm ib_buffer_pool exists in the datadir with a fresh timestamp and plausible size — tens of megabytes for a large pool — because a shutdown that runs out of time or crashes mid-dump leaves you a stale or absent file and a cold start you thought was warm. Point innodb_buffer_pool_filename at a fast, always-mounted path if your datadir layout is exotic; the file is small but it must be readable at the earliest moment of startup. Give the load I/O headroom: the reload is a burst of random reads competing with live traffic’s reads, so on saturated storage we schedule restarts for traffic troughs even though the load is background. And know the abort lever — if a load turns out to be hurting more than helping, because the dump was stale or the working set shifted while the server was down, you can stop it mid-flight:
-- the load is hurting more than helping? stop it
SET GLOBAL innodb_buffer_pool_load_abort = ON;
-- how warm is the pool right now, honestly?
SHOW STATUS LIKE 'Innodb_buffer_pool_pages_total';
SHOW STATUS LIKE 'Innodb_buffer_pool_pages_data';
-- data/total ratio tells you fill; the read-hit rate
-- tells you warmth:
SHOW GLOBAL STATUS LIKE 'Innodb_buffer_pool_read_requests';
SHOW GLOBAL STATUS LIKE 'Innodb_buffer_pool_reads';
-- reads / read_requests is your miss rate; watch it decay
-- over the minutes after startup
The miss-rate pair is the honest warmth metric and it belongs on your restart runbook’s verification step: Innodb_buffer_pool_reads counts fetches that had to go to disk, Innodb_buffer_pool_read_requests counts all logical reads, and the ratio collapsing from double-digit percentages back under one percent is what recovered actually looks like. Up means nothing; warm is the finish line.
When is a warm restart not enough?
Dump and load shrinks the hangover but does not abolish physics, and three scenarios still need real planning. Failovers: the replica you promote has been warming its buffer pool around its replay workload, not the primary’s read workload, so its cache is warm with the wrong pages — we now dump-load on the candidate before planned switchovers and accept a measured degradation window on unplanned ones. Restores: a server rebuilt from backup has no dump at all, its pool starts empty by definition, and the first hours after a restore are a cache-warming incident layered on top of whatever caused the restore — the operational sequencing in the mariabackup guide covers the restore side, and this article covers the minute after it finishes. And very large pools on modest storage: reloading 25% of a 200 GB pool is still 50 GB of random reads, which can take longer than the incident window you were trying to avoid — at that scale the honest options are lowering dump_pct, accepting a longer trough-hour restart, or questioning whether restarts should be routed around entirely with a replica-promotion dance instead of bouncing the writer in place. That last pattern composes naturally with the parallel-replication tuning in the optimistic parallel replication notes, because a promotion-based restart strategy lives or dies on how caught-up the candidate stays.
Results, for the record: with dump at shutdown, load at startup, and 25% dump_pct, our 64 GB pool returns to a sub-one-percent miss rate in about two minutes after a planned restart, versus forty minutes cold. The kernel patch cadence stopped being a latency event and went back to being a ninety-second non-event, which is what it always should have been.
Where MonPG fits
For MonPG product capabilities and setup information, see the MariaDB monitoring page. Use the diagnostics in this article to identify the measurements and operational checks your deployment needs.