MonPG Engineering avatar MonPG Engineering Engineering Team MariaDB 7 min read

MariaDB InnoDB Redo Log Sizing: Checkpoint Stalls and the Write Waves We Fixed

Every forty seconds, like clockwork, write throughput on our busiest primary collapsed to near zero for two to three seconds, then recovered. The culprit was not a query, a lock, or the disks. It was a redo log sized for a workload the server had outgrown two years earlier.

MariaDB

The graph was beautiful in the way only pathological graphs are: a sawtooth with a forty-second period. Writes per second on the orders primary would run healthy at around 18,000 rows a second, then fall off a cliff to almost nothing for two or three seconds, then snap back. CPU was fine. Replication lag was fine. The disks showed brief, violent flush bursts exactly aligned with the dips. Application teams blamed the network, then each other, and the incident review opened with the sentence I have learned to dread: the database is slow sometimes. It took a morning with InnoDB’s status counters to see the real shape of it. The redo log was 512 MB in total, the server was generating redo at roughly 200 MB a minute during the day, and InnoDB was hitting the checkpoint age ceiling every forty seconds, forcing an aggressive flush burst that stalled user threads. The database was not slow sometimes. It was being strangled on a schedule by its own crash-recovery insurance policy.

The redo log is the least discussed file in the datadir relative to how directly it gates write throughput, and its sizing advice is scattered across three eras of defaults. These are the notes from fixing ours.

What does the redo log actually throttle?

Every modification to an InnoDB page is first appended to the redo log — the write-ahead journal that makes crash recovery possible — and the dirty page itself is flushed to the tablespace later, at InnoDB’s convenience. The redo log is a circular buffer of fixed size, and a log record can only be overwritten once the dirty pages it describes have been flushed. That creates a hard coupling: if your redo log fills faster than the page cleaner can flush, InnoDB must stop accepting new changes until it catches up. That stop is the stall. The name of the ceiling is the checkpoint age — the gap between the newest log sequence number and the oldest log record still needed — and InnoDB compares it against thresholds derived from the log capacity: soft limits trigger background flushing, hard limits trigger what the source code calls a sync preflush, which is the polite term for everything waits while InnoDB panics pages to disk. An undersized redo log does not make your server slower on average. It makes it periodically, violently halted, which is far worse for tail latency and far harder to find in a p50 dashboard:

-- how fast are you generating redo? sample twice, subtract
SHOW GLOBAL STATUS LIKE 'Innodb_os_log_written';
-- wait 60 seconds, sample again; the delta is bytes/minute of redo

-- how much redo capacity do you have?
SELECT @@innodb_log_file_size;
-- 10.5+: a single ib_logfile0 of this size; earlier versions
-- multiply by @@innodb_log_files_in_group

-- checkpoint pressure, in one number
SHOW GLOBAL STATUS LIKE 'Innodb_lsn_current';
SHOW GLOBAL STATUS LIKE 'Innodb_lsn_flushed';
SHOW GLOBAL STATUS LIKE 'Innodb_checkpoint_age';  -- 10.8+

Two version facts matter before you touch anything. MariaDB 10.5 rewrote the redo log format, consolidated everything into a single ib_logfile0, improved crash recovery enough that the old keep-the-log-small-for-fast-recovery anxiety largely stopped applying, and raised the default to 96 MB — which was generous in 2020 and is trivially small for a busy server today. And since the 10.5 rewrite, resizing no longer requires the old ritual of clean shutdown, manually deleting log files, and praying on restart; the server rebuilds the redo log itself. Enterprise builds of 10.5 and 10.6 go further and let you SET GLOBAL innodb_log_file_size live, rebuilding the log crash-safely without a restart, and that capability has been flowing into community releases since. Check what your build supports rather than assuming the MySQL 8.0.30 innodb_redo_log_capacity playbooks apply — the divergence between the two InnoDB implementations is real and growing, as the InnoDB divergence notes cover at length.

How big should the redo log actually be?

The sizing rule that survives contact with production is embarrassingly simple: the redo log should hold thirty to sixty minutes of redo generation at your peak write rate, because that is enough slack for InnoDB’s background flushing to ride through bursts without ever approaching the sync-preflush cliff. Measure, do not guess. Sample Innodb_os_log_written sixty seconds apart during the busiest hour of the busiest day you have — for us that was the evening batch window, not the daytime peak, which is exactly the sort of thing you learn by measuring instead of assuming. Our peak came out at 210 MB a minute, so sixty minutes of headroom meant roughly 12 GB against the 512 MB we had: we were running at about 2% of the headroom the workload wanted. The counterintuitive part, and the part that makes senior engineers nervous, is that a redo log can safely be larger than the buffer pool, and often should be — the MariaDB documentation says as much. The historical argument for small logs was crash-recovery time, and 10.5’s recovery rewrite gutted it; recovery now streams the log without loading it wholesale into memory. The remaining cost of a big log is disk space and a slightly longer checkpoint horizon, both trivial next to the cost of stalls:

-- from measurement to setting: peak redo rate x 60 minutes
-- 210 MB/min x 60 = 12.6 GB, rounded to a clean number

-- my.cnf: 10.5+, one number, no file-count arithmetic
[mariadb]
innodb_log_file_size = 12G

-- Enterprise 10.5/10.6 and newer community releases: live resize
SET GLOBAL innodb_log_file_size = 12884901888;
-- then watch it rebuild:
SHOW GLOBAL STATUS LIKE 'Innodb_log_%';

One caveat for the Galera operators: redo flush behavior and flow control interact, and a node that stalls on checkpoint flushes will happily drag the whole cluster into flow control with it. If your sawtooth appears on a Galera node, fix the redo log before you tune wsrep — the flow control notes describe the symptom chain from the cluster side, and this article is the same incident viewed from the log side.

How do you resize without turning a tuning task into an outage?

On anything 10.5 or newer the procedure is boring, which is the highest compliment an operational procedure can earn. If your build supports SET GLOBAL innodb_log_file_size, run it during a traffic trough and watch InnoDB rebuild the log in the background; the server stays up, writes continue, and the resize is crash-safe by design. If it does not, edit the config, restart during your normal window — the restart itself is the maintenance, and on a warm server it takes seconds; the redo rebuild happens as part of shutdown and startup without the file-deletion dance older documentation still describes. What you must not do is take sizing advice written for pre-10.5 MariaDB or for MySQL and apply it blind: the file layout changed, the variable semantics changed, and the recovery tradeoffs changed. If you are mid-upgrade between LTS releases, sequence the resize after the upgrade so you are reasoning about one variable at a time — the upgrade ordering discipline from the 10.6-to-11.4 field notes applies verbatim. And keep the binlog separate in your head: redo is InnoDB’s private journal, binary logs are the server’s replication and point-in-time record, and sizing one does nothing for the other — the retention side of that pair is its own incident class, covered in the binlog purge notes.

Our result, for the record: innodb_log_file_size from 512 MB to 12 GB, applied via live resize on a Tuesday afternoon, and the sawtooth vanished on the first checkpoint cycle. Write throughput’s floor rose from near-zero to within 15% of its peak, p99 write latency dropped by a factor of four, and the flush bursts flattened into a steady background murmur. Total time to fix, once found: one hour. Total time to find: three weeks, because nobody was graphing checkpoint age.

Where MonPG fits

For MonPG product capabilities and setup information, see the MariaDB monitoring page. Use the diagnostics in this article to identify the measurements and operational checks your deployment needs.

Related documentation