innodb_flush_neighbors: The Spinning-Disk Setting We Carried Onto NVMe
Our 8.0 config had innodb_flush_neighbors=1, inherited verbatim from a 5.7 my.cnf written for spinning disks in 2015. On NVMe it was pure write amplification: 18% of checkpoint flush IOPS were pages nobody asked to write. Setting it to 0 smoothed the checkpoint spikes and cut write volume measurably.
The config file was older than the disks by three hardware generations. During a storage review on our busiest write node — prompted by checkpoint-flush latency spikes that showed up as write p99 sawteeth every few minutes — I opened my.cnf and found innodb_flush_neighbors=1 sitting between settings dated by their comments to 2015, when the instance lived on a RAID-10 array of 10k RPM drives. The server now runs on local NVMe, the RAID array is a memory, and the setting had crossed two major upgrades and a hardware refresh untouched, because config files migrate by copying and nobody deletes lines that are not obviously broken. Flush neighbors was not obviously broken; it was quietly expensive. Once we measured, roughly 18% of the pages InnoDB flushed during checkpoint bursts were neighbors — dirty pages adjacent to the ones actually due, written early and in bulk because a setting designed to save seek time on spinning rust was coalescing writes that NVMe never needed coalesced. We set it to 0, which is also the 8.0 default we had overridden without knowing it, and the sawteeth flattened: same checkpoint work, less write volume, tighter tail latency. The whole episode is a small story about one variable and a large story about config archaeology.
What does innodb_flush_neighbors actually do?
It controls whether InnoDB, when flushing a dirty page, also flushes other dirty pages that are physically adjacent in the same extent. At 1, flushing a page triggers a sweep of its neighbors in the extent and any dirty ones go out in the same batch; at 2, a milder variant, the neighbors are gathered but the flush happens in a single write rather than forcing per-page operations; at 0, InnoDB flushes exactly the pages the flush logic selected and no more. The rationale is entirely rotational: on a spinning disk, the expensive part of a write is the seek, and once the head is over a track, writing the neighboring pages is nearly free — so batching adjacent dirty pages converts many small random writes into fewer large sequential ones, a clear win in 2010. On flash, the premise inverts. NVMe has no head and no seek; a random 16KB write costs the same as any other 16KB write, and writing pages early and extra is not free — it is write amplification, consuming IOPS the workload could have used, competing with foreground page reads during checkpoint bursts, and spending the drive’s write endurance on pages that would have been written later anyway, possibly after being dirtied again and written once instead of twice. The defaults tell the story: 5.7 defaulted to 1, from the spinning era; 8.0 defaults to 0, because the industry turned over underneath the setting. Our explicit =1 line had pinned the old default across the turnover.
-- the setting, and the flush machinery it feeds
SELECT @@innodb_flush_neighbors, @@innodb_io_capacity,
@@innodb_io_capacity_max, @@innodb_lru_scan_depth;
-- pages flushed per second, total — watch this drop after the fix
SHOW GLOBAL STATUS LIKE 'Innodb_buffer_pool_pages_flushed';
-- pages actually requested by flush list vs LRU flushing:
SHOW GLOBAL STATUS LIKE 'Innodb_buffer_pool_wait_free';
-- dirty pages outstanding (drives the checkpoint bursts):
SHOW GLOBAL STATUS LIKE 'Innodb_buffer_pool_pages_dirty';
-- apply the fix live, then persist
SET GLOBAL innodb_flush_neighbors = 0;
SET PERSIST innodb_flush_neighbors = 0;
How do you measure whether it is costing you anything?
Compare flushed pages against the pages that needed flushing, and let the storage layer confirm. The InnoDB side gives you Innodb_buffer_pool_pages_flushed as a rate; the redo side gives you how much dirty-page generation the checkpoint machinery must keep pace with, and when the flush rate persistently outruns what checkpoint age and LRU pressure require, the surplus is neighbor flushing and read-ahead of the write variety. The storage side is blunter and more convincing: iostat write throughput and write IOPS during a checkpoint burst, before and after flipping to 0, with the workload held as constant as a production system allows. Our numbers: same batch window, write IOPS at the burst peak down 14%, write megabytes per second down 18%, and the p99 latency of application writes during the burst from 9ms to 4ms, because the flush bursts no longer saturated the queue while foreground reads waited behind neighbor pages nobody wanted. Do not skip the before measurement; the temptation with a setting this obviously wrong is to flip it fleet-wide on principle, and principle does not tell you what you saved or catch the one instance where something else was masking the effect. The wider flushing context — io_capacity, LRU scan depth, and how hard InnoDB works to keep free pages available — is covered in the io_capacity flushing notes, and flush neighbors should be understood as a volume knob bolted onto that machinery, not a substitute for sizing it.
Is there any storage left where 1 is correct?
Yes, and saying so matters because the honest rule is match the setting to the seek cost, not always set 0. Genuine spinning disks still exist in database service: capacity-optimized HDD arrays on older SANs, cold archive replicas, budget tiers of some hosted database products where the volume is backed by magnetic storage without telling you loudly. On those, neighbor flushing still converts seeks into streaming, and the setting still pays. There is also a middle case that took us a conversation to place: SATA SSDs behind a RAID controller with a small write cache can behave rotationally enough at the queue level that batching helps modestly — though even there, 2 is usually the better value than 1, and measuring beats theorizing. Everything else in a modern fleet — local NVMe, instance store, network SSD volumes with real IOPS provisioning — wants 0, and the 8.0 default already says so; the failure mode is never the default, it is the inherited explicit line that overrides it. Which is the actual lesson of the incident: the variable did not hurt us because we misunderstood InnoDB, it hurt us because a my.cnf is an archaeological record, and lines copied forward without review outlive the hardware they were tuned for. The same review that caught this one found two other 5.7-era lines of similar vintage, and the SET PERSIST drift notes cover the runtime half of the same disease — values changed live and never reconciled with the file.
How do you keep old defaults from riding forward forever?
One practice: at every major upgrade and every storage change, diff your effective configuration against the new version’s defaults and force a written verdict on every explicit divergence — keep, change, or delete — with the reason recorded in the file itself as a comment. That is fifteen minutes of work per instance against the reference sheets, and it is the only mechanism we have found that reliably retires settings whose premises expired. Flush neighbors passed three upgrades because nothing about it looked alarming; the words flush and neighbors do not announce I am for spinning disks. The config comment we left behind says it plainly now: innodb_flush_neighbors=0, NVMe since 2021, default since 8.0, revisit only if this instance ever lands on rotational storage. Settings that name their hardware assumption in the file are settings the next engineer can re-litigate in seconds. The flush method question — O_DIRECT and its fsync choreography — is the neighboring topic, literally and figuratively, and the flush method tuning notes cover it with the same evidence-first approach. Storage settings are bets on hardware; hardware changes; the file should say what it is betting on.
Where MonPG stands on MySQL
For MonPG product capabilities and setup information, see the MySQL monitoring page. Use the diagnostics in this article to identify the measurements and operational checks your deployment needs.