Enabled semi-synchronous replication to guarantee zero data loss and watched commit latency quadruple? Here is the deep architecture of MySQL binary log commits, AFTER_SYNC vs AFTER_COMMIT, and cross-AZ network latency.
Database Topic Archive
MySQL Articles
InnoDB internals, replication, locking, query performance, and MySQL operations notes.
During a high-concurrency inventory reservation event, MySQL CPU hit 100% with lock wait timeouts everywhere. Here is the deep architectural mechanics of InnoDB next-key locks, gap locks, and deadlock detection cascades.
Our 8.0 config had innodb_flush_neighbors=1, inherited verbatim from a 5.7 my.cnf written for spinning disks in 2015. On NVMe it was pure write amplification: 18% of checkpoint flush IOPS were pages nobody asked to write. Setting it to 0 smoothed…
Every morning at 06:05 the first report query failed with 'MySQL server has gone away' and succeeded on retry. The pool held connections idle past the default 8-hour wait_timeout, and the server had closed them silently overnight. The fix was…
After a switch flap, the entire application fleet got 'Host is blocked because of many connection errors' — all at once, because every app server sits behind one NAT address. TRUNCATE performance_schema.host_cache unblocked us; understanding the error counter kept it…
The 5.7-to-8.0 upgrade went clean — until Monday, when a PHP billing worker started throwing error 2059 and a Perl cron from 2016 answered with 1251. The default authentication plugin changed, and nobody had read that line of the release…
Our cross-region replica link peaked at 310 Mbps during batch windows and lag followed like a tide. Turning on binlog_transaction_compression cut the stream to 115 Mbps for a 3% CPU cost — and then Debezium stopped reading us entirely, which…
We doubled the RAM, set innodb_buffer_pool_size=96G in my.cnf, restarted — and got 24G. Eight months earlier an on-call engineer had run SET PERSIST during an OOM scare, and mysqld-auto.cnf had been silently overriding the option file ever since.
After our 5.7-to-8.0 upgrade the error log ballooned from 40MB to 1.8GB a day, so we installed log_filter_dragnet — and our first broad rule quietly deleted the startup banners that would have been the timeline of a later crash investigation.…
We flipped innodb_flush_log_at_trx_commit to 2 and commit p99 dropped from 11ms to 0.8ms — a beautiful graph. Eleven weeks later a kernel panic proved what that graph cost: 47 orders the app had confirmed and InnoDB had not.