One Group Replication member came back from a 40-minute network partition with 61,000 transactions queued — and writes on the two healthy members crawled for nine minutes while it caught up. The quota math, the thresholds that matter, and what…
Database Topic Archive
MySQL Articles
InnoDB internals, replication, locking, query performance, and MySQL operations notes.
A compliance-mandated password rotation took the app down for 11 minutes at 02:20 — half the pods had the new credential, half had the old one, and the retry storms did the rest. MySQL 8.0.14's dual passwords make that outage…
The replica stopped with Error 1032 at 21:14, and the on-call runbook said SET GLOBAL sql_slave_skip_counter=1. Six skips later the replica was serving rows that didn't exist on the primary. What these errors actually mean and how to fix the…
Our 5.7 box showed a 38% query cache hit rate and still served most reads from disk — the cache was inflating the numbers while a single mutex capped throughput. What 8.0 removed, how to find dependent queries, and what…
A mysqldump restore died at 61% with ERROR 1153 because one BLOB row exceeded the client default — and the dump itself had been 'successful'. The real defaults per tool, the replication variant, and how to size it without guessing…
A nightly DELETE of 9M rows stalled commits for 40 seconds while every other transaction queued behind it — and lsof showed a 3.1GB deleted-but-open file in /tmp. How the binlog cache works, what the spill costs at commit, and…
We encrypted forty tables for a compliance deadline and nearly lost them to a keyring file that sat in the datadir, unbacked-up. How InnoDB's two-layer keys work, what rotation really rewrites, and the operational rules that keep TDE boring.
The table created fine, then INSERTs started failing with ERROR 1118 only for rows with long JSON payloads. Why InnoDB enforces half a page per row, how DYNAMIC stores columns off-page, the two different 1118 errors, and fixes that survive…
A failover rehearsal exposed 214 rows on a replica that existed nowhere else, and replication had been 'healthy' the whole time. How drift happens under green dashboards, how pt-table-checksum proves it, and how to repair without a rebuild.
I killed a runaway UPDATE at 02:10 and the server kept burning CPU until 04:40 rolling it back. What KILL actually sets, how to measure rollback progress, why restarting makes it worse, and the query shapes that never get killed…