In a 3-node MariaDB Galera cluster, an automated backup on Node 3 caused write throughput across the whole cluster to collapse from 2,800 to 12 QPS. Here is how Galera flow control works and how to design resilient replication topologies.
Database Topic Archive
MariaDB Articles
Galera clustering, InnoDB and Aria internals, replication, optimizer behavior, and MariaDB operations notes.
A routine rolling restart of the application tier turned into twelve minutes of ERROR 1040: every app node reconnected at once, idle connections from the old processes had not timed out yet, and the DBA could not even get a…
The monthly finance report had run for two years without incident, right up until the month the data crossed an invisible line: its GROUP BY stopped fitting in memory, spilled to an on-disk temp table, and filled the 8 GB…
After the third 'the site felt slow around lunch' report in a month, we stopped arguing about whether it happened and turned on the slow query log properly. The first digest took eleven minutes to produce and named one query…
The kernel patch reboot took ninety seconds. The hangover took forty minutes: p99 latency eight times normal while InnoDB re-read its working set from disk, one cold page at a time. The fix had been compiled into the server the…
Every forty seconds, like clockwork, write throughput on our busiest primary collapsed to near zero for two to three seconds, then recovered. The culprit was not a query, a lock, or the disks. It was a redo log sized for…
A customer named José — or rather José, which is how his name had been stored for two years — finally filed the ticket that started our utf8mb4 migration. Then the first ALTER failed with ERROR 1071 because a 255-character…
We added two audit columns as INVISIBLE so nothing would see them until the rollout finished. Nothing saw them — including the nightly loader, which died at 02:40 with ERROR 1136 because its INSERT listed no column names, and the…
The primary ran out of disk at 04:12 on a Saturday: 340 GB of binary logs, accumulating since a migration eleven months earlier, on a server where expire_logs_days had never been set. Then the cleanup broke a replica nobody remembered…
A one-line ALTER to widen a VARCHAR queued behind a long analytics transaction — and because metadata locks queue fairly, every new query on the table stacked up behind the ALTER. The whole application hung on a schema change that…