Binlog Transaction Compression: Cutting Replica Bandwidth 63% with ZSTD
Our cross-region replica link peaked at 310 Mbps during batch windows and lag followed like a tide. Turning on binlog_transaction_compression cut the stream to 115 Mbps for a 3% CPU cost — and then Debezium stopped reading us entirely, which taught us the compatibility half of the feature.
The lag graph had a daily rhythm, and the rhythm was the batch window. Every night between 02:00 and 04:30 our reporting pipelines rewrote large slices of the warehouse-staging tables, the primary’s binary logs swelled with row events carrying wide JSON columns, and the cross-region replica in the failover site fell behind — forty minutes at the worst of it, recovering by breakfast. The link between regions was the constraint, priced by the megabit and already peaking at 310 Mbps during those windows, so the options were a bigger pipe, less data, or a smaller stream. We already knew why the stream was wide: row-based replication ships before-and-after images, and the binlog row image notes cover that story, including the FULL versus MINIMAL trade we had already taken where the schema allowed. What we had not taken seriously was a feature added in MySQL 8.0.20: binlog transaction compression, which ZSTD-compresses the payload of each replicated transaction inside the binlog itself. We enabled it one Tuesday, and the same batch window peaked at 115 Mbps. Then our Debezium connector threw parse errors and stopped consuming entirely, which is the half of this feature the release notes mention and the failure modes section should underline.
What exactly gets compressed, and what does not?
The transaction payload is compressed as a unit; the binlog’s structural scaffolding is not. With binlog_transaction_compression=ON, the server collects a transaction’s row events into a single compressed Transaction_payload_event carrying a ZSTD frame, while the GTID event, the query headers, the anonymous metadata, and the binlog’s own rotation and heartbeat machinery stay uncompressed so the stream remains navigable. That split matters when you estimate savings: you do not compress your binlog by 63%, you compress the row payload, and your realized number is a function of how much of your stream is fat row events versus everything else. Our stream was unusually compressible — the JSON staging columns were repetitive across rows and the batch jobs touched millions of them — so 63% is near the friendly end of the distribution; a workload of short random-key updates on narrow tables should expect far less, and the only honest number is the one from your own stream. The setting is dynamic, applies to new transactions as they are logged, and has a tunable: binlog_transaction_compression_level_zstd defaults to 3, and we tested levels 1 through 7 on a replayed capture before settling at 3, because level 7 bought us four more percentage points of bandwidth for nearly double the compression CPU.
-- the knobs, all dynamic
SET PERSIST binlog_transaction_compression = ON;
SET PERSIST binlog_transaction_compression_level_zstd = 3;
-- confirm what replicas will actually receive
SELECT @@binlog_transaction_compression,
@@binlog_transaction_compression_level_zstd;
-- did the stream shrink? watch bytes written per file rotation:
SHOW BINARY LOGS;
-- and binlog generation rate over time:
SHOW GLOBAL STATUS LIKE 'Binlog_bytes_written'; -- 8.4+ name; on 8.0
-- compare file growth per hour against the pre-change baseline
-- compressed transactions flowing to replicas, per connection:
SELECT * FROM performance_schema.replication_connection_status;
What are the compatibility traps, in the order we hit them?
Trap one, the one that stopped Debezium: anything that reads the binlog outside MySQL’s own replication protocol must understand the Transaction_payload_event, and in 2026 the major CDC connectors do, but their older deployed versions do not — our connector was two minor releases behind and its binlog parser hit an event type it had never seen and died cleanly, which is the good outcome; silently skipping compressed transactions would have been the catastrophic one. The rule we adopted: every downstream consumer — Debezium, Maxwell, Canal, in-house parsers, the analytics team’s mysqlbinlog scripts — gets verified against a compressed stream on a staging replica before production flips. Trap two: version floors. Replicas must run 8.0.20 or later to decompress; an older replica in the chain, including one you forgot about behind a relay, will fail to apply. mysqlbinlog itself needs to be from 8.0.20+ to decode the events for point-in-time work, which matters the night you reach for the PITR workflow with a laptop that has an old client package. Trap three: compression and encryption compose but multiply cost — if you also run binlog_encryption, each payload is compressed then encrypted, and the CPU stacks accordingly. None of these are reasons to avoid the feature; they are reasons to enable it with the same ceremony as a schema change: staged on one replica pair first, consumers checked, runbook updated.
How do you know it is paying for itself?
Three measurements, taken before and after, on the same batch windows. First, bandwidth at the link: graph the replication stream’s Mbps against the prior weeks and confirm the reduction survives a full batch cycle, not just a quiet Tuesday. Second, replica lag: the point of the exercise was the 02:00-04:30 tide, so Seconds_Behind_Source during the window is the scoreboard — ours went from a 40-minute peak to under 4 minutes, which changed the failover site’s RTO math more than the bandwidth bill did. The replication lag diagnosis notes are the companion for telling bandwidth-bound lag from applier-bound lag, because compression fixes only the former. Third, the bill you pay: compression CPU on the source, decompression CPU on the replica’s I/O receiver, and neither is free. We measured 3% additional CPU on the source at level 3 and less on the replica, an excellent trade against a network upgrade — but on a primary already running at 70% CPU in peaks, that 3% deserves a capacity conversation before the flip, not after. One measurement that surprised us: binlog files on disk shrank too, since the compressed payload is what gets written, which relaxed the retention pressure the binlog retention notes describe. Smaller files, less replication bandwidth, faster PITR transfers — the feature pays in three currencies, provided your consumers can read the new event type.
Where MonPG stands on MySQL
For MonPG product capabilities and setup information, see the MySQL monitoring page. Use the diagnostics in this article to identify the measurements and operational checks your deployment needs.