MonPG Engineering avatar MonPG Engineering Engineering Team MySQL 6 min read

Group Replication Flow Control: The Throttle That Stalled My Whole Cluster

One Group Replication member came back from a 40-minute network partition with 61,000 transactions queued — and writes on the two healthy members crawled for nine minutes while it caught up. The quota math, the thresholds that matter, and what to graph.

MySQL

The alert said writes were slow, and nothing was wrong — which is the worst kind of alert. At 14:07 on a Tuesday, p99 write latency across a three-node MySQL 8.0 Group Replication cluster climbed past two seconds and stayed there for nine minutes. The slow query log was empty. There were no lock waits, no replication errors, CPU was at 30 percent on every node, and the primary’s error log had nothing but routine chatter. The cause, when I finally pulled performance_schema.replication_group_member_stats, was on member-03: COUNT_TRANSACTIONS_REMOTE_IN_APPLIER_QUEUE sat at 61,000 and counting down slowly. The member had spent forty minutes behind a network partition earlier that morning, rejoined cleanly, and was now chewing through its backlog — and flow control, doing exactly what it was designed to do, had decided the whole group should write no faster than its slowest member could catch up. Every application write on the two perfectly healthy members was being throttled to protect the one that was behind. That incident rewired how I think about Group Replication: the group’s write throughput is a shared resource with a governor on it, and if you don’t know the governor’s math, it will introduce itself during an incident.

What is flow control actually protecting?

It keeps every member’s transaction backlog bounded, so that any member can take over writes with minimal catch-up time and the certification information on each node can be garbage-collected instead of growing forever. Without it, a slow member falls arbitrarily far behind: its applier queue grows without limit, a failover to that member means minutes of catch-up during which it cannot serve current data, and the certification database on every member — which must be retained until all members have applied the corresponding transactions — balloons in memory. Flow control’s mechanism in the default QUOTA mode is simple to state and subtle in effect: every group_replication_flow_control_period (default one second), members exchange their queue statistics; if any member’s applier queue exceeds group_replication_flow_control_applier_threshold or its certifier queue exceeds group_replication_flow_control_certifier_threshold — both default to 25,000 transactions — the group computes a reduced write quota and every member caps its new writes to that quota until the queues drain below the threshold again. The key phrase is every member: the throttle is group-wide by design, because a group is only as current as its stalest node.

-- the four columns that explain any flow-control incident
SELECT MEMBER_ID,
       COUNT_TRANSACTIONS_IN_QUEUE               AS certifier_queue,
       COUNT_TRANSACTIONS_REMOTE_IN_APPLIER_QUEUE AS applier_queue,
       COUNT_TRANSACTIONS_REMOTE_APPLIED,
       COUNT_TRANSACTIONS_LOCAL_PROPOSED
FROM performance_schema.replication_group_member_stats;

-- the governor's configuration
SELECT @@group_replication_flow_control_mode,
       @@group_replication_flow_control_applier_threshold,
       @@group_replication_flow_control_certifier_threshold,
       @@group_replication_flow_control_period;

Why did one recovering member stall writes on all three?

Because the quota is recomputed each period from the slowest member’s demonstrated apply capacity, so a member applying at 800 transactions per second drags the group’s write ceiling toward that rate for as long as its queue stays above the threshold. In our incident, the healthy members were proposing roughly 2,400 writes per second at midday, the recovering member was applying at about 800 — single-threaded applier, cold buffer pool after its restart, and a storage layer still warming its readahead — and the quota math did the rest: 61,000 queued transactions against a 25,000 threshold meant sustained throttling, and the queue was draining at only 800 minus the writes the quota still allowed through, which is why nine minutes passed before latency snapped back. Two details in the defaults soften this and are worth knowing precisely. group_replication_flow_control_hold_percent (default 10) lets a member keep a percentage of its quota in reserve so a throttled writer is not starved to zero, and the quota is released gradually rather than all at once so the group does not oscillate between full throttle and none. But no default changes the structural fact: in QUOTA mode, your cluster’s write latency during any member’s catch-up window is a function of your slowest member’s applier, and a failover tool that advertises Group Replication as “synchronous-ish” rarely mentions the governor in the brochure.

Should you raise the thresholds or disable flow control entirely?

Neither, at first — fix the lagging member’s apply capacity, because flow control is almost always reporting a real problem, and then size the thresholds to your actual failover tolerance rather than to silence the throttle. The capacity fixes, in the order that paid off for us: parallel applier workers on every member (replica_parallel_workers set to match the workload’s schema parallelism), buffer pools large enough that a recovering member isn’t reading its catch-up from disk, and an honest look at whether one member has weaker storage than its peers — ours did, and equalizing the volume class cut the apply gap more than any tuning. Only then touch the governor. Raising the thresholds from 25,000 to, say, 250,000 is legitimate when your members apply fast in absolute terms and a large queue represents seconds rather than minutes of divergence — the threshold is denominated in transactions, not time, and a 100,000-transaction queue is eight seconds of divergence at 12,000 trx/s and eight minutes at 200. Switching group_replication_flow_control_mode to DISABLED is the move I argue against in every design review: it converts bounded, visible slowness into unbounded, invisible divergence, and the member that falls too far behind no longer catches up at all — it needs to be re-provisioned from scratch, which on 8.0 at least means the distributed-recovery clone path covered in the clone plugin notes rather than a manual rebuild. The throttle was never the enemy; it was the messenger with an inconvenient message.

How do you watch flow control before it throttles you?

Graph the applier queue per member, alert on divergence between members rather than on absolute lag, and rehearse the incident once before it happens for real. The monitoring is genuinely cheap: replication_group_member_stats exposes COUNT_TRANSACTIONS_REMOTE_IN_APPLIER_QUEUE per member, and the useful derived signal is the spread — the max member’s queue minus the min’s — because a healthy group’s queues move together and a diverging member stands out minutes before the threshold trips. Alert at a fraction of the threshold, not at it: if the threshold is 25,000, a page at 10,000 queued buys you the catch-up window to act before the governor does. Pair the queue graph with the member states from replication_group_members, because a member flapping between ONLINE and RECOVERING is the usual precursor — the replication lag diagnosis flow applies to GR members too, with the group views replacing SHOW REPLICA STATUS. And rehearse: partition a staging member for ten minutes, let it rejoin, and watch your own dashboards during the throttle. The first time you see flow control engage should be on a drill, not on a Tuesday at 14:07 with three services timing out. If you are still weighing GR against classic async for a new build, the group replication versus async comparison covers the trade this governor embodies.

Where MonPG stands on MySQL

For MonPG product capabilities and setup information, see the MySQL monitoring page. Use the diagnostics in this article to identify the measurements and operational checks your deployment needs.

Related documentation