MonPG Engineering avatar MonPG Engineering Engineering Team MySQL 7 min read

Host Is Blocked: max_connect_errors, the host_cache, and the NAT Trap

After a switch flap, the entire application fleet got 'Host is blocked because of many connection errors' — all at once, because every app server sits behind one NAT address. TRUNCATE performance_schema.host_cache unblocked us; understanding the error counter kept it from recurring.

MySQL

The alert said the application could not connect to MySQL, which is the kind of alert that makes you check the wrong things first. mysqld was up, replication was fine, the hand-rolled health check from the monitoring host succeeded — but forty application pods were all logging the same line: Host is blocked because of many connection errors; unblock with mysqladmin flush-hosts. The trigger turned out to be a ten-second switch flap during rack maintenance. During the flap, hundreds of pooled connections died mid-handshake, the pool’s aggressive reconnect loop hammered the server with half-open attempts, and each failed handshake incremented a per-host error counter in MySQL’s host cache. When that counter crossed max_connect_errors — default 100, reached in seconds by a reconnecting fleet — MySQL stopped accepting connections from that host entirely. And here is the part that turned a blip into an outage: all forty pods sit behind one NAT address, so as far as the host cache was concerned, one host was misbehaving, and blocking it blocked the entire application tier at once. One TRUNCATE on performance_schema.host_cache restored service in four seconds; the two hours after that were spent understanding a mechanism I had only ever met as a punchline in forum posts.

What is the host cache actually for?

Two jobs, one structure. The host cache is primarily a DNS cache: when a client connects, MySQL resolves the client IP to a hostname for privilege matching against mysql.user host patterns, and the cache spares a DNS lookup on every connection. The same rows double as a security scoreboard: the cache tracks connection errors per host — handshake failures, auth failures from clients that never completed login, protocol violations — and once a host’s error count exceeds max_connect_errors, the server refuses new connections from it with the blocked message until the cache is flushed or the entry is otherwise cleared. The design intent is brute-force and misbehaving-client protection: a host spewing broken handshakes is probably hostile or broken, and refusing it protects the server’s accept loop. The modern exposure is that the per-host granularity assumed clients have distinct addresses, and between NAT gateways, Kubernetes egress, load balancers in TCP mode, and corporate proxies, whole fleets now present as one host. The cache itself is inspectable, which is how we reconstructed the incident: performance_schema.host_cache exposes IP, HOST, the error counters, SUM_CONNECT_ERRORS, and a COUNT_HOST_BLOCKED_ERRORS that told us this was not the first time we had flirted with the threshold.

-- the scoreboard: who has errors, who has been blocked
SELECT ip, host, host_validated,
       sum_connect_errors, count_host_blocked_errors
FROM performance_schema.host_cache
WHERE sum_connect_errors > 0
ORDER BY sum_connect_errors DESC;

-- the threshold and the cache size
SELECT @@max_connect_errors, @@host_cache_size, @@skip_name_resolve;

-- the immediate unblock (requires RELOAD or, better, the table is
-- truncatable directly with just the right privilege)
TRUNCATE TABLE performance_schema.host_cache;

-- the early-warning metric, before blocking starts
SHOW GLOBAL STATUS LIKE 'Connection_errors_%';

How does a healthy fleet cross the threshold in seconds?

By multiplying clients by retry loops. The counter increments on errors that happen during connection establishment — a TCP connection that opens and dies before the handshake completes, an aborted greeting, a TLS failure — and crucially, errors during an established session do not count; it is the connect path only. That means a network flap is the perfect storm: the switch blip severs five hundred pooled connections, the pool’s reconnect policy fires all of them at once into a network that is still converging, each attempt dies mid-handshake and increments the counter, and at the default max_connect_errors of 100 the shared NAT address is blocked in under three seconds. Our error log told the story afterward — a burst of aborted-connection notes in exactly the flap window, the kind of line the error log triage notes teach you to count rather than read. The reconnect policy deserves its own sentence: pools that retry instantly with no backoff turn a two-second network event into a self-inflicted block, and adding jittered exponential backoff to the pool config was one of the three permanent fixes. The connection pool sizing notes cover pool behavior under failure more broadly; the short version is that retry storms are a pool configuration problem before they are a database problem.

What do you change so it never pages you again?

Three fixes, in increasing order of controversy. First, operational: TRUNCATE TABLE performance_schema.host_cache is the unblock — it is instant, it clears DNS entries too, and unlike the older FLUSH HOSTS and mysqladmin flush-hosts forms it does not need a working application connection, just one privileged session you keep on a jump host that is not behind the NAT. Practice it before you need it; the night of our incident, the runbook said flush-hosts and the on-call had to find a path in. Second, architectural: stop letting the whole fleet share one apparent address where you can — direct pod networking, or at least per-node NAT pools, restores the per-host granularity the mechanism assumes. Third, the knob: raising max_connect_errors. It is dynamic, it applies immediately, and it is the fix everyone reaches for — but understand what you are trading. The counter exists to shed hostile or broken clients, and setting it to a million is functionally disabling the protection; we set it to 10,000 on instances fronting the NAT’ed fleet, which is about two orders of magnitude above what a retry storm with backoff produces, and left it at default everywhere else. There is also skip_name_resolve to consider: with it ON, the cache keys on raw IPs and skips DNS entirely, which sidesteps a related failure mode where slow or failing reverse DNS makes every connection wait — but it does not disable error counting, and it forces all your mysql.user host patterns to be IPs or IP wildcards, so adopt it deliberately, not as an incident reflex.

How do you watch the counter before it becomes a block?

The host cache is fully observable and almost nobody graphs it. The two signals that matter: the Connection_errors_* status counters as rates, which tell you handshake failures are happening at all, and the per-host sum_connect_errors from performance_schema.host_cache as a gauge against your configured max_connect_errors, which tells you how close any single host is to the cliff. We alert at 50% of the threshold on any host that fronts more than one application, which in practice means the NAT addresses; the alert has fired twice since, both times during network maintenance, both times resolved by the retry backoff doing its job before the threshold did. The thread states visible during a flap are worth knowing too — a pile of unauthenticated connections in the processlist is what a retry storm looks like from the inside, and the thread states notes decode that view. The lesson I took from the whole episode is the shape of the failure: MySQL’s protection mechanisms are designed around assumptions about client identity that modern networking quietly breaks, and the fix is rarely the knob alone — it is the knob, plus the pool policy, plus the network topology, plus an alert that understands all three.

Where MonPG stands on MySQL

For MonPG product capabilities and setup information, see the MySQL monitoring page. Use the diagnostics in this article to identify the measurements and operational checks your deployment needs.

Related documentation