MySQL Dual Passwords: Zero-Downtime Credential Rotation Done Right
A compliance-mandated password rotation took the app down for 11 minutes at 02:20 — half the pods had the new credential, half had the old one, and the retry storms did the rest. MySQL 8.0.14's dual passwords make that outage a design choice, not an inevitability.
The rotation that took us down was executed exactly the way the old runbook said to execute it. Compliance had mandated a 90-day rotation for the application’s database user, so at 02:00 on the scheduled night I ran ALTER USER with the new password, pushed the new secret to the orchestrator, and triggered the rolling restart that would carry the credential to forty-odd pods. The rollout took eleven minutes. For those eleven minutes, the pods that had not yet restarted were connecting with a password the server no longer accepted — Access denied, immediate application retry, connection storm, the error log scrolling Access denied lines faster than I could read them, and three health-checked services flapping in the load balancer. Nothing was misconfigured; the procedure itself guaranteed an outage, because it contained a window where the fleet necessarily held two different credentials and the server honored only one. MySQL 8.0.14 shipped the fix for exactly this and almost nobody I talk to uses it: dual passwords, where a user account can hold a current and an old password simultaneously, both valid, until you explicitly retire the old one. Since that night, credential rotation here is a non-event, and this is the runbook plus the traps.
What does RETAIN CURRENT PASSWORD actually do?
It stores the new password as the account’s primary credential and keeps the previous one as a secondary, and the server accepts either one until you run ALTER USER … DISCARD OLD PASSWORD. That is the whole mechanism, and it inverts the old failure mode: instead of a cutover where the server and the fleet must change in the same instant, you change the server first, roll the fleet at whatever pace the deploy system wants, and close the window deliberately at the end. The feature is entirely server-side, which means no client upgrade is required — a five-year-old connector authenticating with the old password keeps working during the window, because as far as the protocol is concerned it simply presented a valid credential. It arrived in MySQL 8.0.14, so it survives the 5.7-to-8.0 upgrade as a genuine operational dividend, and it composes with the rest of the account machinery: password validation applies to the new password when you set it, and password_last_changed updates as usual. The one thing the catalog does not give you is visibility — mysql.user shows the primary credential’s metadata, and there is no clean column listing which accounts currently hold a retained secondary password. That gap becomes trap number one later; for now, hold the fact that you must keep your own ledger of in-flight rotations.
-- step 1: new password becomes primary, old stays valid
ALTER USER 'app'@'%' IDENTIFIED BY 'kX9!mZ2vQ7Lp' RETAIN CURRENT PASSWORD;
-- check the account state (note: the secondary password is not listed here)
SELECT user, host, plugin, password_last_changed, password_expired,
password_reuse_history, password_require_current
FROM mysql.user WHERE user = 'app';
-- step 3 (after the fleet has rolled): retire the old credential
ALTER USER 'app'@'%' DISCARD OLD PASSWORD;
How does the zero-downtime runbook actually run?
In four deliberate steps — retain, roll, soak, discard — with the soak being the step that separates a rotation from a gamble. Step one, retain: run ALTER USER … IDENTIFIED BY … RETAIN CURRENT PASSWORD at any quiet moment; it is a metadata change and effectively instant, and from this point both passwords work. Step two, roll: push the new secret through your normal deploy path — ours takes about forty minutes to drain through every pod, cron worker, and batch job, and the pace no longer matters because stragglers authenticate fine on the old credential. Step three, soak: this is the part I now refuse to skip. Wait at least one full business cycle — a day for a user-facing service, a week if monthly jobs use the same account — and verify that everything which should have picked up the new password has restarted since the roll; a forgotten cron box that never restarts will happily authenticate on the secondary password for months, which is precisely the stale credential you are trying to eliminate. Step four, discard: ALTER USER … DISCARD OLD PASSWORD, during business hours, with the error log on a second screen and an alert on Access denied lines for that user. If a straggler surfaces, the rollback is graceful — ALTER USER back to the previous password with RETAIN CURRENT PASSWORD re-establishes both credentials, you fix the straggler, and you run the discard again. Compare that to the old runbook’s rollback, which was another eleven minutes of partial outage.
What are the traps I stepped on?
Four, each with its own scar. First, the invisibility of retained credentials: because no catalog view lists secondary passwords, the only record of an open rotation window is your own, and during an audit we found a service account whose old password — retained “temporarily” during a rotation fourteen months earlier — still worked. The rule now: a rotation is a ticket, and the ticket is not closed until DISCARD OLD PASSWORD has run. Second, replication accounts have no dual-password path of their own: the replica stores its copy of the credentials in the replication connection settings, and CHANGE REPLICATION SOURCE TO SOURCE_PASSWORD = … must be run on each replica to match — the safe order is ALTER USER with RETAIN on the primary first (both passwords now valid), update every replica’s stored password (they immediately authenticate with the new one; the old still works for any replica you missed), then discard on the primary only after every replica checks in healthy. Third, connection pools hide stragglers: a pool that established its connections before the roll keeps them alive indefinitely, so “the app restarted” is not proof the old credential is unused — long-lived pools mean the soak step needs an explicit check of connection age, not just pod age; the connection pool notes cover why those pools outlive every deploy. Fourth, password_require_current interacts: if the account or the global policy requires the current password to change it, your automation must present the old one in the ALTER — build that into the rotation script before it fails at 02:00.
How do you prove the old credential is dead before discarding it?
You cannot prove a negative directly, so you get as close as the instrumentation allows: enumerate who is connecting, age their connections, and make the discard itself a monitored canary. The enumeration comes from performance_schema: session_account_connect_attrs carries the program name and client attributes each connection advertised, which answers “which applications and versions are connected right now” better than any spreadsheet, and the processlist’s connection ages answer “did everything restart since the roll.” On instances running an audit plugin — MySQL Enterprise Audit or the Percona equivalent — filter connect events by the account and you get the definitive list of source hosts still authenticating during the window, though you cannot see which password they used, so the audit answers “who connects” rather than “with which credential”; that distinction is exactly why the soak-plus-canary pattern exists. The canary is the discard run deliberately during staffed hours: alert on Access denied for the user, watch for fifteen minutes, and treat any hit as a finding with a graceful rollback, not an incident. And then the unglamorous closeout: update the ledger, close the ticket, and put the next rotation on the calendar — the error log triage notes cover why Access denied floods deserve their own alert rule rather than drowning in the general noise. Rotations stopped being scary here the day they became boring, and boring is the goal.
Where MonPG stands on MySQL
For MonPG product capabilities and setup information, see the MySQL monitoring page. Use the diagnostics in this article to identify the measurements and operational checks your deployment needs.