I once watched a Profiler trace against a production instance take the box from 2,000 batches a second to timeout errors in ninety seconds. Extended Events is the replacement, but only if you write sessions with the same discipline Profiler…
Database Topic Archive
SQL Server Articles
Wait statistics, tempdb, Query Store, Always On availability groups, and SQL Server operations notes.
Every morning at 8:15 the buffer cache hit ratio sagged, page life expectancy fell off a cliff, and someone asked if the server needed more RAM. The server needed an index. How to read buffer pool counters without the folklore,…
A killed six-hour batch job once rolled back for nearly five hours, holding its locks the whole time. With accelerated database recovery the same class of kill returns the database in seconds — but the persisted version store becomes a…
A 09:20 deploy flipped a hot query from index seek to table scan and p95 latency went from 40 ms to 2.8 seconds. Query Store flagged the regression by 09:35, a forced plan held the line for two days, and…
The events table stalled at 4,100 inserts per second on a 96-core box with idle disks, and every waiter was queued on PAGELATCH_EX for the same page. Why ever-increasing keys funnel inserts into one latch queue, and what actually breaks…
CPU was at 34 percent, the disks were bored, and 380 sessions were queued on PAGELATCH_EX against page 2:1:1 in tempdb. Why one data file serializes allocation, how many files to add, and the version store case files alone will…
One statistics update at 9:14 AM recompiled a stored procedure against the biggest customer's parameters, and every small customer went from 50 ms to six seconds. The patterns, the fixes, and their honest costs.
A failover that should have taken two minutes took forty-seven because the log had 30,000 virtual log files. How autogrowth wrote that layout, what log_reuse_wait_desc is telling you, and how to rebuild the layout exactly once.
The nightly purge had been deleting for six hours and generating 400 GB of log when they called me; the sliding-window replacement runs in 200 milliseconds. Partition functions, aligned indexes, SWITCH mechanics, and when partitioning is the wrong tool.
The index existed and the seek was in the plan, but the estimate said one row and the table said fourteen million. How auto-update thresholds, sampling, and the ascending-key trap produce bad estimates, and how to triage them with sys.dm_db_stats_properties.