Infrastructure

Azure Managed Disks: diagnose I/O saturation before resizing

A production runbook to separate disk limits, VM I/O caps, burst credits and guest contention before resizing a disk or changing its performance tier.

03 Sept 2026 azuremanaged-disksvirtual-machinesstorageiopslatencyburstingobservabilityrunbookvalidationrollbackproduction

A database running on an Azure VM sees write latency climb from a few milliseconds to several dozen during its nightly processing window. The data disk is a Premium SSD, CPU remains stable, and the first proposal is to expand the disk. That may raise disk performance, but it will not help if the active ceiling belongs to the VM, burst credits have just run out, or the queue is forming inside the guest OS.

This runbook covers a VM hosting a data engine, indexer or batch workload with a reproducible I/O profile. Its goal is to identify the first effective ceiling, then choose between workload tuning, a performance tier, a VM resize or an architectural change. The expected outcome is not simply “more IOPS.” It is a testable change with an explicit return path.

Freeze the workload and impact before changing capacity

An hourly average can hide a five-minute bottleneck. Start with a narrow UTC window, one disk or LUN, and a recognizable business operation. Preserve a comparable healthy execution with the same volume, concurrency and data path.

yaml disk-io-incident-scope.yml
incident:
vm: vm-ledger-prod-02
region: westeurope
window_utc:
  start: 2026-09-03T00:55:00Z
  end: 2026-09-03T01:25:00Z

workload:
job: ledger-close
expected_records: <count>
concurrency: 8
affected_operation: transaction-log-write

storage_path:
mount: /var/lib/ledger
lun: 2
managed_disk: disk-ledger-log-prod-02
caching: None

symptom:
application_p95_ms: <value>
guest_await_ms: <value>
queue_depth: <value>

preserve:
- deployment_and_resize_timeline
- vm_size_and_disk_sku
- disk_and_vm_metrics_by_minute
- guest_iostat_sample
- healthy_control_window

Do not start a destructive benchmark against a production volume. Application counters, Azure metrics and a bounded guest sample are usually enough to locate the limit without adding a second workload during the incident.

Read the entire I/O path

One disk request can encounter several ceilings. A managed disk has its own IOPS and throughput limits. The VM size also imposes aggregate limits, with separate cached and uncached paths. Host caching changes the read path, while writes still need to be interpreted against the configured cache mode and the actual workload.

The diagnosis should answer these questions in order:

  • is the application issuing more operations, or smaller operations, than before?
  • do queue depth and latency rise inside the guest?
  • does a disk or LUN reach its consumed IOPS or bandwidth percentage?
  • does the VM reach its aggregate cached or uncached ceiling?
  • was the workload running above baseline by spending burst credits?

High latency without a consumption metric near its ceiling does not prove an IOPS shortage. CPU steal, memory pressure, swap, application locks, filesystem behavior and platform incidents still need to be considered.

Inventory the effective attachment

The ticket may describe a P30 while the workload actually uses another LUN, a temporarily elevated tier, or a cache mode that differs from infrastructure as code. Export effective state before applying theoretical limits.

bash 01-inventory-vm-and-disks.sh
RG="rg-ledger-prod"
VM="vm-ledger-prod-02"
DISK="disk-ledger-log-prod-02"

az vm show --resource-group "$RG" --name "$VM" --query "{size:hardwareProfile.vmSize,storage:storageProfile}" --output json

az disk show --resource-group "$RG" --name "$DISK" --query "{id:id,sku:sku.name,sizeGiB:diskSizeGb,tier:tier,iops:diskIopsReadWrite,mbps:diskMbpsReadWrite,bursting:burstingEnabled,managedBy:managedBy}" --output json

Match managedBy, the LUN, mount point and disk resource ID. On Linux, lsblk, findmnt and the stable links under /dev/disk/azure/ help avoid confusing a changeable device name with the volume the application actually uses.

Correlate latency, queue depth, IOPS and throughput

IOPS and throughput are not interchangeable. A small-write workload may exhaust IOPS without approaching the bandwidth limit; large sequential reads can do the opposite. Queue depth shows requests accumulating, but it must be read alongside latency and the concurrency the application is expected to sustain.

In Azure Monitor, chart at least the following metrics at the same granularity:

  • Data Disk IOPS Consumed Percentage and Data Disk Bandwidth Consumed Percentage, split by LUN;
  • VM Uncached IOPS Consumed Percentage and VM Uncached Bandwidth Consumed Percentage for a disk without host caching;
  • the cached equivalents when the path uses host caching;
  • Data Disk Latency, Data Disk Queue Depth, and read/write rates;
  • disk and VM burst credits when applicable.
bash 02-read-storage-metrics.sh
VM_ID=$(az vm show -g "$RG" -n "$VM" --query id -o tsv)

az monitor metrics list --resource "$VM_ID" --metric "Data Disk IOPS Consumed Percentage"          "Data Disk Bandwidth Consumed Percentage"          "VM Uncached IOPS Consumed Percentage"          "VM Uncached Bandwidth Consumed Percentage"          "Data Disk Queue Depth" --start-time "2026-09-03T00:55:00Z" --end-time "2026-09-03T01:25:00Z" --interval PT1M --aggregation Average Maximum --output json

Split metrics by LUN in the portal or through the metrics API where that dimension is available. An average across four disks can show 25% while one LUN is pinned at 100%.

Separate a disk cap from a VM cap

The decisive observation is the first utilization metric that reaches its ceiling while queue depth and latency rise.

text disk-bottleneck-decision.txt
One disk reaches 100%, VM remains below its limit
Disk-level cap is the leading hypothesis
Check IOPS versus bandwidth and the affected LUN
Consider a higher performance tier or workload distribution

VM cached or uncached metric reaches 100%, disks remain below their limits
VM aggregate storage cap is the leading hypothesis
A larger disk alone will not remove that cap
Evaluate VM size and cached/uncached path

Disk and VM both reach 100%
Both boundaries are active
Size the pair, not one resource in isolation

Queue and latency rise, Azure utilization stays below limits
Continue in guest and application layers
Check locks, filesystem, memory pressure, IO scheduler and platform health

Throughput reaches 100%, IOPS stays low
Larger IO size or sequential traffic is consuming bandwidth
Do not size the fix from IOPS alone

This avoids a common failure mode: raising a disk tier to 10,000 IOPS behind a VM whose uncached path is already capped. The change costs more and the disk graph looks comfortable, but business latency remains unchanged.

Check burst credits before treating the peak as baseline

Some disks and VM sizes can temporarily exceed their nominal performance. With credit-based bursting, a run begins quickly and slows as the bucket drains. The degradation appears progressive even though the workload stays constant.

Compare target IOPS with maximum burst IOPS, then inspect IOPS and bandwidth credit usage. Credit metrics are emitted less frequently than standard I/O metrics, so they should not be read as minute-perfect signals. Check the VM level as well: credits available on a disk cannot help if the VM can no longer carry the aggregate burst.

Credit-based bursting fits short peaks, not a sustained workload that has become normal. On-demand bursting can absorb spikes on eligible Premium SSDs with variable cost. A temporary performance tier offers defined capacity and more predictable billing for a planned window. The decision therefore depends on duration, recurrence and budget, not only the maximum observed value.

Choose the narrowest effective change

Once the ceiling is proven, apply one hypothesis at a time:

  • disk cap: temporarily raise a compatible Premium SSD tier, adjust IOPS and throughput on Premium SSD v2 or Ultra Disk, or distribute the workload;
  • VM cap: select a size whose cached and uncached limits cover the attached disks with headroom;
  • depleted burst: reduce the peak, stagger concurrency, or provision a sustainable baseline;
  • guest issue: fix filesystem, scheduler, memory or application behavior before buying Azure capacity;
  • new load: keep extra capacity only when useful work, rather than a loop or retry storm, explains the increase.

Do not change VM size, disk tier, host cache and application concurrency together. That removes causal evidence and makes rollback harder. Host caching in particular must be validated against the data engine’s write guarantees; it is not a universal accelerator.

Validate with real work and keep the return path

Prepare rollback before the change. For a Premium SSD tier, verify downgrade restrictions and timing. For a VM size, confirm regional capacity, compatibility, any restart requirement and the ability to return to the previous size. For an application change, preserve the previous version and configuration.

yaml disk-change-gates.yml
candidate_change:
type: temporary_performance_tier
target: disk-ledger-log-prod-02
from: P30
to: P40
owner: platform-oncall
expiry: 2026-09-04T06:00:00Z

success_gates:
- same useful workload volume completes
- application p95 returns below agreed threshold
- queue depth drains after the peak
- disk and VM consumed percentages retain headroom
- no new filesystem or database errors appear

rollback_triggers:
- latency does not improve under comparable load
- bottleneck moves to the VM without business gain
- error rate or data durability signal degrades
- observed workload differs from the frozen scope

rollback:
action: restore_previous_tier_when_allowed
evidence_to_keep:
  - before_after_metrics
  - effective_resource_configuration
  - workload_identifier_and_volume
  - decision_and_expiry

Validation should replay comparable useful work, not merely run fio against an empty volume. Measure business duration, latency, queue depth, disk and VM ceilings, then verify that extra capacity is not hiding retry amplification.

Conclusion

Azure I/O saturation is not fixed by enlarging the first visible disk. Freeze the workload, map the actual LUN, correlate IOPS, throughput, latency and queue depth, then compare disk and VM ceilings. Burst credits often explain why a short run succeeds and the same sustained process slows later.

The final decision should name the proven boundary: raise a tier, resize the VM, smooth the workload, fix the guest OS, or leave storage unchanged. The change is validated when the same unit of work meets its latency objective with headroom. Roll it back when the improvement is not measurable or merely moves saturation to another layer.