Infrastructure

Azure Monitor Managed Prometheus: contain a cardinality spike before raising quotas

A production runbook for isolating the metric, target and labels behind an AKS time-series spike before filtering collection, adding capacity or rolling back.

11 Sept 2026 azureazure-monitorprometheusaksobservabilitycardinalitymetricsautomationrunbookrollbackproduction

AKS dashboards slow down, some metrics develop gaps, and the Azure Monitor workspace moves toward its ingestion limits. The immediate response is often to request more capacity. Yet one release may have attached request_id, user_id, an unnormalized URL or an ephemeral pod identifier to Prometheus labels. Every new combination becomes a time series, turning a local instrumentation change into a shared observability incident.

The running case is an API on AKS collected by Azure Monitor Managed Prometheus. After deployment, active series and ingested events accelerate while Grafana queries take longer. This runbook drives a decision: remove a label, narrow a scrape scope, repair instrumentation, distribute collection, request a limit increase or roll back the release. The objective is not to make a dashboard green by silently discarding the metrics operators need.

Freeze the incident window and diagnostic unit

Start with one workspace, cluster, release and UTC interval. Do not mix gradual traffic growth with a sharp post-deployment break.

yaml prometheus-cardinality-incident.yml
incident: inc-20260911-004
workspace: amw-platform-prod
cluster: aks-platform-prod
window_utc:
start: 2026-09-11T06:30:00Z
end: 2026-09-11T07:30:00Z
candidate_release: orders-api-2026.09.11.2
symptoms:
- active_series_rising
- events_per_minute_rising
- slow_promql_queries
- gaps_on_selected_targets

preserve:
- workspace_utilization_before_and_after_release
- deployed_scrape_and_relabel_configuration
- target_list_and_up_state
- metrics_and_labels_added_by_release
- previous_manifest_and_image
- affected_dashboards_and_rules

Capture the configuration actually loaded by the collectors, not only the repository version. Git may contain the intended ConfigMap while the cluster still runs an earlier revision, or an additional scrape object may select the same target twice.

Separate workspace pressure from collection failure

Missing data does not automatically prove a limit was reached. Inspect workspace platform metrics first, then the DCR and ama-metrics collectors. Metric names and dimensions can evolve, so enumerate definitions on the resource before automating a query.

bash 01-ingestion-inventory.sh
az monitor metrics list-definitions \
--resource /subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-observability-prod/providers/Microsoft.Monitor/accounts/amw-platform-prod \
--output table

kubectl get pods -n kube-system -o wide
kubectl get configmap ama-metrics-settings-configmap -n kube-system -o yaml
kubectl get configmap ama-metrics-prometheus-config -n kube-system -o yaml

At workspace scope, compare active time-series utilization, received events and ingestion errors. On the DCR, look for rejected requests and response-code dimensions. In the cluster, separate four states: target not discovered, scrape failing, samples produced but rejected, or ingestion accepted while PromQL queries become expensive.

Enable collector debug mode only for a short, owned window because it can add resource pressure. Collector logs, target status, and the up, scrape_samples_scraped and scrape_series_added metrics are usually enough to locate the offending job.

Prove the cardinality break

Find the change in slope first, then restrict analysis to a few candidate metric families. A workspace-wide query can worsen an incident that already involves query pressure.

promql 02-cardinality-by-metric.promql
topk(
20,
count by (__name__) (
  {__name__=~"http_.+|rpc_.+|orders_.+"}
)
)

Compare the same window and step before and after the release. A family that jumps from hundreds to tens of thousands of series is a candidate, but it does not yet identify the responsible label.

promql 03-candidate-dimensions.promql
topk(20, count by (route) (http_server_request_duration_seconds_count))

topk(20, count by (status_code) (http_server_request_duration_seconds_count))

count(count by (request_id) (http_server_request_duration_seconds_count))

Use the final query only after bounding the metric and time window. If request_id has almost one value per request, it has no operational value as a label. The identifier belongs in correlated traces or logs, not in the time-series index. Apply the same test to user_id, session_id, raw URLs, error messages and unbounded commit identifiers.

Find the source before filtering

The same symptom can originate in three layers: the application exposes too many labels, scrape configuration enriches series, or multiple collection objects scrape the same endpoint. Locate the first layer that introduces the dimension.

text cardinality-cause-tree.txt
Label is already present on /metrics
Repair application instrumentation
Keep detailed identifiers in traces or logs
Version the correction with the release

Label appears only after discovery or relabeling
Inspect ServiceMonitor, PodMonitor or custom scrape configuration
Review labelmap, target_label and copied Kubernetes labels
Remove only the rule that expands the dimension

The same target appears more than once
Compare job, instance, endpoint and discovery objects
Disable duplicate collection only after proving it
Verify rules and dashboards use the retained series

Volume rises without a new label
Compare replicas, targets, scrape frequency and exposed metrics
Separate expected scale from a scope change

Do not drop a label merely because it has many values. pod, namespace or status_code may support an alert or diagnosis. Ask whether the dimension drives an operational decision and whether its value domain is bounded.

Build the smallest reversible correction

The durable fix usually belongs in instrumentation: replace raw URLs with normalized routes and move unique identifiers to traces. When ingestion must be stabilized before an application release is ready, use narrowly scoped relabeling on the affected job.

yaml 04-candidate-relabeling.yml
scrape_configs:
- job_name: orders-api
  scrape_interval: 30s
  kubernetes_sd_configs:
  - role: pod
  metric_relabel_configs:
  - action: labeldrop
    regex: "request_id|user_id|session_id"
  - source_labels: [__name__]
    action: keep
    regex: "http_server_.+|process_.+|runtime_.+|orders_.+"

Adapt the fragment to the collection mechanism in use; do not replace the whole ConfigMap blindly. Export the current manifest, validate YAML away from production, and apply the candidate to a canary cluster or scrape job. A labeldrop does not remove historical series and can merge samples that differed only by the removed label. Check for collisions and changed aggregation semantics.

Increasing the scrape interval can reduce events per minute, but it does not fix an unbounded dimension. Conversely, dropping a label reduces future series without necessarily fixing throughput caused by too many targets. Match each limit to its actual cause.

Protect alerts, dashboards and recording rules

Before cutover, find queries that depend on removed labels or filtered metrics. Cheaper collection that breaks the availability alert is not recovery.

yaml observability-validation-contract.yml
consumers_to_test:
- service_dashboard
- availability_alert
- saturation_alert
- recording_rules
- capacity_reports

signals_to_keep:
- up_by_job_and_instance
- latency_by_normalized_route
- errors_by_status_class
- saturation_by_workload
- trace_id_correlation_in_traces

signals_to_reject_as_labels:
- request_id
- user_id
- session_id
- parameterized_url

Replay representative PromQL queries with the same interval, then compare output, duration and series count. Test recording rules separately: they move query cost and may keep requesting a dimension that no longer exists.

Decide between correction, capacity and rollback

A limit increase is reasonable when growth is expected, measured and useful. It should not become the default rollback for broken instrumentation.

text prometheus-cardinality-decision.txt
Repair instrumentation
An unbounded identifier is exposed as a label
The responsible release is identified
A bounded route or class is available as replacement

Apply a temporary collection filter
The affected job and labels are proven
Critical alerts do not depend on the removed label
Previous manifest is available for rollback

Request more capacity or distribute collection
Growth comes from legitimate workloads
Labels have proven operational value
Projected utilization remains controlled after the increase

Roll back the release
Cardinality spikes immediately with new instrumentation
An application fix cannot fit the incident window
Previous image and metric schema remain compatible

Keep the change blocked
Series origin is still unknown
Data is already rejected or targets are unstable
The candidate correction removes an on-call signal

Validate across several scrape cycles

Keep the canary through several scrape intervals and representative load. Stabilization is not instantaneous: old series remain active for a while, but new combinations should stop appearing.

text prometheus-validation-gates.txt
Keep the correction
New-series growth returns to a stable slope
Events per minute stay below the operational threshold
No new ingestion rejection
Critical targets remain up with regular scrapes
Critical dashboards and alerts are valid
Representative PromQL cost and duration improve
Detailed identifiers remain available in traces or logs

Roll back collection configuration
Series collide or duplicate after labeldrop
A critical alert is empty or semantically wrong
A target disappears after the scope change
No measurable effect on the intended limit
The correction cannot be tied to a proven cause

Rollback means reapplying the captured collection manifest, then verifying collector pickup and target recovery. If the application release is rolled back, also prove that the previous /metrics endpoint no longer emits the offending label. An image rollback without that check is incomplete.

Conclusion

A Managed Prometheus cardinality spike is first a production data-model problem. Bound the workspace, cluster and release, separate ingestion, scraping and query pressure, then prove which metric and label create the series before touching quotas.

The closing decision must remain reversible: repair instrumentation, temporarily filter a valueless label, remove duplicate collection, scale legitimate growth or roll back the release. Recovery is valid only when new-series growth stabilizes without losing alerts, targets or diagnostic power.