Infrastructure

Azure AKS: diagnose HPA oscillation before raising replica limits

A production runbook for separating noisy metrics, miscalibrated requests, slow startup and node capacity before changing Horizontal Pod Autoscaler bounds on AKS.

29 Sept 2026 azureakskuberneteshpaautoscalingmetrics-serverprometheusobservabilitycapacitycanaryrunbookrollbackproduction

An API on AKS moves from 4 to 12 pods, returns to 4 a few minutes later, then scales up again at the next peak. Latency degrades during warm-up, downstream connections multiply, and the team proposes raising minReplicas or maxReplicas to stabilize production.

That oscillation does not prove the bound is too low. The HPA may follow a metric that does not represent demand, calculate utilization from unrealistic CPU requests, wait for missing metrics, or request pods the cluster cannot schedule. The running case is an AKS Deployment scaling on CPU, taking 90 seconds to start, with a database that supports at most 120 application connections. The runbook must end with a decision: repair the signal, bound the scaling rate, add capacity, or roll back without hiding the problem behind permanent replicas.

Freeze a timeline before changing the bounds

Preserve a UTC window covering at least two complete scale-up and scale-down cycles. Export the HPA, Deployment, events, pods, and nodes. Record the release, Kubernetes version, metric provider, and any correlated deployment or traffic peak.

bash 01-capture-hpa-incident.sh
NS="checkout-prod"
HPA="checkout-api"
DEPLOY="checkout-api"

kubectl -n "$NS" get hpa "$HPA" -o yaml > hpa.yaml
kubectl -n "$NS" describe hpa "$HPA" > hpa.describe.txt
kubectl -n "$NS" get deploy "$DEPLOY" -o yaml > deployment.yaml
kubectl -n "$NS" get pods -l app=checkout-api -o wide > pods.txt
kubectl get nodes -o wide > nodes.txt
kubectl -n "$NS" get events --sort-by=.metadata.creationTimestamp > events.txt

Do not change requests, CPU target, stabilization windows, and bounds at the same time. The oscillation might disappear without revealing which contract was wrong. Stop the analysis if another controller also writes spec.replicas, such as a KEDA ScaledObject or deployment automation: two competing loops are not fixed by a different threshold.

Read the HPA as a control loop

The basic calculation compares the current value to the target, then applies that ratio to the replica count. For an average target, the order of magnitude is:

text hpa-ratio.txt
desiredReplicas = ceil(currentReplicas * currentMetric / targetMetric)

Example
current replicas: 4
observed average CPU: 90% of requests
target: 60%
raw recommendation: ceil(4 * 90 / 60) = 6

The visible result is not the formula alone. Initializing pods, missing metrics, tolerance, recent recommendations, minReplicas, maxReplicas, and behavior policies alter or delay the action. With multiple metrics, the HPA chooses the largest recommendation. If one metric is unavailable while the others recommend scaling down, it can skip the reduction.

Read status.conditions, currentMetrics, desiredReplicas, lastScaleTime, and events. AbleToScale, ScalingActive, and ScalingLimited answer different questions: whether the controller can act, whether the metric is usable, and whether a limit constrains the recommendation.

Prove the signal the controller consumes

An application dashboard does not prove what the controller reads. Query the metric API from the cluster and compare it with HPA status at the same instant.

bash 02-read-hpa-signal.sh
NS="checkout-prod"
HPA="checkout-api"

kubectl -n "$NS" get hpa "$HPA" -o jsonpath='{range .status.currentMetrics[*]}{.type}{"	"}{.resource.name}{"	"}{.resource.current.averageUtilization}{"
"}{end}'

kubectl get --raw '/apis/metrics.k8s.io/v1beta1/namespaces/'"$NS"'/pods' | jq -r '.items[] | select(.metadata.labels.app == "checkout-api") | [.metadata.name,.containers[].usage.cpu] | @tsv'

kubectl -n "$NS" top pods -l app=checkout-api --containers

For a custom or external metric, query custom.metrics.k8s.io or external.metrics.k8s.io. Verify its name, selector, unit, freshness, and cardinality. A series aggregated on the wrong label can scale one service from another service’s load. A cumulative counter used as a rate can keep increasing without representing current demand.

If kubectl top is empty or intermittent, inspect Metrics Server and its errors before editing the HPA. For a Prometheus adapter, verify the mapping rule, generated query, and delay from collection to exposure. Each recommendation should be traceable to a reproducible value, not merely a graph with a similar shape.

Verify requests and the per-replica envelope

A CPU averageUtilization target is calculated against requests. If a pod requests 100m and normally consumes 80m, the HPA already sees 80% even when the node has ample capacity. Conversely, an oversized request can delay useful scale-out.

bash 03-compare-requests-and-usage.sh
NS="checkout-prod"
DEPLOY="checkout-api"

kubectl -n "$NS" get deploy "$DEPLOY" -o jsonpath='{range .spec.template.spec.containers[*]}{.name}{"	request="}{.resources.requests.cpu}{"	limit="}{.resources.limits.cpu}{"
"}{end}'

kubectl -n "$NS" top pods -l app=checkout-api --containers
kubectl -n "$NS" get pods -l app=checkout-api -o jsonpath='{range .items[*]}{.metadata.name}{"	ready="}{.status.containerStatuses[0].ready}{"	restarts="}{.status.containerStatuses[0].restartCount}{"
"}{end}'

Calibrate requests from an observed distribution and load test, not from one snapshot. Measure the business capacity of one replica as well: requests per second, concurrency, latency, connection pool, and warm-up cost. Twelve pods with twenty connections each already exceed the downstream limit in the running case, even when the cluster can run them.

Separate oscillation, startup, and node shortage

A healthy scale-up can resemble oscillation when new pods are Running but remain unready. During warm-up, old pods carry the traffic, the metric remains high, and the HPA asks for more replicas. When all pods become ready, the average drops sharply and starts a scale-down.

Check startupProbe and readinessProbe, actual time to Ready, cache loading, and dependency errors. A pod must not receive traffic before it is ready, but its initialization must also fit the scale-up budget.

Then compare desiredReplicas, available replicas, and Pending pods. If the HPA asks for 12 pods but 4 remain Pending, the fault is scheduling or node capacity: excessive requests, affinity, taints, quota, node-pool bounds, or Cluster Autoscaler delay. Raising maxReplicas creates no capacity.

Reconstruct cycles with Prometheus

Exact names depend on the kube-state-metrics export. The useful view overlays recommendation, actual state, availability, and business signal.

promql 04-hpa-oscillation.promql
# Difference between recommendation and current replicas
kube_horizontalpodautoscaler_status_desired_replicas{
namespace="checkout-prod", horizontalpodautoscaler="checkout-api"
}
-
kube_horizontalpodautoscaler_status_current_replicas{
namespace="checkout-prod", horizontalpodautoscaler="checkout-api"
}

# Number of desired-replica changes over 15 minutes
changes(kube_horizontalpodautoscaler_status_desired_replicas{
namespace="checkout-prod", horizontalpodautoscaler="checkout-api"
}[15m])

# Unavailable Deployment pods
kube_deployment_status_replicas_unavailable{
namespace="checkout-prod", deployment="checkout-api"
}

Add CPU usage and requests, time to Ready, Pending pods, p95 latency, error rate, throughput, and downstream saturation. Oscillation matters when it degrades service or consumes a constrained resource. A replica-count change by itself is not a verdict.

Canary a behavior, not a magic value

Create a canary HPA on an isolated Deployment or namespace with the same load profile. Start by limiting scale-down velocity while keeping scale-up responsive. The following settings are an example to measure, not a universal default.

yaml 05-hpa-canary.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: checkout-api-canary
namespace: checkout-canary
spec:
scaleTargetRef:
  apiVersion: apps/v1
  kind: Deployment
  name: checkout-api-canary
minReplicas: 3
maxReplicas: 12
metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 60
behavior:
  scaleUp:
    stabilizationWindowSeconds: 0
    policies:
      - type: Percent
        value: 100
        periodSeconds: 60
  scaleDown:
    stabilizationWindowSeconds: 600
    policies:
      - type: Pods
        value: 1
        periodSeconds: 60

Replay a ramp, plateau, short spike, and return to idle. The canary must meet the latency target, produce no Pending pods, stay within the downstream connection budget, and return to baseline without repeated cycles. Also test temporary metric loss: behavior must remain explainable and must not cause an unsafe reduction.

Decide, validate, or roll back

Repair the signal when the metric does not represent demand or arrives too late. Recalibrate requests when they distort relative utilization. Increase scale-down stabilization or limit its velocity when short peaks and warm-up create cycles. Add node capacity only when useful pods remain Pending and scheduling constraints are intentional.

Raise minReplicas when the minimum capacity required for availability and startup time is measured. Raise maxReplicas only after proving that the cluster and dependencies support it. A higher bound without a downstream budget moves the incident to the database or partner API.

Roll back if the canary increases latency, errors, unavailable pods, downstream connections, or cost without reducing the cycles. Restore the known HPA manifest and requests, stop the canary, then verify that currentReplicas, desiredReplicas, and available replicas converge for one complete observation window.

Conclusion

An oscillating HPA is a control loop receiving a signal, applying a contract, and acting on finite capacity. The pod count is only its visible output.

Freeze the timeline, prove the metric, calibrate requests, measure warm-up, and verify scheduling before changing the bounds. The production decision then becomes defensible: repair the signal, dampen scale-down, provide nodes, raise a proven limit, or return to the last stable behavior.