Infrastructure

Azure AKS: diagnose a stuck rollout before forcing rollback

A production runbook for locating an AKS rollout stalled across Deployment, ReplicaSet, scheduling, image and readiness, then resuming or returning to a known revision with explicit stop conditions.

26 Aug 2026 azureakskubernetesdeploymentrolloutreadinessschedulingobservabilitydevopsautomationrunbookrollbackproduction

An AKS deployment can stall after the new image is published, while a few old pods still serve traffic and the delivery pipeline waits on kubectl rollout status. The tempting responses are immediate: delete pods, extend the timeout, rerun the pipeline, or force a rollback. Each one changes evidence and can turn partial degradation into a wider outage.

The use case is an API managed by a Kubernetes Deployment, with multiple replicas and a RollingUpdate strategy. The new revision never becomes available. This runbook finds the first object that stopped progressing, separates admission, capacity, image, startup and readiness failures, then decides whether to resume with a bounded correction or return to a known revision. A green pipeline is not the primary objective; controlled service recovery is.

Freeze the rollout and its impact

Name the change, the time window and the user-visible symptom first. A stalled rollout with every old pod still available is materially different from one that has already reduced useful capacity.

yaml rollout-incident-card.yml
scope:
cluster: aks-prod-weu
namespace: payments
deployment: payments-api
change: release-20260826.3
started_utc: 2026-08-26T06:42:00Z
expected:
replicas: 6
max_surge: 2
max_unavailable: 1
evidence:
previous_revision: 41
candidate_revision: 42
previous_image_digest: sha256:<known-good>
candidate_image_digest: sha256:<candidate>
stop_conditions:
- available replicas fall below 5
- error rate crosses the incident threshold
- no new pod reaches Ready during the observation window
rollback_requires:
- revision 41 still exists
- configuration remains compatible
- database changes are backward compatible

Keep image digests, not only tags. Record changes outside the Deployment too: a ConfigMap, Secret, admission policy, schema migration, or dependency. kubectl rollout undo will not restore any of them.

Read the Deployment, ReplicaSet and Pod chain

Do not start with application logs. Find the first control-plane object that failed to produce its expected state.

bash 01-capture-rollout-state.sh
NAMESPACE="payments"
DEPLOYMENT="payments-api"
EVIDENCE="aks-rollout-evidence"

mkdir -p "$EVIDENCE"

kubectl -n "$NAMESPACE" get deployment "$DEPLOYMENT" -o yaml > "$EVIDENCE/deployment.yaml"
kubectl -n "$NAMESPACE" describe deployment "$DEPLOYMENT" > "$EVIDENCE/deployment-describe.txt"
kubectl -n "$NAMESPACE" rollout history deployment/"$DEPLOYMENT" > "$EVIDENCE/rollout-history.txt"
kubectl -n "$NAMESPACE" get rs -o wide > "$EVIDENCE/replicasets.txt"
kubectl -n "$NAMESPACE" get pods -o wide > "$EVIDENCE/pods.txt"
kubectl -n "$NAMESPACE" get events --sort-by=.lastTimestamp > "$EVIDENCE/events.txt"

Read observedGeneration, updatedReplicas, readyReplicas, availableReplicas, plus the Progressing and Available conditions. Then identify the ReplicaSet that owns the candidate revision and compare desired, created and ready pods.

The first split is operationally useful: no new pod exists, a pod remains Pending, the container cannot start, or it runs without becoming Ready. Each state points to a different layer and a different owner.

No pod exists: admission, quota or manifest

When the new ReplicaSet exists but creates no pod, inspect its events before touching cluster capacity. A quota, LimitRange, admission policy or invalid field may reject creation. When there is no new ReplicaSet at all, return to Deployment conditions and the exact object applied by the delivery system.

bash 02-check-admission-and-quota.sh
kubectl -n "$NAMESPACE" get deployment "$DEPLOYMENT" -o jsonpath='{.metadata.generation}{" "}{.status.observedGeneration}{"
"}'

kubectl -n "$NAMESPACE" get rs -o custom-columns='NAME:.metadata.name,REVISION:.metadata.annotations.deployment.kubernetes.io/revision,DESIRED:.spec.replicas,CURRENT:.status.replicas,READY:.status.readyReplicas'

kubectl -n "$NAMESPACE" get resourcequota,limitrange
kubectl -n "$NAMESPACE" get events --sort-by=.lastTimestamp | tail -n 80

Fix an admission rejection in the manifest or through a reviewed, narrowly scoped exception. Disabling a webhook or policy cluster-wide to release one deployment removes the same safeguard from every other workload. If the admission service is unavailable, treat its failure policy as a separate platform incident.

Pending pod: prove the scheduling constraint

A Pending pod does not automatically mean insufficient CPU. kubectl describe pod records the scheduler decision: unavailable resources, an untolerated taint, impossible affinity, topology constraints, an unattachable volume, or exhausted pod slots on nodes.

bash 03-diagnose-pending-pod.sh
POD="<new-revision-pod>"

kubectl -n "$NAMESPACE" describe pod "$POD"
kubectl get nodes -o custom-columns='NODE:.metadata.name,READY:.status.conditions[?(@.type=="Ready")].status,TAINTS:.spec.taints,CAPACITY:.status.capacity.pods,ALLOCATABLE_CPU:.status.allocatable.cpu,ALLOCATABLE_MEMORY:.status.allocatable.memory'
kubectl top nodes
kubectl -n "$NAMESPACE" get pvc
kubectl get events --all-namespaces --sort-by=.lastTimestamp | tail -n 120

Correct the constraint the event actually proves. Temporary capacity can be defensible when the rollout budget remains intact and scale-out is observable. Removing affinity, tolerations or resource requests without understanding their purpose can instead place the pod on a node that violates the workload contract.

A PodDisruptionBudget does not, by itself, block the normal replacement performed by a Deployment. It matters when a node drain, node-pool upgrade or other voluntary eviction is happening at the same time. Do not delete a PDB merely because the rollout is slow; first find the eviction event that makes it relevant.

Waiting or restarting container: separate image from startup

For ErrImagePull or ImagePullBackOff, compare the digest reference, kubelet access to the registry and pod events. For CrashLoopBackOff, capture the last terminated state and previous container logs before deleting anything.

bash 04-check-container-start.sh
POD="<new-revision-pod>"
CONTAINER="api"

kubectl -n "$NAMESPACE" get pod "$POD" -o jsonpath='{range .status.containerStatuses[*]}{.name}{" image="}{.image}{" imageID="}{.imageID}{" waiting="}{.state.waiting.reason}{" last="}{.lastState.terminated.reason}{"
"}{end}'

kubectl -n "$NAMESPACE" describe pod "$POD"
kubectl -n "$NAMESPACE" logs "$POD" -c "$CONTAINER" --previous --timestamps --tail=300
kubectl -n "$NAMESPACE" logs "$POD" -c "$CONTAINER" --timestamps --tail=300

An unavailable image is not fixed by increasing progressDeadlineSeconds. A process exiting on invalid configuration is not fixed by deleting the pod; the ReplicaSet will recreate the same state. Correct the demonstrated failure in the image reference, ACR access, startup command or configuration.

Running but not Ready: test the probe from inside

A Running container can stay out of service because its startupProbe or readinessProbe fails. Read the probe type, port, path, delays and events. Then verify what the process actually listens on and which dependency the readiness endpoint exercises.

bash 05-diagnose-readiness.sh
kubectl -n "$NAMESPACE" get pod "$POD" -o jsonpath='{range .status.conditions[*]}{.type}{"="}{.status}{" reason="}{.reason}{"
"}{end}'

kubectl -n "$NAMESPACE" get pod "$POD" -o jsonpath='{.spec.containers[?(@.name=="api")].startupProbe}{"
"}{.spec.containers[?(@.name=="api")].readinessProbe}{"
"}'

kubectl -n "$NAMESPACE" exec "$POD" -c "$CONTAINER" -- sh -c 'wget -qSO- http://127.0.0.1:8080/ready || true'

kubectl -n "$NAMESPACE" get endpointslice -l kubernetes.io/service-name=payments-api -o wide

Do not stretch probe timings without measurement. If startup is genuinely slower, a bounded startupProbe adjustment can be correct. If /ready depends on an unstable external service, decide whether that dependency should remove the pod from traffic or only degrade one feature. Preserve that design intent during incident response.

Validate one corrected revision

Once the cause is isolated, avoid bundling more changes into the recovery. Apply one reviewed correction, watch the candidate ReplicaSet and retain old capacity until the new pods prove they can serve.

bash 06-validate-rollout.sh
kubectl -n "$NAMESPACE" rollout status deployment/"$DEPLOYMENT" --timeout=10m

kubectl -n "$NAMESPACE" get deployment "$DEPLOYMENT" -o custom-columns='NAME:.metadata.name,DESIRED:.spec.replicas,UPDATED:.status.updatedReplicas,READY:.status.readyReplicas,AVAILABLE:.status.availableReplicas,UNAVAILABLE:.status.unavailableReplicas'

kubectl -n "$NAMESPACE" get pods -o custom-columns='POD:.metadata.name,REVISION:.metadata.labels.pod-template-hash,READY:.status.containerStatuses[*].ready,RESTARTS:.status.containerStatuses[*].restartCount,IMAGE_ID:.status.containerStatuses[*].imageID'

kubectl -n "$NAMESPACE" get events --sort-by=.lastTimestamp | tail -n 80

Add a check from the user path: error rate, latency, dependencies and one representative transaction. A completed Kubernetes rollout proves availability according to probes. It does not prove that the business path works.

Decide resume, hold or rollback

Resume when one isolated cause and one bounded correction make candidate pods Ready, capacity remains above the stop threshold and the application check passes. Hold when several layers remain uncertain or when recovery would require disabling a cluster-wide control.

When impact grows, return to a known revision only after proving it is still retained and compatible with external changes.

bash 07-controlled-rollback.sh
kubectl -n "$NAMESPACE" rollout history deployment/"$DEPLOYMENT"
kubectl -n "$NAMESPACE" rollout undo deployment/"$DEPLOYMENT" --to-revision=41
kubectl -n "$NAMESPACE" rollout status deployment/"$DEPLOYMENT" --timeout=10m

kubectl -n "$NAMESPACE" get deployment,rs,pods -o wide
kubectl -n "$NAMESPACE" get events --sort-by=.lastTimestamp | tail -n 80

Deployment rollback restores the pod template, not a SQL migration, policy, replaced secret or deleted object. Run those return paths through their own procedures and dependency order. If the previous revision is no longer compatible, stop the automation and handle recovery as an explicit change.

Conclusion

A stuck AKS rollout is a chain-of-objects diagnosis, not a pipeline timeout. Deployment, ReplicaSet, Pod, scheduler, container runtime and readiness each provide evidence that narrows the next action.

The runbook must end in a decision: resume with one validated correction, keep the previous revision serving while evidence is completed, or execute a rollback whose dependencies remain compatible. Forcing another run before locating the first blocked state does not resolve the incident; it only overwrites its timeline.