Networking
Azure AKS: diagnose outbound SNAT exhaustion before adding a NAT Gateway
A production runbook for separating DNS, remote dependency, network policy, conntrack and Azure SNAT pressure on AKS before changing the cluster outbound path.
Several AKS workloads start reporting intermittent timeouts to a public API. Retries hide part of the failure, but latency rises and scheduled jobs miss their window. The destination is healthy, DNS still resolves, and adding replicas makes the symptom worse. The first proposal is to attach a NAT Gateway and obtain more outbound capacity.
That change may be appropriate. It can also move the cluster to a new source IP, invalidate partner allowlists and preserve the application behavior that exhausted the previous path. This runbook starts from the failing pod and proves whether Azure SNAT is the bottleneck before changing the AKS outbound architecture. It ends with one of three decisions: repair connection behavior, add measured egress capacity, or roll back the last workload or network change.
Freeze one comparable failure window
Record the UTC window, cluster, node pool, deployment revision, affected namespace, destination FQDN and port, observed error, retry count and source IP seen by the destination. Preserve one successful request and one failed request with correlation IDs when the remote service provides them.
Do not begin with a cluster-wide restart. SNAT pressure is often uneven across nodes. Restarting redistributes pods and releases connections, which can temporarily remove the evidence without removing the cause.
RG="rg-platform-prod"
AKS="aks-platform-prod"
az aks show --resource-group "$RG" --name "$AKS" --query "{nodeResourceGroup:nodeResourceGroup,outboundType:networkProfile.outboundType,loadBalancerSku:networkProfile.loadBalancerSku,loadBalancerProfile:networkProfile.loadBalancerProfile,natGatewayProfile:networkProfile.natGatewayProfile}" --output json > aks-egress-profile.json
kubectl get deploy,pods -A -o wide > workload-placement.txt
kubectl get events -A --sort-by=.lastTimestamp > kubernetes-events.txt The incident statement should remain testable: “between 14:05 and 14:18 UTC, pods from revision checkout-7f8c9 on two nodes timed out to api.partner.example:443; unaffected nodes completed the same probe.”
Prove the effective outbound path
Read networkProfile.outboundType before assuming that the AKS load balancer performs SNAT. A cluster can use loadBalancer, a managed or user-assigned NAT Gateway, userDefinedRouting, or an explicitly restricted outbound model. With UDR, the effective next hop may be Azure Firewall or another network virtual appliance, so Load Balancer SNAT metrics are not evidence for that path.
For loadBalancer, identify the AKS-managed load balancer and its outbound frontends in the node resource group. For a NAT Gateway, identify the gateway attached to every node-pool subnet involved in the incident. For UDR, capture the route table, effective 0.0.0.0/0 route, next hop and the public IP that the egress device actually uses.
NODE_RG="$(jq -r .nodeResourceGroup aks-egress-profile.json)"
az network lb list --resource-group "$NODE_RG" --query "[].{name:name,id:id,frontends:frontendIpConfigurations[].publicIpAddress.id,outboundRules:outboundRules[].name}" --output json > load-balancers.json
az network nat gateway list --query "[].{name:name,resourceGroup:resourceGroup,id:id,publicIps:publicIpAddresses[].id,subnets:subnets[].id}" --output json > nat-gateways.json
az network route-table list --query "[].{name:name,resourceGroup:resourceGroup,subnets:subnets[].id,routes:routes[].{prefix:addressPrefix,nextHop:nextHopType,nextHopIp:nextHopIpAddress}}" --output json > route-tables.json Compare that inventory with the source IP observed at the remote dependency. A mismatch means that the path model is incomplete: another NAT device, proxy, firewall or subnet association is participating.
Exclude failures that only resemble SNAT exhaustion
Test from the affected workload context, not from an administrator workstation. Use the same namespace, node, DNS policy, service account and destination when possible. Separate these signals:
- DNS: slow lookup, alternating answers or a private address reached from only part of the cluster;
- remote dependency: HTTP
429,503, TLS rejection or an allowlist denial with a completed TCP connection; - NetworkPolicy, NSG, UDR or firewall: consistent deny or timeout tied to a destination, subnet or rule;
- node conntrack or ephemeral ports: local connection pressure before Azure performs SNAT;
- Azure SNAT: failures opening new outbound flows to public endpoints while existing connections may continue.
Capture resolution, TCP/TLS connection time and application response separately. A single curl exit code collapses too many layers.
NS="checkout"
POD="checkout-7f8c9-abcde"
HOST="api.partner.example"
kubectl exec -n "$NS" "$POD" -- getent ahostsv4 "$HOST"
kubectl exec -n "$NS" "$POD" -- sh -c 'time curl -sS -o /dev/null -w "connect=%{time_connect} tls=%{time_appconnect} total=%{time_total} code=%{http_code}\n" https://api.partner.example/health'
kubectl get networkpolicy -A -o yaml > network-policies.yaml
kubectl get pod -n "$NS" "$POD" -o wide > affected-pod.txt If the application image has no diagnostic tools, use an approved ephemeral debug container or a predeployed probe with the same scheduling and network context. Do not permanently expand the production image for an incident.
Correlate Azure SNAT metrics by backend node
For a Standard Load Balancer outbound path, read UsedSnatPorts, AllocatedSnatPorts and SnatConnectionCount at one-minute granularity. Split by backend IP and protocol where the monitoring surface supports it. An average across the whole backend pool can hide one saturated node.
LB_ID="<aks-load-balancer-resource-id>"
START="2026-09-30T14:00:00Z"
END="2026-09-30T14:25:00Z"
az monitor metrics list --resource "$LB_ID" --metric UsedSnatPorts AllocatedSnatPorts SnatConnectionCount --interval PT1M --start-time "$START" --end-time "$END" --output json > load-balancer-snat-metrics.json Evidence is strong when the failing node approaches its allocated ports, failed SNAT connection counts rise in the same window, new public flows fail, and the workload shows high connection churn to a small set of destination tuples. Port usage below allocation does not by itself clear the network: verify the correct frontend, backend node, protocol and outbound device.
For NAT Gateway or a firewall path, use the metrics of that actual device and preserve the same node-to-pod correlation. Do not compare NAT Gateway capacity figures with Load Balancer port allocation as if they described the same mechanism.
Find the workload creating the flows
Map the affected backend IP to its AKS node, then list every pod placed on it. Compare current placement with the deployment change and the start of the incident.
NODE="aks-userpool-12345678-vmss00000a"
kubectl get node "$NODE" -o wide
kubectl get pods -A --field-selector spec.nodeName="$NODE" -o wide
kubectl top pods -A --containers --sort-by=cpu
kubectl get deploy,statefulset,job,cronjob -A -o wide > workload-owners.txt Look for a rollout that disabled keep-alive, reduced connection-pool reuse, shortened client lifetime, increased worker concurrency, synchronized jobs or added aggressive retries. A replica increase can multiply connection creation even when request volume is stable. Quantify new connections per second and concurrent connections by destination; request count alone is insufficient.
If node-level inspection is approved, capture a short, bounded connection or packet sample during the failure window. Filter to public IPv4 destinations and retain only the fields required for flow counts. Avoid collecting payloads or credentials.
Apply the smallest correction first
Prefer a workload correction when connection churn is abnormal: reuse HTTP clients and database pools, bound concurrency, add exponential backoff with jitter, respect server retry guidance and stop retries when their deadline can no longer succeed. Validate on one revision and keep the previous image digest available.
Add network capacity only after measuring legitimate concurrency. For a Load Balancer path, review explicit outbound rules, allocated ports per backend and the number of outbound IPs. For sustained high-scale egress, a NAT Gateway may provide a clearer capacity and source-IP model. With BYO networking, use the supported user-assigned design for the cluster subnet; do not attach an unrelated gateway without reconciling the AKS outbound configuration.
Treat an outbound-path change as a production migration:
- freeze the old and new public source IPs and update approved destination allowlists;
- inventory every node-pool subnet and route-table association;
- canary a representative workload and test DNS, TLS, dependency authorization and sustained connections;
- monitor connection failures, latency, port usage and application retries;
- retain the previous outbound configuration until the validation window closes.
More ports are not a substitute for bounded clients. Capacity can absorb an expected peak; it should not legitimize an unbounded retry or connection loop.
Validate, expand or roll back
Approve the workload fix when connection creation falls, existing throughput is preserved, timeouts disappear on previously affected nodes and Azure metrics retain operating headroom through a representative peak.
Approve a Load Balancer or NAT Gateway change only when the effective path, source IPs, subnet coverage, dependency allowlists and capacity model are documented; the canary passes; and the previous path remains reproducible.
Roll back the workload revision when the incident aligns with a connection-management change and the previous digest restores stable reuse. Roll back the network migration when source identity, routing or a downstream allowlist differs from the approved plan. If the old path no longer has enough capacity, contain demand first instead of failing back into known exhaustion.
Stop and escalate when the evidence points to an NVA or firewall owned by another team, when the destination tuple cannot be identified, or when different node pools use different egress paths that the change plan does not cover.
Conclusion
An intermittent AKS outbound timeout is not proof that the cluster needs a NAT Gateway. The useful diagnosis connects one pod failure to one node, one destination tuple and the actual egress device, then correlates that path with connection behavior and port pressure.
The production decision becomes explicit: repair connection reuse, add justified SNAT capacity, migrate the outbound path with a canary, or roll back the change that created the pressure. In every case, validation includes the application, the network device and the source identity seen by the dependency.