Networking
AKS: diagnose intermittent DNS before restarting CoreDNS
A production runbook for separating pod resolver, CoreDNS, service path, upstream DNS and node-specific failures before restarting or rolling back AKS DNS.
An AKS workload starts reporting intermittent ENOTFOUND, SERVFAIL or DNS timeouts. A retry usually succeeds, readiness remains green, and only some pods appear affected. Restarting CoreDNS is tempting because it is fast and visible. It can also erase the failing cohort, clear useful counters and hide a node, query-shape or upstream problem until the next peak.
The running case is an API in namespace orders-prod that resolves an internal dependency and an Azure service name. Failures increased after a deployment, but no single layer has been proven responsible. This runbook preserves the evidence, tests the same names from controlled cohorts, separates the Kubernetes DNS service from its upstream resolvers, and permits restart or rollback only when the failed layer is known.
Freeze a DNS failure contract
Do not begin with “DNS is flaky.” Capture the exact question and the runtime context. A short name, a cluster service FQDN and an external FQDN travel through different search and forwarding paths.
cluster: aks-platform-prod
window_utc: <start>/<end>
client:
namespace: orders-prod
workload: deployment/orders-api
pod: <pod-name>
node: <node-name>
questions:
- name: payments.payments-prod.svc.cluster.local
type: A
- name: <internal-zone-name>
type: A
- name: <azure-service-fqdn>
type: A
observed_errors: [timeout, SERVFAIL]
recent_changes:
- application image or resolver library
- CoreDNS ConfigMap or autoscaling
- NetworkPolicy or Cilium policy
- node image, pool or routing
- upstream DNS or private resolver
stop_condition: no cluster-wide restart until the failing layer is identified Preserve the application release, CoreDNS Deployment revision, ConfigMap resource version, node image versions and policy manifests. Record whether failures are concentrated by node, namespace, name suffix, response type or time window.
Reproduce from the workload and a control cohort
A successful lookup from an administrator workstation proves nothing about the pod path. Read the resolver configuration inside an affected pod, then compare it with a healthy pod using the same image and namespace.
kubectl -n orders-prod exec <affected-pod> -- cat /etc/resolv.conf
kubectl -n orders-prod exec <affected-pod> -- getent hosts payments.payments-prod.svc.cluster.local
kubectl -n orders-prod exec <affected-pod> -- getent hosts <internal-zone-name>
kubectl -n orders-prod get pod -o wide
kubectl -n orders-prod get networkpolicy
kubectl -n orders-prod get events --sort-by=.lastTimestamp If the application image has no diagnostic resolver, use an approved ephemeral container or diagnostic pod with the same namespace and scheduling constraints. Do not add packages to a production container during the incident.
Run a small, time-bounded series of queries rather than one nslookup. Keep latency, return code and answer for each attempt. Compare four cohorts: affected pod, healthy pod on the same node, healthy pod on another node, and a diagnostic pod in another namespace. This matrix distinguishes workload configuration, namespace policy, node path and cluster-wide failure.
Measure query amplification before blaming capacity
Read search, nameserver, options ndots and timeout from /etc/resolv.conf. A client that resolves a dotted but non-absolute name may try several search suffixes before the intended name. One application request can therefore produce several DNS questions, especially when libraries retry independently.
Capture the exact names sent by the application if tracing or resolver logs already expose them. Compare request rate with release traffic and DNS error rate. Do not enable verbose CoreDNS logging across the whole cluster without a retention and volume bound; a narrow canary or short capture window is safer.
Prefer FQDNs for external dependencies where the application and library support them. Treat changes to ndots, search domains or client retry behavior as application configuration changes: canary them and retain a rollback. They are not universal CoreDNS fixes.
Prove the Kubernetes DNS service path
Inspect the service, its ready endpoints and the placement of CoreDNS pods before touching the deployment.
kubectl -n kube-system get service kube-dns -o wide
kubectl -n kube-system get endpointslice -l kubernetes.io/service-name=kube-dns -o wide
kubectl -n kube-system get deployment coredns -o wide
kubectl -n kube-system get pods -l k8s-app=kube-dns -o wide
kubectl -n kube-system describe deployment coredns
kubectl -n kube-system get events --sort-by=.lastTimestamp Verify that ready EndpointSlices match healthy CoreDNS pods and that failures are not tied to one endpoint or node. Inspect restarts, readiness failures, throttling, memory pressure and uneven placement. If cluster DNS uses a different service or managed add-on layout, discover the deployed objects first rather than assuming names.
A healthy CoreDNS pod does not prove that pod-to-service traffic reaches it. Check the active AKS network data plane and policy implementation. UDP 53 is the common path, but large or truncated responses can fall back to TCP 53. A policy that permits UDP only can look intermittent because small answers succeed and larger answers fail.
Separate cluster names from upstream names
Test a Kubernetes service FQDN and an upstream name in the same sample window.
- If cluster service names fail, focus on the pod resolver, DNS service, endpoints, policy and node data plane.
- If cluster names succeed but every upstream suffix fails, inspect CoreDNS forwarding and its upstream reachability.
- If one suffix fails, inspect the matching forwarding rule, private resolver, custom DNS server or zone delegation.
- If only one response family or large answer fails, test TCP fallback and path MTU evidence before changing timeouts.
Read the deployed CoreDNS configuration and any supported custom override without editing it live:
kubectl -n kube-system get configmap coredns -o yaml
kubectl -n kube-system get configmap coredns-custom -o yaml 2>/dev/null || true
kubectl -n kube-system logs deployment/coredns --since=20m --prefix
kubectl -n kube-system top pods -l k8s-app=kube-dns The optional custom ConfigMap may not exist. That absence is normal. Look for forwarding changes, malformed server blocks, reload errors, loops and timeouts. Correlate them with the original UTC window and name suffix. Logs outside that window are context, not proof.
For custom or hybrid DNS, query the upstream from an approved diagnostic path that shares CoreDNS network reachability. Verify both UDP and TCP 53, response latency, recursion behavior and the expected zone answer. Private Endpoint may be one consumer of a private zone, but it is not the diagnostic model: the evidence must follow the actual suffix and forwarding chain.
Isolate node-specific failures
When failures follow a node, compare that node with a healthy peer: image version, pressure conditions, recent reboots, network agent health, effective routes and policy events. Check whether every pod on the node is affected or only pods created by one workload.
Do not drain the node before recording its pod list, timestamps and DNS test matrix. A drain moves clients and can make the symptom disappear without identifying the faulty path. If containment requires cordoning, stop new placement first, keep the node available for bounded evidence, then drain only with workload disruption and capacity validated.
Node-local caches, if deliberately deployed, add another resolver hop. Prove whether the pod points to the node-local address, whether that listener is healthy, and whether it can reach cluster DNS or upstream. Do not assume NodeLocal DNSCache is enabled merely because the failure is node-specific.
Apply the narrowest reversible correction
Choose the correction from evidence:
- roll back an application release when query amplification or resolver behavior changed;
- roll back a CoreDNS override when errors began with its resource version;
- restore both UDP and TCP 53 in the exact policy path when policy verdicts prove the denial;
- repair one upstream forwarding rule or resolver path when only its suffix fails;
- scale CoreDNS only when sustained utilization, queueing or dropped-query evidence proves capacity pressure;
- recycle one unhealthy pod or node only after preserving the cohort evidence and validating remaining capacity.
A cluster-wide CoreDNS restart is an operational change, not a diagnostic command. If it is required to restore service, record it as containment, preserve the previous ReplicaSet and configuration, and continue the root-cause investigation. A successful restart does not validate the cause.
Gate recovery and rollback
Use the original names, affected namespaces and node cohorts for validation. Include a cluster service, each relevant upstream suffix, one expected negative answer and at least one TCP-capable lookup.
PROMOTE OR RESUME when
Repeated queries meet the agreed success and latency window
Results pass from every required node and namespace cohort
Cluster and upstream names follow their expected paths
UDP and TCP DNS are permitted where required
CoreDNS has stable ready endpoints without resource pressure
The original application request succeeds without hidden retry growth
HOLD when
Success depends on retries or one resolver endpoint
A node, suffix or response type remains unexplained
Verbose logging or temporary access is still required
ROLL BACK when
The candidate change increases SERVFAIL, timeout or query volume
The correction widens policy beyond DNS dependencies
The previous known-good application or DNS configuration restores the contract Remove temporary logging, diagnostic pods and emergency policy changes after the observation window. Keep the test matrix and decision record with the incident so that a later recurrence starts with known discriminators.
Conclusion
Intermittent DNS in AKS is not a reason to restart CoreDNS by default. It is a path to decompose: client resolver and query shape, Kubernetes DNS service, CoreDNS endpoints, network policy and data plane, upstream forwarding, then node-specific behavior.
The production decision is explicit. Resume when the same query contract succeeds across required cohorts without retry amplification. Roll back the application, DNS configuration or policy that changed the failing layer. If the evidence still depends on a restart or a lucky retry, the incident is contained, not resolved.