Networking

AKS: diagnose a NetworkPolicy-blocked flow before loosening cluster traffic

A production runbook for proving whether AKS traffic is denied by Kubernetes or Cilium policy, using selectors, both flow directions, DNS checks, bounded evidence and rollback.

15 Sept 2026 azureakskubernetesnetwork-policyciliumacnsobservabilitysecurityrunbookrollbackproduction

An AKS deployment completes, pods are ready, but one service can no longer call another. The quickest workaround is often an allow-all policy, a deleted default deny, or a wider namespace selector. That may restore the request while also removing the boundary the cluster was designed to enforce.

The use case is an orders-api pod in namespace apps-prod calling payments-api in payments-prod on TCP 8443. The failure appeared after a release that changed workload labels and network policies. This runbook must decide whether policy enforcement is the cause, which side of the flow is missing an allowance, and whether a narrow correction can be promoted or must be rolled back.

Freeze one failed flow

Treat connectivity as a tuple, not as “the service is unreachable.” Record the real source workload, destination, port, protocol and UTC window. Keep a working control flow from the same source if possible.

yaml network-policy-incident.yml
cluster: aks-platform-prod
source:
namespace: apps-prod
workload: deployment/orders-api
pod_label: app=orders-api
destination:
namespace: payments-prod
service: payments-api
port: 8443
protocol: TCP
incident_window: 2026-09-15T05:40:00Z/2026-09-15T06:10:00Z

change_context:
- application release
- label change
- NetworkPolicy change

decision:
promote_only_if: exact-flow-allowed-and-unrelated-flow-still-denied
rollback_on: selector-expands-or-required-boundary-cannot-be-proven

Save the application commit, policy manifest commit and deployed image digest. A timeout alone does not identify policy: DNS, service endpoints, readiness, TLS and application listeners can produce the same symptom.

Prove the active data plane and policy types

Read the cluster configuration before choosing tooling. AKS can use different network and policy combinations. Cilium-specific commands and ACNS flow logs are useful only when the cluster actually uses that data plane and the relevant observability features are enabled.

bash 01-read-aks-network-profile.sh
RG="rg-platform-prod"
CLUSTER="aks-platform-prod"

az aks show --resource-group "$RG" --name "$CLUSTER" --query '{networkPlugin:networkProfile.networkPlugin,networkPluginMode:networkProfile.networkPluginMode,networkPolicy:networkProfile.networkPolicy,networkDataplane:networkProfile.networkDataplane}' --output yaml

kubectl get networkpolicy -A
kubectl get ciliumnetworkpolicy,ciliumclusterwidenetworkpolicy -A 2>/dev/null || true

Do not assume a standard Kubernetes NetworkPolicy is the only enforcement layer. On Cilium clusters, namespaced and cluster-wide Cilium policies can coexist with standard policies. Azure subnet NSGs, UDRs and firewalls remain separate controls, but they should be investigated only after the in-cluster decision is qualified.

Reconstruct both sides of the allowance

Network policies are additive; there is no first-match ordering to inspect. A selected pod becomes isolated independently for ingress and egress. For a pod-to-pod connection, source egress and destination ingress must both allow the flow when both endpoints are isolated in those directions.

text policy-decision-model.txt
Source pod egress
Which policies select orders-api?
Do their combined egress rules allow payments-api on TCP 8443?
Is DNS allowed before the service name is resolved?

Destination pod ingress
Which policies select payments-api?
Do their combined ingress rules allow the source namespace and pod labels?
Does the Service targetPort match the listener?

Result
Allowed only when every isolated direction permits the tuple
Policy order and manifest filename do not create precedence

This model avoids a common mistake: fixing destination ingress while source egress still denies the connection, then concluding that the policy engine is inconsistent.

Compare selectors with deployed labels

Policies act on current labels, not intended labels. Inspect the exact pods involved and the namespace labels used by selectors.

bash 02-inventory-selectors.sh
kubectl get pod -n apps-prod -l app=orders-api --show-labels -o wide
kubectl get pod -n payments-prod -l app=payments-api --show-labels -o wide
kubectl get namespace apps-prod payments-prod --show-labels

kubectl get networkpolicy -n apps-prod -o yaml > apps-prod-networkpolicies.yml
kubectl get networkpolicy -n payments-prod -o yaml > payments-prod-networkpolicies.yml

kubectl get service payments-api -n payments-prod -o yaml
kubectl get endpointslice -n payments-prod -l kubernetes.io/service-name=payments-api -o wide

Compare podSelector, namespaceSelector, policyTypes, ports and protocols. Pay special attention to a release that renamed app, component, team or environment labels. An empty EndpointSlice or a mismatched targetPort is a service incident, not evidence of a policy denial.

Reproduce from the real security context

Run the smallest possible test from an affected pod. A temporary debug pod with different labels may not be selected by the same policies and can produce a false success.

bash 03-replay-flow.sh
SOURCE_POD=$(kubectl get pod -n apps-prod -l app=orders-api -o jsonpath='{.items[0].metadata.name}')

kubectl exec -n apps-prod "$SOURCE_POD" -- getent hosts payments-api.payments-prod.svc.cluster.local
kubectl exec -n apps-prod "$SOURCE_POD" -- sh -c 'nc -vz -w 3 payments-api.payments-prod.svc.cluster.local 8443'

# Keep one control using the same pod and protocol.
kubectl exec -n apps-prod "$SOURCE_POD" -- sh -c 'nc -vz -w 3 health-api.shared.svc.cluster.local 8443'

Separate name resolution from TCP connection and application response. If DNS fails after an egress default deny, verify the cluster DNS or LocalDNS allowance before changing the business-service policy.

Read flow evidence when ACNS is available

On a Cilium data plane with Advanced Container Networking Services configured, on-demand Hubble observations or scoped container network logs can expose source, destination, protocol and verdict. Confirm feature state first; the absence of rows is not a deny verdict when capture was never enabled or did not select the flow.

text flow-evidence-checklist.txt
For the incident tuple
Source and destination pod identities
Source and destination namespace
Destination port and protocol
FORWARDED, DROPPED or policy-denied verdict
Node carrying the source pod
Policy or rule identity when exposed

If no flow is visible
Confirm Cilium data plane
Confirm ACNS and capture mode
Confirm filters include the tuple
Repeat a timestamped bounded request
Fall back to selectors, endpoints and application evidence

Use stored logs for history only when collection existed during the incident. For a live replay, prefer narrowly filtered on-demand observation. Avoid enabling broad flow capture during an outage without a volume and retention plan.

Build the narrow correction

Add an allowance for the required identity and port instead of removing isolation. The example below targets one source workload and one destination workload using explicit namespace labels. Adapt label keys to the labels your admission and deployment process actually guarantees.

yaml allow-orders-to-payments.yml
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-orders-api
namespace: payments-prod
spec:
podSelector:
  matchLabels:
    app: payments-api
policyTypes:
- Ingress
ingress:
- from:
  - namespaceSelector:
      matchLabels:
        kubernetes.io/metadata.name: apps-prod
    podSelector:
      matchLabels:
        app: orders-api
  ports:
  - protocol: TCP
    port: 8443

If source egress is isolated, add the matching egress allowance there as a separate reviewed rule. Do not rely on ipBlock for pod or node IPs on Azure CNI Powered by Cilium; use namespace and pod identities for in-cluster workloads.

Validate positive and negative canaries

Apply the candidate from version control to a canary namespace or bounded workload. A successful allowed request is only half the test. Confirm that an unrelated source remains denied and that DNS, health probes and required platform traffic still work.

text policy-canary-gates.txt
Positive canary
orders-api resolves payments-api
TCP 8443 connects
Application request returns expected status
No new DNS or reset regression

Negative canary
unrelated workload remains denied on TCP 8443
orders-api remains denied on an unapproved port
cross-namespace traffic outside the contract remains denied

Operational gate
Policy manifest is reviewed and versioned
Flow evidence matches the intended identities
Rollback manifest and owner are recorded

Observe at least one representative request cycle after promotion. If the policy was correct and a label drift caused the outage, decide whether the durable fix belongs in the workload labels, the policy selector or the admission checks. Do not leave an emergency compatibility label indefinitely.

Decide, promote or roll back

yaml network-policy-decision.yml
promote_narrow_allowance:
when:
- failed_tuple_is_reproduced
- selected_policies_explain_the_deny
- positive_and_negative_canaries_pass

fix_workload_or_service:
when:
- labels_drifted_from_the_deployment_contract
- endpoints_or_target_port_are_wrong
- application_or_tls_fails_after_tcp_connects

hold_change:
when:
- active_policy_engine_or_selector_scope_is_unclear
- no_negative_canary_proves_the_boundary

rollback:
action:
- revert_the_candidate_manifest_commit
- reapply_the_previous_policy_set
- repeat_the_same_positive_and_negative_tests
- confirm_unrelated_traffic_did_not_gain_access

Conclusion

A blocked AKS flow is not a reason to loosen the cluster. The useful proof connects one tuple to the active data plane, the policies selecting both endpoints, their deployed labels, service endpoints and an observed verdict when flow telemetry is available.

Promote only the smallest allowance that passes both positive and negative canaries. If policy is not the cause, keep the security boundary intact and move the incident to DNS, service discovery, TLS or the application. Either way, close with a reproducible validation and a versioned return path.