Networking
AKS: diagnose a NetworkPolicy-blocked flow before loosening cluster traffic
A production runbook for proving whether AKS traffic is denied by Kubernetes or Cilium policy, using selectors, both flow directions, DNS checks, bounded evidence and rollback.
An AKS deployment completes, pods are ready, but one service can no longer call another. The quickest workaround is often an allow-all policy, a deleted default deny, or a wider namespace selector. That may restore the request while also removing the boundary the cluster was designed to enforce.
The use case is an orders-api pod in namespace apps-prod calling payments-api in payments-prod on TCP 8443. The failure appeared after a release that changed workload labels and network policies. This runbook must decide whether policy enforcement is the cause, which side of the flow is missing an allowance, and whether a narrow correction can be promoted or must be rolled back.
Freeze one failed flow
Treat connectivity as a tuple, not as “the service is unreachable.” Record the real source workload, destination, port, protocol and UTC window. Keep a working control flow from the same source if possible.
cluster: aks-platform-prod
source:
namespace: apps-prod
workload: deployment/orders-api
pod_label: app=orders-api
destination:
namespace: payments-prod
service: payments-api
port: 8443
protocol: TCP
incident_window: 2026-09-15T05:40:00Z/2026-09-15T06:10:00Z
change_context:
- application release
- label change
- NetworkPolicy change
decision:
promote_only_if: exact-flow-allowed-and-unrelated-flow-still-denied
rollback_on: selector-expands-or-required-boundary-cannot-be-proven Save the application commit, policy manifest commit and deployed image digest. A timeout alone does not identify policy: DNS, service endpoints, readiness, TLS and application listeners can produce the same symptom.
Prove the active data plane and policy types
Read the cluster configuration before choosing tooling. AKS can use different network and policy combinations. Cilium-specific commands and ACNS flow logs are useful only when the cluster actually uses that data plane and the relevant observability features are enabled.
RG="rg-platform-prod"
CLUSTER="aks-platform-prod"
az aks show --resource-group "$RG" --name "$CLUSTER" --query '{networkPlugin:networkProfile.networkPlugin,networkPluginMode:networkProfile.networkPluginMode,networkPolicy:networkProfile.networkPolicy,networkDataplane:networkProfile.networkDataplane}' --output yaml
kubectl get networkpolicy -A
kubectl get ciliumnetworkpolicy,ciliumclusterwidenetworkpolicy -A 2>/dev/null || true Do not assume a standard Kubernetes NetworkPolicy is the only enforcement layer. On Cilium clusters, namespaced and cluster-wide Cilium policies can coexist with standard policies. Azure subnet NSGs, UDRs and firewalls remain separate controls, but they should be investigated only after the in-cluster decision is qualified.
Reconstruct both sides of the allowance
Network policies are additive; there is no first-match ordering to inspect. A selected pod becomes isolated independently for ingress and egress. For a pod-to-pod connection, source egress and destination ingress must both allow the flow when both endpoints are isolated in those directions.
Source pod egress
Which policies select orders-api?
Do their combined egress rules allow payments-api on TCP 8443?
Is DNS allowed before the service name is resolved?
Destination pod ingress
Which policies select payments-api?
Do their combined ingress rules allow the source namespace and pod labels?
Does the Service targetPort match the listener?
Result
Allowed only when every isolated direction permits the tuple
Policy order and manifest filename do not create precedence This model avoids a common mistake: fixing destination ingress while source egress still denies the connection, then concluding that the policy engine is inconsistent.
Compare selectors with deployed labels
Policies act on current labels, not intended labels. Inspect the exact pods involved and the namespace labels used by selectors.
kubectl get pod -n apps-prod -l app=orders-api --show-labels -o wide
kubectl get pod -n payments-prod -l app=payments-api --show-labels -o wide
kubectl get namespace apps-prod payments-prod --show-labels
kubectl get networkpolicy -n apps-prod -o yaml > apps-prod-networkpolicies.yml
kubectl get networkpolicy -n payments-prod -o yaml > payments-prod-networkpolicies.yml
kubectl get service payments-api -n payments-prod -o yaml
kubectl get endpointslice -n payments-prod -l kubernetes.io/service-name=payments-api -o wide Compare podSelector, namespaceSelector, policyTypes, ports and protocols. Pay special attention to a release that renamed app, component, team or environment labels. An empty EndpointSlice or a mismatched targetPort is a service incident, not evidence of a policy denial.
Reproduce from the real security context
Run the smallest possible test from an affected pod. A temporary debug pod with different labels may not be selected by the same policies and can produce a false success.
SOURCE_POD=$(kubectl get pod -n apps-prod -l app=orders-api -o jsonpath='{.items[0].metadata.name}')
kubectl exec -n apps-prod "$SOURCE_POD" -- getent hosts payments-api.payments-prod.svc.cluster.local
kubectl exec -n apps-prod "$SOURCE_POD" -- sh -c 'nc -vz -w 3 payments-api.payments-prod.svc.cluster.local 8443'
# Keep one control using the same pod and protocol.
kubectl exec -n apps-prod "$SOURCE_POD" -- sh -c 'nc -vz -w 3 health-api.shared.svc.cluster.local 8443' Separate name resolution from TCP connection and application response. If DNS fails after an egress default deny, verify the cluster DNS or LocalDNS allowance before changing the business-service policy.
Read flow evidence when ACNS is available
On a Cilium data plane with Advanced Container Networking Services configured, on-demand Hubble observations or scoped container network logs can expose source, destination, protocol and verdict. Confirm feature state first; the absence of rows is not a deny verdict when capture was never enabled or did not select the flow.
For the incident tuple
Source and destination pod identities
Source and destination namespace
Destination port and protocol
FORWARDED, DROPPED or policy-denied verdict
Node carrying the source pod
Policy or rule identity when exposed
If no flow is visible
Confirm Cilium data plane
Confirm ACNS and capture mode
Confirm filters include the tuple
Repeat a timestamped bounded request
Fall back to selectors, endpoints and application evidence Use stored logs for history only when collection existed during the incident. For a live replay, prefer narrowly filtered on-demand observation. Avoid enabling broad flow capture during an outage without a volume and retention plan.
Build the narrow correction
Add an allowance for the required identity and port instead of removing isolation. The example below targets one source workload and one destination workload using explicit namespace labels. Adapt label keys to the labels your admission and deployment process actually guarantees.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-orders-api
namespace: payments-prod
spec:
podSelector:
matchLabels:
app: payments-api
policyTypes:
- Ingress
ingress:
- from:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: apps-prod
podSelector:
matchLabels:
app: orders-api
ports:
- protocol: TCP
port: 8443 If source egress is isolated, add the matching egress allowance there as a separate reviewed rule. Do not rely on ipBlock for pod or node IPs on Azure CNI Powered by Cilium; use namespace and pod identities for in-cluster workloads.
Validate positive and negative canaries
Apply the candidate from version control to a canary namespace or bounded workload. A successful allowed request is only half the test. Confirm that an unrelated source remains denied and that DNS, health probes and required platform traffic still work.
Positive canary
orders-api resolves payments-api
TCP 8443 connects
Application request returns expected status
No new DNS or reset regression
Negative canary
unrelated workload remains denied on TCP 8443
orders-api remains denied on an unapproved port
cross-namespace traffic outside the contract remains denied
Operational gate
Policy manifest is reviewed and versioned
Flow evidence matches the intended identities
Rollback manifest and owner are recorded Observe at least one representative request cycle after promotion. If the policy was correct and a label drift caused the outage, decide whether the durable fix belongs in the workload labels, the policy selector or the admission checks. Do not leave an emergency compatibility label indefinitely.
Decide, promote or roll back
promote_narrow_allowance:
when:
- failed_tuple_is_reproduced
- selected_policies_explain_the_deny
- positive_and_negative_canaries_pass
fix_workload_or_service:
when:
- labels_drifted_from_the_deployment_contract
- endpoints_or_target_port_are_wrong
- application_or_tls_fails_after_tcp_connects
hold_change:
when:
- active_policy_engine_or_selector_scope_is_unclear
- no_negative_canary_proves_the_boundary
rollback:
action:
- revert_the_candidate_manifest_commit
- reapply_the_previous_policy_set
- repeat_the_same_positive_and_negative_tests
- confirm_unrelated_traffic_did_not_gain_access Conclusion
A blocked AKS flow is not a reason to loosen the cluster. The useful proof connects one tuple to the active data plane, the policies selecting both endpoints, their deployed labels, service endpoints and an observed verdict when flow telemetry is available.
Promote only the smallest allowance that passes both positive and negative canaries. If policy is not the cause, keep the security boundary intact and move the incident to DNS, service discovery, TLS or the application. Either way, close with a reproducible validation and a versioned return path.