Infrastructure
Azure AKS: diagnose Workload Identity before bringing back a secret
A production runbook for isolating a Microsoft Entra Workload ID failure across the pod, ServiceAccount, webhook, OIDC token, managed identity and RBAC, then validating or rolling back without a client secret.
An AKS deployment completes, the pods become Ready, but the application can no longer read Key Vault, publish to Service Bus or access Storage. Logs show a token exchange failure, a 401 or a 403. The fastest response appears to be mounting a client secret back into Kubernetes or widening the managed identity role. Both actions bypass the diagnosis and permanently increase the risk surface.
This use case is an AKS workload authenticated with Microsoft Entra Workload ID. A Helm chart change renamed the ServiceAccount, moved the deployment to another namespace, changed an annotation or followed a cluster recreation. The runbook must locate the failure in pod mutation, projected token, federation trust, selected identity or authorization on the target resource. The expected output is a decision: fix the declaration, restore the previous release, repair one bounded federation or address a proven RBAC denial.
Freeze the scope before changing identity
Start by tying one pod to one identity and one target resource. A 403 alone cannot distinguish a missing token, a rejected OIDC exchange and a valid token without sufficient permissions.
incident:
cluster: aks-prod-weu
namespace: payments
deployment: settlement-api
pod: settlement-api-7c9f6d8f9b-abcde
service_account: settlement-api
first_failure_utc: 2026-08-15T14:20:00Z
expected_identity:
type: user_assigned_managed_identity
client_id: <expected-client-id>
tenant_id: <expected-tenant-id>
target:
service: Azure Key Vault
resource_id: <exact-resource-id>
operation: secrets/get
recent_changes:
- helm_release
- namespace_or_service_account
- federated_credential
- cluster_oidc_issuer
- role_assignment Keep the complete error, timestamp, pod name and correlation identifier from the target service. Also check whether every pod fails, only the new revision, or only one namespace. This matrix immediately narrows the search.
Prove that the pod requests Workload Identity
The ServiceAccount must carry the expected client annotation, and the pod template must carry azure.workload.identity/use: "true". The deployment must actually reference that ServiceAccount. Inspect rendered objects, not only Helm values.
NS=payments
DEPLOY=settlement-api
kubectl -n $NS get deploy $DEPLOY -o jsonpath='{.spec.template.spec.serviceAccountName}{"\n"}{.spec.template.metadata.labels}{"\n"}'
SA=$(kubectl -n $NS get deploy $DEPLOY -o jsonpath='{.spec.template.spec.serviceAccountName}')
kubectl -n $NS get serviceaccount $SA -o yaml
kubectl -n $NS get pods -l app=settlement-api -o custom-columns='NAME:.metadata.name,SA:.spec.serviceAccountName,WI:.metadata.labels.azure.workload.identity/use'
kubectl -n $NS get events --sort-by=.lastTimestamp | tail -n 40 A correct annotation on an unused ServiceAccount changes nothing. Updating an annotation after pod startup does not rewrite an existing pod either. Plan a controlled restart or a new revision, then inspect the pod that was actually created.
Verify mutation and the projected token
For an eligible pod, the webhook injects Azure variables and the volume containing the projected service account token. If these elements are absent, investigate the label, webhook or missed admission before looking at RBAC.
NS=payments
POD=settlement-api-7c9f6d8f9b-abcde
kubectl -n $NS get pod $POD -o jsonpath='{range .spec.containers[*]}{.name}{"\n"}{range .env[*]}{.name}={.value}{"\n"}{end}{end}' | grep -E 'AZURE_(CLIENT_ID|TENANT_ID|FEDERATED_TOKEN_FILE)|^settlement'
kubectl -n $NS get pod $POD -o jsonpath='{.spec.volumes}'
kubectl -n $NS exec $POD -- sh -c '
test -n "$AZURE_FEDERATED_TOKEN_FILE" &&
test -r "$AZURE_FEDERATED_TOKEN_FILE" &&
printf "client=%s tenant=%s token_file=readable\n" "$AZURE_CLIENT_ID" "$AZURE_TENANT_ID"
' Do not copy the raw token into the incident record. To inspect its claims, decode only the payload locally and retain only iss, sub, aud, iat and exp. The expected subject has the form system:serviceaccount:<namespace>:<serviceaccount>. Compare it exactly, including case and the issuer’s trailing slash.
Compare the three claims with the Entra federation
The trust relationship is a triplet: cluster issuer, Kubernetes subject and audience. A federated credential can be created successfully with the wrong namespace; the mistake appears only when token exchange occurs.
RG=rg-identities-prod
IDENTITY=mi-settlement-prod
AKS_RG=rg-aks-prod
AKS=aks-prod-weu
NS=payments
SA=settlement-api
ISSUER=$(az aks show -g $AKS_RG -n $AKS --query oidcIssuerProfile.issuerUrl -o tsv)
EXPECTED_SUBJECT="system:serviceaccount:$NS:$SA"
printf 'issuer=%s\nsubject=%s\naudience=api://AzureADTokenExchange\n' "$ISSUER" "$EXPECTED_SUBJECT"
az identity federated-credential list --resource-group $RG --identity-name $IDENTITY --query '[].{name:name,issuer:issuer,subject:subject,audiences:audiences}' -o table If the cluster was recreated, its issuer may have changed even though the cluster name did not. If the namespace or ServiceAccount was renamed, the historical subject no longer matches. The standard audience for direct federation is api://AzureADTokenExchange. Do not edit several fields as a test. Identify the exact mismatch and prepare its reversal.
A new federated credential can require a short propagation delay. That is not a reason for unbounded retries. Use a bounded window, retain its creation time and stop testing if the three claims do not match.
Separate token exchange from Azure authorization
Once an Entra token is issued, authorization on the target remains a separate step. An OIDC exchange error belongs to the federation chain. A 403 returned by Key Vault, Storage or Service Bus with the correct identity belongs to the role, data plane or service policy.
Projected variables or token absent
Check pod label, referenced ServiceAccount and webhook mutation
Token present, Entra exchange rejected
Compare issuer, subject, audience, tenant and client ID
Entra token issued, target returns 403
Prove principalId, operation and exact RBAC scope
Only one revision fails
Compare rendered manifests and restore the previous release
Every workload on the cluster fails
Check OIDC issuer, cluster configuration and shared change
Only one target resource fails
Inspect that resource's role, firewall and logs without changing federation Do not grant subscription-level Contributor as a connectivity test. It hides the cause, may not cover the data plane and leaves privilege that is hard to unwind. Query assignments for the identity at the exact scope and correlate them with target audit logs.
Apply a reversible correction
The fix must follow the evidence. If Helm renamed the ServiceAccount, restore the previous name or deliberately add the new federation. If the issuer belongs to the old cluster, add trust for the active cluster before removing the old one. If the client ID is wrong, correct the annotation and replace pods progressively.
change:
observed_mismatch: subject
intended_fix: restore_previous_service_account_name
scope: deployment/settlement-api
rollout: one_replica_then_progressive
preconditions:
- current_manifest_exported
- expected_identity_principal_recorded
- target_role_scope_recorded
- previous_helm_revision_available
stop_conditions:
- token_exchange_error_on_canary
- unexpected_client_id
- authorization_failure_on_previously_healthy_target
rollback:
command: helm rollback settlement-api <previous-revision> -n payments
preserve: [pod_events, application_logs, target_audit_logs, rendered_manifests] Do not add a “temporary” client secret to the same deployment. Credential chains may then select different mechanisms across environments, and a successful request no longer proves that Workload Identity works.
Validate before removing obsolete objects
Validate one fresh pod, then the full revision. Minimum evidence includes injected variables, a readable projected token, matching claims, the expected Entra identity, a successful business operation and target-side logs. Also test one intentionally forbidden operation to prove the role was not widened.
Positive validation
Fresh pod created with expected ServiceAccount
Workload identity label present
Projected variables and token present
Issuer, subject and audience match
Expected client ID and tenant
Key Vault read or business operation succeeds
No target-side errors during the validation window
Negative validation
An operation outside the role remains denied
No client secret added to the pod or pipeline
No broad role added at subscription scope
Deferred cleanup
Old federation removed only after observation
Old ServiceAccount removed after proving no consumers remain
Before and after evidence attached to the change The safest rollback restores the last known combination of deployment, ServiceAccount and federation. Deleting the previous trust object immediately removes that option and turns a narrow correction into an irreversible migration.
Conclusion
An AKS Workload Identity incident is a trust-chain diagnosis, not a generic permissions problem. The pod must request mutation, receive a projected token, present matching claims to the federation, select the expected identity and hold the minimum role on the target.
The decision then becomes explicit: fix the manifest when injection is absent, repair issuer-subject-audience when exchange fails, address RBAC only after proving identity, or restore the previous release when a change moved the trust boundary. Service returns without a long-lived secret, and the team keeps an operational validation and rollback path.