Automation

Azure Container Registry: diagnose incomplete image replication before retrying a multi-region rollout

A production runbook for separating ACR replication, regional routing, digests, identity and network paths before retrying or rolling back a container rollout.

13 Sept 2026 azurecontainer-registryacrgeo-replicationakscontainersdevopsci-cdobservabilitykqlautomationrunbookrollbackproduction

A pipeline pushes orders-api:2026.09.13.4 to a geo-replicated Azure Container Registry, then deploys that tag to two AKS clusters. The cluster closest to the build runner starts successfully. The other alternates between manifest unknown, ImagePullBackOff, and successful pulls several minutes later. Retrying the rollout may hide the incident, but it does not prove that every region serves the same artifact or that the next release is safe.

ACR replicates images asynchronously across active replicas. A completed push in one region therefore does not guarantee that an immediate pull routed to another region can already see the tag. This runbook separates five causes: replication still in progress, a rewritten tag, unexpected replica routing, an identity or network denial, and regional degradation. Its outcome is a decision to wait within a bound, promote a verified digest, temporarily exclude a replica from global routing, repair the client path, or roll back the release.

Freeze the artifact and timeline

A tag is a mutable pointer. Start by recording the digest produced by the pipeline, the UTC push completion time, the runner and cluster regions, one healthy pod, and one failing pod. Preserve the complete kubelet error: unauthorized, manifest unknown, blob unknown, 429, and timeout each open a different diagnostic branch.

yaml acr-multiregion-incident.yml
incident: inc-acr-20260913-021
registry: acrplatformprod
repository: orders-api
tag: 2026.09.13.4
expected_digest: sha256:<digest-from-build>
push_completed_utc: 2026-09-13T14:02:18Z

producer:
runtime: azure-devops-runner-weu
region: westeurope

consumers:
- cluster: aks-orders-weu
  region: westeurope
  result: success
- cluster: aks-orders-neu
  region: northeurope
  result: manifest-unknown

preserve:
- pipeline run and push output
- immutable digest from the build
- pod events and node identity
- registry replication status
- DNS answers from each runtime
- ACR login and repository events
- last known good deployment digest

Freeze tag rewrites during diagnosis. If two pipelines push the same tag through different replicas, regions can temporarily resolve that tag to different digests. A retry that still uses the tag then introduces another variable instead of producing evidence.

Read the declared replica state

ACR geo-replication requires the Premium tier and uses an active-active model: every available replica can accept reads and writes. The global endpoint selects a region from the client’s network profile and service health. Do not assume that the runner pushes to the registry’s home region.

bash 01-acr-replication-state.sh
ACR="acrplatformprod"
RG="rg-platform-prod"

az acr show --name "$ACR" --resource-group "$RG" --query "{loginServer:loginServer,sku:sku.name,provisioningState:provisioningState,dataEndpointEnabled:dataEndpointEnabled}" --output yaml

az acr replication list --registry "$ACR" --query "[].{region:location,status:provisioningState,globalRouting:regionEndpointEnabled,zoneRedundancy:zoneRedundancy}" --output table

az acr check-health --name "$ACR" --ignore-errors --yes

A healthy provisioningState is necessary, but it does not prove that the expected digest is already visible from every client path. Check Resource Health and ACR metrics for the same window. An Online replica can still be throttled; health-aware failover does not reroute traffic merely because requests return 429.

Prove the digest, not only the tag

Query the manifest from a controlled context, then pull the image by digest. The digest connects the build, registry, and deployment without relying on a mutable tag.

bash 02-verify-acr-digest.sh
ACR="acrplatformprod"
LOGIN_SERVER="$ACR.azurecr.io"
REPOSITORY="orders-api"
TAG="2026.09.13.4"
EXPECTED="sha256:<digest-from-build>"

az acr manifest show-metadata --registry "$ACR" --name "$REPOSITORY:$TAG" --query "{digest:digest,created:createdTime,lastUpdate:lastUpdateTime}" --output yaml

docker pull "$LOGIN_SERVER/$REPOSITORY@$EXPECTED"
docker image inspect "$LOGIN_SERVER/$REPOSITORY@$EXPECTED" --format '{{json .RepoDigests}}'

Run the pull from a runner in every consumer region or from a diagnostic pod subject to the same DNS, egress rules, and identity as the workload. Success from an administrator workstation does not validate the AKS path. Failure by tag followed by success by digest indicates tag visibility or mutation. The same failure by digest points toward replication, routing, network, or availability.

Identify the replica path actually served

The global acrplatformprod.azurecr.io endpoint is routed. Compare DNS answers from the build runner and every cluster. Different answers are expected; a long-lived client cache can prevent a workload from following a replica exclusion or failover.

bash 03-acr-runtime-path.sh
LOGIN_SERVER="acrplatformprod.azurecr.io"

getent ahosts "$LOGIN_SERVER"
dig "$LOGIN_SERVER" CNAME +short
dig "$LOGIN_SERVER" A +short

curl -sS -D- -o /dev/null "https://$LOGIN_SERVER/v2/"

kubectl -n orders describe pod <failing-pod>
kubectl -n orders get events --sort-by=.lastTimestamp --field-selector involvedObject.name=<failing-pod>

A 401 response from /v2/ without a token proves that DNS, TCP, TLS, and the registry response path work; it is not proof that workload authentication failed. If the registry uses dedicated data endpoints, the firewall or proxy must allow the data endpoint for every replicated region. An optional Private Endpoint changes DNS and network routing, but it does not change asynchronous replication or the need to validate the digest.

Read ACR events with their region

ACR diagnostic settings can populate ContainerRegistryRepositoryEvents and ContainerRegistryLoginEvents. Use them to compare operations, regions, identities, and results across the pipeline and AKS event window.

kusto 04-acr-regional-evidence.kql
let StartTime = datetime(2026-09-13T13:50:00Z);
let EndTime = datetime(2026-09-13T14:30:00Z);
union isfuzzy=true ContainerRegistryRepositoryEvents, ContainerRegistryLoginEvents
| where TimeGenerated between (StartTime .. EndTime)
| where LoginServer startswith "acrplatformprod"
| where isempty(Repository) or Repository == "orders-api"
| project TimeGenerated,
        Region,
        OperationName,
        Repository,
        Tag,
        Digest,
        Identity,
        CallerIpAddress,
        ResultType,
        ResultDescription,
        DurationMs
| order by TimeGenerated asc

Map the columns to the schema present in your workspace. Look for a sequence, not an isolated line: a completed push in one region, manifest reads in another, transient failures, then success. If events show Unauthorized for one identity, do not classify the incident as replication lag. If pulls succeed but slow down and return 429, measure load per replica before adding concurrent retries.

Add a multi-region promotion gate

The pipeline should not treat docker push as the end of global promotion. It should publish the digest, then execute a canary in every target region before opening the rollout.

yaml acr-multiregion-promotion-gate.yml
artifact:
repository: orders-api
tag: 2026.09.13.4
digest: sha256:<immutable-digest>

regional_canaries:
- region: westeurope
  pull_by_digest: required
  start_container: required
- region: northeurope
  pull_by_digest: required
  start_container: required

promote_when:
- every region pulls the expected digest
- the tag resolves to the same digest where tags remain in use
- no unauthorized, manifest-unknown or blob-unknown event remains
- pull latency and throttling stay inside the deployment budget

stop_when:
- a region observes another digest
- replication or Resource Health is degraded
- the data endpoint is unreachable from a target runtime
- retries increase load without changing the result

rollback:
deployment_digest: sha256:<last-known-good>
keep_failed_digest_for_investigation: true

The canary must start the container, not merely fetch its manifest. That validates layers, extraction, and runtime constraints. Deploy Kubernetes manifests by digest afterward. Keep the tag as a human-facing label, not as the release identity.

Decide between waiting, repair, exclusion, and rollback

If all replicas are healthy and the same digest becomes visible within the accepted window, wait with bounded backoff and resume one region at a time. If the tag resolves to different digests, stop concurrent producers, select the authorized artifact, and publish one tag without opening the rollout before canaries pass.

If one replica alone serves errors and Resource Health or logs identify it, you can temporarily exclude it from global routing with az acr replication update --global-endpoint-routing false. This does not delete the replica or stop synchronization, but it shifts client traffic. Validate capacity in the remaining regions, DNS cache behavior, and a re-enable plan. Regional endpoints, where used, do not receive this automatic rerouting.

Repair identity or network controls when the digest is present but the runtime receives Unauthorized, a firewall denial, or a timeout. Do not broaden AcrPull, IP ranges, or public access to address an unqualified manifest unknown error.

Roll back the workload to the last known digest when propagation exceeds the change budget, the observed artifact differs from the approved build, or no standby region has proved its capacity. A workload rollback should not immediately delete the failed digest. Retain it with the logs required for analysis.

Conclusion

An intermittent multi-region ACR pull is not resolved by a lucky retry. The investigation must connect the build digest, push region, serving replica, node identity, ACR events, and regional canary result.

The resulting decision is explicit: wait for bounded propagation, correct a mutable tag, repair the runtime path, temporarily remove a replica from routing, or return to the previous digest. Resume deployment only when every region pulls and starts the exact approved artifact while rollback remains immediately executable.