Cloud
Azure Container Apps: diagnose a stuck revision before redeploying
A production runbook for separating image pull, startup, probes, configuration, secrets, resources and platform state when a Container Apps revision never becomes healthy.
The pipeline has published a new Azure Container Apps image, but the revision remains in Processing, moves to Degraded, or ends in Failed. The previous revision is still serving production. Under pressure, teams often replay the exact deployment or keep increasing CPU, memory and probe delays until the platform accepts the container.
Those actions can hide the cause and produce more unusable revisions. A stuck revision can mean that the image cannot be pulled, the process exits, a probe disagrees with the actual port, a secret reference is missing, a volume cannot mount, or a startup dependency is unreachable. This runbook keeps the previous revision as a control and decides whether to fix one parameter, rebuild the candidate from stable, restore the last manifest, or escalate a platform incident.
Freeze the candidate and protect the stable revision
First name the exact revision, image and change. Do not collapse the ARM deployment state, revision provisioning state, running state and live traffic into one status.
incident: inc-20260916-004
resource_group: rg-orders-prod
container_app: ca-orders-api-prod
environment: cae-platform-prod
revision_mode: multiple
stable_revision: ca-orders-api-prod--r118
candidate_revision: ca-orders-api-prod--r119
stable_image: acrprod.azurecr.io/orders-api@sha256:<stable-digest>
candidate_image: acrprod.azurecr.io/orders-api@sha256:<candidate-digest>
candidate_created_at_utc: <timestamp>
observed_state:
arm_deployment: <succeeded|failed|unknown>
provisioning_state: <Provisioning|Provisioned|Failed>
running_state: <Processing|Running|Degraded|Failed|Unknown>
health_state: <Healthy|Unhealthy|None>
replicas: <count>
traffic_weight: <percent>
change_scope:
- image digest
- environment variables
- secrets or secret references
- probes and target port
- cpu and memory
- scale rules
- volumes
stop_conditions:
- stable revision is not healthy
- candidate already receives unexpected production traffic
- image digest or deployment manifest cannot be identified
- evidence collection would expose secret values An ARM Succeeded result proves that Azure accepted the configuration request, not that the container is ready. Conversely, a degraded candidate is not a reason to modify the stable revision. In multiple revisions mode, keep the candidate at 0 percent until its startup path is explained.
Read all four states before reading logs
Capture the app, revision list and candidate details. The provisioningError property is especially valuable because it can retain the first platform message before application logs exist.
RG="rg-orders-prod"
APP="ca-orders-api-prod"
CANDIDATE="ca-orders-api-prod--r119"
az containerapp show --resource-group "$RG" --name "$APP" --query '{revisionMode:properties.configuration.activeRevisionsMode,latestRevision:properties.latestRevisionName,traffic:properties.configuration.ingress.traffic,provisioningState:properties.provisioningState}' --output json
az containerapp revision list --resource-group "$RG" --name "$APP" --query '[].{name:name,created:properties.createdTime,active:properties.active,provisioning:properties.provisioningState,running:properties.runningState,health:properties.healthState,replicas:properties.replicas,traffic:properties.trafficWeight}' --output table
az containerapp revision show --resource-group "$RG" --name "$APP" --revision "$CANDIDATE" --query '{name:name,provisioning:properties.provisioningState,provisioningError:properties.provisioningError,running:properties.runningState,health:properties.healthState,replicas:properties.replicas,template:properties.template}' --output json Interpret the combination, not one field. Provisioning with Processing means the platform is still waiting for a usable candidate. Provisioned with Degraded means the resource exists but no ready replica satisfies the contract. Failed with a pull, mount or startup error gives the investigation a direct branch. If properties remain inconsistent or unknown across several apps in the same environment, preserve that signal for platform escalation.
Correlate system events with container output
System logs describe what Container Apps is trying to do. Console logs expose stdout and stderr when the process starts long enough to write. Query a narrow window around candidate creation.
let Revision = "ca-orders-api-prod--r119";
let Start = datetime(<candidate-created-at-utc>);
union isfuzzy=true
(
ContainerAppSystemLogs_CL
| where TimeGenerated >= Start
| where RevisionName_s == Revision
| project TimeGenerated, Source="system", Reason=Reason_s, Detail=Log_s
),
(
ContainerAppConsoleLogs_CL
| where TimeGenerated >= Start
| where RevisionName_s == Revision
| project TimeGenerated, Source="console", Reason="application", Detail=Log_s
)
| order by TimeGenerated asc Classify the first useful signal, not the last repeated error:
ErrImagePullor registry authentication: verify reference, digest, pull identity,AcrPull, DNS and registry egress.ContainerCrashingor a nonzero exit: read the first startup, entrypoint, arguments and required variables.Deployment Progress Deadline Exceededor zero ready replicas: compare startup and readiness probes with real startup time and the listening port.- volume or secret error: verify name, reference and mount without printing the value.
OOMKilledor repeated initialization restarts: compare limits, expected use and stable behavior before adding resources.
No console logs is evidence too. The pull might fail, the runtime might reject configuration, or the process might exit before logger initialization. Do not declare the application healthy because it wrote nothing.
Compare candidate and stable field by field
Compare effective templates, not only application commits. Export both revisions, reduce them to versioned fields and inspect the diff.
RG="rg-orders-prod"
APP="ca-orders-api-prod"
STABLE="ca-orders-api-prod--r118"
CANDIDATE="ca-orders-api-prod--r119"
az containerapp revision show -g "$RG" -n "$APP" --revision "$STABLE" --query 'properties.template' -o json > stable-template.json
az containerapp revision show -g "$RG" -n "$APP" --revision "$CANDIDATE" --query 'properties.template' -o json > candidate-template.json
jq -S . stable-template.json > stable-template.sorted.json
jq -S . candidate-template.json > candidate-template.sorted.json
diff -u stable-template.sorted.json candidate-template.sorted.json || true Start with differences that can prevent the first replica from becoming ready: image or digest, command, args, variables, secretRef, resources, ports, probes, volumes and scale rules. App-wide configuration can also affect startup without appearing in the template. Capture ingress, registries, named secrets and Dapr configuration separately when relevant.
Prove image pull without replaying deployment
A mutable tag does not prove which image was requested. Resolve the candidate to a digest and verify that the registry contains that manifest for the intended architecture. Then inspect the Container Apps pull identity and its AcrPull scope.
RG="rg-orders-prod"
APP="ca-orders-api-prod"
ACR="acrprod"
IMAGE="orders-api@sha256:<candidate-digest>"
az acr manifest show-metadata --registry "$ACR" --name "$IMAGE" --query '{digest:digest,createdAt:createdAt,imageSize:imageSize}' --output yaml
az containerapp show -g "$RG" -n "$APP" --query '{identity:identity,registries:properties.configuration.registries[].{server:server,identity:identity}}' --output json For a private registry, also test DNS and egress from the same network environment. A successful pull from a laptop or public runner does not prove that the Container Apps environment can reach ACR. Do not immediately replace managed identity with static credentials: that changes the security model and obscures the original failure.
Treat port, startup and probes as one contract
An application can start correctly and still remain unready because the port or probe does not describe its behavior. Compare four facts: process listening port, ingress targetPort, every probe port and actual initialization time.
Candidate startup contract
Process listens on: 8080
Ingress targetPort: 8080
Startup probe: TCP 8080, budget matches cold start
Readiness probe: HTTP /ready on 8080
Liveness probe: HTTP /live on 8080
Ready means
Process accepts requests
Required configuration is loaded
Critical local initialization is complete
Ready does not require
Every optional downstream system is available
An unbounded migration has completed
A write operation has succeeded Do not increase every delay blindly. Measure time to the first listening message and then to the first readiness success. If the candidate never binds the port, a more tolerant probe cannot repair a wrong entrypoint or crash. If startup is legitimately longer, give the startup probe an explicit budget while keeping readiness strict.
Check configuration, secrets and volumes without leaking data
List secret names and template references, never values in the ticket or logs. A valid-looking reference with a typo, a removed secret or a changed mount path can block the revision.
RG="rg-orders-prod"
APP="ca-orders-api-prod"
CANDIDATE="ca-orders-api-prod--r119"
az containerapp secret list -g "$RG" -n "$APP" --query '[].name' --output tsv | sort
az containerapp revision show -g "$RG" -n "$APP" --revision "$CANDIDATE" --query 'properties.template.{containers:containers[].{name:name,env:env[].{name:name,secretRef:secretRef},volumeMounts:volumeMounts},volumes:volumes}' --output json Validate identity permissions if the application loads Key Vault, Storage or another dependency during bootstrap. However, avoid making readiness depend on an optional remote system. An external outage could then prevent every replica from becoming ready and turn partial degradation into full unavailability.
Build a minimal candidate instead of replaying blindly
An identical retry adds evidence only when the hypothesis is a confirmed, bounded transient failure. Otherwise, copy the stable revision and apply one corrected difference. revision copy preserves a known base and makes the delta explicit.
RG="rg-orders-prod"
APP="ca-orders-api-prod"
STABLE="ca-orders-api-prod--r118"
FIXED_IMAGE="acrprod.azurecr.io/orders-api@sha256:<fixed-digest>"
az containerapp revision copy --resource-group "$RG" --name "$APP" --revision "$STABLE" --image "$FIXED_IMAGE" --output json Keep the new candidate out of public traffic. Validate state, system logs, startup, probes and a targeted smoke test through a label or the architecture’s planned validation path. Deactivate the failed candidate only after preserving its provisioningError, template and log window.
Decide fix, rollback or escalation
FIX A CANDIDATE
Cause is proven and bounded: digest, command, port, probe, secretRef or resources.
Copy stable, apply one fix, validate without traffic, then promote.
ROLL BACK THE MANIFEST
Several changes are mixed or the delta cannot be explained.
Redeploy the previously validated manifest and digests; keep stable serving.
RETRY ONCE
A dated external transient incident is confirmed with no configuration drift.
Retry with the same digest and compare traces; do not loop.
ESCALATE THE PLATFORM
Several apps in the environment fail, states remain inconsistent or no actionable error exists.
Preserve region, environment ID, revision, UTC timestamps, correlation IDs and provisioning errors.
STOP
Stable is degraded, candidate receives unexpected traffic or rollback is not executable.
Protect service before creating another revision. Final validation uses the same contract as diagnosis: Provisioned, Running, Healthy, a ready replica, no system-log loop, a correlated smoke test, and then observed progressive traffic. If errors return during promotion, move stable back to 100 percent without immediately deleting the candidate.
Conclusion
A stuck Container Apps revision is not a reason to redeploy until green. It is a failure that can be located between accepted configuration, pulled image, started process, satisfied probe and authorized traffic.
Keep stable as the control, capture state and the first system event, compare effective templates, then create a candidate with one proven change. The runbook must end with a clear outcome: promote after validation, restore the known manifest, retry once on confirmed transient evidence, or escalate with an actionable evidence pack.