Cloud
Azure Container Apps with Dapr: diagnose service invocation before redeploying
A production runbook for separating app ID, Dapr sidecars, application port, mTLS, resiliency and revisions when service invocation fails in Azure Container Apps.
An orders-api service running in Azure Container Apps invokes payments-api through the Dapr service invocation API. After a new revision is released, calls return 500, 503, or time out. Both containers appear healthy, and redeploying occasionally restores the flow for a few minutes.
An immediate redeployment removes useful evidence and conflates several failure modes: a wrong app ID, a missing sidecar, an incorrect application port, a protocol mismatch, an unready target, an aggressive resiliency policy, or a bad revision. The runbook should first locate the break between the caller, its sidecar, the target sidecar, and the target application.
Freeze one representative invocation
Select an operation that reproduces the symptom without creating an irreversible business effect. Preserve UTC time, source revision, target revision, method, status, latency, and a correlation ID. A test from an administrator workstation is not enough: the production path starts inside the calling container.
incident: inc-20260817-001
environment: <container-apps-environment>
source_app: orders-api
source_revision: <revision>
source_dapr_app_id: orders-api
target_app: payments-api
target_revision: <revision>
target_dapr_app_id: payments-api
method: GET /health/dependencies
first_failure_utc: <timestamp>
last_known_healthy_utc: <timestamp>
correlation_id: <redacted-id>
preserve:
- caller application logs
- caller and target Dapr sidecar logs
- target application logs
- revision state and traffic weights
- Dapr configuration and resiliency changes
- deployment and secret or identity changes Freeze new revisions and Dapr component changes while evidence is collected. Do not rely on one failed request. Measure the error rate by revision so you can tell whether the fault follows a version, a replica, or the entire environment.
Verify the effective Dapr contract
In Container Apps, appId is the logical name used by callers. appPort is where the target sidecar reaches the application, while appProtocol defines that local exchange. Inspect the effective platform configuration rather than trusting only the IaC repository.
az containerapp show \
--resource-group <resource-group> \
--name orders-api \
--query '{dapr:properties.configuration.dapr,revisionsMode:properties.configuration.activeRevisionsMode}'
az containerapp show \
--resource-group <resource-group> \
--name payments-api \
--query '{dapr:properties.configuration.dapr,ingress:properties.configuration.ingress}'
az containerapp revision list \
--resource-group <resource-group> \
--name payments-api \
--query '[].{name:name,active:properties.active,health:properties.healthState,created:properties.createdTime,traffic:properties.trafficWeight}' \
--output table Compare app ID, application port, protocol, log level, and API logging. Then prove that the target process is listening on the declared port and on an interface the sidecar can reach. A container can pass its own probe while Dapr is configured for a different port.
The Dapr app ID is not an ingress FQDN and is not necessarily the business name assumed in code. An application rename, a stale Bicep or pipeline parameter, or duplicate logical IDs can break discovery without looking like a conventional network outage.
Test all four path segments
A Dapr invocation crosses four segments: source application to local sidecar, target discovery, target sidecar to target application, and the response path. Test them in that order.
From the same source revision, call the local sidecar when the image, or a controlled diagnostic revision, provides curl:
curl --fail-with-body \
--max-time 10 \
-H 'X-Correlation-ID: <incident-id>' \
'http://localhost:3500/v1.0/invoke/payments-api/method/health/dependencies' Read the result together with logs:
- no response on
localhost:3500points first to the source sidecar or its startup; - an app ID resolution error points to logical identity or target availability;
- connection failure on the application port points to
appPort, the listener, readiness, or target protocol; - an application status with a target trace proves Dapr delivered the request;
- a timeout after multiple attempts requires inspecting resiliency before blaming the network.
Do not permanently bypass Dapr with an ingress FQDN just because a direct call succeeds. That test omits service discovery, mTLS, retries, and telemetry that belong to the production contract.
Correlate sidecars, applications, and revisions
Console logs contain application container and Dapr sidecar output. System logs provide revision provisioning and component context. Start with the frozen invocation window, then widen it only when needed.
let start = datetime(<start-utc>);
let stop = datetime(<stop-utc>);
ContainerAppConsoleLogs_CL
| where TimeGenerated between (start .. stop)
| where ContainerAppName_s in ("orders-api", "payments-api")
| where Log_s has_any (
"<incident-id>",
"ERR_DIRECT_INVOKE",
"ERR_HEALTH_NOT_READY",
"failed to invoke",
"connection refused",
"deadline exceeded",
"retry"
)
| project TimeGenerated, ContainerAppName_s, RevisionName_s,
ReplicaName_s, ContainerName_s, Log_s
| order by TimeGenerated asc Add ContainerAppSystemLogs_CL evidence for Dapr component creation errors, revision provisioning, and traffic changes. When logs are routed to Azure Monitor rather than Log Analytics, adapt table and column names to the configured destination mode.
Build one timeline. A source trace with no target trace suggests failure before the target application. A target-sidecar trace without an application log suggests a sidecar-to-application problem. A completed target handler followed by a caller timeout means the team must check the response, deadline, and existing side effects before retrying.
Separate resiliency from amplification
Retries can hide a short failure, but they can also duplicate a write or keep a saturated dependency under load. Record the effective resiliency policy, timeout, attempt count, and methods to which the policy applies.
Before replaying a write, prove:
- which idempotency key the target actually enforces;
- the downstream state for the correlation ID;
- how many attempts already occurred;
- the caller’s maximum acceptable deadline;
- how the circuit breaker opens and recovers.
A lower error rate combined with more attempts and higher latency is not recovery. It is an incident displaced into the resiliency layer.
Keep mTLS, components, and Azure identity distinct
Dapr protects interservice communication with mTLS, but a 401 or 403 may come from the target application, a Dapr component, or an Azure dependency reached by the handler. Identify which layer emitted the status before changing RBAC.
Successful service invocation does not prove that the target can read Key Vault, publish to Service Bus, or access Storage. If the call reaches the handler and then fails, continue with the revision’s actual managed identity and dependency logs. If no trace reaches the handler, broader Azure roles will not repair the Dapr path.
Compare revisions with a no-effect canary
When the incident follows a new revision, keep the previous revision available long enough to compare a no-effect request on every controllable path. Use an authenticated diagnostic endpoint that checks the listener and dependencies without writing. Compare status, latency, Dapr logs, and the application trace.
The canary is useful only when it:
- originates in the same environment and caller class as production traffic;
- explicitly invokes the intended app ID and method;
- leaves a correlation trail in both applications;
- creates no business side effect;
- runs long enough to cross multiple replicas.
One successful request may have reached the only healthy replica. Group results by revision and replica before changing weights or deactivating a version.
Decide, validate, or roll back
FIX CONFIGURATION
Wrong app ID, application port, or protocol proven in effective configuration.
Change one value, create a revision, and run the canary.
FIX APPLICATION
Dapr delivers the request and the target returns the error.
Fix the listener, handler, idempotency, or failing dependency.
HOLD AND OBSERVE
Transient fault remains unexplained or evidence is incomplete.
Preserve revisions, add targeted telemetry, and do not widen retries.
ROLL BACK
Failure follows the new revision or the latest Dapr configuration.
Restore known-good configuration, remove the bad revision from the active path,
and verify that partial effects will not be replayed.
STOP
Duplicate writes, downstream saturation, or a circuit breaker that remains open.
Stop the calling operation or use the designed degraded mode before retrying. Final validation requires a normal error rate, latency inside budget, complete source-to-target traces, no duplicate effect, and tests across multiple replicas. Preserve before-and-after configuration and the return command. If the fix fails, rollback must restore the last proven combination of application revision and Dapr configuration, not merely restart containers.
Conclusion
A Dapr invocation error is not just a 503 between two containers. The path contains two applications, two sidecars, a logical identity, a protocol, a resiliency policy, and revisions that can change independently.
By freezing one invocation, testing each segment, and correlating evidence by revision, the team can make a clean decision: fix the Dapr contract, fix the application, hold for stronger evidence, or restore the known-good combination. The outcome is not a redeployment that appears to help, but an explained and repeatable service path.