Cloud
Azure App Service: diagnose Health Check failures before restarting instances
A production runbook for separating a faulty health endpoint, instance degradation, redirects, authentication and dependency failures before restarting or changing App Service Health Check.
An Azure App Service remains reachable, yet Health Check reports one or more unhealthy instances. Traffic becomes uneven, latency rises on the remaining workers, or the platform repeatedly takes the same instance out of rotation. Restarting the app may clear the symptom, but it also destroys the process state and timing evidence needed to distinguish a bad worker from a bad health contract.
The running case is a production API scaled to several App Service instances. Health Check calls /health/ready; the endpoint validates the database and a messaging dependency. After a release, only some instances fail, then the whole app intermittently returns non-2xx responses. This runbook drives one decision: repair the endpoint, isolate an instance, roll back the application, or restore the previous Health Check configuration without masking a shared dependency failure.
Freeze the health contract and incident window
Health Check is a traffic decision, not merely a monitoring URL. Record what the endpoint promises before changing its path or failure threshold.
incident: inc-20260910-006
app: app-orders-prod
slot: production
plan: asp-orders-prod
region: westeurope
instances_expected: 4
health_path: /health/ready
first_unhealthy_utc: 2026-09-10T14:12:00Z
release: orders-api-2026.09.10.3
health_contract:
healthy_status: 200-299
anonymous_or_platform_authenticated: explicit
checks:
- process_ready
- database_read
- service_bus_connection
excludes:
- noncritical_reporting_api
- destructive_business_write
preserve_before_action:
- health_check_configuration
- instance_level_metric
- application_and_dependency_logs
- deployment_and_configuration_timeline
- process_memory_threads_and_restarts
- previous_revision_and_rollback_command Do not begin by increasing the failure threshold. A larger threshold delays removal; it does not make the endpoint more representative. Also note whether the site runs on one instance or several. With one instance, Health Check can report failure but cannot reroute traffic to another healthy worker.
Read the deployed configuration
Capture the site configuration and the two settings that control when workers are excluded. App Service pings the configured path on each instance at one-minute intervals. A response outside 200-299, a timeout, an authentication challenge or a redirect can all make the platform view differ from a manual test.
set -eu
RG="rg-app-prod"
APP="app-orders-prod"
az webapp config show \
--resource-group "$RG" \
--name "$APP" \
--query '{healthCheckPath:healthCheckPath,alwaysOn:alwaysOn,http20Enabled:http20Enabled,ftpsState:ftpsState}' \
--output json
az webapp config appsettings list \
--resource-group "$RG" \
--name "$APP" \
--query "[?name=='WEBSITE_HEALTHCHECK_MAXPINGFAILURES' || name=='WEBSITE_HEALTHCHECK_MAXUNHEALTHYWORKERPERCENT'].{name:name,value:value}" \
--output table Keep slot behavior in the snapshot. Health Check configuration is exchanged during a slot swap, so production and staging should use compatible paths and semantics. A swap can otherwise promote good code with an invalid destination-side probe contract.
Separate a probe failure from an instance failure
Test the endpoint through the public or private application route, then correlate it with per-instance evidence. A successful client request only proves that one healthy worker answered. A global failure may point to the endpoint or a shared dependency; a failure on one instance points more strongly to local process, filesystem, memory or worker state.
One instance unhealthy
Compare instance dimension, process restarts, memory, threads and local logs
Check whether the release or warm-up completed on that worker
Preserve diagnostics before restart or replacement
All instances unhealthy at the same time
Check endpoint code, shared dependency, DNS, identity and configuration
Do not treat worker replacement as the first correction
Remember that the platform avoids removing every instance at once
Manual request returns 200, Health Check fails
Check default hostname behavior, redirects and authentication
Confirm the manual test uses the exact configured path
Compare HTTPS Only and application-level redirect rules
Health endpoint is green while users fail
The probe is too shallow or bypasses the failing dependency
Add a representative read without turning health into a full transaction
Keep liveness and readiness responsibilities separate The health endpoint must be cheap and deterministic. It should prove readiness to serve the request class represented by the app, but it should not create data, consume messages or fail because an optional analytics service is slow.
Check redirects and authentication explicitly
App Service Health Check expects 200-299 and does not follow redirects. A redirect from the default hostname to a custom domain, an application-enforced HTTPS redirect, or a login challenge can mark every worker unhealthy even though a browser eventually shows a green page.
set -eu
DEFAULT_HOST="app-orders-prod.azurewebsites.net"
CUSTOM_HOST="orders.example.com"
PATH_TO_TEST="/health/ready"
curl --silent --show-error --output /dev/null \
--write-out 'default status=%{http_code} redirect=%{redirect_url} time=%{time_total}\n' \
"https://$DEFAULT_HOST$PATH_TO_TEST"
curl --silent --show-error --output /dev/null \
--write-out 'custom status=%{http_code} redirect=%{redirect_url} time=%{time_total}\n' \
"https://$CUSTOM_HOST$PATH_TO_TEST" If App Service Authentication protects the app, verify its supported integration with Health Check. If the application implements its own authentication, the path must either allow the platform request or validate the documented internal token. Do not solve the issue by exposing a verbose diagnostic endpoint that returns dependency names, secrets or connection details.
Correlate health, runtime and dependencies
Read HealthCheckStatus with the Instance dimension, then compare the same window with requests, exceptions, dependency failures, memory and restart signals. Health Check status appears only after the configured failure threshold is reached, so keep the earlier application events as well.
let StartTime = datetime(2026-09-10T14:00:00Z);
let EndTime = datetime(2026-09-10T15:00:00Z);
let AppResource = "app-orders-prod";
AzureMetrics
| where TimeGenerated between (StartTime .. EndTime)
| where Resource =~ AppResource
| where MetricName in ("HealthCheckStatus", "Http5xx", "MemoryWorkingSet", "Requests")
| extend Instance = tostring(column_ifexists("Instance", "not-exported"))
| summarize Average=avg(Average), Maximum=max(Maximum), Total=sum(Total) by bin(TimeGenerated, 5m), MetricName, Instance
| order by TimeGenerated asc, Instance asc Table and dimension shape can differ with the chosen diagnostic path. Preserve the raw metric export if the workspace projection does not carry the instance dimension. Then correlate application telemetry separately: endpoint duration, failed dependency, revision, role instance and exception type.
Prove whether a dependency belongs in readiness
A health endpoint can create a self-inflicted outage when it fails on a dependency that is degraded but noncritical. It can also stay green while the app has lost the database required for every request. Classify each check against user impact and recovery behavior.
dependencies:
primary_database:
critical: true
probe: bounded_read
timeout_ms: 500
unhealthy_on: repeated_failure
service_bus:
critical_for: write_commands
probe: connection_or_management_read
timeout_ms: 500
degraded_mode: accept_reads_reject_new_commands
reporting_api:
critical: false
probe: telemetry_only
unhealthy_on: never_by_itself
endpoint_rules:
- no_business_write
- no_secret_or_topology_in_response
- total_budget_below_platform_timeout
- cache_only_when_staleness_is_visible
- emit_dependency_and_instance_reason_in_internal_logs When the shared dependency is the fault, removing workers from rotation only concentrates traffic on the remaining instances. Prefer a bounded degraded mode, circuit breaker or dependency recovery when the business contract allows it. Change probe semantics only when the current contract is demonstrably wrong.
Choose a correction with a return path
Avoid stacking a restart, threshold change and endpoint rewrite in one incident. Select the smallest action that tests the proven hypothesis.
Repair the endpoint
Redirect, authentication or wrong status code explains the platform failure
Dependency policy is too broad or timeout budget is invalid
Validate on a slot before applying the configuration to production
Isolate or restart one instance
Failure is tied to one worker and diagnostics are preserved
Other instances have enough measured capacity
Post-restart validation can identify the same instance and revision
Roll back the application
Failure begins with a release and follows the new revision
Endpoint and business errors regress together
Previous revision and configuration are known healthy
Restore previous Health Check configuration
Path or threshold changed independently of the release
Previous path remains representative and secure
Rollback will not hide a real application outage
Block further action
All workers fail because a shared dependency is unavailable
Remaining capacity is unknown
Logs cannot distinguish probe failure from application failure Health Check configuration changes restart the app. Treat the configuration rollback itself as a production action: use a staging slot when available, capture settings before and after, and do not edit the path merely to silence the metric.
Validate recovery across several probe windows
One 200 is not recovery. Keep the candidate through several one-minute probe cycles and normal traffic. Prove that every expected instance is healthy, traffic is distributed, dependency errors remain bounded and the endpoint still detects a deliberate negative case in a safe environment.
Keep the correction
HealthCheckStatus stable for every expected instance
User error rate and latency return to baseline
Endpoint remains within its response-time budget
No instance restart loop or repeated exclusion
Critical dependency failure still makes readiness fail
Optional dependency failure no longer removes healthy capacity
Roll back or hold
Only the synthetic request is green
Metric improves while user errors remain
One instance repeatedly leaves rotation
Remaining workers approach capacity limits
Endpoint exposes sensitive diagnostic detail
Probe contract differs between production and staging Conclusion
App Service Health Check is useful only when its endpoint represents the traffic decision the platform must make. When workers turn unhealthy, freeze the contract, read the deployed configuration, separate instance-local failure from shared dependency failure, and check redirects and authentication before restarting anything.
The closing decision should be narrow and observable: repair the endpoint, isolate one proven worker, restore the previous configuration, roll back the release, or keep the app under a controlled degraded mode while the shared dependency recovers. Service is restored only when both the platform metric and the user path agree across several probe windows.