Cloud

Azure Managed Redis: contain a connection storm before scaling or failing over

A production runbook to separate client reconnects, server pressure, DNS/TLS and network saturation, then validate a fix, scaling decision or rollback.

29 Aug 2026 azuremanaged-redisredisconnectionsresilienceobservabilitynetworkingtlskqlscalingfailoverrunbookrollbackproduction

An API starts accumulating Redis timeouts just after new application instances are deployed. Connected clients rise, server load follows, and retries turn a short interruption into a sustained incident. Scaling Azure Managed Redis or forcing a failover looks reasonable. Either action can close more connections and amplify the storm when the client creates too many sockets, uses aggressive timeouts, or fails to reuse long-lived connections.

The production case is a service that uses Azure Managed Redis to accelerate application reads. This runbook separates four causes before action: legitimate Redis pressure, a client reconnect loop, a DNS/TLS/network failure, and a release regression. The outcome is a bounded decision: fix the client, add capacity, let high availability operate, roll back the release, or hold reduced traffic until the evidence is sufficient.

Freeze the incident before creating another outage

Preserve a UTC window, the application version, affected instance group, and the endpoint actually used. Do not trigger a reboot or failover to “clear” connections until the team knows whether those connections are the cause or the consequence.

yaml managed-redis-connection-incident.yml
incident:
started_at_utc: 2026-08-29T05:40:00Z
service: catalog-api-prod
cache: redis-catalog-prod
region: westeurope
endpoint: <cache-name>.<region>.redis.azure.net
port: 10000

symptom:
- redis dependency p95 rises after a deployment
- connected clients increase faster than application traffic
- timeouts trigger repeated reconnect attempts
- fallback reads begin loading the system of record

last_changes:
application_release: catalog-api-2026.08.29.3
replica_count: 12-to-30
redis_client_package: <version>
cache_configuration_changed: false

preserve:
- deployment and autoscale events
- cache metrics with dimensions
- client dependency traces by role instance
- DNS, TCP and TLS test results
- Azure Resource Health and Activity Log events

stop_conditions:
- downstream fallback exceeds its accepted load
- reconnect rate keeps rising after mitigation
- server load remains high while useful operations fall
- a canary opens more connections than its envelope

Separate a continuous failure, where no client can connect, from an intermittent one. Continuous failure points toward the endpoint, DNS, Private Endpoint, firewall, port, or TLS. Intermittent degradation requires correlating maintenance, server pressure, connected clients, and reconnect behavior.

Read service and client evidence on the same window

Redis load alone is not enough. A client storm can raise server load, while a saturated server can cause the initial timeouts. Compare at least:

  • server load and CPU, using maximum over the short incident window and average for trend;
  • connected clients, connection churn when exposed, useful operations, and errors;
  • application latency and timeouts by version and role instance;
  • deployment, autoscale, maintenance, and Azure health events;
  • load on the fallback system if the application temporarily bypasses the cache.

Discover the metric definitions exposed by the resource instead of copying names from a different Redis service.

bash 01-managed-redis-service-evidence.sh
RG="rg-catalog-prod"
CACHE="redis-catalog-prod"

CACHE_ID=$(az redisenterprise show \
--resource-group "$RG" \
--name "$CACHE" \
--query id \
--output tsv)

az redisenterprise show \
--resource-group "$RG" \
--name "$CACHE" \
--output json

az monitor metrics list-definitions \
--resource "$CACHE_ID" \
--query "[].{metric:name.value, display:name.localizedValue, unit:unit}" \
--output table

az redisenterprise test-connection \
--resource-group "$RG" \
--name "$CACHE" \
--auth entra

The connection test proves the path used by the Azure CLI identity. It does not replace a test from the application’s runtime network and identity. Preserve the high-availability configuration, SKU, cluster policy, and resource changes visible in Activity Log.

Correlate timeouts with application instances

If Application Insights receives Redis dependency telemetry, use it to identify the release and instances driving the slope. Map the query to the schema emitted by the actual client; some Redis clients require explicit instrumentation.

kusto 02-managed-redis-client-dependencies.kql
let StartTime = datetime(2026-08-29T05:30:00Z);
let EndTime = datetime(2026-08-29T06:30:00Z);
dependencies
| where timestamp between (StartTime .. EndTime)
| where type =~ "Redis" or target contains ".redis.azure.net"
| summarize
  Calls=count(),
  Failures=countif(success == false),
  P50=percentile(duration, 50),
  P95=percentile(duration, 95),
  P99=percentile(duration, 99),
  ResultCodes=make_set(resultCode, 10)
by cloud_RoleName, cloud_RoleInstance, operation_Name, bin(timestamp, 2m)
| extend FailureRate = todouble(Failures) / Calls
| order by timestamp asc, FailureRate desc

One or two noisy instances point toward a local defect: CPU, thread pool, sockets, DNS resolution, or client lifecycle. A simultaneous rise across every instance, with Redis load already high before errors, supports a capacity hypothesis. If failures begin exactly when replicas scale out while useful traffic stays flat, investigate client amplification first.

Application logs should carry the client version, role instance, reconnect reason, attempt count, and active-connection counter. Never record Entra tokens, access keys, or cached values.

Calculate the connection envelope

Pod or VM count is not the Redis connection count. Each process can create several clients, and cluster policy can require multiple physical connections. Measure one healthy instance, then calculate the service envelope.

yaml redis-connection-envelope.yml
runtime:
replicas: 30
processes_per_replica: 2
client_objects_per_process: 1
measured_physical_connections_per_client: <measure-on-healthy-instance>
expected_headroom: <capacity-plan>

verify:
- the client object is long-lived and shared
- dependency injection does not create one client per request
- old clients are disposed after a controlled replacement
- cluster topology refresh does not multiply clients without bound
- autoscaling accounts for the connection envelope
- reconnects use backoff and jitter

hard_stop:
- connection count grows while replica count is stable
- one request creates a new connection
- failed connects immediately restart without delay
- a rollout would exceed the tested connection envelope

Creating and closing connections is expensive for the server. Reuse long-lived connections instead of creating a client per operation. Aggressive connect timeouts can accelerate a connect-fail-retry loop; for Azure Managed Redis, five seconds is a more defensible starting point than a very short timeout, then tune it from journey measurements. Stagger reconnects across instances.

Prove DNS, TCP, and TLS from the runtime

Networking is one branch of the diagnosis, not the only lens. Test from the same subnet, container, or a representative runner. Use the service hostname rather than a remembered public IP. Azure Managed Redis clients target the instance hostname on port 10000. With a Private Endpoint, resolution must lead to the expected private path without configuring the privatelink hostname directly.

bash 03-managed-redis-runtime-path.sh
CACHE_HOST="<cache-name>.<region>.redis.azure.net"
CACHE_PORT="10000"

getent ahosts "$CACHE_HOST"
nslookup "$CACHE_HOST"

timeout 5 bash -c "exec 3<>/dev/tcp/$CACHE_HOST/$CACHE_PORT"
openssl s_client \
-connect "$CACHE_HOST:$CACHE_PORT" \
-servername "$CACHE_HOST" \
-brief </dev/null

Interpret each layer independently. Wrong resolution is not a Redis defect. Failed TCP with correct DNS requires checking NSGs, UDRs, firewalls, proxies, or Private Endpoint state. Failed TLS requires checking SNI, certificate chain, clock, and interception. Healthy TLS followed by intermittent timeouts returns the investigation to client behavior, load, or service state.

Contain the storm without moving the incident

The first mitigation should reduce new connections and protect dependencies:

  • freeze a rollout or scale-out event that adds processes faster than they stabilize;
  • reuse the shared client and stop per-request construction;
  • add exponential backoff and jitter to reconnects;
  • bound concurrent retries under an end-to-end time budget;
  • shed optional cache use only when the fallback was designed and tested;
  • protect the fallback database or API with its own limits and circuit breakers.

Do not make every instance reconnect at once. A global restart, manual failover, or shorter timeout can create exactly that synchronization. For .NET clients using StackExchange.Redis, let the multiplexer manage reconnection and review the force-reconnect pattern before adding an application loop.

Decide between a client fix, scaling, failover, or rollback

text managed-redis-decision-matrix.txt
Fix or roll back the client
The slope starts with a release, rollout or scale-out
Connections grow faster than replicas and useful traffic
A few instances concentrate timeouts or reconnects
Server load rises after the connection wave
The previous version has a known envelope

Scale Azure Managed Redis
Connections are legitimate, stable and near the tested limit
Operations, throughput and server load rise with useful traffic
No leak, retry storm or client bottleneck explains the symptom
The capacity change has an owner and validation window

Let high availability operate or follow the qualified failover procedure
Resource Health or maintenance explains a node interruption
The client can reconnect progressively
The remaining node can carry the expected load
Failover is not being used as a generic reset

Fix the network path
DNS, TCP or TLS fails continuously from the runtime
Redis metrics remain quiet during failures
The test separates endpoint, routing, filtering and certificate

Hold reduced traffic and collect evidence
Client or service evidence is missing
The fallback approaches its own limit
No canary reproduces the behavior yet

Failover closes in-flight connections. It fits a qualified availability event, not an application connection leak. Scaling is justified by sustained useful load, not by connections created in error.

Canary, validate, and keep rollback open

Deploy the client correction to one instance or a small traffic share. Compare it with a control under the same load. Validation must cover the connected-client slope, reconnect rate, timeouts, p95/p99 latency, server load, useful operations, and fallback pressure.

yaml managed-redis-change-gate.yml
candidate:
application_release: catalog-api-2026.08.29.4
change:
  - reuse one long-lived client per process
  - add bounded reconnect backoff with jitter
canary_replicas: 1
observation_window: <derived-from-traffic-cycle>

promote_when:
- physical connections stay inside the measured envelope
- reconnect rate returns to baseline
- redis dependency p95 and failures recover
- server load does not rise without useful operations
- fallback load remains inside its guardrail

rollback_when:
- connections keep growing on the canary
- recovery requires repeated forced reconnects
- request success falls despite lower connection count
- downstream load becomes unsafe

rollback:
application_release: catalog-api-2026.08.29.2
keep_capacity_change: false-unless-legitimate-load-was-proven
preserve_evidence: true

If temporary scaling bought time, reduce it only after the client is stable across a representative traffic cycle. If the client is fixed but useful load remains above the capacity baseline, retain the capacity and document the new baseline.

Conclusion

An Azure Managed Redis connection storm is a chain incident: client lifecycle, rollout, DNS/TLS, Redis capacity, high availability, and fallback behavior interact. Looking only at server load makes the wrong change easy.

An operational decision comes from a shared timeline, a measured connection envelope, and a canary. Fix or roll back the release when it amplifies connections, scale when useful load justifies it, and reserve failover for a qualified availability event. In every path, validate recovery without closing the return route.