Infrastructure
Azure Functions Flex Consumption: diagnose stalled scale-out before raising the instance ceiling
A production runbook for separating demand, concurrency, scale groups, the instance ceiling, regional quota and downstream saturation before changing Function App capacity.
A Function App on Flex Consumption handles its normal load, but falls behind during a burst. Queue depth rises, p95 latency degrades and the instance count appears to stop growing. Raising maximumInstanceCount looks like the obvious fix. It will not help when per-instance concurrency, the active scale group, regional memory quota or a downstream dependency is the actual constraint.
This runbook builds an evidence-based decision: raise the ceiling, tune concurrency, pre-provision instances, repair a dependency, or restore the previous configuration. It covers HTTP and event-driven functions while accounting for the per-function scaling model used by Flex Consumption.
Freeze the symptom and capacity budget
Do not start from the ceiling value. Describe the demand the system must absorb, the affected window and the boundary the platform is meant to protect.
app: func-orders-prod
region: westeurope
plan: flex-consumption
incident_window: 2026-09-07T06:40:00Z/2026-09-07T07:10:00Z
symptom:
trigger: service-bus-orders
backlog: rising
p95_duration: above-service-budget
failures: intermittent-timeouts
capacity_contract:
maximum_instance_count: 100
instance_memory_mb: 2048
per_instance_concurrency: current-iac-value
always_ready: current-iac-value
protected_downstream: orders-database
decision:
promote_only_if: backlog-drains-without-downstream-regression
rollback_on: higher-error-rate-or-dependency-saturation Retain the application commit, IaC commit, last configuration change and a comparable normal-traffic window. Without that baseline, reaching 40 instances may look inadequate even when it is faster than the previous release and deliberately bounded by database capacity.
Read the deployed configuration
Flex Consumption keeps the ceiling, memory size, HTTP concurrency and always-ready instances under functionAppConfig.scaleAndConcurrency. Read the Azure resource and compare it with IaC. A portal screenshot or pipeline variable is not deployed-state evidence.
RG="rg-orders-prod"
APP="func-orders-prod"
az resource show \
--resource-group "$RG" \
--name "$APP" \
--resource-type "Microsoft.Web/sites" \
--api-version "2023-12-01" \
--query "properties.functionAppConfig.scaleAndConcurrency" \
--output json > flex-scale-current.json
jq '{maximumInstanceCount, instanceMemoryMB, triggers, alwaysReady}' \
flex-scale-current.json
# Compare with the approved Bicep, ARM or Terraform resource.
git diff --no-index expected-scale.json flex-scale-current.json || true Confirm region, runtime, version, trigger type and host storage as well. Drift in maximumInstanceCount is one possible finding, not a general explanation for lag.
Identify the scale group carrying demand
Flex Consumption does not necessarily operate the app as one pool. HTTP functions share a group, Event Grid blob functions another, Durable Functions another, while other triggers can scale independently. The instance ceiling applies to on-demand instances in each group, not to a simple app-level total.
For every affected function
Function name and trigger type
Expected group: http, blob, durable or function:<name>
Demand signal: requests, backlog, partitions or orchestrations
Effective per-instance concurrency
Distinct instances during the burst
Mean and p95 execution time
Dependencies called per execution
Do not infer a global ceiling from
the total instance count for the app
one aggregated metric
the configured maximum without regional quota
a backlog that does not prove trigger health This map prevents two common mistakes: raising a ceiling for the wrong group and counting instances that never execute the function under pressure.
Prove demand, concurrency and instances in telemetry
Application Insights can connect processed volume, failures, duration and observed instances. Adapt the query to the telemetry model and role names emitted by the app.
let Start = datetime(2026-09-07T06:30:00Z);
let End = datetime(2026-09-07T07:20:00Z);
requests
| where timestamp between (Start .. End)
| where cloud_RoleName == "func-orders-prod"
| summarize executions=count(),
failed=countif(success == false),
p50_ms=percentile(duration, 50),
p95_ms=percentile(duration, 95),
active_instances=dcount(cloud_RoleInstance),
sample_instances=make_set(cloud_RoleInstance, 8)
by bin(timestamp, 1m), operation_Name
| extend failure_rate=round(100.0 * todouble(failed) / executions, 2)
| order by timestamp asc Read the curve instead of one sample. Rising demand with more instances and stable p95 indicates useful scale-out. Flat instances with growing demand point toward concurrency, ceiling, quota or the trigger signal. More instances combined with worse p95 and errors point first to code, initialization or a dependency.
Separate ceiling, scale rate and regional quota
The configured maximum is a ceiling, not reserved capacity. The platform adds instances along its scale curve, can briefly throttle scale-out requests and retries automatically. The subscription’s regional memory quota can stop growth below the configured value, particularly when several Flex apps consume capacity in the same region.
Application ceiling
Azure value matches IaC
Value supports target throughput at current concurrency
No unqualified recent change
Regional quota
Total memory available in the region
Other active Flex Function Apps during the window
Headroom for independently scaling groups
Scale rate
Sudden burst or progressive demand
Always-ready instances available for predictable load
Startup and dependency initialization duration
Transient throttling separated from a persistent ceiling A short burst that ends before instances arrive often calls for targeted always-ready capacity or demand smoothing. A sustained plateau exactly at the maximum supports investigating the ceiling. A lower plateau requires quota, scale group and trigger health checks first.
Check concurrency and downstream saturation
Concurrency changes how many instances are required. Set it too high and each instance can exhaust CPU, memory, connections or the thread pool. Set it too low and the same throughput requires more instances. HTTP per-instance concurrency is configurable on Flex Consumption; for other triggers, the control depends on the extension and concurrency model.
dependencies
| where timestamp between (datetime(2026-09-07T06:30:00Z) .. datetime(2026-09-07T07:20:00Z))
| where cloud_RoleName == "func-orders-prod"
| summarize calls=count(),
failed=countif(success == false),
p95_ms=percentile(duration, 95),
result_codes=make_set(resultCode, 8)
by bin(timestamp, 1m), target, type
| extend failure_rate=round(100.0 * todouble(failed) / calls, 2)
| order by timestamp asc If the database, API or broker slows as instances arrive, a higher ceiling accelerates the incident. Bound concurrency, protect the dependency and measure completed throughput rather than executions started.
Test one lever on a canary
Do not change the ceiling, memory, concurrency and always-ready count together. Build an equivalent canary Function App, replay representative load and change one IaC parameter at a time.
control:
code: same-artifact
maximum_instance_count: current
concurrency: current
always_ready: current
candidate_a:
change: maximum-instance-count-only
prove:
- active-instances-rise-beyond-old-ceiling
- backlog-drains-faster
- dependency-p95-remains-within-budget
candidate_b:
change: per-instance-concurrency-only
prove:
- throughput-improves-per-instance
- cpu-memory-and-errors-remain-stable
candidate_c:
change: always-ready-only
prove:
- predictable-burst-avoids-cold-start
- idle-cost-is-accepted Use the same messages, partitions, payload sizes and downstream limits as production. An HTTP synthetic test does not qualify a Service Bus trigger.
Decide, validate or roll back
raise_instance_ceiling:
when:
- active_group_reaches_current_ceiling
- regional_memory_quota_has_headroom
- downstream_accepts_added_parallelism
- canary_drains_backlog_without_regression
tune_concurrency_or_always_ready:
when:
- per_instance_pressure_or_cold_start_explains_delay
- one_parameter_canary_meets_capacity_budget
hold_and_fix_dependency:
when:
- more_instances_increase_dependency_latency_or_errors
- trigger_health_or_scale_group_is_not_proven
rollback:
action:
- restore_previous_iac_values
- redeploy_configuration_only
- replay_the_same_load_window
- confirm_backlog_errors_and_dependency_pressure_return_to_baseline After promotion, observe at least one comparable load cycle. Rollback restores the previous value set, not a number invented during the incident. Reconcile business effects from delayed or repeated messages before closure.
Conclusion
An apparently stalled Flex Consumption scale-out is not automatically a low instance ceiling. Evidence must connect the active group, its demand, concurrency, running instances, regional quota and downstream capacity.
Raise maximumInstanceCount only when the group actually reaches it and a canary proves better throughput without moving the bottleneck. Otherwise tune the control that matches the symptom, keep one variable per experiment and retain the previous configuration as the return path.