Networking

Azure App Service VNet Integration: diagnose subnet exhaustion before retrying scale

A production runbook for measuring VNet Integration subnet capacity, separating address pressure from delegation and path failures, then validating resize or migration with rollback.

03 Oct 2026 azureapp-servicevnet-integrationsubnetscalingnetworkingcapacityautomationobservabilityrunbookrollbackproduction

An Azure App Service plan remains reachable, but a scale-out or scale-up operation does not complete. Application metrics still show load, existing instances answer, and no NSG rule has changed. The common response is to retry the operation, raise the target again, or investigate the application. The actual constraint may sit in the VNet Integration subnet: it no longer has enough addresses to create the next worker cohort.

The running case is a production plan with eight instances integrated with a legacy /28 subnet. Eleven addresses remain usable after Azure reserves five. A size change can keep the eight old instances alive while eight new ones start. The transient demand is then sixteen addresses, before accounting for another plan or the specific Windows Containers multiplier. The runbook must reach one of four decisions: wait for proven temporary pressure to clear, lower the scale target, migrate to a correctly sized subnet, or roll back the change that initiated the operation.

Freeze the capacity contract

Do not start by disconnecting VNet Integration. That action restarts the app and removes useful incident state. Capture the plan, subnet, operation and expected capacity first.

text vnet-integration-capacity-contract.txt
UTC window: 2026-10-03 05:35Z to present
App / slot: api-orders-prod / production
App Service plan: asp-orders-prod
Active instances before incident: 8
Requested operation: scale up P1v3 to P2v3
VNet Integration subnet: snet-appsvc-integration-prod
Prefix: 10.42.8.0/28

Symptom
Scale operation incomplete or failed
Existing instances still reachable
No confirmed DNS, UDR or NSG change

Decision to reach
Wait for proven address release
Return to the previous size
Migrate to a subnet with validated headroom
Escalate a platform incident with evidence

Keep the Azure operation ID, Activity Log events, the last integration change time and the autoscale timeline. A blind retry can create a second attempt that obscures the original cause.

Prove which plan consumes which subnet

VNet Integration is an outbound path. A Private Endpoint in the same architecture may protect ingress or access to a dependency, but it contributes no capacity to the integration subnet. Inventory the apps and plans actually connected to the subnet instead of inferring usage from nearby Private Endpoints.

bash inventory-vnet-integration.sh
RG_APP="rg-orders-prod"
APP="api-orders-prod"
PLAN="asp-orders-prod"

az webapp vnet-integration list --resource-group "$RG_APP" --name "$APP" --output json

az appservice plan show --resource-group "$RG_APP" --name "$PLAN" --query '{id:id,sku:sku.name,capacity:sku.capacity,kind:kind,reserved:reserved}' --output json

az monitor autoscale list --resource-group "$RG_APP" --output json

Use Azure Resource Graph to find other apps that declare the same subnet. The query is a configuration inventory, not a free-address counter. Group the results by App Service plan and verify every integration afterwards.

kusto 01-apps-by-integration-subnet.kql
let subnet = tolower('/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-net-prod/providers/Microsoft.Network/virtualNetworks/vnet-prod/subnets/snet-appsvc-integration-prod');
resources
| where type =~ 'microsoft.web/sites'
| extend integrationSubnet = tolower(tostring(properties.virtualNetworkSubnetId))
| where integrationSubnet == subnet
| project subscriptionId, resourceGroup, app=name,
        appServicePlan=tostring(properties.serverFarmId), integrationSubnet

Include slots, integrations exposed through site networking properties, and every plan joined to the subnet. With Multi Plan Subnet Join, headroom is shared: add the instances from every plan, not only the application reporting the incident.

Calculate the envelope, including the transient peak

Azure reserves five addresses in every subnet. For multitenant App Service, a plan instance normally consumes one address in each integration subnet it uses. During a scale up or down in size, old and new instances can coexist, temporarily doubling the plan’s address use. Platform upgrades also require headroom. Addresses released after an operation can remain allocated briefly and, in rare cases, for up to twelve hours.

Write down the calculation instead of treating the CIDR as the answer.

text address-envelope.txt
Subnet /28
Total addresses                              16
Addresses reserved by Azure                   5
Usable addresses                             11

Plan asp-orders-prod
Current instances                             8
New cohort during scale-up                     8
Transient plan peak                           16

Headroom at peak: 11 - 16 = -5
Decision: the subnet cannot safely carry this operation

For a shared subnet, the minimum calculation is:

peak addresses required = stable instances of other plans + 2 × instances of the plan changing size + operational reserve.

Adapt the formula when several plans can scale at once. For Windows Containers, add the extra address consumed by each app for every plan instance. Treat the recommendation to allocate twice the planned maximum as operating capacity for cohort overlap and platform work, not as paperwork. Microsoft recommends starting production designs with a /26, but the correct prefix still depends on plan count, containerized apps and integration count.

Separate exhaustion, delegation and path failure

A nearly full subnet does not explain every failure. Verify the delegation, associations and prefix before declaring address exhaustion.

bash inspect-integration-subnet.sh
RG_NET="rg-net-prod"
VNET="vnet-prod"
SUBNET="snet-appsvc-integration-prod"

az network vnet subnet show --resource-group "$RG_NET" --vnet-name "$VNET" --name "$SUBNET" --query '{prefixes:addressPrefixes,delegations:delegations[].serviceName,nsg:networkSecurityGroup.id,routeTable:routeTable.id,serviceAssociationLinks:serviceAssociationLinks[].id}' --output json

Keep the diagnostic branches separate:

  • a new worker fails while existing instances remain healthy: capacity or platform operation is likely;
  • no instance receives a private address: inspect integration, delegation and the Service Association Link;
  • instances have addresses but a dependency is unreachable: investigate DNS, UDR, NSG, firewall and return routing;
  • only a Windows Container app fails after adding apps: recalculate the container-specific multiplier;
  • the numbers appear sufficient immediately after scale-in: prove whether address release is delayed before retrying.

WEBSITE_PRIVATE_IP can confirm that an instance received an address from the integration subnet. The value can change, so never add individual values to an allowlist. Destinations should allow the intended integration prefix, subject to the security controls defined for that flow.

Choose a correction that does not move the incident

If the pressure is temporary and more compute is not required, wait for observed address release and stop concurrent automatic retries. If autoscale requests a target outside the envelope, temporarily lower its maximum with an owner, expiry and exit threshold. That containment is not a substitute for a correctly sized subnet.

An assigned VNet Integration subnet cannot be resized in place. The durable correction is usually a new dedicated subnet, delegated to Microsoft.Web/serverFarms, with the required NSG, UDR and DNS dependencies. Validate the parent address space too: a new /26 that overlaps an on-premises route creates a second, less obvious incident.

bash create-target-integration-subnet.sh
az network vnet subnet create --resource-group "rg-net-prod" --vnet-name "vnet-prod" --name "snet-appsvc-integration-prod-v2" --address-prefixes "10.42.9.0/26" --delegations "Microsoft.Web/serverFarms"

az network vnet subnet show --resource-group "rg-net-prod" --vnet-name "vnet-prod" --name "snet-appsvc-integration-prod-v2" --output json

Do not copy every legacy subnet rule mechanically. Rebuild the expected outbound contract: DNS servers, private prefixes, any network appliance, NAT Gateway, allowed destinations and the return path. A rule that is no longer justified should not survive merely because it existed on v1.

Canary the migration and retain rollback

When the plan still has a second integration available, use a representative canary app on the same plan to qualify the new subnet. Otherwise, use an equivalent validation plan and document the difference. The canary must test more than /health: DNS resolution from the app, TCP access to private dependencies, managed identity, Internet egress when all-traffic routing is enabled, the source address seen by an external dependency, and telemetry.

text subnet-migration-decision.txt
PROMOTE
Plan and slot inventory is complete
Peak envelope fits the new prefix
Microsoft.Web/serverFarms delegation is verified
DNS, UDR, NSG, NAT and return routing pass from the canary
Identity and application dependencies pass
Canary autoscale completes with positive headroom

STOP
A shared app or plan remains unknown
Canary DNS or routing differs from production
Scale creates 5xx, timeout or identity errors
Headroom excludes worker cohort overlap

ROLL BACK
Reconnect the app to the previous subnet in the change window
Restore the validated capacity and autoscale maximum
Revert UDR, NSG or NAT changes used only by the migration
Confirm health, outbound path and no new scale attempt

Changing the integration restarts the app. Treat the cutover as a production change with a window, synthetic probes and a stop condition. Keep the previous subnet intact until the new cohort and the first scaling operation have both passed. A network rollback that leaves the larger capacity target in place can reproduce the same failure immediately.

Conclusion

A blocked App Service scale operation is not necessarily a compute quota or code problem. With VNet Integration, the subnet belongs to the plan’s capacity envelope. It must absorb stable instances, size-change overlap, platform operations and every other plan that shares it.

The production decision must therefore be measurable. Retry only when an address was actually released and the calculated peak fits. Otherwise, temporarily lower the target or migrate to a subnet sized and tested from a representative app. Rollback is equally explicit: previous integration, previous capacity, previous network rules and application validation before autoscale becomes autonomous again.