Automation
Azure App Configuration: diagnose feature flag drift before rollback
A production runbook for qualifying inconsistent feature flag evaluations across instances by separating version, refresh, cache, targeting, telemetry and rollback.
The CheckoutV2 feature flag has moved from a 10% to a 25% production rollout. Minutes later, some users still see the old journey, others switch variants between requests, and part of the application fleet never appears to receive the new allocation. Disabling the flag may contain the impact, but it does not reveal whether the fault sits in the published definition, application refresh, targeting context or telemetry.
This runbook treats the flag as a production dependency. It proves the requested state, the state loaded by each instance and the state actually evaluated before choosing to hold, repair refresh, reduce exposure or roll back.
Freeze the change and symptom
Start with a record that binds the change to a UTC window, label and affected population. A flag name is ambiguous once environments, labels and variants exist.
incident:
flag: CheckoutV2
store: <app-configuration-store>
label: prod
changed_at_utc: 2026-08-28T05:40:00Z
change: rollout_10_to_25_percent
expected_variant: checkout-v2
previous_known_good: rollout_10_percent
scope:
services:
- web-frontend
- checkout-api
regions:
- westeurope
- northeurope
symptom:
- inconsistent variant between requests
- instances still evaluating the previous allocation
stop_conditions:
- payment error rate exceeds threshold
- assignment changes for the same targeting ID
- flag state cannot be tied to an ETag
- rollback state is not known Record the latest application deployment, App Configuration and feature management library versions, and any identity, network or replica change. Timing does not prove causality, but it limits the comparisons that matter.
Prove the published definition
Read the exact key and label used by the application. Preserve its raw value, ETag and timestamp as evidence. A portal view or an unqualified screenshot is not enough.
set -euo pipefail
store="<app-configuration-store>"
flag="CheckoutV2"
label="prod"
key=".appconfig.featureflag/$flag"
az appconfig kv show --name "$store" --key "$key" --label "$label" --auth-mode login --output json > feature-flag-current.json
jq '{key, label, etag, last_modified, content_type, value}' feature-flag-current.json Confirm that the definition matches the intended change: enabled state, variants, percentages, groups, users, time window and optional distribution seed. A valid JSON document can still be operationally wrong when the wrong label changed or a user override takes precedence over percentage allocation.
Compare it with the known state before the change. If that reference only exists in chat or a ticket, rollback is not executable yet. Prepare a restorable definition and have it reviewed before any mutation.
Separate publication, loading and evaluation
A feature flag moves through three distinct states:
- the definition published in App Configuration;
- the definition loaded and cached by each process;
- the result evaluated for a specific user context.
A correct store value does not prove that every instance consumes it. Conversely, two users can legitimately receive different variants when targeting requires it.
Expose a temporary read-only diagnostic that returns no sensitive data or full targeting lists. It should support instance comparison without turning the application into a configuration console.
{
"service": "checkout-api",
"region": "westeurope",
"instance": "checkout-api-7f8c9",
"deploymentVersion": "2026.08.28.1",
"flag": "CheckoutV2",
"label": "prod",
"configurationEtag": "<etag>",
"lastRefreshUtc": "2026-08-28T05:42:18Z",
"evaluation": {
"targetingIdHash": "<stable-hash>",
"enabled": true,
"variant": "checkout-v2",
"assignmentReason": "Percentile"
}
} Compare at least one healthy and one affected instance in every region. A different ETag points toward refresh or store access. The same ETag with a different result for an identical stable context points toward targeting, distribution seed or context construction. If state changes inside one request, verify that the application uses a feature manager snapshot when request-level consistency is required.
Diagnose refresh without causing a second incident
Refresh behavior depends on the provider and framework. Document the implementation that production actually runs: interval, request-driven trigger, middleware, optional sentinel and behavior when App Configuration is unavailable.
Check that:
- every instance loads the same endpoint, label and key filters;
- the path that triggers refresh is genuinely exercised;
- a refresh interval is not mistaken for an immediate update guarantee;
- 401, 403, 429, DNS and TLS failures to App Configuration are visible;
- the application retains a last known good state instead of silently selecting a dangerous default;
- when a sentinel is used, it changes after the keys composing the release, not before them.
Do not restart the whole fleet to clear caches. That removes evidence and can load the wrong definition everywhere at once. Select one canary instance, invoke the application’s supported refresh path, then compare its ETag and evaluation with the rest of the fleet.
Correlate evaluations in Application Insights
Flag telemetry should answer two questions: which definition was evaluated and what result was served? When enabled, a FeatureEvaluation event can include the flag name, enabled state, variant, assignment reason, ETag and allocation identifier.
customEvents
| where timestamp between (datetime(2026-08-28T05:20:00Z) .. datetime(2026-08-28T06:20:00Z))
| where name == "FeatureEvaluation"
| extend Flag = tostring(customDimensions.FeatureName),
Enabled = tobool(customDimensions.Enabled),
Variant = tostring(customDimensions.Variant),
Reason = tostring(customDimensions.VariantAssignmentReason),
ETag = tostring(customDimensions.ETag),
AllocationId = tostring(customDimensions.AllocationID),
Role = cloud_RoleName,
Instance = cloud_RoleInstance
| where Flag == "CheckoutV2"
| summarize Evaluations=count(),
Instances=dcount(Instance),
Etags=make_set(ETag, 10)
by bin(timestamp, 5m), Role, Enabled, Variant, Reason, AllocationId
| order by timestamp asc Adapt dimension names to the emitted schema. Missing events do not prove that the flag was not evaluated; telemetry may be disabled, filtered or sampled. Correlate evaluation events with refresh traces, service version, region and business metrics. Do not log a raw user identity when a stable pseudonymous identifier is sufficient.
Decide hold, repair or rollback
HOLD
The definition, ETags and allocations match the rollout.
Observed differences are explained by intended targeting.
Technical and business metrics remain inside the accepted window.
REPAIR REFRESH
The store contains the correct definition, but instances retain an old ETag.
Fix the trigger or filter on one canary instance.
Expand only after convergence is proven.
REDUCE ROLLOUT
The definition is consistent, but the variant degrades metrics.
Return to the known percentage, retain the same control population,
and verify the resulting allocation before continuing.
ROLL BACK
Restore the reviewed known definition with a conditional write.
Prove the new ETag, wait for instance convergence,
then verify evaluations and business metrics.
STOP AND INVESTIGATE
The prior state is unknown, several labels changed,
or evaluations cannot be tied to an application version. Avoid a blind overwrite during rollback. Another operator may have published a correction after evidence collection. Use an ETag precondition or the concurrency mechanism supported by the delivery path, record the actor and retain the replaced definition.
After rollback, success is not “the portal shows 10%”. Instances must converge on the restored ETag, assignments must become stable and affected metrics must return to the expected window. Keep a synthetic canary that evaluates fixed contexts before each future rollout increase.
Conclusion
Feature flag drift is not always fixed by flipping a switch. The published definition, per-instance ETag, targeting context and served result need to be joined as one operational chain. That evidence separates normal rollout variance from a stale cache, wrong label or genuine feature regression.
The resulting decision is testable: hold an explained allocation, repair refresh on a canary, reduce exposure in stages or restore the known state with concurrency control. Rollback is complete only when the fleet and its metrics prove convergence.