Cloud

Azure WAF Bot Manager: qualify a bot before adding an allow rule

A production runbook for separating legitimate automation, spoofed identities and hostile bot traffic before changing Azure WAF Bot Manager actions.

01 Oct 2026 azurewaffront-doorapplication-gatewaybot-managersecuritykqlobservabilityautomationrunbookrollbackproduction

A partner crawler starts receiving 403 responses through Azure Front Door. Its operator asks for an allow rule and points to a recognizable user-agent. At the same time, WAF logs classify other requests with the same client library as unknown bots. Allowing that string would restore the integration quickly, but it could also create a reusable bypass for unrelated traffic.

The running case is a production API protected by Azure WAF Bot Manager. A partner reads a bounded catalog endpoint with an automated client, while search crawlers, health probes, SDKs and hostile scanners reach the same hostname. The runbook must decide whether the request is legitimate, whether Bot Manager is the blocking layer, and whether to keep blocking, change a managed rule action, introduce a narrow exception or leave authorization to the application.

Freeze the automation contract

Do not begin with the user-agent or the WAF rule. Describe the automated client as a production dependency with an owner, purpose, route, authentication method, traffic envelope and revocation path.

yaml 01-bot-client-contract.yml
client:
name: partner-catalog-reader
owner: partner-integrations
purpose: read published catalog changes
hostname: api.example.com
methods: [GET]
paths: [/v1/catalog, /v1/catalog/delta]
authentication: oauth2-client-credentials
source_ranges: [203.0.113.0/28]
user_agent: partner-catalog/3.x
expected_rate_per_minute: 60

evidence:
waf_transaction_id: required
application_client_id: required
source_ip: observed_not_assumed
owner_confirmation: required

decision:
no_identity_match: keep_blocking
valid_identity_and_wrong_path: deny
valid_identity_and_expected_path: canary_then_decide
rollback: restore_previous_waf_policy

A user-agent is descriptive metadata, not identity. Source ranges are stronger only when the partner controls them and publishes a change process. Application credentials, mTLS or signed requests provide better proof, but WAF might not be able to validate all of them. Record which control is enforced at the edge and which one is enforced by the origin.

Prove that Bot Manager made the decision

Bot Manager classifies traffic as bad, good or unknown. That classification is useful security evidence; it is not an authorization decision for your application. Preserve the rule set version, rule ID, action, hostname, path, client IP and transaction ID before changing policy.

kusto 02-bot-manager-events.kql
let Window = 2h;
let Host = "api.example.com";
AzureDiagnostics
| where TimeGenerated > ago(Window)
| where Category in ("FrontDoorWebApplicationFirewallLog", "ApplicationGatewayFirewallLog")
| extend Hostname=tostring(column_ifexists("host_s", column_ifexists("hostname_s", ""))),
       Path=tostring(column_ifexists("requestUri_s", "")),
       RuleSet=tostring(column_ifexists("ruleSetType_s", "")),
       RuleId=tostring(column_ifexists("ruleId_s", "")),
       Action=tostring(column_ifexists("action_s", "")),
       ClientIp=tostring(column_ifexists("clientIp_s", "")),
       UserAgent=tostring(column_ifexists("userAgent_s", "")),
       TransactionId=tostring(column_ifexists("trackingReference_s", column_ifexists("transactionId_g", "")))
| where Hostname == Host or isempty(Hostname)
| where RuleSet has "Bot" or RuleId startswith "Bot"
| summarize Requests=count(),
          Clients=dcount(ClientIp),
          Paths=make_set(Path, 15),
          UserAgents=make_set(UserAgent, 10),
          FirstSeen=min(TimeGenerated),
          LastSeen=max(TimeGenerated)
by RuleSet, RuleId, Action
| order by Requests desc

Diagnostic schemas vary by WAF resource and table mode, so verify the available columns in the workspace before relying on the query. The invariant is the evidence chain, not a specific column name.

Rule families matter. Bot100xxx signals bad-bot intelligence or identity falsification, Bot200xxx identifies known good-bot categories, and Bot300xxx covers unknown automation such as crawlers, HTTP clients, service agents or monitoring services. A request classified as Bot300300 because it uses a general-purpose SDK is not automatically malicious or trusted. It needs application context.

Correlate edge classification with application identity

Use the WAF transaction or tracking reference to find the same request in Front Door, Application Gateway and origin logs. Then compare source IP, path, method, application client ID, response and rate. A match on user-agent alone is insufficient.

text bot-correlation-matrix.txt
Expected partner request
Source belongs to the approved range
OAuth client ID is partner-catalog-reader
Method and path match the contract
Rate stays within the agreed envelope
WAF and origin timestamps correlate

Possible spoofing
User-agent matches but client ID does not
Source is outside the approved range
Request probes unrelated paths
Rate or method differs from the contract
Same identity string appears across many networks

Insufficient evidence
WAF transaction cannot be tied to an origin request
Authentication fields are not logged safely
Forwarded client IP is ambiguous
Partner cannot confirm a test window

Do not log tokens or secrets to make correlation easier. Log a stable client identifier, authentication outcome, route, correlation ID and bounded source information. If the request never reached the origin, run a controlled partner test with a unique correlation header and capture the corresponding WAF event.

Read rule ordering before creating an exception

Azure WAF custom rules are evaluated before managed rule sets. An Allow match can therefore stop later managed inspection. A narrowly written rule may still bypass protections that were never part of the incident.

Before approving an exception, inventory:

  • the policy and Bot Manager version attached to the affected hostname;
  • custom-rule priorities and terminating actions;
  • Bot Manager overrides for the matched rule or group;
  • the Default Rule Set action that would normally inspect the request;
  • other hostnames sharing the same WAF policy.

Changing an entire unknown-bot category from Block to Allow is rarely proportional for one partner. A user-agent-only custom allow rule is weaker still. Prefer controls that preserve managed inspection: authenticate the client at the application, keep the bot signal in logs, rate-limit the bounded route, or isolate automated integrations behind a dedicated hostname and policy when their traffic model is materially different.

Replay positive and negative controls

The candidate policy must be tested with traffic that should pass and traffic that must not inherit the exception. Detection-only observation is useful, but the final canary must exercise the exact action and rule ordering intended for production.

yaml 03-bot-policy-replay.yml
cases:
- name: approved_partner
  source: approved_range
  client_id: partner-catalog-reader
  path: /v1/catalog/delta
  expected: reach_origin_and_authorize

- name: spoofed_user_agent
  source: unapproved_range
  client_id: none
  path: /v1/catalog/delta
  expected: no_privileged_access

- name: valid_client_wrong_path
  source: approved_range
  client_id: partner-catalog-reader
  path: /admin/export
  expected: deny

- name: hostile_payload_on_allowed_path
  source: approved_test_range
  client_id: test-identity
  path: /v1/catalog
  expected: managed_rules_still_evaluated

- name: volume_breach
  source: approved_test_range
  client_id: test-identity
  path: /v1/catalog
  expected: rate_control_visible

promotion_requires:
- every request tied to a WAF transaction
- origin authentication result recorded
- negative controls retain protection
- no unrelated hostname or route affected
- previous policy version ready for rollback

Replay from a controlled source and keep payloads harmless. The hostile-payload case is a negative control for rule ordering, not an attack exercise against production. If it passes because a custom Allow short-circuits the managed rules, reject the design.

Choose the smallest defensible change

Keep blocking when the traffic matches bad-bot intelligence, falsifies a known identity, cannot be tied to an application identity, or operates outside the declared route and rate. A recognizable product name does not change that decision.

Keep the Bot Manager action at Log when classification is useful but application authentication should decide access. Pair it with route-level authorization, quotas and alerts on identity mismatch. This is often the right outcome for internal automation and partner SDKs classified as unknown.

Use a policy exception only when the client contract is stable, the match conditions are narrow, negative controls prove that managed protections remain effective, and the exception has an owner and review date. If those conditions cannot coexist in one shared policy, separate the hostname or policy rather than widening a global rule.

Validate and roll back

Canary the change on one route, hostname or policy association. During the observation window, compare Bot Manager events, legitimate 403 responses, origin authentication failures, request rate, latency and managed-rule detections with the pre-change baseline.

Promote only when the approved client succeeds with its real identity, spoofed variants gain no access, unrelated bot categories are unchanged and managed-rule evidence remains visible. Roll back immediately if the exception matches an unexpected client, suppresses managed-rule inspection, affects another hostname or increases unauthenticated origin traffic.

Rollback means restoring the previous policy version or association, confirming propagation, replaying the same matrix and revoking any temporary application credential or source exception. Reverting WAF does not undo an origin action already accepted during the window; review those requests by correlation ID.

Conclusion

Bot Manager tells you how Azure classifies automated traffic. It does not tell you whether that client is authorized to use your API. Treat classification, identity and application permission as separate controls.

The production decision should be explicit: keep blocking, retain logging and enforce identity at the origin, introduce a bounded exception, or isolate the automation behind a dedicated policy. An allow rule is acceptable only when the same replay proves both recovery for the intended client and continued protection against its simplest impersonation.