Automation

Azure Resource Graph: prove inventory completeness before automated remediation

A production runbook for proving Azure Resource Graph scope, identity, pagination and result completeness before an inventory drives automated changes.

18 Sept 2026 azureresource-graphinventoryautomationrbacpaginationobservabilityguardrailsrunbookrollbackproduction

A platform job inventories production resources with Azure Resource Graph, selects those missing an ownership tag, then prepares a remediation batch. One morning, the candidate count drops by thirty percent. The job is green and the query still returns rows. That does not prove the estate improved: the runtime identity may see fewer subscriptions, the query may be truncated, or the client may have stopped after the first page.

An incomplete inventory is dangerous precisely because it can look successful. This runbook turns completeness into an explicit gate before any write. The running case is a multi-subscription tagging workflow, but the same controls apply to cost inventories, security posture checks, decommissioning candidates and change preparation.

Freeze the inventory contract

Stop state-changing steps while keeping the read-only query available. Record the identity, tenant, management-group scope, expected subscriptions, query version, client version and previous successful totals. A row count without this context is not reproducible evidence.

yaml resource-graph-inventory-contract.yml
inventory: prod-owner-tag-candidates
runtime_identity: mi-platform-inventory-prod
tenant_id: <tenant-id>
management_groups:
- mg-production
expected_subscriptions_source: platform-subscription-registry
query_version: 8f2c1d7
client: azure-cli-2.x

write_gate:
mode: disabled
require:
  - expected-subscriptions-covered
  - every-page-consumed
  - result-not-partial
  - control-totals-within-baseline
  - positive-and-negative-canary

rollback_assets:
- previous-query
- previous-scope-manifest
- previous-identity-assignment
- per-resource-change-journal

Use a maintained subscription registry as the expected set. Deriving the expected set from the same Resource Graph call creates a circular check: a missing subscription disappears from both the inventory and its validation.

Prove the identity and subscription scope

Resource Graph only returns resources the principal can read. It does not add placeholder rows for subscriptions hidden by RBAC. A successful response can therefore be complete for the caller and incomplete for the operating model.

Run the diagnostic under the job’s real managed identity or service principal. Compare its visible and enabled subscriptions with the approved registry, then verify Reader-equivalent control-plane access at the intended scope. Refresh the authentication context after subscription or role changes before retesting.

bash 01-compare-visible-subscriptions.sh
az account show \
--query '{tenant:tenantId, subscription:id, user:user}' \
--output json

az account list --all \
--query "[?state=='Enabled'].{id:id,name:name,tenant:tenantId}" \
--output json > visible-subscriptions.json

az role assignment list \
--assignee-object-id <runtime-principal-object-id> \
--all --include-inherited \
--query "[].{scope:scope,role:roleDefinitionName}" \
--output json > runtime-rbac.json

Do not solve an unexplained gap by granting broad tenant access. Identify the exact missing subscription or management-group branch and restore the intended least-privilege assignment.

Make the query pageable by design

Use a deterministic sort and project a scalar resource ID. Avoid limit, take or sample in a production inventory: they intentionally bound the result and can prevent a continuation token from being returned. Queries whose projected columns are only dynamic or null also cannot be treated as safely pageable.

kusto 02-owner-tag-candidates.kql
Resources
| where tostring(tags.Environment) =~ "prod"
| extend Owner = tostring(tags.Owner)
| where isempty(Owner)
| project id = tolower(id), subscriptionId, resourceGroup,
        type = tolower(type), name, location
| order by id asc

Keep the query and scope identical across pages. A continuation token captures paging context; changing the query, subscriptions or management group between requests invalidates the inventory contract.

Inspect response metadata, not only rows

The REST response exposes count, totalRecords, resultTruncated and, when available, $skipToken. Preserve them for every page. Continue until no token remains, and fail closed if the response says it is truncated without giving the client a safe continuation path.

text resource-graph-page-gate.txt
For every response page
Record HTTP status and request correlation ID
Record count, totalRecords and resultTruncated
Record whether $skipToken is present
Record Resource Graph quota headers
Record x-ms-tenant-subscription-limit-hit when present
Append rows only after schema validation

Reject the inventory when
resultTruncated is true and no safe continuation exists
a page repeats a previous token
duplicate resource IDs appear unexpectedly
covered subscriptions differ from the approved registry
retry budget expires after throttling

Treat partial-scope execution as an explicit degraded mode, not a convenience flag. If allowPartialScopes is enabled, surface that fact in the evidence and keep writes disabled until the omitted scope is known and accepted. Likewise, distinguish Resource Graph throttling from Azure Resource Manager throttling and back off using the returned headers instead of starting overlapping jobs.

Build independent control totals

Completeness cannot be proved by comparing a query with itself. Produce totals by subscription and resource type, then compare them with the last known-good inventory and with an independent subscription registry. Investigate both decreases and unexplained increases.

kusto 03-inventory-control-totals.kql
Resources
| summarize resources = count(),
          resourceGroups = dcount(resourceGroup)
by subscriptionId, type = tolower(type)
| order by subscriptionId asc, type asc

For one resource in each critical subscription, compare the Resource Graph row with a direct ARM GET. This is a bounded freshness check, not a reason to replace the scalable inventory with thousands of ARM calls. A recently changed resource can take time to appear in the index, so measure the observation window and retry before classifying it as permanently absent.

Normalize IDs to lowercase, sort them and calculate a digest of the complete set. Store the page count, row count, subscription count and digest beside the query version. These values make a silent first-page regression visible in the next run.

Separate selection from execution

The query should produce candidates, not authorize changes. Materialize an immutable candidate snapshot, attach the completeness evidence, then require a separate approval or policy gate before remediation. Re-querying during execution can silently change the target set.

yaml remediation-batch-evidence.yml
batch_id: owner-tags-20260918-01
inventory_digest: sha256:<digest>
query_version: 8f2c1d7
pages: 4
rows: 3274
subscriptions_expected: 18
subscriptions_observed: 18
partial_scope: false
truncated: false
candidate_rows: 42

execution:
mode: canary
max_targets: 1
require_owner_value_from: approved-cmdb
idempotency_key: batch-id-plus-resource-id
journal_before_and_after: true

Reject free-form tag values inferred from resource names. The inventory identifies a missing field; it does not prove the correct owner. Resolve that value from an approved source or route the candidate for human review.

Validate with positive and negative canaries

Choose one approved candidate and one control resource that must remain untouched. Apply the bounded change to the candidate, read it back through ARM, then wait for Resource Graph to reflect the new state. Confirm that the control resource remains unchanged and that a second execution produces no write.

Validation fails if the canary disappears from the inventory without the ARM state changing, if the negative control is selected, if the job cannot explain every page, or if a rerun writes again. These failures point to inventory, selection or idempotency defects rather than a tagging issue.

Decide resume, hold or rollback

Resume remediation only when the runtime identity covers the approved subscription set, every page has been consumed, no partial-scope or truncation signal remains, control totals are explainable, and both canaries pass.

Hold in read-only mode when the inventory is useful but freshness, subscription coverage or paging evidence is incomplete. Repair the narrow boundary: RBAC, authentication context, scope manifest, query projection, continuation loop or throttle handling.

Rollback when a changed query, identity or client release caused the gap. Restore the previous version, disable the schedule and regenerate the inventory before writing again. If changes already ran, use the per-resource journal to restore previous values only on touched resources; never “roll back” by applying a second broad query to a scope that is still unproven.

Conclusion

Azure Resource Graph can return a valid response that is still unsafe as an automation inventory. Completeness depends on identity, subscription scope, query shape, pagination, partial-result signals and observable control totals.

Make those properties a write gate. When the evidence is complete, promote one idempotent canary and resume in bounded batches. When it is not, keep the query read-only, repair the failed boundary and preserve a precise rollback journal.