Infrastructure

Azure Monitor DCR: diagnose missing logs before rolling back a transformation

A production runbook for qualifying missing logs after a Data Collection Rule change, separating source, stream, KQL transformation and destination, then validating or rolling back with a canary.

18 Aug 2026 azureazure-monitordata-collection-rulelog-analyticskqlobservabilityautomationrunbookrollbackproduction

An operations team reduces noise in a Log Analytics table with a KQL transformation in a Data Collection Rule (DCR). Deployment succeeds and ingestion costs appear to drop. During the next incident, however, useful events are missing. Removing the transformation immediately may restore volume, but it does not prove which records were lost or whether the transformation caused the gap.

This runbook covers a source using Azure Monitor Agent or the Logs Ingestion API, a DCR data flow, a transformKql transformation and a destination table. It separates four failure families: the source stopped sending, the client uses the wrong stream, the transformation filters or rejects records, or the destination rejects the output schema. The outcome is an evidence-backed keep, repair or rollback decision.

Freeze the flow before changing the DCR

A DCR is more than a KQL query. The full path connects a source, declared stream, data flow, transformation, destination and table. Capture that contract for the incident window.

text dcr-incident-scope.txt
Incident: inc-20260818-04
UTC window: 2026-08-18 05:30 - 06:30
DCR resource: /subscriptions/.../dataCollectionRules/dcr-app-events-prod
Immutable ID: dcr-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
Source: api-ingestion-app-events
Expected stream: Custom-AppEventsRawData
Destination: law-operations-prod
Table: AppEvents_CL
Configuration version before and after
Change commit or deployment ID
Expected volume and last known event

Evidence to retain
Complete DCR configuration before and after
Redacted canary payload
HTTP response and client request ID
DCR received, dropped and error metrics
DCRLogErrors records
Canary row in the target table or evidence of absence

For the Logs Ingestion API, verify the immutableId, not just the DCR resource name. Pin the exact stream sent by the client as well. A valid DCR will not process a payload posted to a different stream.

Export the configuration before making a repair.

bash 01-export-dcr.sh
az monitor data-collection rule show \
--resource-group rg-monitoring-prod \
--name dcr-app-events-prod \
--output json > dcr-app-events-prod.json

az monitor data-collection rule show \
--resource-group rg-monitoring-prod \
--name dcr-app-events-prod \
--query '{id:id,immutableId:immutableId,dataFlows:dataFlows,destinations:destinations}' \
--output jsonc

Prove that the source reaches ingestion

A missing target-table row does not prove a transformation problem. Start at the system boundary.

For a Logs Ingestion API client, retain the HTTP status, UTC timestamp, record count and a unique x-ms-client-request-id. A successful HTTP response proves request acceptance, not that every record reached the table. For Azure Monitor Agent, verify the DCR association, agent health and source continuity on the affected machine.

Then compare DCR metrics for the same stream and window: rows received, rows dropped, transformation errors, ingestion bytes and transformation duration.

text dcr-metric-reading.txt
Received rows fall to zero
Check producer, DCR association, endpoint and stream

Received rows stay stable while dropped rows rise
Check the KQL filter and transformation errors

Transformation errors rise
Compare the actual input shape with transformKql columns and conversions

Ingestion bytes stay stable while the target table falls
Check destination, table, ingestion delay and the read query

Transformation duration rises
Look for an expensive expression or a change in input shape

Do not compare row counts alone. A transformation may intentionally remove noise. The useful signal is the relationship between what the source sends, what the DCR receives, what it drops and what the table retains.

Read DCR processing errors

DCR errors are available only when resource logs for the rule are routed to a workspace. Enable the error category before a risky change; enabling it after an incident cannot recover past errors.

Query DCRLogErrors for the resource and incident window. Detailed columns may vary by collection scenario, so start with the common evidence and preserve the raw error description.

kusto 02-dcr-processing-errors.kql
let StartTime = datetime(2026-08-18T05:30:00Z);
let EndTime = datetime(2026-08-18T06:30:00Z);
let DcrResourceId = "/subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/rg-monitoring-prod/providers/microsoft.insights/dataCollectionRules/dcr-app-events-prod";
DCRLogErrors
| where TimeGenerated between (StartTime .. EndTime)
| where _ResourceId =~ DcrResourceId
| project TimeGenerated, OperationName, ResultType, ResultDescription, _ResourceId
| order by TimeGenerated desc

A conversion failure, missing required column or output mismatch points toward transformation or schema. A delivery error points toward the destination. No error does not prove the transformation is healthy: a valid where clause that filters every record may produce no processing error.

Review the transformation as a schema contract

In an Azure Monitor transformation, source represents each incoming record. The query must return at most one record for each input and its result must match the output stream schema. A short transformation can silently remove an event class after a casing, default-value or JSON-shape change.

kusto current-transform.kql
source
| where tostring(Severity) in ("Warning", "Error", "Critical")
| extend TimeGenerated = todatetime(EventTime),
       Service = tostring(ServiceName),
       CorrelationId = tostring(CorrelationId),
       Payload = tostring(Payload)
| project TimeGenerated, Service, Severity, CorrelationId, Payload

Review every stage against real, redacted samples:

  • actual Severity values, including casing and nulls;
  • whether EventTime remains convertible to a timestamp;
  • whether a producer renamed ServiceName or nested Payload;
  • whether projected columns exactly match the target table;
  • whether the filter removes an event class required by alerts or investigations.

Do not add summarize, join or multi-record logic to compensate for a source problem. Standard ingestion transformations operate on individual records and support only a subset of KQL.

Send a canary across every meaningful branch

A useful canary contains more than one record expected to pass. Include one record expected to be filtered and one close to a schema boundary. Use unique identifiers, no sensitive data and no business side effects.

json dcr-canary-payload.json
[
{
  "EventTime": "2026-08-18T06:10:00Z",
  "ServiceName": "billing-api",
  "Severity": "Error",
  "CorrelationId": "dcr-canary-20260818-pass",
  "Payload": "controlled validation event"
},
{
  "EventTime": "2026-08-18T06:10:01Z",
  "ServiceName": "billing-api",
  "Severity": "Information",
  "CorrelationId": "dcr-canary-20260818-filter",
  "Payload": "expected to be filtered"
},
{
  "EventTime": "2026-08-18T06:10:02Z",
  "ServiceName": "billing-api",
  "Severity": null,
  "CorrelationId": "dcr-canary-20260818-null",
  "Payload": "schema boundary"
}
]

Wait for the observed ingestion delay on this flow, then query the expected identifiers. The first record must arrive. The second must be absent for a documented reason. The third must follow the explicit policy for null severity. If diagnosis remains ambiguous, replay the same set against the previous and current transformation outside production.

Decide between repair and rollback

The decision must name both the evidence and its exit test.

yaml dcr-change-decision.yml
decision:
action: rollback_transform
reason:
  - received rows stayed stable after deployment
  - dropped rows increased on the target stream
  - Error canary was absent with the new transformation
  - Error canary arrived with the previous version
scope:
  dcr: dcr-app-events-prod
  data_flow: Custom-AppEventsRawData to AppEvents_CL
validation:
  - passing canary visible in the target table
  - Information record still filtered as intended
  - no new DCRLogErrors entries
  - dependent alerts replayed over a controlled window
stop_conditions:
  - ingestion volume exceeds baseline
  - sensitive data appears in Payload
  - transformation errors persist after rollback

Repair forward when the defect is narrow, understood and tested, such as a nullable conversion or a newly introduced severity value. Roll back when critical events disappear, the real schema is unknown or several producers are affected. If the source sends nothing, leave the transformation alone and restore the producer or DCR association first.

After rollback, retain the faulty configuration, metrics, errors and canary identifiers. Validate dependent alerts and dashboards too. Restoring records without restoring detection is only a partial recovery.

Conclusion

Missing logs after a DCR change cannot be diagnosed from the target table alone. Follow the flow through emission, endpoint or association, stream, received rows, transformation, errors and destination.

Keep the change only when a canary proves both retained and filtered records, metrics remain within baseline and downstream consumers work. Otherwise roll back the bounded transformation, preserve the evidence pack and repair the schema contract before another production deployment.