Infrastructure

Azure Monitor Private Link: diagnose ingestion loss before reopening public access

A production runbook to isolate AMPLS, DCE, DCR, DNS, associations and Azure Monitor Agent when telemetry disappears after network hardening, then validate or roll back without broadly reopening public access.

11 Aug 2026 azureazure-monitorprivate-linkamplsdata-collection-endpointdcedata-collection-ruledcrazure-monitor-agentdnsobservabilitynetworkingkqlrunbookrollbackproduction

A platform team disables public access to Azure Monitor after connecting its network to an Azure Monitor Private Link Scope. The deployment is green, the private endpoint is approved and the machines remain reachable. Twenty minutes later, heartbeats become sparse and operating system logs disappear. Re-enabling public workspace access may restore the flow, but it does not reveal whether the break sits in DNS, the AMPLS boundary, the Data Collection Endpoint, the DCR association or the agent itself.

Azure Monitor Private Link is not simply an endpoint in front of a workspace. An Azure Monitor Agent path has at least two functions: retrieve configuration and ingest telemetry. Those functions can use different endpoints. This runbook drives one operational decision: repair the private path with a bounded change, keep a controlled access mode temporarily, or roll back the last network hardening change with enough evidence to revisit it safely.

State exactly what stopped arriving

Separate absence, delay and partial loss first. Missing heartbeats from every machine do not point to the same failure as one missing DCR stream. An empty query can also result from the wrong time window or a transformation while ingestion remains healthy.

yaml monitor-private-link-incident-scope.yml
incident:
started_at_utc: "<timestamp>"
last_known_good_change: "<change-id>"
affected_networks:
  - "<vnet-or-on-prem-segment>"
representative_sources:
  - "<vm-or-arc-resource-id>"

signals:
heartbeat_missing: true
all_dcr_streams_missing: false
one_stream_missing: true
query_access_healthy: true

expected_path:
configuration: "source -> DNS -> DCE -> Azure Monitor configuration service"
ingestion: "source -> DNS -> DCE or workspace ingestion endpoint"
query: "operator or service -> AMPLS query endpoint -> workspace"

change_window:
- AMPLS access mode
- DCE public network access
- AMPLS scoped resources
- private DNS links or forwarders
- DCR or DCR association
- route, firewall or proxy policy

Select at least one affected source and one healthy control. You will replay the same checks against both after the correction. Without that control, a natural recovery in ingestion can look like a successful fix.

Map AMPLS, DCE, DCR and destination

AMPLS defines the private boundary around Azure Monitor resources. A DCE exposes endpoints used for collection and, depending on the scenario, configuration retrieval. A DCR describes sources, transformations and destinations. Associations connect a machine or cluster to its DCR and can also connect it to the DCE used to fetch configuration.

Inventory those objects before changing a network rule. The objective is to prove that the affected source belongs to the intended resource chain, in the correct region, and that every required monitoring resource is included in the private scope.

bash 01-freeze-monitor-private-link-state.sh
SUBSCRIPTION_ID="<subscription-id>"
RG="<observability-resource-group>"
AMPLS="<ampls-name>"
DCE="<dce-name>"
DCR="<dcr-name>"
SOURCE_ID="<vm-or-arc-resource-id>"

az account set --subscription "$SUBSCRIPTION_ID"

az monitor private-link-scope show --resource-group "$RG" --name "$AMPLS" --output json > ampls-before.json

az monitor data-collection endpoint show --resource-group "$RG" --name "$DCE" --output json > dce-before.json

az monitor data-collection rule show --resource-group "$RG" --name "$DCR" --output json > dcr-before.json

az monitor data-collection rule association list --resource "$SOURCE_ID" --output json > associations-before.json

Add the AMPLS scoped resources, the private endpoint and its IP addresses, private DNS zones and VNet links to that snapshot. Keep machine-readable outputs in the incident record. Portal screenshots are poor material for a clean before-and-after comparison.

Test DNS from the network that is actually failing

AMPLS changes name resolution for several Azure Monitor endpoints. Some are shared and some are resource-specific. Correct resolution from an administrator workstation proves nothing for a VM using a custom resolver, an on-premises forwarder or another VNet.

From the affected source, retrieve the endpoints declared by the DCE and DCR, then resolve those names with the DNS configuration used by the machine itself.

bash 02-check-dce-dns-path.sh
DCE_CONFIG_ENDPOINT="<configuration-endpoint-from-dce>"
DCE_INGEST_ENDPOINT="<ingestion-endpoint-from-dce-or-dcr>"

printf '%s
' "$DCE_CONFIG_ENDPOINT" "$DCE_INGEST_ENDPOINT"

getent ahosts "<configuration-fqdn>"
getent ahosts "<ingestion-fqdn>"

curl --connect-timeout 5 -sv "https://<configuration-fqdn>/" -o /dev/null
curl --connect-timeout 5 -sv "https://<ingestion-fqdn>/" -o /dev/null

An HTTP error can be expected without authentication; this test first proves name resolution, route, TCP connection and TLS negotiation. A timeout points to the network path. A certificate for another host suggests a proxy or TLS inspection issue. A public address after switching to private-only access points to DNS zones or forwarding. An unreachable private address brings the investigation back to routes, NSGs or the firewall.

Do not add a manual DNS record before identifying the missing zone or link. Azure Monitor shared endpoints make duplicate private zones especially difficult to operate.

Verify configuration and ingestion separately

An agent extension can report successful provisioning while the agent no longer receives configuration. Conversely, the agent may still know its DCR but be unable to reach the ingestion endpoint. Classify the symptom with a small decision matrix.

text monitor-path-decision-matrix.txt
Extension provisioned, no recent heartbeat, all streams missing
Check agent process, DCE configuration association, configuration endpoint and identity.

Heartbeat present, one stream missing
Check DCR source, stream, transformation, destination and processing errors.

All sources in one VNet fail after AMPLS change
Check DNS path, AMPLS access mode, scoped resources, route and firewall.

Only one source fails
Check its DCR associations, agent logs, local DNS and machine identity.

Queries fail but ingestion continues
Diagnose query access separately; do not rewrite the collection path.

Logs ingestion API fails while AMA is healthy
Check client identity, DCR immutable ID, stream name and the DCE or DCR ingestion endpoint used by that client.

This separation prevents two common false fixes: reinstalling agents when an entire network segment can no longer resolve the DCE, or editing the DCR when agents cannot retrieve it.

Use heartbeat and DCR metrics as evidence

Heartbeat confirms that the agent still sends data, but it does not prove that every business-critical stream is healthy. Compare the latest arrival by resource, then inspect expected tables and ingestion latency over the same window.

kusto 03-monitor-private-link-ingestion-check.kql
let Window = 2h;
let ExpectedResources = dynamic([
"<affected-resource-id>",
"<healthy-control-resource-id>"
]);
Heartbeat
| where TimeGenerated > ago(Window)
| where Category == "Azure Monitor Agent"
| where _ResourceId in~ (ExpectedResources)
| summarize
  LastHeartbeat=max(TimeGenerated),
  Heartbeats=count(),
  P95IngestionDelay=percentile(ingestion_time() - TimeGenerated, 95)
by _ResourceId
| extend Status=case(
  LastHeartbeat < ago(10m), "missing",
  P95IngestionDelay > 10m, "late",
  "healthy")
| order by Status asc, _ResourceId asc

Add an equivalent query for every critical table and review DCR processing error logs or metrics. If heartbeat returns while one stream remains absent, the private path is unlikely to remain the primary suspect. Move back to the DCR contract, its transformation and destination.

Choose the smallest correction

Match the correction to the evidence. Add a missing DCE or workspace to the scope when the AMPLS boundary is incomplete. Repair a DCE association when agents cannot fetch configuration. Fix a zone link or forwarding rule when the source gets the wrong address. Open a firewall flow only for the proven FQDN, port and source segment, following the approved network model.

Avoid changing the AMPLS mode, DCR, DNS zones and agent at the same time. Observability incidents are especially sensitive to fixes that cannot be attributed. Apply one change, allow for the expected propagation time, and replay the same evidence.

Validate, hold under watch or roll back

Validation must prove that the private path is restored, not merely that a few rows have reappeared.

text monitor-private-link-validation-rollback.txt
Validate
Affected and control sources resolve the expected endpoints.
DNS answers use the intended private path.
TCP and TLS succeed without a hosts-file override.
DCE and DCR associations match the approved design.
Heartbeat returns within the agreed delay.
Every critical stream resumes without abnormal ingestion latency.
DCR processing errors remain within the expected baseline.
Public access remains in the intended state.

Rollback
Restore the previous AMPLS access mode or scoped-resource set.
Restore the previous DCE association or private DNS link.
Revert the last route or firewall policy change.
Do not delete the new objects until before/after evidence is retained.
Replay DNS, connectivity, heartbeat and critical-stream checks.
Set a deadline to remove any temporary public-access exception.

If temporary rollback to a more open mode is required to restore observability, bound it by network, resource, owner and duration. Reopening access without a deadline turns a collection incident into silent security debt.

Conclusion

Ingestion loss after enabling Azure Monitor Private Link should be handled as a broken chain. AMPLS defines the boundary, DCE carries the endpoints, DCR defines the flow, associations target the sources and DNS selects the path actually taken. A deployed state on each object does not prove that the chain works.

The decision should rest on four replayable proofs: resolution from the affected source, access to the expected endpoints, configuration retrieval and recovery of critical streams. With those proofs, the team can repair precisely or roll back cleanly instead of using broad public access as a permanent diagnostic tool.