A production runbook for measuring split-by dimension cardinality, containing notification noise, canarying stable grouping and rolling back an Azure Monitor log alert without losing coverage.
A production runbook for qualifying OpenTelemetry tail sampling across trace affinity, late spans, memory pressure, policy coverage, Azure Monitor evidence, canary rollout and rollback.
A production runbook for isolating the metric, target and labels behind an AKS time-series spike before filtering collection, adding capacity or rolling back.
A production runbook for separating a faulty health endpoint, instance degradation, redirects, authentication and dependency failures before restarting or changing App Service Health Check.
A production runbook for qualifying missing logs after a Data Collection Rule change, separating source, stream, KQL transformation and destination, then validating or rolling back with a canary.
A production runbook for classifying dead-lettered messages, proving the cause, checking idempotency and replaying with a canary without creating a second incident.
A production runbook to separate ingestion pressure, throttling, hot partitions, consumer failures and checkpoint drift before adding capacity or replaying events.
A production runbook to isolate AMPLS, DCE, DCR, DNS, associations and Azure Monitor Agent when telemetry disappears after network hardening, then validate or roll back without broadly reopening public access.
A production runbook for building and qualifying an Azure Monitor burn-rate alert by separating the SLI, error budget, short and long windows, telemetry quality, notification, validation and rollback.
A production runbook for canarying a KQL transformation in an Azure Monitor Data Collection Rule, comparing volume and schema, detecting rejected records, then validating or rolling back without an observability gap.
A production runbook for qualifying Azure Workbooks drift with KQL, Log Analytics, Application Insights, dimensions, ingestion latency, validation and rollback before changing alerts or dashboards.
A production runbook for qualifying apparent Application Insights telemetry loss with sampling, ingestion, SDK configuration, KQL, alerts, validation and rollback before changing thresholds.
A production runbook for qualifying an Azure Monitor Action Group by separating rules, receivers, webhooks, escalation paths, KQL evidence, validation and rollback before reducing notifications.
A production runbook for qualifying a delayed Azure Monitor alert by separating application timestamps, Log Analytics ingestion, KQL query windows, evaluation frequency, action groups, validation and rollback.
A production runbook for qualifying an Azure Managed Grafana dashboard or alert with Azure Monitor datasource, managed identity, Log Analytics permissions, variables, traces, validation and rollback before changing KQL queries.
A production runbook for qualifying an Azure Monitor alert that fired but did not notify anyone, with action groups, receivers, processing rules, webhooks, evidence, validation and rollback.
A production runbook for qualifying missing Azure Monitor logs with Diagnostic Settings, DCRs, ingestion, KQL, cost controls, validation and rollback before changing alerts.
A production runbook for deciding an Azure rollback after deployment with Azure Monitor, KQL, impact correlation, regression evidence, validation and controlled recovery.
A production runbook for qualifying an Azure Monitor alert storm after deployment by separating real signal, noise, regression, threshold drift, action group behavior and rollback.