Incident Console

Choose a symptom. Leave with an action path.

Naxaya turns field notes into a compact response surface: evidence, diagnosis, bounded action, rollback and the exact notes to open next.

Cloud

Application Gateway returns 502

Open atlas

Separate backend health, DNS resolution, TLS settings and private network reachability before changing the application.

01

Evidence

Backend health state, probe result, gateway diagnostic logs, DNS answer from the gateway path and TLS/SNI settings.

02

First checks

  • Check backend health state
  • Resolve backend name from the gateway path
  • Validate TLS/SNI and probe configuration
03

Bounded action

Change only the failing boundary: probe, backend FQDN, certificate binding or route. Retest the same request path after each change.

04

Rollback

Restore previous probe/backend settings and keep the captured failing timestamp for comparison.

copy packs Incident exports
Short handoff
[Incident] Application Gateway returns 502
Context: Separate backend health, DNS resolution, TLS settings and private network reachability before changing the application.
Evidence to confirm: Backend health state, probe result, gateway diagnostic logs, DNS answer from the gateway path and TLS/SNI settings.
Immediate checks: Check backend health state | Resolve backend name from the gateway path | Validate TLS/SNI and probe configuration
Proposed action: Change only the failing boundary: probe, backend FQDN, certificate binding or route. Retest the same request path after each change.
Rollback: Restore previous probe/backend settings and keep the captured failing timestamp for comparison.
Post-incident review
Handled symptom: Application Gateway returns 502
Initial hypothesis: Separate backend health, DNS resolution, TLS settings and private network reachability before changing the application.
Evidence used: Backend health state, probe result, gateway diagnostic logs, DNS answer from the gateway path and TLS/SNI settings.
Checks performed: Check backend health state | Resolve backend name from the gateway path | Validate TLS/SNI and probe configuration
Decision / action: Change only the failing boundary: probe, backend FQDN, certificate binding or route. Retest the same request path after each change.
Rollback plan: Restore previous probe/backend settings and keep the captured failing timestamp for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Cloud

Internal APIM returns an error on a private API

Open atlas

Correlate Application Gateway/WAF and APIM logs, then separate DNS, TLS, policy, identity and private backend reachability before changing policies or opening access.

01

Evidence

Correlate Application Gateway/WAF and APIM logs, then separate DNS, TLS, policy, identity and private backend reachability before changing policies or opening access.

02

First checks

  • Check whether WAF blocked the request
  • Confirm APIM received the same path
  • Validate backend DNS and TLS from the APIM path
  • Replay with a correlation ID
03

Bounded action

Run the first checks in order: Check whether WAF blocked the request | Confirm APIM received the same path | Validate backend DNS and TLS from the APIM path | Replay with a correlation ID. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] Internal APIM returns an error on a private API
Context: Correlate Application Gateway/WAF and APIM logs, then separate DNS, TLS, policy, identity and private backend reachability before changing policies or opening access.
Evidence to confirm: Correlate Application Gateway/WAF and APIM logs, then separate DNS, TLS, policy, identity and private backend reachability before changing policies or opening access.
Immediate checks: Check whether WAF blocked the request | Confirm APIM received the same path | Validate backend DNS and TLS from the APIM path | Replay with a correlation ID
Proposed action: Run the first checks in order: Check whether WAF blocked the request | Confirm APIM received the same path | Validate backend DNS and TLS from the APIM path | Replay with a correlation ID. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: Internal APIM returns an error on a private API
Initial hypothesis: Correlate Application Gateway/WAF and APIM logs, then separate DNS, TLS, policy, identity and private backend reachability before changing policies or opening access.
Evidence used: Correlate Application Gateway/WAF and APIM logs, then separate DNS, TLS, policy, identity and private backend reachability before changing policies or opening access.
Checks performed: Check whether WAF blocked the request | Confirm APIM received the same path | Validate backend DNS and TLS from the APIM path | Replay with a correlation ID
Decision / action: Run the first checks in order: Check whether WAF blocked the request | Confirm APIM received the same path | Validate backend DNS and TLS from the APIM path | Replay with a correlation ID. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Networking

Private Endpoint name still resolves publicly

Open atlas

Confirm the CNAME chain, Private DNS Zone association and hybrid forwarding from the consuming network.

01

Evidence

nslookup from the workload subnet, CNAME chain, private DNS zone links, resolver forwarding path and cached answers.

02

First checks

  • Run nslookup from the workload network
  • Check privatelink CNAME
  • Verify Private DNS Zone links and forwarders
03

Bounded action

Fix zone association or forwarding first, then clear caches and retest from the consuming network.

04

Rollback

Restore the previous link or forwarder and document the public/private answer difference.

copy packs Incident exports
Short handoff
[Incident] Private Endpoint name still resolves publicly
Context: Confirm the CNAME chain, Private DNS Zone association and hybrid forwarding from the consuming network.
Evidence to confirm: nslookup from the workload subnet, CNAME chain, private DNS zone links, resolver forwarding path and cached answers.
Immediate checks: Run nslookup from the workload network | Check privatelink CNAME | Verify Private DNS Zone links and forwarders
Proposed action: Fix zone association or forwarding first, then clear caches and retest from the consuming network.
Rollback: Restore the previous link or forwarder and document the public/private answer difference.
Post-incident review
Handled symptom: Private Endpoint name still resolves publicly
Initial hypothesis: Confirm the CNAME chain, Private DNS Zone association and hybrid forwarding from the consuming network.
Evidence used: nslookup from the workload subnet, CNAME chain, private DNS zone links, resolver forwarding path and cached answers.
Checks performed: Run nslookup from the workload network | Check privatelink CNAME | Verify Private DNS Zone links and forwarders
Decision / action: Fix zone association or forwarding first, then clear caches and retest from the consuming network.
Rollback plan: Restore the previous link or forwarder and document the public/private answer difference.
To improve: detection, runbook, guardrail, ownership and communication delay.
Cloud

Azure Storage private endpoint returns 403, times out or produces no request logs

Open atlas

Separate Storage subresource DNS, Private Endpoint approval, firewall rules, runtime identity and Storage logs before opening public access or broadening RBAC.

01

Evidence

Separate Storage subresource DNS, Private Endpoint approval, firewall rules, runtime identity and Storage logs before opening public access or broadening RBAC.

02

First checks

  • Resolve the exact Storage subresource from the workload network
  • Check Private Endpoint status and private DNS zone group
  • Replay with a client request ID
  • Correlate Storage logs for 403, caller IP and requester identity
03

Bounded action

Run the first checks in order: Resolve the exact Storage subresource from the workload network | Check Private Endpoint status and private DNS zone group | Replay with a client request ID | Correlate Storage logs for 403, caller IP and requester identity. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] Azure Storage private endpoint returns 403, times out or produces no request logs
Context: Separate Storage subresource DNS, Private Endpoint approval, firewall rules, runtime identity and Storage logs before opening public access or broadening RBAC.
Evidence to confirm: Separate Storage subresource DNS, Private Endpoint approval, firewall rules, runtime identity and Storage logs before opening public access or broadening RBAC.
Immediate checks: Resolve the exact Storage subresource from the workload network | Check Private Endpoint status and private DNS zone group | Replay with a client request ID | Correlate Storage logs for 403, caller IP and requester identity
Proposed action: Run the first checks in order: Resolve the exact Storage subresource from the workload network | Check Private Endpoint status and private DNS zone group | Replay with a client request ID | Correlate Storage logs for 403, caller IP and requester identity. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: Azure Storage private endpoint returns 403, times out or produces no request logs
Initial hypothesis: Separate Storage subresource DNS, Private Endpoint approval, firewall rules, runtime identity and Storage logs before opening public access or broadening RBAC.
Evidence used: Separate Storage subresource DNS, Private Endpoint approval, firewall rules, runtime identity and Storage logs before opening public access or broadening RBAC.
Checks performed: Resolve the exact Storage subresource from the workload network | Check Private Endpoint status and private DNS zone group | Replay with a client request ID | Correlate Storage logs for 403, caller IP and requester identity
Decision / action: Run the first checks in order: Resolve the exact Storage subresource from the workload network | Check Private Endpoint status and private DNS zone group | Replay with a client request ID | Correlate Storage logs for 403, caller IP and requester identity. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Cloud

Azure SQL private endpoint returns timeouts, firewall errors or no SQL logs

Open atlas

Separate SQL private DNS, Private Endpoint state, public access, firewall rules, runtime identity and SQL diagnostics before changing schema, code or broad permissions.

01

Evidence

Separate SQL private DNS, Private Endpoint state, public access, firewall rules, runtime identity and SQL diagnostics before changing schema, code or broad permissions.

02

First checks

  • Resolve the SQL FQDN from the workload network
  • Check Private Endpoint status and privatelink.database.windows.net records
  • Replay with the real runtime identity
  • Correlate SQL diagnostics for firewall and login errors
03

Bounded action

Run the first checks in order: Resolve the SQL FQDN from the workload network | Check Private Endpoint status and privatelink.database.windows.net records | Replay with the real runtime identity | Correlate SQL diagnostics for firewall and login errors. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] Azure SQL private endpoint returns timeouts, firewall errors or no SQL logs
Context: Separate SQL private DNS, Private Endpoint state, public access, firewall rules, runtime identity and SQL diagnostics before changing schema, code or broad permissions.
Evidence to confirm: Separate SQL private DNS, Private Endpoint state, public access, firewall rules, runtime identity and SQL diagnostics before changing schema, code or broad permissions.
Immediate checks: Resolve the SQL FQDN from the workload network | Check Private Endpoint status and privatelink.database.windows.net records | Replay with the real runtime identity | Correlate SQL diagnostics for firewall and login errors
Proposed action: Run the first checks in order: Resolve the SQL FQDN from the workload network | Check Private Endpoint status and privatelink.database.windows.net records | Replay with the real runtime identity | Correlate SQL diagnostics for firewall and login errors. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: Azure SQL private endpoint returns timeouts, firewall errors or no SQL logs
Initial hypothesis: Separate SQL private DNS, Private Endpoint state, public access, firewall rules, runtime identity and SQL diagnostics before changing schema, code or broad permissions.
Evidence used: Separate SQL private DNS, Private Endpoint state, public access, firewall rules, runtime identity and SQL diagnostics before changing schema, code or broad permissions.
Checks performed: Resolve the SQL FQDN from the workload network | Check Private Endpoint status and privatelink.database.windows.net records | Replay with the real runtime identity | Correlate SQL diagnostics for firewall and login errors
Decision / action: Run the first checks in order: Resolve the SQL FQDN from the workload network | Check Private Endpoint status and privatelink.database.windows.net records | Replay with the real runtime identity | Correlate SQL diagnostics for firewall and login errors. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Cloud

Azure Service Bus private endpoint times out, denies access or lets backlog grow

Open atlas

Separate Service Bus private DNS, Private Endpoint state, public access, managed identity or SAS, queue metrics and processing logs before touching queues or redeploying consumers.

01

Evidence

Separate Service Bus private DNS, Private Endpoint state, public access, managed identity or SAS, queue metrics and processing logs before touching queues or redeploying consumers.

02

First checks

  • Resolve the Service Bus FQDN from the workload network
  • Check Private Endpoint approval and public network access
  • Verify the real sender or receiver identity
  • Correlate backlog, dead-letter and Service Bus errors
03

Bounded action

Run the first checks in order: Resolve the Service Bus FQDN from the workload network | Check Private Endpoint approval and public network access | Verify the real sender or receiver identity | Correlate backlog, dead-letter and Service Bus errors. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] Azure Service Bus private endpoint times out, denies access or lets backlog grow
Context: Separate Service Bus private DNS, Private Endpoint state, public access, managed identity or SAS, queue metrics and processing logs before touching queues or redeploying consumers.
Evidence to confirm: Separate Service Bus private DNS, Private Endpoint state, public access, managed identity or SAS, queue metrics and processing logs before touching queues or redeploying consumers.
Immediate checks: Resolve the Service Bus FQDN from the workload network | Check Private Endpoint approval and public network access | Verify the real sender or receiver identity | Correlate backlog, dead-letter and Service Bus errors
Proposed action: Run the first checks in order: Resolve the Service Bus FQDN from the workload network | Check Private Endpoint approval and public network access | Verify the real sender or receiver identity | Correlate backlog, dead-letter and Service Bus errors. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: Azure Service Bus private endpoint times out, denies access or lets backlog grow
Initial hypothesis: Separate Service Bus private DNS, Private Endpoint state, public access, managed identity or SAS, queue metrics and processing logs before touching queues or redeploying consumers.
Evidence used: Separate Service Bus private DNS, Private Endpoint state, public access, managed identity or SAS, queue metrics and processing logs before touching queues or redeploying consumers.
Checks performed: Resolve the Service Bus FQDN from the workload network | Check Private Endpoint approval and public network access | Verify the real sender or receiver identity | Correlate backlog, dead-letter and Service Bus errors
Decision / action: Run the first checks in order: Resolve the Service Bus FQDN from the workload network | Check Private Endpoint approval and public network access | Verify the real sender or receiver identity | Correlate backlog, dead-letter and Service Bus errors. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Cloud

A synthetic probe fails on an Azure private path

Open atlas

Separate DNS, TLS, Application Gateway health, WAF blocks and runner network before changing routing or application code.

01

Evidence

Separate DNS, TLS, Application Gateway health, WAF blocks and runner network before changing routing or application code.

02

First checks

  • Resolve the hostname from the probe network
  • Check TLS/SNI with the real hostname
  • Correlate probe run with WAF and gateway logs
03

Bounded action

Run the first checks in order: Resolve the hostname from the probe network | Check TLS/SNI with the real hostname | Correlate probe run with WAF and gateway logs. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] A synthetic probe fails on an Azure private path
Context: Separate DNS, TLS, Application Gateway health, WAF blocks and runner network before changing routing or application code.
Evidence to confirm: Separate DNS, TLS, Application Gateway health, WAF blocks and runner network before changing routing or application code.
Immediate checks: Resolve the hostname from the probe network | Check TLS/SNI with the real hostname | Correlate probe run with WAF and gateway logs
Proposed action: Run the first checks in order: Resolve the hostname from the probe network | Check TLS/SNI with the real hostname | Correlate probe run with WAF and gateway logs. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: A synthetic probe fails on an Azure private path
Initial hypothesis: Separate DNS, TLS, Application Gateway health, WAF blocks and runner network before changing routing or application code.
Evidence used: Separate DNS, TLS, Application Gateway health, WAF blocks and runner network before changing routing or application code.
Checks performed: Resolve the hostname from the probe network | Check TLS/SNI with the real hostname | Correlate probe run with WAF and gateway logs
Decision / action: Run the first checks in order: Resolve the hostname from the probe network | Check TLS/SNI with the real hostname | Correlate probe run with WAF and gateway logs. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Networking

Azure traffic leaves through one path and returns through another

Open atlas

Compare DNS target, effective routes, route table associations, firewall evidence, NAT identity and return path before changing UDRs or bypassing inspection.

01

Evidence

Compare DNS target, effective routes, route table associations, firewall evidence, NAT identity and return path before changing UDRs or bypassing inspection.

02

First checks

  • Resolve the destination from the source path
  • Compare effective routes on source and destination NICs
  • Check firewall or appliance logs for both directions
  • Confirm NAT or outbound source identity
03

Bounded action

Run the first checks in order: Resolve the destination from the source path | Compare effective routes on source and destination NICs | Check firewall or appliance logs for both directions | Confirm NAT or outbound source identity. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] Azure traffic leaves through one path and returns through another
Context: Compare DNS target, effective routes, route table associations, firewall evidence, NAT identity and return path before changing UDRs or bypassing inspection.
Evidence to confirm: Compare DNS target, effective routes, route table associations, firewall evidence, NAT identity and return path before changing UDRs or bypassing inspection.
Immediate checks: Resolve the destination from the source path | Compare effective routes on source and destination NICs | Check firewall or appliance logs for both directions | Confirm NAT or outbound source identity
Proposed action: Run the first checks in order: Resolve the destination from the source path | Compare effective routes on source and destination NICs | Check firewall or appliance logs for both directions | Confirm NAT or outbound source identity. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: Azure traffic leaves through one path and returns through another
Initial hypothesis: Compare DNS target, effective routes, route table associations, firewall evidence, NAT identity and return path before changing UDRs or bypassing inspection.
Evidence used: Compare DNS target, effective routes, route table associations, firewall evidence, NAT identity and return path before changing UDRs or bypassing inspection.
Checks performed: Resolve the destination from the source path | Compare effective routes on source and destination NICs | Check firewall or appliance logs for both directions | Confirm NAT or outbound source identity
Decision / action: Run the first checks in order: Resolve the destination from the source path | Compare effective routes on source and destination NICs | Check firewall or appliance logs for both directions | Confirm NAT or outbound source identity. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Cloud

Azure Container Apps private ingress fails or reaches the wrong revision

Open atlas

Separate private DNS, Application Gateway handoff, Container Apps ingress mode, revision traffic and console logs before rolling back or changing traffic weights.

01

Evidence

Separate private DNS, Application Gateway handoff, Container Apps ingress mode, revision traffic and console logs before rolling back or changing traffic weights.

02

First checks

  • Resolve the hostname from the caller network
  • Check ingress target port and active revisions
  • Correlate system and console logs
03

Bounded action

Run the first checks in order: Resolve the hostname from the caller network | Check ingress target port and active revisions | Correlate system and console logs. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] Azure Container Apps private ingress fails or reaches the wrong revision
Context: Separate private DNS, Application Gateway handoff, Container Apps ingress mode, revision traffic and console logs before rolling back or changing traffic weights.
Evidence to confirm: Separate private DNS, Application Gateway handoff, Container Apps ingress mode, revision traffic and console logs before rolling back or changing traffic weights.
Immediate checks: Resolve the hostname from the caller network | Check ingress target port and active revisions | Correlate system and console logs
Proposed action: Run the first checks in order: Resolve the hostname from the caller network | Check ingress target port and active revisions | Correlate system and console logs. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: Azure Container Apps private ingress fails or reaches the wrong revision
Initial hypothesis: Separate private DNS, Application Gateway handoff, Container Apps ingress mode, revision traffic and console logs before rolling back or changing traffic weights.
Evidence used: Separate private DNS, Application Gateway handoff, Container Apps ingress mode, revision traffic and console logs before rolling back or changing traffic weights.
Checks performed: Resolve the hostname from the caller network | Check ingress target port and active revisions | Correlate system and console logs
Decision / action: Run the first checks in order: Resolve the hostname from the caller network | Check ingress target port and active revisions | Correlate system and console logs. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Cloud

AKS private ingress returns 502 or reaches no service endpoints

Open atlas

Separate private DNS, Application Gateway health, ingress controller routing, Kubernetes service selectors, endpoint slices and pod readiness before rolling back a deployment.

01

Evidence

Separate private DNS, Application Gateway health, ingress controller routing, Kubernetes service selectors, endpoint slices and pod readiness before rolling back a deployment.

02

First checks

  • Resolve the hostname from the caller network
  • Check Application Gateway backend health and host header
  • Verify ingress, service and endpoint slices
  • Correlate controller and application logs
03

Bounded action

Run the first checks in order: Resolve the hostname from the caller network | Check Application Gateway backend health and host header | Verify ingress, service and endpoint slices | Correlate controller and application logs. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] AKS private ingress returns 502 or reaches no service endpoints
Context: Separate private DNS, Application Gateway health, ingress controller routing, Kubernetes service selectors, endpoint slices and pod readiness before rolling back a deployment.
Evidence to confirm: Separate private DNS, Application Gateway health, ingress controller routing, Kubernetes service selectors, endpoint slices and pod readiness before rolling back a deployment.
Immediate checks: Resolve the hostname from the caller network | Check Application Gateway backend health and host header | Verify ingress, service and endpoint slices | Correlate controller and application logs
Proposed action: Run the first checks in order: Resolve the hostname from the caller network | Check Application Gateway backend health and host header | Verify ingress, service and endpoint slices | Correlate controller and application logs. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: AKS private ingress returns 502 or reaches no service endpoints
Initial hypothesis: Separate private DNS, Application Gateway health, ingress controller routing, Kubernetes service selectors, endpoint slices and pod readiness before rolling back a deployment.
Evidence used: Separate private DNS, Application Gateway health, ingress controller routing, Kubernetes service selectors, endpoint slices and pod readiness before rolling back a deployment.
Checks performed: Resolve the hostname from the caller network | Check Application Gateway backend health and host header | Verify ingress, service and endpoint slices | Correlate controller and application logs
Decision / action: Run the first checks in order: Resolve the hostname from the caller network | Check Application Gateway backend health and host header | Verify ingress, service and endpoint slices | Correlate controller and application logs. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Infrastructure

An AKS node pool upgrade is blocked while draining a pod protected by a PDB

Open atlas

Separate unavailable replicas, readiness failures, pod placement, surge capacity and an over-constrained PodDisruptionBudget before deleting the safeguard or forcing the upgrade.

01

Evidence

Separate unavailable replicas, readiness failures, pod placement, surge capacity and an over-constrained PodDisruptionBudget before deleting the safeguard or forcing the upgrade.

02

First checks

  • Capture the blocked node and eviction events
  • Inspect PDB disruptionsAllowed and selectors
  • Verify replacement capacity, replica readiness and topology
  • Resume only with positive disruption headroom
03

Bounded action

Run the first checks in order: Capture the blocked node and eviction events | Inspect PDB disruptionsAllowed and selectors | Verify replacement capacity, replica readiness and topology | Resume only with positive disruption headroom. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An AKS node pool upgrade is blocked while draining a pod protected by a PDB
Context: Separate unavailable replicas, readiness failures, pod placement, surge capacity and an over-constrained PodDisruptionBudget before deleting the safeguard or forcing the upgrade.
Evidence to confirm: Separate unavailable replicas, readiness failures, pod placement, surge capacity and an over-constrained PodDisruptionBudget before deleting the safeguard or forcing the upgrade.
Immediate checks: Capture the blocked node and eviction events | Inspect PDB disruptionsAllowed and selectors | Verify replacement capacity, replica readiness and topology | Resume only with positive disruption headroom
Proposed action: Run the first checks in order: Capture the blocked node and eviction events | Inspect PDB disruptionsAllowed and selectors | Verify replacement capacity, replica readiness and topology | Resume only with positive disruption headroom. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An AKS node pool upgrade is blocked while draining a pod protected by a PDB
Initial hypothesis: Separate unavailable replicas, readiness failures, pod placement, surge capacity and an over-constrained PodDisruptionBudget before deleting the safeguard or forcing the upgrade.
Evidence used: Separate unavailable replicas, readiness failures, pod placement, surge capacity and an over-constrained PodDisruptionBudget before deleting the safeguard or forcing the upgrade.
Checks performed: Capture the blocked node and eviction events | Inspect PDB disruptionsAllowed and selectors | Verify replacement capacity, replica readiness and topology | Resume only with positive disruption headroom
Decision / action: Run the first checks in order: Capture the blocked node and eviction events | Inspect PDB disruptionsAllowed and selectors | Verify replacement capacity, replica readiness and topology | Resume only with positive disruption headroom. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Cloud

Azure Functions private HTTP endpoint returns 403, 503 or no request logs

Open atlas

Separate private DNS, Private Endpoint reachability, access restrictions, Functions runtime state, private storage and Application Insights evidence before redeploying code or opening public access.

01

Evidence

Separate private DNS, Private Endpoint reachability, access restrictions, Functions runtime state, private storage and Application Insights evidence before redeploying code or opening public access.

02

First checks

  • Resolve the hostname from the caller network
  • Replay with a correlation ID
  • Check Function App access restrictions and Private Endpoint status
  • Correlate requests, traces and exceptions
03

Bounded action

Run the first checks in order: Resolve the hostname from the caller network | Replay with a correlation ID | Check Function App access restrictions and Private Endpoint status | Correlate requests, traces and exceptions. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] Azure Functions private HTTP endpoint returns 403, 503 or no request logs
Context: Separate private DNS, Private Endpoint reachability, access restrictions, Functions runtime state, private storage and Application Insights evidence before redeploying code or opening public access.
Evidence to confirm: Separate private DNS, Private Endpoint reachability, access restrictions, Functions runtime state, private storage and Application Insights evidence before redeploying code or opening public access.
Immediate checks: Resolve the hostname from the caller network | Replay with a correlation ID | Check Function App access restrictions and Private Endpoint status | Correlate requests, traces and exceptions
Proposed action: Run the first checks in order: Resolve the hostname from the caller network | Replay with a correlation ID | Check Function App access restrictions and Private Endpoint status | Correlate requests, traces and exceptions. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: Azure Functions private HTTP endpoint returns 403, 503 or no request logs
Initial hypothesis: Separate private DNS, Private Endpoint reachability, access restrictions, Functions runtime state, private storage and Application Insights evidence before redeploying code or opening public access.
Evidence used: Separate private DNS, Private Endpoint reachability, access restrictions, Functions runtime state, private storage and Application Insights evidence before redeploying code or opening public access.
Checks performed: Resolve the hostname from the caller network | Replay with a correlation ID | Check Function App access restrictions and Private Endpoint status | Correlate requests, traces and exceptions
Decision / action: Run the first checks in order: Resolve the hostname from the caller network | Replay with a correlation ID | Check Function App access restrictions and Private Endpoint status | Correlate requests, traces and exceptions. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Cloud

Azure WAF blocks a legitimate request

Open atlas

Start from blocked requests, rule ID and URI before deciding between exclusion, custom rule or application fix.

01

Evidence

Blocked URI, ruleId, match variable, client IP, hostname, request ID and exact time window.

02

First checks

  • List blocked URIs in KQL
  • Identify ruleId and match field
  • Validate false-positive scope
03

Bounded action

Create the smallest exclusion or custom rule that covers the false positive without disabling the rule globally.

04

Rollback

Remove the exclusion/custom rule and verify expected blocking returns for the same rule family.

copy packs Incident exports
Short handoff
[Incident] Azure WAF blocks a legitimate request
Context: Start from blocked requests, rule ID and URI before deciding between exclusion, custom rule or application fix.
Evidence to confirm: Blocked URI, ruleId, match variable, client IP, hostname, request ID and exact time window.
Immediate checks: List blocked URIs in KQL | Identify ruleId and match field | Validate false-positive scope
Proposed action: Create the smallest exclusion or custom rule that covers the false positive without disabling the rule globally.
Rollback: Remove the exclusion/custom rule and verify expected blocking returns for the same rule family.
Post-incident review
Handled symptom: Azure WAF blocks a legitimate request
Initial hypothesis: Start from blocked requests, rule ID and URI before deciding between exclusion, custom rule or application fix.
Evidence used: Blocked URI, ruleId, match variable, client IP, hostname, request ID and exact time window.
Checks performed: List blocked URIs in KQL | Identify ruleId and match field | Validate false-positive scope
Decision / action: Create the smallest exclusion or custom rule that covers the false positive without disabling the rule globally.
Rollback plan: Remove the exclusion/custom rule and verify expected blocking returns for the same rule family.
To improve: detection, runbook, guardrail, ownership and communication delay.
Automation

Terraform state lock is stuck

Open atlas

Prove that no apply is still running before using force-unlock, then restart with a clean plan.

01

Evidence

Lock ID, lock owner, CI run, backend target, pending plan and whether an apply is still active.

02

First checks

  • Identify lock owner
  • Check CI job status
  • Run plan after unlock
03

Bounded action

Unlock only after proving no apply is running, then start with a fresh plan before any apply.

04

Rollback

Return to previous commit or restore the last validated state version if drift was introduced.

copy packs Incident exports
Short handoff
[Incident] Terraform state lock is stuck
Context: Prove that no apply is still running before using force-unlock, then restart with a clean plan.
Evidence to confirm: Lock ID, lock owner, CI run, backend target, pending plan and whether an apply is still active.
Immediate checks: Identify lock owner | Check CI job status | Run plan after unlock
Proposed action: Unlock only after proving no apply is running, then start with a fresh plan before any apply.
Rollback: Return to previous commit or restore the last validated state version if drift was introduced.
Post-incident review
Handled symptom: Terraform state lock is stuck
Initial hypothesis: Prove that no apply is still running before using force-unlock, then restart with a clean plan.
Evidence used: Lock ID, lock owner, CI run, backend target, pending plan and whether an apply is still active.
Checks performed: Identify lock owner | Check CI job status | Run plan after unlock
Decision / action: Unlock only after proving no apply is running, then start with a fresh plan before any apply.
Rollback plan: Return to previous commit or restore the last validated state version if drift was introduced.
To improve: detection, runbook, guardrail, ownership and communication delay.
Infrastructure

Azure Monitor fires an alert storm after deployment

Open atlas

Separate real service impact, noisy dimensions, threshold drift, action group behavior and rollback before silencing notifications or changing rules.

01

Evidence

Separate real service impact, noisy dimensions, threshold drift, action group behavior and rollback before silencing notifications or changing rules.

02

First checks

  • Group fired alerts by rule and target
  • Compare first alert time with the deployment window
  • Check service logs before changing thresholds
  • Keep one user-symptom alert active
03

Bounded action

Run the first checks in order: Group fired alerts by rule and target | Compare first alert time with the deployment window | Check service logs before changing thresholds | Keep one user-symptom alert active. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] Azure Monitor fires an alert storm after deployment
Context: Separate real service impact, noisy dimensions, threshold drift, action group behavior and rollback before silencing notifications or changing rules.
Evidence to confirm: Separate real service impact, noisy dimensions, threshold drift, action group behavior and rollback before silencing notifications or changing rules.
Immediate checks: Group fired alerts by rule and target | Compare first alert time with the deployment window | Check service logs before changing thresholds | Keep one user-symptom alert active
Proposed action: Run the first checks in order: Group fired alerts by rule and target | Compare first alert time with the deployment window | Check service logs before changing thresholds | Keep one user-symptom alert active. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: Azure Monitor fires an alert storm after deployment
Initial hypothesis: Separate real service impact, noisy dimensions, threshold drift, action group behavior and rollback before silencing notifications or changing rules.
Evidence used: Separate real service impact, noisy dimensions, threshold drift, action group behavior and rollback before silencing notifications or changing rules.
Checks performed: Group fired alerts by rule and target | Compare first alert time with the deployment window | Check service logs before changing thresholds | Keep one user-symptom alert active
Decision / action: Run the first checks in order: Group fired alerts by rule and target | Compare first alert time with the deployment window | Check service logs before changing thresholds | Keep one user-symptom alert active. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Infrastructure

A secret rotation, federated identity or managed identity change breaks an application or pipeline consumer

Open atlas

Separate preparation, cutover, revocation, workload identity federation and managed identity diagnostics; validate the real execution identity, private path and authentication errors before deleting the old value or broadening access.

01

Evidence

Separate preparation, cutover, revocation, workload identity federation and managed identity diagnostics; validate the real execution identity, private path and authentication errors before deleting the old value or broadening access.

02

First checks

  • List real consumers
  • Verify the runtime identity and vault read access
  • Check OIDC claims or private DNS depending on the path
  • Watch 401/403/500, sign-in failures or Key Vault denials
03

Bounded action

Run the first checks in order: List real consumers | Verify the runtime identity and vault read access | Check OIDC claims or private DNS depending on the path | Watch 401/403/500, sign-in failures or Key Vault denials. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] A secret rotation, federated identity or managed identity change breaks an application or pipeline consumer
Context: Separate preparation, cutover, revocation, workload identity federation and managed identity diagnostics; validate the real execution identity, private path and authentication errors before deleting the old value or broadening access.
Evidence to confirm: Separate preparation, cutover, revocation, workload identity federation and managed identity diagnostics; validate the real execution identity, private path and authentication errors before deleting the old value or broadening access.
Immediate checks: List real consumers | Verify the runtime identity and vault read access | Check OIDC claims or private DNS depending on the path | Watch 401/403/500, sign-in failures or Key Vault denials
Proposed action: Run the first checks in order: List real consumers | Verify the runtime identity and vault read access | Check OIDC claims or private DNS depending on the path | Watch 401/403/500, sign-in failures or Key Vault denials. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: A secret rotation, federated identity or managed identity change breaks an application or pipeline consumer
Initial hypothesis: Separate preparation, cutover, revocation, workload identity federation and managed identity diagnostics; validate the real execution identity, private path and authentication errors before deleting the old value or broadening access.
Evidence used: Separate preparation, cutover, revocation, workload identity federation and managed identity diagnostics; validate the real execution identity, private path and authentication errors before deleting the old value or broadening access.
Checks performed: List real consumers | Verify the runtime identity and vault read access | Check OIDC claims or private DNS depending on the path | Watch 401/403/500, sign-in failures or Key Vault denials
Decision / action: Run the first checks in order: List real consumers | Verify the runtime identity and vault read access | Check OIDC claims or private DNS depending on the path | Watch 401/403/500, sign-in failures or Key Vault denials. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Automation

An Azure Functions poison queue keeps growing

Open atlas

Separate contract failures, transient dependencies, identity, resource pressure and partial side effects before replaying Storage Queue messages.

01

Evidence

Separate contract failures, transient dependencies, identity, resource pressure and partial side effects before replaying Storage Queue messages.

02

First checks

  • Peek a redacted sample without consuming messages
  • Correlate message IDs with invocation and dependency logs
  • Verify downstream state and idempotency before replay
03

Bounded action

Run the first checks in order: Peek a redacted sample without consuming messages | Correlate message IDs with invocation and dependency logs | Verify downstream state and idempotency before replay. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An Azure Functions poison queue keeps growing
Context: Separate contract failures, transient dependencies, identity, resource pressure and partial side effects before replaying Storage Queue messages.
Evidence to confirm: Separate contract failures, transient dependencies, identity, resource pressure and partial side effects before replaying Storage Queue messages.
Immediate checks: Peek a redacted sample without consuming messages | Correlate message IDs with invocation and dependency logs | Verify downstream state and idempotency before replay
Proposed action: Run the first checks in order: Peek a redacted sample without consuming messages | Correlate message IDs with invocation and dependency logs | Verify downstream state and idempotency before replay. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An Azure Functions poison queue keeps growing
Initial hypothesis: Separate contract failures, transient dependencies, identity, resource pressure and partial side effects before replaying Storage Queue messages.
Evidence used: Separate contract failures, transient dependencies, identity, resource pressure and partial side effects before replaying Storage Queue messages.
Checks performed: Peek a redacted sample without consuming messages | Correlate message IDs with invocation and dependency logs | Verify downstream state and idempotency before replay
Decision / action: Run the first checks in order: Peek a redacted sample without consuming messages | Correlate message IDs with invocation and dependency logs | Verify downstream state and idempotency before replay. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Cloud

An Azure Service Bus dead-letter queue keeps growing

Open atlas

Classify dead-letter reasons, correlate consumer and downstream evidence, verify idempotency, then use a manifest, canary and stop conditions before replay.

01

Evidence

Classify dead-letter reasons, correlate consumer and downstream evidence, verify idempotency, then use a manifest, canary and stop conditions before replay.

02

First checks

  • Record the exact queue or subscription and DLQ growth window
  • Peek a redacted sample grouped by dead-letter reason
  • Check consumer errors, lock loss and downstream state
  • Prove idempotency before a one-message canary
03

Bounded action

Run the first checks in order: Record the exact queue or subscription and DLQ growth window | Peek a redacted sample grouped by dead-letter reason | Check consumer errors, lock loss and downstream state | Prove idempotency before a one-message canary. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An Azure Service Bus dead-letter queue keeps growing
Context: Classify dead-letter reasons, correlate consumer and downstream evidence, verify idempotency, then use a manifest, canary and stop conditions before replay.
Evidence to confirm: Classify dead-letter reasons, correlate consumer and downstream evidence, verify idempotency, then use a manifest, canary and stop conditions before replay.
Immediate checks: Record the exact queue or subscription and DLQ growth window | Peek a redacted sample grouped by dead-letter reason | Check consumer errors, lock loss and downstream state | Prove idempotency before a one-message canary
Proposed action: Run the first checks in order: Record the exact queue or subscription and DLQ growth window | Peek a redacted sample grouped by dead-letter reason | Check consumer errors, lock loss and downstream state | Prove idempotency before a one-message canary. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An Azure Service Bus dead-letter queue keeps growing
Initial hypothesis: Classify dead-letter reasons, correlate consumer and downstream evidence, verify idempotency, then use a manifest, canary and stop conditions before replay.
Evidence used: Classify dead-letter reasons, correlate consumer and downstream evidence, verify idempotency, then use a manifest, canary and stop conditions before replay.
Checks performed: Record the exact queue or subscription and DLQ growth window | Peek a redacted sample grouped by dead-letter reason | Check consumer errors, lock loss and downstream state | Prove idempotency before a one-message canary
Decision / action: Run the first checks in order: Record the exact queue or subscription and DLQ growth window | Peek a redacted sample grouped by dead-letter reason | Check consumer errors, lock loss and downstream state | Prove idempotency before a one-message canary. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Automation

An Azure Policy remediation changes unexpected production resources

Open atlas

Pin the definition and assignment, materialize the eligible target set, verify the managed identity, then use a canary and bounded batches before expanding or compensating.

01

Evidence

Pin the definition and assignment, materialize the eligible target set, verify the managed identity, then use a canary and bounded batches before expanding or compensating.

02

First checks

  • Identify the exact definition reference and assignment parameters
  • Export non-compliant target resource IDs
  • Review the assignment identity and effective RBAC scope
  • Correlate remediation deployments with Activity Log
03

Bounded action

Run the first checks in order: Identify the exact definition reference and assignment parameters | Export non-compliant target resource IDs | Review the assignment identity and effective RBAC scope | Correlate remediation deployments with Activity Log. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An Azure Policy remediation changes unexpected production resources
Context: Pin the definition and assignment, materialize the eligible target set, verify the managed identity, then use a canary and bounded batches before expanding or compensating.
Evidence to confirm: Pin the definition and assignment, materialize the eligible target set, verify the managed identity, then use a canary and bounded batches before expanding or compensating.
Immediate checks: Identify the exact definition reference and assignment parameters | Export non-compliant target resource IDs | Review the assignment identity and effective RBAC scope | Correlate remediation deployments with Activity Log
Proposed action: Run the first checks in order: Identify the exact definition reference and assignment parameters | Export non-compliant target resource IDs | Review the assignment identity and effective RBAC scope | Correlate remediation deployments with Activity Log. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An Azure Policy remediation changes unexpected production resources
Initial hypothesis: Pin the definition and assignment, materialize the eligible target set, verify the managed identity, then use a canary and bounded batches before expanding or compensating.
Evidence used: Pin the definition and assignment, materialize the eligible target set, verify the managed identity, then use a canary and bounded batches before expanding or compensating.
Checks performed: Identify the exact definition reference and assignment parameters | Export non-compliant target resource IDs | Review the assignment identity and effective RBAC scope | Correlate remediation deployments with Activity Log
Decision / action: Run the first checks in order: Identify the exact definition reference and assignment parameters | Export non-compliant target resource IDs | Review the assignment identity and effective RBAC scope | Correlate remediation deployments with Activity Log. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Automation

An automation entry point behaves like a remote console

Open atlas

Bound inputs, templates and repository structure before exposing operations to more users.

01

Evidence

Inputs accepted by the template, permissions, inventory scope, repository branch and audit trail.

02

First checks

  • List accepted inputs
  • Remove arbitrary command fields
  • Review job template permissions
03

Bounded action

Replace arbitrary inputs with bounded choices and isolate job templates by operational intent.

04

Rollback

Disable the exposed template or revert to the previous approved template version.

copy packs Incident exports
Short handoff
[Incident] An automation entry point behaves like a remote console
Context: Bound inputs, templates and repository structure before exposing operations to more users.
Evidence to confirm: Inputs accepted by the template, permissions, inventory scope, repository branch and audit trail.
Immediate checks: List accepted inputs | Remove arbitrary command fields | Review job template permissions
Proposed action: Replace arbitrary inputs with bounded choices and isolate job templates by operational intent.
Rollback: Disable the exposed template or revert to the previous approved template version.
Post-incident review
Handled symptom: An automation entry point behaves like a remote console
Initial hypothesis: Bound inputs, templates and repository structure before exposing operations to more users.
Evidence used: Inputs accepted by the template, permissions, inventory scope, repository branch and audit trail.
Checks performed: List accepted inputs | Remove arbitrary command fields | Review job template permissions
Decision / action: Replace arbitrary inputs with bounded choices and isolate job templates by operational intent.
Rollback plan: Disable the exposed template or revert to the previous approved template version.
To improve: detection, runbook, guardrail, ownership and communication delay.
AI

A private AI agent can act but nobody can explain the action

Open atlas

Tie sources, identities, tool calls, logs and human validation before increasing autonomy.

01

Evidence

Source documents, identity, tool calls, prompt context, logs and human approval point.

02

First checks

  • List approved sources
  • Trace tool calls
  • Define human approval points
03

Bounded action

Reduce tool scope, require approval on sensitive actions and trace every tool call to a source.

04

Rollback

Disable the tool integration or force human validation until the action path is explainable.

copy packs Incident exports
Short handoff
[Incident] A private AI agent can act but nobody can explain the action
Context: Tie sources, identities, tool calls, logs and human validation before increasing autonomy.
Evidence to confirm: Source documents, identity, tool calls, prompt context, logs and human approval point.
Immediate checks: List approved sources | Trace tool calls | Define human approval points
Proposed action: Reduce tool scope, require approval on sensitive actions and trace every tool call to a source.
Rollback: Disable the tool integration or force human validation until the action path is explainable.
Post-incident review
Handled symptom: A private AI agent can act but nobody can explain the action
Initial hypothesis: Tie sources, identities, tool calls, logs and human validation before increasing autonomy.
Evidence used: Source documents, identity, tool calls, prompt context, logs and human approval point.
Checks performed: List approved sources | Trace tool calls | Define human approval points
Decision / action: Reduce tool scope, require approval on sensitive actions and trace every tool call to a source.
Rollback plan: Disable the tool integration or force human validation until the action path is explainable.
To improve: detection, runbook, guardrail, ownership and communication delay.
Cloud

An Azure Event Hubs consumer keeps falling behind

Open atlas

Separate namespace throttling, hot partitions, consumer processing, checkpoint progress and downstream saturation before adding capacity or replaying events.

01

Evidence

Separate namespace throttling, hot partitions, consumer processing, checkpoint progress and downstream saturation before adding capacity or replaying events.

02

First checks

  • Compare ingress, egress and throttled requests
  • Measure progress and checkpoint age per partition
  • Check consumer ownership, errors and downstream latency
  • Prove sequence bounds and idempotency before replay
03

Bounded action

Run the first checks in order: Compare ingress, egress and throttled requests | Measure progress and checkpoint age per partition | Check consumer ownership, errors and downstream latency | Prove sequence bounds and idempotency before replay. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An Azure Event Hubs consumer keeps falling behind
Context: Separate namespace throttling, hot partitions, consumer processing, checkpoint progress and downstream saturation before adding capacity or replaying events.
Evidence to confirm: Separate namespace throttling, hot partitions, consumer processing, checkpoint progress and downstream saturation before adding capacity or replaying events.
Immediate checks: Compare ingress, egress and throttled requests | Measure progress and checkpoint age per partition | Check consumer ownership, errors and downstream latency | Prove sequence bounds and idempotency before replay
Proposed action: Run the first checks in order: Compare ingress, egress and throttled requests | Measure progress and checkpoint age per partition | Check consumer ownership, errors and downstream latency | Prove sequence bounds and idempotency before replay. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An Azure Event Hubs consumer keeps falling behind
Initial hypothesis: Separate namespace throttling, hot partitions, consumer processing, checkpoint progress and downstream saturation before adding capacity or replaying events.
Evidence used: Separate namespace throttling, hot partitions, consumer processing, checkpoint progress and downstream saturation before adding capacity or replaying events.
Checks performed: Compare ingress, egress and throttled requests | Measure progress and checkpoint age per partition | Check consumer ownership, errors and downstream latency | Prove sequence bounds and idempotency before replay
Decision / action: Run the first checks in order: Compare ingress, egress and throttled requests | Measure progress and checkpoint age per partition | Check consumer ownership, errors and downstream latency | Prove sequence bounds and idempotency before replay. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Automation

An Azure Automation Hybrid Runbook Worker stops picking up jobs

Open atlas

Separate dispatch, heartbeat, extension health, outbound HTTPS, capacity, runtime and identity before retrying a production job.

01

Evidence

Separate dispatch, heartbeat, extension health, outbound HTTPS, capacity, runtime and identity before retrying a production job.

02

First checks

  • Preserve the job ID, streams and UTC window
  • Compare HybridWorkerPing and extension state across the group
  • Read local worker logs before restarting services
  • Complete a no-side-effect canary before retry
03

Bounded action

Run the first checks in order: Preserve the job ID, streams and UTC window | Compare HybridWorkerPing and extension state across the group | Read local worker logs before restarting services | Complete a no-side-effect canary before retry. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An Azure Automation Hybrid Runbook Worker stops picking up jobs
Context: Separate dispatch, heartbeat, extension health, outbound HTTPS, capacity, runtime and identity before retrying a production job.
Evidence to confirm: Separate dispatch, heartbeat, extension health, outbound HTTPS, capacity, runtime and identity before retrying a production job.
Immediate checks: Preserve the job ID, streams and UTC window | Compare HybridWorkerPing and extension state across the group | Read local worker logs before restarting services | Complete a no-side-effect canary before retry
Proposed action: Run the first checks in order: Preserve the job ID, streams and UTC window | Compare HybridWorkerPing and extension state across the group | Read local worker logs before restarting services | Complete a no-side-effect canary before retry. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An Azure Automation Hybrid Runbook Worker stops picking up jobs
Initial hypothesis: Separate dispatch, heartbeat, extension health, outbound HTTPS, capacity, runtime and identity before retrying a production job.
Evidence used: Separate dispatch, heartbeat, extension health, outbound HTTPS, capacity, runtime and identity before retrying a production job.
Checks performed: Preserve the job ID, streams and UTC window | Compare HybridWorkerPing and extension state across the group | Read local worker logs before restarting services | Complete a no-side-effect canary before retry
Decision / action: Run the first checks in order: Preserve the job ID, streams and UTC window | Compare HybridWorkerPing and extension state across the group | Read local worker logs before restarting services | Complete a no-side-effect canary before retry. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Networking

Azure Bastion cannot open an SSH or RDP session

Open atlas

Separate the client HTTPS path, Bastion health, AzureBastionSubnet rules, routing, target NSGs, guest listener and identity before exposing administration ports.

01

Evidence

Separate the client HTTPS path, Bastion health, AzureBastionSubnet rules, routing, target NSGs, guest listener and identity before exposing administration ports.

02

First checks

  • Freeze one failed attempt with UTC time and target private IP
  • Check Bastion provisioning state and the complete AzureBastionSubnet rule set
  • Inspect target NIC effective NSGs and routes
  • Verify the guest listener and authentication separately
03

Bounded action

Run the first checks in order: Freeze one failed attempt with UTC time and target private IP | Check Bastion provisioning state and the complete AzureBastionSubnet rule set | Inspect target NIC effective NSGs and routes | Verify the guest listener and authentication separately. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] Azure Bastion cannot open an SSH or RDP session
Context: Separate the client HTTPS path, Bastion health, AzureBastionSubnet rules, routing, target NSGs, guest listener and identity before exposing administration ports.
Evidence to confirm: Separate the client HTTPS path, Bastion health, AzureBastionSubnet rules, routing, target NSGs, guest listener and identity before exposing administration ports.
Immediate checks: Freeze one failed attempt with UTC time and target private IP | Check Bastion provisioning state and the complete AzureBastionSubnet rule set | Inspect target NIC effective NSGs and routes | Verify the guest listener and authentication separately
Proposed action: Run the first checks in order: Freeze one failed attempt with UTC time and target private IP | Check Bastion provisioning state and the complete AzureBastionSubnet rule set | Inspect target NIC effective NSGs and routes | Verify the guest listener and authentication separately. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: Azure Bastion cannot open an SSH or RDP session
Initial hypothesis: Separate the client HTTPS path, Bastion health, AzureBastionSubnet rules, routing, target NSGs, guest listener and identity before exposing administration ports.
Evidence used: Separate the client HTTPS path, Bastion health, AzureBastionSubnet rules, routing, target NSGs, guest listener and identity before exposing administration ports.
Checks performed: Freeze one failed attempt with UTC time and target private IP | Check Bastion provisioning state and the complete AzureBastionSubnet rule set | Inspect target NIC effective NSGs and routes | Verify the guest listener and authentication separately
Decision / action: Run the first checks in order: Freeze one failed attempt with UTC time and target private IP | Check Bastion provisioning state and the complete AzureBastionSubnet rule set | Inspect target NIC effective NSGs and routes | Verify the guest listener and authentication separately. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Cloud

An application exhausts its Azure SQL connection pool

Open atlas

Separate local connection retention, SQL sessions and waits, network failures, managed identity token acquisition and autoscaling multiplication before raising the pool limit or database tier.

01

Evidence

Separate local connection retention, SQL sessions and waits, network failures, managed identity token acquisition and autoscaling multiplication before raising the pool limit or database tier.

02

First checks

  • Tie one timeout to a revision and role instance
  • Calculate the global pool capacity envelope
  • Compare SQL sessions with active requests and waits
  • Canary the correction with stop and rollback conditions
03

Bounded action

Run the first checks in order: Tie one timeout to a revision and role instance | Calculate the global pool capacity envelope | Compare SQL sessions with active requests and waits | Canary the correction with stop and rollback conditions. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An application exhausts its Azure SQL connection pool
Context: Separate local connection retention, SQL sessions and waits, network failures, managed identity token acquisition and autoscaling multiplication before raising the pool limit or database tier.
Evidence to confirm: Separate local connection retention, SQL sessions and waits, network failures, managed identity token acquisition and autoscaling multiplication before raising the pool limit or database tier.
Immediate checks: Tie one timeout to a revision and role instance | Calculate the global pool capacity envelope | Compare SQL sessions with active requests and waits | Canary the correction with stop and rollback conditions
Proposed action: Run the first checks in order: Tie one timeout to a revision and role instance | Calculate the global pool capacity envelope | Compare SQL sessions with active requests and waits | Canary the correction with stop and rollback conditions. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An application exhausts its Azure SQL connection pool
Initial hypothesis: Separate local connection retention, SQL sessions and waits, network failures, managed identity token acquisition and autoscaling multiplication before raising the pool limit or database tier.
Evidence used: Separate local connection retention, SQL sessions and waits, network failures, managed identity token acquisition and autoscaling multiplication before raising the pool limit or database tier.
Checks performed: Tie one timeout to a revision and role instance | Calculate the global pool capacity envelope | Compare SQL sessions with active requests and waits | Canary the correction with stop and rollback conditions
Decision / action: Run the first checks in order: Tie one timeout to a revision and role instance | Calculate the global pool capacity envelope | Compare SQL sessions with active requests and waits | Canary the correction with stop and rollback conditions. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Cloud

An ACR image fails signature verification before AKS promotion

Open atlas

Separate mutable tags, missing signatures, untrusted publisher identity, trust-store rotation and ACR access before bypassing admission or deploying an unverified image.

01

Evidence

Separate mutable tags, missing signatures, untrusted publisher identity, trust-store rotation and ACR access before bypassing admission or deploying an unverified image.

02

First checks

  • Resolve the candidate tag to one immutable digest
  • Inventory signatures with Notation
  • Verify repository scope, trust store and publisher identity
  • Compare the verified digest with the AKS runtime imageID
03

Bounded action

Run the first checks in order: Resolve the candidate tag to one immutable digest | Inventory signatures with Notation | Verify repository scope, trust store and publisher identity | Compare the verified digest with the AKS runtime imageID. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An ACR image fails signature verification before AKS promotion
Context: Separate mutable tags, missing signatures, untrusted publisher identity, trust-store rotation and ACR access before bypassing admission or deploying an unverified image.
Evidence to confirm: Separate mutable tags, missing signatures, untrusted publisher identity, trust-store rotation and ACR access before bypassing admission or deploying an unverified image.
Immediate checks: Resolve the candidate tag to one immutable digest | Inventory signatures with Notation | Verify repository scope, trust store and publisher identity | Compare the verified digest with the AKS runtime imageID
Proposed action: Run the first checks in order: Resolve the candidate tag to one immutable digest | Inventory signatures with Notation | Verify repository scope, trust store and publisher identity | Compare the verified digest with the AKS runtime imageID. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An ACR image fails signature verification before AKS promotion
Initial hypothesis: Separate mutable tags, missing signatures, untrusted publisher identity, trust-store rotation and ACR access before bypassing admission or deploying an unverified image.
Evidence used: Separate mutable tags, missing signatures, untrusted publisher identity, trust-store rotation and ACR access before bypassing admission or deploying an unverified image.
Checks performed: Resolve the candidate tag to one immutable digest | Inventory signatures with Notation | Verify repository scope, trust store and publisher identity | Compare the verified digest with the AKS runtime imageID
Decision / action: Run the first checks in order: Resolve the candidate tag to one immutable digest | Inventory signatures with Notation | Verify repository scope, trust store and publisher identity | Compare the verified digest with the AKS runtime imageID. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Infrastructure

An AKS Deployment rollout stalls before the new revision becomes available

Open atlas

Locate the first blocked object across Deployment, ReplicaSet, admission, scheduling, image startup and readiness before deleting pods, extending timeouts or forcing rollback.

01

Evidence

Locate the first blocked object across Deployment, ReplicaSet, admission, scheduling, image startup and readiness before deleting pods, extending timeouts or forcing rollback.

02

First checks

  • Capture Deployment conditions and rollout history
  • Compare desired, created and ready pods on the candidate ReplicaSet
  • Read pod and namespace events before changing capacity
  • Prove external configuration compatibility before rollback
03

Bounded action

Run the first checks in order: Capture Deployment conditions and rollout history | Compare desired, created and ready pods on the candidate ReplicaSet | Read pod and namespace events before changing capacity | Prove external configuration compatibility before rollback. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An AKS Deployment rollout stalls before the new revision becomes available
Context: Locate the first blocked object across Deployment, ReplicaSet, admission, scheduling, image startup and readiness before deleting pods, extending timeouts or forcing rollback.
Evidence to confirm: Locate the first blocked object across Deployment, ReplicaSet, admission, scheduling, image startup and readiness before deleting pods, extending timeouts or forcing rollback.
Immediate checks: Capture Deployment conditions and rollout history | Compare desired, created and ready pods on the candidate ReplicaSet | Read pod and namespace events before changing capacity | Prove external configuration compatibility before rollback
Proposed action: Run the first checks in order: Capture Deployment conditions and rollout history | Compare desired, created and ready pods on the candidate ReplicaSet | Read pod and namespace events before changing capacity | Prove external configuration compatibility before rollback. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An AKS Deployment rollout stalls before the new revision becomes available
Initial hypothesis: Locate the first blocked object across Deployment, ReplicaSet, admission, scheduling, image startup and readiness before deleting pods, extending timeouts or forcing rollback.
Evidence used: Locate the first blocked object across Deployment, ReplicaSet, admission, scheduling, image startup and readiness before deleting pods, extending timeouts or forcing rollback.
Checks performed: Capture Deployment conditions and rollout history | Compare desired, created and ready pods on the candidate ReplicaSet | Read pod and namespace events before changing capacity | Prove external configuration compatibility before rollback
Decision / action: Run the first checks in order: Capture Deployment conditions and rollout history | Compare desired, created and ready pods on the candidate ReplicaSet | Read pod and namespace events before changing capacity | Prove external configuration compatibility before rollback. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Automation

Azure Logic Apps runs accumulate 429 responses and retries

Open atlas

Separate Logic Apps resource limits, connector throttling and destination saturation before adding retries or raising concurrency.

01

Evidence

Separate Logic Apps resource limits, connector throttling and destination saturation before adding retries or raising concurrency.

02

First checks

  • Freeze one workflow, action, connection and UTC window
  • Correlate run and destination request IDs
  • Measure retry and concurrency amplification
  • Run a bounded canary before staged recovery
03

Bounded action

Run the first checks in order: Freeze one workflow, action, connection and UTC window | Correlate run and destination request IDs | Measure retry and concurrency amplification | Run a bounded canary before staged recovery. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] Azure Logic Apps runs accumulate 429 responses and retries
Context: Separate Logic Apps resource limits, connector throttling and destination saturation before adding retries or raising concurrency.
Evidence to confirm: Separate Logic Apps resource limits, connector throttling and destination saturation before adding retries or raising concurrency.
Immediate checks: Freeze one workflow, action, connection and UTC window | Correlate run and destination request IDs | Measure retry and concurrency amplification | Run a bounded canary before staged recovery
Proposed action: Run the first checks in order: Freeze one workflow, action, connection and UTC window | Correlate run and destination request IDs | Measure retry and concurrency amplification | Run a bounded canary before staged recovery. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: Azure Logic Apps runs accumulate 429 responses and retries
Initial hypothesis: Separate Logic Apps resource limits, connector throttling and destination saturation before adding retries or raising concurrency.
Evidence used: Separate Logic Apps resource limits, connector throttling and destination saturation before adding retries or raising concurrency.
Checks performed: Freeze one workflow, action, connection and UTC window | Correlate run and destination request IDs | Measure retry and concurrency amplification | Run a bounded canary before staged recovery
Decision / action: Run the first checks in order: Freeze one workflow, action, connection and UTC window | Correlate run and destination request IDs | Measure retry and concurrency amplification | Run a bounded canary before staged recovery. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Automation

Feature flag evaluations differ across application instances

Open atlas

Separate the published definition, per-instance ETag, refresh path, targeting context and evaluation telemetry before restarting the fleet or rolling back the flag.

01

Evidence

Separate the published definition, per-instance ETag, refresh path, targeting context and evaluation telemetry before restarting the fleet or rolling back the flag.

02

First checks

  • Freeze the flag, label, change window and previous known state
  • Compare the published ETag with healthy and affected instances
  • Replay one stable targeting context on a canary
  • Correlate FeatureEvaluation events with refresh traces and business metrics
03

Bounded action

Run the first checks in order: Freeze the flag, label, change window and previous known state | Compare the published ETag with healthy and affected instances | Replay one stable targeting context on a canary | Correlate FeatureEvaluation events with refresh traces and business metrics. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] Feature flag evaluations differ across application instances
Context: Separate the published definition, per-instance ETag, refresh path, targeting context and evaluation telemetry before restarting the fleet or rolling back the flag.
Evidence to confirm: Separate the published definition, per-instance ETag, refresh path, targeting context and evaluation telemetry before restarting the fleet or rolling back the flag.
Immediate checks: Freeze the flag, label, change window and previous known state | Compare the published ETag with healthy and affected instances | Replay one stable targeting context on a canary | Correlate FeatureEvaluation events with refresh traces and business metrics
Proposed action: Run the first checks in order: Freeze the flag, label, change window and previous known state | Compare the published ETag with healthy and affected instances | Replay one stable targeting context on a canary | Correlate FeatureEvaluation events with refresh traces and business metrics. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: Feature flag evaluations differ across application instances
Initial hypothesis: Separate the published definition, per-instance ETag, refresh path, targeting context and evaluation telemetry before restarting the fleet or rolling back the flag.
Evidence used: Separate the published definition, per-instance ETag, refresh path, targeting context and evaluation telemetry before restarting the fleet or rolling back the flag.
Checks performed: Freeze the flag, label, change window and previous known state | Compare the published ETag with healthy and affected instances | Replay one stable targeting context on a canary | Correlate FeatureEvaluation events with refresh traces and business metrics
Decision / action: Run the first checks in order: Freeze the flag, label, change window and previous known state | Compare the published ETag with healthy and affected instances | Replay one stable targeting context on a canary | Correlate FeatureEvaluation events with refresh traces and business metrics. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
AI

An AI agent loses constraints as the conversation grows

Open atlas

Separate system instructions, history, retrieval, tool schemas and tool outputs before raising context limits or changing models.

01

Evidence

Separate system instructions, history, retrieval, tool schemas and tool outputs before raising context limits or changing models.

02

First checks

  • Freeze one healthy and one degraded trace for the same journey
  • Measure context growth by segment and turn
  • Verify approvals and open decisions survive compaction
  • Canary the bounded policy before promotion
03

Bounded action

Run the first checks in order: Freeze one healthy and one degraded trace for the same journey | Measure context growth by segment and turn | Verify approvals and open decisions survive compaction | Canary the bounded policy before promotion. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An AI agent loses constraints as the conversation grows
Context: Separate system instructions, history, retrieval, tool schemas and tool outputs before raising context limits or changing models.
Evidence to confirm: Separate system instructions, history, retrieval, tool schemas and tool outputs before raising context limits or changing models.
Immediate checks: Freeze one healthy and one degraded trace for the same journey | Measure context growth by segment and turn | Verify approvals and open decisions survive compaction | Canary the bounded policy before promotion
Proposed action: Run the first checks in order: Freeze one healthy and one degraded trace for the same journey | Measure context growth by segment and turn | Verify approvals and open decisions survive compaction | Canary the bounded policy before promotion. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An AI agent loses constraints as the conversation grows
Initial hypothesis: Separate system instructions, history, retrieval, tool schemas and tool outputs before raising context limits or changing models.
Evidence used: Separate system instructions, history, retrieval, tool schemas and tool outputs before raising context limits or changing models.
Checks performed: Freeze one healthy and one degraded trace for the same journey | Measure context growth by segment and turn | Verify approvals and open decisions survive compaction | Canary the bounded policy before promotion
Decision / action: Run the first checks in order: Freeze one healthy and one degraded trace for the same journey | Measure context growth by segment and turn | Verify approvals and open decisions survive compaction | Canary the bounded policy before promotion. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Cloud

Azure Managed Redis connections surge while application requests time out

Open atlas

Separate legitimate server pressure, client reconnect amplification, runtime DNS/TLS and release drift before scaling or triggering a failover.

01

Evidence

Separate legitimate server pressure, client reconnect amplification, runtime DNS/TLS and release drift before scaling or triggering a failover.

02

First checks

  • Freeze the release, replica count and incident window
  • Compare connected clients, server load and useful operations
  • Measure the connection envelope per role instance
  • Canary the client fix with explicit rollback
03

Bounded action

Run the first checks in order: Freeze the release, replica count and incident window | Compare connected clients, server load and useful operations | Measure the connection envelope per role instance | Canary the client fix with explicit rollback. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] Azure Managed Redis connections surge while application requests time out
Context: Separate legitimate server pressure, client reconnect amplification, runtime DNS/TLS and release drift before scaling or triggering a failover.
Evidence to confirm: Separate legitimate server pressure, client reconnect amplification, runtime DNS/TLS and release drift before scaling or triggering a failover.
Immediate checks: Freeze the release, replica count and incident window | Compare connected clients, server load and useful operations | Measure the connection envelope per role instance | Canary the client fix with explicit rollback
Proposed action: Run the first checks in order: Freeze the release, replica count and incident window | Compare connected clients, server load and useful operations | Measure the connection envelope per role instance | Canary the client fix with explicit rollback. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: Azure Managed Redis connections surge while application requests time out
Initial hypothesis: Separate legitimate server pressure, client reconnect amplification, runtime DNS/TLS and release drift before scaling or triggering a failover.
Evidence used: Separate legitimate server pressure, client reconnect amplification, runtime DNS/TLS and release drift before scaling or triggering a failover.
Checks performed: Freeze the release, replica count and incident window | Compare connected clients, server load and useful operations | Measure the connection envelope per role instance | Canary the client fix with explicit rollback
Decision / action: Run the first checks in order: Freeze the release, replica count and incident window | Compare connected clients, server load and useful operations | Measure the connection envelope per role instance | Canary the client fix with explicit rollback. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Infrastructure

An Azure Monitor alert fires but no action group notifies on-call

Open atlas

Separate signal evaluation, fired-alert processing, scope, filters, schedule and action group delivery before changing the alert rule.

01

Evidence

Separate signal evaluation, fired-alert processing, scope, filters, schedule and action group delivery before changing the alert rule.

02

First checks

  • Freeze one fired alert ID and its expected action groups
  • List enabled alert processing rules and match their scopes
  • Replay the schedule in its configured time zone
  • Generate a new bounded alert after propagation
03

Bounded action

Run the first checks in order: Freeze one fired alert ID and its expected action groups | List enabled alert processing rules and match their scopes | Replay the schedule in its configured time zone | Generate a new bounded alert after propagation. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An Azure Monitor alert fires but no action group notifies on-call
Context: Separate signal evaluation, fired-alert processing, scope, filters, schedule and action group delivery before changing the alert rule.
Evidence to confirm: Separate signal evaluation, fired-alert processing, scope, filters, schedule and action group delivery before changing the alert rule.
Immediate checks: Freeze one fired alert ID and its expected action groups | List enabled alert processing rules and match their scopes | Replay the schedule in its configured time zone | Generate a new bounded alert after propagation
Proposed action: Run the first checks in order: Freeze one fired alert ID and its expected action groups | List enabled alert processing rules and match their scopes | Replay the schedule in its configured time zone | Generate a new bounded alert after propagation. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An Azure Monitor alert fires but no action group notifies on-call
Initial hypothesis: Separate signal evaluation, fired-alert processing, scope, filters, schedule and action group delivery before changing the alert rule.
Evidence used: Separate signal evaluation, fired-alert processing, scope, filters, schedule and action group delivery before changing the alert rule.
Checks performed: Freeze one fired alert ID and its expected action groups | List enabled alert processing rules and match their scopes | Replay the schedule in its configured time zone | Generate a new bounded alert after propagation
Decision / action: Run the first checks in order: Freeze one fired alert ID and its expected action groups | List enabled alert processing rules and match their scopes | Replay the schedule in its configured time zone | Generate a new bounded alert after propagation. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Cloud

An Azure Database for PostgreSQL read replica falls behind before promotion

Open atlas

Separate time lag, WAL byte gap, primary log retention, replica replay capacity and target connection readiness before planned switchover or forced promotion.

01

Evidence

Separate time lag, WAL byte gap, primary log retention, replica replay capacity and target connection readiness before planned switchover or forced promotion.

02

First checks

  • Freeze the promotion mode and accepted RPO
  • Compare lag seconds, lag bytes and a business marker
  • Check primary WAL storage and replica saturation
  • Validate virtual endpoints, identity and write path before promotion
03

Bounded action

Run the first checks in order: Freeze the promotion mode and accepted RPO | Compare lag seconds, lag bytes and a business marker | Check primary WAL storage and replica saturation | Validate virtual endpoints, identity and write path before promotion. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An Azure Database for PostgreSQL read replica falls behind before promotion
Context: Separate time lag, WAL byte gap, primary log retention, replica replay capacity and target connection readiness before planned switchover or forced promotion.
Evidence to confirm: Separate time lag, WAL byte gap, primary log retention, replica replay capacity and target connection readiness before planned switchover or forced promotion.
Immediate checks: Freeze the promotion mode and accepted RPO | Compare lag seconds, lag bytes and a business marker | Check primary WAL storage and replica saturation | Validate virtual endpoints, identity and write path before promotion
Proposed action: Run the first checks in order: Freeze the promotion mode and accepted RPO | Compare lag seconds, lag bytes and a business marker | Check primary WAL storage and replica saturation | Validate virtual endpoints, identity and write path before promotion. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An Azure Database for PostgreSQL read replica falls behind before promotion
Initial hypothesis: Separate time lag, WAL byte gap, primary log retention, replica replay capacity and target connection readiness before planned switchover or forced promotion.
Evidence used: Separate time lag, WAL byte gap, primary log retention, replica replay capacity and target connection readiness before planned switchover or forced promotion.
Checks performed: Freeze the promotion mode and accepted RPO | Compare lag seconds, lag bytes and a business marker | Check primary WAL storage and replica saturation | Validate virtual endpoints, identity and write path before promotion
Decision / action: Run the first checks in order: Freeze the promotion mode and accepted RPO | Compare lag seconds, lag bytes and a business marker | Check primary WAL storage and replica saturation | Validate virtual endpoints, identity and write path before promotion. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Automation

An Azure Automation runbook returns 403 after a managed identity change

Open atlas

Separate cloud sandbox and Hybrid Worker execution, Automation account and VM identities, persisted Az context, control-plane and data-plane permissions before broadening RBAC.

01

Evidence

Separate cloud sandbox and Hybrid Worker execution, Automation account and VM identities, persisted Az context, control-plane and data-plane permissions before broadening RBAC.

02

First checks

  • Freeze one job ID, execution target and denied operation
  • Reset Az context and prove the runtime principal
  • Match the exact action to role and scope
  • Run positive and negative canaries before promotion
03

Bounded action

Run the first checks in order: Freeze one job ID, execution target and denied operation | Reset Az context and prove the runtime principal | Match the exact action to role and scope | Run positive and negative canaries before promotion. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An Azure Automation runbook returns 403 after a managed identity change
Context: Separate cloud sandbox and Hybrid Worker execution, Automation account and VM identities, persisted Az context, control-plane and data-plane permissions before broadening RBAC.
Evidence to confirm: Separate cloud sandbox and Hybrid Worker execution, Automation account and VM identities, persisted Az context, control-plane and data-plane permissions before broadening RBAC.
Immediate checks: Freeze one job ID, execution target and denied operation | Reset Az context and prove the runtime principal | Match the exact action to role and scope | Run positive and negative canaries before promotion
Proposed action: Run the first checks in order: Freeze one job ID, execution target and denied operation | Reset Az context and prove the runtime principal | Match the exact action to role and scope | Run positive and negative canaries before promotion. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An Azure Automation runbook returns 403 after a managed identity change
Initial hypothesis: Separate cloud sandbox and Hybrid Worker execution, Automation account and VM identities, persisted Az context, control-plane and data-plane permissions before broadening RBAC.
Evidence used: Separate cloud sandbox and Hybrid Worker execution, Automation account and VM identities, persisted Az context, control-plane and data-plane permissions before broadening RBAC.
Checks performed: Freeze one job ID, execution target and denied operation | Reset Az context and prove the runtime principal | Match the exact action to role and scope | Run positive and negative canaries before promotion
Decision / action: Run the first checks in order: Freeze one job ID, execution target and denied operation | Reset Az context and prove the runtime principal | Match the exact action to role and scope | Run positive and negative canaries before promotion. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Networking

An Azure Front Door origin becomes unhealthy and traffic shifts unexpectedly

Open atlas

Separate the health probe contract, DNS/TLS, optional Private Link, backend capacity and routing changes before draining, failing over or restoring traffic.

01

Evidence

Separate the health probe contract, DNS/TLS, optional Private Link, backend capacity and routing changes before draining, failing over or restoring traffic.

02

First checks

  • Freeze the route, origin group and UTC window
  • Read failed probes by origin and POP
  • Verify hostname, SNI, certificate and health response
  • Prove standby capacity before changing weights
03

Bounded action

Run the first checks in order: Freeze the route, origin group and UTC window | Read failed probes by origin and POP | Verify hostname, SNI, certificate and health response | Prove standby capacity before changing weights. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An Azure Front Door origin becomes unhealthy and traffic shifts unexpectedly
Context: Separate the health probe contract, DNS/TLS, optional Private Link, backend capacity and routing changes before draining, failing over or restoring traffic.
Evidence to confirm: Separate the health probe contract, DNS/TLS, optional Private Link, backend capacity and routing changes before draining, failing over or restoring traffic.
Immediate checks: Freeze the route, origin group and UTC window | Read failed probes by origin and POP | Verify hostname, SNI, certificate and health response | Prove standby capacity before changing weights
Proposed action: Run the first checks in order: Freeze the route, origin group and UTC window | Read failed probes by origin and POP | Verify hostname, SNI, certificate and health response | Prove standby capacity before changing weights. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An Azure Front Door origin becomes unhealthy and traffic shifts unexpectedly
Initial hypothesis: Separate the health probe contract, DNS/TLS, optional Private Link, backend capacity and routing changes before draining, failing over or restoring traffic.
Evidence used: Separate the health probe contract, DNS/TLS, optional Private Link, backend capacity and routing changes before draining, failing over or restoring traffic.
Checks performed: Freeze the route, origin group and UTC window | Read failed probes by origin and POP | Verify hostname, SNI, certificate and health response | Prove standby capacity before changing weights
Decision / action: Run the first checks in order: Freeze the route, origin group and UTC window | Read failed probes by origin and POP | Verify hostname, SNI, certificate and health response | Prove standby capacity before changing weights. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
AI

An AI agent recalls context from another session or team

Open atlas

Separate conversation history, compacted summaries, durable memory, retrieval caches, tool results and runtime identity before enabling persistence for more users.

01

Evidence

Separate conversation history, compacted summaries, durable memory, retrieval caches, tool results and runtime identity before enabling persistence for more users.

02

First checks

  • Freeze the user, team, environment and session boundaries
  • Seed distinct synthetic canaries in two scopes
  • Enable each state layer independently
  • Repeat under concurrency, expiry and revocation
03

Bounded action

Run the first checks in order: Freeze the user, team, environment and session boundaries | Seed distinct synthetic canaries in two scopes | Enable each state layer independently | Repeat under concurrency, expiry and revocation. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An AI agent recalls context from another session or team
Context: Separate conversation history, compacted summaries, durable memory, retrieval caches, tool results and runtime identity before enabling persistence for more users.
Evidence to confirm: Separate conversation history, compacted summaries, durable memory, retrieval caches, tool results and runtime identity before enabling persistence for more users.
Immediate checks: Freeze the user, team, environment and session boundaries | Seed distinct synthetic canaries in two scopes | Enable each state layer independently | Repeat under concurrency, expiry and revocation
Proposed action: Run the first checks in order: Freeze the user, team, environment and session boundaries | Seed distinct synthetic canaries in two scopes | Enable each state layer independently | Repeat under concurrency, expiry and revocation. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An AI agent recalls context from another session or team
Initial hypothesis: Separate conversation history, compacted summaries, durable memory, retrieval caches, tool results and runtime identity before enabling persistence for more users.
Evidence used: Separate conversation history, compacted summaries, durable memory, retrieval caches, tool results and runtime identity before enabling persistence for more users.
Checks performed: Freeze the user, team, environment and session boundaries | Seed distinct synthetic canaries in two scopes | Enable each state layer independently | Repeat under concurrency, expiry and revocation
Decision / action: Run the first checks in order: Freeze the user, team, environment and session boundaries | Seed distinct synthetic canaries in two scopes | Enable each state layer independently | Repeat under concurrency, expiry and revocation. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Infrastructure

An Azure Update Manager maintenance window ends with part of the fleet still noncompliant

Open atlas

Separate dynamic-scope membership, schedule orchestration, exhausted window budget, patch failures and pending restarts before forcing an out-of-window installation.

01

Evidence

Separate dynamic-scope membership, schedule orchestration, exhausted window budget, patch failures and pending restarts before forcing an out-of-window installation.

02

First checks

  • Freeze the maintenance contract and expected inventory
  • Compare matching machines with maintenance and installation runs
  • Separate no run, partial run, failed update and pending restart
  • Prove service capacity before resuming a bounded set
03

Bounded action

Run the first checks in order: Freeze the maintenance contract and expected inventory | Compare matching machines with maintenance and installation runs | Separate no run, partial run, failed update and pending restart | Prove service capacity before resuming a bounded set. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An Azure Update Manager maintenance window ends with part of the fleet still noncompliant
Context: Separate dynamic-scope membership, schedule orchestration, exhausted window budget, patch failures and pending restarts before forcing an out-of-window installation.
Evidence to confirm: Separate dynamic-scope membership, schedule orchestration, exhausted window budget, patch failures and pending restarts before forcing an out-of-window installation.
Immediate checks: Freeze the maintenance contract and expected inventory | Compare matching machines with maintenance and installation runs | Separate no run, partial run, failed update and pending restart | Prove service capacity before resuming a bounded set
Proposed action: Run the first checks in order: Freeze the maintenance contract and expected inventory | Compare matching machines with maintenance and installation runs | Separate no run, partial run, failed update and pending restart | Prove service capacity before resuming a bounded set. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An Azure Update Manager maintenance window ends with part of the fleet still noncompliant
Initial hypothesis: Separate dynamic-scope membership, schedule orchestration, exhausted window budget, patch failures and pending restarts before forcing an out-of-window installation.
Evidence used: Separate dynamic-scope membership, schedule orchestration, exhausted window budget, patch failures and pending restarts before forcing an out-of-window installation.
Checks performed: Freeze the maintenance contract and expected inventory | Compare matching machines with maintenance and installation runs | Separate no run, partial run, failed update and pending restart | Prove service capacity before resuming a bounded set
Decision / action: Run the first checks in order: Freeze the maintenance contract and expected inventory | Compare matching machines with maintenance and installation runs | Separate no run, partial run, failed update and pending restart | Prove service capacity before resuming a bounded set. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Cloud

An Azure App Service certificate is renewed but the old certificate is still served

Open atlas

Separate issuance, domain validation, Key Vault version, App Service import, hostname binding and the certificate observed with SNI before syncing or rebinding.

01

Evidence

Separate issuance, domain validation, Key Vault version, App Service import, hostname binding and the certificate observed with SNI before syncing or rebinding.

02

First checks

  • Read the served fingerprint with the real hostname and SNI
  • Compare candidate SAN and expiration with App Service inventory
  • Check Key Vault version and import state when applicable
  • Preserve the current binding and rollback thumbprint
03

Bounded action

Run the first checks in order: Read the served fingerprint with the real hostname and SNI | Compare candidate SAN and expiration with App Service inventory | Check Key Vault version and import state when applicable | Preserve the current binding and rollback thumbprint. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An Azure App Service certificate is renewed but the old certificate is still served
Context: Separate issuance, domain validation, Key Vault version, App Service import, hostname binding and the certificate observed with SNI before syncing or rebinding.
Evidence to confirm: Separate issuance, domain validation, Key Vault version, App Service import, hostname binding and the certificate observed with SNI before syncing or rebinding.
Immediate checks: Read the served fingerprint with the real hostname and SNI | Compare candidate SAN and expiration with App Service inventory | Check Key Vault version and import state when applicable | Preserve the current binding and rollback thumbprint
Proposed action: Run the first checks in order: Read the served fingerprint with the real hostname and SNI | Compare candidate SAN and expiration with App Service inventory | Check Key Vault version and import state when applicable | Preserve the current binding and rollback thumbprint. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An Azure App Service certificate is renewed but the old certificate is still served
Initial hypothesis: Separate issuance, domain validation, Key Vault version, App Service import, hostname binding and the certificate observed with SNI before syncing or rebinding.
Evidence used: Separate issuance, domain validation, Key Vault version, App Service import, hostname binding and the certificate observed with SNI before syncing or rebinding.
Checks performed: Read the served fingerprint with the real hostname and SNI | Compare candidate SAN and expiration with App Service inventory | Check Key Vault version and import state when applicable | Preserve the current binding and rollback thumbprint
Decision / action: Run the first checks in order: Read the served fingerprint with the real hostname and SNI | Compare candidate SAN and expiration with App Service inventory | Check Key Vault version and import state when applicable | Preserve the current binding and rollback thumbprint. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Networking

An Azure flow is denied although the NSG allows it

Open atlas

Separate Azure Virtual Network Manager Security Admin rules, network group membership, NSG evaluation and actual service reachability before opening local rules.

01

Evidence

Separate Azure Virtual Network Manager Security Admin rules, network group membership, NSG evaluation and actual service reachability before opening local rules.

02

First checks

  • Capture the exact source, destination and port tuple
  • List effective Security Admin rules on the VNet
  • Run IP Flow Verify and identify the deciding rule
  • Confirm the decision in VNet flow logs
03

Bounded action

Run the first checks in order: Capture the exact source, destination and port tuple | List effective Security Admin rules on the VNet | Run IP Flow Verify and identify the deciding rule | Confirm the decision in VNet flow logs. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An Azure flow is denied although the NSG allows it
Context: Separate Azure Virtual Network Manager Security Admin rules, network group membership, NSG evaluation and actual service reachability before opening local rules.
Evidence to confirm: Separate Azure Virtual Network Manager Security Admin rules, network group membership, NSG evaluation and actual service reachability before opening local rules.
Immediate checks: Capture the exact source, destination and port tuple | List effective Security Admin rules on the VNet | Run IP Flow Verify and identify the deciding rule | Confirm the decision in VNet flow logs
Proposed action: Run the first checks in order: Capture the exact source, destination and port tuple | List effective Security Admin rules on the VNet | Run IP Flow Verify and identify the deciding rule | Confirm the decision in VNet flow logs. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An Azure flow is denied although the NSG allows it
Initial hypothesis: Separate Azure Virtual Network Manager Security Admin rules, network group membership, NSG evaluation and actual service reachability before opening local rules.
Evidence used: Separate Azure Virtual Network Manager Security Admin rules, network group membership, NSG evaluation and actual service reachability before opening local rules.
Checks performed: Capture the exact source, destination and port tuple | List effective Security Admin rules on the VNet | Run IP Flow Verify and identify the deciding rule | Confirm the decision in VNet flow logs
Decision / action: Run the first checks in order: Capture the exact source, destination and port tuple | List effective Security Admin rules on the VNet | Run IP Flow Verify and identify the deciding rule | Confirm the decision in VNet flow logs. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Networking

An Azure Traffic Manager profile is degraded and failover appears inconsistent across clients

Open atlas

Separate the deployed probe contract, endpoint monitor state, backend evidence, DNS answers, TTL and secondary capacity before disabling the primary endpoint or forcing failover.

01

Evidence

Separate the deployed probe contract, endpoint monitor state, backend evidence, DNS answers, TTL and secondary capacity before disabling the primary endpoint or forcing failover.

02

First checks

  • Capture profile, endpoint, probe and TTL configuration
  • Replay the exact probe contract from external viewpoints
  • Correlate ProbeHealthStatusEvents with backend logs
  • Compare DNS answers and prove secondary readiness
03

Bounded action

Run the first checks in order: Capture profile, endpoint, probe and TTL configuration | Replay the exact probe contract from external viewpoints | Correlate ProbeHealthStatusEvents with backend logs | Compare DNS answers and prove secondary readiness. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An Azure Traffic Manager profile is degraded and failover appears inconsistent across clients
Context: Separate the deployed probe contract, endpoint monitor state, backend evidence, DNS answers, TTL and secondary capacity before disabling the primary endpoint or forcing failover.
Evidence to confirm: Separate the deployed probe contract, endpoint monitor state, backend evidence, DNS answers, TTL and secondary capacity before disabling the primary endpoint or forcing failover.
Immediate checks: Capture profile, endpoint, probe and TTL configuration | Replay the exact probe contract from external viewpoints | Correlate ProbeHealthStatusEvents with backend logs | Compare DNS answers and prove secondary readiness
Proposed action: Run the first checks in order: Capture profile, endpoint, probe and TTL configuration | Replay the exact probe contract from external viewpoints | Correlate ProbeHealthStatusEvents with backend logs | Compare DNS answers and prove secondary readiness. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An Azure Traffic Manager profile is degraded and failover appears inconsistent across clients
Initial hypothesis: Separate the deployed probe contract, endpoint monitor state, backend evidence, DNS answers, TTL and secondary capacity before disabling the primary endpoint or forcing failover.
Evidence used: Separate the deployed probe contract, endpoint monitor state, backend evidence, DNS answers, TTL and secondary capacity before disabling the primary endpoint or forcing failover.
Checks performed: Capture profile, endpoint, probe and TTL configuration | Replay the exact probe contract from external viewpoints | Correlate ProbeHealthStatusEvents with backend logs | Compare DNS answers and prove secondary readiness
Decision / action: Run the first checks in order: Capture profile, endpoint, probe and TTL configuration | Replay the exact probe contract from external viewpoints | Correlate ProbeHealthStatusEvents with backend logs | Compare DNS answers and prove secondary readiness. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Automation

An Azure Automation runbook works in its current runtime but fails in the candidate environment

Open atlas

Separate language version, package resolution, identity, serialization and Hybrid Worker readiness before relinking production runbooks.

01

Evidence

Separate language version, package resolution, identity, serialization and Hybrid Worker readiness before relinking production runbooks.

02

First checks

  • Freeze current and candidate runtime manifests
  • Run positive and negative draft tests with the candidate
  • Qualify every target Hybrid Worker
  • Canary one bounded effect and confirm an empty second run
03

Bounded action

Run the first checks in order: Freeze current and candidate runtime manifests | Run positive and negative draft tests with the candidate | Qualify every target Hybrid Worker | Canary one bounded effect and confirm an empty second run. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An Azure Automation runbook works in its current runtime but fails in the candidate environment
Context: Separate language version, package resolution, identity, serialization and Hybrid Worker readiness before relinking production runbooks.
Evidence to confirm: Separate language version, package resolution, identity, serialization and Hybrid Worker readiness before relinking production runbooks.
Immediate checks: Freeze current and candidate runtime manifests | Run positive and negative draft tests with the candidate | Qualify every target Hybrid Worker | Canary one bounded effect and confirm an empty second run
Proposed action: Run the first checks in order: Freeze current and candidate runtime manifests | Run positive and negative draft tests with the candidate | Qualify every target Hybrid Worker | Canary one bounded effect and confirm an empty second run. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An Azure Automation runbook works in its current runtime but fails in the candidate environment
Initial hypothesis: Separate language version, package resolution, identity, serialization and Hybrid Worker readiness before relinking production runbooks.
Evidence used: Separate language version, package resolution, identity, serialization and Hybrid Worker readiness before relinking production runbooks.
Checks performed: Freeze current and candidate runtime manifests | Run positive and negative draft tests with the candidate | Qualify every target Hybrid Worker | Canary one bounded effect and confirm an empty second run
Decision / action: Run the first checks in order: Freeze current and candidate runtime manifests | Run positive and negative draft tests with the candidate | Qualify every target Hybrid Worker | Canary one bounded effect and confirm an empty second run. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Infrastructure

Azure Monitor Managed Prometheus cardinality spikes after an AKS deployment

Open atlas

Separate workspace ingestion pressure, scrape failures, duplicate targets and unbounded labels before dropping metrics, raising quotas or rolling back the release.

01

Evidence

Separate workspace ingestion pressure, scrape failures, duplicate targets and unbounded labels before dropping metrics, raising quotas or rolling back the release.

02

First checks

  • Freeze workspace, cluster, release and UTC window
  • Compare active-series and event-ingestion trends
  • Bound PromQL to identify the metric and label
  • Canary relabeling without breaking alerts
03

Bounded action

Run the first checks in order: Freeze workspace, cluster, release and UTC window | Compare active-series and event-ingestion trends | Bound PromQL to identify the metric and label | Canary relabeling without breaking alerts. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] Azure Monitor Managed Prometheus cardinality spikes after an AKS deployment
Context: Separate workspace ingestion pressure, scrape failures, duplicate targets and unbounded labels before dropping metrics, raising quotas or rolling back the release.
Evidence to confirm: Separate workspace ingestion pressure, scrape failures, duplicate targets and unbounded labels before dropping metrics, raising quotas or rolling back the release.
Immediate checks: Freeze workspace, cluster, release and UTC window | Compare active-series and event-ingestion trends | Bound PromQL to identify the metric and label | Canary relabeling without breaking alerts
Proposed action: Run the first checks in order: Freeze workspace, cluster, release and UTC window | Compare active-series and event-ingestion trends | Bound PromQL to identify the metric and label | Canary relabeling without breaking alerts. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: Azure Monitor Managed Prometheus cardinality spikes after an AKS deployment
Initial hypothesis: Separate workspace ingestion pressure, scrape failures, duplicate targets and unbounded labels before dropping metrics, raising quotas or rolling back the release.
Evidence used: Separate workspace ingestion pressure, scrape failures, duplicate targets and unbounded labels before dropping metrics, raising quotas or rolling back the release.
Checks performed: Freeze workspace, cluster, release and UTC window | Compare active-series and event-ingestion trends | Bound PromQL to identify the metric and label | Canary relabeling without breaking alerts
Decision / action: Run the first checks in order: Freeze workspace, cluster, release and UTC window | Compare active-series and event-ingestion trends | Bound PromQL to identify the metric and label | Canary relabeling without breaking alerts. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Infrastructure

Azure VMSS repeatedly replaces or reimages unhealthy instances

Open atlas

Separate the health signal, grace period, orchestration service state, current VMSS model and persistence contract before changing Replace, Reimage or Restart.

01

Evidence

Separate the health signal, grace period, orchestration service state, current VMSS model and persistence contract before changing Replace, Reimage or Restart.

02

First checks

  • Freeze the repair timeline and useful capacity
  • Read automaticRepairsPolicy and orchestration service state
  • Replay the health contract on healthy and affected instances
  • Build one canary from the current model before resuming
03

Bounded action

Run the first checks in order: Freeze the repair timeline and useful capacity | Read automaticRepairsPolicy and orchestration service state | Replay the health contract on healthy and affected instances | Build one canary from the current model before resuming. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] Azure VMSS repeatedly replaces or reimages unhealthy instances
Context: Separate the health signal, grace period, orchestration service state, current VMSS model and persistence contract before changing Replace, Reimage or Restart.
Evidence to confirm: Separate the health signal, grace period, orchestration service state, current VMSS model and persistence contract before changing Replace, Reimage or Restart.
Immediate checks: Freeze the repair timeline and useful capacity | Read automaticRepairsPolicy and orchestration service state | Replay the health contract on healthy and affected instances | Build one canary from the current model before resuming
Proposed action: Run the first checks in order: Freeze the repair timeline and useful capacity | Read automaticRepairsPolicy and orchestration service state | Replay the health contract on healthy and affected instances | Build one canary from the current model before resuming. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: Azure VMSS repeatedly replaces or reimages unhealthy instances
Initial hypothesis: Separate the health signal, grace period, orchestration service state, current VMSS model and persistence contract before changing Replace, Reimage or Restart.
Evidence used: Separate the health signal, grace period, orchestration service state, current VMSS model and persistence contract before changing Replace, Reimage or Restart.
Checks performed: Freeze the repair timeline and useful capacity | Read automaticRepairsPolicy and orchestration service state | Replay the health contract on healthy and affected instances | Build one canary from the current model before resuming
Decision / action: Run the first checks in order: Freeze the repair timeline and useful capacity | Read automaticRepairsPolicy and orchestration service state | Replay the health contract on healthy and affected instances | Build one canary from the current model before resuming. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Automation

A newly pushed ACR image is not visible from every deployment region

Open atlas

Separate asynchronous geo-replication, mutable tags, global endpoint routing, runtime identity and data endpoint reachability before retrying a multi-region rollout.

01

Evidence

Separate asynchronous geo-replication, mutable tags, global endpoint routing, runtime identity and data endpoint reachability before retrying a multi-region rollout.

02

First checks

  • Freeze the build digest and push completion time
  • List replica state and inspect Resource Health
  • Pull the expected digest from every target region
  • Correlate ACR events by region, identity and result
03

Bounded action

Run the first checks in order: Freeze the build digest and push completion time | List replica state and inspect Resource Health | Pull the expected digest from every target region | Correlate ACR events by region, identity and result. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] A newly pushed ACR image is not visible from every deployment region
Context: Separate asynchronous geo-replication, mutable tags, global endpoint routing, runtime identity and data endpoint reachability before retrying a multi-region rollout.
Evidence to confirm: Separate asynchronous geo-replication, mutable tags, global endpoint routing, runtime identity and data endpoint reachability before retrying a multi-region rollout.
Immediate checks: Freeze the build digest and push completion time | List replica state and inspect Resource Health | Pull the expected digest from every target region | Correlate ACR events by region, identity and result
Proposed action: Run the first checks in order: Freeze the build digest and push completion time | List replica state and inspect Resource Health | Pull the expected digest from every target region | Correlate ACR events by region, identity and result. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: A newly pushed ACR image is not visible from every deployment region
Initial hypothesis: Separate asynchronous geo-replication, mutable tags, global endpoint routing, runtime identity and data endpoint reachability before retrying a multi-region rollout.
Evidence used: Separate asynchronous geo-replication, mutable tags, global endpoint routing, runtime identity and data endpoint reachability before retrying a multi-region rollout.
Checks performed: Freeze the build digest and push completion time | List replica state and inspect Resource Health | Pull the expected digest from every target region | Correlate ACR events by region, identity and result
Decision / action: Run the first checks in order: Freeze the build digest and push completion time | List replica state and inspect Resource Health | Pull the expected digest from every target region | Correlate ACR events by region, identity and result. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Networking

An AKS service-to-service flow is blocked after a NetworkPolicy change

Open atlas

Separate source egress, destination ingress, deployed selectors, service endpoints and observed Cilium verdicts before adding an allow-all rule or removing default deny.

01

Evidence

Separate source egress, destination ingress, deployed selectors, service endpoints and observed Cilium verdicts before adding an allow-all rule or removing default deny.

02

First checks

  • Freeze the source, destination, port and UTC window
  • Read the AKS data plane and every policy type
  • Compare selectors with deployed pod and namespace labels
  • Run positive and negative canaries before promotion
03

Bounded action

Run the first checks in order: Freeze the source, destination, port and UTC window | Read the AKS data plane and every policy type | Compare selectors with deployed pod and namespace labels | Run positive and negative canaries before promotion. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An AKS service-to-service flow is blocked after a NetworkPolicy change
Context: Separate source egress, destination ingress, deployed selectors, service endpoints and observed Cilium verdicts before adding an allow-all rule or removing default deny.
Evidence to confirm: Separate source egress, destination ingress, deployed selectors, service endpoints and observed Cilium verdicts before adding an allow-all rule or removing default deny.
Immediate checks: Freeze the source, destination, port and UTC window | Read the AKS data plane and every policy type | Compare selectors with deployed pod and namespace labels | Run positive and negative canaries before promotion
Proposed action: Run the first checks in order: Freeze the source, destination, port and UTC window | Read the AKS data plane and every policy type | Compare selectors with deployed pod and namespace labels | Run positive and negative canaries before promotion. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An AKS service-to-service flow is blocked after a NetworkPolicy change
Initial hypothesis: Separate source egress, destination ingress, deployed selectors, service endpoints and observed Cilium verdicts before adding an allow-all rule or removing default deny.
Evidence used: Separate source egress, destination ingress, deployed selectors, service endpoints and observed Cilium verdicts before adding an allow-all rule or removing default deny.
Checks performed: Freeze the source, destination, port and UTC window | Read the AKS data plane and every policy type | Compare selectors with deployed pod and namespace labels | Run positive and negative canaries before promotion
Decision / action: Run the first checks in order: Freeze the source, destination, port and UTC window | Read the AKS data plane and every policy type | Compare selectors with deployed pod and namespace labels | Run positive and negative canaries before promotion. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Infrastructure

OpenTelemetry tail sampling reduces volume but incident traces become incomplete or disappear

Open atlas

Separate trace affinity, decision-window sizing, pending-trace memory, late spans, policy behavior and export evidence before widening the sampling rollout.

01

Evidence

Separate trace affinity, decision-window sizing, pending-trace memory, late spans, policy behavior and export evidence before widening the sampling rollout.

02

First checks

  • Freeze the current and candidate Collector configuration digests
  • Prove every span for one trace reaches the same sampling shard
  • Replay error, latency and late-span canaries
  • Correlate Collector decisions with complete operations in Azure Monitor
03

Bounded action

Run the first checks in order: Freeze the current and candidate Collector configuration digests | Prove every span for one trace reaches the same sampling shard | Replay error, latency and late-span canaries | Correlate Collector decisions with complete operations in Azure Monitor. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] OpenTelemetry tail sampling reduces volume but incident traces become incomplete or disappear
Context: Separate trace affinity, decision-window sizing, pending-trace memory, late spans, policy behavior and export evidence before widening the sampling rollout.
Evidence to confirm: Separate trace affinity, decision-window sizing, pending-trace memory, late spans, policy behavior and export evidence before widening the sampling rollout.
Immediate checks: Freeze the current and candidate Collector configuration digests | Prove every span for one trace reaches the same sampling shard | Replay error, latency and late-span canaries | Correlate Collector decisions with complete operations in Azure Monitor
Proposed action: Run the first checks in order: Freeze the current and candidate Collector configuration digests | Prove every span for one trace reaches the same sampling shard | Replay error, latency and late-span canaries | Correlate Collector decisions with complete operations in Azure Monitor. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: OpenTelemetry tail sampling reduces volume but incident traces become incomplete or disappear
Initial hypothesis: Separate trace affinity, decision-window sizing, pending-trace memory, late spans, policy behavior and export evidence before widening the sampling rollout.
Evidence used: Separate trace affinity, decision-window sizing, pending-trace memory, late spans, policy behavior and export evidence before widening the sampling rollout.
Checks performed: Freeze the current and candidate Collector configuration digests | Prove every span for one trace reaches the same sampling shard | Replay error, latency and late-span canaries | Correlate Collector decisions with complete operations in Azure Monitor
Decision / action: Run the first checks in order: Freeze the current and candidate Collector configuration digests | Prove every span for one trace reaches the same sampling shard | Replay error, latency and late-span canaries | Correlate Collector decisions with complete operations in Azure Monitor. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Infrastructure

Key Vault returns 403 after an Azure RBAC cutover

Open atlas

Separate missed runtime identities, data-plane role mapping, assignment scope and propagation from network failures before granting a broad role or reverting the permission model.

01

Evidence

Separate missed runtime identities, data-plane role mapping, assignment scope and propagation from network failures before granting a broad role or reverting the permission model.

02

First checks

  • Confirm the active Key Vault permission model
  • Resolve the failing runtime principal object ID
  • Compare required operations with direct and inherited data-plane roles
  • Correlate 403 responses with Key Vault AuditEvent logs
03

Bounded action

Run the first checks in order: Confirm the active Key Vault permission model | Resolve the failing runtime principal object ID | Compare required operations with direct and inherited data-plane roles | Correlate 403 responses with Key Vault AuditEvent logs. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] Key Vault returns 403 after an Azure RBAC cutover
Context: Separate missed runtime identities, data-plane role mapping, assignment scope and propagation from network failures before granting a broad role or reverting the permission model.
Evidence to confirm: Separate missed runtime identities, data-plane role mapping, assignment scope and propagation from network failures before granting a broad role or reverting the permission model.
Immediate checks: Confirm the active Key Vault permission model | Resolve the failing runtime principal object ID | Compare required operations with direct and inherited data-plane roles | Correlate 403 responses with Key Vault AuditEvent logs
Proposed action: Run the first checks in order: Confirm the active Key Vault permission model | Resolve the failing runtime principal object ID | Compare required operations with direct and inherited data-plane roles | Correlate 403 responses with Key Vault AuditEvent logs. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: Key Vault returns 403 after an Azure RBAC cutover
Initial hypothesis: Separate missed runtime identities, data-plane role mapping, assignment scope and propagation from network failures before granting a broad role or reverting the permission model.
Evidence used: Separate missed runtime identities, data-plane role mapping, assignment scope and propagation from network failures before granting a broad role or reverting the permission model.
Checks performed: Confirm the active Key Vault permission model | Resolve the failing runtime principal object ID | Compare required operations with direct and inherited data-plane roles | Correlate 403 responses with Key Vault AuditEvent logs
Decision / action: Run the first checks in order: Confirm the active Key Vault permission model | Resolve the failing runtime principal object ID | Compare required operations with direct and inherited data-plane roles | Correlate 403 responses with Key Vault AuditEvent logs. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Networking

Virtual WAN traffic breaks or becomes asymmetric after Routing Intent is enabled

Open atlas

Separate policy scope, effective routes, private prefixes, default-route propagation, security next hop, SNAT and return paths before widening the rollout or deleting Routing Intent.

01

Evidence

Separate policy scope, effective routes, private prefixes, default-route propagation, security next hop, SNAT and return paths before widening the rollout or deleting Routing Intent.

02

First checks

  • Freeze the former route associations and propagations
  • Verify every contract prefix on the security next hop
  • Probe forward and return paths through each affected hub
  • Correlate canaries with Firewall allows and denies
03

Bounded action

Run the first checks in order: Freeze the former route associations and propagations | Verify every contract prefix on the security next hop | Probe forward and return paths through each affected hub | Correlate canaries with Firewall allows and denies. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] Virtual WAN traffic breaks or becomes asymmetric after Routing Intent is enabled
Context: Separate policy scope, effective routes, private prefixes, default-route propagation, security next hop, SNAT and return paths before widening the rollout or deleting Routing Intent.
Evidence to confirm: Separate policy scope, effective routes, private prefixes, default-route propagation, security next hop, SNAT and return paths before widening the rollout or deleting Routing Intent.
Immediate checks: Freeze the former route associations and propagations | Verify every contract prefix on the security next hop | Probe forward and return paths through each affected hub | Correlate canaries with Firewall allows and denies
Proposed action: Run the first checks in order: Freeze the former route associations and propagations | Verify every contract prefix on the security next hop | Probe forward and return paths through each affected hub | Correlate canaries with Firewall allows and denies. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: Virtual WAN traffic breaks or becomes asymmetric after Routing Intent is enabled
Initial hypothesis: Separate policy scope, effective routes, private prefixes, default-route propagation, security next hop, SNAT and return paths before widening the rollout or deleting Routing Intent.
Evidence used: Separate policy scope, effective routes, private prefixes, default-route propagation, security next hop, SNAT and return paths before widening the rollout or deleting Routing Intent.
Checks performed: Freeze the former route associations and propagations | Verify every contract prefix on the security next hop | Probe forward and return paths through each affected hub | Correlate canaries with Firewall allows and denies
Decision / action: Run the first checks in order: Freeze the former route associations and propagations | Verify every contract prefix on the security next hop | Probe forward and return paths through each affected hub | Correlate canaries with Firewall allows and denies. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Networking

AKS pods see intermittent DNS timeouts or SERVFAIL responses

Open atlas

Separate pod resolver behavior, query amplification, the Kubernetes DNS service, CoreDNS endpoints, network policy, upstream forwarding and node-specific paths before restarting CoreDNS.

01

Evidence

Separate pod resolver behavior, query amplification, the Kubernetes DNS service, CoreDNS endpoints, network policy, upstream forwarding and node-specific paths before restarting CoreDNS.

02

First checks

  • Freeze the exact names, errors and affected pod/node cohort
  • Compare repeated lookups across nodes and namespaces
  • Inspect the kube-dns service, EndpointSlices and CoreDNS pods
  • Test cluster names separately from each upstream suffix
03

Bounded action

Run the first checks in order: Freeze the exact names, errors and affected pod/node cohort | Compare repeated lookups across nodes and namespaces | Inspect the kube-dns service, EndpointSlices and CoreDNS pods | Test cluster names separately from each upstream suffix. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] AKS pods see intermittent DNS timeouts or SERVFAIL responses
Context: Separate pod resolver behavior, query amplification, the Kubernetes DNS service, CoreDNS endpoints, network policy, upstream forwarding and node-specific paths before restarting CoreDNS.
Evidence to confirm: Separate pod resolver behavior, query amplification, the Kubernetes DNS service, CoreDNS endpoints, network policy, upstream forwarding and node-specific paths before restarting CoreDNS.
Immediate checks: Freeze the exact names, errors and affected pod/node cohort | Compare repeated lookups across nodes and namespaces | Inspect the kube-dns service, EndpointSlices and CoreDNS pods | Test cluster names separately from each upstream suffix
Proposed action: Run the first checks in order: Freeze the exact names, errors and affected pod/node cohort | Compare repeated lookups across nodes and namespaces | Inspect the kube-dns service, EndpointSlices and CoreDNS pods | Test cluster names separately from each upstream suffix. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: AKS pods see intermittent DNS timeouts or SERVFAIL responses
Initial hypothesis: Separate pod resolver behavior, query amplification, the Kubernetes DNS service, CoreDNS endpoints, network policy, upstream forwarding and node-specific paths before restarting CoreDNS.
Evidence used: Separate pod resolver behavior, query amplification, the Kubernetes DNS service, CoreDNS endpoints, network policy, upstream forwarding and node-specific paths before restarting CoreDNS.
Checks performed: Freeze the exact names, errors and affected pod/node cohort | Compare repeated lookups across nodes and namespaces | Inspect the kube-dns service, EndpointSlices and CoreDNS pods | Test cluster names separately from each upstream suffix
Decision / action: Run the first checks in order: Freeze the exact names, errors and affected pod/node cohort | Compare repeated lookups across nodes and namespaces | Inspect the kube-dns service, EndpointSlices and CoreDNS pods | Test cluster names separately from each upstream suffix. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Networking

Azure Firewall IDPS blocks a legitimate flow after a signature is moved to deny

Open atlas

Separate signature evidence, traffic direction, private-range classification, TLS inspection and application outcome before disabling IDPS or adding a broad bypass.

01

Evidence

Separate signature evidence, traffic direction, private-range classification, TLS inspection and application outcome before disabling IDPS or adding a broad bypass.

02

First checks

  • Freeze the signature override and affected five-tuple
  • Query AZFWIdpsSignature for source, destination and action
  • Replay one malicious and one legitimate canary
  • Return only the signature to Alert if the legitimate flow regresses
03

Bounded action

Run the first checks in order: Freeze the signature override and affected five-tuple | Query AZFWIdpsSignature for source, destination and action | Replay one malicious and one legitimate canary | Return only the signature to Alert if the legitimate flow regresses. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] Azure Firewall IDPS blocks a legitimate flow after a signature is moved to deny
Context: Separate signature evidence, traffic direction, private-range classification, TLS inspection and application outcome before disabling IDPS or adding a broad bypass.
Evidence to confirm: Separate signature evidence, traffic direction, private-range classification, TLS inspection and application outcome before disabling IDPS or adding a broad bypass.
Immediate checks: Freeze the signature override and affected five-tuple | Query AZFWIdpsSignature for source, destination and action | Replay one malicious and one legitimate canary | Return only the signature to Alert if the legitimate flow regresses
Proposed action: Run the first checks in order: Freeze the signature override and affected five-tuple | Query AZFWIdpsSignature for source, destination and action | Replay one malicious and one legitimate canary | Return only the signature to Alert if the legitimate flow regresses. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: Azure Firewall IDPS blocks a legitimate flow after a signature is moved to deny
Initial hypothesis: Separate signature evidence, traffic direction, private-range classification, TLS inspection and application outcome before disabling IDPS or adding a broad bypass.
Evidence used: Separate signature evidence, traffic direction, private-range classification, TLS inspection and application outcome before disabling IDPS or adding a broad bypass.
Checks performed: Freeze the signature override and affected five-tuple | Query AZFWIdpsSignature for source, destination and action | Replay one malicious and one legitimate canary | Return only the signature to Alert if the legitimate flow regresses
Decision / action: Run the first checks in order: Freeze the signature override and affected five-tuple | Query AZFWIdpsSignature for source, destination and action | Replay one malicious and one legitimate canary | Return only the signature to Alert if the legitimate flow regresses. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Automation

An Azure Automation job may run a version that does not match the approved commit

Open atlas

Tie the source commit, sync job, Draft and Published hashes, no-effect test and execution marker together before rerunning or broadening production scope.

01

Evidence

Tie the source commit, sync job, Draft and Published hashes, no-effect test and execution marker together before rerunning or broadening production scope.

02

First checks

  • Freeze the approved commit, script hash and target account
  • Read the last sync job and its streams before retrying
  • Export and hash Draft and Published separately
  • Run Plan and one reversible canary before full execution
03

Bounded action

Run the first checks in order: Freeze the approved commit, script hash and target account | Read the last sync job and its streams before retrying | Export and hash Draft and Published separately | Run Plan and one reversible canary before full execution. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An Azure Automation job may run a version that does not match the approved commit
Context: Tie the source commit, sync job, Draft and Published hashes, no-effect test and execution marker together before rerunning or broadening production scope.
Evidence to confirm: Tie the source commit, sync job, Draft and Published hashes, no-effect test and execution marker together before rerunning or broadening production scope.
Immediate checks: Freeze the approved commit, script hash and target account | Read the last sync job and its streams before retrying | Export and hash Draft and Published separately | Run Plan and one reversible canary before full execution
Proposed action: Run the first checks in order: Freeze the approved commit, script hash and target account | Read the last sync job and its streams before retrying | Export and hash Draft and Published separately | Run Plan and one reversible canary before full execution. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An Azure Automation job may run a version that does not match the approved commit
Initial hypothesis: Tie the source commit, sync job, Draft and Published hashes, no-effect test and execution marker together before rerunning or broadening production scope.
Evidence used: Tie the source commit, sync job, Draft and Published hashes, no-effect test and execution marker together before rerunning or broadening production scope.
Checks performed: Freeze the approved commit, script hash and target account | Read the last sync job and its streams before retrying | Export and hash Draft and Published separately | Run Plan and one reversible canary before full execution
Decision / action: Run the first checks in order: Freeze the approved commit, script hash and target account | Read the last sync job and its streams before retrying | Export and hash Draft and Published separately | Run Plan and one reversible canary before full execution. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Infrastructure

One Azure Monitor log alert creates hundreds of independently evaluated alert instances

Open atlas

Separate real affected cohorts, split-by dimension cardinality, alert lifecycle and notification routing before raising limits, muting the action group or changing the KQL rule.

01

Evidence

Separate real affected cohorts, split-by dimension cardinality, alert lifecycle and notification routing before raising limits, muting the action group or changing the KQL rule.

02

First checks

  • Export the deployed scheduled query rule and freeze a UTC window
  • Count rows, distinct dimension values and their combinations
  • Compare query combinations, active instances and delivered notifications
  • Canary stable dimensions through a complete fire-and-resolve cycle
03

Bounded action

Run the first checks in order: Export the deployed scheduled query rule and freeze a UTC window | Count rows, distinct dimension values and their combinations | Compare query combinations, active instances and delivered notifications | Canary stable dimensions through a complete fire-and-resolve cycle. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] One Azure Monitor log alert creates hundreds of independently evaluated alert instances
Context: Separate real affected cohorts, split-by dimension cardinality, alert lifecycle and notification routing before raising limits, muting the action group or changing the KQL rule.
Evidence to confirm: Separate real affected cohorts, split-by dimension cardinality, alert lifecycle and notification routing before raising limits, muting the action group or changing the KQL rule.
Immediate checks: Export the deployed scheduled query rule and freeze a UTC window | Count rows, distinct dimension values and their combinations | Compare query combinations, active instances and delivered notifications | Canary stable dimensions through a complete fire-and-resolve cycle
Proposed action: Run the first checks in order: Export the deployed scheduled query rule and freeze a UTC window | Count rows, distinct dimension values and their combinations | Compare query combinations, active instances and delivered notifications | Canary stable dimensions through a complete fire-and-resolve cycle. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: One Azure Monitor log alert creates hundreds of independently evaluated alert instances
Initial hypothesis: Separate real affected cohorts, split-by dimension cardinality, alert lifecycle and notification routing before raising limits, muting the action group or changing the KQL rule.
Evidence used: Separate real affected cohorts, split-by dimension cardinality, alert lifecycle and notification routing before raising limits, muting the action group or changing the KQL rule.
Checks performed: Export the deployed scheduled query rule and freeze a UTC window | Count rows, distinct dimension values and their combinations | Compare query combinations, active instances and delivered notifications | Canary stable dimensions through a complete fire-and-resolve cycle
Decision / action: Run the first checks in order: Export the deployed scheduled query rule and freeze a UTC window | Count rows, distinct dimension values and their combinations | Compare query combinations, active instances and delivered notifications | Canary stable dimensions through a complete fire-and-resolve cycle. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
AI

An AI gateway returns 429 after a token-limit policy change

Open atlas

Separate APIM policy scope, counter isolation, gateway-local accounting, token estimation, backend throttling and client retries before raising the budget.

01

Evidence

Separate APIM policy scope, counter isolation, gateway-local accounting, token estimation, backend throttling and client retries before raising the budget.

02

First checks

  • Preserve one rejected request with gateway, region and Retry-After
  • Prove whether APIM or the model backend produced the 429
  • Verify the effective policy and a non-empty counter key
  • Canary the correction without resetting production counters
03

Bounded action

Run the first checks in order: Preserve one rejected request with gateway, region and Retry-After | Prove whether APIM or the model backend produced the 429 | Verify the effective policy and a non-empty counter key | Canary the correction without resetting production counters. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An AI gateway returns 429 after a token-limit policy change
Context: Separate APIM policy scope, counter isolation, gateway-local accounting, token estimation, backend throttling and client retries before raising the budget.
Evidence to confirm: Separate APIM policy scope, counter isolation, gateway-local accounting, token estimation, backend throttling and client retries before raising the budget.
Immediate checks: Preserve one rejected request with gateway, region and Retry-After | Prove whether APIM or the model backend produced the 429 | Verify the effective policy and a non-empty counter key | Canary the correction without resetting production counters
Proposed action: Run the first checks in order: Preserve one rejected request with gateway, region and Retry-After | Prove whether APIM or the model backend produced the 429 | Verify the effective policy and a non-empty counter key | Canary the correction without resetting production counters. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An AI gateway returns 429 after a token-limit policy change
Initial hypothesis: Separate APIM policy scope, counter isolation, gateway-local accounting, token estimation, backend throttling and client retries before raising the budget.
Evidence used: Separate APIM policy scope, counter isolation, gateway-local accounting, token estimation, backend throttling and client retries before raising the budget.
Checks performed: Preserve one rejected request with gateway, region and Retry-After | Prove whether APIM or the model backend produced the 429 | Verify the effective policy and a non-empty counter key | Canary the correction without resetting production counters
Decision / action: Run the first checks in order: Preserve one rejected request with gateway, region and Retry-After | Prove whether APIM or the model backend produced the 429 | Verify the effective policy and a non-empty counter key | Canary the correction without resetting production counters. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Networking

An Azure application path fails intermittently while manual connectivity tests remain healthy

Open atlas

Use Connection Monitor to preserve source-specific loss and latency, then correlate one failing interval with DNS, effective routes, security decisions, firewall evidence and service health before changing NSGs or UDRs.

01

Evidence

Use Connection Monitor to preserve source-specific loss and latency, then correlate one failing interval with DNS, effective routes, security decisions, firewall evidence and service health before changing NSGs or UDRs.

02

First checks

  • Freeze the exact source, FQDN, port and UTC window
  • Compare failed-check percentage and RTT per source
  • Run Connection Troubleshoot during the failing interval
  • Canary one bounded correction with a refusal control
03

Bounded action

Run the first checks in order: Freeze the exact source, FQDN, port and UTC window | Compare failed-check percentage and RTT per source | Run Connection Troubleshoot during the failing interval | Canary one bounded correction with a refusal control. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An Azure application path fails intermittently while manual connectivity tests remain healthy
Context: Use Connection Monitor to preserve source-specific loss and latency, then correlate one failing interval with DNS, effective routes, security decisions, firewall evidence and service health before changing NSGs or UDRs.
Evidence to confirm: Use Connection Monitor to preserve source-specific loss and latency, then correlate one failing interval with DNS, effective routes, security decisions, firewall evidence and service health before changing NSGs or UDRs.
Immediate checks: Freeze the exact source, FQDN, port and UTC window | Compare failed-check percentage and RTT per source | Run Connection Troubleshoot during the failing interval | Canary one bounded correction with a refusal control
Proposed action: Run the first checks in order: Freeze the exact source, FQDN, port and UTC window | Compare failed-check percentage and RTT per source | Run Connection Troubleshoot during the failing interval | Canary one bounded correction with a refusal control. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An Azure application path fails intermittently while manual connectivity tests remain healthy
Initial hypothesis: Use Connection Monitor to preserve source-specific loss and latency, then correlate one failing interval with DNS, effective routes, security decisions, firewall evidence and service health before changing NSGs or UDRs.
Evidence used: Use Connection Monitor to preserve source-specific loss and latency, then correlate one failing interval with DNS, effective routes, security decisions, firewall evidence and service health before changing NSGs or UDRs.
Checks performed: Freeze the exact source, FQDN, port and UTC window | Compare failed-check percentage and RTT per source | Run Connection Troubleshoot during the failing interval | Canary one bounded correction with a refusal control
Decision / action: Run the first checks in order: Freeze the exact source, FQDN, port and UTC window | Compare failed-check percentage and RTT per source | Run Connection Troubleshoot during the failing interval | Canary one bounded correction with a refusal control. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
AI

An AI agent proposes an unrequested production action after reading MCP tool output

Open atlas

Separate authenticated transport from field-level trust, preserve tool-result provenance, enforce write authorization outside the model and replay adversarial results before reopening production actions.

01

Evidence

Separate authenticated transport from field-level trust, preserve tool-result provenance, enforce write authorization outside the model and replay adversarial results before reopening production actions.

02

First checks

  • Freeze the user intent, raw tool result and proposed next call
  • Identify which untrusted field influenced the action
  • Verify approval and arguments in an external policy gate
  • Replay hostile results with writes disabled before a bounded canary
03

Bounded action

Run the first checks in order: Freeze the user intent, raw tool result and proposed next call | Identify which untrusted field influenced the action | Verify approval and arguments in an external policy gate | Replay hostile results with writes disabled before a bounded canary. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An AI agent proposes an unrequested production action after reading MCP tool output
Context: Separate authenticated transport from field-level trust, preserve tool-result provenance, enforce write authorization outside the model and replay adversarial results before reopening production actions.
Evidence to confirm: Separate authenticated transport from field-level trust, preserve tool-result provenance, enforce write authorization outside the model and replay adversarial results before reopening production actions.
Immediate checks: Freeze the user intent, raw tool result and proposed next call | Identify which untrusted field influenced the action | Verify approval and arguments in an external policy gate | Replay hostile results with writes disabled before a bounded canary
Proposed action: Run the first checks in order: Freeze the user intent, raw tool result and proposed next call | Identify which untrusted field influenced the action | Verify approval and arguments in an external policy gate | Replay hostile results with writes disabled before a bounded canary. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An AI agent proposes an unrequested production action after reading MCP tool output
Initial hypothesis: Separate authenticated transport from field-level trust, preserve tool-result provenance, enforce write authorization outside the model and replay adversarial results before reopening production actions.
Evidence used: Separate authenticated transport from field-level trust, preserve tool-result provenance, enforce write authorization outside the model and replay adversarial results before reopening production actions.
Checks performed: Freeze the user intent, raw tool result and proposed next call | Identify which untrusted field influenced the action | Verify approval and arguments in an external policy gate | Replay hostile results with writes disabled before a bounded canary
Decision / action: Run the first checks in order: Freeze the user intent, raw tool result and proposed next call | Identify which untrusted field influenced the action | Verify approval and arguments in an external policy gate | Replay hostile results with writes disabled before a bounded canary. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
AI

An agent stays available through model fallback but changes its tool decisions or refusal behavior

Open atlas

Separate endpoint availability from response-schema, tool, source, refusal, approval, state and trace compatibility before enabling automatic multi-model routing.

01

Evidence

Separate endpoint availability from response-schema, tool, source, refusal, approval, state and trace compatibility before enabling automatic multi-model routing.

02

First checks

  • Freeze both model attempts under one logical run ID
  • Validate structured output and normalized tool arguments
  • Reconcile any unknown primary result before fallback
  • Canary new read-only sessions with route-level traces
03

Bounded action

Run the first checks in order: Freeze both model attempts under one logical run ID | Validate structured output and normalized tool arguments | Reconcile any unknown primary result before fallback | Canary new read-only sessions with route-level traces. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An agent stays available through model fallback but changes its tool decisions or refusal behavior
Context: Separate endpoint availability from response-schema, tool, source, refusal, approval, state and trace compatibility before enabling automatic multi-model routing.
Evidence to confirm: Separate endpoint availability from response-schema, tool, source, refusal, approval, state and trace compatibility before enabling automatic multi-model routing.
Immediate checks: Freeze both model attempts under one logical run ID | Validate structured output and normalized tool arguments | Reconcile any unknown primary result before fallback | Canary new read-only sessions with route-level traces
Proposed action: Run the first checks in order: Freeze both model attempts under one logical run ID | Validate structured output and normalized tool arguments | Reconcile any unknown primary result before fallback | Canary new read-only sessions with route-level traces. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An agent stays available through model fallback but changes its tool decisions or refusal behavior
Initial hypothesis: Separate endpoint availability from response-schema, tool, source, refusal, approval, state and trace compatibility before enabling automatic multi-model routing.
Evidence used: Separate endpoint availability from response-schema, tool, source, refusal, approval, state and trace compatibility before enabling automatic multi-model routing.
Checks performed: Freeze both model attempts under one logical run ID | Validate structured output and normalized tool arguments | Reconcile any unknown primary result before fallback | Canary new read-only sessions with route-level traces
Decision / action: Run the first checks in order: Freeze both model attempts under one logical run ID | Validate structured output and normalized tool arguments | Reconcile any unknown primary result before fallback | Canary new read-only sessions with route-level traces. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Infrastructure

An AKS Horizontal Pod Autoscaler repeatedly scales up and down

Open atlas

Separate noisy or missing metrics, miscalibrated resource requests, pod warm-up, competing controllers and node capacity before raising replica bounds.

01

Evidence

Separate noisy or missing metrics, miscalibrated resource requests, pod warm-up, competing controllers and node capacity before raising replica bounds.

02

First checks

  • Freeze two complete scaling cycles
  • Compare the HPA status with the metric API
  • Measure requests, readiness and downstream capacity per replica
  • Canary one behavior change with explicit rollback
03

Bounded action

Run the first checks in order: Freeze two complete scaling cycles | Compare the HPA status with the metric API | Measure requests, readiness and downstream capacity per replica | Canary one behavior change with explicit rollback. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An AKS Horizontal Pod Autoscaler repeatedly scales up and down
Context: Separate noisy or missing metrics, miscalibrated resource requests, pod warm-up, competing controllers and node capacity before raising replica bounds.
Evidence to confirm: Separate noisy or missing metrics, miscalibrated resource requests, pod warm-up, competing controllers and node capacity before raising replica bounds.
Immediate checks: Freeze two complete scaling cycles | Compare the HPA status with the metric API | Measure requests, readiness and downstream capacity per replica | Canary one behavior change with explicit rollback
Proposed action: Run the first checks in order: Freeze two complete scaling cycles | Compare the HPA status with the metric API | Measure requests, readiness and downstream capacity per replica | Canary one behavior change with explicit rollback. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An AKS Horizontal Pod Autoscaler repeatedly scales up and down
Initial hypothesis: Separate noisy or missing metrics, miscalibrated resource requests, pod warm-up, competing controllers and node capacity before raising replica bounds.
Evidence used: Separate noisy or missing metrics, miscalibrated resource requests, pod warm-up, competing controllers and node capacity before raising replica bounds.
Checks performed: Freeze two complete scaling cycles | Compare the HPA status with the metric API | Measure requests, readiness and downstream capacity per replica | Canary one behavior change with explicit rollback
Decision / action: Run the first checks in order: Freeze two complete scaling cycles | Compare the HPA status with the metric API | Measure requests, readiness and downstream capacity per replica | Canary one behavior change with explicit rollback. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Cloud

An Azure deployment fails because a resource or inherited scope is locked

Open atlas

Separate management locks, RBAC, Policy and deny assignments, identify the exact destructive operation, then bound any unlock window before rerun or rollback.

01

Evidence

Separate management locks, RBAC, Policy and deny assignments, identify the exact destructive operation, then bound any unlock window before rerun or rollback.

02

First checks

  • Freeze the failed operation, target resource and correlation ID
  • Reconstruct inherited CanNotDelete and ReadOnly locks
  • Confirm whether the plan really requires a delete
  • Recreate and verify the lock before closing the change
03

Bounded action

Run the first checks in order: Freeze the failed operation, target resource and correlation ID | Reconstruct inherited CanNotDelete and ReadOnly locks | Confirm whether the plan really requires a delete | Recreate and verify the lock before closing the change. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An Azure deployment fails because a resource or inherited scope is locked
Context: Separate management locks, RBAC, Policy and deny assignments, identify the exact destructive operation, then bound any unlock window before rerun or rollback.
Evidence to confirm: Separate management locks, RBAC, Policy and deny assignments, identify the exact destructive operation, then bound any unlock window before rerun or rollback.
Immediate checks: Freeze the failed operation, target resource and correlation ID | Reconstruct inherited CanNotDelete and ReadOnly locks | Confirm whether the plan really requires a delete | Recreate and verify the lock before closing the change
Proposed action: Run the first checks in order: Freeze the failed operation, target resource and correlation ID | Reconstruct inherited CanNotDelete and ReadOnly locks | Confirm whether the plan really requires a delete | Recreate and verify the lock before closing the change. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An Azure deployment fails because a resource or inherited scope is locked
Initial hypothesis: Separate management locks, RBAC, Policy and deny assignments, identify the exact destructive operation, then bound any unlock window before rerun or rollback.
Evidence used: Separate management locks, RBAC, Policy and deny assignments, identify the exact destructive operation, then bound any unlock window before rerun or rollback.
Checks performed: Freeze the failed operation, target resource and correlation ID | Reconstruct inherited CanNotDelete and ReadOnly locks | Confirm whether the plan really requires a delete | Recreate and verify the lock before closing the change
Decision / action: Run the first checks in order: Freeze the failed operation, target resource and correlation ID | Reconstruct inherited CanNotDelete and ReadOnly locks | Confirm whether the plan really requires a delete | Recreate and verify the lock before closing the change. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Networking

AKS pods intermittently time out to public dependencies

Open atlas

Separate DNS, the remote dependency, network policy, node connection pressure and the effective Azure egress device before adding SNAT capacity or changing the cluster outbound path.

01

Evidence

Separate DNS, the remote dependency, network policy, node connection pressure and the effective Azure egress device before adding SNAT capacity or changing the cluster outbound path.

02

First checks

  • Capture the failing pod, node and destination tuple
  • Confirm networkProfile.outboundType and the source IP seen downstream
  • Correlate per-node SNAT metrics with connection churn
  • Canary connection reuse or egress capacity before rollout
03

Bounded action

Run the first checks in order: Capture the failing pod, node and destination tuple | Confirm networkProfile.outboundType and the source IP seen downstream | Correlate per-node SNAT metrics with connection churn | Canary connection reuse or egress capacity before rollout. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] AKS pods intermittently time out to public dependencies
Context: Separate DNS, the remote dependency, network policy, node connection pressure and the effective Azure egress device before adding SNAT capacity or changing the cluster outbound path.
Evidence to confirm: Separate DNS, the remote dependency, network policy, node connection pressure and the effective Azure egress device before adding SNAT capacity or changing the cluster outbound path.
Immediate checks: Capture the failing pod, node and destination tuple | Confirm networkProfile.outboundType and the source IP seen downstream | Correlate per-node SNAT metrics with connection churn | Canary connection reuse or egress capacity before rollout
Proposed action: Run the first checks in order: Capture the failing pod, node and destination tuple | Confirm networkProfile.outboundType and the source IP seen downstream | Correlate per-node SNAT metrics with connection churn | Canary connection reuse or egress capacity before rollout. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: AKS pods intermittently time out to public dependencies
Initial hypothesis: Separate DNS, the remote dependency, network policy, node connection pressure and the effective Azure egress device before adding SNAT capacity or changing the cluster outbound path.
Evidence used: Separate DNS, the remote dependency, network policy, node connection pressure and the effective Azure egress device before adding SNAT capacity or changing the cluster outbound path.
Checks performed: Capture the failing pod, node and destination tuple | Confirm networkProfile.outboundType and the source IP seen downstream | Correlate per-node SNAT metrics with connection churn | Canary connection reuse or egress capacity before rollout
Decision / action: Run the first checks in order: Capture the failing pod, node and destination tuple | Confirm networkProfile.outboundType and the source IP seen downstream | Correlate per-node SNAT metrics with connection churn | Canary connection reuse or egress capacity before rollout. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Automation

A Microsoft Sentinel rule misses a target or automates a response from stale watchlist data

Open atlas

Separate source freshness, snapshot completeness, bulk-update deletion semantics, SearchKey normalization and KQL join behavior before allowing a watchlist to drive incidents or playbooks.

01

Evidence

Separate source freshness, snapshot completeness, bulk-update deletion semantics, SearchKey normalization and KQL join behavior before allowing a watchlist to drive incidents or playbooks.

02

First checks

  • Prove the alias and source snapshot independently of TimeGenerated
  • Compare added, modified and removed keys with the active version
  • Reject blank, duplicate or expired SearchKey values
  • Shadow the analytics rule before enabling the playbook
03

Bounded action

Run the first checks in order: Prove the alias and source snapshot independently of TimeGenerated | Compare added, modified and removed keys with the active version | Reject blank, duplicate or expired SearchKey values | Shadow the analytics rule before enabling the playbook. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] A Microsoft Sentinel rule misses a target or automates a response from stale watchlist data
Context: Separate source freshness, snapshot completeness, bulk-update deletion semantics, SearchKey normalization and KQL join behavior before allowing a watchlist to drive incidents or playbooks.
Evidence to confirm: Separate source freshness, snapshot completeness, bulk-update deletion semantics, SearchKey normalization and KQL join behavior before allowing a watchlist to drive incidents or playbooks.
Immediate checks: Prove the alias and source snapshot independently of TimeGenerated | Compare added, modified and removed keys with the active version | Reject blank, duplicate or expired SearchKey values | Shadow the analytics rule before enabling the playbook
Proposed action: Run the first checks in order: Prove the alias and source snapshot independently of TimeGenerated | Compare added, modified and removed keys with the active version | Reject blank, duplicate or expired SearchKey values | Shadow the analytics rule before enabling the playbook. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: A Microsoft Sentinel rule misses a target or automates a response from stale watchlist data
Initial hypothesis: Separate source freshness, snapshot completeness, bulk-update deletion semantics, SearchKey normalization and KQL join behavior before allowing a watchlist to drive incidents or playbooks.
Evidence used: Separate source freshness, snapshot completeness, bulk-update deletion semantics, SearchKey normalization and KQL join behavior before allowing a watchlist to drive incidents or playbooks.
Checks performed: Prove the alias and source snapshot independently of TimeGenerated | Compare added, modified and removed keys with the active version | Reject blank, duplicate or expired SearchKey values | Shadow the analytics rule before enabling the playbook
Decision / action: Run the first checks in order: Prove the alias and source snapshot independently of TimeGenerated | Compare added, modified and removed keys with the active version | Reject blank, duplicate or expired SearchKey values | Shadow the analytics rule before enabling the playbook. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Infrastructure

A self-hosted Azure DevOps agent shows signs of compromise but can still receive production jobs

Open atlas

Stop scheduling without erasing evidence, reconstruct affected runs, map exposed identities and validate a clean replacement before reopening the pool.

01

Evidence

Stop scheduling without erasing evidence, reconstruct affected runs, map exposed identities and validate a clean replacement before reopening the pool.

02

First checks

  • Quarantine the agent and pause production entry points
  • Freeze run IDs, commits, agent diagnostics and audit events
  • Map credentials and downstream targets used in the incident window
  • Canary a clean replacement without protected resources
03

Bounded action

Run the first checks in order: Quarantine the agent and pause production entry points | Freeze run IDs, commits, agent diagnostics and audit events | Map credentials and downstream targets used in the incident window | Canary a clean replacement without protected resources. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] A self-hosted Azure DevOps agent shows signs of compromise but can still receive production jobs
Context: Stop scheduling without erasing evidence, reconstruct affected runs, map exposed identities and validate a clean replacement before reopening the pool.
Evidence to confirm: Stop scheduling without erasing evidence, reconstruct affected runs, map exposed identities and validate a clean replacement before reopening the pool.
Immediate checks: Quarantine the agent and pause production entry points | Freeze run IDs, commits, agent diagnostics and audit events | Map credentials and downstream targets used in the incident window | Canary a clean replacement without protected resources
Proposed action: Run the first checks in order: Quarantine the agent and pause production entry points | Freeze run IDs, commits, agent diagnostics and audit events | Map credentials and downstream targets used in the incident window | Canary a clean replacement without protected resources. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: A self-hosted Azure DevOps agent shows signs of compromise but can still receive production jobs
Initial hypothesis: Stop scheduling without erasing evidence, reconstruct affected runs, map exposed identities and validate a clean replacement before reopening the pool.
Evidence used: Stop scheduling without erasing evidence, reconstruct affected runs, map exposed identities and validate a clean replacement before reopening the pool.
Checks performed: Quarantine the agent and pause production entry points | Freeze run IDs, commits, agent diagnostics and audit events | Map credentials and downstream targets used in the incident window | Canary a clean replacement without protected resources
Decision / action: Run the first checks in order: Quarantine the agent and pause production entry points | Freeze run IDs, commits, agent diagnostics and audit events | Map credentials and downstream targets used in the incident window | Canary a clean replacement without protected resources. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Networking

An App Service plan stays online but scaling fails on a VNet Integration subnet

Open atlas

Separate address exhaustion, delayed release, delegation and outbound path failures; calculate worker cohort overlap before retrying scale or migrating the integration.

01

Evidence

Separate address exhaustion, delayed release, delegation and outbound path failures; calculate worker cohort overlap before retrying scale or migrating the integration.

02

First checks

  • Freeze the failed operation and current instance count
  • Inventory every app and plan joined to the subnet
  • Calculate usable addresses and transient worker overlap
  • Canary a larger delegated subnet with explicit rollback
03

Bounded action

Run the first checks in order: Freeze the failed operation and current instance count | Inventory every app and plan joined to the subnet | Calculate usable addresses and transient worker overlap | Canary a larger delegated subnet with explicit rollback. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An App Service plan stays online but scaling fails on a VNet Integration subnet
Context: Separate address exhaustion, delayed release, delegation and outbound path failures; calculate worker cohort overlap before retrying scale or migrating the integration.
Evidence to confirm: Separate address exhaustion, delayed release, delegation and outbound path failures; calculate worker cohort overlap before retrying scale or migrating the integration.
Immediate checks: Freeze the failed operation and current instance count | Inventory every app and plan joined to the subnet | Calculate usable addresses and transient worker overlap | Canary a larger delegated subnet with explicit rollback
Proposed action: Run the first checks in order: Freeze the failed operation and current instance count | Inventory every app and plan joined to the subnet | Calculate usable addresses and transient worker overlap | Canary a larger delegated subnet with explicit rollback. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An App Service plan stays online but scaling fails on a VNet Integration subnet
Initial hypothesis: Separate address exhaustion, delayed release, delegation and outbound path failures; calculate worker cohort overlap before retrying scale or migrating the integration.
Evidence used: Separate address exhaustion, delayed release, delegation and outbound path failures; calculate worker cohort overlap before retrying scale or migrating the integration.
Checks performed: Freeze the failed operation and current instance count | Inventory every app and plan joined to the subnet | Calculate usable addresses and transient worker overlap | Canary a larger delegated subnet with explicit rollback
Decision / action: Run the first checks in order: Freeze the failed operation and current instance count | Inventory every app and plan joined to the subnet | Calculate usable addresses and transient worker overlap | Canary a larger delegated subnet with explicit rollback. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Cloud

Azure WAF appears to enforce different policy behavior after a deployment

Open atlas

Separate ARM rollout state, gateway, listener and path-rule associations, deployed rule priority and real request variance before waiting or restoring the whole policy.

01

Evidence

Separate ARM rollout state, gateway, listener and path-rule associations, deployed rule priority and real request variance before waiting or restoring the whole policy.

02

First checks

  • Freeze one accepted and one blocked request
  • Read policy and gateway provisioning states
  • Inventory policy resource IDs at every association scope
  • Replay an allowed and denied canary after correction or rollback
03

Bounded action

Run the first checks in order: Freeze one accepted and one blocked request | Read policy and gateway provisioning states | Inventory policy resource IDs at every association scope | Replay an allowed and denied canary after correction or rollback. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] Azure WAF appears to enforce different policy behavior after a deployment
Context: Separate ARM rollout state, gateway, listener and path-rule associations, deployed rule priority and real request variance before waiting or restoring the whole policy.
Evidence to confirm: Separate ARM rollout state, gateway, listener and path-rule associations, deployed rule priority and real request variance before waiting or restoring the whole policy.
Immediate checks: Freeze one accepted and one blocked request | Read policy and gateway provisioning states | Inventory policy resource IDs at every association scope | Replay an allowed and denied canary after correction or rollback
Proposed action: Run the first checks in order: Freeze one accepted and one blocked request | Read policy and gateway provisioning states | Inventory policy resource IDs at every association scope | Replay an allowed and denied canary after correction or rollback. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: Azure WAF appears to enforce different policy behavior after a deployment
Initial hypothesis: Separate ARM rollout state, gateway, listener and path-rule associations, deployed rule priority and real request variance before waiting or restoring the whole policy.
Evidence used: Separate ARM rollout state, gateway, listener and path-rule associations, deployed rule priority and real request variance before waiting or restoring the whole policy.
Checks performed: Freeze one accepted and one blocked request | Read policy and gateway provisioning states | Inventory policy resource IDs at every association scope | Replay an allowed and denied canary after correction or rollback
Decision / action: Run the first checks in order: Freeze one accepted and one blocked request | Read policy and gateway provisioning states | Inventory policy resource IDs at every association scope | Replay an allowed and denied canary after correction or rollback. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Automation

A previously allowed Azure deployment is denied after its Policy exemption expires

Open atlas

Tie the first denial to expiresOn, the exact assignment, initiative references and current resource state before renewing the exemption or disabling enforcement.

01

Evidence

Tie the first denial to expiresOn, the exact assignment, initiative references and current resource state before renewing the exemption or disabling enforcement.

02

First checks

  • Freeze the first denied run and correlation ID
  • Read the exemption at its exact scope and compare expiresOn in UTC
  • Match assignment and initiative reference IDs
  • Canary the same artifact with an out-of-scope denial control
03

Bounded action

Run the first checks in order: Freeze the first denied run and correlation ID | Read the exemption at its exact scope and compare expiresOn in UTC | Match assignment and initiative reference IDs | Canary the same artifact with an out-of-scope denial control. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] A previously allowed Azure deployment is denied after its Policy exemption expires
Context: Tie the first denial to expiresOn, the exact assignment, initiative references and current resource state before renewing the exemption or disabling enforcement.
Evidence to confirm: Tie the first denial to expiresOn, the exact assignment, initiative references and current resource state before renewing the exemption or disabling enforcement.
Immediate checks: Freeze the first denied run and correlation ID | Read the exemption at its exact scope and compare expiresOn in UTC | Match assignment and initiative reference IDs | Canary the same artifact with an out-of-scope denial control
Proposed action: Run the first checks in order: Freeze the first denied run and correlation ID | Read the exemption at its exact scope and compare expiresOn in UTC | Match assignment and initiative reference IDs | Canary the same artifact with an out-of-scope denial control. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: A previously allowed Azure deployment is denied after its Policy exemption expires
Initial hypothesis: Tie the first denial to expiresOn, the exact assignment, initiative references and current resource state before renewing the exemption or disabling enforcement.
Evidence used: Tie the first denial to expiresOn, the exact assignment, initiative references and current resource state before renewing the exemption or disabling enforcement.
Checks performed: Freeze the first denied run and correlation ID | Read the exemption at its exact scope and compare expiresOn in UTC | Match assignment and initiative reference IDs | Canary the same artifact with an out-of-scope denial control
Decision / action: Run the first checks in order: Freeze the first denied run and correlation ID | Read the exemption at its exact scope and compare expiresOn in UTC | Match assignment and initiative reference IDs | Canary the same artifact with an out-of-scope denial control. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
AI

An AI agent trace contains sensitive data after an instrumentation change

Open atlas

Tie release, span, unsafe field, export route and destination together before disabling observability; contain the narrow path, handle historical copies and canary redaction.

01

Evidence

Tie release, span, unsafe field, export route and destination together before disabling observability; contain the narrow path, handle historical copies and canary redaction.

02

First checks

  • Freeze trace metadata without copying the raw value
  • Map every collector, exporter and destination copy
  • Disable the unsafe enrichment while keeping minimum events
  • Replay synthetic markers through every input path
03

Bounded action

Run the first checks in order: Freeze trace metadata without copying the raw value | Map every collector, exporter and destination copy | Disable the unsafe enrichment while keeping minimum events | Replay synthetic markers through every input path. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An AI agent trace contains sensitive data after an instrumentation change
Context: Tie release, span, unsafe field, export route and destination together before disabling observability; contain the narrow path, handle historical copies and canary redaction.
Evidence to confirm: Tie release, span, unsafe field, export route and destination together before disabling observability; contain the narrow path, handle historical copies and canary redaction.
Immediate checks: Freeze trace metadata without copying the raw value | Map every collector, exporter and destination copy | Disable the unsafe enrichment while keeping minimum events | Replay synthetic markers through every input path
Proposed action: Run the first checks in order: Freeze trace metadata without copying the raw value | Map every collector, exporter and destination copy | Disable the unsafe enrichment while keeping minimum events | Replay synthetic markers through every input path. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An AI agent trace contains sensitive data after an instrumentation change
Initial hypothesis: Tie release, span, unsafe field, export route and destination together before disabling observability; contain the narrow path, handle historical copies and canary redaction.
Evidence used: Tie release, span, unsafe field, export route and destination together before disabling observability; contain the narrow path, handle historical copies and canary redaction.
Checks performed: Freeze trace metadata without copying the raw value | Map every collector, exporter and destination copy | Disable the unsafe enrichment while keeping minimum events | Replay synthetic markers through every input path
Decision / action: Run the first checks in order: Freeze trace metadata without copying the raw value | Map every collector, exporter and destination copy | Disable the unsafe enrichment while keeping minimum events | Replay synthetic markers through every input path. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
Automation

A canceled Azure DevOps deployment leaves the production target in an uncertain partial state

Open atlas

Reconcile the run, agent, Azure control plane and application before rerunning; prove each checkpoint, artifact and side effect, then choose resume, compensation or rollback.

01

Evidence

Reconcile the run, agent, Azure control plane and application before rerunning; prove each checkpoint, artifact and side effect, then choose resume, compensation or rollback.

02

First checks

  • Freeze every writer targeting the environment
  • Pin the run, commit and artifact digest
  • Read target-side operations across the cancellation window
  • Block rerun while a critical checkpoint remains unknown
03

Bounded action

Run the first checks in order: Freeze every writer targeting the environment | Pin the run, commit and artifact digest | Read target-side operations across the cancellation window | Block rerun while a critical checkpoint remains unknown. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] A canceled Azure DevOps deployment leaves the production target in an uncertain partial state
Context: Reconcile the run, agent, Azure control plane and application before rerunning; prove each checkpoint, artifact and side effect, then choose resume, compensation or rollback.
Evidence to confirm: Reconcile the run, agent, Azure control plane and application before rerunning; prove each checkpoint, artifact and side effect, then choose resume, compensation or rollback.
Immediate checks: Freeze every writer targeting the environment | Pin the run, commit and artifact digest | Read target-side operations across the cancellation window | Block rerun while a critical checkpoint remains unknown
Proposed action: Run the first checks in order: Freeze every writer targeting the environment | Pin the run, commit and artifact digest | Read target-side operations across the cancellation window | Block rerun while a critical checkpoint remains unknown. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: A canceled Azure DevOps deployment leaves the production target in an uncertain partial state
Initial hypothesis: Reconcile the run, agent, Azure control plane and application before rerunning; prove each checkpoint, artifact and side effect, then choose resume, compensation or rollback.
Evidence used: Reconcile the run, agent, Azure control plane and application before rerunning; prove each checkpoint, artifact and side effect, then choose resume, compensation or rollback.
Checks performed: Freeze every writer targeting the environment | Pin the run, commit and artifact digest | Read target-side operations across the cancellation window | Block rerun while a critical checkpoint remains unknown
Decision / action: Run the first checks in order: Freeze every writer targeting the environment | Pin the run, commit and artifact digest | Read target-side operations across the cancellation window | Block rerun while a critical checkpoint remains unknown. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.
AI

An Azure AI Search indexer succeeds partially and the agent corpus becomes incomplete

Open atlas

Separate source access, incremental tracking, skillset failures and target-index writes before rerunning selected documents or clearing the high-water mark.

01

Evidence

Separate source access, incremental tracking, skillset failures and target-index writes before rerunning selected documents or clearing the high-water mark.

02

First checks

  • Freeze the execution history and tracking states
  • Reconcile failed source IDs with index keys and versions
  • Group errors by stage, code and skill
  • Canary targeted recovery before any full reset
03

Bounded action

Run the first checks in order: Freeze the execution history and tracking states | Reconcile failed source IDs with index keys and versions | Group errors by stage, code and skill | Canary targeted recovery before any full reset. Open the linked notes before changing production.

04

Rollback

Stop the change, restore the last known safe state and keep the captured evidence for comparison.

copy packs Incident exports
Short handoff
[Incident] An Azure AI Search indexer succeeds partially and the agent corpus becomes incomplete
Context: Separate source access, incremental tracking, skillset failures and target-index writes before rerunning selected documents or clearing the high-water mark.
Evidence to confirm: Separate source access, incremental tracking, skillset failures and target-index writes before rerunning selected documents or clearing the high-water mark.
Immediate checks: Freeze the execution history and tracking states | Reconcile failed source IDs with index keys and versions | Group errors by stage, code and skill | Canary targeted recovery before any full reset
Proposed action: Run the first checks in order: Freeze the execution history and tracking states | Reconcile failed source IDs with index keys and versions | Group errors by stage, code and skill | Canary targeted recovery before any full reset. Open the linked notes before changing production.
Rollback: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Post-incident review
Handled symptom: An Azure AI Search indexer succeeds partially and the agent corpus becomes incomplete
Initial hypothesis: Separate source access, incremental tracking, skillset failures and target-index writes before rerunning selected documents or clearing the high-water mark.
Evidence used: Separate source access, incremental tracking, skillset failures and target-index writes before rerunning selected documents or clearing the high-water mark.
Checks performed: Freeze the execution history and tracking states | Reconcile failed source IDs with index keys and versions | Group errors by stage, code and skill | Canary targeted recovery before any full reset
Decision / action: Run the first checks in order: Freeze the execution history and tracking states | Reconcile failed source IDs with index keys and versions | Group errors by stage, code and skill | Canary targeted recovery before any full reset. Open the linked notes before changing production.
Rollback plan: Stop the change, restore the last known safe state and keep the captured evidence for comparison.
To improve: detection, runbook, guardrail, ownership and communication delay.