Separate backend health, DNS resolution, TLS settings and private network reachability before changing the application.
01 Evidence
Backend health state, probe result, gateway diagnostic logs, DNS answer from the gateway path and TLS/SNI settings.
02 First checks
- Check backend health state
- Resolve backend name from the gateway path
- Validate TLS/SNI and probe configuration
03 Bounded action
Change only the failing boundary: probe, backend FQDN, certificate binding or route. Retest the same request path after each change.
04 Rollback
Restore previous probe/backend settings and keep the captured failing timestamp for comparison.
Cloud
Internal APIM returns an error on a private API
Open atlas Correlate Application Gateway/WAF and APIM logs, then separate DNS, TLS, policy, identity and private backend reachability before changing policies or opening access.
01 Evidence
Correlate Application Gateway/WAF and APIM logs, then separate DNS, TLS, policy, identity and private backend reachability before changing policies or opening access.
02 First checks
- Check whether WAF blocked the request
- Confirm APIM received the same path
- Validate backend DNS and TLS from the APIM path
- Replay with a correlation ID
03 Bounded action
Run the first checks in order: Check whether WAF blocked the request | Confirm APIM received the same path | Validate backend DNS and TLS from the APIM path | Replay with a correlation ID. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Networking
Private Endpoint name still resolves publicly
Open atlas Confirm the CNAME chain, Private DNS Zone association and hybrid forwarding from the consuming network.
01 Evidence
nslookup from the workload subnet, CNAME chain, private DNS zone links, resolver forwarding path and cached answers.
02 First checks
- Run nslookup from the workload network
- Check privatelink CNAME
- Verify Private DNS Zone links and forwarders
03 Bounded action
Fix zone association or forwarding first, then clear caches and retest from the consuming network.
04 Rollback
Restore the previous link or forwarder and document the public/private answer difference.
Cloud
Azure Storage private endpoint returns 403, times out or produces no request logs
Open atlas Separate Storage subresource DNS, Private Endpoint approval, firewall rules, runtime identity and Storage logs before opening public access or broadening RBAC.
01 Evidence
Separate Storage subresource DNS, Private Endpoint approval, firewall rules, runtime identity and Storage logs before opening public access or broadening RBAC.
02 First checks
- Resolve the exact Storage subresource from the workload network
- Check Private Endpoint status and private DNS zone group
- Replay with a client request ID
- Correlate Storage logs for 403, caller IP and requester identity
03 Bounded action
Run the first checks in order: Resolve the exact Storage subresource from the workload network | Check Private Endpoint status and private DNS zone group | Replay with a client request ID | Correlate Storage logs for 403, caller IP and requester identity. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Cloud
Azure SQL private endpoint returns timeouts, firewall errors or no SQL logs
Open atlas Separate SQL private DNS, Private Endpoint state, public access, firewall rules, runtime identity and SQL diagnostics before changing schema, code or broad permissions.
01 Evidence
Separate SQL private DNS, Private Endpoint state, public access, firewall rules, runtime identity and SQL diagnostics before changing schema, code or broad permissions.
02 First checks
- Resolve the SQL FQDN from the workload network
- Check Private Endpoint status and privatelink.database.windows.net records
- Replay with the real runtime identity
- Correlate SQL diagnostics for firewall and login errors
03 Bounded action
Run the first checks in order: Resolve the SQL FQDN from the workload network | Check Private Endpoint status and privatelink.database.windows.net records | Replay with the real runtime identity | Correlate SQL diagnostics for firewall and login errors. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Cloud
Azure Service Bus private endpoint times out, denies access or lets backlog grow
Open atlas Separate Service Bus private DNS, Private Endpoint state, public access, managed identity or SAS, queue metrics and processing logs before touching queues or redeploying consumers.
01 Evidence
Separate Service Bus private DNS, Private Endpoint state, public access, managed identity or SAS, queue metrics and processing logs before touching queues or redeploying consumers.
02 First checks
- Resolve the Service Bus FQDN from the workload network
- Check Private Endpoint approval and public network access
- Verify the real sender or receiver identity
- Correlate backlog, dead-letter and Service Bus errors
03 Bounded action
Run the first checks in order: Resolve the Service Bus FQDN from the workload network | Check Private Endpoint approval and public network access | Verify the real sender or receiver identity | Correlate backlog, dead-letter and Service Bus errors. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Cloud
A synthetic probe fails on an Azure private path
Open atlas Separate DNS, TLS, Application Gateway health, WAF blocks and runner network before changing routing or application code.
01 Evidence
Separate DNS, TLS, Application Gateway health, WAF blocks and runner network before changing routing or application code.
02 First checks
- Resolve the hostname from the probe network
- Check TLS/SNI with the real hostname
- Correlate probe run with WAF and gateway logs
03 Bounded action
Run the first checks in order: Resolve the hostname from the probe network | Check TLS/SNI with the real hostname | Correlate probe run with WAF and gateway logs. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Networking
Azure traffic leaves through one path and returns through another
Open atlas Compare DNS target, effective routes, route table associations, firewall evidence, NAT identity and return path before changing UDRs or bypassing inspection.
01 Evidence
Compare DNS target, effective routes, route table associations, firewall evidence, NAT identity and return path before changing UDRs or bypassing inspection.
02 First checks
- Resolve the destination from the source path
- Compare effective routes on source and destination NICs
- Check firewall or appliance logs for both directions
- Confirm NAT or outbound source identity
03 Bounded action
Run the first checks in order: Resolve the destination from the source path | Compare effective routes on source and destination NICs | Check firewall or appliance logs for both directions | Confirm NAT or outbound source identity. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Cloud
Azure Container Apps private ingress fails or reaches the wrong revision
Open atlas Separate private DNS, Application Gateway handoff, Container Apps ingress mode, revision traffic and console logs before rolling back or changing traffic weights.
01 Evidence
Separate private DNS, Application Gateway handoff, Container Apps ingress mode, revision traffic and console logs before rolling back or changing traffic weights.
02 First checks
- Resolve the hostname from the caller network
- Check ingress target port and active revisions
- Correlate system and console logs
03 Bounded action
Run the first checks in order: Resolve the hostname from the caller network | Check ingress target port and active revisions | Correlate system and console logs. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Cloud
AKS private ingress returns 502 or reaches no service endpoints
Open atlas Separate private DNS, Application Gateway health, ingress controller routing, Kubernetes service selectors, endpoint slices and pod readiness before rolling back a deployment.
01 Evidence
Separate private DNS, Application Gateway health, ingress controller routing, Kubernetes service selectors, endpoint slices and pod readiness before rolling back a deployment.
02 First checks
- Resolve the hostname from the caller network
- Check Application Gateway backend health and host header
- Verify ingress, service and endpoint slices
- Correlate controller and application logs
03 Bounded action
Run the first checks in order: Resolve the hostname from the caller network | Check Application Gateway backend health and host header | Verify ingress, service and endpoint slices | Correlate controller and application logs. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Infrastructure
An AKS node pool upgrade is blocked while draining a pod protected by a PDB
Open atlas Separate unavailable replicas, readiness failures, pod placement, surge capacity and an over-constrained PodDisruptionBudget before deleting the safeguard or forcing the upgrade.
01 Evidence
Separate unavailable replicas, readiness failures, pod placement, surge capacity and an over-constrained PodDisruptionBudget before deleting the safeguard or forcing the upgrade.
02 First checks
- Capture the blocked node and eviction events
- Inspect PDB disruptionsAllowed and selectors
- Verify replacement capacity, replica readiness and topology
- Resume only with positive disruption headroom
03 Bounded action
Run the first checks in order: Capture the blocked node and eviction events | Inspect PDB disruptionsAllowed and selectors | Verify replacement capacity, replica readiness and topology | Resume only with positive disruption headroom. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Cloud
Azure Functions private HTTP endpoint returns 403, 503 or no request logs
Open atlas Separate private DNS, Private Endpoint reachability, access restrictions, Functions runtime state, private storage and Application Insights evidence before redeploying code or opening public access.
01 Evidence
Separate private DNS, Private Endpoint reachability, access restrictions, Functions runtime state, private storage and Application Insights evidence before redeploying code or opening public access.
02 First checks
- Resolve the hostname from the caller network
- Replay with a correlation ID
- Check Function App access restrictions and Private Endpoint status
- Correlate requests, traces and exceptions
03 Bounded action
Run the first checks in order: Resolve the hostname from the caller network | Replay with a correlation ID | Check Function App access restrictions and Private Endpoint status | Correlate requests, traces and exceptions. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Cloud
Azure WAF blocks a legitimate request
Open atlas Start from blocked requests, rule ID and URI before deciding between exclusion, custom rule or application fix.
01 Evidence
Blocked URI, ruleId, match variable, client IP, hostname, request ID and exact time window.
02 First checks
- List blocked URIs in KQL
- Identify ruleId and match field
- Validate false-positive scope
03 Bounded action
Create the smallest exclusion or custom rule that covers the false positive without disabling the rule globally.
04 Rollback
Remove the exclusion/custom rule and verify expected blocking returns for the same rule family.
Automation
Terraform state lock is stuck
Open atlas Prove that no apply is still running before using force-unlock, then restart with a clean plan.
01 Evidence
Lock ID, lock owner, CI run, backend target, pending plan and whether an apply is still active.
02 First checks
- Identify lock owner
- Check CI job status
- Run plan after unlock
03 Bounded action
Unlock only after proving no apply is running, then start with a fresh plan before any apply.
04 Rollback
Return to previous commit or restore the last validated state version if drift was introduced.
Infrastructure
Azure Monitor fires an alert storm after deployment
Open atlas Separate real service impact, noisy dimensions, threshold drift, action group behavior and rollback before silencing notifications or changing rules.
01 Evidence
Separate real service impact, noisy dimensions, threshold drift, action group behavior and rollback before silencing notifications or changing rules.
02 First checks
- Group fired alerts by rule and target
- Compare first alert time with the deployment window
- Check service logs before changing thresholds
- Keep one user-symptom alert active
03 Bounded action
Run the first checks in order: Group fired alerts by rule and target | Compare first alert time with the deployment window | Check service logs before changing thresholds | Keep one user-symptom alert active. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Infrastructure
A secret rotation, federated identity or managed identity change breaks an application or pipeline consumer
Open atlas Separate preparation, cutover, revocation, workload identity federation and managed identity diagnostics; validate the real execution identity, private path and authentication errors before deleting the old value or broadening access.
01 Evidence
Separate preparation, cutover, revocation, workload identity federation and managed identity diagnostics; validate the real execution identity, private path and authentication errors before deleting the old value or broadening access.
02 First checks
- List real consumers
- Verify the runtime identity and vault read access
- Check OIDC claims or private DNS depending on the path
- Watch 401/403/500, sign-in failures or Key Vault denials
03 Bounded action
Run the first checks in order: List real consumers | Verify the runtime identity and vault read access | Check OIDC claims or private DNS depending on the path | Watch 401/403/500, sign-in failures or Key Vault denials. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Automation
An Azure Functions poison queue keeps growing
Open atlas Separate contract failures, transient dependencies, identity, resource pressure and partial side effects before replaying Storage Queue messages.
01 Evidence
Separate contract failures, transient dependencies, identity, resource pressure and partial side effects before replaying Storage Queue messages.
02 First checks
- Peek a redacted sample without consuming messages
- Correlate message IDs with invocation and dependency logs
- Verify downstream state and idempotency before replay
03 Bounded action
Run the first checks in order: Peek a redacted sample without consuming messages | Correlate message IDs with invocation and dependency logs | Verify downstream state and idempotency before replay. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Cloud
An Azure Service Bus dead-letter queue keeps growing
Open atlas Classify dead-letter reasons, correlate consumer and downstream evidence, verify idempotency, then use a manifest, canary and stop conditions before replay.
01 Evidence
Classify dead-letter reasons, correlate consumer and downstream evidence, verify idempotency, then use a manifest, canary and stop conditions before replay.
02 First checks
- Record the exact queue or subscription and DLQ growth window
- Peek a redacted sample grouped by dead-letter reason
- Check consumer errors, lock loss and downstream state
- Prove idempotency before a one-message canary
03 Bounded action
Run the first checks in order: Record the exact queue or subscription and DLQ growth window | Peek a redacted sample grouped by dead-letter reason | Check consumer errors, lock loss and downstream state | Prove idempotency before a one-message canary. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Automation
An Azure Policy remediation changes unexpected production resources
Open atlas Pin the definition and assignment, materialize the eligible target set, verify the managed identity, then use a canary and bounded batches before expanding or compensating.
01 Evidence
Pin the definition and assignment, materialize the eligible target set, verify the managed identity, then use a canary and bounded batches before expanding or compensating.
02 First checks
- Identify the exact definition reference and assignment parameters
- Export non-compliant target resource IDs
- Review the assignment identity and effective RBAC scope
- Correlate remediation deployments with Activity Log
03 Bounded action
Run the first checks in order: Identify the exact definition reference and assignment parameters | Export non-compliant target resource IDs | Review the assignment identity and effective RBAC scope | Correlate remediation deployments with Activity Log. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Automation
An automation entry point behaves like a remote console
Open atlas Bound inputs, templates and repository structure before exposing operations to more users.
01 Evidence
Inputs accepted by the template, permissions, inventory scope, repository branch and audit trail.
02 First checks
- List accepted inputs
- Remove arbitrary command fields
- Review job template permissions
03 Bounded action
Replace arbitrary inputs with bounded choices and isolate job templates by operational intent.
04 Rollback
Disable the exposed template or revert to the previous approved template version.
AI
A private AI agent can act but nobody can explain the action
Open atlas Tie sources, identities, tool calls, logs and human validation before increasing autonomy.
01 Evidence
Source documents, identity, tool calls, prompt context, logs and human approval point.
02 First checks
- List approved sources
- Trace tool calls
- Define human approval points
03 Bounded action
Reduce tool scope, require approval on sensitive actions and trace every tool call to a source.
04 Rollback
Disable the tool integration or force human validation until the action path is explainable.
Cloud
An Azure Event Hubs consumer keeps falling behind
Open atlas Separate namespace throttling, hot partitions, consumer processing, checkpoint progress and downstream saturation before adding capacity or replaying events.
01 Evidence
Separate namespace throttling, hot partitions, consumer processing, checkpoint progress and downstream saturation before adding capacity or replaying events.
02 First checks
- Compare ingress, egress and throttled requests
- Measure progress and checkpoint age per partition
- Check consumer ownership, errors and downstream latency
- Prove sequence bounds and idempotency before replay
03 Bounded action
Run the first checks in order: Compare ingress, egress and throttled requests | Measure progress and checkpoint age per partition | Check consumer ownership, errors and downstream latency | Prove sequence bounds and idempotency before replay. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Automation
An Azure Automation Hybrid Runbook Worker stops picking up jobs
Open atlas Separate dispatch, heartbeat, extension health, outbound HTTPS, capacity, runtime and identity before retrying a production job.
01 Evidence
Separate dispatch, heartbeat, extension health, outbound HTTPS, capacity, runtime and identity before retrying a production job.
02 First checks
- Preserve the job ID, streams and UTC window
- Compare HybridWorkerPing and extension state across the group
- Read local worker logs before restarting services
- Complete a no-side-effect canary before retry
03 Bounded action
Run the first checks in order: Preserve the job ID, streams and UTC window | Compare HybridWorkerPing and extension state across the group | Read local worker logs before restarting services | Complete a no-side-effect canary before retry. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Networking
Azure Bastion cannot open an SSH or RDP session
Open atlas Separate the client HTTPS path, Bastion health, AzureBastionSubnet rules, routing, target NSGs, guest listener and identity before exposing administration ports.
01 Evidence
Separate the client HTTPS path, Bastion health, AzureBastionSubnet rules, routing, target NSGs, guest listener and identity before exposing administration ports.
02 First checks
- Freeze one failed attempt with UTC time and target private IP
- Check Bastion provisioning state and the complete AzureBastionSubnet rule set
- Inspect target NIC effective NSGs and routes
- Verify the guest listener and authentication separately
03 Bounded action
Run the first checks in order: Freeze one failed attempt with UTC time and target private IP | Check Bastion provisioning state and the complete AzureBastionSubnet rule set | Inspect target NIC effective NSGs and routes | Verify the guest listener and authentication separately. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Cloud
An application exhausts its Azure SQL connection pool
Open atlas Separate local connection retention, SQL sessions and waits, network failures, managed identity token acquisition and autoscaling multiplication before raising the pool limit or database tier.
01 Evidence
Separate local connection retention, SQL sessions and waits, network failures, managed identity token acquisition and autoscaling multiplication before raising the pool limit or database tier.
02 First checks
- Tie one timeout to a revision and role instance
- Calculate the global pool capacity envelope
- Compare SQL sessions with active requests and waits
- Canary the correction with stop and rollback conditions
03 Bounded action
Run the first checks in order: Tie one timeout to a revision and role instance | Calculate the global pool capacity envelope | Compare SQL sessions with active requests and waits | Canary the correction with stop and rollback conditions. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Cloud
An ACR image fails signature verification before AKS promotion
Open atlas Separate mutable tags, missing signatures, untrusted publisher identity, trust-store rotation and ACR access before bypassing admission or deploying an unverified image.
01 Evidence
Separate mutable tags, missing signatures, untrusted publisher identity, trust-store rotation and ACR access before bypassing admission or deploying an unverified image.
02 First checks
- Resolve the candidate tag to one immutable digest
- Inventory signatures with Notation
- Verify repository scope, trust store and publisher identity
- Compare the verified digest with the AKS runtime imageID
03 Bounded action
Run the first checks in order: Resolve the candidate tag to one immutable digest | Inventory signatures with Notation | Verify repository scope, trust store and publisher identity | Compare the verified digest with the AKS runtime imageID. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Infrastructure
An AKS Deployment rollout stalls before the new revision becomes available
Open atlas Locate the first blocked object across Deployment, ReplicaSet, admission, scheduling, image startup and readiness before deleting pods, extending timeouts or forcing rollback.
01 Evidence
Locate the first blocked object across Deployment, ReplicaSet, admission, scheduling, image startup and readiness before deleting pods, extending timeouts or forcing rollback.
02 First checks
- Capture Deployment conditions and rollout history
- Compare desired, created and ready pods on the candidate ReplicaSet
- Read pod and namespace events before changing capacity
- Prove external configuration compatibility before rollback
03 Bounded action
Run the first checks in order: Capture Deployment conditions and rollout history | Compare desired, created and ready pods on the candidate ReplicaSet | Read pod and namespace events before changing capacity | Prove external configuration compatibility before rollback. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Automation
Azure Logic Apps runs accumulate 429 responses and retries
Open atlas Separate Logic Apps resource limits, connector throttling and destination saturation before adding retries or raising concurrency.
01 Evidence
Separate Logic Apps resource limits, connector throttling and destination saturation before adding retries or raising concurrency.
02 First checks
- Freeze one workflow, action, connection and UTC window
- Correlate run and destination request IDs
- Measure retry and concurrency amplification
- Run a bounded canary before staged recovery
03 Bounded action
Run the first checks in order: Freeze one workflow, action, connection and UTC window | Correlate run and destination request IDs | Measure retry and concurrency amplification | Run a bounded canary before staged recovery. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Automation
Feature flag evaluations differ across application instances
Open atlas Separate the published definition, per-instance ETag, refresh path, targeting context and evaluation telemetry before restarting the fleet or rolling back the flag.
01 Evidence
Separate the published definition, per-instance ETag, refresh path, targeting context and evaluation telemetry before restarting the fleet or rolling back the flag.
02 First checks
- Freeze the flag, label, change window and previous known state
- Compare the published ETag with healthy and affected instances
- Replay one stable targeting context on a canary
- Correlate FeatureEvaluation events with refresh traces and business metrics
03 Bounded action
Run the first checks in order: Freeze the flag, label, change window and previous known state | Compare the published ETag with healthy and affected instances | Replay one stable targeting context on a canary | Correlate FeatureEvaluation events with refresh traces and business metrics. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
AI
An AI agent loses constraints as the conversation grows
Open atlas Separate system instructions, history, retrieval, tool schemas and tool outputs before raising context limits or changing models.
01 Evidence
Separate system instructions, history, retrieval, tool schemas and tool outputs before raising context limits or changing models.
02 First checks
- Freeze one healthy and one degraded trace for the same journey
- Measure context growth by segment and turn
- Verify approvals and open decisions survive compaction
- Canary the bounded policy before promotion
03 Bounded action
Run the first checks in order: Freeze one healthy and one degraded trace for the same journey | Measure context growth by segment and turn | Verify approvals and open decisions survive compaction | Canary the bounded policy before promotion. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Cloud
Azure Managed Redis connections surge while application requests time out
Open atlas Separate legitimate server pressure, client reconnect amplification, runtime DNS/TLS and release drift before scaling or triggering a failover.
01 Evidence
Separate legitimate server pressure, client reconnect amplification, runtime DNS/TLS and release drift before scaling or triggering a failover.
02 First checks
- Freeze the release, replica count and incident window
- Compare connected clients, server load and useful operations
- Measure the connection envelope per role instance
- Canary the client fix with explicit rollback
03 Bounded action
Run the first checks in order: Freeze the release, replica count and incident window | Compare connected clients, server load and useful operations | Measure the connection envelope per role instance | Canary the client fix with explicit rollback. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Infrastructure
An Azure Monitor alert fires but no action group notifies on-call
Open atlas Separate signal evaluation, fired-alert processing, scope, filters, schedule and action group delivery before changing the alert rule.
01 Evidence
Separate signal evaluation, fired-alert processing, scope, filters, schedule and action group delivery before changing the alert rule.
02 First checks
- Freeze one fired alert ID and its expected action groups
- List enabled alert processing rules and match their scopes
- Replay the schedule in its configured time zone
- Generate a new bounded alert after propagation
03 Bounded action
Run the first checks in order: Freeze one fired alert ID and its expected action groups | List enabled alert processing rules and match their scopes | Replay the schedule in its configured time zone | Generate a new bounded alert after propagation. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Cloud
An Azure Database for PostgreSQL read replica falls behind before promotion
Open atlas Separate time lag, WAL byte gap, primary log retention, replica replay capacity and target connection readiness before planned switchover or forced promotion.
01 Evidence
Separate time lag, WAL byte gap, primary log retention, replica replay capacity and target connection readiness before planned switchover or forced promotion.
02 First checks
- Freeze the promotion mode and accepted RPO
- Compare lag seconds, lag bytes and a business marker
- Check primary WAL storage and replica saturation
- Validate virtual endpoints, identity and write path before promotion
03 Bounded action
Run the first checks in order: Freeze the promotion mode and accepted RPO | Compare lag seconds, lag bytes and a business marker | Check primary WAL storage and replica saturation | Validate virtual endpoints, identity and write path before promotion. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Automation
An Azure Automation runbook returns 403 after a managed identity change
Open atlas Separate cloud sandbox and Hybrid Worker execution, Automation account and VM identities, persisted Az context, control-plane and data-plane permissions before broadening RBAC.
01 Evidence
Separate cloud sandbox and Hybrid Worker execution, Automation account and VM identities, persisted Az context, control-plane and data-plane permissions before broadening RBAC.
02 First checks
- Freeze one job ID, execution target and denied operation
- Reset Az context and prove the runtime principal
- Match the exact action to role and scope
- Run positive and negative canaries before promotion
03 Bounded action
Run the first checks in order: Freeze one job ID, execution target and denied operation | Reset Az context and prove the runtime principal | Match the exact action to role and scope | Run positive and negative canaries before promotion. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Networking
An Azure Front Door origin becomes unhealthy and traffic shifts unexpectedly
Open atlas Separate the health probe contract, DNS/TLS, optional Private Link, backend capacity and routing changes before draining, failing over or restoring traffic.
01 Evidence
Separate the health probe contract, DNS/TLS, optional Private Link, backend capacity and routing changes before draining, failing over or restoring traffic.
02 First checks
- Freeze the route, origin group and UTC window
- Read failed probes by origin and POP
- Verify hostname, SNI, certificate and health response
- Prove standby capacity before changing weights
03 Bounded action
Run the first checks in order: Freeze the route, origin group and UTC window | Read failed probes by origin and POP | Verify hostname, SNI, certificate and health response | Prove standby capacity before changing weights. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
AI
An AI agent recalls context from another session or team
Open atlas Separate conversation history, compacted summaries, durable memory, retrieval caches, tool results and runtime identity before enabling persistence for more users.
01 Evidence
Separate conversation history, compacted summaries, durable memory, retrieval caches, tool results and runtime identity before enabling persistence for more users.
02 First checks
- Freeze the user, team, environment and session boundaries
- Seed distinct synthetic canaries in two scopes
- Enable each state layer independently
- Repeat under concurrency, expiry and revocation
03 Bounded action
Run the first checks in order: Freeze the user, team, environment and session boundaries | Seed distinct synthetic canaries in two scopes | Enable each state layer independently | Repeat under concurrency, expiry and revocation. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Infrastructure
An Azure Update Manager maintenance window ends with part of the fleet still noncompliant
Open atlas Separate dynamic-scope membership, schedule orchestration, exhausted window budget, patch failures and pending restarts before forcing an out-of-window installation.
01 Evidence
Separate dynamic-scope membership, schedule orchestration, exhausted window budget, patch failures and pending restarts before forcing an out-of-window installation.
02 First checks
- Freeze the maintenance contract and expected inventory
- Compare matching machines with maintenance and installation runs
- Separate no run, partial run, failed update and pending restart
- Prove service capacity before resuming a bounded set
03 Bounded action
Run the first checks in order: Freeze the maintenance contract and expected inventory | Compare matching machines with maintenance and installation runs | Separate no run, partial run, failed update and pending restart | Prove service capacity before resuming a bounded set. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Cloud
An Azure App Service certificate is renewed but the old certificate is still served
Open atlas Separate issuance, domain validation, Key Vault version, App Service import, hostname binding and the certificate observed with SNI before syncing or rebinding.
01 Evidence
Separate issuance, domain validation, Key Vault version, App Service import, hostname binding and the certificate observed with SNI before syncing or rebinding.
02 First checks
- Read the served fingerprint with the real hostname and SNI
- Compare candidate SAN and expiration with App Service inventory
- Check Key Vault version and import state when applicable
- Preserve the current binding and rollback thumbprint
03 Bounded action
Run the first checks in order: Read the served fingerprint with the real hostname and SNI | Compare candidate SAN and expiration with App Service inventory | Check Key Vault version and import state when applicable | Preserve the current binding and rollback thumbprint. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Networking
An Azure flow is denied although the NSG allows it
Open atlas Separate Azure Virtual Network Manager Security Admin rules, network group membership, NSG evaluation and actual service reachability before opening local rules.
01 Evidence
Separate Azure Virtual Network Manager Security Admin rules, network group membership, NSG evaluation and actual service reachability before opening local rules.
02 First checks
- Capture the exact source, destination and port tuple
- List effective Security Admin rules on the VNet
- Run IP Flow Verify and identify the deciding rule
- Confirm the decision in VNet flow logs
03 Bounded action
Run the first checks in order: Capture the exact source, destination and port tuple | List effective Security Admin rules on the VNet | Run IP Flow Verify and identify the deciding rule | Confirm the decision in VNet flow logs. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Networking
An Azure Traffic Manager profile is degraded and failover appears inconsistent across clients
Open atlas Separate the deployed probe contract, endpoint monitor state, backend evidence, DNS answers, TTL and secondary capacity before disabling the primary endpoint or forcing failover.
01 Evidence
Separate the deployed probe contract, endpoint monitor state, backend evidence, DNS answers, TTL and secondary capacity before disabling the primary endpoint or forcing failover.
02 First checks
- Capture profile, endpoint, probe and TTL configuration
- Replay the exact probe contract from external viewpoints
- Correlate ProbeHealthStatusEvents with backend logs
- Compare DNS answers and prove secondary readiness
03 Bounded action
Run the first checks in order: Capture profile, endpoint, probe and TTL configuration | Replay the exact probe contract from external viewpoints | Correlate ProbeHealthStatusEvents with backend logs | Compare DNS answers and prove secondary readiness. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Automation
An Azure Automation runbook works in its current runtime but fails in the candidate environment
Open atlas Separate language version, package resolution, identity, serialization and Hybrid Worker readiness before relinking production runbooks.
01 Evidence
Separate language version, package resolution, identity, serialization and Hybrid Worker readiness before relinking production runbooks.
02 First checks
- Freeze current and candidate runtime manifests
- Run positive and negative draft tests with the candidate
- Qualify every target Hybrid Worker
- Canary one bounded effect and confirm an empty second run
03 Bounded action
Run the first checks in order: Freeze current and candidate runtime manifests | Run positive and negative draft tests with the candidate | Qualify every target Hybrid Worker | Canary one bounded effect and confirm an empty second run. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Infrastructure
Azure Monitor Managed Prometheus cardinality spikes after an AKS deployment
Open atlas Separate workspace ingestion pressure, scrape failures, duplicate targets and unbounded labels before dropping metrics, raising quotas or rolling back the release.
01 Evidence
Separate workspace ingestion pressure, scrape failures, duplicate targets and unbounded labels before dropping metrics, raising quotas or rolling back the release.
02 First checks
- Freeze workspace, cluster, release and UTC window
- Compare active-series and event-ingestion trends
- Bound PromQL to identify the metric and label
- Canary relabeling without breaking alerts
03 Bounded action
Run the first checks in order: Freeze workspace, cluster, release and UTC window | Compare active-series and event-ingestion trends | Bound PromQL to identify the metric and label | Canary relabeling without breaking alerts. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Infrastructure
Azure VMSS repeatedly replaces or reimages unhealthy instances
Open atlas Separate the health signal, grace period, orchestration service state, current VMSS model and persistence contract before changing Replace, Reimage or Restart.
01 Evidence
Separate the health signal, grace period, orchestration service state, current VMSS model and persistence contract before changing Replace, Reimage or Restart.
02 First checks
- Freeze the repair timeline and useful capacity
- Read automaticRepairsPolicy and orchestration service state
- Replay the health contract on healthy and affected instances
- Build one canary from the current model before resuming
03 Bounded action
Run the first checks in order: Freeze the repair timeline and useful capacity | Read automaticRepairsPolicy and orchestration service state | Replay the health contract on healthy and affected instances | Build one canary from the current model before resuming. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Automation
A newly pushed ACR image is not visible from every deployment region
Open atlas Separate asynchronous geo-replication, mutable tags, global endpoint routing, runtime identity and data endpoint reachability before retrying a multi-region rollout.
01 Evidence
Separate asynchronous geo-replication, mutable tags, global endpoint routing, runtime identity and data endpoint reachability before retrying a multi-region rollout.
02 First checks
- Freeze the build digest and push completion time
- List replica state and inspect Resource Health
- Pull the expected digest from every target region
- Correlate ACR events by region, identity and result
03 Bounded action
Run the first checks in order: Freeze the build digest and push completion time | List replica state and inspect Resource Health | Pull the expected digest from every target region | Correlate ACR events by region, identity and result. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Networking
An AKS service-to-service flow is blocked after a NetworkPolicy change
Open atlas Separate source egress, destination ingress, deployed selectors, service endpoints and observed Cilium verdicts before adding an allow-all rule or removing default deny.
01 Evidence
Separate source egress, destination ingress, deployed selectors, service endpoints and observed Cilium verdicts before adding an allow-all rule or removing default deny.
02 First checks
- Freeze the source, destination, port and UTC window
- Read the AKS data plane and every policy type
- Compare selectors with deployed pod and namespace labels
- Run positive and negative canaries before promotion
03 Bounded action
Run the first checks in order: Freeze the source, destination, port and UTC window | Read the AKS data plane and every policy type | Compare selectors with deployed pod and namespace labels | Run positive and negative canaries before promotion. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Infrastructure
OpenTelemetry tail sampling reduces volume but incident traces become incomplete or disappear
Open atlas Separate trace affinity, decision-window sizing, pending-trace memory, late spans, policy behavior and export evidence before widening the sampling rollout.
01 Evidence
Separate trace affinity, decision-window sizing, pending-trace memory, late spans, policy behavior and export evidence before widening the sampling rollout.
02 First checks
- Freeze the current and candidate Collector configuration digests
- Prove every span for one trace reaches the same sampling shard
- Replay error, latency and late-span canaries
- Correlate Collector decisions with complete operations in Azure Monitor
03 Bounded action
Run the first checks in order: Freeze the current and candidate Collector configuration digests | Prove every span for one trace reaches the same sampling shard | Replay error, latency and late-span canaries | Correlate Collector decisions with complete operations in Azure Monitor. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Infrastructure
Key Vault returns 403 after an Azure RBAC cutover
Open atlas Separate missed runtime identities, data-plane role mapping, assignment scope and propagation from network failures before granting a broad role or reverting the permission model.
01 Evidence
Separate missed runtime identities, data-plane role mapping, assignment scope and propagation from network failures before granting a broad role or reverting the permission model.
02 First checks
- Confirm the active Key Vault permission model
- Resolve the failing runtime principal object ID
- Compare required operations with direct and inherited data-plane roles
- Correlate 403 responses with Key Vault AuditEvent logs
03 Bounded action
Run the first checks in order: Confirm the active Key Vault permission model | Resolve the failing runtime principal object ID | Compare required operations with direct and inherited data-plane roles | Correlate 403 responses with Key Vault AuditEvent logs. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Networking
Virtual WAN traffic breaks or becomes asymmetric after Routing Intent is enabled
Open atlas Separate policy scope, effective routes, private prefixes, default-route propagation, security next hop, SNAT and return paths before widening the rollout or deleting Routing Intent.
01 Evidence
Separate policy scope, effective routes, private prefixes, default-route propagation, security next hop, SNAT and return paths before widening the rollout or deleting Routing Intent.
02 First checks
- Freeze the former route associations and propagations
- Verify every contract prefix on the security next hop
- Probe forward and return paths through each affected hub
- Correlate canaries with Firewall allows and denies
03 Bounded action
Run the first checks in order: Freeze the former route associations and propagations | Verify every contract prefix on the security next hop | Probe forward and return paths through each affected hub | Correlate canaries with Firewall allows and denies. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Networking
AKS pods see intermittent DNS timeouts or SERVFAIL responses
Open atlas Separate pod resolver behavior, query amplification, the Kubernetes DNS service, CoreDNS endpoints, network policy, upstream forwarding and node-specific paths before restarting CoreDNS.
01 Evidence
Separate pod resolver behavior, query amplification, the Kubernetes DNS service, CoreDNS endpoints, network policy, upstream forwarding and node-specific paths before restarting CoreDNS.
02 First checks
- Freeze the exact names, errors and affected pod/node cohort
- Compare repeated lookups across nodes and namespaces
- Inspect the kube-dns service, EndpointSlices and CoreDNS pods
- Test cluster names separately from each upstream suffix
03 Bounded action
Run the first checks in order: Freeze the exact names, errors and affected pod/node cohort | Compare repeated lookups across nodes and namespaces | Inspect the kube-dns service, EndpointSlices and CoreDNS pods | Test cluster names separately from each upstream suffix. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Networking
Azure Firewall IDPS blocks a legitimate flow after a signature is moved to deny
Open atlas Separate signature evidence, traffic direction, private-range classification, TLS inspection and application outcome before disabling IDPS or adding a broad bypass.
01 Evidence
Separate signature evidence, traffic direction, private-range classification, TLS inspection and application outcome before disabling IDPS or adding a broad bypass.
02 First checks
- Freeze the signature override and affected five-tuple
- Query AZFWIdpsSignature for source, destination and action
- Replay one malicious and one legitimate canary
- Return only the signature to Alert if the legitimate flow regresses
03 Bounded action
Run the first checks in order: Freeze the signature override and affected five-tuple | Query AZFWIdpsSignature for source, destination and action | Replay one malicious and one legitimate canary | Return only the signature to Alert if the legitimate flow regresses. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Automation
An Azure Automation job may run a version that does not match the approved commit
Open atlas Tie the source commit, sync job, Draft and Published hashes, no-effect test and execution marker together before rerunning or broadening production scope.
01 Evidence
Tie the source commit, sync job, Draft and Published hashes, no-effect test and execution marker together before rerunning or broadening production scope.
02 First checks
- Freeze the approved commit, script hash and target account
- Read the last sync job and its streams before retrying
- Export and hash Draft and Published separately
- Run Plan and one reversible canary before full execution
03 Bounded action
Run the first checks in order: Freeze the approved commit, script hash and target account | Read the last sync job and its streams before retrying | Export and hash Draft and Published separately | Run Plan and one reversible canary before full execution. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Infrastructure
One Azure Monitor log alert creates hundreds of independently evaluated alert instances
Open atlas Separate real affected cohorts, split-by dimension cardinality, alert lifecycle and notification routing before raising limits, muting the action group or changing the KQL rule.
01 Evidence
Separate real affected cohorts, split-by dimension cardinality, alert lifecycle and notification routing before raising limits, muting the action group or changing the KQL rule.
02 First checks
- Export the deployed scheduled query rule and freeze a UTC window
- Count rows, distinct dimension values and their combinations
- Compare query combinations, active instances and delivered notifications
- Canary stable dimensions through a complete fire-and-resolve cycle
03 Bounded action
Run the first checks in order: Export the deployed scheduled query rule and freeze a UTC window | Count rows, distinct dimension values and their combinations | Compare query combinations, active instances and delivered notifications | Canary stable dimensions through a complete fire-and-resolve cycle. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
AI
An AI gateway returns 429 after a token-limit policy change
Open atlas Separate APIM policy scope, counter isolation, gateway-local accounting, token estimation, backend throttling and client retries before raising the budget.
01 Evidence
Separate APIM policy scope, counter isolation, gateway-local accounting, token estimation, backend throttling and client retries before raising the budget.
02 First checks
- Preserve one rejected request with gateway, region and Retry-After
- Prove whether APIM or the model backend produced the 429
- Verify the effective policy and a non-empty counter key
- Canary the correction without resetting production counters
03 Bounded action
Run the first checks in order: Preserve one rejected request with gateway, region and Retry-After | Prove whether APIM or the model backend produced the 429 | Verify the effective policy and a non-empty counter key | Canary the correction without resetting production counters. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Networking
An Azure application path fails intermittently while manual connectivity tests remain healthy
Open atlas Use Connection Monitor to preserve source-specific loss and latency, then correlate one failing interval with DNS, effective routes, security decisions, firewall evidence and service health before changing NSGs or UDRs.
01 Evidence
Use Connection Monitor to preserve source-specific loss and latency, then correlate one failing interval with DNS, effective routes, security decisions, firewall evidence and service health before changing NSGs or UDRs.
02 First checks
- Freeze the exact source, FQDN, port and UTC window
- Compare failed-check percentage and RTT per source
- Run Connection Troubleshoot during the failing interval
- Canary one bounded correction with a refusal control
03 Bounded action
Run the first checks in order: Freeze the exact source, FQDN, port and UTC window | Compare failed-check percentage and RTT per source | Run Connection Troubleshoot during the failing interval | Canary one bounded correction with a refusal control. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
AI
An AI agent proposes an unrequested production action after reading MCP tool output
Open atlas Separate authenticated transport from field-level trust, preserve tool-result provenance, enforce write authorization outside the model and replay adversarial results before reopening production actions.
01 Evidence
Separate authenticated transport from field-level trust, preserve tool-result provenance, enforce write authorization outside the model and replay adversarial results before reopening production actions.
02 First checks
- Freeze the user intent, raw tool result and proposed next call
- Identify which untrusted field influenced the action
- Verify approval and arguments in an external policy gate
- Replay hostile results with writes disabled before a bounded canary
03 Bounded action
Run the first checks in order: Freeze the user intent, raw tool result and proposed next call | Identify which untrusted field influenced the action | Verify approval and arguments in an external policy gate | Replay hostile results with writes disabled before a bounded canary. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
AI
An agent stays available through model fallback but changes its tool decisions or refusal behavior
Open atlas Separate endpoint availability from response-schema, tool, source, refusal, approval, state and trace compatibility before enabling automatic multi-model routing.
01 Evidence
Separate endpoint availability from response-schema, tool, source, refusal, approval, state and trace compatibility before enabling automatic multi-model routing.
02 First checks
- Freeze both model attempts under one logical run ID
- Validate structured output and normalized tool arguments
- Reconcile any unknown primary result before fallback
- Canary new read-only sessions with route-level traces
03 Bounded action
Run the first checks in order: Freeze both model attempts under one logical run ID | Validate structured output and normalized tool arguments | Reconcile any unknown primary result before fallback | Canary new read-only sessions with route-level traces. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Infrastructure
An AKS Horizontal Pod Autoscaler repeatedly scales up and down
Open atlas Separate noisy or missing metrics, miscalibrated resource requests, pod warm-up, competing controllers and node capacity before raising replica bounds.
01 Evidence
Separate noisy or missing metrics, miscalibrated resource requests, pod warm-up, competing controllers and node capacity before raising replica bounds.
02 First checks
- Freeze two complete scaling cycles
- Compare the HPA status with the metric API
- Measure requests, readiness and downstream capacity per replica
- Canary one behavior change with explicit rollback
03 Bounded action
Run the first checks in order: Freeze two complete scaling cycles | Compare the HPA status with the metric API | Measure requests, readiness and downstream capacity per replica | Canary one behavior change with explicit rollback. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Cloud
An Azure deployment fails because a resource or inherited scope is locked
Open atlas Separate management locks, RBAC, Policy and deny assignments, identify the exact destructive operation, then bound any unlock window before rerun or rollback.
01 Evidence
Separate management locks, RBAC, Policy and deny assignments, identify the exact destructive operation, then bound any unlock window before rerun or rollback.
02 First checks
- Freeze the failed operation, target resource and correlation ID
- Reconstruct inherited CanNotDelete and ReadOnly locks
- Confirm whether the plan really requires a delete
- Recreate and verify the lock before closing the change
03 Bounded action
Run the first checks in order: Freeze the failed operation, target resource and correlation ID | Reconstruct inherited CanNotDelete and ReadOnly locks | Confirm whether the plan really requires a delete | Recreate and verify the lock before closing the change. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Networking
AKS pods intermittently time out to public dependencies
Open atlas Separate DNS, the remote dependency, network policy, node connection pressure and the effective Azure egress device before adding SNAT capacity or changing the cluster outbound path.
01 Evidence
Separate DNS, the remote dependency, network policy, node connection pressure and the effective Azure egress device before adding SNAT capacity or changing the cluster outbound path.
02 First checks
- Capture the failing pod, node and destination tuple
- Confirm networkProfile.outboundType and the source IP seen downstream
- Correlate per-node SNAT metrics with connection churn
- Canary connection reuse or egress capacity before rollout
03 Bounded action
Run the first checks in order: Capture the failing pod, node and destination tuple | Confirm networkProfile.outboundType and the source IP seen downstream | Correlate per-node SNAT metrics with connection churn | Canary connection reuse or egress capacity before rollout. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Automation
A Microsoft Sentinel rule misses a target or automates a response from stale watchlist data
Open atlas Separate source freshness, snapshot completeness, bulk-update deletion semantics, SearchKey normalization and KQL join behavior before allowing a watchlist to drive incidents or playbooks.
01 Evidence
Separate source freshness, snapshot completeness, bulk-update deletion semantics, SearchKey normalization and KQL join behavior before allowing a watchlist to drive incidents or playbooks.
02 First checks
- Prove the alias and source snapshot independently of TimeGenerated
- Compare added, modified and removed keys with the active version
- Reject blank, duplicate or expired SearchKey values
- Shadow the analytics rule before enabling the playbook
03 Bounded action
Run the first checks in order: Prove the alias and source snapshot independently of TimeGenerated | Compare added, modified and removed keys with the active version | Reject blank, duplicate or expired SearchKey values | Shadow the analytics rule before enabling the playbook. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Infrastructure
A self-hosted Azure DevOps agent shows signs of compromise but can still receive production jobs
Open atlas Stop scheduling without erasing evidence, reconstruct affected runs, map exposed identities and validate a clean replacement before reopening the pool.
01 Evidence
Stop scheduling without erasing evidence, reconstruct affected runs, map exposed identities and validate a clean replacement before reopening the pool.
02 First checks
- Quarantine the agent and pause production entry points
- Freeze run IDs, commits, agent diagnostics and audit events
- Map credentials and downstream targets used in the incident window
- Canary a clean replacement without protected resources
03 Bounded action
Run the first checks in order: Quarantine the agent and pause production entry points | Freeze run IDs, commits, agent diagnostics and audit events | Map credentials and downstream targets used in the incident window | Canary a clean replacement without protected resources. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Networking
An App Service plan stays online but scaling fails on a VNet Integration subnet
Open atlas Separate address exhaustion, delayed release, delegation and outbound path failures; calculate worker cohort overlap before retrying scale or migrating the integration.
01 Evidence
Separate address exhaustion, delayed release, delegation and outbound path failures; calculate worker cohort overlap before retrying scale or migrating the integration.
02 First checks
- Freeze the failed operation and current instance count
- Inventory every app and plan joined to the subnet
- Calculate usable addresses and transient worker overlap
- Canary a larger delegated subnet with explicit rollback
03 Bounded action
Run the first checks in order: Freeze the failed operation and current instance count | Inventory every app and plan joined to the subnet | Calculate usable addresses and transient worker overlap | Canary a larger delegated subnet with explicit rollback. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Cloud
Azure WAF appears to enforce different policy behavior after a deployment
Open atlas Separate ARM rollout state, gateway, listener and path-rule associations, deployed rule priority and real request variance before waiting or restoring the whole policy.
01 Evidence
Separate ARM rollout state, gateway, listener and path-rule associations, deployed rule priority and real request variance before waiting or restoring the whole policy.
02 First checks
- Freeze one accepted and one blocked request
- Read policy and gateway provisioning states
- Inventory policy resource IDs at every association scope
- Replay an allowed and denied canary after correction or rollback
03 Bounded action
Run the first checks in order: Freeze one accepted and one blocked request | Read policy and gateway provisioning states | Inventory policy resource IDs at every association scope | Replay an allowed and denied canary after correction or rollback. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Automation
A previously allowed Azure deployment is denied after its Policy exemption expires
Open atlas Tie the first denial to expiresOn, the exact assignment, initiative references and current resource state before renewing the exemption or disabling enforcement.
01 Evidence
Tie the first denial to expiresOn, the exact assignment, initiative references and current resource state before renewing the exemption or disabling enforcement.
02 First checks
- Freeze the first denied run and correlation ID
- Read the exemption at its exact scope and compare expiresOn in UTC
- Match assignment and initiative reference IDs
- Canary the same artifact with an out-of-scope denial control
03 Bounded action
Run the first checks in order: Freeze the first denied run and correlation ID | Read the exemption at its exact scope and compare expiresOn in UTC | Match assignment and initiative reference IDs | Canary the same artifact with an out-of-scope denial control. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
AI
An AI agent trace contains sensitive data after an instrumentation change
Open atlas Tie release, span, unsafe field, export route and destination together before disabling observability; contain the narrow path, handle historical copies and canary redaction.
01 Evidence
Tie release, span, unsafe field, export route and destination together before disabling observability; contain the narrow path, handle historical copies and canary redaction.
02 First checks
- Freeze trace metadata without copying the raw value
- Map every collector, exporter and destination copy
- Disable the unsafe enrichment while keeping minimum events
- Replay synthetic markers through every input path
03 Bounded action
Run the first checks in order: Freeze trace metadata without copying the raw value | Map every collector, exporter and destination copy | Disable the unsafe enrichment while keeping minimum events | Replay synthetic markers through every input path. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
Automation
A canceled Azure DevOps deployment leaves the production target in an uncertain partial state
Open atlas Reconcile the run, agent, Azure control plane and application before rerunning; prove each checkpoint, artifact and side effect, then choose resume, compensation or rollback.
01 Evidence
Reconcile the run, agent, Azure control plane and application before rerunning; prove each checkpoint, artifact and side effect, then choose resume, compensation or rollback.
02 First checks
- Freeze every writer targeting the environment
- Pin the run, commit and artifact digest
- Read target-side operations across the cancellation window
- Block rerun while a critical checkpoint remains unknown
03 Bounded action
Run the first checks in order: Freeze every writer targeting the environment | Pin the run, commit and artifact digest | Read target-side operations across the cancellation window | Block rerun while a critical checkpoint remains unknown. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.
AI
An Azure AI Search indexer succeeds partially and the agent corpus becomes incomplete
Open atlas Separate source access, incremental tracking, skillset failures and target-index writes before rerunning selected documents or clearing the high-water mark.
01 Evidence
Separate source access, incremental tracking, skillset failures and target-index writes before rerunning selected documents or clearing the high-water mark.
02 First checks
- Freeze the execution history and tracking states
- Reconcile failed source IDs with index keys and versions
- Group errors by stage, code and skill
- Canary targeted recovery before any full reset
03 Bounded action
Run the first checks in order: Freeze the execution history and tracking states | Reconcile failed source IDs with index keys and versions | Group errors by stage, code and skill | Canary targeted recovery before any full reset. Open the linked notes before changing production.
04 Rollback
Stop the change, restore the last known safe state and keep the captured evidence for comparison.