Cloud 2 notes
Application Gateway returns 502
Separate backend health, DNS resolution, TLS settings and private network reachability before changing the application.
First checks - Check backend health state
- Resolve backend name from the gateway path
- Validate TLS/SNI and probe configuration
- 01 Azure Application Gateway: diagnose 502 errors without mixing DNS, TLS and backend health
- 02 Azure Private Endpoint: build a validation matrix before production
Cloud 4 notes
Internal APIM returns an error on a private API
Correlate Application Gateway/WAF and APIM logs, then separate DNS, TLS, policy, identity and private backend reachability before changing policies or opening access.
First checks - Check whether WAF blocked the request
- Confirm APIM received the same path
- Validate backend DNS and TLS from the APIM path
- Replay with a correlation ID
- 01 Azure internal APIM: diagnose a private API before changing policies
- 02 KQL snippet: correlate WAF and APIM on an Azure private API
- 03 Azure Application Gateway: diagnose 502 errors without mixing DNS, TLS and backend health
- 04 Azure APIM: diagnose backend TLS before disabling certificate validation
Networking 4 notes
Private Endpoint name still resolves publicly
Confirm the CNAME chain, Private DNS Zone association and hybrid forwarding from the consuming network.
First checks - Run nslookup from the workload network
- Check privatelink CNAME
- Verify Private DNS Zone links and forwarders
- 01 Azure snippet: check Private Endpoint DNS resolution
- 02 Azure snippet: detect Private Endpoint DNS drift
- 03 Azure hybrid DNS: when to use Private Resolver, on-premises forwarders and private zones
- 04 Azure Private Endpoint: detect Terraform, DNS, and network drift before incident
Cloud 3 notes
Azure Storage private endpoint returns 403, times out or produces no request logs
Separate Storage subresource DNS, Private Endpoint approval, firewall rules, runtime identity and Storage logs before opening public access or broadening RBAC.
First checks - Resolve the exact Storage subresource from the workload network
- Check Private Endpoint status and private DNS zone group
- Replay with a client request ID
- Correlate Storage logs for 403, caller IP and requester identity
- 01 Azure Storage: diagnose a private endpoint without opening the account
- 02 KQL snippet: isolate Azure Storage 403 on a private endpoint
- 03 Azure Private Endpoint: detect Terraform, DNS, and network drift before incident
Cloud 3 notes
Azure SQL private endpoint returns timeouts, firewall errors or no SQL logs
Separate SQL private DNS, Private Endpoint state, public access, firewall rules, runtime identity and SQL diagnostics before changing schema, code or broad permissions.
First checks - Resolve the SQL FQDN from the workload network
- Check Private Endpoint status and privatelink.database.windows.net records
- Replay with the real runtime identity
- Correlate SQL diagnostics for firewall and login errors
- 01 Azure SQL: diagnose a private endpoint before changing the database
- 02 KQL snippet: isolate Azure SQL private endpoint errors
- 03 Azure Private Endpoint: detect Terraform, DNS, and network drift before incident
Cloud 3 notes
Azure Service Bus private endpoint times out, denies access or lets backlog grow
Separate Service Bus private DNS, Private Endpoint state, public access, managed identity or SAS, queue metrics and processing logs before touching queues or redeploying consumers.
First checks - Resolve the Service Bus FQDN from the workload network
- Check Private Endpoint approval and public network access
- Verify the real sender or receiver identity
- Correlate backlog, dead-letter and Service Bus errors
- 01 Azure Service Bus: diagnose a private endpoint before touching queues
- 02 KQL snippet: isolate Service Bus private endpoint errors
- 03 Azure Private Endpoint: detect Terraform, DNS, and network drift before incident
Cloud 4 notes
A synthetic probe fails on an Azure private path
Separate DNS, TLS, Application Gateway health, WAF blocks and runner network before changing routing or application code.
First checks - Resolve the hostname from the probe network
- Check TLS/SNI with the real hostname
- Correlate probe run with WAF and gateway logs
- 01 Azure: make private paths verifiable with synthetic probes
- 02 KQL snippet: track synthetic probes for an Azure private path
- 03 Azure Application Gateway: diagnose 502 errors without mixing DNS, TLS and backend health
- 04 Azure Private Endpoint: detect Terraform, DNS, and network drift before incident
Networking 5 notes
Azure traffic leaves through one path and returns through another
Compare DNS target, effective routes, route table associations, firewall evidence, NAT identity and return path before changing UDRs or bypassing inspection.
First checks - Resolve the destination from the source path
- Compare effective routes on source and destination NICs
- Check firewall or appliance logs for both directions
- Confirm NAT or outbound source identity
- 01 Azure route asymmetry: diagnose before changing a UDR
- 02 Azure UDR: diagnose an NVA black hole before changing routes
- 03 Azure NSG, UDR and Firewall: diagnose blocked traffic before adding a rule
- 04 Azure Firewall: diagnose a shadowed rule before opening traffic
- 05 Azure UDR and NAT Gateway: diagnose egress before opening the firewall
Cloud 3 notes
Azure Container Apps private ingress fails or reaches the wrong revision
Separate private DNS, Application Gateway handoff, Container Apps ingress mode, revision traffic and console logs before rolling back or changing traffic weights.
First checks - Resolve the hostname from the caller network
- Check ingress target port and active revisions
- Correlate system and console logs
- 01 Azure Container Apps: diagnose private ingress before changing revisions
- 02 KQL snippet: diagnose Container Apps private ingress and revisions
- 03 Azure: make private paths verifiable with synthetic probes
Cloud 3 notes
AKS private ingress returns 502 or reaches no service endpoints
Separate private DNS, Application Gateway health, ingress controller routing, Kubernetes service selectors, endpoint slices and pod readiness before rolling back a deployment.
First checks - Resolve the hostname from the caller network
- Check Application Gateway backend health and host header
- Verify ingress, service and endpoint slices
- Correlate controller and application logs
- 01 Azure AKS: diagnose private ingress before changing deployments
- 02 KQL snippet: correlate AKS private ingress and application logs
- 03 Azure Application Gateway: diagnose 502 errors without mixing DNS, TLS and backend health
Infrastructure 1 note
An AKS node pool upgrade is blocked while draining a pod protected by a PDB
Separate unavailable replicas, readiness failures, pod placement, surge capacity and an over-constrained PodDisruptionBudget before deleting the safeguard or forcing the upgrade.
First checks - Capture the blocked node and eviction events
- Inspect PDB disruptionsAllowed and selectors
- Verify replacement capacity, replica readiness and topology
- Resume only with positive disruption headroom
- 01 Azure AKS: unblock a node pool upgrade stopped by a PDB
Cloud 3 notes
Azure Functions private HTTP endpoint returns 403, 503 or no request logs
Separate private DNS, Private Endpoint reachability, access restrictions, Functions runtime state, private storage and Application Insights evidence before redeploying code or opening public access.
First checks - Resolve the hostname from the caller network
- Replay with a correlation ID
- Check Function App access restrictions and Private Endpoint status
- Correlate requests, traces and exceptions
- 01 Azure Functions: diagnose a private HTTP endpoint before changing code
- 02 KQL snippet: correlate an Azure Functions private HTTP endpoint
- 03 Azure: make private paths verifiable with synthetic probes
Cloud 5 notes
Azure WAF blocks a legitimate request
Start from blocked requests, rule ID and URI before deciding between exclusion, custom rule or application fix.
First checks - List blocked URIs in KQL
- Identify ruleId and match field
- Validate false-positive scope
- 01 KQL snippet: list Azure WAF blocked URIs quickly
- 02 WAF and KQL: identify a false positive before creating an exclusion
- 03 Azure WAF: add an OWASP/CRS exclusion without weakening all protection
- 04 Azure WAF: frame an emergency custom rule without losing evidence
- 05 Azure snippet: audit WAF custom rule priorities
Automation 4 notes
Terraform state lock is stuck
Prove that no apply is still running before using force-unlock, then restart with a clean plan.
First checks - Identify lock owner
- Check CI job status
- Run plan after unlock
- 01 Terraform snippet: diagnose a stuck state lock before force-unlock
- 02 Terraform Azure: secure a private state backend without breaking CI
- 03 Terraform on Azure: diagnose a partial apply before rerunning production
- 04 Terraform on Azure: contain concurrent applies before state divergence
Infrastructure 2 notes
Azure Monitor fires an alert storm after deployment
Separate real service impact, noisy dimensions, threshold drift, action group behavior and rollback before silencing notifications or changing rules.
First checks - Group fired alerts by rule and target
- Compare first alert time with the deployment window
- Check service logs before changing thresholds
- Keep one user-symptom alert active
- 01 Azure Monitor: diagnose an alert storm after deployment
- 02 Monitoring: turn an alert into an actionable operations runbook
Infrastructure 8 notes
A secret rotation, federated identity or managed identity change breaks an application or pipeline consumer
Separate preparation, cutover, revocation, workload identity federation and managed identity diagnostics; validate the real execution identity, private path and authentication errors before deleting the old value or broadening access.
First checks - List real consumers
- Verify the runtime identity and vault read access
- Check OIDC claims or private DNS depending on the path
- Watch 401/403/500, sign-in failures or Key Vault denials
- 01 Service identity and secret rotation: a production runbook, not an isolated task
- 02 KQL snippet: detect authentication errors after secret rotation
- 03 Azure managed identity: diagnose private access before changing permissions
- 04 KQL snippet: diagnose Key Vault denial with managed identity
- 05 Azure Workload Identity Federation: diagnose CI authentication before bringing back a secret
- 06 Azure DevOps WIF: diagnose issuer or subject drift before bringing back a secret
- 07 Azure Managed Identity: diagnose principal drift after resource recreation
- 08 Azure Managed Grafana: diagnose API access before recreating a token
Automation 1 note
An Azure Functions poison queue keeps growing
Separate contract failures, transient dependencies, identity, resource pressure and partial side effects before replaying Storage Queue messages.
First checks - Peek a redacted sample without consuming messages
- Correlate message IDs with invocation and dependency logs
- Verify downstream state and idempotency before replay
- 01 Azure Functions: diagnose a poison queue before replaying messages
Cloud 1 note
An Azure Service Bus dead-letter queue keeps growing
Classify dead-letter reasons, correlate consumer and downstream evidence, verify idempotency, then use a manifest, canary and stop conditions before replay.
First checks - Record the exact queue or subscription and DLQ growth window
- Peek a redacted sample grouped by dead-letter reason
- Check consumer errors, lock loss and downstream state
- Prove idempotency before a one-message canary
- 01 Azure Service Bus: diagnose the dead-letter queue before a bounded replay
Automation 1 note
An Azure Policy remediation changes unexpected production resources
Pin the definition and assignment, materialize the eligible target set, verify the managed identity, then use a canary and bounded batches before expanding or compensating.
First checks - Identify the exact definition reference and assignment parameters
- Export non-compliant target resource IDs
- Review the assignment identity and effective RBAC scope
- Correlate remediation deployments with Activity Log
- 01 Azure Policy: bound a remediation task before it rewrites production
Automation 2 notes
An automation entry point behaves like a remote console
Bound inputs, templates and repository structure before exposing operations to more users.
First checks - List accepted inputs
- Remove arbitrary command fields
- Review job template permissions
- 01 AWX: design job templates that do not become a dangerous remote console
- 02 Ansible in production: structure an operations repository before exposing it in AWX
AI 2 notes
A private AI agent can act but nobody can explain the action
Tie sources, identities, tool calls, logs and human validation before increasing autonomy.
First checks - List approved sources
- Trace tool calls
- Define human approval points
- 01 Private-network AI agent: which controls to keep around data, actions and logs
- 02 AgentOps: diagnose an AI agent that calls the wrong tool
Cloud 1 note
An Azure Event Hubs consumer keeps falling behind
Separate namespace throttling, hot partitions, consumer processing, checkpoint progress and downstream saturation before adding capacity or replaying events.
First checks - Compare ingress, egress and throttled requests
- Measure progress and checkpoint age per partition
- Check consumer ownership, errors and downstream latency
- Prove sequence bounds and idempotency before replay
- 01 Azure Event Hubs: diagnose consumer lag before scaling or replaying
Automation 1 note
An Azure Automation Hybrid Runbook Worker stops picking up jobs
Separate dispatch, heartbeat, extension health, outbound HTTPS, capacity, runtime and identity before retrying a production job.
First checks - Preserve the job ID, streams and UTC window
- Compare HybridWorkerPing and extension state across the group
- Read local worker logs before restarting services
- Complete a no-side-effect canary before retry
- 01 Azure Automation: diagnose a Hybrid Runbook Worker that stops picking up jobs
Networking 1 note
Azure Bastion cannot open an SSH or RDP session
Separate the client HTTPS path, Bastion health, AzureBastionSubnet rules, routing, target NSGs, guest listener and identity before exposing administration ports.
First checks - Freeze one failed attempt with UTC time and target private IP
- Check Bastion provisioning state and the complete AzureBastionSubnet rule set
- Inspect target NIC effective NSGs and routes
- Verify the guest listener and authentication separately
- 01 Azure Bastion: diagnose SSH or RDP failures before opening NSGs
Cloud 1 note
An application exhausts its Azure SQL connection pool
Separate local connection retention, SQL sessions and waits, network failures, managed identity token acquisition and autoscaling multiplication before raising the pool limit or database tier.
First checks - Tie one timeout to a revision and role instance
- Calculate the global pool capacity envelope
- Compare SQL sessions with active requests and waits
- Canary the correction with stop and rollback conditions
- 01 Azure SQL: diagnose connection pool exhaustion before scaling
Cloud 1 note
An ACR image fails signature verification before AKS promotion
Separate mutable tags, missing signatures, untrusted publisher identity, trust-store rotation and ACR access before bypassing admission or deploying an unverified image.
First checks - Resolve the candidate tag to one immutable digest
- Inventory signatures with Notation
- Verify repository scope, trust store and publisher identity
- Compare the verified digest with the AKS runtime imageID
- 01 Azure ACR and AKS: verify an image signature before production promotion
Infrastructure 1 note
An AKS Deployment rollout stalls before the new revision becomes available
Locate the first blocked object across Deployment, ReplicaSet, admission, scheduling, image startup and readiness before deleting pods, extending timeouts or forcing rollback.
First checks - Capture Deployment conditions and rollout history
- Compare desired, created and ready pods on the candidate ReplicaSet
- Read pod and namespace events before changing capacity
- Prove external configuration compatibility before rollback
- 01 Azure AKS: diagnose a stuck rollout before forcing rollback
Automation 1 note
Azure Logic Apps runs accumulate 429 responses and retries
Separate Logic Apps resource limits, connector throttling and destination saturation before adding retries or raising concurrency.
First checks - Freeze one workflow, action, connection and UTC window
- Correlate run and destination request IDs
- Measure retry and concurrency amplification
- Run a bounded canary before staged recovery
- 01 Azure Logic Apps: contain connector throttling before adding retries
Automation 1 note
Feature flag evaluations differ across application instances
Separate the published definition, per-instance ETag, refresh path, targeting context and evaluation telemetry before restarting the fleet or rolling back the flag.
First checks - Freeze the flag, label, change window and previous known state
- Compare the published ETag with healthy and affected instances
- Replay one stable targeting context on a canary
- Correlate FeatureEvaluation events with refresh traces and business metrics
- 01 Azure App Configuration: diagnose feature flag drift before rollback
AI 1 note
An AI agent loses constraints as the conversation grows
Separate system instructions, history, retrieval, tool schemas and tool outputs before raising context limits or changing models.
First checks - Freeze one healthy and one degraded trace for the same journey
- Measure context growth by segment and turn
- Verify approvals and open decisions survive compaction
- Canary the bounded policy before promotion
- 01 AgentOps: diagnose context-window saturation before raising token limits
Cloud 1 note
Azure Managed Redis connections surge while application requests time out
Separate legitimate server pressure, client reconnect amplification, runtime DNS/TLS and release drift before scaling or triggering a failover.
First checks - Freeze the release, replica count and incident window
- Compare connected clients, server load and useful operations
- Measure the connection envelope per role instance
- Canary the client fix with explicit rollback
- 01 Azure Managed Redis: contain a connection storm before scaling or failing over
Infrastructure 1 note
An Azure Monitor alert fires but no action group notifies on-call
Separate signal evaluation, fired-alert processing, scope, filters, schedule and action group delivery before changing the alert rule.
First checks - Freeze one fired alert ID and its expected action groups
- List enabled alert processing rules and match their scopes
- Replay the schedule in its configured time zone
- Generate a new bounded alert after propagation
- 01 Azure Monitor: diagnose suppressed notifications before changing the alert
Cloud 1 note
An Azure Database for PostgreSQL read replica falls behind before promotion
Separate time lag, WAL byte gap, primary log retention, replica replay capacity and target connection readiness before planned switchover or forced promotion.
First checks - Freeze the promotion mode and accepted RPO
- Compare lag seconds, lag bytes and a business marker
- Check primary WAL storage and replica saturation
- Validate virtual endpoints, identity and write path before promotion
- 01 Azure PostgreSQL: diagnose read replica lag before promotion
Automation 1 note
An Azure Automation runbook returns 403 after a managed identity change
Separate cloud sandbox and Hybrid Worker execution, Automation account and VM identities, persisted Az context, control-plane and data-plane permissions before broadening RBAC.
First checks - Freeze one job ID, execution target and denied operation
- Reset Az context and prove the runtime principal
- Match the exact action to role and scope
- Run positive and negative canaries before promotion
- 01 Azure Automation: prove the runbook identity before broadening RBAC
Networking 1 note
An Azure Front Door origin becomes unhealthy and traffic shifts unexpectedly
Separate the health probe contract, DNS/TLS, optional Private Link, backend capacity and routing changes before draining, failing over or restoring traffic.
First checks - Freeze the route, origin group and UTC window
- Read failed probes by origin and POP
- Verify hostname, SNI, certificate and health response
- Prove standby capacity before changing weights
- 01 Azure Front Door: diagnose an unhealthy origin before changing routing
AI 1 note
An AI agent recalls context from another session or team
Separate conversation history, compacted summaries, durable memory, retrieval caches, tool results and runtime identity before enabling persistence for more users.
First checks - Freeze the user, team, environment and session boundaries
- Seed distinct synthetic canaries in two scopes
- Enable each state layer independently
- Repeat under concurrency, expiry and revocation
- 01 AgentOps: validate session and memory isolation before production
Infrastructure 1 note
An Azure Update Manager maintenance window ends with part of the fleet still noncompliant
Separate dynamic-scope membership, schedule orchestration, exhausted window budget, patch failures and pending restarts before forcing an out-of-window installation.
First checks - Freeze the maintenance contract and expected inventory
- Compare matching machines with maintenance and installation runs
- Separate no run, partial run, failed update and pending restart
- Prove service capacity before resuming a bounded set
- 01 Azure Update Manager: diagnose a maintenance window before forcing patching
Cloud 1 note
An Azure App Service certificate is renewed but the old certificate is still served
Separate issuance, domain validation, Key Vault version, App Service import, hostname binding and the certificate observed with SNI before syncing or rebinding.
First checks - Read the served fingerprint with the real hostname and SNI
- Compare candidate SAN and expiration with App Service inventory
- Check Key Vault version and import state when applicable
- Preserve the current binding and rollback thumbprint
- 01 Azure App Service: diagnose TLS renewal before rebinding production
Networking 3 notes
An Azure flow is denied although the NSG allows it
Separate Azure Virtual Network Manager Security Admin rules, network group membership, NSG evaluation and actual service reachability before opening local rules.
First checks - Capture the exact source, destination and port tuple
- List effective Security Admin rules on the VNet
- Run IP Flow Verify and identify the deciding rule
- Confirm the decision in VNet flow logs
- 01 Azure Virtual Network Manager: diagnose a Security Admin rule before changing NSGs
- 02 Azure Network Watcher: validate Flow Logs before opening an NSG rule
- 03 Azure Firewall: diagnose a shadowed rule before opening traffic
Networking 2 notes
An Azure Traffic Manager profile is degraded and failover appears inconsistent across clients
Separate the deployed probe contract, endpoint monitor state, backend evidence, DNS answers, TTL and secondary capacity before disabling the primary endpoint or forcing failover.
First checks - Capture profile, endpoint, probe and TTL configuration
- Replay the exact probe contract from external viewpoints
- Correlate ProbeHealthStatusEvents with backend logs
- Compare DNS answers and prove secondary readiness
- 01 Azure Traffic Manager: diagnose a degraded profile before forcing failover
- 02 Azure Service Health: qualify a regional signal before failover
Automation 1 note
An Azure Automation runbook works in its current runtime but fails in the candidate environment
Separate language version, package resolution, identity, serialization and Hybrid Worker readiness before relinking production runbooks.
First checks - Freeze current and candidate runtime manifests
- Run positive and negative draft tests with the candidate
- Qualify every target Hybrid Worker
- Canary one bounded effect and confirm an empty second run
- 01 Azure Automation: validate a runtime migration before cutting production runbooks over
Infrastructure 1 note
Azure Monitor Managed Prometheus cardinality spikes after an AKS deployment
Separate workspace ingestion pressure, scrape failures, duplicate targets and unbounded labels before dropping metrics, raising quotas or rolling back the release.
First checks - Freeze workspace, cluster, release and UTC window
- Compare active-series and event-ingestion trends
- Bound PromQL to identify the metric and label
- Canary relabeling without breaking alerts
- 01 Azure Monitor Managed Prometheus: contain a cardinality spike before raising quotas
Infrastructure 1 note
Azure VMSS repeatedly replaces or reimages unhealthy instances
Separate the health signal, grace period, orchestration service state, current VMSS model and persistence contract before changing Replace, Reimage or Restart.
First checks - Freeze the repair timeline and useful capacity
- Read automaticRepairsPolicy and orchestration service state
- Replay the health contract on healthy and affected instances
- Build one canary from the current model before resuming
- 01 Azure VMSS: diagnose an automatic repair loop before changing the action
Automation 1 note
A newly pushed ACR image is not visible from every deployment region
Separate asynchronous geo-replication, mutable tags, global endpoint routing, runtime identity and data endpoint reachability before retrying a multi-region rollout.
First checks - Freeze the build digest and push completion time
- List replica state and inspect Resource Health
- Pull the expected digest from every target region
- Correlate ACR events by region, identity and result
- 01 Azure Container Registry: diagnose incomplete image replication before retrying a multi-region rollout
Networking 1 note
An AKS service-to-service flow is blocked after a NetworkPolicy change
Separate source egress, destination ingress, deployed selectors, service endpoints and observed Cilium verdicts before adding an allow-all rule or removing default deny.
First checks - Freeze the source, destination, port and UTC window
- Read the AKS data plane and every policy type
- Compare selectors with deployed pod and namespace labels
- Run positive and negative canaries before promotion
- 01 AKS: diagnose a NetworkPolicy-blocked flow before loosening cluster traffic
Infrastructure 1 note
OpenTelemetry tail sampling reduces volume but incident traces become incomplete or disappear
Separate trace affinity, decision-window sizing, pending-trace memory, late spans, policy behavior and export evidence before widening the sampling rollout.
First checks - Freeze the current and candidate Collector configuration digests
- Prove every span for one trace reaches the same sampling shard
- Replay error, latency and late-span canaries
- Correlate Collector decisions with complete operations in Azure Monitor
- 01 OpenTelemetry Collector: validate tail sampling before losing incident traces
Infrastructure 1 note
Key Vault returns 403 after an Azure RBAC cutover
Separate missed runtime identities, data-plane role mapping, assignment scope and propagation from network failures before granting a broad role or reverting the permission model.
First checks - Confirm the active Key Vault permission model
- Resolve the failing runtime principal object ID
- Compare required operations with direct and inherited data-plane roles
- Correlate 403 responses with Key Vault AuditEvent logs
- 01 Azure Key Vault: migrate access policies to Azure RBAC without breaking workloads
Networking 1 note
Virtual WAN traffic breaks or becomes asymmetric after Routing Intent is enabled
Separate policy scope, effective routes, private prefixes, default-route propagation, security next hop, SNAT and return paths before widening the rollout or deleting Routing Intent.
First checks - Freeze the former route associations and propagations
- Verify every contract prefix on the security next hop
- Probe forward and return paths through each affected hub
- Correlate canaries with Firewall allows and denies
- 01 Azure Virtual WAN: validate Routing Intent before enforcing production inspection
Networking 1 note
AKS pods see intermittent DNS timeouts or SERVFAIL responses
Separate pod resolver behavior, query amplification, the Kubernetes DNS service, CoreDNS endpoints, network policy, upstream forwarding and node-specific paths before restarting CoreDNS.
First checks - Freeze the exact names, errors and affected pod/node cohort
- Compare repeated lookups across nodes and namespaces
- Inspect the kube-dns service, EndpointSlices and CoreDNS pods
- Test cluster names separately from each upstream suffix
- 01 AKS: diagnose intermittent DNS before restarting CoreDNS
Networking 1 note
Azure Firewall IDPS blocks a legitimate flow after a signature is moved to deny
Separate signature evidence, traffic direction, private-range classification, TLS inspection and application outcome before disabling IDPS or adding a broad bypass.
First checks - Freeze the signature override and affected five-tuple
- Query AZFWIdpsSignature for source, destination and action
- Replay one malicious and one legitimate canary
- Return only the signature to Alert if the legitimate flow regresses
- 01 Azure Firewall IDPS: validate a signature before switching it to deny
Automation 1 note
An Azure Automation job may run a version that does not match the approved commit
Tie the source commit, sync job, Draft and Published hashes, no-effect test and execution marker together before rerunning or broadening production scope.
First checks - Freeze the approved commit, script hash and target account
- Read the last sync job and its streams before retrying
- Export and hash Draft and Published separately
- Run Plan and one reversible canary before full execution
- 01 Azure Automation: prove the published version before running a runbook
Infrastructure 1 note
One Azure Monitor log alert creates hundreds of independently evaluated alert instances
Separate real affected cohorts, split-by dimension cardinality, alert lifecycle and notification routing before raising limits, muting the action group or changing the KQL rule.
First checks - Export the deployed scheduled query rule and freeze a UTC window
- Count rows, distinct dimension values and their combinations
- Compare query combinations, active instances and delivered notifications
- Canary stable dimensions through a complete fire-and-resolve cycle
- 01 Azure Monitor: contain log alert fan-out before changing the KQL rule
AI 1 note
An AI gateway returns 429 after a token-limit policy change
Separate APIM policy scope, counter isolation, gateway-local accounting, token estimation, backend throttling and client retries before raising the budget.
First checks - Preserve one rejected request with gateway, region and Retry-After
- Prove whether APIM or the model backend produced the 429
- Verify the effective policy and a non-empty counter key
- Canary the correction without resetting production counters
- 01 Azure API Management AI gateway: diagnose token-limit 429s before raising the quota
Networking 1 note
An Azure application path fails intermittently while manual connectivity tests remain healthy
Use Connection Monitor to preserve source-specific loss and latency, then correlate one failing interval with DNS, effective routes, security decisions, firewall evidence and service health before changing NSGs or UDRs.
First checks - Freeze the exact source, FQDN, port and UTC window
- Compare failed-check percentage and RTT per source
- Run Connection Troubleshoot during the failing interval
- Canary one bounded correction with a refusal control
- 01 Azure Network Watcher: diagnose an intermittent path before changing NSG or UDR
AI 1 note
An AI agent proposes an unrequested production action after reading MCP tool output
Separate authenticated transport from field-level trust, preserve tool-result provenance, enforce write authorization outside the model and replay adversarial results before reopening production actions.
First checks - Freeze the user intent, raw tool result and proposed next call
- Identify which untrusted field influenced the action
- Verify approval and arguments in an external policy gate
- Replay hostile results with writes disabled before a bounded canary
- 01 AgentOps: neutralize indirect prompt injection in MCP tool output before a production action
AI 1 note
An agent stays available through model fallback but changes its tool decisions or refusal behavior
Separate endpoint availability from response-schema, tool, source, refusal, approval, state and trace compatibility before enabling automatic multi-model routing.
First checks - Freeze both model attempts under one logical run ID
- Validate structured output and normalized tool arguments
- Reconcile any unknown primary result before fallback
- Canary new read-only sessions with route-level traces
- 01 AgentOps: validate a fallback model before enabling multi-model routing
Infrastructure 1 note
An AKS Horizontal Pod Autoscaler repeatedly scales up and down
Separate noisy or missing metrics, miscalibrated resource requests, pod warm-up, competing controllers and node capacity before raising replica bounds.
First checks - Freeze two complete scaling cycles
- Compare the HPA status with the metric API
- Measure requests, readiness and downstream capacity per replica
- Canary one behavior change with explicit rollback
- 01 Azure AKS: diagnose HPA oscillation before raising replica limits
Cloud 1 note
An Azure deployment fails because a resource or inherited scope is locked
Separate management locks, RBAC, Policy and deny assignments, identify the exact destructive operation, then bound any unlock window before rerun or rollback.
First checks - Freeze the failed operation, target resource and correlation ID
- Reconstruct inherited CanNotDelete and ReadOnly locks
- Confirm whether the plan really requires a delete
- Recreate and verify the lock before closing the change
- 01 Azure Resource Locks: diagnose a blocked deployment before removing the lock
Networking 2 notes
AKS pods intermittently time out to public dependencies
Separate DNS, the remote dependency, network policy, node connection pressure and the effective Azure egress device before adding SNAT capacity or changing the cluster outbound path.
First checks - Capture the failing pod, node and destination tuple
- Confirm networkProfile.outboundType and the source IP seen downstream
- Correlate per-node SNAT metrics with connection churn
- Canary connection reuse or egress capacity before rollout
- 01 Azure AKS: diagnose outbound SNAT exhaustion before adding a NAT Gateway
- 02 Azure NAT Gateway: diagnose SNAT exhaustion and outbound IP drift
Automation 1 note
A Microsoft Sentinel rule misses a target or automates a response from stale watchlist data
Separate source freshness, snapshot completeness, bulk-update deletion semantics, SearchKey normalization and KQL join behavior before allowing a watchlist to drive incidents or playbooks.
First checks - Prove the alias and source snapshot independently of TimeGenerated
- Compare added, modified and removed keys with the active version
- Reject blank, duplicate or expired SearchKey values
- Shadow the analytics rule before enabling the playbook
- 01 Microsoft Sentinel: validate a watchlist before it drives automated response
Infrastructure 1 note
A self-hosted Azure DevOps agent shows signs of compromise but can still receive production jobs
Stop scheduling without erasing evidence, reconstruct affected runs, map exposed identities and validate a clean replacement before reopening the pool.
First checks - Quarantine the agent and pause production entry points
- Freeze run IDs, commits, agent diagnostics and audit events
- Map credentials and downstream targets used in the incident window
- Canary a clean replacement without protected resources
- 01 Azure DevOps: contain a compromised self-hosted agent before reopening the pool
Networking 1 note
An App Service plan stays online but scaling fails on a VNet Integration subnet
Separate address exhaustion, delayed release, delegation and outbound path failures; calculate worker cohort overlap before retrying scale or migrating the integration.
First checks - Freeze the failed operation and current instance count
- Inventory every app and plan joined to the subnet
- Calculate usable addresses and transient worker overlap
- Canary a larger delegated subnet with explicit rollback
- 01 Azure App Service VNet Integration: diagnose subnet exhaustion before retrying scale
Cloud 1 note
Azure WAF appears to enforce different policy behavior after a deployment
Separate ARM rollout state, gateway, listener and path-rule associations, deployed rule priority and real request variance before waiting or restoring the whole policy.
First checks - Freeze one accepted and one blocked request
- Read policy and gateway provisioning states
- Inventory policy resource IDs at every association scope
- Replay an allowed and denied canary after correction or rollback
- 01 Azure WAF: diagnose inconsistent policy enforcement after deployment
Automation 1 note
A previously allowed Azure deployment is denied after its Policy exemption expires
Tie the first denial to expiresOn, the exact assignment, initiative references and current resource state before renewing the exemption or disabling enforcement.
First checks - Freeze the first denied run and correlation ID
- Read the exemption at its exact scope and compare expiresOn in UTC
- Match assignment and initiative reference IDs
- Canary the same artifact with an out-of-scope denial control
- 01 Azure Policy: diagnose an expired exemption before disabling the assignment
AI 1 note
An AI agent trace contains sensitive data after an instrumentation change
Tie release, span, unsafe field, export route and destination together before disabling observability; contain the narrow path, handle historical copies and canary redaction.
First checks - Freeze trace metadata without copying the raw value
- Map every collector, exporter and destination copy
- Disable the unsafe enrichment while keeping minimum events
- Replay synthetic markers through every input path
- 01 AgentOps: contain sensitive data in agent traces before disabling observability
Automation 1 note
A canceled Azure DevOps deployment leaves the production target in an uncertain partial state
Reconcile the run, agent, Azure control plane and application before rerunning; prove each checkpoint, artifact and side effect, then choose resume, compensation or rollback.
First checks - Freeze every writer targeting the environment
- Pin the run, commit and artifact digest
- Read target-side operations across the cancellation window
- Block rerun while a critical checkpoint remains unknown
- 01 Azure DevOps: diagnose a canceled deployment before rerunning production
AI 1 note
An Azure AI Search indexer succeeds partially and the agent corpus becomes incomplete
Separate source access, incremental tracking, skillset failures and target-index writes before rerunning selected documents or clearing the high-water mark.
First checks - Freeze the execution history and tracking states
- Reconcile failed source IDs with index keys and versions
- Group errors by stage, code and skill
- Canary targeted recovery before any full reset
- 01 Azure AI Search: diagnose a partial indexer failure before resetting it