Failure Atlas

Start from the symptom, not the category.

A diagnostic map for operational situations: 502 errors, private DNS drift, WAF blocks, Terraform locks, automation guardrails and private AI controls.

Cloud 4 notes

Internal APIM returns an error on a private API

Correlate Application Gateway/WAF and APIM logs, then separate DNS, TLS, policy, identity and private backend reachability before changing policies or opening access.

First checks
  • Check whether WAF blocked the request
  • Confirm APIM received the same path
  • Validate backend DNS and TLS from the APIM path
  • Replay with a correlation ID
Networking 4 notes

Private Endpoint name still resolves publicly

Confirm the CNAME chain, Private DNS Zone association and hybrid forwarding from the consuming network.

First checks
  • Run nslookup from the workload network
  • Check privatelink CNAME
  • Verify Private DNS Zone links and forwarders
Cloud 3 notes

Azure Storage private endpoint returns 403, times out or produces no request logs

Separate Storage subresource DNS, Private Endpoint approval, firewall rules, runtime identity and Storage logs before opening public access or broadening RBAC.

First checks
  • Resolve the exact Storage subresource from the workload network
  • Check Private Endpoint status and private DNS zone group
  • Replay with a client request ID
  • Correlate Storage logs for 403, caller IP and requester identity
Cloud 3 notes

Azure SQL private endpoint returns timeouts, firewall errors or no SQL logs

Separate SQL private DNS, Private Endpoint state, public access, firewall rules, runtime identity and SQL diagnostics before changing schema, code or broad permissions.

First checks
  • Resolve the SQL FQDN from the workload network
  • Check Private Endpoint status and privatelink.database.windows.net records
  • Replay with the real runtime identity
  • Correlate SQL diagnostics for firewall and login errors
Cloud 3 notes

Azure Service Bus private endpoint times out, denies access or lets backlog grow

Separate Service Bus private DNS, Private Endpoint state, public access, managed identity or SAS, queue metrics and processing logs before touching queues or redeploying consumers.

First checks
  • Resolve the Service Bus FQDN from the workload network
  • Check Private Endpoint approval and public network access
  • Verify the real sender or receiver identity
  • Correlate backlog, dead-letter and Service Bus errors
Cloud 4 notes

A synthetic probe fails on an Azure private path

Separate DNS, TLS, Application Gateway health, WAF blocks and runner network before changing routing or application code.

First checks
  • Resolve the hostname from the probe network
  • Check TLS/SNI with the real hostname
  • Correlate probe run with WAF and gateway logs
Networking 5 notes

Azure traffic leaves through one path and returns through another

Compare DNS target, effective routes, route table associations, firewall evidence, NAT identity and return path before changing UDRs or bypassing inspection.

First checks
  • Resolve the destination from the source path
  • Compare effective routes on source and destination NICs
  • Check firewall or appliance logs for both directions
  • Confirm NAT or outbound source identity
Cloud 3 notes

Azure Container Apps private ingress fails or reaches the wrong revision

Separate private DNS, Application Gateway handoff, Container Apps ingress mode, revision traffic and console logs before rolling back or changing traffic weights.

First checks
  • Resolve the hostname from the caller network
  • Check ingress target port and active revisions
  • Correlate system and console logs
Cloud 3 notes

AKS private ingress returns 502 or reaches no service endpoints

Separate private DNS, Application Gateway health, ingress controller routing, Kubernetes service selectors, endpoint slices and pod readiness before rolling back a deployment.

First checks
  • Resolve the hostname from the caller network
  • Check Application Gateway backend health and host header
  • Verify ingress, service and endpoint slices
  • Correlate controller and application logs
Infrastructure 1 note

An AKS node pool upgrade is blocked while draining a pod protected by a PDB

Separate unavailable replicas, readiness failures, pod placement, surge capacity and an over-constrained PodDisruptionBudget before deleting the safeguard or forcing the upgrade.

First checks
  • Capture the blocked node and eviction events
  • Inspect PDB disruptionsAllowed and selectors
  • Verify replacement capacity, replica readiness and topology
  • Resume only with positive disruption headroom
Cloud 3 notes

Azure Functions private HTTP endpoint returns 403, 503 or no request logs

Separate private DNS, Private Endpoint reachability, access restrictions, Functions runtime state, private storage and Application Insights evidence before redeploying code or opening public access.

First checks
  • Resolve the hostname from the caller network
  • Replay with a correlation ID
  • Check Function App access restrictions and Private Endpoint status
  • Correlate requests, traces and exceptions
Cloud 5 notes

Azure WAF blocks a legitimate request

Start from blocked requests, rule ID and URI before deciding between exclusion, custom rule or application fix.

First checks
  • List blocked URIs in KQL
  • Identify ruleId and match field
  • Validate false-positive scope
Automation 4 notes

Terraform state lock is stuck

Prove that no apply is still running before using force-unlock, then restart with a clean plan.

First checks
  • Identify lock owner
  • Check CI job status
  • Run plan after unlock
Infrastructure 2 notes

Azure Monitor fires an alert storm after deployment

Separate real service impact, noisy dimensions, threshold drift, action group behavior and rollback before silencing notifications or changing rules.

First checks
  • Group fired alerts by rule and target
  • Compare first alert time with the deployment window
  • Check service logs before changing thresholds
  • Keep one user-symptom alert active
Infrastructure 8 notes

A secret rotation, federated identity or managed identity change breaks an application or pipeline consumer

Separate preparation, cutover, revocation, workload identity federation and managed identity diagnostics; validate the real execution identity, private path and authentication errors before deleting the old value or broadening access.

First checks
  • List real consumers
  • Verify the runtime identity and vault read access
  • Check OIDC claims or private DNS depending on the path
  • Watch 401/403/500, sign-in failures or Key Vault denials
Automation 1 note

An Azure Functions poison queue keeps growing

Separate contract failures, transient dependencies, identity, resource pressure and partial side effects before replaying Storage Queue messages.

First checks
  • Peek a redacted sample without consuming messages
  • Correlate message IDs with invocation and dependency logs
  • Verify downstream state and idempotency before replay
Cloud 1 note

An Azure Service Bus dead-letter queue keeps growing

Classify dead-letter reasons, correlate consumer and downstream evidence, verify idempotency, then use a manifest, canary and stop conditions before replay.

First checks
  • Record the exact queue or subscription and DLQ growth window
  • Peek a redacted sample grouped by dead-letter reason
  • Check consumer errors, lock loss and downstream state
  • Prove idempotency before a one-message canary
Automation 1 note

An Azure Policy remediation changes unexpected production resources

Pin the definition and assignment, materialize the eligible target set, verify the managed identity, then use a canary and bounded batches before expanding or compensating.

First checks
  • Identify the exact definition reference and assignment parameters
  • Export non-compliant target resource IDs
  • Review the assignment identity and effective RBAC scope
  • Correlate remediation deployments with Activity Log
Cloud 1 note

An Azure Event Hubs consumer keeps falling behind

Separate namespace throttling, hot partitions, consumer processing, checkpoint progress and downstream saturation before adding capacity or replaying events.

First checks
  • Compare ingress, egress and throttled requests
  • Measure progress and checkpoint age per partition
  • Check consumer ownership, errors and downstream latency
  • Prove sequence bounds and idempotency before replay
Automation 1 note

An Azure Automation Hybrid Runbook Worker stops picking up jobs

Separate dispatch, heartbeat, extension health, outbound HTTPS, capacity, runtime and identity before retrying a production job.

First checks
  • Preserve the job ID, streams and UTC window
  • Compare HybridWorkerPing and extension state across the group
  • Read local worker logs before restarting services
  • Complete a no-side-effect canary before retry
Networking 1 note

Azure Bastion cannot open an SSH or RDP session

Separate the client HTTPS path, Bastion health, AzureBastionSubnet rules, routing, target NSGs, guest listener and identity before exposing administration ports.

First checks
  • Freeze one failed attempt with UTC time and target private IP
  • Check Bastion provisioning state and the complete AzureBastionSubnet rule set
  • Inspect target NIC effective NSGs and routes
  • Verify the guest listener and authentication separately
Cloud 1 note

An application exhausts its Azure SQL connection pool

Separate local connection retention, SQL sessions and waits, network failures, managed identity token acquisition and autoscaling multiplication before raising the pool limit or database tier.

First checks
  • Tie one timeout to a revision and role instance
  • Calculate the global pool capacity envelope
  • Compare SQL sessions with active requests and waits
  • Canary the correction with stop and rollback conditions
Cloud 1 note

An ACR image fails signature verification before AKS promotion

Separate mutable tags, missing signatures, untrusted publisher identity, trust-store rotation and ACR access before bypassing admission or deploying an unverified image.

First checks
  • Resolve the candidate tag to one immutable digest
  • Inventory signatures with Notation
  • Verify repository scope, trust store and publisher identity
  • Compare the verified digest with the AKS runtime imageID
Infrastructure 1 note

An AKS Deployment rollout stalls before the new revision becomes available

Locate the first blocked object across Deployment, ReplicaSet, admission, scheduling, image startup and readiness before deleting pods, extending timeouts or forcing rollback.

First checks
  • Capture Deployment conditions and rollout history
  • Compare desired, created and ready pods on the candidate ReplicaSet
  • Read pod and namespace events before changing capacity
  • Prove external configuration compatibility before rollback
Automation 1 note

Azure Logic Apps runs accumulate 429 responses and retries

Separate Logic Apps resource limits, connector throttling and destination saturation before adding retries or raising concurrency.

First checks
  • Freeze one workflow, action, connection and UTC window
  • Correlate run and destination request IDs
  • Measure retry and concurrency amplification
  • Run a bounded canary before staged recovery
Automation 1 note

Feature flag evaluations differ across application instances

Separate the published definition, per-instance ETag, refresh path, targeting context and evaluation telemetry before restarting the fleet or rolling back the flag.

First checks
  • Freeze the flag, label, change window and previous known state
  • Compare the published ETag with healthy and affected instances
  • Replay one stable targeting context on a canary
  • Correlate FeatureEvaluation events with refresh traces and business metrics
AI 1 note

An AI agent loses constraints as the conversation grows

Separate system instructions, history, retrieval, tool schemas and tool outputs before raising context limits or changing models.

First checks
  • Freeze one healthy and one degraded trace for the same journey
  • Measure context growth by segment and turn
  • Verify approvals and open decisions survive compaction
  • Canary the bounded policy before promotion
Cloud 1 note

Azure Managed Redis connections surge while application requests time out

Separate legitimate server pressure, client reconnect amplification, runtime DNS/TLS and release drift before scaling or triggering a failover.

First checks
  • Freeze the release, replica count and incident window
  • Compare connected clients, server load and useful operations
  • Measure the connection envelope per role instance
  • Canary the client fix with explicit rollback
Infrastructure 1 note

An Azure Monitor alert fires but no action group notifies on-call

Separate signal evaluation, fired-alert processing, scope, filters, schedule and action group delivery before changing the alert rule.

First checks
  • Freeze one fired alert ID and its expected action groups
  • List enabled alert processing rules and match their scopes
  • Replay the schedule in its configured time zone
  • Generate a new bounded alert after propagation
Cloud 1 note

An Azure Database for PostgreSQL read replica falls behind before promotion

Separate time lag, WAL byte gap, primary log retention, replica replay capacity and target connection readiness before planned switchover or forced promotion.

First checks
  • Freeze the promotion mode and accepted RPO
  • Compare lag seconds, lag bytes and a business marker
  • Check primary WAL storage and replica saturation
  • Validate virtual endpoints, identity and write path before promotion
Automation 1 note

An Azure Automation runbook returns 403 after a managed identity change

Separate cloud sandbox and Hybrid Worker execution, Automation account and VM identities, persisted Az context, control-plane and data-plane permissions before broadening RBAC.

First checks
  • Freeze one job ID, execution target and denied operation
  • Reset Az context and prove the runtime principal
  • Match the exact action to role and scope
  • Run positive and negative canaries before promotion
Networking 1 note

An Azure Front Door origin becomes unhealthy and traffic shifts unexpectedly

Separate the health probe contract, DNS/TLS, optional Private Link, backend capacity and routing changes before draining, failing over or restoring traffic.

First checks
  • Freeze the route, origin group and UTC window
  • Read failed probes by origin and POP
  • Verify hostname, SNI, certificate and health response
  • Prove standby capacity before changing weights
AI 1 note

An AI agent recalls context from another session or team

Separate conversation history, compacted summaries, durable memory, retrieval caches, tool results and runtime identity before enabling persistence for more users.

First checks
  • Freeze the user, team, environment and session boundaries
  • Seed distinct synthetic canaries in two scopes
  • Enable each state layer independently
  • Repeat under concurrency, expiry and revocation
Infrastructure 1 note

An Azure Update Manager maintenance window ends with part of the fleet still noncompliant

Separate dynamic-scope membership, schedule orchestration, exhausted window budget, patch failures and pending restarts before forcing an out-of-window installation.

First checks
  • Freeze the maintenance contract and expected inventory
  • Compare matching machines with maintenance and installation runs
  • Separate no run, partial run, failed update and pending restart
  • Prove service capacity before resuming a bounded set
Cloud 1 note

An Azure App Service certificate is renewed but the old certificate is still served

Separate issuance, domain validation, Key Vault version, App Service import, hostname binding and the certificate observed with SNI before syncing or rebinding.

First checks
  • Read the served fingerprint with the real hostname and SNI
  • Compare candidate SAN and expiration with App Service inventory
  • Check Key Vault version and import state when applicable
  • Preserve the current binding and rollback thumbprint
Networking 3 notes

An Azure flow is denied although the NSG allows it

Separate Azure Virtual Network Manager Security Admin rules, network group membership, NSG evaluation and actual service reachability before opening local rules.

First checks
  • Capture the exact source, destination and port tuple
  • List effective Security Admin rules on the VNet
  • Run IP Flow Verify and identify the deciding rule
  • Confirm the decision in VNet flow logs
Networking 2 notes

An Azure Traffic Manager profile is degraded and failover appears inconsistent across clients

Separate the deployed probe contract, endpoint monitor state, backend evidence, DNS answers, TTL and secondary capacity before disabling the primary endpoint or forcing failover.

First checks
  • Capture profile, endpoint, probe and TTL configuration
  • Replay the exact probe contract from external viewpoints
  • Correlate ProbeHealthStatusEvents with backend logs
  • Compare DNS answers and prove secondary readiness
Automation 1 note

An Azure Automation runbook works in its current runtime but fails in the candidate environment

Separate language version, package resolution, identity, serialization and Hybrid Worker readiness before relinking production runbooks.

First checks
  • Freeze current and candidate runtime manifests
  • Run positive and negative draft tests with the candidate
  • Qualify every target Hybrid Worker
  • Canary one bounded effect and confirm an empty second run
Infrastructure 1 note

Azure Monitor Managed Prometheus cardinality spikes after an AKS deployment

Separate workspace ingestion pressure, scrape failures, duplicate targets and unbounded labels before dropping metrics, raising quotas or rolling back the release.

First checks
  • Freeze workspace, cluster, release and UTC window
  • Compare active-series and event-ingestion trends
  • Bound PromQL to identify the metric and label
  • Canary relabeling without breaking alerts
Infrastructure 1 note

Azure VMSS repeatedly replaces or reimages unhealthy instances

Separate the health signal, grace period, orchestration service state, current VMSS model and persistence contract before changing Replace, Reimage or Restart.

First checks
  • Freeze the repair timeline and useful capacity
  • Read automaticRepairsPolicy and orchestration service state
  • Replay the health contract on healthy and affected instances
  • Build one canary from the current model before resuming
Automation 1 note

A newly pushed ACR image is not visible from every deployment region

Separate asynchronous geo-replication, mutable tags, global endpoint routing, runtime identity and data endpoint reachability before retrying a multi-region rollout.

First checks
  • Freeze the build digest and push completion time
  • List replica state and inspect Resource Health
  • Pull the expected digest from every target region
  • Correlate ACR events by region, identity and result
Networking 1 note

An AKS service-to-service flow is blocked after a NetworkPolicy change

Separate source egress, destination ingress, deployed selectors, service endpoints and observed Cilium verdicts before adding an allow-all rule or removing default deny.

First checks
  • Freeze the source, destination, port and UTC window
  • Read the AKS data plane and every policy type
  • Compare selectors with deployed pod and namespace labels
  • Run positive and negative canaries before promotion
Infrastructure 1 note

OpenTelemetry tail sampling reduces volume but incident traces become incomplete or disappear

Separate trace affinity, decision-window sizing, pending-trace memory, late spans, policy behavior and export evidence before widening the sampling rollout.

First checks
  • Freeze the current and candidate Collector configuration digests
  • Prove every span for one trace reaches the same sampling shard
  • Replay error, latency and late-span canaries
  • Correlate Collector decisions with complete operations in Azure Monitor
Infrastructure 1 note

Key Vault returns 403 after an Azure RBAC cutover

Separate missed runtime identities, data-plane role mapping, assignment scope and propagation from network failures before granting a broad role or reverting the permission model.

First checks
  • Confirm the active Key Vault permission model
  • Resolve the failing runtime principal object ID
  • Compare required operations with direct and inherited data-plane roles
  • Correlate 403 responses with Key Vault AuditEvent logs
Networking 1 note

Virtual WAN traffic breaks or becomes asymmetric after Routing Intent is enabled

Separate policy scope, effective routes, private prefixes, default-route propagation, security next hop, SNAT and return paths before widening the rollout or deleting Routing Intent.

First checks
  • Freeze the former route associations and propagations
  • Verify every contract prefix on the security next hop
  • Probe forward and return paths through each affected hub
  • Correlate canaries with Firewall allows and denies
Networking 1 note

AKS pods see intermittent DNS timeouts or SERVFAIL responses

Separate pod resolver behavior, query amplification, the Kubernetes DNS service, CoreDNS endpoints, network policy, upstream forwarding and node-specific paths before restarting CoreDNS.

First checks
  • Freeze the exact names, errors and affected pod/node cohort
  • Compare repeated lookups across nodes and namespaces
  • Inspect the kube-dns service, EndpointSlices and CoreDNS pods
  • Test cluster names separately from each upstream suffix
Networking 1 note

Azure Firewall IDPS blocks a legitimate flow after a signature is moved to deny

Separate signature evidence, traffic direction, private-range classification, TLS inspection and application outcome before disabling IDPS or adding a broad bypass.

First checks
  • Freeze the signature override and affected five-tuple
  • Query AZFWIdpsSignature for source, destination and action
  • Replay one malicious and one legitimate canary
  • Return only the signature to Alert if the legitimate flow regresses
Automation 1 note

An Azure Automation job may run a version that does not match the approved commit

Tie the source commit, sync job, Draft and Published hashes, no-effect test and execution marker together before rerunning or broadening production scope.

First checks
  • Freeze the approved commit, script hash and target account
  • Read the last sync job and its streams before retrying
  • Export and hash Draft and Published separately
  • Run Plan and one reversible canary before full execution
Infrastructure 1 note

One Azure Monitor log alert creates hundreds of independently evaluated alert instances

Separate real affected cohorts, split-by dimension cardinality, alert lifecycle and notification routing before raising limits, muting the action group or changing the KQL rule.

First checks
  • Export the deployed scheduled query rule and freeze a UTC window
  • Count rows, distinct dimension values and their combinations
  • Compare query combinations, active instances and delivered notifications
  • Canary stable dimensions through a complete fire-and-resolve cycle
AI 1 note

An AI gateway returns 429 after a token-limit policy change

Separate APIM policy scope, counter isolation, gateway-local accounting, token estimation, backend throttling and client retries before raising the budget.

First checks
  • Preserve one rejected request with gateway, region and Retry-After
  • Prove whether APIM or the model backend produced the 429
  • Verify the effective policy and a non-empty counter key
  • Canary the correction without resetting production counters
Networking 1 note

An Azure application path fails intermittently while manual connectivity tests remain healthy

Use Connection Monitor to preserve source-specific loss and latency, then correlate one failing interval with DNS, effective routes, security decisions, firewall evidence and service health before changing NSGs or UDRs.

First checks
  • Freeze the exact source, FQDN, port and UTC window
  • Compare failed-check percentage and RTT per source
  • Run Connection Troubleshoot during the failing interval
  • Canary one bounded correction with a refusal control
AI 1 note

An AI agent proposes an unrequested production action after reading MCP tool output

Separate authenticated transport from field-level trust, preserve tool-result provenance, enforce write authorization outside the model and replay adversarial results before reopening production actions.

First checks
  • Freeze the user intent, raw tool result and proposed next call
  • Identify which untrusted field influenced the action
  • Verify approval and arguments in an external policy gate
  • Replay hostile results with writes disabled before a bounded canary
AI 1 note

An agent stays available through model fallback but changes its tool decisions or refusal behavior

Separate endpoint availability from response-schema, tool, source, refusal, approval, state and trace compatibility before enabling automatic multi-model routing.

First checks
  • Freeze both model attempts under one logical run ID
  • Validate structured output and normalized tool arguments
  • Reconcile any unknown primary result before fallback
  • Canary new read-only sessions with route-level traces
Infrastructure 1 note

An AKS Horizontal Pod Autoscaler repeatedly scales up and down

Separate noisy or missing metrics, miscalibrated resource requests, pod warm-up, competing controllers and node capacity before raising replica bounds.

First checks
  • Freeze two complete scaling cycles
  • Compare the HPA status with the metric API
  • Measure requests, readiness and downstream capacity per replica
  • Canary one behavior change with explicit rollback
Cloud 1 note

An Azure deployment fails because a resource or inherited scope is locked

Separate management locks, RBAC, Policy and deny assignments, identify the exact destructive operation, then bound any unlock window before rerun or rollback.

First checks
  • Freeze the failed operation, target resource and correlation ID
  • Reconstruct inherited CanNotDelete and ReadOnly locks
  • Confirm whether the plan really requires a delete
  • Recreate and verify the lock before closing the change
Networking 2 notes

AKS pods intermittently time out to public dependencies

Separate DNS, the remote dependency, network policy, node connection pressure and the effective Azure egress device before adding SNAT capacity or changing the cluster outbound path.

First checks
  • Capture the failing pod, node and destination tuple
  • Confirm networkProfile.outboundType and the source IP seen downstream
  • Correlate per-node SNAT metrics with connection churn
  • Canary connection reuse or egress capacity before rollout
Automation 1 note

A Microsoft Sentinel rule misses a target or automates a response from stale watchlist data

Separate source freshness, snapshot completeness, bulk-update deletion semantics, SearchKey normalization and KQL join behavior before allowing a watchlist to drive incidents or playbooks.

First checks
  • Prove the alias and source snapshot independently of TimeGenerated
  • Compare added, modified and removed keys with the active version
  • Reject blank, duplicate or expired SearchKey values
  • Shadow the analytics rule before enabling the playbook
Infrastructure 1 note

A self-hosted Azure DevOps agent shows signs of compromise but can still receive production jobs

Stop scheduling without erasing evidence, reconstruct affected runs, map exposed identities and validate a clean replacement before reopening the pool.

First checks
  • Quarantine the agent and pause production entry points
  • Freeze run IDs, commits, agent diagnostics and audit events
  • Map credentials and downstream targets used in the incident window
  • Canary a clean replacement without protected resources
Networking 1 note

An App Service plan stays online but scaling fails on a VNet Integration subnet

Separate address exhaustion, delayed release, delegation and outbound path failures; calculate worker cohort overlap before retrying scale or migrating the integration.

First checks
  • Freeze the failed operation and current instance count
  • Inventory every app and plan joined to the subnet
  • Calculate usable addresses and transient worker overlap
  • Canary a larger delegated subnet with explicit rollback
Cloud 1 note

Azure WAF appears to enforce different policy behavior after a deployment

Separate ARM rollout state, gateway, listener and path-rule associations, deployed rule priority and real request variance before waiting or restoring the whole policy.

First checks
  • Freeze one accepted and one blocked request
  • Read policy and gateway provisioning states
  • Inventory policy resource IDs at every association scope
  • Replay an allowed and denied canary after correction or rollback
Automation 1 note

A previously allowed Azure deployment is denied after its Policy exemption expires

Tie the first denial to expiresOn, the exact assignment, initiative references and current resource state before renewing the exemption or disabling enforcement.

First checks
  • Freeze the first denied run and correlation ID
  • Read the exemption at its exact scope and compare expiresOn in UTC
  • Match assignment and initiative reference IDs
  • Canary the same artifact with an out-of-scope denial control
AI 1 note

An AI agent trace contains sensitive data after an instrumentation change

Tie release, span, unsafe field, export route and destination together before disabling observability; contain the narrow path, handle historical copies and canary redaction.

First checks
  • Freeze trace metadata without copying the raw value
  • Map every collector, exporter and destination copy
  • Disable the unsafe enrichment while keeping minimum events
  • Replay synthetic markers through every input path
Automation 1 note

A canceled Azure DevOps deployment leaves the production target in an uncertain partial state

Reconcile the run, agent, Azure control plane and application before rerunning; prove each checkpoint, artifact and side effect, then choose resume, compensation or rollback.

First checks
  • Freeze every writer targeting the environment
  • Pin the run, commit and artifact digest
  • Read target-side operations across the cancellation window
  • Block rerun while a critical checkpoint remains unknown
AI 1 note

An Azure AI Search indexer succeeds partially and the agent corpus becomes incomplete

Separate source access, incremental tracking, skillset failures and target-index writes before rerunning selected documents or clearing the high-water mark.

First checks
  • Freeze the execution history and tracking states
  • Reconcile failed source IDs with index keys and versions
  • Group errors by stage, code and skill
  • Canary targeted recovery before any full reset