AI
Azure AI Search: diagnose a partial indexer failure before resetting it
A production runbook for separating source, tracking, skillset and document failures in Azure AI Search before rerunning, targeting recovery or resetting the indexer.
An operations agent starts returning incomplete procedures after the nightly Azure AI Search indexer reports a partially successful run. Most documents were processed, some failed, and the index still answers queries. The tempting response is to reset the indexer and rebuild everything.
That action clears incremental tracking and turns a bounded document failure into a full reprocessing operation. It can increase enrichment load, overwrite working documents and leave the corpus in another mixed state without proving why the first items failed. The running case is a runbook corpus stored in Azure Blob Storage, enriched by a skillset and consumed by a Microsoft Foundry agent. The goal is to restore completeness without losing the evidence required to choose an incremental rerun, targeted recovery, quarantine or full reset.
Freeze the affected run and user impact
Start from one execution, not from the latest status badge. Record the indexer, index, data source, skillset, run window, release or configuration change, processed and failed counts, and the retrieval journeys affected.
incident:
detected_utc: 2026-10-06T14:20:00Z
search_service: search-ops-prod
indexer: runbooks-blob-indexer
target_index: runbooks-v12
datasource: runbooks-prod
skillset: runbooks-enrichment-v7
execution_start_utc: 2026-10-06T13:00:03Z
execution_status: transientFailure
items_processed: 1842
items_failed: 17
impact:
missing_document_keys: pending
affected_journeys: [incident_triage, rollback_lookup]
agent_mode: read_only_with_corpus_warning
change_window:
source_release: docs-2026-10-06.2
indexer_definition_digest: <sha256>
skillset_definition_digest: <sha256> Do not claim that 17 failures mean 17 missing search documents. A failed source item can produce several search documents, overwrite an older document, or fail after part of its enrichment tree has completed. Conversely, the index might still contain a stale version from an earlier successful run. Validate impact by document key and expected version.
Read execution history before changing state
The indexer-level status tells whether the object can run. Each execution has its own result, error list, warning list, item counts and tracking states. Capture the status response before a reset adds a new history entry or a later scheduled run displaces useful evidence.
SERVICE="search-ops-prod"
INDEXER="runbooks-blob-indexer"
API_VERSION="2024-07-01"
TOKEN="<entra-access-token>"
curl -sS \
-H "Authorization: Bearer $TOKEN" \
"https://$SERVICE.search.windows.net/indexers/$INDEXER/status?api-version=$API_VERSION" \
> indexer-status-before-action.json
jq '{
status,
lastResult,
executionHistory: [.executionHistory[] | {
status, startTime, endTime, itemsProcessed, itemsFailed,
initialTrackingState, finalTrackingState, errors, warnings
}]
}' indexer-status-before-action.json Keep the raw response in restricted incident evidence. Group the failures by error code, message, document key and skill name. Repeated failures on the same stage suggest a deterministic mapping or enrichment defect; scattered timeouts across stages suggest dependency or capacity pressure.
Warnings matter when they remove content that an agent relies on. A truncated field, unsupported document, missing skill input or projection warning can preserve a green run while degrading retrieval.
Separate the four failure planes
Treat the pipeline as four contracts:
1. Source access
Identity, network, throttling, object availability and change detection
2. Incremental tracking
High-water mark, timestamps, delete detection and previous run completion
3. Transformation
Parsing, field mappings, skill inputs, custom skill responses and projections
4. Target index
Schema compatibility, key validity, vector dimensions, document size and write capacity
Evidence required for every failed item
source identifier -> document key -> failed stage -> error code
-> expected version -> indexed version -> retrieval visibility A 403 from the source is not repaired by changing an index field. A vector dimension mismatch is not repaired by opening a firewall. A custom skill timeout does not prove that the source document is corrupt. Keep these boundaries explicit so the smallest corrective action remains visible.
For custom skills, correlate the indexer execution window with dependency logs using the operation or document identifier. Measure response status and latency distribution, not only average duration. One slow class of documents can consume the execution window while ordinary documents succeed.
Prove what is missing from the corpus
Build a reconciliation set from the expected source manifest, not from the failed counter alone. Each source item should have a stable key, content version, access scope and expected index projection.
{
"sourceId": "runbooks/network/dns-042.md",
"sourceVersion": "sha256:<source-digest>",
"documentKey": "runbooks-network-dns-042",
"expectedProjection": "runbook-chunks",
"expectedChunkCount": 6,
"indexedVersion": "sha256:<indexed-digest>",
"indexedChunkCount": 4,
"retrievalState": "quarantined",
"failureStage": "skill.custom-metadata",
"errorCode": "WebApiSkillResponseError"
} Query the index by the exact keys and version fields used by the application. Then run representative retrieval probes for the affected operational journeys. A matching document count is insufficient if ACL metadata, language, source version or a critical chunk is absent.
While completeness is unknown, make the degradation visible to the consumer. Exclude known-bad keys through a filter, route sensitive journeys to the last qualified index alias, or put the agent in read-only mode with a corpus freshness warning. Do not let fluent answers hide an incomplete evidence base.
Choose the narrowest recovery
A normal rerun preserves incremental tracking and is appropriate for transient failures that the source and skill dependencies can now process. Scheduled indexers also retry work over time, so prove whether another bounded run is enough before changing state.
Targeted document recovery is preferable when the failing keys are known and the selective reset capability is approved for the environment. Validate how source identifiers map to index keys, especially when one blob produces multiple documents. A wrong key can report a clean operation while leaving the failed content untouched.
Resetting the full indexer clears its high-water mark and causes the next run to reprocess all documents. It is justified only when tracking state is invalid, a pipeline-wide change requires complete regeneration, or reconciliation proves that the affected set cannot be bounded. A reset does not itself run the indexer, cannot be undone, and does not remove orphaned documents that no longer exist in the source.
normal_rerun:
when: [failures_are_transient, tracking_state_is_credible]
targeted_recovery:
when: [exact_keys_are_known, key_mapping_is_verified, selective_reset_is_approved]
controls: [snapshot_status, canary_two_documents, resume_regular_tracking]
full_reset:
when: [tracking_state_is_invalid, pipeline_wide_regeneration_is_required]
requires:
- capacity_window
- source_and_skill_dependency_readiness
- qualified_index_or_alias_return_path
- orphan_cleanup_plan Do not raise maxFailedItems as the first fix. That setting changes whether processing continues; it does not make failed content valid. Use it only with an explicit completeness objective, alert threshold and quarantine path.
Canary the correction outside the production corpus
Reproduce one healthy item and two representative failed items against a disposable index or a versioned candidate index. Keep source bytes, indexer definition, skillset, index schema and dependency versions pinned.
The canary must prove that expected chunks exist, source version and access metadata are correct, no stale projection remains, the agent retrieves the corrected evidence, an out-of-scope document stays unavailable, and processing latency remains inside the operating window.
Promote through an index alias or equivalent application configuration only after these checks pass. If the existing index is repaired in place, retain an export of definitions and the previous qualified index until reconciliation completes.
Validate recovery and keep a rollback path
After the bounded run, compare the failed key set, item counts, tracking transition, indexed versions and retrieval probes. Run one additional incremental execution: it should process only genuinely new or changed content, not replay the same corpus unexpectedly.
Promote
Failed set is empty or every exclusion is explicitly accepted
Expected versions and projections match the source manifest
Retrieval and access-control probes pass
Next incremental run resumes from a credible tracking state
Keep quarantine
Some document keys or projections remain unexplained
Skill dependency behavior is still intermittent
Corpus freshness cannot be stated to consumers
Roll back
Switch the application to the last qualified index or alias
Restore the previous indexer, skillset and schema definitions
Keep new ingestion paused for the affected path
Preserve failed source items and execution evidence for replay Rollback is not another blind reset. It restores a known corpus and known pipeline definitions while the failed source set remains isolated. If the application cannot switch indexes, reduce the affected retrieval scope and disable write-capable agent journeys until evidence is complete.
Conclusion
A partially successful Azure AI Search run is a corpus integrity incident, not merely a scheduler error. The useful diagnosis links the source item, tracking state, failed stage, target document and retrieval impact before changing indexer state.
The decision is deliberately narrow: rerun when the failure is transient, recover selected documents when keys are proven, quarantine when completeness is unknown, and reset only when incremental state or the whole pipeline truly requires regeneration. Production is recovered when the next incremental run behaves normally and the agent retrieves a complete, authorized and versioned corpus.