Automation
Azure DevOps: diagnose a canceled deployment before rerunning production
A production runbook for reconstructing the effects of a canceled Azure DevOps deployment, proving target state, and choosing an idempotent resume, compensation or rollback.
An Azure DevOps pipeline is canceled during a production deployment. The run says Canceled, but the new image may already be referenced, an application setting has changed, and traffic may not have moved yet. Rerunning feels like the quickest recovery. It can also layer a second release over a target state nobody has established.
The working scenario is a deployment job that releases an Azure API, applies configuration, and runs a smoke test. Cancellation lands between those steps. This runbook drives one bounded decision: resume from a proven checkpoint, compensate a partial effect, restore the last known-good release, or keep writes frozen while the target remains ambiguous.
Freeze the run before another write
Record the organization, project, pipeline, run ID, stage, environment, commit and artifact. Capture the UTC time when cancellation was requested and the timestamp of the last evidence produced by the agent. Do not replace these identifiers with “latest run” or “current release.”
incident: INC-DEPLOY-20261005-01
pipeline: deploy-orders-api
run_id: 18427
stage: production
environment: prod-orders
source_commit: 8d2c1f4
artifact:
name: orders-api
digest: sha256:<digest>
cancellation_requested_utc: 2026-10-05T15:42:18Z
last_pipeline_evidence_utc: 2026-10-05T15:42:31Z
change_window_end_utc: 2026-10-05T16:30:00Z
writers_frozen:
- scheduled_release
- manual_hotfix
- configuration_automation
decision_owner: platform-on-call Freeze the other writers that target the environment: concurrent pipelines, scheduled jobs and manual changes. The freeze protects the investigation. It does not assume the process already in flight has stopped.
Separate run status from target state
Canceled is the Azure DevOps run outcome, not a distributed transaction rollback. A remote command can succeed before the agent receives the signal. A step guarded by always() can still run during the cancellation grace period. Conversely, a disconnected agent may never upload its final result.
Azure DevOps
Run, stage, job and task: queued | inProgress | completed
Final result: canceled
Agent
Cancellation signal received or not
Child process stopped or still active
Last known log and heartbeat
Azure target
ARM operation accepted, completed or running
Configuration actually written
Version or revision receiving traffic
Migrations or messages already produced
Application
Version observed by probes
Health, errors and business side effects
Compatibility with schema and dependencies The run provides a command timeline. The target provides operational truth. Reconcile both before choosing a recovery action.
Capture the commit, artifact and timeline
Environment deployment history helps identify the pipelines and commits that targeted production. Combine it with run details and published artifacts. A branch-based rerun can rebuild different bytes, so the deployed digest remains the stronger reference.
ORG="https://dev.azure.com/example"
PROJECT="platform-prod"
RUN_ID="18427"
az pipelines runs show --org "$ORG" --project "$PROJECT" --id "$RUN_ID" --query '{id:id,name:name,state:state,result:result,branch:sourceBranch,commit:sourceVersion,created:createdDate,finished:finishedDate}' --output json
az pipelines runs artifact list --org "$ORG" --project "$PROJECT" --run-id "$RUN_ID" --output table Classify every task as not started, started without completion evidence, completed with evidence, or completed but not validated. A script success line is insufficient for an asynchronous operation. Keep its operation ID and read the terminal state from Azure.
Reconstruct effects from the target
List every possible pipeline write: ARM deployment, image reference, settings, secret references, traffic weights, migrations, messages or cache invalidation. Read the effective target state and timestamp for each write. Do not issue corrective commands during this pass.
SUBSCRIPTION="<subscription-id>"
RESOURCE_GROUP="rg-orders-prod"
START="2026-10-05T15:30:00Z"
END="2026-10-05T16:00:00Z"
az monitor activity-log list --subscription "$SUBSCRIPTION" --resource-group "$RESOURCE_GROUP" --start-time "$START" --end-time "$END" --query '[].{time:eventTimestamp,status:status.localizedValue,operation:operationName.localizedValue,caller:caller,correlationId:correlationId,resourceId:resourceId}' --output table
az deployment group list --resource-group "$RESOURCE_GROUP" --query '[].{name:name,state:properties.provisioningState,time:properties.timestamp,correlationId:properties.correlationId}' --output table The Activity Log proves control-plane operations. It does not, by itself, prove the served version, a data-plane write or a successful business outcome. Add service state, service logs and a read-only probe.
Correlate identity and timing in KQL
When the Activity Log is exported to Log Analytics, filter by the incident window, resource group, service connection principal and correlation IDs. Look for writes after the cancellation request as well: they indicate either an operation that continued or another change source.
let Start = datetime(2026-10-05T15:30:00Z);
let CancelRequested = datetime(2026-10-05T15:42:18Z);
let End = datetime(2026-10-05T16:00:00Z);
let DeploymentCaller = "<service-connection-principal>";
AzureActivity
| where TimeGenerated between (Start .. End)
| where ResourceGroup =~ "rg-orders-prod"
| where Caller =~ DeploymentCaller
| where OperationNameValue endswith "/write" or OperationNameValue endswith "/action"
| extend AfterCancellation = TimeGenerated > CancelRequested
| project TimeGenerated, AfterCancellation, ActivityStatusValue,
OperationNameValue, ResourceId, CorrelationId, Caller
| order by TimeGenerated asc A post-cancellation event is not automatically wrong. An operation accepted before the signal may complete later. It does mean recovery must wait for that operation to reach a terminal state.
Build a checkpoint matrix
Turn the combined timeline into a decision matrix. Every step needs target-side evidence, an idempotency property and a compensation. “Probably safe to replay” is not an operational property.
checkpoints:
- name: publish_image
target_evidence: registry_digest_exists
observed: true
idempotent_key: image_digest
compensation: none
- name: apply_configuration
target_evidence: app_settings_fingerprint
observed: true
idempotent_key: configuration_version
compensation: restore_previous_fingerprint
- name: route_traffic
target_evidence: active_revision_and_weight
observed: false
idempotent_key: revision_name
compensation: restore_previous_weights
- name: migrate_schema
target_evidence: migration_ledger
observed: unknown
idempotent_key: migration_id
compensation: forward_fix_or_tested_down_migration
terminal_rule: unknown_blocks_rerun An unknown checkpoint blocks a full rerun. Resolve it with a direct read, or treat it as potentially applied when stronger evidence is unavailable.
Choose resume, compensation or rollback
A resume is defensible when the artifact is identical, previous effects are proven and remaining steps are idempotent. Compensation fits an isolated partial write with a tested inverse. Rollback restores a known coherent state; it is not a blind execution of the previous pipeline.
Resume from a checkpoint
Same commit and digest
Previous states proven on the target
Remaining steps are bounded and idempotent
No concurrent writer
Compensate, then validate
Partial effect identified and isolated
Inverse operation tested
No irreversible migration or business effect
Validation runs before traffic resumes
Roll back
Active release is inconsistent or probes regress
Last healthy artifact and matching configuration are available
Schema compatibility is confirmed
The same validation matrix runs after restoration
Keep writes frozen
A critical checkpoint remains unknown
An Azure operation is still running
Multiple writers changed the target
Artifact, identity or rollback cannot be attributed If a data migration or external call might have succeeded, use an idempotency key, business ledger or purpose-built reconciliation. Pipeline status cannot prove the absence of a side effect.
Make the next cancellation observable
For future runs, use a deployment job bound to an environment, publish commit and digest before the first write, and record a checkpoint only after the target confirms state. An always() step can preserve evidence during the cancellation timeout, but it should not launch automatic rollback without reading effective state first.
Serialize writers, separate deployment from traffic movement, attach stable identifiers to remote operations, and make compensations independently callable. Test cancellation in preproduction at several points: before the first write, during an asynchronous operation, and after traffic changes.
Conclusion
A canceled Azure DevOps run proves neither that production was untouched nor that it was rolled back. It creates a break in the timeline that must be reconciled across the pipeline, agent, Azure control plane and application.
The final decision should name a proven state: resume the same artifact from an idempotent checkpoint, compensate one bounded effect, restore the last coherent release, or keep production frozen. Validation ends only when version, configuration, traffic, dependencies and business effects agree.