Automation

Azure DevOps: diagnose a canceled deployment before rerunning production

A production runbook for reconstructing the effects of a canceled Azure DevOps deployment, proving target state, and choosing an idempotent resume, compensation or rollback.

05 Oct 2026 azureazure-devopspipelinesdeploymentautomationdevopsenvironmentsactivity-logobservabilityidempotencyrunbookrollbackproduction

An Azure DevOps pipeline is canceled during a production deployment. The run says Canceled, but the new image may already be referenced, an application setting has changed, and traffic may not have moved yet. Rerunning feels like the quickest recovery. It can also layer a second release over a target state nobody has established.

The working scenario is a deployment job that releases an Azure API, applies configuration, and runs a smoke test. Cancellation lands between those steps. This runbook drives one bounded decision: resume from a proven checkpoint, compensate a partial effect, restore the last known-good release, or keep writes frozen while the target remains ambiguous.

Freeze the run before another write

Record the organization, project, pipeline, run ID, stage, environment, commit and artifact. Capture the UTC time when cancellation was requested and the timestamp of the last evidence produced by the agent. Do not replace these identifiers with “latest run” or “current release.”

yaml canceled-deployment-incident.yml
incident: INC-DEPLOY-20261005-01
pipeline: deploy-orders-api
run_id: 18427
stage: production
environment: prod-orders
source_commit: 8d2c1f4
artifact:
name: orders-api
digest: sha256:<digest>
cancellation_requested_utc: 2026-10-05T15:42:18Z
last_pipeline_evidence_utc: 2026-10-05T15:42:31Z
change_window_end_utc: 2026-10-05T16:30:00Z
writers_frozen:
- scheduled_release
- manual_hotfix
- configuration_automation
decision_owner: platform-on-call

Freeze the other writers that target the environment: concurrent pipelines, scheduled jobs and manual changes. The freeze protects the investigation. It does not assume the process already in flight has stopped.

Separate run status from target state

Canceled is the Azure DevOps run outcome, not a distributed transaction rollback. A remote command can succeed before the agent receives the signal. A step guarded by always() can still run during the cancellation grace period. Conversely, a disconnected agent may never upload its final result.

text deployment-state-layers.txt
Azure DevOps
Run, stage, job and task: queued | inProgress | completed
Final result: canceled

Agent
Cancellation signal received or not
Child process stopped or still active
Last known log and heartbeat

Azure target
ARM operation accepted, completed or running
Configuration actually written
Version or revision receiving traffic
Migrations or messages already produced

Application
Version observed by probes
Health, errors and business side effects
Compatibility with schema and dependencies

The run provides a command timeline. The target provides operational truth. Reconcile both before choosing a recovery action.

Capture the commit, artifact and timeline

Environment deployment history helps identify the pipelines and commits that targeted production. Combine it with run details and published artifacts. A branch-based rerun can rebuild different bytes, so the deployed digest remains the stronger reference.

bash 01-freeze-run-and-artifact.sh
ORG="https://dev.azure.com/example"
PROJECT="platform-prod"
RUN_ID="18427"

az pipelines runs show --org "$ORG" --project "$PROJECT" --id "$RUN_ID" --query '{id:id,name:name,state:state,result:result,branch:sourceBranch,commit:sourceVersion,created:createdDate,finished:finishedDate}' --output json

az pipelines runs artifact list --org "$ORG" --project "$PROJECT" --run-id "$RUN_ID" --output table

Classify every task as not started, started without completion evidence, completed with evidence, or completed but not validated. A script success line is insufficient for an asynchronous operation. Keep its operation ID and read the terminal state from Azure.

Reconstruct effects from the target

List every possible pipeline write: ARM deployment, image reference, settings, secret references, traffic weights, migrations, messages or cache invalidation. Read the effective target state and timestamp for each write. Do not issue corrective commands during this pass.

bash 02-read-azure-changes.sh
SUBSCRIPTION="<subscription-id>"
RESOURCE_GROUP="rg-orders-prod"
START="2026-10-05T15:30:00Z"
END="2026-10-05T16:00:00Z"

az monitor activity-log list --subscription "$SUBSCRIPTION" --resource-group "$RESOURCE_GROUP" --start-time "$START" --end-time "$END" --query '[].{time:eventTimestamp,status:status.localizedValue,operation:operationName.localizedValue,caller:caller,correlationId:correlationId,resourceId:resourceId}' --output table

az deployment group list --resource-group "$RESOURCE_GROUP" --query '[].{name:name,state:properties.provisioningState,time:properties.timestamp,correlationId:properties.correlationId}' --output table

The Activity Log proves control-plane operations. It does not, by itself, prove the served version, a data-plane write or a successful business outcome. Add service state, service logs and a read-only probe.

Correlate identity and timing in KQL

When the Activity Log is exported to Log Analytics, filter by the incident window, resource group, service connection principal and correlation IDs. Look for writes after the cancellation request as well: they indicate either an operation that continued or another change source.

kusto 03-correlate-canceled-deployment.kql
let Start = datetime(2026-10-05T15:30:00Z);
let CancelRequested = datetime(2026-10-05T15:42:18Z);
let End = datetime(2026-10-05T16:00:00Z);
let DeploymentCaller = "<service-connection-principal>";
AzureActivity
| where TimeGenerated between (Start .. End)
| where ResourceGroup =~ "rg-orders-prod"
| where Caller =~ DeploymentCaller
| where OperationNameValue endswith "/write" or OperationNameValue endswith "/action"
| extend AfterCancellation = TimeGenerated > CancelRequested
| project TimeGenerated, AfterCancellation, ActivityStatusValue,
  OperationNameValue, ResourceId, CorrelationId, Caller
| order by TimeGenerated asc

A post-cancellation event is not automatically wrong. An operation accepted before the signal may complete later. It does mean recovery must wait for that operation to reach a terminal state.

Build a checkpoint matrix

Turn the combined timeline into a decision matrix. Every step needs target-side evidence, an idempotency property and a compensation. “Probably safe to replay” is not an operational property.

yaml deployment-checkpoints.yml
checkpoints:
- name: publish_image
  target_evidence: registry_digest_exists
  observed: true
  idempotent_key: image_digest
  compensation: none

- name: apply_configuration
  target_evidence: app_settings_fingerprint
  observed: true
  idempotent_key: configuration_version
  compensation: restore_previous_fingerprint

- name: route_traffic
  target_evidence: active_revision_and_weight
  observed: false
  idempotent_key: revision_name
  compensation: restore_previous_weights

- name: migrate_schema
  target_evidence: migration_ledger
  observed: unknown
  idempotent_key: migration_id
  compensation: forward_fix_or_tested_down_migration

terminal_rule: unknown_blocks_rerun

An unknown checkpoint blocks a full rerun. Resolve it with a direct read, or treat it as potentially applied when stronger evidence is unavailable.

Choose resume, compensation or rollback

A resume is defensible when the artifact is identical, previous effects are proven and remaining steps are idempotent. Compensation fits an isolated partial write with a tested inverse. Rollback restores a known coherent state; it is not a blind execution of the previous pipeline.

text canceled-deployment-decision.txt
Resume from a checkpoint
Same commit and digest
Previous states proven on the target
Remaining steps are bounded and idempotent
No concurrent writer

Compensate, then validate
Partial effect identified and isolated
Inverse operation tested
No irreversible migration or business effect
Validation runs before traffic resumes

Roll back
Active release is inconsistent or probes regress
Last healthy artifact and matching configuration are available
Schema compatibility is confirmed
The same validation matrix runs after restoration

Keep writes frozen
A critical checkpoint remains unknown
An Azure operation is still running
Multiple writers changed the target
Artifact, identity or rollback cannot be attributed

If a data migration or external call might have succeeded, use an idempotency key, business ledger or purpose-built reconciliation. Pipeline status cannot prove the absence of a side effect.

Make the next cancellation observable

For future runs, use a deployment job bound to an environment, publish commit and digest before the first write, and record a checkpoint only after the target confirms state. An always() step can preserve evidence during the cancellation timeout, but it should not launch automatic rollback without reading effective state first.

Serialize writers, separate deployment from traffic movement, attach stable identifiers to remote operations, and make compensations independently callable. Test cancellation in preproduction at several points: before the first write, during an asynchronous operation, and after traffic changes.

Conclusion

A canceled Azure DevOps run proves neither that production was untouched nor that it was rolled back. It creates a break in the timeline that must be reconciled across the pipeline, agent, Azure control plane and application.

The final decision should name a proven state: resume the same artifact from an idempotent checkpoint, compensate one bounded effect, restore the last coherent release, or keep production frozen. Validation ends only when version, configuration, traffic, dependencies and business effects agree.