Automation
Azure DevOps: contain a secret exposed in logs before rerunning the pipeline
A production runbook for bounding a secret leak in Azure DevOps logs, preserving evidence, revoking access, validating redaction and deciding recovery or rollback.
An Azure DevOps pipeline fails and a sensitive value appears in clear text in a task log. The instinctive response is to hide the line, delete the run, or rotate every secret in the project. Those actions may reduce visible exposure, but they do not establish which credential leaked, where it was copied, or whether immediate revocation will interrupt a production workload.
The running case is a deployment pipeline that reads a value from Azure Key Vault, passes it to a script, and accidentally prints an expanded command. The same value may also have reached a diagnostic artifact, a self-hosted agent workspace, or an external log platform. This runbook contains the incident, proves its scope, replaces the compromised credential, and permits recovery only after a clean canary run.
Freeze writes without destroying evidence
Stop creating new exposure first. Pause the affected automatic trigger, suspend stages that can write to production, and block unapproved manual reruns. Do not immediately delete the run: its timeline, execution identity, and artifacts are needed to bound the incident.
Capture an incident record that never contains the secret value:
incident_id: sec-2026-09-20-01
detected_utc: <timestamp>
pipeline:
project: <project>
definition_id: <id>
run_id: <id>
commit: <sha>
stage: <stage>
job: <job>
task: <task>
suspected_secret:
source: azure-key-vault
logical_name: <name-without-value>
version_or_revision: <version>
fingerprint: <keyed-fingerprint>
containment:
automatic_triggers_paused: true
production_writes_blocked: true
evidence_access_restricted: true
decision_owner: <role>
rollback_owner: <role> Keep evidence access within a restricted group. A compromised log must not be pasted into a general ticket, chat channel, or broadly readable report. Reference the run and timestamps instead of copying the sensitive line.
Identify the secret without publishing it again
The variable name printed by a task is not sufficient evidence. Its value may originate in a Variable Group, Key Vault reference, task output, or local transformation. Reconstruct the value path from pipeline definition to the process that wrote it.
To compare the same value across authorized locations without recording it again, compute an HMAC with an ephemeral incident key. A plain unsalted hash remains vulnerable when a credential has low entropy.
set -euo pipefail
read -rsp "Secret value: " SECRET_VALUE
printf '
'
read -rsp "Incident HMAC key: " INCIDENT_KEY
printf '
'
FINGERPRINT=$(printf '%s' "${SECRET_VALUE}" | openssl dgst -sha256 -hmac "${INCIDENT_KEY}" -binary | base64)
printf 'fingerprint=%s
' "${FINGERPRINT}"
unset SECRET_VALUE INCIDENT_KEY FINGERPRINT Run this only on a controlled investigation workstation with shell history and terminal collection disabled. The fingerprint supports comparisons within already authorized evidence; it does not justify bulk downloads of logs or artifacts.
Bound the time window and every copy
Start with the first run that contains the value, then find the last known-clean run with the same commit, definition, templates, and secret source. Include retries, parallel jobs, and other pipelines that consume the same Variable Group or Key Vault secret.
Classify locations instead of treating the task log as the only copy:
Location Check
Job log Line, timestamp, task and access boundary
Diagnostic artifact Content, retention and authorized downloads
Agent workspace Temporary files, cache and local permissions
Task output Variables propagated to later jobs or stages
External log platform Ingestion, index, export and retention policy
Notification or webhook Sent payload and recipients
Manual fork, copy or export Owner and approved location
For every copy
Keep the reference and fingerprint, never the value in the ticket
Restrict access before cleanup
Record deletion, expiry or inability to retract
Treat any downloaded copy as outside pipeline control Use available audit events to identify views, downloads, or changes during the exposure window. Missing audit evidence does not prove that nobody read the value; it only defines the observations you actually have.
Qualify impact before rotating
A visible value is not always an active credential, but it must be treated as compromised until use is ruled out. Establish its type, issuer, expiration, permissions, targets, and active consumers.
Separate four questions:
- can the value still authenticate;
- which verbs and resources can it reach;
- which workloads still consume this version;
- which evidence would reveal use after exposure.
For an application credential, correlate sign-in and target-resource logs with the incident window, expected identity, known network sources, and normal operations. For an API key or service-specific token, use that provider’s audit trail. Do not infer compromise, or the absence of compromise, from the pipeline result alone.
Fix the leak path before minting a new value
Rotating while the pipeline still prints expanded arguments creates a second leak. Correct the exposure path on a controlled branch before issuing the replacement.
Common causes are explicit: set -x, Write-Host or echo around an expanded command, an environment dump, serialized configuration, an error message containing a full URL, or an overbroad debug artifact. Azure DevOps masking is a final safety layer. It should not be expected to recognize every substring, encoding, split, or transformation of a secret.
Replace command-line arguments with a channel supported by the program: standard input, a short-lived file with strict permissions, an environment variable consumed without printing, or federated/managed identity when the backend supports it. Disable tracing around secret retrieval and restore it only after local values are destroyed.
Add a negative test using a harmless, unique canary value. The test must fail when the raw value, its expected encodings, or a URL containing it appears in generated logs or artifacts.
Revoke in an order the service can survive
Prepare a new version, move one canary consumer, and validate access through the real runtime identity and network path. Once proven, migrate consumers in bounded cohorts and watch for denials caused by the old value.
If the credential permits a sensitive write or unauthorized use is visible, revoke it immediately even when that creates a controlled outage. If no suspicious use is observed and critical consumers still depend on the old value, the incident owner may approve a short overlap. That exception needs a deadline, telemetry, and an explicit stop condition.
Do not disable an entire identity when revoking one credential or version is sufficient. Conversely, do not preserve an old credential merely because an unknown consumer might exist: identify its owner, migrate it, or explicitly accept that it will stop.
Validate a clean canary run
Recovery starts with a bounded manual run against a target that cannot create an irreversible business effect. Use the replacement value, the logging fix, and the same agent class as production.
The canary is acceptable when
The canary value appears in no log or artifact
The new value fingerprint is absent from controlled locations
Authentication uses the expected version and identity
Denials from the old credential are understood
No critical consumer still depends on the old version
Access to compromised evidence remains restricted
The negative disclosure test blocks regression
Decision
RESUME: criteria pass and the old credential is revoked
HOLD: scope or consumers remain unknown
ROLL BACK CODE: the fix breaks the pipeline without reversing revocation Code rollback and credential rollback are different decisions. The team may restore the last working script only if it cannot recreate the disclosure. A compromised credential must not be re-enabled to make a pipeline pass.
Conclusion
A secret printed by Azure DevOps is an execution-chain incident, not a cosmetic logging defect. The response must stop new runs, preserve evidence, identify the value without copying it, bound its replicas and privileges, repair the logging path, and revoke according to actual impact.
The recovery decision is concrete: a canary succeeds with the replacement credential, no sensitive representation reaches outputs, the old value is revoked, and critical consumers are accounted for. While any of those conditions remains ambiguous, the pipeline stays blocked. Lack of evidence is not permission to deploy again.