Automation

Azure DevOps: diagnose workspace contamination before recreating a self-hosted agent

A production runbook for separating persistent sources, outputs, caches and machine state on a self-hosted Azure DevOps agent before cleanup, quarantine or recreation.

04 Sept 2026 azure-devopspipelinesself-hosted-agentworkspacecacheci-cdautomationdevopsobservabilitysecurityrunbookrollbackproduction

An Azure DevOps pipeline succeeds on one agent, fails on another, then turns green after a retry without any commit change. The visible symptom may be a stale binary, a generated file left behind, an incomplete dependency restore, an inconsistent cache or configuration inherited from the previous job. Recreating the agent feels fast, but it also destroys the state that could explain the incident.

This runbook covers a self-hosted pool used for private builds and deployments. Unlike a Microsoft-hosted agent that is replaced for each job, its working directory and machine state can persist across runs. The operational decision is whether to clean sources, outputs or the full workspace, repair the pipeline, quarantine one agent, or rebuild it from a known image.

Freeze a comparable pair of runs

Start with one healthy run and one failed run that should have produced the same result. A pool name is not enough. Record the exact agent, commit, parameters, input artifacts and tool versions.

text workspace-incident-scope.txt
Incident: inc-20260904-002
Pipeline: build-orders-api
Pool: private-linux-prod
Healthy run: 20260904.17 / agent ci-linux-02
Failed run: 20260904.18 / agent ci-linux-04
Expected commit: 7d0d0d0
Mode: build and publish one immutable artifact

Compare
Agent.Name, Agent.Version and Agent.OS
Pipeline.Workspace and Build.SourcesDirectory
actual checked-out commit and submodules
runtime, SDK, package manager and tool versions
fingerprints of inputs and produced artifact
files present before restore and before build
non-secret variables controlling the build

Stop during evidence collection
repeated retries on the same agent
manual deletion of the whole work directory
opportunistic dependency upgrades
reuse of an artifact with unknown provenance

Two runs are comparable only when their inputs are comparable. If the pipeline rebuilds from a moving branch, downloads latest, or resolves dependencies without a lockfile, stabilize those inputs first.

Locate the persistence boundary

A self-hosted workspace contains areas with different lifecycles. Sources usually live under s, binaries under b, staged artifacts under a, and test results under TestResults. The last two are cleaned between runs by default, but that does not mean sources and outputs are clean.

bash 01-workspace-inventory.sh
set -eu

printf 'agent=%s version=%s os=%s
' "${AGENT_NAME:-unknown}" "${AGENT_VERSION:-unknown}" "${AGENT_OS:-unknown}"

printf 'workspace=%s
sources=%s
binaries=%s
' "${PIPELINE_WORKSPACE:-unknown}" "${BUILD_SOURCESDIRECTORY:-unknown}" "${BUILD_BINARIESDIRECTORY:-unknown}"

git -C "$BUILD_SOURCESDIRECTORY" rev-parse HEAD
git -C "$BUILD_SOURCESDIRECTORY" status --short --untracked-files=all

find "$PIPELINE_WORKSPACE" -maxdepth 2 -type f -printf '%TY-%Tm-%TdT%TH:%TM:%TSZ %s %p
' | sort | tail -n 200

Keep this inventory free of secrets. Do not publish file contents, variables or persisted credentials. Paths, timestamps, sizes and hashes are usually enough to prove that residual state exists.

Separate four contamination classes

A Git cleanup does not remove every persistent state. Classify the residue before selecting a cleanup boundary.

text workspace-contamination-classes.txt
Git sources
ignored or untracked file still present
submodule on the wrong revision
incomplete Git LFS content
modified local Git configuration

Build outputs
binary compiled from an older commit
dist, bin or obj reused without evidence
manifest generated before dependency restore

Explicit caches
cache key too broad or missing the lockfile
cache shared across branches, architectures or runtimes
restored content never validated

State outside the workspace
globally installed package
credential or configuration in the agent account home
container, volume, process or service left running
multiple agents sharing the same work directory

The last group is the deceptive one. workspace.clean: all does not clean the user home, a daemon, a Docker volume or a machine-wide tool. If the clean canary still fails, expand the diagnosis to the image and service account instead of adding more deletion commands.

Prove the hypothesis with a clean canary

Create a diagnostic run on an isolated agent or one temporarily removed from normal traffic. Keep the same inputs as the failed run and change only the cleanup policy.

yaml azure-pipelines-clean-canary.yml
jobs:
- job: clean_workspace_canary
displayName: Clean workspace canary
pool:
  name: private-linux-prod
  demands:
  - Agent.Name -equals ci-linux-canary
workspace:
  clean: all
steps:
- checkout: self
  clean: true
  fetchDepth: 0

- bash: |
    set -eu
    test -z "$(git status --porcelain --untracked-files=all)"
    git rev-parse HEAD
    sha256sum package-lock.json
  displayName: Prove clean inputs

- script: npm ci
  displayName: Restore from lockfile

- script: npm run build
  displayName: Build candidate

- bash: |
    set -eu
    find dist -type f -print0 | sort -z | xargs -0 sha256sum > artifact.sha256
  displayName: Fingerprint outputs

checkout.clean: true resets the Git working tree before fetching. workspace.clean: all deletes the pipeline workspace before the job. Together, they test the local-residue hypothesis, but they should not become the permanent fix until the source and performance cost of the contamination are understood.

Give every state an owner

The pipeline must not rely on landing on the same agent again. Any state needed by a later job should be published as an artifact, restored through an explicitly versioned cache, or rebuilt from locked inputs.

yaml workspace-policy.yml
workspace_policy:
sources:
  owner: checkout
  validation: exact commit and clean status
build_outputs:
  owner: current run
  reuse: forbidden unless artifact digest is verified
dependency_cache:
  key_parts:
    - operating_system
    - architecture
    - runtime_version
    - lockfile_hash
  fallback: clean restore
machine_state:
  owner: agent image
  drift_detection:
    - agent version
    - tool manifest
    - running services
    - container and volume inventory
secrets:
  persistence: forbidden
  validation: no credential file after job

Then select the narrowest cleanup. resources targets sources, outputs targets binaries, and all targets the whole workspace. When a toolchain writes into the user home or launches services, the correction belongs in the script, job container or agent image, not in checkout.

Validate without hiding recurrence

Run the clean canary first, then two consecutive runs on the same agent, followed by a run on another pool member. This sequence tests reproducibility and the absence of an implicit dependency on one machine.

text workspace-validation-gates.txt
Promote the correction
same commit and inputs across all three runs
final artifact identical or differences explained
no canary file from the previous run
no persisted credential
build time and cache volume remain acceptable
diagnostics remain available if the issue returns

Quarantine the agent
clean run fails only on this agent
toolchain, service or local storage has drifted
state cannot be inspected without risk

Rebuild the agent
evidence collected and cause tied to the image
known image, tool manifest and enrollment procedure available
negative and positive canaries planned before pool return

Rollback
remove the new policy if it deletes a required cache
restore the previous YAML from version control
keep the suspect agent out of the pool
return to the last validated agent image

Rollback does not mean reintroducing a contaminated workspace. It restores the previous pipeline policy while keeping the suspect agent away from production until state ownership is fixed.

Conclusion

A self-hosted Azure DevOps agent provides private network access, reusable caches and a controlled toolchain, but that persistence must stay explainable. When identical runs diverge, freeze the inputs, locate the residue, run a clean canary and compare agents before rebuilding the machine.

The final decision becomes precise: clean one defined area, repair a cache or build, quarantine one agent, rebuild a known image, or restore the previous policy. The pipeline regains an essential property: its result depends on declared inputs, not on the worker’s invisible history.