Infrastructure
Azure DevOps: contain a compromised self-hosted agent before reopening the pool
A production runbook for quarantining a suspicious Azure DevOps self-hosted agent, preserving evidence, bounding exposed identities and validating a clean replacement before the pool receives production jobs again.
A self-hosted Azure DevOps agent runs code pulled from repositories and can reach the private systems needed by a pipeline. That is precisely why an unexpected process, a modified task, an outbound connection to an unknown host or a secret appearing in a log must be treated as more than a broken runner.
The running case is a Linux agent in a production pool. An endpoint alert reports a shell spawned by a build task, the process contacts an address outside the expected egress path, and the same host recently ran deployments using a service connection and a private package feed. The agent is still online. The immediate goal is not to prove the entire attack chain. It is to stop new work, preserve useful evidence and bound which identities may have crossed the host before deciding whether the pool can reopen.
Freeze the execution path, not the evidence
Start with a short incident contract. A suspicious host, an agent identity and a pipeline identity are different scopes. Record them before an emergency rotation or deletion makes the timeline harder to reconstruct.
Incident window: 2026-10-02 13:10Z to present
Organization / project: naxaya-devops / platform
Pool / agent: prod-private-linux / ado-prod-07
Host: vm-ado-prod-07
Trigger: unexpected child process and outbound destination
Last known-good run: deploy-api #2418
Runs after signal: #2419, #2420
Immediate boundaries
Stop new jobs reaching this agent
Do not delete the agent record or wipe the workspace
Pause production pipelines that can select the pool
Preserve run IDs, commit SHAs and UTC timestamps
Decision to reach
False positive and controlled return
Rebuild and credential rotation
Wider pool containment and incident escalation Disable or otherwise remove the agent from scheduling, then pause the production entry points that can still target the pool. Do not rely on a demand that happens to exclude one machine: another YAML change may bypass it. If other agents share the same image, cache, bootstrap token or administrator path, treat the pool as a possible common scope until evidence narrows it.
Network isolation must preserve the response channel and follow the organisation’s forensic procedure. A hard power-off, VM deletion or broad subnet deny can destroy volatile evidence, strand the responder or interrupt healthy agents sharing the path.
Preserve the control-plane timeline first
Azure DevOps retains evidence that does not depend on the suspect host. Capture it before changing permissions:
- pipeline run IDs, stages, jobs, task versions and commit SHAs;
- queue time, assignment time and the agent name for every run in the window;
- changes to YAML, protected branches, variable groups, secure files and service connections;
- agent additions or deletions, pool permission changes and pipeline authorisations from the audit log;
- approvals, checks and environment deployment history.
For each affected run
run_id, pipeline_id, commit_sha
queued_at_utc, started_at_utc, finished_at_utc
pool_name, agent_name, job_name
repositories and task versions
service connections and protected resources used
artifacts produced or promoted
For each administrative change
event_id, actor, timestamp_utc
resource, previous scope, new scope
related approval or ticket Export or retain this material according to the incident process. Avoid pasting raw logs into an open ticket: masking is useful but is not proof that every token, command-line argument or structured secret was removed.
Collect host evidence without running the pipeline again
The host can answer useful questions without executing another job. Capture the agent diagnostics, service journal, running processes, recent network state and a bounded inventory of changed files. Hash evidence before moving it to approved storage.
INCIDENT_DIR="/var/tmp/ado-ir-20261002"
AGENT_HOME="/opt/azdo-agent"
sudo install -d -m 0700 "$INCIDENT_DIR"
sudo systemctl status 'vsts.agent.*' --no-pager > "$INCIDENT_DIR/systemd-status.txt"
sudo journalctl -u 'vsts.agent.*' --since '2026-10-02 13:00:00 UTC' --no-pager > "$INCIDENT_DIR/agent-journal.txt"
ps auxwwf > "$INCIDENT_DIR/processes.txt"
ss -tpna > "$INCIDENT_DIR/sockets.txt"
sudo find "$AGENT_HOME/_diag" -maxdepth 1 -type f -newermt '2026-10-02 13:00:00 UTC' -print > "$INCIDENT_DIR/diag-files.txt"
sudo find "$AGENT_HOME/_work" -xdev -type f -newermt '2026-10-02 13:00:00 UTC' -printf '%TY-%Tm-%TdT%TH:%TM:%TSZ %s %p
' > "$INCIDENT_DIR/work-files.txt"
sudo sha256sum "$INCIDENT_DIR"/* > "$INCIDENT_DIR/SHA256SUMS" Adapt paths and commands to the image. Do not archive the entire work directory by default: it may contain source code, artifacts and live credentials. Do not execute the suspected script to reproduce the alert. If memory capture, disk snapshots or endpoint tooling are required, hand them to the team authorised to collect them.
Build an exposure map before rotating anything
An agent registration credential lets the runner communicate with its pool. Jobs may receive completely different access: workload identity federation, service connections, variable-group secrets, secure files, registry tokens, SSH keys or a managed identity attached to the host. Rotation must follow observed exposure, not the longest list of credentials the team can remember.
Credential or trust Evidence of use Immediate control
Agent registration Agent session in window Keep agent quarantined
WIF service connection Token issued for run Disable or restrict connection
Managed identity Azure sign-in / resource Block host path; review role use
Variable-group secret Pipeline and task inputs Revoke value; rotate at source
Secure file Download event / task Remove authorisation; replace file
Registry or feed token Client and service logs Revoke token; inspect packages
SSH deployment key Target authentication log Remove key; review target activity
Record for every item
owner, scope, last use, expiry, replacement dependency
runs that received it and systems it could reach Start with credentials proven present in affected jobs, then expand to resources reachable through the same identity. For federated access, review token issuance, subject and audience as well as the service connection definition. For a managed identity, inspect Azure sign-in, activity and resource logs for the incident window; there may be no static secret to rotate, but the compromised host may still have requested tokens.
Avoid a blanket rotation with no dependency map. It can turn one security incident into several availability incidents while leaving the actual access path untouched.
Distinguish a malicious run from a contaminated runner
The same alert can come from different layers:
- an authorised build step invoked a scanner or package installer with an unexpected child process;
- a commit, template or third-party task changed and executed unwanted code;
- a dependency or artifact was replaced upstream;
- the agent image, bootstrap process or persistent workspace was modified;
- an operator used the host outside the pipeline.
Compare the suspicious run with a known-good run of the same pipeline. Diff the resolved YAML, commit, templates, task versions, downloaded artifacts, environment and egress destinations. Then compare the host with a clean agent built from the same image. A change that follows the commit points toward the workload; a change present before job checkout points toward the image or host. Neither observation alone clears the other layer.
Rebuild trust from a clean boundary
When arbitrary execution or credential access on the host is plausible, an in-place cleanup is not a trust reset. Build a replacement from the last approved image and bootstrap source. Give it a new agent registration, patch it, restore only declared tools and attach it to a restricted canary pool.
Do not copy _work, tool caches, SSH material or agent configuration from the quarantined machine. Pin bootstrap artifacts and task versions where the delivery model permits it. Run the service under a non-interactive, least-privileged identity distinct from the administrator that registers the agent. Restrict which pipelines may use the pool instead of reopening it to every project.
trigger: none
pool:
name: prod-private-linux-canary
steps:
- checkout: none
- bash: |
set -euo pipefail
id
uname -a
test -w "$(Agent.TempDirectory)"
getent hosts dev.azure.com
displayName: Validate runtime without protected resources The first canary must not load variable groups, secure files, service connections or production repositories. Validate the agent session, expected OS state, DNS, proxy and egress. Next, use a read-only repository and a disposable artifact. Add one protected resource at a time only after its replacement or continued trust has been approved.
Decide whether the pool can reopen
Reopening is a security decision with operational acceptance criteria:
REOPEN WITH CLEAN AGENTS
Incident window and affected runs are bounded
Suspect hosts remain quarantined
Exposed credentials are revoked or replaced
Clean image and bootstrap provenance are verified
Canary passes without unexpected egress
Pool and protected-resource permissions are restricted
Detection and audit coverage are active
KEEP THE POOL CLOSED
Unknown jobs or identities remain in the window
The same indicator appears on another agent
A produced artifact cannot be trusted
Rotation is incomplete or target activity is unexplained
ROLL BACK THE RETURN TO SERVICE
Pause pipelines and remove pool authorisation
Disable newly introduced credentials
Quarantine the replacement image cohort
Return deployments to the validated standby path Keep the original host outside the pool until the evidence retention decision is complete. Closing the alert is not the same as authorising the execution boundary again.
Conclusion
A suspicious self-hosted agent sits at the intersection of source code, deployment identities, private networking and production targets. Containment therefore starts by stopping scheduling without erasing the host, then reconstructing the control-plane timeline and mapping only the access that could have crossed the runner.
The final decision is deliberately narrow: reopen the pool with clean agents and rotated access only when the incident window, artifacts and downstream actions are explained; otherwise keep it closed and use the validated standby path. The rollback is equally explicit: withdraw pool authorisation before another production job can turn uncertainty into a second incident.