Infrastructure

Azure DevOps: diagnose a queued job before adding self-hosted agents

A production runbook for separating parallel-job capacity, pool authorization, demands, capabilities and agent eligibility before scaling an Azure DevOps self-hosted pool.

22 Sept 2026 azure-devopsself-hosted-agentpipelinesautomationdevopscapabilitiesobservabilityrunbookrollbackproduction

A production pipeline has been waiting for an agent for twenty minutes. The pool shows several self-hosted agents, so adding another VM appears to be the fastest fix. That response can increase cost without moving the job: the organization may have exhausted its parallel-job capacity, the pipeline may not be authorized to use the pool, or no online agent may satisfy the job’s demands.

The running case is a deployment pipeline targeting a private Azure environment from the prod-linux pool. Other jobs still run, but one stage remains queued after a pipeline change. The goal is to identify the scheduling constraint, restore one bounded execution and decide whether the durable action is capacity, authorization, capability repair, agent upgrade or pipeline rollback.

Freeze the scheduling contract

Start with the queued run, not the pool dashboard. Record the run ID, job name, queue time, requested pool, branch and pipeline revision. Preserve the exact message shown on the pending job; “waiting for an agent”, “waiting for a parallel job” and “resource authorization required” point to different control planes.

yaml queued-job-contract.yml
organization: https://dev.azure.com/example
project: platform
pipeline: deploy-orders
run_id: 18427
job: deploy_prod
queued_at_utc: 2026-09-22T05:42:18Z
pool: prod-linux
pipeline_revision: 4d83c7a
expected_demands:
- Agent.OS -equals Linux
- terraform
- prod-network
change_window:
start_utc: 2026-09-22T05:30:00Z
end_utc: 2026-09-22T07:00:00Z
success: one canary and one production job are assigned to eligible agents
rollback: restore the previous pipeline revision and capability contract

Do not restart agents or edit capabilities yet. Those actions change the evidence the scheduler used. Compare the queued run with one recent successful run from the same pipeline and pool: job YAML, demands, tasks, agent name, agent version and queue duration.

Separate authorization, concurrency and eligibility

A self-hosted pool containing idle machines does not prove that the job can run. Azure DevOps must pass three gates in order:

  1. The YAML pipeline is authorized to use the pool.
  2. The organization has an available parallel job.
  3. At least one enabled, online agent satisfies every explicit and task-generated demand.

Check the pending job details and pool security first. A resource authorization prompt is not a capacity issue. Authorize the individual pipeline when that access is intended; opening the pool to every pipeline widens the execution boundary and is not an incident fix.

Then inspect Organization settings > Pipelines > Parallel jobs and the pool consumption view. If all concurrency slots are occupied, a fourth registered agent cannot receive a job. Identify the running jobs, their owners and expected completion times. Cancel a job only when it is proven abandoned or safely restartable.

Finally, treat agent eligibility separately. One matching agent may be busy while five non-matching agents remain idle. That is a scheduling bottleneck, not necessarily a fleet shortage.

Inventory agents with their real capabilities

Use the Azure DevOps CLI to read the pool and include capabilities and assigned requests. Keep this step read-only.

bash 01-inventory-agent-pool.sh
ORG="https://dev.azure.com/example"
POOL="prod-linux"

az extension add --name azure-devops --only-show-errors
az devops configure --defaults organization="$ORG"

POOL_ID=$(az pipelines pool list --pool-name "$POOL" --query "[0].id" --output tsv)

az pipelines agent list --pool-id "$POOL_ID" --include-capabilities true --include-assigned-request true --include-last-completed-request true --output json > agent-pool-snapshot.json

jq '.[] | {
id,
name,
enabled,
status,
version,
assignedRequest,
lastCompletedRequest,
systemCapabilities,
userCapabilities
}' agent-pool-snapshot.json

Interpret the result as a scheduler would. An agent must be enabled, online, unassigned and compatible with the full demand set. The Agent.Version, Agent.OS and Agent.OSArchitecture values are capabilities too. User capabilities such as terraform or prod-network are labels asserted by operators; they do not prove that the binary, route or permission still works.

Do not publish secrets as user capabilities. System capabilities can be derived from environment variables when the agent starts. Use VSO_AGENT_IGNORE for sensitive or volatile variables that must not be stored or reused as capabilities.

Diff demands against capabilities

Demands may be written in the pool block or introduced automatically by a task. Build an explicit matrix from the queued job and every candidate agent.

text demand-capability-matrix.txt
Demand                         agent-01       agent-02       agent-03
Agent.OS = Linux               match          match          match
terraform exists               match          missing        match
prod-network exists            match          match          missing
minimum task agent version     outdated       match          match
enabled and online             yes            yes            no
currently assigned             yes            no             no

Eligible now
none

Scheduling constraint
agent-01 matches but is busy
agent-02 lacks terraform
agent-03 is offline and lacks prod-network

Review the pipeline diff that preceded the queue. A renamed demand, a value change, a new task or an exact Agent.Version comparison can eliminate the entire pool. Remember that exists and equals are the supported demand operations; string comparisons that encode ranges or shell logic do not create a meaningful scheduler rule.

If software was installed after an agent started, restart the agent only after preserving its diagnostics: capability discovery happens at startup. Do not add a user capability merely to make the scheduler green when the underlying executable is absent.

Check version and task compatibility

A task can require a newer agent even when the YAML has no explicit version demand. Compare the queued job’s task changes with the Agent.Version capability of eligible agents. Also check whether automatic minor updates are enabled and whether the host operating system supports the target agent generation.

Upgrade one non-critical agent first. Bring it online, confirm its capabilities, run the canary, then roll through the pool. Updating every agent at once removes the known-good scheduling path. If the new task forced the requirement and the maintenance window cannot absorb an agent upgrade, rollback the task or pipeline revision instead of falsifying the version capability.

Prove scheduling with a side-effect-free canary

A canary should use the same pool and intentional demands but perform no deployment. It proves authorization and assignment without mixing scheduler recovery with production access.

yaml agent-scheduling-canary.yml
trigger: none

pool:
name: prod-linux
demands:
- Agent.OS -equals Linux
- terraform
- prod-network

steps:
- checkout: none
- bash: |
  set -euo pipefail
  echo "agent=$AGENT_NAME"
  echo "version=$AGENT_VERSION"
  command -v terraform
  terraform version
displayName: Validate agent scheduling contract

The canary must reach Running, identify the expected agent and validate the capability behind each important label. A successful canary does not authorize the deployment automatically. Requeue one bounded production job and observe assignment time, agent identity and final cleanup before releasing the backlog.

Choose the smallest corrective action

Use the evidence to select one change:

text queued-job-decision.txt
Authorize the pipeline
The pending job explicitly requests pool authorization
The pipeline is approved for this execution boundary
Access is granted to this pipeline, not opened globally

Free or buy parallel capacity
Eligible idle agents exist
Every concurrency slot is consumed
Queue history proves sustained capacity pressure, not one stuck run

Repair capability or upgrade an agent
The required capability is legitimate
The executable or compatible agent version is missing
One-agent canary and rollback are available

Rollback the pipeline
A recent YAML or task change introduced an accidental demand
The previous revision ran on the current pool
Restoring it is lower risk than changing every agent

Add or autoscale agents
Parallel capacity is available
Matching agents are saturated over a representative window
Image, network, identity and cleanup behavior are already validated

Scaling is the final branch, not the first. New agents need the same network routes, toolchain, service identity boundaries, update policy and workspace cleanup as the existing fleet. Otherwise the pool grows while the scheduling contract diverges.

Validate recovery and retain rollback

Recovery is complete when the original job is assigned for the understood reason, not when an operator presses Run pipeline repeatedly. Record the matched demands, selected agent, queue duration, parallel-slot state and pipeline authorization decision.

Rollback the smallest changed layer. Remove an accidental demand by restoring the prior YAML revision. Revert a task upgrade if agent compatibility cannot be completed. Disable a newly introduced agent image if its canary fails. Revoke a pipeline authorization only if it was granted to the wrong execution boundary. Do not delete agents or open the pool globally to make the queue disappear.

Conclusion

An Azure DevOps job can remain queued while healthy agents appear idle because scheduling depends on authorization, parallel capacity and exact agent eligibility. Adding machines helps only when matching agents are genuinely saturated and concurrency is available.

The production decision should follow the constraint: authorize narrowly, release proven capacity, restore a real capability, upgrade through a canary, rollback the pipeline revision or scale a validated pool. Once the cause and return path are explicit, queue recovery becomes an operable change instead of a blind increase in runners.