Automation

Azure Automation: diagnose a Hybrid Runbook Worker that stops picking up jobs

A production runbook for separating queueing, heartbeat, extension, network, capacity, identity and runtime failures before retrying a Hybrid Worker job.

14 Aug 2026 azureazure-automationhybrid-runbook-workerrunbookautomationmanaged-identitynetworkingobservabilityincidentrollbackproduction

An Azure Automation job sits in the queue, fails because no worker is available, or moves to Running without producing output. The runbook worked the previous day. The obvious responses are to retry the job, restart the VM, reinstall the extension, or add another worker. Any of them may restore service, but they also blur the critical boundary between a dispatch failure, a stalled runtime, and a script that has already started changing production.

This runbook covers an extension-based Hybrid Runbook Worker group hosted on Azure VMs or Azure Arc-enabled servers. The group runs tasks that need private network reachability: local rotation, application maintenance, collection, bounded restart, or infrastructure change. The operational goal is to decide whether the job is safe to resume, the worker needs repair, capacity needs adjustment, the outbound path is broken, or the job must remain blocked until its state is known.

Freeze the incident before retrying

Preserve the job ID, published runbook version, parameters, target group, and UTC timestamps. Queued, Running, Failed, and Suspended describe different failure stages. Record potential side effects as well. A job can call an API, write a file, or restart a service before emitting its first output line.

text hybrid-worker-incident-scope.txt
Incident: inc-20260814-017
Automation Account: aa-platform-prod
Runbook: Invoke-PrivateMaintenance
Job ID: 00000000-0000-0000-0000-000000000000
Worker group: hrw-prod-weu
Requested at: 2026-08-14T15:10:00Z
Last known output: none
Expected target: srv-app-07 / service api-orders

Questions before action
Was the job dispatched to a worker?
Did a runtime process start?
Could a local or remote write already have completed?
Can the same group still run a canary job?
What recently changed on the VM, Arc agent, extension, proxy or firewall?

Temporary holds
Do not retry the business job
Do not reinstall the extension
Do not widen RBAC or firewall rules
Do not delete local logs

Collect Output, Error, Warning, and Verbose streams even when they are empty. No output is still evidence: the worker may never have received the job, the runtime may have failed to start, or the script may perform its first action before logging.

Separate dispatch from execution

Make the first branch explicit. When no worker picked up the job, investigate group membership, heartbeat, extension health, and connectivity to Azure Automation. When a worker did pick it up, investigate the runtime, modules, execution account, local resources, and script dependencies.

text hybrid-worker-failure-matrix.txt
Queued, then worker unavailable
First checks: heartbeat, hwd/HybridWorkerService, extension, outbound TCP 443

Running with no application checkpoint
First checks: runtime launch, CPU quota, local account, module or binary

Running after an application checkpoint
First checks: target dependency, lock, timeout, idempotency before stop

Failed or Suspended with output
First checks: runbook contract, identity, permissions, exit code

Several jobs fail on one group
First checks: worker fleet, capacity or shared network path

One runbook fails across several workers
First checks: script, module, input or runbook-specific dependency

This prevents every symptom from becoming a VM problem. A healthy worker can dispatch a script whose module is missing. Conversely, publishing another runbook version will not restore a stopped local service or a proxy that blocks the Automation endpoint.

Check the group, worker and extension

An extension-based Hybrid Runbook Worker depends on machine health, the machine system-assigned managed identity, the extension, and group registration. Compare all workers in the group rather than focusing on one host. Capture extension version and provisioning state, last ping, OS, resource pressure, and recent changes.

bash 01-hybrid-worker-extension-inventory.sh
RG="rg-automation-prod"
VM="vm-hrw-prod-01"

az vm get-instance-view --resource-group "$RG" --name "$VM" --query '{power:instanceView.statuses[].displayStatus,extensions:instanceView.extensions[].{name:name,status:statuses[0].displayStatus,message:statuses[0].message}}' --output json

az vm identity show --resource-group "$RG" --name "$VM" --query '{type:type,principalId:principalId,tenantId:tenantId}' --output json

For an Arc-enabled server, inspect the equivalent Connected Machine resource and its extensions. Keep two identities separate: the machine identity supports extension operation, while the runbook must authenticate its own actions deliberately. Before assigning another role, prove which identity the script actually receives and the narrow scope of the missing permission.

Use the HybridWorkerPing metric to identify heartbeat gaps. Align the gap with VM health, extension events, and change windows. A worker that was powered off, removed from the group, or stopped pinging is not repaired by replaying the jobs that targeted it.

Prove outbound access to Azure Automation

The worker initiates outbound HTTPS connections. A VM that accepts RDP or SSH can still be unable to receive jobs. Read AutomationHybridServiceUrl from the Automation account properties or extension settings, then test that exact hostname from the worker. A successful test from an administrator workstation says nothing about the worker path.

powershell 02-test-automation-egress.ps1
$AutomationHost = "replace-with-account-endpoint.azure-automation.net"

Resolve-DnsName $AutomationHost
Test-NetConnection $AutomationHost -Port 443

Get-NetIPConfiguration | Select-Object InterfaceAlias,IPv4Address,DNSServer
Get-NetRoute -AddressFamily IPv4 |
Where-Object DestinationPrefix -eq "0.0.0.0/0" |
Select-Object InterfaceAlias,NextHop,RouteMetric

# Preserve proxy, firewall and TLS denials from the same UTC window.

Inspect DNS, system proxy, TLS inspection, UDR, NSG, and firewall state without opening all HTTPS traffic. If policy relies on FQDN rules or service tags, compare the resolved destination and effective path with the deployed rule. The proof is not merely that a port answers. The worker must resolve the expected name, leave through the intended route, and complete TLS without incompatible interception.

Read local evidence before restarting

On Windows, check HybridWorkerService, the Microsoft-SMA/Operational event log, and extension logs under C:\WindowsAzure\Logs\Plugins\Microsoft.Azure.Automation.HybridWorker.HybridWorkerForWindows*. On Linux, check hwd.service, /home/hweautomation/run/worker.log, and /var/log/azure/Microsoft.Azure.Automation.HybridWorker.HybridWorkerForLinux.

bash 03-linux-hybrid-worker-evidence.sh
sudo systemctl status hwd.service --no-pager
sudo journalctl -u hwd.service --since "2026-08-14 14:55:00 UTC" --until "2026-08-14 15:30:00 UTC" --no-pager

sudo tail -n 250 /home/hweautomation/run/worker.log
sudo find /var/log/azure/Microsoft.Azure.Automation.HybridWorker.HybridWorkerForLinux -type f -mmin -180 -maxdepth 3 -print

ps -eo pid,ppid,user,%cpu,%mem,etime,cmd --sort=-%cpu | head -n 25
df -h
free -m

Copy the logs and record their hashes before a restart. Look for a ping break, job download failure, process creation error, CPU limit, disk pressure, missing module, or local permission denial. The Linux extension runtime uses the hweautomation account. A command that succeeds as root does not prove that the job can read a file, load a binary, or write to the expected directory.

Distinguish saturation, runtime and identity

Every active worker polls the service periodically and accepts a bounded number of jobs. A burst of scheduled work can therefore look like an outage when the group is saturated or poorly balanced. Measure job arrival rate, queue age, concurrent executions, CPU, memory, and disk. Add capacity only after heartbeat and network health are proven and the workload genuinely exceeds the group profile.

If dispatch succeeds but execution stalls, compare the local environment with the runbook contract: PowerShell or Python version, modules, environment variables, local account, private dependency reachability, and Azure identity. Do not reinstall the extension to fix an application module. Do not grant Contributor to hide that the runtime obtained a different identity than expected.

text hybrid-worker-cause-proof.txt
Accepted cause: hwd stopped after an OS update
Evidence
HybridWorkerPing absent since 14:58Z
No job dispatched to this worker after 14:58Z
Outbound TCP 443 and DNS resolution are healthy
journalctl records the service startup failure
Other workers in the group complete a canary

Rejected cause: runbook code
The same published version completes as a canary
No business side effect is found for the blocked job

Bounded action
Correct the systemd unit, then restart hwd.service
Do not retry the business job before worker validation

Validate with a no-side-effect canary

After the repair, do not use the incident job as a test. Keep a canary runbook that writes a correlation ID, reports worker and runtime details, tests one read-only dependency, and exits. Target the same group and observe the complete cycle: pickup, process start, output, completion, and continued heartbeat.

The canary must fail closed. It should not restart a service, change a role, or consume business work. If it passes on one worker and fails on another, remove the failing host from rotation and repair it independently. If it fails everywhere, return to the shared layer: Automation account, worker group, network path, or runbook publication.

Decide resume, repair or rollback

End recovery with an explicit decision. Search for partial effects using the job ID and UTC window before retrying business work. A non-idempotent job is safe to retry only when target state proves the first action did not complete or when an idempotency key protects the resume path.

text hybrid-worker-recovery-decision.txt
Resume the job
Worker and heartbeat are healthy
Canary completes on the same group
No partial effect, or idempotent resume is proven
Identity, modules and dependencies match the contract

Repair without replay
Extension, local service or outbound path explains the failure
Business state for the incident job remains uncertain
Evidence is preserved before restart

Remove one worker from rotation
Failure is isolated to one machine
Remaining workers can absorb the workload
No RBAC or firewall widening is required

Rollback
Revert the last proven extension, proxy or OS configuration change
Restore the previous outbound path
Replay the canary and verify HybridWorkerPing
Keep the business job blocked while its effects remain unknown

Conclusion

A Hybrid Runbook Worker that stops picking up jobs should be diagnosed as a production chain: Azure Automation dispatch, heartbeat, extension, local service, outbound HTTPS, capacity, runtime, identity, and business dependency. Restarting or retrying too early mixes these layers and can replay an action that already completed partially.

A sound recovery leaves verifiable evidence: a cause supported by logs, a bounded repair, a no-side-effect canary, and a documented decision for the original job. The worker is operable again when job pickup and safe resume are proven, not merely when its status turns green.