Automation
Azure Automation: rotate a webhook before the production trigger expires
A production runbook for inventorying Azure Automation webhook consumers, introducing a second endpoint, proving one execution per event, cutting over and revoking the old URL with a tested rollback.
An Azure Automation webhook is approaching expiration. It starts a published runbook from an alerting or orchestration platform, but the team no longer knows every consumer of the URL. Extending the date looks harmless. Replacing the webhook looks cleaner. Either move can break the trigger, preserve an exposed URL for years, or run the same production action twice during cutover.
The running case is a webhook that starts Invoke-PlatformRemediation after an approved monitoring event. The goal is not merely to obtain a new URL. It is to prove who calls the endpoint, what the runbook will accept, how duplicate requests are contained, and when the old credential can be disabled without losing the rollback path.
Treat the URL as a credential and the trigger as a contract
An Automation webhook URL contains the token that authorizes invocation. A caller that knows it can submit a POST without a separate Azure identity. The URL is shown when the webhook is created and must be captured then; do not assume it can be recovered later from the resource inventory.
Freeze the operational contract before changing anything:
automation_account: aa-platform-prod
runbook: Invoke-PlatformRemediation
old_webhook: remediation-prod-v1
old_expiry_utc: <timestamp>
expected_callers:
- alerting-platform
- approved-orchestrator
execution_target: hybrid-worker-prod
required_fields: [eventId, targetId, action, requestedAt]
deduplication_key: eventId
maximum_event_age: 10m
success_evidence: one accepted request produces one traceable job
rollback: re-enable old caller configuration while old webhook remains enabled Record the published runbook version, fixed webhook parameters, Hybrid Worker group, managed identity and target scope. The webhook authorizes the start, but the runbook identity authorizes the effect. Both boundaries must remain stable during rotation.
Inventory consumers from configuration and evidence
Do not rely on the webhook name to identify callers; the external system only needs the URL. Search approved secret stores, alert receivers, Logic Apps, deployment variables and orchestration configuration for the reference. Do not print the URL into a terminal transcript, ticket or CI log while searching.
Then compare declared consumers with runtime evidence. Inventory the webhook metadata and recent jobs without exposing request secrets:
$scope = @{
ResourceGroupName = '<resource-group>'
AutomationAccountName = 'aa-platform-prod'
}
$webhook = Get-AzAutomationWebhook @scope -Name 'remediation-prod-v1'
$webhook | Select-Object Name, IsEnabled, ExpiryTime, LastInvokedTime, RunOn,
@{n='Runbook';e={$_.RunbookName}}, Parameters
Get-AzAutomationJob @scope -StartTime (Get-Date).AddDays(-14) |
Where-Object RunbookName -eq 'Invoke-PlatformRemediation' |
Select-Object JobId, CreationTime, StartTime, EndTime, Status Job history proves execution, not caller ownership. Correlate the accepted eventId and a non-secret caller identifier emitted by the runbook with the source platform. If the current runbook cannot attribute a request without recording the token or full payload, fix that observability gap before rotating a critical trigger.
Reject stale, malformed and duplicate requests inside the runbook
The webhook URL alone does not prove business intent. The runbook should accept a narrow payload, validate allowed targets and actions, reject events outside a short age window, and claim a deduplication key before any side effect.
param([object] $WebhookData)
if (-not $WebhookData) { throw 'WebhookData is required' }
$request = $WebhookData.RequestBody | ConvertFrom-Json -Depth 10
$allowedActions = @('restart-approved-service', 'refresh-approved-cache')
if ([string]::IsNullOrWhiteSpace($request.eventId)) { throw 'eventId is required' }
if ($request.action -notin $allowedActions) { throw 'action is not allowed' }
if ($request.targetId -notlike '/subscriptions/<approved-subscription>/*') {
throw 'target is outside the approved scope'
}
$requestedAt = [DateTimeOffset]::Parse($request.requestedAt)
if ([DateTimeOffset]::UtcNow - $requestedAt -gt [TimeSpan]::FromMinutes(10)) {
throw 'event is stale'
}
# Atomically create or claim eventId in the approved state store.
# If it already exists, emit duplicate_ignored and exit before any write. Use a state store whose conditional create is atomic. A process-local variable, job history query or eventual log search is not a deduplication lock. Retain the key long enough to cover caller retries and operator replay. Emit eventId, webhook generation, validation result, job ID and final outcome, but never the webhook URL or a payload containing secrets.
Azure Automation records runbook input parameters with the job. Keep sensitive values out of the request body and headers. If the operation needs stronger caller authentication, stateful delivery or job tracking, use an authenticated API or queue in front of the runbook rather than stretching a bearer URL beyond its trust model.
Create a second webhook without changing the first
Create a new, separately named webhook linked to the same published runbook, with the same fixed parameters and worker target. Set a deliberate expiry that matches the ownership and review cycle. Capture the returned URI directly into the approved secret store without displaying it in pipeline output.
$webhookParameters = @{
ResourceGroupName = '<resource-group>'
AutomationAccountName = 'aa-platform-prod'
RunbookName = 'Invoke-PlatformRemediation'
Name = 'remediation-prod-v2'
RunOn = 'hybrid-worker-prod'
Parameters = @{ Environment = 'prod'; Mode = 'bounded' }
ExpiryTime = [DateTimeOffset]::UtcNow.AddMonths(12)
IsEnabled = $true
Force = $true
}
$new = New-AzAutomationWebhook @webhookParameters
# Send $new.WebhookURI directly to the approved secret update step.
# Do not write the object, URI or deployment output to logs. Creating the second endpoint is the rollback mechanism. Keep v1 enabled while v2 is tested, but do not let normal traffic fan out to both. The overlap window must have an owner, an end time and monitoring for duplicate event IDs.
If the existing webhook is still trusted and only its expiry is wrong, extending it before expiration may be a valid emergency measure. It is not a rotation: the credential remains unchanged and every unknown holder keeps access. Use extension to avoid an outage only with a scheduled replacement and a short, reviewed horizon.
Test one request through the new path
Use a synthetic event whose target is a canary or whose action resolves to a dry run. The source platform must send it through the same secret resolution and HTTP path as production. A manual curl proves reachability but not the real integration.
Verify four linked records:
- the source emitted one event with a unique
eventId; - the new webhook accepted it and started one Automation job;
- the runbook recorded validation, deduplication and the expected worker target;
- the canary produced no production side effect.
Resend the same event intentionally. The second request may start another job because the webhook endpoint is a trigger, but the runbook must recognize the same eventId and exit before the action. Also test an expired timestamp, an unknown action and an out-of-scope target. A rotation that validates only the happy path preserves the most dangerous failure modes.
Cut over one consumer cohort at a time
Update the smallest caller cohort to reference the new secret version. Observe a full operating window before moving the next cohort. During the overlap, track accepted requests by webhook generation and deduplication result.
For each eventId retain
source system and source event timestamp
webhook generation: v1 or v2
Automation job ID and worker group
validation result and deduplication result
requested action and bounded target identifier
side-effect result and post-action validation
Stop the cutover when
one event reaches both webhook generations
a caller cannot identify which secret version it loaded
requests arrive without a stable eventId
the runbook records a wider target or action than the contract Do not infer completion from LastInvokedTime alone. A low-frequency consumer may be healthy but silent. Obtain an explicit owner confirmation or a controlled event from every declared consumer before revoking v1.
Disable, observe, then remove the old webhook
Once all consumers are proven on v2, disable the old webhook. Disabling is preferable to immediate deletion because it gives a short, explicit rollback window without keeping the credential active.
$scope = @{
ResourceGroupName = '<resource-group>'
AutomationAccountName = 'aa-platform-prod'
}
Set-AzAutomationWebhook @scope -Name 'remediation-prod-v1' -IsEnabled $false
Get-AzAutomationWebhook @scope -Name 'remediation-prod-v1' |
Select-Object Name, IsEnabled, ExpiryTime, LastInvokedTime Watch source delivery failures, job starts and the absence of new v1 correlations for the agreed window. Remove the old secret from every consumer and then delete the disabled webhook. If the URL was exposed, skip the long overlap: disable it, contain the caller, and use an authenticated fallback while the new path is validated.
Decide validation or rollback
Promote the rotation when every known caller uses v2, repeated and duplicate tests behave as designed, job attribution is complete, and the old webhook remains unused while disabled. Remove temporary diagnostics and store the next expiry owner and alert with the service record.
Rollback means restoring the caller reference to v1 only while that webhook is still trusted and enabled. Roll back the caller configuration, not the runbook contract, then investigate why v2 failed. Never re-enable a URL that was rotated because it leaked. In that case, the rollback path is a separately authenticated trigger or a newly created webhook with a new token.
Conclusion
Rotating an Azure Automation webhook is a credential change and a delivery change at the same time. The safe sequence is inventory, guard the runbook, create a parallel endpoint, test normal and duplicate requests, cut over by cohort, disable, observe and delete.
The production decision is evidence-based. Keep v2 only when each caller and each event can be traced to one bounded effect. Roll back the caller reference when the new path fails but the old credential remains trusted. If ownership, deduplication or attribution is missing, hold the rotation before a silent trigger outage becomes an unsafe replay.