Infrastructure
OpenTelemetry Collector: validate tail sampling before losing incident traces
A production runbook for qualifying OpenTelemetry tail sampling across trace affinity, late spans, memory pressure, policy coverage, Azure Monitor evidence, canary rollout and rollback.
An observability team wants to reduce trace ingestion without making incidents harder to explain. Keeping every failed or slow trace and only a baseline of successful traffic sounds safer than fixed-rate sampling. After rollout, however, some errors arrive without their upstream request, long traces disappear, and Collector memory rises during traffic peaks. The policy looks correct; the trace path is not.
The running case is a Kubernetes platform that sends OTLP data through gateway Collectors to Azure Monitor and Application Insights. The team is replacing simple head sampling with the OpenTelemetry Collector tail_sampling processor. This runbook decides whether the candidate can be promoted, must remain in shadow mode, or should be rolled back before it removes the evidence needed for the next incident.
Freeze the evidence contract before the percentage
Tail sampling is an operational policy, not just a cost setting. Write down what must remain diagnosable before choosing thresholds.
change: tail-sampling-v3
scope:
services: [checkout-api, payment-worker]
environment: production
exporter: azure-monitor
must_keep:
- trace_with_error_status
- trace_over_2000_ms
- synthetic_canary
- explicitly_debugged_trace
baseline:
successful_traces_percent: 5
required_evidence:
- root_and_child_spans_share_trace_id
- error_trace_keeps_upstream_request
- late_span_follows_original_decision
- collector_drops_are_explained
- ingestion_volume_stays_within_budget
rollback: restore-pipelines-v2-and-restart-canary-pool Record the current Collector version, configuration digest, replica count, incoming trace rate, exported trace rate and memory headroom. A lower Azure Monitor bill is not a successful change if failed traces become fragmented or absent.
Prove trace affinity before reading policies
The tail sampler groups spans by trace_id while it waits for a decision. All spans for one trace must therefore reach the same sampling Collector. A round-robin Service in front of several stateful processors can split a trace across replicas: each replica sees an incomplete story and makes a locally valid decision.
Use two explicit tiers when the gateway must scale horizontally:
Applications and agents
-> stateless ingress Collectors
-> load-balancing exporter keyed by trace ID
-> stateful tail-sampling Collectors
-> Azure Monitor / Application Insights
Invariant
Every span with the same trace_id reaches the same sampling shard
Failure to reject
Kubernetes Service distributes spans directly across tail samplers Do not infer affinity from a complete trace found once. Emit a synthetic trace with a root span, several children and one delayed child, then inspect receiver and exporter telemetry on every replica. Repeat while replicas scale and restart. If the same trace_id appears as a new trace on multiple samplers, policy tuning cannot repair the topology.
Size the decision window from observed traces
decision_wait controls how long the processor accumulates a trace before evaluation. A short window reduces memory but can decide before a slow child arrives. A long window retains more spans and raises state pressure. Measure span arrival delay from the real path, including queues and batch exporters, instead of copying a default into production.
Build a distribution for the delay between the first and last observed span. Test normal requests, slow dependencies, asynchronous work and a Collector restart. Choose a candidate window that covers the diagnostic cases in the contract, then prove its memory cost at peak new-trace rate.
num_traces is also a stop condition. When the pending set exceeds capacity, the oldest traces may leave the decision window before the team expects. expected_new_traces_per_sec helps allocation; it does not create capacity or replace load testing. Track accepted spans, refused spans, sampling decisions, processor latency and process memory together.
Make policies readable and testable
Start with a small policy set. The order and interaction of several matchers are easier to misunderstand than one explicit decision table.
processors:
tail_sampling/canary:
decision_wait: 30s
num_traces: 50000
expected_new_traces_per_sec: 1000
decision_cache:
sampled_cache_size: 100000
non_sampled_cache_size: 100000
policies:
- name: keep-errors
type: status_code
status_code:
status_codes: [ERROR]
- name: keep-slow
type: latency
latency:
threshold_ms: 2000
- name: keep-canaries
type: string_attribute
string_attribute:
key: test.kind
values: [tail-sampling-canary]
- name: successful-baseline
type: probabilistic
probabilistic:
sampling_percentage: 5
service:
pipelines:
traces/canary:
receivers: [otlp]
processors: [memory_limiter, k8sattributes, tail_sampling/canary, batch]
exporters: [azuremonitor] Treat the values as a candidate, not universal sizing. Enrichment required by a policy must run before tail sampling. Processors that need the original request context also belong before it, because reassembled spans leave the sampler in new batches.
Avoid an implicit “keep everything unusual” rule. Name the attributes, allowed values and owner. Never put unbounded user input, account identifiers or secrets into policy attributes merely to make sampling selective.
Test decisions, not only configuration syntax
A Collector that starts successfully has only proved that the YAML parses. Replay a deterministic matrix through the whole topology.
Case Expected result
Fast success Baseline policy may keep or drop the complete trace
HTTP 500 with child dependency Root, dependency and error remain together
Slow success above 2000 ms Complete trace is retained by keep-slow
Explicit canary Always retained with expected attributes
Delayed error span Follows the original decision, never forms a new fragment
Trace longer than decision_wait Known and quantified behavior, with no false success claim
Collector replica added Trace remains on one sampling shard
Sampling replica restarted Loss is bounded and visible in Collector telemetry
Traffic burst Memory stays below the stop threshold
Unknown service Follows the declared default, not an accidental allow rule Positive cases prove that required traces survive. Negative cases prove that ordinary traffic is actually reduced and that a malformed attribute cannot force every trace through. Keep the generated trace IDs and expected span counts as release evidence.
Qualify late spans and decision caches
A late span can arrive after its trace was sampled or dropped. Without a retained decision, the processor may treat it as a new trace and evaluate it again. That creates fragments or contradictory outcomes.
Set sampled and non-sampled decision caches from observed late-span volume and cardinality, then test both paths. The cache should outlive the typical late arrival and be large enough for the decision rate, but its size must still fit the memory budget. A cache is not a substitute for fixing systematically delayed exporters.
Make one canary send its final error span after the normal decision window. Verify that it inherits the earlier outcome while the decision is cached. Then repeat beyond the planned cache horizon and document the behavior. If that second case would hide a real incident pattern, extend the architecture or change where sampling occurs before promotion.
Correlate Collector decisions with Azure Monitor
Collector health shows whether the pipeline processed data. Azure Monitor shows whether the diagnostic contract arrived. Use both.
First compare decision counts by policy with receiver and exporter rates. A rise in processor drops, refused spans, queue saturation or memory limiter actions invalidates a clean-looking sampling ratio. If the Collector build supports policy attribution, enable it on the canary only first and confirm the added attributes and cardinality before a wider rollout.
In Application Insights, search for the retained canary operation IDs and count the expected layers:
let StartUtc = datetime(<start-utc>);
let EndUtc = datetime(<end-utc>);
let CanaryOperationIds = dynamic(["<fast-trace-id>", "<error-trace-id>", "<slow-trace-id>"]);
union isfuzzy=true
(requests | project timestamp, operation_Id, itemType="request", name, success),
(dependencies | project timestamp, operation_Id, itemType="dependency", name, success),
(exceptions | project timestamp, operation_Id, itemType="exception", name=type, success=false)
| where timestamp between (StartUtc .. EndUtc)
| where operation_Id in (CanaryOperationIds)
| summarize Items=count(), Types=make_set(itemType), Failed=countif(success == false) by operation_Id
| order by operation_Id asc The query is a validation aid, not a universal schema contract. Adapt table and field names to the ingestion mode in use. The release evidence must link each generated trace ID, expected span count, Collector decision and Azure Monitor result.
Also compare the retained percentage and ingestion volume over time. Metrics are not sampled like traces, so keep independent service metrics and synthetic checks as the control signal. Otherwise a policy can make the trace view look healthy by removing the failures used to calculate it.
Roll out with a shadow path and bounded canary
Mirror a small, non-sensitive slice to a candidate pipeline or route only named canary services through it. Do not send the duplicate stream to the same billable destination without a plan: use a test Application Insights resource, a debug exporter with strict limits or short-lived local evidence.
Run the matrix at normal load, peak load, during a rolling restart and while adding a replica. Compare the candidate against the current path by trace ID, not aggregate percentage alone. Promotion requires representative error and latency traces to remain complete, not just similar counts.
Expand one service at a time. Keep the previous Collector configuration and deployment artifact immutable. The rollback switch must restore the old trace pipeline independently of application deployment.
Decide promotion, hold or rollback
Promote when trace affinity survives scaling, the measured decision window covers required cases, memory remains bounded, late spans follow a consistent decision, policy tests pass, and Azure Monitor contains the expected complete canaries. Continue to watch Collector drops and ingestion cost after each service cohort.
Hold in shadow mode when the policy produces the expected percentage but span completeness, decision attribution or peak memory remains ambiguous. More traffic is not the next test; a narrower experiment is.
Roll back immediately when one trace is split across samplers, required error or slow traces disappear, memory limiting starts dropping evidence, or a late span creates a contradictory fragment. Restore the previous pipeline, restart only the affected canary pool if required, and verify new trace IDs end to end. Do not claim that rollback reconstructs telemetry already dropped.
Conclusion
Tail sampling is safe when the decision path is as observable as the traces it filters. The contract spans routing, state, time, memory, policy semantics, export and the destination query.
The production decision is therefore concrete: promote only with trace affinity, tested late-span behavior, bounded state and complete canaries in Azure Monitor; hold while any evidence is ambiguous; return to the previous pipeline at the first unexplained loss. Cost control remains useful because it preserves the traces that operations will actually need.