When to use this pattern
Use this pattern when errors begin during or shortly after a rollout.
Investigation flow
- Identify the rollout interval, service, environment, version, and affected instances.
- Compare failed requests divided by total requests for each version in equivalent windows.
- Inspect endpoint mix, request volume, and dependencies for differences between cohorts.
- Use traces and logs to look for a release-specific mechanism, then apply the rollout policy.
Required fields
| Field or dimension | Purpose |
|---|---|
| service, environment, and version | Tie telemetry to the intended release cohort. |
| deployment ID and rollout interval | Separate staged rollout events from a single global release timestamp. |
| failed and total requests | Normalize error count by exposure to traffic. |
Worked example: A canary makes the cohort visible
The new version receives 100 requests and has 8 failures; the old version has 5 failures among 1,000 requests. The observed rates are 8% and 0.5%. Investigate the difference while checking traffic mix and the small canary sample; counts alone would make the older version look almost equally affected.
| Evidence | Observation |
|---|---|
| New cohort | 8 / 100 requests failed = 8% observed error rate |
| Old cohort | 5 / 1,000 requests failed = 0.5% observed error rate |
| Next check | Match endpoints, regions, and workload before attributing the difference to code. |
Limitations and false matches
- Small cohorts produce noisy estimates and need more evidence.
- Feature flags, configuration, and dependency changes can coincide with a release.
- A rollback improving the symptom supports a hypothesis but does not explain the failure mechanism on its own.
Verification checklist
- Confirm requests from two versions are distinguishable in the same environment.
- Check the denominator and error definition are consistent.
- Record the decision and compare service health after the chosen action.
Supported by
Documented examples, not an exhaustive compatibility list. Features require suitable instrumentation and configuration; availability can depend on the runtime, backend, and subscription.
- Datadog — Version-based service views support release comparisons.
- Grafana — Annotations display events alongside metric time series.
Related signals
Related concepts
Related patterns
Related guides
FAQ
Is error rate the same as change failure rate?
No. Error rate describes requests. Change failure rate describes deployments requiring intervention; use the DORA reference for that delivery metric.