When to use this pattern

Use this pattern when errors begin during or shortly after a rollout.

Investigation flow

  1. Identify the rollout interval, service, environment, version, and affected instances.
  2. Compare failed requests divided by total requests for each version in equivalent windows.
  3. Inspect endpoint mix, request volume, and dependencies for differences between cohorts.
  4. Use traces and logs to look for a release-specific mechanism, then apply the rollout policy.

Required fields

Fields that make the connection possible
Field or dimensionPurpose
service, environment, and versionTie telemetry to the intended release cohort.
deployment ID and rollout intervalSeparate staged rollout events from a single global release timestamp.
failed and total requestsNormalize error count by exposure to traffic.

Worked example: A canary makes the cohort visible

Illustrative scenario

The new version receives 100 requests and has 8 failures; the old version has 5 failures among 1,000 requests. The observed rates are 8% and 0.5%. Investigate the difference while checking traffic mix and the small canary sample; counts alone would make the older version look almost equally affected.

Example observations and the next comparison
EvidenceObservation
New cohort8 / 100 requests failed = 8% observed error rate
Old cohort5 / 1,000 requests failed = 0.5% observed error rate
Next checkMatch endpoints, regions, and workload before attributing the difference to code.

Limitations and false matches

  • Small cohorts produce noisy estimates and need more evidence.
  • Feature flags, configuration, and dependency changes can coincide with a release.
  • A rollback improving the symptom supports a hypothesis but does not explain the failure mechanism on its own.

Verification checklist

  • Confirm requests from two versions are distinguishable in the same environment.
  • Check the denominator and error definition are consistent.
  • Record the decision and compare service health after the chosen action.

Supported by

Documented examples, not an exhaustive compatibility list. Features require suitable instrumentation and configuration; availability can depend on the runtime, backend, and subscription.

  • Datadog — Version-based service views support release comparisons.
  • Grafana — Annotations display events alongside metric time series.

Related signals

Related concepts

Related patterns

Related guides

FAQ

Is error rate the same as change failure rate?

No. Error rate describes requests. Change failure rate describes deployments requiring intervention; use the DORA reference for that delivery metric.