Prerequisites
- Python with prometheus-client and opentelemetry-sdk for the synthetic console example.
- For a real rollout: deployment events, counters scraped over time, version identity, and comparable traffic cohorts.
Install dependencies
python3 -m pip install prometheus-client opentelemetry-sdk
Runnable example
import json
from datetime import datetime, timezone
from prometheus_client import Counter, CollectorRegistry, generate_latest
from opentelemetry.sdk.resources import Resource
registry = CollectorRegistry()
requests = Counter("checkout_requests_total", "Completed requests",
["service", "environment", "version", "outcome"], registry=registry)
for version, errors in [("1.0.0", 2), ("1.1.0", 12)]:
resource = Resource.create({"service.name": "checkout",
"service.version": version, "deployment.environment.name": "demo"})
print(json.dumps({"event": "deployment", "timestamp": datetime.now(timezone.utc).isoformat(),
"service": "checkout", "environment": "demo", "version": version}))
requests.labels("checkout", "demo", version, "success").inc(100 - errors)
requests.labels("checkout", "demo", version, "error").inc(errors)
print(json.dumps({"version": resource.attributes["service.version"],
"requests": 100, "errors": errors, "error_fraction": errors / 100}))
print(generate_latest(registry).decode())
Error fraction per version
sum by (service, environment, version) (
rate(checkout_requests_total{service="checkout",environment="demo",outcome="error"}[5m])
)
/
sum by (service, environment, version) (
rate(checkout_requests_total{service="checkout",environment="demo"}[5m])
)
Connect your backend
The example produces synthetic counts, version resources, and deployment events; it does not deploy software or expose an HTTP server. In an application, initialize one resource per process, attach it to the existing tracer/meter provider, and increment the counter once per completed request. Copy the process release identity into the bounded metric version label. OTel resource attributes do not automatically become Prometheus labels. Record real rollout start/end events with service, environment, version and UTC time. Keep instance identity when scraping individual counters; aggregate after calculating rates. Overlay those events on your dashboard and compare old and new cohorts over the same window.
Verification checklist
- Run the example: version 1.0.0 has 2 errors per 100 requests and version 1.1.0 has 12. These synthetic fractions are not Prometheus rates.
- For scraped production counters, query the error fraction using the PromQL example below. Confirm enough samples, nonzero request rate, and the expected version labels.
- Compare request rate, route/region mix, latency and dependency health in old and new cohorts; a timing match alone does not establish cause.
- Confirm release identity on a known log or trace and correlate the dashboard window with the deployment event.
Common failures
| Symptom | Check |
|---|---|
| Version absent from metrics | Explicitly map release metadata to metric labels; backend resource-to-label mapping depends on your exporter. |
| Misleading rollout comparison | Low-volume cohorts, changed traffic mix and overlapping windows can distort a ratio. Compare denominators and matched routes/regions. |
| Restart-related spikes | Use rate on each counter series before summing, so instance resets can be handled. Do not divide raw cumulative counters for a time-window view. |
| Zero or missing traffic | An absent numerator or zero denominator is not proof of a healthy release. Show request rate and missing-data state beside the error fraction. |