When to use this pattern
Use this pattern after an aggregate latency or error signal identifies an affected service and interval.
Investigation flow
- Choose a service, environment, endpoint, and comparable time window on the metric chart.
- Open an exemplar when the metric and backend supply one; otherwise search traces using the same scope.
- Inspect the request path and its relevant logs.
- Check several requests and unaffected cohorts before generalizing from a sample.
Required fields
| Field or dimension | Purpose |
|---|---|
| service and environment | Keep chart and trace search in the same population. |
| time window and endpoint | Narrow the observation to the operation that changed. |
| exemplar trace ID | Identify a selected observation without adding an unbounded metric label. |
Worked example: Investigate a checkout latency spike
A checkout latency histogram rises for the new version. One exemplar opens a trace with a slow inventory call. More traces show the same boundary while the older version remains stable. The repeated pattern supports investigating the inventory path; the first trace alone would be weak evidence.
| Evidence | Observation |
|---|---|
| Metric | Checkout request latency, production, version v2 |
| Exemplar | A selected trace reference on the affected series |
| Comparison | Several v2 traces versus v1 traces under comparable traffic |
Limitations and false matches
- A percentile represents a distribution, not one request whose duration equals that percentile.
- Sampling and retention can make an exemplar link unavailable.
- Metrics derived from sampled traces describe the collected sample unless the pipeline accounts for sampling.
Verification checklist
- Emit a test operation and confirm an exemplar can open its trace.
- Check that trace-search filters preserve chart dimensions and time range.
- Keep a scoped search available when no exemplar exists.
Supported by
Documented examples, not an exhaustive compatibility list. Features require suitable instrumentation and configuration; availability can depend on the runtime, backend, and subscription.
- Grafana with Prometheus and Tempo — Configured exemplars provide selected metric-to-trace links.
Related signals
Related concepts
Related patterns
Related guides
- What is Observability?
- What is MTTR?
- Grafana: Metrics to Traces with Exemplars
- Troubleshoot: Exemplars Missing in Grafana
FAQ
Does an exemplar explain the whole spike?
No. It supplies one recorded example. Compare multiple requests and the metric population before drawing conclusions.