Definition

Metrics summarize numerical observations over time, such as request counts, latency distributions, and resource usage.

Question it answers

Which service or cohort changed, and how large is the impact?

Role in an investigation

Start an investigation with the affected population and time range. Compare request rate, errors, and latency before selecting a representative trace. Keep the aggregation dimensions visible so a fleet-wide average does not conceal a failing instance.

Correlation fields

Fields that make the connection possible
Field or dimensionPurpose
service and environment dimensionsMatch the scope of the chart to the scope of other telemetry.
time window and aggregationCompare equivalent intervals and understand whether values are rates, totals, or distributions.
exemplar trace referenceLink selected observations to recorded requests without creating per-trace metric series.

Narrow a latency regression

Checkout p95 latency rises while request volume stays steady. Splitting by version isolates the newer cohort; an exemplar supplies one request to inspect. Check more requests before treating that example as representative of the full regression.

Limitations

  • Aggregates omit per-request details and can hide differences between cohorts.
  • Percentiles generally cannot be averaged across instances to obtain a fleet percentile.
  • Exemplar availability depends on collection, sampling, storage, and retention.

Supported by

Documented examples, not an exhaustive compatibility list. Features require suitable instrumentation and configuration; availability can depend on the runtime, backend, and subscription.

Investigation patterns

Related concepts

Related guides

FAQ

Does every metric point have a trace?

No. Metrics may describe many operations or infrastructure state. Exemplars link selected observations only.