When to use this pattern

Use this pattern when a service is slow or failing and its local resource measurements do not explain the symptom.

Investigation flow

  1. Scope the affected service, environment, operation, and interval.
  2. Inspect client spans and observed service edges to identify slow or failing dependencies.
  3. Compare downstream timing, retry behavior, and error responses with healthy requests.
  4. Check the dependency directly and distinguish upstream symptoms from a shared underlying failure.

Required fields

Fields that make the connection possible
Field or dimensionPurpose
service identity and dependency identityAvoid conflating different environments or similarly named backends.
trace context and client/server span relationshipFollow recorded calls across instrumented boundaries.
operation and time windowSeparate one failing endpoint from healthy traffic.

Worked example: Trace a cascading inventory failure

Illustrative scenario

Checkout latency rises while its CPU remains normal. Client spans show repeated inventory retries; inventory traces show database waits. The observed chain directs the investigation toward the database, but separate database health evidence is still needed to test that explanation.

Example observations and the next comparison
EvidenceObservation
CheckoutLong inventory client spans and retries
InventoryDatabase client spans dominate recorded request duration
Dependency checkCompare database wait indicators and unaffected callers.

Limitations and false matches

  • A service graph shows observed edges and may omit uninstrumented or unsampled calls.
  • Missing server spans can make transport and processing time hard to separate.
  • A dependency can be slow because of its own downstream dependency or overload from callers.

Verification checklist

  • Generate a known cross-service request and verify the expected edge appears.
  • Check context propagation and whether both ends of the edge are collected.
  • Compare the observed graph with the deployment architecture to identify blind spots.

Supported by

Documented examples, not an exhaustive compatibility list. Features require suitable instrumentation and configuration; availability can depend on the runtime, backend, and subscription.

  • Grafana Tempo — A configured processor derives observed service edges from recorded spans.

Related signals

Related concepts

Related patterns

Related guides

FAQ

Does a service graph show every dependency?

No. It is built from available observations. Static architecture information can reveal dependencies missing from recorded traces.