When to use this pattern
Use this pattern when a service is slow or failing and its local resource measurements do not explain the symptom.
Investigation flow
- Scope the affected service, environment, operation, and interval.
- Inspect client spans and observed service edges to identify slow or failing dependencies.
- Compare downstream timing, retry behavior, and error responses with healthy requests.
- Check the dependency directly and distinguish upstream symptoms from a shared underlying failure.
Required fields
| Field or dimension | Purpose |
|---|---|
| service identity and dependency identity | Avoid conflating different environments or similarly named backends. |
| trace context and client/server span relationship | Follow recorded calls across instrumented boundaries. |
| operation and time window | Separate one failing endpoint from healthy traffic. |
Worked example: Trace a cascading inventory failure
Checkout latency rises while its CPU remains normal. Client spans show repeated inventory retries; inventory traces show database waits. The observed chain directs the investigation toward the database, but separate database health evidence is still needed to test that explanation.
| Evidence | Observation |
|---|---|
| Checkout | Long inventory client spans and retries |
| Inventory | Database client spans dominate recorded request duration |
| Dependency check | Compare database wait indicators and unaffected callers. |
Limitations and false matches
- A service graph shows observed edges and may omit uninstrumented or unsampled calls.
- Missing server spans can make transport and processing time hard to separate.
- A dependency can be slow because of its own downstream dependency or overload from callers.
Verification checklist
- Generate a known cross-service request and verify the expected edge appears.
- Check context propagation and whether both ends of the edge are collected.
- Compare the observed graph with the deployment architecture to identify blind spots.
Supported by
Documented examples, not an exhaustive compatibility list. Features require suitable instrumentation and configuration; availability can depend on the runtime, backend, and subscription.
- Grafana Tempo — A configured processor derives observed service edges from recorded spans.
Related signals
Related concepts
Related patterns
- Trace to Logs
- Alerts to Incidents
- Service to Database
- Queue Producer to Consumer
- User Session to Backend Trace
Related guides
FAQ
Does a service graph show every dependency?
No. It is built from available observations. Static architecture information can reveal dependencies missing from recorded traces.