When to use this pattern

Use this pattern when application degradation may be associated with resource pressure, restarts, or placement changes.

Investigation flow

  1. Identify the affected application instances and the resources hosting them during the incident.
  2. Match infrastructure and application observations by stable resource identity and environment.
  3. Compare resource pressure or lifecycle events with request errors and latency on those instances.
  4. Compare unaffected instances and inspect application traces, logs, or profiles for the failure mechanism.

Required fields

Fields that make the connection possible
Field or dimensionPurpose
host, container, or pod identityMatch the actual runtime during the event, not a later reused display name.
service and environmentAssociate the runtime with the application under investigation.
event time and measurement intervalAccount for rescheduling and aggregation differences.

Worked example: A node problem affects only part of the service

Illustrative scenario

Checkout errors cluster on pods that were running on one node. Node pressure and restart events overlap those errors, while pods on other nodes stay healthy. This comparison narrows the hypothesis; logs and traces establish whether restarts interrupted requests or another condition caused the failures.

Example observations and the next comparison
EvidenceObservation
InfrastructureResource pressure and restarts on node A
ApplicationErrors concentrated on instances placed on node A at the time
Control groupInstances on other nodes under comparable load remain healthy.

Limitations and false matches

  • Rescheduled instances and reused host names can produce false joins.
  • A host-wide measurement includes other workloads on that host.
  • CPU saturation, CPU throttling, and time spent waiting describe different mechanisms.

Verification checklist

  • Check resource metadata on both infrastructure and application telemetry.
  • Confirm instance-to-host mapping for the incident interval.
  • Verify a cross-environment query cannot merge production and staging resources.

Supported by

Documented examples, not an exhaustive compatibility list. Features require suitable instrumentation and configuration; availability can depend on the runtime, backend, and subscription.

  • OpenTelemetry — Resource attributes identify the service and runtime producing telemetry.
  • Datadog — Consistent env, service, and version tags connect supported telemetry views.

Related signals

Related concepts

Related patterns

Related guides

FAQ

Can matching timestamps establish infrastructure causation?

No. Resource identity, affected-versus-unaffected comparisons, and a plausible failure mechanism are also needed.