When to use this pattern
Use this pattern when application degradation may be associated with resource pressure, restarts, or placement changes.
Investigation flow
- Identify the affected application instances and the resources hosting them during the incident.
- Match infrastructure and application observations by stable resource identity and environment.
- Compare resource pressure or lifecycle events with request errors and latency on those instances.
- Compare unaffected instances and inspect application traces, logs, or profiles for the failure mechanism.
Required fields
| Field or dimension | Purpose |
|---|---|
| host, container, or pod identity | Match the actual runtime during the event, not a later reused display name. |
| service and environment | Associate the runtime with the application under investigation. |
| event time and measurement interval | Account for rescheduling and aggregation differences. |
Worked example: A node problem affects only part of the service
Checkout errors cluster on pods that were running on one node. Node pressure and restart events overlap those errors, while pods on other nodes stay healthy. This comparison narrows the hypothesis; logs and traces establish whether restarts interrupted requests or another condition caused the failures.
| Evidence | Observation |
|---|---|
| Infrastructure | Resource pressure and restarts on node A |
| Application | Errors concentrated on instances placed on node A at the time |
| Control group | Instances on other nodes under comparable load remain healthy. |
Limitations and false matches
- Rescheduled instances and reused host names can produce false joins.
- A host-wide measurement includes other workloads on that host.
- CPU saturation, CPU throttling, and time spent waiting describe different mechanisms.
Verification checklist
- Check resource metadata on both infrastructure and application telemetry.
- Confirm instance-to-host mapping for the incident interval.
- Verify a cross-environment query cannot merge production and staging resources.
Supported by
Documented examples, not an exhaustive compatibility list. Features require suitable instrumentation and configuration; availability can depend on the runtime, backend, and subscription.
- OpenTelemetry — Resource attributes identify the service and runtime producing telemetry.
- Datadog — Consistent env, service, and version tags connect supported telemetry views.
Related signals
Related concepts
Related patterns
Related guides
FAQ
Can matching timestamps establish infrastructure causation?
No. Resource identity, affected-versus-unaffected comparisons, and a plausible failure mechanism are also needed.