Capability

Incident evidence & outage attribution

Your site went down. Was it you, or your vendors? Correlated timelines give the answer - with confidence levels, not guesswork.

Who this is for

On-call engineers, incident commanders and platform teams who need to route an active incident to the right owner in minutes - your codebase or a specific vendor - and defend that routing in the postmortem.

The core problem

An outage you caused and an outage your vendor caused look identical from inside your own monitoring. Both present as your service failing. Opening two tabs - your dashboard and the vendor’s status page - is weaker than it feels: the status page is human-written, approximate, and owned by the counterparty.

How attribution works

  • Shared timeline: your incident window and the dependency’s independent observations on one axis.
  • Quorum verdicts: multi-region failures confirm vendor-side degradation; single-region disagreement suggests a path problem.
  • Confidence levels: the engine reports how strongly the timelines overlap - never a bare “vendor did it.”
  • No causation claims: correlated failure is strong evidence for where to look first, not proof of cause.

Practical example

Checkout errors spike 14:02–14:19. The payment dependency shows timeouts from two regions across the same window; your database and queue telemetry stay flat. Attribution routes the incident to the vendor, the deploy rollback is stood down, and the fault report attaches to the vendor ticket.

What to do next

Read the Dependency Gap, configure monitoring, or compile the evidence.

Know when your dependencies fail. Prove what happened.

Monitor the external APIs your product relies on, correlate their failures with your incidents, and generate verifiable evidence when a vendor causes downtime.