Skip to content

Glossary

Latency

Time from probe dispatch to complete response - the leading indicator of degradation.

Definition

Probe latency is the wall-clock time between a scheduled check issuing its request and receiving the complete response (including any policy-validated redirect hops), recorded per observation and aggregated as a mean and 95th percentile per window.

The problem

Availability says whether a request completed; only latency says what clients experienced while it did. Most dependency incidents spend minutes-to-hours in “up but slow” territory before, during, and after the visible outage - and availability dashboards show green the whole time.

Why it matters

Timeouts are latency failures viewed from the client: a route degrading to 3 seconds turns into “outage” the moment it crosses the caller’s deadline. Watching the p95 is what separates a capacity conversation from a postmortem about a cliff nobody saw.

Practical example

A record showing p95 climbing while the mean holds describes long-tail failure (one shard, one region). RELIASTRA charts the mean per bucket with failed buckets breaking the line, and prints the p95 threshold beside it.

How RELIASTRA approaches it

Latency is measured from outside both networks, on the same schedule as availability, so a drift is timestamped in the same series that will later be cited as evidence. An aggregate mean of zero means “no successful response recorded,” and is printed as no data, never as a fast response.

Know what you depend on. Prove what it did.

RELIASTRA observes the external services your product relies on, attributes their failures, and produces evidence you can act on. Every new organization starts on a 14-day Pro trial.