Two nines — 99% availability
Allowed downtime: 3d 15h 36m per year, 7h 18m per month, and 14m 24s per day. This leaves a multi-hour interruption budget and is usually reserved for workloads where extended downtime is acceptable.
Three nines — 99.9% availability
Allowed downtime: 8h 45m 36s per year, 43m 49s per month, and 1m 26s per day. A single long incident can consume the annual allowance, so alerting and a practiced recovery path matter.
Four nines — 99.99% availability
Allowed downtime: 52m 35s per year, 4m 22s per month, and 8.6s per day. This normally requires redundancy, automated failover, disciplined changes, and maintenance that does not interrupt service.
Five nines — 99.999% availability
Allowed downtime: 5m 15s per year, 26.3s per month, and 0.86s per day. This is a very small end-to-end error budget; component availability claims alone do not prove the whole service meets it.
Six nines — 99.9999% availability
Allowed downtime: 31.5s per year, 2.6s per month, and 0.086s per day. This is an exceptional target that requires failure detection and recovery measured in seconds or less across the full service path.
The "nines" naming convention
The "nines" shorthand counts the number of 9s in the availability percentage. 99.9% has three 9s — three nines. 99.99% has four 9s — four nines. The naming convention breaks down above six nines, where it becomes impractical to express availability as a simple percentage.
Each additional nine reduces allowed downtime by approximately 90%. Going from three nines to four nines cuts annual allowed downtime from 8.75 hours to 52 minutes — a 10× reduction.
Composite availability
When a system depends on multiple components, the overall availability is the product of each component's availability. If service A has 99.9% availability and service B has 99.9% availability, the combined system has approximately 99.8% availability (0.999 × 0.999 = 0.998).
This is why distributed systems are harder to keep available than single-server systems — each dependency is a potential failure point that multiplies downtime risk.
SLA vs SLO — which tier to commit to?
External SLAs should always be set below internal SLO targets. If your engineering team targets 99.99% internally, committing to 99.99% externally leaves no buffer — any SLO breach immediately triggers a contractual violation.
A common practice is to set the external SLA one tier below the internal SLO. Target 99.99% internally, commit to 99.9% externally. The gap is your operational buffer.
See SLO vs SLA vs SLI for a full explanation of the difference.