Mean Time to Recovery MTTR
Mean Time to Recovery (MTTR) measures the average time required to restore a service or system to normal operation after a failure or incident. It is one of the four DORA metrics and a critical reliability KPI. MTTR directly determines the amount of downtime caused by each incident and the impact on users and revenue. Shorter MTTR requires strong observability, clear incident response processes, and empowered on-call engineers.
MTTR is meaningfully different from Mean Time Between Failures (MTBF); MTTR focuses on recovery speed while MTBF focuses on failure prevention. Both matter for overall availability.
- PagerDutyIncident lifecycle tracking from alert to resolution
- DatadogIncident management and MTTR reporting
- OpsgenieOn-call management and incident timeline tracking
- FireHydrantIncident orchestration with MTTR analytics
- Observability depth (logs, metrics, traces) enabling fast diagnosis
- On-call response time and runbook quality
- Incident command process clarity and escalation procedures
- Automated rollback capability for deployment-caused incidents
- Blast radius of failures (isolated microservices recover faster than monoliths)
DORA elite teams recover in under 1 hour; high performers in under 24 hours; medium performers in less than 1 week; low performers in over 1 week.
How different roles think about this metric
Each function reads MTTR through a different lens and takes different actions when it changes.
Common Questions About Mean Time to Recovery
Click any question to expand the answer.
What is the difference between MTTR and MTTD (Mean Time to Detect)?
How can observability improve MTTR?
What is a blameless post-incident review?
How does runbook quality affect MTTR?
Related Metrics
Metrics that are commonly analyzed alongside MTTR.
Role guides that include this metric
See how each role uses MTTR in context with the full set of metrics they own.
See What’s Actually Moving Your MTTR
askotter connects your data sources and applies causal analysis to tell you exactly why your metrics are changing, not just that they changed.
Book a demo