MTBF and MTTR are the two numbers behind most equipment reliability discussions: how long a machine runs between failures, and how long it takes to get it running again. Together they determine availability, the fraction of time equipment is ready to work.
This guide gives the definitions and formulas, warns about the most common way MTBF is misread, and works through a packaging machine example that compares two ways to cut downtime. It also covers how to collect data that makes the numbers trustworthy.
Why These Numbers Matter
They Turn Downtime Into a Measurable Quantity
MTBF and MTTR separate how often equipment fails from how long it takes to get back to work, which point to different fixes.
They Drive Availability
Availability is set by the ratio of MTBF to MTTR. Improving either one raises it.
They Guide Spares and Staffing
MTTR shows where repair time goes, and MTBF shows how often the crew and spares will be needed.
They Are Easy to Misread
An MTBF is an average, not a guaranteed life. Misreading it leads to wrong maintenance intervals.
Definitions
| Term | Formula | Meaning |
|---|---|---|
| MTBF | Total operating time / number of failures | Average operating time between failures of a repairable item. |
| MTTR | Total repair time / number of repairs | Average time to restore an item after a failure. |
| MTTF | Total operating time / number of items failed | Average time to failure of a non-repairable item, such as a light bulb. |
| Failure rate (λ) | 1 / MTBF (constant-rate case) | Failures per operating hour. |
| Inherent availability | MTBF / (MTBF + MTTR) | Fraction of time up, counting only failures and repairs. |
| Reliability R(t) | e−t / MTBF | Probability of no failure over time t, if the failure rate is constant. |
The Most Common Misreading
An MTBF of 450 hours does not mean an item will last 450 hours. With a constant failure rate, the probability of surviving a full MTBF is e−1, about 36.8%, so roughly 63% of items fail before reaching it. MTBF is an average of a spread of failure times, not a guaranteed life.
It also assumes failures happen at a roughly constant rate. Equipment that wears out, or that fails mostly early in life, needs a different model, such as Weibull analysis. See the Product Lifecycle Reliability Analysis tool and the Weibull Analysis entry.
Worked Example: A Packaging Machine
A packaging machine ran for 4,050 hours over six months and failed 9 times. Repairs took a total of 72 hours. The numbers are illustrative.
| Measure | Calculation | Result |
|---|---|---|
| MTBF | 4,050 / 9 | 450 hours |
| MTTR | 72 / 9 | 8 hours |
| Availability | 450 / (450 + 8) | 98.25% |
| Reliability over 100 hours | e−100/450 | 80.1% |
| Expected downtime per 8,760-hour year | 8,760 × (1 − 0.9825) | about 153 hours |
Two improvement options are compared, each of which improves the MTBF-to-MTTR ratio from 56 to 113:
- Halve MTTR from 8 to 4 hours, through spares kits, a better procedure, and pre-staged tools.
- Double MTBF from 450 to 900 hours, through better preventive tasks or a design change.
The lesson is that neither number is automatically the best lever. Compare the cost and effort of each option. Halving MTTR by reorganizing the spares crib may cost far less than redesigning a component to double its life, and in the packaging example it delivers the same 76 hours less downtime a year.
Try your own numbers in the Reliability and Availability Calculator, and track them monthly with the Maintenance KPI Calculator.
Getting Good Data
- Define a failure. Decide whether a short stop, a quality reject, or a minor jam counts, and apply the same rule every time.
- Separate repair time from waiting time. MTTR that includes waiting for parts hides the reason for delay. Record active repair, waiting for parts, and waiting for people separately.
- Use operating time, not calendar time. Equipment that runs one shift a day accumulates hours differently from one that runs continuously.
- Track by failure mode. An overall MTBF averages very different problems. Break it down with a Pareto of causes.
- Use enough data. Nine failures give a rough estimate. Report the sample size along with the MTBF.
Self-Assessment Questions
- Do we have a clear, written definition of a failure?
- Do we record operating hours and repair hours for each event?
- Do we separate active repair from waiting time?
- Do we analyze by failure mode, not just overall?
- Do we test whether a constant failure rate is reasonable before relying on MTBF?
Common Mistakes
Treating MTBF as a Guaranteed Life
About 63% of items fail before reaching the MTBF when the failure rate is constant. Do not set replacement intervals from MTBF alone.
Mixing Failure Modes
Averaging unrelated problems hides the ones that matter. Analyze by mode.
Including Waiting in MTTR Without Saying So
It is fine to measure it, but separate it so you can fix delays in spares and staffing.
Comparing Numbers With Different Definitions
MTBF from two sites means little if they define failure differently.
MTBF, MTTR, and Availability: Frequently Asked Questions
What is the difference between MTBF and MTTR?
MTBF, mean time between failures, is the average operating time between failures of a repairable item, and measures how often it fails. MTTR, mean time to repair, is the average time to restore it after a failure, and measures how long each failure lasts. Availability combines them: MTBF divided by MTBF plus MTTR.
Does an MTBF of 1,000 hours mean the equipment lasts 1,000 hours?
No. With a constant failure rate, the probability of running a full MTBF without failure is about 37%, so most items fail before reaching it. MTBF is an average of a spread of failure times, and equipment with wear-out behavior needs a different model such as Weibull analysis.
How do I improve availability?
Availability depends on the ratio of MTBF to MTTR, so either fewer failures or faster repairs improve it. Compare the cost of each option, since shortening repair time through spares, procedures, and staging is often cheaper than extending time between failures through design changes.
Sources and Further Reading
- NIST/SEMATECH e-Handbook of Statistical Methods, reliability chapter.
- Patrick D. T. O'Connor and Andre Kleyner, Practical Reliability Engineering.
- IEC 60050-192, International Electrotechnical Vocabulary: Dependability.
- ASQ Certified Reliability Engineer Body of Knowledge.