Reliability-Centered Maintenance is a disciplined way to decide what maintenance an asset really needs. Instead of applying the same preventive schedule to every machine, RCM asks what each asset must do, how it can fail, what each failure costs in safety, environment, and production, and which task, if any, is worth doing.
This guide walks through the seven questions, why failure does not always follow age, how to choose between condition monitoring, scheduled tasks, failure-finding, redesign, and run to failure, and how to size an inspection interval from the P-F interval. A worked example on a cooling-water pump puts costs on the decision.
Before You Start
Why RCM Matters in Energy Operations
Maintenance Effort Goes Where Failure Hurts
RCM starts from what each asset must do and what its failure costs in safety, environment, and production, so effort follows consequences rather than habit.
Stops Pointless Time-Based Work
Many failures do not correlate with age. Replacing parts on a calendar can add infant-mortality failures instead of preventing them.
Handles Hidden Failures
Protective devices such as relief valves, trips, and standby pumps can fail silently. RCM forces a plan to find those failures before the demand arrives.
Creates a Traceable Basis
Every task links to a failure mode and a consequence, which makes the maintenance program defensible and easier to update.
What Reliability-Centered Maintenance Is
Reliability-Centered Maintenance (RCM) is a structured process for deciding what maintenance each asset needs to keep doing what its users want it to do. It grew out of work in commercial aviation, documented by F. Stanley Nowlan and Howard Heap in 1978, and was adopted widely in power generation, oil and gas, and other asset-intensive industries. The SAE standard JA1011 sets the criteria that a process must meet to be called RCM.
RCM answers seven questions about every asset in its operating context:
- What are the functions and performance standards of the asset?
- In what ways can it fail to fulfil those functions (functional failures)?
- What causes each functional failure (failure modes)?
- What happens when each failure mode occurs (failure effects)?
- How does each failure matter (failure consequences)?
- What can be done to predict or prevent each failure?
- What should be done if no suitable proactive task can be found?
Failure Does Not Always Follow Age
Nowlan and Heap studied how the probability of failure changes with age in aircraft components and found six patterns. Only some show wear-out, so an age-based replacement helps only those.
| Pattern | Shape | Share of components studied | Implication |
|---|---|---|---|
| A | Bathtub: early failures, constant, then wear-out | 4% | Age limit useful only after the wear-out zone starts. |
| B | Constant, then a wear-out zone | 2% | Age-based replacement can work. |
| C | Slowly increasing failure probability | 5% | No clear age limit; condition monitoring helps. |
| D | Low at first, then constant | 7% | Age limits do not help. |
| E | Constant (random) at all ages | 14% | Age limits do not help. |
| F | High infant mortality, then constant or slowly rising | 68% | Intrusive maintenance can create early failures. |
These percentages come from a study of commercial aircraft components. Industrial and energy equipment can differ, but the lesson holds: do not assume that every asset gets less reliable as it gets older.
Choosing a Maintenance Task
| Task type | What it does | Use when |
|---|---|---|
| On-condition (predictive) | Detects a developing failure by monitoring condition, such as vibration, oil analysis, or thermography | A warning period (the P-F interval) exists and is long enough to act on. |
| Scheduled restoration | Overhauls an item at a fixed interval | There is a clear wear-out age and most items survive to it. |
| Scheduled discard | Replaces an item at a fixed interval | The item has a known life and cannot be usefully restored. |
| Failure-finding | Periodically tests a protective or standby function | The failure is hidden until the function is demanded. |
| Redesign | Changes the design, procedure, or operating limits | The failure is unacceptable and no task reduces the risk enough. |
| Run to failure | Accepts the failure and repairs afterward | The consequences are small and a task costs more than the failure. |
For safety and environmental consequences, a task must reduce the risk to an acceptable level, or the item must be redesigned. For purely economic consequences, a task is worthwhile only if it costs less than the failure it prevents.
The P-F Interval
Condition monitoring works because many failures give a warning. The P-F interval is the time between the earliest point at which a developing failure can be detected (P) and the point at which it becomes a functional failure (F). A common rule is to inspect at no more than half the P-F interval, which gives at least two chances to spot the fault and time to plan the repair.
Worked Example: A Cooling-Water Pump Bearing
A plant's cooling-water pump has a bearing failure roughly once every three years. A failure stops the unit for about 18 hours at an assumed cost of $5,000 per hour, plus a $6,000 emergency repair. The numbers are illustrative.
| Item | Value |
|---|---|
| Cost of one unplanned failure | 18 h × $5,000 + $6,000 = $96,000 |
| Failure rate | 1 every 3 years |
| Expected annual cost of failure | $96,000 / 3 = $32,000 per year |
| Observed P-F interval (vibration trend) | About 6 weeks (42 days) |
| Chosen monitoring interval (half of 42 days) | 21 days, about 17 checks per year |
| Cost of monitoring | 17 checks × 0.5 h × $75/h ≈ $640 per year |
| Cost of a planned repair after detection | $6,000 (no lost production, scheduled during a planned window) |
If monitoring catches every developing failure, the annual cost drops to $6,000 / 3 + $640 = about $2,640, a saving of roughly $29,000 per year. That assumes perfect detection. If monitoring catches only 80% of failures, the cost is 0.8 × $2,000 + 0.2 × $32,000 + $640 = about $8,640, a saving of about $23,000. The task pays for itself either way, but the second figure is the one to budget with.
Use the Reliability and Availability Calculator to see what a change in MTBF or repair time does to availability, and record the decision in the RCM Worksheet Template.
Self-Assessment Questions
- Do we know the function and required performance of each critical asset in its operating context?
- Have we identified hidden failures in protective and standby equipment, and tested them?
- Is each preventive task linked to a specific failure mode and consequence?
- Do we have evidence of a P-F interval before choosing a condition-monitoring interval?
- Do we review the program with failure data, or was it written once and filed?
Common Mistakes
Analyzing Every Asset in Detail
A full RCM on every item is impractical. Rank assets by criticality first and analyze the ones that matter.
Assuming Age Causes Failure
Fixed overhauls on random-failure items add cost and infant-mortality risk. Check the failure pattern before choosing an interval.
Ignoring Hidden Failures
Relief valves, trips, and standby equipment rarely announce that they are broken. Give them failure-finding tasks.
Treating the Analysis as Finished
Equipment, operating context, and failure history change. Feed real failure data back into the analysis.
Reliability-Centered Maintenance (RCM): Frequently Asked Questions
What is Reliability-Centered Maintenance?
Reliability-Centered Maintenance is a structured process, defined by standards such as SAE JA1011, for deciding what maintenance each asset needs in its operating context. It identifies functions, functional failures, failure modes, effects, and consequences, then selects tasks that are technically feasible and worth doing, or decides on redesign or run to failure.
What is the P-F interval?
The P-F interval is the time between the earliest point at which a developing failure can be detected (potential failure, P) and the point at which it becomes a functional failure (F). A condition-monitoring task should run more often than the P-F interval, and a common starting point is half of it.
Is RCM the same as preventive maintenance?
No. Preventive maintenance is a category of tasks. RCM is the decision process that determines which tasks, if any, are justified for each failure mode, including condition monitoring, failure-finding, redesign, and deliberate run to failure, based on consequences.
Sources and Further Reading
- F. Stanley Nowlan and Howard F. Heap, Reliability-Centered Maintenance, United Airlines and U.S. Department of Defense, 1978.
- John Moubray, Reliability-centred Maintenance (RCM II).
- SAE International, JA1011: Evaluation Criteria for Reliability-Centered Maintenance (RCM) Processes.
- Campbell and Reyes-Picknell, Uptime: Strategies for Excellence in Maintenance Management.