Written by David Rodgers

Quality and Operations Perspective

Written by David Rodgers, Lean Six Sigma Black Belt and ASQ-certified quality leader. This guide applies quality and process-improvement methods to energy and utility operations from a quality and operations perspective. The author is not a licensed professional engineer, process safety specialist, or reliability engineer.

Last editorial review: September 24, 2026. Educational content only: not medical, legal, or regulatory advice. Follow your organization's policies and the requirements that apply to you, and have subject-matter experts review any change to a live process.

  • Lean Six Sigma Black Belt
  • ASQ CQE
  • ASQ CMQ/OE
  • Quality systems and process improvement

Reliability-Centered Maintenance is a disciplined way to decide what maintenance an asset really needs. Instead of applying the same preventive schedule to every machine, RCM asks what each asset must do, how it can fail, what each failure costs in safety, environment, and production, and which task, if any, is worth doing.

This guide walks through the seven questions, why failure does not always follow age, how to choose between condition monitoring, scheduled tasks, failure-finding, redesign, and run to failure, and how to size an inspection interval from the P-F interval. A worked example on a cooling-water pump puts costs on the decision.

Open the Reliability Calculator Get the RCM Worksheet

Before You Start

Educational content. This guide applies quality and reliability methods to energy operations. It is not engineering, legal, safety-compliance, or regulatory advice, and it does not replace your site's procedures, applicable regulations, or the judgment of qualified engineers and safety professionals.

Why RCM Matters in Energy Operations

Maintenance Effort Goes Where Failure Hurts

RCM starts from what each asset must do and what its failure costs in safety, environment, and production, so effort follows consequences rather than habit.

Stops Pointless Time-Based Work

Many failures do not correlate with age. Replacing parts on a calendar can add infant-mortality failures instead of preventing them.

Handles Hidden Failures

Protective devices such as relief valves, trips, and standby pumps can fail silently. RCM forces a plan to find those failures before the demand arrives.

Creates a Traceable Basis

Every task links to a failure mode and a consequence, which makes the maintenance program defensible and easier to update.

What Reliability-Centered Maintenance Is

Reliability-Centered Maintenance (RCM) is a structured process for deciding what maintenance each asset needs to keep doing what its users want it to do. It grew out of work in commercial aviation, documented by F. Stanley Nowlan and Howard Heap in 1978, and was adopted widely in power generation, oil and gas, and other asset-intensive industries. The SAE standard JA1011 sets the criteria that a process must meet to be called RCM.

RCM answers seven questions about every asset in its operating context:

  1. What are the functions and performance standards of the asset?
  2. In what ways can it fail to fulfil those functions (functional failures)?
  3. What causes each functional failure (failure modes)?
  4. What happens when each failure mode occurs (failure effects)?
  5. How does each failure matter (failure consequences)?
  6. What can be done to predict or prevent each failure?
  7. What should be done if no suitable proactive task can be found?

Failure Does Not Always Follow Age

Nowlan and Heap studied how the probability of failure changes with age in aircraft components and found six patterns. Only some show wear-out, so an age-based replacement helps only those.

PatternShapeShare of components studiedImplication
ABathtub: early failures, constant, then wear-out4%Age limit useful only after the wear-out zone starts.
BConstant, then a wear-out zone2%Age-based replacement can work.
CSlowly increasing failure probability5%No clear age limit; condition monitoring helps.
DLow at first, then constant7%Age limits do not help.
EConstant (random) at all ages14%Age limits do not help.
FHigh infant mortality, then constant or slowly rising68%Intrusive maintenance can create early failures.

These percentages come from a study of commercial aircraft components. Industrial and energy equipment can differ, but the lesson holds: do not assume that every asset gets less reliable as it gets older.

Choosing a Maintenance Task

Task typeWhat it doesUse when
On-condition (predictive)Detects a developing failure by monitoring condition, such as vibration, oil analysis, or thermographyA warning period (the P-F interval) exists and is long enough to act on.
Scheduled restorationOverhauls an item at a fixed intervalThere is a clear wear-out age and most items survive to it.
Scheduled discardReplaces an item at a fixed intervalThe item has a known life and cannot be usefully restored.
Failure-findingPeriodically tests a protective or standby functionThe failure is hidden until the function is demanded.
RedesignChanges the design, procedure, or operating limitsThe failure is unacceptable and no task reduces the risk enough.
Run to failureAccepts the failure and repairs afterwardThe consequences are small and a task costs more than the failure.

For safety and environmental consequences, a task must reduce the risk to an acceptable level, or the item must be redesigned. For purely economic consequences, a task is worthwhile only if it costs less than the failure it prevents.

The P-F Interval

Condition monitoring works because many failures give a warning. The P-F interval is the time between the earliest point at which a developing failure can be detected (P) and the point at which it becomes a functional failure (F). A common rule is to inspect at no more than half the P-F interval, which gives at least two chances to spot the fault and time to plan the repair.

Equipment condition Time P: failure detectable F: functional failure P-F interval check at least this often (half)
The task interval must be shorter than the P-F interval. Half the interval is a common starting point, then adjusted for how consistent the warning period is.

Worked Example: A Cooling-Water Pump Bearing

A plant's cooling-water pump has a bearing failure roughly once every three years. A failure stops the unit for about 18 hours at an assumed cost of $5,000 per hour, plus a $6,000 emergency repair. The numbers are illustrative.

ItemValue
Cost of one unplanned failure18 h × $5,000 + $6,000 = $96,000
Failure rate1 every 3 years
Expected annual cost of failure$96,000 / 3 = $32,000 per year
Observed P-F interval (vibration trend)About 6 weeks (42 days)
Chosen monitoring interval (half of 42 days)21 days, about 17 checks per year
Cost of monitoring17 checks × 0.5 h × $75/h ≈ $640 per year
Cost of a planned repair after detection$6,000 (no lost production, scheduled during a planned window)

If monitoring catches every developing failure, the annual cost drops to $6,000 / 3 + $640 = about $2,640, a saving of roughly $29,000 per year. That assumes perfect detection. If monitoring catches only 80% of failures, the cost is 0.8 × $2,000 + 0.2 × $32,000 + $640 = about $8,640, a saving of about $23,000. The task pays for itself either way, but the second figure is the one to budget with.

Use the Reliability and Availability Calculator to see what a change in MTBF or repair time does to availability, and record the decision in the RCM Worksheet Template.

Self-Assessment Questions

  • Do we know the function and required performance of each critical asset in its operating context?
  • Have we identified hidden failures in protective and standby equipment, and tested them?
  • Is each preventive task linked to a specific failure mode and consequence?
  • Do we have evidence of a P-F interval before choosing a condition-monitoring interval?
  • Do we review the program with failure data, or was it written once and filed?

Common Mistakes

Analyzing Every Asset in Detail

A full RCM on every item is impractical. Rank assets by criticality first and analyze the ones that matter.

Assuming Age Causes Failure

Fixed overhauls on random-failure items add cost and infant-mortality risk. Check the failure pattern before choosing an interval.

Ignoring Hidden Failures

Relief valves, trips, and standby equipment rarely announce that they are broken. Give them failure-finding tasks.

Treating the Analysis as Finished

Equipment, operating context, and failure history change. Feed real failure data back into the analysis.

Reliability-Centered Maintenance (RCM): Frequently Asked Questions

What is Reliability-Centered Maintenance?

Reliability-Centered Maintenance is a structured process, defined by standards such as SAE JA1011, for deciding what maintenance each asset needs in its operating context. It identifies functions, functional failures, failure modes, effects, and consequences, then selects tasks that are technically feasible and worth doing, or decides on redesign or run to failure.

What is the P-F interval?

The P-F interval is the time between the earliest point at which a developing failure can be detected (potential failure, P) and the point at which it becomes a functional failure (F). A condition-monitoring task should run more often than the P-F interval, and a common starting point is half of it.

Is RCM the same as preventive maintenance?

No. Preventive maintenance is a category of tasks. RCM is the decision process that determines which tasks, if any, are justified for each failure mode, including condition monitoring, failure-finding, redesign, and deliberate run to failure, based on consequences.

Sources and Further Reading

  • F. Stanley Nowlan and Howard F. Heap, Reliability-Centered Maintenance, United Airlines and U.S. Department of Defense, 1978.
  • John Moubray, Reliability-centred Maintenance (RCM II).
  • SAE International, JA1011: Evaluation Criteria for Reliability-Centered Maintenance (RCM) Processes.
  • Campbell and Reyes-Picknell, Uptime: Strategies for Excellence in Maintenance Management.