Skip to content
LKaizeN

What is RCM (Reliability Centered Maintenance)

RCM is not just another type of maintenance: it is a method for deciding which maintenance strategy fits each failure mode, by answering seven questions.

Reading time
8 minutes
Sources
1 book, 1 guide, 1 thesis
Tool
Reading only

In one line

RCM (Reliability Centered Maintenance) does not compete with preventive, predictive or corrective maintenance: it is the method that decides which of those three fits each specific failure mode of a piece of equipment.

What it is

RCM was born in the aviation industry in the 1960s and was later formalized for general industrial use (the most cited version is RCM-II, by John Moubray). The core idea: instead of defining a generic maintenance plan "for equipment X", you analyze the equipment function by function, identify its possible failure modes, and answer seven questions for each failure mode:

  1. What are the functions of the asset?
  2. In what ways can it fail to fulfill those functions (functional failures)?
  3. What causes each functional failure (failure modes)?
  4. What happens when each failure occurs (effects)?
  5. How much does each failure matter (consequences: safety, environment, production, cost)?
  6. What can be done to predict or prevent each failure (proactive task)?
  7. What is done if there is no reasonable proactive task (default action, including redesign)?

What it is for

García Garrido describes it as a way to build the maintenance plan from the real consequence of each failure, not from habit or the manufacturer's generic manual. This changes the result: two identical pieces of equipment in two different plants can end up with different maintenance plans if the consequence of their failing is different (one feeds a single line, the other has backup equipment).

How it is applied

In practice, an RCM process follows this flow:

  1. Select the system or equipment to analyze (you do not do the whole plant at once)
  2. Define its functions (including the secondary ones: containment, safety, control)
  3. List the functional failures — every way in which the system stops fulfilling a function
  4. For each functional failure, identify specific failure modes (not "the motor fails", but "bearing X degrades due to lack of lubrication")
  5. Evaluate the consequence of each failure mode and classify it (hidden/evident; safety/environment/operational/non-operational)
  6. Depending on the consequence, choose the task: predictive if there is a condition that can be measured in advance, time-based preventive if the failure mode has a known age, redesign if no task is technically feasible or cost-effective, or run to failure if the consequence is low and there is no cost-effective task

Real example

A thesis from the Universidad Politécnica Salesiana (Cuenca, Ecuador) applied this logic to the vehicle fleet of a public utility company (EMMAIPC-EP), which started from 49 automotive components already identified as generators of repeated failures (clutch, engine, radiator, brakes, tires, and so on).

The study followed two steps, in the order they appear in this article:

  1. Prioritization with Pareto: for each vehicle in the fleet, its historical failures were ordered and the 80/20 rule was applied — in the case of the vehicle identified as UMA-1013, the analysis found that 9 components concentrated 80% of the failure consequences, out of the total number of recorded items.
  2. Criticality analysis (equivalent to RCM question 5, consequences): for each prioritized component, a Risk Priority Number (RPN) was calculated by combining failure severity, probability of occurrence, probability of non-detection, safety/environmental consequences and repair cost. Any component with an RPN greater than 100 was defined as critical.

Of the 49 initial components across the whole fleet, only 9 exceeded that threshold and were classified as critical: clutch, leaf springs, engine, radiator, hydraulic hose, wiring harness, tires, brakes and body. The resulting maintenance plan focused exclusively on those 9 — exactly the RCM logic of dedicating the analysis effort to where the consequence of failing is high, instead of spreading it evenly across all 49.

Benefits when it is implemented well

  • The maintenance plan is justified failure mode by failure mode, not because "the manual says so"
  • It prioritizes maintenance resources where the consequence of failing is high, instead of spreading them evenly
  • It reduces unnecessary preventive tasks on components whose failure mode has no relation to time in use

Limitations to keep in mind

  • A complete and rigorous RCM analysis consumes specialists' time — which is why in practice many plants use simplified versions ("light RCM") for non-critical assets, and reserve the full analysis for the assets with the highest failure consequence
  • It depends on having good knowledge of the equipment's real failure modes — if the analysis is done without data or plant experience, the result is only as good as a generic list
  • It is not a one-time task: failure modes and their consequences change if the process changes, so the analysis needs periodic review

In summary

RCM is a decision method, not a maintenance task in itself. Its contribution is to force you to ask "what happens if this fails?" before choosing how to maintain it, instead of applying the same criterion to all equipment equally.

More on Reliability and TPM