Reliability & Risk-Based Maintenance

FMEA for Maintenance Engineers

Failure Mode and Effects Analysis (FMEA) helps maintenance teams think systematically about how equipment can fail, what the consequences could be, what controls already exist and which actions deserve priority before the next breakdown occurs.

Published 18 August 2026 · Reviewed and updated 19 August 2026

What maintenance FMEA is trying to answer

A useful FMEA asks five practical questions: What function must the asset perform? How can that function fail? What effect would the failure have? What could cause it? What controls currently prevent or detect it? The value comes from the discussion and evidence behind those answers—not from the score alone.

Start with the equipment function

Define the required function before listing failure modes. For an AHU supply fan, the function might be “deliver the required airflow at the specified operating condition.” A failure mode is then a loss or degradation of that function, such as no airflow, insufficient airflow, unstable airflow or excessive vibration.

Build the FMEA row by row

  1. Function: What must the component or system do?
  2. Failure mode: In what way can the function be lost or degraded?
  3. Failure effect: What happens locally, to the system and to operations or safety?
  4. Potential cause: What mechanism or condition can create the failure?
  5. Existing controls: What currently prevents the cause or detects the failure?
  6. Action: What change would reduce occurrence, improve detection or reduce consequence?

Example: motor-driven fan

Example FMEA logic:
Function: deliver required airflow.
Failure mode: fan stops during operation.
Effect: loss of ventilation/cooling and possible production impact.
Potential causes: motor bearing seizure, VFD trip, loose power connection, mechanical overload.
Existing controls: motor protection, VFD alarms, periodic inspection, current monitoring.
Possible actions: improve bearing-condition checks, verify protection settings against approved design/OEM data, trend repeat VFD alarms, inspect terminations during planned isolation, and define escalation criteria for abnormal current or vibration.

Severity, occurrence and detection

Many FMEA systems use numerical rankings for severity, occurrence and detection. Teams may multiply these values to create a Risk Priority Number (RPN), but the scoring scale must be defined consistently within the organization. A high-severity failure should not be ignored simply because an overall RPN appears moderate.

Do not treat RPN as a universal risk-acceptance rule. Safety, environmental, legal, quality and business-critical consequences may require action regardless of the numerical ranking. Site risk procedures and applicable standards take priority.

Use evidence instead of guessing occurrence

Use breakdown history, work orders, condition-monitoring trends, repeat-failure records, alarm history, inspection findings and OEM information where available. If evidence is weak, mark the uncertainty rather than presenting an invented frequency as fact.

Improve detection controls carefully

Detection does not always mean adding more PPM tasks. It may mean improving the quality of an existing inspection, adding a measurable acceptance limit, trending a condition parameter, using an alarm already available in the control system, or changing the inspection frequency based on evidence.

Convert FMEA findings into maintenance actions

Good actions are specific and verifiable. “Check motor regularly” is weak. “Record DE/NDE bearing vibration monthly, define alert/action limits using approved engineering criteria, review the trend after every alarm, and update the PPM checklist” is much stronger because the control can be audited.

FMEA and RCA are different

FMEA is primarily proactive: it asks what could fail and how to control it. RCA is primarily reactive: it investigates why an actual event occurred. They work well together: repeated RCA findings can update an FMEA, while a good FMEA can identify where stronger preventive or detective controls are needed.

Practical maintenance FMEA checklist

  • Use a clearly defined asset/system boundary.
  • State the required function before the failure mode.
  • Separate failure effects from failure causes.
  • Use historical evidence for occurrence where possible.
  • Record current prevention and detection controls accurately.
  • Define scoring criteria before comparing items.
  • Do not use RPN alone to accept high-consequence risk.
  • Assign every action to an owner with a target date.
  • Re-score only after the action is actually implemented and verified.
  • Feed significant actions into PPM, inspection, spares, training or design-improvement processes.
Safety note. FMEA is an engineering prioritization tool, not a substitute for formal risk assessment, PTW/LOTO, statutory inspection, OEM instructions or competent-person review. Hazardous energy, pressure systems, rotating machinery, chemicals, lifting and other high-risk work must follow approved site procedures.

Related engineering guides

← Back to Engineering Guides