FTA in Maintenance: Using Fault Tree Analysis to Prevent Equipment Failures

Calendar
Duration:
18 min
calendar today
Published on
September 11, 2026
Featured Image

Fault tree analysis in maintenance (FTA) is a top-down, deductive method for mapping every possible path to a specific equipment failure — so you can eliminate those paths before the failure occurs. Starting from a clearly defined undesired event (the "top event"), such as "conveyor stops unexpectedly" or "pump fails to deliver pressure," an FTA works backward through logical relationships to every contributing cause: worn components, failed sensors, skipped inspections, or operator errors.

Originally developed in 1962 by Bell Labs for the US Air Force's Minuteman missile program, FTA is now codified under IEC 61025 and used across nuclear, aerospace, chemical, and heavy industrial maintenance. This guide covers what fault tree analysis in maintenance looks like in practice, how it compares to FMEA, how to build one step by step, and how to connect findings to your Computerized Maintenance Management System for measurable failure prevention.

Key Takeaways

  • FTA is deductive: It starts with the failure you want to prevent and works backward through logical gates to every contributing root cause.
  • Use the TRACE Framework: Target, Root, Analyze, Control, and Embed — a 5-step method for implementing FTA in any maintenance program.
  • FTA and FMEA complement each other: FTA reveals critical failure paths for complex events; FMEA systematically covers every component's failure modes.
  • A CMMS turns FTA into action: Work orders, PM schedules, and downtime tracking keep the fault tree accurate and corrective actions moving.

What Is Fault Tree Analysis (FTA)?

Fault tree analysis top-down concept from top event to basic causes | Cryotos

Fault tree analysis (FTA) is a structured, diagram-based reliability tool that identifies all possible causes of a specific equipment failure by working backward from the failure event through a logical tree of contributing factors. Each branch of the tree represents a cause; each node uses a logic gate (AND or OR) to show how causes combine to produce the event above them. Applying fault tree analysis in maintenance gives teams a structured, defensible reason for every PM task on their schedule.

Unlike reactive troubleshooting — where you diagnose after a breakdown — FTA is proactive. You run it before failures occur, using historical failure data, engineering schematics, and operator knowledge to map the full risk profile for a critical asset.

Where FTA Came From

Bell Telephone Laboratories engineer H.A. Watson developed FTA in 1962 to analyze the reliability of the Minuteman intercontinental ballistic missile's launch control system. Boeing refined it further, and by the 1970s, the nuclear industry had adopted FTA as a mandatory safety analysis tool following events like the Three Mile Island accident. Today, it sits alongside root cause analysis, RCM, and FMEA as a core reliability engineering method, recognised by ASQ as a fundamental quality and safety tool.

What FTA Produces: The Fault Tree Diagram

The output of an FTA is a fault tree diagram — a visual, hierarchical map with the top event at the apex and basic events (individual component failures or human errors) at the leaves. Between them sit intermediate events connected by AND and OR gates.

  • Top event: The specific failure you are analyzing (e.g., "Hydraulic press fails to cycle").
  • Intermediate events: Sub-failures that combine through gates to produce the event above them.
  • Basic events: Root-level failures that cannot be broken down further (e.g., "Pressure relief valve stuck open").
  • Minimal cut sets: The smallest combinations of basic events that, together, will cause the top event — these become your maintenance priorities.

A quantitative FTA also calculates probability. If a valve has a failure rate of 0.002 failures/hour and a motor has 0.001 failures/hour, and they are connected by an AND gate, the combined probability is the product of both rates — giving you data to justify inspection frequency changes.

Key Symbols and Gates in a Fault Tree

Fault tree symbols and logic gates AND OR basic event legend | Cryotos

Reading a fault tree requires understanding its symbols. Every diagram uses a standard set drawn from IEC 61025, and knowing what each one means is the difference between a working analysis and a wall decoration.

Gate Types

  • AND gate: All inputs must occur simultaneously for the output event to occur. Example: Both the temperature sensor AND the cooling fan must fail for the motor to overheat. AND gates lower combined probability — two simultaneous rare failures are rarer still.
  • OR gate: Any single input is sufficient to cause the output event. Example: Either a power surge OR a faulty capacitor can stop the motor. OR gates raise combined probability — more paths means more risk.
  • Inhibit gate: A special AND gate where the output occurs only when the input occurs AND a conditional event (shown as an oval) is also true. Example: "Vibration causes bearing failure" only if lubrication is insufficient.
  • Priority AND gate: All inputs must occur, but in a specific sequence. Used for time-dependent failure chains.

Event Types

  • Basic event (circle): A root cause that requires no further development — a specific component failure, human error, or environmental condition.
  • Undeveloped event (diamond): A cause that could be expanded further but is left as-is due to insufficient data or irrelevance to the current analysis scope.
  • Intermediate event (rectangle): A failure that is both an output of the gates below it and an input to the gate above it.
  • Conditioning event (oval): A constraint on an Inhibit or Priority AND gate — represents a condition that must be true for the gate logic to apply.

Maintenance teams working without dedicated fault tree software often build trees in Visio, Lucidchart, or Excel. Purpose-built tools like Isograph FaultTree+ or Relyence can calculate cut sets and failure probabilities automatically, saving hours of manual arithmetic.

FTA vs FMEA: Key Differences

FTA and FMEA are both failure analysis tools, but they work in opposite directions and answer different questions. FTA asks "what causes this specific failure?" FMEA asks "what failures can this component cause?" Understanding the difference helps you choose the right tool — or use both together.

AttributeFault Tree Analysis (FTA)FMEA (Failure Mode & Effects Analysis)
Analysis directionTop-down (deductive)Bottom-up (inductive)
Starting pointA specific undesired failure eventEach component's individual failure modes
Best forComplex system failures, safety-critical eventsComponent-level coverage, design review
OutputFault tree diagram, minimal cut setsFMEA table with RPN scores
Quantitative?Yes — failure probability per pathYes — RPN (Severity × Occurrence × Detection)
Industry standardIEC 61025IEC 60812 / AIAG FMEA-4
CMMS integrationWork orders for critical cut sets, PM triggersRisk register, PM checklist by failure mode

When to Use Each Method

Use FTA when you have a specific high-consequence failure event you need to prevent — a press that could injure an operator, a chiller that keeps tripping, a packaging line that loses batches. FTA is also the right tool when multiple failures interact: you need to see which combinations of two or three simultaneous faults are actually dangerous.

Use FMEA when you are systematically reviewing a new piece of equipment, a new process, or a revised design. FMEA ensures you have not missed any failure mode across all components, not just the ones you already suspect. The best reliability programs run FMEA first to catalogue failure modes, then use FTA on the highest-risk events that FMEA surfaces. The 5 Whys technique complements both — it drills into individual basic events once the tree has identified the critical path.

Step-by-Step: How to Build a Fault Tree Analysis in Maintenance

Seven step process to build a fault tree analysis in maintenance | Cryotos

A fault tree analysis for a production asset typically takes one to three working days for a cross-functional team. The quality of the output depends almost entirely on how precisely you define the top event and how well you gather system knowledge before you start drawing.

Step 1: Define the Top Event

The top event must be specific, observable, and failure-focused. "Motor fails" is too vague. "Drive motor on Line 3 conveyor stops during production shift" is right. A precise definition keeps the tree from sprawling into an unmanageable diagram. Involve operators — they often know the exact failure mode that causes production loss.

Step 2: Gather System Documentation

Collect P&IDs, electrical schematics, maintenance records, OEM manuals, and any previous incident reports. Historical failure data from your root cause analysis investigation checklist is particularly valuable: it tells you which basic events have actually occurred before, giving you real frequency data rather than estimates.

Step 3: Identify Immediate Causes (Level 1)

Ask: "What events, in isolation or combination, directly cause the top event?" List these as the first level of the tree. Apply AND or OR gates based on the physics: does any single cause suffice (OR), or must multiple failures coincide (AND)?

Step 4: Expand Each Cause to Basic Events

For each intermediate event, repeat the process. Ask what causes it. Keep expanding until you reach events that need no further breakdown — basic events like "bearing fails," "relay contacts weld shut," or "operator bypasses interlock." Each basic event should be something you can directly measure, inspect, or prevent.

Step 5: Calculate Failure Probability

For quantitative FTA, assign a failure rate to each basic event using historical maintenance data, OEM specifications, or industry databases (MIL-HDBK-217, OREDA). Propagate probabilities up through the tree: OR gates add probabilities (approximately), AND gates multiply them. The result is a calculated top event probability per operating hour or per year.

Step 6: Identify Minimal Cut Sets

A minimal cut set is the smallest group of basic events that, if they all fail simultaneously, will cause the top event. Single-event cut sets (where one failure alone triggers the top event) are your highest-priority maintenance actions. Two- and three-event cut sets represent secondary priorities. Most fault tree software identifies cut sets automatically; manual methods use Boolean algebra.

Step 7: Develop and Document Corrective Actions

For every minimal cut set, assign a corrective action: adjust inspection frequency, add a redundant component, install condition monitoring, or update an operator procedure. Document who owns each action, with a due date. The FTA is not complete until every single-event cut set has a countermeasure assigned.

Calculate your equipment failure rate to get the data you need for a quantitative FTA — knowing your current failure rate per hour or per cycle makes Step 5 significantly faster and more accurate.

The TRACE Framework: Implementing FTA in Your Maintenance Program

TRACE framework Target Root Analyze Control Embed for FTA | Cryotos

Running a one-time FTA is useful. Building a system where fault tree analysis in maintenance continuously improves your PM program is transformative. The TRACE Framework is a five-stage method for doing exactly that, designed for maintenance teams who do not have dedicated reliability engineers on staff.

T — Target the Right Failure Events

Target: Identify which assets and failure modes deserve a fault tree analysis. Not every piece of equipment warrants a full FTA. Focus on assets where failure causes: significant production downtime, safety hazards, regulatory non-compliance, or high repair costs. A criticality ranking exercise — scoring each asset on consequence × frequency — produces a shortlist. Start with the top three to five assets on that list.

R — Root Out Every Failure Path

Root: Build the fault tree with the full cross-functional team. Involve the maintenance technicians who service the equipment, the operators who run it daily, and the engineers who designed the process. Each group sees failure through a different lens. Technicians know what physically fails first; operators know what behaviors precede failures; engineers know the system interdependencies. A tree built by one person in isolation misses half the branches.

A — Analyze Probability and Priority

Analyze: Calculate failure probability and identify your minimal cut sets. Even a rough qualitative ranking (High / Medium / Low likelihood for each basic event) gives you enough information to prioritize corrective actions. If you have historical maintenance records in a CMMS, pull the actual failure frequency for each component — this turns a qualitative analysis into a quantitative one with far more actionable output. Flag all single-event cut sets as immediate action items.

C — Control Through Targeted Actions

Control: Assign a specific corrective action to every critical cut set. Each action gets an owner, a deadline, and a budget code. Actions fall into four types:

  • Eliminate: Remove the failure path entirely — add redundancy, change the design, or replace a chronically failing component.
  • Detect earlier: Install condition monitoring (vibration, temperature, oil analysis) to catch the basic event before it causes the intermediate event above it.
  • Reduce frequency: Increase inspection or replacement intervals for components with high individual failure rates.
  • Protect: Add interlocks, alarms, or procedural barriers so that even if the basic event occurs, it cannot propagate to the top event.

E — Embed in Your CMMS and PM Program

Embed: Integrate every corrective action into your maintenance management system so it actually gets done. Create work orders for one-time fixes. Add recurring tasks to preventive maintenance software schedules, driven by the inspection frequencies the FTA recommends. Store the fault tree diagram in your asset's document management record. Set a review trigger — every six months, or after any actual failure — so the tree stays current as the system evolves.

Using Fault Tree Analysis in Maintenance to Strengthen Your PM Program

Most preventive maintenance programs are built on OEM recommendations and collective experience. Fault tree analysis in maintenance gives you a third, more precise data source: a map of which failure paths are actually critical, so you focus PM resources where they reduce the most risk.

FTA-Driven PM Frequency Adjustments

A standard OEM recommendation might say "inspect bearing every 1,000 operating hours." But if your FTA shows that a bearing failure is a single-event cut set — meaning bearing failure alone causes the top event — you have a strong quantitative case for increasing that to every 500 hours, or for adding continuous vibration monitoring between manual inspections.

Conversely, components that appear only in multi-event cut sets (requiring two or three simultaneous failures) may not need more frequent PM. The FTA tells you it is acceptable to inspect those at the OEM-recommended interval and redirect the saved labour hours to higher-risk components. This is the kind of risk-based maintenance optimization that the Society for Maintenance & Reliability Professionals (SMRP) advocates as a core competency for modern maintenance teams.

Critical Component Spare Parts Planning

FTA also informs your spare parts strategy. Components that appear in single-event cut sets should be stocked on-site — their failure alone stops production, so waiting for a supplier is not acceptable. Components in multi-event cut sets can often be managed with longer lead times. This is a direct, data-driven input into your MRO inventory planning, reducing both stockout risk and excess inventory carrying cost.

When FTA findings drive PM schedules and spare parts decisions, the result is a maintenance program built on evidence rather than habit. Teams that apply this approach consistently report measurable reductions in unplanned downtime — the same direction as Cryotos customers who use data-driven PM scheduling to achieve up to 30% reduction in equipment downtime.

How a CMMS Helps You Act on Fault Tree Findings

A fault tree diagram on paper is a starting point. A CMMS turns fault tree analysis in maintenance findings into scheduled tasks, recorded inspections, and a continuous feedback loop that keeps the analysis accurate over time.

  • Work order generation: Every corrective action from the FTA becomes a work order in the CMMS, assigned to a technician with a due date and a link to the relevant fault tree documentation. Work order management features track completion, technician notes, and actual time spent — giving you audit data that satisfies regulatory and insurance requirements.
  • PM schedule integration: FTA-recommended inspection frequencies feed directly into the PM calendar. Dynamic scheduling adjusts intervals based on actual usage hours or meter readings, so a conveyor that ran double shifts this month gets its bearing checked sooner than one that idled.
  • Downtime tracking as feedback: Every breakdown that does occur gets logged against the asset in the CMMS. Over time, this data tells you whether your FTA-driven corrective actions are working: if the same basic event keeps triggering, the countermeasure needs revisiting. Real-time downtime tracking by asset lets you compare actual failure frequency against your FTA probability estimates.
  • Document management: The fault tree diagram, cut set analysis, and corrective action register all live inside the asset's CMMS record, accessible on a mobile device at the machine. When a new technician joins the team, they have the full failure analysis history — not a folder in someone's desk drawer.

Cryotos CMMS also captures 5 Whys analysis directly within work orders, so individual breakdown investigations contribute structured data back to your fault trees — creating a continuous improvement loop between reactive investigations and proactive FTA planning.

Common Mistakes Maintenance Teams Make in Fault Tree Analysis

Fault tree analysis in maintenance produces poor results when teams rush the setup or skip the integration steps. These are the four mistakes that consistently undermine the analysis.

  • Defining the top event too broadly: "Equipment failure" is not a top event. "Packaging machine stops mid-cycle during a production run" is. Broad top events produce enormous trees that are impossible to act on. The more precisely you define the failure, the more focused and useful the tree becomes.
  • Stopping at the symptom, not the root cause: Teams often stop expanding a branch when they reach a familiar component — "seal fails" — without asking why the seal fails. Was it the wrong material? Incorrect torque on installation? Contaminated lubricant? Each of these has a different countermeasure. Basic events must represent true root-level causes.
  • Treating FTA as a one-time exercise: A fault tree built in 2022 for a machine that has since had control system upgrades, new operators, and a changed shift pattern may no longer reflect reality. FTA should be reviewed after every significant equipment change and after every actual occurrence of the top event.
  • Not connecting findings to the CMMS: The most common FTA failure mode is a complete and accurate fault tree that never translates into maintenance actions. Without work orders and PM schedule entries in the CMMS, findings stay in a presentation file and the failures they predict keep happening.

Frequently Asked Questions

What is the difference between a qualitative and a quantitative fault tree analysis?

A qualitative FTA identifies the logical failure paths and minimal cut sets without assigning probabilities — it tells you which combinations of failures can cause the top event. A quantitative FTA assigns failure rates (from historical data or reference databases) to each basic event and calculates a numerical probability for the top event and each cut set. Qualitative FTA is faster and works well when failure rate data is limited; quantitative FTA gives you the numbers to justify maintenance budget decisions and compare design alternatives.

How long does it take to complete a fault tree analysis for a production asset?

A focused fault tree analysis in maintenance on a single asset with a well-defined top event typically takes one to three days for a cross-functional team of three to five people. Complex systems with many sub-systems and interdependencies can take one to two weeks. The largest time investment is usually gathering accurate system documentation and historical failure data — having a well-maintained CMMS with complete maintenance histories cuts this significantly.

Can I use fault tree analysis for safety risk assessment, not just equipment failures?

Yes — and it is actually where FTA originated. Safety-focused FTA uses the same methodology but defines safety incidents (an injury, a fire, a toxic release) as the top event. For maintenance teams working under ISO 45001 or PSM regulations, FTA is a recognised method for demonstrating that safety-critical systems have been systematically analysed. Many maintenance departments run parallel trees: one for production failure risk and one for the associated safety event if the same failure occurs during maintenance.

How does fault tree analysis integrate with reliability-centered maintenance (RCM)?

FTA and RCM are complementary, not competing methods. RCM is a framework for selecting the right maintenance strategy for each failure mode (preventive, predictive, run-to-failure, or redesign). FTA is an analytical tool for understanding how specific failure modes combine to produce system-level failures. In an RCM programme, FTA is typically used during the functional failure analysis stage to identify which failure modes are critical and which can be accepted — feeding directly into the RCM decision logic for task selection.

What is a minimal cut set and why does it matter for maintenance planning?

A minimal cut set is the smallest combination of basic events (component failures or human errors) that, if they all occur, will guarantee the top event occurs. Single-element minimal cut sets — where one failure alone causes the top event — represent your most critical maintenance priorities, since there is no redundancy protecting the system. Multi-element minimal cut sets are lower priority because multiple simultaneous failures are required. Ranking your maintenance actions by cut set size helps you direct limited maintenance resources to where they provide the greatest risk reduction.

Preventing equipment failures starts with understanding exactly how they happen. Fault tree analysis gives maintenance teams the structured map to do that — and Schedule a free demo to see how Cryotos CMMS turns your FTA findings into scheduled work orders, PM tasks, and real-time downtime tracking that keeps critical assets running.

Want to Try Cryotos CMMS Today?

Get Free Demo

Let AI Take Control of Your Maintenance

Cryotos AI predicts failures, automates work orders, and simplifies maintenance—before problems slow you down.

Try AI-Powered CMMS
🡢