Fault Tree Diagram: How to Draw One for Equipment Failure Analysis

Calendar
Duration:
22 min
calendar today
Published on
September 11, 2026
Featured Image

A fault tree diagram is a top-down, tree-shaped visual that maps every possible cause of a specific equipment failure using logic gates and event symbols. It starts with the failure you want to prevent — the "top event" — and branches downward through AND gates, OR gates, and failure events until it reaches the root-level causes your maintenance team can actually control.

Drawing a fault tree diagram is the core skill behind fault tree analysis (FTA). Once you can read the symbols, apply the gate logic, and identify minimal cut sets from the completed tree, you have a precise map of every failure path threatening a critical asset — one that feeds directly into work orders and PM schedules in your Computerized Maintenance Management System. This guide covers every symbol and gate, a step-by-step drawing process, three real equipment failure examples, and how a completed fault tree diagram translates into your maintenance workflow.

Key Takeaways

  • A fault tree diagram uses two types of building blocks: events (the circles, rectangles, and diamonds that represent failures) and gates (the AND and OR symbols that define how failures combine).
  • The DEFINE Method gives structure to the drawing process: Define, Establish, Frame, Iterate, Name cut sets, and Embed in your workflow.
  • Minimal cut sets are the diagram's most actionable output: single-event cut sets are your highest-priority maintenance actions, because one failure alone causes the top event.
  • A CMMS turns the completed diagram into action: each minimal cut set becomes a work order or PM task, not a line in a presentation file.

What Is a Fault Tree Diagram?

Top-down fault tree diagram structure from top event through logic gates to basic events | Cryotos

A fault tree diagram is a graphical model that represents all the logical relationships between events and conditions that can cause a specific system failure. The diagram takes the shape of an inverted tree: the undesired top event sits at the root (apex), and the contributing causes branch downward through a hierarchy of gates and events until every path terminates at a basic, root-level cause.

Standardised under IEC 61025, the fault tree diagram is one of the most widely used tools in root cause analysis, reliability engineering, and safety analysis. For maintenance teams, it translates system-level risk into a visual that every technician, engineer, and plant manager can read — and act on.

The Anatomy of a Fault Tree Diagram

Every fault tree diagram, regardless of system complexity, is built from the same structural elements arranged in a consistent hierarchy.

  • Top event (apex): The specific, well-defined failure the tree is built around — e.g., "Centrifugal pump fails to deliver flow." It sits alone at the top of the diagram.
  • Intermediate events: Sub-failures that are both the output of the gates below them and the input to the gate above them. They form the middle layers of the tree and are shown in rectangles.
  • Gate level: The layer of AND/OR gates directly below the top event or any intermediate event. Gates define the logical relationship between causes.
  • Basic events (leaves): The lowest level of the tree — root-level failures that require no further breakdown. Each basic event is something you can directly measure, inspect, replace, or prevent.

What a Completed Fault Tree Diagram Tells You

A completed fault tree diagram tells you three things your maintenance program cannot function without: which combinations of failures cause your critical failure event, how probable those combinations are, and which single failure paths have no redundancy protecting you.

The diagram's most important output is the set of minimal cut sets — the smallest combinations of basic events that, together, guarantee the top event occurs. A single-element minimal cut set means one component failure alone stops your asset. Those single-event paths are your first maintenance priority. Multi-event cut sets, where two or three simultaneous failures are required, are secondary. The diagram makes this priority ranking visible at a glance.

Fault Tree Diagram Symbols and Gates: A Visual Reference

Fault tree diagrams use a standardised set of symbols drawn from IEC 61025. Every symbol has a specific shape, a defined meaning, and a precise place in the diagram's hierarchy. Misusing a symbol — especially confusing AND and OR gate logic — produces a tree that points to the wrong maintenance actions.

Symbol NameShapeWhat It MeansWhen to Use in Equipment FTA
Top EventRectangle at apexThe specific undesired failure being analysedAlways — exactly one per diagram, at the top
AND GateFlat-bottomed D-shape (shield)ALL inputs must occur simultaneously for the output to happenWhen two or more failures must coincide — e.g., both the sensor and backup alarm fail
OR GateCurved D-shape (pointed bottom, curved top)ANY single input is sufficient to cause the output eventWhen any one failure path alone triggers the event above — the most common gate type
Basic EventCircleA root-level failure that needs no further breakdownComponent failures, human errors, environmental inputs — the actionable leaf nodes
Intermediate EventRectangle (body of tree)A sub-failure that results from the gate below and feeds the gate aboveMulti-stage failure chains — sits between the top event and basic events
Undeveloped EventDiamondA cause not expanded due to insufficient data or scopeEvents outside the analysis boundary, or causes with no available failure rate data
Conditioning EventOvalA condition that must be true for the gate logic above it to applyOnly with Inhibit or Priority AND gates — e.g., "only when operating above 80°C"
Inhibit GateHexagon with conditioning ovalInput causes output ONLY IF the conditioning event is also trueContext-dependent failures — e.g., "vibration causes bearing failure only if lubrication is low"
Priority AND GateAND gate shape with "P" insideAll inputs must occur, but in a specific time sequenceTime-ordered failure chains where sequence matters — e.g., overload must precede thermal trip
Transfer Gate (Triangle)Triangle (in/out pair)Links the current branch to another section of the same treeLarge trees spanning multiple pages — keeps each sheet readable

The AND/OR gate distinction is the most critical choice in any fault tree diagram. The underlying logic comes directly from Boolean algebra: AND gates multiply failure probabilities (reducing combined risk), while OR gates add them (increasing combined risk). Getting this wrong misidentifies your most critical failure paths.

Step-by-Step: How to Draw a Fault Tree Diagram

Step-by-step process for drawing a fault tree diagram | Cryotos

Drawing a fault tree diagram for an equipment failure follows a consistent seven-step process. The first two steps — defining the top event and gathering documentation — determine the quality of every branch that follows. Teams that rush them produce sprawling trees that point to the wrong maintenance actions.

Step 1: Choose and Define the Top Event Precisely

The top event must be specific, observable, and unambiguous. "Equipment failure" is not a valid top event. "Centrifugal pump on Line 2 fails to deliver flow during production shift" is. A precise top event bounds the diagram — it tells you exactly which failure paths belong in the tree and which fall outside scope. Use the exact failure description that operators and technicians would recognize from their own experience of the event.

Step 2: Gather System Documentation Before Drawing

Collect P&IDs, electrical schematics, OEM manuals, maintenance records, and any previous incident reports before you draw a single symbol. Your root cause analysis investigation checklist is particularly useful here: it surfaces which component failures have actually occurred before, giving you real-world evidence rather than engineered guesses for your basic events.

Step 3: Draw the Top Event and Its Gate

Place the top event rectangle at the top of your diagram. Directly below it, draw the first gate. Ask the core question: can any single sub-failure alone cause this top event, or must multiple sub-failures coincide? If any single cause suffices, draw an OR gate. If multiple failures must happen simultaneously, draw an AND gate. This first gate is the most consequential decision in the entire diagram.

Step 4: Identify Level 1 Intermediate Events

From the gate, branch to the Level 1 intermediate events — the immediate, direct causes of the top event. Involve the full cross-functional team here: maintenance technicians, process engineers, and operators each contribute failure modes the others miss. List every plausible direct cause before selecting the gate logic that connects them. Intermediate events are always written in rectangles using the format "Event [component] does [failure]" — specific, not vague.

Step 5: Expand Each Intermediate Event Downward

For each intermediate event, repeat Steps 3 and 4: place a gate and identify the events immediately below it. Continue expanding until you reach basic events — root-level failures that require no further breakdown. A basic event is anything you can directly inspect, measure, replace, or train against: "seal fails due to material incompatibility," "operator bypasses pressure interlock," "bearing runs dry due to lube line blockage." When you genuinely cannot expand a cause further due to scope or data constraints, draw a diamond (undeveloped event) and move on.

Step 6: Check Gate Logic at Every Level

After completing the draft, walk back up the tree and verify gate logic at every node. Ask: "If only one of these inputs occurs, does the event above still happen?" If yes, the gate should be OR. If you genuinely need all inputs to occur simultaneously, AND is correct. This review step catches the most common drawing error — using AND gates too liberally, which underestimates failure probability and creates false confidence in the system's resilience.

Step 7: Identify Minimal Cut Sets

A minimal cut set is the smallest group of basic events that, if they all occur, guarantees the top event. For qualitative analysis, trace every path from the top event through gates to the basic events. Any path that reaches basic events exclusively through OR gates creates a single-element cut set — your highest-priority maintenance action, since one failure alone triggers the top event. For quantitative analysis, assign failure rates to each basic event and propagate probabilities through the gate logic to calculate the top event frequency.

Calculate your equipment's failure rate to get the probability data you need for Step 7 — turning your fault tree diagram from a qualitative map into a quantitative risk model with calculable top-event frequency.

The DEFINE Method: A Drawing Framework for Maintenance Teams

The DEFINE method six-stage framework for drawing a fault tree diagram | Cryotos

Most fault tree diagram guides are written for reliability engineers. The DEFINE Method is a six-stage drawing framework built specifically for maintenance teams who need to produce their first fault tree diagram without a reliability engineering background — and have it produce actionable maintenance outputs, not a document that sits in a folder.

D — Define the Top Event with Precision

Define: Write one sentence that specifies the exact failure, the exact asset, and the exact condition under which it matters. Include the asset name, the failure mode, and the operational context. "Pump fails" becomes "Cooling water pump CP-02 fails to maintain minimum flow rate during peak production hours." This sentence is the foundation of every branch in the diagram. If the top event is wrong, the entire tree points to the wrong maintenance actions.

E — Establish the System Boundary

Establish: Draw a boundary around what the analysis covers before touching the diagram. What subsystems are in scope? What failure causes are outside scope (e.g., acts of nature, utility outages)? Who and what can cause the basic events — only the mechanical system, or also operators and contractors? Establishing this boundary prevents the tree from sprawling into an unmanageable diagram. Draw the boundary on a separate sheet if needed, and reference it throughout the drawing session.

F — Frame Level 1 with Correct Gate Logic

Frame: Place the top event rectangle, draw the first gate, and identify all Level 1 intermediate events before expanding further. The first gate decision — AND or OR — shapes everything below it. Make this decision with the engineering team, not alone. Ask: "Has this top event ever occurred when only one of these causes was present?" If yes, it is an OR gate. Resist the temptation to draw AND gates to reduce the apparent probability of the top event — the diagram should reflect reality, not wishful thinking.

I — Iterate Downward to Basic Events

Iterate: Expand each intermediate event one level at a time, involving the team at each stage. Use the test: "Can this event be caused by something more specific that the team can act on?" If yes, keep expanding. If the event is already a root-level cause — a specific component failure, human error, or environmental input — place a circle (basic event) and stop. Work through the tree in breadth-first order: fully expand one level across all branches before going deeper. This keeps the overall structure visible as you build.

N — Name and Validate All Cut Sets

Name: Trace every path from top event to basic events and list the minimal cut sets. For each cut set, record: the basic events it contains, whether it is a single- or multi-event cut set, and the estimated failure probability if data is available. Validate the cut sets with the maintenance team. If a cut set describes a failure scenario that the team says "would never happen in practice," revisit your gate logic — it likely has an AND gate where an OR gate is more accurate. Every single-event cut set must have an assigned corrective action before the analysis is complete.

E — Embed in Your Maintenance Workflow

Embed: Convert every minimal cut set into a maintenance action and connect it to your live workflow system. Single-event cut sets become priority PM tasks or condition monitoring setpoints. Multi-event cut sets become secondary inspection items. Store the completed fault tree diagram in your asset's document record inside the CMMS. Set a review date — the diagram should be revisited after any significant equipment change or after the top event actually occurs. A fault tree diagram that exists only on paper or in a shared drive folder is not a maintenance program — it is a completed exercise.

Fault Tree Diagram Examples for Equipment Failure Analysis

The fastest way to master fault tree diagram drawing is to work through realistic examples. The three examples below cover common industrial failure scenarios using the symbol set and gate logic from this guide. Each example shows how the top event, intermediate events, and basic events connect — and what the minimal cut sets reveal about maintenance priorities.

Example 1: Centrifugal Pump Fails to Deliver Flow

Top event: "Centrifugal pump CP-02 fails to deliver minimum flow rate." Two Level 1 causes connect via OR gate — "Pump mechanical failure" OR "No power to pump motor."

Expanding "Pump mechanical failure" via OR gate: "Impeller damaged" OR "Seal failure causing cavitation" OR "Bearing seizure." Expanding "Bearing seizure" via AND gate: "Lubricant film breakdown" AND "High bearing temperature" — both must occur simultaneously.

Minimal Cut Sets and Maintenance Actions

  • Single-event cut set — Impeller damaged: One failure stops flow. Stock a replacement on-site. Apply the 5 Whys to confirm root cause (material selection? foreign object ingress? cavitation?).
  • Single-event cut set — Seal failure causing cavitation: Add seal condition inspection to every PM interval.
  • Two-event cut set — Lubricant breakdown AND high temperature: Add bearing temperature alarm to PM checklist; secondary priority.

Example 2: Conveyor Belt Stops Unexpectedly

Top event: "Line 3 conveyor belt stops during production run." Level 1 via OR gate: "Drive motor trips" OR "Belt tension failure" OR "Control system fault."

Expanding "Drive motor trips" via OR gate: "Motor overload protection activates" OR "Motor winding failure." Expanding overload activation via AND gate: "Belt load exceeds rated capacity" AND "Overload relay not reset after previous trip."

Minimal Cut Sets and Maintenance Actions

  • Single-event cut set — Motor winding failure: Schedule periodic motor winding resistance testing. Failure alone stops the conveyor.
  • Single-event cut set — Belt tension failure: Add belt tension check to every PM cycle.
  • Two-event cut set — Overload AND relay not reset: Lower probability but flags a human-factor gap — add formal relay reset sign-off to the work order template.

Example 3: HVAC Chiller Unit Trips

Top event: "Chiller unit CHW-01 trips on high discharge pressure." Level 1 via OR gate: "Condenser heat rejection failure" OR "Refrigerant overcharge" OR "Expansion valve fault."

Expanding "Condenser heat rejection failure" via OR gate: "Condenser coils fouled" OR "Condenser fan failure" OR "Ambient temperature exceeds design limit." Fan failure expands via AND gate: "Fan motor fails" AND "Backup fan relay does not engage."

Minimal Cut Sets and Maintenance Actions

  • Single-event cut set — Condenser coils fouled: Schedule quarterly coil cleaning in preventive maintenance software — direct and immediately actionable.
  • Single-event cut set — Refrigerant overcharge: Add refrigerant charge verification to the semi-annual service checklist.
  • Two-event cut set — Primary fan AND backup relay failure: Test backup fan relay monthly and log as a standalone work order — low probability but a critical redundancy gap.

Fault Tree Diagram vs FMEA vs Reliability Block Diagram

Fault tree diagrams, FMEA worksheets, and reliability block diagrams (RBDs) are all failure analysis tools, but they produce fundamentally different visual outputs and answer different questions. Choosing the right tool — or combining all three — depends on what kind of analysis your reliability program needs at a given stage.

AttributeFault Tree Diagram (FTA)FMEA WorksheetReliability Block Diagram (RBD)
Analysis directionTop-down (deductive)Bottom-up (inductive)Component-to-system (inductive)
Primary questionWhat causes this specific failure?What can each component fail to do?What system reliability does this configuration produce?
Visual formatInverted tree with gates and eventsTabular worksheet (rows per failure mode)Block diagram with series/parallel paths
Key outputMinimal cut sets, top-event probabilityRPN scores (Severity × Occurrence × Detection)System reliability figure, critical paths
Best forSpecific failure prevention, safety analysisSystematic failure mode coverage, design reviewSystem design optimisation, redundancy planning
Handles interactions?Yes — AND/OR gate logic shows combinationsNo — treats each failure mode independentlyPartial — series/parallel shows redundancy, not logic
Industry standardIEC 61025IEC 60812 / AIAG FMEA-4IEC 61078

When to Use Each Method

A FMEA worksheet is the right starting point when you are systematically reviewing all failure modes of a new asset or a redesigned process. It ensures completeness — no failure mode goes unexamined. A reliability block diagram (RBD) is the right tool when you need to compare system configurations — series vs. parallel, redundant vs. non-redundant — and calculate the overall system reliability figure for a given configuration. A fault tree diagram is the right tool when FMEA has surfaced a high-risk failure event and you need to understand exactly how it can happen and which combinations of causes are most critical.

In a mature maintenance program, all three tools work together: FMEA catalogues failure modes, RBD quantifies system-level reliability, and fault tree diagrams target the highest-risk events for deep analysis. Each adds something the others cannot.

Common Mistakes When Drawing a Fault Tree Diagram

Fault tree diagrams are only as useful as the accuracy of their gate logic and the specificity of their basic events. These five mistakes consistently appear in first-time fault tree diagrams drawn by maintenance teams — and each one undermines the analysis it produces.

  • Using AND gates to reduce apparent probability: AND gates require all inputs to occur simultaneously, which makes the combined probability smaller. Teams sometimes draw AND gates where OR gates are correct, because AND makes the failure seem less likely. The result is a tree that underestimates risk. If any single input would cause the output on its own, the gate is OR.
  • Mixing intermediate events with basic events: An intermediate event (rectangle) must have at least one gate below it. A basic event (circle) has no gate and no further breakdown. When teams place a rectangle at the bottom of a branch with no gate below it, they signal that a cause is unexpanded — but the reader does not know whether it is truly a basic event or an undeveloped one. Use the diamond for undeveloped events; use the circle only for true root-level causes.
  • Writing vague basic events: "Motor problem," "sensor issue," or "operator error" are not basic events — they are categories. A valid basic event is specific enough to assign a failure rate to: "Motor winding short-circuit due to moisture ingress," "Proximity sensor face contaminated with weld spatter," "Operator bypasses pressure interlock per informal workaround." Vague basic events produce unmeasurable probabilities and unactionable corrective measures.
  • Drawing the tree alone without field input: The most accurate fault trees are built with maintenance technicians who service the equipment, operators who run it, and engineers who designed the process. Engineers alone miss the informal workarounds and known failure patterns that technicians see daily. Technicians alone miss the system interdependencies that engineers track. Any fault tree built by a single person in isolation is missing at least a third of its branches.
  • Not validating cut sets against operating experience: Once the tree is drawn, review the minimal cut sets with the team. If a cut set describes a scenario the team says "would never happen," the gate logic is wrong — revisit it. If a cut set describes a scenario the team says "happens all the time but we don't count it as a failure," you have found an undocumented production problem worth investigating separately.

Using Your Fault Tree Diagram in a CMMS Workflow

A fault tree diagram earns its value when it moves from a completed diagram into your maintenance management system as active work. The output of a fault tree is a prioritised list of failure paths — each one corresponding to a specific maintenance action, inspection frequency, or equipment design change. Work order management tools translate those actions into assigned tasks with due dates, technician ownership, and completion tracking.

The workflow has four stages once the diagram is finalised. First, every single-event minimal cut set generates a corrective work order if an immediate fix is needed, or a recurring PM task if the action is an inspection or replacement at a defined interval. Second, the FTA-recommended inspection frequencies override the OEM-default frequencies for components in single-event cut sets — these are the components where failure alone stops the asset, so OEM defaults are often insufficient. Third, downtime tracking in the CMMS provides a feedback loop: every actual breakdown is logged against the asset, and the failure mode is categorised against the fault tree's basic events. If the same basic event keeps appearing, the corrective action is insufficient and needs to be revisited. Fourth, the completed fault tree diagram and its cut set register are stored as documents inside the asset's CMMS record — accessible on a mobile device at the machine by any technician, not buried in a shared drive folder only reliability engineers know how to navigate.

For organisations working under OSHA's Process Safety Management (PSM) standard, fault tree diagrams stored and maintained in the CMMS also serve as documentation that process hazard analyses have been completed and kept current — a direct audit requirement. Connecting the diagram to actual work orders proves that the analysis translated into maintenance action, not just a completed document.

Frequently Asked Questions

What software do maintenance teams use to draw fault tree diagrams?

Purpose-built fault tree software includes Isograph FaultTree+, Relyence FTA, and Windchill Quality Solutions — these calculate minimal cut sets and failure probabilities automatically. General-purpose diagramming tools like Lucidchart, Visio, and Miro work well for drawing the visual tree when manual probability calculation is acceptable. For smaller teams doing qualitative analysis only, a whiteboard and manual drawing during a cross-functional workshop is a valid starting point — the important part is the team knowledge captured during the session, not the specific software.

How many levels deep should a fault tree diagram go?

A fault tree should expand until every branch terminates at a basic event — a root-level cause that is specific enough to assign a corrective action to. In practice, most maintenance fault trees for a single piece of equipment run three to five levels deep. Very complex systems with multiple subsystems may require six to eight levels. Use transfer gates (triangles) to split large trees across multiple pages when any single page exceeds 15 to 20 event boxes — beyond that, the diagram becomes unreadable and the team loses track of the overall structure.

Can a fault tree diagram have more than one top event?

No — a single fault tree diagram has exactly one top event. If you need to analyse multiple undesired failure events on the same asset, draw separate fault trees, one per top event. This is standard practice in reliability engineering: a pump might have one fault tree for "fails to deliver flow" and a separate tree for "delivers contaminated flow." These are distinct failure modes with different causal paths and different maintenance priorities. Attempting to combine them into one tree at a shared top event produces an unmanageable diagram and logical errors in the gate structure.

How do I calculate failure probability from a fault tree diagram?

For quantitative fault tree analysis, assign a failure rate (failures per operating hour or per year) to each basic event from historical maintenance records, OEM specifications, or industry databases such as MIL-HDBK-217 or OREDA. Propagate rates upward through gates using the following logic: for AND gates, multiply the probabilities of all input events; for OR gates, add the probabilities (using the exact formula: 1 − (1−P1)(1−P2)... for precision). The calculated probability at the top event gives you the expected failure frequency. For most maintenance teams, qualitative cut set analysis is sufficient — quantitative calculation adds value when you need to compare design alternatives or justify capital expenditure on redundancy.

How often should a fault tree diagram be updated?

A fault tree diagram should be reviewed and updated after any of the following: a significant equipment modification (new sensors, control system changes, hardware upgrades), a change in operating conditions (different shift patterns, new production requirements, changed raw materials), the actual occurrence of the top event, or the occurrence of any basic event that had been assumed unlikely. In practice, most industrial maintenance teams review active fault trees once or twice per year as part of a formal reliability review cycle — and trigger an immediate review whenever an unexpected failure occurs that the existing tree did not predict.

A well-drawn fault tree diagram is the clearest picture your maintenance team has of which failure paths are actually threatening your critical assets. Schedule a free demo to see how Cryotos CMMS stores your fault tree diagrams, converts minimal cut sets into work orders, and tracks whether your corrective actions are actually reducing failure frequency over time.

Want to Try Cryotos CMMS Today?

Get Free Demo

Let AI Take Control of Your Maintenance

Cryotos AI predicts failures, automates work orders, and simplifies maintenance—before problems slow you down.

Try AI-Powered CMMS
🡢