Resources · Incident reporting
Root cause analysis explained
Root cause analysis is the discipline of asking why an incident was possible, rather than stopping at who was involved.
Root cause analysis is a structured method for finding the underlying conditions that allowed an incident to happen, rather than the immediate action that triggered it. The common methods are 5 Whys, cause and effect (fishbone) diagrams, fault tree analysis and barrier analysis. The test of a genuine root cause is simple: if you removed it, could this incident still happen?
What is a root cause, exactly?
Most incidents have three layers. There is the immediate cause, which is what happened at the moment of the event. There are underlying causes, the conditions that made the immediate cause possible. And there are root causes, the organisational decisions or absences that allowed those conditions to exist.
An operator reaches into a machine and is injured. The immediate cause is the reach. The underlying cause is that the guard interlock had been defeated. The root cause is that defeating it was the only way to clear a jam that happened several times a shift, everyone knew, and no one had raised it because the machine was meeting its output target.
Stop at the first layer and you retrain the operator. The next operator does the same thing, for the same reason. Stop at the second and you replace the interlock. It gets defeated again within a fortnight. Only the third layer changes anything, and it is the only one that costs money and requires a decision, which is why investigations so often stop before reaching it.
The 5 Whys method, and where it goes wrong
5 Whys is the most used technique because it needs no training and no software. State the problem, ask why it happened, then ask why of the answer, and keep going until you reach something you can actually change. Five is a rule of thumb, not a rule.
It has two well known failure modes. The first is that a single chain of whys suggests incidents have one cause, when most have several running in parallel. The fix is to allow the chain to branch rather than forcing one line.
The second is more damaging. Ask why often enough without discipline and you arrive at “the operator did not follow the procedure”, declare that the root cause, and stop. That is not a root cause. It is a restatement of the immediate cause in blaming language. The useful next question is why not following the procedure was possible, or easier, or faster, or normal. If the honest answer is that the procedure cannot be followed as written while hitting the target, you have found something worth fixing.
Cause and effect diagrams and when to reach for them
A cause and effect diagram, also called a fishbone or Ishikawa diagram, puts the problem at the head and groups possible causes along branches. In manufacturing the branches are commonly the six Ms: machine, method, material, manpower, measurement and environment. In service settings people often use people, process, policy, place and technology instead. The labels matter less than the fact that they force you to look in more than one direction.
Its value over 5 Whys is breadth. It is a group exercise that surfaces contributing factors nobody would have reached along a single chain, and it makes visible how many things had to line up. Its weakness is that it produces candidates, not conclusions. A fishbone with forty branches and no evidence attached is a brainstorm, not an analysis. Use it to generate the list, then test the branches against what you actually know.
Fault trees, barrier analysis and the bigger events
For serious or complex incidents, two more formal methods earn their overhead.
Fault tree analysis works backwards from the unwanted event through logic gates, showing which combinations of failures were necessary to produce it. It is the right tool when several things had to fail together and you need to understand which combinations remain possible.
Barrier analysis asks a different question: what was supposed to stop this, and why did it not? Every hazard is separated from the people it could harm by a set of barriers, some physical, some procedural, some human. Listing them and marking each as absent, failed or bypassed produces a very direct set of actions, and it has the useful property of pointing at systems rather than individuals.
Neither is necessary for a trip in a corridor. Both are worth the time when the consequence was severe or could easily have been.
Turning findings into actions that hold
An analysis that produces no change is an expensive way of writing history. Three things separate the investigations that stick from the ones that do not.
- Actions target the layer you found. If the root cause was organisational, a toolbox talk is not the answer. Retraining is the default action precisely because it is the cheapest and least disruptive, which is also why it so rarely works.
- Each action has one owner and a date. Actions assigned to a department are assigned to nobody. This is where most corrective action programmes quietly fail.
- Effectiveness gets checked later. Closing an action means the task was done. It does not mean the cause was removed. A review some months on, asking whether the condition has actually gone, is what turns a fix into a verified fix.
That last step is the same logic that sits behind formal corrective and preventive action programmes in quality management. See CAPA explained for how the same discipline is applied to non-conformances, and incident investigation basics for the surrounding process.
Frequently asked questions
What is the difference between a root cause and an immediate cause?
The immediate cause is what happened at the moment of the incident, such as a slip or a wrong setting. The root cause is the organisational condition that made it possible, such as an unworkable procedure or a control that was never maintained. Fixing the immediate cause rarely prevents a recurrence.
Is “human error” ever a root cause?
No. Human error is a starting point for an investigation, not a conclusion. The useful question is what made the error likely, or easy, or unnoticed. Ending an analysis at human error usually means the investigation stopped at the point where it became uncomfortable.
How many whys should I ask in a 5 Whys analysis?
As many as it takes to reach something you can change and that would prevent recurrence. Five is a memorable average, not a target. Some chains resolve in three, others need eight, and most benefit from branching rather than following one line.
Which root cause analysis method should I use?
5 Whys for straightforward incidents, a cause and effect diagram when you need breadth and a group view, and fault tree or barrier analysis for serious events where several failures combined. Match the effort to the potential consequence, not to the actual outcome.
Can an incident have more than one root cause?
Usually it has several. Serious incidents almost always require multiple contributing conditions to line up, which is why single-chain methods can mislead. Recording all of them is more honest and produces a better set of actions.
When should root cause analysis be done?
Soon enough that evidence and memory are intact, which usually means days rather than weeks. The severity of what could have happened, rather than what did, is the better guide to how much analysis is warranted.
How do you know the root cause is right?
Apply the removal test. If this condition had not existed, could the incident still have happened in the same way? If yes, you have not reached the root cause. If the incident becomes impossible or clearly much less likely, you have.
Sources
- Health and Safety Executive, investigating accidents and incidents (HSG245). https://www.hse.gov.uk/pubns/books/hsg245.htm
- Health and Safety Executive, managing for health and safety (HSG65). https://www.hse.gov.uk/pubns/books/hsg65.htm
- James Reason, Managing the Risks of Organizational Accidents, Ashgate, 1997 (origin of the barrier and defences model).
- Kaoru Ishikawa, Guide to Quality Control, Asian Productivity Organization, 1976 (cause and effect diagrams).
Find the pattern before the next investigation
Consistent categories and cause codes across every report turn one-off investigations into trends you can act on.
Book a demo