Structured Troubleshooting
A repeatable method: symptom, split-half, evidence, root cause, verification.
- The method as the transferable skill
- Split-half searching
- Evidence versus elimination
- Cognitive traps under pressure
- Knowing when to escalate
What it is
Everything in the previous fourteen chapters was one method demonstrated on different equipment.
Chapter 3 walked a control string and found the device holding the supply voltage. Chapter 9 walked an I/O chain and found the link where the story changed. Chapter 11 paired a gauge reading with an actuator, Chapter 13 paired a frequency with an amplitude. Every one of those is the same move: narrow the search space with an observation, and choose the observation that narrows it most.
That method is the transferable part of this trade. Equipment knowledge is local and it perishes — you will change plants, and the machines will change under you. The method moves with you.
The claim this chapter makes
Given a machine neither of them has seen, a technician with the method will out-diagnose a technician with more equipment knowledge and no method.
That is not a slight on knowledge. Knowledge makes you faster on machines you know. But most of a career is spent in front of things you have not met, and on those, structure beats familiarity.
Split-half is the most powerful single technique in fault-finding and it works on anything with a chain structure: a control string, a pipe run, a signal path, a conveyor line, a process. The question it asks of every measurement is "what will this rule out?" — and a measurement that cannot rule anything out is not worth taking.
How it works
The method
1. Make it safe. Every energy source inventoried, isolated and proved, or a deliberate, authorized decision to work live because the fault only exists while running. Chapters 1 and 2, and it is not a preliminary — it is the first step of the diagnosis, because it forces you to understand the machine's energy before you touch it.
2. Define the symptom precisely. Not "it's not working". What stopped, when, under what conditions, how often, and what changed. Ask the operator: they have watched this machine for longer than you have, and the phrasing they use is frequently the diagnosis. "It starts and then stops" eliminated half of Chapter 8's rung. "One, then three good" named the head in Filler 02 before any measurement.
3. Take the free information first. Alarm history, fault logs, work orders, the last shift's notes, module LEDs, filter bowls, the sensor's own indicator, oil color, sound, smell, temperature. All of it costs nothing and most of it gets skipped in favor of a meter, which is how a five-minute job becomes an afternoon.
4. Form a model, then predict. Work out what the reading should be before you take it. Ohm's law on a relay coil, force from pressure and area, flow from what an exhaust can pass, a frequency from a running speed. A measurement that matches your prediction confirms both the circuit and your understanding. One that does not means either the machine or your model is wrong — and you have learned something either way, which is not true of a number you had no expectation for.
5. Halve the search. Do not walk it. Measure where the answer splits the remaining possibilities most evenly, then repeat.
6. Separate evidence from elimination — and value both. What a measurement rules out is half the diagnosis, and the half most often left out of the story. Conveyor 04's balanced motor currents proved nothing about the fault and eliminated an entire family of them.
7. Find the root cause, not the failed part. A blown fuse, a tripped overload, a burned contact and a failed bearing are all consequences. Chapter 5's overload was reporting a real thermal condition; Chapter 12's third bearing was reporting a belt tensioned by feel. Replacing the component without answering why buys one shift and returns the fault to somebody else's night.
8. Verify under real conditions. The same cycle, the same load, the same position, the same temperature that produced the fault. A repair confirmed at standstill is not confirmed. This is also where you find out whether you fixed the thing you thought you fixed.
9. Record it. What the symptom was, what you checked, what it ruled out, what it turned out to be, and what it cost. This is the step with no immediate payoff and the largest compound return: it is the next person's step 3, and when the next person is you in eight months, it is the difference between a diagnosis and a fresh start.
Where the method actually breaks
It is rarely the technique that fails. It is the thinking around it, and the failures are consistent enough to be worth naming.
Anchoring. The first hypothesis gets privileged, and everything after it is read as support. The alarm text is the most common anchor of all — motor speed mismatch invites you to suspect the encoder, which is exactly the trap Conveyor 04 is built on.
Confirmation bias. Once you have a theory you start taking measurements that could confirm it rather than measurements that could kill it. The correction is mechanical: for each hypothesis, ask what observation would prove it wrong, and go and make that observation.
Availability. The last three faults you saw feel more likely than they are. If the previous two callouts on this line were sensors, the third one gets diagnosed as a sensor.
Sunk cost. Two hours into a theory, abandoning it feels like wasting two hours. It is not — continuing is what wastes the next two.
Pressure. Somebody is standing behind you and the line is down, and the method feels like a luxury. It is the opposite: under pressure is precisely when unstructured searching turns into part-swapping, and part-swapping is slower than it feels.
Knowing when to stop
Not every fault is yours to solve, and recognizing that is part of the method rather than a failure of it.
Escalate when the fault is outside your authorization, when it needs equipment or access you do not have, when it crosses into somebody else's system — the Palletiser 01 case — or when you have stopped narrowing the search and started repeating measurements. The last one is the hardest to notice and the most useful to name.
When you hand over, hand over the evidence, not just the symptom: what you checked, what it ruled out, what you would do next. That is worth far more than a description of the fault, and it is exactly the structure of a Fault Log entry.
What normally fails
- Symptom
- An hour spent measuring, and the search space is no smaller than when you started
- Likely cause
- Measurements chosen for availability rather than for what they would rule out
- How common
- Very common
The signature of unstructured searching. Every measurement should have an answer to "what does this eliminate?" before it is taken. If it does not, it is activity rather than progress — and it feels like work, which is why it can continue for a long time.
- Symptom
- The diagnosis matches the alarm text, and the repair does not fix it
- Likely cause
- Anchoring on what the machine called the fault rather than on what it did
- How common
- Very common
Alarm text names a symptom and is frequently written by someone describing what the control system detected, not what failed. Conveyor 04's MOTOR SPEED MISMATCH is accurate and points at nothing useful. Read the alarm as evidence, never as a conclusion.
- Symptom
- A fault that keeps coming back on a different shift
- Likely cause
- The failed component was replaced and the cause was never established
- How common
- Very common
Every chapter has a version of this: the overload that was reset, the fuse that was replaced, the bearing that was fitted, the seal that was renewed. The component is a consequence, and a consequence recurs until its cause is addressed.
- Symptom
- A repair that tested fine and failed within the shift
- Likely cause
- Verified at standstill, or under conditions that did not reproduce the fault
- How common
- Common
Faults that appear under load, at temperature, at one position, or only when another machine runs will all pass a static test. Verification has to reproduce the conditions that produced the symptom, which sometimes means waiting.
- Symptom
- Two technicians independently diagnose the same machine from scratch
- Likely cause
- Nothing was recorded the first time
- How common
- Very common
The most expensive habit in maintenance and the least visible, because the cost lands on somebody else. Every unrecorded diagnosis is thrown away, and plants where nothing is written down solve the same faults repeatedly for years.
- Symptom
- A theory pursued long after the evidence stopped supporting it
- Likely cause
- Sunk cost — abandoning it would mean the last two hours were wasted
- How common
- Common
They were wasted either way; the only question is whether the next two join them. A useful habit is to state the theory out loud with the evidence for it, because a theory that sounds thin when spoken usually is.
How to troubleshoot it
The method, as a sequence. This is the one to carry.
Make it safe, and understand the energy while you do it
SafetyInventory every source, isolate, prove. If the fault only exists while running, make that a deliberate, authorized decision rather than a default. The inventory itself is diagnostic — you cannot list a machine's energy sources without learning how it works.
Define the symptom until it is specific
What, when, under what conditions, how often, and what changed. Ask the operator and listen to their phrasing. Keep going until the description would let somebody else reproduce the fault.
Take everything that is free
Alarm history, logs, work orders, previous shift's notes, indicator LEDs, filter bowls, oil, sound, smell, heat, and the drawing's revision block. Most of it is skipped in favor of an instrument, and most callouts contain their answer somewhere in it.
Predict the reading before you take it
From the model that applies: Ohm's law, pressure times area, flow through an exhaust, a frequency as a multiple of running speed. The prediction is what turns a number into information.
Choose the measurement that eliminates the most
SafetyNot the easiest one, and not the one nearest the symptom. Measure where the result splits the remaining possibilities most evenly, then repeat on what is left.
Ask what would prove your theory wrong
Then go and check that. This is the single most effective correction for confirmation bias, and it costs one measurement. A theory that survives an honest attempt to kill it is worth acting on.
Keep going past the failed component
The part in your hand is a consequence. Ask what loaded it, starved it, heated it, contaminated it or let it move. Stop when the answer is something that will not recur on its own.
Verify under the conditions that produced the fault
Same load, same cycle, same position, same temperature. Then watch it for longer than feels necessary, because the faults that come back within the shift are the ones that were verified in a hurry.
Write it down before you leave
Symptom, what you checked, what it ruled out, root cause, how you proved it, what it cost. Five minutes, and it is the only step whose value compounds — for the next person, for the machine's history, and for you when somebody asks you to walk them through the hardest fault you have found.
Common technician mistakes
Letting the alarm text set the hypothesis
WhyThe alarm is the first thing you read, it is written in confident capitals, and it names a component often enough to be believable. But it describes what the control system detected, which is a symptom — and a symptom named precisely is still a symptom. Conveyor 04 exists to teach this one specific reflex.
Taking the measurement that is easiest to reach
WhyThe accessible test point is right there and the informative one means moving a guard or a ladder. So the easy measurement gets taken, produces a number that rules nothing out, and the search space is unchanged — but it feels like progress, which is what allows it to be repeated for an hour.
Looking for evidence that fits
WhyOnce a theory exists, every subsequent observation gets read through it, and ambiguous readings resolve in its favor without anybody deciding to cheat. The fix is procedural rather than moral: name what would disprove it, and go and look for that instead.
Abandoning the method under pressure
WhySomebody is waiting, the line is losing money, and structure feels like a luxury you cannot afford. It is the reverse — unstructured searching under pressure becomes part-swapping, which is slower, more expensive, and far more likely to leave the fault in place.
Stopping when the machine runs
WhyThe machine running is an extremely convincing signal, and it is the point at which everybody's attention leaves. But it only proves the symptom has gone, which is not the same as the cause having gone — and the difference shows up on somebody else's shift.
Not writing it down because you will remember
WhyYou will, for about two weeks. The cost of forgetting is invisible at the time and lands as a fresh diagnosis months later, usually on the same machine and often on you. It is the cheapest step in the method and the first one dropped.
Hands-on challenge
Scenario
A machine you have never seen — capstone
You are covering nights at a site you started at last week. A machine you have never seen — an unfamiliar make of cartoner — has stopped twice in three hours. There are no drawings in the panel, the day-shift technician who knows it is not contactable, and the operator has been running it for two years.
You have your instruments, this method, and nothing else.
Write down, in order, what you would do — and for each step, say what it would rule out. Then answer three questions specifically:
- Which three pieces of information would you try to get from the operator first, and why those three?
- The machine has an HMI showing a fault history you do not understand the codes for. What can you still extract from it?
- At what point would you stop and escalate, and what exactly would you hand over?
There is no single correct answer here. A good one is a sequence where every step narrows the search, and where none of the steps depend on having seen this machine before.
Show how to approach it
There is no equipment knowledge available to you here, which is the point. Work the method and notice how much of it does not depend on knowing the machine.
- Safety first, and it is also reconnaissance. You cannot inventory the energy sources without learning what the machine has — pneumatics, hydraulics, a drive, stored gravity. Chapter 1's list is doing double duty.
- The operator’s phrasing is the highest-value evidence available. “Every fourth one”, “only after about an hour”, “it started after the weekend” — each of those has been a whole diagnosis in a mission on this site.
- The free information is entirely equipment-independent. Alarm history, work orders, the last shift’s notes, filter bowls, oil, LEDs, sound, smell, heat. None of it requires you to know the machine.
- Periodicity, correlation and timing are your best tools on unfamiliar equipment. A fault on a cycle of four names a station. A fault that appears after an hour is thermal. A fault that started after a shutdown belongs to whatever happened during it.
- Split-half needs a chain, not a schematic. Whatever the machine does, it does in a sequence — so measure in the middle of the sequence and ask which half the fault is in.
- Escalate deliberately, not apologetically. If the answer crosses into somebody else’s equipment or authorization, hand over the evidence and what it ruled out — which is worth far more than the symptom.
Knowledge check
Five questions. Every one of them is about the method rather than about any piece of equipment, which is the whole argument of the course.
Question 1 of 5