Found The Easy Fault: Keep Looking Or Stop Decision Tree
Why this matters
You found a fault quickly. It explains the symptom. The customer wants you done. Do you stop and verify, or keep looking for a second fault? Stop too soon and you may have addressed a symptom of a deeper problem that will resurface as a callback. Keep looking when one fault was the whole story and you waste the customer's time. This decision tree gives you a way to make the call without guessing.
The principle comes from the multi-cause analysis tradition in reliability engineering (ISO 14224) and the systematic-completion guidance in NFPA 70B: a single found fault is not a complete diagnosis until you have ruled out coexisting faults that share the same symptom.
Symptom presentation
- You arrive with a complaint, find an obvious cause within minutes (blown fuse, tripped switch, disconnected wire, visibly failed component)
- Replacing or restoring the obvious cause clears the original symptom
- The customer is satisfied and ready to sign
- But you have not actually swept the system for coexisting conditions
The question is whether to call the job done or to keep looking.
Quick checks before calling done
- Why did the obvious cause fail? A fuse blew for a reason. A switch tripped for a reason. A wire came loose for a reason. The cause-of-cause matters.
- Were there other indicators on the panel or in the conversation that the easy fault does not explain?
- Is the failure mode consistent with the obvious cause being primary, or is the obvious cause itself a downstream effect?
- Are there safety implications that you have not verified post-restore?
- How long has the symptom been present? A long-standing intermittent paired with a recently-discovered obvious cause is often a coincidence.
Decision tree
Branch A: Stop, the easy fault was the whole story
Take this branch if all of the following are true:
- The obvious cause fully explains every symptom the customer reported
- The failure mode of the obvious cause is consistent with normal wear or a clear external event (lightning, power surge, foreign object, age)
- Post-restore operation is clean across at least one full cycle
- No coexisting indicators (codes, residual noise, residual heat, residual leak) remain
- The customer's reported pattern matches a single-cause story (started suddenly, no prior issues)
Examples: a single blown fuse from a documented power surge during a storm last night; a contactor with normal end-of-life contact erosion replaced and the unit running cleanly; a clogged condensate line cleared and the safety switch reset with normal drain flow restored.
Branch B: Keep looking, the easy fault may be downstream of a deeper cause
Take this branch if any of the following:
- The obvious cause is a protective device (fuse, breaker, overload, safety switch) and you have not yet established why it tripped
- The failure mode is not consistent with normal wear (fuse blown but no visible short; contactor pitted but unit is new; bearing seized but lubrication looks correct)
- The customer reported multiple symptoms and the obvious cause explains only one
- Panel diagnostics show codes that the obvious cause does not explain
- The symptom has been present off and on for a long time but the obvious cause looks recent
- Post-restore operation shows any anomaly: unusual current draw, slow response, brief glitch, intermittent code
Keep-looking work: trace the cause of the protective trip. Read the load side of the cleared fuse. Measure currents on each phase. Look at the contactor that came in contact with the load that took the fuse out. Trace each unexplained code.
Branch C: Stop now, return to investigate
Sometimes the right call is to restore service, document the open question, and schedule a return for deeper investigation. Take this branch if:
- The customer's immediate need requires restored service now (cooler load at risk, occupied space)
- Continued deep-dive in the current visit would extend downtime past the customer's tolerance
- You suspect a second cause but do not have the time, tools, or parts to chase it in this visit
- You can monitor the system between visits via remote telemetry or the customer's own watch
Write the open question into the customer record. "Restored service after blown fuse; cause of fuse failure not isolated; return scheduled to investigate." This is a legitimate completion state.
Confirming the call
The test of "is this the whole story" is the post-restore observation period. Watch the equipment through at least one full operating cycle in the conditions normal operation imposes. A unit that ran fine on the bench but tripped on the first real load cycle is a unit with a second fault. The minutes you spend watching are the cheapest insurance against a callback you will own.
Read the full panel after the restore. Codes that persist after the obvious cause is cleared are codes that the obvious cause did not generate.
A blown protective device that you replace without finding the cause of the trip is a malpractice pattern. Fuses, breakers, and overloads exist to detect specific fault conditions. Replacing one without diagnosing the underlying fault sets up the next failure as a more energetic event with the protection now reset and ready to clear again. Always investigate the cause of any protective trip before considering the call complete.
Next steps
When you stop with the call closed, document the observed cause, the fix, the post-restore observation, and your rationale for declaring the issue resolved. When you keep looking and find a second fault, document the relationship: the deeper cause and the protective response that masked it. Both are valuable for the next tech.
Build a habit of post-fix verification rounds. After every repair, walk the panel, observe at least one operating cycle, check thermals on related components, and read fault history one more time. The discipline turns "found and fixed" into "found, fixed, and confirmed."
References
- ISO 14224 Reliability and maintenance data collection, multi-cause failure analysis
- ISO 13379-1 Condition monitoring and diagnostics of machines, data interpretation and diagnostics techniques
- NFPA 70B Recommended Practice for Electrical Equipment Maintenance, protective device investigation
- IEC 60812 Failure modes and effects analysis
- OSHA 29 CFR 1910 Subpart S, Electrical safety standards, protective device guidance
- ACCA Standard 4 Maintenance of Residential HVAC Systems