The Fault That Only Happens When You're Gone: Decision Tree
Why this matters
The fault that clears the moment you arrive is the one that eats your margin and your reputation. You drive out, everything runs, you leave, and the call comes back the same week. Each round trip costs you a slot you cannot bill and a customer who trusts you a little less. The skill that separates a good diagnostician here is refusing to chase the symptom and instead chasing the condition that was present when it failed and absent when you tested. This tree is how you do that without guessing.
Start here: stabilize any hazard first
Before any of the diagnostic work below, confirm the intermittent is not a safety event in disguise. A breaker that trips "sometimes," a faint gas smell that "comes and goes," a unit that "occasionally" gets hot, or water near live electrical is not an intermittent to monitor. De-energize, shut off the gas or water, and verify it is safe before you do anything else. An intermittent hazard is still a hazard. Only a fault with no safety dimension belongs on the monitor-and-recreate path.
The core idea: what was different
An intermittent fault is not random. Something was different when it failed. Your whole job is to name that variable. Walk the four families of "different" in order, because they get easier to the harder:
- Different temperature - the fault appears cold (first startup) or hot (after running for an hour). Thermal expansion opens a hairline crack, a marginal connection, or a tolerance.
- Different load or demand - the fault appears only under peak draw, full flow, or a specific cycle stage. The system is fine until you ask the most of it.
- Different time or condition - only overnight, only when it rains, only when another large load runs, only on a certain cycle. The trigger is outside the unit.
- Different mechanical position - only when a door is shut, a panel is on, a wire is flexed, a moving part reaches one spot in its travel.
Ask the customer which of these matches. Their answer is the most valuable data you will get all day.
Branch by who can reproduce it
If the customer can reliably make it happen (it always fails on the third wash, every morning, whenever the dryer runs): you have a recipe. Run that exact recipe with your meters already connected. Do not test your way; test their way. The fault you can recreate is a fault you can find.
If it happens on a schedule you can wait out (fails every evening, fails after an hour of runtime): stay or come back at that window with instrumentation in place. Billing a return visit timed to the fault is honest work and far cheaper than three blind trips.
If no one can predict it (truly random, no pattern): you are on the monitor-and-document path. Trying to "find it" cold is how you replace good parts and still get the callback.
Recreate it on purpose
Once you know the variable, attack it directly while watching the meter:
- Suspect thermal: heat the suspect connection or component with a heat source, or chill it with freeze spray, and watch for the reading to jump. Wiggle-test wiring and connectors under load. Flex the harness where it bends. A reading that moves when you move or heat one spot has just named the fault.
- Suspect load: put the system under its worst-case demand and hold it there. Watch voltage sag, pressure drop, temperature climb, or amperage spike at the moment of failure. The component that cannot hold under load is failing even if it tests fine at rest.
- Suspect external condition: check what else shares the circuit, the supply, or the space. A neighbor load, a shared neutral, a marginal supply voltage, or weather intrusion is often the real cause and the unit is innocent.
When you cannot recreate it: monitor
If the fault will not come out for you, do not pretend you found it. Set up to catch it the next time it happens:
- Leave a logging meter, a min/max recorder, or a current clamp with memory on the suspect circuit or line. The next fault leaves a fingerprint you can read on your return.
- Have the customer note the exact time and conditions the next time it fails. Time-stamping the failure lets you correlate it to weather, other loads, or a cycle stage.
- Document what you tested, what passed, and what you are watching. Write it down so the next visit starts where this one ended, not from zero.
Tell the customer plainly that an intermittent that will not reproduce gets monitored, not blindly parts-swapped, because swapping good parts wastes their money and still leaves the real fault in place.
Decide: monitor, replace on suspicion, or escalate
| Situation | Right move |
|---|---|
| Fault recreated, one component clearly moves the reading | Replace that component, verify the recipe no longer fails |
| Pattern known, instrument in place, fault not yet caught | Monitor through one more cycle, then decide on data |
| No pattern, low consequence if it recurs | Document, monitor, set expectation for a return |
| No pattern, high consequence (no heat in winter, no water) | Replace the most-suspect single item, instrument, and watch |
The judgment to bank: you have not diagnosed an intermittent until you can make it happen on demand or you have a recorder waiting for it. Anything in between is a guess wearing a work order.
References
- Trade-standard practice for intermittent-fault isolation and load testing
- Manufacturer documentation on operating tolerances and cycle sequencing
- OSHA general guidance on energized-work and lockout/tagout for hazard branches (29 CFR 1910)
- See related: Everything Tests Good But It Fails; The Ghost Fault: Document and Monitor