The Ghost Fault: Document and Monitor
Why this matters
A ghost fault is a failure that is real but will not show up while you are standing there. You cannot fix what you cannot see, and you cannot see it on demand. The trap is treating "I cannot make it happen" as "there is nothing wrong," then either walking away or blind-swapping parts. The professional path is neither: you build a trap that catches the fault the next time it appears, and you document so well that the return visit starts with evidence instead of a fresh guess. Done right, the second trip is short and certain. This card is how to build that trap.
Accept what a ghost fault is
A ghost fault is intermittent by definition: present sometimes, absent now. That means three things you must internalize:
- It is real. A customer describing a clear, repeatable symptom is not imagining it. Believe the report even when the meter disagrees.
- It has a trigger. Something is different when it fails - temperature, load, time, a condition, a position. Random is just a trigger you have not named.
- You will not find it by looking harder right now. You find it by being ready when it returns. Shift your goal from "fix today" to "capture next time."
Document like the next tech is a stranger
Assume the person on the return visit (maybe future you, maybe someone else) knows nothing about today. Write down:
- The exact complaint, in the customer's words and conditions. When it fails, what they were doing, what else was running, time of day, weather.
- Everything you tested and the actual readings, not "checked, OK." The next tech needs to know the values were good so they do not re-walk your ground.
- What you ruled out and why. Each eliminated cause saves the next visit time.
- Your leading suspicion and the evidence for and against it.
- What you left in place to catch it and how to read it.
A clean record turns a second blind diagnosis into a five-minute confirmation. A sloppy record means you pay for the same diagnosis twice.
Build the trap: instrument the fault
Monitoring is not "call me if it happens again." It is leaving a tool that records the failure when you are not there. Match the recorder to the suspected trigger:
| Suspected trigger | What to leave watching |
|---|---|
| Voltage sag or spike | A logging meter or min/max recorder on the supply |
| Current draw at failure | A clamp meter with memory on the load |
| Temperature-related | A recorder noting the reading across heat-up and cool-down |
| Pressure swing | A gauge with a drag pointer or a logging pressure tool |
| Timing or sequence | A tool that timestamps the event so it ties to a cycle stage |
The goal is a fingerprint. When the fault fires, the recorder captures the value at that instant, and that single captured reading usually names the cause that hours of live testing could not.
Enlist the customer as your sensor
The customer is on site every hour you are not. Brief them to be your data collector:
- Note the exact time the next failure happens. Time-stamping lets you correlate it to weather, other loads, or a cycle.
- Note the conditions - what was running, what they were doing, hot or cold, recent or after long runtime.
- Do not reset or fiddle if it is safe to leave as-is, so you can see the failed state on arrival.
A customer who hands you "it failed at 6:40 this morning, right after the other big load kicked on" has just done half your diagnosis.
Set the trigger for the return
Before you leave, agree on what brings you back and what it will accomplish:
- Call the moment it acts up, or after the recorder has run through one failure window.
- The return visit reads the captured data and acts on it, not on a new guess.
- Be clear about cost and expectation for that next step so there is no surprise.
What documenting and monitoring is NOT
It is not a polite way to give up. It is not a token part thrown at the problem to look busy. And it is not skipping the boring causes - you still rule out a changed setting, a marginal supply, or normal behavior before you call something a ghost. A true ghost fault is one you have honestly chased, cannot reproduce, and have now set a trap for.
The judgment to bank: you do not beat a ghost fault by staring at it. You beat it by writing down everything you know and leaving a recorder that catches it in the act, so the next visit is confirmation, not a fresh hunt.
References
- Trade-standard practice for intermittent-fault data logging and documentation
- Manufacturer documentation on operating tolerances and error-capture features
- OSHA general guidance on safe energized monitoring and lockout/tagout where applicable (29 CFR 1910)
- See related: Monitor vs Replace on an Intermittent; The Fault That Only Happens When You're Gone