The Swap Fixes It, But You Still Don't Know Why: Decision Tree
Why this matters
Sometimes a substitution test does exactly what it is supposed to: you swap in a known-good part, following a real hypothesis, and the fault clears cleanly. The part was confirmed as the cause. But "the part was the cause" and "you know why the part failed" are two different findings, and the second one does not automatically come with the first. You can run a textbook-correct substitution test and still be standing there unable to say why this specific part failed on this specific unit at this specific time. This tree is for that exact spot: the diagnosis is confirmed, the root cause is not, and you have to decide what to do about the gap honestly, not paper over it.
Start here: this is not the same problem as "did the swap actually work"
If you are unsure whether the swap really fixed the fault or just coincided with it stopping, that is a different question with its own tree (see the related article on verifying a swap before you credit it). Here, assume the confirmation is solid: you proved the mechanism, the part was genuinely bad, and the fault is genuinely gone. The open question is narrower and specifically about the "why," not the "what."
Check the obvious root causes first, quickly
Before accepting that the cause is unknown, rule out the common ones. This should take minutes, not the rest of the visit.
If the part shows a wear pattern consistent with simple age or normal duty cycle (even wear, expected lifespan for the component and its runtime, no external damage), the root cause is ordinary wear-out. Document it as such and move on; not every failure needs a deeper investigation.
If there is visible evidence of an external stressor (heat discoloration suggesting a nearby overheating source, contamination, a physical impact, a connection that shows arcing or corrosion), you likely have your answer. Address the external cause, not just the part, or the replacement inherits the same risk.
If neither is present and the part failed with no obvious wear pattern and no visible external cause, continue.
Decide how much further investigation is actually warranted
Not every unexplained failure justifies extended diagnostic time on this visit. Weigh it honestly:
If this is a low-cost, commonly-replaced component with no history of repeat failures on this unit or similar units, further investigation likely costs more in billable time than the risk it is protecting against. Document the failure and the fact that no obvious external cause was found, and move on. An unexplained but isolated failure on an inexpensive part is a normal cost of doing business, not a mystery worth solving on the clock.
If this is a component that is expensive, labor-intensive to replace, or safety-relevant, or if this exact unit or a fleet of similar units has failed the same way before, the unexplained cause is worth more attention. Continue to the next check.
If deeper investigation is warranted, look upstream and at conditions, not just the part
- Check the supply feeding the part: voltage, pressure, or flow outside normal range can stress a component into premature failure without leaving damage on the component itself that you would recognize.
- Check for a recent event: a power event, a recent modification, a recent repair by someone else, an environmental change (new equipment nearby, a changed duty cycle, a seasonal extreme). See the related article on correlating a fault with a recent power event.
- Check whether this is the first occurrence or part of a pattern across similar units you service. A single unexplained failure is an anomaly; the same failure on a second or third unit is a signal that something systemic is going on.
If you genuinely cannot determine the cause on this visit
This is a legitimate outcome, not a failure of diagnosis. Some root causes require monitoring over time, data you do not have access to on a single visit, or conditions that only recur occasionally.
- Say so directly to the customer. "The part was confirmed as the cause and it's fixed, but I can't tell you yet why it failed. I want to keep an eye on it" is honest and it is a stronger position than guessing at a cause to sound complete.
- Document the unresolved question explicitly in the record, not just the fix. A future tech, possibly you, benefits enormously from seeing "cause of failure undetermined, monitor for recurrence" instead of a record that implies the investigation was closed and complete.
- Set a concrete follow-up trigger where practical: ask the customer to note if the same symptom returns, or flag the unit for a check at the next scheduled visit.
Do not manufacture a cause to sound thorough
The failure mode to avoid is worse than admitting you do not know: inventing a plausible-sounding cause you have not actually confirmed, just because "I don't know" feels unsatisfying to say out loud. A guessed cause written down as if confirmed misleads the next person who reads the record and treats it as settled fact, and it can send a future diagnosis in the wrong direction entirely.
Decision recap
- Confirm the swap itself was a real, verified fix before asking about root cause at all.
- Rule out ordinary wear-out and obvious external damage first, quickly.
- Weigh whether further investigation is worth the time, based on cost, safety relevance, and repeat-failure history.
- If warranted, check upstream supply conditions, recent events, and fleet-wide patterns.
- If the cause still cannot be determined, say so plainly, document it as an open question, and set a follow-up trigger rather than guessing.
References
- Trade-standard root-cause methodology
- Manufacturer documentation on component failure analysis
- See related: Swap Confirmed the Fault, Now What (decision tree)
- See related: Documenting What You Tested, Even When a Swap Fixes It
- See related: Correlating a Fault with a Recent Power Event