Intermittent Fault Cannot Reproduce Decision Tree
Why this matters
An intermittent fault that the tech cannot reproduce on the visit is the highest-risk diagnostic situation in field service, because every action the tech takes from this point is a probability bet. Replacing the most-likely-failed part on a guess is sometimes right and sometimes wrong; doing nothing and walking away is sometimes right and almost always reads to the customer as incompetence. The protocol below provides a structured path from "I cannot make it fail" to a defensible outcome that either captures the fault, defers with documentation, or replaces a high-probability suspect with a written rationale.
Symptom presentation
The customer reports a problem that happens often enough to be a real complaint but not on a schedule the tech can predict. Common patterns: an appliance that fails once a week, a circuit that trips during a thunderstorm, a noise that only happens after the system has been running for a few hours, a smell that comes and goes. The tech arrives, runs the system, observes nothing wrong.
The customer's frustration at this point is high. They invited a service call because the problem is real to them, and the tech is finding nothing. The tech's diagnostic credibility is on the line in the customer's eyes.
Quick checks before declaring no-fault-found
Read the prior service history. A unit with two prior visits for the same symptom is not a no-fault-found call; it is a structured investigation call. Treat as such.
Ask the customer to walk you through the most recent occurrence in detail. Time of day, temperature, humidity, what else was running, what they had just done before noticing it. Customer recall on intermittent faults is the highest-value diagnostic data on the visit.
Inspect for evidence the fault has happened. Burn marks, water stains, corrosion, intermittent fault codes in the unit's event log, the unit's diagnostic counters showing fault accumulations, error codes the customer reported but did not photograph. Many modern appliances, HVAC systems, and electrical panels log faults that did not produce a customer-visible error indicator.
Isolation tree
Branch A: the fault is event-log captured. Read the unit's event log per the platform service manual. A logged fault gives you the actual failure mode without needing to reproduce it. Walk the tree for that failure mode, repair, and document. The fact that the fault did not reproduce on the visit does not mean it cannot be diagnosed from the log.
Branch B: the fault is correlated with an environmental condition. Temperature, humidity, supply voltage, gas pressure, water pressure, and time-of-day patterns all narrow the search. A breaker that only trips on humid days is a moisture-intrusion call (insulation breakdown in a damp wall cavity, water in a junction box). An HVAC system that only fails on hot days is a charge or refrigerant-flow call. A washer that only fails on heavy loads is a suspension or off-balance-sensor call. Match the pattern to the most-likely subsystem and walk the corresponding decision tree.
Branch C: the fault is correlated with concurrent events. The customer says it fails when the dishwasher and the washer run together, when the central AC compressor starts, when a storm passes through, when a neighbor's pool pump turns on. These point at supply-side issues (low voltage, supply pressure transient, harmonic interference) that the tech needs to capture with instrumentation. Install a data logger (voltage logger on the circuit, temperature logger in the cavity, pressure logger on the supply line) and leave it in place for a week.
Branch D: the fault has no discernible pattern. The customer says it is random. Two possibilities: it has a pattern the customer has not observed, or it is a connection / mechanical fault that opens randomly with vibration or thermal cycling. Inspect for loose connections at junction boxes, breaker terminals, appliance terminals, and the platform's main connector blocks. Re-torque per the manufacturer's spec (most appliance terminals are 12 to 22 in-lb; consult the model service tag). Document.
Branch E: the customer reports the fault has been getting worse. Increasing frequency over weeks is the strongest indicator that the fault will reproduce soon and that returning in 2 to 3 weeks will allow capture. Schedule a follow-up visit at a defined interval, deploy a data logger if available, and set the customer expectation that the next visit is the diagnostic visit, not the repair visit.
Branch F: the customer reports the fault has been the same frequency for a year or more. This is often a chronic upstream condition (water hardness, supply voltage drop in a long branch circuit, gas-pressure marginal) that the appliance is responding to. The investigation moves upstream of the appliance.
Branch G: the customer reports the fault stops when one specific thing happens (turning the unit off and back on, opening a window, rebooting the smart home system). The reset pattern is diagnostic. A fault that clears with a unit-power-cycle is a control-board or firmware issue. A fault that clears with a window-opening is an air-pressure / ventilation issue. A fault that clears with a smart-home reboot is a communications or integration issue.
Confirming or deferring
If the visit ends without a captured fault, the close-out is a documented next step, not a "we will see what happens." The customer needs:
- A written summary of what was tested and what was ruled out.
- A defined follow-up trigger (call us when the fault happens, take a photo of any error code, time-stamp the next occurrence).
- A data logger left in place if the shop has the inventory.
- A scheduled follow-up visit at a defined interval if the customer wants the system monitored proactively.
The shop's policy on no-charge intermittent visits varies. A short-list of high-trust customers and low-cost time blocks generally absorbs the visit; a customer on a third intermittent-no-find visit may receive a higher quote that includes data-logger placement and analysis.
Remediation by branch
Branch A: standard repair per the captured fault.
Branch B: standard repair per the matched-pattern decision tree.
Branch C: data logger placement, return visit to interpret the log, then standard repair.
Branch D: connection retorque, photographed and documented, with follow-up if the fault recurs.
Branch E: scheduled return visit at a defined interval.
Branch F: upstream investigation (water test, voltage measurement, gas-pressure measurement) followed by upstream contractor referral if warranted.
References
- IEEE Std 100 (Authoritative Dictionary of IEEE Standards Terms; intermittent-fault definitions).
- ASHRAE Handbook (HVAC Applications, 2023, Chapter on commissioning and diagnostics).
- NFPA 70 (NEC) Article 110 (general requirements for examination, identification, installation, and use of electrical equipment).
- OSHA 29 CFR 1910 (general safety; technician documentation requirements during diagnostic visits).
- FTC Used Car Rule and Magnuson-Moss Warranty Act, 15 USC 2301 et seq. (consumer-facing warranty disclosure framework).