One Bad Batch vs a Design Flaw: Decision Tree
Why this matters
Once you have noticed the same failure two or three times, the next question decides everything: is this one bad production run, or is every unit of this design carrying the same weakness. Call it a batch issue when it is really a design flaw, and you keep installing the same problem into new jobs indefinitely, chasing "bad luck" that is not luck at all. Call it a design flaw when it is really one bad batch, and you needlessly distrust or avoid a part that is fine outside that narrow window. This tree separates the two using evidence you can actually gather in the field.
Start here: what evidence do you have
You need at least the failure mode, and ideally the manufacture date, batch or lot code, and install date, for each failed unit. If you do not have this yet, stop and go collect it (see the related article on reading a fleet-wide pattern) before trying to conclude anything. A decision made without dates and codes is a guess dressed up as an investigation.
If failures cluster around a specific manufacture date or batch code
If you have three or more failed units, all sharing the same failure mode, and their manufacture dates or batch codes fall within a narrow window, while units of the same design made before or after that window are not showing the same failure, this points strongly to a bad batch.
- Confirm the boundary. Check a few units made just before and just after the suspect window, if any are accessible. Clean performance on both sides of the window strengthens the batch conclusion; failures bleeding outside the window weaken it and push you back toward a design-flaw investigation.
- Check for a known batch or lot recall or advisory before doing further independent investigation. See the related article on the recall question.
- If no known issue is published, document the pattern with specifics (failure mode, count, date range) and report it to the supplier or manufacturer. A single well-documented report with real dates gets taken far more seriously than "we have had some fail."
- Proactively check or flag other installed units from the same batch window, even ones that have not failed yet, since a batch defect often means every unit from that run carries elevated risk even before it manifests.
If failures are spread across a wide range of manufacture dates
If the failed units share the same failure mode but their manufacture dates or batch codes are scattered across a wide range, with no tight clustering, this points toward a design flaw rather than a production defect: something inherent to how the part is designed, not how one run of it was made.
- Rule out a shared external cause first. Confirm the failures are not actually explained by a shared installation method, a shared site condition, or a shared usage pattern, any of which can look like a design flaw across scattered dates but is actually an environmental or install cause instead. Cross-check against the fleet-pattern guidance.
- Check whether the failure mode matches a known, documented weak point for that design before assuming you have found something new. Manufacturers and trade groups sometimes publish known limitations or recommended modifications for a design's known weak point.
- If confirmed as a design-level issue, the fix is systemic: use a different part or design where practical, add a modification or mitigation where the manufacturer supports one, or at minimum flag it in your own shop's install standards so new installs anticipate the weak point rather than discovering it later.
- Communicate it as a known limitation, not a one-off repair, to any customer with the same design installed, so they understand why you are recommending a change rather than just another like-for-like replacement.
If the evidence does not clearly sort into either bucket
If you have fewer than three confirmed instances, or the failure modes are not actually the same across units, or you cannot get manufacture dates or batch codes, you do not yet have enough to call it either way.
- Keep tracking. Add each new instance with as much detail as you can get.
- Do not act as though you have confirmed a pattern (do not warn other customers, do not change your install standard, do not escalate to a supplier) based on incomplete evidence. A false alarm costs credibility the next time you do have a real pattern to report.
- Reassess once you cross the three-to-five instance threshold with consistent failure mode, or once you obtain the dates and codes that were missing.
Decision recap
- Confirm you have real evidence: failure mode, count, manufacture date or batch code, install date.
- Tight clustering around a date or batch window, with clean units on either side, means bad batch.
- Wide scatter across dates with the same failure mode, and no shared install or site cause, means design flaw.
- Incomplete or inconsistent evidence means keep collecting, not concluding.
- Batch issues get reported with specifics and other batch units get flagged; design flaws get addressed in your install standard and communicated as a known limitation.
References
- Manufacturer documentation on batch, lot, and serial-date coding
- Trade-standard practice for defect investigation and reporting
- See related: Reading a Fleet-Wide Failure Pattern
- See related: The Recall Question, Is This a Known Issue