The denominator problem
Failure rate looks like the simplest metric in the business. Failures divided by units. The difficulty is entirely in the second term, and getting it wrong is the default rather than the exception.
If you divide this month's claims by this month's sales, you are comparing a group of units that have been in the field for six to eighteen months against a group that shipped last week. The first group has had every opportunity to fail. The second has had almost none. The ratio between them means nothing, and it is wrong in a specific, predictable direction: it always understates failure, and it understates it most when you're growing.
The fix is to hold the population fixed. Pick a cohort — a month of sales, a production batch, a shipment — and follow it. Count every failure that cohort produces, and divide by the size of that cohort. Now the numerator and denominator describe the same units, and the number becomes something you can compare, trend and take to a supplier.
A supplier who is shown a calendar-month rate has an easy out: your denominator includes units that couldn't possibly have failed yet, so the figure is not evidence of anything. A cohort rate closes that argument. It's also the only form in which failure data is comparable across your own product generations — which is how you find out whether the "improved" version you switched to last year is actually better.
One caveat worth being honest about: cohorts need maturity. A cohort that's four weeks old will show a low rate because most of its failures haven't happened yet. Compare cohorts at the same age — three-month rate against three-month rate, twelve against twelve — or the comparison reintroduces the same bias by a different route.
What to actually measure
Four numbers between them tell you almost everything worth knowing. Each answers a different question, and each drives a different decision.
Failures per cohort, measured at a fixed age. The headline number. Tells you whether a product line is acceptable and whether it is getting better or worse over time.
How long units survive before failing. Separates manufacturing defects from wear-out, which are two entirely different problems with different owners.
Which failure modes account for the volume. Almost always concentrated — a handful of codes drive most of the cost, and they're the only ones worth engineering effort.
Spread between batches of the same SKU. High variance points at supplier process control rather than product design, and it's the strongest evidence for a credit claim.
Notice what isn't on this list: total return count, return rate across the whole catalogue, and anything expressed as a single business-wide percentage. Aggregate numbers are for board slides. They can't tell you which SKU to renegotiate, which supplier to escalate to, or which fault to engineer out, because the signal you need lives at the level of a product and a batch.
When things fail, not just how often
Two SKUs can have identical failure rates and require completely different responses. What separates them is when in a unit's life the failure arrives. Plotting failures against months since purchase makes the distinction obvious, and it's the single most useful chart in quality analytics.
Manufacturing and assembly defects. These units left the factory faulty. This is the supplier's cost, and the batch data to prove it is sitting in your fault codes.
Design life or duty cycle, not defects. The answer here is a spec change, a different component grade, or honest positioning — not a warranty claim against your supplier.
The spike just before month twelve is not a product behaviour. It's a customer behaviour: people file claims they've been putting off once they notice the warranty is about to expire. It's worth knowing about, because it means a chunk of what looks like month-eleven failure actually occurred earlier, and it will distort your early-life numbers if you don't capture the date the fault was first noticed as well as the date the claim was raised.
Which faults are worth your attention
Failure modes are never evenly distributed. Rank fault codes by volume and the curve is almost always steep — which is good news, because it means a small amount of engineering or supplier attention addresses most of the cost.
Two things make this chart possible, and neither is analytics software. The fault code has to be mandatory at diagnosis, and the code list has to be short enough that technicians use it consistently. Thirty codes will be applied inconsistently by three technicians and the distribution becomes noise. Eight to twelve codes, clearly distinguishable, is the range where this stays reliable.
Thresholds that trigger action
Failure data that produces a monthly report changes nothing. Failure data attached to a threshold changes behaviour, because it converts a number into an obligation. The thresholds below are a starting point — the levels matter less than the fact that crossing one triggers something specific.
| Signal | Threshold | What it triggers |
|---|---|---|
| Cohort rate vs SKU baseline | 1.5× baseline | Flag for review at next supplier call; check whether it's one batch or all |
| Cohort rate vs SKU baseline | 3× baseline | Hold remaining stock, pull batch numbers, open a formal claim |
| Single batch vs sibling batches | 2× spread | Process-control issue at the supplier — escalate to their QA, not sales |
| Early-life share of failures | Above 40% | Manufacturing defect pattern; request incoming inspection data |
| One fault code | Above 30% of claims | Engineering or component review — this is a design problem, not bad luck |
| No-fault-found rate | Above 8% | Product instructions, listing accuracy or setup experience needs work |
| Warranty cost per unit sold | Above target margin % | Pricing, spec or supplier decision — the SKU is not viable as sold |
The last row is the one that most often gets skipped and most often matters. A SKU can have an unremarkable failure rate and still be unprofitable once you load in freight both ways, bench labour, parts and the share of claims where no supplier credit is recoverable. Tracking warranty cost per unit sold, rather than failure rate alone, is what turns quality data into a commercial decision.
From data to decision
Everything above exists to support four decisions. If your failure tracking isn't feeding at least one of them, it's reporting rather than analysis.
Batch data, fault codes and serials, presented as a quality finding rather than a request.
Wear-out failures and single dominant fault codes point at component grade or design.
When warranty cost per unit exceeds the margin, the product is a liability at that price.
No-fault-found volume is a listing, instructions or onboarding problem you can solve cheaply.
This can feel like a project that only pays off eventually. It isn't. Make the fault code mandatory, record the batch or supplier reference against each claim, and capture the date the customer first noticed the fault. One quarter of that is enough to produce a Pareto chart and a batch comparison — which is enough for a supplier conversation that recovers real money.
If your products contain lithium cells, the inbound leg of every one of these claims is a dangerous-goods shipment with its own documentation and retention requirements — see dangerous goods returns. And the workflow that captures the fault code in the first place is covered in warranty management.
Frequently asked questions
What counts as a good failure rate?
There's no universal figure, and any benchmark quoted without a product category attached is meaningless. Simple accessories with no moving parts might sit under 0.5%; power tools, appliances and anything with a battery and a motor commonly run 2–5% over a full warranty period. The number that matters is your own trend and your own batch variance. A 4% rate that's stable and fully recoverable from a supplier is a healthier position than a 2% rate that's climbing and unattributed.
How do I set a baseline when I've only just started tracking?
Use the first two or three mature cohorts as a provisional baseline and treat it as provisional. You'll have a usable reference within a quarter and a reliable one within two. In the meantime, batch-to-batch comparison is available immediately and doesn't need a baseline at all — if one batch fails at three times the rate of the batch either side of it, that's actionable on its own, whatever the absolute numbers turn out to be.
Should no-fault-found claims count in the failure rate?
Track them, but report them separately. They aren't product failures, so including them inflates your quality numbers and weakens any supplier claim built on them. They're also genuinely useful on their own: a rising no-fault-found rate usually means a listing is overselling, the instructions are unclear, or setup is harder than customers expect. All three are cheaper to fix than a manufacturing defect, and all three are invisible if the number is buried inside a single failure-rate figure.
Can I do this in a spreadsheet?
You can, and starting there is better than not starting. What you need per claim is a SKU, a batch or supplier reference, a fault code, the date of sale, the date the fault was noticed and the date the claim was raised. With those six fields, cohort rates and Pareto charts are pivot-table work. The point at which it stops scaling is usually the batch reference — capturing it reliably at claim time is hard without it being a required field in the workflow, and a spreadsheet can't enforce that.
How far back should I hold failure data?
Longer than feels necessary. Wear-out patterns only become visible across two to three years, and product-generation comparisons need history on the generation you've already discontinued. If you sell dangerous goods, the retention requirements attached to transport documentation will usually be the binding constraint anyway, and they run well beyond the life of the claim. Plan for years, not months.
My supplier disputes my numbers. What now?
Almost always a granularity problem rather than a disagreement about facts. Aggregate rates are easy to dispute because they mix batches, fault types and customer-caused damage. A claim scoped to one batch, one fault code and a list of serial numbers, with the failure rate stated against units shipped in that batch, is verifiable against the supplier's own production records — which means the cheapest thing they can do is check it and issue the credit. Narrow the claim until it's easier to verify than to argue with.