ReturnMate logoReturnMate
Quality

Product Failure Rate Tracking

Almost every merchant who tells us their failure rate is quoting a number that understates it by a factor of two or three. Not because anyone is being dishonest — because the obvious way to calculate it is the wrong way, and every dashboard that reports "returns as a percentage of sales" reproduces the same error. This is how to measure failure properly, and what to do with the answer.

10 min readUpdated 22 Aug 2026Quality analytics
Two ways to measure the same product — 18V brushless drill, same underlying data
Calendar methodCohort method
Calendar month — what most dashboards show
34 claims received in August1,890 units sold in August=1.8%

The numerator and denominator describe different populations. The units failing in August were mostly sold in autumn last year; the units sold in August have barely been switched on. This ratio is not a failure rate — it's a coincidence of two unrelated monthly totals.

Cohort lifetime — what's actually true
116 failures to date, Jan–Mar cohort2,140 units sold in that cohort=5.4%

One population, followed over time. Every unit in the denominator has had the same opportunity to fail, and every failure counted belongs to a unit in that denominator. This number is comparable between batches, suppliers and product generations.

Same product, same claims, same period. The calendar method understates this SKU's failure rate by 3× — and it's the number nearly every returns dashboard reports by default.
Illustrative figures based on a typical seasonal-demand power tool. The gap widens whenever sales are growing or seasonal, because a rising denominator dilutes failures that came from a smaller cohort months earlier. For fast-growing merchants the calendar method can understate failure by 4–5×.
§ 01

The denominator problem

Failure rate looks like the simplest metric in the business. Failures divided by units. The difficulty is entirely in the second term, and getting it wrong is the default rather than the exception.

If you divide this month's claims by this month's sales, you are comparing a group of units that have been in the field for six to eighteen months against a group that shipped last week. The first group has had every opportunity to fail. The second has had almost none. The ratio between them means nothing, and it is wrong in a specific, predictable direction: it always understates failure, and it understates it most when you're growing.

The fix is to hold the population fixed. Pick a cohort — a month of sales, a production batch, a shipment — and follow it. Count every failure that cohort produces, and divide by the size of that cohort. Now the numerator and denominator describe the same units, and the number becomes something you can compare, trend and take to a supplier.

Why this matters commercially, not just statistically

A supplier who is shown a calendar-month rate has an easy out: your denominator includes units that couldn't possibly have failed yet, so the figure is not evidence of anything. A cohort rate closes that argument. It's also the only form in which failure data is comparable across your own product generations — which is how you find out whether the "improved" version you switched to last year is actually better.

One caveat worth being honest about: cohorts need maturity. A cohort that's four weeks old will show a low rate because most of its failures haven't happened yet. Compare cohorts at the same age — three-month rate against three-month rate, twelve against twelve — or the comparison reintroduces the same bias by a different route.

§ 02

What to actually measure

Four numbers between them tell you almost everything worth knowing. Each answers a different question, and each drives a different decision.

M1Cohort failure rate

Failures per cohort, measured at a fixed age. The headline number. Tells you whether a product line is acceptable and whether it is getting better or worse over time.

M2Time to failure

How long units survive before failing. Separates manufacturing defects from wear-out, which are two entirely different problems with different owners.

M3Fault distribution

Which failure modes account for the volume. Almost always concentrated — a handful of codes drive most of the cost, and they're the only ones worth engineering effort.

M4Batch variance

Spread between batches of the same SKU. High variance points at supplier process control rather than product design, and it's the strongest evidence for a credit claim.

Notice what isn't on this list: total return count, return rate across the whole catalogue, and anything expressed as a single business-wide percentage. Aggregate numbers are for board slides. They can't tell you which SKU to renegotiate, which supplier to escalate to, or which fault to engineer out, because the signal you need lives at the level of a product and a batch.

§ 03

When things fail, not just how often

Two SKUs can have identical failure rates and require completely different responses. What separates them is when in a unit's life the failure arrives. Plotting failures against months since purchase makes the distinction obvious, and it's the single most useful chart in quality analytics.

Failures by months since purchase — 18V brushless drill, Jan–Mar cohortn = 116 failures · illustrative
Early life, months 0–3 — a third of failures

Manufacturing and assembly defects. These units left the factory faulty. This is the supplier's cost, and the batch data to prove it is sitting in your fault codes.

Wear-out, months 16+ — rising steadily

Design life or duty cycle, not defects. The answer here is a spec change, a different component grade, or honest positioning — not a warranty claim against your supplier.

The spike just before month twelve is not a product behaviour. It's a customer behaviour: people file claims they've been putting off once they notice the warranty is about to expire. It's worth knowing about, because it means a chunk of what looks like month-eleven failure actually occurred earlier, and it will distort your early-life numbers if you don't capture the date the fault was first noticed as well as the date the claim was raised.

§ 04

Which faults are worth your attention

Failure modes are never evenly distributed. Rank fault codes by volume and the curve is almost always steep — which is good news, because it means a small amount of engineering or supplier attention addresses most of the cost.

Fault code distribution — 18V brushless drill, Jan–Mar cohortn = 116 · illustrative
F-012Motor / BMS fault31%31%
F-034Chuck assembly22%53%
F-001Won't power on18%71%
F-051Firmware / electronics11%82%
80% of failures
F-021Switch / trigger7%89%
F-099No fault found6%95%
All other codes5%100%
Three codes account for 71% of failures. That's the entire agenda for your next supplier review — and it's a far more productive conversation than presenting a total claim count. Note F-099: six per cent of units came back with nothing wrong, which is a returns-experience or product-instructions problem, not a quality one.

Two things make this chart possible, and neither is analytics software. The fault code has to be mandatory at diagnosis, and the code list has to be short enough that technicians use it consistently. Thirty codes will be applied inconsistently by three technicians and the distribution becomes noise. Eight to twelve codes, clearly distinguishable, is the range where this stays reliable.

§ 05

Thresholds that trigger action

Failure data that produces a monthly report changes nothing. Failure data attached to a threshold changes behaviour, because it converts a number into an obligation. The thresholds below are a starting point — the levels matter less than the fact that crossing one triggers something specific.

SignalThresholdWhat it triggers
Cohort rate vs SKU baseline1.5× baselineFlag for review at next supplier call; check whether it's one batch or all
Cohort rate vs SKU baseline3× baselineHold remaining stock, pull batch numbers, open a formal claim
Single batch vs sibling batches2× spreadProcess-control issue at the supplier — escalate to their QA, not sales
Early-life share of failuresAbove 40%Manufacturing defect pattern; request incoming inspection data
One fault codeAbove 30% of claimsEngineering or component review — this is a design problem, not bad luck
No-fault-found rateAbove 8%Product instructions, listing accuracy or setup experience needs work
Warranty cost per unit soldAbove target margin %Pricing, spec or supplier decision — the SKU is not viable as sold

The last row is the one that most often gets skipped and most often matters. A SKU can have an unremarkable failure rate and still be unprofitable once you load in freight both ways, bench labour, parts and the share of claims where no supplier credit is recoverable. Tracking warranty cost per unit sold, rather than failure rate alone, is what turns quality data into a commercial decision.

§ 06

From data to decision

Everything above exists to support four decisions. If your failure tracking isn't feeding at least one of them, it's reporting rather than analysis.

DECISION 01
Claim from the supplier

Batch data, fault codes and serials, presented as a quality finding rather than a request.

DECISION 02
Change the specification

Wear-out failures and single dominant fault codes point at component grade or design.

DECISION 03
Reprice or delist

When warranty cost per unit exceeds the margin, the product is a liability at that price.

DECISION 04
Fix the experience

No-fault-found volume is a listing, instructions or onboarding problem you can solve cheaply.

You need one quarter of clean data, not a year

This can feel like a project that only pays off eventually. It isn't. Make the fault code mandatory, record the batch or supplier reference against each claim, and capture the date the customer first noticed the fault. One quarter of that is enough to produce a Pareto chart and a batch comparison — which is enough for a supplier conversation that recovers real money.

If your products contain lithium cells, the inbound leg of every one of these claims is a dangerous-goods shipment with its own documentation and retention requirements — see dangerous goods returns. And the workflow that captures the fault code in the first place is covered in warranty management.

§ 07

Frequently asked questions

What counts as a good failure rate?

There's no universal figure, and any benchmark quoted without a product category attached is meaningless. Simple accessories with no moving parts might sit under 0.5%; power tools, appliances and anything with a battery and a motor commonly run 2–5% over a full warranty period. The number that matters is your own trend and your own batch variance. A 4% rate that's stable and fully recoverable from a supplier is a healthier position than a 2% rate that's climbing and unattributed.

How do I set a baseline when I've only just started tracking?

Use the first two or three mature cohorts as a provisional baseline and treat it as provisional. You'll have a usable reference within a quarter and a reliable one within two. In the meantime, batch-to-batch comparison is available immediately and doesn't need a baseline at all — if one batch fails at three times the rate of the batch either side of it, that's actionable on its own, whatever the absolute numbers turn out to be.

Should no-fault-found claims count in the failure rate?

Track them, but report them separately. They aren't product failures, so including them inflates your quality numbers and weakens any supplier claim built on them. They're also genuinely useful on their own: a rising no-fault-found rate usually means a listing is overselling, the instructions are unclear, or setup is harder than customers expect. All three are cheaper to fix than a manufacturing defect, and all three are invisible if the number is buried inside a single failure-rate figure.

Can I do this in a spreadsheet?

You can, and starting there is better than not starting. What you need per claim is a SKU, a batch or supplier reference, a fault code, the date of sale, the date the fault was noticed and the date the claim was raised. With those six fields, cohort rates and Pareto charts are pivot-table work. The point at which it stops scaling is usually the batch reference — capturing it reliably at claim time is hard without it being a required field in the workflow, and a spreadsheet can't enforce that.

How far back should I hold failure data?

Longer than feels necessary. Wear-out patterns only become visible across two to three years, and product-generation comparisons need history on the generation you've already discontinued. If you sell dangerous goods, the retention requirements attached to transport documentation will usually be the binding constraint anyway, and they run well beyond the life of the claim. Plan for years, not months.

My supplier disputes my numbers. What now?

Almost always a granularity problem rather than a disagreement about facts. Aggregate rates are easy to dispute because they mix batches, fault types and customer-caused damage. A claim scoped to one batch, one fault code and a list of serial numbers, with the failure rate stated against units shipped in that batch, is verifiable against the supplier's own production records — which means the cheapest thing they can do is check it and issue the credit. Narrow the claim until it's easier to verify than to argue with.