The Bathtub Curve That Was Actually a Cliff: Weibull Analysis of Field Returns and the Burn-In Decision

Every reliability deck has the same slide: the bathtub curve. Infant mortality on the left, a long flat useful life in the middle, wear-out on the right. It is a comforting shape. It implies that if you burn in your product long enough, you scrape off the early failures and ship the flat part.

Then the returns come in, and the curve is not a bathtub. It is a cliff. Units work for six weeks, then a cohort fails within a narrow window. A burn-in screen would not have caught them, because at burn-in time they were fine. The failure mechanism was not latent in the component. It was latent in the assembly process, the firmware state machine, or the connector interface, and it needed field conditions to express itself.

This article is about how to tell the difference using Weibull analysis of field returns, and how to make a defensible burn-in decision at 1k–10k unit volumes, where you have enough returns to see a shape but not enough to be statistically luxurious about it.

What the Weibull parameters actually tell you

The two-parameter Weibull distribution describes time-to-failure with a shape parameter β (beta) and a scale parameter η (eta). Beta is the one that matters for the burn-in decision.

  • β < 1: decreasing failure rate. Classic infant mortality. Failures cluster early and the survivors get better. This is the regime a burn-in screen is designed to address.
  • β ≈ 1: constant failure rate. Random failures, no memory. Burn-in does not help; you are just removing hours from the useful life.
  • β > 1: increasing failure rate. Wear-out or cumulative damage. Failures accelerate with time or cycles. Burn-in does not help either, and can make it worse by consuming life.

A cliff in field returns typically shows up as a high beta (often 2–5) with a scale parameter that corresponds to a specific field exposure: a number of thermal cycles, a number of power cycles, a number of connector mating events, or a calendar time at a particular humidity. The shape is not telling you the units are wearing out in the classical sense. It is telling you that a damage mechanism is accumulating and crossing a threshold.

The practical consequence: if your field return Weibull plot shows β > 1, a burn-in screen is the wrong tool. You need to find the accumulating damage mechanism and either eliminate it or screen for the precursor, not for the failure.

Sample size at 1k–10k units: what you can and cannot conclude

At 1k–10k shipped units, a 1% return rate gives you 10–100 returns. That is enough to fit a Weibull if the failures are not heavily censored and if you know the exposure time for each unit. It is not enough to resolve a bimodal distribution into two clean populations without strong assumptions.

Three things make the analysis usable at this volume:

  1. You need the exposure time for every unit, not just the failed ones. If you only plot failures, you are fitting a distribution to a biased sample. Units still in the field are right-censored observations and they carry information.
  2. You need to stratify by manufacturing lot, date code, and field deployment cohort. A cliff is often a lot-specific or date-code-specific event. Pooling lots can smear a cliff into something that looks like a gentle slope.
  3. You need to be honest about the confidence interval on beta. With 20 failures, the confidence interval on beta is wide. A point estimate of β = 2.5 might have a lower bound below 1. That does not mean the cliff is not real; it means you should not over-claim the shape from a single fit.

If you have fewer than about 10 failures in a stratum, treat the Weibull fit as a hypothesis, not a conclusion. Use it to decide what to measure next, not to justify a burn-in recipe.

Why a burn-in screen would not have caught the cliff

Burn-in works by accelerating mechanisms that are already present at time zero. It is effective against:

  • Latent semiconductor defects that are activated by temperature and voltage.
  • Contamination and particle defects that produce early opens or shorts.
  • Weak solder joints that fail under thermal cycling.
  • Component lots with a subpopulation that is marginal from the start.

It is not effective against mechanisms that require field-specific conditions to initiate or accumulate:

  • Firmware watchdog behavior. A watchdog that resets the MCU under a specific combination of brownout and I/O state will not show up in a burn-in rack that holds the unit at nominal voltage. The cliff appears when field units see a particular power-sequencing pattern.
  • Connector contamination. A connector that passes continuity test at the factory can develop intermittent contact after a number of mating cycles or after exposure to a particular humidity profile. Burn-in does not mate the connector.
  • Solder voiding under thermal pads. A void that is electrically invisible at room temperature can become a thermal problem only when the unit is under sustained load in a warm enclosure. Burn-in at low duty cycle will not reproduce it.
  • Electrochemical migration. Requires humidity and bias. A dry burn-in rack will not trigger it.

In each of these cases, the failure is not latent in the component at time zero. It is latent in the interaction between the assembly, the firmware, and the field environment. Burn-in tests the component. It does not test the interaction.

The manufacturing decisions that produce a cliff

Three decisions show up repeatedly in field returns that look like cliffs rather than bathtubs.

1. Solder voiding under thermal pads

Voiding under a thermal pad is not a defect that shows up in a functional test. It shows up as a thermal resistance that is higher than the design assumed. At low duty cycle, the unit is fine. At sustained load, the junction temperature rises, the void grows or the thermal interface degrades, and the unit fails. The failure time depends on the field duty cycle, not on calendar time. That produces a high-beta Weibull when plotted against calendar time, and a much lower beta when plotted against cumulative energy or cumulative thermal cycles.

The fix is not burn-in. The fix is process control on the reflow profile and stencil design, plus a thermal measurement on a sample of units under load. If you cannot measure junction temperature in the field, measure it on the bench at the worst-case duty cycle and compare to the design margin.

2. Connector contamination and mating-cycle damage

A connector that passes continuity test at the factory can fail after a number of mating cycles or after exposure to a particular humidity profile. The failure is intermittent, which makes it hard to catch in a functional test. The field return rate looks like a cliff because the damage accumulates with mating cycles, and mating cycles correlate with deployment time.

The fix is to define a mating-cycle specification, test to it, and control the connector interface in manufacturing. If the connector is not sealed, define the humidity exposure and test it. Burn-in does not exercise the connector.

3. Firmware watchdog behavior

A watchdog that resets the MCU under a specific combination of brownout and I/O state will not show up in a burn-in rack that holds the unit at nominal voltage. The cliff appears when field units see a particular power-sequencing pattern. The failure is not a component failure; it is a firmware state-machine failure that requires a specific input sequence.

The fix is to reproduce the field power sequence on the bench, not to burn in the unit. If you cannot reproduce it, instrument the field units to capture the state at reset, and use that to drive the bench test.

What a burn-in decision should actually be based on

A burn-in screen is justified when three conditions hold:

  1. The field return Weibull shows β < 1 in the early-life region, after stratifying by lot and date code.
  2. The failure mechanism is one that burn-in accelerates: temperature, voltage, or both.
  3. The cost of burn-in (capital, floor space, cycle time, yield loss) is less than the cost of the returns it prevents.

If any of those conditions fails, burn-in is not the right response. If β > 1, burn-in is actively harmful because it consumes useful life without removing the failure mechanism.

At 1k–10k unit volumes, the burn-in decision is often made on a small number of returns. The temptation is to burn in everything because it feels safer. The evidence-first approach is to plot the returns, stratify by lot, and look at the shape. If the shape is a cliff, the burn-in rack is not where the problem is.

How to falsify the burn-in hypothesis

If someone proposes a burn-in screen to address a field failure, the falsification test is straightforward:

  • Take units from the same lot that produced the field failures.
  • Run them through the proposed burn-in profile.
  • Then run them through the field-equivalent stress that produced the cliff.
  • Compare the failure time to units that did not receive burn-in.

If burn-in does not shift the failure time, it is not addressing the mechanism. If it shifts the failure time earlier (because it consumed life), it is making the problem worse. If it shifts the failure time later, you have evidence that the mechanism is latent at time zero and burn-in is removing it.

This test is cheap relative to a full burn-in deployment. It should be run before committing to a burn-in recipe.

FAQ

How many field returns do I need to fit a Weibull?

You can fit a Weibull with as few as 5–10 failures if you have the exposure time for all units, but the confidence interval on beta will be wide. At 1k–10k shipped units, a 1% return rate gives you 10–100 returns, which is enough to see the shape but not enough to resolve a bimodal distribution without strong assumptions. Treat fits with fewer than about 10 failures per stratum as hypotheses.

What burn-in duration and temperature should I use?

That depends on the failure mechanism and the component class. JEDEC JESD47 and JESD22-A104 define test methods for qualification and burn-in, but they do not prescribe a single recipe for all products. The recipe should be derived from the acceleration factor for the specific mechanism and the field exposure you are trying to screen. If you do not know the mechanism, you do not know the acceleration factor, and the burn-in recipe is a guess.

Can burn-in make field failures worse?

Yes. If the failure mechanism is wear-out (β > 1), burn-in consumes useful life and can shift the cliff earlier. If the mechanism is not accelerated by temperature or voltage, burn-in does not remove it and the cycle time and yield loss are pure cost.

What if my field returns show a cliff but I cannot reproduce it on the bench?

Instrument the field units to capture the state at failure, then use that data to drive the bench test. The cliff is usually a specific sequence or exposure that the bench test is not reproducing. Burn-in will not help until you can reproduce the mechanism.

Sources and further reading

The Weibull distribution and its parameters are standard reliability engineering material; the two-parameter form and the interpretation of beta are covered in any reliability textbook and in the NIST/SEMATECH Engineering Statistics Handbook. JEDEC JESD47 covers qualification and burn-in test methods for solid-state devices, and JESD22-A104 covers temperature cycling. IPC-9592A covers requirements for power conversion devices and includes reliability and burn-in provisions. These documents define methods; they do not tell you which method applies to your specific field failure. That determination comes from your own return data and bench measurements.

If you are making a burn-in decision at 1k–10k unit volumes, the sequence is: plot the returns, stratify by lot, look at the shape, and only then decide whether burn-in addresses the mechanism. The bathtub curve is a useful default. It is not a diagnosis.