The Golden Board That Disagreed With the Fixture: Test Limit Correlation, Guardbanding, and GR&R at the CM

Your prototype passed. Your golden board passed. The fixture says the golden board failed. That is not a board problem. That is a measurement-system problem, and at 1k–10k unit volumes it is the kind of problem that eats a week of engineering time and produces a shipment hold that nobody can explain to the customer.

This article is about the specific failure mode where a production test fixture and a bench reference disagree about the same physical unit. It covers what GR&R actually measures, how guardbanding interacts with test limits, and what to demand from a contract manufacturer (CM) when the golden board and the fixture stop agreeing.

What GR&R is, and what it is not

GR&R stands for Gauge Repeatability and Reproducibility. It is a measurement-systems analysis (MSA) technique. The AIAG Measurement Systems Analysis reference manual is the common automotive-industry source for the method and its acceptance criteria; the manual is distributed by AIAG and is not freely available online, so treat any specific percentage thresholds you see quoted as coming from that manual rather than from a public web page.

The two components:

  • Repeatability — variation when one operator measures the same part with the same gauge multiple times. This is the equipment’s own noise.
  • Reproducibility — variation when different operators (or different setups, or different fixtures) measure the same part. This is the appraiser/setup contribution.

GR&R is usually expressed as a percentage of total variation or of the tolerance band. The classic AIAG guidance treats under 10% as generally acceptable, 10–30% as conditionally acceptable depending on the application and the cost of the measurement, and over 30% as unacceptable. Those bands are guidance, not physics. They were written with continuous measurements in mind — micrometers, calipers, torque gauges — not pass/fail functional testers.

That distinction matters. A pass/fail fixture does not produce a continuous reading you can feed into a standard GR&R ANOVA. It produces a binary outcome. The right way to characterize a binary measurement system is an attribute agreement study: multiple known-good and known-bad units, multiple operators, multiple trials, and a look at the agreement rate and the misclassification rate. If your CM hands you a GR&R number computed from a pass/fail fixture as if it were a continuous gauge, ask what they actually measured.

Guardbanding: the part nobody puts in the test spec

Guardbanding is the practice of tightening the test limits relative to the product specification so that measurement uncertainty does not let a marginal unit pass. If the product spec says a rail must be 3.30 V ± 3%, and your measurement system has 20 mV of uncertainty, you do not test at 3.201 V and 3.399 V. You test inside those limits by some amount that reflects the uncertainty.

The simple form is:

guardbanded limit = spec limit − k × U

where U is the expanded measurement uncertainty and k is a coverage factor. For a 95% coverage assumption with a normal distribution, k = 2 is common. The exact value depends on the uncertainty budget and the risk you are willing to accept.

Two things go wrong here at low volume:

  1. No uncertainty budget exists. The fixture was built, the limits were copied from the bench script, and nobody wrote down what the fixture’s measurement uncertainty actually is. Without U, guardbanding is a guess.
  2. The guardband is applied asymmetrically or not at all. One test gets tightened, another does not, and the fixture’s pass/fail behavior becomes inconsistent with the bench in ways that only show up on marginal units.

Guardbanding is not free. Every millivolt you tighten is a unit you might reject that would have worked in the field. At 1k–10k units, a 1% false-reject rate is 10–100 units. That is real money and real schedule. The point of an uncertainty budget is to make that trade explicit rather than accidental.

The golden board is a correlation artifact, not a truth source

A golden board (or golden unit) is a physical reference used to check that a test system is behaving the same way over time. It is not a calibrated standard. It is a board that was measured on a reference system and whose behavior is treated as known for correlation purposes.

When the golden board disagrees with the fixture, the useful question is not “which one is right?” It is “what changed?” The candidates, roughly in order of how often they turn out to be the answer:

  • Fixture contact resistance or probe wear. Pogo pins wear. Contact resistance rises. A rail that measured 3.28 V on the bench measures 3.24 V in the fixture because the fixture is dropping voltage in the measurement path.
  • Grounding and return-path differences. The bench used a short ground lead. The fixture uses a long harness. The measurement includes different common-mode and IR-drop contributions.
  • Temperature. The bench was at 22 °C. The fixture is in a room that swings 18–28 °C. A temperature coefficient of 100 ppm/°C on a 3.3 V rail is 330 µV/°C, which is small — but on a current-sense resistor or a bandgap reference it can be enough to move a marginal unit across a limit.
  • Instrument differences. The bench used a 6.5-digit DMM. The fixture uses a 5.5-digit module in a PXI chassis. Both are fine; they are not the same measurement.
  • Software and sequencing. The fixture runs tests in a different order, or with different settling times, or with a different load applied during the measurement.
  • The golden board itself drifted. It happens. A golden board that lives in a drawer and gets handled is not a stable reference.

None of these are exotic. All of them are invisible if the only correlation check is “does the golden board pass?”

What a correlation procedure actually needs

If you are specifying a fixture correlation procedure for a CM, the useful minimum is:

  1. A written reference measurement. The golden board’s expected values, with the instrument, settings, temperature, and date recorded. Not “it passed on the bench.” Actual numbers.
  2. A defined correlation tolerance. How much disagreement between bench and fixture is acceptable before the fixture is considered out of correlation? This should be derived from the measurement uncertainty, not picked because 5% sounded reasonable.
  3. A re-correlation interval. Time-based, event-based (after fixture repair, after probe replacement, after a software change), or both. The interval should be short enough that drift is caught before it affects shipments.
  4. A drift record. The golden board’s measured values over time. If the fixture reading on the golden board is walking in one direction, that is a trend, and trends are actionable before they become failures.
  5. A defined response when correlation fails. Stop, investigate, re-correlate, and document. Not “adjust the limit until the golden board passes.”

The last point is the one that gets violated most often. Adjusting the limit to make the golden board pass is not correlation. It is calibration by wishful thinking.

Why this is worse at 1k–10k units

At high volume, the economics of test engineering are different. You can afford a dedicated test engineer, a proper fixture, a full uncertainty budget, and a statistical process control (SPC) system that watches the test data. At 1k–10k units, the fixture is often built by the CM from a test spec you wrote, using instruments they already have, and the correlation check is whatever the CM’s process says it is.

That is not a criticism of CMs. It is a statement about where the engineering effort lives. At low volume, the CM’s test process is often a template applied to your product. The template may be perfectly reasonable for a different product with different tolerances and a different measurement uncertainty. It may not be reasonable for yours.

The practical consequence: at 1k–10k units, you are more likely to discover a correlation problem through returns than through test data. A unit that passed the fixture and failed in the field is a correlation escape. A unit that failed the fixture and passed on the bench is a false reject. Both are symptoms of the same underlying issue: the fixture and the bench do not agree about what “pass” means.

What to ask the CM

These are the questions that surface the problem before it surfaces in returns:

  • What is the measurement uncertainty of the fixture for each critical measurement? If the answer is “we don’t have that,” that is the answer.
  • How were the test limits derived? From the product spec, from the bench script, or from a guardband calculation?
  • What is the correlation procedure between the bench reference and the fixture? How often is it run? What is the acceptance criterion?
  • What happens when the golden board fails correlation? Is there a documented stop-and-investigate step, or does someone adjust a limit?
  • For pass/fail tests, what attribute agreement study has been done? How many known-good and known-bad units, how many operators, how many trials?
  • Is the fixture’s test data logged per unit? Can you see the actual measured values, not just pass/fail?

The last question is the one that pays off fastest. If the fixture logs actual values, you can plot them, watch for drift, and compare distributions between the bench and the fixture. If it only logs pass/fail, you are blind to everything except the outcome.

The uncomfortable part

Guardbanding and correlation are not glamorous. They do not show up in a demo. They are the difference between a prototype that works and a product that ships without a returns problem. At 1k–10k units, the volume is high enough that a systematic correlation error produces a measurable return rate, and low enough that nobody has built the infrastructure to catch it automatically.

The golden board disagreeing with the fixture is not a crisis. It is a signal. The crisis is when nobody notices, or when the response is to adjust the limit until the disagreement goes away. The useful response is to treat the disagreement as data: measure the uncertainty, write down the correlation procedure, and make the guardband an explicit decision rather than an accident.

FAQ

Is GR&R the right tool for a pass/fail fixture?
Not in its standard continuous-measurement form. Use an attribute agreement study instead. GR&R assumes a continuous measurement you can decompose into variance components. A binary outcome does not support that decomposition in the same way.

How much guardband is enough?
Enough to cover the measurement uncertainty at the coverage factor you have chosen. If you do not have an uncertainty budget, you do not have a defensible guardband. Start there.

How often should the golden board be re-correlated?
There is no universal interval. It depends on the stability of the fixture, the handling of the golden board, and the consequence of drift. A common starting point is at every fixture setup change, after any repair or probe replacement, and on a fixed calendar interval that is short enough to catch drift before it affects a shipment. The interval should be justified by data, not by convenience.

What if the CM says the fixture is fine and the golden board is bad?
That is possible. Golden boards drift, get damaged, or get replaced with a unit that was never properly characterized. The way to resolve it is to measure both against a third reference — a calibrated instrument, not another fixture — and compare. If you cannot do that, you cannot resolve the disagreement, and you should treat the fixture’s results as uncertain until you can.

Does this matter at 1k units?
Yes. At 1k units, a 1% correlation escape rate is 10 units. Whether that is acceptable depends on what the unit does and what a field failure costs. The point is to make the number visible rather than to assume it is zero.