Gage R&R from first principles: how much of what you see is the part?

2026-08-24 · long read · no math background needed

Sooner or later a customer asks for it: a gage R&R, the spreadsheet with the shaded cells, constants like 0.8862 that nobody in the room can explain. Most labs run the study the way they run a fire drill — because the audit requires it. This note rebuilds the whole thing from one idea — some of what you see is the part; the rest is the measurement — and walks a complete gage R&R study, small enough to check with a pencil, so that every constant and every percentage has a reason attached to it.

Contents

  1. The question a gage R&R actually answers
  2. Repeatability and reproducibility — named by what changed
  3. The experiment: two operators, five parts, two trials
  4. Repeatability, from the ranges
  5. Reproducibility, from the operator averages
  6. The part variation and the total
  7. The verdict: %GRR, ndc, and what the thresholds really are
  8. Percent of what? Tolerance and variation answer different questions
  9. ANOVA: what modern software does instead
  10. When the gage is a scanner and its software
  11. A passed study is a snapshot — the short version

1 · The question a gage R&R actually answers

Measure ten parts from a process and the numbers scatter. Two things put that scatter there: the parts genuinely differ (part-to-part variation), and the act of measuring adds noise of its own — the instrument, the operator's touch, the fixturing, the temperature, the software. You never see a part's value directly; you see it through a measurement system, and the system's fingerprints are on every reading.

A gage R&R study is an experiment designed to pull those two contributions apart: how much of the variation I observe is the parts, and how much is the measuring? That's the entire question; the K-constants, percentages and thresholds are bookkeeping in service of it. (Terminology: the AIAG manuals spell it gage, much of the world writes gauge, and GRR is the same study either way. R&R is one tool inside measurement system analysis, MSA, which also covers bias, linearity and stability — more on those at the end.)

The study leans on one mathematical fact, and it is the only one you need all day:

Equation 1Independent sources of scatter add like the sides of a right triangle
TV ²  =  PV ² + GRR ²
TVtotal variation: the typical spread of the readings you actually observe. All spreads here are measured as standard deviations — the usual "typical distance from the average" number.
PVpart variation: the spread the parts alone would show, measured by a perfect instrument.
GRRthe measurement system's own spread — the thing the study exists to estimate.

In words: spreads do not add like lengths; they add like the perpendicular sides of a right triangle (Pythagoras, wearing a lab coat). The consequence is lopsided: a measurement system a third the size of the part variation inflates what you see by only about 5%, because √3² + 1² = 3.16, barely more than 3. Small noise hides — which is exactly why finding it takes a designed experiment.

2 · Repeatability and reproducibility — named by what changed

A gage R&R study splits measurement system variation into two named pieces, and the names are best understood by asking what changed between two readings.

Repeatability is the scatter you get when nothing changed: same part, same operator, same instrument, same setup, readings minutes apart. Whatever spread remains is the measurement system arguing with itself — the instrument's noise floor plus the tiny inconsistencies of a single person's technique. The AIAG manual calls this equipment variation, EV.

Reproducibility is the extra scatter that appears when the thing that changed is who — a different operator, with their own touch, their own reading of the vernier, their own habit of seating the part. It is measured from how far apart the averages land when different operators measure the same parts, and the manual calls it appraiser variation, AV.

The definitions are operational, not philosophical. If your "operators" are three shifts, or three fixtures, or three software templates, the same arithmetic measures reproducibility across whatever you varied. That flexibility is worth remembering — it is what will let the method survive contact with automated scanning in section 10.

3 · The experiment: two operators, five parts, two trials

Here is a complete gage R&R study you can check by hand — a deliberately small teaching example, its numbers constructed to be easy to follow rather than taken from anywhere: two operators, five parts, two trials each, 20 readings. (Real studies more often use ten parts, three operators, three trials; the machinery is identical, only the constants change.)

The setting: a turned shaft, drawing calls the diameter 25.00 ± 0.10 mm, measured with a micrometer reading to 0.01 mm. The five parts were pulled to span the process — one runs near the top of the tolerance and one just beyond it, which is deliberate on both counts: a gage R&R study spans the process, not the print, and it needs genuine part-to-part variation, or there is nothing for the gage to distinguish. Both operators measure every part twice, in randomized order, blind to their own earlier reading and to the other operator's. The blindness is not ceremony; an operator who remembers writing 25.08 will find 25.08 again, and the study will flatter the gage.

partOperator A (mm)Operator B (mm)
trial 1trial 2meanrangetrial 1trial 2meanrange
124.9424.9524.9450.0124.9624.9424.9500.02
225.0025.0025.0000.0025.0225.0125.0150.01
325.0425.0225.0300.0225.0325.0325.0300.00
425.0825.0925.0850.0125.1025.0825.0900.02
525.1225.1125.1150.0125.1325.1325.1300.00
operator mean 25.03525.043
mean range 0.0100.010
24.90 (low limit) 25.10 (high limit) P1P2P3P4P5 24.9525.0025.0525.15 op A op B
All twenty readings, one row per part. The distances between the clusters are the parts (PV); the width of each cluster is the measurement system (EV and AV together). The whole study is just an honest accounting of "clusters far apart, clusters tight."

Look at the picture before any arithmetic. The part clusters are clearly separated — good. Each cluster is a couple of hundredths wide — that width is the measurement system. And operator B's rings sit very slightly right of operator A's dots on most parts — that small systematic offset is reproducibility, about to be caught by the math.

4 · Repeatability, from the ranges

Each row of the table contains a pair of readings where nothing changed — same part, same operator. The difference within such a pair can only be measurement error. The range — the bigger reading minus the smaller — captures it: 0.01, 0.00, 0.02, 0.01, 0.01 for operator A; 0.02, 0.01, 0.00, 0.02, 0.00 for operator B. Average all ten and you get the mean range,  = 0.010 mm.

A range is not yet a standard deviation, but for small samples the two are tied together by a known conversion factor. Statisticians tabulate the factor (they call it d2*); the AIAG MSA manual publishes its reciprocals as the K-constants on the study form, derived from those d2* tables. For pairs — two trials — the constant is K1 = 0.8862.

Equation 2Repeatability (equipment variation)
EV  =  × K1  =  0.010 × 0.8862  =  0.008862 mm
("R-bar") the average of all ten within-pair ranges — the typical disagreement of a reading with its own repeat.
K1the range-to-standard-deviation conversion factor: 0.8862 when each range comes from 2 trials (0.5908 for 3 trials), from the d2*-based tables in the AIAG MSA manual.
EVequipment variation — repeatability, expressed as a standard deviation: about 0.0089 mm, or 8.9 µm.

In words: collect all the "nothing changed" disagreements, average them, and rescale the average into standard-deviation units. The result says: hand this micrometer and this technique the same shaft twice, and the readings typically wander by about nine microns' worth of standard deviation. That is the noise floor everything else sits on.

5 · Reproducibility, from the operator averages

Now the second piece. Operator A's grand average over all ten readings is 25.035 mm; operator B's is 25.043 mm. Two different operators measured the same five parts, so in a perfect world the averages would agree. The difference, DIFF = 25.043 − 25.035 = 0.008 mm, is the between-operator signal.

But there's a subtlety, and it is the one genuinely clever step in the whole worksheet. Each operator's average was built from ten noisy readings — so even two clones of the same operator would show some difference, purely from equipment noise. If we converted DIFF straight into a standard deviation, we would count repeatability twice. So the form subtracts the share of the difference that repeatability alone would have produced:

Equation 3Reproducibility (appraiser variation), with the double-count removed
AV  =  √(DIFF × K2)² − EV ²n × r
DIFFthe range of the operator averages — with two operators, simply the bigger minus the smaller: 0.008 mm.
K2the same kind of range-to-standard-deviation factor as K1, but for the number of operators: 0.7071 for 2 operators (0.5231 for 3), again from the AIAG manual's d2*-based tables.
n, rparts and trials: 5 and 2. Their product, 10, is how many readings went into each operator's average — and averaging 10 readings shrinks the equipment noise in that average by a factor of 10 in the squared units.
AVappraiser variation — reproducibility as a standard deviation. If the subtraction goes negative, AV is set to zero: the operators' disagreement was no bigger than equipment noise alone explains, so there is no separate operator effect to report.

In words: convert the operator disagreement into standard-deviation units, then subtract (in squares, as Equation 1 demands) the part of it that mere instrument noise would have produced anyway. What survives is genuinely attributable to the people. The floor at zero is honest bookkeeping, not a fudge — you cannot observe negative variation.

Run the numbers, and check each step:

Combine the two pieces the Pythagorean way, exactly as Equation 1 promised, and you have the headline number, the total gage R&R — the measurement system's combined spread, gage repeatability and reproducibility:

GRR  =  √EV ² + AV ²  =  √0.0000785350 + 0.0000241459  =  0.010133 mm

About ten microns of standard deviation, of which the instrument's own noise is the larger share and the operator difference the smaller. Already useful: if you wanted to improve this system, you now know that swapping operators buys less than fixing the instrument or the technique.

6 · The part variation and the total

Now the part-to-part variation. The parts get the same treatment as the operators. Average each part across everything — both operators, both trials, four readings per part: 24.9475, 25.0075, 25.0300, 25.0875, 25.1225 mm. The range of those five part averages is Rp = 25.1225 − 24.9475 = 0.175 mm, and one more K-constant converts a range of five values into a standard deviation:

Equation 4Part variation
PV  =  Rp × K3  =  0.175 × 0.4030  =  0.070525 mm
Rpthe range of the part averages: biggest part minus smallest part, each seen through the mean of its four readings.
K3the conversion factor for a range of 5 values: 0.4030, from the same AIAG d2*-based tables (it depends on how many parts you used — 0.3146 for the classic 10-part study).
PVpart variation as a standard deviation: 0.070525 mm, about 70.5 µm. Averaging four readings per part has already beaten most of the measurement noise out of these part averages, which is why this is a fair estimate of the parts alone.

In words: the parts themselves spread about seven times wider than the measurement system's ten microns. Now Equation 1 runs forward instead of backward: total variation is the Pythagorean sum, TV = √GRR ² + PV ² = √0.010133² + 0.070525² = 0.071249 mm — barely more than PV alone, because small noise hides.

Everything now goes into one summary table, each component expressed as a percentage of total variation. Note the built-in self-check: because of Equation 1, %GRR² + %PV² must equal 100² before rounding — a check worth running on any report you're handed. Here, 14.22² + 98.98² = 9,999.2, the small shortfall being purely the rounding of the two percentages.

componentsymbolstd. dev. (mm)≈ µm% of TV
repeatability (equipment)EV0.0088628.912.4%
reproducibility (appraisers)AV0.0049144.96.9%
total gage R&R (measurement system)GRR0.01013310.114.2%
part variationPV0.07052570.599.0%
total variationTV0.07124971.2

7 · The verdict: %GRR, ndc, and what the thresholds really are

The AIAG manual's acceptance guidelines for the total gage R&R, quoted in every customer requirement you will ever receive, read: under 10% of total variation, the measurement system is generally acceptable; 10% to 30%, it may be acceptable for some applications, depending on the importance of the measurement, the cost of the gage, the cost of rework — a decision to be justified and, in practice, approved by the customer; over 30%, unacceptable — improve the system. Our study, at %GRR = 14.2, lands in the conditional band.

Alongside %GRR the form reports one more figure, the number of distinct categories:

Equation 5Number of distinct categories
ndc  =  1.41 × PVGRR  =  1.41 × 0.0705250.010133  =  1.41 × 6.96  =  9.81  →  9
ndcroughly, how many distinct groups the measurement system can reliably sort this study's parts into — 9 here, after dropping the fraction (the manual says to round down). The AIAG guideline asks for 5 or more.
1.41this is √2 in disguise. Telling two parts apart means comparing two noisy readings, and the noise of a difference is √2 times the noise of a single reading — so the factor converts "spread of parts over spread of noise" into "how many separable bins."

In words: an ndc of 1 means the gage can only say "small-ish or large-ish." An ndc of 9 means it can resolve the current part spread into about nine honest steps — comfortably enough to run a control chart on. Note what ndc does not say: nothing about the tolerance. It compares the gage to the parts you happened to study, which is a warning we are about to take seriously.

So: conditional %GRR, healthy ndc. But hold the thresholds at arm's length. Nothing physical happens at 10.0% or 30.0% — a system at 9.9% and one at 10.1% are, for every practical purpose, the same system. These are decision conventions, like speed limits: round numbers drawing enforceable lines through a continuum of risk. The MSA manual itself presents them as guidelines whose application depends on what the measurement is for, not as laws of nature. Treat a 29% system on a safety-critical characteristic with more suspicion than a 31% system sorting cosmetic shims, and you are using the manual as intended.

8 · Percent of what? Tolerance and variation answer different questions

Here is the part most studies get handed to an auditor without anyone noticing. That 14.2% had total variation in the denominator. There is a second convention: compare the measurement system variation to the tolerance instead. Take six standard deviations of GRR — the conventional 99.73% width of its scatter — against the full tolerance width of 0.20 mm:

%GRRtolerance  =  6 × 0.0101330.20  =  0.06080.20  =  30.4%

Same study, same twenty readings: 14.2% one way, 30.4% the other — conditionally acceptable by one convention, at the unacceptable line by the other. Neither number is wrong. They answer different questions:

Our example shows why the distinction bites. The five parts spread 0.175 mm — chosen, quite properly, to exercise the gage across the whole process — so TV is large and the variation-based percentage looks decent. But the drawing only grants 0.20 mm, and a system that eats 30% of it will misclassify parts near the limits routinely: reading 25.09 on a part whose true diameter is 25.11 is a measurement error of about two of this gage's standard deviations — the kind of miss its normal scatter produces every few dozen readings, well inside the six-standard-deviation spread the 30% was built from. Flip the situation — a superbly capable process making nearly identical parts — and the trap reverses: TV collapses, %GRR of TV looks catastrophic, and a good gage fails because the parts gave it nothing to distinguish. When the studied parts do not represent the process, the manual's own advice is to lean on the tolerance-based figure or an independent estimate of process variation.

Two footnotes for report-readers. Older editions of the MSA manual used 5.15 standard deviations (a 99% width) where this note uses 6 — if a legacy spreadsheet's percentages run curiously low, that multiplier is usually why. And the measurement-uncertainty world — VDA 5 and ISO 22514-7 on measurement-process capability, ISO 14253-1 on conformance decisions — attacks the same question by budgeting an expanded uncertainty against the tolerance and shrinking the acceptance zone; different vocabulary, same worry about decisions made near the limits.

9 · ANOVA: what modern software does instead

The average-and-range method you just walked is the pencil-and-paper method, and its virtue is that you can see every moving part. It is not what good software uses. The MSA manual's preferred method is ANOVA — analysis of variance — which takes the same 20 readings and partitions their variation by an algebra we won't walk here, with three practical advantages.

First, ANOVA uses all the information — every reading's distance from every relevant average, not just highest-minus-lowest ranges — so its estimates are steadier for the same data. Second, the big one: it can detect the operator-by-part interaction, where operators disagree differently on different parts. Operator B reads high, but only on the parts with a heavy burr; the range method smears that into general noise, while ANOVA isolates it as its own line item — usually the most diagnostic number on the report, because it points at a method problem (an ambiguous measuring position, an unstated cleaning step) rather than a hardware problem. Third, ANOVA extends gracefully: unbalanced studies, more factors, confidence intervals on every component.

The practical advice is unglamorous: do the range method once by hand so the machinery is demystified, then let software do ANOVA forever after — and read the interaction line first. The two methods disagree slightly on the same data; they are different estimators of the same quantities, and that disagreement is itself a small lesson in how much statistical wobble lives in a 20-reading gage R&R study.

10 · When the gage is a scanner and its software

Now the question this site keeps circling: what does any of this mean when the "gage" is a structured-light scanner, a robot program, and an inspection template — when no human touches a micrometer and the "reading" is the output of a software evaluation?

Start with a fact that surprises people: the MSA manual is already on the right side of this. It defines a measurement system broadly — the instrument and the standards, operations, methods, fixtures, software, personnel and environment used to get the number. By that definition, a scanning cell's measurement system includes the sensor and its calibration, the spray or surface prep (which is its own error budget), the fixturing and part loading, the scan strategy, the meshing parameters, the alignment choice and fitting rule, the datum scheme, and the evaluation template's version. The software choices are not the context of the measurement; they are components of the gage.

That has three concrete consequences for running an honest gage R&R study on a scanner:

  1. Repeat the physical act, not the computation. Software is deterministic: re-running the evaluation on the same mesh returns the same number to the last digit, and a "study" built that way reports a glorious 0% GRR while measuring nothing. A trial must repeat what actually varies — re-scan the part, and between trials take it out of the fixture and load it again. Scanner repeatability is re-scan plus re-load, or it is fiction.
  2. Decide what "operator" means before the study. In a manual scan the operator genuinely handles the device and the spray can — classic reproducibility. In an automated cell the humans left; what plays the operator's role is whatever still differs between nominally identical measurements: who loads the part and how, which fixture nest, morning warm-up versus afternoon, this week's template versus last week's. Reproducibility is defined by what you chose to vary — section 2's lesson — so choose deliberately: across loaders, shifts, re-created setups.
  3. Freeze the evaluation, and version it. A scanner gage R&R study is only meaningful if the alignment, fitting rules, and datum strategy are locked for its duration and recorded with the result — because changing any of them changes the reported number more than most operators ever could. Switching a fit or an alignment from best-fit to datum-based doesn't add scatter; it moves the answer, systematically, for every part. If the template changes, the measurement system changed, and the old study describes a gage that no longer exists.

One more transfer from the manual worth keeping: a gage R&R belongs to a characteristic, not to an instrument. The same cell can be excellent on a bore diameter and marginal on a flatness across a sprayed surface — pass/fail is per measurement, and the study list should follow the drawing's critical characteristics, not the hardware inventory. (For what these evaluation choices look like in working automated routines, the worked demonstrations on this site are exactly that.)

11 · A passed study is a snapshot — the short version

Last, the honest limits — the paragraph that should be stapled to every GRR certificate. A gage R&R is a snapshot: these parts, these operators, this day, this temperature, this state of wear, this version of the template. Twenty readings estimate EV from just ten ranges, and estimates that thin carry real statistical wobble — run the identical study next week and expect a visibly different %GRR from the same healthy system. And the study says nothing about bias (a gage can repeat beautifully around the wrong value — that's a calibration and bias study), linearity (fine at 25 mm, biased at 80), or stability (fine in March, drifted by August — that's a control chart on a reference part, running continuously). Passing 14.2% did not confer a property on the gage; it recorded the behavior of a whole measurement system at one moment. The systems that stay trustworthy are the ones somebody keeps watching.

Further reading

If your scanning cell has a GRR certificate but nobody can say which alignment, fitting rule, and template version it certifies, that is a solvable problem.

This is the work I offer. Metrology Maven turns approved inspection methods into pipelines that run unattended — starting with a fixed-fee assessment, from $3,500. How engagements work →