Sooner or later a customer asks for it: a gage R&R, the spreadsheet with the shaded cells, constants like 0.8862 that nobody in the room can explain. Most labs run the study the way they run a fire drill — because the audit requires it. This note rebuilds the whole thing from one idea — some of what you see is the part; the rest is the measurement — and walks a complete gage R&R study, small enough to check with a pencil, so that every constant and every percentage has a reason attached to it.
Contents
Measure ten parts from a process and the numbers scatter. Two things put that scatter there: the parts genuinely differ (part-to-part variation), and the act of measuring adds noise of its own — the instrument, the operator's touch, the fixturing, the temperature, the software. You never see a part's value directly; you see it through a measurement system, and the system's fingerprints are on every reading.
A gage R&R study is an experiment designed to pull those two contributions apart: how much of the variation I observe is the parts, and how much is the measuring? That's the entire question; the K-constants, percentages and thresholds are bookkeeping in service of it. (Terminology: the AIAG manuals spell it gage, much of the world writes gauge, and GRR is the same study either way. R&R is one tool inside measurement system analysis, MSA, which also covers bias, linearity and stability — more on those at the end.)
The study leans on one mathematical fact, and it is the only one you need all day:
| TV | total variation: the typical spread of the readings you actually observe. All spreads here are measured as standard deviations — the usual "typical distance from the average" number. |
| PV | part variation: the spread the parts alone would show, measured by a perfect instrument. |
| GRR | the measurement system's own spread — the thing the study exists to estimate. |
In words: spreads do not add like lengths; they add like the perpendicular sides of a right triangle (Pythagoras, wearing a lab coat). The consequence is lopsided: a measurement system a third the size of the part variation inflates what you see by only about 5%, because √3² + 1² = 3.16, barely more than 3. Small noise hides — which is exactly why finding it takes a designed experiment.
A gage R&R study splits measurement system variation into two named pieces, and the names are best understood by asking what changed between two readings.
Repeatability is the scatter you get when nothing changed: same part, same operator, same instrument, same setup, readings minutes apart. Whatever spread remains is the measurement system arguing with itself — the instrument's noise floor plus the tiny inconsistencies of a single person's technique. The AIAG manual calls this equipment variation, EV.
Reproducibility is the extra scatter that appears when the thing that changed is who — a different operator, with their own touch, their own reading of the vernier, their own habit of seating the part. It is measured from how far apart the averages land when different operators measure the same parts, and the manual calls it appraiser variation, AV.
The definitions are operational, not philosophical. If your "operators" are three shifts, or three fixtures, or three software templates, the same arithmetic measures reproducibility across whatever you varied. That flexibility is worth remembering — it is what will let the method survive contact with automated scanning in section 10.
Here is a complete gage R&R study you can check by hand — a deliberately small teaching example, its numbers constructed to be easy to follow rather than taken from anywhere: two operators, five parts, two trials each, 20 readings. (Real studies more often use ten parts, three operators, three trials; the machinery is identical, only the constants change.)
The setting: a turned shaft, drawing calls the diameter 25.00 ± 0.10 mm, measured with a micrometer reading to 0.01 mm. The five parts were pulled to span the process — one runs near the top of the tolerance and one just beyond it, which is deliberate on both counts: a gage R&R study spans the process, not the print, and it needs genuine part-to-part variation, or there is nothing for the gage to distinguish. Both operators measure every part twice, in randomized order, blind to their own earlier reading and to the other operator's. The blindness is not ceremony; an operator who remembers writing 25.08 will find 25.08 again, and the study will flatter the gage.
| part | Operator A (mm) | Operator B (mm) | ||||||
|---|---|---|---|---|---|---|---|---|
| trial 1 | trial 2 | mean | range | trial 1 | trial 2 | mean | range | |
| 1 | 24.94 | 24.95 | 24.945 | 0.01 | 24.96 | 24.94 | 24.950 | 0.02 |
| 2 | 25.00 | 25.00 | 25.000 | 0.00 | 25.02 | 25.01 | 25.015 | 0.01 |
| 3 | 25.04 | 25.02 | 25.030 | 0.02 | 25.03 | 25.03 | 25.030 | 0.00 |
| 4 | 25.08 | 25.09 | 25.085 | 0.01 | 25.10 | 25.08 | 25.090 | 0.02 |
| 5 | 25.12 | 25.11 | 25.115 | 0.01 | 25.13 | 25.13 | 25.130 | 0.00 |
| operator mean X̄ | 25.035 | 25.043 | ||||||
| mean range R̄ | 0.010 | 0.010 | ||||||
Look at the picture before any arithmetic. The part clusters are clearly separated — good. Each cluster is a couple of hundredths wide — that width is the measurement system. And operator B's rings sit very slightly right of operator A's dots on most parts — that small systematic offset is reproducibility, about to be caught by the math.
Each row of the table contains a pair of readings where nothing changed — same part, same operator. The difference within such a pair can only be measurement error. The range — the bigger reading minus the smaller — captures it: 0.01, 0.00, 0.02, 0.01, 0.01 for operator A; 0.02, 0.01, 0.00, 0.02, 0.00 for operator B. Average all ten and you get the mean range, R̄ = 0.010 mm.
A range is not yet a standard deviation, but for small samples the two are tied together by a known conversion factor. Statisticians tabulate the factor (they call it d2*); the AIAG MSA manual publishes its reciprocals as the K-constants on the study form, derived from those d2* tables. For pairs — two trials — the constant is K1 = 0.8862.
| R̄ | ("R-bar") the average of all ten within-pair ranges — the typical disagreement of a reading with its own repeat. |
| K1 | the range-to-standard-deviation conversion factor: 0.8862 when each range comes from 2 trials (0.5908 for 3 trials), from the d2*-based tables in the AIAG MSA manual. |
| EV | equipment variation — repeatability, expressed as a standard deviation: about 0.0089 mm, or 8.9 µm. |
In words: collect all the "nothing changed" disagreements, average them, and rescale the average into standard-deviation units. The result says: hand this micrometer and this technique the same shaft twice, and the readings typically wander by about nine microns' worth of standard deviation. That is the noise floor everything else sits on.
Now the second piece. Operator A's grand average over all ten readings is 25.035 mm; operator B's is 25.043 mm. Two different operators measured the same five parts, so in a perfect world the averages would agree. The difference, X̄DIFF = 25.043 − 25.035 = 0.008 mm, is the between-operator signal.
But there's a subtlety, and it is the one genuinely clever step in the whole worksheet. Each operator's average was built from ten noisy readings — so even two clones of the same operator would show some difference, purely from equipment noise. If we converted X̄DIFF straight into a standard deviation, we would count repeatability twice. So the form subtracts the share of the difference that repeatability alone would have produced:
| X̄DIFF | the range of the operator averages — with two operators, simply the bigger minus the smaller: 0.008 mm. |
| K2 | the same kind of range-to-standard-deviation factor as K1, but for the number of operators: 0.7071 for 2 operators (0.5231 for 3), again from the AIAG manual's d2*-based tables. |
| n, r | parts and trials: 5 and 2. Their product, 10, is how many readings went into each operator's average — and averaging 10 readings shrinks the equipment noise in that average by a factor of 10 in the squared units. |
| AV | appraiser variation — reproducibility as a standard deviation. If the subtraction goes negative, AV is set to zero: the operators' disagreement was no bigger than equipment noise alone explains, so there is no separate operator effect to report. |
In words: convert the operator disagreement into standard-deviation units, then subtract (in squares, as Equation 1 demands) the part of it that mere instrument noise would have produced anyway. What survives is genuinely attributable to the people. The floor at zero is honest bookkeeping, not a fudge — you cannot observe negative variation.
Run the numbers, and check each step:
Combine the two pieces the Pythagorean way, exactly as Equation 1 promised, and you have the headline number, the total gage R&R — the measurement system's combined spread, gage repeatability and reproducibility:
GRR = √EV ² + AV ² = √0.0000785350 + 0.0000241459 = 0.010133 mm
About ten microns of standard deviation, of which the instrument's own noise is the larger share and the operator difference the smaller. Already useful: if you wanted to improve this system, you now know that swapping operators buys less than fixing the instrument or the technique.
Now the part-to-part variation. The parts get the same treatment as the operators. Average each part across everything — both operators, both trials, four readings per part: 24.9475, 25.0075, 25.0300, 25.0875, 25.1225 mm. The range of those five part averages is Rp = 25.1225 − 24.9475 = 0.175 mm, and one more K-constant converts a range of five values into a standard deviation:
| Rp | the range of the part averages: biggest part minus smallest part, each seen through the mean of its four readings. |
| K3 | the conversion factor for a range of 5 values: 0.4030, from the same AIAG d2*-based tables (it depends on how many parts you used — 0.3146 for the classic 10-part study). |
| PV | part variation as a standard deviation: 0.070525 mm, about 70.5 µm. Averaging four readings per part has already beaten most of the measurement noise out of these part averages, which is why this is a fair estimate of the parts alone. |
In words: the parts themselves spread about seven times wider than the measurement system's ten microns. Now Equation 1 runs forward instead of backward: total variation is the Pythagorean sum, TV = √GRR ² + PV ² = √0.010133² + 0.070525² = 0.071249 mm — barely more than PV alone, because small noise hides.
Everything now goes into one summary table, each component expressed as a percentage of total variation. Note the built-in self-check: because of Equation 1, %GRR² + %PV² must equal 100² before rounding — a check worth running on any report you're handed. Here, 14.22² + 98.98² = 9,999.2, the small shortfall being purely the rounding of the two percentages.
| component | symbol | std. dev. (mm) | ≈ µm | % of TV |
|---|---|---|---|---|
| repeatability (equipment) | EV | 0.008862 | 8.9 | 12.4% |
| reproducibility (appraisers) | AV | 0.004914 | 4.9 | 6.9% |
| total gage R&R (measurement system) | GRR | 0.010133 | 10.1 | 14.2% |
| part variation | PV | 0.070525 | 70.5 | 99.0% |
| total variation | TV | 0.071249 | 71.2 | — |
The AIAG manual's acceptance guidelines for the total gage R&R, quoted in every customer requirement you will ever receive, read: under 10% of total variation, the measurement system is generally acceptable; 10% to 30%, it may be acceptable for some applications, depending on the importance of the measurement, the cost of the gage, the cost of rework — a decision to be justified and, in practice, approved by the customer; over 30%, unacceptable — improve the system. Our study, at %GRR = 14.2, lands in the conditional band.
Alongside %GRR the form reports one more figure, the number of distinct categories:
| ndc | roughly, how many distinct groups the measurement system can reliably sort this study's parts into — 9 here, after dropping the fraction (the manual says to round down). The AIAG guideline asks for 5 or more. |
| 1.41 | this is √2 in disguise. Telling two parts apart means comparing two noisy readings, and the noise of a difference is √2 times the noise of a single reading — so the factor converts "spread of parts over spread of noise" into "how many separable bins." |
In words: an ndc of 1 means the gage can only say "small-ish or large-ish." An ndc of 9 means it can resolve the current part spread into about nine honest steps — comfortably enough to run a control chart on. Note what ndc does not say: nothing about the tolerance. It compares the gage to the parts you happened to study, which is a warning we are about to take seriously.
So: conditional %GRR, healthy ndc. But hold the thresholds at arm's length. Nothing physical happens at 10.0% or 30.0% — a system at 9.9% and one at 10.1% are, for every practical purpose, the same system. These are decision conventions, like speed limits: round numbers drawing enforceable lines through a continuum of risk. The MSA manual itself presents them as guidelines whose application depends on what the measurement is for, not as laws of nature. Treat a 29% system on a safety-critical characteristic with more suspicion than a 31% system sorting cosmetic shims, and you are using the manual as intended.
Here is the part most studies get handed to an auditor without anyone noticing. That 14.2% had total variation in the denominator. There is a second convention: compare the measurement system variation to the tolerance instead. Take six standard deviations of GRR — the conventional 99.73% width of its scatter — against the full tolerance width of 0.20 mm:
%GRRtolerance = 6 × 0.0101330.20 = 0.06080.20 = 30.4%
Same study, same twenty readings: 14.2% one way, 30.4% the other — conditionally acceptable by one convention, at the unacceptable line by the other. Neither number is wrong. They answer different questions:
Our example shows why the distinction bites. The five parts spread 0.175 mm — chosen, quite properly, to exercise the gage across the whole process — so TV is large and the variation-based percentage looks decent. But the drawing only grants 0.20 mm, and a system that eats 30% of it will misclassify parts near the limits routinely: reading 25.09 on a part whose true diameter is 25.11 is a measurement error of about two of this gage's standard deviations — the kind of miss its normal scatter produces every few dozen readings, well inside the six-standard-deviation spread the 30% was built from. Flip the situation — a superbly capable process making nearly identical parts — and the trap reverses: TV collapses, %GRR of TV looks catastrophic, and a good gage fails because the parts gave it nothing to distinguish. When the studied parts do not represent the process, the manual's own advice is to lean on the tolerance-based figure or an independent estimate of process variation.
Two footnotes for report-readers. Older editions of the MSA manual used 5.15 standard deviations (a 99% width) where this note uses 6 — if a legacy spreadsheet's percentages run curiously low, that multiplier is usually why. And the measurement-uncertainty world — VDA 5 and ISO 22514-7 on measurement-process capability, ISO 14253-1 on conformance decisions — attacks the same question by budgeting an expanded uncertainty against the tolerance and shrinking the acceptance zone; different vocabulary, same worry about decisions made near the limits.
The average-and-range method you just walked is the pencil-and-paper method, and its virtue is that you can see every moving part. It is not what good software uses. The MSA manual's preferred method is ANOVA — analysis of variance — which takes the same 20 readings and partitions their variation by an algebra we won't walk here, with three practical advantages.
First, ANOVA uses all the information — every reading's distance from every relevant average, not just highest-minus-lowest ranges — so its estimates are steadier for the same data. Second, the big one: it can detect the operator-by-part interaction, where operators disagree differently on different parts. Operator B reads high, but only on the parts with a heavy burr; the range method smears that into general noise, while ANOVA isolates it as its own line item — usually the most diagnostic number on the report, because it points at a method problem (an ambiguous measuring position, an unstated cleaning step) rather than a hardware problem. Third, ANOVA extends gracefully: unbalanced studies, more factors, confidence intervals on every component.
The practical advice is unglamorous: do the range method once by hand so the machinery is demystified, then let software do ANOVA forever after — and read the interaction line first. The two methods disagree slightly on the same data; they are different estimators of the same quantities, and that disagreement is itself a small lesson in how much statistical wobble lives in a 20-reading gage R&R study.
Now the question this site keeps circling: what does any of this mean when the "gage" is a structured-light scanner, a robot program, and an inspection template — when no human touches a micrometer and the "reading" is the output of a software evaluation?
Start with a fact that surprises people: the MSA manual is already on the right side of this. It defines a measurement system broadly — the instrument and the standards, operations, methods, fixtures, software, personnel and environment used to get the number. By that definition, a scanning cell's measurement system includes the sensor and its calibration, the spray or surface prep (which is its own error budget), the fixturing and part loading, the scan strategy, the meshing parameters, the alignment choice and fitting rule, the datum scheme, and the evaluation template's version. The software choices are not the context of the measurement; they are components of the gage.
That has three concrete consequences for running an honest gage R&R study on a scanner:
One more transfer from the manual worth keeping: a gage R&R belongs to a characteristic, not to an instrument. The same cell can be excellent on a bore diameter and marginal on a flatness across a sprayed surface — pass/fail is per measurement, and the study list should follow the drawing's critical characteristics, not the hardware inventory. (For what these evaluation choices look like in working automated routines, the worked demonstrations on this site are exactly that.)
Last, the honest limits — the paragraph that should be stapled to every GRR certificate. A gage R&R is a snapshot: these parts, these operators, this day, this temperature, this state of wear, this version of the template. Twenty readings estimate EV from just ten ranges, and estimates that thin carry real statistical wobble — run the identical study next week and expect a visibly different %GRR from the same healthy system. And the study says nothing about bias (a gage can repeat beautifully around the wrong value — that's a calibration and bias study), linearity (fine at 25 mm, biased at 80), or stability (fine in March, drifted by August — that's a control chart on a reference part, running continuously). Passing 14.2% did not confer a property on the gage; it recorded the behavior of a whole measurement system at one moment. The systems that stay trustworthy are the ones somebody keeps watching.
If your scanning cell has a GRR certificate but nobody can say which alignment, fitting rule, and template version it certifies, that is a solvable problem.
This is the work I offer. Metrology Maven turns approved inspection methods into pipelines that run unattended — starting with a fixed-fee assessment, from $3,500. How engagements work →