Every calibration certificate in your lab carries a line like U = 3 µm (k = 2) — a measurement uncertainty statement. Most of the people who read that line every week were never shown where it comes from — it arrives from the calibration lab like weather. This note builds one of those statements from scratch: a complete measurement uncertainty budget for a pin diameter, with numbers deliberately chosen so that you can check every step on paper. By the end, U, k, Type A, Type B, and the guard bands on your acceptance decisions will all be one connected piece of machinery.
Contents
A diameter reported as 10.0020 mm, full stop, makes a strangely incomplete claim. Taken literally it says the diameter is that value — which nobody believes, including the person who measured it. Measure the same pin tomorrow, or on the neighboring instrument, and you will get a slightly different number. A measurement result without a stated measurement uncertainty is a point with no error bars: it tells you where the dart landed, but not how big the cluster of darts would be if you kept throwing.
A complete result — 10.0020 mm ± 0.0019 mm (k = 2) — makes a claim with a defined shape: the value of the thing I set out to measure lies, with roughly 95% confidence, somewhere between 10.0001 and 10.0039 mm. Not "the answer is 10.0020"; rather, "here is the interval I am prepared to defend, and here is how sure I am." That is the entire content of the statement, and everything in this note exists to manufacture those two numbers honestly.
The claim's shape explains something that otherwise causes grief: two labs can disagree and both be right. If your lab reports 10.0020 ± 0.0019 mm and a customer's lab reports 10.0035 ± 0.0030 mm, the headline numbers differ by 1.5 µm — but check the intervals: yours runs 10.0001 to 10.0039 and contains their 10.0035; theirs runs 10.0005 to 10.0065 and contains your 10.0020. The two claims are consistent; the disagreement is well within what the measurement uncertainty statements predicted. Without those statements, the same 1.5 µm looks like a dispute, and disputes without error bars are settled by whoever shouts loudest.
One more idea completes the picture: traceability. Your micrometer was calibrated against gauge blocks, the gauge blocks against a national laboratory's interferometry, and that against the definition of the metre itself. Each link in that chain is a comparison, and each comparison carries its own doubt. The measurement uncertainty on a calibration certificate is the accumulated, honestly bookkept doubt of the whole chain above it — which is why accredited labs are required to state it, and why a certificate without a measurement uncertainty is not a calibration but a rumor.
Two words get used interchangeably on shop floors and mean rigorously different things.
Error is the difference between your measured value and the true value. It is a perfectly good concept with one fatal flaw: you can never know it, because you never know the true value. If you knew your error was +1.3 µm you would simply subtract 1.3 µm and have zero error. What you can know, you correct: a systematic effect you have measured — a documented instrument bias, a known thermal offset — stops being error and becomes a correction. What remains, by construction, is the part you cannot pin down.
Measurement uncertainty is the honest response to that situation: instead of pretending to know the error, you quantify your doubt about it. Formally — this is the framing of the GUM, the Guide to the Expression of Uncertainty in Measurement (JCGM 100:2008), the document the entire system rests on — measurement uncertainty characterizes the spread of values that could reasonably be attributed to the thing you measured. It is a standard deviation: a width, not an offset. Error is a single unknowable number; uncertainty is a knowable statement about how large it might plausibly be.
The GUM's second great contribution is bookkeeping discipline. Every influence on the final measurement uncertainty gets evaluated, expressed in a common currency called standard uncertainty (one standard deviation's worth of doubt, in the units of the measurement), and combined by a fixed rule. The evaluations come in exactly two flavors, and the names are duller than they deserve: Type A, doubt you evaluate statistically from your own repeated measurements, and Type B, doubt you evaluate by any other means — certificates, datasheets, physics, experience. The distinction is about how you found out, not about random versus systematic; both types end up as standard deviations, and both feed the same machinery. We take them in turn, with real numbers.
Here is the running example for the rest of this note. It is a teaching example with deliberately convenient numbers — no real part, no real lab — built so every value is checkable by hand. You measure a nominally 10 mm steel pin with a digital micrometer that reads to 0.001 mm, ten times, reseating the pin each time. The readings, in millimetres:
10.004 10.001 10.003 9.999 10.002
10.005 10.000 10.004 10.002 10.000
Step one: the value you will report is the mean.
| x1, x2, … | the individual measured values — here, the ten diameters above. Add them up and you get exactly 100.020 mm; try it. |
| x̄ | ("x-bar") their average — the reported result. |
| n | how many readings — here, 10. |
In words: ten slightly different answers, one representative value. The question this section answers is: how much should the scatter of those ten make you doubt it?
Step two: measure the scatter. For each reading, take its deviation from the mean, square it, and add the squares. The table is the whole calculation — every row is checkable in your head:
| reading (mm) | deviation from mean 10.0020 (µm) | deviation squared (µm²) |
|---|---|---|
| 10.004 | +2 | 4 |
| 10.001 | −1 | 1 |
| 10.003 | +1 | 1 |
| 9.999 | −3 | 9 |
| 10.002 | 0 | 0 |
| 10.005 | +3 | 9 |
| 10.000 | −2 | 4 |
| 10.004 | +2 | 4 |
| 10.002 | 0 | 0 |
| 10.000 | −2 | 4 |
| sum: 100.020 | 0 (always — good check) | 36 |
| s | the typical size of one reading's wobble about the mean — the instrument-plus-process repeatability, in micrometres. |
| n − 1 | nine, not ten. The deviations were measured from the sample's own mean, which by definition sits in the middle of them — so they underestimate the true scatter slightly. Dividing by n − 1 instead of n pays that borrowing back. (Statisticians call it Bessel's correction; your spreadsheet's STDEV does it automatically.) |
In words: a single reading of this pin on this micrometer typically wanders about 2 µm from the centre. That is the scatter of one throw.
Step three — and this is the part people meet last, if ever: you are not reporting one throw. You are reporting the mean of ten, and means are steadier than the individual measured values they came from, because the up-wobbles and down-wobbles partially cancel. The standard uncertainty of the mean shrinks by the square root of the number of readings:
| uA | the Type A standard uncertainty: how much the reported mean would wobble if you repeated the whole ten-reading exercise many times. |
| √n | the square root of the number of readings — for 10 readings, about 3.16. |
In words: averaging ten readings bought the doubt down from 2 µm to 0.63 µm. One honest condition attaches: cancellation only works on effects that wobble between your repeats. Anything that pushed all ten readings the same way — a miscalibrated micrometer, a warm lab — sits untouched underneath, invisible to this arithmetic. That is exactly why Type B exists.
Repeating the measurement interrogates your process about its own scatter. It cannot interrogate the instrument's calibration, the finiteness of its display, or the temperature of the room — those influences either don't vary across your ten minutes of repeats, or vary too slowly to show. For them, you evaluate the doubt by other means: Type B. Three contributions cover most everyday dimensional measurement uncertainty budgets, and our example has all three.
The micrometer reads in steps of 0.001 mm. When it displays 10.004, the underlying value could be anything from 10.0035 to 10.0045 — the display cannot say where in that band, and every position in it is equally believable. That "equally believable anywhere within ±a" shape is called a rectangular probability distribution, and it converts to the common currency by a fixed rule:
| a | the half-width of the band you are sure the value lies in. For a 1 µm display step, the half-width is 0.5 µm. |
| √3 | about 1.73. Dividing by it is not a convention pulled from the air: it is the actual standard deviation of a flat, rectangular spread. A rectangle of total doubt is "worth" about 58% of its half-width, once expressed as a standard deviation. |
In words: whenever your knowledge has the form "somewhere in ±a, no idea where," the standard-uncertainty equivalent is the half-width divided by 1.73. You will use this little rule constantly — resolution, temperature bands, datasheet limits all arrive in this shape.
The micrometer's own certificate says its indication error is within U = 1.2 µm (k = 2). This is the moment the mysterious certificate line becomes a working input: the lab is telling you the standard-deviation-worth of doubt, already multiplied by 2 for their 95% statement. To use it in your own budget, divide the 2 back out:
| U | the expanded uncertainty printed on the certificate — a 95%-confidence half-width (Section 7 builds this object properly). |
| k | the coverage factor the lab used, stated right next to it. Almost always 2. |
| ucal | the calibration doubt back in standard-deviation form, ready to combine. Treated as a normal (bell-shaped) probability distribution, because that is what the lab's own combination produced. |
In words: a certificate's U and k are a compressed file; dividing by k unzips it. This single habit — U over k, into the budget — is how traceability physically propagates from the national laboratory down to your part.
Dimensional metrology's reference temperature is 20 °C by international agreement (ISO 1) — every length on every drawing means "at 20 °C." Steel expands by about 11.5 × 10⁻⁶ of its length per degree, so our 10 mm pin changes by about 0.115 µm per °C. Say the lab is controlled to within ±2 °C of 20 °C, and that is all you know — the temperature could be anywhere in that band. That is a rectangular half-width of 0.115 × 2 = 0.23 µm, and Equation 4 already tells us what to do with it: utemp = 0.23 / 1.73 ≈ 0.13 µm.
We now hold four separate doubts, all in the same currency: 0.63, 0.29, 0.60, and 0.13 µm. How do they combine? The tempting answer — add them, get 1.65 µm — is wrong in an instructive way. Straight addition describes a conspiracy: it is the doubt you would have if every source pushed its hardest in the same direction at the same time. But these sources have nothing to do with one another. The repeatability scatter does not know what the thermostat is doing; the display's rounding does not consult the calibration error. Sometimes they push together; just as often they partially cancel.
For independent sources, statistics gives an exact rule: it is the squares of standard deviations (their variances) that add. Geometrically, independent doubts combine like steps taken at right angles:
| uc | the combined standard uncertainty: all sources of doubt merged into one standard deviation for the final result. |
| uA², ures², … | each source's standard uncertainty, squared: 0.400 + 0.083 + 0.360 + 0.018 = 0.861 µm². (Squares taken before rounding; square the rounded values instead and you get 0.858 — the same 0.93 after the square root. The arithmetic is forgiving, which is a healthy property for arithmetic you will do at a desk.) |
In words: square every doubt, add the squares, take the square root. Two conditions bought this simplicity: every input was already a standard uncertainty (one common currency — that was the point of Equations 3, 4 and 5), and the sources are independent. Correlated sources — two terms fed by the same thermometer, say — need extra care, which the GUM provides and this note does not.
Notice what the quadrature did: 0.63, 0.29, 0.60 and 0.13 merged into 0.93 — far less than the 1.65 of the conspiracy theory, and dominated by the two big terms. The 0.13 µm temperature term contributed 0.018 of the 0.861 total: almost nothing. This is the mercy of quadrature — small terms vanish into big ones — and also its trap, because it tempts you to wave small terms in without scrutiny. The discipline of listing them anyway is what the bookkeeping is for. The table that holds this bookkeeping has a name you have seen on certificates: an uncertainty budget.
One piece of the GUM's machinery has been working quietly behind everything above; it deserves a name. Formally, every measurement uncertainty evaluation begins with a measurement model: the measurand — the quantity you set out to know — is an output quantity y, calculated from input quantities through a function, y = f(x1, x2, …, xn). The inputs are everything the result depends on: the indicated values, the calibration correction, the temperature of the room. The model answers a blunt question — where does every number enter the calculation? — and an influence can only reach y through an input, so a forgotten input is doubt silently left out of the budget.
Each input reaches the output with a certain leverage, and that leverage has a name: the sensitivity coefficient. It is how much y moves per unit move of one input — the partial derivative of f, said plainly: nudge one input, hold the rest still, watch the output. I have already used one without naming it. In Section 4, the input "lab temperature" was uncertain by ±2 °C — doubt in degrees, useless to a budget kept in micrometres. The sensitivity coefficient did the conversion: our 10 mm steel pin changes by about 0.115 µm per °C, so the ±2 °C band became 0.115 × 2 = 0.23 µm. That multiplication is the whole job: translating an input's measurement uncertainty into the currency of the output.
The pin diameter itself is a direct reading: the micrometer's indication, its calibration correction and its resolution all enter the model in millimetres and add straight into the result, so each sensitivity coefficient is exactly 1. That is why Section 5 could add squared uncertainties with no further ceremony. The GUM's general rule — the law of propagation of uncertainty — weights each squared standard uncertainty by its squared sensitivity coefficient before the adding; with every coefficient 1, Equation 6's quadrature sum is the whole law.
One honest caveat. The law of propagation treats f as locally straight — it is a first-order approximation. For nearly linear models and modest doubts — most everyday dimensional work — it is excellent. When the model is strongly non-linear it can mislead, and the GUM's own supplement (JCGM 101) takes a different road: a Monte Carlo method feeds each input's whole probability distribution into the model, recalculates y many times, and reads the output's probability distribution — and its 95% interval — directly.
None of this is academic bookkeeping. ISO/IEC 17025, the standard accredited laboratories operate under, requires a laboratory to identify the contributions to measurement uncertainty and evaluate them by documented methods — which in practice means writing this model down. A budget with no model behind it is a list of plausible numbers; the model makes it defensible.
The combined standard uncertainty is one standard deviation of doubt. If the combined doubt is roughly bell-shaped — and it usually is, because adding several independent wobbles of comparable size pushes any pile of probability distributions toward the bell curve — then an interval of ±1 standard deviation contains the value only about 68% of the time. No one wants to sign a certificate that is wrong one time in three. So the final step widens the interval to a defendable confidence:
| U | the expanded uncertainty — the half-width actually printed after the ±. Capital U; the lower-case u's are its standard-deviation ingredients. |
| k | the coverage factor: how many standard deviations wide you make the claim. For a bell shape, k = 1 covers about 68%, k = 2 about 95%, k = 3 about 99.7%. Industry has settled on k = 2 as the default, which is why the certificate says so explicitly — it lets the reader undo it (Equation 5). |
In words: k = 2 buys roughly 95% confidence: across many measurements reported this way, about 19 intervals in 20 will contain the value they claim to contain. (When a budget leans on very few repeat readings, the GUM has machinery — effective degrees of freedom — that says to use a larger k; with ten readings and several comparable Type B terms, k = 2 is the standard regime.)
And with that, the budget closes. Here is the whole measurement, one table — the finished measurement uncertainty budget. Every line traces back to a section above, and the reader-check is built in: square the fourth column, sum it, and the square root of the sum must give the combined row.
| source of doubt | type | evaluated from | standard uncertainty u (µm) | u² (µm²) |
|---|---|---|---|---|
| repeatability of the mean | A | 10 readings, s = 2.0 µm, ÷√10 | 0.63 | 0.400 |
| display resolution 1 µm | B | rectangular, 0.5 ÷ √3 | 0.29 | 0.083 |
| micrometer calibration | B | certificate U = 1.2 µm, k = 2, ÷2 | 0.60 | 0.360 |
| temperature ±2 °C, steel, 10 mm | B | rectangular, 0.23 ÷ √3 | 0.13 | 0.018 |
| sum of squares | 0.861 | |||
| combined standard uncertainty uc = √0.861 | 0.93 | |||
| expanded uncertainty U = 2 × uc (k = 2, ≈95%) | 1.9 | |||
The certified result: diameter = 10.0020 mm ± 0.0019 mm (k = 2). That line — the one that used to arrive like weather — is now something you calculated, and could defend term by term. Also worth noticing: the measurement uncertainty budget tells you where improvement lives. Take a hundred readings instead of ten and the Type A term shrinks to 0.2 µm — but uc only drops to about 0.7, because the calibration term does not care how many times you repeat. Past a point, the only way to a better measurement is a better instrument, calibrated better. The budget shows you exactly where that point is.
Measurement uncertainty stops being philosophy the moment a part must pass or fail. Suppose the pin's drawing says 10.000 ± 0.010 mm. Our result, 10.0020, sits well inside — easy call. But what about a measured 10.0095? The point estimate is inside the limit; the upper end of its measurement uncertainty interval, 10.0095 + 0.0019 = 10.0114, is not. Ship it, and there is a real chance the pin's actual diameter is beyond the limit. The measurement cannot tell you it conforms — it can only tell you it might.
ISO 14253-1 — the standard on decision rules for proving conformance and nonconformance — resolves this with a rule of admirable bluntness. By default, to prove conformance, the whole measurement claim must fit inside the specification: the acceptance zone is the tolerance zone shrunk by a guard band of width U at each end. Symmetrically, to prove nonconformance, the result must sit beyond the limit by more than U. In between lies a gray strip where the measurement honestly cannot decide, and what happens there is not physics but contract — something for customer and supplier to agree in writing (which is why accredited labs are now required to state which decision rule they applied).
Do the calculation on what the guard bands cost and the economics of measurement appear from nowhere. The tolerance band is 20 µm wide; two guard bands of 1.9 µm consume 3.8 µm — 19% of the tolerance is spent on measurement uncertainty before manufacturing gets a say. If the measurement uncertainty were 5 µm, half the band would be gone. This is why the old gauge-maker's rule of thumb wanted the measurement roughly four to ten times finer than the tolerance it polices, and why the process-capability world has whole frameworks — VDA 5, ISO 22514-7, and the gauge R&R studies of AIAG's MSA manual, each asking a related but distinct question — for deciding whether a measurement process is good enough for the tolerance it serves. Better measurement is not pedantry. It is purchased tolerance, handed back to production.
Everything so far treated a measurement as instrument-touches-part. Increasingly the number on your report is calculated by software from a scan: a structured-light sensor collects a few million points, and "the diameter" is whatever the evaluation pipeline distills from them. The measurement uncertainty machinery still applies — but the list of contributors changes character, because several of the biggest ones are now choices rather than physics:
Which exposes the trap in vendor datasheets. A scanner's "accuracy: 0.02 mm" — or, more honestly, its MPE from a standardized acceptance test — is a statement about the instrument's behavior measuring calibrated spheres and bars under test conditions. It is not the measurement uncertainty of your feature, on your part, with your spray, your sampling, your fitting rule and your alignment; it is at best one term among the several above. Task-specific measurement uncertainty has to be evaluated for the task — the clean route is measuring a calibrated part of similar geometry through your exact pipeline and watching what comes out (the approach ISO 15530-3 formalizes for CMMs, and VDA 5 generalizes), and working demonstrations of scan-based evaluation with checkable numbers live on the examples page. A single spec-sheet number standing in for all of that is not a measurement uncertainty statement. It is a hope.
A closing caution, because the measurement uncertainty budget's greatest danger is looking finished. The budget in Section 7 is not the uncertainty of the measurement; it is the uncertainty of the model of the measurement that its author thought to write down. Nothing in the quadrature machinery can represent a term that was never listed. And the terms that go unlisted are, reliably, the embarrassing ones:
The budget is not helpless against this — it is falsifiable, which is its quiet superpower. The budget makes a prediction: two competent measurements of the same part should agree within their combined uncertainties, most of the time. So test it. Compare operators, compare instruments, compare your lab against a calibration lab's value on a known part. When disagreements run larger than the budget predicts, the budget is missing a term, and the size of the surprise tells you roughly how big the missing term is. The practitioners' version of this whole section fits in one sentence: the biggest term in your measurement uncertainty budget is usually the one you left out — and the discipline of writing the budget down, then checking it against reality, is the only reliable way to find it.
If your inspection reports carry numbers with no measurement uncertainty statement — or one nobody in the building can defend — that is a solvable problem.
This is the work I offer. Metrology Maven turns approved inspection methods into pipelines that run unattended — starting with a fixed-fee assessment, from $3,500. How engagements work →