GaugeConnection
Are you a quality consultant? Get a public profile, client referrals, and free access to all quality tools.
Learn more

Attribute Agreement Analysis

Table of Contents

Variable data versus attribute data

The distinction determines which study you can run, and it is decided by what the inspection actually records, not by what is being inspected.

Aspect Variable data Attribute data
What is recorded A measured value on a continuous scale, such as 25.012 mm A category, such as pass or fail, or a grade such as A, B or C
Applicable study Gauge R and R, bias, linearity, stability Attribute agreement analysis
Headline result Percent GRR and ndc Percent agreement and kappa

A useful test: if you cannot meaningfully subtract one result from another, the data is attribute data. Pass minus fail is not a quantity, so no variance can be decomposed and the entire Gauge R and R machinery does not apply.

Why you cannot just run a Gauge R and R

Gauge R and R works by splitting the total observed variance into part variation, repeatability and reproducibility. Every one of those terms is a variance, and a variance requires a numeric scale with meaningful distances between values. Category labels have no such scale.

People sometimes code pass as 1 and fail as 0 and feed that into an R and R calculation. Do not do this. The arithmetic will produce numbers, and those numbers will be meaningless, because the distance between the codes is an artifact of the coding rather than a property of the measurement. An auditor who understands the method will recognize it immediately.

The three comparisons

A complete attribute study has each appraiser inspect the same set of parts more than once, in a randomized order, without knowing the previous result. That design supports three distinct questions.

Within appraiser

Does one appraiser reach the same verdict when shown the same part again? This is the attribute equivalent of repeatability. Poor within-appraiser agreement means the criteria are ambiguous or the defect is near the limit of what can be seen.

Between appraisers

Do different appraisers reach the same verdict on the same part? This is the attribute equivalent of reproducibility. Poor between-appraiser agreement usually points at inconsistent training or an unwritten standard that each person has interpreted differently.

Versus the standard

Do the appraisers reach the correct verdict, judged against an independently established reference decision for each part? This is the one that matters most, and the one that requires the most preparation.

Kappa: agreement corrected for chance

Raw percent agreement overstates how good a process is, because two people guessing randomly will still agree a large fraction of the time. With only two categories, pure guessing produces about 50 percent agreement. Cohen's kappa removes that baseline.

Cohen's kappa

κ = (Po − Pe) ÷ (1 − Pe)

where Po is the observed proportion of agreement and Pe is the proportion of agreement expected by chance alone, calculated from how often each appraiser uses each category.

The numerator is the agreement above chance actually achieved. The denominator is the agreement above chance that was available to achieve. Kappa is therefore the fraction of the achievable improvement over guessing that the appraisers captured. Kappa of 1 is perfect agreement and kappa of 0 is no better than chance. Negative values are possible and indicate systematic disagreement.

Worked example

Illustrative data. Two appraisers each classify the same 50 parts as pass or fail.

  • They reach the same verdict on 45 of the 50 parts, so Po = 45 / 50 = 0.90.
  • Both appraisers call about 80 percent of parts pass and 20 percent fail, so the chance agreement is Pe = (0.80 x 0.80) + (0.20 x 0.20) = 0.64 + 0.04 = 0.68.
  • kappa = (0.90 - 0.68) / (1 - 0.68) = 0.22 / 0.32 = 0.69.

Ninety percent agreement sounds strong. Kappa of 0.69 tells a different and more honest story: most of that agreement was available for free because the process rejects few parts. This gap between raw agreement and kappa is exactly why the correction exists, and it grows wider as the categories become more unbalanced.

Interpreting kappa

Kappa Common interpretation
κ ≥ 0.90Good agreement. The inspection is generally considered acceptable on this criterion.
0.70 ≤ κ < 0.90Marginal. Usually treated as needing improvement to the criteria, the training, or the inspection conditions.
κ < 0.70Unacceptable. The inspection is not reliably distinguishing conforming from non-conforming parts.

Treat a low kappa as a process finding rather than a people finding. It almost always traces back to criteria that were never written down precisely, boundary samples that do not exist, or lighting and fixturing that make the defect genuinely hard to see. Retraining without fixing those conditions rarely moves the number.

What you can do in GaugeConnection today

Until that changes, the practical steps are:

  • Run the study outside this tool against your own documented procedure, and retain the records where your quality system requires them.
  • Before running it, write down the acceptance criteria and produce physical boundary samples. Most attribute studies fail because the standard was never made concrete, and no statistic will repair that.
  • Where an attribute check can be replaced by a measurement, consider doing so. Variable data gives you far more information per part and supports the full range of MSA studies.
The one thing to remember
If your inspection produces a verdict rather than a number, a Gauge R and R is not the wrong answer, it is not an available answer. Use attribute agreement analysis, and read kappa rather than raw percent agreement.

Rejoining the server...

Rejoin failed... trying again in seconds.

Failed to rejoin.
Please retry or reload the page.

The session has been paused by the server.

Failed to resume the session.
Please retry or reload the page.