Attribute Agreement Analysis for Human and Vision Inspection
Measure repeatability, reproducibility, reference agreement, and class-specific risk before replacing or combining inspectors.
Three experienced inspectors and a vision system classify the same cosmetic defect differently. Overall agreement looks high because most parts are obvious passes, but the disputed boundary class contains nearly all customer risk. Attribute agreement analysis must expose within-appraiser, between-appraiser, and reference disagreement by class—not hide them in one percentage.
This is a vendor-neutral engineering method. The worked example is hypothetical and must be replaced by measurements from the real product, line, and risk assessment.
What you will learn
Identify the physical, optical, data, or process limit behind attribute agreement analysis for inspection.
Convert the inspection need into measurable acceptance criteria.
Compare practical architectures and their trade-offs.
Commission the method using repeatable evidence.
Validate the final system under representative production variation.
Technical foundation
Repeatability and reproducibility are distinct
Within-appraiser agreement asks whether the same inspector or system repeats its classification. Between-appraiser agreement asks whether different decision makers agree. Agreement to a reference asks whether decisions are correct relative to an accepted standard.
Percent agreement can hide prevalence effects
If 98% of samples are acceptable, an always-pass method achieves 98% overall agreement while missing every defect. Report confusion counts and class-specific sensitivity and specificity; kappa-style statistics adjust for chance agreement but are also affected by prevalence and interpretation.
The reference is another measurement system
A master classification may come from expert adjudication, destructive test, downstream function, or another instrument. Its definition, uncertainty, independence, and resolution process must be documented instead of treated as infallible.
Related guides on this publication: Machine Vision Threshold Selection with Precision–Recall Curves and Industrial Vision Dataset Design: Avoid Leakage and False Confidence and Machine Vision FAT and SAT: Build an Acceptance Test That Works.
Engineering workflow
1. Define categories and reference
Evaluate: pass, defect classes, severity, boundary rules, examples, adjudication, and intended action.
Why it matters: ambiguous category definitions create disagreement that no statistic fixes.
Measure or calculate: write operational definitions and qualify the reference process. Preserve settings, sample identity, operating state, and the calculation method so alternatives remain comparable.
Trade-off: finer categories support diagnosis but require more samples and expertise. Common failure: merging safety-critical and cosmetic defects into one fail class.
2. Select representative samples
Evaluate: class balance, boundary samples, nuisance variation, lots, positions, defect severities, duplicates, and sample stability.
Why it matters: production prevalence alone may provide too few critical defects to assess.
Measure or calculate: stratify the study and retain counts so prevalence-weighted operation can be estimated separately. Preserve settings, sample identity, operating state, and the calculation method so alternatives remain comparable.
Trade-off: enrichment improves class precision but overall agreement no longer reflects line prevalence directly. Common failure: using only obvious passes and failures.
3. Design blinded repeated trials
Evaluate: appraisers, automated recipes, random order, repetitions, washout, lighting, display, and independence.
Why it matters: memory and discussion inflate agreement.
Measure or calculate: randomise coded samples and separate repetitions sufficiently for the task. Preserve settings, sample identity, operating state, and the calculation method so alternatives remain comparable.
Trade-off: long washout reduces recall but extends the study. Common failure: letting inspectors see prior or machine results.
4. Calculate complementary metrics
Evaluate: within-appraiser agreement, between-appraiser agreement, reference agreement, confusion matrix, class recall, false decisions, kappa, and confidence intervals.
Why it matters: no single score describes all error modes.
Measure or calculate: report raw numerators and denominators by class plus uncertainty intervals. Preserve settings, sample identity, operating state, and the calculation method so alternatives remain comparable.
Trade-off: many metrics add interpretation work but prevent false reassurance. Common failure: publishing kappa without the confusion counts.
5. Investigate disagreement
Evaluate: sample images, defect morphology, category boundary, presentation, operator comments, machine scores, and reference review.
Why it matters: disagreement is evidence about the measurement process.
Measure or calculate: adjudicate without erasing original labels and code root causes. Preserve settings, sample identity, operating state, and the calculation method so alternatives remain comparable.
Trade-off: expert review improves understanding but can introduce authority bias. Common failure: changing the reference until the preferred system agrees.
6. Approve and monitor the process
Evaluate: acceptance criteria, training, recipe lock, requalification, audit sampling, drift, new defects, and escalation.
Why it matters: agreement at launch does not guarantee continued performance.
Measure or calculate: repeat blinded audits and trend class-specific errors and review rates. Preserve settings, sample identity, operating state, and the calculation method so alternatives remain comparable.
Trade-off: frequent audits cost labour but detect category or process drift. Common failure: monitoring only total agreement after release.
Worked example
Hypothetical binary study: Forty stable samples include 20 reference passes and 20 reference fails. One vision recipe classifies each sample twice. It agrees with itself on 38 of 40 repeated pairs. Against the reference on the first trial it passes 18 of 20 passes and detects 17 of 20 fails.
Repeatability = 38 / 40 = 95.0%
Specificity to reference = 18 / 20 = 90.0%
Defect recall to reference = 17 / 20 = 85.0%
Overall reference agreement = (18 + 17) / 40 = 87.5%
These values are descriptive, not automatic acceptance. Confidence intervals are wide for only 20 samples per class, and the enriched 50% defect mix differs from production prevalence. The five reference disagreements and two repeatability failures require review by defect type and boundary severity.
Practical decision aid
| Result | Question answered | Frequent misinterpretation |
|---|---|---|
| Within-appraiser agreement | does one decision maker repeat? | high repeatability does not prove correctness |
| Between-appraiser agreement | do decision makers classify alike? | consensus can still be biased |
| Agreement to reference | does each method match the standard? | reference may be uncertain |
| Class recall | which defects are detected? | must state denominator by class |
| Specificity | how often good parts are accepted? | does not quantify defect escapes |
| Kappa statistic | agreement beyond chance model | not a substitute for raw counts and risk |
Use the table to choose the next controlled experiment, not as a universal product recommendation. A component or algorithm is acceptable only when the complete inspection cell meets pre-agreed technical and operational criteria.
Common mistakes and how to prevent them
Study dominated by easy passes. overall agreement conceals defect failure. Prevent it by stratifying classes and boundary cases.
No repeated presentations. repeatability cannot be estimated. Prevent it by including randomised blind repeats.
Unstable physical samples. the measurand changes between trials. Prevent it by using stable artefacts or controlled imagery appropriately.
Reference chosen by one opinion. bias is treated as truth. Prevent it by documenting independent adjudication.
Showing machine outputs to people. human decisions become anchored. Prevent it by blinding methods during collection.
Collapsing all defects together. critical class weakness disappears. Prevent it by reporting class-specific confusion counts.
Validate under production conditions
Run a pre-agreed blinded study with stable representative samples across defect classes, severity boundaries, acceptable extremes, lots, appearance nuisance, and presentation conditions. Include repeated randomised trials for every appraiser and the locked vision recipe. Preserve all original classifications, raw counts, confidence intervals, and adjudication outcomes. After release, conduct periodic blind audits on production-prevalence samples plus enriched rare-defect challenge sets.
Use representative acceptable parts, confirmed defects, boundary samples, and nuisance variation. Repeat complete part presentations rather than processing one stored image many times. Include start-up, warm-up, maximum speed, changeover, maintenance, environmental limits, communication faults, and long-duration operation where relevant.
Define acceptance criteria before reviewing final results. Preserve raw counts and denominators for false accepts, false rejects, invalid acquisitions, timing overruns, and manually reviewed cases. After release, trend leading indicators and conduct labelled audits so deterioration is detected before a customer escape.
Key takeaways
Separate repeatability, reproducibility, and reference agreement.
Balance the study for diagnostic power but retain real prevalence context.
Report class-specific counts and errors alongside kappa.
Treat the reference classification as a qualified measurement process.
Preserve blinded repeats and ongoing audit evidence.
Follow this Hashnode blog for more practical industrial machine-vision engineering, and connect with Kivanc Ekici on LinkedIn. To explore relevant technology information, visit Algomedi.
Frequently asked questions
What is attribute agreement analysis?
It evaluates consistency of categorical inspections within appraisers, between appraisers, and against a reference classification.
Is kappa better than percent agreement?
It adds a chance-agreement model, but it is affected by prevalence and must be interpreted with raw confusion counts and class risks.
How many samples are needed?
There is no universal count. Size each critical class and boundary subgroup to achieve useful uncertainty on the error rates, not merely a total sample target.
Can stored images be used for the study?
They isolate decision consistency, but a complete system study also needs repeated physical presentations to include acquisition and handling variation.
How should human and machine disagreements be resolved?
Use a documented independent adjudication process, preserve original decisions, and investigate whether the issue is definition, reference, presentation, or algorithm.

