Design of Experiments for Machine Vision Lighting
Use randomised, blocked, and factorial trials to identify lighting settings that maximise defect separation and robustness.
A team compares four lights by changing one setting at a time and keeping whichever image looks best. Unfortunately, angle interacts with polarization, intensity interacts with exposure, and the last trials occur after the lamp and part have warmed. A design of experiments can separate main effects, interactions, drift, and noise—if the response is an inspection metric rather than visual preference.
This is a vendor-neutral engineering method. The worked example is hypothetical and must be replaced by measurements from the real product, line, and risk assessment.
What you will learn
Identify the physical, optical, data, or process limit behind design of experiments for vision lighting.
Convert the inspection need into measurable acceptance criteria.
Compare practical architectures and their trade-offs.
Commission the method using repeatable evidence.
Validate the final system under representative production variation.
Technical foundation
DOE estimates effects efficiently
A factorial design varies factors together so main effects and interactions can be estimated from planned combinations. Fractional designs reduce runs for screening but deliberately confound some effects, which must be understood before conclusions.
Randomisation and blocking address nuisance variation
Random run order helps distribute unknown time trends. Blocking groups known nuisance conditions such as material lot, day, camera, or operator so their variation is not mistaken for lighting effects.
The response must represent the decision
Mean image brightness is rarely enough. Responses can include defect-to-background contrast, separation relative to variation, saturation fraction, spatial uniformity, feature error, invalid rate, false accepts, false rejects, and cycle time. Multiple responses may require constraints or a stated desirability rule.
Related guides on this publication: Machine Vision Lighting for Reflective and Shiny Parts and Flat-Field Correction for Machine Vision and Machine Vision Threshold Selection with Precision–Recall Curves.
Engineering workflow
1. State objective and responses
Evaluate: defect classes, acceptable backgrounds, minimum separation, saturation, uniformity, measurement bias, false decisions, and cycle time.
Why it matters: the experiment optimises whatever is measured.
Measure or calculate: predefine response formulas and acceptance constraints from inspection risk. Preserve settings, sample identity, operating state, and the calculation method so alternatives remain comparable.
Trade-off: part-level error is meaningful but needs more samples than image contrast. Common failure: choosing the winning image by eye after the runs.
2. Select factors and safe ranges
Evaluate: light geometry, angle, distance, diffuser, polarization, wavelength, intensity, pulse, exposure, aperture, gain, and part pose.
Why it matters: effects and interactions exist only within the tested ranges.
Measure or calculate: use engineering trials to set feasible low and high levels without clipping or unsafe operation. Preserve settings, sample identity, operating state, and the calculation method so alternatives remain comparable.
Trade-off: wide ranges reveal effects but may include unrealistic settings. Common failure: including exposure as a hidden automatic variable.
3. Choose design and blocks
Evaluate: comparative, screening, full or fractional factorial, centre points, response surface, replications, material lots, and days.
Why it matters: design resolution determines which effects can be distinguished.
Measure or calculate: document the design matrix, alias structure, randomisation, and sample allocation. Preserve settings, sample identity, operating state, and the calculation method so alternatives remain comparable.
Trade-off: fractional screening saves runs but confounds interactions. Common failure: selecting a fraction without checking likely optical interactions.
4. Execute controlled randomised trials
Evaluate: warm-up, fixed camera settings, coded samples, run order, cleaning, repositioning, environmental record, and raw images.
Why it matters: execution drift can dominate calculated effects.
Measure or calculate: follow the random schedule, record deviations, and repeat physical presentations. Preserve settings, sample identity, operating state, and the calculation method so alternatives remain comparable.
Trade-off: strict control improves inference but takes line time. Common failure: running all low settings first and high settings last.
5. Analyse effects and interactions
Evaluate: effect plots, residuals, variance, block effects, interaction plots, model adequacy, outliers, and practical magnitude.
Why it matters: statistical significance alone does not prove engineering value.
Measure or calculate: inspect residual structure and calculate predicted response with uncertainty. Preserve settings, sample identity, operating state, and the calculation method so alternatives remain comparable.
Trade-off: richer models capture curvature but need more data. Common failure: dropping inconvenient runs without documented cause.
6. Confirm and stress the selected region
Evaluate: confirmation runs, settings tolerance, acceptable extremes, defect boundary, hot and cold state, ageing, and production duration.
Why it matters: an optimum point may be fragile or overfit to the experiment.
Measure or calculate: run independent samples at the selected setting and deliberate perturbations around it. Preserve settings, sample identity, operating state, and the calculation method so alternatives remain comparable.
Trade-off: a robust plateau may sacrifice peak contrast but improve production reliability. Common failure: releasing the model optimum without confirmation.
Worked example
Hypothetical 2³ factorial screen: Factors are light angle A, polarizer B, and exposure C at low/high levels, with two replicates of eight combinations. Average defect-separation score is 1.8 at low A and 3.0 at high A. The standard convention for the high-minus-low main effect gives:
Main effect of angle A = mean(high A) - mean(low A)
Main effect A = 3.0 - 1.8 = +1.2 score units
Total image sets = 2³ combinations × 2 replicates = 16
The +1.2 effect applies only over the chosen factor ranges and does not show whether A interacts with the polarizer. The interaction contrast, residuals, material-lot blocks, saturation constraint, and independent confirmation parts must be reviewed before selecting high angle.
Practical decision aid
| Design | Best use | Important limitation |
|---|---|---|
| Paired comparison | one primary lighting change | does not map multiple interactions |
| Full factorial | main effects and all interactions | run count grows quickly |
| Fractional factorial | screen many factors efficiently | selected effects are confounded |
| Centre points | detect curvature in numeric factors | do not identify the full curved surface |
| Response surface | optimise important numeric factors | requires a valid local model |
| Randomised block | separate known nuisance sources | factor and block interactions may remain |
Use the table to choose the next controlled experiment, not as a universal product recommendation. A component or algorithm is acceptable only when the complete inspection cell meets pre-agreed technical and operational criteria.
Common mistakes and how to prevent them
One-factor-at-a-time testing. interactions remain invisible. Prevent it by using factorial combinations.
Response is visual preference. choice is subjective and weakly linked to risk. Prevent it by predefining inspection metrics.
No randomisation. warm-up or drift aliases with settings. Prevent it by randomising within practical blocks.
Camera auto settings enabled. lighting factors are silently compensated. Prevent it by locking exposure gain and processing.
No replicate physical presentations. repeatability and residual variance are unknown. Prevent it by including genuine replicates.
No confirmation run. the selected optimum is overfit. Prevent it by testing new parts and perturbations.
Validate under production conditions
After analysing the planned experiment, run independent confirmation samples from new lots and defect severities at the selected settings. Perturb angle, distance, intensity, exposure, part pose, focus, polarization alignment, ambient light, temperature, and lamp warm-up within realistic limits. Compare predictions with observed contrast, saturation, measurement, false decisions, and timing. Preserve the design matrix, raw images, deviations, model, residuals, and approval rationale.
Use representative acceptable parts, confirmed defects, boundary samples, and nuisance variation. Repeat complete part presentations rather than processing one stored image many times. Include start-up, warm-up, maximum speed, changeover, maintenance, environmental limits, communication faults, and long-duration operation where relevant.
Define acceptance criteria before reviewing final results. Preserve raw counts and denominators for false accepts, false rejects, invalid acquisitions, timing overruns, and manually reviewed cases. After release, trend leading indicators and conduct labelled audits so deterioration is detected before a customer escape.
Key takeaways
Optimise a defined inspection response, not image appearance.
Use factorial designs to reveal lighting interactions.
Randomise runs and block known nuisance sources.
Inspect practical effect size and residuals, not p-values alone.
Confirm the selected settings on new samples and perturbations.
Follow this Hashnode blog for more practical industrial machine-vision engineering, and connect with Kivanc Ekici on LinkedIn. For related machine-vision and automation information, visit ITAGE.
Frequently asked questions
Why use DOE instead of changing one lighting setting at a time?
DOE can estimate interactions and separates planned factor effects more efficiently, while one-factor-at-a-time trials can miss combinations that matter.
What should the response variable be?
Use metrics tied to the decision: class separation, feature bias, saturation, uniformity, false decisions, invalid rate, and timing as appropriate.
When is a fractional factorial suitable?
For screening many factors when higher-order interactions can be defensibly confounded and the alias structure is understood.
Why randomise the trial order?
It reduces the chance that warm-up, drift, contamination, or other time trends are mistaken for factor effects.
What is a confirmation run?
An independent test at the selected settings, preferably with new samples and deliberate nuisance perturbations, used to verify predicted performance.

