- Question it answers
- Which factors affect the response, and what settings are best?
- Data needed
- Responses from a planned, randomized, replicated factorial design
- Key output
- Effects, interactions, an ANOVA table, a model, and best settings
- Design
- 2k full factorial; coded −1/+1
- Assumptions
- Independent, normal residuals with constant spread
- Excel
- Coded columns, AVERAGEIF, Regression tool
- Minitab
- Stat > DOE > Factorial > Create and Analyze Factorial Design
- Why it matters
- It finds interactions and the best settings with few runs
The Idea in Plain Language
A designed experiment changes several inputs on purpose, in a planned pattern, and measures the output. Analyzing it means answering three questions: which factors matter (main effects), do factors depend on each other (interactions), and what settings give the best result.
The simplest and most useful design is a two-level full factorial: every combination of the high and low settings of each factor. With three factors that is 2 × 2 × 2 = 8 runs. Coding the levels as −1 (low) and +1 (high) makes the arithmetic easy, and the same arithmetic scales to any number of factors.
| Term | Meaning | Calculation |
|---|---|---|
| Main effect | Average change in the response when a factor goes from low to high | Mean at high − mean at low |
| Interaction | The effect of one factor depends on the level of another | Half the difference between the effect of A at high B and at low B |
| Replicate | A repeat of the same settings to measure pure error | Needed to test significance |
| Coefficient (coded) | Half the effect: the change per coded unit | Effect / 2 |
Worked Example: A Three-Factor Experiment
A chemical process was tested at two levels of each of three factors, with two replicates of every combination (16 runs, in random order). The response is yield (%).
| Factor | Low (−1) | High (+1) |
|---|---|---|
| Temperature | 160°C | 180°C |
| Pressure | 20 bar | 30 bar |
| Catalyst | 1.0 % | 2.0 % |
| Run | A | B | C | Yield 1 | Yield 2 | Mean |
|---|---|---|---|---|---|---|
| 1 | -1 | -1 | -1 | 73.2 | 69.8 | 71.50 |
| 2 | +1 | -1 | -1 | 70.1 | 77.2 | 73.65 |
| 3 | -1 | +1 | -1 | 69.7 | 69.5 | 69.60 |
| 4 | +1 | +1 | -1 | 78.1 | 78.4 | 78.25 |
| 5 | -1 | -1 | +1 | 74.3 | 74.5 | 74.40 |
| 6 | +1 | -1 | +1 | 76.9 | 76.4 | 76.65 |
| 7 | -1 | +1 | +1 | 70.8 | 71.9 | 71.35 |
| 8 | +1 | +1 | +1 | 82.5 | 82.3 | 82.40 |
- Main effect of temperature (A): mean of the four high-temperature cells = 77.74; mean of the four low = 71.71; effect = 77.74 − 71.71 = 6.03 points.
- Other main effects the same way: pressure 1.35, catalyst 2.95.
- Interaction A × B: multiply the A and B sign columns to get a new column, then take mean(+1) − mean(−1) = 3.83. Other interactions: A × C 0.62, B × C 0.00, A × B × C 0.58.
- Sum of squares for a term = N × effect² / 4 = 16 × 6.03² / 4 = 145.20 for temperature. The error sum of squares is 31.82 on 8 df (pure error from the replicates), so MSE = 3.978 and s = 1.994.
- Test each term: F = SS / MSE, so for temperature F = 145.20 / 3.978 = 36.5 with 1 and 8 df, p < 0.001.
- Standard error of an effect = 2s / √N = 0.997. An effect is significant at 5% if it is larger than t × SE = 2.306 × 0.997 = 2.30.
| Source | Effect | Sum of squares | F | P-value |
|---|---|---|---|---|
| Temperature (A) | +6.03 | 145.20 | 36.51 | < 0.001 |
| Pressure (B) | +1.35 | 7.29 | 1.83 | 0.213 |
| Catalyst (C) | +2.95 | 34.81 | 8.75 | 0.018 |
| A × B | +3.83 | 58.52 | 14.71 | 0.005 |
| A × C | +0.62 | 1.56 | 0.39 | 0.548 |
| B × C | +0.00 | 0.00 | 0.00 | 1.000 |
| A × B × C | +0.58 | 1.32 | 0.33 | 0.580 |
| Error (pure) | 31.82 | |||
| Total | 280.53 |
Reading the Plots
Main effects plot. The steeper the line, the bigger the effect. Temperature and catalyst raise the yield when moved from low to high; pressure has little effect on its own.
Interaction plot. Parallel lines mean no interaction. Lines that are not parallel mean the effect of one factor depends on the other.
A Prediction Model and Checking It
Dropping the terms that are not significant (but keeping a main effect when its interaction is in the model) gives a reduced model with R² = 87.6% (adjusted 83.1%) and s = 1.78.
Before trusting the model, check the residuals, exactly as for ANOVA and regression:
Best settings. The model predicts the highest yield at temperature high, pressure high, catalyst high: 81.8%, with a 95% confidence interval for the mean of 79.6 to 84.0 and a 95% prediction interval for one run of 77.3 to 86.3.
Run It in Excel and Minitab
ExcelStep by step
- Enter the design in coded form: columns A, B, C with −1 and +1, then compute interaction columns , , , . Put the yield in the last column.
- Effects: gives the temperature effect (6.03). Repeat for each column.
- All at once: with Y as the response and the seven coded columns as X. The coefficients are half the effects, and the ANOVA table and p-values come with them.
- Sum of squares for a term: . Critical effect: (2.30).
- Plots: make a Line with Markers chart of the cell means for the interaction plot. Excel has no built-in factorial plots or Pareto of effects.
MinitabStep by step
- Plan: . Choose 2-level factorial (default generators), 3 factors, full design, 2 replicates, and name the factors and levels. Minitab randomizes the run order.
- Enter the responses in the worksheet column.
- Analyze: . Choose the response, and under Terms include up to three-way interactions. Under Graphs choose a Pareto of effects and the four-in-one residual plot.
- Plots: for main effects and interaction plots, and for the cell means.
- Reduce the model: run again with only the significant terms (Minitab also offers stepwise selection under Stepwise), then use to find the best settings.
- Predict: gives the fit and intervals for chosen settings.
Factorial Regression: Yield versus Temperature, Pressure, Catalyst
Analysis of Variance
Source DF Adj SS Adj MS F-Value P-Value
Model 7 248.710 35.530 8.93 0.003
Temperature (A) 1 145.203 145.203 36.51 0.000
Pressure (B) 1 7.290 7.290 1.83 0.213
Catalyst (C) 1 34.810 34.810 8.75 0.018
A x B 1 58.523 58.523 14.71 0.005
A x C 1 1.562 1.562 0.39 0.548
B x C 1 0.000 0.000 0.00 1.000
A x B x C 1 1.323 1.323 0.33 0.580
Error 8 31.820 3.978
Total 15 280.530
Model Summary
S R-sq R-sq(adj)
1.9944 88.66% 78.73%
Coded Coefficients
Term Coef SE Coef T-Value P-Value
Constant 74.725 0.499 149.87 0.000
Temperature (A) 3.012 0.499 6.04 0.000
Pressure (B) 0.675 0.499 1.35 0.213
Catalyst (C) 1.475 0.499 2.96 0.018
A x B 1.913 0.499 3.84 0.005
A x C 0.313 0.499 0.63 0.548
B x C 0.000 0.499 0.00 1.000
A x B x C 0.287 0.499 0.58 0.580
Reading and Reporting
- Look at the Pareto of effects and the p-values together, and name the factors that matter.
- Check interactions before main effects. A significant interaction changes the story.
- Check the residuals for normality, constant spread, and patterns in run order.
- Give the model, its R², and the range of settings it covers.
- Predict, then confirm. Run the recommended settings and compare with the interval.
Common Mistakes
| Mistake | Why it misleads | Better |
|---|---|---|
| No replicates and no way to estimate error | Cannot test significance | Replicate, or use center points, or the normal plot of effects |
| Running experiments in standard order | Time trends get mixed with factor effects | Randomize the run order |
| Interpreting a main effect when it is part of a significant interaction | The average hides opposite effects | Read the interaction plot |
| Using levels too close together | The effect is lost in the noise | Choose levels wide enough to matter but safe |
| Dropping a main effect whose interaction stays in the model | Breaks hierarchy and distorts coefficients | Keep the parents |
| Extrapolating beyond the tested levels | The model has no information there | Run a new experiment in the new region |
Try It Yourself
A two-factor experiment gave these mean responses: A low, B low = 20; A high, B low = 30; A low, B high = 25; A high, B high = 45.
- Find the main effects of A and B.
- Find the A × B interaction effect.
Show the answer
Effect of A = (30 + 45)/2 − (20 + 25)/2 = 37.5 − 22.5 = 15. Effect of B = (25 + 45)/2 − (20 + 30)/2 = 35 − 25 = 10.
Interaction = [(20 + 45)/2 − (30 + 25)/2] = 32.5 − 27.5 = 5. A is worth 10 units at low B but 20 at high B, so the effects reinforce each other.
Analyzing Designed Experiments: Frequently Asked Questions
What is a full factorial experiment?
An experiment that runs every combination of the levels of all the factors. A two-level design with k factors needs 2k runs and estimates all main effects and all interactions.
Why code the factors as −1 and +1?
It puts every factor on the same scale, makes the effects easy to calculate and compare, and makes the columns orthogonal so the estimates do not interfere with each other.
What if I cannot replicate the experiment?
Use center points to estimate error and check curvature, or use a normal probability plot of the effects, where the real effects stand out from the line of small ones. Replication is better when you can afford it.
What is an interaction?
It means the effect of one factor depends on the level of another. On an interaction plot the lines are not parallel.
What do I do when the model fails the residual checks?
Look for outliers and run-order trends, consider transforming the response, and check that the experiment was run as planned. Curvature may call for adding center points and a response surface design.
How is analyzing a DOE different from ANOVA?
It is the same mathematics. A factorial ANOVA tests each term, and the DOE view adds coded coefficients, effect plots, and a prediction equation for finding the best settings.
Sources and Further Reading
- Douglas C. Montgomery, Design and Analysis of Experiments, Wiley.
- George E. P. Box, J. Stuart Hunter, and William G. Hunter, Statistics for Experimenters, Wiley.
- NIST/SEMATECH, e-Handbook of Statistical Methods, Process Improvement (itl.nist.gov/div898/handbook).
- Minitab Support, “Methods and formulas for Analyze Factorial Design” (support.minitab.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.