Question it answers
Which factors affect the response, and what settings are best?
Data needed
Responses from a planned, randomized, replicated factorial design
Key output
Effects, interactions, an ANOVA table, a model, and best settings
Design
2k full factorial; coded −1/+1
Assumptions
Independent, normal residuals with constant spread
Excel
Coded columns, AVERAGEIF, Regression tool
Minitab
Stat > DOE > Factorial > Create and Analyze Factorial Design
Why it matters
It finds interactions and the best settings with few runs

The Idea in Plain Language

A designed experiment changes several inputs on purpose, in a planned pattern, and measures the output. Analyzing it means answering three questions: which factors matter (main effects), do factors depend on each other (interactions), and what settings give the best result.

The simplest and most useful design is a two-level full factorial: every combination of the high and low settings of each factor. With three factors that is 2 × 2 × 2 = 8 runs. Coding the levels as −1 (low) and +1 (high) makes the arithmetic easy, and the same arithmetic scales to any number of factors.

TermMeaningCalculation
Main effectAverage change in the response when a factor goes from low to highMean at high − mean at low
InteractionThe effect of one factor depends on the level of anotherHalf the difference between the effect of A at high B and at low B
ReplicateA repeat of the same settings to measure pure errorNeeded to test significance
Coefficient (coded)Half the effect: the change per coded unitEffect / 2
Why it matters. One-factor-at-a-time testing cannot find interactions and needs more runs for the same precision. A factorial design uses every run to estimate every effect.

Worked Example: A Three-Factor Experiment

A chemical process was tested at two levels of each of three factors, with two replicates of every combination (16 runs, in random order). The response is yield (%).

FactorLow (−1)High (+1)
Temperature160°C180°C
Pressure20 bar30 bar
Catalyst1.0 %2.0 %
RunABCYield 1Yield 2Mean
1-1-1-173.269.871.50
2+1-1-170.177.273.65
3-1+1-169.769.569.60
4+1+1-178.178.478.25
5-1-1+174.374.574.40
6+1-1+176.976.476.65
7-1+1+170.871.971.35
8+1+1+182.582.382.40
  1. Main effect of temperature (A): mean of the four high-temperature cells = 77.74; mean of the four low = 71.71; effect = 77.74 − 71.71 = 6.03 points.
  2. Other main effects the same way: pressure 1.35, catalyst 2.95.
  3. Interaction A × B: multiply the A and B sign columns to get a new column, then take mean(+1) − mean(−1) = 3.83. Other interactions: A × C 0.62, B × C 0.00, A × B × C 0.58.
  4. Sum of squares for a term = N × effect² / 4 = 16 × 6.03² / 4 = 145.20 for temperature. The error sum of squares is 31.82 on 8 df (pure error from the replicates), so MSE = 3.978 and s = 1.994.
  5. Test each term: F = SS / MSE, so for temperature F = 145.20 / 3.978 = 36.5 with 1 and 8 df, p < 0.001.
  6. Standard error of an effect = 2s / √N = 0.997. An effect is significant at 5% if it is larger than t × SE = 2.306 × 0.997 = 2.30.
SourceEffectSum of squaresFP-value
Temperature (A)+6.03145.2036.51< 0.001
Pressure (B)+1.357.291.830.213
Catalyst (C)+2.9534.818.750.018
A × B+3.8358.5214.710.005
A × C+0.621.560.390.548
B × C+0.000.000.001.000
A × B × C+0.581.320.330.580
Error (pure)31.82
Total280.53
Temperature (A) 6.03 A x B 3.83 Catalyst (C) 2.95 Pressure (B) 1.35 A x C 0.62 A x B x C 0.58 B x C 0.00 Significance line 2.30 Absolute effect on yield (percentage points)
Effects larger than the dashed line are statistically significant. Everything below it is indistinguishable from noise.
Conclusion. Temperature (A), Catalyst (C), A × B are significant at 5%. The A × B interaction means the effect of temperature depends on pressure, so pressure matters even though its own effect is small. The other terms are small enough to treat as noise and pool into the error.

Reading the Plots

Main effects plot. The steeper the line, the bigger the effect. Temperature and catalyst raise the yield when moved from low to high; pressure has little effect on its own.

71 72 73 74 75 76 77 78 79 80 Temperature Pressure Catalyst Low level High level Mean yield (%)
A steep line means a large main effect, and a flat line means a small one.

Interaction plot. Parallel lines mean no interaction. Lines that are not parallel mean the effect of one factor depends on the other.

70 72 74 76 78 80 Pressure 20 bar Pressure 30 bar 160&deg;C 180&deg;C Temperature Mean yield (%)
Raising the temperature helps much more at 30 bar than at 20 bar, so the lines are not parallel.
Interaction trap. When an interaction is significant, do not interpret the main effects alone. Pressure has almost no average effect, yet it matters because it changes how temperature performs.

A Prediction Model and Checking It

Dropping the terms that are not significant (but keeping a main effect when its interaction is in the model) gives a reduced model with R² = 87.6% (adjusted 83.1%) and s = 1.78.

Yield = 74.73 + 3.01·A + 0.67·B + 1.47·C + 1.91·A×B (A, B, C coded −1 to +1)

Before trusting the model, check the residuals, exactly as for ANOVA and regression:

-2 -1 0 1 2 Normal score (expected z) Normal probability plot of residuals Residual (percentage points)
Residuals follow a straight line (Shapiro-Wilk p = 0.28).
70 73 76 79 Fitted yield (%) Residuals versus fitted values Residual (points)
No pattern and a steady spread across the fitted values.

Best settings. The model predicts the highest yield at temperature high, pressure high, catalyst high: 81.8%, with a 95% confidence interval for the mean of 79.6 to 84.0 and a 95% prediction interval for one run of 77.3 to 86.3.

Confirm it. A prediction is a hypothesis. Run the best settings several times and check that the results fall inside the prediction interval before changing the process. Predictions only hold inside the range tested.

Run It in Excel and Minitab

ExcelStep by step

  1. Enter the design in coded form: columns A, B, C with −1 and +1, then compute interaction columns =A2*B2, =A2*C2, =B2*C2, =A2*B2*C2. Put the yield in the last column.
  2. Effects: =AVERAGEIF(A2:A17, 1, Y2:Y17) - AVERAGEIF(A2:A17, -1, Y2:Y17) gives the temperature effect (6.03). Repeat for each column.
  3. All at once: Data > Data Analysis > Regression with Y as the response and the seven coded columns as X. The coefficients are half the effects, and the ANOVA table and p-values come with them.
  4. Sum of squares for a term: =16*effect^2/4. Critical effect: =T.INV.2T(0.05, 8)*2*s/SQRT(16) (2.30).
  5. Plots: make a Line with Markers chart of the cell means for the interaction plot. Excel has no built-in factorial plots or Pareto of effects.

MinitabStep by step

  1. Plan: Stat > DOE > Factorial > Create Factorial Design. Choose 2-level factorial (default generators), 3 factors, full design, 2 replicates, and name the factors and levels. Minitab randomizes the run order.
  2. Enter the responses in the worksheet column.
  3. Analyze: Stat > DOE > Factorial > Analyze Factorial Design. Choose the response, and under Terms include up to three-way interactions. Under Graphs choose a Pareto of effects and the four-in-one residual plot.
  4. Plots: Stat > DOE > Factorial > Factorial Plots for main effects and interaction plots, and Stat > DOE > Factorial > Cube Plots for the cell means.
  5. Reduce the model: run Stat > DOE > Factorial > Analyze Factorial Design again with only the significant terms (Minitab also offers stepwise selection under Stepwise), then use Stat > DOE > Factorial > Response Optimizer to find the best settings.
  6. Predict: Stat > DOE > Factorial > Predict gives the fit and intervals for chosen settings.
Minitab session window: Analyze Factorial Design, full model (typed excerpt, simplified)
Factorial Regression: Yield versus Temperature, Pressure, Catalyst

Analysis of Variance

  Source                DF   Adj SS    Adj MS  F-Value  P-Value
  Model                  7   248.710    35.530     8.93    0.003
  Temperature (A)        1   145.203   145.203    36.51    0.000
  Pressure (B)           1     7.290     7.290     1.83    0.213
  Catalyst (C)           1    34.810    34.810     8.75    0.018
  A x B                  1    58.523    58.523    14.71    0.005
  A x C                  1     1.562     1.562     0.39    0.548
  B x C                  1     0.000     0.000     0.00    1.000
  A x B x C              1     1.323     1.323     0.33    0.580
  Error                  8    31.820     3.978
  Total                 15   280.530

Model Summary

      S    R-sq  R-sq(adj)
 1.9944  88.66%   78.73%

Coded Coefficients

  Term                    Coef  SE Coef  T-Value  P-Value
  Constant                   74.725    0.499   149.87  0.000
  Temperature (A)           3.012    0.499     6.04    0.000
  Pressure (B)              0.675    0.499     1.35    0.213
  Catalyst (C)              1.475    0.499     2.96    0.018
  A x B                     1.913    0.499     3.84    0.005
  A x C                     0.313    0.499     0.63    0.548
  B x C                     0.000    0.499     0.00    1.000
  A x B x C                 0.287    0.499     0.58    0.580

Reading and Reporting

  1. Look at the Pareto of effects and the p-values together, and name the factors that matter.
  2. Check interactions before main effects. A significant interaction changes the story.
  3. Check the residuals for normality, constant spread, and patterns in run order.
  4. Give the model, its R², and the range of settings it covers.
  5. Predict, then confirm. Run the recommended settings and compare with the interval.
A sentence you can use. A 23 factorial with two replicates showed that temperature (a), catalyst (c), a x b significantly affect yield (R² = 88%); the predicted best settings give 81.8% (95% prediction interval 77.3 to 86.3), to be confirmed with a verification run.

Common Mistakes

MistakeWhy it misleadsBetter
No replicates and no way to estimate errorCannot test significanceReplicate, or use center points, or the normal plot of effects
Running experiments in standard orderTime trends get mixed with factor effectsRandomize the run order
Interpreting a main effect when it is part of a significant interactionThe average hides opposite effectsRead the interaction plot
Using levels too close togetherThe effect is lost in the noiseChoose levels wide enough to matter but safe
Dropping a main effect whose interaction stays in the modelBreaks hierarchy and distorts coefficientsKeep the parents
Extrapolating beyond the tested levelsThe model has no information thereRun a new experiment in the new region

Try It Yourself

A two-factor experiment gave these mean responses: A low, B low = 20; A high, B low = 30; A low, B high = 25; A high, B high = 45.

  • Find the main effects of A and B.
  • Find the A × B interaction effect.
Show the answer

Effect of A = (30 + 45)/2 − (20 + 25)/2 = 37.5 − 22.5 = 15. Effect of B = (25 + 45)/2 − (20 + 30)/2 = 35 − 25 = 10.

Interaction = [(20 + 45)/2 − (30 + 25)/2] = 32.5 − 27.5 = 5. A is worth 10 units at low B but 20 at high B, so the effects reinforce each other.

Analyzing Designed Experiments: Frequently Asked Questions

What is a full factorial experiment?

An experiment that runs every combination of the levels of all the factors. A two-level design with k factors needs 2k runs and estimates all main effects and all interactions.

Why code the factors as &minus;1 and +1?

It puts every factor on the same scale, makes the effects easy to calculate and compare, and makes the columns orthogonal so the estimates do not interfere with each other.

What if I cannot replicate the experiment?

Use center points to estimate error and check curvature, or use a normal probability plot of the effects, where the real effects stand out from the line of small ones. Replication is better when you can afford it.

What is an interaction?

It means the effect of one factor depends on the level of another. On an interaction plot the lines are not parallel.

What do I do when the model fails the residual checks?

Look for outliers and run-order trends, consider transforming the response, and check that the experiment was run as planned. Curvature may call for adding center points and a response surface design.

How is analyzing a DOE different from ANOVA?

It is the same mathematics. A factorial ANOVA tests each term, and the DOE view adds coded coefficients, effect plots, and a prediction equation for finding the best settings.

Sources and Further Reading

  • Douglas C. Montgomery, Design and Analysis of Experiments, Wiley.
  • George E. P. Box, J. Stuart Hunter, and William G. Hunter, Statistics for Experimenters, Wiley.
  • NIST/SEMATECH, e-Handbook of Statistical Methods, Process Improvement (itl.nist.gov/div898/handbook).
  • Minitab Support, “Methods and formulas for Analyze Factorial Design” (support.minitab.com).

This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.