- Question it answers
- Does each of two factors change the average, and does one factor's effect depend on the other?
- Data needed
- One continuous response, two categorical factors, and replicates in every combination
- Key output
- F and p for each factor and for the interaction; R-squared
- Read first
- The interaction. If it is significant, do not read the main effects alone
- Assumptions
- Independent observations, normal residuals, similar spread in every cell, ideally balanced
- Excel
- Data Analysis > Anova: Two-Factor With Replication
- Minitab
- Stat > ANOVA > General Linear Model > Fit General Linear Model
- Prerequisite
- One-way ANOVA
The Idea in Plain Language
The one-way ANOVA page compared three filling heads. In a real plant, the heads run on more than one shift, and the shift might matter too. A two-way ANOVA studies two factors in the same experiment and answers three questions at once:
- Does the first factor (head) change the average fill weight?
- Does the second factor (shift) change it?
- Does the effect of one factor depend on the level of the other? This is the interaction.
The interaction is the reason to use a two-way design. Two one-way analyses cannot see it. Many real problems are interactions: a setting that works on one machine and fails on another, a method that helps one shift and not the other, a material that behaves differently at high temperature.
When to Use Two-Way ANOVA
| Your situation | Use | Why |
|---|---|---|
| One continuous response and two categorical factors, with several measurements in each combination | Two-way ANOVA | Tests both factors and their interaction |
| Only one factor | One-way ANOVA | Simpler, with the same logic |
| Three or more factors, or factors with numeric levels you set on purpose | Factorial design analysis (DOE) | See the Design of Experiments guide |
| A nuisance factor you cannot control (batch, day) but want to remove | Randomized block design | A two-way model without an interaction term, with the block as the second factor |
| A numeric input such as temperature and a category such as machine | General linear model with a covariate | Fit in Minitab’s General Linear Model |
| Data are not normal and groups are small | Transform the response, or use a general linear model with care | There is no simple rank-based two-way test in Excel or Minitab’s basic menu |
How It Works
The model says each observation is the grand mean, plus an effect of the first factor, plus an effect of the second, plus an interaction effect, plus random error:
| Source | Sum of squares (balanced design) | Degrees of freedom | F |
|---|---|---|---|
| Factor A (a levels) | b n Σ (meani − grand mean)² | a − 1 | MSA / MSE |
| Factor B (b levels) | a n Σ (meanj − grand mean)² | b − 1 | MSB / MSE |
| Interaction A × B | n Σ (cell mean − row mean − column mean + grand mean)² | (a − 1)(b − 1) | MSAB / MSE |
| Error | Σ (value − cell mean)² | ab(n − 1) | |
| Total | Σ (value − grand mean)² | abn − 1 |
Here a = 3 heads, b = 2 shifts, and n = 4 fills per combination, so 24 fills in all. Each F is the mean square for that effect divided by the error mean square, and each is compared with the F distribution to get a p-value.
Hypotheses and Assumptions
| Test | Null hypothesis |
|---|---|
| Factor A | All head means are equal (averaged over the shifts) |
| Factor B | The shift means are equal (averaged over the heads) |
| Interaction | The effect of the head is the same on both shifts |
The assumptions are the same as for one-way ANOVA: independent observations, roughly normal residuals, and similar variance in every combination of the two factors. Check them with residual plots and a test for equal variances. The design also works best when it is balanced, with the same number of measurements in every cell. Excel’s tool requires a balanced design. Minitab handles unbalanced designs with its General Linear Model.
Worked Example
The same packaging engineer now has fills from the three heads on both shifts. The target is still 500 g.
| Head | Shift | Fill 1 | Fill 2 | Fill 3 | Fill 4 |
|---|---|---|---|---|---|
| Head 1 | Day | 498.0 | 498.9 | 499.1 | 498.4 |
| Head 1 | Night | 498.8 | 497.9 | 498.5 | 498.4 |
| Head 2 | Day | 500.5 | 501.4 | 500.4 | 500.9 |
| Head 2 | Night | 500.8 | 500.4 | 501.1 | 500.1 |
| Head 3 | Day | 499.5 | 498.9 | 498.6 | 499.0 |
| Head 3 | Night | 500.4 | 501.3 | 501.1 | 500.8 |
| Shift | Head 1 | Head 2 | Head 3 | Shift mean |
|---|---|---|---|---|
| Day | 498.60 | 500.80 | 499.00 | 499.47 |
| Night | 498.40 | 500.60 | 500.90 | 499.97 |
| Head mean | 498.50 | 500.70 | 499.95 | 499.72 |
- Grand mean = 499.717 g.
- Head (A). SS = 2 × 4 × Σ(head mean − 499.717)² = 20.013, df = 2.
- Shift (B). SS = 3 × 4 × Σ(shift mean − 499.717)² = 1.500, df = 1.
- Interaction. For each cell take (cell mean − head mean − shift mean + grand mean), square it, add the six, and multiply by 4: SS = 5.880, df = 2.
- Error. Add the squared distances of each fill from its own cell mean: SS = 3.240, df = 18, MSE = 0.180.
- Divide each mean square by MSE to get F, and look up the p-values.
| Source | SS | df | MS | F | p |
|---|---|---|---|---|---|
| Head | 20.013 | 2 | 10.007 | 55.59 | 0.000 |
| Shift | 1.500 | 1 | 1.500 | 8.33 | 0.0098 |
| Head × Shift | 5.880 | 2 | 2.940 | 16.33 | 0.00009 |
| Error | 3.240 | 18 | 0.180 | ||
| Total | 30.633 | 23 |
Run It in Excel and Minitab
ExcelStep by step
- Arrange the data with one column per head, and the rows for each shift stacked in blocks of four. Put the head names in row 1 (B1:D1), and the shift name in column A on the first row of each block (A2 “Day”, A6 “Night”). The data fill B2:D9.
- Choose .
- Set Input Range to and Rows per sample to . Leave Alpha at 0.05.
- Choose an output location and click OK. Read the last table: Excel calls the row factor Sample (here Shift), the column factor Columns (here Head), and adds Interaction and Within (error).
- Build the interaction plot by averaging each cell ( and so on) and inserting a line chart with one line per head.
The tool needs the same number of replicates in every cell. It has no post-hoc tests and no residual plots, so use the simple-effects calculation below or Minitab for follow-up.
MinitabStep by step
- Stack the data: one column for Weight, one for Head, one for Shift.
- Choose . Set Responses to Weight and Factors to Head and Shift.
- Click Model. Select both factors in the list, set Terms through order to 2, and click Add. The model now shows Head, Shift, and Head*Shift.
- Click Graphs, tick Four in one residual plots, and click OK. Click Comparisons and choose Tukey on the term Head*Shift if you want all cell comparisons.
- Draw the interaction plot with (Responses: Weight; Factors: Head, Shift).
Use if you prefer the classic layout for a balanced design.
What the output looks like
Anova: Two-Factor With Replication ANOVA Source of Variation SS df MS F P-value F crit Sample (Shift) 1.5000 1 1.5000 8.3333 0.009822 4.4139 Columns (Head) 20.0133 2 10.0067 55.5926 1.98e-08 3.5546 Interaction 5.8800 2 2.9400 16.3333 0.000090 3.5546 Within 3.2400 18 0.1800 Total 30.6333 23
Factor Information Factor Type Levels Values Head Fixed 3 Head 1, Head 2, Head 3 Shift Fixed 2 Day, Night Analysis of Variance Source DF Adj SS Adj MS F-Value P-Value Head 2 20.0133 10.0067 55.59 0.000 Shift 1 1.5000 1.5000 8.33 0.010 Head*Shift 2 5.8800 2.9400 16.33 0.000 Error 18 3.2400 0.1800 Total 23 30.6333 Model Summary S R-sq R-sq(adj) R-sq(pred) 0.42426 89.42% 86.49% 81.20%
Reading the Result: Interaction First
- Look at the interaction. If it is not significant, read the main effects as the average effect of each factor. If it is significant, as here, do not interpret the main effects in isolation.
- Look at the interaction plot. Heads 1 and 2 are flat across shifts. Head 3 changes a lot.
- Follow up with simple effects. Compare the shifts separately for each head, using the error mean square from the full model.
| Head | Day mean | Night mean | Night − Day | t | Adjusted p |
|---|---|---|---|---|---|
| Head 1 | 498.60 | 498.40 | -0.20 | -0.67 | 1.000 |
| Head 2 | 500.80 | 500.60 | -0.20 | -0.67 | 1.000 |
| Head 3 | 499.00 | 500.90 | +1.90 | 6.33 | 0.000 |
Each comparison uses SE = √(2 × MSE / n) = √(2 × 0.180 / 4) = 0.300 on 18 degrees of freedom. The p-values are multiplied by 3 (Bonferroni) because three comparisons were made.
This is why an interaction matters in practice. Averaged over the heads, the shift effect is only 0.5 g, and a manager might call it small and move on. In fact, one head has a problem on one shift and the others have none. The fix is to investigate Head 3 on nights, not to adjust the whole line.
When the Assumptions Do Not Hold
| Problem | What to do |
|---|---|
| Unequal spread across cells | Transform the response (for example, log), or investigate why the spread differs; check with the residual-versus-fits plot |
| Non-normal residuals | Check for outliers and skewness; consider a transformation; with balanced data and moderate sample sizes ANOVA is fairly robust |
| Unbalanced data (different numbers per cell) | Use a general linear model (Minitab) with adjusted sums of squares; Excel’s tool cannot handle it |
| No replicates in the cells | Drop the interaction term, or collect replicates; with one value per cell the interaction cannot be tested |
| Observations not independent | Randomize the run order; add the structure to the model (blocks or repeated measures) |
Common Mistakes
| Mistake | Why it misleads | Better |
|---|---|---|
| Reading main effects when the interaction is significant | The average hides different behavior at different levels | Interpret the interaction and the simple effects |
| Running two one-way ANOVAs instead of one two-way | The interaction is invisible and the error is larger | Fit both factors in one model |
| Testing every cell pair without adjustment | False alarms add up | Use Tukey on the interaction or a stated number of planned comparisons |
| No replicates | Interaction and error cannot be separated | Plan at least two, preferably more, measurements per cell |
| Ignoring the run order | Drift is mistaken for a factor effect | Randomize, and plot residuals against run order |
| Calling an effect “not significant” to mean “absent” | A small design may not detect it | Check the power and report the size of the effect |
Try It Yourself
Seal strength (N) was measured on three seals at each combination of sealing temperature and pressure.
| Temperature | Pressure | Seal 1 | Seal 2 | Seal 3 |
|---|---|---|---|---|
| Low | Low | 20 | 22 | 21 |
| Low | High | 25 | 24 | 26 |
| High | Low | 26 | 27 | 25 |
| High | High | 30 | 31 | 29 |
- Which effects are significant at α = 0.05?
- Does the effect of temperature depend on the pressure?
- What would you tell the process owner?
Show the answer
Temperature: F = 75.0, p = 0.00002. Pressure: F = 48.0, p = 0.00012. Interaction: F = 0.00, p = 1.000. Both main effects are significant and the interaction is not.
The lines on an interaction plot would be parallel: raising the temperature adds about 5 N at either pressure, and raising the pressure adds about 4 N at either temperature. Because there is no interaction, the two settings can be chosen independently, and the strongest seal comes from high temperature with high pressure.
Two-Way ANOVA and Interactions: Frequently Asked Questions
What is an interaction in ANOVA?
An interaction means that the effect of one factor depends on the level of the other. In the example, the shift makes no difference for Heads 1 and 2 but a large difference for Head 3. On an interaction plot, interacting factors give lines that are not parallel.
If the interaction is significant, can I still look at the main effects?
Look at them with care. A main effect is an average over the levels of the other factor, and when there is an interaction that average can hide opposite or unequal effects. Interpret the interaction first, then compare levels of one factor within each level of the other.
How many replicates do I need?
At least two per cell to estimate the interaction, and more to detect smaller effects. Use a power calculation: decide the smallest difference that matters, estimate the standard deviation, and size the replicates to detect it.
What is the difference between fixed and random factors?
A factor is fixed when you chose its levels on purpose and they are the only ones of interest, such as three particular filling heads. It is random when the levels are a sample from a larger population, such as operators drawn from a pool. Random factors change the F tests and the variance components. The examples here use fixed factors.
What if my design is unbalanced?
Use a general linear model, as in Minitab’s Fit General Linear Model, which uses adjusted sums of squares. Excel’s two-factor tool needs the same number of replicates in every cell.
How is a randomized block design related?
It is a two-way analysis in which the second factor is a nuisance variable, such as batch or day, that you block on to remove its variation. It is usually analyzed without an interaction term.
Sources and Further Reading
- Douglas C. Montgomery, Design and Analysis of Experiments, Wiley, chapters on factorial designs.
- NIST/SEMATECH, e-Handbook of Statistical Methods, section on two-way ANOVA (itl.nist.gov/div898/handbook).
- Michael H. Kutner, Christopher J. Nachtsheim, John Neter, and William Li, Applied Linear Statistical Models, McGraw-Hill.
- Minitab Support, “Methods and formulas for General Linear Model” (support.minitab.com).
- Microsoft Support, “Analysis ToolPak” documentation (support.microsoft.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.