Question it answers
Does each of two factors change the average, and does one factor's effect depend on the other?
Data needed
One continuous response, two categorical factors, and replicates in every combination
Key output
F and p for each factor and for the interaction; R-squared
Read first
The interaction. If it is significant, do not read the main effects alone
Assumptions
Independent observations, normal residuals, similar spread in every cell, ideally balanced
Excel
Data Analysis > Anova: Two-Factor With Replication
Minitab
Stat > ANOVA > General Linear Model > Fit General Linear Model
Prerequisite
One-way ANOVA

The Idea in Plain Language

The one-way ANOVA page compared three filling heads. In a real plant, the heads run on more than one shift, and the shift might matter too. A two-way ANOVA studies two factors in the same experiment and answers three questions at once:

  1. Does the first factor (head) change the average fill weight?
  2. Does the second factor (shift) change it?
  3. Does the effect of one factor depend on the level of the other? This is the interaction.

The interaction is the reason to use a two-way design. Two one-way analyses cannot see it. Many real problems are interactions: a setting that works on one machine and fails on another, a method that helps one shift and not the other, a material that behaves differently at high temperature.

No interaction LowHigh Lines are parallel: each factor adds the same amount at every level of the other. Interaction, same order LowHigh Lines converge: the effect of one factor is bigger at one level of the other. Interaction, crossing LowHigh Lines cross: the best level of one factor depends on the other. Main effects can mislead.
Plot the average for each combination of the two factors. Parallel lines mean the factors act independently. Lines that converge or cross mean they interact.

When to Use Two-Way ANOVA

Your situationUseWhy
One continuous response and two categorical factors, with several measurements in each combinationTwo-way ANOVATests both factors and their interaction
Only one factorOne-way ANOVASimpler, with the same logic
Three or more factors, or factors with numeric levels you set on purposeFactorial design analysis (DOE)See the Design of Experiments guide
A nuisance factor you cannot control (batch, day) but want to removeRandomized block designA two-way model without an interaction term, with the block as the second factor
A numeric input such as temperature and a category such as machineGeneral linear model with a covariateFit in Minitab’s General Linear Model
Data are not normal and groups are smallTransform the response, or use a general linear model with careThere is no simple rank-based two-way test in Excel or Minitab’s basic menu
Replicates matter. To estimate the interaction you need more than one measurement in every combination of the two factors. The example uses four fills per head and shift. With one measurement per cell, the interaction cannot be separated from the error.

How It Works

The model says each observation is the grand mean, plus an effect of the first factor, plus an effect of the second, plus an interaction effect, plus random error:

yijk = μ + αi + βj + (αβ)ij + εijk
SourceSum of squares (balanced design)Degrees of freedomF
Factor A (a levels)b n Σ (meani − grand mean)²a − 1MSA / MSE
Factor B (b levels)a n Σ (meanj − grand mean)²b − 1MSB / MSE
Interaction A × Bn Σ (cell mean − row mean − column mean + grand mean)²(a − 1)(b − 1)MSAB / MSE
ErrorΣ (value − cell mean)²ab(n − 1)
TotalΣ (value − grand mean)²abn − 1

Here a = 3 heads, b = 2 shifts, and n = 4 fills per combination, so 24 fills in all. Each F is the mean square for that effect divided by the error mean square, and each is compared with the F distribution to get a p-value.

498 498.5 499 499.5 500 500.5 501 501.5 Day Night Shift Mean fill weight (g) Head 1 Head 2 Head 3
The interaction plot for the example. Heads 1 and 2 fill the same weight on both shifts. Head 3 fills almost 2 g heavier on nights. That difference is what the interaction term detects.
Head (A) 20.01 (65.3%) Head x Shift (A x B) 5.88 (19.2%) Shift (B) 1.50 (4.9%) Error 3.24 (10.6%) Sum of squares
Where the variation comes from. Most of it is the difference between heads, and a fifth is the interaction. The shift on its own explains very little.

Hypotheses and Assumptions

TestNull hypothesis
Factor AAll head means are equal (averaged over the shifts)
Factor BThe shift means are equal (averaged over the heads)
InteractionThe effect of the head is the same on both shifts

The assumptions are the same as for one-way ANOVA: independent observations, roughly normal residuals, and similar variance in every combination of the two factors. Check them with residual plots and a test for equal variances. The design also works best when it is balanced, with the same number of measurements in every cell. Excel’s tool requires a balanced design. Minitab handles unbalanced designs with its General Linear Model.

Worked Example

The same packaging engineer now has fills from the three heads on both shifts. The target is still 500 g.

Fill weights in grams (4 fills per head and shift)
HeadShiftFill 1Fill 2Fill 3Fill 4
Head 1Day498.0498.9499.1498.4
Head 1Night498.8497.9498.5498.4
Head 2Day500.5501.4500.4500.9
Head 2Night500.8500.4501.1500.1
Head 3Day499.5498.9498.6499.0
Head 3Night500.4501.3501.1500.8
Cell means, margins, and grand mean (g)
ShiftHead 1Head 2Head 3Shift mean
Day498.60500.80499.00499.47
Night498.40500.60500.90499.97
Head mean498.50500.70499.95499.72
  1. Grand mean = 499.717 g.
  2. Head (A). SS = 2 × 4 × Σ(head mean − 499.717)² = 20.013, df = 2.
  3. Shift (B). SS = 3 × 4 × Σ(shift mean − 499.717)² = 1.500, df = 1.
  4. Interaction. For each cell take (cell mean − head mean − shift mean + grand mean), square it, add the six, and multiply by 4: SS = 5.880, df = 2.
  5. Error. Add the squared distances of each fill from its own cell mean: SS = 3.240, df = 18, MSE = 0.180.
  6. Divide each mean square by MSE to get F, and look up the p-values.
SourceSSdfMSFp
Head20.013210.00755.590.000
Shift1.50011.5008.330.0098
Head × Shift5.88022.94016.330.00009
Error3.240180.180
Total30.63323
Read the interaction first. The interaction is significant (F = 16.33, p = 0.00009), so the effect of the head depends on the shift, and the shift effect is not the same for every head. The shift is also significant on its own (p = 0.010), but that average hides what is really going on, as the next section shows.

Run It in Excel and Minitab

ExcelStep by step

  1. Arrange the data with one column per head, and the rows for each shift stacked in blocks of four. Put the head names in row 1 (B1:D1), and the shift name in column A on the first row of each block (A2 “Day”, A6 “Night”). The data fill B2:D9.
  2. Choose Data > Data Analysis > Anova: Two-Factor With Replication.
  3. Set Input Range to A1:D9 and Rows per sample to 4. Leave Alpha at 0.05.
  4. Choose an output location and click OK. Read the last table: Excel calls the row factor Sample (here Shift), the column factor Columns (here Head), and adds Interaction and Within (error).
  5. Build the interaction plot by averaging each cell (=AVERAGE(B2:B5) and so on) and inserting a line chart with one line per head.

The tool needs the same number of replicates in every cell. It has no post-hoc tests and no residual plots, so use the simple-effects calculation below or Minitab for follow-up.

MinitabStep by step

  1. Stack the data: one column for Weight, one for Head, one for Shift.
  2. Choose Stat > ANOVA > General Linear Model > Fit General Linear Model. Set Responses to Weight and Factors to Head and Shift.
  3. Click Model. Select both factors in the list, set Terms through order to 2, and click Add. The model now shows Head, Shift, and Head*Shift.
  4. Click Graphs, tick Four in one residual plots, and click OK. Click Comparisons and choose Tukey on the term Head*Shift if you want all cell comparisons.
  5. Draw the interaction plot with Stat > ANOVA > Interactions Plot (Responses: Weight; Factors: Head, Shift).

Use Stat > ANOVA > Balanced ANOVA if you prefer the classic layout for a balanced design.

What the output looks like

Excel Data Analysis ToolPak output (Anova: Two-Factor With Replication, ANOVA table only)
Anova: Two-Factor With Replication

ANOVA
Source of Variation          SS   df       MS        F    P-value   F crit
Sample (Shift)           1.5000    1   1.5000   8.3333   0.009822   4.4139
Columns (Head)          20.0133    2  10.0067  55.5926   1.98e-08   3.5546
Interaction              5.8800    2   2.9400  16.3333   0.000090   3.5546
Within                   3.2400   18   0.1800

Total                   30.6333   23
Minitab session window (typed excerpt, simplified)
Factor Information

Factor  Type   Levels  Values
Head    Fixed       3  Head 1, Head 2, Head 3
Shift   Fixed       2  Day, Night

Analysis of Variance

Source       DF   Adj SS   Adj MS  F-Value  P-Value
Head          2  20.0133  10.0067    55.59    0.000
Shift         1   1.5000   1.5000     8.33    0.010
Head*Shift    2   5.8800   2.9400    16.33    0.000
Error        18   3.2400   0.1800
Total        23  30.6333

Model Summary

       S    R-sq  R-sq(adj)  R-sq(pred)
0.42426  89.42%     86.49%      81.20%

Reading the Result: Interaction First

  1. Look at the interaction. If it is not significant, read the main effects as the average effect of each factor. If it is significant, as here, do not interpret the main effects in isolation.
  2. Look at the interaction plot. Heads 1 and 2 are flat across shifts. Head 3 changes a lot.
  3. Follow up with simple effects. Compare the shifts separately for each head, using the error mean square from the full model.
HeadDay meanNight meanNight − DaytAdjusted p
Head 1498.60498.40-0.20-0.671.000
Head 2500.80500.60-0.20-0.671.000
Head 3499.00500.90+1.906.330.000

Each comparison uses SE = √(2 × MSE / n) = √(2 × 0.180 / 4) = 0.300 on 18 degrees of freedom. The p-values are multiplied by 3 (Bonferroni) because three comparisons were made.

A sentence you can use. A two-way ANOVA showed a significant head by shift interaction, F(2, 18) = 16.33, p < 0.001. Head 3 filled 1.9 g heavier on the night shift than on the day shift (adjusted p < 0.001), while Heads 1 and 2 showed no shift effect. The model explained 89% of the variation in fill weight.

This is why an interaction matters in practice. Averaged over the heads, the shift effect is only 0.5 g, and a manager might call it small and move on. In fact, one head has a problem on one shift and the others have none. The fix is to investigate Head 3 on nights, not to adjust the whole line.

When the Assumptions Do Not Hold

ProblemWhat to do
Unequal spread across cellsTransform the response (for example, log), or investigate why the spread differs; check with the residual-versus-fits plot
Non-normal residualsCheck for outliers and skewness; consider a transformation; with balanced data and moderate sample sizes ANOVA is fairly robust
Unbalanced data (different numbers per cell)Use a general linear model (Minitab) with adjusted sums of squares; Excel’s tool cannot handle it
No replicates in the cellsDrop the interaction term, or collect replicates; with one value per cell the interaction cannot be tested
Observations not independentRandomize the run order; add the structure to the model (blocks or repeated measures)

Common Mistakes

MistakeWhy it misleadsBetter
Reading main effects when the interaction is significantThe average hides different behavior at different levelsInterpret the interaction and the simple effects
Running two one-way ANOVAs instead of one two-wayThe interaction is invisible and the error is largerFit both factors in one model
Testing every cell pair without adjustmentFalse alarms add upUse Tukey on the interaction or a stated number of planned comparisons
No replicatesInteraction and error cannot be separatedPlan at least two, preferably more, measurements per cell
Ignoring the run orderDrift is mistaken for a factor effectRandomize, and plot residuals against run order
Calling an effect “not significant” to mean “absent”A small design may not detect itCheck the power and report the size of the effect

Try It Yourself

Seal strength (N) was measured on three seals at each combination of sealing temperature and pressure.

TemperaturePressureSeal 1Seal 2Seal 3
LowLow202221
LowHigh252426
HighLow262725
HighHigh303129
  • Which effects are significant at α = 0.05?
  • Does the effect of temperature depend on the pressure?
  • What would you tell the process owner?
Show the answer

Temperature: F = 75.0, p = 0.00002. Pressure: F = 48.0, p = 0.00012. Interaction: F = 0.00, p = 1.000. Both main effects are significant and the interaction is not.

The lines on an interaction plot would be parallel: raising the temperature adds about 5 N at either pressure, and raising the pressure adds about 4 N at either temperature. Because there is no interaction, the two settings can be chosen independently, and the strongest seal comes from high temperature with high pressure.

Two-Way ANOVA and Interactions: Frequently Asked Questions

What is an interaction in ANOVA?

An interaction means that the effect of one factor depends on the level of the other. In the example, the shift makes no difference for Heads 1 and 2 but a large difference for Head 3. On an interaction plot, interacting factors give lines that are not parallel.

If the interaction is significant, can I still look at the main effects?

Look at them with care. A main effect is an average over the levels of the other factor, and when there is an interaction that average can hide opposite or unequal effects. Interpret the interaction first, then compare levels of one factor within each level of the other.

How many replicates do I need?

At least two per cell to estimate the interaction, and more to detect smaller effects. Use a power calculation: decide the smallest difference that matters, estimate the standard deviation, and size the replicates to detect it.

What is the difference between fixed and random factors?

A factor is fixed when you chose its levels on purpose and they are the only ones of interest, such as three particular filling heads. It is random when the levels are a sample from a larger population, such as operators drawn from a pool. Random factors change the F tests and the variance components. The examples here use fixed factors.

What if my design is unbalanced?

Use a general linear model, as in Minitab’s Fit General Linear Model, which uses adjusted sums of squares. Excel’s two-factor tool needs the same number of replicates in every cell.

How is a randomized block design related?

It is a two-way analysis in which the second factor is a nuisance variable, such as batch or day, that you block on to remove its variation. It is usually analyzed without an interaction term.

Sources and Further Reading

  • Douglas C. Montgomery, Design and Analysis of Experiments, Wiley, chapters on factorial designs.
  • NIST/SEMATECH, e-Handbook of Statistical Methods, section on two-way ANOVA (itl.nist.gov/div898/handbook).
  • Michael H. Kutner, Christopher J. Nachtsheim, John Neter, and William Li, Applied Linear Statistical Models, McGraw-Hill.
  • Minitab Support, “Methods and formulas for General Linear Model” (support.minitab.com).
  • Microsoft Support, “Analysis ToolPak” documentation (support.microsoft.com).

This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.