Question it answers
Which groups differ, and can I trust the ANOVA model?
When
After a significant ANOVA, or for planned comparisons with a control
Key output
Adjusted p-values, intervals for differences, and grouping letters
Default choice
Tukey for all pairs, Dunnett against a control, Games-Howell for unequal variances
Assumption checks
Residual probability plot, residuals versus fits, Levene’s test
Excel
Formulas with a q table; no built-in post-hoc test
Minitab
One-Way > Comparisons; Test for Equal Variances
Prerequisite
One-way ANOVA

Why You Need a Post-Hoc Test

A significant one-way ANOVA says that at least one group mean is different. It does not say which one. To find out, you compare the groups in pairs. The trouble is the one you met on the One-Way ANOVA page: every comparison carries a risk of a false alarm, and the risks add up.

A post-hoc comparison (post-hoc means “after this”) compares the groups in pairs while keeping the overall chance of at least one false alarm at 5%. That overall chance is called the family-wise error rate. Different methods control it in different ways, and each suits a different question.

-1 0 1 2 3 4 No difference Head 2 − Head 1 Head 3 − Head 1 Head 2 − Head 3 Difference in mean fill weight (g), with Tukey 95% simultaneous interval
Tukey’s method gives an interval for every pairwise difference, built so that all the intervals together have 95% confidence. An interval that excludes zero (gold) marks a significant difference.

Choosing a Method

MethodBest forHow it controls the errorWatch for
Tukey HSDAll pairwise comparisonsUses the studentized range distribution, so the family-wise error is 5% for all pairsAssumes equal variances; Tukey-Kramer handles unequal group sizes
Fisher LSDA few planned comparisons after a significant FOrdinary t-tests with the pooled error; no adjustmentError rate grows with the number of comparisons; use only after a significant ANOVA and few groups
BonferroniA small, fixed number of planned comparisonsDivides α by the number of comparisonsConservative when there are many comparisons
DunnettComparing every group with one controlUses the multivariate t distribution for the comparisons with the control onlyDoes not compare the non-control groups with each other
SchefféAny comparison, including complex contrasts, chosen after seeing the dataControls error for all possible contrastsVery conservative for simple pairs
Games-HowellPairwise comparisons when variances are unequalWelch-type test with the studentized rangeNeeds reasonable sample sizes (about 5 or more per group)
Hsu MCBFinding the best group (largest or smallest mean)Compares each group with the best of the othersAnswers a narrower question than all-pairs
Rule of thumb. Compare everything with everything: Tukey. Compare to a control: Dunnett. Unequal variances: Games-Howell. A handful of comparisons decided before the data: Bonferroni. Decide the method before looking at the results.

Worked Example: Which Filling Heads Differ?

The one-way ANOVA on the three filling heads gave F = 15.52 and p = 0.0002. The error mean square is MSE = 0.4503 on 15 degrees of freedom, with n = 6 fills per head. The head means are 498.533, 500.667, and 499.317 g.

MethodCritical valueStandard errorSmallest significant difference
Fisher LSDt = 2.1314√(2 MSE / n) = 0.38740.826 g
Tukey HSDq = 3.6734√(MSE / n) = 0.27401.006 g
Bonferroni (3 comparisons)t = 2.69370.38741.044 g
Scheffé√((k − 1) F) = 2.71380.38741.051 g

Any difference larger than the smallest significant difference is declared significant. The methods differ in how demanding that threshold is. Fisher LSD is the lowest because it makes no adjustment, and Tukey and Bonferroni are higher to protect the family-wise error rate.

PairDifference (g)Fisher LSDBonferroniTukeyScheffé
Head 2 − Head 1+2.133YesYesYesYes
Head 3 − Head 1+0.783NoNoNoNo
Head 2 − Head 3+1.350YesYesYesYes
PairTukey adjusted pTukey 95% interval (g)Dunnett adjusted p (control Head 1)Dunnett 95% interval (g)
Head 2 − Head 10.00021.13 to 3.140.00011.19 to 3.08
Head 3 − Head 10.141-0.22 to 1.790.108-0.16 to 1.73
Head 2 − Head 30.0090.34 to 2.36not comparednot compared
Conclusion. Head 2 fills significantly heavier than both Head 1 (2.13 g) and Head 3 (1.35 g). Heads 1 and 3 cannot be distinguished (0.78 g, adjusted p = 0.14). In the Minitab letter display that is Head 2 = A, Head 3 = B, Head 1 = B. All four methods agree here because the differences are well separated from the thresholds. They can disagree when a difference sits close to a threshold.

Checking the ANOVA Assumptions

Post-hoc tests rest on the same assumptions as the ANOVA itself. The residuals (each value minus its group mean) are the tool for checking them.

-2 -1 0 1 2 Normal score (expected z) Normal probability plot of residuals Residual (g)
Points close to the line suggest normal residuals.
498.5 499.3 500.7 Fitted value (group mean, g) Residuals versus fitted values Residual (g)
Equal vertical spread at each fitted value, and no pattern, suggests equal variances.
Pattern you seeWhat it suggestsWhat to do
Points follow the line in the probability plotResiduals are close to normalNothing; proceed
S-shaped or curved probability plotSkewed or heavy-tailed residualsTry a transformation (log, square root); or a rank-based test
One point far from the lineAn outlierCheck for a recording error; report with and without the point
Residuals fan out (funnel) as fitted values riseVariance grows with the meanTransform the response (often log), or use Welch ANOVA with Games-Howell
A trend or cycle in residuals against run orderObservations are not independent, or the process driftedRandomize the run order; add time or block to the model
One group with much larger spreadUnequal variancesWelch’s ANOVA and Games-Howell; investigate why that group is more variable
CheckResult for the exampleReading
Shapiro-Wilk test of the residualsp = 0.23No evidence against normality
Levene’s test (median-based)p = 0.79No evidence of unequal variances
Bartlett’s testp = 0.93Same conclusion; sensitive to non-normality, so Levene is preferred
Largest sd / smallest sd1.20Well under the rule of thumb of 2
Plots first, tests second. A significance test of normality can flag trivial departures when the sample is large, and miss serious ones when it is small. Judge the plots, and use the tests as a second opinion.

Run It in Excel and Minitab

ExcelStep by step

  1. Run Anova: Single Factor (see the One-Way ANOVA page) and note MSE (Within Groups MS), the error df, and n per group.
  2. Compute the group means with =AVERAGE(range) and each pairwise difference with =ABS(mean1-mean2).
  3. Tukey: look up q for your number of groups and error df in the table below, then =q*SQRT(MSE/n) gives the HSD. A difference larger than the HSD is significant.
  4. Fisher LSD: =T.INV.2T(0.05, df)*SQRT(2*MSE/n). Bonferroni: replace 0.05 with 0.05 divided by the number of comparisons.
  5. Residual plot: subtract each group mean from its values, sort the residuals, and plot them against =NORM.S.INV((rank-0.5)/N) in a scatter chart.
  6. Levene’s test: compute each value’s absolute distance from its group median in new columns, then run Anova: Single Factor on those columns.

Excel has no built-in Tukey, Dunnett, or Games-Howell test and no studentized range function, which is why the table is provided.

MinitabStep by step

  1. Choose Stat > ANOVA > One-Way and set the Response and Factor.
  2. Click Comparisons. Tick Tukey (all pairs), or Dunnett (choose the control level), or Fisher, or Hsu MCB. Keep Grouping information, Tests and intervals, and Interval plot for differences of means ticked.
  3. Click Graphs and tick Four in one under residual plots, then OK.
  4. For unequal variances, click Options and clear Assume equal variances. The comparisons become Games-Howell.
  5. Run Stat > ANOVA > Test for Equal Variances for Bartlett’s and Levene’s tests.
  6. Read the Grouping Information table: means that share a letter are not significantly different.

The Assistant (Assistant > Hypothesis Tests > One-Way ANOVA) adds a report card that flags assumption problems.

Studentized range critical values q (α = 0.05) for Excel work
Error dfk = 2k = 3k = 4k = 5k = 6
103.1513.8774.3274.6544.912
153.0143.6734.0764.3674.595
202.9503.5783.9584.2324.445
302.8883.4863.8454.1024.301
602.8293.3993.7373.9774.163

What the Minitab output looks like

Minitab: Tukey grouping and interval table for the example (typed excerpt, simplified)
Grouping Information Using the Tukey Method and 95% Confidence

Factor  N     Mean  Grouping
Head 2   6  500.667  A
Head 3   6  499.317  B
Head 1   6  498.533  B

Means that do not share a letter are significantly different.

Difference of    Difference    Adjusted
Levels             of Means     P-Value
Head 2 - Head 1       2.133       0.000
Head 3 - Head 1       0.783       0.141
Head 2 - Head 3       1.350       0.009

Reading and Reporting

  1. State the method and why. “Tukey’s HSD was used to compare all pairs while holding the family-wise error at 5%.”
  2. Report the differences with their intervals, not only the p-values. The interval shows how large the difference could be.
  3. Use the letters or the interval plot to summarize which groups are alike and which differ.
  4. Say what it means in practice. A 2.1 g difference between heads matters when the tolerance is ±3 g.
A sentence you can use. Tukey comparisons showed that Head 2 filled 2.13 g heavier than Head 1 (95% CI 1.13 to 3.14 g) and 1.35 g heavier than Head 3 (95% CI 0.34 to 2.36 g); Heads 1 and 3 did not differ significantly.

Common Mistakes

MistakeWhy it misleadsBetter
Running t-tests on all pairs without adjustmentThe family-wise false-alarm rate climbs well above 5%Use Tukey, or Bonferroni for a few planned comparisons
Choosing the method after seeing the resultsYou can pick the one that gives the answer you wantDecide the comparisons and the method before analysis
Using Tukey when variances are clearly unequalThe error rate is wrongUse Games-Howell
Reading “no significant difference” as “the same”A small sample may not detect a real differenceLook at the interval; check the power
Comparing only the best and worst groupThe extremes of many groups always look differentUse a method that accounts for all the comparisons
Running post-hoc tests after a non-significant ANOVAThe overall test found nothing to localizeStop, or plan the comparisons separately

Try It Yourself

A one-way ANOVA on four suppliers (A, B, C, D) with 6 observations each gave a significant F. The error mean square is MSE = 4.0 on 20 degrees of freedom. The supplier means are A = 40, B = 43, C = 47, and D = 41.

  • Use Tukey’s method with q = 3.958 to decide which suppliers differ.
  • Which supplier would you investigate first?
Show the answer

HSD = q √(MSE / n) = 3.958 × √(4.0 / 6) = 3.23. Differences larger than 3.23 are significant.

PairDifferenceSignificant?
A − B3No
A − C7Yes
A − D1No
B − C4Yes
B − D2No
C − D6Yes

Supplier C differs from A, B, and D. A, B, and D cannot be separated (B and A differ by 3, which is just under the threshold). Investigate supplier C first.

Post-Hoc Comparisons and ANOVA Assumptions: Frequently Asked Questions

When do I need a post-hoc test?

After a significant ANOVA, when you want to know which groups differ. If the ANOVA is not significant, there is usually nothing to localize. If you have only two groups, no post-hoc test is needed, because the t-test already compares them.

What is the difference between Tukey and Fisher LSD?

Fisher LSD uses ordinary t-tests with the pooled error and makes no adjustment, so false alarms accumulate as the number of comparisons grows. Tukey adjusts so that the chance of any false alarm across all pairs stays at 5%. Tukey is the safer default.

Can I use a post-hoc test without running ANOVA first?

Yes. Tukey and Dunnett tests control the family-wise error without needing a significant F. The “ANOVA first” rule matters mainly for Fisher LSD.

What if the variances are unequal?

Use Welch’s ANOVA and the Games-Howell comparisons, which do not assume equal variances. In Minitab, clear Assume equal variances in the One-Way Options. Then find out why the spread differs, because that can be the most useful finding.

What do the letters in Minitab’s grouping table mean?

Groups that share a letter are not significantly different. Groups that share no letter are. Letters are assigned in order of decreasing mean, starting with A for the largest.

How is a confidence interval for a difference read?

If the interval excludes zero, the difference is significant at the family-wise level. The interval also shows how large the difference could plausibly be, which is what matters for the decision.

Sources and Further Reading

  • Douglas C. Montgomery, Design and Analysis of Experiments, Wiley, section on multiple comparisons.
  • NIST/SEMATECH, e-Handbook of Statistical Methods, section on multiple comparisons (itl.nist.gov/div898/handbook).
  • John W. Tukey, “Comparing Individual Means in the Analysis of Variance,” Biometrics, 1949.
  • Charles W. Dunnett, “A Multiple Comparison Procedure for Comparing Several Treatments with a Control,” Journal of the American Statistical Association, 1955.
  • Minitab Support, “Methods and formulas for comparisons in One-Way ANOVA” (support.minitab.com).

This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.