- Question it answers
- Which groups differ, and can I trust the ANOVA model?
- When
- After a significant ANOVA, or for planned comparisons with a control
- Key output
- Adjusted p-values, intervals for differences, and grouping letters
- Default choice
- Tukey for all pairs, Dunnett against a control, Games-Howell for unequal variances
- Assumption checks
- Residual probability plot, residuals versus fits, Levene’s test
- Excel
- Formulas with a q table; no built-in post-hoc test
- Minitab
- One-Way > Comparisons; Test for Equal Variances
- Prerequisite
- One-way ANOVA
Why You Need a Post-Hoc Test
A significant one-way ANOVA says that at least one group mean is different. It does not say which one. To find out, you compare the groups in pairs. The trouble is the one you met on the One-Way ANOVA page: every comparison carries a risk of a false alarm, and the risks add up.
A post-hoc comparison (post-hoc means “after this”) compares the groups in pairs while keeping the overall chance of at least one false alarm at 5%. That overall chance is called the family-wise error rate. Different methods control it in different ways, and each suits a different question.
Choosing a Method
| Method | Best for | How it controls the error | Watch for |
|---|---|---|---|
| Tukey HSD | All pairwise comparisons | Uses the studentized range distribution, so the family-wise error is 5% for all pairs | Assumes equal variances; Tukey-Kramer handles unequal group sizes |
| Fisher LSD | A few planned comparisons after a significant F | Ordinary t-tests with the pooled error; no adjustment | Error rate grows with the number of comparisons; use only after a significant ANOVA and few groups |
| Bonferroni | A small, fixed number of planned comparisons | Divides α by the number of comparisons | Conservative when there are many comparisons |
| Dunnett | Comparing every group with one control | Uses the multivariate t distribution for the comparisons with the control only | Does not compare the non-control groups with each other |
| Scheffé | Any comparison, including complex contrasts, chosen after seeing the data | Controls error for all possible contrasts | Very conservative for simple pairs |
| Games-Howell | Pairwise comparisons when variances are unequal | Welch-type test with the studentized range | Needs reasonable sample sizes (about 5 or more per group) |
| Hsu MCB | Finding the best group (largest or smallest mean) | Compares each group with the best of the others | Answers a narrower question than all-pairs |
Worked Example: Which Filling Heads Differ?
The one-way ANOVA on the three filling heads gave F = 15.52 and p = 0.0002. The error mean square is MSE = 0.4503 on 15 degrees of freedom, with n = 6 fills per head. The head means are 498.533, 500.667, and 499.317 g.
| Method | Critical value | Standard error | Smallest significant difference |
|---|---|---|---|
| Fisher LSD | t = 2.1314 | √(2 MSE / n) = 0.3874 | 0.826 g |
| Tukey HSD | q = 3.6734 | √(MSE / n) = 0.2740 | 1.006 g |
| Bonferroni (3 comparisons) | t = 2.6937 | 0.3874 | 1.044 g |
| Scheffé | √((k − 1) F) = 2.7138 | 0.3874 | 1.051 g |
Any difference larger than the smallest significant difference is declared significant. The methods differ in how demanding that threshold is. Fisher LSD is the lowest because it makes no adjustment, and Tukey and Bonferroni are higher to protect the family-wise error rate.
| Pair | Difference (g) | Fisher LSD | Bonferroni | Tukey | Scheffé |
|---|---|---|---|---|---|
| Head 2 − Head 1 | +2.133 | Yes | Yes | Yes | Yes |
| Head 3 − Head 1 | +0.783 | No | No | No | No |
| Head 2 − Head 3 | +1.350 | Yes | Yes | Yes | Yes |
| Pair | Tukey adjusted p | Tukey 95% interval (g) | Dunnett adjusted p (control Head 1) | Dunnett 95% interval (g) |
|---|---|---|---|---|
| Head 2 − Head 1 | 0.0002 | 1.13 to 3.14 | 0.0001 | 1.19 to 3.08 |
| Head 3 − Head 1 | 0.141 | -0.22 to 1.79 | 0.108 | -0.16 to 1.73 |
| Head 2 − Head 3 | 0.009 | 0.34 to 2.36 | not compared | not compared |
Checking the ANOVA Assumptions
Post-hoc tests rest on the same assumptions as the ANOVA itself. The residuals (each value minus its group mean) are the tool for checking them.
| Pattern you see | What it suggests | What to do |
|---|---|---|
| Points follow the line in the probability plot | Residuals are close to normal | Nothing; proceed |
| S-shaped or curved probability plot | Skewed or heavy-tailed residuals | Try a transformation (log, square root); or a rank-based test |
| One point far from the line | An outlier | Check for a recording error; report with and without the point |
| Residuals fan out (funnel) as fitted values rise | Variance grows with the mean | Transform the response (often log), or use Welch ANOVA with Games-Howell |
| A trend or cycle in residuals against run order | Observations are not independent, or the process drifted | Randomize the run order; add time or block to the model |
| One group with much larger spread | Unequal variances | Welch’s ANOVA and Games-Howell; investigate why that group is more variable |
| Check | Result for the example | Reading |
|---|---|---|
| Shapiro-Wilk test of the residuals | p = 0.23 | No evidence against normality |
| Levene’s test (median-based) | p = 0.79 | No evidence of unequal variances |
| Bartlett’s test | p = 0.93 | Same conclusion; sensitive to non-normality, so Levene is preferred |
| Largest sd / smallest sd | 1.20 | Well under the rule of thumb of 2 |
Run It in Excel and Minitab
ExcelStep by step
- Run Anova: Single Factor (see the One-Way ANOVA page) and note MSE (Within Groups MS), the error df, and n per group.
- Compute the group means with and each pairwise difference with .
- Tukey: look up q for your number of groups and error df in the table below, then gives the HSD. A difference larger than the HSD is significant.
- Fisher LSD: . Bonferroni: replace 0.05 with 0.05 divided by the number of comparisons.
- Residual plot: subtract each group mean from its values, sort the residuals, and plot them against in a scatter chart.
- Levene’s test: compute each value’s absolute distance from its group median in new columns, then run Anova: Single Factor on those columns.
Excel has no built-in Tukey, Dunnett, or Games-Howell test and no studentized range function, which is why the table is provided.
MinitabStep by step
- Choose and set the Response and Factor.
- Click Comparisons. Tick Tukey (all pairs), or Dunnett (choose the control level), or Fisher, or Hsu MCB. Keep Grouping information, Tests and intervals, and Interval plot for differences of means ticked.
- Click Graphs and tick Four in one under residual plots, then OK.
- For unequal variances, click Options and clear Assume equal variances. The comparisons become Games-Howell.
- Run for Bartlett’s and Levene’s tests.
- Read the Grouping Information table: means that share a letter are not significantly different.
The Assistant () adds a report card that flags assumption problems.
| Error df | k = 2 | k = 3 | k = 4 | k = 5 | k = 6 |
|---|---|---|---|---|---|
| 10 | 3.151 | 3.877 | 4.327 | 4.654 | 4.912 |
| 15 | 3.014 | 3.673 | 4.076 | 4.367 | 4.595 |
| 20 | 2.950 | 3.578 | 3.958 | 4.232 | 4.445 |
| 30 | 2.888 | 3.486 | 3.845 | 4.102 | 4.301 |
| 60 | 2.829 | 3.399 | 3.737 | 3.977 | 4.163 |
What the Minitab output looks like
Grouping Information Using the Tukey Method and 95% Confidence Factor N Mean Grouping Head 2 6 500.667 A Head 3 6 499.317 B Head 1 6 498.533 B Means that do not share a letter are significantly different. Difference of Difference Adjusted Levels of Means P-Value Head 2 - Head 1 2.133 0.000 Head 3 - Head 1 0.783 0.141 Head 2 - Head 3 1.350 0.009
Reading and Reporting
- State the method and why. “Tukey’s HSD was used to compare all pairs while holding the family-wise error at 5%.”
- Report the differences with their intervals, not only the p-values. The interval shows how large the difference could be.
- Use the letters or the interval plot to summarize which groups are alike and which differ.
- Say what it means in practice. A 2.1 g difference between heads matters when the tolerance is ±3 g.
Common Mistakes
| Mistake | Why it misleads | Better |
|---|---|---|
| Running t-tests on all pairs without adjustment | The family-wise false-alarm rate climbs well above 5% | Use Tukey, or Bonferroni for a few planned comparisons |
| Choosing the method after seeing the results | You can pick the one that gives the answer you want | Decide the comparisons and the method before analysis |
| Using Tukey when variances are clearly unequal | The error rate is wrong | Use Games-Howell |
| Reading “no significant difference” as “the same” | A small sample may not detect a real difference | Look at the interval; check the power |
| Comparing only the best and worst group | The extremes of many groups always look different | Use a method that accounts for all the comparisons |
| Running post-hoc tests after a non-significant ANOVA | The overall test found nothing to localize | Stop, or plan the comparisons separately |
Try It Yourself
A one-way ANOVA on four suppliers (A, B, C, D) with 6 observations each gave a significant F. The error mean square is MSE = 4.0 on 20 degrees of freedom. The supplier means are A = 40, B = 43, C = 47, and D = 41.
- Use Tukey’s method with q = 3.958 to decide which suppliers differ.
- Which supplier would you investigate first?
Show the answer
HSD = q √(MSE / n) = 3.958 × √(4.0 / 6) = 3.23. Differences larger than 3.23 are significant.
| Pair | Difference | Significant? |
|---|---|---|
| A − B | 3 | No |
| A − C | 7 | Yes |
| A − D | 1 | No |
| B − C | 4 | Yes |
| B − D | 2 | No |
| C − D | 6 | Yes |
Supplier C differs from A, B, and D. A, B, and D cannot be separated (B and A differ by 3, which is just under the threshold). Investigate supplier C first.
Post-Hoc Comparisons and ANOVA Assumptions: Frequently Asked Questions
When do I need a post-hoc test?
After a significant ANOVA, when you want to know which groups differ. If the ANOVA is not significant, there is usually nothing to localize. If you have only two groups, no post-hoc test is needed, because the t-test already compares them.
What is the difference between Tukey and Fisher LSD?
Fisher LSD uses ordinary t-tests with the pooled error and makes no adjustment, so false alarms accumulate as the number of comparisons grows. Tukey adjusts so that the chance of any false alarm across all pairs stays at 5%. Tukey is the safer default.
Can I use a post-hoc test without running ANOVA first?
Yes. Tukey and Dunnett tests control the family-wise error without needing a significant F. The “ANOVA first” rule matters mainly for Fisher LSD.
What if the variances are unequal?
Use Welch’s ANOVA and the Games-Howell comparisons, which do not assume equal variances. In Minitab, clear Assume equal variances in the One-Way Options. Then find out why the spread differs, because that can be the most useful finding.
What do the letters in Minitab’s grouping table mean?
Groups that share a letter are not significantly different. Groups that share no letter are. Letters are assigned in order of decreasing mean, starting with A for the largest.
How is a confidence interval for a difference read?
If the interval excludes zero, the difference is significant at the family-wise level. The interval also shows how large the difference could plausibly be, which is what matters for the decision.
Sources and Further Reading
- Douglas C. Montgomery, Design and Analysis of Experiments, Wiley, section on multiple comparisons.
- NIST/SEMATECH, e-Handbook of Statistical Methods, section on multiple comparisons (itl.nist.gov/div898/handbook).
- John W. Tukey, “Comparing Individual Means in the Analysis of Variance,” Biometrics, 1949.
- Charles W. Dunnett, “A Multiple Comparison Procedure for Comparing Several Treatments with a Control,” Journal of the American Statistical Association, 1955.
- Minitab Support, “Methods and formulas for comparisons in One-Way ANOVA” (support.minitab.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.