- Question it answers
- Do three or more groups have the same average?
- Data needed
- One continuous response and one factor with 3 or more levels
- Null hypothesis
- All group means are equal
- Key output
- F statistic, p-value, and R-squared
- Assumptions
- Independent observations, normal residuals, similar spread in each group
- Excel
- Data Analysis > Anova: Single Factor
- Minitab
- Stat > ANOVA > One-Way
- Next step
- A post-hoc comparison to find which groups differ
The Idea in Plain Language
Three filling heads on one line are each set to deliver 500 g. You weigh six fills from each head. The averages come out at 498.5 g, 500.7 g, and 499.3 g. Are the heads really different, or is this the ordinary scatter you get between any two fills from the same head?
Analysis of variance (ANOVA) answers by comparing two kinds of variation. The first is how far the group averages sit from each other: the signal. The second is how much individual fills vary within a single head: the noise. If the signal is much larger than the noise, the groups probably differ. If it is about the same size, the differences are what chance alone would produce.
Despite the name, one-way ANOVA tests means. It uses variances to do it, which is where the name comes from. “One-way” means one factor, here the filling head, with three or more levels.
Why not run several t-tests?
With three groups you could run three pairwise t-tests. Each test carries a 5% risk of a false alarm, and the risks add up. The more groups you compare, the more likely it is that at least one pair looks different by chance. ANOVA asks the single question first: is there any difference at all?
When to Use One-Way ANOVA
| Your situation | Use | Why |
|---|---|---|
| One continuous response, one factor with 3 or more levels, and you want to compare the means | One-way ANOVA | This is the case it was built for |
| Only 2 groups | Two-sample t-test (ANOVA gives the same p-value, with F = t²) | Simpler to report; see Hypothesis Testing |
| Two factors at once (for example head and shift) | Two-way ANOVA | Shows the effect of each factor and their interaction |
| Data are skewed or have outliers and the groups are small | Kruskal-Wallis or Mood’s median | Rank-based tests do not need normal data |
| You want to compare variation, not means | Test for equal variances (Levene) | ANOVA compares averages only |
| The input is continuous (temperature, speed) | Regression | It uses the numeric information that grouping throws away |
| The response is pass/fail or a count | Chi-square or proportion tests | ANOVA needs a continuous response |
How It Works: Signal and Noise
ANOVA measures the total variation in all 18 fills and splits it into two parts.
| Quantity | Formula | In words |
|---|---|---|
| Between-group sum of squares (SSB) | Σ ni (x̄i − x̄)² | How far each group mean is from the grand mean, weighted by group size |
| Within-group sum of squares (SSW) | Σ Σ (xij − x̄i)² | How far each value is from its own group mean |
| Total sum of squares (SST) | SSB + SSW | Total variation around the grand mean |
| Degrees of freedom | Between k − 1; within N − k; total N − 1 | k groups, N observations in all |
| Mean squares | MSB = SSB / (k − 1); MSW = SSW / (N − k) | Variation per degree of freedom: signal and noise |
| F statistic | F = MSB / MSW | Signal divided by noise |
| p-value | P(Fk−1, N−k ≥ F observed) | Chance of a signal this large if all means were equal |
| R-squared | SSB / SST | Share of the total variation explained by the factor |
If the group means are all equal, the signal and the noise estimate the same thing, and F is close to 1. The larger F is, the less plausible it is that all the means are equal.
Hypotheses and Assumptions
| Hypothesis | Statement |
|---|---|
| Null (H0) | μ1 = μ2 = μ3: all group means are equal |
| Alternative (H1) | At least one mean is different. It does not say all means differ, or which one |
| Assumption | What it means | How to check | In the example |
|---|---|---|---|
| Independent observations | One fill does not influence another | Think about how the data were collected; plot residuals in time order | Fills were taken in random order from each head |
| Normal residuals | The scatter within each group is roughly bell-shaped | Normal probability plot of residuals; Anderson-Darling or Shapiro-Wilk test | Shapiro-Wilk p = 0.23: no sign of non-normality |
| Equal variances | Each group has about the same spread | Levene’s test (Minitab and the calculator use the median-based version); the largest standard deviation should be under about twice the smallest | Levene p = 0.79; sd ratio 1.20 |
| No extreme outliers | One wild value can move a mean and inflate the noise | Dot plot or box plot of each group | No outliers in the dot plot |
| Continuous response | A measured value, not a count or a pass/fail | Look at the data type | Weight in grams |
ANOVA is fairly forgiving. With groups of similar size, moderate departures from normality or equal variances change the p-value very little. Unequal group sizes combined with unequal variances are the dangerous case. See the section on what to do when the assumptions fail.
Worked Example by Hand
A packaging engineer wants to know whether three filling heads deliver the same average weight. The target is 500 g.
| Fill | Head 1 | Head 2 | Head 3 |
|---|---|---|---|
| 1 | 498.2 | 500.4 | 499.0 |
| 2 | 499.1 | 501.2 | 499.8 |
| 3 | 497.6 | 499.8 | 498.6 |
| 4 | 498.8 | 500.9 | 500.2 |
| 5 | 499.5 | 501.6 | 499.4 |
| 6 | 498.0 | 500.1 | 498.9 |
| Head | n | Mean | Standard deviation | Sum of squares within the head |
|---|---|---|---|---|
| Head 1 | 6 | 498.533 | 0.720 | 2.593 |
| Head 2 | 6 | 500.667 | 0.686 | 2.353 |
| Head 3 | 6 | 499.317 | 0.601 | 1.808 |
- State the hypotheses. H0: all three head means are equal. H1: at least one differs. Choose α = 0.05.
- Grand mean. The average of all 18 fills is 499.506 g.
- Between-head sum of squares. SSB = 6(498.533 − 499.506)² + 6(500.667 − 499.506)² + 6(499.317 − 499.506)² = 5.671 + 8.089 + 0.214 = 13.974.
- Within-head sum of squares. Add the three within-head sums: 2.593 + 2.353 + 1.808 = 6.755.
- Degrees of freedom. Between = 3 − 1 = 2. Within = 18 − 3 = 15. Total = 18 − 1 = 17.
- Mean squares. MSB = 13.974 / 2 = 6.987. MSW = 6.755 / 15 = 0.450.
- F statistic. F = 6.987 / 0.450 = 15.52.
- Decide. The critical value for 2 and 15 degrees of freedom at α = 0.05 is 3.68. Since 15.52 > 3.68, reject H0. The p-value is 0.0002.
| Source | SS | df | MS | F | p |
|---|---|---|---|---|---|
| Between heads | 13.974 | 2 | 6.987 | 15.52 | 0.0002 |
| Within heads (error) | 6.755 | 15 | 0.450 | ||
| Total | 20.729 | 17 |
Run It in Excel and Minitab
ExcelStep by step
- Put each head in its own column with the head name in row 1, and the six weights below it (here ).
- Check that Data Analysis appears on the Data tab. If not: , choose Excel Add-ins, click Go, and tick Analysis ToolPak.
- Choose and click OK.
- Set Input Range to , Grouped By to Columns, tick Labels in first row, and leave Alpha at 0.05.
- Pick an output location, click OK, and read the ANOVA table at the bottom. Compare P-value with alpha, or F with F crit.
With formulas instead. gives SSW, gives SST, and SSB is the difference. Then gives the p-value and the critical F.
Excel has no built-in Tukey test or equal-variance test for several groups. Use the HSD formula in the post-hoc page, or Minitab.
MinitabStep by step
- Stack the data: one column for the weights (Weight) and one for the head (Head).
- Choose . Leave Response data are in one column for all factor levels selected.
- Set Response to Weight and Factor to Head.
- Click Comparisons, tick Tukey, and keep Grouping information and Tests and intervals ticked.
- Click Graphs and tick Individual value plot, Boxplot of data, and Four in one residual plots. Click OK twice.
- Check the equal-variance assumption with (Response: Weight; Factors: Head).
Unequal variances? In the One-Way dialog click Options and clear Assume equal variances. Minitab then runs Welch’s ANOVA, and the comparisons use Games-Howell.
The Assistant () runs the same analysis with a report card that flags assumption problems.
What the output looks like
The highlighted numbers are the ones to read first: F, the p-value, and R-squared.
Anova: Single Factor SUMMARY Groups Count Sum Average Variance Head 1 6 2991.2 498.5333 0.5187 Head 2 6 3004.0 500.6667 0.4707 Head 3 6 2995.9 499.3167 0.3617 ANOVA Source of Variation SS df MS F P-value F crit Between Groups 13.9744 2 6.9872 15.5157 0.000223 3.6823 Within Groups 6.7550 15 0.4503 Total 20.7294 17
Method Null hypothesis All means are equal Alternative hypothesis At least one mean is different Significance level α = 0.05 Equal variances were assumed for the analysis. Factor Information Factor Levels Values Head 3 Head 1, Head 2, Head 3 Analysis of Variance Source DF Adj SS Adj MS F-Value P-Value Head 2 13.974 6.987 15.52 0.000 Error 15 6.755 0.450 Total 17 20.729 Model Summary S R-sq R-sq(adj) R-sq(pred) 0.67107 67.41% 63.07% 53.08% Means Factor N Mean StDev 95% CI Head 1 6 498.533 0.720 (497.949, 499.117) Head 2 6 500.667 0.686 (500.083, 501.251) Head 3 6 499.317 0.601 (498.733, 499.901) Pooled StDev = 0.67107 Tukey Pairwise Comparisons Grouping Information Using the Tukey Method and 95% Confidence Factor N Mean Grouping Head 2 6 500.667 A Head 3 6 499.317 B Head 1 6 498.533 B Means that do not share a letter are significantly different. Tukey Simultaneous Tests for Differences of Means Difference Difference SE of Adjusted of Levels of Means Difference 95% CI T-Value P-Value Head 2 - Head 1 2.133 0.387 ( 1.127, 3.140) 5.51 0.000 Head 3 - Head 1 0.783 0.387 (-0.223, 1.790) 2.02 0.141 Head 3 - Head 2 -1.350 0.387 (-2.356, -0.344) -3.48 0.009
Reading and Reporting the Result
- Look at the p-value first. p = 0.0002 is below 0.05, so at least one mean differs.
- Check the effect size. R-squared is 67%: the head explains about two thirds of the variation. Adjusted R-squared (63%) allows for the number of groups; predicted R-squared (53%) estimates how well the model would predict new data.
- Find out which groups differ. Tukey’s test shows Head 2 is 2.13 g heavier than Head 1 (95% CI 1.13 to 3.14 g, adjusted p < 0.001) and 1.35 g heavier than Head 3 (adjusted p = 0.009). Heads 1 and 3 cannot be told apart (adjusted p = 0.141).
- Translate it into practice. A head that fills 2.1 g heavier than another uses a large part of a typical ±3 g tolerance. Statistically clear and practically important are not the same thing, so compare the difference with what matters to the customer.
When the Assumptions Do Not Hold
| Problem | What it does | What to do |
|---|---|---|
| Unequal variances | Inflates false alarms, especially when group sizes differ | Welch’s ANOVA with Games-Howell comparisons; or transform the data; investigate why the spread differs, which is often the real finding |
| Non-normal residuals, small groups | Distorts the p-value | Kruskal-Wallis or Mood’s median test (Minitab: ); or transform the response |
| Outliers | Pull a mean and inflate the noise | Find the cause first. Correct a recording error; do not delete a real value to get a better p-value. Report results with and without it |
| Observations not independent | Makes every p-value too small | Redesign the collection: randomize, or use a model that includes the structure (blocks, repeated measures) |
| Very unequal group sizes | Makes the test sensitive to unequal variances | Aim for balanced groups; use Welch’s ANOVA when you cannot |
Common Mistakes
| Mistake | Why it misleads | Better |
|---|---|---|
| Running many t-tests instead of ANOVA | False alarms add up | One ANOVA, then a post-hoc test that controls the overall error rate |
| Concluding “all groups differ” | A significant F says only that at least one mean differs | Use Tukey or another post-hoc comparison to say which |
| Reading a significant p-value as a large effect | With enough data a tiny difference is significant | Report the differences, intervals, and R-squared |
| Reading a non-significant p-value as no difference | A small sample may not detect a real difference | Check the sample size and power |
| Skipping the residual plots | Assumption problems stay hidden | Always look at the four-in-one residual plots |
| Using ANOVA on counts or pass/fail data | The response is not continuous | Use proportion or chi-square tests |
| Comparing groups with different conditions | Machine, shift, and operator get mixed up with the factor | Randomize and block, or add the other factor to the model |
Try It Yourself
Tensile strength (in MPa) of a seal was measured on five samples from each of three suppliers.
| Sample | Supplier X | Supplier Y | Supplier Z |
|---|---|---|---|
| 1 | 41 | 45 | 42 |
| 2 | 43 | 47 | 41 |
| 3 | 42 | 46 | 43 |
| 4 | 44 | 44 | 40 |
| 5 | 40 | 48 | 44 |
- Test whether the supplier means differ at α = 0.05.
- Which suppliers differ, if any?
- Is the conclusion about the suppliers, or about the samples?
Show the answer
Means: X = 42.0, Y = 46.0, Z = 42.0. SSB = 53.3, SSW = 30.0, df = 2 and 12, F = 10.67, p = 0.0022. Reject H0: the supplier means are not all equal.
Tukey: Y is higher than X (adjusted p = 0.005) and higher than Z (adjusted p = 0.005); X and Z do not differ (adjusted p = 1.000).
The conclusion concerns the supplier averages, based on random samples. If the samples were not drawn at random from each supplier’s production, the result cannot be generalized.
One-Way ANOVA: Frequently Asked Questions
What is the difference between ANOVA and a t-test?
A t-test compares two means. ANOVA compares three or more. With exactly two groups the two give the same p-value, and F equals t squared. ANOVA uses one test for all the groups, which avoids the false alarms that come from running many t-tests.
What does a significant ANOVA result tell me?
That at least one group mean differs from the others, at the risk level you chose. It does not say which means differ or by how much. Follow it with a post-hoc comparison such as Tukey, and report the sizes of the differences.
How many observations do I need per group?
It depends on the difference you need to detect and the noise in the data. As a very rough guide, five to ten per group can detect large differences, and detecting small differences takes far more. Use the Sample Size and Confidence Calculator or a power calculation, and decide the difference that matters before you collect data.
Do the groups need the same number of observations?
No, but balanced groups make the test more robust to unequal variances and non-normality, and make comparisons more efficient. If group sizes are very unequal, check the variances and consider Welch’s ANOVA.
What does R-squared mean in ANOVA?
It is the share of the total variation in the response that is explained by the factor. In the example, 67% of the variation in fill weight is explained by the head. A high R-squared does not prove cause. A low one can still be important if the effect is large compared with the tolerance.
When should I use Kruskal-Wallis instead?
When the data are clearly non-normal and the groups are small, when the response is ranked or ordinal, or when outliers are real and cannot be removed. It compares the groups using ranks instead of means, and has less power than ANOVA when ANOVA’s assumptions are met.
Sources and Further Reading
- Douglas C. Montgomery, Design and Analysis of Experiments, Wiley, chapter on the analysis of variance.
- NIST/SEMATECH, e-Handbook of Statistical Methods, section on one-way ANOVA (itl.nist.gov/div898/handbook).
- Michael H. Kutner, Christopher J. Nachtsheim, John Neter, and William Li, Applied Linear Statistical Models, McGraw-Hill.
- David S. Moore, George P. McCabe, and Bruce A. Craig, Introduction to the Practice of Statistics, Freeman.
- Minitab Support, “Methods and formulas for One-Way ANOVA” (support.minitab.com).
- Microsoft Support, “Load the Analysis ToolPak in Excel” (support.microsoft.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.