Question it answers
Do three or more groups have the same average?
Data needed
One continuous response and one factor with 3 or more levels
Null hypothesis
All group means are equal
Key output
F statistic, p-value, and R-squared
Assumptions
Independent observations, normal residuals, similar spread in each group
Excel
Data Analysis > Anova: Single Factor
Minitab
Stat > ANOVA > One-Way
Next step
A post-hoc comparison to find which groups differ

The Idea in Plain Language

Three filling heads on one line are each set to deliver 500 g. You weigh six fills from each head. The averages come out at 498.5 g, 500.7 g, and 499.3 g. Are the heads really different, or is this the ordinary scatter you get between any two fills from the same head?

Analysis of variance (ANOVA) answers by comparing two kinds of variation. The first is how far the group averages sit from each other: the signal. The second is how much individual fills vary within a single head: the noise. If the signal is much larger than the noise, the groups probably differ. If it is about the same size, the differences are what chance alone would produce.

Despite the name, one-way ANOVA tests means. It uses variances to do it, which is where the name comes from. “One-way” means one factor, here the filling head, with three or more levels.

Why not run several t-tests?

With three groups you could run three pairwise t-tests. Each test carries a 5% risk of a false alarm, and the risks add up. The more groups you compare, the more likely it is that at least one pair looks different by chance. ANOVA asks the single question first: is there any difference at all?

Chance of at least one false alarm at a 5% risk per test 3 groups: 3 pairwise tests 14.3% 5 groups: 10 pairwise tests 40.1% 6 groups: 15 pairwise tests 53.7%
A false alarm is claiming a difference that is not there. Testing every pair separately makes the chance of at least one false alarm climb quickly, which is the reason to start with one overall test.

When to Use One-Way ANOVA

Your situationUseWhy
One continuous response, one factor with 3 or more levels, and you want to compare the meansOne-way ANOVAThis is the case it was built for
Only 2 groupsTwo-sample t-test (ANOVA gives the same p-value, with F = t²)Simpler to report; see Hypothesis Testing
Two factors at once (for example head and shift)Two-way ANOVAShows the effect of each factor and their interaction
Data are skewed or have outliers and the groups are smallKruskal-Wallis or Mood’s medianRank-based tests do not need normal data
You want to compare variation, not meansTest for equal variances (Levene)ANOVA compares averages only
The input is continuous (temperature, speed)RegressionIt uses the numeric information that grouping throws away
The response is pass/fail or a countChi-square or proportion testsANOVA needs a continuous response
Not sure? Use the Stat Dojo method selector. Answer three questions about your data and it points you to the right test, the Excel route, and the Minitab menu.

How It Works: Signal and Noise

497 g 498 g 499 g 500 g 501 g 502 g 498.53 Head 1 500.67 Head 2 499.32 Head 3 Grand mean 499.51
Six fills from each head. The gold bars are the group averages, and the dashed line is the average of all 18 fills. The averages differ more than the spread within each head would suggest.

ANOVA measures the total variation in all 18 fills and splits it into two parts.

Total sum of squares = 20.73 13.97 67.4% Between heads (signal) 6.75 32.6% Within heads (noise)
Of the total variation, about two thirds is explained by which head filled the container. The rest is variation within each head.
QuantityFormulaIn words
Between-group sum of squares (SSB)Σ ni (x̄i − x̄)²How far each group mean is from the grand mean, weighted by group size
Within-group sum of squares (SSW)Σ Σ (xij − x̄i)²How far each value is from its own group mean
Total sum of squares (SST)SSB + SSWTotal variation around the grand mean
Degrees of freedomBetween k − 1; within N − k; total N − 1k groups, N observations in all
Mean squaresMSB = SSB / (k − 1); MSW = SSW / (N − k)Variation per degree of freedom: signal and noise
F statisticF = MSB / MSWSignal divided by noise
p-valueP(Fk−1, N−k ≥ F observed)Chance of a signal this large if all means were equal
R-squaredSSB / SSTShare of the total variation explained by the factor

If the group means are all equal, the signal and the noise estimate the same thing, and F is close to 1. The larger F is, the less plausible it is that all the means are equal.

0 3 6 9 12 15 18 Critical F = 3.68 (5% in the tail) Observed F = 15.52 p-value = 0.0002 F value (between-group variation divided by within-group variation); df = 2, 15
The F distribution for 2 and 15 degrees of freedom. The shaded tail holds 5% of the area. An observed F beyond the critical value is rejected as chance.

Hypotheses and Assumptions

HypothesisStatement
Null (H0)μ1 = μ2 = μ3: all group means are equal
Alternative (H1)At least one mean is different. It does not say all means differ, or which one
AssumptionWhat it meansHow to checkIn the example
Independent observationsOne fill does not influence anotherThink about how the data were collected; plot residuals in time orderFills were taken in random order from each head
Normal residualsThe scatter within each group is roughly bell-shapedNormal probability plot of residuals; Anderson-Darling or Shapiro-Wilk testShapiro-Wilk p = 0.23: no sign of non-normality
Equal variancesEach group has about the same spreadLevene’s test (Minitab and the calculator use the median-based version); the largest standard deviation should be under about twice the smallestLevene p = 0.79; sd ratio 1.20
No extreme outliersOne wild value can move a mean and inflate the noiseDot plot or box plot of each groupNo outliers in the dot plot
Continuous responseA measured value, not a count or a pass/failLook at the data typeWeight in grams

ANOVA is fairly forgiving. With groups of similar size, moderate departures from normality or equal variances change the p-value very little. Unequal group sizes combined with unequal variances are the dangerous case. See the section on what to do when the assumptions fail.

Worked Example by Hand

A packaging engineer wants to know whether three filling heads deliver the same average weight. The target is 500 g.

Fill weights in grams (6 fills per head)
FillHead 1Head 2Head 3
1498.2500.4499.0
2499.1501.2499.8
3497.6499.8498.6
4498.8500.9500.2
5499.5501.6499.4
6498.0500.1498.9
HeadnMeanStandard deviationSum of squares within the head
Head 16498.5330.7202.593
Head 26500.6670.6862.353
Head 36499.3170.6011.808
  1. State the hypotheses. H0: all three head means are equal. H1: at least one differs. Choose α = 0.05.
  2. Grand mean. The average of all 18 fills is 499.506 g.
  3. Between-head sum of squares. SSB = 6(498.533 − 499.506)² + 6(500.667 − 499.506)² + 6(499.317 − 499.506)² = 5.671 + 8.089 + 0.214 = 13.974.
  4. Within-head sum of squares. Add the three within-head sums: 2.593 + 2.353 + 1.808 = 6.755.
  5. Degrees of freedom. Between = 3 − 1 = 2. Within = 18 − 3 = 15. Total = 18 − 1 = 17.
  6. Mean squares. MSB = 13.974 / 2 = 6.987. MSW = 6.755 / 15 = 0.450.
  7. F statistic. F = 6.987 / 0.450 = 15.52.
  8. Decide. The critical value for 2 and 15 degrees of freedom at α = 0.05 is 3.68. Since 15.52 > 3.68, reject H0. The p-value is 0.0002.
SourceSSdfMSFp
Between heads13.97426.98715.520.0002
Within heads (error)6.755150.450
Total20.72917
Conclusion. The heads do not all deliver the same average weight (F(2, 15) = 15.52, p = 0.0002). The head explains 67% of the variation in fill weight. ANOVA does not say which heads differ; that is the job of a post-hoc comparison.

Run It in Excel and Minitab

ExcelStep by step

  1. Put each head in its own column with the head name in row 1, and the six weights below it (here A1:C7).
  2. Check that Data Analysis appears on the Data tab. If not: File > Options > Add-ins, choose Excel Add-ins, click Go, and tick Analysis ToolPak.
  3. Choose Data > Data Analysis > Anova: Single Factor and click OK.
  4. Set Input Range to A1:C7, Grouped By to Columns, tick Labels in first row, and leave Alpha at 0.05.
  5. Pick an output location, click OK, and read the ANOVA table at the bottom. Compare P-value with alpha, or F with F crit.

With formulas instead. =DEVSQ(A2:A7)+DEVSQ(B2:B7)+DEVSQ(C2:C7) gives SSW, =DEVSQ(A2:C7) gives SST, and SSB is the difference. Then =F.DIST.RT(F,2,15) gives the p-value and =F.INV.RT(0.05,2,15) the critical F.

Excel has no built-in Tukey test or equal-variance test for several groups. Use the HSD formula in the post-hoc page, or Minitab.

MinitabStep by step

  1. Stack the data: one column for the weights (Weight) and one for the head (Head).
  2. Choose Stat > ANOVA > One-Way. Leave Response data are in one column for all factor levels selected.
  3. Set Response to Weight and Factor to Head.
  4. Click Comparisons, tick Tukey, and keep Grouping information and Tests and intervals ticked.
  5. Click Graphs and tick Individual value plot, Boxplot of data, and Four in one residual plots. Click OK twice.
  6. Check the equal-variance assumption with Stat > ANOVA > Test for Equal Variances (Response: Weight; Factors: Head).

Unequal variances? In the One-Way dialog click Options and clear Assume equal variances. Minitab then runs Welch’s ANOVA, and the comparisons use Games-Howell.

The Assistant (Assistant > Hypothesis Tests > One-Way ANOVA) runs the same analysis with a report card that flags assumption problems.

What the output looks like

The highlighted numbers are the ones to read first: F, the p-value, and R-squared.

Excel Data Analysis ToolPak output (Anova: Single Factor)
Anova: Single Factor

SUMMARY
Groups     Count         Sum      Average    Variance
Head 1         6      2991.2     498.5333      0.5187
Head 2         6      3004.0     500.6667      0.4707
Head 3         6      2995.9     499.3167      0.3617

ANOVA
Source of Variation           SS   df        MS         F    P-value   F crit
Between Groups           13.9744    2    6.9872   15.5157   0.000223   3.6823
Within Groups             6.7550   15    0.4503

Total                    20.7294   17
Minitab session window (typed excerpt, simplified)
Method
Null hypothesis         All means are equal
Alternative hypothesis  At least one mean is different
Significance level      α = 0.05

Equal variances were assumed for the analysis.

Factor Information

Factor  Levels  Values
Head         3  Head 1, Head 2, Head 3

Analysis of Variance

Source  DF   Adj SS  Adj MS  F-Value  P-Value
Head     2   13.974   6.987   15.52    0.000
Error   15    6.755   0.450
Total   17   20.729

Model Summary

       S    R-sq  R-sq(adj)  R-sq(pred)
0.67107  67.41%     63.07%      53.08%

Means

Factor  N     Mean  StDev       95% CI
Head 1   6  498.533  0.720  (497.949, 499.117)
Head 2   6  500.667  0.686  (500.083, 501.251)
Head 3   6  499.317  0.601  (498.733, 499.901)

Pooled StDev = 0.67107

Tukey Pairwise Comparisons

Grouping Information Using the Tukey Method and 95% Confidence

Factor  N     Mean  Grouping
Head 2   6  500.667  A
Head 3   6  499.317  B
Head 1   6  498.533  B

Means that do not share a letter are significantly different.

Tukey Simultaneous Tests for Differences of Means

Difference            Difference      SE of                          Adjusted
of Levels              of Means  Difference      95% CI           T-Value   P-Value
Head 2 - Head 1        2.133      0.387  ( 1.127,  3.140)     5.51     0.000
Head 3 - Head 1        0.783      0.387  (-0.223,  1.790)     2.02     0.141
Head 3 - Head 2       -1.350      0.387  (-2.356, -0.344)    -3.48     0.009
Same answer, different layout. Both programs report F = 15.52 and p ≈ 0.0002. Minitab adds the model summary, the group means with intervals, and the Tukey comparisons. Excel gives the table only.

Reading and Reporting the Result

  1. Look at the p-value first. p = 0.0002 is below 0.05, so at least one mean differs.
  2. Check the effect size. R-squared is 67%: the head explains about two thirds of the variation. Adjusted R-squared (63%) allows for the number of groups; predicted R-squared (53%) estimates how well the model would predict new data.
  3. Find out which groups differ. Tukey’s test shows Head 2 is 2.13 g heavier than Head 1 (95% CI 1.13 to 3.14 g, adjusted p < 0.001) and 1.35 g heavier than Head 3 (adjusted p = 0.009). Heads 1 and 3 cannot be told apart (adjusted p = 0.141).
  4. Translate it into practice. A head that fills 2.1 g heavier than another uses a large part of a typical ±3 g tolerance. Statistically clear and practically important are not the same thing, so compare the difference with what matters to the customer.
A sentence you can use. A one-way ANOVA showed that mean fill weight differed among the three filling heads, F(2, 15) = 15.52, p = 0.0002, R² = 67%. Tukey comparisons showed that Head 2 filled 2.13 g heavier than Head 1 and 1.35 g heavier than Head 3, while Heads 1 and 3 did not differ significantly.

When the Assumptions Do Not Hold

ProblemWhat it doesWhat to do
Unequal variancesInflates false alarms, especially when group sizes differWelch’s ANOVA with Games-Howell comparisons; or transform the data; investigate why the spread differs, which is often the real finding
Non-normal residuals, small groupsDistorts the p-valueKruskal-Wallis or Mood’s median test (Minitab: Stat > Nonparametric > Kruskal-Wallis); or transform the response
OutliersPull a mean and inflate the noiseFind the cause first. Correct a recording error; do not delete a real value to get a better p-value. Report results with and without it
Observations not independentMakes every p-value too smallRedesign the collection: randomize, or use a model that includes the structure (blocks, repeated measures)
Very unequal group sizesMakes the test sensitive to unequal variancesAim for balanced groups; use Welch’s ANOVA when you cannot

Common Mistakes

MistakeWhy it misleadsBetter
Running many t-tests instead of ANOVAFalse alarms add upOne ANOVA, then a post-hoc test that controls the overall error rate
Concluding “all groups differ”A significant F says only that at least one mean differsUse Tukey or another post-hoc comparison to say which
Reading a significant p-value as a large effectWith enough data a tiny difference is significantReport the differences, intervals, and R-squared
Reading a non-significant p-value as no differenceA small sample may not detect a real differenceCheck the sample size and power
Skipping the residual plotsAssumption problems stay hiddenAlways look at the four-in-one residual plots
Using ANOVA on counts or pass/fail dataThe response is not continuousUse proportion or chi-square tests
Comparing groups with different conditionsMachine, shift, and operator get mixed up with the factorRandomize and block, or add the other factor to the model

Try It Yourself

Tensile strength (in MPa) of a seal was measured on five samples from each of three suppliers.

SampleSupplier XSupplier YSupplier Z
1414542
2434741
3424643
4444440
5404844
  • Test whether the supplier means differ at α = 0.05.
  • Which suppliers differ, if any?
  • Is the conclusion about the suppliers, or about the samples?
Show the answer

Means: X = 42.0, Y = 46.0, Z = 42.0. SSB = 53.3, SSW = 30.0, df = 2 and 12, F = 10.67, p = 0.0022. Reject H0: the supplier means are not all equal.

Tukey: Y is higher than X (adjusted p = 0.005) and higher than Z (adjusted p = 0.005); X and Z do not differ (adjusted p = 1.000).

The conclusion concerns the supplier averages, based on random samples. If the samples were not drawn at random from each supplier’s production, the result cannot be generalized.

One-Way ANOVA: Frequently Asked Questions

What is the difference between ANOVA and a t-test?

A t-test compares two means. ANOVA compares three or more. With exactly two groups the two give the same p-value, and F equals t squared. ANOVA uses one test for all the groups, which avoids the false alarms that come from running many t-tests.

What does a significant ANOVA result tell me?

That at least one group mean differs from the others, at the risk level you chose. It does not say which means differ or by how much. Follow it with a post-hoc comparison such as Tukey, and report the sizes of the differences.

How many observations do I need per group?

It depends on the difference you need to detect and the noise in the data. As a very rough guide, five to ten per group can detect large differences, and detecting small differences takes far more. Use the Sample Size and Confidence Calculator or a power calculation, and decide the difference that matters before you collect data.

Do the groups need the same number of observations?

No, but balanced groups make the test more robust to unequal variances and non-normality, and make comparisons more efficient. If group sizes are very unequal, check the variances and consider Welch’s ANOVA.

What does R-squared mean in ANOVA?

It is the share of the total variation in the response that is explained by the factor. In the example, 67% of the variation in fill weight is explained by the head. A high R-squared does not prove cause. A low one can still be important if the effect is large compared with the tolerance.

When should I use Kruskal-Wallis instead?

When the data are clearly non-normal and the groups are small, when the response is ranked or ordinal, or when outliers are real and cannot be removed. It compares the groups using ranks instead of means, and has less power than ANOVA when ANOVA’s assumptions are met.

Sources and Further Reading

  • Douglas C. Montgomery, Design and Analysis of Experiments, Wiley, chapter on the analysis of variance.
  • NIST/SEMATECH, e-Handbook of Statistical Methods, section on one-way ANOVA (itl.nist.gov/div898/handbook).
  • Michael H. Kutner, Christopher J. Nachtsheim, John Neter, and William Li, Applied Linear Statistical Models, McGraw-Hill.
  • David S. Moore, George P. McCabe, and Bruce A. Craig, Introduction to the Practice of Statistics, Freeman.
  • Minitab Support, “Methods and formulas for One-Way ANOVA” (support.minitab.com).
  • Microsoft Support, “Load the Analysis ToolPak in Excel” (support.microsoft.com).

This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.