- Question it answers
- Is the variation different between two or more processes, or from a target?
- Data needed
- Measured values from each process
- Key output
- Ratio of standard deviations with an interval, and a p-value
- Default test
- Levene’s (robust); F or Bartlett only for clearly normal data
- Assumptions
- Independent observations; normality for F, Bartlett, and chi-square
- Excel
- F.TEST, F.DIST.RT, CHISQ.DIST.RT; Levene by formula
- Minitab
- Stat > Basic Statistics > 2 Variances; Stat > ANOVA > Test for Equal Variances
- Why it matters
- Variation drives capability and the validity of other tests
The Idea in Plain Language
Quality is mostly about variation. Two suppliers can have the same average fill weight and still be very different: one is steady and predictable, the other is scattered. Tests for variances ask whether the spread of one process differs from a target, or from another process.
The usual statistic is the variance (the standard deviation squared), because it behaves well mathematically. To compare two processes you take the ratio of their variances: if the processes are equally variable, the ratio is about 1. The F distribution tells you how far from 1 a ratio can wander by chance.
When to Use Which Test
| Your situation | Use | Notes |
|---|---|---|
| Two groups, normal data | F test | Exactly right for normal data, and badly affected by non-normality |
| Two or more groups, any shape | Levene’s test (median-based, Brown-Forsythe) | Robust; the best general choice |
| Three or more groups, normal data | Bartlett’s test | More powerful than Levene for normal data, but sensitive to non-normality |
| One group against a target standard deviation | Chi-square test for one variance | Needs normal data |
| Non-normal data | Levene’s, or a bootstrap | F and Bartlett can give false alarms |
| A confidence interval for variation | See Confidence Intervals | Chi-square for one, F for a ratio |
How It Works
| Test | Statistic | Distribution under H0 |
|---|---|---|
| F test (two variances) | F = s1² / s2² (larger on top for a two-sided test) | F with n1 − 1 and n2 − 1 df |
| Interval for the variance ratio | (F / Fupper) to (F / Flower) | Take square roots for the standard deviation ratio |
| Chi-square (one variance) | χ² = (n − 1) s² / σ0² | Chi-square with n − 1 df |
| Levene (Brown-Forsythe) | One-way ANOVA on |x − group median| | F with k − 1 and N − k df |
| Bartlett | Compares the pooled variance with the individual ones | Chi-square with k − 1 df |
Levene’s test works because a group with larger spread will have larger absolute deviations from its center, so an ordinary ANOVA on those deviations picks up the difference. Using the median instead of the mean for the center makes it robust to skewed data.
Worked Example 1: Do Two Suppliers Have the Same Variation?
Fill weights from two suppliers, 15 bottles each. The means are 500.15 g and 500.35 g. The standard deviations are 0.674 g (Supplier A) and 1.236 g (Supplier B). Normality checks give p = 0.48 and 0.34, so the F test is reasonable.
- Hypotheses. H0: σB = σA. H1: σB ≠ σA. α = 0.05.
- Variances. sA² = 0.4541, sB² = 1.5284.
- F = 1.5284 / 0.4541 = 3.37 with 14 and 14 df.
- Critical value (two-sided, 2.5% in the upper tail) = 2.98. Since 3.37 > 2.98, reject H0. The two-sided p-value is 0.030.
- Interval for the ratio of standard deviations: 1.06 to 3.17, with an estimate of 1.83.
- Levene’s test (median-based) gives F = 5.18, p = 0.031, the same conclusion without needing normality.
With three groups. Adding a third supplier with a tight spread, Bartlett’s test gives χ² = 28.77 (p < 0.0001) and Levene’s gives F = 10.16 (p = 0.0003). Both flag unequal variances. If they disagreed, you would trust Levene’s unless the data are clearly normal.
Worked Example 2: One Variance Against a Target
A process standard says the fill-weight standard deviation must not exceed 1.2 g. A sample of 20 fills has s = 1.511 g. Is the process too variable?
- Hypotheses. H0: σ = 1.2. H1: σ > 1.2. α = 0.05.
- Statistic. χ² = (20 − 1) × 1.511² / 1.2² = 30.14 on 19 df.
- One-sided p-value = 0.050. (Two-sided: 0.100.)
- Interval: the 95% interval for σ is 1.15 to 2.21 g, which includes 1.2 at the low end.
Run It in Excel and Minitab
ExcelStep by step
- Put the data for each supplier in its own column. Compute the variances with and the counts with .
- F test: F = larger variance divided by smaller. Two-sided p-value: (0.0301). Critical value: .
- Built-in: returns the two-sided p-value directly. The ToolPak also has , which reports a one-sided p-value; double it for a two-sided test.
- One variance: , then for the upper-tail p-value.
- Levene’s test: in new columns compute for each value, then run Anova: Single Factor on those columns.
- Interval for the ratio: use the F critical values as in the table above.
Excel has no built-in Levene’s or Bartlett’s test, so use the steps above, or Minitab.
MinitabStep by step
- Two variances: . Choose Each sample is in its own column (or both in one column with a Subscripts column). Click Options to set the confidence level and the ratio.
- Minitab reports both the F test and Levene’s test, with an interval for the ratio, and draws an individual value plot or boxplot if you choose it under Graphs.
- Three or more groups: . Set the Response and the Factor. Minitab shows Bartlett’s test and Levene’s test, with intervals for each group.
- One variance: . Enter the sample statistics or the column, and the hypothesized standard deviation.
- In the output, choose the test that matches your data: F or Bartlett for normal data, Levene’s otherwise.
The Assistant () picks a suitable method and reports a power check.
Method
σ₁: StDev of Supplier B
σ₂: StDev of Supplier A
Ratio: σ₁/σ₂
F method was used. This method is accurate for normal data only.
Descriptive Statistics
Variable N StDev Variance 95% CI for StDevs
Supplier B 15 1.236 1.528 (0.905, 1.950)
Supplier A 15 0.674 0.454 (0.493, 1.063)
Ratio of standard deviations
Estimated 95% CI for
Ratio Ratio
1.83460 (1.0630, 3.1663)
Test
Null hypothesis H₀: σ₁ / σ₂ = 1
Alternative hypothesis H₁: σ₁ / σ₂ ≠ 1
Significance level α = 0.05
Test
Method Statistic DF1 DF2 P-Value
F 3.37 14 14 0.030
Levene's 5.18 1 28 0.031Reading and Reporting
- Report the standard deviations (not only the variances) and their ratio with an interval.
- Say which test you used, and whether the data looked normal.
- If the variances differ, say what it means: a less predictable process, a different measurement system, a mixture of sources.
- Carry the finding forward: use Welch’s methods for comparing means, and note the capability consequences.
Common Mistakes
| Mistake | Why it misleads | Better |
|---|---|---|
| Using the F test or Bartlett on non-normal data | False alarms are common | Use Levene’s test |
| Testing variances as a gate before every t-test | The two-step procedure distorts error rates | Use Welch’s t-test, which does not assume equal variances |
| Comparing standard deviations by eye | Small samples vary a lot | Test, and give an interval for the ratio |
| Ignoring that a standard deviation needs a large sample | The interval is very wide | Use more data than for a mean |
| Reading “no significant difference in variance” as “equal” | A small sample may miss a real difference | Look at the interval and the power |
| Treating a variance difference as a nuisance | It may be the most important finding | Investigate the cause |
Try It Yourself
Two operators measured the same 12 parts. Operator 1’s measurements had a standard deviation of 0.040 mm and Operator 2’s had 0.075 mm.
- Test whether the operators differ in variation, assuming normal data.
- Give a 95% interval for the ratio of standard deviations.
Show the answer
F = (0.075/0.04)² = 3.52 on 11 and 11 df; two-sided p = 0.0480. The operators differ in variation: Operator 2 is more variable.
The 95% interval for the ratio of standard deviations is 1.01 to 3.49. Operator 2’s spread is between about 1.0 and 3.5 times Operator 1’s.
Tests for Variances: Frequently Asked Questions
Why not use the F test for everything?
It is exact for normal data but very sensitive to non-normality: a skewed distribution can make it reject far more often than 5%. Levene’s test is robust, so it is the safer default.
Do I need to test for equal variances before a t-test?
No. Use Welch’s t-test, which does not assume equal variances. Testing first and then choosing a method changes the error rates and gains little.
What is the difference between Levene’s test and Brown-Forsythe?
Levene’s original test uses deviations from the group mean. The Brown-Forsythe version uses deviations from the group median, which is more robust to skew. Minitab’s Levene’s test is the median-based one.
Why does the F test put the larger variance on top?
For a two-sided test it is a convenience: you then only need the upper tail of the F distribution and double the p-value. The result is the same as putting either variance on top and using both tails.
How many observations do I need to compare variances?
More than for means. The standard error of a standard deviation is about s divided by √(2(n − 1)), so to detect a ratio of 1.5 you need roughly 40 to 50 per group for reasonable power.
What if the variances are different?
First find out why: a different machine, a mixed population, a different gauge. Then use methods that do not assume equal variances, such as Welch’s t-test or Welch’s ANOVA, and treat the difference in spread as a result in its own right.
Sources and Further Reading
- NIST/SEMATECH, e-Handbook of Statistical Methods, sections on tests for equal variances (itl.nist.gov/div898/handbook).
- Morton B. Brown and Alan B. Forsythe, “Robust tests for the equality of variances,” Journal of the American Statistical Association, 1974.
- Maurice S. Bartlett, “Properties of sufficiency and statistical tests,” Proceedings of the Royal Society A, 1937.
- Douglas C. Montgomery and George C. Runger, Applied Statistics and Probability for Engineers, Wiley.
- Minitab Support, “Methods and formulas for 2 Variances and Test for Equal Variances” (support.minitab.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.