Question it answers
Is the variation different between two or more processes, or from a target?
Data needed
Measured values from each process
Key output
Ratio of standard deviations with an interval, and a p-value
Default test
Levene’s (robust); F or Bartlett only for clearly normal data
Assumptions
Independent observations; normality for F, Bartlett, and chi-square
Excel
F.TEST, F.DIST.RT, CHISQ.DIST.RT; Levene by formula
Minitab
Stat > Basic Statistics > 2 Variances; Stat > ANOVA > Test for Equal Variances
Why it matters
Variation drives capability and the validity of other tests

The Idea in Plain Language

Quality is mostly about variation. Two suppliers can have the same average fill weight and still be very different: one is steady and predictable, the other is scattered. Tests for variances ask whether the spread of one process differs from a target, or from another process.

The usual statistic is the variance (the standard deviation squared), because it behaves well mathematically. To compare two processes you take the ratio of their variances: if the processes are equally variable, the ratio is about 1. The F distribution tells you how far from 1 a ratio can wander by chance.

498 g 499 g 500 g 501 g 502 g 503 g 500.15 Supplier A 500.35 Supplier B Grand mean 500.25
Two suppliers with nearly the same average. Supplier B’s weights are spread over about 4 g, and Supplier A’s over about 2.6 g.
Why it matters. A t-test or ANOVA that assumes equal variances can be wrong if the variances differ. And in capability work, variation, not the average, usually decides whether a process meets its limits. Differences in spread are often the most valuable finding in a comparison.

When to Use Which Test

Your situationUseNotes
Two groups, normal dataF testExactly right for normal data, and badly affected by non-normality
Two or more groups, any shapeLevene’s test (median-based, Brown-Forsythe)Robust; the best general choice
Three or more groups, normal dataBartlett’s testMore powerful than Levene for normal data, but sensitive to non-normality
One group against a target standard deviationChi-square test for one varianceNeeds normal data
Non-normal dataLevene’s, or a bootstrapF and Bartlett can give false alarms
A confidence interval for variationSee Confidence IntervalsChi-square for one, F for a ratio
Default to Levene’s test. Minitab shows the F test (or Bartlett) and Levene’s test together. If they disagree, and the data are not clearly normal, trust Levene’s.

How It Works

TestStatisticDistribution under H0
F test (two variances)F = s1² / s2² (larger on top for a two-sided test)F with n1 − 1 and n2 − 1 df
Interval for the variance ratio(F / Fupper) to (F / Flower)Take square roots for the standard deviation ratio
Chi-square (one variance)χ² = (n − 1) s² / σ0²Chi-square with n − 1 df
Levene (Brown-Forsythe)One-way ANOVA on |x − group median|F with k − 1 and N − k df
BartlettCompares the pooled variance with the individual onesChi-square with k − 1 df

Levene’s test works because a group with larger spread will have larger absolute deviations from its center, so an ordinary ANOVA on those deviations picks up the difference. Using the median instead of the mean for the center makes it robust to skewed data.

Worked Example 1: Do Two Suppliers Have the Same Variation?

Fill weights from two suppliers, 15 bottles each. The means are 500.15 g and 500.35 g. The standard deviations are 0.674 g (Supplier A) and 1.236 g (Supplier B). Normality checks give p = 0.48 and 0.34, so the F test is reasonable.

  1. Hypotheses. H0: σB = σA. H1: σB ≠ σA. α = 0.05.
  2. Variances. sA² = 0.4541, sB² = 1.5284.
  3. F = 1.5284 / 0.4541 = 3.37 with 14 and 14 df.
  4. Critical value (two-sided, 2.5% in the upper tail) = 2.98. Since 3.37 > 2.98, reject H0. The two-sided p-value is 0.030.
  5. Interval for the ratio of standard deviations: 1.06 to 3.17, with an estimate of 1.83.
  6. Levene’s test (median-based) gives F = 5.18, p = 0.031, the same conclusion without needing normality.
0 3 6 Critical F = 2.98 (5% in the tail) Observed F = 3.37 two-sided p = 0.030 F value (between-group variation divided by within-group variation); df = 14, 14
The variance ratio is far enough above 1 to be beyond the 97.5th percentile of the F distribution.
Conclusion. Supplier B’s fill weights are more variable than Supplier A’s (SD 1.24 g against 0.67 g; ratio 1.83, 95% CI 1.06 to 3.17; F(14, 14) = 3.37, p = 0.030; Levene p = 0.031). The averages are close, so the difference that matters is the spread.

With three groups. Adding a third supplier with a tight spread, Bartlett’s test gives χ² = 28.77 (p < 0.0001) and Levene’s gives F = 10.16 (p = 0.0003). Both flag unequal variances. If they disagreed, you would trust Levene’s unless the data are clearly normal.

Worked Example 2: One Variance Against a Target

A process standard says the fill-weight standard deviation must not exceed 1.2 g. A sample of 20 fills has s = 1.511 g. Is the process too variable?

  1. Hypotheses. H0: σ = 1.2. H1: σ > 1.2. α = 0.05.
  2. Statistic. χ² = (20 − 1) × 1.511² / 1.2² = 30.14 on 19 df.
  3. One-sided p-value = 0.050. (Two-sided: 0.100.)
  4. Interval: the 95% interval for σ is 1.15 to 2.21 g, which includes 1.2 at the low end.
0 10 20 30 40 50 Observed = 30.1 Chi-square statistic (n − 1) s² / σ₀²
The observed statistic is in the upper tail, but only just beyond the 5% cut-off.
Conclusion. The sample standard deviation of 1.51 g is above the 1.2 g standard, and the one-sided test is borderline (p = 0.050). The interval (1.15 to 2.21 g) is wide, because a standard deviation estimated from 20 values is not precise. This test needs normal data, so check the probability plot.

Run It in Excel and Minitab

ExcelStep by step

  1. Put the data for each supplier in its own column. Compute the variances with =VAR.S(range) and the counts with =COUNT(range).
  2. F test: F = larger variance divided by smaller. Two-sided p-value: =2*F.DIST.RT(F, df1, df2) (0.0301). Critical value: =F.INV.RT(0.025, df1, df2).
  3. Built-in: =F.TEST(array1, array2) returns the two-sided p-value directly. The ToolPak also has Data > Data Analysis > F-Test Two-Sample for Variances, which reports a one-sided p-value; double it for a two-sided test.
  4. One variance: =(n-1)*s^2/sigma0^2, then =CHISQ.DIST.RT(chi, n-1) for the upper-tail p-value.
  5. Levene’s test: in new columns compute =ABS(A2-MEDIAN($A$2:$A$16)) for each value, then run Anova: Single Factor on those columns.
  6. Interval for the ratio: use the F critical values as in the table above.

Excel has no built-in Levene’s or Bartlett’s test, so use the steps above, or Minitab.

MinitabStep by step

  1. Two variances: Stat > Basic Statistics > 2 Variances. Choose Each sample is in its own column (or both in one column with a Subscripts column). Click Options to set the confidence level and the ratio.
  2. Minitab reports both the F test and Levene’s test, with an interval for the ratio, and draws an individual value plot or boxplot if you choose it under Graphs.
  3. Three or more groups: Stat > ANOVA > Test for Equal Variances. Set the Response and the Factor. Minitab shows Bartlett’s test and Levene’s test, with intervals for each group.
  4. One variance: Stat > Basic Statistics > 1 Variance. Enter the sample statistics or the column, and the hypothesized standard deviation.
  5. In the output, choose the test that matches your data: F or Bartlett for normal data, Levene’s otherwise.

The Assistant (Assistant > Hypothesis Tests > Standard Deviation tests) picks a suitable method and reports a power check.

Minitab session window: 2 Variances (typed excerpt, simplified)
Method

σ₁: StDev of Supplier B
σ₂: StDev of Supplier A
Ratio: σ₁/σ₂
F method was used. This method is accurate for normal data only.

Descriptive Statistics

Variable      N  StDev  Variance  95% CI for StDevs
Supplier B   15  1.236     1.528  (0.905, 1.950)
Supplier A   15  0.674     0.454  (0.493, 1.063)

Ratio of standard deviations

Estimated      95% CI for
     Ratio           Ratio
1.83460  (1.0630, 3.1663)

Test

Null hypothesis         H₀: σ₁ / σ₂ = 1
Alternative hypothesis  H₁: σ₁ / σ₂ ≠ 1
Significance level      α = 0.05

                 Test
Method      Statistic  DF1  DF2  P-Value
F               3.37   14   14    0.030
Levene's        5.18    1   28    0.031

Reading and Reporting

  1. Report the standard deviations (not only the variances) and their ratio with an interval.
  2. Say which test you used, and whether the data looked normal.
  3. If the variances differ, say what it means: a less predictable process, a different measurement system, a mixture of sources.
  4. Carry the finding forward: use Welch’s methods for comparing means, and note the capability consequences.
A sentence you can use. Supplier B was more variable than Supplier A (SD 1.24 g against 0.67 g; ratio 1.83, 95% CI 1.06 to 3.17; Levene’s test p = 0.031).

Common Mistakes

MistakeWhy it misleadsBetter
Using the F test or Bartlett on non-normal dataFalse alarms are commonUse Levene’s test
Testing variances as a gate before every t-testThe two-step procedure distorts error ratesUse Welch’s t-test, which does not assume equal variances
Comparing standard deviations by eyeSmall samples vary a lotTest, and give an interval for the ratio
Ignoring that a standard deviation needs a large sampleThe interval is very wideUse more data than for a mean
Reading “no significant difference in variance” as “equal”A small sample may miss a real differenceLook at the interval and the power
Treating a variance difference as a nuisanceIt may be the most important findingInvestigate the cause

Try It Yourself

Two operators measured the same 12 parts. Operator 1’s measurements had a standard deviation of 0.040 mm and Operator 2’s had 0.075 mm.

  • Test whether the operators differ in variation, assuming normal data.
  • Give a 95% interval for the ratio of standard deviations.
Show the answer

F = (0.075/0.04)² = 3.52 on 11 and 11 df; two-sided p = 0.0480. The operators differ in variation: Operator 2 is more variable.

The 95% interval for the ratio of standard deviations is 1.01 to 3.49. Operator 2’s spread is between about 1.0 and 3.5 times Operator 1’s.

Tests for Variances: Frequently Asked Questions

Why not use the F test for everything?

It is exact for normal data but very sensitive to non-normality: a skewed distribution can make it reject far more often than 5%. Levene’s test is robust, so it is the safer default.

Do I need to test for equal variances before a t-test?

No. Use Welch’s t-test, which does not assume equal variances. Testing first and then choosing a method changes the error rates and gains little.

What is the difference between Levene&rsquo;s test and Brown-Forsythe?

Levene’s original test uses deviations from the group mean. The Brown-Forsythe version uses deviations from the group median, which is more robust to skew. Minitab’s Levene’s test is the median-based one.

Why does the F test put the larger variance on top?

For a two-sided test it is a convenience: you then only need the upper tail of the F distribution and double the p-value. The result is the same as putting either variance on top and using both tails.

How many observations do I need to compare variances?

More than for means. The standard error of a standard deviation is about s divided by √(2(n − 1)), so to detect a ratio of 1.5 you need roughly 40 to 50 per group for reasonable power.

What if the variances are different?

First find out why: a different machine, a mixed population, a different gauge. Then use methods that do not assume equal variances, such as Welch’s t-test or Welch’s ANOVA, and treat the difference in spread as a result in its own right.

Sources and Further Reading

  • NIST/SEMATECH, e-Handbook of Statistical Methods, sections on tests for equal variances (itl.nist.gov/div898/handbook).
  • Morton B. Brown and Alan B. Forsythe, “Robust tests for the equality of variances,” Journal of the American Statistical Association, 1974.
  • Maurice S. Bartlett, “Properties of sufficiency and statistical tests,” Proceedings of the Royal Society A, 1937.
  • Douglas C. Montgomery and George C. Runger, Applied Statistics and Probability for Engineers, Wiley.
  • Minitab Support, “Methods and formulas for 2 Variances and Test for Equal Variances” (support.minitab.com).

This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.