Question it answers
How many observations do I need to detect a difference that matters?
You must decide
The smallest important difference, the standard deviation, α, and the target power
Key output
Sample size per group, or the power of a planned test
Rule of thumb
Halve the difference, and you need about four times the data
Typical targets
α = 0.05, power 0.80 (0.90 for important decisions)
Excel
NORM.S.INV formulas; no exact t-power
Minitab
Stat > Power and Sample Size
Online
Sample Size and Confidence Calculator

The Idea in Plain Language

Before you collect data, ask: how much do I need? Too little, and a real improvement can slip past unnoticed. Too much, and you waste time and money measuring something you already know. Power is the probability that your test will detect a real effect of a given size. It is 1 − β, where β is the chance of a Type II error (missing a real change).

Power is set by four things that trade off against each other. If you fix any three, the fourth is determined:

FactorEffect on powerUnder your control?
Effect size: the smallest difference that matters (δ)Bigger difference, more powerYou choose what matters; you do not choose what is real
Variation: the standard deviation (σ)Less noise, more powerPartly: better measurement, tighter process, blocking
Sample size (n)More data, more powerYes
Significance level (α)Larger α, more power (and more false alarms)Yes, but it is a risk choice
Critical value H0 trueH1 true (a real shift) Type I error (alpha): a false alarm Type II error (beta): a missed shift Power (1 - beta): the shift is detected Test statistic (in standard errors), small sample
A small sample: the curves overlap, and the test often misses the shift.
Critical value H0 trueH1 true (a real shift) Type I error (alpha): a false alarm Type II error (beta): a missed shift Power (1 - beta): the shift is detected Test statistic (in standard errors), larger sample
A larger sample: the curves separate and the shift is detected reliably.
The question that matters is the effect size. “How many samples do I need?” has no answer until you say what difference is worth detecting. Choose the smallest difference that would change a decision: a tolerance, a cost break-even, or a customer requirement.

Planning a Sample Size, Step by Step

  1. State the question and the test (one mean, two means, proportions, ANOVA).
  2. Set the smallest difference that matters (δ), in the units of the data.
  3. Estimate the standard deviation from past data, a pilot, or the specification range. If you must guess, guess high.
  4. Set α and the target power (commonly 0.05 and 0.80; use 0.90 for important decisions).
  5. Calculate n, and round up. Add an allowance for lost or unusable data.
  6. Check that you can afford it. If not, change the plan: a bigger effect size, less noise, or a different design, rather than quietly accepting low power.
TestSample size per group (approximation)Notes
One mean, against a targetn = ((zα/2 + zβ) σ / δ)²Add 1 or 2 for the t distribution
Two meansn = 2 ((zα/2 + zβ) σ / δ)²Per group; add 1 or 2 for t
Two proportionsn = (zα/2 √(2 p̄ q̄) + zβ √(p1q1 + p2q2))² / (p1 − p2)²Per group; rates near 0 need large samples
Estimating a mean to within ±En = (zα/2 σ / E)²Add a little for t
Estimating a proportion to within ±En = zα/2² p(1 − p) / E²Use p = 0.5 if unknown (worst case)

Here zα/2 = 1.96 for α = 0.05, and zβ = 0.84 for 80% power or 1.28 for 90%. Software uses the exact noncentral t distribution and returns slightly larger values than the formulas.

Worked Example 1: Comparing Two Means

A team will compare fill weights from two heads. A difference of 1.0 g would matter, and past data suggest σ = 1.2 g. How many fills per head give 80% power at α = 0.05?

  1. Standardize the effect: d = δ / σ = 1.0 / 1.2 = 0.83.
  2. Formula: n = 2 ((1.96 + 0.84) × 1.2 / 1.0)² = 22.6, so about 23 per group.
  3. Exact (noncentral t): 24 per group gives power 0.807. For 90% power, 32 per group.
  4. Total: twice that, since there are two groups.
Effect size d (difference / σ)n per group, 80% powern per group, 90% power
0.2394527
0.56486
0.82634
1.01723
1.5911
2.067
0% 20% 40% 60% 80% 100% 0 10 20 30 40 50 60 Target 80% d = 0.5 d = 0.8 d = 1.0 d = 1.5 Sample size per group Power
Power rises with sample size and effect size. A large effect (d = 1.5) is found with about nine per group, while detecting a half-standard-deviation shift needs about 64.
Halving the difference you want to detect quadruples the sample. From d = 1.0 to d = 0.5 the requirement goes from 17 to 64 per group. Decide carefully what difference is worth chasing.

Worked Example 2: Comparing Two Proportions

A project aims to cut a defect rate from 4% to 2%. How many units per group are needed to detect that with 80% power?

Formula: n = (1.96 √(2 × 0.03 × 0.97) + 0.84 √(0.04 × 0.96 + 0.02 × 0.98))² / (0.04 − 0.02)² = 1,141 per group (1,527 for 90% power).

0% 20% 40% 60% 80% 100% 0 200 400 600 800 1000 1200 1400 1600 Target 80% 4% to 2% 4% to 1% Sample size per group Power
Defect-rate comparisons need large samples. Detecting 4% to 2% needs more than a thousand per group, and a bigger drop to 1% needs about 424.

Rates are expensive to compare because each unit gives only a pass or a fail, which carries little information. If you can measure a continuous quantity (a dimension, a time) in place of counting defects, you can often cut the sample size dramatically.

Worked Example 3: Estimating a Mean to a Given Precision

You want to estimate a mean cycle time to within ± 0.5 s (95% confidence), with σ = 1.5 s. Formula: n = (1.96 × 1.5 / 0.5)² = 35. Allowing for the t distribution, 38 observations give a half-width of at most 0.5 s.

This is a precision question, not a hypothesis test, so no power is needed. See Confidence Intervals.

What About Power After the Fact?

Suppose the 25-fill test on the P-Values page had been planned to detect a 0.8 g shift with σ = 1.8 g (d = 0.44). Its power was 0.57, and it would have needed 42 fills for 80%.

Do not compute “observed power” from your own result. It is a mathematical function of the p-value and adds no information. Calculate power before the study, using the effect size you decided mattered. After the study, use the confidence interval to say what the data can and cannot rule out.

Run It in Excel and Minitab

ExcelStep by step

  1. Enter α, power, σ, and δ in cells.
  2. Two means, per group: =2*((NORM.S.INV(1-alpha/2)+NORM.S.INV(power))*sigma/delta)^2, rounded up, plus 1 or 2 for the t distribution.
  3. One mean: drop the leading 2.
  4. Power for a given n (approximate): =NORM.S.DIST(delta/(sigma*SQRT(2/n))-NORM.S.INV(1-alpha/2),TRUE).
  5. Two proportions: enter the formula from the table above, using =NORM.S.INV() for the z values.
  6. Precision of a mean: =(NORM.S.INV(1-alpha/2)*sigma/E)^2.

Excel cannot compute the exact power of a t-test, because it has no noncentral t function. The formulas above run a little low (they understate n by 1 or 2). Use Minitab or the site’s calculator for exact values.

MinitabStep by step

  1. Choose Stat > Power and Sample Size > 2-Sample t (or 1-Sample t, Paired t, 1 Proportion, 2 Proportions, One-Way ANOVA, and others).
  2. Fill in two of the three boxes: Sample sizes, Differences, Power values. Leave the one you want calculated blank. Enter the Standard deviation.
  3. To see the trade-offs, enter several differences or powers separated by spaces, and click Graph to draw the power curve.
  4. Click Options to choose the alternative hypothesis (not equal, less than, greater than) and α.
  5. Click OK and read the table: sample size is per group.
  6. For a design with several factors, use Stat > Power and Sample Size > 2-Level Factorial Design or the equivalent for your design.

The Assistant (Assistant > Hypothesis Tests) shows a power report with each test, so you can see what difference your data could have detected.

Minitab: Power and Sample Size, 2-Sample t (typed excerpt, simplified)
Power and Sample Size

2-Sample t Test

Testing mean 1 = mean 2 (versus ≠)
Calculating power for mean 1 = mean 2 + difference
α = 0.05  Assumed standard deviation = 1.2

        Sample
Difference    Size    Power
         1      24  0.8068

The sample size is for each group.
Minitab: Power and Sample Size, 2 Proportions (typed excerpt, simplified)
Power and Sample Size

2 Proportions

Testing comparison p = 0.02 (versus ≠)
Calculating power for comparison p = 0.02
α = 0.05

Comparison  Sample  Target
         p    Size   Power  Actual Power
     0.02    1141     0.8        0.8001

The sample size is for each group.
Try it online. The Sample Size and Confidence Calculator and the Control Test Sample Size Calculator on this site cover common cases.

Reporting a Sample-Size Plan

  1. State the test, α, and the power you aimed for.
  2. State the effect size, and why it matters.
  3. State where the standard deviation came from.
  4. Report the n per group and the total, and what you allowed for lost data.
A sentence you can use. With α = 0.05, a standard deviation of 1.2 g (from the last 30 days of data), and a smallest difference of interest of 1.0 g, 24 fills per head give 80% power in a two-sided two-sample t-test.

Common Mistakes

MistakeWhy it misleadsBetter
Using the sample size you had instead of planning oneThe test may be unable to detect anything that mattersCalculate n for the effect that matters
Using a guessed standard deviation without a checkA low guess leaves the test underpoweredUse real data or a pilot; guess high
Rounding downPower falls below the targetAlways round up
Computing power from the observed resultIt adds no informationUse the interval instead
Ignoring the per-group versus total differenceA factor of two errorCheck which one the software reports
Powering for a big effect to save money, then claiming to have tested small onesA negative result says nothing about small effectsSay what size of effect the test could detect
Planning for the primary question but then slicing into subgroupsEach subgroup is underpoweredPlan for the subgroups, or treat them as exploratory

Try It Yourself

You want to detect a drop in average response time of 2 seconds with a standard deviation of 3 seconds, using a paired design. How many pairs give 90% power at α = 0.05?

Show the answer

For a paired test, the effect size is d = 2 / 3 = 0.67 (if 3 s is the standard deviation of the differences). The exact calculation gives 26 pairs for 90% power (20 for 80%). The approximation ((1.96 + 1.28) × 3 / 2)² = 23.6 gives about 24, to which you add 1 or 2 for the t distribution.

In a paired design, use the standard deviation of the differences between pairs, not of the individual values. Pairing helps because the differences usually vary less than the raw values.

Sample Size and Power: Frequently Asked Questions

What power should I aim for?

80% is the usual minimum. Use 90% or more when missing a real effect is costly or the study cannot be repeated. A power of 50% means a real effect of that size is missed half the time.

How do I choose the effect size?

Ask what difference would change a decision: a tolerance, a cost break-even, a customer requirement, or the smallest improvement worth implementing. Do not use the size of effect you hope to find, or the size you expect, unless that is also what matters.

What if I do not know the standard deviation?

Use data from a similar process, a pilot of 10 to 20 observations, or an estimate from the specification range (range divided by about 6). When in doubt, use a larger value. A sample plan that relies on an optimistic guess is likely to be underpowered.

Is a bigger sample always better?

Past a point, the gain is small, and a very large sample makes trivial differences significant. Plan for the smallest difference that matters, and no more.

What is the relationship between power and the confidence interval?

They are linked. A study with high power for a difference δ will give a confidence interval narrower than about δ, so it can show whether a difference of that size is present. Planning for interval width is an alternative to planning for power.

Does a sample size calculation guarantee a significant result?

No. With 80% power, there is still a 20% chance of missing a real effect of the planned size, and no chance of finding a difference that is not there. It sets the odds before you start.

Sources and Further Reading

  • Jacob Cohen, Statistical Power Analysis for the Behavioral Sciences, Lawrence Erlbaum.
  • Douglas C. Montgomery and George C. Runger, Applied Statistics and Probability for Engineers, Wiley, chapters on sample size.
  • NIST/SEMATECH, e-Handbook of Statistical Methods, sections on sample sizes (itl.nist.gov/div898/handbook).
  • Minitab Support, “Methods and formulas for Power and Sample Size” (support.minitab.com).
  • John M. Hoenig and Dennis M. Heisey, “The abuse of power,” The American Statistician, 2001.

This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.