Question it answers
Is a defect rate on target, or different between two processes?
Data needed
Counts of events out of the number of units, for each sample
Key output
Rate(s), difference with interval, z or exact p-value
Watch for
Small counts: use exact tests; large samples are needed for rare events
Assumptions
Independent units; random sampling; for the normal method, enough events
Excel
NORM.S.DIST, BINOM.DIST, CHISQ.TEST formulas
Minitab
Stat > Basic Statistics > 1 Proportion, 2 Proportions
Next step
Chi-square tests for larger tables

The Idea in Plain Language

Many quality questions are about rates: the fraction of parts that fail, orders that ship late, invoices with errors. Each unit is a yes or a no, and you want to know whether the rate differs from a target, or between two processes. Tests for proportions compare the observed fraction with what chance would produce, in the same way a t-test does for means.

Proportions behave differently from measurements in two ways. Each unit carries very little information (one bit), so you need large samples, especially for rare events. And the spread depends on the rate itself: the standard error of a proportion p is √(p(1 − p)/n), so it is largest near 50% and shrinks toward 0 and 100%.

If you can measure instead of count, do. A continuous measurement of each part (a dimension, a time) carries far more information than pass or fail, and needs a much smaller sample for the same power. See Sample Size and Power.

When to Use Which Test

Your situationUseWhy
One rate against a target1-proportion testIs the defect rate above 2%?
Two rates from independent samples2-proportion testIs Line 1 worse than Line 2?
Small counts (fewer than about 5 events or non-events)Exact tests (binomial, Fisher’s)The normal approximation is unreliable
Three or more rates, or two categories with several levelsChi-square testCompares the whole table
Counts of defects per unit (several per unit possible)Poisson rate testNot pass/fail
A rate over timep chart (SPC)Tests for a shift as data arrive
A measured valuet-testMore power per unit

How It Works

TestStatisticNotes
1-proportion (normal)z = (p̂ − p0) / √(p0(1 − p0)/n)Uses the hypothesized rate for the standard error. Needs n p0 and n(1 − p0) both at least about 5 to 10
1-proportion (exact)Binomial probability of the observed count or something more extremeAlways valid; used by Minitab by default
2-proportionz = (p̂1 − p̂2) / √(p̂(1 − p̂)(1/n1 + 1/n2)), where p̂ is the pooled ratePooled for the test, because H0 says the rates are equal; unpooled for the interval
2-proportion (exact)Fisher’s exact test, from the hypergeometric distributionUse when counts are small
Interval for a rateWilson or exact (see the Confidence Intervals page)The simple Wald interval is poor for small counts
Effect measureFormulaMeaning
Difference in ratesp1 − p2Percentage points; easiest to interpret
Relative risk (ratio)p1 / p2How many times higher
Odds ratio[p1/(1 − p1)] / [p2/(1 − p2)]Used in logistic regression; close to the ratio for rare events

Worked Example 1: One Proportion

A supplier’s contract allows a 2% defect rate. In a sample of 400 parts, 14 were defective (3.5%). Is the rate above 2%?

  1. Hypotheses. H0: p = 0.02. H1: p > 0.02 (one-sided: only a higher rate is a problem). α = 0.05.
  2. Check the approximation. n p0 = 8 and n(1 − p0) = 392. The first is borderline, so compare with the exact test.
  3. Normal test. SE = √(0.02 × 0.98 / 400) = 0.0070. z = (0.035 − 0.02) / 0.0070 = 2.14. One-sided p = 0.0161; two-sided p = 0.0321.
  4. Exact test. The binomial probability of 14 or more defectives when p = 0.02 is 0.0327 (one-sided); two-sided 0.0458.
  5. Interval. The 95% exact interval for the defect rate is 1.93% to 5.80% (Wilson: 2.10% to 5.79%; Wald: 1.70% to 5.30%).
Conclusion. The defect rate in the sample is 3.5%, above the 2% limit (exact one-sided p = 0.033; 95% exact CI 1.9% to 5.8%). The evidence is moderate. The normal approximation gives a smaller p-value (0.016) than the exact test, which is why the exact test is preferred when expected counts are small (here, 8 expected defectives).

Worked Example 2: Two Proportions

Two packaging lines were checked. Line 1 had 38 defective packs in 520 (7.3%). Line 2 had 21 in 480 (4.4%). Is Line 1 worse?

Line 1 (38 of 520) 7.3% Line 2 (21 of 480) 4.4% Defect rate (percent)
The observed defect rates. Is a gap of this size more than chance?
  1. Hypotheses. H0: p1 = p2. H1: p1 ≠ p2. α = 0.05.
  2. Pooled rate p̂ = (38 + 21) / (520 + 480) = 0.0590.
  3. Standard error = √(0.0590 × 0.9410 × (1/520 + 1/480)) = 0.0149.
  4. z = (0.0731 − 0.0437) / 0.0149 = 1.97, so p = 0.0493. (As a chi-square test of the 2 × 2 table, χ² = 3.87 = z², the same p-value. Fisher’s exact test gives 0.0595.)
  5. Interval for the difference (unpooled SE = 0.0147): 2.93 ± 2.89 = 0.0 to 5.8 percentage points.
  6. Effect sizes. Relative risk = 1.67: Line 1 produces 1.7 times as many defects per pack. Odds ratio = 1.72.
-4 -3 -2 -1 0 1 2 3 4 Observed z = 1.97 z statistic (p = 0.049)
The observed z is about two standard errors from zero, right at the edge of the rejection region.
-1 0 1 2 3 4 5 6 7 No difference Line 1 − Line 2 (normal approximation) Difference in defect rate (percentage points), 95% interval
The interval for the difference only just excludes zero: the true difference could be close to nothing, or as large as about 6 points.
Conclusion. The evidence that Line 1 is worse is borderline. Its defect rate was 7.3% against 4.4% (difference 2.9 percentage points, 95% CI 0.04 to 5.8). The z test gives p = 0.049, just under 0.05, but Fisher’s exact test gives p = 0.060, just over. When the two methods straddle the cut-off, the honest statement is that the data are inconclusive: collect more units before acting.

Run It in Excel and Minitab

ExcelStep by step

  1. One proportion (normal): =(p-p0)/SQRT(p0*(1-p0)/n) for z, then =2*(1-NORM.S.DIST(ABS(z),TRUE)) (two-sided) or =1-NORM.S.DIST(z,TRUE) (upper).
  2. One proportion (exact): upper tail =1-BINOM.DIST(x-1,n,p0,TRUE) = 0.0327.
  3. Two proportions: pooled =(x1+x2)/(n1+n2); z =(p1-p2)/SQRT(pp*(1-pp)*(1/n1+1/n2)); p-value as above.
  4. As a table: put the counts in a 2 × 2 block (defective and good for each line), build the expected counts, and use =CHISQ.TEST(actual, expected). It gives 0.0493, the same as the z test.
  5. Fisher’s exact: no built-in function; use =HYPGEOM.DIST() by hand, or Minitab.
  6. Interval: difference ± =NORM.S.INV(0.975)*SQRT(p1*(1-p1)/n1+p2*(1-p2)/n2).

Excel has no proportion-test menu. The formulas above are the standard route.

MinitabStep by step

  1. One proportion: Stat > Basic Statistics > 1 Proportion. Choose Summarized data and enter the number of events and trials (14 and 400), or select the raw column.
  2. Tick Perform hypothesis test and enter the hypothesized proportion (0.02). Click Options to set the confidence level and the alternative; choose the exact or the normal-approximation method.
  3. Two proportions: Stat > Basic Statistics > 2 Proportions. Choose Summarized data and enter the events and trials for both samples.
  4. Under Options for two proportions, choose whether the test uses the pooled estimate or estimates the proportions separately; Minitab also reports Fisher’s exact p-value.
  5. To plan the sample, use Stat > Power and Sample Size > 2 Proportions.

The Assistant (Assistant > Hypothesis Tests > 2-Sample % Defective) adds a power check and tells you whether the sample is large enough for the approximation.

Minitab session window: 1 Proportion (typed excerpt, simplified)
Method

Success: event proportion
Exact method is used for this analysis.

Descriptive Statistics

  N  Event  Sample p        95% CI for p
400   14  0.035000  (0.019264, 0.058027)

Test

Null hypothesis         H₀: p = 0.02
Alternative hypothesis  H₁: p ≠ 0.02

Exact P-Value
       0.0458
Minitab session window: 2 Proportions (typed excerpt, simplified)
Method

p₁: proportion where Sample 1 = Event
p₂: proportion where Sample 2 = Event
Difference: p₁ - p₂

Descriptive Statistics

Sample    N  Event  Sample p
Sample 1  520   38  0.073077
Sample 2  480   21  0.043750

Estimation for Difference

          95% CI for
Difference  Difference
0.029327  (0.000427, 0.058227)

Test

Null hypothesis         H₀: p₁ - p₂ = 0
Alternative hypothesis  H₁: p₁ - p₂ ≠ 0

Method                          Z-Value  P-Value
Normal approximation              1.97    0.0493
Fisher's exact                              0.0595

Reading and Reporting

  1. Report the counts and the rates, not just the percentages.
  2. Report the difference (or the ratio) with an interval.
  3. Say which method: exact or normal approximation, pooled or not.
  4. Say what it means in practice: defects per thousand, cost, or the contract limit.
  5. Check the sample was large enough for the method.
A sentence you can use. The defect rate was 7.3% on Line 1 (38/520) and 4.4% on Line 2 (21/480); the difference of 2.9 percentage points (95% CI 0.0 to 5.8) was borderline (two-proportion z = 1.97, p = 0.049; Fisher’s exact p = 0.060).

Common Mistakes

MistakeWhy it misleadsBetter
Using the normal approximation with few eventsThe p-value and interval are unreliableUse exact tests, or collect more data
Comparing percentages without the countsA small count looks like a trendShow the counts and the interval
Using the sample rate in the standard error for a 1-proportion testThe test assumes the hypothesized rateUse p0 for the test, p̂ for the interval
Testing many defect types and reporting the significant oneFalse alarms accumulatePlan the comparison, or adjust
Treating units from the same batch as independentClustering makes the sample effectively smallerSample across batches; account for clustering
Ignoring small samples when the rate is near zeroZero defects in 20 units does not show a zero rateUse the rule of three: the 95% upper limit is about 3/n
Reading “not significant” as “same rate”Rates need big samplesReport the interval and the power

The rule of three. If you see zero defects in n units, the 95% upper confidence limit on the rate is about 3/n. Zero defects in 100 units means the rate could still be as high as 3%.

Try It Yourself

An audit found 9 errors in 150 invoices before a process change and 3 errors in 140 invoices after it.

  • Calculate the two rates, and test whether the change reduced the error rate. Use a two-proportion z test and Fisher’s exact test.
  • Is the normal approximation reliable here?
Show the answer

Rates: 6.0% before and 2.1% after. Pooled rate = 0.0414; z = 1.65; two-sided p = 0.099. Fisher’s exact p = 0.141.

With only 3 events after the change, the expected count in that cell (about 5.8 under the null) is borderline, so rely on Fisher’s exact test. The evidence of a reduction is suggestive, not conclusive; keep collecting data.

Tests for Proportions: Frequently Asked Questions

When should I use the exact test instead of the normal approximation?

Whenever the expected number of events or non-events under the null hypothesis is less than about 5 to 10. Exact tests are always valid; the normal approximation is only an approximation. Minitab uses the exact method by default for one proportion.

Why is the standard error different for the test and the interval?

The test assumes the null hypothesis is true, so it uses the hypothesized rate (or the pooled rate for two groups). The interval describes the observed rates, so it uses the sample rates.

What is the difference between a chi-square test and a two-proportion test?

For a 2 × 2 table they are the same test: the chi-square statistic equals z squared and the p-values are identical. The chi-square test also extends to larger tables.

How big a sample do I need for a defect rate?

More than you think. Detecting a fall from 4% to 2% needs more than a thousand units per group for 80% power. Use the Sample Size and Power page, and consider measuring a continuous quantity instead.

What if I have zero defects?

A sample with zero defects does not prove a zero rate. The 95% upper limit is about 3 divided by the sample size, so 0 of 300 means the rate could be up to about 1%.

Can I compare rates over time?

Yes, with a p chart, which tests whether the rate changes from subgroup to subgroup, and signals a shift when it happens. See the SPC Control Charts guide.

Sources and Further Reading

  • NIST/SEMATECH, e-Handbook of Statistical Methods, sections on proportions and attribute tests (itl.nist.gov/div898/handbook).
  • Alan Agresti, Categorical Data Analysis, Wiley.
  • Lawrence D. Brown, T. Tony Cai, and Anirban DasGupta, “Interval estimation for a binomial proportion,” Statistical Science, 2001.
  • Douglas C. Montgomery and George C. Runger, Applied Statistics and Probability for Engineers, Wiley.
  • Minitab Support, “Methods and formulas for 1 Proportion and 2 Proportions” (support.minitab.com).

This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.