Question it answers
How likely is a value, a defective, or a number of defects?
Data needed
A process mean and spread, or a defect rate
Key output
A probability or a parts-per-million figure
Pick by data
Measurements: normal; defectives: binomial; defects: Poisson
Assumptions
Independence; a steady rate; a fitting shape
Excel
NORM.DIST, NORM.INV, BINOM.DIST, POISSON.DIST
Minitab
Calc > Probability Distributions; Distribution Plot
Why it matters
Capability and sample plans rest on the distribution

The Idea in Plain Language

A distribution describes how likely each possible value is. Once you know the distribution behind a process, you can answer questions that data alone cannot: how often will a part be out of spec, what is the chance of two defects in a sample, how many units will fail?

Three distributions cover most Six Sigma work. The normal describes continuous measurements that cluster around an average. The binomial counts defectives in a sample (each item is good or bad). The Poisson counts defects on a unit or events in an interval.

DistributionWhat it modelsParametersExample
NormalContinuous measurements, symmetricMean μ, standard deviation σFill weight, shaft diameter
BinomialNumber of defectives in n itemsn, defect probability pFailures in a sample of 20
PoissonNumber of defects or events in an area or timeAverage rate λScratches per panel, calls per hour
Why it matters. Capability, control charts, and sample sizes all rest on a distribution. If you use the wrong one, the predicted defect rate can be off by a factor of ten or more.

The Normal Distribution

The bell curve is fully described by its mean and standard deviation. The z score, z = (x − μ) / σ, says how many standard deviations a value sits from the mean, and the standard normal table converts z to a probability.

WithinShare of values
± 1σ68.27%
± 2σ95.45%
± 3σ99.73%
± 6σ99.9999998%

Worked example. A filler runs at a mean of 500 g with a standard deviation of 2 g. The specification is 495 to 504 g.

  1. z for the lower limit: (495 − 500) / 2 = -2.50. Area below = 0.0062.
  2. z for the upper limit: (504 − 500) / 2 = 2.00. Area above = 0.0228.
  3. Out of spec = 0.0062 + 0.0228 = 0.0290, or about 28,960 parts per million.
  4. The value below which 95% of fills fall: 500 + 1.645 × 2 = 503.29 g.
492 496 500 504 508 LSL 495 USL 504 Mean 500 Fill weight (g)
The shaded tails are the parts outside the specification. The upper tail is larger because the upper limit is closer to the mean in standard deviations.
Conclusion. About 97.10% of fills will be in spec, and 28,960 per million will not. The mean is not centered in the spec window (499.5 is the center), which is why the tail above is bigger. Moving the mean to 499.5 would balance it.

The Binomial Distribution

Use the binomial when you inspect n items, each independently good or bad with the same defect probability p, and you count the defectives. The probability of exactly k defectives is C(n, k) pk (1 − p)n−k. The mean is np and the standard deviation is √(np(1 − p)).

Worked example. A process makes 5% defective parts. You inspect 20 parts.

  1. None defective: P(0) = (1 − 0.05)20 = 0.3585.
  2. Two or more defectives: 1 − P(0) − P(1) = 1 − 0.3585 − 0.3774 = 0.2642.
  3. Expected defectives = 20 × 0.05 = 1, with standard deviation 0.97.
0.00 0.20 0.40 0 1 2 3 4 5 6 7 Defectives in a sample of 20 Probability
Even at a 5% defect rate, a sample of 20 contains no defectives about a third of the time. A clean sample does not prove a good process.
Conclusion. There is a 36% chance that a sample of 20 shows no defectives and a 26% chance of two or more. If an acceptance rule is “accept when 0 or 1 defectives,” it passes 74% of lots from this 5% process.

The Poisson Distribution

Use the Poisson when you count defects on a unit, or events in a period, and they occur independently at a steady average rate λ. The probability of k is e−λλk/k!. The mean and the variance are both λ.

Worked example. Painted panels average 2.4 blemishes each.

  1. A perfect panel: P(0) = e−2.4 = 0.0907.
  2. One or fewer: P(0) + P(1) = 0.0907 + 0.2177 = 0.3084.
  3. Five or more: 1 − P(0 to 4) = 0.0959.
  4. Standard deviation = √2.4 = 1.55.
0.00 0.10 0.20 0.30 0 1 2 3 4 5 6 7 8 9 10 Defects on one unit Probability
Defect counts are lopsided when the average is small. The Poisson approaches a bell shape as λ grows.
Conclusion. Only 9.1% of panels are perfect, and 9.6% have five or more blemishes. This is how defects per unit relates to first-pass yield: yield = e−DPU, here 9.1%.

Other Distributions You Will Meet

DistributionShape and useWhere in the dojo
LognormalRight-skewed; the log of the data is normal. Repair times, particle sizesNormality tests, capability on skewed data
ExponentialTime between random events; constant failure rateReliability
WeibullFlexible life distribution; shape parameter sets the failure patternReliability and Weibull
UniformEvery value equally likely; rounding errorMeasurement
Student’s tLike the normal with heavier tails; for means with small samplest-Tests
Chi-squareSkewed, from sums of squares; variances and countsChi-Square Tests
FRatio of two variances; ANOVA and regressionOne-Way ANOVA

To find out which distribution fits your data, plot it and use a probability plot. The Normality Tests page shows how.

Normal approximations. The binomial is close to normal when np and n(1 − p) are both at least 10 (some texts accept 5). The Poisson is close to normal when λ is large, about 20 or more. For small counts, use the exact distribution.

Run It in Excel and Minitab

ExcelStep by step

  1. Normal: =NORM.DIST(x, mean, sd, TRUE) gives the area below x. =NORM.DIST(495, 500, 2, TRUE) = 0.0062. Upper area: =1-NORM.DIST(x, mean, sd, TRUE).
  2. Percentile: =NORM.INV(0.95, 500, 2) = 503.29. For z scores use =NORM.S.DIST(z, TRUE) and =NORM.S.INV(p).
  3. Binomial: =BINOM.DIST(k, n, p, FALSE) is exactly k; with TRUE it is k or fewer. =BINOM.DIST(0, 20, 0.05, FALSE) = 0.3585.
  4. Poisson: =POISSON.DIST(k, lambda, FALSE) is exactly k; with TRUE it is k or fewer. =1-POISSON.DIST(4, 2.4, TRUE) = 0.0959.
  5. To see which family fits your data, use a histogram and a probability plot (Normality Tests page).

MinitabStep by step

  1. Calc > Probability Distributions > Normal. Choose Cumulative probability (area below x) or Inverse cumulative probability (the x for a given area). Enter the mean, standard deviation, and the input constant.
  2. Calc > Probability Distributions > Binomial and Calc > Probability Distributions > Poisson. Choose Probability for exactly k or Cumulative probability for k or fewer.
  3. Use Graph > Probability Distribution Plot to draw a distribution and shade an area (View Probability).
  4. Stat > Quality Tools > Individual Distribution Identification tests your data against many distributions at once and gives a probability plot for each.
Minitab session window: Probability Distributions (typed excerpt, simplified)
Cumulative Distribution Function

Normal with mean = 500 and standard deviation = 2

  x  P( X <= x )
495      0.006210

Probability Density Function

Binomial with n = 20 and p = 0.05

 x  P( X = x )
 0   0.358486

Cumulative Distribution Function

Poisson with mean = 2.4

 x  P( X <= x )
 4   0.904131

Reading and Reporting

  1. Say which distribution you assumed and why: measurements (normal), defectives (binomial), defects (Poisson).
  2. Check the assumption with a plot or a goodness-of-fit test before quoting a tail probability.
  3. Quote tail probabilities as parts per million when they are small, and note they depend on the model fitting in the tails.
  4. Remember the inputs are estimates. A defect rate predicted from a sample has uncertainty; see confidence intervals.

Common Mistakes

MistakeWhy it misleadsBetter
Assuming normality without lookingSkewed data give wildly wrong tail areasPlot, and run a probability plot
Using the binomial when defects are not independentClusters of failures break the modelCheck for batches and common causes
Using the binomial for defects per unitA unit can have several defectsUse Poisson for counts of defects
Using Poisson when the rate changes over timeThe average is not steadyChart the rate first (u chart)
Trusting the far tail of a fitted normalReal processes have heavier tailsTreat ppm figures as rough, and verify with data
Confusing “P(X ≤ k)” and “P(X = k)”Cumulative versus exactCheck the Excel TRUE/FALSE argument

Try It Yourself

A machine produces bolts with a mean length of 50.0 mm and a standard deviation of 0.15 mm. The upper limit is 50.3 mm. Separately, a line averages 0.8 defects per board.

  • What fraction of bolts exceed the limit?
  • What is the chance a board has no defects, and the chance it has two or more?
Show the answer

z = (50.3 − 50.0) / 0.15 = 2.00. The fraction above the limit is 0.0228, or 22,750 ppm.

With λ = 0.8, P(0) = e−0.8 = 0.4493. P(2 or more) = 1 − P(0) − P(1) = 0.1912.

Distributions: Normal, Binomial, Poisson: Frequently Asked Questions

How do I know if my data are normal?

Plot a histogram and a normal probability plot, and run a normality test such as Anderson-Darling. If the points follow the line and the p-value is above 0.05, normal is a reasonable model. See the Normality Tests page.

When do I use binomial and when Poisson?

Use the binomial when each item is classified good or bad and you count defectives out of n items. Use Poisson when you count defects on a unit or events in a time or space and there is no fixed upper limit.

What is a z score?

It is the number of standard deviations a value is from the mean: z = (x − mean) / standard deviation. It lets you use one table for every normal distribution.

Why does Six Sigma talk about a 1.5 sigma shift?

It is a convention that allows for long-term drift in the process mean. A 6 sigma process then has about 3.4 defects per million, rather than 0.002 per million with no shift. It is a convention, not a law of nature.

What if my data are not normal?

Check for mixed sources or outliers. Then consider a transformation (Box-Cox), a better-fitting distribution (lognormal, Weibull), or methods that do not assume normality.

Sources and Further Reading

  • NIST/SEMATECH, e-Handbook of Statistical Methods, Probability Distributions (itl.nist.gov/div898/handbook).
  • Douglas C. Montgomery and George C. Runger, Applied Statistics and Probability for Engineers, Wiley.
  • Minitab Support, “Probability distributions” and “Individual Distribution Identification” (support.minitab.com).
  • Microsoft Support, “NORM.DIST, BINOM.DIST, and POISSON.DIST functions” (support.microsoft.com).

This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.