- Question it answers
- How likely is a value, a defective, or a number of defects?
- Data needed
- A process mean and spread, or a defect rate
- Key output
- A probability or a parts-per-million figure
- Pick by data
- Measurements: normal; defectives: binomial; defects: Poisson
- Assumptions
- Independence; a steady rate; a fitting shape
- Excel
- NORM.DIST, NORM.INV, BINOM.DIST, POISSON.DIST
- Minitab
- Calc > Probability Distributions; Distribution Plot
- Why it matters
- Capability and sample plans rest on the distribution
The Idea in Plain Language
A distribution describes how likely each possible value is. Once you know the distribution behind a process, you can answer questions that data alone cannot: how often will a part be out of spec, what is the chance of two defects in a sample, how many units will fail?
Three distributions cover most Six Sigma work. The normal describes continuous measurements that cluster around an average. The binomial counts defectives in a sample (each item is good or bad). The Poisson counts defects on a unit or events in an interval.
| Distribution | What it models | Parameters | Example |
|---|---|---|---|
| Normal | Continuous measurements, symmetric | Mean μ, standard deviation σ | Fill weight, shaft diameter |
| Binomial | Number of defectives in n items | n, defect probability p | Failures in a sample of 20 |
| Poisson | Number of defects or events in an area or time | Average rate λ | Scratches per panel, calls per hour |
The Normal Distribution
The bell curve is fully described by its mean and standard deviation. The z score, z = (x − μ) / σ, says how many standard deviations a value sits from the mean, and the standard normal table converts z to a probability.
| Within | Share of values |
|---|---|
| ± 1σ | 68.27% |
| ± 2σ | 95.45% |
| ± 3σ | 99.73% |
| ± 6σ | 99.9999998% |
Worked example. A filler runs at a mean of 500 g with a standard deviation of 2 g. The specification is 495 to 504 g.
- z for the lower limit: (495 − 500) / 2 = -2.50. Area below = 0.0062.
- z for the upper limit: (504 − 500) / 2 = 2.00. Area above = 0.0228.
- Out of spec = 0.0062 + 0.0228 = 0.0290, or about 28,960 parts per million.
- The value below which 95% of fills fall: 500 + 1.645 × 2 = 503.29 g.
The Binomial Distribution
Use the binomial when you inspect n items, each independently good or bad with the same defect probability p, and you count the defectives. The probability of exactly k defectives is C(n, k) pk (1 − p)n−k. The mean is np and the standard deviation is √(np(1 − p)).
Worked example. A process makes 5% defective parts. You inspect 20 parts.
- None defective: P(0) = (1 − 0.05)20 = 0.3585.
- Two or more defectives: 1 − P(0) − P(1) = 1 − 0.3585 − 0.3774 = 0.2642.
- Expected defectives = 20 × 0.05 = 1, with standard deviation 0.97.
The Poisson Distribution
Use the Poisson when you count defects on a unit, or events in a period, and they occur independently at a steady average rate λ. The probability of k is e−λλk/k!. The mean and the variance are both λ.
Worked example. Painted panels average 2.4 blemishes each.
- A perfect panel: P(0) = e−2.4 = 0.0907.
- One or fewer: P(0) + P(1) = 0.0907 + 0.2177 = 0.3084.
- Five or more: 1 − P(0 to 4) = 0.0959.
- Standard deviation = √2.4 = 1.55.
Other Distributions You Will Meet
| Distribution | Shape and use | Where in the dojo |
|---|---|---|
| Lognormal | Right-skewed; the log of the data is normal. Repair times, particle sizes | Normality tests, capability on skewed data |
| Exponential | Time between random events; constant failure rate | Reliability |
| Weibull | Flexible life distribution; shape parameter sets the failure pattern | Reliability and Weibull |
| Uniform | Every value equally likely; rounding error | Measurement |
| Student’s t | Like the normal with heavier tails; for means with small samples | t-Tests |
| Chi-square | Skewed, from sums of squares; variances and counts | Chi-Square Tests |
| F | Ratio of two variances; ANOVA and regression | One-Way ANOVA |
To find out which distribution fits your data, plot it and use a probability plot. The Normality Tests page shows how.
Run It in Excel and Minitab
ExcelStep by step
- Normal: gives the area below x. = 0.0062. Upper area: .
- Percentile: = 503.29. For z scores use and .
- Binomial: is exactly k; with TRUE it is k or fewer. = 0.3585.
- Poisson: is exactly k; with TRUE it is k or fewer. = 0.0959.
- To see which family fits your data, use a histogram and a probability plot (Normality Tests page).
MinitabStep by step
- . Choose Cumulative probability (area below x) or Inverse cumulative probability (the x for a given area). Enter the mean, standard deviation, and the input constant.
- and . Choose Probability for exactly k or Cumulative probability for k or fewer.
- Use Graph > Probability Distribution Plot to draw a distribution and shade an area (View Probability).
- tests your data against many distributions at once and gives a probability plot for each.
Cumulative Distribution Function Normal with mean = 500 and standard deviation = 2 x P( X <= x ) 495 0.006210 Probability Density Function Binomial with n = 20 and p = 0.05 x P( X = x ) 0 0.358486 Cumulative Distribution Function Poisson with mean = 2.4 x P( X <= x ) 4 0.904131
Reading and Reporting
- Say which distribution you assumed and why: measurements (normal), defectives (binomial), defects (Poisson).
- Check the assumption with a plot or a goodness-of-fit test before quoting a tail probability.
- Quote tail probabilities as parts per million when they are small, and note they depend on the model fitting in the tails.
- Remember the inputs are estimates. A defect rate predicted from a sample has uncertainty; see confidence intervals.
Common Mistakes
| Mistake | Why it misleads | Better |
|---|---|---|
| Assuming normality without looking | Skewed data give wildly wrong tail areas | Plot, and run a probability plot |
| Using the binomial when defects are not independent | Clusters of failures break the model | Check for batches and common causes |
| Using the binomial for defects per unit | A unit can have several defects | Use Poisson for counts of defects |
| Using Poisson when the rate changes over time | The average is not steady | Chart the rate first (u chart) |
| Trusting the far tail of a fitted normal | Real processes have heavier tails | Treat ppm figures as rough, and verify with data |
| Confusing “P(X ≤ k)” and “P(X = k)” | Cumulative versus exact | Check the Excel TRUE/FALSE argument |
Try It Yourself
A machine produces bolts with a mean length of 50.0 mm and a standard deviation of 0.15 mm. The upper limit is 50.3 mm. Separately, a line averages 0.8 defects per board.
- What fraction of bolts exceed the limit?
- What is the chance a board has no defects, and the chance it has two or more?
Show the answer
z = (50.3 − 50.0) / 0.15 = 2.00. The fraction above the limit is 0.0228, or 22,750 ppm.
With λ = 0.8, P(0) = e−0.8 = 0.4493. P(2 or more) = 1 − P(0) − P(1) = 0.1912.
Distributions: Normal, Binomial, Poisson: Frequently Asked Questions
How do I know if my data are normal?
Plot a histogram and a normal probability plot, and run a normality test such as Anderson-Darling. If the points follow the line and the p-value is above 0.05, normal is a reasonable model. See the Normality Tests page.
When do I use binomial and when Poisson?
Use the binomial when each item is classified good or bad and you count defectives out of n items. Use Poisson when you count defects on a unit or events in a time or space and there is no fixed upper limit.
What is a z score?
It is the number of standard deviations a value is from the mean: z = (x − mean) / standard deviation. It lets you use one table for every normal distribution.
Why does Six Sigma talk about a 1.5 sigma shift?
It is a convention that allows for long-term drift in the process mean. A 6 sigma process then has about 3.4 defects per million, rather than 0.002 per million with no shift. It is a convention, not a law of nature.
What if my data are not normal?
Check for mixed sources or outliers. Then consider a transformation (Box-Cox), a better-fitting distribution (lognormal, Weibull), or methods that do not assume normality.
Sources and Further Reading
- NIST/SEMATECH, e-Handbook of Statistical Methods, Probability Distributions (itl.nist.gov/div898/handbook).
- Douglas C. Montgomery and George C. Runger, Applied Statistics and Probability for Engineers, Wiley.
- Minitab Support, “Probability distributions” and “Individual Distribution Identification” (support.minitab.com).
- Microsoft Support, “NORM.DIST, BINOM.DIST, and POISSON.DIST functions” (support.microsoft.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.