- Question it answers
- Is the data (or the residual) pattern close enough to normal for my method?
- Null hypothesis
- The data come from a normal distribution
- Key output
- Probability plot, Anderson-Darling or Shapiro-Wilk p-value
- Read first
- The probability plot, then the p-value
- If not normal
- Find the cause, transform, use a nonparametric or distribution-specific method
- Excel
- Histogram, SKEW, KURT, Jarque-Bera, NORM.S.INV plot
- Minitab
- Stat > Basic Statistics > Normality Test
- Level
- Foundation for most tests
The Idea in Plain Language
Many statistical methods assume that the data, or the errors around a fitted model, follow a normal distribution: the symmetric bell curve. When the assumption is badly wrong, p-values, intervals, and capability indices can mislead. A normality test asks whether the data are consistent with a normal distribution. A normal probability plot shows how they differ if they are not.
The two are partners. The plot shows how the data depart from normal, and whether it matters. The test gives a single p-value. When the two disagree, trust the plot and your judgement of what matters for the decision.
When Normality Matters
| Method | Does it need normal data? | Notes |
|---|---|---|
| t-tests and ANOVA | Residuals roughly normal | Fairly robust, especially with balanced groups and moderate sample sizes |
| Regression | Residuals roughly normal | Check the residuals, not the raw Y |
| Capability indices (Cp, Cpk, Pp, Ppk) | Yes: they assume normality to turn the index into a defect rate | Use a non-normal method if the data are skewed |
| Control charts for individuals | Reasonably sensitive to skew | Consider a transformation |
| Confidence interval for a mean | Robust for large samples (central limit theorem) | Small and skewed samples need care |
| Tolerance intervals | Yes: very sensitive | Use a distribution that fits |
| Nonparametric tests, proportions, and counts | No | That is their point |
How the Tests Work
Every normality test has the same null hypothesis: the data come from a normal distribution. A small p-value is evidence against normality. A large p-value means only that the data do not show a clear departure.
| Test | What it measures | Notes |
|---|---|---|
| Anderson-Darling | The distance between the data’s cumulative distribution and the normal, giving extra weight to the tails | The default in Minitab; sensitive to tail departures, which matter for capability |
| Shapiro-Wilk | How well the ordered data match the ordered values expected from a normal distribution | Very good power for small and moderate samples (up to a few thousand) |
| Ryan-Joiner | The correlation between the data and normal scores in the probability plot | Similar to Shapiro-Wilk; available in Minitab |
| Kolmogorov-Smirnov / Lilliefors | The largest gap between the two cumulative curves | Less powerful than the others; rarely the best choice |
| Jarque-Bera | Skewness and kurtosis together | Easy to calculate in Excel; needs a large sample |
| Sample size | What to expect |
|---|---|
| Small (under 20) | Tests have little power: a skewed sample can still pass. Rely on the plot, and on what you know about the process |
| Moderate (20 to 100) | Tests work well; use plot and test together |
| Large (several hundred or more) | Tests flag trivial departures that do not matter. Judge by the plot and by how much the departure affects the method |
Reading a Probability Plot
On a normal probability plot, each value is plotted against the value a perfectly normal sample would have at that rank. If the data are normal, the points fall along a straight line.
| What you see | What it means | What to do |
|---|---|---|
| Points close to the line | Consistent with normal | Proceed |
| A smooth curve, bending up at the right end | Right skew: a long tail of high values | Try a log or square-root transformation; time and cost data are often like this |
| A curve bending down at the right end | Left skew | Look at the cause, such as a ceiling or a limit; transform |
| An S shape | Heavy tails (more extreme values than normal) | Check for outliers; consider a robust or nonparametric method |
| A step or gap in the middle | Two groups mixed together (two machines, shifts, or materials) | Find the groups and analyze them separately: this is a finding, not a nuisance |
| Vertical stripes of points | Rounded or coarse measurements | Improve the gauge resolution, or accept it, noting the limits |
| One point far from the line | An outlier | Investigate the cause before deciding what to do |
Worked Example: Cycle Times
A packaging step has 30 recorded cycle times. The team wants to calculate a confidence interval for the mean and a capability index, both of which lean on normality.
| 26.8 | 45.2 | 17.5 | 55.5 | 70.9 | 31.9 |
| 49.6 | 45.3 | 32.8 | 26.0 | 14.5 | 81.0 |
| 39.1 | 141.2 | 60.5 | 46.0 | 28.8 | 80.2 |
| 31.1 | 87.7 | 32.8 | 43.9 | 18.7 | 129.1 |
| 38.2 | 33.0 | 60.0 | 25.6 | 58.7 | 22.3 |
The mean is 49.1 s, the median is 41.5 s, and the standard deviation is 30.4 s. The mean is well above the median, and the skewness is 1.63. Both point to a long tail on the right.
| Test (raw cycle times) | Statistic | p-value | Reading |
|---|---|---|---|
| Anderson-Darling | A² = 1.427 | < 0.005 | Reject normality |
| Shapiro-Wilk | W = 0.843 | < 0.005 | Reject normality |
| Jarque-Bera | JB = 17.64 | < 0.005 | Reject normality |
The plot and all three tests agree: the cycle times are not normal. A mean and a standard deviation are poor summaries here, and a capability index based on them would be unreliable.
Try a transformation
Taking the natural logarithm of each time pulls in the long tail. The Box-Cox method, which searches for the best power transformation, gives λ = -0.19 with a 95% interval of -0.79 to 0.39. The interval includes 0, and λ = 0 means a log transformation, which is also easy to explain.
| Test (log of cycle times) | Statistic | p-value | Reading |
|---|---|---|---|
| Anderson-Darling | A² = 0.171 | 0.924 | No evidence against normality |
| Shapiro-Wilk | W = 0.984 | 0.919 | No evidence against normality |
What to Do When the Data Are Not Normal
| Option | When to use it | Notes |
|---|---|---|
| First, ask why | Always | Skew may come from a natural limit (times cannot be negative), a mixture of two sources, an outlier, or a process that is out of control. Understand the cause before fixing the shape |
| Transform the data | Skewed data; want to keep using normal methods | Log, square root, or Box-Cox. Analyze on the transformed scale and back-transform the results |
| Use a nonparametric test | Group comparisons with small, non-normal samples | See Nonparametric Tests |
| Fit another distribution | Capability or reliability of non-normal data | Weibull, lognormal, or others; Minitab’s Individual Distribution Identification compares them |
| Use the central limit theorem | Inference about means with a large sample | Works for averages, not for individuals; skewed data need more observations |
| Split the data | A step or two humps in the plot | Analyze each source separately |
| Use a robust method | Outliers that are real | Medians, trimmed means, or rank-based methods |
Run It in Excel and Minitab
ExcelStep by step
- Put the data in A2:A31 with a heading in A1.
- Draw a histogram with , and look at the shape.
- and give skewness and excess kurtosis. Values near 0 suggest symmetry. Here skewness is 1.63.
- Jarque-Bera test: then for the p-value.
- Probability plot: sort the data in column B (), compute the normal score in column C with , and draw a scatter chart of B against C.
- Transform: in a new column use or , then repeat the plot.
Excel has no built-in Anderson-Darling or Shapiro-Wilk test and no Box-Cox tool. For formal tests use Minitab, or a statistics add-in.
MinitabStep by step
- Choose . Set Variable to Cycle Time and choose the test (Anderson-Darling is the default).
- Read the probability plot with the test statistic and p-value in the box. Points inside the confidence bands on either side of the line are consistent with normal.
- For a plot with other distributions, use , or compare many at once with .
- To transform, use (Box-Cox) or , and store the transformed data.
- For capability of non-normal data, use .
- To see skewness and kurtosis, use .
Minitab reports the Anderson-Darling p-value as < 0.005 when it is very small. The Assistant () also checks normality for you.
What the output looks like
=SKEW(A2:A31) 1.628 =KURT(A2:A31) 2.771 (excess kurtosis) Jarque-Bera = n/6*(skew^2 + kurt^2/4) 22.84 =CHISQ.DIST.RT(JB, 2) 0.0000
Normality Test: Cycle Time Anderson-Darling Normality Test Mean 49.130 StDev 30.368 N 30 AD 1.427 P-Value <0.005
Normality Test: ln(Cycle Time) Anderson-Darling Normality Test Mean 3.738 StDev 0.558 N 30 AD 0.171 P-Value 0.924
Reading and Reporting
- Look at the plot first. Say what shape you see.
- Use the p-value as a second opinion. Say which test you used.
- Say what you did about it: transformed, used a nonparametric method, or checked that the method is robust.
- Back-transform results. Report geometric means or medians for log-transformed data, with a note on the scale.
Common Mistakes
| Mistake | Why it misleads | Better |
|---|---|---|
| Treating “p > 0.05” as proof of normality | A small sample may not detect a clear departure | Look at the plot; consider the sample size |
| Rejecting a method because a huge sample fails the test | Tests flag trivial departures with large n | Judge the size of the departure on the plot |
| Testing the raw data when the method assumes normal residuals | Regression and ANOVA need normal errors, not a normal Y | Test the residuals |
| Testing groups pooled together | Different group means make the pooled data look non-normal | Test the residuals, or test within each group |
| Deleting outliers to pass the test | It distorts the analysis | Find the cause; report with and without |
| Transforming without explaining | Results on a log scale are hard to interpret | Back-transform and say what the transformation was |
| Ignoring a step in the plot | Two populations mixed together are the real story | Find the sources |
Try It Yourself
Three data sets were tested for normality with the Anderson-Darling test. Sample sizes and p-values are shown.
| Data set | n | p-value |
|---|---|---|
| A | 18 | 0.42 |
| B | 45 | 0.003 |
| C | 300 | 0.07 |
- Which data sets do you consider non-normal at α = 0.05?
- What else would you look at before deciding what to do with data set C?
Show the answer
Data set B (p = 0.003) is non-normal. Data set A (p = 0.42) shows no evidence against normality, though with only 18 values the test has little power, so look at the plot. Data set C (p = 0.07) is not significant at 0.05, but is borderline.
For C, with n = 300 the test is powerful, so a p-value near 0.07 means any departure is probably small. Look at the probability plot and ask whether the departure is large enough to affect the method. For means with n = 300, the central limit theorem makes most methods robust. For capability or tolerance of individual items, a heavy tail would still matter.
Normality Tests and Probability Plots: Frequently Asked Questions
Which normality test should I use?
Anderson-Darling and Shapiro-Wilk are both good choices. Anderson-Darling is the Minitab default and is sensitive to the tails. Shapiro-Wilk has excellent power for small and moderate samples. Whatever you use, look at the probability plot as well.
My data fail the normality test. Can I still use a t-test?
Often yes, if the sample is not tiny and the departure is mild. The t-test is fairly robust, especially with balanced groups. For small samples and clear skew, transform the data or use a nonparametric test.
Do I test the data or the residuals?
For regression and ANOVA, test the residuals. The assumption is that the errors are normal, and raw data from several groups or X values will look non-normal even when the errors are fine.
What does the p-value of a normality test mean?
It is the chance of seeing a departure from normality at least this large if the data were truly normal. A small value is evidence against normality. A large value does not prove normality: it only means there is no clear evidence against it.
What is a Box-Cox transformation?
A family of power transformations, y to the power lambda, with lambda chosen to make the data as normal as possible. Lambda of 0 is the log, 0.5 the square root, and 1 means no change. Use a convenient value inside the confidence interval for lambda.
Can I make non-normal capability numbers valid?
Yes, by transforming the data and the specification limits, or by fitting a distribution that matches the data and calculating capability from it. See the Process Capability guide.
Sources and Further Reading
- NIST/SEMATECH, e-Handbook of Statistical Methods, sections on normal probability plots, Anderson-Darling, and Shapiro-Wilk (itl.nist.gov/div898/handbook).
- Ralph B. D’Agostino and Michael A. Stephens (eds.), Goodness-of-Fit Techniques, Marcel Dekker, 1986.
- Samuel S. Shapiro and Martin B. Wilk, “An analysis of variance test for normality,” Biometrika, 1965.
- George E. P. Box and David R. Cox, “An analysis of transformations,” Journal of the Royal Statistical Society B, 1964.
- Minitab Support, “Methods and formulas for Normality Test” and “Box-Cox Transformation” (support.minitab.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.