Question it answers
Is the data (or the residual) pattern close enough to normal for my method?
Null hypothesis
The data come from a normal distribution
Key output
Probability plot, Anderson-Darling or Shapiro-Wilk p-value
Read first
The probability plot, then the p-value
If not normal
Find the cause, transform, use a nonparametric or distribution-specific method
Excel
Histogram, SKEW, KURT, Jarque-Bera, NORM.S.INV plot
Minitab
Stat > Basic Statistics > Normality Test
Level
Foundation for most tests

The Idea in Plain Language

Many statistical methods assume that the data, or the errors around a fitted model, follow a normal distribution: the symmetric bell curve. When the assumption is badly wrong, p-values, intervals, and capability indices can mislead. A normality test asks whether the data are consistent with a normal distribution. A normal probability plot shows how they differ if they are not.

The two are partners. The plot shows how the data depart from normal, and whether it matters. The test gives a single p-value. When the two disagree, trust the plot and your judgement of what matters for the decision.

Normal Right-skewed Heavy tails Two groups
Normal probability plots. Points on the dashed line mean normal data. A curve means skew, an S-bend means heavy tails, and a step means a mixture of two populations.

When Normality Matters

MethodDoes it need normal data?Notes
t-tests and ANOVAResiduals roughly normalFairly robust, especially with balanced groups and moderate sample sizes
RegressionResiduals roughly normalCheck the residuals, not the raw Y
Capability indices (Cp, Cpk, Pp, Ppk)Yes: they assume normality to turn the index into a defect rateUse a non-normal method if the data are skewed
Control charts for individualsReasonably sensitive to skewConsider a transformation
Confidence interval for a meanRobust for large samples (central limit theorem)Small and skewed samples need care
Tolerance intervalsYes: very sensitiveUse a distribution that fits
Nonparametric tests, proportions, and countsNoThat is their point
Averages are friendlier than individuals. Even if single values are skewed, averages of several values tend toward a bell shape, because of the central limit theorem. That is why many methods are robust for means with a reasonable sample, but capability and tolerance statements about individual items are not.

How the Tests Work

Every normality test has the same null hypothesis: the data come from a normal distribution. A small p-value is evidence against normality. A large p-value means only that the data do not show a clear departure.

TestWhat it measuresNotes
Anderson-DarlingThe distance between the data’s cumulative distribution and the normal, giving extra weight to the tailsThe default in Minitab; sensitive to tail departures, which matter for capability
Shapiro-WilkHow well the ordered data match the ordered values expected from a normal distributionVery good power for small and moderate samples (up to a few thousand)
Ryan-JoinerThe correlation between the data and normal scores in the probability plotSimilar to Shapiro-Wilk; available in Minitab
Kolmogorov-Smirnov / LillieforsThe largest gap between the two cumulative curvesLess powerful than the others; rarely the best choice
Jarque-BeraSkewness and kurtosis togetherEasy to calculate in Excel; needs a large sample
Sample sizeWhat to expect
Small (under 20)Tests have little power: a skewed sample can still pass. Rely on the plot, and on what you know about the process
Moderate (20 to 100)Tests work well; use plot and test together
Large (several hundred or more)Tests flag trivial departures that do not matter. Judge by the plot and by how much the departure affects the method

Reading a Probability Plot

On a normal probability plot, each value is plotted against the value a perfectly normal sample would have at that rank. If the data are normal, the points fall along a straight line.

What you seeWhat it meansWhat to do
Points close to the lineConsistent with normalProceed
A smooth curve, bending up at the right endRight skew: a long tail of high valuesTry a log or square-root transformation; time and cost data are often like this
A curve bending down at the right endLeft skewLook at the cause, such as a ceiling or a limit; transform
An S shapeHeavy tails (more extreme values than normal)Check for outliers; consider a robust or nonparametric method
A step or gap in the middleTwo groups mixed together (two machines, shifts, or materials)Find the groups and analyze them separately: this is a finding, not a nuisance
Vertical stripes of pointsRounded or coarse measurementsImprove the gauge resolution, or accept it, noting the limits
One point far from the lineAn outlierInvestigate the cause before deciding what to do

Worked Example: Cycle Times

A packaging step has 30 recorded cycle times. The team wants to calculate a confidence interval for the mean and a capability index, both of which lean on normality.

Cycle times in seconds (30 cycles, in time order across rows)
26.845.217.555.570.931.9
49.645.332.826.014.581.0
39.1141.260.546.028.880.2
31.187.732.843.918.7129.1
38.233.060.025.658.722.3

The mean is 49.1 s, the median is 41.5 s, and the standard deviation is 30.4 s. The mean is well above the median, and the skewness is 1.63. Both point to a long tail on the right.

0 2 4 6 8 10 12 14 10 30 50 70 90 110 130 150 Cycle time (s) Cycle time: skewed right
Most cycles are short, with a few long ones.
-2 -1 0 1 2 Normal score (expected z) Normal probability plot of residuals Standardized cycle time
The points curve upward at the right end.
Test (raw cycle times)Statisticp-valueReading
Anderson-DarlingA² = 1.427< 0.005Reject normality
Shapiro-WilkW = 0.843< 0.005Reject normality
Jarque-BeraJB = 17.64< 0.005Reject normality

The plot and all three tests agree: the cycle times are not normal. A mean and a standard deviation are poor summaries here, and a capability index based on them would be unreliable.

Try a transformation

Taking the natural logarithm of each time pulls in the long tail. The Box-Cox method, which searches for the best power transformation, gives λ = -0.19 with a 95% interval of -0.79 to 0.39. The interval includes 0, and λ = 0 means a log transformation, which is also easy to explain.

0 2 4 6 8 10 12 2.5 3 3.5 4 4.5 5 ln(cycle time) After the log transformation: symmetric
The logged times are symmetric.
-2 -1 0 1 2 Normal score (expected z) Normal probability plot of residuals Standardized ln(cycle time)
The points now follow the line.
Test (log of cycle times)Statisticp-valueReading
Anderson-DarlingA² = 0.1710.924No evidence against normality
Shapiro-WilkW = 0.9840.919No evidence against normality
Conclusion. The raw cycle times are strongly right-skewed (Anderson-Darling p < 0.005). After a log transformation they are consistent with a normal distribution (p = 0.92). Work on the log scale for methods that need normality, and transform back to report results: the geometric mean is 42.0 s, a better center for skewed time data than the arithmetic mean of 49.1 s.

What to Do When the Data Are Not Normal

OptionWhen to use itNotes
First, ask whyAlwaysSkew may come from a natural limit (times cannot be negative), a mixture of two sources, an outlier, or a process that is out of control. Understand the cause before fixing the shape
Transform the dataSkewed data; want to keep using normal methodsLog, square root, or Box-Cox. Analyze on the transformed scale and back-transform the results
Use a nonparametric testGroup comparisons with small, non-normal samplesSee Nonparametric Tests
Fit another distributionCapability or reliability of non-normal dataWeibull, lognormal, or others; Minitab’s Individual Distribution Identification compares them
Use the central limit theoremInference about means with a large sampleWorks for averages, not for individuals; skewed data need more observations
Split the dataA step or two humps in the plotAnalyze each source separately
Use a robust methodOutliers that are realMedians, trimmed means, or rank-based methods
Do not delete data to make it normal. Removing real values to improve a p-value is a form of data fabrication. If a value is a recording error, correct it and say so. If it is real, it belongs in the analysis.

Run It in Excel and Minitab

ExcelStep by step

  1. Put the data in A2:A31 with a heading in A1.
  2. Draw a histogram with Insert > Chart > Histogram, and look at the shape.
  3. =SKEW(A2:A31) and =KURT(A2:A31) give skewness and excess kurtosis. Values near 0 suggest symmetry. Here skewness is 1.63.
  4. Jarque-Bera test: =n/6*(SKEW^2+KURT^2/4) then =CHISQ.DIST.RT(JB,2) for the p-value.
  5. Probability plot: sort the data in column B (=SMALL($A$2:$A$31,ROW()-1)), compute the normal score in column C with =NORM.S.INV((ROW()-1-0.5)/30), and draw a scatter chart of B against C.
  6. Transform: in a new column use =LN(A2) or =SQRT(A2), then repeat the plot.

Excel has no built-in Anderson-Darling or Shapiro-Wilk test and no Box-Cox tool. For formal tests use Minitab, or a statistics add-in.

MinitabStep by step

  1. Choose Stat > Basic Statistics > Normality Test. Set Variable to Cycle Time and choose the test (Anderson-Darling is the default).
  2. Read the probability plot with the test statistic and p-value in the box. Points inside the confidence bands on either side of the line are consistent with normal.
  3. For a plot with other distributions, use Graph > Probability Plot, or compare many at once with Stat > Quality Tools > Individual Distribution Identification.
  4. To transform, use Stat > Control Charts > Box-Cox Transformation (Box-Cox) or Stat > Quality Tools > Johnson Transformation, and store the transformed data.
  5. For capability of non-normal data, use Stat > Quality Tools > Capability Analysis > Nonnormal.
  6. To see skewness and kurtosis, use Stat > Basic Statistics > Display Descriptive Statistics.

Minitab reports the Anderson-Darling p-value as < 0.005 when it is very small. The Assistant (Assistant > Hypothesis Tests) also checks normality for you.

What the output looks like

Excel formula results
=SKEW(A2:A31)                     1.628
=KURT(A2:A31)                     2.771   (excess kurtosis)
Jarque-Bera = n/6*(skew^2 + kurt^2/4)   22.84
=CHISQ.DIST.RT(JB, 2)               0.0000
Minitab: probability plot box, raw cycle times (typed excerpt)
Normality Test: Cycle Time

Anderson-Darling Normality Test

Mean           49.130
StDev          30.368
N                  30
AD              1.427
P-Value        <0.005
Minitab: probability plot box, after Box-Cox with lambda = 0 (log) (typed excerpt)
Normality Test: ln(Cycle Time)

Anderson-Darling Normality Test

Mean            3.738
StDev           0.558
N                  30
AD              0.171
P-Value         0.924

Reading and Reporting

  1. Look at the plot first. Say what shape you see.
  2. Use the p-value as a second opinion. Say which test you used.
  3. Say what you did about it: transformed, used a nonparametric method, or checked that the method is robust.
  4. Back-transform results. Report geometric means or medians for log-transformed data, with a note on the scale.
A sentence you can use. The cycle times were strongly right-skewed (Anderson-Darling A² = 1.43, p < 0.005, n = 30); after a natural-log transformation they were consistent with a normal distribution (A² = 0.17, p = 0.92), so subsequent analyses used the log scale.

Common Mistakes

MistakeWhy it misleadsBetter
Treating “p > 0.05” as proof of normalityA small sample may not detect a clear departureLook at the plot; consider the sample size
Rejecting a method because a huge sample fails the testTests flag trivial departures with large nJudge the size of the departure on the plot
Testing the raw data when the method assumes normal residualsRegression and ANOVA need normal errors, not a normal YTest the residuals
Testing groups pooled togetherDifferent group means make the pooled data look non-normalTest the residuals, or test within each group
Deleting outliers to pass the testIt distorts the analysisFind the cause; report with and without
Transforming without explainingResults on a log scale are hard to interpretBack-transform and say what the transformation was
Ignoring a step in the plotTwo populations mixed together are the real storyFind the sources

Try It Yourself

Three data sets were tested for normality with the Anderson-Darling test. Sample sizes and p-values are shown.

Data setnp-value
A180.42
B450.003
C3000.07
  • Which data sets do you consider non-normal at α = 0.05?
  • What else would you look at before deciding what to do with data set C?
Show the answer

Data set B (p = 0.003) is non-normal. Data set A (p = 0.42) shows no evidence against normality, though with only 18 values the test has little power, so look at the plot. Data set C (p = 0.07) is not significant at 0.05, but is borderline.

For C, with n = 300 the test is powerful, so a p-value near 0.07 means any departure is probably small. Look at the probability plot and ask whether the departure is large enough to affect the method. For means with n = 300, the central limit theorem makes most methods robust. For capability or tolerance of individual items, a heavy tail would still matter.

Normality Tests and Probability Plots: Frequently Asked Questions

Which normality test should I use?

Anderson-Darling and Shapiro-Wilk are both good choices. Anderson-Darling is the Minitab default and is sensitive to the tails. Shapiro-Wilk has excellent power for small and moderate samples. Whatever you use, look at the probability plot as well.

My data fail the normality test. Can I still use a t-test?

Often yes, if the sample is not tiny and the departure is mild. The t-test is fairly robust, especially with balanced groups. For small samples and clear skew, transform the data or use a nonparametric test.

Do I test the data or the residuals?

For regression and ANOVA, test the residuals. The assumption is that the errors are normal, and raw data from several groups or X values will look non-normal even when the errors are fine.

What does the p-value of a normality test mean?

It is the chance of seeing a departure from normality at least this large if the data were truly normal. A small value is evidence against normality. A large value does not prove normality: it only means there is no clear evidence against it.

What is a Box-Cox transformation?

A family of power transformations, y to the power lambda, with lambda chosen to make the data as normal as possible. Lambda of 0 is the log, 0.5 the square root, and 1 means no change. Use a convenient value inside the confidence interval for lambda.

Can I make non-normal capability numbers valid?

Yes, by transforming the data and the specification limits, or by fitting a distribution that matches the data and calculating capability from it. See the Process Capability guide.

Sources and Further Reading

  • NIST/SEMATECH, e-Handbook of Statistical Methods, sections on normal probability plots, Anderson-Darling, and Shapiro-Wilk (itl.nist.gov/div898/handbook).
  • Ralph B. D’Agostino and Michael A. Stephens (eds.), Goodness-of-Fit Techniques, Marcel Dekker, 1986.
  • Samuel S. Shapiro and Martin B. Wilk, “An analysis of variance test for normality,” Biometrika, 1965.
  • George E. P. Box and David R. Cox, “An analysis of transformations,” Journal of the Royal Statistical Society B, 1964.
  • Minitab Support, “Methods and formulas for Normality Test” and “Box-Cox Transformation” (support.minitab.com).

This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.