- Question it answers
- Can I trust this model, and what is wrong if not?
- Data needed
- A fitted regression, ANOVA, or DOE model with its residuals
- Key output
- Residual plots, flagged points, and a decision to refit or transform
- Four assumptions
- Linear, constant variance, normal, independent
- Measures
- Standardized residual, leverage, Cook’s distance
- Excel
- Regression tool residuals, formulas for leverage and Cook’s D
- Minitab
- Fit Regression Model > Graphs, Storage, Options (Durbin-Watson)
- Why it matters
- R² cannot tell you the model is wrong; residuals can
The Idea in Plain Language
A residual is what is left over after the model has done its best: the observed value minus the value the model predicted. If the model captured the real pattern, the residuals should look like random noise. If they do not, the model is wrong in a way the R² will not show, and its p-values, intervals, and predictions cannot be trusted.
Every regression, ANOVA, and designed experiment makes the same four assumptions about the errors, and residuals are how you check them:
| Assumption | What it means | How to check it |
|---|---|---|
| Linearity (the model form is right) | No curve or other pattern left behind | Residuals versus fitted values (or versus each predictor) |
| Constant variance | Spread of residuals is the same everywhere | Residuals versus fitted values: no funnel |
| Normality | Residuals are roughly bell-shaped | Normal probability plot of the residuals, histogram |
| Independence | One error tells nothing about the next | Residuals versus observation order; Durbin-Watson |
The Patterns and What They Mean
| Pattern | Diagnosis | Remedy |
|---|---|---|
| Random scatter around zero | Assumptions look fine | Proceed |
| U shape or arch | The relationship is curved | Add a squared term, or transform x |
| Funnel (spread grows or shrinks) | Unequal variance | Transform y (log or square root), or use weighted least squares |
| Long run above then below zero, or a trend in order | Errors are correlated, often over time | Add time or a missing variable; use a time-series model |
| One or two points far from the rest | Outliers | Check the data; do not delete without a reason |
| Points off the line in the probability plot tails | Non-normal errors | Transform y; check for outliers; see Normality Tests |
| Two clusters of residuals | A hidden group (machine, shift) | Add that factor to the model |
Standardized Residuals, Leverage, and Influence
Raw residuals are in the units of y and have different variances at different x. Three standardized measures make the checking systematic:
| Measure | Formula | Flag when |
|---|---|---|
| Standardized residual | ri = ei / (s √(1 − hi)) | |r| > 2 (about 1 in 20 by chance); |r| > 3 is a strong flag |
| Leverage | hi = 1/n + (xi − x̄)² / Sxx (simple regression) | h > 3p / n, where p is the number of model terms including the constant |
| Cook’s distance | Di = (ri² / p) × hi / (1 − hi) | D > 1 is a common rule; also look at D > 4/n |
| Deleted (studentized) residual | ti = ri √((n − p − 1) / (n − p − ri²)) | Compare with a t distribution; stable when one point is extreme |
Worked Example 1: Finding an Outlier and an Influential Point
A line was fitted to 22 points: y = 18.67 + 1.158 x, with s = 3.46 and R² = 88.4%. On the summary table alone this looks like an adequate fit.
- Standardized residuals: only point 8 exceeds 2 (+2.65).
- Leverage cut-off = 3p / n = 3 × 2 / 22 = 0.273. Point 22 has h = 0.587 because its x of 48 is far from the mean of 20.9.
- Cook’s distance for point 22 = (-1.91² / 2) × 0.587 / (1 − 0.587) = 2.59, far above 1. For point 8 it is 0.19.
- Refit without point 22: y = 14.86 + 1.363 x, s = 3.21, R² = 84.8%. The slope moves from 1.16 to 1.36.
| Point | x | y | Fit | Residual | Std resid | Leverage | Cook’s D | Flag |
|---|---|---|---|---|---|---|---|---|
| 2 | 29.0 | 58.1 | 52.25 | +5.85 | +1.78 | 0.094 | 0.16 | |
| 8 | 18.2 | 48.7 | 39.74 | +8.96 | +2.65 | 0.051 | 0.19 | R |
| 22 | 48.0 | 70.0 | 74.25 | -4.25 | -1.91 | 0.587 | 2.59 | X D |
Worked Example 2: Curvature and the Fix
A response measured at 30 settings of x. A straight line gives R² = 96.7%, which sounds good, with s = 2.08.
- The squared term has t = -12.5, p < 0.001.
- Adding it changes R² from 96.7% to 99.5% and drops s from 2.08 to 0.81.
- Model: y = 7.65 + 3.341 x − 0.0671 x².
Worked Example 3: Unequal Variance and a Transformation
Output (y) rises with load (x), but the scatter widens as the load increases: the residual standard deviation is 16.7 for the lower half of the fitted values and 31.0 for the upper half, a ratio of 1.9 (Levene’s test p 0.010). That is the funnel in the third panel above.
Taking the log of both variables (appropriate when errors are proportional to the size of the value) gives residual standard deviations of 0.24 and 0.20, a ratio of 0.9 (Levene’s test p = 0.61). The spread is now even.
Run It in Excel and Minitab
ExcelStep by step
- Run and tick Residuals, Standardized Residuals, Residual Plots, and Line Fit Plots. Excel’s “standardized residual” divides by a simple estimate and is not the same as Minitab’s.
- Residuals versus fitted: plot the residual column against the predicted values with .
- Leverage (simple regression): . Point 22 gives 0.587.
- Studentized residual: . Cook’s D: .
- Normal probability plot: sort the residuals, compute , and scatter the two columns.
- Durbin-Watson: (the drifting example gives 0.51). Values well below 2 mean positive correlation.
MinitabStep by step
- . Under Graphs, choose Four in one (normal plot, residuals versus fits, histogram, residuals versus order), or choose each plot separately, including residuals versus each predictor.
- Under Storage, tick Residuals, Standardized residuals, Deleted residuals, Leverages (Hi), and Cook’s distance to put them in the worksheet for plotting or sorting.
- Under Options, tick Durbin-Watson statistic to test independence.
- Read the session window: the Fits and Diagnostics for Unusual Observations table lists points with a large standardized residual (R) or unusual X (X).
- For a designed experiment use the same plots under .
Regression Analysis: y versus x
Model Summary
S R-sq R-sq(adj)
3.4628 88.38% 87.80%
Fits and Diagnostics for Unusual Observations
Obs y Fit Resid Std Resid
8 48.7 39.74 +8.96 +2.65 R
22 70.0 74.25 -4.25 -1.91 X
R Large residual
X Unusual X
Stored columns (leverage and Cook's distance) for obs 22: HI = 0.5869 COOK = 2.5865Reading and Reporting
- Always show the residual plots beside any regression, ANOVA, or DOE result, or say that you checked them.
- Go through the four assumptions in order: linearity, constant variance, normality, independence.
- Investigate every flagged point: a typing error, a different process condition, or a real extreme value?
- Report what you changed and why: a transformation, an added term, or a removed point with its cause.
- Limit predictions to the range of the data. A leverage point is a warning that the model is least certain there.
Common Mistakes
| Mistake | Why it misleads | Better |
|---|---|---|
| Trusting R² and p-values without looking at residuals | A curved or unstable model can show a high R² | Always plot residuals |
| Deleting points because they are large residuals | Data should be removed for a documented cause, not a convenient one | Investigate, then decide; show the fit both ways |
| Ignoring leverage | A far-out x can control the whole line while its residual looks small | Check leverage and Cook’s distance |
| Testing normality of the raw y instead of the residuals | The assumption is about the errors | Check the residuals |
| Using the normality test alone with a big sample | Tiny departures become significant; with small samples nothing is | Use the probability plot and judgment |
| Skipping the order plot | Drift and correlation go unseen | Plot residuals against time or run order |
Try It Yourself
A simple regression with 20 points flags an observation with standardized residual 2.4 and leverage 0.45.
- Is its leverage high? (Use the 3p/n rule.)
- Calculate its Cook’s distance. Is it influential?
Show the answer
p = 2 and n = 20, so the cut-off is 3 × 2 / 20 = 0.30. Leverage 0.45 is above it: a high-leverage point.
Cook’s D = (2.4² / 2) × 0.45 / (1 − 0.45) = 2.36, which is above 1. The point is influential: check it and refit without it to see how much the model changes.
Residual Analysis and Model Checking: Frequently Asked Questions
What are residuals?
The difference between each observed value and the value the model predicts. They represent what the model did not explain, and if the model is adequate they should look like random noise.
What should a good residual plot look like?
A horizontal band of points scattered randomly around zero, with no curve, no funnel, no clusters, and no trend over the observation order.
What is the difference between an outlier and a high-leverage point?
An outlier is unusual in the response given its x. A high-leverage point is unusual in x. A point with both can be influential, changing the fitted line a lot.
Should I remove outliers?
Only with a documented reason, such as a measurement or entry error. Otherwise keep the point, report the fit with and without it, and look for the cause.
What is Cook’s distance?
A measure of how much the fitted model would change if one point were removed. It combines the size of the residual with the leverage of the point. Values above 1, or well above 4/n, deserve attention.
What if the residuals are not normal?
With a reasonable sample the tests are fairly robust to mild departures. Check for outliers and skew, consider a transformation, and use the residual probability plot rather than a formal test alone.
Sources and Further Reading
- R. Dennis Cook and Sanford Weisberg, Residuals and Influence in Regression, Chapman and Hall.
- Douglas C. Montgomery, Elizabeth A. Peck, and G. Geoffrey Vining, Introduction to Linear Regression Analysis, Wiley.
- NIST/SEMATECH, e-Handbook of Statistical Methods, Process Modeling (itl.nist.gov/div898/handbook).
- Minitab Support, “Methods and formulas for residuals and diagnostic measures” (support.minitab.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.