- Question it answers
- How does Y change as X changes, and what will Y be at a new X?
- Data needed
- One continuous response (Y) and one continuous predictor (X)
- Key output
- Slope with interval, R², S, and prediction intervals
- Null hypothesis
- The true slope is zero
- Assumptions
- Linear, independent, normal, equal-variance errors (L-I-N-E)
- Excel
- Data Analysis > Regression; SLOPE, INTERCEPT, FORECAST.LINEAR
- Minitab
- Stat > Regression > Regression > Fit Regression Model
- Prerequisite
- Correlation
The Idea in Plain Language
Correlation tells you that curing temperature and bond strength move together. Regression goes further: it fits a straight line through the points, so you can say how much the strength changes for each extra degree, and predict the strength at a temperature you have not tested.
The line is chosen by least squares: of all possible lines, it is the one that makes the total of the squared vertical distances from the points to the line as small as it can be. Those distances are the residuals, the part of each result that the line does not explain.
When to Use Simple Linear Regression
| Your situation | Use | Why |
|---|---|---|
| One continuous response (Y) and one continuous predictor (X) in a roughly straight-line pattern | Simple linear regression | Gives the slope, the intercept, and predictions with intervals |
| You only want the strength of the relationship | Correlation | Simpler, with no choice of Y and X |
| Two or more predictors | Multiple regression | Separates the effect of each predictor |
| A curved pattern | Regression with a squared term, or a transformation | A straight line would be a poor fit |
| A pass/fail response | Logistic regression | Predicts a probability |
| The predictor is a category | One-way ANOVA | Compares group means |
How It Works
The model says the response is a straight-line function of X plus random error ε. The sample gives estimates b0 and b1.
| Quantity | Formula | In words |
|---|---|---|
| Slope b1 | Sxy / Sxx | Average change in Y for one more unit of X |
| Intercept b0 | ȳ − b1 x̄ | Where the line meets the Y axis (often outside the data) |
| Residual | ei = yi − (b0 + b1 xi) | What the line misses for each point |
| SSE | Σ ei² | Unexplained variation |
| SSR | Syy − SSE | Variation explained by the line |
| MSE and S | MSE = SSE / (n − 2); S = √MSE | Typical size of a residual, in units of Y |
| R² | SSR / Syy | Share of variation explained |
| F test | F = SSR / MSE, df = 1 and n − 2 | Tests whether the line explains more than chance |
| t test for the slope | t = b1 / SE(b1), SE(b1) = √(MSE / Sxx) | Tests H0: the true slope is zero (same p as the F test) |
| Interval for the mean of Y at x0 | Ŷ ± t √(MSE (1/n + (x0 − x̄)²/Sxx)) | Where the average strength at x0 lies |
| Prediction interval for one new Y | Ŷ ± t √(MSE (1 + 1/n + (x0 − x̄)²/Sxx)) | Where one new bond at x0 will fall |
Assumptions: Think L-I-N-E
| Letter | Assumption | How to check | In the example |
|---|---|---|---|
| L | Linear: the mean of Y is a straight-line function of X | Scatter plot; residuals versus fitted values show no curve | Plot is straight |
| I | Independent errors | Think about the data collection; residuals versus run order show no pattern | Coupons were cured and tested in random order |
| N | Normal errors | Normal probability plot of residuals | Points close to the line |
| E | Equal variance: the scatter is the same at every X | Residuals versus fitted values show no funnel | Even spread |
The most important checks are the plot of the data and the residual-versus-fitted plot. If the residuals curve, the line is the wrong model. If they fan out, the variance is not constant. Judge the plots first, and use tests as a second opinion.
Worked Example by Hand
The same 12 coupons as on the Correlation page. From that page, Sxx = 3575.00, Sxy = 711.50, Syy = 161.777, x̄ = 147.50, and ȳ = 25.583.
- Slope. b1 = 711.50 / 3575.00 = 0.1990 N per °C.
- Intercept. b0 = 25.583 − 0.1990 × 147.50 = -3.772 N. (A bond cured at 0 °C is far outside the data, so the intercept is only an anchor for the line.)
- Fitted values and residuals for each coupon, as in the table below. The residuals sum to zero.
- Error variation. SSE = 20.173, so MSE = 20.173 / 10 = 2.017 and S = √2.017 = 1.420 N. A typical coupon sits about 1.4 N from the line.
- Explained variation. SSR = 161.777 − 20.173 = 141.603. R² = 141.603 / 161.777 = 0.875.
- Test the slope. SE(b1) = √(2.017 / 3575.00) = 0.0238, so t = 0.1990 / 0.0238 = 8.38 on 10 df (p = 0.000008). The F test gives F = 70.19, which equals t².
- Interval for the slope. 0.1990 ± 2.228 × 0.0238 gives a 95% interval of 0.146 to 0.252 N per °C.
| Bond | x (°C) | y (N) | Fitted | Residual | Residual² |
|---|---|---|---|---|---|
| 1 | 120 | 21.2 | 20.11 | +1.09 | 1.19 |
| 2 | 125 | 19.4 | 21.11 | -1.71 | 2.91 |
| 3 | 130 | 22.7 | 22.10 | +0.60 | 0.36 |
| 4 | 135 | 24.9 | 23.10 | +1.80 | 3.26 |
| 5 | 140 | 22.9 | 24.09 | -1.19 | 1.42 |
| 6 | 145 | 23.0 | 25.09 | -2.09 | 4.35 |
| 7 | 150 | 26.9 | 26.08 | +0.82 | 0.67 |
| 8 | 155 | 28.6 | 27.08 | +1.52 | 2.32 |
| 9 | 160 | 26.7 | 28.07 | -1.37 | 1.88 |
| 10 | 165 | 29.5 | 29.07 | +0.43 | 0.19 |
| 11 | 170 | 29.2 | 30.06 | -0.86 | 0.74 |
| 12 | 175 | 32.0 | 31.06 | +0.94 | 0.89 |
| Sum | -0.00 | 20.17 |
| Source | SS | df | MS | F | p |
|---|---|---|---|---|---|
| Regression | 141.603 | 1 | 141.603 | 70.19 | 0.000008 |
| Residual | 20.173 | 10 | 2.017 | ||
| Total | 161.777 | 11 |
Predicting: Confidence Interval or Prediction Interval?
Once you have the line, you will want to use it. There are two different questions, and they have two different intervals.
| Question | Interval | At 150 °C |
|---|---|---|
| What is the average strength of all bonds cured at this temperature? | 95% confidence interval for the mean | 26.08 N, 25.16 to 27.00 |
| What strength will one new bond have? | 95% prediction interval | 26.08 N, 22.78 to 29.38 |
The prediction interval is wider because it adds the scatter of an individual bond around the average. Both are narrowest at the mean temperature (147.5 °C) and widen toward the ends. Outside the studied range the intervals grow quickly, and they also assume the straight line still holds.
Run It in Excel and Minitab
ExcelStep by step
- Put the temperatures in A2:A13 and the strengths in B2:B13, with headings in row 1.
- Choose . Set Input Y Range to and Input X Range to , and tick Labels.
- Tick Residuals, Residual Plots, and Normal Probability Plots, set the confidence level (95%), and click OK.
- Read R Square, Standard Error, Significance F, and the Coefficients table with its p-values and intervals.
- For a quick fit without the full output, use , , , and .
- To predict: returns 26.081. For an interval, calculate ± t × √(MSE(1 + 1/n + (x0 − x̄)²/Sxx)).
Excel’s regression tool does not give prediction intervals, and its residual plots are basic. A scatter chart with a trendline (Display equation and R-squared) draws the line quickly.
MinitabStep by step
- Choose . Set Responses to Strength and Continuous predictors to Temp.
- Click Graphs and tick Four in one residual plots. Click Results to see the full tables and unusual observations.
- Click OK and read the Regression Equation, the Coefficients, the Model Summary, and the Analysis of Variance.
- To draw the line with its bands, choose and tick Display confidence interval and Display prediction interval under Options.
- To predict, choose , enter Temp = 150, and read the Fit, the 95% CI, and the 95% PI.
Minitab flags unusual observations (large residuals, high leverage) automatically. The Assistant () adds a report card with these checks.
What the output looks like
SUMMARY OUTPUT
Regression Statistics
Multiple R 0.935576
R Square 0.875302
Adjusted R Square 0.862832
Standard Error 1.420325
Observations 12
ANOVA
df SS MS F Signif F
Regression 1 141.6034 141.6034 70.1937 7.84e-06
Residual 10 20.1732 2.0173
Total 11 161.7767
Coeff Std Err t Stat P-value Lower 95% Upper 95%
Intercept -3.7723 3.5277 -1.069 0.3101 -11.6325 4.0880
Temp 0.1990 0.0238 8.378 7.84e-06 0.1461 0.2519Regression Equation Strength = -3.77 + 0.1990 Temp Coefficients Term Coef SE Coef T-Value P-Value VIF Constant -3.77 3.53 -1.07 0.310 Temp 0.1990 0.0238 8.38 0.000 1.00 Model Summary S R-sq R-sq(adj) R-sq(pred) 1.42033 87.53% 86.28% 82.45% Analysis of Variance Source DF Adj SS Adj MS F-Value P-Value Regression 1 141.60 141.60 70.19 0.000 Temp 1 141.60 141.60 70.19 0.000 Error 10 20.17 2.017 Total 11 161.78 Prediction for Strength Settings: Temp = 150 Fit SE Fit 95% CI 95% PI 26.081 0.414 (25.158, 27.004) (22.784, 29.377)
Reading and Reporting the Result
- Start with the plots. A straight pattern and random residuals mean the line is a fair description.
- Read the slope with its interval. 0.199 N per °C, 95% CI 0.146 to 0.252. The interval excludes zero, so the relationship is not chance.
- Read R² honestly. 88% explained. Adjusted R² (86%) corrects for model size. Predicted R² (82%) estimates the fit on new data and is the most honest of the three.
- Use S. 1.42 N tells you the size of a typical miss. If the process tolerance is ±2 N, a prediction good to ±2.8 N is not precise enough to set the temperature by.
- Respect the range. Say which temperatures the line covers, and do not extrapolate.
When the Assumptions Do Not Hold
| What you see in the residual plots | Meaning | What to do |
|---|---|---|
| A curve (U shape or arch) in residuals versus fitted values | The relationship is not straight | Add a squared term, transform X or Y, or fit a different model |
| A funnel: residuals widen as fitted values grow | Variance is not constant | Transform Y (often log); weighted regression |
| One point far from the others | An outlier or an influential point | Check the data; refit without it and compare; report both |
| A trend or cycle in residuals against run order | Errors are not independent | Randomize; model the time structure; investigate drift |
| Curved probability plot | Errors are skewed | Transform Y; check for outliers |
See also Multiple Regression, where an omitted variable is a common reason for a pattern in the residuals.
Common Mistakes
| Mistake | Why it misleads | Better |
|---|---|---|
| Fitting a line without plotting | A curve or outlier makes the line meaningless | Plot first, and check residuals |
| Extrapolating | The pattern may not continue | Stay inside the studied range |
| Reading R² as proof of a good model | A curved relationship can give a high R²; a good model can have a low one | Read the residual plots, S, and predicted R² |
| Using the confidence interval for a single item | It is too narrow for one new value | Use the prediction interval |
| Swapping X and Y | The regression of Y on X differs from X on Y | Put the thing you want to predict in Y |
| Concluding that X causes Y | Regression on observational data shows association | Use an experiment to test cause |
| Ignoring outliers and leverage | One influential point can tilt the line | Look at unusual observations and refit |
Try It Yourself
A tool-wear study measured the wear (in 0.01 mm, y) after each hour of use (x).
| Hours (x) | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| Wear (y) | 3.1 | 4.9 | 7.2 | 8.8 | 11.1 | 12.9 |
- Find the least squares line and R².
- Is the slope significant at α = 0.05?
- Predict the wear after 4.5 hours and say what the prediction interval is for.
Show the answer
Slope = 1.977, intercept = 1.080, R² = 0.9984. The slope test gives t = 49.7 with p = 0.000001, so it is highly significant: each extra hour adds about 1.98 units of wear.
At 4.5 hours the fitted wear is 9.98. A prediction interval, not a confidence interval, tells you where the wear of a single tool at that time will fall.
Simple Linear Regression: Frequently Asked Questions
What is the difference between R-squared and adjusted R-squared?
R-squared always rises when you add a term to the model, even a useless one. Adjusted R-squared penalizes extra terms, so it can fall. With a single predictor the two are close. Predicted R-squared goes further and estimates how well the model would predict new data.
How do I know if a straight line is appropriate?
Look at the scatter plot and the plot of residuals against fitted values. If the points follow a straight band and the residuals scatter randomly around zero, a line is reasonable. A curve in the residuals means it is not.
What does the intercept mean?
It is the predicted Y when X is zero. If zero lies outside the range of your data, as it does here, the intercept is only a mathematical anchor for the line and has no practical meaning.
Why is the prediction interval so much wider than the confidence interval?
The confidence interval describes uncertainty about the average response at a given X. The prediction interval must also cover the natural scatter of a single new observation around that average, so it is always wider.
What is a significant slope?
A slope whose confidence interval excludes zero, so the data show that Y changes with X. It does not mean the line predicts well. Check R-squared, S, and the prediction interval for that.
When should I force the line through zero?
Only when a physical reason requires that Y is exactly zero when X is zero, and zero is within or near your data. Forcing it otherwise distorts the slope and the fit statistics. Fit the intercept normally and see whether it is close to zero.
Sources and Further Reading
- NIST/SEMATECH, e-Handbook of Statistical Methods, section on linear least squares regression (itl.nist.gov/div898/handbook).
- Douglas C. Montgomery, Elizabeth A. Peck, and G. Geoffrey Vining, Introduction to Linear Regression Analysis, Wiley.
- Michael H. Kutner, Christopher J. Nachtsheim, John Neter, and William Li, Applied Linear Statistical Models, McGraw-Hill.
- Minitab Support, “Methods and formulas for Fit Regression Model” (support.minitab.com).
- Microsoft Support, documentation for the Analysis ToolPak Regression tool, SLOPE, INTERCEPT, and FORECAST.LINEAR (support.microsoft.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.