Question it answers
How does Y change as X changes, and what will Y be at a new X?
Data needed
One continuous response (Y) and one continuous predictor (X)
Key output
Slope with interval, R², S, and prediction intervals
Null hypothesis
The true slope is zero
Assumptions
Linear, independent, normal, equal-variance errors (L-I-N-E)
Excel
Data Analysis > Regression; SLOPE, INTERCEPT, FORECAST.LINEAR
Minitab
Stat > Regression > Regression > Fit Regression Model
Prerequisite
Correlation

The Idea in Plain Language

Correlation tells you that curing temperature and bond strength move together. Regression goes further: it fits a straight line through the points, so you can say how much the strength changes for each extra degree, and predict the strength at a temperature you have not tested.

The line is chosen by least squares: of all possible lines, it is the one that makes the total of the squared vertical distances from the points to the line as small as it can be. Those distances are the residuals, the part of each result that the line does not explain.

15 20 25 30 35 120 140 160 180 Curing temperature (°C) Bond strength (N)
The least squares line for the 12 bonds. Each red segment is a residual. The line is strength = -3.77 + 0.1990 × temperature.
The slope is the answer to the practical question. Here every extra degree is associated with about 0.2 N more strength. That is a more useful statement than “r = 0.94”.

When to Use Simple Linear Regression

Your situationUseWhy
One continuous response (Y) and one continuous predictor (X) in a roughly straight-line patternSimple linear regressionGives the slope, the intercept, and predictions with intervals
You only want the strength of the relationshipCorrelationSimpler, with no choice of Y and X
Two or more predictorsMultiple regressionSeparates the effect of each predictor
A curved patternRegression with a squared term, or a transformationA straight line would be a poor fit
A pass/fail responseLogistic regressionPredicts a probability
The predictor is a categoryOne-way ANOVACompares group means

How It Works

Y = β0 + β1 X + ε

The model says the response is a straight-line function of X plus random error ε. The sample gives estimates b0 and b1.

QuantityFormulaIn words
Slope b1Sxy / SxxAverage change in Y for one more unit of X
Intercept b0ȳ − b1 x̄Where the line meets the Y axis (often outside the data)
Residualei = yi − (b0 + b1 xi)What the line misses for each point
SSEΣ ei²Unexplained variation
SSRSyy − SSEVariation explained by the line
MSE and SMSE = SSE / (n − 2); S = √MSETypical size of a residual, in units of Y
R²SSR / SyyShare of variation explained
F testF = SSR / MSE, df = 1 and n − 2Tests whether the line explains more than chance
t test for the slopet = b1 / SE(b1), SE(b1) = √(MSE / Sxx)Tests H0: the true slope is zero (same p as the F test)
Interval for the mean of Y at x0Ŷ ± t √(MSE (1/n + (x0 − x̄)²/Sxx))Where the average strength at x0 lies
Prediction interval for one new YŶ ± t √(MSE (1 + 1/n + (x0 − x̄)²/Sxx))Where one new bond at x0 will fall
Total sum of squares = 161.78 141.60 87.5% Explained by temperature (regression) 20.17 12.5% Unexplained (residual)
The total variation in strength splits into the part explained by the line and the part left over. R² is the explained share, 87.5%.

Assumptions: Think L-I-N-E

LetterAssumptionHow to checkIn the example
LLinear: the mean of Y is a straight-line function of XScatter plot; residuals versus fitted values show no curvePlot is straight
IIndependent errorsThink about the data collection; residuals versus run order show no patternCoupons were cured and tested in random order
NNormal errorsNormal probability plot of residualsPoints close to the line
EEqual variance: the scatter is the same at every XResiduals versus fitted values show no funnelEven spread
-2 -1 0 1 2 Normal score (expected z) Normal probability plot of residuals Residual (N)
A straight pattern in the probability plot supports normal errors.
20 22 24 26 28 30 32 Fitted strength (N) Residuals versus fitted values Residual (N)
Random scatter around zero, with no curve or funnel, supports a straight line and equal variance.

The most important checks are the plot of the data and the residual-versus-fitted plot. If the residuals curve, the line is the wrong model. If they fan out, the variance is not constant. Judge the plots first, and use tests as a second opinion.

Worked Example by Hand

The same 12 coupons as on the Correlation page. From that page, Sxx = 3575.00, Sxy = 711.50, Syy = 161.777, x̄ = 147.50, and ȳ = 25.583.

  1. Slope. b1 = 711.50 / 3575.00 = 0.1990 N per °C.
  2. Intercept. b0 = 25.583 − 0.1990 × 147.50 = -3.772 N. (A bond cured at 0 °C is far outside the data, so the intercept is only an anchor for the line.)
  3. Fitted values and residuals for each coupon, as in the table below. The residuals sum to zero.
  4. Error variation. SSE = 20.173, so MSE = 20.173 / 10 = 2.017 and S = √2.017 = 1.420 N. A typical coupon sits about 1.4 N from the line.
  5. Explained variation. SSR = 161.777 − 20.173 = 141.603. R² = 141.603 / 161.777 = 0.875.
  6. Test the slope. SE(b1) = √(2.017 / 3575.00) = 0.0238, so t = 0.1990 / 0.0238 = 8.38 on 10 df (p = 0.000008). The F test gives F = 70.19, which equals t².
  7. Interval for the slope. 0.1990 ± 2.228 × 0.0238 gives a 95% interval of 0.146 to 0.252 N per °C.
Fitted values and residuals
Bondx (°C)y (N)FittedResidualResidual²
112021.220.11+1.091.19
212519.421.11-1.712.91
313022.722.10+0.600.36
413524.923.10+1.803.26
514022.924.09-1.191.42
614523.025.09-2.094.35
715026.926.08+0.820.67
815528.627.08+1.522.32
916026.728.07-1.371.88
1016529.529.07+0.430.19
1117029.230.06-0.860.74
1217532.031.06+0.940.89
Sum-0.0020.17
SourceSSdfMSFp
Regression141.6031141.60370.190.000008
Residual20.173102.017
Total161.77711
Conclusion. Bond strength increases by an estimated 0.199 N for each extra °C of curing temperature (95% CI 0.146 to 0.252, p < 0.001). Temperature explains 88% of the variation in strength, and the typical residual is 1.4 N.

Predicting: Confidence Interval or Prediction Interval?

Once you have the line, you will want to use it. There are two different questions, and they have two different intervals.

15 20 25 30 35 120 140 160 180 Curing temperature (°C) Bond strength (N)
The darker band is the 95% confidence interval for the mean strength at each temperature. The lighter, wider band is the 95% prediction interval for one new bond.
QuestionIntervalAt 150 °C
What is the average strength of all bonds cured at this temperature?95% confidence interval for the mean26.08 N, 25.16 to 27.00
What strength will one new bond have?95% prediction interval26.08 N, 22.78 to 29.38

The prediction interval is wider because it adds the scatter of an individual bond around the average. Both are narrowest at the mean temperature (147.5 °C) and widen toward the ends. Outside the studied range the intervals grow quickly, and they also assume the straight line still holds.

Do not extrapolate. The data cover 120 to 175 °C. At 200 °C the model gives 36.0 N with a prediction interval of 31.7 to 40.3 N, but the adhesive may degrade at that temperature, and the line says nothing about it.

Run It in Excel and Minitab

ExcelStep by step

  1. Put the temperatures in A2:A13 and the strengths in B2:B13, with headings in row 1.
  2. Choose Data > Data Analysis > Regression. Set Input Y Range to B1:B13 and Input X Range to A1:A13, and tick Labels.
  3. Tick Residuals, Residual Plots, and Normal Probability Plots, set the confidence level (95%), and click OK.
  4. Read R Square, Standard Error, Significance F, and the Coefficients table with its p-values and intervals.
  5. For a quick fit without the full output, use =SLOPE(B2:B13,A2:A13), =INTERCEPT(B2:B13,A2:A13), =RSQ(B2:B13,A2:A13), and =STEYX(B2:B13,A2:A13).
  6. To predict: =FORECAST.LINEAR(150,B2:B13,A2:A13) returns 26.081. For an interval, calculate ± t × √(MSE(1 + 1/n + (x0 − x̄)²/Sxx)).

Excel’s regression tool does not give prediction intervals, and its residual plots are basic. A scatter chart with a trendline (Display equation and R-squared) draws the line quickly.

MinitabStep by step

  1. Choose Stat > Regression > Regression > Fit Regression Model. Set Responses to Strength and Continuous predictors to Temp.
  2. Click Graphs and tick Four in one residual plots. Click Results to see the full tables and unusual observations.
  3. Click OK and read the Regression Equation, the Coefficients, the Model Summary, and the Analysis of Variance.
  4. To draw the line with its bands, choose Stat > Regression > Regression > Fitted Line Plot and tick Display confidence interval and Display prediction interval under Options.
  5. To predict, choose Stat > Regression > Regression > Predict, enter Temp = 150, and read the Fit, the 95% CI, and the 95% PI.

Minitab flags unusual observations (large residuals, high leverage) automatically. The Assistant (Assistant > Regression > Simple Regression) adds a report card with these checks.

What the output looks like

Excel Data Analysis ToolPak output (Regression)
SUMMARY OUTPUT

Regression Statistics
Multiple R              0.935576
R Square                0.875302
Adjusted R Square       0.862832
Standard Error          1.420325
Observations                  12

ANOVA
              df        SS        MS        F   Signif F
Regression     1  141.6034  141.6034  70.1937   7.84e-06
Residual      10   20.1732    2.0173
Total         11  161.7767

               Coeff  Std Err  t Stat   P-value  Lower 95%  Upper 95%
Intercept    -3.7723   3.5277  -1.069    0.3101   -11.6325     4.0880
Temp          0.1990   0.0238   8.378  7.84e-06     0.1461     0.2519
Minitab session window (typed excerpt, simplified)
Regression Equation

Strength = -3.77 + 0.1990 Temp

Coefficients

Term        Coef  SE Coef  T-Value  P-Value   VIF
Constant   -3.77    3.53    -1.07    0.310
Temp      0.1990  0.0238     8.38    0.000  1.00

Model Summary

      S    R-sq  R-sq(adj)  R-sq(pred)
1.42033  87.53%     86.28%      82.45%

Analysis of Variance

Source       DF  Adj SS  Adj MS  F-Value  P-Value
Regression    1  141.60  141.60    70.19    0.000
  Temp        1  141.60  141.60    70.19    0.000
Error        10   20.17   2.017
Total        11  161.78

Prediction for Strength

Settings: Temp = 150

   Fit  SE Fit        95% CI              95% PI
26.081   0.414  (25.158, 27.004)  (22.784, 29.377)

Reading and Reporting the Result

  1. Start with the plots. A straight pattern and random residuals mean the line is a fair description.
  2. Read the slope with its interval. 0.199 N per °C, 95% CI 0.146 to 0.252. The interval excludes zero, so the relationship is not chance.
  3. Read R² honestly. 88% explained. Adjusted R² (86%) corrects for model size. Predicted R² (82%) estimates the fit on new data and is the most honest of the three.
  4. Use S. 1.42 N tells you the size of a typical miss. If the process tolerance is ±2 N, a prediction good to ±2.8 N is not precise enough to set the temperature by.
  5. Respect the range. Say which temperatures the line covers, and do not extrapolate.
A sentence you can use. Simple linear regression showed that bond strength increased by 0.199 N for each additional °C of curing temperature (95% CI 0.146 to 0.252, p < 0.001; R² = 88%, S = 1.42 N, n = 12, 120 to 175 °C).

When the Assumptions Do Not Hold

What you see in the residual plotsMeaningWhat to do
A curve (U shape or arch) in residuals versus fitted valuesThe relationship is not straightAdd a squared term, transform X or Y, or fit a different model
A funnel: residuals widen as fitted values growVariance is not constantTransform Y (often log); weighted regression
One point far from the othersAn outlier or an influential pointCheck the data; refit without it and compare; report both
A trend or cycle in residuals against run orderErrors are not independentRandomize; model the time structure; investigate drift
Curved probability plotErrors are skewedTransform Y; check for outliers

See also Multiple Regression, where an omitted variable is a common reason for a pattern in the residuals.

Common Mistakes

MistakeWhy it misleadsBetter
Fitting a line without plottingA curve or outlier makes the line meaninglessPlot first, and check residuals
ExtrapolatingThe pattern may not continueStay inside the studied range
Reading R² as proof of a good modelA curved relationship can give a high R²; a good model can have a low oneRead the residual plots, S, and predicted R²
Using the confidence interval for a single itemIt is too narrow for one new valueUse the prediction interval
Swapping X and YThe regression of Y on X differs from X on YPut the thing you want to predict in Y
Concluding that X causes YRegression on observational data shows associationUse an experiment to test cause
Ignoring outliers and leverageOne influential point can tilt the lineLook at unusual observations and refit

Try It Yourself

A tool-wear study measured the wear (in 0.01 mm, y) after each hour of use (x).

Hours (x)123456
Wear (y)3.14.97.28.811.112.9
  • Find the least squares line and R².
  • Is the slope significant at α = 0.05?
  • Predict the wear after 4.5 hours and say what the prediction interval is for.
Show the answer

Slope = 1.977, intercept = 1.080, R² = 0.9984. The slope test gives t = 49.7 with p = 0.000001, so it is highly significant: each extra hour adds about 1.98 units of wear.

At 4.5 hours the fitted wear is 9.98. A prediction interval, not a confidence interval, tells you where the wear of a single tool at that time will fall.

Simple Linear Regression: Frequently Asked Questions

What is the difference between R-squared and adjusted R-squared?

R-squared always rises when you add a term to the model, even a useless one. Adjusted R-squared penalizes extra terms, so it can fall. With a single predictor the two are close. Predicted R-squared goes further and estimates how well the model would predict new data.

How do I know if a straight line is appropriate?

Look at the scatter plot and the plot of residuals against fitted values. If the points follow a straight band and the residuals scatter randomly around zero, a line is reasonable. A curve in the residuals means it is not.

What does the intercept mean?

It is the predicted Y when X is zero. If zero lies outside the range of your data, as it does here, the intercept is only a mathematical anchor for the line and has no practical meaning.

Why is the prediction interval so much wider than the confidence interval?

The confidence interval describes uncertainty about the average response at a given X. The prediction interval must also cover the natural scatter of a single new observation around that average, so it is always wider.

What is a significant slope?

A slope whose confidence interval excludes zero, so the data show that Y changes with X. It does not mean the line predicts well. Check R-squared, S, and the prediction interval for that.

When should I force the line through zero?

Only when a physical reason requires that Y is exactly zero when X is zero, and zero is within or near your data. Forcing it otherwise distorts the slope and the fit statistics. Fit the intercept normally and see whether it is close to zero.

Sources and Further Reading

  • NIST/SEMATECH, e-Handbook of Statistical Methods, section on linear least squares regression (itl.nist.gov/div898/handbook).
  • Douglas C. Montgomery, Elizabeth A. Peck, and G. Geoffrey Vining, Introduction to Linear Regression Analysis, Wiley.
  • Michael H. Kutner, Christopher J. Nachtsheim, John Neter, and William Li, Applied Linear Statistical Models, McGraw-Hill.
  • Minitab Support, “Methods and formulas for Fit Regression Model” (support.minitab.com).
  • Microsoft Support, documentation for the Analysis ToolPak Regression tool, SLOPE, INTERCEPT, and FORECAST.LINEAR (support.microsoft.com).

This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.