Question it answers
Which inputs drive the response, and by how much, with the others held constant?
Data needed
One continuous response and two or more predictors, with enough observations
Key output
Coefficients with intervals, adjusted and predicted R², S, VIF
Watch for
Correlated predictors (VIF above 5) and overfitting
Assumptions
Linear, independent, normal, equal-variance errors
Excel
Data Analysis > Regression with adjacent X columns
Minitab
Stat > Regression > Regression > Fit Regression Model; Best Subsets
Prerequisite
Simple linear regression

The Idea in Plain Language

Real processes have more than one input. Bond strength depends on curing temperature, but also on the clamping pressure and the dwell time. Multiple regression fits an equation with several predictors at once, and it answers a question that one-predictor regression cannot: what is the effect of this input with the others held constant?

Y = β0 + β1 X1 + β2 X2 + … + βk Xk + ε

Each coefficient is the change in the response for a one-unit change in that predictor while the other predictors stay fixed. That “holding the others constant” is the whole point. It is what lets you separate inputs that move together.

The classic surprise. An input can look important on its own and become unimportant once the others are in the model. This page’s example shows it: dwell time is strongly related to strength by itself, but adds nothing once temperature and pressure are included.

When to Use Multiple Regression

Your situationUseWhy
One continuous response and two or more predictorsMultiple regressionSeparates the effect of each predictor
One predictorSimple linear regressionThe starting point
Predictors are categories (machine, shift)ANOVA, or regression with dummy variablesA category with k levels becomes k − 1 indicator columns
A mix of continuous predictors and categoriesGeneral linear modelIn Minitab, Fit Regression Model accepts categorical predictors directly
Planned experiment with set factor levelsDOE analysisDesigned data give clean, uncorrelated estimates
A pass/fail responseLogistic regressionPredicts a probability
Observational versus designed data. With data collected from normal production, the predictors tend to move together, and the coefficients can be unstable. With a designed experiment, you set the factors independently, and the estimates are clean. Use multiple regression to find candidates, and an experiment to confirm them.

How It Works

The least squares estimates minimize the sum of squared residuals, as before. In matrix form, with X holding a column of ones and one column per predictor, the coefficients are b = (X′X)−1 X′y. Software does this; what matters is how to read the results.

QuantityWhat it meansHow to use it
Coefficient bjChange in Y per unit of Xj, others held constantThe effect size, in the units of Y per unit of Xj
t test for each coefficientDoes this predictor add anything once the others are in the model?Small p-value: keep. Large p-value: candidate to remove
F test for the whole modelDo the predictors together explain more than chance?Significant F: at least one predictor matters
R²Share of variation explainedAlways rises when you add a predictor
Adjusted R²R² penalized for the number of predictorsRises only if a new predictor earns its place
Predicted R²Fit to new data, from leave-one-out residuals (PRESS)A big gap below R² signals overfitting
STypical size of a residualCompare with the tolerance
VIF = 1 / (1 − Rj²)How much a predictor’s variance is inflated by correlation with the other predictorsAbove 5 is a concern; above 10 is serious

The t-tests and the F test assume the same L-I-N-E conditions as simple regression: a linear relationship, independent errors, normal errors, and equal variance. Check them with residual plots.

Worked Example: Three Inputs and One Surprise

The adhesive study now has 20 bonds, with curing temperature, clamping pressure, and dwell time recorded for each. The question is which inputs drive strength.

Bond strength study: 20 bonds
BondTemp (°C)Pressure (kPa)Dwell (s)Strength (N)
1140385.430.6
2130345.027.8
3125323.427.0
4155406.735.1
5145447.534.1
6175387.238.4
7165467.139.3
8165407.537.1
9125468.432.2
10150426.337.3
11155426.536.8
12125385.730.8
13145324.931.6
14140446.735.1
15130484.133.9
16165445.338.6
17145385.533.4
18125323.026.9
19140343.028.9
20170427.340.3

Step 1: Look at each input alone

PredictorSimple regression slopep-valueR² alone
Temp0.2220.00075%
Pressure0.5740.00147%
Dwell1.8430.00148%

On its own, each of the three looks significant. Dwell time looks as important as pressure.

Step 2: Fit all three together

TermCoefficientSEtp-value95% CIVIF
Constant-8.23672.6983-3.050.008-13.957 to -2.517
Temp0.18100.016411.050.0000.146 to 0.2161.36
Pressure0.37200.05876.340.0000.248 to 0.4961.63
Dwell0.14520.20740.700.494-0.295 to 0.5852.03

The full model has R² = 95.2%, adjusted R² = 94.3%, and S = 1.00 N. The overall F test is F = 105.7, p < 0.001. But the dwell-time coefficient has p = 0.49: once temperature and pressure are in the model, dwell adds nothing.

Temp 11.05 Pressure 6.34 Dwell 0.70 Critical t = 2.12 Absolute t value (coefficient divided by its standard error)
Standardized effects. Bars that pass the critical-t line are significant. Dwell time falls far short.

Why did dwell time lose its effect?

Dwell time and pressure are correlated (r = 0.62). In this process, higher pressure goes with longer dwell. So on its own, dwell time looked important only because it was standing in for pressure. When pressure is in the model, dwell has no extra information to give.

2 4 6 8 10 30 35 40 45 50 Pressure (kPa) Dwell time (s)
Pressure and dwell time move together, so each carries some of the same information about strength.
This is the reason to fit all candidates together. Looking at one input at a time would have suggested lengthening the dwell time. The full model shows that effort would be wasted.

Building the Model: Which Predictors to Keep?

More predictors are not always better. A predictor that adds nothing makes the model harder to explain, can reduce its predictive power, and uses up degrees of freedom. A good model is the simplest one that explains the data well.

ModelR²Adjusted R²Predicted R²Mallows CpS (N)
Temp74.7%73.3%68.8%68.42.164
Pressure46.9%44.0%33.2%160.93.134
Dwell48.0%45.2%34.2%157.13.100
Temp + Pressure95.1%94.5%93.7%2.50.984
Temp + Dwell83.1%81.2%77.8%42.21.817
Pressure + Dwell58.5%53.7%42.0%124.12.849
Temp + Pressure + Dwell95.2%94.3%93.3%4.01.000
  • Temp + Pressure has the highest adjusted R² (94.5%), the highest predicted R² (93.7%), and the smallest S (0.984 N). Adding Dwell raises R² slightly but lowers the other measures.
  • Mallows Cp should be close to the number of parameters (including the constant). For Temp + Pressure it is about 2.5, near 3, so the model has little bias.
  • Best subsets shows every combination, so you can compare them as above. Stepwise methods add or remove terms automatically; they are quick, but can pick up chance patterns, so use them to suggest candidates, and then judge with subject knowledge and a confirmation test.
  • Keep a term for a reason. A term that is significant and makes engineering sense belongs in the model. A term that is significant by chance, or makes no physical sense, does not.

The reduced model

TermCoefficientSEtp-value95% CIVIF
Constant-9.06692.3873-3.800.001-14.104 to -4.030
Temp0.18610.014512.860.0000.156 to 0.2171.36
Pressure0.39560.04738.360.0000.296 to 0.4951.63

Strength = -9.07 + 0.1861 Temp + 0.3956 Pressure. Each extra °C adds 0.186 N and each extra kPa adds 0.396 N, with the other held constant.

-2 -1 0 1 2 Normal score (expected z) Normal probability plot of residuals Residual (N)
Residuals close to the line support normal errors.
28 30 32 34 36 38 40 Fitted strength (N) Residuals versus fitted values Residual (N)
No curve and no funnel in the residuals supports the model.
Predicting. For a bond cured at 150 °C and 40 kPa, the model gives 34.67 N. The 95% confidence interval for the average of all such bonds is 34.19 to 35.15 N, and the 95% prediction interval for one new bond is 32.54 to 36.80 N.

Multicollinearity: When Predictors Overlap

When predictors are strongly correlated with each other, the model cannot tell their effects apart. The overall fit may be fine, but individual coefficients become unstable: large standard errors, wrong signs, and results that change when a single point is removed.

SignWhat to checkWhat to do
VIF above 5 (or tolerance below 0.2)Which predictors are correlatedDrop one of the pair, combine them into one measure, or collect data that breaks the link
A coefficient with the wrong signCorrelations between predictorsCheck the VIFs; do not interpret the sign until it is resolved
Overall F significant but no individual t significantMany correlated predictorsUse fewer predictors, or principal components or partial least squares
Coefficients change a lot when a term is added or removedCollinearity or an influential pointCompare models; look at unusual observations

In the example the largest VIF is 2.03 (Dwell), a mild overlap, well under 5. Even so, it was enough to make dwell time look important on its own. A designed experiment avoids the problem, because the factor settings are chosen to be uncorrelated.

Categorical Predictors and Interactions

A category such as machine (A, B, C) enters a regression as indicator, or dummy, variables: for k levels there are k − 1 columns of 0 and 1, and the coefficient of each is the difference from the reference level. Minitab builds them for you when you list the variable under Categorical predictors. In Excel you create the columns yourself.

An interaction term is the product of two predictors (for example Temp × Pressure). It lets the effect of one depend on the level of the other. Add one only when there is a reason and the data have enough points, and keep the main effects in the model when you keep the interaction. See Two-Way ANOVA and Interactions.

Run It in Excel and Minitab

ExcelStep by step

  1. Put the predictors in adjacent columns: Temp in A, Pressure in B, Dwell in C, and Strength in D, with headings in row 1. Excel needs the X columns side by side.
  2. Choose Data > Data Analysis > Regression. Set Input Y Range to D1:D21 and Input X Range to A1:C21, and tick Labels.
  3. Tick Residuals, Residual Plots, and Normal Probability Plots, then click OK.
  4. Read Adjusted R Square, Significance F, and the p-value of each coefficient. To drop a term, delete its column and run the tool again.
  5. VIF: for each predictor, regress it on the other predictors and use =1/(1-R Square) from that run.
  6. To predict, multiply the coefficients by the new settings, or use =TREND(known_y, known_x, new_x).

Excel gives no VIF, no predicted R-squared, no best subsets, and no prediction intervals, so it is best for a quick fit. Use Minitab for model building.

MinitabStep by step

  1. Choose Stat > Regression > Regression > Fit Regression Model. Set Responses to Strength and Continuous predictors to Temp, Pressure, and Dwell (list categories under Categorical predictors).
  2. Click Graphs and tick Four in one. Click Results and keep the coefficient table with VIF.
  3. Read the Coefficients (VIF is in the last column), the Model Summary (S, R-sq, adjusted, predicted), and the Analysis of Variance.
  4. To compare models, choose Stat > Regression > Regression > Best Subsets and select the predictors. For automatic selection, use Fit Regression Model > Stepwise.
  5. To predict, choose Stat > Regression > Regression > Predict, enter the settings, and read the Fit, 95% CI, and 95% PI.
  6. To add an interaction, open Model, select two predictors, and click Add under Terms through order.

The Assistant (Assistant > Regression > Multiple Regression) reports unusual points and suggests terms.

What the output looks like

Excel Data Analysis ToolPak output (full model)
SUMMARY OUTPUT

Regression Statistics
Multiple R              0.975696
R Square                0.951983
Adjusted R Square       0.942980
Standard Error          0.999597
Observations                  20

ANOVA
              df        SS        MS         F   Signif F
Regression     3   316.961   105.654    105.74   9.23e-11
Residual      16    15.987    0.9992
Total         19   332.948

               Coeff  Std Err  t Stat   P-value
Intercept    -8.2367   2.6983  -3.053    0.0076
Temp          0.1810   0.0164  11.051    0.0000
Pressure      0.3720   0.0587   6.340    0.0000
Dwell         0.1452   0.2074   0.700    0.4939
Minitab session window: full model (typed excerpt, simplified)
Regression Equation

Strength = -8.24 + 0.1810 Temp + 0.3720 Pressure + 0.1452 Dwell

Coefficients

Term        Coef  SE Coef  T-Value  P-Value   VIF
Constant  -8.2367   2.6983    -3.05    0.008       
Temp       0.1810   0.0164    11.05    0.000   1.36
Pressure   0.3720   0.0587     6.34    0.000   1.63
Dwell      0.1452   0.2074     0.70    0.494   2.03

Model Summary

      S    R-sq  R-sq(adj)  R-sq(pred)
0.99960  95.20%     94.30%      93.26%

Analysis of Variance

Source       DF  Adj SS  Adj MS  F-Value  P-Value
Regression    3  316.96  105.65   105.74    0.000
Error        16   15.99   0.999
Total        19  332.95
Minitab session window: reduced model with Temp and Pressure (typed excerpt, simplified)
Regression Equation

Strength = -9.07 + 0.1861 Temp + 0.3956 Pressure

Coefficients

Term        Coef  SE Coef  T-Value  P-Value   VIF
Constant  -9.0669   2.3873    -3.80    0.001       
Temp       0.1861   0.0145    12.86    0.000   1.09
Pressure   0.3956   0.0473     8.36    0.000   1.09

Model Summary

      S    R-sq  R-sq(adj)  R-sq(pred)
0.98450  95.05%     94.47%      93.72%

Analysis of Variance

Source       DF  Adj SS  Adj MS  F-Value  P-Value
Regression    2  316.47  158.24   163.26    0.000
Error        17   16.48   0.969
Total        19  332.95

Reading and Reporting the Result

  1. Check the overall F first. It is significant, so the predictors together explain more than chance.
  2. Read each coefficient with its interval and p-value, and interpret it as the effect with the others held constant.
  3. Use adjusted and predicted R², not R² alone. Here 94.5% and 93.7% for the reduced model.
  4. Remove terms that add nothing, and refit. Compare the two models, as above.
  5. Look at the residual plots for the final model.
  6. State the range of each predictor in the data, and do not predict outside it.
A sentence you can use. Multiple regression on 20 bonds showed that bond strength increased by 0.186 N per °C of curing temperature (95% CI 0.156 to 0.217) and by 0.396 N per kPa of clamping pressure (95% CI 0.296 to 0.495), each with the other held constant. Dwell time did not add to the model (p = 0.49). The model explained 94% of the variation (adjusted R²; S = 0.98 N).

Common Mistakes

MistakeWhy it misleadsBetter
Judging each input by its separate regressionCorrelated inputs stand in for each otherFit them together and read the partial effects
Adding predictors to push R² upR² always rises; the model overfitsUse adjusted and predicted R²
Ignoring multicollinearityCoefficients become unstable and unreliableCheck VIF; remove or combine overlapping predictors
Trusting automatic selection blindlyIt can pick chance patternsUse it for candidates; confirm with subject knowledge and new data
Reading coefficients as causesObservational data show associationConfirm with a designed experiment
Comparing coefficients across different unitsA bigger number may be a smaller effectCompare standardized effects or t values
ExtrapolatingThe model is valid only within the dataState the ranges used
Skipping the residual plotsCurves, funnels, and outliers go unseenAlways look at the four-in-one plot

Try It Yourself

A regression of weld penetration (mm) on current, speed, and wire feed rate for 25 welds gave the output below.

Output (typed excerpt)
Term        Coef  SE Coef  T-Value  P-Value   VIF
Constant  -2.410    0.980    -2.46    0.023
Current    0.0310   0.0042     7.38    0.000   1.8
Speed     -0.0520   0.0150    -3.47    0.002   1.4
Feed       0.0140   0.0120     1.17    0.255   9.6

S = 0.21   R-sq = 88.0%   R-sq(adj) = 86.3%   R-sq(pred) = 80.1%
  • Which predictors are significant at α = 0.05?
  • What is odd about the feed rate?
  • What would you do next?
Show the answer

Current and speed are significant. Penetration rises 0.031 mm per unit of current and falls 0.052 mm per unit of speed, each with the others held constant.

Feed has p = 0.255 and a VIF of 9.6. It is strongly correlated with the other predictors (probably with current, since the machine feeds more wire at a higher current), so its own effect cannot be separated. Do not conclude that feed rate does not matter.

Next steps: check which predictor feed is correlated with, and refit without it (or without the one it duplicates) and compare adjusted and predicted R². If feed rate matters physically, run a designed experiment that sets current and feed independently.

Multiple Regression: Frequently Asked Questions

How many predictors can I use?

A common guide is at least 10 to 15 observations per predictor, so 20 observations supports only one or two reliably. With too many predictors the model fits the noise (overfitting), and predicted R-squared drops well below R-squared. Fewer, well-chosen predictors are better.

What is the difference between R-squared, adjusted R-squared, and predicted R-squared?

R-squared is the share of variation explained and always rises when you add a predictor. Adjusted R-squared corrects for the number of predictors. Predicted R-squared estimates the fit on new data using leave-one-out residuals, and is the best guard against overfitting.

A predictor is significant alone but not in the model. Why?

It is probably correlated with another predictor that carries the same information. Alone, it stands in for the other one. In the model, with the other predictor present, it adds nothing. This is why correlated predictors should be fitted together.

What VIF is too high?

Above 5 is a concern and above 10 is serious. A VIF of 5 means the variance of that coefficient is five times larger than it would be if the predictor were uncorrelated with the others.

Should I use stepwise regression?

As a screening aid, yes. As a final answer, be careful. Automatic selection can choose terms that fit chance patterns, and it ignores engineering knowledge. Compare the candidate models with adjusted and predicted R-squared and the residual plots, and confirm with new data.

Can I include a categorical variable?

Yes. Minitab’s Fit Regression Model accepts categorical predictors directly. In Excel, create 0/1 indicator columns for all but one level.

Sources and Further Reading

  • NIST/SEMATECH, e-Handbook of Statistical Methods, sections on process modeling and linear regression (itl.nist.gov/div898/handbook).
  • Douglas C. Montgomery, Elizabeth A. Peck, and G. Geoffrey Vining, Introduction to Linear Regression Analysis, Wiley.
  • Michael H. Kutner, Christopher J. Nachtsheim, John Neter, and William Li, Applied Linear Statistical Models, McGraw-Hill.
  • Minitab Support, “Methods and formulas for Fit Regression Model” and “Best Subsets” (support.minitab.com).
  • Microsoft Support, documentation for the Analysis ToolPak Regression tool and TREND (support.microsoft.com).

This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.