- Question it answers
- Which inputs drive the response, and by how much, with the others held constant?
- Data needed
- One continuous response and two or more predictors, with enough observations
- Key output
- Coefficients with intervals, adjusted and predicted R², S, VIF
- Watch for
- Correlated predictors (VIF above 5) and overfitting
- Assumptions
- Linear, independent, normal, equal-variance errors
- Excel
- Data Analysis > Regression with adjacent X columns
- Minitab
- Stat > Regression > Regression > Fit Regression Model; Best Subsets
- Prerequisite
- Simple linear regression
The Idea in Plain Language
Real processes have more than one input. Bond strength depends on curing temperature, but also on the clamping pressure and the dwell time. Multiple regression fits an equation with several predictors at once, and it answers a question that one-predictor regression cannot: what is the effect of this input with the others held constant?
Each coefficient is the change in the response for a one-unit change in that predictor while the other predictors stay fixed. That “holding the others constant” is the whole point. It is what lets you separate inputs that move together.
When to Use Multiple Regression
| Your situation | Use | Why |
|---|---|---|
| One continuous response and two or more predictors | Multiple regression | Separates the effect of each predictor |
| One predictor | Simple linear regression | The starting point |
| Predictors are categories (machine, shift) | ANOVA, or regression with dummy variables | A category with k levels becomes k − 1 indicator columns |
| A mix of continuous predictors and categories | General linear model | In Minitab, Fit Regression Model accepts categorical predictors directly |
| Planned experiment with set factor levels | DOE analysis | Designed data give clean, uncorrelated estimates |
| A pass/fail response | Logistic regression | Predicts a probability |
How It Works
The least squares estimates minimize the sum of squared residuals, as before. In matrix form, with X holding a column of ones and one column per predictor, the coefficients are b = (X′X)−1 X′y. Software does this; what matters is how to read the results.
| Quantity | What it means | How to use it |
|---|---|---|
| Coefficient bj | Change in Y per unit of Xj, others held constant | The effect size, in the units of Y per unit of Xj |
| t test for each coefficient | Does this predictor add anything once the others are in the model? | Small p-value: keep. Large p-value: candidate to remove |
| F test for the whole model | Do the predictors together explain more than chance? | Significant F: at least one predictor matters |
| R² | Share of variation explained | Always rises when you add a predictor |
| Adjusted R² | R² penalized for the number of predictors | Rises only if a new predictor earns its place |
| Predicted R² | Fit to new data, from leave-one-out residuals (PRESS) | A big gap below R² signals overfitting |
| S | Typical size of a residual | Compare with the tolerance |
| VIF = 1 / (1 − Rj²) | How much a predictor’s variance is inflated by correlation with the other predictors | Above 5 is a concern; above 10 is serious |
The t-tests and the F test assume the same L-I-N-E conditions as simple regression: a linear relationship, independent errors, normal errors, and equal variance. Check them with residual plots.
Worked Example: Three Inputs and One Surprise
The adhesive study now has 20 bonds, with curing temperature, clamping pressure, and dwell time recorded for each. The question is which inputs drive strength.
| Bond | Temp (°C) | Pressure (kPa) | Dwell (s) | Strength (N) |
|---|---|---|---|---|
| 1 | 140 | 38 | 5.4 | 30.6 |
| 2 | 130 | 34 | 5.0 | 27.8 |
| 3 | 125 | 32 | 3.4 | 27.0 |
| 4 | 155 | 40 | 6.7 | 35.1 |
| 5 | 145 | 44 | 7.5 | 34.1 |
| 6 | 175 | 38 | 7.2 | 38.4 |
| 7 | 165 | 46 | 7.1 | 39.3 |
| 8 | 165 | 40 | 7.5 | 37.1 |
| 9 | 125 | 46 | 8.4 | 32.2 |
| 10 | 150 | 42 | 6.3 | 37.3 |
| 11 | 155 | 42 | 6.5 | 36.8 |
| 12 | 125 | 38 | 5.7 | 30.8 |
| 13 | 145 | 32 | 4.9 | 31.6 |
| 14 | 140 | 44 | 6.7 | 35.1 |
| 15 | 130 | 48 | 4.1 | 33.9 |
| 16 | 165 | 44 | 5.3 | 38.6 |
| 17 | 145 | 38 | 5.5 | 33.4 |
| 18 | 125 | 32 | 3.0 | 26.9 |
| 19 | 140 | 34 | 3.0 | 28.9 |
| 20 | 170 | 42 | 7.3 | 40.3 |
Step 1: Look at each input alone
| Predictor | Simple regression slope | p-value | R² alone |
|---|---|---|---|
| Temp | 0.222 | 0.000 | 75% |
| Pressure | 0.574 | 0.001 | 47% |
| Dwell | 1.843 | 0.001 | 48% |
On its own, each of the three looks significant. Dwell time looks as important as pressure.
Step 2: Fit all three together
| Term | Coefficient | SE | t | p-value | 95% CI | VIF |
|---|---|---|---|---|---|---|
| Constant | -8.2367 | 2.6983 | -3.05 | 0.008 | -13.957 to -2.517 | |
| Temp | 0.1810 | 0.0164 | 11.05 | 0.000 | 0.146 to 0.216 | 1.36 |
| Pressure | 0.3720 | 0.0587 | 6.34 | 0.000 | 0.248 to 0.496 | 1.63 |
| Dwell | 0.1452 | 0.2074 | 0.70 | 0.494 | -0.295 to 0.585 | 2.03 |
The full model has R² = 95.2%, adjusted R² = 94.3%, and S = 1.00 N. The overall F test is F = 105.7, p < 0.001. But the dwell-time coefficient has p = 0.49: once temperature and pressure are in the model, dwell adds nothing.
Why did dwell time lose its effect?
Dwell time and pressure are correlated (r = 0.62). In this process, higher pressure goes with longer dwell. So on its own, dwell time looked important only because it was standing in for pressure. When pressure is in the model, dwell has no extra information to give.
Building the Model: Which Predictors to Keep?
More predictors are not always better. A predictor that adds nothing makes the model harder to explain, can reduce its predictive power, and uses up degrees of freedom. A good model is the simplest one that explains the data well.
| Model | R² | Adjusted R² | Predicted R² | Mallows Cp | S (N) |
|---|---|---|---|---|---|
| Temp | 74.7% | 73.3% | 68.8% | 68.4 | 2.164 |
| Pressure | 46.9% | 44.0% | 33.2% | 160.9 | 3.134 |
| Dwell | 48.0% | 45.2% | 34.2% | 157.1 | 3.100 |
| Temp + Pressure | 95.1% | 94.5% | 93.7% | 2.5 | 0.984 |
| Temp + Dwell | 83.1% | 81.2% | 77.8% | 42.2 | 1.817 |
| Pressure + Dwell | 58.5% | 53.7% | 42.0% | 124.1 | 2.849 |
| Temp + Pressure + Dwell | 95.2% | 94.3% | 93.3% | 4.0 | 1.000 |
- Temp + Pressure has the highest adjusted R² (94.5%), the highest predicted R² (93.7%), and the smallest S (0.984 N). Adding Dwell raises R² slightly but lowers the other measures.
- Mallows Cp should be close to the number of parameters (including the constant). For Temp + Pressure it is about 2.5, near 3, so the model has little bias.
- Best subsets shows every combination, so you can compare them as above. Stepwise methods add or remove terms automatically; they are quick, but can pick up chance patterns, so use them to suggest candidates, and then judge with subject knowledge and a confirmation test.
- Keep a term for a reason. A term that is significant and makes engineering sense belongs in the model. A term that is significant by chance, or makes no physical sense, does not.
The reduced model
| Term | Coefficient | SE | t | p-value | 95% CI | VIF |
|---|---|---|---|---|---|---|
| Constant | -9.0669 | 2.3873 | -3.80 | 0.001 | -14.104 to -4.030 | |
| Temp | 0.1861 | 0.0145 | 12.86 | 0.000 | 0.156 to 0.217 | 1.36 |
| Pressure | 0.3956 | 0.0473 | 8.36 | 0.000 | 0.296 to 0.495 | 1.63 |
Strength = -9.07 + 0.1861 Temp + 0.3956 Pressure. Each extra °C adds 0.186 N and each extra kPa adds 0.396 N, with the other held constant.
Multicollinearity: When Predictors Overlap
When predictors are strongly correlated with each other, the model cannot tell their effects apart. The overall fit may be fine, but individual coefficients become unstable: large standard errors, wrong signs, and results that change when a single point is removed.
| Sign | What to check | What to do |
|---|---|---|
| VIF above 5 (or tolerance below 0.2) | Which predictors are correlated | Drop one of the pair, combine them into one measure, or collect data that breaks the link |
| A coefficient with the wrong sign | Correlations between predictors | Check the VIFs; do not interpret the sign until it is resolved |
| Overall F significant but no individual t significant | Many correlated predictors | Use fewer predictors, or principal components or partial least squares |
| Coefficients change a lot when a term is added or removed | Collinearity or an influential point | Compare models; look at unusual observations |
In the example the largest VIF is 2.03 (Dwell), a mild overlap, well under 5. Even so, it was enough to make dwell time look important on its own. A designed experiment avoids the problem, because the factor settings are chosen to be uncorrelated.
Categorical Predictors and Interactions
A category such as machine (A, B, C) enters a regression as indicator, or dummy, variables: for k levels there are k − 1 columns of 0 and 1, and the coefficient of each is the difference from the reference level. Minitab builds them for you when you list the variable under Categorical predictors. In Excel you create the columns yourself.
An interaction term is the product of two predictors (for example Temp × Pressure). It lets the effect of one depend on the level of the other. Add one only when there is a reason and the data have enough points, and keep the main effects in the model when you keep the interaction. See Two-Way ANOVA and Interactions.
Run It in Excel and Minitab
ExcelStep by step
- Put the predictors in adjacent columns: Temp in A, Pressure in B, Dwell in C, and Strength in D, with headings in row 1. Excel needs the X columns side by side.
- Choose . Set Input Y Range to and Input X Range to , and tick Labels.
- Tick Residuals, Residual Plots, and Normal Probability Plots, then click OK.
- Read Adjusted R Square, Significance F, and the p-value of each coefficient. To drop a term, delete its column and run the tool again.
- VIF: for each predictor, regress it on the other predictors and use from that run.
- To predict, multiply the coefficients by the new settings, or use .
Excel gives no VIF, no predicted R-squared, no best subsets, and no prediction intervals, so it is best for a quick fit. Use Minitab for model building.
MinitabStep by step
- Choose . Set Responses to Strength and Continuous predictors to Temp, Pressure, and Dwell (list categories under Categorical predictors).
- Click Graphs and tick Four in one. Click Results and keep the coefficient table with VIF.
- Read the Coefficients (VIF is in the last column), the Model Summary (S, R-sq, adjusted, predicted), and the Analysis of Variance.
- To compare models, choose and select the predictors. For automatic selection, use .
- To predict, choose , enter the settings, and read the Fit, 95% CI, and 95% PI.
- To add an interaction, open Model, select two predictors, and click Add under Terms through order.
The Assistant () reports unusual points and suggests terms.
What the output looks like
SUMMARY OUTPUT
Regression Statistics
Multiple R 0.975696
R Square 0.951983
Adjusted R Square 0.942980
Standard Error 0.999597
Observations 20
ANOVA
df SS MS F Signif F
Regression 3 316.961 105.654 105.74 9.23e-11
Residual 16 15.987 0.9992
Total 19 332.948
Coeff Std Err t Stat P-value
Intercept -8.2367 2.6983 -3.053 0.0076
Temp 0.1810 0.0164 11.051 0.0000
Pressure 0.3720 0.0587 6.340 0.0000
Dwell 0.1452 0.2074 0.700 0.4939Regression Equation
Strength = -8.24 + 0.1810 Temp + 0.3720 Pressure + 0.1452 Dwell
Coefficients
Term Coef SE Coef T-Value P-Value VIF
Constant -8.2367 2.6983 -3.05 0.008
Temp 0.1810 0.0164 11.05 0.000 1.36
Pressure 0.3720 0.0587 6.34 0.000 1.63
Dwell 0.1452 0.2074 0.70 0.494 2.03
Model Summary
S R-sq R-sq(adj) R-sq(pred)
0.99960 95.20% 94.30% 93.26%
Analysis of Variance
Source DF Adj SS Adj MS F-Value P-Value
Regression 3 316.96 105.65 105.74 0.000
Error 16 15.99 0.999
Total 19 332.95Regression Equation
Strength = -9.07 + 0.1861 Temp + 0.3956 Pressure
Coefficients
Term Coef SE Coef T-Value P-Value VIF
Constant -9.0669 2.3873 -3.80 0.001
Temp 0.1861 0.0145 12.86 0.000 1.09
Pressure 0.3956 0.0473 8.36 0.000 1.09
Model Summary
S R-sq R-sq(adj) R-sq(pred)
0.98450 95.05% 94.47% 93.72%
Analysis of Variance
Source DF Adj SS Adj MS F-Value P-Value
Regression 2 316.47 158.24 163.26 0.000
Error 17 16.48 0.969
Total 19 332.95Reading and Reporting the Result
- Check the overall F first. It is significant, so the predictors together explain more than chance.
- Read each coefficient with its interval and p-value, and interpret it as the effect with the others held constant.
- Use adjusted and predicted R², not R² alone. Here 94.5% and 93.7% for the reduced model.
- Remove terms that add nothing, and refit. Compare the two models, as above.
- Look at the residual plots for the final model.
- State the range of each predictor in the data, and do not predict outside it.
Common Mistakes
| Mistake | Why it misleads | Better |
|---|---|---|
| Judging each input by its separate regression | Correlated inputs stand in for each other | Fit them together and read the partial effects |
| Adding predictors to push R² up | R² always rises; the model overfits | Use adjusted and predicted R² |
| Ignoring multicollinearity | Coefficients become unstable and unreliable | Check VIF; remove or combine overlapping predictors |
| Trusting automatic selection blindly | It can pick chance patterns | Use it for candidates; confirm with subject knowledge and new data |
| Reading coefficients as causes | Observational data show association | Confirm with a designed experiment |
| Comparing coefficients across different units | A bigger number may be a smaller effect | Compare standardized effects or t values |
| Extrapolating | The model is valid only within the data | State the ranges used |
| Skipping the residual plots | Curves, funnels, and outliers go unseen | Always look at the four-in-one plot |
Try It Yourself
A regression of weld penetration (mm) on current, speed, and wire feed rate for 25 welds gave the output below.
Term Coef SE Coef T-Value P-Value VIF Constant -2.410 0.980 -2.46 0.023 Current 0.0310 0.0042 7.38 0.000 1.8 Speed -0.0520 0.0150 -3.47 0.002 1.4 Feed 0.0140 0.0120 1.17 0.255 9.6 S = 0.21 R-sq = 88.0% R-sq(adj) = 86.3% R-sq(pred) = 80.1%
- Which predictors are significant at α = 0.05?
- What is odd about the feed rate?
- What would you do next?
Show the answer
Current and speed are significant. Penetration rises 0.031 mm per unit of current and falls 0.052 mm per unit of speed, each with the others held constant.
Feed has p = 0.255 and a VIF of 9.6. It is strongly correlated with the other predictors (probably with current, since the machine feeds more wire at a higher current), so its own effect cannot be separated. Do not conclude that feed rate does not matter.
Next steps: check which predictor feed is correlated with, and refit without it (or without the one it duplicates) and compare adjusted and predicted R². If feed rate matters physically, run a designed experiment that sets current and feed independently.
Multiple Regression: Frequently Asked Questions
How many predictors can I use?
A common guide is at least 10 to 15 observations per predictor, so 20 observations supports only one or two reliably. With too many predictors the model fits the noise (overfitting), and predicted R-squared drops well below R-squared. Fewer, well-chosen predictors are better.
What is the difference between R-squared, adjusted R-squared, and predicted R-squared?
R-squared is the share of variation explained and always rises when you add a predictor. Adjusted R-squared corrects for the number of predictors. Predicted R-squared estimates the fit on new data using leave-one-out residuals, and is the best guard against overfitting.
A predictor is significant alone but not in the model. Why?
It is probably correlated with another predictor that carries the same information. Alone, it stands in for the other one. In the model, with the other predictor present, it adds nothing. This is why correlated predictors should be fitted together.
What VIF is too high?
Above 5 is a concern and above 10 is serious. A VIF of 5 means the variance of that coefficient is five times larger than it would be if the predictor were uncorrelated with the others.
Should I use stepwise regression?
As a screening aid, yes. As a final answer, be careful. Automatic selection can choose terms that fit chance patterns, and it ignores engineering knowledge. Compare the candidate models with adjusted and predicted R-squared and the residual plots, and confirm with new data.
Can I include a categorical variable?
Yes. Minitab’s Fit Regression Model accepts categorical predictors directly. In Excel, create 0/1 indicator columns for all but one level.
Sources and Further Reading
- NIST/SEMATECH, e-Handbook of Statistical Methods, sections on process modeling and linear regression (itl.nist.gov/div898/handbook).
- Douglas C. Montgomery, Elizabeth A. Peck, and G. Geoffrey Vining, Introduction to Linear Regression Analysis, Wiley.
- Michael H. Kutner, Christopher J. Nachtsheim, John Neter, and William Li, Applied Linear Statistical Models, McGraw-Hill.
- Minitab Support, “Methods and formulas for Fit Regression Model” and “Best Subsets” (support.minitab.com).
- Microsoft Support, documentation for the Analysis ToolPak Regression tool and TREND (support.microsoft.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.