- Question it answers
- Which inputs change the chance of a pass/fail outcome, and by how much?
- Data needed
- A 0/1 response and one or more predictors, with enough events
- Key output
- Odds ratios with intervals, and predicted probabilities
- Model
- logit(p) = b0 + b1x1 + …
- Assumptions
- Independent cases; linear in the logit; enough events
- Excel
- Solver on the log-likelihood; EXP for odds ratios
- Minitab
- Stat > Regression > Binary Logistic Regression
- Why it matters
- Most defect data are pass/fail
The Idea in Plain Language
Ordinary regression predicts a number. Logistic regression predicts the probability of a yes/no outcome: a weld that fails or holds, a part that passes or fails, a customer who returns or stays. The inputs can be measurements, categories, or both.
A straight line cannot do this, because it would predict probabilities below 0 and above 1. Logistic regression instead models the log of the odds as a straight line, and the S-shaped logistic curve turns that back into a probability that stays between 0 and 1.
| Term | Definition | Example |
|---|---|---|
| Probability p | Chance of the event | 0.20 |
| Odds | p / (1 − p) | 0.20 / 0.80 = 0.25, or 1 to 4 |
| Log odds (logit) | ln(odds) = b0 + b1x1 + … | ln(0.25) = −1.39 |
| Probability from logit | p = 1 / (1 + e−logit) | 1 / (1 + e1.39) = 0.20 |
| Odds ratio | eb: the factor the odds are multiplied by for a one-unit increase in x | e0.12 = 1.13 per degree |
Worked Example: Weld Failures
100 welds were tested. Each was made at a cure temperature between 170 and 210 °C using material from Supplier A or Supplier B, and was recorded as failed (1) or held (0). 39 failed and 61 held.
| Term | Coefficient | SE | z | P-value | Odds ratio | 95% interval for the odds ratio |
|---|---|---|---|---|---|---|
| Constant | -30.915 | 6.104 | -5.06 | < 0.001 | ||
| Temperature (per °C) | 0.1525 | 0.0309 | 4.94 | < 0.001 | 1.165 | 1.096 to 1.237 |
| Supplier B (vs A) | 1.526 | 0.619 | 2.46 | 0.014 | 4.60 | 1.37 to 15.48 |
- Model: logit(failure) = -30.92 + 0.1525 × Temperature + 1.53 × (Supplier B).
- Temperature: each extra degree multiplies the odds of failure by e0.1525 = 1.165. Ten degrees multiplies them by 4.59 (95% interval 2.51 to 8.42).
- Supplier: at any temperature, Supplier B has 4.60 times the odds of failure of Supplier A (interval 1.37 to 15.48).
Predicting by Hand and Testing the Model
To predict, put the settings in the equation, convert the logit to odds, and the odds to a probability:
| Settings | Logit | Odds = elogit | Probability = odds / (1 + odds) |
|---|---|---|---|
| Supplier A at 190 °C | -1.942 | 0.143 | 12.5% |
| Supplier B at 190 °C | -0.416 | 0.660 | 39.8% |
| Supplier A at 200 °C | -0.417 | 0.659 | 39.7% |
The failure probability is 50% at about 203 °C for Supplier A and 193 °C for Supplier B, where the logit is zero.
Is the model better than nothing? The likelihood ratio test compares the model with one that uses only the overall failure rate:
- Null model log-likelihood = 39 ln(39/100) + 61 ln(61/100) = -66.87, deviance 133.75.
- Fitted model log-likelihood = -40.68, deviance 81.36.
- G = 2 (-40.68 − (-66.87)) = 52.39 on 2 df, p < 0.001.
- Deviance R² = 1 − 81.36/133.75 = 39.2%. Dropping temperature alone costs G = 40.83 (p < 0.001); dropping supplier costs G = 6.69 (p = 0.010). AIC = 87.4.
Checking the Fit
Residual plots are not much use for a yes/no outcome. Three better checks:
- Hosmer-Lemeshow test: sort the welds by predicted probability, split into ten groups, and compare observed and expected failures. χ² = 4.71 on 8 df, p = 0.79. A small p-value would signal poor fit; this one gives no evidence of it.
- Discrimination: the area under the ROC curve (concordance) is 0.89. 0.5 is a coin toss and 1.0 is perfect separation; 0.7 to 0.8 is acceptable and above 0.8 is good.
- Classification: predicting failure when the probability is 0.5 or more gives 83 of 100 correct (83%), with sensitivity 77% (failures caught) and specificity 87% (good welds passed).
| Predicted failure | Predicted hold | |
|---|---|---|
| Actually failed | 30 | 9 |
| Actually held | 8 | 53 |
Run It in Excel and Minitab
ExcelStep by step
- Excel has no logistic regression tool. Enter the response as 0/1 in one column and the predictors next to it, and start with guesses for the coefficients in three cells.
- Probability: for each row, using the coefficient cells.
- Log-likelihood per row: ; sum the column.
- Fit: : maximize the summed log-likelihood by changing the three coefficient cells (GRG Nonlinear). The result matches the table above (-30.92, 0.1525, 1.53).
- Odds ratio: . Likelihood ratio p: (< 0.001).
- For anything beyond a quick check use Minitab or a statistics package.
MinitabStep by step
- Enter the response (for example Fail/Hold, or 1/0), temperature, and supplier in columns.
- . Choose Response in binary response/frequency format, enter the response, and choose the Response event (the outcome you want to model, such as Fail).
- Put Temperature under Continuous predictors and Supplier under Categorical predictors. Minitab uses the first level (alphabetically) as the reference.
- Under Options, choose a confidence level and the unit for odds ratios. Under Results, include the goodness-of-fit tests. Under Graphs, choose residual plots if wanted.
- Predict: for the probability at chosen settings, and for a single-predictor S-curve.
- Quicker: does not cover logistic models, so use the menu above.
Binary Logistic Regression: Fail versus Temp, Supplier
Method
Link function Logit
Rows used 100
Response Information
Variable Value Count
Fail 1 39 (Event)
0 61
Total 100
Deviance Table
Source DF Adj Dev Adj Mean Chi-Square P-Value
Regression 2 52.39 26.19 52.39 0.000
Temp 1 40.83 40.83 40.83 0.000
Supplier 1 6.69 6.69 6.69 0.010
Error 97 81.36 0.84
Total 99 133.75
Model Summary
Deviance Deviance
R-Sq R-Sq(adj) AIC
39.17% (adjusted value also reported) 87.36
Coefficients
Term Coef SE Coef Z-Value P-Value
Constant -30.915 6.104 -5.06 0.000
Temp 0.1525 0.0309 4.94 0.000
Supplier
B 1.526 0.619 2.46 0.014
Odds Ratios for Continuous Predictors
Odds Ratio 95% CI
Temp 1.1647 (1.0963, 1.2374)
Odds Ratios for Categorical Predictors
Level A Level B Odds Ratio 95% CI
B A 4.5997 (1.3667, 15.4805)
Goodness-of-Fit Tests
Test DF Chi-Square P-Value
Deviance 97 81.36 0.873
Pearson 97 85.83 0.784
Hosmer-Lemeshow 8 4.71 0.788Reading and Reporting
- State the event you are modeling (failure, not pass) so the direction of every odds ratio is clear.
- Report odds ratios with intervals, and say per what unit (per degree, per 10 degrees).
- Translate to probabilities for typical settings: most readers think in probabilities, not odds.
- Report the fit: the likelihood-ratio test, a goodness-of-fit test, and the ROC area.
- Say how many events there were. Models need at least about 10 events (and 10 non-events) per predictor.
Common Mistakes
| Mistake | Why it misleads | Better |
|---|---|---|
| Reading an odds ratio as a ratio of probabilities | Odds ratios and risk ratios agree only when the event is rare | Convert to probabilities for typical cases |
| Modeling the wrong event | All the directions flip | State and check the response event |
| Too few events for the number of predictors | Unstable coefficients; overfitting | Aim for 10 events per predictor |
| Perfect separation (a predictor splits the outcome exactly) | The coefficient runs off to infinity | Collect more data, combine levels, or use exact or penalized methods |
| Fitting a straight line to a 0/1 outcome | Predicts probabilities outside 0 to 1 | Use logistic regression |
| Turning a continuous response into pass/fail to analyze | Throws away information | Analyze the measurement if it exists |
| Extrapolating beyond the tested range | The S-curve is only supported by the data you have | Stay within the tested range |
Try It Yourself
A model for the chance a unit fails gives logit = −6 + 0.08 x, where x is the operating hours in hundreds.
- Find the probability of failure at x = 90.
- By what factor do the odds change for each extra 10 units of x?
- At what x is the probability 50%?
Show the answer
Logit at x = 90: −6 + 0.08 × 90 = 1.2. Odds = e1.2 = 3.32, so probability = 3.32 / (1 + 3.32) = 0.769.
Ten units multiply the odds by e0.8 = 2.23.
The probability is 50% when the logit is 0: x = 6 / 0.08 = 75.
Binary Logistic Regression: Frequently Asked Questions
What is logistic regression used for?
To model the probability of a yes/no outcome, such as pass or fail, from one or more inputs that can be continuous or categorical. It also gives odds ratios that measure each input’s effect.
What is an odds ratio?
The factor by which the odds of the event are multiplied for a one-unit increase in a continuous predictor, or for one level of a category compared with the reference level. An odds ratio of 1 means no effect.
How is it different from linear regression?
Linear regression predicts a continuous response and assumes constant variance and normal errors. Logistic regression predicts a probability, fits by maximum likelihood, and uses a binomial error model.
What is the Hosmer-Lemeshow test?
A goodness-of-fit test that groups cases by predicted probability and compares the observed and expected number of events in each group. A small p-value suggests the model fits poorly.
How many observations do I need?
A common guide is at least 10 events and 10 non-events for every predictor in the model. Fewer gives unstable estimates.
What does the ROC area mean?
It is the probability that a randomly chosen event has a higher predicted probability than a randomly chosen non-event. 0.5 is no better than chance, and 1.0 is perfect discrimination.
Sources and Further Reading
- David W. Hosmer, Stanley Lemeshow, and Rodney X. Sturdivant, Applied Logistic Regression, Wiley.
- Alan Agresti, Categorical Data Analysis, Wiley.
- NIST/SEMATECH, e-Handbook of Statistical Methods (itl.nist.gov/div898/handbook).
- Minitab Support, “Methods and formulas for Fit Binary Logistic Model” (support.minitab.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.