- Question it answers
- Do two measured variables move together in a straight-line pattern, and how tightly?
- Data needed
- Two continuous variables measured on the same units
- Key output
- r from −1 to +1, its confidence interval, and r²
- Null hypothesis
- The true correlation is zero
- Plot first
- A scatter plot, to see curves and outliers
- Excel
- =CORREL, =RSQ, and Data Analysis > Correlation
- Minitab
- Stat > Basic Statistics > Correlation
- Next step
- Simple linear regression to predict one from the other
The Idea in Plain Language
When the oven runs hotter, does the adhesive bond get stronger? When the line speeds up, do defects rise? Correlation answers a narrow question: do two measured variables move together in a straight-line pattern, and how tightly?
The most common measure is the Pearson correlation coefficient, r. It runs from −1 to +1. A value near +1 means that high values of one variable go with high values of the other. Near −1 means high goes with low. Near 0 means there is no straight-line relationship. The square of r, r², is the share of the variation in one variable that moves along with the other.
When to Use Correlation
| Your situation | Use | Why |
|---|---|---|
| Two continuous variables, a roughly straight-line pattern, no strong outliers | Pearson correlation | Measures the strength of the linear pattern |
| A curved but always rising or falling pattern, ranked data, or outliers | Spearman rank correlation | Uses ranks, so it measures any steady rise or fall and resists outliers |
| You want to predict one variable from the other | Simple linear regression | Gives the equation, the slope, and prediction intervals |
| Several inputs | Multiple regression | Separates the effect of each input |
| One variable is a category (machine, shift) | One-way ANOVA | Compares group averages |
| Both variables are categories | Chi-square test of association | Tests whether the categories are related |
How It Works
Each point is described by how far it sits from the mean temperature and from the mean strength. Points in the lower-left and upper-right quadrants (both below, or both above, the means) pull the correlation up. Points in the other two quadrants pull it down.
| Quantity | Formula | In words |
|---|---|---|
| Sxx | Σ (x − x̄)² | Variation in x around its mean |
| Syy | Σ (y − ȳ)² | Variation in y around its mean |
| Sxy | Σ (x − x̄)(y − ȳ) | How x and y vary together |
| Correlation r | Sxy / √(Sxx Syy) | Together-variation scaled to run from −1 to +1 |
| Test statistic | t = r √(n − 2) / √(1 − r²), df = n − 2 | Tests H0: the true correlation is zero |
| Confidence interval | tanh( arctanh(r) ± 1.96 / √(n − 3) ) | Fisher’s z transform gives a range for the true correlation |
| |r| | Rough description | r² |
|---|---|---|
| 0.00 to 0.30 | Weak or none | Under 9% |
| 0.30 to 0.70 | Moderate | 9% to 49% |
| 0.70 to 0.90 | Strong | 49% to 81% |
| 0.90 to 1.00 | Very strong | Over 81% |
These labels are conventions, not rules. In a measurement-system study, r = 0.9 is poor agreement. In a field study of people, r = 0.5 can be important. Judge r against what matters in your setting.
Assumptions and Pitfalls
| Issue | What goes wrong | What to do |
|---|---|---|
| A curved relationship | r can be near zero even when the relationship is strong | Plot first; transform, fit a curve, or use regression with a squared term |
| Outliers | One extreme point can create or destroy a correlation | Find the cause; report r with and without it; use Spearman |
| Restricted range | A narrow range of x shrinks r | Do not compare r values from studies with different ranges |
| Mixed groups | Two groups with different levels can create a correlation that exists within neither | Plot by group; stratify by machine, shift, or material |
| Lurking variable | A third variable drives both, such as production volume | Think about what else changes together; use a designed experiment to test cause |
| Non-independent observations | Time series data trend together and give high r that means little | Look at changes instead of levels, or use time-series methods |
| Many correlations tested | Test enough pairs and some will be “significant” by chance | Decide the pairs in advance; adjust for the number of tests |
For the significance test and the confidence interval, the standard method assumes that the two variables together follow a roughly bivariate normal distribution. With small samples, check the scatter plot for outliers and skew.
Worked Example by Hand
An adhesive is cured at different temperatures, and the bond strength is measured on 12 test coupons. The question is whether higher curing temperature goes with higher strength.
| Bond | x (°C) | y (N) | x − x̄ | y − ȳ | (x − x̄)² | (y − ȳ)² | (x − x̄)(y − ȳ) |
|---|---|---|---|---|---|---|---|
| 1 | 120 | 21.2 | -27.50 | -4.38 | 756.25 | 19.21 | +120.54 |
| 2 | 125 | 19.4 | -22.50 | -6.18 | 506.25 | 38.23 | +139.12 |
| 3 | 130 | 22.7 | -17.50 | -2.88 | 306.25 | 8.31 | +50.46 |
| 4 | 135 | 24.9 | -12.50 | -0.68 | 156.25 | 0.47 | +8.54 |
| 5 | 140 | 22.9 | -7.50 | -2.68 | 56.25 | 7.20 | +20.12 |
| 6 | 145 | 23.0 | -2.50 | -2.58 | 6.25 | 6.67 | +6.46 |
| 7 | 150 | 26.9 | +2.50 | +1.32 | 6.25 | 1.73 | +3.29 |
| 8 | 155 | 28.6 | +7.50 | +3.02 | 56.25 | 9.10 | +22.63 |
| 9 | 160 | 26.7 | +12.50 | +1.12 | 156.25 | 1.25 | +13.96 |
| 10 | 165 | 29.5 | +17.50 | +3.92 | 306.25 | 15.34 | +68.54 |
| 11 | 170 | 29.2 | +22.50 | +3.62 | 506.25 | 13.08 | +81.38 |
| 12 | 175 | 32.0 | +27.50 | +6.42 | 756.25 | 41.17 | +176.46 |
| Sum | 3575.00 | 161.78 | +711.50 |
- Means. x̄ = 147.50 °C and ȳ = 25.583 N.
- Sums of squares and products from the table: Sxx = 3575.00, Syy = 161.777, Sxy = 711.50.
- Correlation. r = 711.50 / √(3575.00 × 161.777) = 0.936. Then r² = 0.875, so 87.5% of the variation in strength goes along with temperature.
- Test whether the true correlation is zero. t = 0.936 × √(12 − 2) / √(1 − 0.875) = 8.38 on 10 degrees of freedom, so p = 0.000008.
- Confidence interval. arctanh(0.936) = 1.701, standard error 1 / √(12 − 3) = 0.333. Back-transforming 1.701 ± 1.96 × 0.333 gives a 95% interval of 0.78 to 0.98.
- Rank version. Spearman’s rho, the correlation of the ranks, is 0.944 (p = 0.000004), which agrees because the pattern is a steady rise.
Run It in Excel and Minitab
ExcelStep by step
- Put the temperatures in A2:A13 and the strengths in B2:B13, with headings in row 1.
- returns r (0.9356). returns r².
- For the p-value: with r in a cell and n = 12.
- For the 95% interval: for the lower limit, and plus for the upper limit.
- For several variables at once, choose to get a matrix of r values.
- Spearman: rank each column with , then apply CORREL to the two rank columns.
- Draw the scatter plot with and add a trendline if you want to see the line.
Excel’s CORREL gives only r. It does not give a p-value or an interval, so use the formulas above.
MinitabStep by step
- Enter Temp and Strength in two columns.
- Draw the scatter plot first: (Y variable: Strength; X variable: Temp).
- Choose and select both variables.
- Under Options, choose the method (Pearson or Spearman) and keep the confidence interval ticked.
- Click OK and read the correlation, the interval, and the p-value in the session window.
- For many variables, select all of them to get a matrix; draws the scatter plots together.
Option names can vary slightly between Minitab versions. The Assistant () also reports the correlation with a diagnostic report card.
What the output looks like
=CORREL(A2:A13,B2:B13) 0.935576 =RSQ(B2:B13,A2:A13) 0.875302 =T.DIST.2T(t, 10), with t = 8.3782 7.84e-06
Method Correlation type Pearson Number of rows used 12 Correlations Pearson correlation 95% CI for ρ P-Value 0.936 (0.781, 0.982) 0.000
Reading and Reporting the Result
- Look at the plot, then at r. Here r = 0.94: strong and positive.
- Use the interval. The 95% interval 0.78 to 0.98 shows how well the sample pins down the true correlation. With only 12 points the interval is wide, even though r is high.
- The p-value is a yes/no on “is it zero?”. It does not say the correlation is useful. A large sample makes a trivial r significant.
- Translate r². 88% of the variation in strength goes along with temperature. The other 12% comes from other things.
- State the limits. Report the range of x that was studied. Do not claim a relationship outside it.
When the Usual Method Does Not Fit
| Problem | Better approach |
|---|---|
| Curved relationship | Plot; transform one variable (log, square root), or fit a curve in regression |
| Outliers or skewed data | Spearman rank correlation, and report what happens with and without the outliers |
| Ordered categories (ratings 1 to 5) | Spearman rank correlation |
| Groups mixed together | Calculate r within each group, or fit regression with the group as a factor |
| Time-ordered data that trend | Compare changes from period to period, or use time-series methods |
| Need to show cause | A designed experiment, with the factor set on purpose and runs randomized; see the DOE guide |
Common Mistakes
| Mistake | Why it misleads | Better |
|---|---|---|
| Calculating r without plotting | Curves, clusters, and outliers are invisible | Plot first, every time |
| Saying “X causes Y” from r | Correlation does not identify cause | Say “is associated with”, and test cause with an experiment |
| Treating r = 0 as “no relationship” | It means no linear relationship | Look for curves |
| Reporting only the p-value | Says nothing about strength | Give r, the interval, and n |
| Using the correlation of averages | Averaging hides the scatter and inflates r | Use the individual data |
| Extrapolating beyond the data | The pattern may not continue | State the range studied |
| Comparing r across different ranges | Range changes r | Compare slopes, or use the same range |
Try It Yourself
The number of setup changes per week (x) and the weekly output in thousands of units (y) were recorded for six weeks.
| Week | Setup changes (x) | Output (y) |
|---|---|---|
| 1 | 2 | 30 |
| 2 | 4 | 28 |
| 3 | 5 | 24 |
| 4 | 7 | 21 |
| 5 | 9 | 18 |
| 6 | 11 | 12 |
- Calculate the correlation coefficient and decide whether it is significant at α = 0.05.
- Can you say that setup changes reduce output? Why or why not?
Show the answer
r = -0.989 and r² = 0.978. The test gives p = 0.0002, so the negative correlation is significant: more setup changes go with lower output.
It does not show cause. Both could be driven by a third factor, such as the product mix: complicated orders need more changeovers and also run slower. With only six points the interval for the true correlation is also wide. A designed trial, or a look at output per hour of run time, would test the cause.
Correlation: Frequently Asked Questions
What is a good correlation coefficient?
It depends on the field and on what you need. In many manufacturing studies, r above 0.7 is strong. In a measurement-system comparison, you may need r above 0.99. Look at r², the interval, and whether the relationship is useful for decisions, not at a fixed cutoff.
Does a significant correlation mean the relationship is strong?
No. The p-value only says the true correlation is unlikely to be exactly zero. With a large sample, a tiny correlation of 0.1 can be significant. Read the size of r and the interval.
What is the difference between Pearson and Spearman?
Pearson measures straight-line association between the values. Spearman applies the same calculation to the ranks, so it measures how steadily one variable rises or falls with the other, whatever the shape, and it is less affected by outliers.
What is the difference between correlation and regression?
Correlation measures how tightly two variables move together and treats them alike. Regression builds an equation to predict one (the response) from the other (the predictor), and gives the slope, intercept, and prediction intervals.
How many data points do I need?
At least 10 to 15 to get a useful picture, and more for a narrow interval. The interval for r shrinks with the square root of the sample, so doubling the data does not halve the uncertainty. Plot the data and look at the interval.
Can I correlate more than two variables?
You can calculate a matrix of pairwise correlations, but they do not separate the effect of one variable from another. Use multiple regression for that.
Sources and Further Reading
- NIST/SEMATECH, e-Handbook of Statistical Methods, sections on scatter plots and correlation (itl.nist.gov/div898/handbook).
- David S. Moore, George P. McCabe, and Bruce A. Craig, Introduction to the Practice of Statistics, Freeman, chapter on correlation and regression.
- Douglas C. Montgomery and George C. Runger, Applied Statistics and Probability for Engineers, Wiley.
- Minitab Support, “Methods and formulas for Correlation” (support.minitab.com).
- Microsoft Support, documentation for CORREL, PEARSON, RSQ, and FISHER (support.microsoft.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.