Question it answers
Do two measured variables move together in a straight-line pattern, and how tightly?
Data needed
Two continuous variables measured on the same units
Key output
r from −1 to +1, its confidence interval, and r²
Null hypothesis
The true correlation is zero
Plot first
A scatter plot, to see curves and outliers
Excel
=CORREL, =RSQ, and Data Analysis > Correlation
Minitab
Stat > Basic Statistics > Correlation
Next step
Simple linear regression to predict one from the other

The Idea in Plain Language

When the oven runs hotter, does the adhesive bond get stronger? When the line speeds up, do defects rise? Correlation answers a narrow question: do two measured variables move together in a straight-line pattern, and how tightly?

The most common measure is the Pearson correlation coefficient, r. It runs from −1 to +1. A value near +1 means that high values of one variable go with high values of the other. Near −1 means high goes with low. Near 0 means there is no straight-line relationship. The square of r, r², is the share of the variation in one variable that moves along with the other.

Strong positive: r = +0.92 Moderate positive: r = +0.55 No relationship: r = +0.00 Strong negative: r = -0.85 Curved (U shape): r = +0.00 One outlier: r = +0.42
Six patterns and their r values. A coefficient summarizes a straight-line pattern only. The U-shaped cloud is strongly related but has r near zero, and one outlier can create a correlation that is not there.
Correlation is not causation. Two things can move together because one causes the other, because the other causes the first, because a third factor drives both, or by chance. A correlation is a reason to investigate, not a conclusion.

When to Use Correlation

Your situationUseWhy
Two continuous variables, a roughly straight-line pattern, no strong outliersPearson correlationMeasures the strength of the linear pattern
A curved but always rising or falling pattern, ranked data, or outliersSpearman rank correlationUses ranks, so it measures any steady rise or fall and resists outliers
You want to predict one variable from the otherSimple linear regressionGives the equation, the slope, and prediction intervals
Several inputsMultiple regressionSeparates the effect of each input
One variable is a category (machine, shift)One-way ANOVACompares group averages
Both variables are categoriesChi-square test of associationTests whether the categories are related
Always plot first. Draw the scatter plot before you calculate anything. The plot shows curves, outliers, clusters, and restricted ranges that a single number hides.

How It Works

Each point is described by how far it sits from the mean temperature and from the mean strength. Points in the lower-left and upper-right quadrants (both below, or both above, the means) pull the correlation up. Points in the other two quadrants pull it down.

15 20 25 30 35 120 140 160 180 mean x = 147.5 mean y = 25.6 Curing temperature (°C) Bond strength (N)
The 12 bonds. Nearly all the points sit in the lower-left and upper-right quadrants formed by the two means, so the correlation is strongly positive.
QuantityFormulaIn words
SxxΣ (x − x̄)²Variation in x around its mean
SyyΣ (y − ȳ)²Variation in y around its mean
SxyΣ (x − x̄)(y − ȳ)How x and y vary together
Correlation rSxy / √(Sxx Syy)Together-variation scaled to run from −1 to +1
Test statistict = r √(n − 2) / √(1 − r²), df = n − 2Tests H0: the true correlation is zero
Confidence intervaltanh( arctanh(r) ± 1.96 / √(n − 3) )Fisher’s z transform gives a range for the true correlation
|r|Rough descriptionr²
0.00 to 0.30Weak or noneUnder 9%
0.30 to 0.70Moderate9% to 49%
0.70 to 0.90Strong49% to 81%
0.90 to 1.00Very strongOver 81%

These labels are conventions, not rules. In a measurement-system study, r = 0.9 is poor agreement. In a field study of people, r = 0.5 can be important. Judge r against what matters in your setting.

Assumptions and Pitfalls

IssueWhat goes wrongWhat to do
A curved relationshipr can be near zero even when the relationship is strongPlot first; transform, fit a curve, or use regression with a squared term
OutliersOne extreme point can create or destroy a correlationFind the cause; report r with and without it; use Spearman
Restricted rangeA narrow range of x shrinks rDo not compare r values from studies with different ranges
Mixed groupsTwo groups with different levels can create a correlation that exists within neitherPlot by group; stratify by machine, shift, or material
Lurking variableA third variable drives both, such as production volumeThink about what else changes together; use a designed experiment to test cause
Non-independent observationsTime series data trend together and give high r that means littleLook at changes instead of levels, or use time-series methods
Many correlations testedTest enough pairs and some will be “significant” by chanceDecide the pairs in advance; adjust for the number of tests

For the significance test and the confidence interval, the standard method assumes that the two variables together follow a roughly bivariate normal distribution. With small samples, check the scatter plot for outliers and skew.

Worked Example by Hand

An adhesive is cured at different temperatures, and the bond strength is measured on 12 test coupons. The question is whether higher curing temperature goes with higher strength.

15 20 25 30 35 120 140 160 180 Curing temperature (°C) Bond strength (N)
Bond strength against curing temperature. The points rise from the lower left to the upper right.
Working table: curing temperature (x) and bond strength (y)
Bondx (°C)y (N)x − x̄y − ȳ(x − x̄)²(y − ȳ)²(x − x̄)(y − ȳ)
112021.2-27.50-4.38756.2519.21+120.54
212519.4-22.50-6.18506.2538.23+139.12
313022.7-17.50-2.88306.258.31+50.46
413524.9-12.50-0.68156.250.47+8.54
514022.9-7.50-2.6856.257.20+20.12
614523.0-2.50-2.586.256.67+6.46
715026.9+2.50+1.326.251.73+3.29
815528.6+7.50+3.0256.259.10+22.63
916026.7+12.50+1.12156.251.25+13.96
1016529.5+17.50+3.92306.2515.34+68.54
1117029.2+22.50+3.62506.2513.08+81.38
1217532.0+27.50+6.42756.2541.17+176.46
Sum3575.00161.78+711.50
  1. Means. x̄ = 147.50 °C and ȳ = 25.583 N.
  2. Sums of squares and products from the table: Sxx = 3575.00, Syy = 161.777, Sxy = 711.50.
  3. Correlation. r = 711.50 / √(3575.00 × 161.777) = 0.936. Then r² = 0.875, so 87.5% of the variation in strength goes along with temperature.
  4. Test whether the true correlation is zero. t = 0.936 × √(12 − 2) / √(1 − 0.875) = 8.38 on 10 degrees of freedom, so p = 0.000008.
  5. Confidence interval. arctanh(0.936) = 1.701, standard error 1 / √(12 − 3) = 0.333. Back-transforming 1.701 ± 1.96 × 0.333 gives a 95% interval of 0.78 to 0.98.
  6. Rank version. Spearman’s rho, the correlation of the ranks, is 0.944 (p = 0.000004), which agrees because the pattern is a steady rise.
Conclusion. Curing temperature and bond strength have a strong positive linear correlation (r = 0.94, 95% CI 0.78 to 0.98, p < 0.001, n = 12). The correlation does not by itself show that temperature causes the change, but it is a strong reason to study the relationship with a regression and, if possible, a designed experiment.

Run It in Excel and Minitab

ExcelStep by step

  1. Put the temperatures in A2:A13 and the strengths in B2:B13, with headings in row 1.
  2. =CORREL(A2:A13,B2:B13) returns r (0.9356). =RSQ(B2:B13,A2:A13) returns r².
  3. For the p-value: =T.DIST.2T(ABS(r)*SQRT((n-2)/(1-r^2)), n-2) with r in a cell and n = 12.
  4. For the 95% interval: =FISHERINV(FISHER(r)-NORM.S.INV(0.975)/SQRT(n-3)) for the lower limit, and plus for the upper limit.
  5. For several variables at once, choose Data > Data Analysis > Correlation to get a matrix of r values.
  6. Spearman: rank each column with =RANK.AVG(A2,$A$2:$A$13,1), then apply CORREL to the two rank columns.
  7. Draw the scatter plot with Insert > Scatter and add a trendline if you want to see the line.

Excel’s CORREL gives only r. It does not give a p-value or an interval, so use the formulas above.

MinitabStep by step

  1. Enter Temp and Strength in two columns.
  2. Draw the scatter plot first: Graph > Scatterplot > Simple (Y variable: Strength; X variable: Temp).
  3. Choose Stat > Basic Statistics > Correlation and select both variables.
  4. Under Options, choose the method (Pearson or Spearman) and keep the confidence interval ticked.
  5. Click OK and read the correlation, the interval, and the p-value in the session window.
  6. For many variables, select all of them to get a matrix; Graph > Matrix Plot draws the scatter plots together.

Option names can vary slightly between Minitab versions. The Assistant (Assistant > Regression) also reports the correlation with a diagnostic report card.

What the output looks like

Excel formula results
=CORREL(A2:A13,B2:B13)           0.935576
=RSQ(B2:B13,A2:A13)             0.875302
=T.DIST.2T(t, 10), with t = 8.3782      7.84e-06
Minitab session window (typed excerpt, simplified)
Method

Correlation type         Pearson
Number of rows used           12

Correlations

Pearson correlation       95% CI for ρ       P-Value
0.936             (0.781, 0.982)      0.000

Reading and Reporting the Result

  1. Look at the plot, then at r. Here r = 0.94: strong and positive.
  2. Use the interval. The 95% interval 0.78 to 0.98 shows how well the sample pins down the true correlation. With only 12 points the interval is wide, even though r is high.
  3. The p-value is a yes/no on “is it zero?”. It does not say the correlation is useful. A large sample makes a trivial r significant.
  4. Translate r². 88% of the variation in strength goes along with temperature. The other 12% comes from other things.
  5. State the limits. Report the range of x that was studied. Do not claim a relationship outside it.
A sentence you can use. Bond strength was strongly and positively correlated with curing temperature (Pearson r = 0.94, 95% CI 0.78 to 0.98, p < 0.001, n = 12, temperatures 120 to 175 °C).

When the Usual Method Does Not Fit

ProblemBetter approach
Curved relationshipPlot; transform one variable (log, square root), or fit a curve in regression
Outliers or skewed dataSpearman rank correlation, and report what happens with and without the outliers
Ordered categories (ratings 1 to 5)Spearman rank correlation
Groups mixed togetherCalculate r within each group, or fit regression with the group as a factor
Time-ordered data that trendCompare changes from period to period, or use time-series methods
Need to show causeA designed experiment, with the factor set on purpose and runs randomized; see the DOE guide

Common Mistakes

MistakeWhy it misleadsBetter
Calculating r without plottingCurves, clusters, and outliers are invisiblePlot first, every time
Saying “X causes Y” from rCorrelation does not identify causeSay “is associated with”, and test cause with an experiment
Treating r = 0 as “no relationship”It means no linear relationshipLook for curves
Reporting only the p-valueSays nothing about strengthGive r, the interval, and n
Using the correlation of averagesAveraging hides the scatter and inflates rUse the individual data
Extrapolating beyond the dataThe pattern may not continueState the range studied
Comparing r across different rangesRange changes rCompare slopes, or use the same range

Try It Yourself

The number of setup changes per week (x) and the weekly output in thousands of units (y) were recorded for six weeks.

WeekSetup changes (x)Output (y)
1230
2428
3524
4721
5918
61112
  • Calculate the correlation coefficient and decide whether it is significant at α = 0.05.
  • Can you say that setup changes reduce output? Why or why not?
Show the answer

r = -0.989 and r² = 0.978. The test gives p = 0.0002, so the negative correlation is significant: more setup changes go with lower output.

It does not show cause. Both could be driven by a third factor, such as the product mix: complicated orders need more changeovers and also run slower. With only six points the interval for the true correlation is also wide. A designed trial, or a look at output per hour of run time, would test the cause.

Correlation: Frequently Asked Questions

What is a good correlation coefficient?

It depends on the field and on what you need. In many manufacturing studies, r above 0.7 is strong. In a measurement-system comparison, you may need r above 0.99. Look at r², the interval, and whether the relationship is useful for decisions, not at a fixed cutoff.

Does a significant correlation mean the relationship is strong?

No. The p-value only says the true correlation is unlikely to be exactly zero. With a large sample, a tiny correlation of 0.1 can be significant. Read the size of r and the interval.

What is the difference between Pearson and Spearman?

Pearson measures straight-line association between the values. Spearman applies the same calculation to the ranks, so it measures how steadily one variable rises or falls with the other, whatever the shape, and it is less affected by outliers.

What is the difference between correlation and regression?

Correlation measures how tightly two variables move together and treats them alike. Regression builds an equation to predict one (the response) from the other (the predictor), and gives the slope, intercept, and prediction intervals.

How many data points do I need?

At least 10 to 15 to get a useful picture, and more for a narrow interval. The interval for r shrinks with the square root of the sample, so doubling the data does not halve the uncertainty. Plot the data and look at the interval.

Can I correlate more than two variables?

You can calculate a matrix of pairwise correlations, but they do not separate the effect of one variable from another. Use multiple regression for that.

Sources and Further Reading

  • NIST/SEMATECH, e-Handbook of Statistical Methods, sections on scatter plots and correlation (itl.nist.gov/div898/handbook).
  • David S. Moore, George P. McCabe, and Bruce A. Craig, Introduction to the Practice of Statistics, Freeman, chapter on correlation and regression.
  • Douglas C. Montgomery and George C. Runger, Applied Statistics and Probability for Engineers, Wiley.
  • Minitab Support, “Methods and formulas for Correlation” (support.minitab.com).
  • Microsoft Support, documentation for CORREL, PEARSON, RSQ, and FISHER (support.microsoft.com).

This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.