Question it answers
Is an average different from a target, from another average, or after a change?
Data needed
Measured values: one sample, two independent samples, or paired measurements
Key output
t, degrees of freedom, p-value, and the interval for the difference
Default for two samples
Welch’s test (no equal-variance assumption)
Assumptions
Independent observations, roughly normal data or means, no extreme outliers
Excel
Data Analysis > t-Test tools; T.TEST; T.DIST.2T
Minitab
Stat > Basic Statistics > 1-Sample t, 2-Sample t, Paired t
Prerequisite
P-values and error types

The Idea in Plain Language

The t-test is the workhorse of improvement work. It answers the question: is this average different from a target, or from another average, by more than random scatter would produce? It compares the difference you see with how much that difference would wobble from sample to sample, which is the standard error. The ratio is the t statistic. A large t means the difference is large compared with the noise.

There are three forms, and choosing the right one is most of the skill:

FormQuestionData
One-sample tIs the average different from a target or standard?One sample of measurements, and a target value
Two-sample tAre the averages of two separate groups different?Two independent samples (different units in each)
Paired tDid the same units change?Two measurements on each unit (before and after, or matched pairs)
Paired or two-sample? If each measurement in one group has a natural partner in the other (the same operator, part, or machine), use the paired test. It removes the variation between units and is usually far more powerful. If the groups are unrelated, use the two-sample test. Using the wrong one is a common error.

When to Use It

Your situationUseWhy
Compare one average with a targetOne-sample tTests whether the process is on target
Compare two independent averagesTwo-sample t (Welch)Does not assume equal spread
Compare before and after on the same unitsPaired tRemoves unit-to-unit variation
Three or more averagesOne-way ANOVAOne test, controlled false-alarm rate
Skewed data, small samplesNonparametric testsNo normality needed
Compare variation, not averagesTests for variancest-tests compare means only
Counts, rates, pass/failTests for proportionst-tests need measured values

How It Works

Testt statisticDegrees of freedomStandard error
One-sample(x̄ − μ0) / SEn − 1s / √n
Two-sample (Welch)(x̄1 − x̄2) / SESatterthwaite: (v1 + v2)² / (v1²/(n1−1) + v2²/(n2−1)), with vi = si²/ni√(s1²/n1 + s2²/n2)
Two-sample (pooled)(x̄1 − x̄2) / SEn1 + n2 − 2sp √(1/n1 + 1/n2)
Pairedd̄ / SEn − 1sd / √n, using the differences

The confidence interval is estimate ± tcrit × SE. A two-sided test at α = 0.05 rejects the null exactly when the 95% interval excludes the null value. Use Welch’s version by default for two samples: it performs almost as well as the pooled test when the variances are equal, and much better when they are not.

AssumptionCheckIf it fails
Independent observations (and independent groups, for two-sample)How the data were collectedUse the paired test, or redesign the study
Roughly normal data (or differences, for paired)Probability plot, histogram; matters most with n under about 15Transform, or use a nonparametric test; with n of 30 or more the t-test is robust
No extreme outliersDot plot or box plotFind the cause before deciding
Equal spread (pooled test only)Standard deviations, Levene’s testUse Welch’s test

Worked Example 1: One-Sample t

A packing step should take 45 seconds. Ten timed cycles gave: 44.1, 47.3, 46.8, 43.9, 48.2, 45.7, 49.1, 46.2, 44.8, 47.5. Is the average different from 45 s?

  1. Hypotheses. H0: μ = 45. H1: μ ≠ 45. α = 0.05.
  2. Summary. n = 10, x̄ = 46.36, s = 1.742, SE = 1.742 / √10 = 0.551.
  3. Test statistic. t = (46.36 − 45) / 0.551 = 2.47 on 9 df.
  4. p-value = 0.036, and the 95% interval for the mean is 45.11 to 47.61 s.
  5. Decide. p = 0.036 < 0.05, so reject H0: the mean cycle time is above the 45 s target, by about 1.4 s. The interval 45.1 to 47.6 s excludes 45, but is wide, so the true excess could be anywhere from about 0.1 s to 2.6 s.

Excel: =T.DIST.2T(ABS((AVERAGE(A2:A11)-45)/(STDEV.S(A2:A11)/SQRT(10))),9). Minitab: Stat > Basic Statistics > 1-Sample t. See the P-Values page for the full output.

Worked Example 2: Two-Sample t (Welch)

Changeover times (minutes) were recorded on two lines. Line A: 12 changeovers. Line B: 10 changeovers. Is the average time different?

20 min 22 min 24 min 26 min 28 min 30 min 23.39 Line A 26.54 Line B Overall mean 24.82
Changeover times by line. Line B is slower, and the spreads are similar.
LinenMeanStd devVariance of the mean (s²/n)
A1223.3921.0490.0917
B1026.5401.1300.1276
  1. Hypotheses. H0: μA − μB = 0. H1: ≠ 0. α = 0.05.
  2. Standard error = √(0.0917 + 0.1276) = 0.4683.
  3. Difference = 23.392 − 26.540 = -3.148 min, so t = -3.148 / 0.4683 = -6.72.
  4. Degrees of freedom (Satterthwaite) = 18.7. The p-value is < 0.001.
  5. Interval for the difference: -3.148 ± 2.095 × 0.4683 = -4.13 to -2.17 min.
  6. Size of the effect: the pooled standard deviation is 1.086, so d = 2.9 standard deviations, a very large difference.
-6 -4 -2 0 2 4 6 Observed t = -6.72 Critical t = ±2.10 t statistic
The observed t is far beyond the critical values, so a gap this large between two lines that were really the same would be very unlikely.
Conclusion. Line A changeovers were 3.1 minutes shorter on average than Line B (95% CI 2.2 to 4.1 min; Welch t(18.7) = -6.72, p < 0.001). The pooled test gives t = -6.77, p < 0.001: here the two agree because the standard deviations are similar.

Worked Example 3: Paired t

A new work instruction is meant to shorten handling time. Eight operators were timed before and after, so each operator provides a pair.

30 35 40 45 50 55 Before (mean 43.5)After (mean 40.6) Handling time (s)
Each line is one operator. Every line falls, by 1 to 6 seconds. The consistent drop in each pair is what the paired test uses.
Operator12345678Mean
Before (s)423851454047364943.50
After (s)393746443645354340.62
Difference (before − after)+3+1+5+1+4+2+1+62.875
  1. Work with the differences. Mean d̄ = 2.875 s, standard deviation sd = 1.959 s.
  2. Standard error = 1.959 / √8 = 0.693.
  3. t = 2.875 / 0.693 = 4.15 on 7 df, p = 0.0043.
  4. Interval for the mean reduction: 1.24 to 4.51 s.
Why pairing matters. Treating the same data as two independent samples gives p = 0.26, which would miss the improvement, because the operators differ a lot from one another (36 to 51 s) and that spread swamps the change. Pairing compares each operator with themselves.

Run It in Excel and Minitab

ExcelStep by step

  1. Two-sample: put each line in its own column, then choose Data > Data Analysis > t-Test: Two-Sample Assuming Unequal Variances. Set the two ranges, Hypothesized Mean Difference to 0, tick Labels, and keep Alpha at 0.05.
  2. Paired: put before and after in two columns and choose Data > Data Analysis > t-Test: Paired Two Sample for Means.
  3. Read t Stat, P(T<=t) two-tail, and t Critical two-tail. Use the two-tail p-value unless you chose a direction in advance.
  4. Quick p-value without the tool: =T.TEST(range1, range2, 2, 3) (Welch, two-sided) or =T.TEST(range1, range2, 2, 1) (paired). The last number is the type: 1 paired, 2 pooled, 3 Welch.
  5. One-sample: use the formula in Example 1.
  6. Interval: difference ± =T.INV.2T(0.05, df) × SE.

Excel’s tool rounds the degrees of freedom down and gives no confidence interval for the difference, so compute it from the output.

MinitabStep by step

  1. Two-sample: stack the data in one Weight column and one Line column, then choose Stat > Basic Statistics > 2-Sample t. Choose Both samples are in one column (or each in its own column).
  2. Click Options. Leave Assume equal variances cleared (Welch), set the confidence level, and choose the alternative (not equal).
  3. Click Graphs to tick Individual value plot and Boxplot.
  4. Paired: choose Stat > Basic Statistics > Paired t and pick the Before and After columns.
  5. One-sample: Stat > Basic Statistics > 1-Sample t.
  6. Read the interval for the difference, the T-Value, DF, and P-Value. Plan the sample with Stat > Power and Sample Size.

The Assistant (Assistant > Hypothesis Tests > 2-Sample t) adds a power check, an unusual-data check, and a plain-language summary.

Excel ToolPak output: t-Test Two-Sample Assuming Unequal Variances
t-Test: Two-Sample Assuming Unequal Variances

                                    Line A    Line B
Mean                               23.3917   26.5400
Variance                            1.1008    1.2760
Observations                            12        10
Hypothesized Mean Difference             0
df                                      18
t Stat                             -6.7224
P(T<=t) one-tail                  1.33e-06
t Critical one-tail                 1.7341
P(T<=t) two-tail                  2.66e-06
t Critical two-tail                 2.1009
Minitab session window: 2-Sample t (typed excerpt, simplified)
Method

μ₁: mean of Line A
μ₂: mean of Line B
Difference: μ₁ - μ₂

Equal variances are not assumed for this analysis.

Descriptive Statistics

Sample   N   Mean  StDev  SE Mean
Line A  12  23.392  1.049    0.303
Line B  10  26.540  1.130    0.357

Estimation for Difference

          95% CI for
Difference  Difference
  -3.148  (-4.130, -2.167)

Test

Null hypothesis         H₀: μ₁ - μ₂ = 0
Alternative hypothesis  H₁: μ₁ - μ₂ ≠ 0

T-Value  DF  P-Value
  -6.72  19    0.000
Minitab session window: Paired t (typed excerpt, simplified)
Descriptive Statistics

Sample   N   Mean  StDev  SE Mean
Before   8  43.500  5.372    1.899
After    8  40.625  4.373    1.546

Estimation for Paired Difference

  Mean  StDev  SE Mean        95% CI for
                         μ_difference
2.875  1.959    0.693  (1.237, 4.513)

Test

Null hypothesis         H₀: μ_difference = 0
Alternative hypothesis  H₁: μ_difference ≠ 0

T-Value  P-Value
   4.15    0.004

Reading and Reporting

  1. Report the means, the difference, and its interval. The interval is the most useful single result.
  2. Give t, the degrees of freedom, and the exact p-value.
  3. Name the test: one-sample, Welch two-sample, or paired.
  4. Judge importance against the tolerance or the cost, and mention the effect size.
  5. State the sample sizes and how the data were collected.
A sentence you can use. Mean changeover time was 23.4 min on Line A (n = 12, SD = 1.05) and 26.5 min on Line B (n = 10, SD = 1.13); the difference of -3.1 min (95% CI -4.1 to -2.2) was significant (Welch t(18.7) = -6.72, p < 0.001).

Common Mistakes

MistakeWhy it misleadsBetter
Using the two-sample test on paired dataUnit-to-unit variation hides the effectPair the data
Using the paired test on unrelated groupsThere is no natural pairingUse the two-sample test
Assuming equal variances without a reasonPooled p-values are wrong when spreads differ and group sizes differUse Welch
Running many t-tests on several groupsFalse alarms accumulateANOVA first, then a post-hoc test
Reading “not significant” as “equal”A small sample may miss a real differenceLook at the interval; check the power
Ignoring outliers or strong skew in small samplesThe t statistic is distortedPlot first; consider a nonparametric test
Testing several measures and reporting the significant oneInflated false alarmsDecide the primary measure first

Try It Yourself

Two suppliers delivered coatings. Thickness (microns) from 6 samples each: Supplier P: 52, 55, 51, 54, 53, 56. Supplier Q: 58, 57, 60, 56, 59, 61.

  • Which test applies? Calculate the t statistic and the p-value.
  • Give a 95% interval for the difference.
Show the answer

The groups are independent, so use the two-sample (Welch) test. Means: P = 53.50, Q = 58.50; standard deviations 1.87 and 1.87. t = -4.63, df = 10.0, p = 0.0009, so Q is thicker.

The 95% interval for P − Q is -7.41 to -2.59 microns. The coating from Q is about 5.0 microns thicker, which you would compare with the specification.

One- and Two-Sample t-Tests: Frequently Asked Questions

Should I use the pooled or the Welch t-test?

Welch’s test, by default. It does not assume equal variances, and when the variances are equal it gives nearly the same answer as the pooled test. The pooled test is only better when you are sure the variances are equal and the samples are small.

How do I know if my data are paired?

Ask whether each value in one group has a natural partner in the other: the same unit measured twice, or a matched pair chosen on purpose. If you could shuffle the order of the rows in one group without losing information, the data are not paired.

How normal does my data need to be?

The t-test needs the sampling distribution of the average to be roughly normal. With 30 or more observations, or for symmetric data with at least 10 to 15, it is robust. For small, skewed samples, transform the data or use a nonparametric test.

What is the difference between a one-sided and a two-sided test?

A two-sided test looks for a difference in either direction. A one-sided test looks in one direction only, and has more power there. Use one-sided only if you decided in advance that the other direction is irrelevant.

How many observations do I need?

Enough to detect the smallest difference that matters. For two means, about 16 per group detects a difference of one standard deviation with 80% power. Use the Sample Size and Power page or Minitab’s Power and Sample Size menu.

What is the effect size for a t-test?

Cohen’s d is the difference between means divided by the pooled standard deviation. About 0.2 is small, 0.5 medium, and 0.8 large, but judge importance against your tolerance and costs.

Sources and Further Reading

  • Douglas C. Montgomery and George C. Runger, Applied Statistics and Probability for Engineers, Wiley, chapters on inference for means.
  • NIST/SEMATECH, e-Handbook of Statistical Methods, sections on t-tests (itl.nist.gov/div898/handbook).
  • B. L. Welch, “The generalization of Student’s problem when several different population variances are involved,” Biometrika, 1947.
  • Minitab Support, “Methods and formulas for 2-Sample t and Paired t” (support.minitab.com).
  • Microsoft Support, documentation for the Analysis ToolPak t-tests and the T.TEST function (support.microsoft.com).

This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.