- Question it answers
- Is an average different from a target, from another average, or after a change?
- Data needed
- Measured values: one sample, two independent samples, or paired measurements
- Key output
- t, degrees of freedom, p-value, and the interval for the difference
- Default for two samples
- Welch’s test (no equal-variance assumption)
- Assumptions
- Independent observations, roughly normal data or means, no extreme outliers
- Excel
- Data Analysis > t-Test tools; T.TEST; T.DIST.2T
- Minitab
- Stat > Basic Statistics > 1-Sample t, 2-Sample t, Paired t
- Prerequisite
- P-values and error types
The Idea in Plain Language
The t-test is the workhorse of improvement work. It answers the question: is this average different from a target, or from another average, by more than random scatter would produce? It compares the difference you see with how much that difference would wobble from sample to sample, which is the standard error. The ratio is the t statistic. A large t means the difference is large compared with the noise.
There are three forms, and choosing the right one is most of the skill:
| Form | Question | Data |
|---|---|---|
| One-sample t | Is the average different from a target or standard? | One sample of measurements, and a target value |
| Two-sample t | Are the averages of two separate groups different? | Two independent samples (different units in each) |
| Paired t | Did the same units change? | Two measurements on each unit (before and after, or matched pairs) |
When to Use It
| Your situation | Use | Why |
|---|---|---|
| Compare one average with a target | One-sample t | Tests whether the process is on target |
| Compare two independent averages | Two-sample t (Welch) | Does not assume equal spread |
| Compare before and after on the same units | Paired t | Removes unit-to-unit variation |
| Three or more averages | One-way ANOVA | One test, controlled false-alarm rate |
| Skewed data, small samples | Nonparametric tests | No normality needed |
| Compare variation, not averages | Tests for variances | t-tests compare means only |
| Counts, rates, pass/fail | Tests for proportions | t-tests need measured values |
How It Works
| Test | t statistic | Degrees of freedom | Standard error |
|---|---|---|---|
| One-sample | (x̄ − μ0) / SE | n − 1 | s / √n |
| Two-sample (Welch) | (x̄1 − x̄2) / SE | Satterthwaite: (v1 + v2)² / (v1²/(n1−1) + v2²/(n2−1)), with vi = si²/ni | √(s1²/n1 + s2²/n2) |
| Two-sample (pooled) | (x̄1 − x̄2) / SE | n1 + n2 − 2 | sp √(1/n1 + 1/n2) |
| Paired | d̄ / SE | n − 1 | sd / √n, using the differences |
The confidence interval is estimate ± tcrit × SE. A two-sided test at α = 0.05 rejects the null exactly when the 95% interval excludes the null value. Use Welch’s version by default for two samples: it performs almost as well as the pooled test when the variances are equal, and much better when they are not.
| Assumption | Check | If it fails |
|---|---|---|
| Independent observations (and independent groups, for two-sample) | How the data were collected | Use the paired test, or redesign the study |
| Roughly normal data (or differences, for paired) | Probability plot, histogram; matters most with n under about 15 | Transform, or use a nonparametric test; with n of 30 or more the t-test is robust |
| No extreme outliers | Dot plot or box plot | Find the cause before deciding |
| Equal spread (pooled test only) | Standard deviations, Levene’s test | Use Welch’s test |
Worked Example 1: One-Sample t
A packing step should take 45 seconds. Ten timed cycles gave: 44.1, 47.3, 46.8, 43.9, 48.2, 45.7, 49.1, 46.2, 44.8, 47.5. Is the average different from 45 s?
- Hypotheses. H0: μ = 45. H1: μ ≠ 45. α = 0.05.
- Summary. n = 10, x̄ = 46.36, s = 1.742, SE = 1.742 / √10 = 0.551.
- Test statistic. t = (46.36 − 45) / 0.551 = 2.47 on 9 df.
- p-value = 0.036, and the 95% interval for the mean is 45.11 to 47.61 s.
- Decide. p = 0.036 < 0.05, so reject H0: the mean cycle time is above the 45 s target, by about 1.4 s. The interval 45.1 to 47.6 s excludes 45, but is wide, so the true excess could be anywhere from about 0.1 s to 2.6 s.
Excel: . Minitab: . See the P-Values page for the full output.
Worked Example 2: Two-Sample t (Welch)
Changeover times (minutes) were recorded on two lines. Line A: 12 changeovers. Line B: 10 changeovers. Is the average time different?
| Line | n | Mean | Std dev | Variance of the mean (s²/n) |
|---|---|---|---|---|
| A | 12 | 23.392 | 1.049 | 0.0917 |
| B | 10 | 26.540 | 1.130 | 0.1276 |
- Hypotheses. H0: μA − μB = 0. H1: ≠ 0. α = 0.05.
- Standard error = √(0.0917 + 0.1276) = 0.4683.
- Difference = 23.392 − 26.540 = -3.148 min, so t = -3.148 / 0.4683 = -6.72.
- Degrees of freedom (Satterthwaite) = 18.7. The p-value is < 0.001.
- Interval for the difference: -3.148 ± 2.095 × 0.4683 = -4.13 to -2.17 min.
- Size of the effect: the pooled standard deviation is 1.086, so d = 2.9 standard deviations, a very large difference.
Worked Example 3: Paired t
A new work instruction is meant to shorten handling time. Eight operators were timed before and after, so each operator provides a pair.
| Operator | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | Mean |
|---|---|---|---|---|---|---|---|---|---|
| Before (s) | 42 | 38 | 51 | 45 | 40 | 47 | 36 | 49 | 43.50 |
| After (s) | 39 | 37 | 46 | 44 | 36 | 45 | 35 | 43 | 40.62 |
| Difference (before − after) | +3 | +1 | +5 | +1 | +4 | +2 | +1 | +6 | 2.875 |
- Work with the differences. Mean d̄ = 2.875 s, standard deviation sd = 1.959 s.
- Standard error = 1.959 / √8 = 0.693.
- t = 2.875 / 0.693 = 4.15 on 7 df, p = 0.0043.
- Interval for the mean reduction: 1.24 to 4.51 s.
Run It in Excel and Minitab
ExcelStep by step
- Two-sample: put each line in its own column, then choose . Set the two ranges, Hypothesized Mean Difference to 0, tick Labels, and keep Alpha at 0.05.
- Paired: put before and after in two columns and choose .
- Read t Stat, P(T<=t) two-tail, and t Critical two-tail. Use the two-tail p-value unless you chose a direction in advance.
- Quick p-value without the tool: (Welch, two-sided) or (paired). The last number is the type: 1 paired, 2 pooled, 3 Welch.
- One-sample: use the formula in Example 1.
- Interval: difference ± × SE.
Excel’s tool rounds the degrees of freedom down and gives no confidence interval for the difference, so compute it from the output.
MinitabStep by step
- Two-sample: stack the data in one Weight column and one Line column, then choose . Choose Both samples are in one column (or each in its own column).
- Click Options. Leave Assume equal variances cleared (Welch), set the confidence level, and choose the alternative (not equal).
- Click Graphs to tick Individual value plot and Boxplot.
- Paired: choose and pick the Before and After columns.
- One-sample: .
- Read the interval for the difference, the T-Value, DF, and P-Value. Plan the sample with .
The Assistant () adds a power check, an unusual-data check, and a plain-language summary.
t-Test: Two-Sample Assuming Unequal Variances
Line A Line B
Mean 23.3917 26.5400
Variance 1.1008 1.2760
Observations 12 10
Hypothesized Mean Difference 0
df 18
t Stat -6.7224
P(T<=t) one-tail 1.33e-06
t Critical one-tail 1.7341
P(T<=t) two-tail 2.66e-06
t Critical two-tail 2.1009Method
μ₁: mean of Line A
μ₂: mean of Line B
Difference: μ₁ - μ₂
Equal variances are not assumed for this analysis.
Descriptive Statistics
Sample N Mean StDev SE Mean
Line A 12 23.392 1.049 0.303
Line B 10 26.540 1.130 0.357
Estimation for Difference
95% CI for
Difference Difference
-3.148 (-4.130, -2.167)
Test
Null hypothesis H₀: μ₁ - μ₂ = 0
Alternative hypothesis H₁: μ₁ - μ₂ ≠ 0
T-Value DF P-Value
-6.72 19 0.000Descriptive Statistics
Sample N Mean StDev SE Mean
Before 8 43.500 5.372 1.899
After 8 40.625 4.373 1.546
Estimation for Paired Difference
Mean StDev SE Mean 95% CI for
μ_difference
2.875 1.959 0.693 (1.237, 4.513)
Test
Null hypothesis H₀: μ_difference = 0
Alternative hypothesis H₁: μ_difference ≠ 0
T-Value P-Value
4.15 0.004Reading and Reporting
- Report the means, the difference, and its interval. The interval is the most useful single result.
- Give t, the degrees of freedom, and the exact p-value.
- Name the test: one-sample, Welch two-sample, or paired.
- Judge importance against the tolerance or the cost, and mention the effect size.
- State the sample sizes and how the data were collected.
Common Mistakes
| Mistake | Why it misleads | Better |
|---|---|---|
| Using the two-sample test on paired data | Unit-to-unit variation hides the effect | Pair the data |
| Using the paired test on unrelated groups | There is no natural pairing | Use the two-sample test |
| Assuming equal variances without a reason | Pooled p-values are wrong when spreads differ and group sizes differ | Use Welch |
| Running many t-tests on several groups | False alarms accumulate | ANOVA first, then a post-hoc test |
| Reading “not significant” as “equal” | A small sample may miss a real difference | Look at the interval; check the power |
| Ignoring outliers or strong skew in small samples | The t statistic is distorted | Plot first; consider a nonparametric test |
| Testing several measures and reporting the significant one | Inflated false alarms | Decide the primary measure first |
Try It Yourself
Two suppliers delivered coatings. Thickness (microns) from 6 samples each: Supplier P: 52, 55, 51, 54, 53, 56. Supplier Q: 58, 57, 60, 56, 59, 61.
- Which test applies? Calculate the t statistic and the p-value.
- Give a 95% interval for the difference.
Show the answer
The groups are independent, so use the two-sample (Welch) test. Means: P = 53.50, Q = 58.50; standard deviations 1.87 and 1.87. t = -4.63, df = 10.0, p = 0.0009, so Q is thicker.
The 95% interval for P − Q is -7.41 to -2.59 microns. The coating from Q is about 5.0 microns thicker, which you would compare with the specification.
One- and Two-Sample t-Tests: Frequently Asked Questions
Should I use the pooled or the Welch t-test?
Welch’s test, by default. It does not assume equal variances, and when the variances are equal it gives nearly the same answer as the pooled test. The pooled test is only better when you are sure the variances are equal and the samples are small.
How do I know if my data are paired?
Ask whether each value in one group has a natural partner in the other: the same unit measured twice, or a matched pair chosen on purpose. If you could shuffle the order of the rows in one group without losing information, the data are not paired.
How normal does my data need to be?
The t-test needs the sampling distribution of the average to be roughly normal. With 30 or more observations, or for symmetric data with at least 10 to 15, it is robust. For small, skewed samples, transform the data or use a nonparametric test.
What is the difference between a one-sided and a two-sided test?
A two-sided test looks for a difference in either direction. A one-sided test looks in one direction only, and has more power there. Use one-sided only if you decided in advance that the other direction is irrelevant.
How many observations do I need?
Enough to detect the smallest difference that matters. For two means, about 16 per group detects a difference of one standard deviation with 80% power. Use the Sample Size and Power page or Minitab’s Power and Sample Size menu.
What is the effect size for a t-test?
Cohen’s d is the difference between means divided by the pooled standard deviation. About 0.2 is small, 0.5 medium, and 0.8 large, but judge importance against your tolerance and costs.
Sources and Further Reading
- Douglas C. Montgomery and George C. Runger, Applied Statistics and Probability for Engineers, Wiley, chapters on inference for means.
- NIST/SEMATECH, e-Handbook of Statistical Methods, sections on t-tests (itl.nist.gov/div898/handbook).
- B. L. Welch, “The generalization of Student’s problem when several different population variances are involved,” Biometrika, 1947.
- Minitab Support, “Methods and formulas for 2-Sample t and Paired t” (support.minitab.com).
- Microsoft Support, documentation for the Analysis ToolPak t-tests and the T.TEST function (support.microsoft.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.