Question it answers
Is what I see real, or could chance explain it?
Null hypothesis
No change, no difference, or on target
Key output
Test statistic, p-value, and a confidence interval
You set in advance
The significance level (α), the direction, and the sample size
Two errors
Type I (false alarm, α) and Type II (missed signal, β)
Excel
T.DIST.2T, T.INV.2T, NORM.S.DIST formulas
Minitab
Stat > Basic Statistics > 1-Sample t; Power and Sample Size
Next step
Sample size and power, then the specific tests

The Idea in Plain Language

A statistical test is a structured way to decide whether what you see in a sample is a real feature of the process or just the luck of the draw. You start with a skeptical position, the null hypothesis (H0): nothing has changed, there is no difference, the process is on target. The alternative hypothesis (H1) says there is. Then you ask how surprising your data would be if the null hypothesis were true.

The answer to that question is the p-value: the probability of getting a result at least as extreme as the one you observed, if the null hypothesis were true. A small p-value says the data would be surprising under the null, so the null becomes hard to believe. It is the courtroom logic: the defendant is presumed innocent, and you convict only if the evidence would be very unlikely for an innocent person.

What a p-value is not. It is not the probability that the null hypothesis is true, and it is not the probability that you made a mistake. It is a statement about the data, assuming the null. Section 5 lists the most common misreadings.

The Steps of Every Test

  1. State the question in practical terms, and the hypotheses: H0 (no change) and H1 (change, or change in a stated direction).
  2. Choose the risk you will accept. The significance level α (usually 0.05) is the chance of a false alarm you will tolerate.
  3. Decide the sample size before collecting data, so the test can see the difference that matters (power).
  4. Collect the data in a way that makes the observations independent and representative.
  5. Calculate the test statistic and the p-value. The statistic measures how far the sample is from the null, in standard errors.
  6. Decide and report. If p ≤ α, reject H0; otherwise, fail to reject it. Report the size of the effect with a confidence interval, not only the verdict.
“Fail to reject” is not “accept”. A non-significant result means the data did not provide enough evidence of a change. It does not prove there is none, especially with a small sample.

Two Ways to Be Wrong: Type I and Type II Errors

Null hypothesis is really true (no change)Null hypothesis is really false (a real change)
You reject H0 (call it a difference)Type I error: a false alarm. Probability = αCorrect: you found the real change. Probability = power (1 − β)
You fail to reject H0Correct. Probability = 1 − αType II error: a missed signal. Probability = β
Critical value H0 trueH1 true (a real shift) Type I error (alpha): a false alarm Type II error (beta): a missed shift Power (1 - beta): the shift is detected Test statistic (in standard errors of the mean)
If the process has really shifted (right curve), the test detects it whenever the statistic lands beyond the critical value. The red tail is the false-alarm risk when nothing has changed. The gold area is the risk of missing a real shift.
To getYou canBut
Fewer false alarms (smaller α)Use a stricter level, such as 0.01Power falls, so you miss more real changes
More power (smaller β)Collect more data, reduce measurement noise, or look for a larger effectMore data costs time and money
A real effect detected reliablyPlan the sample size in advanceYou must decide what size of effect matters

Both errors cost something, and the costs differ. Missing a real defect-rate increase in a safety part is worse than a false alarm, so you would accept a larger α and demand more power. Raising prices on a false signal may be worse, so you would use a smaller α. Choose the risks before you see the data. See Sample Size and Power.

Worked Example: Is the Filling Line on Target?

A line should fill 500 g. A sample of 25 fills has a mean of 499.2 g and a standard deviation of 1.8 g. Is the line off target?

  1. Hypotheses. H0: μ = 500 g. H1: μ ≠ 500 g (two-sided, since a shift either way matters). α = 0.05.
  2. Standard error. SE = s / √n = 1.8 / √25 = 0.360 g.
  3. Test statistic. t = (x̄ − μ0) / SE = (499.2 − 500) / 0.360 = -2.22, with 24 degrees of freedom.
  4. p-value. The chance of a |t| at least 2.22 when the mean is really 500 g is 0.036, the shaded area below.
  5. Decide. 0.036 < 0.05, so reject H0. The 95% confidence interval for the mean is 498.46 to 499.94 g, which excludes 500.
-4 -3 -2 -1 0 1 2 3 4 Observed t = -2.22 t statistic (p = 0.036 is the shaded area, both tails)
The p-value is the area in the tails beyond the observed statistic. Here it is about 3.6%, so a sample this far from 500 g would be unusual if the line were on target.
Conclusion. The mean fill weight is 499.2 g, which is 0.8 g below target (95% CI 498.46 to 499.94 g; t(24) = -2.22, p = 0.036). The shortfall is small (0.44 standard deviations), so the practical question is whether 0.8 g matters for your tolerance and cost.

The same data with different sample sizes

The p-value depends on how much data you have as well as how big the difference is. Keep the same mean (499.2 g) and standard deviation (1.8 g), and change only the sample size:

Sample sizeStandard errortp-valueDecision at 0.05
90.600-1.330.2191Fail to reject
250.360-2.220.0359Reject
1000.180-4.44< 0.0001Reject

The 0.8 g shortfall is the same each time. With 9 fills it cannot be told from chance, and with 100 fills it is overwhelming. A p-value measures the strength of evidence, not the size of the effect.

Six Misreadings of the p-Value

Common claimWhy it is wrongA better statement
“p = 0.036 means there is a 3.6% chance the null is true.”The p-value assumes the null is true; it does not give its probability“If the line were on target, a result this far off would occur 3.6% of the time.”
“p < 0.05 means the effect is important.”A large sample makes a trivial effect significantReport the size of the difference and its interval
“p > 0.05 means there is no difference.”A small sample may miss a real effect“The data did not show a difference; the interval is …”
“p = 0.049 and p = 0.051 are completely different.”They are practically the same evidenceReport the p-value, and judge the whole picture
“A smaller p-value means a bigger effect.”It depends on sample size and spread as wellLook at the effect size
“We tested until we got p < 0.05.”Repeated testing and peeking inflate false alarmsFix the sample size and the test in advance
Chance of at least one false alarm at a 5% risk per test 1 groups: 1 pairwise tests 5.0% 5 groups: 5 pairwise tests 22.6% 10 groups: 10 pairwise tests 40.1% 20 groups: 20 pairwise tests 64.2%
Run enough tests at the 5% level and false alarms become likely. If you test 20 things, you should expect about one “significant” result by chance alone. Adjust for the number of tests, or decide the key test in advance.

Statistical and Practical Significance

SituationWhat it meansWhat to do
Significant and large enough to matterA real, important effectAct, and plan how to confirm it
Significant but tinyReal but unimportant (often a very large sample)Do not spend effort on it unless it is cheap
Not significant, but the interval includes an important differenceInconclusive: the sample was too smallCollect more data
Not significant, and the interval is tightly around zeroNo important differenceConclude that any difference is small

Always state the smallest difference that would matter before the test. A useful quick measure of size is the effect in standard deviations (Cohen’s d): about 0.2 is small, 0.5 medium, and 0.8 large, but your tolerance and costs matter more than these labels.

Run It in Excel and Minitab

ExcelStep by step

  1. Calculate the summary values with =AVERAGE(range), =STDEV.S(range), and =COUNT(range).
  2. Standard error: =s/SQRT(n). Test statistic: =(xbar-mu0)/SE.
  3. Two-sided p-value: =T.DIST.2T(ABS(t), n-1) gives 0.0359 for the example. One-sided: =T.DIST.RT(t, n-1) (upper tail).
  4. Critical value: =T.INV.2T(0.05, n-1). Interval half-width: =CONFIDENCE.T(0.05, s, n).
  5. For a p-value from a z statistic: =2*(1-NORM.S.DIST(ABS(z),TRUE)).

Excel has no single “one-sample t-test” command, so use the formulas above. For two samples, see the t-test page.

MinitabStep by step

  1. Choose Stat > Basic Statistics > 1-Sample t.
  2. From the dropdown, choose Summarized data (or One or more samples, each in a column if you have the raw values). Enter the sample size, mean, and standard deviation.
  3. Tick Perform hypothesis test and enter the hypothesized mean (500).
  4. Click Options to set the confidence level and the alternative (not equal, less than, or greater than).
  5. Click OK, and read the interval, the T-Value, and the P-Value in the session window.
  6. To plan or check power, choose Stat > Power and Sample Size > 1-Sample t.

The Assistant (Assistant > Hypothesis Tests > 1-Sample t) adds a power check and a report card.

Minitab session window (typed excerpt, simplified)
Descriptive Statistics

N   Mean  StDev  SE Mean        95% CI for μ
25  499.200  1.800    0.360  (498.457, 499.943)

μ: mean of Fill Weight

Test

Null hypothesis         H₀: μ = 500
Alternative hypothesis  H₁: μ ≠ 500

T-Value  P-Value
  -2.22    0.036

Reading and Reporting

  1. Say what you tested and why the sample is representative.
  2. Give the estimate and its interval. The interval shows the size and the uncertainty.
  3. Give the test statistic, degrees of freedom, and exact p-value, not just “p < 0.05”.
  4. Say what it means in practice, against the tolerance, the cost, or the target.
  5. State the limits: how the data were collected, and what you did not test.
A sentence you can use. The mean fill weight was 499.2 g (n = 25, SD = 1.8 g), 0.8 g below the 500 g target (95% CI 498.46 to 499.94 g; one-sample t(24) = -2.22, p = 0.036).

Common Mistakes

MistakeWhy it misleadsBetter
Choosing α or the direction after seeing the dataIt inflates false alarmsDecide both before the test
Using a one-sided test to get a smaller p-valueHides a shift in the other directionUse one-sided only when the other direction is truly irrelevant, and decide in advance
Ignoring powerA non-significant result may be a missed effectPlan the sample size
Testing many things and reporting the significant onesFalse alarms accumulateAdjust, or state the planned tests
Reporting only the p-valueNo size, no uncertaintyAdd the estimate and its interval
Treating non-significant as “no effect”Absence of evidence is not evidence of absenceLook at the interval and the power
Treating the assumptions as optionalThe p-value is only as good as the methodCheck independence, shape, and spread

Try It Yourself

A supplier claims an average of at most 2.0% moisture. Your 16 samples average 2.3% with a standard deviation of 0.5%. Test the claim at α = 0.05, using a one-sided test because only higher moisture is a problem.

  • State the hypotheses and calculate t and the p-value.
  • What would a Type I error and a Type II error mean here?
Show the answer

H0: μ ≤ 2.0. H1: μ > 2.0. SE = 0.5 / √16 = 0.125, so t = (2.3 − 2.0) / 0.125 = 2.40 on 15 df, and the one-sided p-value is 0.0149. Since p < 0.05, reject H0: the moisture is above 2.0%.

A Type I error would be accusing the supplier of excess moisture when the true mean is within spec (a false complaint). A Type II error would be accepting the lots when the true moisture is above spec (a missed problem). Which is worse decides how you set α and the sample size.

P-Values and Error Types: Frequently Asked Questions

What is a good significance level?

0.05 is the convention, but it is a choice, not a law. Use a smaller level when a false alarm is costly (changing a validated process), and a larger one when missing a real change is costly (a safety problem), and decide before you see the data.

What does p = 0.04 really mean?

If the null hypothesis were true, you would see a result at least this extreme about 4% of the time. It does not mean there is a 4% chance the null is true, and it says nothing about the size of the effect.

Should I use a one-sided or two-sided test?

Use two-sided unless you decided beforehand, for a practical reason, that a change in the other direction is irrelevant. A one-sided test has more power in its direction and none in the other.

What is the relationship between a confidence interval and a test?

They agree. A two-sided test at level α rejects the null hypothesis exactly when the (1 − α) confidence interval excludes the null value. The interval adds the size and the uncertainty.

Why do big samples make everything significant?

The standard error shrinks as the sample grows, so even a tiny difference becomes many standard errors. Statistical significance is then guaranteed for any real difference, however small. Judge importance by the effect size.

How is power related to Type II error?

Power is one minus the Type II error rate (β). A test with 80% power misses a real effect of the planned size 20% of the time. Power rises with sample size, effect size, and α, and falls with noise.

Sources and Further Reading

  • Ronald L. Wasserstein and Nicole A. Lazar, “The ASA Statement on p-Values: Context, Process, and Purpose,” The American Statistician, 2016.
  • NIST/SEMATECH, e-Handbook of Statistical Methods, sections on hypothesis tests (itl.nist.gov/div898/handbook).
  • Douglas C. Montgomery and George C. Runger, Applied Statistics and Probability for Engineers, Wiley.
  • David S. Moore, George P. McCabe, and Bruce A. Craig, Introduction to the Practice of Statistics, Freeman.
  • Minitab Support, “Methods and formulas for 1-Sample t” (support.minitab.com).

This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.