- Question it answers
- Is what I see real, or could chance explain it?
- Null hypothesis
- No change, no difference, or on target
- Key output
- Test statistic, p-value, and a confidence interval
- You set in advance
- The significance level (α), the direction, and the sample size
- Two errors
- Type I (false alarm, α) and Type II (missed signal, β)
- Excel
- T.DIST.2T, T.INV.2T, NORM.S.DIST formulas
- Minitab
- Stat > Basic Statistics > 1-Sample t; Power and Sample Size
- Next step
- Sample size and power, then the specific tests
The Idea in Plain Language
A statistical test is a structured way to decide whether what you see in a sample is a real feature of the process or just the luck of the draw. You start with a skeptical position, the null hypothesis (H0): nothing has changed, there is no difference, the process is on target. The alternative hypothesis (H1) says there is. Then you ask how surprising your data would be if the null hypothesis were true.
The answer to that question is the p-value: the probability of getting a result at least as extreme as the one you observed, if the null hypothesis were true. A small p-value says the data would be surprising under the null, so the null becomes hard to believe. It is the courtroom logic: the defendant is presumed innocent, and you convict only if the evidence would be very unlikely for an innocent person.
The Steps of Every Test
- State the question in practical terms, and the hypotheses: H0 (no change) and H1 (change, or change in a stated direction).
- Choose the risk you will accept. The significance level α (usually 0.05) is the chance of a false alarm you will tolerate.
- Decide the sample size before collecting data, so the test can see the difference that matters (power).
- Collect the data in a way that makes the observations independent and representative.
- Calculate the test statistic and the p-value. The statistic measures how far the sample is from the null, in standard errors.
- Decide and report. If p ≤ α, reject H0; otherwise, fail to reject it. Report the size of the effect with a confidence interval, not only the verdict.
Two Ways to Be Wrong: Type I and Type II Errors
| Null hypothesis is really true (no change) | Null hypothesis is really false (a real change) | |
|---|---|---|
| You reject H0 (call it a difference) | Type I error: a false alarm. Probability = α | Correct: you found the real change. Probability = power (1 − β) |
| You fail to reject H0 | Correct. Probability = 1 − α | Type II error: a missed signal. Probability = β |
| To get | You can | But |
|---|---|---|
| Fewer false alarms (smaller α) | Use a stricter level, such as 0.01 | Power falls, so you miss more real changes |
| More power (smaller β) | Collect more data, reduce measurement noise, or look for a larger effect | More data costs time and money |
| A real effect detected reliably | Plan the sample size in advance | You must decide what size of effect matters |
Both errors cost something, and the costs differ. Missing a real defect-rate increase in a safety part is worse than a false alarm, so you would accept a larger α and demand more power. Raising prices on a false signal may be worse, so you would use a smaller α. Choose the risks before you see the data. See Sample Size and Power.
Worked Example: Is the Filling Line on Target?
A line should fill 500 g. A sample of 25 fills has a mean of 499.2 g and a standard deviation of 1.8 g. Is the line off target?
- Hypotheses. H0: μ = 500 g. H1: μ ≠ 500 g (two-sided, since a shift either way matters). α = 0.05.
- Standard error. SE = s / √n = 1.8 / √25 = 0.360 g.
- Test statistic. t = (x̄ − μ0) / SE = (499.2 − 500) / 0.360 = -2.22, with 24 degrees of freedom.
- p-value. The chance of a |t| at least 2.22 when the mean is really 500 g is 0.036, the shaded area below.
- Decide. 0.036 < 0.05, so reject H0. The 95% confidence interval for the mean is 498.46 to 499.94 g, which excludes 500.
The same data with different sample sizes
The p-value depends on how much data you have as well as how big the difference is. Keep the same mean (499.2 g) and standard deviation (1.8 g), and change only the sample size:
| Sample size | Standard error | t | p-value | Decision at 0.05 |
|---|---|---|---|---|
| 9 | 0.600 | -1.33 | 0.2191 | Fail to reject |
| 25 | 0.360 | -2.22 | 0.0359 | Reject |
| 100 | 0.180 | -4.44 | < 0.0001 | Reject |
The 0.8 g shortfall is the same each time. With 9 fills it cannot be told from chance, and with 100 fills it is overwhelming. A p-value measures the strength of evidence, not the size of the effect.
Six Misreadings of the p-Value
| Common claim | Why it is wrong | A better statement |
|---|---|---|
| “p = 0.036 means there is a 3.6% chance the null is true.” | The p-value assumes the null is true; it does not give its probability | “If the line were on target, a result this far off would occur 3.6% of the time.” |
| “p < 0.05 means the effect is important.” | A large sample makes a trivial effect significant | Report the size of the difference and its interval |
| “p > 0.05 means there is no difference.” | A small sample may miss a real effect | “The data did not show a difference; the interval is …” |
| “p = 0.049 and p = 0.051 are completely different.” | They are practically the same evidence | Report the p-value, and judge the whole picture |
| “A smaller p-value means a bigger effect.” | It depends on sample size and spread as well | Look at the effect size |
| “We tested until we got p < 0.05.” | Repeated testing and peeking inflate false alarms | Fix the sample size and the test in advance |
Statistical and Practical Significance
| Situation | What it means | What to do |
|---|---|---|
| Significant and large enough to matter | A real, important effect | Act, and plan how to confirm it |
| Significant but tiny | Real but unimportant (often a very large sample) | Do not spend effort on it unless it is cheap |
| Not significant, but the interval includes an important difference | Inconclusive: the sample was too small | Collect more data |
| Not significant, and the interval is tightly around zero | No important difference | Conclude that any difference is small |
Always state the smallest difference that would matter before the test. A useful quick measure of size is the effect in standard deviations (Cohen’s d): about 0.2 is small, 0.5 medium, and 0.8 large, but your tolerance and costs matter more than these labels.
Run It in Excel and Minitab
ExcelStep by step
- Calculate the summary values with , , and .
- Standard error: . Test statistic: .
- Two-sided p-value: gives 0.0359 for the example. One-sided: (upper tail).
- Critical value: . Interval half-width: .
- For a p-value from a z statistic: .
Excel has no single “one-sample t-test” command, so use the formulas above. For two samples, see the t-test page.
MinitabStep by step
- Choose .
- From the dropdown, choose Summarized data (or One or more samples, each in a column if you have the raw values). Enter the sample size, mean, and standard deviation.
- Tick Perform hypothesis test and enter the hypothesized mean (500).
- Click Options to set the confidence level and the alternative (not equal, less than, or greater than).
- Click OK, and read the interval, the T-Value, and the P-Value in the session window.
- To plan or check power, choose .
The Assistant () adds a power check and a report card.
Descriptive Statistics N Mean StDev SE Mean 95% CI for μ 25 499.200 1.800 0.360 (498.457, 499.943) μ: mean of Fill Weight Test Null hypothesis H₀: μ = 500 Alternative hypothesis H₁: μ ≠ 500 T-Value P-Value -2.22 0.036
Reading and Reporting
- Say what you tested and why the sample is representative.
- Give the estimate and its interval. The interval shows the size and the uncertainty.
- Give the test statistic, degrees of freedom, and exact p-value, not just “p < 0.05”.
- Say what it means in practice, against the tolerance, the cost, or the target.
- State the limits: how the data were collected, and what you did not test.
Common Mistakes
| Mistake | Why it misleads | Better |
|---|---|---|
| Choosing α or the direction after seeing the data | It inflates false alarms | Decide both before the test |
| Using a one-sided test to get a smaller p-value | Hides a shift in the other direction | Use one-sided only when the other direction is truly irrelevant, and decide in advance |
| Ignoring power | A non-significant result may be a missed effect | Plan the sample size |
| Testing many things and reporting the significant ones | False alarms accumulate | Adjust, or state the planned tests |
| Reporting only the p-value | No size, no uncertainty | Add the estimate and its interval |
| Treating non-significant as “no effect” | Absence of evidence is not evidence of absence | Look at the interval and the power |
| Treating the assumptions as optional | The p-value is only as good as the method | Check independence, shape, and spread |
Try It Yourself
A supplier claims an average of at most 2.0% moisture. Your 16 samples average 2.3% with a standard deviation of 0.5%. Test the claim at α = 0.05, using a one-sided test because only higher moisture is a problem.
- State the hypotheses and calculate t and the p-value.
- What would a Type I error and a Type II error mean here?
Show the answer
H0: μ ≤ 2.0. H1: μ > 2.0. SE = 0.5 / √16 = 0.125, so t = (2.3 − 2.0) / 0.125 = 2.40 on 15 df, and the one-sided p-value is 0.0149. Since p < 0.05, reject H0: the moisture is above 2.0%.
A Type I error would be accusing the supplier of excess moisture when the true mean is within spec (a false complaint). A Type II error would be accepting the lots when the true moisture is above spec (a missed problem). Which is worse decides how you set α and the sample size.
P-Values and Error Types: Frequently Asked Questions
What is a good significance level?
0.05 is the convention, but it is a choice, not a law. Use a smaller level when a false alarm is costly (changing a validated process), and a larger one when missing a real change is costly (a safety problem), and decide before you see the data.
What does p = 0.04 really mean?
If the null hypothesis were true, you would see a result at least this extreme about 4% of the time. It does not mean there is a 4% chance the null is true, and it says nothing about the size of the effect.
Should I use a one-sided or two-sided test?
Use two-sided unless you decided beforehand, for a practical reason, that a change in the other direction is irrelevant. A one-sided test has more power in its direction and none in the other.
What is the relationship between a confidence interval and a test?
They agree. A two-sided test at level α rejects the null hypothesis exactly when the (1 − α) confidence interval excludes the null value. The interval adds the size and the uncertainty.
Why do big samples make everything significant?
The standard error shrinks as the sample grows, so even a tiny difference becomes many standard errors. Statistical significance is then guaranteed for any real difference, however small. Judge importance by the effect size.
How is power related to Type II error?
Power is one minus the Type II error rate (β). A test with 80% power misses a real effect of the planned size 20% of the time. Power rises with sample size, effect size, and α, and falls with noise.
Sources and Further Reading
- Ronald L. Wasserstein and Nicole A. Lazar, “The ASA Statement on p-Values: Context, Process, and Purpose,” The American Statistician, 2016.
- NIST/SEMATECH, e-Handbook of Statistical Methods, sections on hypothesis tests (itl.nist.gov/div898/handbook).
- Douglas C. Montgomery and George C. Runger, Applied Statistics and Probability for Engineers, Wiley.
- David S. Moore, George P. McCabe, and Bruce A. Craig, Introduction to the Practice of Statistics, Freeman.
- Minitab Support, “Methods and formulas for 1-Sample t” (support.minitab.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.