- Question it answers
- How well does a sample average estimate the process average?
- Data needed
- A random, representative sample
- Key output
- Standard error and an interval for the mean
- Rule
- Standard error = s / √n
- Assumptions
- Independent, representative observations; enough n for the shape
- Excel
- STDEV.S/SQRT(COUNT), CONFIDENCE.T, RAND, GAMMA.INV
- Minitab
- Display Descriptive Statistics (SE Mean); Calc > Random Data
- Why it matters
- It explains why averages, tests, and control charts work
The Idea in Plain Language
You almost never measure everything. You take a sample and use it to learn about the whole population. Two ideas make that work. First, the sample must represent the population. Second, the average of a sample behaves in a predictable way, which is the central limit theorem (CLT).
The CLT says: take many samples of size n from almost any population, and the averages of those samples will pile up in a bell shape centered on the population mean, with a spread of σ/√n. That spread is the standard error. It is why t-tests, control charts for averages, and confidence intervals work even when individual values are not normal.
Sampling Well
| Method | How | Best for | Watch out |
|---|---|---|---|
| Simple random | Every item has an equal chance | Uniform populations, a list to draw from | Needs a complete list |
| Systematic | Every k-th item | Production lines, easy to run | Hidden cycles that match k |
| Stratified | Random samples within groups (shifts, machines) | Known groups that differ | Need to know the groups |
| Cluster | Sample whole groups (cartons, batches) | Cheap access to groups | Less precise than random |
| Rational subgroup | Small samples taken close in time | Control charts: spread within, change between | Subgroups that mix sources |
| Convenience | Whatever is easy | Nothing reliable | Biased; avoid |
For more on choosing a method, see the Sampling Methods entry. For how many to take, see Sample Size and Power.
Seeing the Central Limit Theorem
Repair times in a plant are right-skewed. The population mean is 9.8 hours and the standard deviation is 5.9 hours. We drew 4,000 samples at each size and plotted the averages.
| Sample size n | Predicted standard error σ/√n | Simulated SD of the averages | Skewness of the averages |
|---|---|---|---|
| 1 | 5.86 | 5.83 | 1.10 |
| 5 | 2.62 | 2.57 | 0.47 |
| 30 | 1.07 | 1.06 | 0.21 |
How large must n be? For roughly symmetric data, 5 to 10 is often enough. For moderately skewed data, 30 is a common rule. For extremely skewed data or heavy tails, you may need hundreds. The CLT also needs independent observations.
Standard Error in Practice
One real sample of 25 repairs gives a mean of 8.94 hours and a sample standard deviation of 6.14 hours.
- Standard error = s / √n = 6.14 / √25 = 1.228 hours.
- 95% interval for the mean = 8.94 ± 2.064 × 1.228 = 6.40 to 11.47 hours.
- To halve the standard error to 0.614 you need n = (6.14 / 0.614)² = 100 repairs, four times as many.
Run It in Excel and Minitab
ExcelStep by step
- Standard error: (1.228).
- Interval: gives the margin (2.535).
- Simulate the CLT: in a grid of 30 columns, enter in many rows, then average each row with and make a histogram of the averages. Press F9 to redraw.
- Random sample from a list: add beside each item, sort by it, and take the top n.
MinitabStep by step
- Standard error: shows SE Mean next to the mean and standard deviation.
- Simulate: to fill 30 columns (shape 2.8, scale 3.5), then with Mean to average across the row, then of the averages.
- Random sample from a column: .
- Sample size for a target margin: or with a confidence interval, to see how the margin shrinks.
Descriptive Statistics: Repair Variable N Mean SE Mean StDev Minimum Q1 Median Q3 Maximum Repair 25 8.94 1.228 6.14 2.30 3.90 8.20 11.55 28.50
Reading and Reporting
- Describe how you sampled, and why it should represent the process.
- Report the mean with its standard error or interval, not only the mean.
- Report the standard deviation to describe the process, and the standard error to describe the estimate.
- Say how many observations and whether they were independent.
Common Mistakes
| Mistake | Why it misleads | Better |
|---|---|---|
| Confusing standard deviation with standard error | The first does not shrink with n | State which one, and use each for its purpose |
| Believing n = 30 always makes data normal | Heavy skew or outliers need more | Plot the averages or use a bootstrap |
| Sampling only what is convenient | Bias does not average out | Use random or stratified sampling |
| Using the CLT on non-independent data | Autocorrelation makes the effective n much smaller | Check the time order; subsample or model it |
| Applying the CLT to individual values | It describes averages, not single items | Use the distribution of the individuals for specification limits |
| Ignoring the finite population | When the sample is a big share of a small population, the SE is overstated | Apply the finite population correction |
Try It Yourself
Cycle times have a standard deviation of 8 minutes. You plan to average 16 observations.
- What is the standard error of the average?
- How many observations would you need for a standard error of 1 minute?
Show the answer
Standard error = 8 / √16 = 2.0 minutes.
For 1 minute: n = (8 / 1)² = 64 observations, four times as many to halve the standard error.
Sampling and the Central Limit Theorem: Frequently Asked Questions
What is the central limit theorem in simple terms?
Averages of random samples tend to follow a bell curve, whatever the shape of the individual values, and the larger the sample the closer the fit and the narrower the curve.
What is the standard error?
It is the standard deviation of a statistic, usually the sample mean. For the mean it equals the standard deviation divided by the square root of the sample size, and it describes how far the sample mean is likely to be from the true mean.
Is n = 30 enough?
It is a rule of thumb and works for moderately skewed data. For symmetric data a smaller n is enough. For very skewed data or data with outliers you may need far more.
Why do control charts use subgroup averages?
Averages of subgroups are close to normal by the central limit theorem, so the 3-sigma limits for the chart behave predictably even if individual values are not normal.
Does a bigger sample fix a biased sample?
No. A bigger sample reduces random error but not bias. The estimate just becomes more precisely wrong.
Sources and Further Reading
- NIST/SEMATECH, e-Handbook of Statistical Methods, sections on sampling and the central limit theorem (itl.nist.gov/div898/handbook).
- David S. Moore, George P. McCabe, and Bruce A. Craig, Introduction to the Practice of Statistics, Freeman.
- William G. Cochran, Sampling Techniques, Wiley.
- Minitab Support, “Random Data” and “Row Statistics” (support.minitab.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.