- Question it answers
- What is a typical value, and how much do the values vary?
- Data needed
- A column of numerical measurements
- Key output
- Center, spread, and shape, with a plot
- Default pair
- Mean and SD (symmetric); median and IQR (skewed)
- Assumptions
- None for describing; shape decides which summary fits
- Excel
- AVERAGE, MEDIAN, STDEV.S, QUARTILE.EXC; Descriptive Statistics tool
- Minitab
- Stat > Basic Statistics > Display Descriptive Statistics
- Why it matters
- Every test and chart builds on these numbers
The Idea in Plain Language
Descriptive statistics boil a pile of numbers down to a few that describe it. Two jobs matter most: where the data are centered (location) and how spread out they are (variation). A third, the shape, tells you whether the first two can be trusted.
| Job | Measures | Tells you |
|---|---|---|
| Center | Mean, median, mode | A typical value |
| Spread | Range, interquartile range, variance, standard deviation | How much the values differ |
| Shape | Skewness, kurtosis, and a picture | Symmetric or lopsided, heavy tails or not |
| Position | Minimum, quartiles, percentiles, maximum | Where a value sits in the data |
How Each Measure Works
| Measure | Formula or rule | Use it when |
|---|---|---|
| Mean | x̄ = Σx / n | Data are roughly symmetric; it uses every value |
| Median | Middle value of the sorted data | Data are skewed or have outliers |
| Mode | Most frequent value | Categories or discrete values |
| Range | Maximum − minimum | Very small samples; subgroup charts |
| Interquartile range | Q3 − Q1 | Spread that ignores the extremes |
| Sample variance | s² = Σ(x − x̄)² / (n − 1) | Intermediate step; additive across independent sources |
| Sample standard deviation | s = √s² | The standard measure of spread, in the data’s own units |
| Standard error of the mean | s / √n | How precisely the mean is known |
| Coefficient of variation | s / x̄ × 100% | Compare spread between things with different averages |
A useful rule for roughly bell-shaped data: about 68% of values fall within one standard deviation of the mean, 95% within two, and 99.7% within three.
Worked Example 1: By Hand
Eight order cycle times in minutes: 12, 15, 14, 10, 13, 16, 11, 13.
| x | x − mean | (x − mean)² |
|---|---|---|
| 12 | -1.00 | 1.0000 |
| 15 | +2.00 | 4.0000 |
| 14 | +1.00 | 1.0000 |
| 10 | -3.00 | 9.0000 |
| 13 | +0.00 | 0.0000 |
| 16 | +3.00 | 9.0000 |
| 11 | -2.00 | 4.0000 |
| 13 | +0.00 | 0.0000 |
| Σx = 104 | 0.00 | 28.0000 |
- Mean = 104 / 8 = 13.000.
- Median: sorted values are 10, 11, 12, 13, 13, 14, 15, 16; the middle two are 13 and 13, so the median is 13.
- Range = 16 − 10 = 6.
- Variance = 28.0000 / (8 − 1) = 4.0000.
- Standard deviation = √4.0000 = 2.000 minutes.
Worked Example 2: When the Mean Misleads
A larger sample of 24 order processing times, in minutes. Most orders take 12 to 19 minutes, but a few take much longer.
| Statistic | All 24 orders | Without the slowest |
|---|---|---|
| Mean | 17.15 | 16.11 |
| Median | 15.05 | 14.90 |
| Standard deviation | 6.33 | 3.86 |
| Range | 29.2 | 16.1 |
- Q1 = 13.65, Q3 = 18.17, interquartile range = 4.52.
- Trimmed mean (10% off each end) = 15.93, between the median and the mean.
- Skewness = 2.71: a long right tail. Excess kurtosis = 8.60: heavy tails.
- Standard error of the mean = 1.29; coefficient of variation = 37%.
Run It in Excel and Minitab
ExcelStep by step
- Put the data in a column. Use , , and .
- Spread: (6.327), (40.03), . Use only for a full population.
- Quartiles: (13.65) matches Minitab. uses a slightly different rule and can differ.
- Shape: (2.71), (8.60), and .
- All at once: , tick Summary statistics.
MinitabStep by step
- . Choose the column as Variable.
- Click Statistics to choose what to show: mean, SE of mean, standard deviation, variance, coefficient of variation, minimum, Q1, median, Q3, maximum, range, IQR, skewness, kurtosis, and more.
- Click Graphs to add a histogram, an individual value plot, or a box plot.
- For several groups, enter a By variable.
- Quicker: or gives the numbers, a histogram, a box plot, and intervals together.
Descriptive Statistics: Time Variable N N* Mean SE Mean StDev Minimum Q1 Median Q3 Maximum Time 24 0 17.15 1.29 6.33 11.80 13.65 15.05 18.17 41.00 Variable Variance CoefVar Range IQR Skewness Kurtosis Time 40.03 36.89 29.20 4.52 2.71 8.60
Reading and Reporting
- Plot first. Look at the shape before trusting any single number.
- Pair a center with a spread: mean with standard deviation for symmetric data, median with interquartile range for skewed data.
- Give the sample size, and the units.
- Show outliers, do not hide them. Investigate the cause before deciding to set one aside, and say what you did.
Common Mistakes
| Mistake | Why it misleads | Better |
|---|---|---|
| Reporting only the mean | Hides spread and shape | Report the mean and the standard deviation, and plot the data |
| Using the mean for skewed data | A few large values drag it away from the typical value | Use the median and IQR |
| Using STDEV.P on a sample | Underestimates the spread | Use STDEV.S for samples |
| Deleting an outlier because it is inconvenient | It may be the most important point | Find the cause; document the decision |
| Comparing standard deviations of different-scale things | A 2-gram spread and a 2-kilogram spread are not comparable | Use the coefficient of variation |
| Mixing subgroups in one summary | Two processes blend into a misleading average | Summarize each group, then compare |
Try It Yourself
Five fill weights in grams: 498, 502, 500, 503, 497.
- Find the mean, the median, and the sample standard deviation.
- What would the standard deviation be if you wrongly divided by n?
Show the answer
Mean = 500 g. Median = 500 g. Sample standard deviation = 2.550 g.
Dividing by n gives 2.280 g, which is smaller. That is the population formula, and it understates the spread of a sample.
Descriptive Statistics: Frequently Asked Questions
Should I use the mean or the median?
Use the mean for symmetric data and the median for skewed data or data with outliers. If they differ a lot, the data are skewed, and that is itself worth reporting.
What is the difference between standard deviation and variance?
Variance is the average squared deviation, and the standard deviation is its square root. The standard deviation is in the same units as the data, so it is easier to interpret. Variance is easier to work with when adding independent sources of variation.
What is the difference between STDEV.S and STDEV.P?
STDEV.S divides by n − 1 and estimates the spread of a population from a sample. STDEV.P divides by n and is for when you have the entire population. Use STDEV.S almost always.
Why does Excel give a different quartile from Minitab?
There are several ways to compute quartiles. QUARTILE.EXC matches Minitab; QUARTILE.INC, and the older QUARTILE function, use a slightly different rule. The difference is small for large samples.
Is a coefficient of variation always useful?
It is useful for ratio-scale data that cannot be negative, such as times and weights, when you want to compare spread across different averages. It makes no sense for data that can be zero or negative, such as temperature in degrees Celsius.
Sources and Further Reading
- NIST/SEMATECH, e-Handbook of Statistical Methods, Exploratory Data Analysis (itl.nist.gov/div898/handbook).
- David S. Moore, George P. McCabe, and Bruce A. Craig, Introduction to the Practice of Statistics, Freeman.
- Rick J. Hyndman and Yanan Fan, “Sample Quantiles in Statistical Packages,” The American Statistician, 1996.
- Minitab Support, “Methods and formulas for Display Descriptive Statistics” (support.minitab.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.