- Question it answers
- Do groups differ, without assuming normal data?
- Data needed
- Continuous or ordinal data; independent observations
- Method
- Replace values with ranks, then test the ranks
- Key output
- Rank-sum statistic, p-value, difference in medians with an interval
- Use when
- Skewed or small samples, outliers, or ordinal data
- Excel
- Built from RANK.AVG, SUMIF, and the normal or chi-square functions
- Minitab
- Stat > Nonparametric > Mann-Whitney, 1-Sample Wilcoxon, Kruskal-Wallis
- Prerequisite
- Normality tests
The Idea in Plain Language
Most of the tests you meet first, such as the t-test and ANOVA, work with the actual values and assume the data are roughly normal. Nonparametric tests do not assume a particular distribution. Most of them replace the values with their ranks: the smallest value gets rank 1, the next rank 2, and so on. The test then asks whether the ranks in one group are systematically higher than in another.
Ranks have two useful properties. An extreme value counts only as the next step up, so outliers do little damage. And the pattern of ranks does not depend on the shape of the distribution, so the tests stay valid for skewed data, small samples, and ordinal ratings.
When to Use Which Test
| Question | Parametric test | Nonparametric alternative | Minitab menu |
|---|---|---|---|
| One sample against a target | 1-sample t | 1-sample Wilcoxon signed-rank (or sign test) | Stat > Nonparametric > 1-Sample Wilcoxon |
| Two paired measurements | Paired t | Wilcoxon signed-rank on the differences | Stat > Nonparametric > 1-Sample Wilcoxon |
| Two independent groups | 2-sample t | Mann-Whitney (Wilcoxon rank-sum) | Stat > Nonparametric > Mann-Whitney |
| Three or more independent groups | One-way ANOVA | Kruskal-Wallis (or Mood’s median) | Stat > Nonparametric > Kruskal-Wallis |
| Blocked or repeated measures | Two-way ANOVA with blocks | Friedman test | Stat > Nonparametric > Friedman |
| Association between two variables | Pearson correlation | Spearman rank correlation | Stat > Basic Statistics > Correlation |
| Use a nonparametric test when | Because |
|---|---|
| The data are clearly skewed and the groups are small | Parametric p-values may be wrong; ranks do not depend on the shape |
| There are outliers you cannot remove | Ranks limit their influence |
| The data are ordinal (ratings 1 to 5, severity levels) | The distances between ratings are not meaningful |
| The measurement has a ceiling or floor (values pile up at a limit) | Ranks are less distorted |
| A transformation does not fix the problem | No other simple remedy |
Mann-Whitney Test: Two Independent Groups
A maintenance team compares repair times (hours) after two diagnostic methods. Method A has 8 repairs and Method B has 9. Repair times are skewed, with a few long jobs.
By hand
- Hypotheses. H0: the two methods give the same distribution of repair times. H1: one method tends to give longer times. α = 0.05.
- Rank all 17 values together, from smallest (rank 1) to largest, ignoring the groups. Average the ranks when values tie (there are none here).
- Add the ranks in each group. WA = 44 and WB = 109. As a check they add to N(N + 1)/2 = 153.
- Convert to U. UA = WA − nA(nA + 1)/2 = 44 − 36 = 8, and UB = nAnB − UA = 64. The smaller, 8, is the test statistic.
- Find the p-value. Under H0, WA has mean nA(N + 1)/2 = 72 and standard deviation √(nAnB(N + 1)/12) = 10.39. The normal approximation with continuity correction gives z = (28 − 0.5) / 10.39 = 2.65, p = 0.0081. The exact p-value for these sample sizes is 0.0055.
- Estimate the size of the difference. The Hodges-Lehmann estimate is the median of all 72 pairwise differences (A − B): -13.0 hours, with an approximate 95% interval of -24 to -5 hours.
| Rank | Time (h) | Method |
|---|---|---|
| 1 | 9 | A |
| 2 | 11 | A |
| 3 | 12 | A |
| 4 | 14 | A |
| 5 | 15 | A |
| 6 | 17 | A |
| 7 | 18 | B |
| 8 | 20 | A |
| 9 | 22 | B |
| 10 | 24 | B |
| 11 | 25 | B |
| 12 | 27 | B |
| 13 | 30 | B |
| 14 | 33 | B |
| 15 | 38 | A |
| 16 | 45 | B |
| 17 | 60 | B |
Run Mann-Whitney in Excel and Minitab
ExcelStep by step
- Put the two groups in A2:A9 (Method A) and B2:B10 (Method B). In column D, stack all 17 values; in column E, mark the group (A or B).
- Rank all values together: in column F. It averages the ranks of ties.
- Sum the ranks for Method A: gives WA.
- Calculate z: and the p-value .
- Median of each group: .
Excel has no built-in rank-sum test, so the steps above build it. The normal approximation is good for samples of about 8 or more per group; for smaller samples use exact tables or Minitab.
MinitabStep by step
- Put the data in two columns, one for each method (Method A and Method B).
- Choose . Set First sample and Second sample.
- Set the Confidence level (95%) and the Alternative (not equal, unless you decided on a direction in advance), and click OK.
- Read the difference in medians with its interval, the W-value, and the p-value.
- For several groups, use .
Minitab reports the p-value from the normal approximation, adjusted for ties when there are any. Its interval for the difference may differ slightly from the one shown here in the last digit.
Method
η₁: median of Method A
η₂: median of Method B
Difference: η₁ - η₂
Descriptive Statistics
Sample N Median
Method A 8 14.5
Method B 9 27.0
Estimation for Difference
CI for Achieved
Difference Difference Confidence
-13.0 (-24.0, -5.0) 96.14%
Test
Null hypothesis H₀: η₁ - η₂ = 0
Alternative hypothesis H₁: η₁ - η₂ ≠ 0
Method W-Value P-Value
Not adjusted for ties 44.0 0.0081Wilcoxon Signed-Rank Test: Paired Data
A new fixture is meant to cut handling time. Nine operators were timed with the old and the new fixture. The differences are small in number and not clearly normal, so the team uses a rank test on the paired differences.
| Operator | Before | After | Difference | |Difference| | Rank of |difference| | Signed rank |
|---|---|---|---|---|---|---|
| 1 | 34 | 30 | -4 | 4 | 4 | −4 |
| 2 | 28 | 27 | -1 | 1 | 1 | −1 |
| 3 | 45 | 38 | -7 | 7 | 7 | −7 |
| 4 | 39 | 41 | +2 | 2 | 2 | +2 |
| 5 | 52 | 41 | -11 | 11 | 9 | −9 |
| 6 | 31 | 28 | -3 | 3 | 3 | −3 |
| 7 | 47 | 39 | -8 | 8 | 8 | −8 |
| 8 | 36 | 31 | -5 | 5 | 5 | −5 |
| 9 | 41 | 35 | -6 | 6 | 6 | −6 |
- Difference each pair (after − before). Drop any zero differences.
- Rank the absolute differences from smallest to largest, and give each rank the sign of its difference.
- Add the ranks of the positive and the negative differences. W+ = 2 and W− = 43, which add to n(n + 1)/2 = 45.
- The test statistic is the smaller sum, 2. For n = 9, the exact two-sided p-value is 0.0117, so the change is significant at 0.05. (A paired t-test gives p = 0.0060.)
In Excel: compute the differences, then for the ranks and for W+. In Minitab: make a column of differences and choose with a test median of 0.
Kruskal-Wallis Test: Three or More Groups
Kruskal-Wallis is the rank version of one-way ANOVA. Setup times on three machines were recorded for six runs each.
| Run | X | Rank | Y | Rank | Z | Rank |
|---|---|---|---|---|---|---|
| 1 | 12 | 2 | 20 | 10 | 16 | 6 |
| 2 | 15 | 5 | 25 | 15 | 19 | 9 |
| 3 | 11 | 1 | 22 | 12 | 17 | 7 |
| 4 | 18 | 8 | 30 | 17 | 35 | 18 |
| 5 | 14 | 4 | 27 | 16 | 21 | 11 |
| 6 | 13 | 3 | 24 | 14 | 23 | 13 |
| Rank sum | 23 | 84 | 64 |
With N = 18, rank sums 23, 84, and 64, and six values per machine: H = 12 / (18 × 19) × (23² + 84² + 64²) / 6 − 3 × 19 = 11.31. Compare with the chi-square distribution on k − 1 = 2 degrees of freedom: p = 0.0035. Machine setup times are not all alike. (One-way ANOVA on the same data gives F = 8.48, p = 0.0034, the same conclusion.)
As with ANOVA, a significant result does not say which machines differ. Follow up with Mann-Whitney tests on each pair and adjust the p-values for the number of comparisons (Bonferroni: multiply each by 3), or use Dunn’s test where available.
Mood’s median test is a simpler alternative. It counts how many values in each group fall above the overall median (19.5) and tests the counts with chi-square. Here it gives p = 0.002. It is more robust to extreme outliers but has less power than Kruskal-Wallis.
In Excel: rank with , sum the ranks per machine with , calculate H as above, and get the p-value with . In Minitab: (Response and Factor columns), or .
Assumptions and Limits
| Assumption or issue | Why it matters | What to do |
|---|---|---|
| Independent observations | Every test needs it | Think about how the data were collected |
| Similar shape in the groups | Needed to read the result as a difference in medians; if shapes differ, the test shows only that one group tends to be higher | Compare the dot plots or box plots |
| Many tied values | Ties reduce the information in the ranks | Use the version adjusted for ties; improve measurement resolution |
| Very small samples | Little power; the smallest possible p-value may be above 0.05 | Use exact p-values; collect more data (two groups of 3 can never give p below 0.1) |
| The question is about the mean | Rank tests compare the whole distributions or the medians, not the means | State the question as medians, or use a transformation and a t-test |
| Estimates and intervals | Ranks do not give means and standard errors | Report medians and the Hodges-Lehmann interval |
Reading and Reporting
- Name the test and why: “Because repair times were skewed, a Mann-Whitney test was used.”
- Report the medians, not the means, with the estimated difference and its interval.
- Report the statistic and the p-value (W, U, H, or the signed-rank sum).
- Check the shape assumption before you call it a difference in medians.
- Say what it means in practice, in the units of the data.
Common Mistakes
| Mistake | Why it misleads | Better |
|---|---|---|
| Using a rank test on large, near-normal samples | Loses a little power for no gain | Use the t-test or ANOVA, after checking the plots |
| Reporting means with a rank test | The test does not compare means | Report medians |
| Choosing the test after seeing which gives p < 0.05 | Inflates false alarms | Decide from the data type and shape before testing |
| Ignoring ties | The p-value can be off | Use the tie-adjusted version |
| Comparing groups with different shapes as if only the medians differed | A significant result may be about spread or shape | Plot the groups; describe the difference honestly |
| Many pairwise rank tests with no adjustment | False alarms accumulate | Adjust, or use Kruskal-Wallis first |
| Assuming nonparametric means “no assumptions” | Independence still matters | Check independence and shape |
Try It Yourself
Customer waiting times (minutes) were recorded for two service desks: Desk 1: 4, 6, 3, 8, 5. Desk 2: 9, 12, 7, 15, 10, 11.
- Rank all 11 values together and find the rank sum for Desk 1.
- Is Desk 1 faster at α = 0.05? Use the exact p-value for groups this small.
Show the answer
Ranks of Desk 1: 2, 4, 1, 6, 3. The rank sum is 16 (the minimum possible is 15). The exact two-sided p-value is 0.0087, so Desk 1 is significantly faster at the 5% level: the median waiting time was 5 minutes at Desk 1 and 10.5 at Desk 2.
With groups this small, the exact p-value matters. Normal approximations can be inaccurate when both groups have fewer than about 8 values.
Nonparametric Tests: Frequently Asked Questions
What is the difference between Mann-Whitney and Wilcoxon rank-sum?
They are the same test, described with different statistics. Mann-Whitney is often reported with U, and Wilcoxon rank-sum with W, which is the sum of the ranks of one group. Minitab calls it Mann-Whitney and reports W. The Wilcoxon signed-rank test is a different test, for paired data.
Do nonparametric tests compare medians?
Only if the groups have a similar shape. In general they test whether values in one group tend to be larger than in the other. With similarly shaped distributions, that is the same as a difference in medians. Plot the groups first.
Are nonparametric tests less powerful?
Slightly, when the data really are normal: the Mann-Whitney test needs about 5% more data than the t-test for the same power. When the data are skewed or have outliers, they can be more powerful than the t-test.
What sample size do I need?
The same planning logic applies: choose the smallest difference that matters and size the sample for it. A rough approach is to calculate the sample for the t-test and add about 15%. Very small groups cannot give a significant result however large the difference.
What if I have only a few values?
Use exact p-values, which Minitab gives for small samples with no ties. Remember that with groups of 3 and 3, the smallest possible p-value is 0.10, so a significant result is impossible.
When should I use the sign test?
The sign test counts only which way each difference goes and ignores their size. It needs the fewest assumptions but has the least power. Use it when the data are only ordinal or the size of the differences is not meaningful.
Sources and Further Reading
- NIST/SEMATECH, e-Handbook of Statistical Methods, sections on nonparametric tests (itl.nist.gov/div898/handbook).
- Myles Hollander, Douglas A. Wolfe, and Eric Chicken, Nonparametric Statistical Methods, Wiley.
- W. J. Conover, Practical Nonparametric Statistics, Wiley.
- Frank Wilcoxon, “Individual comparisons by ranking methods,” Biometrics Bulletin, 1945; Henry B. Mann and Donald R. Whitney, Annals of Mathematical Statistics, 1947; William H. Kruskal and W. Allen Wallis, Journal of the American Statistical Association, 1952.
- Minitab Support, “Methods and formulas for Mann-Whitney, 1-Sample Wilcoxon, and Kruskal-Wallis” (support.minitab.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.