Question it answers
Do groups differ, without assuming normal data?
Data needed
Continuous or ordinal data; independent observations
Method
Replace values with ranks, then test the ranks
Key output
Rank-sum statistic, p-value, difference in medians with an interval
Use when
Skewed or small samples, outliers, or ordinal data
Excel
Built from RANK.AVG, SUMIF, and the normal or chi-square functions
Minitab
Stat > Nonparametric > Mann-Whitney, 1-Sample Wilcoxon, Kruskal-Wallis
Prerequisite
Normality tests

The Idea in Plain Language

Most of the tests you meet first, such as the t-test and ANOVA, work with the actual values and assume the data are roughly normal. Nonparametric tests do not assume a particular distribution. Most of them replace the values with their ranks: the smallest value gets rank 1, the next rank 2, and so on. The test then asks whether the ranks in one group are systematically higher than in another.

Ranks have two useful properties. An extreme value counts only as the next step up, so outliers do little damage. And the pattern of ranks does not depend on the shape of the distribution, so the tests stay valid for skewed data, small samples, and ordinal ratings.

Original values: one extreme value stretches the scale 3 5 6 8 120 Ranks: the same order, evenly spaced, so the outlier counts as just one step 1 2 3 4 5
Ranking turns an extreme value into just another step. An outlier of 120 has the same influence on a rank test as a value of 9 would.
Not “assumption-free”. These tests still need independent observations, and for the simplest interpretation, groups whose distributions have a similar shape. They trade some power for robustness: when the data really are normal, they detect a difference slightly less often than the t-test or ANOVA (about 95% as efficiently), but when the data are skewed they can be much better.

When to Use Which Test

QuestionParametric testNonparametric alternativeMinitab menu
One sample against a target1-sample t1-sample Wilcoxon signed-rank (or sign test)Stat > Nonparametric > 1-Sample Wilcoxon
Two paired measurementsPaired tWilcoxon signed-rank on the differencesStat > Nonparametric > 1-Sample Wilcoxon
Two independent groups2-sample tMann-Whitney (Wilcoxon rank-sum)Stat > Nonparametric > Mann-Whitney
Three or more independent groupsOne-way ANOVAKruskal-Wallis (or Mood’s median)Stat > Nonparametric > Kruskal-Wallis
Blocked or repeated measuresTwo-way ANOVA with blocksFriedman testStat > Nonparametric > Friedman
Association between two variablesPearson correlationSpearman rank correlationStat > Basic Statistics > Correlation
Use a nonparametric test whenBecause
The data are clearly skewed and the groups are smallParametric p-values may be wrong; ranks do not depend on the shape
There are outliers you cannot removeRanks limit their influence
The data are ordinal (ratings 1 to 5, severity levels)The distances between ratings are not meaningful
The measurement has a ceiling or floor (values pile up at a limit)Ranks are less distorted
A transformation does not fix the problemNo other simple remedy
With large samples, or near-normal data, prefer the parametric test. It has more power and gives estimates, such as means and intervals, that are easy to explain. Nonparametric tests earn their place when normality fails and the sample is not large.

Mann-Whitney Test: Two Independent Groups

A maintenance team compares repair times (hours) after two diagnostic methods. Method A has 8 repairs and Method B has 9. Repair times are skewed, with a few long jobs.

0 h 10 h 20 h 30 h 40 h 50 h 60 h 70 h 14.50 Method A 27.00 Method B Overall median 22.00
Repair times by method. The gold bars are the medians. Method B is higher and has the long tail.

By hand

  1. Hypotheses. H0: the two methods give the same distribution of repair times. H1: one method tends to give longer times. α = 0.05.
  2. Rank all 17 values together, from smallest (rank 1) to largest, ignoring the groups. Average the ranks when values tie (there are none here).
  3. Add the ranks in each group. WA = 44 and WB = 109. As a check they add to N(N + 1)/2 = 153.
  4. Convert to U. UA = WA − nA(nA + 1)/2 = 44 − 36 = 8, and UB = nAnB − UA = 64. The smaller, 8, is the test statistic.
  5. Find the p-value. Under H0, WA has mean nA(N + 1)/2 = 72 and standard deviation √(nAnB(N + 1)/12) = 10.39. The normal approximation with continuity correction gives z = (28 − 0.5) / 10.39 = 2.65, p = 0.0081. The exact p-value for these sample sizes is 0.0055.
  6. Estimate the size of the difference. The Hodges-Lehmann estimate is the median of all 72 pairwise differences (A − B): -13.0 hours, with an approximate 95% interval of -24 to -5 hours.
All 17 repair times in order, with their ranks and groups
RankTime (h)Method
19A
211A
312A
414A
515A
617A
718B
820A
922B
1024B
1125B
1227B
1330B
1433B
1538A
1645B
1760B
Conclusion. Method A repairs were shorter than Method B repairs (median 14.5 h against 27.0 h; Mann-Whitney W = 44, p = 0.008; estimated median difference -13.0 h, 95% CI -24 to -5 h). For comparison, a Welch t-test on the same data gives p = 0.018: the long right tail inflates the t-test’s standard error, so its evidence is weaker.

Run Mann-Whitney in Excel and Minitab

ExcelStep by step

  1. Put the two groups in A2:A9 (Method A) and B2:B10 (Method B). In column D, stack all 17 values; in column E, mark the group (A or B).
  2. Rank all values together: =RANK.AVG(D2,$D$2:$D$18,1) in column F. It averages the ranks of ties.
  3. Sum the ranks for Method A: =SUMIF(E2:E18,"A",F2:F18) gives WA.
  4. Calculate z: =(ABS(W-n1*(N+1)/2)-0.5)/SQRT(n1*n2*(N+1)/12) and the p-value =2*(1-NORM.S.DIST(z,TRUE)).
  5. Median of each group: =MEDIAN(A2:A9).

Excel has no built-in rank-sum test, so the steps above build it. The normal approximation is good for samples of about 8 or more per group; for smaller samples use exact tables or Minitab.

MinitabStep by step

  1. Put the data in two columns, one for each method (Method A and Method B).
  2. Choose Stat > Nonparametric > Mann-Whitney. Set First sample and Second sample.
  3. Set the Confidence level (95%) and the Alternative (not equal, unless you decided on a direction in advance), and click OK.
  4. Read the difference in medians with its interval, the W-value, and the p-value.
  5. For several groups, use Stat > Nonparametric > Kruskal-Wallis.

Minitab reports the p-value from the normal approximation, adjusted for ties when there are any. Its interval for the difference may differ slightly from the one shown here in the last digit.

Minitab session window (typed excerpt, simplified)
Method

η₁: median of Method A
η₂: median of Method B
Difference: η₁ - η₂

Descriptive Statistics

Sample      N  Median
Method A    8    14.5
Method B    9    27.0

Estimation for Difference

          CI for  Achieved
Difference  Difference  Confidence
  -13.0  (-24.0, -5.0)      96.14%

Test

Null hypothesis         H₀: η₁ - η₂ = 0
Alternative hypothesis  H₁: η₁ - η₂ ≠ 0

Method                W-Value  P-Value
Not adjusted for ties    44.0   0.0081

Wilcoxon Signed-Rank Test: Paired Data

A new fixture is meant to cut handling time. Nine operators were timed with the old and the new fixture. The differences are small in number and not clearly normal, so the team uses a rank test on the paired differences.

Handling time (s) for nine operators before and after a new fixture
OperatorBeforeAfterDifference|Difference|Rank of |difference|Signed rank
13430-444−4
22827-111−1
34538-777−7
43941+222+2
55241-11119−9
63128-333−3
74739-888−8
83631-555−5
94135-666−6
  1. Difference each pair (after − before). Drop any zero differences.
  2. Rank the absolute differences from smallest to largest, and give each rank the sign of its difference.
  3. Add the ranks of the positive and the negative differences. W+ = 2 and W− = 43, which add to n(n + 1)/2 = 45.
  4. The test statistic is the smaller sum, 2. For n = 9, the exact two-sided p-value is 0.0117, so the change is significant at 0.05. (A paired t-test gives p = 0.0060.)
Conclusion. Handling time was lower with the new fixture for 8 of 9 operators (median change -5 s; Wilcoxon signed-rank W = 2, n = 9, p = 0.012).

In Excel: compute the differences, then =RANK.AVG(ABS(D2),ABS_range,1) for the ranks and =SUMIF(diff_range,">0",rank_range) for W+. In Minitab: make a column of differences and choose Stat > Nonparametric > 1-Sample Wilcoxon with a test median of 0.

Kruskal-Wallis Test: Three or More Groups

Kruskal-Wallis is the rank version of one-way ANOVA. Setup times on three machines were recorded for six runs each.

Setup times (min) on three machines, with ranks across all 18 values
RunXRankYRankZRank
11222010166
21552515199
31112212177
418830173518
514427162111
613324142313
Rank sum238464
H = 12 / (N(N + 1)) × Σ Ri² / ni − 3(N + 1)

With N = 18, rank sums 23, 84, and 64, and six values per machine: H = 12 / (18 × 19) × (23² + 84² + 64²) / 6 − 3 × 19 = 11.31. Compare with the chi-square distribution on k − 1 = 2 degrees of freedom: p = 0.0035. Machine setup times are not all alike. (One-way ANOVA on the same data gives F = 8.48, p = 0.0034, the same conclusion.)

As with ANOVA, a significant result does not say which machines differ. Follow up with Mann-Whitney tests on each pair and adjust the p-values for the number of comparisons (Bonferroni: multiply each by 3), or use Dunn’s test where available.

Mood’s median test is a simpler alternative. It counts how many values in each group fall above the overall median (19.5) and tests the counts with chi-square. Here it gives p = 0.002. It is more robust to extreme outliers but has less power than Kruskal-Wallis.

In Excel: rank with =RANK.AVG, sum the ranks per machine with =SUMIF, calculate H as above, and get the p-value with =CHISQ.DIST.RT(H,2). In Minitab: Stat > Nonparametric > Kruskal-Wallis (Response and Factor columns), or Stat > Nonparametric > Mood's Median Test.

Assumptions and Limits

Assumption or issueWhy it mattersWhat to do
Independent observationsEvery test needs itThink about how the data were collected
Similar shape in the groupsNeeded to read the result as a difference in medians; if shapes differ, the test shows only that one group tends to be higherCompare the dot plots or box plots
Many tied valuesTies reduce the information in the ranksUse the version adjusted for ties; improve measurement resolution
Very small samplesLittle power; the smallest possible p-value may be above 0.05Use exact p-values; collect more data (two groups of 3 can never give p below 0.1)
The question is about the meanRank tests compare the whole distributions or the medians, not the meansState the question as medians, or use a transformation and a t-test
Estimates and intervalsRanks do not give means and standard errorsReport medians and the Hodges-Lehmann interval

Reading and Reporting

  1. Name the test and why: “Because repair times were skewed, a Mann-Whitney test was used.”
  2. Report the medians, not the means, with the estimated difference and its interval.
  3. Report the statistic and the p-value (W, U, H, or the signed-rank sum).
  4. Check the shape assumption before you call it a difference in medians.
  5. Say what it means in practice, in the units of the data.
A sentence you can use. Repair time was shorter with Method A than with Method B (median 14.5 h versus 27.0 h; estimated difference -13.0 h, 95% CI -24 to -5 h; Mann-Whitney W = 44, p = 0.008).

Common Mistakes

MistakeWhy it misleadsBetter
Using a rank test on large, near-normal samplesLoses a little power for no gainUse the t-test or ANOVA, after checking the plots
Reporting means with a rank testThe test does not compare meansReport medians
Choosing the test after seeing which gives p < 0.05Inflates false alarmsDecide from the data type and shape before testing
Ignoring tiesThe p-value can be offUse the tie-adjusted version
Comparing groups with different shapes as if only the medians differedA significant result may be about spread or shapePlot the groups; describe the difference honestly
Many pairwise rank tests with no adjustmentFalse alarms accumulateAdjust, or use Kruskal-Wallis first
Assuming nonparametric means “no assumptions”Independence still mattersCheck independence and shape

Try It Yourself

Customer waiting times (minutes) were recorded for two service desks: Desk 1: 4, 6, 3, 8, 5. Desk 2: 9, 12, 7, 15, 10, 11.

  • Rank all 11 values together and find the rank sum for Desk 1.
  • Is Desk 1 faster at α = 0.05? Use the exact p-value for groups this small.
Show the answer

Ranks of Desk 1: 2, 4, 1, 6, 3. The rank sum is 16 (the minimum possible is 15). The exact two-sided p-value is 0.0087, so Desk 1 is significantly faster at the 5% level: the median waiting time was 5 minutes at Desk 1 and 10.5 at Desk 2.

With groups this small, the exact p-value matters. Normal approximations can be inaccurate when both groups have fewer than about 8 values.

Nonparametric Tests: Frequently Asked Questions

What is the difference between Mann-Whitney and Wilcoxon rank-sum?

They are the same test, described with different statistics. Mann-Whitney is often reported with U, and Wilcoxon rank-sum with W, which is the sum of the ranks of one group. Minitab calls it Mann-Whitney and reports W. The Wilcoxon signed-rank test is a different test, for paired data.

Do nonparametric tests compare medians?

Only if the groups have a similar shape. In general they test whether values in one group tend to be larger than in the other. With similarly shaped distributions, that is the same as a difference in medians. Plot the groups first.

Are nonparametric tests less powerful?

Slightly, when the data really are normal: the Mann-Whitney test needs about 5% more data than the t-test for the same power. When the data are skewed or have outliers, they can be more powerful than the t-test.

What sample size do I need?

The same planning logic applies: choose the smallest difference that matters and size the sample for it. A rough approach is to calculate the sample for the t-test and add about 15%. Very small groups cannot give a significant result however large the difference.

What if I have only a few values?

Use exact p-values, which Minitab gives for small samples with no ties. Remember that with groups of 3 and 3, the smallest possible p-value is 0.10, so a significant result is impossible.

When should I use the sign test?

The sign test counts only which way each difference goes and ignores their size. It needs the fewest assumptions but has the least power. Use it when the data are only ordinal or the size of the differences is not meaningful.

Sources and Further Reading

  • NIST/SEMATECH, e-Handbook of Statistical Methods, sections on nonparametric tests (itl.nist.gov/div898/handbook).
  • Myles Hollander, Douglas A. Wolfe, and Eric Chicken, Nonparametric Statistical Methods, Wiley.
  • W. J. Conover, Practical Nonparametric Statistics, Wiley.
  • Frank Wilcoxon, “Individual comparisons by ranking methods,” Biometrics Bulletin, 1945; Henry B. Mann and Donald R. Whitney, Annals of Mathematical Statistics, 1947; William H. Kruskal and W. Allen Wallis, Journal of the American Statistical Association, 1952.
  • Minitab Support, “Methods and formulas for Mann-Whitney, 1-Sample Wilcoxon, and Kruskal-Wallis” (support.minitab.com).

This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.