- Question it answers
- Are two categories related, or do counts fit a stated pattern?
- Data needed
- Counts in categories (not percentages)
- Key output
- χ², degrees of freedom, p-value, Cramér’s V, residuals
- Null hypothesis
- Independence (association test) or the stated proportions (fit test)
- Assumption
- Expected counts of at least 5 in most cells; independent units
- Excel
- CHISQ.TEST, CHISQ.DIST.RT, PivotTable
- Minitab
- Stat > Tables > Chi-Square Test for Association; Goodness-of-Fit
- If counts are small
- Fisher’s exact test
The Idea in Plain Language
Measurements have averages and spreads. Categories have counts. When your data are counts in categories (defect types, shifts, suppliers, weekdays), the chi-square test compares the counts you observed with the counts you would expect if a stated idea were true. The more the two differ, the larger the chi-square statistic, and the harder the idea is to believe.
There are two main uses, and the calculation is the same:
| Test | Question | The idea being tested |
|---|---|---|
| Test of association (independence) | Are two categorical variables related? Does the defect type depend on the shift? | The variables are independent: the pattern of one is the same in every level of the other |
| Goodness of fit | Do the counts match a stated pattern? Are late deliveries spread evenly over the week? | The counts follow the stated proportions |
When to Use It
| Your situation | Use | Why |
|---|---|---|
| Two categorical variables, counts in a table | Chi-square test of association | Tests whether they are related |
| One categorical variable against stated proportions | Chi-square goodness of fit | Tests whether counts match the pattern |
| A 2 × 2 table | Same test as the two-proportion z test | χ² = z² |
| Small expected counts (below about 5) | Fisher’s exact test, or combine categories | The chi-square approximation fails |
| Is the data from a particular distribution (normal, Poisson)? | Goodness of fit, or a normality test | Compare observed counts in bins with the distribution’s |
| A measured response | t-test, ANOVA, or regression | Counts need chi-square; measurements need the others |
How It Works
- Observed counts go in a table.
- Expected counts. For association: expected = (row total × column total) / grand total. For goodness of fit: expected = total × the stated proportion.
- Contribution of each cell = (O − E)² / E. Add them to get χ².
- Degrees of freedom. Association: (rows − 1)(columns − 1). Goodness of fit: categories − 1 (minus one for each parameter estimated from the data).
- Compare χ² with the chi-square distribution, and read the p-value from the upper tail.
| Assumption | Why | If it fails |
|---|---|---|
| Counts of independent observations; each unit in one cell | The formula assumes it | Redesign the sample; do not count the same unit twice |
| Expected counts at least 5 in most cells (none below 1) | The chi-square approximation needs it | Combine categories, collect more data, or use an exact test |
| Categories chosen before looking at the data | Choosing them after hides false alarms | Define categories in advance |
The statistic says whether there is an association, not how strong it is. For strength use Cramér’s V, which runs from 0 (none) to 1 (perfect): V = √(χ² / (n × (k − 1))), where k is the smaller of the number of rows and columns. As a rough guide for small tables, 0.1 is weak, 0.3 moderate, and 0.5 strong. To see where the association lives, look at the standardized residuals: values beyond about ± 2 mark cells that contribute.
Worked Example 1: Does the Defect Type Depend on the Shift?
A plant recorded 280 defects by type and shift. Is the mix of defect types the same on every shift?
| Defect type | Day | Evening | Night | Total |
|---|---|---|---|---|
| Scratch | 40 | 35 | 20 | 95 |
| Dent | 22 | 30 | 38 | 90 |
| Discoloration | 18 | 25 | 52 | 95 |
| Total | 80 | 90 | 110 | 280 |
- Hypotheses. H0: defect type and shift are independent. H1: they are associated. α = 0.05.
- Expected counts. For Scratch on Day: (row total 95 × column total 80) / 280 = 27.14.
- Check the assumption: the smallest expected count is 25.71, well above 5.
- Add the contributions to get χ² = 25.41.
- Degrees of freedom = (3 − 1)(3 − 1) = 4. The critical value at 0.05 is 9.49, and the p-value is 0.00004.
- Strength. Cramér’s V = 0.21, a weak to moderate association.
| Defect type | Day | Evening | Night |
|---|---|---|---|
| Scratch | 27.14 | 30.54 | 37.32 |
| Dent | 25.71 | 28.93 | 35.36 |
| Discoloration | 27.14 | 30.54 | 37.32 |
| Defect type | Day | Evening | Night |
|---|---|---|---|
| Scratch | 6.09 | 0.65 | 8.04 |
| Dent | 0.54 | 0.04 | 0.20 |
| Discoloration | 3.08 | 1.00 | 5.77 |
| χ² total | 25.41 | ||
| Defect type | Day | Evening | Night |
|---|---|---|---|
| Scratch | +3.59 | +1.21 | -4.48 |
| Dent | -1.05 | +0.29 | +0.69 |
| Discoloration | -2.55 | -1.50 | +3.79 |
Worked Example 2: Goodness of Fit
130 late deliveries were logged over several weeks. If lateness had no connection with the weekday, the five weekdays would each have about 26. Do the counts fit that?
| Day | Observed | Expected if even | (O − E)² / E |
|---|---|---|---|
| Mon | 31 | 26 | 0.962 |
| Tue | 18 | 26 | 2.462 |
| Wed | 20 | 26 | 1.385 |
| Thu | 22 | 26 | 0.615 |
| Fri | 39 | 26 | 6.500 |
| Total | 130 | 130 | 11.923 |
- Hypotheses. H0: late deliveries are spread evenly over the five days. H1: they are not.
- Expected = 130 × 0.2 = 26 per day. The contributions add to χ² = 11.92.
- Degrees of freedom = 5 − 1 = 4, and the p-value is 0.0179.
The same test checks whether data follow a distribution. Put the data into bins, calculate the expected count in each bin from the distribution (for example, the normal with the sample mean and standard deviation), and subtract one degree of freedom for each parameter estimated from the data.
Run It in Excel and Minitab
ExcelStep by step
- Enter the observed counts in a block (B2:D4 for the example), with row and column totals computed by .
- Expected counts: in a second block of the same shape, filled across and down (row total × column total / grand total).
- p-value: returns 0.000042 for the association example. It uses the correct degrees of freedom.
- Statistic: . Critical value: .
- Goodness of fit: enter the observed counts and the expected counts, then . The degrees of freedom are categories − 1, so adjust by hand if you estimated parameters.
- Counts from raw data: build the table with a PivotTable (, rows and columns as the two categories, values as the count).
Excel gives the p-value but not Cramér’s V or standardized residuals; calculate them from the formulas on this page.
MinitabStep by step
- Association from a summary table: enter the counts in a block of columns, then choose and select Summarized data in a two-way table. Pick the columns that hold the counts.
- Association from raw data: choose Raw data (categorical variables), and set the Row and Column variables.
- Click Statistics and tick Chi-square analysis, Expected cell counts, Contribution to chi-square, and Standardized residuals.
- Goodness of fit: . Choose the observed counts column, and enter the test proportions (equal by default) or a column of them.
- Read the Pearson chi-square, DF, and P-Value. Minitab warns when expected counts are small.
- For 2 × 2 tables with small counts, also run Fisher’s exact test from .
The Assistant () covers the association test with a power check and a plain-language summary.
Chi-Square Test for Association: Defect, Shift
Rows: Defect Columns: Shift
Day Evening Night All
Scratch 40 35 20 95
Expected 27.14 30.54 37.32
Contribution 6.09 0.65 8.04
Dent 22 30 38 90
Expected 25.71 28.93 35.36
Contribution 0.54 0.04 0.20
Discoloration 18 25 52 95
Expected 27.14 30.54 37.32
Contribution 3.08 1.00 5.77
All 80 90 110 280
Cell Contents: Count, Expected count, Contribution to Chi-square
Chi-Square Test
Chi-Square DF P-Value
Pearson 25.412 4 0.000
Likelihood Ratio 26.122 4 0.000Chi-Square Goodness-of-Fit Test for Observed Counts in Variable: Late
Category Observed Test Prop Expected Contribution
Mon 31 0.20 26.0 0.9615
Tue 18 0.20 26.0 2.4615
Wed 20 0.20 26.0 1.3846
Thu 22 0.20 26.0 0.6154
Fri 39 0.20 26.0 6.5000
N DF Chi-Sq P-Value
130 4 11.9231 0.0179Reading and Reporting
- Give the table of counts (or the percentages with the totals).
- Report χ², the degrees of freedom, n, and the p-value, and the effect size (Cramér’s V).
- Say where the association lives, using residuals or percentages by row.
- Check and say that the expected counts were large enough.
- Say what it means in practice, and what you would investigate.
Common Mistakes
| Mistake | Why it misleads | Better |
|---|---|---|
| Using percentages instead of counts | The test needs the sample size | Use the counts |
| Small expected counts | The p-value is unreliable | Combine categories or use Fisher’s exact test |
| Counting the same unit in several cells | Observations are not independent | One unit, one cell |
| Reading a significant result as a strong association | A large sample makes a weak association significant | Report Cramér’s V |
| Reading a significant result as cause | Association is not cause | Investigate with a designed study |
| Choosing categories after seeing the data | It manufactures patterns | Define categories in advance |
| Many tests on subsets of one table | False alarms accumulate | Test the whole table; follow up with residuals |
| Forgetting to subtract degrees of freedom for estimated parameters in goodness of fit | The p-value is too large | Subtract one per parameter estimated |
Try It Yourself
Customer complaints by region and type: Region A: 30 billing, 20 delivery. Region B: 15 billing, 35 delivery.
- Test whether complaint type depends on region.
- How strong is the association, and what would you check?
Show the answer
This is a 2 × 2 table. χ² = 9.09 (without continuity correction) on 1 df, p = 0.0026. Region A has relatively more billing complaints (60% against 30% in Region B). Cramér’s V = 0.30, a moderate association.
The same test is a two-proportion test of the billing share. Check whether a local process (billing system, invoice template) differs between the regions, and whether the regions have different customer mixes.
Chi-Square Tests: Frequently Asked Questions
What is the difference between a test of association and a test of independence?
They are the same test with two names. Both test whether two categorical variables are unrelated, using the chi-square statistic on a table of counts.
What if some expected counts are below 5?
The chi-square approximation becomes unreliable. Combine sparse categories (if it makes sense), collect more data, or use an exact test such as Fisher’s exact test, which Minitab supplies for 2 × 2 tables.
Does a significant chi-square show a strong relationship?
No. It shows that the variables are probably not independent. With a large sample, a very weak association is significant. Use Cramér’s V for strength, and the residuals for where it lies.
Why is chi-square always one-sided?
Because deviations in any direction make the statistic larger, and large values are the evidence against the null. The p-value is the upper tail only.
Can I use chi-square on percentages?
No. Convert back to counts. The same percentages from a smaller sample would carry less evidence, and the test must reflect that.
How is this related to logistic regression?
Chi-square tests whether two categories are related. Logistic regression goes further: it models the probability of an outcome from several predictors, which can be numeric or categorical, and gives odds ratios.
Sources and Further Reading
- Alan Agresti, Categorical Data Analysis, Wiley.
- NIST/SEMATECH, e-Handbook of Statistical Methods, sections on chi-square tests (itl.nist.gov/div898/handbook).
- Karl Pearson, “On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling,” Philosophical Magazine, 1900.
- Douglas C. Montgomery and George C. Runger, Applied Statistics and Probability for Engineers, Wiley.
- Minitab Support, “Methods and formulas for Chi-Square Test for Association” (support.minitab.com).
This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.