Question it answers
Are two categories related, or do counts fit a stated pattern?
Data needed
Counts in categories (not percentages)
Key output
χ², degrees of freedom, p-value, Cramér’s V, residuals
Null hypothesis
Independence (association test) or the stated proportions (fit test)
Assumption
Expected counts of at least 5 in most cells; independent units
Excel
CHISQ.TEST, CHISQ.DIST.RT, PivotTable
Minitab
Stat > Tables > Chi-Square Test for Association; Goodness-of-Fit
If counts are small
Fisher’s exact test

The Idea in Plain Language

Measurements have averages and spreads. Categories have counts. When your data are counts in categories (defect types, shifts, suppliers, weekdays), the chi-square test compares the counts you observed with the counts you would expect if a stated idea were true. The more the two differ, the larger the chi-square statistic, and the harder the idea is to believe.

There are two main uses, and the calculation is the same:

TestQuestionThe idea being tested
Test of association (independence)Are two categorical variables related? Does the defect type depend on the shift?The variables are independent: the pattern of one is the same in every level of the other
Goodness of fitDo the counts match a stated pattern? Are late deliveries spread evenly over the week?The counts follow the stated proportions
χ² = Σ (Observed − Expected)² / Expected
Counts, not percentages. The chi-square test needs the actual counts, because the evidence depends on how many units you have. 30% of 10 and 30% of 1,000 are very different amounts of information.

When to Use It

Your situationUseWhy
Two categorical variables, counts in a tableChi-square test of associationTests whether they are related
One categorical variable against stated proportionsChi-square goodness of fitTests whether counts match the pattern
A 2 × 2 tableSame test as the two-proportion z testχ² = z²
Small expected counts (below about 5)Fisher’s exact test, or combine categoriesThe chi-square approximation fails
Is the data from a particular distribution (normal, Poisson)?Goodness of fit, or a normality testCompare observed counts in bins with the distribution’s
A measured responset-test, ANOVA, or regressionCounts need chi-square; measurements need the others

How It Works

  1. Observed counts go in a table.
  2. Expected counts. For association: expected = (row total × column total) / grand total. For goodness of fit: expected = total × the stated proportion.
  3. Contribution of each cell = (O − E)² / E. Add them to get χ².
  4. Degrees of freedom. Association: (rows − 1)(columns − 1). Goodness of fit: categories − 1 (minus one for each parameter estimated from the data).
  5. Compare χ² with the chi-square distribution, and read the p-value from the upper tail.
AssumptionWhyIf it fails
Counts of independent observations; each unit in one cellThe formula assumes itRedesign the sample; do not count the same unit twice
Expected counts at least 5 in most cells (none below 1)The chi-square approximation needs itCombine categories, collect more data, or use an exact test
Categories chosen before looking at the dataChoosing them after hides false alarmsDefine categories in advance

The statistic says whether there is an association, not how strong it is. For strength use Cramér’s V, which runs from 0 (none) to 1 (perfect): V = √(χ² / (n × (k − 1))), where k is the smaller of the number of rows and columns. As a rough guide for small tables, 0.1 is weak, 0.3 moderate, and 0.5 strong. To see where the association lives, look at the standardized residuals: values beyond about ± 2 mark cells that contribute.

Worked Example 1: Does the Defect Type Depend on the Shift?

A plant recorded 280 defects by type and shift. Is the mix of defect types the same on every shift?

Observed defect counts
Defect typeDayEveningNightTotal
Scratch40352095
Dent22303890
Discoloration18255295
Total8090110280
40 22 18 Day 35 30 25 Evening 20 38 52 Night Scratch Dent Discoloration Number of defects
The pattern differs by shift: scratches dominate the day shift, discoloration the night shift.
  1. Hypotheses. H0: defect type and shift are independent. H1: they are associated. α = 0.05.
  2. Expected counts. For Scratch on Day: (row total 95 × column total 80) / 280 = 27.14.
  3. Check the assumption: the smallest expected count is 25.71, well above 5.
  4. Add the contributions to get χ² = 25.41.
  5. Degrees of freedom = (3 − 1)(3 − 1) = 4. The critical value at 0.05 is 9.49, and the p-value is 0.00004.
  6. Strength. Cramér’s V = 0.21, a weak to moderate association.
Expected counts if defect type were independent of shift
Defect typeDayEveningNight
Scratch27.1430.5437.32
Dent25.7128.9335.36
Discoloration27.1430.5437.32
Contribution of each cell, (O − E)² / E
Defect typeDayEveningNight
Scratch6.090.658.04
Dent0.540.040.20
Discoloration3.081.005.77
χ² total25.41
0 5 10 15 20 25 30 Critical = 9.49 Observed = 25.4 Chi-square statistic
The observed statistic is far beyond the critical value, so the association is very unlikely to be chance.
Standardized residuals (beyond ± 2 is notable)
Defect typeDayEveningNight
Scratch+3.59+1.21-4.48
Dent-1.05+0.29+0.69
Discoloration-2.55-1.50+3.79
Conclusion. Defect type and shift are associated (χ²(4) = 25.4, p < 0.001, n = 280, V = 0.21). The standardized residuals show where: scratches are over-represented on the day shift (+3.6) and discoloration on the night shift (+3.8), while dents are spread about evenly. That points to different causes on different shifts, which is where the investigation should go.

Worked Example 2: Goodness of Fit

130 late deliveries were logged over several weeks. If lateness had no connection with the weekday, the five weekdays would each have about 26. Do the counts fit that?

31 26 Mon 18 26 Tue 20 26 Wed 22 26 Thu 39 26 Fri Observed Expected if even Late deliveries
Monday and Friday stand out against the even pattern.
Late deliveries by weekday
DayObservedExpected if even(O − E)² / E
Mon31260.962
Tue18262.462
Wed20261.385
Thu22260.615
Fri39266.500
Total13013011.923
  1. Hypotheses. H0: late deliveries are spread evenly over the five days. H1: they are not.
  2. Expected = 130 × 0.2 = 26 per day. The contributions add to χ² = 11.92.
  3. Degrees of freedom = 5 − 1 = 4, and the p-value is 0.0179.
Conclusion. Late deliveries are not evenly spread over the week (χ²(4) = 11.92, p = 0.018). Friday (39) and Monday (31) are well above the 26 expected, and Tuesday (18) is low. Look at weekend dispatch and Friday cut-off practices.

The same test checks whether data follow a distribution. Put the data into bins, calculate the expected count in each bin from the distribution (for example, the normal with the sample mean and standard deviation), and subtract one degree of freedom for each parameter estimated from the data.

Run It in Excel and Minitab

ExcelStep by step

  1. Enter the observed counts in a block (B2:D4 for the example), with row and column totals computed by =SUM().
  2. Expected counts: in a second block of the same shape, =($E2*B$5)/$E$5 filled across and down (row total × column total / grand total).
  3. p-value: =CHISQ.TEST(B2:D4, B8:D10) returns 0.000042 for the association example. It uses the correct degrees of freedom.
  4. Statistic: =SUMPRODUCT((B2:D4-B8:D10)^2/B8:D10). Critical value: =CHISQ.INV.RT(0.05, df).
  5. Goodness of fit: enter the observed counts and the expected counts, then =CHISQ.TEST(observed, expected). The degrees of freedom are categories − 1, so adjust by hand if you estimated parameters.
  6. Counts from raw data: build the table with a PivotTable (Insert > PivotTable, rows and columns as the two categories, values as the count).

Excel gives the p-value but not Cramér’s V or standardized residuals; calculate them from the formulas on this page.

MinitabStep by step

  1. Association from a summary table: enter the counts in a block of columns, then choose Stat > Tables > Chi-Square Test for Association and select Summarized data in a two-way table. Pick the columns that hold the counts.
  2. Association from raw data: choose Raw data (categorical variables), and set the Row and Column variables.
  3. Click Statistics and tick Chi-square analysis, Expected cell counts, Contribution to chi-square, and Standardized residuals.
  4. Goodness of fit: Stat > Tables > Chi-Square Goodness-of-Fit Test (One Variable). Choose the observed counts column, and enter the test proportions (equal by default) or a column of them.
  5. Read the Pearson chi-square, DF, and P-Value. Minitab warns when expected counts are small.
  6. For 2 × 2 tables with small counts, also run Fisher’s exact test from Stat > Basic Statistics > 2 Proportions.

The Assistant (Assistant > Hypothesis Tests > Chi-Square Test) covers the association test with a power check and a plain-language summary.

Minitab session window: Chi-Square Test for Association (typed excerpt, simplified)
Chi-Square Test for Association: Defect, Shift

Rows: Defect   Columns: Shift

                     Day  Evening    Night     All
Scratch               40       35       20      95
  Expected         27.14    30.54    37.32
  Contribution      6.09     0.65     8.04
Dent                  22       30       38      90
  Expected         25.71    28.93    35.36
  Contribution      0.54     0.04     0.20
Discoloration         18       25       52      95
  Expected         27.14    30.54    37.32
  Contribution      3.08     1.00     5.77
All                   80       90      110     280

Cell Contents: Count, Expected count, Contribution to Chi-square

Chi-Square Test

                       Chi-Square  DF  P-Value
Pearson                    25.412   4    0.000
Likelihood Ratio           26.122   4    0.000
Minitab session window: Goodness-of-Fit (typed excerpt, simplified)
Chi-Square Goodness-of-Fit Test for Observed Counts in Variable: Late

Category  Observed  Test Prop  Expected  Contribution
Mon             31       0.20      26.0        0.9615
Tue             18       0.20      26.0        2.4615
Wed             20       0.20      26.0        1.3846
Thu             22       0.20      26.0        0.6154
Fri             39       0.20      26.0        6.5000

  N  DF  Chi-Sq  P-Value
130   4  11.9231   0.0179

Reading and Reporting

  1. Give the table of counts (or the percentages with the totals).
  2. Report χ², the degrees of freedom, n, and the p-value, and the effect size (Cramér’s V).
  3. Say where the association lives, using residuals or percentages by row.
  4. Check and say that the expected counts were large enough.
  5. Say what it means in practice, and what you would investigate.
A sentence you can use. Defect type was associated with shift (χ²(4) = 25.4, n = 280, p < 0.001, Cramér’s V = 0.21); scratches were over-represented on the day shift and discoloration on the night shift.

Common Mistakes

MistakeWhy it misleadsBetter
Using percentages instead of countsThe test needs the sample sizeUse the counts
Small expected countsThe p-value is unreliableCombine categories or use Fisher’s exact test
Counting the same unit in several cellsObservations are not independentOne unit, one cell
Reading a significant result as a strong associationA large sample makes a weak association significantReport Cramér’s V
Reading a significant result as causeAssociation is not causeInvestigate with a designed study
Choosing categories after seeing the dataIt manufactures patternsDefine categories in advance
Many tests on subsets of one tableFalse alarms accumulateTest the whole table; follow up with residuals
Forgetting to subtract degrees of freedom for estimated parameters in goodness of fitThe p-value is too largeSubtract one per parameter estimated

Try It Yourself

Customer complaints by region and type: Region A: 30 billing, 20 delivery. Region B: 15 billing, 35 delivery.

  • Test whether complaint type depends on region.
  • How strong is the association, and what would you check?
Show the answer

This is a 2 × 2 table. χ² = 9.09 (without continuity correction) on 1 df, p = 0.0026. Region A has relatively more billing complaints (60% against 30% in Region B). Cramér’s V = 0.30, a moderate association.

The same test is a two-proportion test of the billing share. Check whether a local process (billing system, invoice template) differs between the regions, and whether the regions have different customer mixes.

Chi-Square Tests: Frequently Asked Questions

What is the difference between a test of association and a test of independence?

They are the same test with two names. Both test whether two categorical variables are unrelated, using the chi-square statistic on a table of counts.

What if some expected counts are below 5?

The chi-square approximation becomes unreliable. Combine sparse categories (if it makes sense), collect more data, or use an exact test such as Fisher’s exact test, which Minitab supplies for 2 × 2 tables.

Does a significant chi-square show a strong relationship?

No. It shows that the variables are probably not independent. With a large sample, a very weak association is significant. Use Cramér’s V for strength, and the residuals for where it lies.

Why is chi-square always one-sided?

Because deviations in any direction make the statistic larger, and large values are the evidence against the null. The p-value is the upper tail only.

Can I use chi-square on percentages?

No. Convert back to counts. The same percentages from a smaller sample would carry less evidence, and the test must reflect that.

How is this related to logistic regression?

Chi-square tests whether two categories are related. Logistic regression goes further: it models the probability of an outcome from several predictors, which can be numeric or categorical, and gives odds ratios.

Sources and Further Reading

  • Alan Agresti, Categorical Data Analysis, Wiley.
  • NIST/SEMATECH, e-Handbook of Statistical Methods, sections on chi-square tests (itl.nist.gov/div898/handbook).
  • Karl Pearson, “On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling,” Philosophical Magazine, 1900.
  • Douglas C. Montgomery and George C. Runger, Applied Statistics and Probability for Engineers, Wiley.
  • Minitab Support, “Methods and formulas for Chi-Square Test for Association” (support.minitab.com).

This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.