Question it answers
What kind of data do I have, and what can I do with it?
Data needed
A data sheet or data collection plan
Key output
A type for each variable, and the methods that fit
Scales
Nominal, ordinal, interval, ratio
Families
Categorical or numerical; discrete or continuous
Excel
ISNUMBER, COUNTIF, PivotTable, Text to Columns
Minitab
Data > Change Data Type; Stat > Tables > Tally
Why it matters
The wrong method for the data type gives wrong answers

The Idea in Plain Language

Before you choose a chart, a test, or a control chart, ask one question: what kind of data is this? The answer decides almost everything that follows. You can average weights, but you cannot average the colors of cars. You can count defects per unit, but a count behaves differently from a measurement.

DataCategorical (attribute)labels or groupsNumerical (variable)measured or countedNominalnames, no orderOrdinalranked categoriesDiscretecounts: 0, 1, 2, ...Continuousany value in a rangeSupplier, defect typeRating 1 to 5, severityDefects per unitWeight, time, length
Data come in two families. Categorical data put things in groups; numerical data are numbers that mean something when you add or compare them.
The rule of thumb. Prefer continuous data when you can get it. A measurement carries far more information than a pass/fail judgment on the same part, so it needs a much smaller sample to see the same change.

The Four Scales of Measurement

ScaleWhat it tells youExampleFair summariesNot meaningful
NominalWhich groupSupplier, defect type, shift, machineCounts, percentages, modeAverage, order
OrdinalWhich group, and the orderSurvey rating 1 to 5, severity High/Med/LowMedian, percentiles, counts per levelGaps between levels, a plain average
IntervalOrder and equal gaps, no true zeroTemperature in °C, calendar dateMean, standard deviationRatios (40°C is not twice as hot as 20°C)
RatioOrder, equal gaps, and a true zeroWeight, time, length, pressure, moneyEverything above, plus ratios and the coefficient of variationNothing: all arithmetic is fair

In practice most process data are ratio (cycle time, weight) or nominal and ordinal (defect type, rating). Interval data are rare outside temperature and dates. The useful split for choosing methods is categorical versus numerical, and within numerical, discrete counts versus continuous measurements.

Why the Type Decides the Method

Data typeTypical questionSummaryChartTest or tool
Continuous, one groupWhat is the average and the spread?Mean, standard deviationHistogram, individuals chartt-test, interval for the mean
Continuous, several groupsDo the groups differ?Mean and SD by groupBox plot, dot plotANOVA, t-test
Continuous, two variablesDoes Y move with X?CorrelationScatter plotRegression
Discrete counts of defectsIs the defect rate changing?Defects per unitc or u chart, ParetoPoisson methods
Pass/failIs the proportion different?Proportion defectivep chart, bar chartTest for proportions
Two categories (nominal)Are they associated?Counts in a tableStacked barChi-square
Ordinal ratingsDo ratings differ between groups?Median, counts per levelBar chart of countsMann-Whitney, Kruskal-Wallis

The method selector on the Stat Dojo home page follows this same logic and gives the Excel and Minitab route for each choice.

Worked Example 1: The Trap of Averaging Ratings

Twenty operators rated a new work instruction from 1 (poor) to 5 (excellent). The ratings are ordinal: a 4 is better than a 3, but nobody can say the gap from 3 to 4 equals the gap from 1 to 2.

Rating12345
Number of operators22286
  • Mean = 3.70. It treats the ratings as if the gaps were equal.
  • Median = 4.0, the middle rating. It uses only the order, so it is safe for ordinal data.
  • Mode = 4, the most common answer (8 of 20).
  • Top-two-box = 70% of operators gave a 4 or 5. A share of operators is easy to explain and honest for ordinal data.
Conclusion. Report the counts, the median, and the share who rated 4 or 5. The mean of 3.70 is common in practice and often harmless, but it assumes equal gaps, so label it as an approximation. Never run a t-test on a few ordinal levels without thinking: use a nonparametric test instead.

Worked Example 2: Measuring Beats Sorting

Two lots of 25 shafts each were checked against an upper diameter limit of 10.0 mm. A go/no-go gauge says 25 of 25 pass in Lot A and 25 of 25 pass in Lot B. As attribute data, the lots look identical.

9.2 mm 9.3 mm 9.4 mm 9.5 mm 9.6 mm 9.7 mm 9.8 mm 9.9 mm 10 mm 10.1 mm 9.50 Lot A 9.85 Lot B Upper limit 10.00
The same two lots as measurements. Lot B runs much closer to the limit, which the pass/fail count cannot show.
  • Lot A: mean 9.50 mm, SD 0.102 mm, largest value 9.71.
  • Lot B: mean 9.85 mm, SD 0.059 mm, largest value 9.95.
Conclusion. The pass/fail data hide a shift of about 0.35 mm that will soon produce defects. Continuous data would have shown the drift long before the first reject. If you only have attribute data, you need many more parts (often 10 to 50 times as many) to see the same change.

Run It in Excel and Minitab

ExcelStep by step

  1. Check the type of each column. Numbers are right-aligned by default; text is left-aligned. =ISNUMBER(A2) returns TRUE for real numbers, which catches numbers stored as text.
  2. Fix numbers stored as text with Data > Text to Columns > Finish, or =VALUE(A2).
  3. Count categories with =COUNTIF(range, "Fail"), or build a Insert > PivotTable with the category in Rows and Count in Values.
  4. Keep ordinal codes consistent with a list: Data > Data Validation > List.
  5. Use the median for ordinal data: =MEDIAN(range).

MinitabStep by step

  1. Minitab marks the type in the column header: a column holding text shows -T after its name, a date column -D, and a numeric column nothing.
  2. Change the type with Data > Change Data Type > Numeric to Text or Data > Change Data Type > Text to Numeric. Use Data > Code > Numeric to Text to replace codes with labels.
  3. Count categories with Stat > Tables > Tally Individual Variables and tick Counts and Percents.
  4. Declare ordered text levels (Low, Medium, High) with Column Properties > Value Order from the column’s right-click menu, so charts list them in order.
  5. Summaries by type: Stat > Basic Statistics > Display Descriptive Statistics for numeric data, Tally for categorical data.
Minitab session window: Tally Individual Variables (typed excerpt)
Tally for Discrete Variables: Rating

Rating  Count  Percent
     1      2    10.00
     2      2    10.00
     3      2    10.00
     4      8    40.00
     5      6    30.00
N=20

Reading and Reporting

  1. State the type of each variable in the data plan before you collect anything.
  2. Choose the summary that fits the type: counts and percentages for categories, median for ordinal, mean and SD for continuous.
  3. Do not recode data downward (a measurement into pass/fail) unless the decision truly is pass/fail. Keep the measurement.
  4. Document the gauge or definition for attribute data so two people classify the same way.

Common Mistakes

MistakeWhy it misleadsBetter
Averaging nominal codes (1 = Plant A, 2 = Plant B)The numbers are labels; the average is meaninglessCount and compare percentages
Averaging a 1 to 5 rating without commentThe gaps between levels are not equalReport counts, median, and top-two-box
Turning a measurement into pass/fail for analysisThrows away most of the informationAnalyze the measurement; report pass/fail separately
Treating counts as continuousSmall counts are skewed and cannot go below zeroUse Poisson methods, or a c or u chart
Numbers stored as textSorts and summaries silently failConvert the type before analysis
No defined classification rule for attribute dataDifferent people judge differentlyWrite an operational definition and test it with an attribute agreement study

Try It Yourself

Classify each variable: (a) the number of scratches on a panel; (b) the color of a wire; (c) the satisfaction score from 1 to 7; (d) the time to resolve a ticket; (e) pass or fail on a leak test.

Show the answer

(a) Discrete numerical (a count). (b) Nominal. (c) Ordinal. (d) Continuous (ratio). (e) Categorical with two levels (nominal), often called binary or attribute data.

A c chart suits (a), a Pareto chart suits (b), a bar chart of counts suits (c), an individuals chart and histogram suit (d), and a p chart suits (e).

Data Types and Scales: Frequently Asked Questions

What is the difference between discrete and continuous data?

Discrete data are counts that can only take certain values, such as 0, 1, 2 defects. Continuous data can take any value in a range, limited only by how finely you can measure, such as weight or time.

Is a 1 to 5 rating scale continuous?

No. It is ordinal. The levels are ordered, but the gaps are not necessarily equal. Many practitioners average the ratings anyway, which is often acceptable for large samples, but the median and the share of top ratings are safer.

What is attribute data?

Attribute data are categories or counts, such as pass/fail or number of defects. Variable data are measurements. Control charts, capability studies, and tests differ between the two.

Why does the data type matter so much?

Each statistical method assumes a kind of data. A t-test needs numerical measurements, a chi-square test needs counts in categories, and a control chart is chosen by whether you measure or count.

Can I convert continuous data to categories?

You can, by grouping measurements into bands, but you lose information and power. Do it only to communicate a result, not to analyze.

Sources and Further Reading

  • S. S. Stevens, “On the Theory of Scales of Measurement,” Science, 1946.
  • NIST/SEMATECH, e-Handbook of Statistical Methods, Exploratory Data Analysis (itl.nist.gov/div898/handbook).
  • David S. Moore, George P. McCabe, and Bruce A. Craig, Introduction to the Practice of Statistics, Freeman.
  • Minitab Support, “Data types in Minitab” (support.minitab.com).

This content is educational. Worked examples use made-up data. Menu names for Minitab follow recent versions of Minitab Statistical Software and can differ slightly in older releases; Excel steps use Microsoft 365 and the Analysis ToolPak.