Types of Data and Variables
Categorical and numerical variables, their subtypes, and why the type of a variable decides which summaries and plots make sense.
Prerequisites: Introduction to Statistics.
A variable is a characteristic that can differ from one individual or observation to the next: a person’s age, a country’s currency, the outcome of a coin toss. Variables come in different types, and the type determines which calculations are meaningful. Averaging postcodes, or ranking eye colours, produces numbers but not information.
Recognizing the type of each variable is usually the first step in any analysis, because it decides which summaries, plots, and models are appropriate.
Intuition
Consider a small table about students:
| Student | Programme | Year of study | Satisfaction (1–5) | Courses taken | Height (cm) |
|---|---|---|---|---|---|
| A | Physics | 2 | 4 | 5 | 172.4 |
| B | History | 1 | 2 | 4 | 165.0 |
| C | Physics | 3 | 5 | 6 | 181.9 |
Every column holds a value per student, but the columns behave differently.
- Programme is a label. There is no order, and no arithmetic makes sense.
- Satisfaction has an order (4 is better than 2), but is the step from 2 to 3 the same size as the step from 4 to 5? The scale does not say.
- Courses taken is a count: whole numbers, and differences are meaningful.
- Height is a measurement that could, in principle, take any value in a range.
These differences are what the classification below makes precise.
Categorical and numerical variables
The first split is between variables whose values are categories and variables whose values are numbers that measure or count something.
Categorical variables
A categorical (or qualitative) variable places each observation into one of a set of groups.
- Nominal variables have categories with no natural order: blood type, programme of study, country, colour. A nominal variable with exactly two categories (yes/no, success/failure) is called binary; binary variables are modelled by the Bernoulli distribution.
- Ordinal variables have categories with a meaningful order but no fixed distance between them: education level, a pain scale from “none” to “severe”, or survey answers from “strongly disagree” to “strongly agree”.
Categories are sometimes stored as numbers, such as 1 for “male” and 2 for “female”, or postcodes. The numbers are codes, not quantities. Software will happily compute their mean, but the result means nothing.
Numerical variables
A numerical (or quantitative) variable takes values that are numbers on a meaningful scale, so that differences between values make sense.
- Discrete variables can take only separate, countable values, typically whole numbers: the number of children in a family, the number of emails received in an hour, the number of heads in ten coin tosses.
- Continuous variables can take any value in an interval, at least in principle: height, time, temperature, weight.
In practice every measurement is rounded, so recorded “continuous” data are technically discrete (heights to the nearest millimetre). The distinction is about the underlying quantity and about which model is natural. It matters in probability: a discrete variable is described by a probability mass function that gives the probability of each value, while a continuous variable is described by a density, and the probability of any single exact value is zero. See random variables.
Interval and ratio scales
Numerical variables can be further divided by whether zero has a real meaning.
- On an interval scale, differences are meaningful but zero is arbitrary. Temperature in degrees Celsius is the standard example: 20 °C is 10 degrees warmer than 10 °C, but it is not “twice as hot”, because 0 °C is just the freezing point of water, not an absence of heat. Calendar years are another example.
- On a ratio scale, zero means “none of the quantity”, so ratios are meaningful as well. A 4 kg parcel weighs twice as much as a 2 kg one; 0 seconds is no time at all. Height, weight, income, counts, and temperature in kelvin are ratio variables.
The practical consequence is that statements like “twice as large” or “a 50% increase” only make sense on a ratio scale.
Why the type matters
The type of a variable determines which summaries and plots are meaningful. The table shows the usual choices; each row allows everything in the rows above it, plus more.
| Type | Examples | Meaningful operations | Centre | Spread | Typical plots |
|---|---|---|---|---|---|
| Nominal | blood type, country | count, equal / not equal | mode | (proportions in each category) | bar chart |
| Ordinal | satisfaction rating, education level | + order | mode, median | range, interquartile range | bar chart (in order) |
| Interval | °C, calendar year | + differences | mean, median | standard deviation, IQR | histogram, box plot |
| Ratio | height, income, counts | + ratios | mean, median | standard deviation, IQR | histogram, box plot |
For example, the median needs the values to be ordered, so it works for ordinal data but not nominal data. The mean and standard deviation need differences between values to be meaningful, so strictly they need an interval or ratio scale. Quantiles, like the median, need only an order.
Plot choice follows the same logic. A bar chart shows how many observations fall in each category and works for any categorical variable. A histogram groups numerical values into intervals along a number line, which only makes sense when the values lie on one. See descriptive statistics.
Worked example
A survey records, for each of 200 people, their mode of transport to work, their satisfaction with it on a 1–5 scale, their number of trips last week, and their typical commute time in minutes.
- Mode of transport (car, bus, bicycle, walk) is nominal. Summarize it with counts or percentages in each category and a bar chart. The most common category is the mode. A “mean mode of transport” has no meaning.
- Satisfaction (1–5) is ordinal. The median and the percentage answering 4 or 5 are safe summaries. A mean satisfaction of 3.6 is common in practice and can be useful for comparisons, but it assumes the steps between levels are equal, which the scale does not guarantee. State that assumption if you use it.
- Number of trips is discrete numerical on a ratio scale. Mean, median, and standard deviation all make sense, and so does “twice as many trips”.
- Commute time is continuous numerical on a ratio scale. A histogram and a box plot show its shape; times are often right-skewed, so report the median alongside the mean.
Common misunderstandings
“If it is stored as a number, it is numerical.” Codes such as postcodes, phone numbers, or “1 = yes, 0 = no” are categorical. The test is whether arithmetic on the values means something.
“Ordinal data can be averaged like any other numbers.” Averaging ratings is widespread and often informative, but it treats the gaps between levels as equal. Two groups with the same mean rating can have very different distributions of answers, so look at the counts too.
“Discrete means categorical.” Counts are discrete but numerical: 4 courses really is twice as many as 2. Discrete and continuous are both subtypes of numerical data.
“The type is a fixed property of the data.” It partly depends on the question. Age in years is numerical, but age groups such as “18–29, 30–44, 45+” are ordinal. Grouping a numerical variable into categories throws information away, so do it only for a reason.
Further reading
- David M. Diez, Mine Çetinkaya-Rundel, and Christopher D. Barr, OpenIntro Statistics, 4th ed., 2019. Free online; the introductory chapter classifies variables with many examples.
- David Freedman, Robert Pisani, and Roger Purves, Statistics, 4th ed., W. W. Norton, 2007. Discusses qualitative and quantitative variables and how to display them.