Types of Data and Variables

Categorical and numerical variables, their subtypes, and why the type of a variable decides which summaries and plots make sense.

Prerequisites: Introduction to Statistics.

A variable is a characteristic that can differ from one individual or observation to the next: a person’s age, a country’s currency, the outcome of a coin toss. Variables come in different types, and the type determines which calculations are meaningful. Averaging postcodes, or ranking eye colours, produces numbers but not information.

Recognizing the type of each variable is usually the first step in any analysis, because it decides which summaries, plots, and models are appropriate.

Intuition

Consider a small table about students:

Student Programme Year of study Satisfaction (1–5) Courses taken Height (cm)
A Physics 2 4 5 172.4
B History 1 2 4 165.0
C Physics 3 5 6 181.9

Every column holds a value per student, but the columns behave differently.

These differences are what the classification below makes precise.

Categorical and numerical variables

The first split is between variables whose values are categories and variables whose values are numbers that measure or count something.

Categorical variables

A categorical (or qualitative) variable places each observation into one of a set of groups.

Categories are sometimes stored as numbers, such as 1 for “male” and 2 for “female”, or postcodes. The numbers are codes, not quantities. Software will happily compute their mean, but the result means nothing.

Numerical variables

A numerical (or quantitative) variable takes values that are numbers on a meaningful scale, so that differences between values make sense.

In practice every measurement is rounded, so recorded “continuous” data are technically discrete (heights to the nearest millimetre). The distinction is about the underlying quantity and about which model is natural. It matters in probability: a discrete variable is described by a probability mass function that gives the probability of each value, while a continuous variable is described by a density, and the probability of any single exact value is zero. See random variables.

Interval and ratio scales

Numerical variables can be further divided by whether zero has a real meaning.

The practical consequence is that statements like “twice as large” or “a 50% increase” only make sense on a ratio scale.

Why the type matters

The type of a variable determines which summaries and plots are meaningful. The table shows the usual choices; each row allows everything in the rows above it, plus more.

Type Examples Meaningful operations Centre Spread Typical plots
Nominal blood type, country count, equal / not equal mode (proportions in each category) bar chart
Ordinal satisfaction rating, education level + order mode, median range, interquartile range bar chart (in order)
Interval °C, calendar year + differences mean, median standard deviation, IQR histogram, box plot
Ratio height, income, counts + ratios mean, median standard deviation, IQR histogram, box plot

For example, the median needs the values to be ordered, so it works for ordinal data but not nominal data. The mean and standard deviation need differences between values to be meaningful, so strictly they need an interval or ratio scale. Quantiles, like the median, need only an order.

Plot choice follows the same logic. A bar chart shows how many observations fall in each category and works for any categorical variable. A histogram groups numerical values into intervals along a number line, which only makes sense when the values lie on one. See descriptive statistics.

Worked example

A survey records, for each of 200 people, their mode of transport to work, their satisfaction with it on a 1–5 scale, their number of trips last week, and their typical commute time in minutes.

  1. Mode of transport (car, bus, bicycle, walk) is nominal. Summarize it with counts or percentages in each category and a bar chart. The most common category is the mode. A “mean mode of transport” has no meaning.
  2. Satisfaction (1–5) is ordinal. The median and the percentage answering 4 or 5 are safe summaries. A mean satisfaction of 3.6 is common in practice and can be useful for comparisons, but it assumes the steps between levels are equal, which the scale does not guarantee. State that assumption if you use it.
  3. Number of trips is discrete numerical on a ratio scale. Mean, median, and standard deviation all make sense, and so does “twice as many trips”.
  4. Commute time is continuous numerical on a ratio scale. A histogram and a box plot show its shape; times are often right-skewed, so report the median alongside the mean.

Common misunderstandings

“If it is stored as a number, it is numerical.” Codes such as postcodes, phone numbers, or “1 = yes, 0 = no” are categorical. The test is whether arithmetic on the values means something.

“Ordinal data can be averaged like any other numbers.” Averaging ratings is widespread and often informative, but it treats the gaps between levels as equal. Two groups with the same mean rating can have very different distributions of answers, so look at the counts too.

“Discrete means categorical.” Counts are discrete but numerical: 4 courses really is twice as many as 2. Discrete and continuous are both subtypes of numerical data.

“The type is a fixed property of the data.” It partly depends on the question. Age in years is numerical, but age groups such as “18–29, 30–44, 45+” are ordinal. Grouping a numerical variable into categories throws information away, so do it only for a reason.

Further reading