Mean, Median, and Mode
Three ways to describe the typical value of a dataset, how they differ, and when to use which.
Prerequisites: Descriptive Statistics.
The mean, median, and mode are three ways to summarize a dataset by a single “typical” value. The mean is the ordinary average; the median is the middle value once the data are sorted; the mode is the most common value. They are called measures of centre (or of central tendency).
They often give similar answers, but not always, and the differences are informative. When they disagree, the choice between them can change the story a set of numbers seems to tell.
Intuition
Each measure answers a slightly different question.
- The mean asks: if the total were shared out equally, how much would each observation get? It is the balance point of the data: imagine the values as equal weights on a ruler; the ruler balances at the mean.
- The median asks: which value splits the data into two equal halves? Half the observations are at or below it, half at or above.
- The mode asks: which value occurs most often?
The balance-point picture explains the most important difference. Moving one weight far out along the ruler shifts the balance point, however many other weights there are. The median only cares about which values are above and which below the middle, not how far away they are.
Definitions
Let be observed values.
The mean (or arithmetic mean, or average) is
For a whole population the mean is written ; the mean of a probability distribution is its expected value.
The median is found by sorting the values from smallest to largest:
- if is odd, the median is the middle value, in position ;
- if is even, it is the average of the two middle values, in positions and .
The mode is the value that occurs most often. A dataset can have two modes (bimodal), several, or effectively none if every value occurs once. For continuous measurements, where exact repeats are rare, “the mode” usually means the location of the highest peak of a histogram or density curve.
Which measures make sense depends on the type of data. The mode works for any variable, including categories such as “most common blood type”. The median needs ordered values. The mean needs numbers whose differences are meaningful.
Worked example
A small firm has seven employees with annual salaries, in thousands of euros:
The 180 is the owner’s salary.
Step 1: mean. The total is , so
Step 2: median. The values are already sorted and is odd, so the median is the 4th value: 38 thousand euros.
Step 3: mode. The value 35 occurs twice and every other value once, so the mode is 35 thousand euros.
Step 4: compare. Six of the seven employees earn less than the mean. The mean is not “typical” of anyone here; the single large salary has dragged it upward. The median, 38, is a much better description of what a typical employee earns.
Step 5: remove the outlier. Without the owner, the six remaining salaries have mean and median . Removing one value moved the mean by more than 20 thousand euros, but the median by only 1.5.
Robustness
A summary is robust if a few extreme or erroneous values cannot change it much. The median is robust: you could raise the three largest salaries above by any amounts at all, and the median would stay at 38. The mean is not: changing a single value by an amount changes the mean by , with no limit.
Robustness is not automatically good. The mean uses the size of every observation, which is exactly what you want when totals matter. If a city wants to know the total water use of its households, the mean (times the number of households) is the relevant figure, and the large users are a real part of it.
A useful way to see the difference: the mean is the number that makes the sum of squared distances as small as possible, while the median makes the sum of absolute distances as small as possible. Squaring gives large distances much more weight, which is why the mean chases outliers. The first fact is the starting point for variance and standard deviation.
Skew and the order of mean and median
In a symmetric distribution, the mean and median coincide (provided the mean exists), and if there is a single peak, the mode is there too. In a skewed distribution they separate.
In the right-skewed curve, the long tail of large values pulls the mean upwards while the median, which only counts how much area lies on each side, moves less. The mode sits at the peak. This is the typical pattern for incomes, house prices, and waiting times: mode < median < mean. For a left-skewed distribution the order is usually reversed.
This ordering is a useful heuristic, not a theorem. It holds for many common distributions but can fail, especially for discrete data or distributions with more than one peak. For example, the dataset
has a long right tail (one value far above the rest), yet its mean is below its median of 3. Among standard distributions, a Poisson distribution with mean 1.7 is right-skewed but has median 2. So use the gap between mean and median as a hint about skew, and confirm by looking at a histogram.
Which one to use
- Roughly symmetric numerical data without extreme values: the mean, usually reported with the standard deviation. It uses all the information and has convenient mathematical properties.
- Skewed data or data with outliers: the median, usually with the interquartile range. Median household income is reported for this reason.
- Totals matter: the mean, since total = mean × number of observations.
- Categorical data, or the most common value is what matters: the mode (the most common shoe size is what a shop should stock most of).
When in doubt, report both the mean and the median. If they differ a lot, say so and show a plot.
Common misunderstandings
“The average is what a typical member experiences.” Only if the distribution is roughly symmetric. In the salary example, six of seven employees earn below the mean.
“Mean > median proves the data are right-skewed.” It is a hint, not a proof; the example above shows the rule can fail.
“The median ignores the data.” It uses every observation to decide what is in the middle; it just ignores how far the outer values are from the centre.
“The mean of a sample is the population mean.” The sample mean is an estimate of the population mean and changes from sample to sample. How much it changes is described by its sampling distribution.
Further reading
- David Freedman, Robert Pisani, and Roger Purves, Statistics, 4th ed., W. W. Norton, 2007. Gives a clear, example-driven discussion of the average and the median, including their behaviour with skewed data.
- David M. Diez, Mine Çetinkaya-Rundel, and Christopher D. Barr, OpenIntro Statistics, 4th ed., 2019. Free online; covers measures of centre, skew, and robust statistics.