Quantiles, Percentiles, and Outliers

Values that cut ordered data into given proportions, the interquartile range, and how to find and handle unusual observations.

Prerequisites: Descriptive Statistics, Mean, Median, and Mode.

A quantile is a value that splits ordered data at a given proportion: the 0.9 quantile, also called the 90th percentile, is the value below which about 90% of the observations lie. The median is the 0.5 quantile. Quantiles describe a whole distribution by a few landmark points, and they are the basis of the box plot and of the most common rule for flagging outliers, observations that sit unusually far from the rest.

Quantiles are everywhere once you look: growth charts for children (“75th percentile for height”), exam results (“top 10%”), and service-level targets (“95% of requests answered within 200 milliseconds”) are all statements about quantiles.

Intuition

Sort the data from smallest to largest and walk along the list. The quantile for proportion pp is the point you have reached when a fraction pp of the data is behind you. Halfway along is the median; a quarter of the way is the first quartile; nine-tenths of the way is the 90th percentile.

Because quantiles depend only on the order of the values, they are not pulled around by extreme values the way the mean and standard deviation are. Except in the tiniest datasets, making the largest value ten times larger does not change the median or the quartiles at all. That robustness is why quantile-based summaries are preferred for skewed data.

Quantiles, percentiles, and quartiles

For a proportion pp between 0 and 1, the pp-quantile is a value qpq_p such that roughly a fraction pp of the data are at or below qpq_p and a fraction 1−p1 - p are at or above it. Named special cases:

The interquartile range is the width of the middle half of the data:

IQR=Q3−Q1.\text{IQR} = Q_3 - Q_1 .

It is a robust measure of spread, the natural companion of the median in the way that the standard deviation is the companion of the mean.

Quantiles of a distribution

For a probability distribution, quantiles are defined through the cumulative distribution function F(x)=P(X≤x)F(x) = P(X \le x) (see probability distributions). The pp-quantile is the smallest xx with F(x)≥pF(x) \ge p. For a continuous distribution with a strictly increasing FF, this is simply the xx that solves F(x)=pF(x) = p. For example, the 0.975 quantile of the standard normal distribution is about 1.96, the number behind the familiar “mean ± 1.96 standard deviations” interval.

Sample quantiles are a matter of convention

With a finite dataset there is usually no value that has exactly a fraction pp of the data below it, so a rule is needed to pick one or to interpolate between neighbours. Many reasonable rules exist, and different software uses different ones by default. A widely used rule, the default in R, in NumPy, and in Excel’s QUARTILE.INC, works like this for sorted data x(1)≤x(2)≤⋯≤x(n)x_{(1)} \le x_{(2)} \le \dots \le x_{(n)}:

  1. Compute the position h=(n−1) p+1h = (n - 1)\,p + 1.
  2. If hh is a whole number, qp=x(h)q_p = x_{(h)}.
  3. Otherwise, interpolate linearly between the two neighbouring sorted values. If h=k+fh = k + f with whole number kk and fraction 0<f<10 < f < 1, then qp=x(k)+f⋅(x(k+1)−x(k))q_p = x_{(k)} + f \cdot (x_{(k+1)} - x_{(k)}).

Other rules, such as Excel’s QUARTILE.EXC or the textbook method of taking the median of each half of the data, can give slightly different answers. The differences shrink as nn grows and rarely matter in practice, but they explain why two programs may report different quartiles for the same small dataset. When it matters, state which method you used.

Worked example

A commuter records the delays, in minutes, of 13 trains, sorted:

2,  3,  4,  5,  5,  6,  7,  8,  9,  10,  12,  14,  31.2,\; 3,\; 4,\; 5,\; 5,\; 6,\; 7,\; 8,\; 9,\; 10,\; 12,\; 14,\; 31.

Step 1: quartiles. With n=13n = 13, the rule above gives positions h=12p+1h = 12p + 1:

The interquartile range is IQR=10−5=5\text{IQR} = 10 - 5 = 5 minutes: the middle half of the delays span 5 minutes.

Step 2: a percentile that needs interpolation. For the 90th percentile, h=12×0.9+1=11.8h = 12 \times 0.9 + 1 = 11.8. This lies 0.8 of the way from x(11)=12x_{(11)} = 12 to x(12)=14x_{(12)} = 14, so q0.9=12+0.8×(14−12)=13.6q_{0.9} = 12 + 0.8 \times (14 - 12) = 13.6 minutes.

Step 3: a different convention. Excel’s QUARTILE.EXC gives Q1=4.5Q_1 = 4.5 and Q3=11Q_3 = 11 for the same data. Neither answer is wrong; they are different definitions.

Step 4: fences. Using the 1.5 × IQR rule described below, the fences are

Q1−1.5 IQR=5−7.5=−2.5,Q3+1.5 IQR=10+7.5=17.5.Q_1 - 1.5\,\text{IQR} = 5 - 7.5 = -2.5, \qquad Q_3 + 1.5\,\text{IQR} = 10 + 7.5 = 17.5 .

No delay is below −2.5-2.5, but 31 is above 17.5, so it is flagged as a possible outlier.

Step 5: how much does it matter? With the 31 included, the mean is about 8.9 minutes and the standard deviation about 7.5. Without it, they become about 7.1 and 3.7: one value doubled the standard deviation. The median moves only from 7 to 6.5, and the IQR (by the same rule) from 5 to 4.5.

Box plots and the 1.5 × IQR rule

A box plot draws the quartiles and flags unusual values. The version in common use today follows John Tukey’s exploratory data analysis:

An annotated box plot of 13 train delays. The box runs from Q1 = 5 to Q3 = 10 with the median line at 7, so the IQR is 5. Dashed fences sit at −2.5 and 17.5. Whiskers extend to 2 and 14, the most extreme data inside the fences. The value 31 lies beyond the upper fence and is drawn as a separate outlier point. The raw data are shown as dots underneath.
Box plot of the train delays 2, 3, 4, 5, 5, 6, 7, 8, 9, 10, 12, 14, 31 minutes, with quartiles by linear interpolation. Fences at Q1 − 1.5 × IQR and Q3 + 1.5 × IQR are dashed; fences are normally not drawn in a box plot and are shown here only for explanation.

The factor 1.5 is a convention, chosen because it works reasonably in practice, not a law of nature. Some people also mark values beyond 3×IQR3 \times \text{IQR} as “far out”. For data from a normal distribution, about 0.7% of values fall outside the 1.5 × IQR fences. So in a large normal sample, flagged points are expected: a sample of 1,000 will typically have about 7, even though nothing is wrong with any of them.

Outliers: what to do and what not to do

An outlier is an observation that lies unusually far from the bulk of the data. There is no universal definition; the 1.5 × IQR rule and z-scores beyond 3 are common screening rules, not tests of truth. A flagged value is a prompt to investigate, not a verdict.

When you find one, ask why it is there:

Good practice:

Common misunderstandings

“Outliers are errors.” Many are real. Whether to keep a point depends on how it arose, not on how far it lies from the others.

“The 90th percentile means scoring 90%.” A percentile is a position in the distribution, not a score. A test result at the 90th percentile was higher than about 90% of the results, whatever the raw score was.

“Quartiles are uniquely defined.” For a finite sample several conventions exist, and software defaults differ. For large datasets the differences are tiny.

“Points outside the whiskers must be removed before analysis.” The whisker rule is a display convention for drawing attention to points. It does not say what to do with them.

Further reading