Introduction to Statistics

What statistics is for, the difference between describing data and drawing conclusions from it, and how this encyclopedia is organized.

Statistics is the discipline of learning from data. It gives us methods to summarize what a set of observations looks like, to measure how much we should trust conclusions drawn from limited or noisy data, and to make decisions when we cannot be certain.

Almost every field that measures something uses statistics: medicine to decide whether a treatment works, economics to estimate unemployment, engineering to control the quality of a product, machine learning to fit models to data. The ideas are the same in each case, which is why they are worth learning once, carefully.

Intuition

Suppose you want to know what fraction of voters in a country support a new policy. You cannot ask everyone, so you ask 1,000 randomly chosen people, and 52% of them say yes. Three questions follow immediately:

  1. What did we see? 520 of 1,000 people said yes. This is a description of the data in hand.
  2. What does it tell us about everyone else? Probably that support in the whole country is somewhere near 52%. But a different random 1,000 people would have given a slightly different answer, so we need to say how near.
  3. Can we act on it? Is support really above 50%, or could 52% easily have come from a country that is evenly split?

These three questions correspond to the main jobs of statistics: describing data, quantifying uncertainty, and drawing conclusions. A careful answer to the second question, which you will meet in the article on confidence intervals, is that a sample of this size typically pins down the true percentage to within about 3 percentage points. So 52% is not convincing evidence of a majority.

Descriptive and inferential statistics

Statistics is traditionally split into two parts.

Descriptive statistics summarizes the data you actually have. It answers questions such as “what is a typical value?”, “how spread out are the values?”, and “are there unusual observations?”. Its tools are numbers like the mean and median and the standard deviation, and pictures like histograms and box plots. Description makes no claim beyond the data. See descriptive statistics.

Inferential statistics uses the data to say something about a larger group or process that was not fully observed. In the language of populations and samples, we observe a sample and want to learn about the population it came from. Because the sample is only part of the picture, every inferential statement comes with a measure of uncertainty: an interval, a standard error, or a p-value.

The distinction matters because the two parts can fail in different ways. A perfectly accurate description of a badly chosen sample can still lead to a wrong inference about the population.

The role of probability

Probability and statistics are closely related but run in opposite directions.

To answer the statistical question, we need the probabilistic one: we can only judge whether 7 heads is surprising by knowing how often a fair coin produces it. This is why inference is built on probability. Random sampling and randomized experiments are what make the connection work: when chance decides who is in a sample, the laws of probability describe how much the sample can differ from the population.

A road map of this encyclopedia

The statistics articles are grouped so that each builds on the ones before it.

Foundations covers the vocabulary and descriptive tools: populations and samples, types of data, descriptive statistics, measures of centre (mean, median, and mode), measures of spread (variance and standard deviation), and quantiles and outliers.

Probability provides the mathematical language of chance: probability itself, conditional probability, independence, Bayes’ theorem, random variables, expected value, and the variance of a random variable.

Distributions describes the most common probability models: probability distributions in general, then the binomial, normal, Poisson, uniform, and exponential distributions.

Inference explains how to go from a sample to the population: sampling distributions, the law of large numbers, the central limit theorem, point estimation, maximum likelihood, confidence intervals, hypothesis testing, and p-values.

Relationships covers how two variables move together: covariance, correlation, and simple linear regression.

A reader new to the subject can follow this order. A reader looking up one idea can start anywhere; each article lists its prerequisites at the top.

Common misunderstandings

“Statistics is just a collection of formulas.” The formulas are the easy part. The hard part is judging whether the data can answer the question at all: how they were collected, what was measured, and what could have gone wrong. A correct calculation on biased data gives a confidently wrong answer.

“With enough data, uncertainty disappears.” More data reduces random error, but not systematic error. A million responses to a badly designed survey are still badly designed; see sampling bias.

“A statistically significant result is an important result.” “Significant” has a narrow technical meaning: the data would be surprising if a particular hypothesis were true. It says nothing about whether an effect is large or matters in practice. See p-values.

“Correlation shows causation.” Two variables can move together because one causes the other, because both are driven by a third factor, or by chance. Establishing cause usually needs a randomized experiment or careful additional reasoning. See correlation.

Further reading