Introduction to Statistics
What statistics is for, the difference between describing data and drawing conclusions from it, and how this encyclopedia is organized.
Statistics is the discipline of learning from data. It gives us methods to summarize what a set of observations looks like, to measure how much we should trust conclusions drawn from limited or noisy data, and to make decisions when we cannot be certain.
Almost every field that measures something uses statistics: medicine to decide whether a treatment works, economics to estimate unemployment, engineering to control the quality of a product, machine learning to fit models to data. The ideas are the same in each case, which is why they are worth learning once, carefully.
Intuition
Suppose you want to know what fraction of voters in a country support a new policy. You cannot ask everyone, so you ask 1,000 randomly chosen people, and 52% of them say yes. Three questions follow immediately:
- What did we see? 520 of 1,000 people said yes. This is a description of the data in hand.
- What does it tell us about everyone else? Probably that support in the whole country is somewhere near 52%. But a different random 1,000 people would have given a slightly different answer, so we need to say how near.
- Can we act on it? Is support really above 50%, or could 52% easily have come from a country that is evenly split?
These three questions correspond to the main jobs of statistics: describing data, quantifying uncertainty, and drawing conclusions. A careful answer to the second question, which you will meet in the article on confidence intervals, is that a sample of this size typically pins down the true percentage to within about 3 percentage points. So 52% is not convincing evidence of a majority.
Descriptive and inferential statistics
Statistics is traditionally split into two parts.
Descriptive statistics summarizes the data you actually have. It answers questions such as “what is a typical value?”, “how spread out are the values?”, and “are there unusual observations?”. Its tools are numbers like the mean and median and the standard deviation, and pictures like histograms and box plots. Description makes no claim beyond the data. See descriptive statistics.
Inferential statistics uses the data to say something about a larger group or process that was not fully observed. In the language of populations and samples, we observe a sample and want to learn about the population it came from. Because the sample is only part of the picture, every inferential statement comes with a measure of uncertainty: an interval, a standard error, or a p-value.
The distinction matters because the two parts can fail in different ways. A perfectly accurate description of a badly chosen sample can still lead to a wrong inference about the population.
The role of probability
Probability and statistics are closely related but run in opposite directions.
- Probability starts from a known model and asks what data it is likely to produce. “If a coin is fair, how likely are 7 or more heads in 10 tosses?”
- Statistics starts from observed data and asks which models are consistent with it. “I saw 7 heads in 10 tosses. Is the coin fair?”
To answer the statistical question, we need the probabilistic one: we can only judge whether 7 heads is surprising by knowing how often a fair coin produces it. This is why inference is built on probability. Random sampling and randomized experiments are what make the connection work: when chance decides who is in a sample, the laws of probability describe how much the sample can differ from the population.
A road map of this encyclopedia
The statistics articles are grouped so that each builds on the ones before it.
Foundations covers the vocabulary and descriptive tools: populations and samples, types of data, descriptive statistics, measures of centre (mean, median, and mode), measures of spread (variance and standard deviation), and quantiles and outliers.
Probability provides the mathematical language of chance: probability itself, conditional probability, independence, Bayes’ theorem, random variables, expected value, and the variance of a random variable.
Distributions describes the most common probability models: probability distributions in general, then the binomial, normal, Poisson, uniform, and exponential distributions.
Inference explains how to go from a sample to the population: sampling distributions, the law of large numbers, the central limit theorem, point estimation, maximum likelihood, confidence intervals, hypothesis testing, and p-values.
Relationships covers how two variables move together: covariance, correlation, and simple linear regression.
A reader new to the subject can follow this order. A reader looking up one idea can start anywhere; each article lists its prerequisites at the top.
Common misunderstandings
“Statistics is just a collection of formulas.” The formulas are the easy part. The hard part is judging whether the data can answer the question at all: how they were collected, what was measured, and what could have gone wrong. A correct calculation on biased data gives a confidently wrong answer.
“With enough data, uncertainty disappears.” More data reduces random error, but not systematic error. A million responses to a badly designed survey are still badly designed; see sampling bias.
“A statistically significant result is an important result.” “Significant” has a narrow technical meaning: the data would be surprising if a particular hypothesis were true. It says nothing about whether an effect is large or matters in practice. See p-values.
“Correlation shows causation.” Two variables can move together because one causes the other, because both are driven by a third factor, or by chance. Establishing cause usually needs a randomized experiment or careful additional reasoning. See correlation.
Further reading
- David Freedman, Robert Pisani, and Roger Purves, Statistics, 4th ed., W. W. Norton, 2007. A classic, non-technical introduction that emphasizes careful reasoning over formulas.
- David M. Diez, Mine Çetinkaya-Rundel, and Christopher D. Barr, OpenIntro Statistics, 4th ed., 2019. Free online; a complete first course with many examples.