Sampling is the science of learning about a large population by studying a smaller subset of it. Almost all research depends on it, because measuring everyone is usually impossible or pointlessly expensive. The central promise of sampling is genuinely remarkable: under the right conditions, a few hundred or a few thousand people can tell you, within a known margin, what millions think. The entire promise hinges on those right conditions.
Why it works at all
The reason a small sample can represent a huge population is randomness, not size. If every member of the population has an equal chance of being selected, then by the mathematics of probability the sample's characteristics will, on average, mirror the population's — and you can even calculate how much they are likely to differ (the margin of error). This is why a national poll of 1,500 people can estimate a country's opinion to within a few percentage points. The sample is not a miniature copy of the population; it is a fair draw from it, and fairness is what lets the maths work.
Types of sampling
Sampling methods fall into two families. Probability sampling — simple random, stratified, cluster — gives every unit a known chance of selection and is the only family that supports valid generalisation with a calculable margin of error. Stratified sampling improves on pure randomness by first dividing the population into groups (say, age bands) and sampling within each, guaranteeing the proportions match. Non-probability sampling — convenience, quota, volunteer, snowball — selects people by ease or judgement rather than chance. It is faster and cheaper, and often the only practical option, but it cannot support honest statistical generalisation, because you do not know who systematically had no chance of being included.
Size versus selection
The most expensive lesson in sampling is that selection beats size. A large sample reduces random error — the luck-of-the-draw wobble — but it does nothing about systematic bias, the consistent skew from a flawed selection method. If your method over-represents a group, surveying more people just pins down the biased answer more precisely. The classic cautionary tale is the 1936 Literary Digest poll, which surveyed over two million Americans and still predicted the wrong election winner, because its sampling lists (telephone and car owners) skewed wealthy during the Depression. A young pollster named George Gallup got it right with a far smaller but properly constructed sample. Bigger was not better; fairer was.
Common Misconceptions
Most people think
"A bigger sample is always more reliable."
Actually
Bigger only reduces random error, not bias. A small representative sample
beats a huge skewed one every time. Once a sample is biased, adding more
people from the same flawed source makes the wrong answer look more
trustworthy, not less.
Most people think
"If the sample is a large fraction of the population, it must be accurate."
Actually
What matters is how the sample was selected, not what percentage of the
population it represents. A well-drawn random sample of 1,000 can describe a
country of millions; a self-selected sample of 100,000 website visitors
cannot even reliably describe that website's users.
Common mistakes
The first mistake is chasing volume while ignoring method — collecting huge numbers of responses from whoever is easiest, then treating the total as if it were representative. The second is generalising from a convenience sample, presenting "our 5,000 survey takers" as "our customers" when the takers are a self-selected, atypical slice. The third is forgetting who is systematically missing — the people without internet access, the customers too busy or too unhappy to respond — and quietly treating conclusions about everyone as if those people were included. Good sampling is humble: it gets the selection as fair as circumstances allow, and is honest about the limits of what the sample can and cannot say.