Research for Busy People

Sampling for Busy People

Who you ask matters far more than how many you ask.

01

The 30 sec read

Sampling is studying a small group to learn about a much larger one. You cannot ask every customer or every voter, so you ask a slice and generalise.

The surprising part is that it works: a few hundred or a couple of thousand well-chosen people can describe millions, as long as the sample is chosen fairly. The key word is fairly. A sample only represents the whole when everyone in the population had a real chance of being included.

The biggest mistake is assuming a bigger sample is always better. A huge sample drawn from the wrong place is more confidently wrong than a small fair one. Who you ask matters more than how many. Get the selection right first; worry about size second.

02

The 2 min read

Sampling is the practice of studying a subset of a group in order to draw conclusions about the whole. It is what makes research affordable: instead of surveying all ten million customers, you survey a well-chosen thousand and generalise. Done properly, the results are remarkably accurate — national polls routinely describe entire countries from a couple of thousand people.

The magic depends entirely on how the sample is selected. The gold standard is random sampling, where every member of the population has a known, equal chance of being picked. Randomness is what makes a sample representative: it stops your own choices, conscious or not, from skewing who gets included. When selection is not random — when you survey whoever is easiest to reach, or whoever volunteers — you get a convenience sample, which can be useful for quick signals but cannot safely be generalised, because the people who are easy to reach are usually different from those who are not.

The most common misconception is that size beats method. It does not. A famous 1936 poll surveyed millions of people and still called a US election wrong, because its list over-represented wealthier households. A properly drawn sample of a few thousand got it right. Size narrows random error; it does nothing to fix a biased selection. A big biased sample just gives you a precise wrong answer.

In practice you rarely get a perfect random sample, so good sampling is about getting as close as you can and being honest about the gaps — knowing who is likely to be missing, and treating conclusions about those groups with extra caution.

03

The 5 min read

Sampling is the science of learning about a large population by studying a smaller subset of it. Almost all research depends on it, because measuring everyone is usually impossible or pointlessly expensive. The central promise of sampling is genuinely remarkable: under the right conditions, a few hundred or a few thousand people can tell you, within a known margin, what millions think. The entire promise hinges on those right conditions.

Why it works at all

The reason a small sample can represent a huge population is randomness, not size. If every member of the population has an equal chance of being selected, then by the mathematics of probability the sample's characteristics will, on average, mirror the population's — and you can even calculate how much they are likely to differ (the margin of error). This is why a national poll of 1,500 people can estimate a country's opinion to within a few percentage points. The sample is not a miniature copy of the population; it is a fair draw from it, and fairness is what lets the maths work.

Types of sampling

Sampling methods fall into two families. Probability sampling — simple random, stratified, cluster — gives every unit a known chance of selection and is the only family that supports valid generalisation with a calculable margin of error. Stratified sampling improves on pure randomness by first dividing the population into groups (say, age bands) and sampling within each, guaranteeing the proportions match. Non-probability sampling — convenience, quota, volunteer, snowball — selects people by ease or judgement rather than chance. It is faster and cheaper, and often the only practical option, but it cannot support honest statistical generalisation, because you do not know who systematically had no chance of being included.

Size versus selection

The most expensive lesson in sampling is that selection beats size. A large sample reduces random error — the luck-of-the-draw wobble — but it does nothing about systematic bias, the consistent skew from a flawed selection method. If your method over-represents a group, surveying more people just pins down the biased answer more precisely. The classic cautionary tale is the 1936 Literary Digest poll, which surveyed over two million Americans and still predicted the wrong election winner, because its sampling lists (telephone and car owners) skewed wealthy during the Depression. A young pollster named George Gallup got it right with a far smaller but properly constructed sample. Bigger was not better; fairer was.

Common Misconceptions

Most people think

"A bigger sample is always more reliable."

Actually

Bigger only reduces random error, not bias. A small representative sample
beats a huge skewed one every time. Once a sample is biased, adding more
people from the same flawed source makes the wrong answer look more
trustworthy, not less.

Most people think

"If the sample is a large fraction of the population, it must be accurate."

Actually

What matters is how the sample was selected, not what percentage of the
population it represents. A well-drawn random sample of 1,000 can describe a
country of millions; a self-selected sample of 100,000 website visitors
cannot even reliably describe that website's users.

Common mistakes

The first mistake is chasing volume while ignoring method — collecting huge numbers of responses from whoever is easiest, then treating the total as if it were representative. The second is generalising from a convenience sample, presenting "our 5,000 survey takers" as "our customers" when the takers are a self-selected, atypical slice. The third is forgetting who is systematically missing — the people without internet access, the customers too busy or too unhappy to respond — and quietly treating conclusions about everyone as if those people were included. Good sampling is humble: it gets the selection as fair as circumstances allow, and is honest about the limits of what the sample can and cannot say.