Signal & Noise

The Open Question Finally Pays for Itself

For decades the honest answer to "what do you think?" cost too much to read at scale. That arithmetic just flipped, and it is quietly changing what a survey looks like.

01

Essay

Asking "what do you think?" in someone's own words has always been the honest way to ask a question, and the expensive one. Every open text field a respondent filled in became someone else's job afterward: read it, decide what it means, assign it to a category, then do that again ten thousand times. That labor, more than any theory about data quality, is why market research standardized on scales and tick-boxes. A five-point Likert item does not capture a feeling better than a sentence does — it is just that a computer can tally checkboxes for free, and a codebook needs a human being with a coffee and a week.

That arithmetic just flipped.

What actually changed

Large language models can now help code open-ended responses at scales and costs that were impractical a few years ago. That is not the same as saying that a model can replace an analyst. It means that some parts of the work can be tested rather than simply abandoned for lack of time.

One peer-reviewed approach pairs responses, asks an LLM which better expresses a defined concept, then fits a statistical model across thousands of comparisons. In that study, pairwise comparisons were more consistent than simple zero-shot ratings across several model sizes, and aligned with a knowledgeable human benchmark. It is a useful method, not a universal verdict on what machines understand.

The conversation instead of the form

The effect shows up earlier than the analysis stage. When the marginal cost of reading a short answer falls, there is less reason to force every question into four options someone else wrote. A survey can ask "tell me about the last time this happened" and get back keywords and themes extracted, not guessed at from a forced pick. The respondent answers in their own frame instead of yours — exactly where a short answer captures a nuance no checklist anticipated.

Who writes the codebook

There is a methodological question hiding under this, and it matters more than which tool you use: is the AI inventing the categories, or applying ones a researcher already defined? The evidence is clearer for the second. In one blinded focus-group study, models matched human analysts closely when applying a predefined codebook, but inductive theme generation was variable and still needed human verification. The practice that seems most defensible: the model drafts candidate categories from theory and a first pass over the data, a researcher merges the overlaps and drops what does not fit, then the model applies that codebook at scale. The researcher does not disappear from this picture — the job moves from reading every verbatim to owning the framework the verbatims get sorted into, which was always the part requiring judgment rather than time.

Two catches worth naming

The first: the same shift that makes it cheap to read open text is making it harder to know whose words you are reading. On crowdsourced platforms, a meaningful share of open-ended responses — estimates run as high as three in ten on some crowdsourcing platforms — are now drafted with an LLM's help rather than typed from scratch. Automated "AI text" checkers are not a clean answer: in one study, two tools agreed in only four out of five cases. Activity signals such as unusually fast entry, copy-and-paste, or leaving the survey page can strengthen the evidence, but they can also create false positives and miss a second device. They are useful signals, not proof.

The second is the one worth worrying about most, because this industry has walked into it before: an open question is not more representative by default, it is more representative of whoever bothers to answer it, and that group shrinks as the expected answer gets longer. Pew Research Center's panel data puts average nonresponse on closed questions at 1-2%, against about 18% for open-ended ones. Open questions asking for multiple sentences had higher nonresponse than those asking for a word or phrase, and the difference was especially clear on mobile. Ask for a paragraph instead of a sentence and more people quietly skip it — disproportionately the busy and the impatient. So the honest version of this story is not "ask for 200 words, because AI can read it." It is closer to the opposite: ask for one sentence, because AI can now do something useful with one sentence, and keeping the ask small is not a compromise on the vision — it is the version that does not quietly bias itself toward whoever had the most patience.

What doesn't go away

None of this retires the closed question. Tracking a metric over eighteen quarters, benchmarking against last year, comparing two markets on the same scale — that still wants a fixed set of options that mean the same thing every time, which is precisely what open text cannot promise. The realistic future is a mix chosen on purpose: closed where you need the same yardstick every time, open where the yardstick itself is what you do not want to assume in advance.

Why this is worth being pleased about

For most of the industry's history, "we would love to ask that as an open question, but nobody could read them all" was a reasonable, permanent excuse. It no longer is. The genuinely new craft problem — asking a short, real question and trusting the model to find the signal in a brief answer, rather than asking a long question and hoping people stick around to finish it — is a better problem to have than the old one, where almost nobody read the verbatims regardless of length.