
Introduction
Formative research is research conducted during development to identify problems and improve a product while change is still cheap, as opposed to summative research, which judges the finished result. There are two moments to evaluate a design: while it's still being shaped, and after it's done. Formative research is the first kind, run during development to find what's wrong and improve it, as opposed to summative research, which judges the finished thing against a standard. The distinction sounds academic and decides everything about how a study should be designed, sized, and reported. This article covers what formative research is for, how it differs from summative evaluation, and why the most valuable research a product team runs is usually the kind that never produces a score.
What is Formative Research?
Formative research is research conducted during the development of a product, feature, service, or program, with the purpose of informing and improving it while change is still cheap: identifying problems, understanding needs, testing early concepts, and shaping design decisions. Its counterpart, summative research, evaluates the finished artifact: how usable is it, how does it compare to the previous version or a competitor, did it meet the goals? The terms come from educational evaluation (Michael Scriven drew the distinction in 1967), and the memorable gloss is that formative evaluation happens when the cook tastes the soup and summative when the guests do. Both are necessary; they are not interchangeable, and a study designed for one does the other badly.
Formative and Summative, Side by Side
Purpose. Formative asks "what's wrong and why?"; summative asks "how good is it?".
Timing. Formative runs early and often, on sketches, prototypes, and betas; summative runs at milestones and after release.
Method and sample. Formative work is mostly qualitative and small (five to eight participants per round, iterated), because its job is problem detection, which saturates quickly. Summative work is mostly quantitative and larger, because its job is measurement, which needs sample for precision: task-success rates, times, SUS scores, benchmarks with intervals.
Output. Formative produces a prioritized list of issues and insights with recommendations; summative produces metrics against a standard and a verdict.
Fidelity. Formative embraces rough artifacts (paper, wireframes, clickable mockups), since roughness invites honest criticism; summative needs the real thing, or something close enough to measure fairly.
The Formative Toolkit
Discovery interviews and field observation before anything is designed; concept and value-proposition tests on early ideas; usability testing on prototypes at every fidelity, with think-aloud narration to expose the reasoning behind each stumble; expert review to catch the obvious before participants' time is spent on it; and first-click and preference tests on specific design questions. The common thread is speed and cheapness relative to the cost of the changes they inform: a recorded prototype study with a handful of target users (a formative Ballpark study, fielded overnight) costs a fraction of the engineering week it can redirect.
Doing It Well
1. Test early enough to matter.
The value of a formative finding falls with every line of code written after it; the first round belongs on the sketch, not the release candidate.
2. Report problems, not scores.
A formative study that leads with "73% task success" has answered a summative question with a formative sample; lead with what went wrong, why, how severe, and what to change.
3. Prioritize by severity and frequency.
Not every finding deserves a fix this round; a severity rubric keeps the list actionable and the arguments short.
4. Close the loop with the next round.
The fix's evidence is the problem's absence next time; formative research is a series, not a study.
5. Switch modes deliberately.
When the question becomes "is it good enough to ship?", change the design: larger sample, fixed tasks, metrics with intervals, a defined baseline or benchmark. Summative questions asked of formative studies produce confident numbers with no precision behind them.
What to Remember
Formative research shapes the thing while it can still be shaped: early, small, qualitative, iterative, and reported as problems to fix rather than scores to celebrate. Summative research judges the finished thing with the sample and metrics judgment requires. Know which question you're asking, design the study for it, and taste the soup before the guests arrive; it is much cheaper to add salt in the kitchen.
Further reading
For the distinction and its consequences:
Articles:
1. Formative vs. Summative Evaluations - Nielsen Norman Group
The two evaluation modes compared on purpose, timing, method, and reporting, with guidance on choosing.
2. Usability Testing 101 - Nielsen Norman Group
The formative workhorse, from planning through moderation to findings.