
Introduction
MaxDiff (maximum difference scaling, or best-worst scaling) is a survey technique for measuring the relative importance of a list of items, such as features, benefits, messages, or attributes, by repeatedly showing respondents small subsets and asking them to pick the most and least important in each. It produces a clean, ratio-like ranking where rating scales produce a pile of "very important" answers, and it has become the standard method for prioritizing anything with more than a handful of options. This article covers how MaxDiff works, why it beats rating scales for prioritization, and how to design one.
What is MaxDiff?
MaxDiff is a choice-based survey method in which respondents see a series of sets, each containing a few items (typically four or five) drawn from a longer list, and in each set choose the item that matters most and the item that matters least to them. Across enough sets, every item is seen and judged against every other several times, and the pattern of best and worst choices is analyzed (usually with the same hierarchical Bayes estimation used in conjoint) to produce a preference score for every item, on a scale where the scores are directly comparable: an item scoring 20 is roughly twice as preferred as one scoring 10. Developed by Jordan Louviere in the late 1980s and early 1990s, the method solves the problem that has plagued importance questions forever: asked to rate features on a five-point scale, respondents rate nearly everything four or five, and the researcher learns nothing about priority.
Why It Beats Rating Scales
It forces trade-offs. You cannot call everything "most important" when you must also name a "least"; the format extracts the discrimination that rating scales let respondents avoid.
It removes scale-use bias. Some people rate everything high, others low, and cultures differ in scale use; choosing best and worst has no scale, so those biases vanish, which makes cross-segment and cross-country comparison honest.
It produces a ratio scale. Scores that support "twice as important" statements, and that sum across the list, rather than an ordinal pile of fours.
It is easy for respondents. Picking the best and worst of four is a natural judgment, quicker and less tiring than rating twenty items or ranking them all, which protects data quality against fatigue.
What It's Used For
Feature and roadmap prioritization. Which of twenty candidate features matter most to which segments: the direct input to prioritization frameworks' impact column.
Message and claim testing. Which of a dozen value-proposition statements resonates most, the quantitative complement to proposition research.
Benefit and attribute importance. What buyers weigh in a category, as an input to positioning and to the attribute selection for a later conjoint study.
Pain-point prioritization. Which of the problems discovered in interviews matter most to the wider audience, closing the loop from qualitative discovery to quantitative priority.
Designing One
1. Build the item list from research.
Ten to thirty items, each a single, clearly worded, mutually distinct thing, in the audience's language; items that overlap or bundle two ideas corrupt the comparisons.
2. Use an experimental design for the sets.
Software generates a balanced design so every item appears equally often and against varied competitors; a typical study shows each respondent ten to fifteen sets of four or five items, with each item seen three or more times.
3. Define "important" concretely.
"Most important when choosing a tool for X" rather than "most important", so respondents judge against a shared context.
4. Sample the audience at size.
A few hundred respondents from the target audience for stable segment-level scores; screened panel recruitment (the route a Ballpark survey study takes) makes the sample practical.
5. Segment the results.
Aggregate scores hide the fact that two segments want opposite things; the segment view is usually the finding.
The Limitations
MaxDiff measures relative importance among the items shown; it says nothing about absolute importance (the "most important" item might still not matter much) or about price, which is conjoint's job. Items missing from the list are invisible. Long lists need many sets, and respondent patience is finite. And, as with every stated-preference method, what people say they prioritize and what they pay for or use can diverge, which is why MaxDiff results are validated against behavior where possible.
In Short
MaxDiff asks respondents to pick the best and worst from small sets, over and over, and turns the choices into a comparable, ratio-scaled ranking of everything on the list. It replaces the useless pile of "very important" ratings with genuine priorities, free of scale bias, at low respondent burden. Build the list from research, design the sets properly, define the context, sample at size, and read the segments; then use conjoint when the question adds a price.
Further reading
For the method and its analysis:
Articles:
1. MaxDiff - Sawtooth Software
The reference treatment: design, estimation, interpretation, and when MaxDiff beats rating and ranking.
2. Writing Survey Questions - Pew Research Center
The question-craft fundamentals for wording the items MaxDiff compares.