Introduction
Strongly disagree, disagree, neither, agree, strongly agree: the five-point scale is so ubiquitous that most people have answered thousands of them without knowing the format has a name, an inventor, and ninety years of methodological argument behind it. The Likert scale is survey research's default instrument for measuring attitudes, and the design decisions hiding inside it (how many points? label them all? include a midpoint?) quietly shape the data it produces. This article unpacks the scale, its craft, and its controversies.
What is a Likert Scale?
A Likert scale measures attitudes by asking respondents to rate their agreement with a statement along a symmetric range, classically five points from "strongly disagree" to "strongly agree", with seven-point variants close behind. Strictly speaking, a single such question is a Likert item; the scale is the sum or average across a set of items designed to measure one underlying attitude. In everyday research usage the distinction has largely dissolved, and "Likert scale" describes the response format itself, including its many cousins measuring satisfaction, frequency, importance, or likelihood rather than agreement.
The format's genius is that it converts something fuzzy, how strongly someone feels, into ordered, comparable, countable data, using a task nearly any respondent can perform in seconds. That combination of psychological accessibility and analytical convenience is why it conquered the survey world.
A Brief History
The scale is named for Rensis Likert, the American social psychologist who introduced it in his 1932 monograph A Technique for the Measurement of Attitudes. Likert's contribution was partly pragmatic: the dominant attitude-measurement approach of the day (Thurstone scaling) required panels of judges to calibrate every statement, and Likert showed that a simpler summed-rating format performed comparably at a fraction of the effort. He went on to found the University of Michigan's Institute for Social Research and to apply his methods to management theory; the scale, meanwhile, escaped academia entirely and became the default idiom of customer feedback, employee engagement, and every questionnaire in between. (He pronounced it "LICK-ert", for what it's worth, a fact that settles bar bets at research conferences.)
The Design Decisions That Matter
1. How many points?
Five and seven dominate for good reason: fewer points discard information respondents can reliably give; many more add precision respondents don't actually possess. Research on reliability generally finds modest gains moving from five to seven and little beyond, while very long scales (0-100 sliders, say) mostly add noise dressed as granularity.
2. Midpoint or no midpoint?
An odd-numbered scale offers a neutral option; an even-numbered one forces a lean. The neutral point is honest, since some people genuinely are neutral, but it also becomes a refuge for the disengaged and the conflict-averse. Forcing a choice extracts more signal at the cost of misrepresenting the genuinely indifferent. There is no universally right answer; there is only knowing which error you would rather make for this question.
3. Label every point.
Scales with only the endpoints labelled leave the interior points open to interpretation, and respondents interpret them differently. Fully labelled scales ("disagree", "somewhat disagree"...) are more reliable and more comparable across people, at the cost of the researcher having to actually write good labels.
4. Keep direction consistent, and statements balanced.
Respondents fall into rhythms. A block where agreement always means positivity invites acquiescence bias, the documented tendency to agree with statements regardless of content. The traditional fix, reverse-worded items, brings its own problem: negatively phrased statements confuse respondents and generate errors. Modern practice favours neutral, single-idea statements ("The checkout process was easy to complete") over strongly valenced ones, and never two ideas in one item ("The app is fast and reliable": which half is the respondent rating?).
Analysing Likert Data
Here lies the format's oldest argument. Likert responses are ordinal: the categories are ordered, but nothing guarantees the psychological distance from "agree" to "strongly agree" equals the distance from "neutral" to "agree". Purists therefore report medians, frequency distributions, and use non-parametric tests. Pragmatists note that summed multi-item scales behave approximately like interval data and analyse means accordingly, which is what most of industry does. A defensible working position: for a single item, prefer the distribution (show how many people chose each option, since a mean of 3.4 can hide a polarised split); for a composite of several items, means and standard analysis are reasonable; and either way, decide the analysis before fielding, not after. When comparing groups or time periods, statistical significance testing applies exactly as it does to any other measure, including all its usual caveats.
Likert Scales in Product Research
Most of the field's standard instruments are Likert machinery under the bonnet: the System Usability Scale is ten five-point Likert items, and countless satisfaction and ease-of-use questions follow the same pattern. Used inside product studies, the format shines as connective tissue. A quick rating after each task in a usability test, or a scale question between open-ended ones in a survey, gives you comparable numbers alongside the rich material. In mixed-method tools like Ballpark, that pairing is the point: a Likert item tells you that ease-of-use dropped between prototypes; the video answer beside it tells you why. Beware only the wall of grids: a dozen consecutive Likert matrices is the express route to survey fatigue and the straight-lined data that comes with it.
The Benefits
The Likert format is fast for respondents, familiar to the point of invisibility, cheap to field at scale, and produces ordered data that supports comparison across questions, groups, and time. Nine decades of methodological literature mean its failure modes are mapped in detail, a luxury few instruments enjoy, and composite Likert scales, properly constructed, achieve reliability that single questions of any format cannot.
The Limitations
It measures self-reported attitude, not behaviour, with all the gaps that implies. It is vulnerable to acquiescence, central-tendency (midpoint-hugging), and extreme-response styles that differ across individuals and cultures, complicating comparisons. The ordinal-versus-interval question never fully goes away. And its very convenience is a trap: because Likert grids are the easiest question type to write, questionnaires silt up with them, measuring ever more attitudes at ever lower quality. The scale is a precision tool that mass adoption keeps mistaking for a default.
The Takeaway
The Likert scale earned its ubiquity: no other format converts feeling into data so cheaply and so legibly. But every one of its design choices (points, midpoint, labels, wording) moves the numbers, which means a Likert question is never neutral infrastructure. Treat the format with the same care you would give the analysis, show distributions rather than hiding behind means, and pair the ratings with methods that capture the why. Rensis Likert gave research a superb instrument; the craft is in refusing to let it become a reflex.
Further reading
For the origins and craft of the format:
Articles:
1. Likert Scale: Definition, Examples, and Analysis - Simply Psychology
A clear, well-referenced introduction covering the scale's structure, its ordinal-data debate, and analysis options, with the historical context of Likert's original 1932 work.
2. Survey Best Practices - Nielsen Norman Group
Practical guidance on question wording, scale construction, and ordering: the craft decisions that determine whether Likert data means anything.
3. Measuring Usability with the System Usability Scale - MeasuringU
The best-known applied Likert instrument in UX, explained: scoring, benchmarks, and what a composite Likert scale looks like when it is done properly.