Glossary

Rating Scales

Glossary

Rating Scales

Rating Scales

Introduction

Rating scales are closed-ended survey questions that ask respondents to place an evaluation on an ordered set of points: satisfaction from 1 to 5, agreement from strongly disagree to strongly agree, likelihood from 0 to 10. They are the most used question type in research and among the most carelessly built, because small choices in the number of points, the labels, and the direction change what the numbers mean. This article covers the main types of rating scale, the design decisions that matter, and how to keep scales comparable across studies and time.

What are Rating Scales?

A rating scale is a survey response format in which the answer options form an ordered continuum and each option carries a value, so that responses can be summarized numerically: a mean satisfaction of 4.1, a top-two-box share of 72%, a distribution across the points. Scales measure attitude, evaluation, frequency, and intensity, and they underlie most of the metrics in this glossary, from CSAT and NPS to the SUS and every "how likely" question in a concept test. Their apparent simplicity hides a set of design decisions with measurable effects on the answers, and the craft is knowing which decisions matter for which purpose.

The Main Types

Agreement scales. The Likert family: a statement, and a symmetric scale from strongly disagree to strongly agree, usually five or seven points. Flexible and familiar; vulnerable to acquiescence, since agreeing is easier than disagreeing.

Satisfaction and evaluation scales. "How satisfied were you..." or "How would you rate..." from very dissatisfied to very satisfied, or poor to excellent. Direct, and dependent on fully labeled points.

Likelihood and intent scales. "How likely are you to..." on 0-10 (NPS) or a five-point verbal scale (purchase intent); the numeric 0-10 format supports fine distinctions and invites cultural differences in how the top is used.

Frequency scales. Never to always, or daily to less than monthly; better replaced by specific behavioral questions ("how many times last week?") wherever memory allows, since "often" means different things to different people.

Semantic differential. Bipolar adjective pairs (simple to complex, boring to exciting) across a scale; useful for impression and brand work, and a natural companion to first-impression studies.

The Design Decisions That Matter

1. Number of points.
Five or seven for most attitude questions: enough to discriminate, few enough to label. Eleven-point scales (0-10) support NPS-style analysis and are harder to label fully. Fewer than five loses sensitivity; more than eleven is false precision.

2. Label every point.
Scales with only the endpoints labeled produce different answers from fully labeled ones and are interpreted differently by different respondents; label all points, in words, for anything you'll compare or track.

3. Keep it balanced.
Equal numbers of positive and negative options, with a neutral midpoint or a deliberate decision to omit one (odd scales allow neutrality; even scales force a lean). Unbalanced scales ("good, very good, excellent, outstanding") manufacture positivity.

4. Fix the direction and keep it fixed.
Low-to-high, negative on the left, consistently across the instrument, so respondents don't misread a reversed scale; and never change it mid-series, because a flipped scale is a new question.

5. Avoid grids where you can.
Matrices of items on the same scale invite straight-lining and drive fatigue; single-item screens (the default in mobile-first survey design) cost seconds and improve attention.

6. Prefer validated multi-item scales for constructs.
A single item measures an item; a construct like usability or trust needs several, ideally from an instrument with known reliability.

Analyzing and Reporting

Rating-scale data is ordinal, and the long-running argument about treating it as interval (means, t-tests) versus ordinal (medians, distributions) resolves in practice: report the distribution and top-box or top-two-box shares, use means for multi-item composites, and flag single-item means as approximate. Keep the reporting convention fixed for the life of a series, because switching from mean to top-box is the most common cause of a false trend. And attach an open question to the ratings that matter, since a 3 tells you where and a sentence tells you why.

Where This Leaves You

Rating scales turn evaluations into numbers, and the numbers mean what the scale's design allows. Choose five or seven points, label all of them, balance the options, fix the direction, avoid grids, use validated multi-item scales for constructs, and report distributions with a fixed convention. Get the scale right before the survey goes out; there is no analysis that repairs a scale that manufactured its own answers.

Further reading

For scale design and evidence:

Articles:

1. Rating Scales in UX Research: The Ultimate Guide - Nielsen Norman Group
Scale types, point counts, labeling, and analysis, with the research behind each recommendation.

2. Writing Survey Questions - Pew Research Center
Balance, wording, and order effects from a leading survey organization.