Glossary

Rating Scales

Glossary

Rating Scales

Rating Scales

Rating scales provide ordered response options for expressing an evaluation, attitude, intensity or frequency. Their meaning depends on the question and the response labels.

A rating of four can look reassuringly precise until someone asks what it means. Four out of five for satisfaction is different from four out of seven for effort, and either can conceal a disagreement between people who had very different experiences. The number becomes useful when it stays attached to the question, the response labels and the circumstances in which it was given.

Rating scales offer an ordered set of responses for expressing an evaluation, attitude, intensity or frequency. Choosing one is a measurement decision: it determines what distinctions participants can express and what the resulting data can reasonably support.

Match the scale to the judgement

A question about satisfaction should offer satisfaction responses. A question about frequency needs a reference period and categories people can distinguish. Asking someone whether they agree that they use a feature frequently adds an unnecessary judgement about agreement to a question that might be answered more directly by asking how often they used it last week.

Some scales run from an absence to a high degree, such as no difficulty to extreme difficulty. Others run between opposing evaluations, such as very dissatisfied to very satisfied. Their middle categories therefore mean different things. A moderate amount of difficulty is not neutrality between difficulty and ease.

A Likert scale is a particular approach associated with agreement items, rather than a name for every numbered response format. Established instruments such as the System Usability Scale also carry their own wording and scoring conventions. Changing those conventions creates a different measure, even if the result still fits on a familiar chart.

Make the options understandable

Choose enough categories to capture distinctions respondents can make, without asking them to pretend to finer judgement than they possess. Five or seven points are common choices, but neither is a universal optimum. Clear labels matter more than the appearance of mathematical sophistication.

Keep direction consistent through a questionnaire unless an established instrument specifically requires otherwise. Separate uncertainty or lack of experience from the substantive scale: someone who has never contacted support cannot meaningfully rate that experience as neutral. The Government Analysis Function’s questionnaire guidance is a useful reference for reviewing response categories and scale presentation.

In a pilot, ask people what adjacent options mean to them. If they cannot explain the difference between “fairly” and “somewhat”, the extra category may be adding apparent precision rather than useful information.

Show what lies behind the summary

In a fictional five-point survey, ten people all choosing three produce the same mean as five people choosing one and five choosing five. The experiences behind those results are plainly different. Report the distribution alongside a summary when that distinction matters, and state any rule used to combine categories into a percentage.

Comparisons also depend on who answered and when. A rating immediately after a difficult task is not interchangeable with an overall judgement sent weeks later. Preserve the question, labels, recruitment and timing when tracking change, and document any unavoidable alteration.

A rating can identify where to look more closely. To understand why a score changed, pair it with relevant task evidence or a focused open-ended question, rather than expecting a number to supply its own explanation.

Further reading

Guides

  1. Questionnaire design guidance — Government Analysis Function
    Discusses the design choices behind survey questions and response formats. Useful when checking whether a scale's wording and options allow people to express the judgement you intend to measure.