Glossary

System Usability Scale (SUS)

Glossary

System Usability Scale (SUS)

System Usability Scale (SUS)

Introduction

The System Usability Scale (SUS) is a ten-item standardized questionnaire that produces a single score from 0 to 100 representing users' perceived usability of a product. Ten short statements, five response options each, one number out of 100 at the end: the System Usability Scale is the closest thing usability measurement has to a universal standard. Created in the 1980s as a self-described "quick and dirty" questionnaire, it outlived hundreds of more sophisticated instruments precisely because it is short, free, and backed by decades of benchmark data. This article explains how SUS works, how to score and interpret it properly, and the traps that catch first-time users of the scale.

What is the System Usability Scale?

The System Usability Scale (SUS) is a ten-item questionnaire for measuring perceived usability. Respondents rate statements like "I thought the system was easy to use" and "I found the system unnecessarily complex" on a five-point agreement scale, and the answers combine into a single score from 0 to 100. The items alternate between positively and negatively worded statements, a deliberate design that forces respondents to read each one rather than straight-lining down the page.

SUS measures perception, not performance. It tells you how usable people felt the system was, which complements rather than replaces the behavioral measures (task success, time, errors) collected in a usability test. The two can disagree, and the disagreement is informative: a product people struggle with but rate highly, or breeze through but rate poorly, is telling you something about expectations.

A Brief History

SUS was created in 1986 by John Brooke, then at Digital Equipment Corporation, who needed a fast way to compare the usability of office systems and openly described the result as a "quick and dirty" scale. Released free of charge, it spread through industry and academia on the strength of costing nothing and taking two minutes. The accumulated result, decades later, is its real moat: thousands of published studies across products and industries, which let researchers place any new score against a meaningful distribution. Jeff Sauro and colleagues at MeasuringU have done much of the work of compiling those benchmarks; across large aggregated datasets, the average SUS score sits around 68, a number worth memorising before interpreting your first result.

Scoring It Correctly

The scoring is fiddly and frequently botched. Each item is converted to a 0-4 contribution: for the positively worded (odd-numbered) items, the score is the response position minus 1; for the negatively worded (even-numbered) items, it is 5 minus the response position. Sum the ten contributions and multiply by 2.5 to reach the 0-100 range. Two errors dominate in practice. The first is forgetting to reverse the negative items, which quietly wrecks the score. The second is reading the result as a percentage: a SUS of 68 does not mean 68% of anything. It means dead average against published benchmarks, roughly the 50th percentile. A 78 is well above average; below 60 signals real trouble. Grading schemes and adjective mappings (from "poor" through "excellent") exist in the literature and help translate the number for stakeholders.

Using SUS in Practice

1. Field it immediately after use.
SUS belongs right after a usage session or task set, before discussion or debrief contaminates the impression. In an unmoderated study (the kind you can run in Ballpark alongside prototype tasks), append the ten items as a closing questionnaire so every participant produces both behavioral data and a score.

2. Use it to compare, not to diagnose.
The score is a thermometer: superb for tracking a product across releases, comparing against a competitor in a benchmarking study, or checking a redesign against its predecessor. It contains no information about what to fix; the ten items were never designed for item-level diagnosis. Pair the number with observation and open questions for the why.

3. Mind the sample size.
SUS scores from five participants carry wide uncertainty; scores stabilize usefully somewhere in the teens and twenties of respondents. Report confidence intervals when comparing versions, and resist celebrating a three-point rise measured on eight people. The usual rules of statistical significance apply in full.

4. Keep the wording intact.
The benchmarks are the point, and the benchmarks assume the standard instrument. Rewording items, dropping the awkward ones, or switching the response scale produces a questionnaire that may be fine but is no longer comparable to the published distribution. (One sanctioned tweak: replacing the word "cumbersome" in item 8, which non-native speakers stumble on, with "awkward".)

The Benefits

SUS is free, takes two minutes, works across nearly any kind of system, and produces a single number stakeholders instantly grasp. Its psychometric properties are well established, its benchmark corpus is unmatched, and it is robust at the modest sample sizes real product teams actually run. As a standardized layer on top of qualitative testing, it turns usability from an anecdote into a trend line.

The Limitations

It is a perception measure with all the biases self-report carries, and it is global: one number for the whole experience, blind to which screen or flow caused the damage. It is not a diagnostic tool, not a substitute for watching users, and not meaningful as a percentage however often it gets read as one. The alternating item polarity, clever in 1986, measurably trips some respondents. And a good score is easy to over-trust: a system can be perfectly usable and still pointless, a distinction no questionnaire will catch.

Where This Leaves You

SUS endures because it solved the right problem cheaply: a standard, comparable, two-minute measure of perceived usability with a benchmark corpus no rival can match. Score it correctly, field it consistently, read 68 as average rather than a grade, and let it do what it does best: tell you whether the needle is moving, while your other methods tell you why.

Further reading

For scoring, benchmarks, and the scale's history:

Articles:

1. Measuring Usability with the System Usability Scale - MeasuringU
The most useful single reference: scoring walkthrough, the benchmark distribution, grading schemes, and answers to the questions every first-time user asks.

2. When to Use Which User-Experience Research Methods - Nielsen Norman Group
Context for where standardized questionnaires sit in the wider research toolkit, alongside the behavioral methods SUS is designed to complement.