Glossary

Usability Testing

Glossary

Usability Testing

Introduction

Every team believes its product is easy to use, right up until they watch a real person try. Usability testing is the practice of closing that gap: putting a design in front of actual users, giving them realistic tasks, and observing where they succeed, stumble, or give up entirely. It is one of the oldest methods in the user research toolkit and still one of the highest-value. This article covers what usability testing is, where it came from, how to run one well, and where its limits lie.

What is Usability Testing?

Usability testing is a method of evaluating a product or prototype by observing real users as they attempt to complete realistic tasks with it. Rather than asking people what they think of a design, you watch what they actually do: where they click, where they hesitate, what they misread, and where they abandon the task altogether. The gap between what people say and what they do is one of the most reliable findings in all of user research, and usability testing exists to capture the "do" side.

Usability testing is easy to confuse with neighbouring methods, so it helps to be precise. A focus group gathers opinions in conversation; a demo is driven by the team; user acceptance testing checks whether software meets a specification. Usability testing has one question at its core: can a representative person accomplish a realistic goal with this design, and what gets in their way? Everything about the method, from task design to recruitment to moderation technique, flows from that question.

Why It Matters

The people who build a product are the worst-placed people on earth to judge whether it is usable. They know where everything lives, what every label means, and what the system is doing behind the scenes. Users know none of this. Designers call the phenomenon the curse of knowledge, and no amount of team experience cures it. Usability testing is the correction mechanism: it replaces the team's assumptions with evidence, usually within the first handful of sessions.

Jakob Nielsen's famous argument that you only need to test with five users is a reminder that the barrier to entry is far lower than teams assume. His reasoning: each additional user in a round observes mostly the same problems as the previous ones, so the insight-per-participant curve flattens quickly. The often-forgotten second half of the argument is that the budget saved should fund more rounds rather than bigger ones. Five users, fix, five more, fix again. The rule also has honest caveats: if your product serves several very different audiences (say, doctors and patients), each audience needs its own five, and quantitative benchmarking needs far larger samples.

A Brief History

Usability testing grew out of human factors engineering in the mid-20th century, when researchers at organisations like Bell Labs and IBM studied how people interacted with complex systems. That work predates the personal computer by decades. Through the 1980s, dedicated usability labs, complete with one-way mirrors and banks of recording equipment, became standard at large software companies, and testing was an expensive, specialist affair. The 1990s brought Jakob Nielsen and the "discount usability" movement, which argued that small, cheap, frequent tests beat large, expensive, occasional ones. Steve Krug's Don't Make Me Think (2000) and Rocket Surgery Made Easy (2010) pushed the method into the mainstream, reframing it as something any team could do monthly with a laptop and a sandwich budget. The 2010s made the lab itself optional: remote and unmoderated testing platforms mean a test that once took weeks of scheduling can be in participants' hands the same afternoon, with recordings back before the end of the day.

Moderated or Unmoderated?

The single biggest methodological choice is whether a researcher is present. Moderated sessions, run in person or over video, let you probe in the moment: when a participant stalls, you can ask what they expected to happen, chase the reasoning behind a wrong turn, and rescue a session when a prototype breaks. They are the right choice for early, exploratory work, fragile prototypes, and complex domains where the "why" matters as much as the "what".

Unmoderated sessions trade that depth for scale, speed, and naturalism. Participants complete tasks alone, on their own devices, in their own environment, while their screen, voice, and clicks are recorded. This removes both the scheduling bottleneck and the subtle pressure of being watched by the person who (participants often assume) designed the thing. The cost is that nobody is there to ask the perfect follow-up, so task and question wording carry the full weight of the session. In practice mature teams use both: moderated sessions to understand a problem space, then unmoderated tests, the kind you can launch in Ballpark against a Figma prototype, to check designs quickly and repeatedly as they iterate.

What to Measure

Even in qualitative sessions, a few simple measures keep findings honest and comparable across rounds. Task success rate is the workhorse: did the participant genuinely complete the task, partially complete it, or fail? Time on task catches designs that technically work but grind. Error counts, and where the errors occur, pinpoint problem steps. Self-reported instruments like the System Usability Scale (SUS), a ten-item questionnaire with decades of published benchmarks, add a standardised score you can track release over release, turning usability from an opinion into a trend line. These metrics make small-sample testing more rigorous; they do not make it statistically predictive, a distinction covered below.

How to Run a Usability Test

1. Decide what you need to learn.
Pick the riskiest assumptions in the design, the flows where failure is most expensive. A test that tries to evaluate everything evaluates nothing. Three to five tasks is a comfortable ceiling for a session.

2. Recruit the right participants.
Your findings are only as good as your participants' relevance. Use screeners to filter for the behaviours and experience levels that match your real audience. Actual behaviour ("How many times have you bought clothing online in the past month?") screens far better than self-assessment ("Are you good with technology?").

3. Write task scenarios, not instructions.
"Find a pair of running shoes under £80 and get them delivered by Friday" tells you whether the navigation works. "Click the Shop menu, then Footwear" tells you whether people can follow instructions. Scenarios should give a goal and a context, never a route, and they should avoid borrowing the interface's own vocabulary, which quietly hands participants the answer.

4. Pilot the test.
Run one throwaway session first. Pilots catch ambiguous task wording, broken prototype links, and impossible time estimates while they are still free to fix.

5. Observe, and resist the urge to help.
The long pauses, the wrong turns, the muttered "where is it?": these uncomfortable moments are the entire point of the exercise. Ask participants to think aloud, answer questions with questions ("what would you expect that to do?"), and let silence do its work.

6. Analyse for patterns, not anecdotes.
One participant struggling is a data point; three struggling in the same place is a finding. Rate each issue by severity (does it block the task or merely slow it?) and frequency, fix the worst, and test again. The second round is where usability testing pays compound interest, because it tells you whether your fixes worked.

Usability Testing in Action: The $300 Million Button

One of the most cited case studies in the field is Jared Spool's account of the $300 million button. A major retailer required shoppers to register before checking out, assuming the form was a minor hurdle. Usability sessions showed the opposite. First-time buyers resented being forced into a relationship just to buy a product, and some abandoned the purchase entirely. Returning customers frequently could not remember their credentials, generating a blizzard of password resets. The fix was almost embarrassingly small: replacing "Register" with "Continue", making account creation optional, and adding a single explanatory sentence. The result was a dramatic, sustained increase in completed purchases worth hundreds of millions in the first year. No analytics dashboard had surfaced the problem. The data showed where people left; only watching real people revealed why.

Common Mistakes

A few failure modes account for most bad tests. Recruiting for convenience (colleagues, friends, whoever is nearby) produces participants who don't share your users' context, and findings that evaporate on contact with reality. Leading tasks and leading follow-ups ("was that easy?") manufacture the answers the team hoped for. Testing too late, when the design is finished and the team is emotionally committed, turns sessions into a formality whose findings arrive too expensive to act on. And treating a single round as "done" misses the method's core loop: test, fix, test again.

The Benefits

Usability testing catches problems while they are still cheap to fix, replaces internal debate with evidence, and builds empathy in a way no report can. There is nothing quite like watching a customer fail at the thing you built. It works at every fidelity, from paper sketches to production software, which means it can inform decisions from the first week of a project rather than auditing them at the end. And because sessions produce video, it generates the single most persuasive artefact in product development: a clip of a real customer struggling, played in a meeting.

The Limitations

Testing tells you what is broken more reliably than why, which is why it pairs well with user interviews and other qualitative methods. Sessions are an artificial situation: participants know they are being observed, and behaviour shifts accordingly (researchers call this the Hawthorne effect). Small samples cannot tell you which of two working designs performs better at scale; that is a job for A/B testing with live traffic. And usability is not desirability: a product can pass every task with flying colours and still be something nobody wants. The method's power is diagnostic rather than statistical. Diagnostic, as it happens, is exactly what most teams are missing.

The Takeaway

Usability testing is the closest thing user research has to a sure bet: almost every team that starts doing it regularly wonders how they shipped anything without it. It does not need a lab, a large budget, or a perfect protocol. It needs a prototype, a handful of representative users, and the willingness to watch quietly while your assumptions come apart.

Further reading

If you want to go deeper on usability testing, these are worth your time:

Articles:

1. Usability Testing 101 - Nielsen Norman Group
The canonical primer: what usability testing is, the elements of a test, and how moderated and unmoderated approaches differ. A sensible first read before running your own sessions.

2. Why You Only Need to Test with 5 Users - Jakob Nielsen
The classic argument for small, frequent tests over large, occasional ones, including the reasoning and the caveats that are often left out when the "five users" rule is quoted.

3. Usability Testing - Interaction Design Foundation
A broader topic overview with links into methods, history, and adjacent evaluation techniques, useful for placing usability testing in the wider research landscape.

Books:

1. Rocket Surgery Made Easy - Steve Krug
The most practical book ever written on do-it-yourself usability testing: how to run a monthly test morning, what to say, and how to turn observations into fixes.

2. Don't Make Me Think - Steve Krug
The book that made usability a mainstream concern. Short, funny, and still the best articulation of why self-evident design wins.