A task-completion rate of 78% invites an obvious question: is that good? The answer depends on what people were trying to do, who took part and what happens when they fail. A missed item in a casual browsing task and a failed attempt to retrieve an essential document should not inherit the same target simply because both can be expressed as percentages.
Benchmarking compares performance against a defined reference, such as an earlier version of the product, a competitor or a suitable external dataset. It makes a result easier to interpret, but only when the reference is relevant enough to support the comparison.
Choose a reference that answers the decision
An internal benchmark is useful for asking whether a product is improving. A competitive study asks a different question: how well people can accomplish comparable goals using alternative products. External norms can provide context for an established measure, although differences in audience, task and study conditions still matter.
Before choosing a reference, identify the outcome that matters. Task success may be central when people need to complete a specific action; perceived usability may help assess the wider experience. Time on task requires interpretation because a quick exit can reflect either efficient completion or early abandonment. Nielsen Norman Group’s benchmarking guide discusses selecting measures and collecting them consistently.
Avoid adopting an industry figure merely because it is available. Check where it came from, how it was measured and whether the comparison would change a sensible product decision. A precise but irrelevant average offers less guidance than a modest, well-documented study of the task your customers actually need to perform.
Design the study so it can be compared
For a fictional comparison of two reporting tools, give participants equivalent goals rather than identical interface instructions. “Download the invoice for March” can work across different navigation systems; “open the Billing tab” may favour whichever product happens to use that label.
Decide whether the same people will try both products or separate groups will use each one. Reusing participants can make individual differences easier to account for, but introduces learning and order effects. Separate groups avoid that carry-over while requiring enough participants to handle variation between them. The experimental design should reflect the question and practical constraints.
Preserve the protocol, recruitment criteria and scoring rules. For a recurring programme, the initial baseline measurement should include enough detail for someone else to repeat the study without reconstructing its assumptions.
Explain the difference, then investigate it
Report sample sizes and uncertainty alongside the headline comparison. A small difference may be too uncertain to support a strong conclusion, while a substantial difference can still need explanation. Study recordings and participant accounts can help identify whether failures arose from navigation, terminology or a missing capability.
Those observations inform the next design decision; the benchmark alone does not tell the team what to build. Nor does an improvement between successive studies prove that a particular release caused it if recruitment or other conditions changed at the same time.
A useful benchmarking programme preserves comparability while remaining open to revising an obsolete measure. When the product’s purpose changes, document why the old benchmark no longer serves it and establish a new reference. Keeping an irrelevant score alive for the sake of a continuous line is a poor substitute for understanding performance.
