Glossary

Baseline Measurement

Glossary

Baseline Measurement

Baseline Measurement

A baseline measurement records a starting condition before a change, providing a documented reference for interpreting later results.

A redesign launches, the completion rate rises, and the team wants to know whether its work helped. That question is much harder to answer if nobody recorded how completion was measured before the release. Even a remembered figure is of limited use when its denominator, time period and definition of success have disappeared.

A baseline measurement records the starting condition against which later results will be compared. It might be a usability study of the existing product, several weeks of behavioural data or the first wave of a recurring survey. Its value lies in making the comparison possible and interpretable, rather than merely supplying a number labelled “before”.

Capture the conditions as well as the result

In a fictional invoice study, a team records whether participants can find and download the correct invoice without help. To repeat that measurement, a future researcher needs the task wording, the participant criteria, the device conditions and the rule for judging success. A percentage without those details cannot distinguish a better design from an easier test.

For an ongoing product metric, record the event definition, exclusions and reporting window. Preserve the size of the sample and the variation in the results, not just their average. Nielsen Norman Group’s guidance on establishing baselines discusses the role of both analytics and research measures in creating a useful starting point.

One study can provide a baseline, but it still carries uncertainty. Several readings may be valuable for a metric that fluctuates with weekdays, seasonal demand or release cycles. The aim is enough context to recognise ordinary variation, not a ritual requirement to collect the same amount of history for every decision.

Keep the comparison fair

When measuring again, preserve the aspects of the study that are intended to remain comparable. Changing the audience, the task, the rating question and the product at once leaves several explanations for any difference. If a necessary change breaks the series, document it and consider collecting overlapping measurements to understand the effect.

A baseline is often used as an internal benchmark. Benchmarking is the broader practice of comparison and may also use competitors or external reference data. Neither an industry average nor a competitor’s result can recreate a missing measurement of your own product before a change.

Separate improvement from attribution

A before-and-after difference shows that a measured result changed. It does not establish that the redesign caused it. A marketing campaign may have brought a different audience, or a temporary outage may have made the baseline unusually poor. If the starting point was selected because it was an extreme low, a later recovery may also partly reflect ordinary fluctuation.

A suitable concurrent comparison group can help distinguish these explanations. Where feasible, an A/B test measures the change against a version that did not receive it during the same period. When that is impractical, report the before-and-after evidence with its limitations and check other explanations using release records, audience data and research.

The useful conclusion may be that the new flow performs better under the measured conditions, while the exact contribution of the redesign remains uncertain. That is still valuable evidence, provided the report does not quietly turn a comparison into a causal claim.

Further reading

Articles

  1. Establishing baselines — Nielsen Norman Group
    Explains why a starting measurement matters and how to make it useful for later comparisons. Read before an improvement removes the opportunity to measure the original experience.

  2. A practical guide to SUS — MeasuringU
    Explains a standard questionnaire that can form part of a usability baseline. Useful when you need a consistent subjective measure alongside task performance.