Glossary

Baseline Measurement

Glossary

Baseline Measurement

Introduction

You can't measure improvement without knowing where you started. Baseline measurement is the deliberate recording of a metric before an intervention (a redesign, a launch, a process change) so that afterwards there is something honest to compare against. Skipping it is one of research's most common and least visible mistakes: teams ship, measure, and then argue about whether the number is good with nothing to anchor the argument. This article covers what a baseline is, why one reading is rarely enough, and how to establish baselines that survive the noise and regression effects waiting to distort them.

What is Baseline Measurement?

A baseline measurement is the value of a metric captured before a change, under the conditions that will later be compared against: task success on the current flow before the redesign, satisfaction scores before the support overhaul, conversion in the weeks before the pricing change. It is the "before" that gives an "after" meaning, and it differs from a benchmark, which is an external or standardised reference point (the published SUS average, an industry conversion rate). Benchmarks tell you how you compare to others; baselines tell you how you compare to your own past, which is the comparison most product decisions actually need.

Why One Reading Is Not a Baseline

A single pre-change measurement carries three problems. Noise: any one reading sits somewhere within the metric's normal variation, and a post-change reading compared against a low pre-change fluke "improves" by chance. Regression to the mean: interventions are usually triggered by bad readings (the quarter satisfaction dropped, the week errors spiked), and extreme readings tend to be followed by ordinary ones regardless of intervention, so the "recovery" after the fix was partly guaranteed. Trend: a metric already rising or falling before the change will continue, and a single point can't reveal the slope. The remedy for all three is the same: several readings over time, establishing the metric's level, spread, and direction before the intervention lands, which turns the comparison into an interrupted time series with a real pre-period rather than a two-point anecdote.

Establishing a Good Baseline

1. Measure what you'll measure after, the same way.
Same instrument, same tasks, same definitions, same recruitment profile; a baseline collected with a different survey or a different sample is a benchmark from another study wearing a baseline's name. Comparative studies built to be re-run (a Ballpark task study with fixed tasks and screener, fielded before and after a redesign) are baselines by construction.

2. Capture the spread, not just the centre.
The baseline's variability is what tells you whether a later change is large; a mean without its distribution or a proportion without its n cannot say whether the "after" is different or merely different-looking. Report intervals both times.

3. Take multiple readings when the metric is ongoing.
Weeks of analytics, several survey waves, repeated benchmark runs: enough history to see the normal band and the trend.

4. Record the context.
What else was true during the baseline period (season, campaigns, releases, incidents): the comparison is only fair if the after-period's context is judged against it.

5. Consider a concurrent control.
Where possible, a holdout group that doesn't receive the change measures what would have happened anyway, absorbing history and maturation; it is the baseline's stronger sibling, and the logic of the A/B test.

Baselines in Research Programmes

Baselines are the foundation of trend analysis and benchmarking programmes: the first wave of a tracked UX metric is the baseline every subsequent wave is read against, which is why the first wave deserves the most care in instrument design, since its choices are frozen for the life of the series. They also discipline goal-setting: a target of "80% task success" means something different when the baseline is 45% than when it's 78%, and KPI targets set without baselines are guesses with deadlines.

The Takeaway

A baseline is the honest "before": measured the same way as the "after", with its spread, over enough readings to reveal the normal band and the trend, with context recorded and, where possible, a concurrent control alongside. Establish it before you ship, not after, and treat the first wave of any tracked metric as the most important measurement in the series. Improvement is a comparison, and comparisons need both halves.

Further reading

For measurement programmes and their foundations:

Articles:

1. UX Benchmarking - Nielsen Norman Group
How to establish the first wave of a repeatable measurement programme and compare honestly against it.

2. Measuring Usability with the System Usability Scale - MeasuringU
A standardised instrument suited to baselines, with the external benchmark that complements them.