Glossary

Validity

Glossary

Validity

Introduction

A study can be flawlessly executed and still measure the wrong thing: the survey that captures politeness instead of satisfaction, the experiment whose effect came from a confounder, the lab finding that evaporates in the wild. Validity is the umbrella question behind all of these: does the research actually measure, and demonstrate, what it claims to? It comes in named varieties, each guarding a different door. This article maps the four that matter most, the threats to each, and the trade-off between clean designs and real-world truth.

What is Validity?

Validity is the extent to which research measures what it intends to measure and supports the conclusions drawn from it. Where reliability asks whether the instrument is consistent, validity asks whether it is right: right construct, right causal story, right generalisation. A measure can be reliable and invalid (the consistently mis-calibrated scale); it cannot be valid without being reliable. Validity is also not a property a study simply has: it is an argument, assembled from design choices and evidence, and the named types below are the argument's chapters.

The Four That Matter Most

Construct validity: are you measuring the thing itself?
The gap between the concept ("user satisfaction", "engagement", "trust") and its operationalisation (a rating item, a click count, a retention curve) is where construct validity lives. Time-in-app as "engagement" measures confusion as happily as delight; a "would you recommend?" item measures politeness alongside loyalty (response bias is largely a construct-validity problem). The defences: multiple converging indicators, validated instruments over improvised ones, and the humility to ask what else this number could be measuring. Face validity (does it look right on inspection?) is this family's weakest member: necessary for participant buy-in, proof of nothing.

Internal validity: did X really cause Y?
The causal chapter: whether the observed effect came from the studied factor rather than confounders, selection, maturation, or measurement artefacts, the full rogues' gallery of correlation-versus-causation. Randomised experiments buy the strongest internal validity; observational designs argue for theirs confounder by confounder.

External validity: does it hold beyond the study?
Generalisation across people, settings, and time: from the sampled to the population (the territory of representative sampling), from the lab task to real use, from this quarter's users to next year's. The recurring product-research version: pristine prototype tasks completed by recruited participants predicting messier field behaviour imperfectly, which is why lab findings graduate through staged rollouts.

Ecological validity: does the study resemble life?
A sub-plot of external validity worth its own name: the realism of tasks, stimuli, and context. Unmoderated studies on participants' own devices, in their own homes, at their own hours (the format a Ballpark study runs by default) trade a little control for a meaningful step toward how the product is actually met, one reason remote testing often generalises better than its lab ancestor.

The Standing Tension

Internal and external validity pull against each other. Tight control (fixed tasks, clean environments, screened samples) isolates causes and strips away the world; realism restores the world and lets confounders back in. There is no design that maximises both, only sequences that earn each in turn: controlled studies to establish that an effect exists, field methods and rollouts to establish that it survives contact with reality. Programmes, not single studies, achieve validity.

Building the Argument

1. Define constructs before choosing metrics, and prefer validated instruments for anything abstract.
2. Match the design to the claim: causal claims need experimental or quasi-experimental backing; descriptive claims need honest sampling.
3. State the generalisation scope out loud: who was studied, in what conditions, and how far the finding is being stretched.
4. Triangulate: convergence across methods with different weaknesses is the strongest validity evidence applied research produces.

The Takeaway

Validity is the master question hiding behind every finding: right measure, right cause, right reach. Treat it as an argument to be built (constructs defined, designs matched to claims, scope stated, methods converging) rather than a box to be ticked, and remember the division of labour with reliability: consistency makes measurement possible; validity makes it true.

Further reading

For the varieties and their threats:

Articles:

1. Reliability vs. Validity in Research - Scribbr
The foundational pairing, with the main validity types and the classic threats to each laid out clearly.

2. When to Use Which User-Experience Research Methods - Nielsen Norman Group
Method selection as validity strategy: matching designs to the claims they can actually support.