Glossary

Experimental Design

Glossary

Experimental Design

Introduction

Most research watches the world; experiments interrogate it. By deliberately changing one thing while holding the rest steady, and assigning who gets the change by chance, experimental design earns the claim every other method can only gesture at: this caused that. Born in agricultural fields and perfected in clinical trials, the logic now runs every A/B test on the internet. This article covers the anatomy of a true experiment, the design choices that matter, and the discipline that separates causal evidence from expensive coincidence.

What is Experimental Design?

Experimental design is the structuring of research so that causal claims become testable: the researcher manipulates one or more factors (the independent variables), measures the outcomes (the dependent variables), and controls everything else, above all through random assignment of participants to conditions. Randomisation is the load-bearing wall: when chance decides who sees which condition, the groups match on every other trait (age, motivation, expertise, the confounders you know about and the ones you don't) in expectation, so a difference in outcomes has one live explanation. This is precisely the leap that observational correlation can never make, and the entire authority of the method rests on it.

A Brief History

The formal machinery comes from Ronald Fisher's work at Rothamsted agricultural station in the 1920s and 30s, where comparing crop treatments across stubbornly variable fields forced the invention of randomisation, replication, and blocking, codified in his 1935 book The Design of Experiments. Medicine adopted and hardened the template into the randomised controlled trial, adding blinding against expectation effects; the landmark 1948 UK streptomycin trial is usually cited as the modern RCT's arrival. The internet then made experimentation cheap and continuous: the A/B test is Fisher's logic running at web scale, thousands of times a day.

The Core Design Choices

Between-subjects or within-subjects. Between-subjects gives each participant one condition: clean, immune to carryover, hungry for sample. Within-subjects gives each participant every condition: statistically efficient (each person is their own control), but exposed to order effects (practice, fatigue, contrast), which counterbalancing (rotating condition order) exists to neutralise. Comparative usability studies face this choice constantly; a Ballpark study comparing two prototypes can randomise which one each participant meets first for exactly this reason.

Control conditions. A treatment difference means nothing without a comparison: the current design, a placebo-like neutral variant, or a no-change group. The control's job is to absorb everything that isn't the treatment (time trends, novelty, measurement effects) so the treatment's contribution stands alone.

One factor or several. Factorial designs vary multiple factors at once (headline × layout), revealing interactions single-factor tests miss, at the price of complexity and sample. The discipline is planning the comparisons in advance rather than mining the cells afterwards.

Running Experiments Honestly

1. Pre-state the hypothesis, metric, and sample.
The full hypothesis-testing contract: primary outcome chosen before data, sample sized by power analysis for the smallest effect worth acting on, and a stopping rule that resists the green dashboard at day two.

2. Randomise mechanically, verify empirically.
Assignment by code, not judgment, and a balance check afterwards: do the arms match on key observables? Randomisation guarantees balance in expectation, and checking catches the implementation bugs that break it in practice.

3. Guard the manipulation.
Confirm the treatment actually differed as intended (manipulation checks), and watch for contamination: participants experiencing both conditions, or the treatment leaking through shared channels.

4. Mind the validity trade.
Tight lab control maximises internal validity and strips away the world; experiments embedded in real products (field experiments) restore realism at some cost in control. Programmes earn both in sequence, not in one study.

The Benefits

Experimental design is the only method that supports causal claims by construction rather than argument. It converts strategy debates into decidable questions, measures effect sizes decision-makers can price, and (institutionalised as an experimentation culture) protects organisations from shipping plausible-sounding changes that quietly make things worse, which uninstrumented intuition does more often than anyone enjoys learning.

The Limitations

Not everything can be randomised: ethics, practicality, and scale exclude whole territories (pricing for existing customers, org-level changes), which is quasi-experimental design's domain. Experiments answer narrow questions narrowly; they tell you the variant won, not why, which is why the strongest programmes pair them with qualitative work that supplies mechanisms and hypotheses. And they inherit every sin of their execution: underpowered samples, peeked-at dashboards, and post-hoc metric shopping produce causal-sounding noise, wearing Fisher's authority.

The Takeaway

Experimental design is manipulation plus randomisation plus control, run to a pre-registered plan: the machinery that turns "these moved together" into "this caused that". Use it for the decisions worth its rigour, feed it hypotheses from methods that watch and listen, and honour the contract it demands, because a sloppy experiment is the most convincing wrong answer research can produce.

Further reading

For the design machinery in depth:

Articles:

1. A Guide to Experimental Design - Scribbr
Variables, randomisation, between/within choices, and worked design examples, laid out step by step.

2. A Refresher on Statistical Significance - Harvard Business Review
The analysis contract every experiment reports through, in decision-maker language.

Books:

1. The Design of Experiments - Ronald A. Fisher
The 1935 original: randomisation, replication, and the famous lady-tasting-tea illustration of inference by design.