Experimental design specifies how an intervention will be compared with alternatives so that the resulting evidence can answer a causal question. It includes the conditions, assignment, outcomes, timing and analysis. The design begins before data collection, when the team decides which comparison would actually resolve its uncertainty.
A product experiment may compare two interfaces, several messages or combinations of features. The visible change is only part of the design. Who receives it, what they can influence and how their outcomes are recorded can matter just as much.
Define the intervention and the unit receiving it
Suppose a fictional collaboration tool wants to test a new way of assigning work. Assigning individual users to versions may be awkward if everyone on a team shares the same tasks. A team-level assignment may better contain the experience, but outcomes within teams are related and the sample-size planning and analysis need to account for that structure.
State exactly what differs between conditions and who is eligible. Keep the comparison concurrent where appropriate so that unrelated changes over time do not become the apparent effect. Check that the intended treatment is actually delivered and recorded.
NIST’s guidance on choosing an experimental design describes design choices for different objectives. The general discipline is to choose the arrangement around the question and sources of variation, rather than selecting a familiar test after the data arrive.
Choose how participants encounter the conditions
A between-subjects design places participants in different conditions. A within-subjects design lets the same people encounter several, which can improve some comparisons while introducing learning, fatigue and carry-over. Counterbalancing can distribute order effects; it does not guarantee that every carry-over problem disappears.
Blocking groups comparable units before assignment within those groups. Factorial designs vary more than one factor and can investigate interactions, such as whether the effect of a message depends on its placement. These approaches require a plan for the comparisons and enough information to estimate them usefully.
Randomisation supports comparability in expectation, not identical observed groups. Check implementation and unexpected allocation patterns without treating ordinary chance differences as proof that randomisation failed.
Plan outcomes, precision and stopping together
Define the primary outcome, important guardrails and the effect worth detecting. Plan the sample and duration using the relevant baseline, variability, assignment unit and statistical method. Include meaningful usage cycles where they affect interpretation.
Specify exclusions, missing-data handling and stopping rules before outcome inspection. A fixed-sample analysis and a sequential analysis permit different forms of interim decision-making. Quality and safety monitoring remain necessary, but repeated searching for a favourable p-value is not a substitute for a stopping plan.
Interpret the experiment that actually ran
Report departures, exposure failures and any plausible interference between units. In the collaboration example, a team using both versions through shared accounts could undermine the intended contrast. That possibility needs investigation before the result is presented as a clean comparison.
Use effect estimates and confidence intervals, alongside hypothesis testing where appropriate. Consider practical consequences and the limits of generalizability. An experiment can estimate a specific intervention’s effect without explaining every mechanism or establishing that the result will persist unchanged in another setting.
