A shop moves delivery costs from the final checkout step to the basket. The change seems more transparent, but the team does not yet know whether it will increase completed purchases, reduce surprises or discourage some shoppers earlier. An A/B test compares the current experience with the alternative through random assignment and a planned analysis of outcomes.
The aim is to estimate the effect of the change, including uncertainty. A higher number in one dashboard column is the beginning of interpretation, not enough on its own to establish that the proposed version is better.
Choose a question the experiment can answer
In the illustrative shop example, the primary outcome might be completed purchases per eligible shopper. Other consequential measures could include revenue, checkout errors and cancellations. Defining them before the experiment helps prevent a favourable secondary result from replacing the original question after the fact.
A/B testing is useful when a specific intervention, measurable outcome and sufficient sample are available. If the team does not understand why customers struggle, usability testing or user interviews may first help shape the proposal. The experiment assesses a change; it does not necessarily explain every reason people respond to it.
Microsoft Research’s account of online experimentation describes the role of controlled experiments in evaluating product ideas. The strength of the comparison depends on its design and execution, not merely on calling two versions A and B.
Keep assignment and measurement trustworthy
Decide which unit receives a version: a user, account or another appropriate unit. Keep assignment consistent where the design requires it, and consider whether users can affect one another. Shared accounts or collaboration can make a nominally individual assignment more complicated than it first appears.
Verify that both versions load as intended and that events mean the same thing in each. An apparent increase in completion is uninformative if one version fires the purchase event twice. Investigate unexpected group-size differences or missing exposure data before interpreting the outcome.
Use a sample and duration planned for the effect worth detecting, baseline rate or variability and chosen method. There is no universal number of participants. Rare outcomes and small effects often require more information, while repeated observations or clustered assignment require an analysis suited to their dependence.
Follow the stopping rule the analysis assumes
A fixed-sample test should not be stopped simply because repeated checks finally show significance. Properly designed sequential methods can support interim decisions, but they require their own rules. Understand which procedure the platform uses before treating its result as a decision signal.
Monitoring for broken tracking or harmful behaviour is different from searching for a winner. Record material interruptions, exclusions and changes so the analysis reflects the experiment that actually ran.
Decide using the effect and its consequences
Report the estimated difference and a confidence interval, along with statistical significance where used. A precise but negligible improvement may not justify maintenance costs; an uncertain result may leave useful benefits and harms unresolved.
For the shop, earlier disclosure could improve completion while changing basket value or cancellations. Those consequences belong in the decision. Also consider whether the result reflects an initial novelty response or a period that misses important usage patterns.
Document what was changed, who was included and what remains unknown. A test can support a well-bounded launch decision without proving that the same effect will persist indefinitely or apply to every customer segment.
