Hypothesis testing evaluates how compatible data are with a specified statistical hypothesis under a chosen model. It gives a research question a formal comparison and a decision procedure, but it does not turn the outcome into an automatic product decision.
A team might ask whether a new checkout changes purchase completion. The statistical hypotheses describe the quantities being compared, while the study design determines whether that comparison can be interpreted causally. The practical decision also needs the size of any change, uncertainty, costs and effects on other outcomes.
State the comparison before choosing the calculation
Define the population, outcome, unit of analysis and conditions. A null hypothesis might specify equal completion rates, with a two-sided alternative allowing either an increase or a decrease. Other questions use different null values or forms of comparison.
Choose a test appropriate to the data and design. Means, proportions, paired observations and clustered units raise different requirements. A one-sided test needs a justified directional question established before inspecting the results; it should not be selected afterwards because the estimate points in a convenient direction.
NIST’s introduction to hypothesis testing describes the statistical framework. In practice, the hypotheses, assumptions and analysis plan should be explicit enough for another researcher to understand what would count as evidence against the null.
Plan what the study needs to detect
Sample-size planning depends on the effect worth detecting, variability or baseline rate, significance level, desired power and the design. Power is the probability of rejecting the null under a specified alternative; it is not a generic property that a study has independently of effect size.
For a fictional checkout experiment, detecting a very small change in an uncommon purchase event may require substantial traffic. A handful of usability sessions can still uncover problems in the checkout, but they answer a different question from a planned estimate of conversion impact.
Use an appropriate stopping rule. A fixed-sample procedure ordinarily requires following the planned collection, while a suitable sequential design can allow interim decisions under its rules. Monitoring technical quality and safety is necessary, but it should remain distinguishable from repeatedly searching for a favourable statistical threshold.
Translate the result without overstating it
Rejecting the null at a selected level indicates a result incompatible enough with that model to meet the decision rule. Failing to reject does not prove the null true. Neither outcome assigns a probability that the observed finding happened by chance.
Report the effect estimate, confidence interval, p-value where used and relevant assumptions. Statistical significance does not measure practical importance. Examine consequential secondary outcomes and disclose multiple comparisons or exploratory analyses.
A hypothesis test can be used with observational data, but its p-value does not resolve correlation versus causation. Causal interpretation depends on the research design and assumptions. The procedure is most useful when the team can explain both what the statistical result establishes and what further judgement the decision requires.
