One participant takes twenty minutes to finish a task that everyone else completes in two. Removing the result would make the chart tidier, but it might also remove the most consequential experience in the study. The first question is what happened, not which deletion rule produces a more comfortable average.
An outlier is an observation unusually distant from other values or from a pattern described by a model. It may be an error, a legitimate extreme case or evidence that the assumed model does not fit the data well.
Investigate the source of the unusual value
In a fictional usability study, a long task time could reflect a timer left running during a break, a recording error or a genuine struggle with the interface. Check the source material and the measurement definition before deciding which explanation applies.
NIST’s guidance on outlier detection distinguishes identifying unusual observations from deciding how to handle them. A statistical flag is a prompt for investigation, not proof that a record is invalid.
Consider the context and distribution. A large purchase may be ordinary for an enterprise account and unusual for an individual plan. Applying one threshold across both groups can discard precisely the variation the analysis needs to understand.
Choose handling that matches the question
Correct a verified error when the intended value can be recovered, preserving a record of the change. If the value is genuine, consider whether the planned analysis adequately represents it. Robust summaries, transformations or models suited to the distribution may be more appropriate than deletion, but each changes the analytical treatment and should be explained.
A median can reduce the influence of an extreme task time on the centre without making that experience disappear. Report the distribution or relevant extremes when they matter to the decision. Central tendency is only one part of the description.
Make sensitive conclusions visible
Define exclusion rules in advance where feasible, and document new decisions made after inspecting the data. If a borderline case materially changes the conclusion, show a suitable sensitivity analysis rather than silently selecting the favourable version.
A small dataset may not support a confident distinction between a rare valid case and a separate process. State that uncertainty. Data cleaning should improve the credibility of the analysis, not remove inconvenient experiences until the evidence matches the team’s expectation.
