Glossary

Confidence Interval

Glossary

Confidence Interval

Confidence Interval

A confidence interval is a range calculated from data by a method with a stated long-run coverage rate, expressing uncertainty around an estimate under its assumptions.

Twelve people out of twenty complete a task. The observed success rate is 60%, but another sample could produce a different result. A confidence interval expresses uncertainty around an estimate using a method with a stated long-run coverage rate, under its assumptions.

The interval helps the team see how much the data resolve. A narrow range may support a fairly precise decision; a wide range may leave materially different outcomes compatible with the evidence. Neither width nor the confidence level protects against a poorly defined measure or biased recruitment.

What a 95% confidence interval means

If the study and calculation were repeated under the same assumptions, a method with 95% coverage would produce intervals containing the target value about 95% of the time. The confidence level describes that procedure. It is not a probability assigned to the fixed population value being inside this particular interval after calculation.

NIST’s explanation of confidence intervals sets out this repeated-sampling interpretation. For everyday reporting, keep the estimate, interval, confidence level and method together, rather than reducing the result to a percentage with unexplained brackets.

An interval for an average or proportion is also different from a range containing individual outcomes. A confidence interval around mean task time does not say that 95% of participants finish within those bounds.

Read the range in terms of the decision

For the illustrative result of 12 successes in 20 independent binary observations, a 95% Wilson interval is approximately 39% to 78%. The observed rate remains 60%; the range shows the limited precision of that small dataset under the model. NIST describes the Wilson method for proportions, which avoids some problems of a simple symmetric approximation.

If the team needs strong evidence that performance exceeds a demanding target, those observations may leave too much uncertainty. If the sessions exposed a clear and inexpensive wording problem, the team may reasonably fix it without waiting to estimate the population success rate precisely. The interval informs the decision; it does not determine which decisions are worth making.

For a comparison, calculate an interval for the difference or relevant effect directly. Looking at whether two separate intervals overlap is not a reliable substitute for analysing the comparison, because the relationship depends on the design and the estimates’ dependence.

Use a method that matches the data

Repeated attempts from the same person are not necessarily independent. Responses from users within the same organisation may be related. Weighting, clustering and sequential analysis can change the appropriate calculation. A familiar formula applied to the wrong structure can give an unjustifiably narrow range.

Larger samples, lower variability and a lower chosen confidence level can narrow intervals, but changing the confidence level after seeing the result misrepresents the planned analysis. Explain the method and assumptions, including any important departures.

Sampling bias and response bias remain outside a conventional sampling interval. Report them separately. A confidence interval is most useful when it makes a sound estimate’s uncertainty visible, rather than giving questionable data an appearance of mathematical authority.

Further reading

Guides

  1. Confidence intervals — NIST
    Explains how intervals express uncertainty under a statistical procedure. Useful when checking the interpretation behind a range, rather than treating it as a guarantee about the result.

  2. Confidence intervals for proportions — NIST
    A technical reference for estimates such as a task-success rate. Read it when a small sample or a proportion near zero or one makes the familiar symmetric interval unreliable.