Introduction
A bathroom scale that reads differently every time you step on it is useless even before you ask whether it's accurate. Reliability is that first hurdle for every research instrument: consistency. Does the survey, the coding scheme, the metric produce the same result under the same conditions, across time, items, and raters? Without it, observed changes are indistinguishable from measurement wobble, and trend lines are fiction. This article covers the main forms of reliability, how each is assessed, and the working relationship between consistency and truth.
What is Reliability?
Reliability is the consistency of a measurement: the degree to which an instrument produces the same results under the same conditions. It is deliberately distinct from validity, which asks whether the instrument measures the right thing; the scale that always reads four pounds heavy is perfectly reliable and consistently wrong. The relationship is asymmetric and foundational: reliability is necessary for validity but never sufficient, because an instrument that can't agree with itself can't be measuring anything in particular, while an instrument that agrees with itself may still be measuring the wrong thing beautifully.
The Main Forms
Test-retest reliability. Stability across time: administer the same measure to the same people twice, and see whether results hold (for constructs assumed stable across the interval). Low test-retest consistency means observed movement in your tracker may be the instrument breathing, not the population changing.
Internal consistency. Coherence across items: do the questions of a multi-item scale (the ten items of the SUS, a five-item satisfaction battery) behave as measures of one underlying thing? The standard index is Cronbach's alpha, with values in the high .70s and above conventionally read as acceptable, and suspiciously high values (.95+) hinting at redundant items rather than virtue.
Inter-rater reliability. Agreement across judges: when two researchers independently code the same interview transcripts, score the same usability sessions for task success, or rate the same issues for severity, how often do they agree beyond chance? Indices like Cohen's kappa quantify it, and the practice behind it (codebooks, definitions, calibration rounds) does as much for quality as the number does.
Parallel-forms reliability. Equivalence across versions: do two forms of the same instrument (rotated question sets, alternate task versions) yield comparable results? The concern behind every "we updated the survey" conversation.
Reliability in Product Research
The concept operationalises directly. Trend lines demand instrument stability: a benchmark programme that reworded its questions mid-stream is comparing two different rulers, which is why core items stay frozen even when imperfect. Behavioural metrics need defined rubrics: "task success" scored by vibe drifts between studies and between raters, so a written rubric plus an occasional double-scored sample is cheap insurance. Coding of qualitative data earns trust through second coders on samples and documented codebooks. And standardised instruments exist largely because reliability is expensive to build: the published consistency of the SUS is precisely what a home-grown five-item alternative lacks, and why questionnaire craft leans on validated scales for anything tracked over time.
Improving It
1. Standardise the conditions.
Same instructions, same tasks, same recruitment profile, same context; every varying condition is noise wearing a lab coat.
2. Write the rubric down.
Definitions of success, severity, and codes, concrete enough that a new rater lands where the old one did. Calibrate on shared examples before scoring alone.
3. Use multiple items for fuzzy constructs.
Single questions carry single-question noise; small validated batteries average it down, which is the whole logic of composite scales.
4. Pilot for consistency, not just clarity.
A pilot that double-codes a few sessions or re-runs a few respondents catches unreliability while it's still an edit.
The Takeaway
Reliability is the boring virtue everything else stands on: consistent instruments, stable conditions, written rubrics, agreeing raters. It never guarantees you're measuring the right thing, but without it you're measuring nothing in particular, with confidence. Make the ruler steady first; then argue about whether it's the right ruler, which is validity's department.
Further reading
For the forms and their assessment:
Articles:
1. Reliability vs. Validity in Research - Scribbr
The two concepts and their relationship, with the main reliability types and how each is evaluated.
2. Measuring Usability with the System Usability Scale - MeasuringU
A standardised instrument whose published reliability is the reason it beats home-grown alternatives for tracking.