Two researchers watch the same usability session. One records the task as successful because the participant reached the final screen; the other records failure because the submitted information was wrong. Their disagreement reveals a measurement problem before any average is calculated.
Reliability concerns consistency under specified measurement conditions. Those conditions might involve repeated measurement over time, different raters or related items in a scale. The relevant form depends on what the team needs to measure and how the result will be used.
Decide which consistency matters
Test–retest reliability concerns stability when the underlying attribute is expected to remain stable across the interval. If a person’s experience genuinely changes between measurements, a different score is not automatically unreliability. The study needs a reasonable account of what should have stayed the same.
Inter-rater reliability concerns consistency between people making assessments. Internal consistency concerns relationships among items intended to measure a common construct. A high internal-consistency coefficient does not establish that a questionnaire is unidimensional or valid; redundant items can also produce high values.
The COSMIN measurement framework distinguishes measurement properties in health-related instruments. Its terminology is useful for understanding why reliability, measurement error and validity require different evidence, although its specific standards should not be applied indiscriminately to every product metric.
Make the scoring rule concrete
Return to the usability example and define success before the next round. Does it require an accurate submission, independent completion and recognition that the task is finished? What happens when the researcher helps or the prototype fails? A scoring guide should make those distinctions clear enough to apply consistently.
Have raters independently assess a suitable sample, then examine disagreements. Choose a measure appropriate to the type of score and design rather than defaulting to a familiar coefficient. Document the rule and any changes, so future comparisons do not unknowingly use a different definition.
For recurring surveys, retain the wording, response options and relevant administration conditions where comparability requires them. If the instrument must change, investigate the effect instead of assuming the old and new series are continuous.
Do not confuse agreement with correctness
Validity asks whether the result supports its intended interpretation. Two raters can agree on a poorly chosen definition, and a stable score can measure the wrong thing. Reliability evidence helps assess measurement quality without settling those questions.
Qualitative analysis also contains different methodological traditions. A shared coding framework may appropriately assess coding consistency, while reflexive thematic analysis treats interpretation differently and does not necessarily seek interchangeable coders. Choose quality criteria that fit the approach rather than imposing inter-rater agreement on all qualitative work.
Report the form of reliability assessed, the sample, conditions and relevant uncertainty. Avoid a universal threshold that declares every measure acceptable. The required consistency depends on the consequence of error and the decision the measurement supports.
