Introduction
"Your responses are anonymous" is the most common promise in research and one of the most carelessly made. True anonymity means nobody, including the research team, can link data to a person; what most studies actually offer is confidentiality, where the link exists but is protected. The gap matters legally, ethically, and practically, and it widens as recordings, rich profiles, and small samples make re-identification easier than intuition suggests. This article covers anonymity versus confidentiality, why de-identification is harder than deleting names, and how to promise only what you can deliver.
What is Anonymity?
Anonymity in research means that data cannot be linked to the individual who provided it: no identifiers are collected, or those collected are irreversibly removed, so that not even the researchers can say who said what. Its more common neighbour is confidentiality: identities are known (or knowable) to the team but protected by access controls, pseudonyms, and disclosure limits. Nearly every study that records video, pays incentives, follows up, or screens on profile data is confidential rather than anonymous, and describing it as anonymous is a false promise that participants rely on when deciding what to say. The distinction sits at the centre of informed consent: people agree to the protection actually on offer, so the offer must be stated accurately.
Why De-Identification Is Hard
Removing names is the easy part; the hard part is quasi-identifiers, ordinary attributes that combine into a fingerprint. Latanya Sweeney's well-known finding, that a large majority of the US population could be uniquely identified from just ZIP code, birth date, and sex, made the point two decades ago, and product research data is richer than that: job title plus company size plus city plus a distinctive quote identifies a person to anyone who knows the industry. Small samples sharpen the problem (a "senior engineer at a fintech in Leeds" in a twelve-person study is a name in disguise), and recordings make it acute: a face, a voice, or a screen showing an inbox is identity itself. Standards like GDPR draw the line accordingly: pseudonymised data (identifiers replaced but re-linkable) remains personal data with full obligations; only genuinely anonymised data, where re-identification is not reasonably likely, leaves the regime.
The Working Toolkit
1. Collect less.
The strongest protection is data that never existed: don't collect names, emails, or precise attributes the study doesn't need, and separate recruitment records (which need identity) from research data (which usually doesn't).
2. Pseudonymise early, key separately.
Participant IDs in the dataset, the ID-to-identity key stored apart with restricted access and a deletion date. This is confidentiality engineering, and it should be described as such.
3. Generalise and suppress in reporting.
Coarsen attributes (industry rather than employer, region rather than city, band rather than exact age), remove distinctive details from quotes, and drop cells small enough to identify; the discipline formalised in statistical disclosure control.
4. Treat recordings as identity.
Consent for recording is consent for a known identity, with specific permissions for each downstream use; clips shared beyond the agreed circle need explicit permission, and a "highlight reel" of faces is not anonymous by any definition. Platforms that hold recordings behind access controls with permissioned sharing (as a Ballpark workspace does) make the confidentiality promise operable rather than aspirational.
5. Set retention and honour it.
Data kept indefinitely is data awaiting a breach; deletion schedules are part of the promise.
Promising Honestly
The consent language should match the mechanism: "anonymous" only for genuinely unlinked collection (an unrecorded survey with no identifiers and no profile join); "confidential" for everything else, with a plain sentence about who can access what and how reporting will be generalised. Participants disclose more when the protection is credible and specific than when it is sweeping and vague, so accuracy here is not only ethical but data-improving. And where anonymity is impossible in principle (employee research in a small team, studies of a niche role), say so, and design the reporting so candour cannot be traced back to the candid.
The Takeaway
Anonymity is the absence of any link to a person; confidentiality is a protected link; most research offers the second while promising the first. Collect less, pseudonymise with a separate key, generalise in reporting, treat recordings as identity, delete on schedule, and write consent that names the protection actually provided. The promise participants remember is the one you can keep.
Further reading
For the standards and their application:
Articles:
1. Anonymisation Guidance - UK Information Commissioner's Office
The regulatory line between anonymised and pseudonymised data, with practical techniques and risk assessment.
2. Ethical Considerations in Research - Scribbr
Anonymity and confidentiality in the wider ethics framework, including consent language.