How do I anonymize qualitative interview data?
Removing names is the easy quarter of the job. What identifies people in an interview is the combination of details: a role, an institution, an unusual event, a timeframe. Work through the transcript looking for combinations rather than words, and generalize the details that make a person locatable.
Updated
The mistake is treating this as a find-and-replace task. Names, places, and organizations are the obvious identifiers and they are rarely what actually identifies someone. In a small professional community, a person is locatable from a role plus a timeframe plus an unusual event, none of which is a proper noun.
So read for combinations. A participant described as a departmental head who took extended leave in 2023 at a mid-sized institution may be identifiable to anyone in the field even with every name removed. The response is to generalize rather than delete: a senior manager, a period of absence, a large organization. You keep the analytical content and lose the locating precision.
The vocabulary from data protection work gives you a name for what you are hunting. Direct identifiers are names, phone numbers, and the like. Indirect identifiers are attributes such as age, occupation, place, and date of birth, which identify no one alone but identify one person once combined, at which point they are called quasi-identifiers (Lulamba et al., 2025, pp. 3-4). That is the same combination problem you meet in a transcript, arriving from structured health datasets rather than from interviews, so treat the categories as a checklist and not as a procedure you can run on prose.
When you judge whether a combination is identifying, the test is not what the transcript contains but what a reader can add to it. Lulamba and colleagues recommend considering the background knowledge available through public records, social media, employers, professional organisations, and personal contacts (Lulamba et al., 2025, p. 9). For a study inside one department, personal contacts are the strongest of those, and they are the ones your ethics committee is least likely to ask about.
Direct quotations need particular attention. A vivid phrase is memorable, and a colleague who was present at the event may recognize both the speaker and the occasion. Where a quotation is essential and identifying, consider paraphrasing and saying that you did.
Do this before analysis rather than before publication. Working with pseudonymized transcripts throughout means the identifiable version exists in one place, encrypted and separately stored, rather than in every file you have opened for eighteen months.
Keep a log of the changes you made. Not the key linking pseudonyms to people, which is stored separately, but a record of the kinds of generalization applied, since a reader needs to know that ages were banded and institutions described by type rather than named.
What is the difference between anonymization and pseudonymization?
Pseudonymization replaces identifiers with codes while a key linking them still exists somewhere. Anonymization removes the possibility of re-identification entirely, including through the key. Most qualitative research is pseudonymized, and calling it anonymized in a consent form is a claim you cannot keep.
The distinction has legal weight in many jurisdictions, because pseudonymized data is still personal data and remains subject to data protection rules, while genuinely anonymous data is not.
There is a third word you also owe participants accurately. Saunders and colleagues separate confidentiality, which means keeping information from everyone outside the research team, from anonymity, which concerns keeping identities secret, and they point out that confidentiality sometimes requires withholding part of what a participant said rather than only removing the names attached to it (Saunders et al., 2015, p. 617). So a passage can be fully pseudonymized and still be a breach, because publishing it hands a specific reader something the participant meant only for you.
Use the accurate word in your ethics application, your consent forms, and your thesis. Participants told their data would be anonymous have been promised something specific, and if a key exists that promise is not accurate.
Can I share anonymized transcripts?
Only if your consent covered it and the anonymization genuinely holds. Rich qualitative data is hard to deidentify without destroying what makes it useful, which is why controlled access, where a named person approves each request, is often the better answer than open deposit.
If you intend to share, say so in the consent form before collecting, in plain language describing what will be shared and with whom. Retrofitting consent afterwards is usually impossible and always awkward.
Where sharing the full transcripts is not appropriate, there are still useful things to deposit: the interview guide, the coding frame, the metadata record, and a description of the dataset. Those make your study findable and partly reusable without exposing anyone.
Sources
- Saunders et al., Anonymising interview data: challenges and compromise in practice, Qualitative Research (2015) · checked 6 August 2026
- Lulamba et al., Ten quick tips for protecting health data using de-identification and perturbation of structured datasets, PLoS Computational Biology (2025) · checked 6 August 2026
- Korstjens and Moser, Series: Practical guidance to qualitative research. Part 4: Trustworthiness and publishing, European Journal of General Practice (2018) · checked 6 August 2026