Medical record anonymization for clinics, hospitals and research groups
A medical record concentrates special-category data —health— and GDPR applies its highest level of protection to it. At the same time, records are the raw material of research, teaching and care improvement. Real, irreversible anonymization is the exit from the reinforced regime: once anonymized, clinical data can be used without explicit consent; while it remains personal, a specific legal basis is required. anonimiza.do processes reports, notes and exports in seconds, free text included, and records every operation.
Which clinical documents carry identifiable data
A record holds more layers of identifiable information than people usually assume:
- Admission, progress and discharge reports, nursing notes and clinical judgement: the narrative content, with names, references to third parties, dates and places embedded in prose.
- Identification data: full name, DNI, NIE or passport, health card, Social Security number, medical record number, phone, email and address.
- Care-setting identifiers: the centre (in a small area it identifies on its own), the doctor or team, ward, room and bed, exact dates of admission, discharge and consultation.
- Diagnostic tests, laboratory and imaging results, with the hospital’s identifying label.
- Exports from the electronic health record to a spreadsheet for a study, with file metadata and identification columns that are hidden but not deleted.
- Research paperwork: informed consents, case report forms, reports to the ethics committee.
When medical records must be anonymized
Patient care, billing and the patient’s own access to their record have their own legal basis. What requires anonymization is secondary use:
Retrospective research
A cross-sectional study on existing records does not need the patient’s identity. With real, irreversible anonymization the data stops being personal and the project no longer depends on GDPR explicit consent, although the Spanish Biomedical Research Act and regional rules may still require informing patients about secondary use.
Longitudinal studies and biological samples
When patients must be followed over time, samples linked to new data, or contacted about a relevant finding, anonymization is not possible: the right technique is double-key pseudonymization, with the mapping table held by someone independent of the research team.
Teaching, clinical sessions and publications
Presenting a case at a session, a conference or in a journal requires that nobody can re-identify the patient by combining age, sex, location and an uncommon condition. With rare diagnoses, removing the name is not enough.
Transfers to other centres, universities or companies
Sharing records with a group at another hospital, a university or a company building a model requires prior anonymization or, if pseudonymized, a processing agreement and, outside the EU, the safeguards of GDPR Chapter V. Anonymized, they can be shared without those restrictions.
Complaints, expert reports and quality audits
The expert, the quality auditor or the committee reviewing an episode need the clinical content, not the identity of the other patients who appear in the same report or export. Third parties are anonymized.
Test environments and record-system migrations
Loading real records into a test environment when changing systems, or sending them to the vendor to reproduce an incident, is a transfer without a legal basis. Anonymize first or generate synthetic data.
What to remove, what to generalise and what to watch
Clinical anonymization is more demanding than for other documents: removing names is not enough when the diagnosis itself singles someone out.
- Direct identifiers, to be removed entirely: name, DNI, NIE, passport, health card, Social Security, medical record number, phone, email, address. Replacing them with a code would be pseudonymization, not anonymization.
- Care-setting identifiers: centre, doctor, ward, bed and exact dates of admission, discharge and consultation.
- Quasi-identifiers, to be generalised: date of birth to age band, postcode to province, occupation to sector, exact dates to month.
- Narrative content: this is where simple processes fail. Structured fields clean up well; progress notes carry names, relatives, neighbours and places in free text.
- Rare diagnoses and dates: with a five-year age band, sex, postcode and an uncommon condition, one person can be unique in a district of 50,000. Check that every combination of quasi-identifiers has at least k identical records, with k of 5 as the reference for clinical data.
- File metadata: author, creation date, device and change history reveal which doctor produced the document.
Recommended clinical anonymization process
- Define the use case: longitudinal follow-up (double-key pseudonymization) or aggregated data for a cross-sectional study (irreversible anonymization). The level depends on the intended use.
- Classify fields into direct identifiers, care-setting identifiers and quasi-identifiers.
- Remove direct identifiers entirely and generalise quasi-identifiers.
- Process the narrative text with detection tuned to clinical Spanish: names, third-party references, locations and dates inside the prose.
- Remove file metadata and assess re-identification risk: k-anonymity over each combination of quasi-identifiers.
- Document the procedure —date, person responsible, techniques, test result— and go through the Research Ethics Committee before the project starts.
What anonimiza.do brings to a healthcare centre or research group
- Detects the 20 data types in the catalogue —name, NIF, NIE, passport, Social Security, address, phone, email, dates— with Spanish formats validated by check digit, and accepts custom data types described in plain language for what is not in the catalogue: medical record number, health card, CIAS code, diagnosis.
- Works on the free text of reports and progress notes, not only on structured fields, and on spreadsheets column by column for electronic health record exports.
- Two modes depending on the study: irreversible anonymization, or reversible pseudonymization that keeps the patient traceable within the batch; plus k-anonymity verification over the dataset.
- OCR for scanned reports and batch processing, with file metadata removed before download.
- Audit log for every document processed: the evidence presented to the Ethics Committee and the data protection authority.
- Data always in the European Union (AWS Frankfurt), with optional deployment in the AWS Spain region; Spanish National Security Framework (ENS) MEDIUM category and a data processing agreement.
- Free plan of 3 documents a month with no card; Professional plan with OCR, batches and API; Enterprise plan with unlimited volume and SSO for hospitals and groups.
Try for free: 3 documents a month
Frequently asked questions
Is patient consent needed to use a record in research if it is anonymized?
Once the data is properly anonymized there is no personal data, and GDPR consent does not apply. Even so, the Spanish Biomedical Research Act and regional rules often require informing patients in advance about secondary use of their data, even anonymized.
Anonymize or pseudonymize a medical record?
If the study never needs to go back to the patient, irreversible anonymization: direct identifiers are removed and quasi-identifiers generalised. If there is longitudinal follow-up, biological samples or a possibility of contacting the patient, double-key pseudonymization with the mapping table held by an independent custodian.
Can I share anonymized records with researchers outside the EU?
Yes. Once genuinely anonymized they can be shared without the restrictions of GDPR Chapter V. If pseudonymized, they remain personal data and international-transfer rules apply.
Is removing the name and ID number enough?
No. The free text of progress notes carries as much personal data as structured fields, and an uncommon diagnosis combined with age, sex and location can identify a single person. You must process the narrative text, generalise quasi-identifiers and check k-anonymity.
How long does it take to anonymize a 50-page record by hand?
Two to four hours, and the result is inconsistent because each person applies different criteria. With a specialised tool the same record is processed in seconds with uniform criteria.
Doesn’t the electronic health record already anonymize on export?
Some systems strip direct identifiers on export, but almost none generalise quasi-identifiers, process free text or assess re-identification risk. Research needs an additional layer.