De-Identification of Electronic Health Records Using Deep Learning and Transformers


Dilmaç F., Alpkocak A.

Applied Sciences (Switzerland), cilt.16, sa.4, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 16 Sayı: 4
  • Basım Tarihi: 2026
  • Doi Numarası: 10.3390/app16041692
  • Dergi Adı: Applied Sciences (Switzerland)
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Compendex, INSPEC, Directory of Open Access Journals
  • Anahtar Kelimeler: electronic health records, de-identification, named entity recognition, large language models, clinical natural language processing
  • Dokuz Eylül Üniversitesi Adresli: Evet

Özet

Adoption of electronic health records (EHRs) has significantly advanced healthcare by enabling extensive data storage and analysis for clinical decisions and research. However, sensitive personally identifiable information (PII) within EHRs presents major challenges concerning patient privacy, data security, and regulatory compliance. Effective automated de-identification techniques for detecting and removing protected health information (PHI) are thus essential. This study presents one of the first focused studies on Turkish EHR de-identification, comparing traditional sequence-based neural architectures with advanced transformer-based large language models (LLMs) for PHI detection. We introduce and publicly release a manually annotated benchmark dataset of TEHRs, covering diverse PHI types, supporting further research in Turkish clinical text. Two methodologies were evaluated: bidirectional long short-term memory (BiLSTM) models (with and without Conditional Random Fields (CRFs)) and six fine-tuned pre-trained LLMs. Experiments demonstrated the superior performance of transformer-based LLMs, achieving a macro F1 score of 92.20%, significantly outperforming traditional methods. Among sequence-based models, BiLSTM + CRF attained an 83.00% F1 score, exceeding the baseline BiLSTM 78.40%. Results highlight the potential of transformer-based models for privacy-preserving Turkish clinical text and underscore the importance of annotated benchmark datasets.