{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T04:41:40Z","timestamp":1777696900640,"version":"3.51.4"},"reference-count":36,"publisher":"SAGE Publications","issue":"1","license":[{"start":{"date-parts":[[2025,5,11]],"date-time":"2025-05-11T00:00:00Z","timestamp":1746921600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["Intelligent Data Analysis: An International Journal"],"published-print":{"date-parts":[[2026,1]]},"abstract":"<jats:p>\n                    Information leakage and model attacks pose risks in the analysis and exchange of medical data, as the language models used to process medical records may retain training data. Traditional models, on the other hand, are often too complicated and use old ineffectual methods for removing personal data. This can compromise the data\u2019s integrity and quality, making it less useful for future tasks, especially when combined with other language models. This paper introduces the Self-Decoded Model of Medical De-identification (SDM-M-DID). The model employs a secure BERT-based encoder to paraphrase sensitive data, ensuring HIPAA compliance. Unlike traditional models that only mask sensitive tokens, the SDM-M-DID decodes its own embeddings to generate an internal representations of these tokens. Then, it integrates this representations with the pre-trained BERT dictionary to rephrase tokens, preserving their semantic role while altering grammar to prevent re-identification. Compared to existing large language models, our model achieves a score of 0.8416 F1 BERTscore, striking an optimal balance between the variability and similarity of deidentified tokens. We conducted experiments on two medical datasets to demonstrate the effectiveness of the model. Metrics show that there is only a\n                    <jats:inline-formula>\n                      <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\" display=\"inline\" overflow=\"scroll\">\n                        <mml:mo>\u00b1<\/mml:mo>\n                        <mml:mn>1<\/mml:mn>\n                      <\/mml:math>\n                    <\/jats:inline-formula>\n                    % to\n                    <jats:inline-formula>\n                      <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\" display=\"inline\" overflow=\"scroll\">\n                        <mml:mo>\u00b1<\/mml:mo>\n                        <mml:mn>2<\/mml:mn>\n                        <mml:mi mathvariant=\"normal\">%<\/mml:mi>\n                      <\/mml:math>\n                    <\/jats:inline-formula>\n                    difference in accuracy between the original datasets and the de-identified datasets. In total, this demonstrates that SDM-M-DID not only effectively preserves data integrity and is not inferior in efficiency to new large language models but even improves it in some cases while using a more secure and less resource-intensive technology.\n                  <\/jats:p>","DOI":"10.1177\/1088467x251332846","type":"journal-article","created":{"date-parts":[[2025,5,12]],"date-time":"2025-05-12T01:11:43Z","timestamp":1747012303000},"page":"127-144","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":0,"title":["SDM-M-DID: Self-decoding model for medical de-identification"],"prefix":"10.1177","volume":"30","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-0862-6652","authenticated-orcid":false,"given":"Bohdan","family":"Budiakov","sequence":"first","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0485-839X","authenticated-orcid":false,"given":"Tong","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Junwen","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-5012-3170","authenticated-orcid":false,"given":"Alexey","family":"Karev","sequence":"additional","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2025,5,11]]},"reference":[{"key":"e_1_3_4_2_2","unstructured":"Vakili T Lamproudis A Henriksson A et\u00a0al. Downstream Task Performance of BERT Models Pre-Trained Using Automatically De-Identified Clinical Data. 2022 4245\u20134252. https:\/\/aclanthology.org\/2022.lrec-1.451."},{"key":"e_1_3_4_3_2","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2288-10-70"},{"key":"e_1_3_4_4_2","doi-asserted-by":"crossref","unstructured":"Grouin C Griffon N N\u00e9v\u00e9ol A. Is it possible to recover personal health information from an automatically de-identified corpus of French EHRs? In: Proceedings of the sixth international workshop on health text mining information analysis. 2015 pp.31\u201339.","DOI":"10.18653\/v1\/W15-2604"},{"key":"e_1_3_4_5_2","doi-asserted-by":"publisher","DOI":"10.3390\/fi13050136"},{"key":"e_1_3_4_6_2","unstructured":"Seyedi S Xiong L Nemati S et\u00a0al. An analysis of protected health information leakage in deep-learning based de-identification algorithms. arXiv preprint arXiv:2101.12099 2021."},{"key":"e_1_3_4_7_2","unstructured":"Larbi IBC Burchardt A Roller R. Which anonymization technique is best for which NLP task?\u2013It depends. A Systematic Study on Clinical Text Processing. arXiv preprint arXiv:2209.00262 2022."},{"key":"e_1_3_4_8_2","unstructured":"Mikolov T. Efficient estimation of word representations in vector space arXiv preprint arXiv:1301.3781 2013; 3781."},{"key":"e_1_3_4_9_2","doi-asserted-by":"crossref","unstructured":"Pennington J Socher R Manning CD. Glove: Global vectors for word representation. In: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) 2014 pp.1532\u20131543.","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_4_10_2","unstructured":"Devlin J. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 2018."},{"key":"e_1_3_4_11_2","doi-asserted-by":"crossref","unstructured":"Peters ME Neumann M Zettlemoyer L et\u00a0al. Dissecting contextual word embeddings: Architecture and representation. arXiv preprint arXiv:1808.08949 2018.","DOI":"10.18653\/v1\/D18-1179"},{"key":"e_1_3_4_12_2","unstructured":"Carlini N Tramer F Wallace E et\u00a0al. Extracting training data from large language models. In: 30th USENIX security symposium (USENIX Security 21) 2021 pp.2633\u20132650."},{"key":"e_1_3_4_13_2","unstructured":"Nakamura Y Hanaoka S Nomura Y et\u00a0al. KART: Parameterization of privacy leakage scenarios from pre-trained language models. arXiv preprint arXiv:2101.00036 2020."},{"key":"e_1_3_4_14_2","doi-asserted-by":"crossref","unstructured":"Liu X He P Chen W et\u00a0al. Multi-task deep neural networks for natural language understanding. arXiv preprint arXiv:1901.11504 2019.","DOI":"10.18653\/v1\/P19-1441"},{"key":"e_1_3_4_15_2","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel C","year":"2020","unstructured":"Raffel C, Shazeer N, Roberts A, et\u00a0al. Exploring the limits of transfer learning with a unified text-to-text transformer. J Mach Learn Res 2020; 21: 1\u201367.","journal-title":"J Mach Learn Res"},{"key":"e_1_3_4_16_2","doi-asserted-by":"crossref","unstructured":"Sun T Shao Y Qiu X et\u00a0al. Colake: Contextualized language and knowledge embedding. arXiv preprint arXiv:2010.00309 2020.","DOI":"10.18653\/v1\/2020.coling-main.327"},{"key":"e_1_3_4_17_2","doi-asserted-by":"crossref","unstructured":"Dong L Mallinson J Reddy S et\u00a0al. Learning to paraphrase for question answering. arXiv preprint arXiv:1708.06022 2017.","DOI":"10.18653\/v1\/D17-1091"},{"key":"e_1_3_4_18_2","first-page":"6000","article-title":"Attention is all you need","author":"Vaswani A","year":"2017","unstructured":"Vaswani A. Attention is all you need. Adv Neural Inf Process Syst 2017: 6000\u20136010.","journal-title":"Adv Neural Inf Process Syst"},{"key":"e_1_3_4_19_2","doi-asserted-by":"publisher","DOI":"10.1089\/blr.2023.29329.aso"},{"key":"e_1_3_4_20_2","doi-asserted-by":"crossref","unstructured":"Berg H Henriksson A Dalianis H. The impact of de-identification on downstream named entity recognition in clinical text. In: Proceedings of the 11th international workshop on health text mining and information analysis 2020 pp.1\u201311.","DOI":"10.18653\/v1\/2020.louhi-1.1"},{"key":"e_1_3_4_21_2","doi-asserted-by":"crossref","unstructured":"Vakili T Dalianis H. Utility preservation of clinical text after De-Identification. In: Proceedings of the 21st workshop on biomedical language processing 2022 pp.383\u2013388.","DOI":"10.18653\/v1\/2022.bionlp-1.38"},{"key":"e_1_3_4_22_2","unstructured":"n2c2 NLP Research Data Sets. https:\/\/portal.dbmi.hms.harvard.edu\/projects\/n2c2-nlp\/."},{"key":"e_1_3_4_23_2","unstructured":"Liu Y. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692364 (2019)."},{"key":"e_1_3_4_24_2","unstructured":"Sanh V. DistilBERT a distilled version of BERT: smaller faster cheaper and lighter. arXiv preprint arXiv:1910.01108 2019."},{"key":"e_1_3_4_25_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41746-021-00455-y"},{"key":"e_1_3_4_26_2","unstructured":"Healthcare Dataset 2024. https:\/\/www.kaggle.com\/datasets\/prasad22\/healthcare-dataset."},{"key":"e_1_3_4_27_2","unstructured":"Loshchilov I Hutter F et\u00a0al. Fixing weight decay regularization in adam. arXiv preprint arXiv:1711.05101 2017; 5."},{"key":"e_1_3_4_28_2","unstructured":"Aung KMM. Comparison of levenshtein distance algorithm and needleman-wunsch distance algorithm for string matching. PhD thesis MERAL Portal 2019."},{"key":"e_1_3_4_29_2","doi-asserted-by":"crossref","unstructured":"Anderson J Tarigan JT Sharif A. Damerau-Levenshtein distance and cosine similarity to select the optimal word in word typing game. In: AIP conference proceedings Vol. 2987 AIP Publishing 2024.","DOI":"10.1063\/5.0199835"},{"key":"e_1_3_4_30_2","unstructured":"Hanna M Bojar O. A Fine-Grained Analysis of BERTScore. In: Proceedings of the sixth conference on machine translation Barrault L Bojar O Bougares F et\u00a0al. (eds) Association for Computational Linguistics Online 2021 pp.507\u2013517. https:\/\/aclanthology.org\/2021.wmt-1.59."},{"key":"e_1_3_4_31_2","doi-asserted-by":"crossref","unstructured":"Liu S-y. Liu Z Huang X et\u00a0al. LLM-FP4: 4-Bit Floating-Point Quantized Transformers. In: Proceedings of the 2023 conference on empirical methods in natural language processing Association for Computational Linguistics 2023 pp.592\u2013605. doi:10.18653\/v1\/2023.emnlp-main.39.","DOI":"10.18653\/v1\/2023.emnlp-main.39"},{"key":"e_1_3_4_32_2","doi-asserted-by":"crossref","unstructured":"Kanakarajan KR Kundumani B Sankarasubbu M. BioELECTRA: pretrained biomedical text encoder using discriminators. In: Proceedings of the 20th workshop on biomedical language processing 2021 pp.143\u2013154.","DOI":"10.18653\/v1\/2021.bionlp-1.16"},{"key":"e_1_3_4_33_2","unstructured":"Kotfic GitHub - kotfic\/i2b2-evaluation-scripts: Repository for managing python tools that model standoff annotations for i2b2 2014 challenge. https:\/\/github.com\/kotfic\/i2b2_evaluation_scripts."},{"key":"e_1_3_4_34_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbi.2015.06.015"},{"key":"e_1_3_4_35_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbi.2017.05.023"},{"key":"e_1_3_4_36_2","doi-asserted-by":"crossref","unstructured":"Beryozkin G Drori Y Gilon O et\u00a0al. A Joint Named-Entity Recognizer for Heterogeneous Tag-sets Using a Tag Hierarchy. 2019 140\u2013150. doi:10.18653\/v1\/P19-1014. https:\/\/aclanthology.org\/P19-1014.","DOI":"10.18653\/v1\/P19-1014"},{"key":"e_1_3_4_37_2","doi-asserted-by":"publisher","DOI":"10.2196\/17622"}],"container-title":["Intelligent Data Analysis: An International Journal"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1088467X251332846","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1088467X251332846","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1088467X251332846","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:21:16Z","timestamp":1777454476000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1088467X251332846"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,11]]},"references-count":36,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1]]}},"alternative-id":["10.1177\/1088467X251332846"],"URL":"https:\/\/doi.org\/10.1177\/1088467x251332846","relation":{},"ISSN":["1088-467X","1571-4128"],"issn-type":[{"value":"1088-467X","type":"print"},{"value":"1571-4128","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,11]]}}}