{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,12]],"date-time":"2026-05-12T23:24:03Z","timestamp":1778628243992,"version":"3.51.4"},"reference-count":64,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2024,2,26]],"date-time":"2024-02-26T00:00:00Z","timestamp":1708905600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100000279","name":"Nuffield Foundation","doi-asserted-by":"publisher","award":["EP\/V047949\/1"],"award-info":[{"award-number":["EP\/V047949\/1"]}],"id":[{"id":"10.13039\/501100000279","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100014013","name":"UKRI\/EPSRC","doi-asserted-by":"publisher","award":["\u00a0"],"award-info":[{"award-number":["\u00a0"]}],"id":[{"id":"10.13039\/100014013","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Digit. Health"],"abstract":"<jats:p>Clinical text and documents contain very rich information and knowledge in healthcare, and their processing using state-of-the-art language technology becomes very important for building intelligent systems for supporting healthcare and social good. This processing includes creating language understanding models and translating resources into other natural languages to share domain-specific cross-lingual knowledge. In this work, we conduct investigations on clinical text machine translation by examining multilingual neural network models using deep learning such as Transformer based structures. Furthermore, to address the language resource imbalance issue, we also carry out experiments using a transfer learning methodology based on massive multilingual pre-trained language models (MMPLMs). The experimental results on three sub-tasks including (1) clinical case (CC), (2) clinical terminology (CT), and (3) ontological concept (OC) show that our models achieved top-level performances in the ClinSpEn-2022 shared task on English-Spanish clinical domain data. Furthermore, our expert-based human evaluations demonstrate that the small-sized pre-trained language model (PLM) outperformed the other two extra-large language models by a large margin in the clinical domain fine-tuning, which finding was never reported in the field. Finally, the transfer learning method works well in our experimental setting using the WMT21fb model to accommodate a new language space Spanish that was not seen at the pre-training stage within WMT21fb itself, which deserves more exploitation for clinical knowledge transformation, e.g. to investigate into more languages. These research findings can shed some light on domain-specific machine translation development, especially in clinical and healthcare fields. Further research projects can be carried out based on our work to improve healthcare text analytics and knowledge transformation. Our data is openly available for research purposes at: <jats:ext-link>https:\/\/github.com\/HECTA-UoM\/ClinicalNMT<\/jats:ext-link>.<\/jats:p>","DOI":"10.3389\/fdgth.2024.1211564","type":"journal-article","created":{"date-parts":[[2024,2,26]],"date-time":"2024-02-26T10:58:02Z","timestamp":1708945082000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":21,"title":["Neural machine translation of clinical text: an empirical investigation into multilingual pre-trained language models and transfer-learning"],"prefix":"10.3389","volume":"6","author":[{"given":"Lifeng","family":"Han","sequence":"first","affiliation":[]},{"given":"Serge","family":"Gladkoff","sequence":"additional","affiliation":[]},{"given":"Gleb","family":"Erofeev","sequence":"additional","affiliation":[]},{"given":"Irina","family":"Sorokina","sequence":"additional","affiliation":[]},{"given":"Betty","family":"Galiano","sequence":"additional","affiliation":[]},{"given":"Goran","family":"Nenadic","sequence":"additional","affiliation":[]}],"member":"1965","published-online":{"date-parts":[[2024,2,26]]},"reference":[{"key":"B1","first-page":"627","article-title":"Topic modelling of Swedish newspaper articles about coronoavirus: A Case Study using latent girichlet allocation method","volume-title":"IEEE 11th International Conference on Healthcare Informatics (ICHI)","author":"Grici\u016bt\u0117","year":"2023"},{"key":"B2","doi-asserted-by":"publisher","first-page":"e22734","DOI":"10.2196\/22734","article-title":"Health, psychosocial, and social issues emanating from the COVID-19 pandemic based on social media comments: text mining, thematic analysis approach","volume":"9","author":"Oyebode","year":"2021","journal-title":"JMIR Med Inform"},{"key":"B3","doi-asserted-by":"publisher","first-page":"1737","DOI":"10.1109\/JBHI.2021.3123192","article-title":"A deep language model for symptom extraction from clinical text and its application to extract COVID-19 symptoms from social media","volume":"26","author":"Luo","year":"2022","journal-title":"IEEE J Biomed Health Inform"},{"key":"B4","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1093\/jamia\/ocz166","article-title":"2018 n2c2 shared task on adverse drug events and medication extraction in electronic health records","volume":"27","author":"Henry","year":"2020","journal-title":"J Am Med Inform Assoc"},{"key":"B5","doi-asserted-by":"publisher","first-page":"e17984","DOI":"10.2196\/17984","article-title":"Clinical text data in machine learning: systematic review","volume":"8","author":"Spasic","year":"2020","journal-title":"JMIR Med Inform"},{"key":"B6","doi-asserted-by":"publisher","first-page":"165","DOI":"10.1146\/annurev-biodatasci-030421-030931","article-title":"Modern clinical text mining: a guide, review","volume":"4","author":"Percha","year":"2021","journal-title":"Annu Rev Biomed Data Sci"},{"key":"B7","doi-asserted-by":"publisher","first-page":"e38122","DOI":"10.2196\/38122","article-title":"Deployment of a free-text analytics platform at a UK national health service research hospital: cogstack at University College London hospitals","volume":"10","author":"Noor","year":"2022","journal-title":"JMIR Med Inform"},{"key":"B8","doi-asserted-by":"publisher","first-page":"15","DOI":"10.1007\/s10994-020-05921-4","article-title":"CPAS: the UK\u2019s national machine learning-based hospital capacity planning system for COVID-19","volume":"110","author":"Qian","year":"2021","journal-title":"Mach Learn"},{"key":"B9","author":"Wu","year":""},{"key":"B10","volume-title":"Span-Based Named Entity Recognition by Generating, Compressing InformationarXiv","author":"Nguyen","year":"2023"},{"key":"B11","author":"Wroge","year":""},{"key":"B12","doi-asserted-by":"publisher","first-page":"73","DOI":"10.1007\/s12539-020-00408-1","article-title":"Classification of COVID-19 by compressed chest ct image through deep learning on a large patients cohort","volume":"13","author":"Zhu","year":"2021","journal-title":"Interdiscip Sci Comput Life Sci"},{"key":"B13","author":"Costa-juss\u00e0","year":""},{"key":"B14","doi-asserted-by":"publisher","first-page":"1275","DOI":"10.1007\/s11606-021-07164-y","article-title":"A research agenda for using machine translation in clinical medicine","volume":"37","author":"Khoong","year":"2022","journal-title":"J Gen Intern Med"},{"key":"B15","author":"Weaver","year":""},{"key":"B16","first-page":"6000","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Conf Neural Inf Process Syst"},{"key":"B17","author":"Devlin","year":""},{"key":"B18","author":"Han","year":""},{"key":"B19","author":"Han","year":""},{"key":"B20","author":"Kuang","year":""},{"key":"B21","author":"Han","year":""},{"key":"B22","author":"Junczys-Dowmunt","year":""},{"key":"B23","author":"Junczys-Dowmunt","year":""},{"key":"B24","author":"Tran","year":""},{"key":"B25","author":"Neves","year":""},{"key":"B26","author":"Almansor","year":""},{"key":"B27","doi-asserted-by":"publisher","first-page":"12141","DOI":"10.1007\/s00521-021-05895-x","article-title":"Towards achieving a delicate blending between rule-based translator and neural machine translator","volume":"33","author":"Islam","year":"2021","journal-title":"Neural Comput Appl"},{"key":"B28","author":"Han","year":""},{"key":"B29","article-title":"Using massive multilingual pre-trained language models towards real zero-shot neural machine translation in clinical domain","volume-title":"arXiv","author":"Han","year":"2022"},{"key":"B30","author":"Han","year":""},{"key":"B31","doi-asserted-by":"publisher","first-page":"596","DOI":"10.1197\/jamia.M3096","article-title":"A text mining approach to the prediction of disease status from clinical discharge summaries","volume":"16","author":"Yang","year":"2009","journal-title":"J Am Med Inform Assoc JAMIA"},{"key":"B32","doi-asserted-by":"publisher","first-page":"859","DOI":"10.1136\/amiajnl-2013-001625","article-title":"Combining rules and machine learning for extraction of temporal expressions and events from clinical narratives","volume":"20","author":"Kova\u010devi\u0107","year":"2013","journal-title":"J Am Med Inform Assoc JAMIA"},{"key":"B33","doi-asserted-by":"publisher","first-page":"S53","DOI":"10.1016\/j.jbi.2015.06.029","article-title":"Combining knowledge-and data-driven methods for de-identification of clinical narratives","volume":"58","author":"Dehghan","year":"2015","journal-title":"J Biomed Inform"},{"key":"B34","doi-asserted-by":"publisher","first-page":"825","DOI":"10.5220\/0010414508250832","article-title":"The role of text analytics in healthcare: a review of recent developments, applications","volume":"5","author":"Elbattah","year":"2021","journal-title":"Healthinf"},{"key":"B35","doi-asserted-by":"publisher","first-page":"56","DOI":"10.1016\/j.jbi.2018.07.018","article-title":"Development of machine translation technology for assisting health communication: a systematic review","volume":"85","author":"Dew","year":"2018","journal-title":"J Biomed Inform"},{"key":"B36","first-page":"382","article-title":"Using machine translation in clinical practice","volume":"59","author":"Randhawa","year":"2013","journal-title":"Can Fam Phys"},{"key":"B37","author":"Soto","year":""},{"key":"B38","first-page":"279","article-title":"SNOMED-CT: the advanced terminology and coding system for eHealth","volume":"121","author":"Donnelly","year":"2006","journal-title":"Stud Health Technol Inform"},{"key":"B39","author":"Mujjiga","year":""},{"key":"B40","doi-asserted-by":"publisher","first-page":"D267","DOI":"10.1093\/nar\/gkh061","article-title":"The unified medical language system (UMLS): integrating biomedical terminology","volume":"32","author":"Bodenreider","year":"2004","journal-title":"Nucleic Acids Res"},{"key":"B41","author":"Finley","year":""},{"key":"B42","author":"Sennrich","year":""},{"key":"B43","author":"Bojar","year":""},{"key":"B44","author":"Yeganova","year":""},{"key":"B45","author":"Alyafeai","year":""},{"key":"B46","author":"Pomares-Quimbaya","year":""},{"key":"B47","author":"Peng","year":""},{"key":"B48","doi-asserted-by":"publisher","first-page":"107856","DOI":"10.1016\/j.compeleceng.2022.107856","article-title":"Transfer learning based on lexical constraint mechanism in low-resource machine translation","volume":"100","author":"Jiang","year":"2022","journal-title":"Comput Electr Eng"},{"key":"B49","author":"Junczys-Dowmunt","year":""},{"key":"B50","author":"Tiedemann","year":""},{"key":"B51","author":"Tiedemann","year":""},{"key":"B52","author":"Lepikhin","year":""},{"key":"B53","author":"Zhang","year":""},{"key":"B54","author":"Villegas","year":""},{"key":"B55","first-page":"1","article-title":"Beyond english-centric multilingual machine translation","volume":"22","author":"Fan","year":"2021","journal-title":"J Mach Learn Res"},{"key":"B56","author":"Papineni","year":""},{"key":"B57","author":"Lin","year":""},{"key":"B58","author":"Banerjee","year":""},{"key":"B59","author":"Post","year":""},{"key":"B60","author":"Rei","year":""},{"key":"B61","author":"Manchanda","year":""},{"key":"B62","author":"Wang","year":""},{"key":"B63","author":"Gladkoff","year":""},{"key":"B64","author":"Han","year":""}],"container-title":["Frontiers in Digital Health"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fdgth.2024.1211564\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,2,26]],"date-time":"2024-02-26T10:58:12Z","timestamp":1708945092000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fdgth.2024.1211564\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,2,26]]},"references-count":64,"alternative-id":["10.3389\/fdgth.2024.1211564"],"URL":"https:\/\/doi.org\/10.3389\/fdgth.2024.1211564","relation":{},"ISSN":["2673-253X"],"issn-type":[{"value":"2673-253X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,2,26]]},"article-number":"1211564"}}