{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T12:26:53Z","timestamp":1784723213755,"version":"3.55.0"},"reference-count":39,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2024,6,3]],"date-time":"2024-06-03T00:00:00Z","timestamp":1717372800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,6,3]],"date-time":"2024-06-03T00:00:00Z","timestamp":1717372800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Med Inform Decis Mak"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Background<\/jats:title>\n                    <jats:p>BERT models have seen widespread use on unstructured text within the clinical domain. However, little to no research has been conducted into classifying unstructured clinical notes on the basis of patient lifestyle indicators, especially in Dutch. This article aims to test the feasibility of deep BERT models on the task of patient lifestyle classification, as well as introducing an experimental framework that is easily reproducible in future research.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Methods<\/jats:title>\n                    <jats:p>This study makes use of unstructured general patient text data from HagaZiekenhuis, a large hospital in The Netherlands. Over 148 000 notes were provided to us, which were each automatically labelled on the basis of the respective patients\u2019 smoking, alcohol usage and drug usage statuses. In this paper we test feasibility of automatically assigning labels, and justify it using hand-labelled input. Ultimately, we compare macro F1-scores of string matching, SGD and several BERT models on the task of classifying smoking, alcohol and drug usage. We test Dutch BERT models and English models with translated input.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>We find that our further pre-trained MedRoBERTa.nl-HAGA model outperformed every other model on smoking (0.93) and drug usage (0.77). Interestingly, our ClinicalBERT model that was merely fine-tuned on translated text performed best on the alcohol task (0.80). In t-SNE visualisations, we show our MedRoBERTa.nl-HAGA model is the best model to differentiate between classes in the embedding space, explaining its superior classification performance.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Conclusions<\/jats:title>\n                    <jats:p>We suggest MedRoBERTa.nl-HAGA to be used as a baseline in future research on Dutch free text patient lifestyle classification. We furthermore strongly suggest further exploring the application of translation to input text in non-English clinical BERT research, as we only translated a subset of the full set and yet achieved very promising results.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1186\/s12911-024-02557-5","type":"journal-article","created":{"date-parts":[[2024,6,3]],"date-time":"2024-06-03T11:02:42Z","timestamp":1717412562000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":19,"title":["Extracting patient lifestyle characteristics from Dutch clinical text with BERT models"],"prefix":"10.1186","volume":"24","author":[{"given":"Hielke","family":"Muizelaar","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Marcel","family":"Haas","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Koert","family":"van Dortmont","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Peter","family":"van der Putten","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Marco","family":"Spruit","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,6,3]]},"reference":[{"key":"2557_CR1","doi-asserted-by":"publisher","first-page":"102086","DOI":"10.1016\/j.artmed.2021.102086","volume":"118","author":"A Kormilitzin","year":"2021","unstructured":"Kormilitzin A, Vaci N, Liu Q, Nevado-Holgado A. Med7: a transferable clinical natural language processing model for electronic health records. Artif Intell Med. 2021;118:102086.","journal-title":"Artif Intell Med."},{"issue":"5","key":"2557_CR2","doi-asserted-by":"publisher","first-page":"1059","DOI":"10.1093\/rheumatology\/kez375","volume":"59","author":"SS Zhao","year":"2020","unstructured":"Zhao SS, Hong C, Cai T, Xu C, Huang J, Ermann J, et al. Incorporating natural language processing to improve classification of axial spondyloarthritis using electronic health records. Rheumatology. 2020;59(5):1059\u201365.","journal-title":"Rheumatology."},{"key":"2557_CR3","doi-asserted-by":"publisher","unstructured":"Zheng C, Lee M, Bansal N, Go AS, Chen C, Harrison TN, et al. Identification of recurrent atrial fibrillation using natural language processing applied to electronic health records. Eur Heart J Qual Care Clin Outcomes. https:\/\/doi.org\/10.1093\/ehjqcco\/qcad021.","DOI":"10.1093\/ehjqcco\/qcad021"},{"key":"2557_CR4","doi-asserted-by":"publisher","first-page":"86","DOI":"10.1038\/s41746-021-00455-y","volume":"4","author":"L Rasmy","year":"2021","unstructured":"Rasmy L, Xiang Y, Xie Z, Tao C, Zhi D. Med-BERT: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction. NPJ Digit Med. 2021;4:86.","journal-title":"NPJ Digit Med."},{"key":"2557_CR5","unstructured":"Devlin J, Chang MW, Lee K, Toutanova K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Association for Computational Linguistics. 2019:4171\u201386."},{"key":"2557_CR6","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention is all you need. Adv Neural Inf Process Syst. 2017;30:5998\u20136008."},{"key":"2557_CR7","unstructured":"Liu Y, Ott M, Goyal N, Du J, Joshi M, Chen D, et al. RoBERTa: a robustly optimized BERT pretraining approach. 2019. Preprint at arXiv:1907.11692."},{"issue":"8","key":"2557_CR8","doi-asserted-by":"publisher","first-page":"e0270595","DOI":"10.1371\/journal.pone.0270595","volume":"17","author":"S Chaichulee","year":"2022","unstructured":"Chaichulee S, Promchai C, Kaewkomon T, Kongkamol C, Ingviya T, Sansupawanich P. Multi-label classification of symptom terms from free-text bilingual adverse drug reaction reports using natural language processing. PLoS ONE. 2022;17(8):e0270595.","journal-title":"PLoS ONE."},{"key":"2557_CR9","first-page":"3255","volume":"2020","author":"P Delobelle","year":"2020","unstructured":"Delobelle P, Winters T, Berendt B. RobBERT: a Dutch RoBERTa-based Language Model. Findings of the Association for Computational Linguistics: EMNLP. 2020;2020:3255\u201365.","journal-title":"Findings of the Association for Computational Linguistics: EMNLP"},{"key":"2557_CR10","first-page":"125","volume":"11","author":"P Delobelle","year":"2021","unstructured":"Delobelle P, Winters T, Berendt B. RobBERTje: A Distilled Dutch BERT Model. Comput Linguist Neth. 2021;11:125\u201340.","journal-title":"Comput Linguist Neth."},{"issue":"1","key":"2557_CR11","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3458754","volume":"3","author":"Y Gu","year":"2021","unstructured":"Gu Y, Tinn R, Cheng H, Lucas M, Usuyama N, Liu X, et al. Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing. ACM Trans Comput Healthc. 2021;3(1):1\u201323.","journal-title":"ACM Trans Comput Healthc."},{"key":"2557_CR12","first-page":"231","volume":"11","author":"L De Bruyne","year":"2021","unstructured":"De Bruyne L, De Clercq O, Hoste V. Prospects for Dutch Emotion Detection: Insights from the New EmotioNL data set. Comput Linguist Neth. 2021;11:231\u201355.","journal-title":"Comput Linguist Neth."},{"issue":"3","key":"2557_CR13","doi-asserted-by":"publisher","first-page":"37","DOI":"10.1167\/tvst.11.3.37","volume":"11","author":"W Hu","year":"2022","unstructured":"Hu W, Wang SY. Predicting Glaucoma Progression Requiring Surgery Using Clinical Free-Text Notes and Transfer Learning With Transformers. Transl Vis Sci Technol. 2022;11(3):37.","journal-title":"Transl Vis Sci Technol."},{"key":"2557_CR14","doi-asserted-by":"crossref","unstructured":"Lan Z, Chen M, Goodman S, Gimpel K, Sharma P, Soricut R. ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. 2020. Preprint at arXiv:1909.11942.","DOI":"10.1109\/SLT48900.2021.9383575"},{"key":"2557_CR15","doi-asserted-by":"crossref","unstructured":"Lewis M, Liu Y, Goyal N, Ghazvininejad M, Abdelrahman M, Levy O, et al. BART: Denoising Sequence-to-Sequence Pretraining for Natural Language Generation, Translation, and Comprehension. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics. 2020. p. 7871\u201380.","DOI":"10.18653\/v1\/2020.acl-main.703"},{"issue":"4","key":"2557_CR16","doi-asserted-by":"publisher","first-page":"146045822211311","DOI":"10.1177\/14604582221131198","volume":"28","author":"AW Olthof","year":"2022","unstructured":"Olthof AW, Van Ooijen PMA, Cornelissen LJ. The natural language processing of radiology requests and reports of chest imaging: Comparing five transformer models\u2019 multilabel classification and a proof-of-concept study. Health Inform J. 2022;28(4):14604582221131198.","journal-title":"Health Inform J."},{"key":"2557_CR17","unstructured":"Verkijk S, Vossen P. MedRoBERTa.nl: A Language Model for Dutch Electronic Health Records. Comput Linguist Neth. 2021;11:141\u201359."},{"key":"2557_CR18","unstructured":"Huang K, Altosaar J, Ranganath R. ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission. 2020. Preprint at arXiv:1904.05342."},{"issue":"4","key":"2557_CR19","doi-asserted-by":"publisher","first-page":"1234","DOI":"10.1093\/bioinformatics\/btz682","volume":"36","author":"J Lee","year":"2020","unstructured":"Lee J, Yoon W, Kim S, Kim D, Kim S, So CH, et al. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics. 2020;36(4):1234\u201340.","journal-title":"Bioinformatics."},{"key":"2557_CR20","doi-asserted-by":"publisher","first-page":"105","DOI":"10.1007\/s10278-022-00712-w","volume":"36","author":"I Banerjee","year":"2023","unstructured":"Banerjee I, Davis MA, Vey BL, Mazaheri S, Khan F, Zavaletta V, et al. Natural Language Processing Model for Identifying Critical Findings-A Multi-Institutional Study. J Digit Imaging. 2023;36:105\u201313.","journal-title":"J Digit Imaging."},{"key":"2557_CR21","doi-asserted-by":"crossref","unstructured":"Michalopoulos G, Wang Y, Kaka H, Chen H, Wong A. UmlsBERT: Clinical Domain Knowledge Augmentation of Contextual Embeddings Using the Unified Medical Language System Metathesaurus. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics. 2021. p. 1744\u201353.","DOI":"10.18653\/v1\/2021.naacl-main.139"},{"issue":"5","key":"2557_CR22","doi-asserted-by":"publisher","first-page":"873","DOI":"10.1093\/jamia\/ocac018","volume":"29","author":"K Xie","year":"2022","unstructured":"Xie K, Gallagher RS, Conrad EC, Garrick CO, Baldassano SN, Bernabei JM, et al. Extracting seizure frequency from epilepsy clinic notes: a machine reading approach to natural language processing. J Am Med Inform Assoc. 2022;29(5):873\u201381.","journal-title":"J Am Med Inform Assoc."},{"key":"2557_CR23","unstructured":"Saadullah A, Neumann G, Dunfield KA, Vechkaeva A, Chapman KA, Wixted MK. MLT-DFKI at CLEF eHealth 2019: Multi-label Classification of ICD-10 Codes with BERT. CEUR Workshop Proceedings. Conference and Labs of the Evaluation Forum Initiative. 2019;2380:67."},{"issue":"21","key":"2557_CR24","doi-asserted-by":"publisher","first-page":"11250","DOI":"10.3390\/app122111250","volume":"12","author":"S Park","year":"2022","unstructured":"Park S, Bong JW, Park I, Lee H, Choi J, Park P, et al. ConBERT: A Concatenation of Bidirectional Transformers for Standardization of Operative Reports from Electronic Medical Records. Appl Sci. 2022;12(21):11250.","journal-title":"Appl Sci."},{"key":"2557_CR25","unstructured":"Wouts J, De Boer J, Voppel A, Brederoo S, Van Splunter S, Sommer I. belabBERT: a Dutch RoBERTa-based language model applied to psychiatric classification. 2021. Preprint at arXiv:2106.01091."},{"key":"2557_CR26","unstructured":"Heath C. Natural Language Processing for lifestyle recognition in discharge summaries. 2022. https:\/\/theses.liacs.nl\/2279. Accessed 12 Nov 2023."},{"key":"2557_CR27","unstructured":"Reuver M. FINDING THE SMOKE SIGNAL: Smoking Status Classification with a Weakly Supervised Paradigm in Sparsely Labelled Dutch Free Text in Electronic Medical Records. 2020. https:\/\/theses.ubn.ru.nl\/handle\/123456789\/10278. Accessed 13 Nov 2023."},{"issue":"3","key":"2557_CR28","doi-asserted-by":"publisher","first-page":"437","DOI":"10.1093\/ehjdh\/ztac031","volume":"3","author":"AR De Boer","year":"2022","unstructured":"De Boer AR, De Groot MCH, Groenhof TKJ, Van Doorn S, Vaartjes I, Bots ML, et al. Data mining to retrieve smoking status from electronic health records in general practice. Eur Heart J Digit Health. 2022;3(3):437\u201344.","journal-title":"Eur Heart J Digit Health."},{"key":"2557_CR29","doi-asserted-by":"publisher","first-page":"100","DOI":"10.1016\/j.jclinepi.2019.11.006","volume":"118","author":"TKJ Groenhof","year":"2020","unstructured":"Groenhof TKJ, Koers LR, Blasse E, De Groot M, Grobbee DE, Bots ML, et al. Data mining information from electronic health records produced high yield and accuracy for current smoking status. J Clin Epidemiol. 2020;118:100\u20136.","journal-title":"J Clin Epidemiol."},{"issue":"3","key":"2557_CR30","doi-asserted-by":"publisher","first-page":"400","DOI":"10.1214\/aoms\/1177729586","volume":"22","author":"H Robbins","year":"1951","unstructured":"Robbins H, Monro S. A Stochastic Approximation Method. Ann Math Stat. 1951;22(3):400\u20137.","journal-title":"Ann Math Stat."},{"key":"2557_CR31","doi-asserted-by":"publisher","first-page":"986","DOI":"10.1007\/978-0-387-30164-8_832","volume-title":"Encyclopedia of Machine Learning","author":"C Sammut","year":"2011","unstructured":"Sammut C, Webb GI. TF-IDF. In: Sammut C, Webb GI, editors. Encyclopedia of Machine Learning. Boston: Springer; 2011. p. 986\u20137."},{"key":"2557_CR32","unstructured":"De Wynter A, Perry DJ. Optimal Subarchitecture Extraction For BERT. 2020. Preprint at arXiv:2010.10499."},{"key":"2557_CR33","unstructured":"Tiedemann J, Thottingal S. OPUS-MT - Building open translation services for the World. Proceedings of the 22nd Annual Conference of the European Association for Machine Translation. European Association for Machine Translation. 2020. p. 479\u201380."},{"key":"2557_CR34","doi-asserted-by":"crossref","unstructured":"Popovi\u0107 M. chrF: character n-gram F-score for automatic MT evaluation. Proceedings of the Tenth Workshop on Statistical Machine Translation. Association for Computational Linguistics. 2015. p. 392\u201395.","DOI":"10.18653\/v1\/W15-3049"},{"key":"2557_CR35","doi-asserted-by":"crossref","unstructured":"Lipton ZC, Elkan C, Naryanaswamy B. Optimal Thresholding of Classifiers to Maximize F1 Measure. In: Calders T, Esposito F, H\u00fcllermeier E, Meo R. Machine Learning and Knowledge Discovery in Databases. Berlin, Heidelberg: Springer; 2014. pp. 225-239.","DOI":"10.1007\/978-3-662-44851-9_15"},{"issue":"86","key":"2557_CR36","first-page":"2579","volume":"9","author":"L Van der Maaten","year":"2008","unstructured":"Van der Maaten L, Hinton G. Visualizing Data using t-SNE. J Mach Learn Res. 2008;9(86):2579\u2013605.","journal-title":"J Mach Learn Res."},{"key":"2557_CR37","doi-asserted-by":"crossref","unstructured":"Wolf T, Debut L, Sanh V, Chaumond J, Delangue C, Moi A, et al. HuggingFace Transformers: State-of-the-Art Natural Language Processing. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Association for Computational Linguistics. 2020. p. 38\u201345.","DOI":"10.18653\/v1\/2020.emnlp-demos.6"},{"key":"2557_CR38","unstructured":"Boeckhout M, Beusink M, Bouter L, Kist I, Rebers S, Van Veen EB, et al. Niet-WMO-plichtig onderzoek en ethische toetsing. Commissioned by the Dutch Ministry of Health, Welfare and Sport. 2020. https:\/\/www.rijksoverheid.nl\/documenten\/rapporten\/2020\/02\/14\/niet-wmo-plichtig-onderzoek-en-ethische-toetsing. Accessed 16 Jan 2024."},{"key":"2557_CR39","unstructured":"Regulation (EU) 2016\/679 of the European Parliament and of the Council of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing Directive 95\/46\/EC (General Data Protection Regulation) (Text with EEA relevance). Off J. 2016;L 119:1\u201388. http:\/\/data.europa.eu\/eli\/reg\/2016\/679\/oj. Accessed 16 Jan 2024."}],"updated-by":[{"DOI":"10.1186\/s12911-024-02575-3","type":"correction","label":"Correction","source":"publisher","updated":{"date-parts":[[2024,6,17]],"date-time":"2024-06-17T00:00:00Z","timestamp":1718582400000}}],"container-title":["BMC Medical Informatics and Decision Making"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12911-024-02557-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12911-024-02557-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12911-024-02557-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,23]],"date-time":"2024-08-23T17:17:40Z","timestamp":1724433460000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcmedinformdecismak.biomedcentral.com\/articles\/10.1186\/s12911-024-02557-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,3]]},"references-count":39,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,12]]}},"alternative-id":["2557"],"URL":"https:\/\/doi.org\/10.1186\/s12911-024-02557-5","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-3831694\/v1","asserted-by":"object"}]},"ISSN":["1472-6947"],"issn-type":[{"value":"1472-6947","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,3]]},"assertion":[{"value":"3 January 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 May 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 June 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 June 2024","order":4,"name":"change_date","label":"Change Date","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Correction","order":5,"name":"change_type","label":"Change Type","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"A Correction to this paper has been published:","order":6,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"https:\/\/doi.org\/10.1186\/s12911-024-02575-3","URL":"https:\/\/doi.org\/10.1186\/s12911-024-02575-3","order":7,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The Science Bureau of the Haga Teaching Hospital on 16 January 2024 confirmed with reference number 20240116 that the manuscript \u201cExtracting Patient Lifestyle Characteristics from Dutch Clinical Text with BERT Models\u201d does not fall under the scope of the Dutch law on Medical Research Involving Human Subjects (Dutch abbreviation: WMO). Therefore, this manuscript is confirmed to be exempt from review by the Medical Research Ethics Committee (MREC). This study was conducted according to the scientific policy of the Haga Teaching hospital.\u00a0Exemption from the Dutch law on Medical Research Involving Human Subjects\u00a0indicates that patient consent is waived for the respective manuscript, as patients\u00a0are not actively involved [\n                      \n                      ]. The clinical notes analysed in this study were\u00a0extracted from the CTCue platform, on which HagaZiekenhuis\u2019 patient data is\u00a0hosted. Patient data were pseudominised and anonymised such that users comply\u00a0with privacy rules and legislation set out in the European Union\u2019s General Data\u00a0Protection Regulation (GDPR) [\n                      \n                      ]. See\n                      \n                      (webpage in Dutch) on more\u00a0information on the de-identification process. No further patient data were included\u00a0in this research.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"151"}}