{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,10]],"date-time":"2026-04-10T07:42:22Z","timestamp":1775806942194,"version":"3.50.1"},"reference-count":63,"publisher":"Springer Science and Business Media LLC","issue":"S1","license":[{"start":{"date-parts":[[2021,12,17]],"date-time":"2021-12-17T00:00:00Z","timestamp":1639699200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,12,17]],"date-time":"2021-12-17T00:00:00Z","timestamp":1639699200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100003329","name":"Ministerio de Econom\u00eda y Competitividad","doi-asserted-by":"publisher","award":["DeepEMR project TIN2017-87548-C2-1-R"],"award-info":[{"award-number":["DeepEMR project TIN2017-87548-C2-1-R"]}],"id":[{"id":"10.13039\/501100003329","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                <jats:title>Background<\/jats:title>\n                <jats:p>The volume of biomedical literature and clinical data is growing at an exponential rate. Therefore, efficient access to data described in unstructured biomedical texts is a crucial task for the biomedical industry and research. Named Entity Recognition (NER) is the first step for information and knowledge acquisition when we deal with unstructured texts. Recent NER approaches use contextualized word representations as input for a downstream classification task. However, distributed word vectors (embeddings) are very limited in Spanish and even more for the biomedical domain.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Methods<\/jats:title>\n                <jats:p>In this work, we develop several biomedical Spanish word representations, and we introduce two Deep Learning approaches for pharmaceutical, chemical, and other biomedical entities recognition in Spanish clinical case texts and biomedical texts, one based on a Bi-STM-CRF model and the other on a BERT-based architecture.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Results<\/jats:title>\n                <jats:p>Several Spanish biomedical embeddigns together with the two deep learning models were evaluated on the PharmaCoNER and CORD-19 datasets. The PharmaCoNER dataset is composed of a set of Spanish clinical cases annotated with drugs, chemical compounds and pharmacological substances; our extended Bi-LSTM-CRF model obtains an F-score of 85.24% on entity identification and classification and the BERT model obtains an F-score of 88.80% . For the entity normalization task, the extended Bi-LSTM-CRF model achieves an F-score of 72.85% and the BERT model achieves 79.97%. The CORD-19 dataset consists of scholarly articles written in English annotated with biomedical concepts such as disorder, species, chemical or drugs, gene and protein, enzyme and anatomy. Bi-LSTM-CRF model and BERT model obtain an F-measure of 78.23% and 78.86% on entity identification and classification, respectively on the CORD-19 dataset.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Conclusion<\/jats:title>\n                <jats:p>These results prove that deep learning models with in-domain knowledge learned from large-scale datasets highly improve named entity recognition performance. Moreover, contextualized representations help to understand complexities and ambiguity inherent to biomedical texts. Embeddings based on word, concepts, senses, etc. other than those for English are required to improve NER tasks in other languages.<\/jats:p>\n              <\/jats:sec>","DOI":"10.1186\/s12859-021-04247-9","type":"journal-article","created":{"date-parts":[[2021,12,17]],"date-time":"2021-12-17T16:04:31Z","timestamp":1639757071000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["Analyzing transfer learning impact in biomedical cross-lingual named entity recognition and normalization"],"prefix":"10.1186","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5082-5930","authenticated-orcid":false,"given":"Renzo M.","family":"Rivera-Zavala","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Paloma","family":"Mart\u00ednez","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,12,17]]},"reference":[{"issue":"suppl 1","key":"4247_CR1","doi-asserted-by":"publisher","first-page":"267","DOI":"10.1093\/nar\/gkh061","volume":"32","author":"O Bodenreider","year":"2004","unstructured":"Bodenreider O. The Unified Medical Language System (UMLS): integrating biomedical terminology. Nucleic Acids Res. 2004;32(suppl 1):267\u201370. https:\/\/doi.org\/10.1093\/nar\/gkh061.","journal-title":"Nucleic Acids Res"},{"key":"4247_CR2","doi-asserted-by":"publisher","first-page":"229","DOI":"10.1136\/jamia.2009.002733","volume":"17","author":"A Aronson","year":"2010","unstructured":"Aronson A, Lang F-M. An overview of metamap: historical perspective and recent advances. J Am Med Inform Assoc JAMIA. 2010;17:229\u201336. https:\/\/doi.org\/10.1136\/jamia.2009.002733.","journal-title":"J Am Med Inform Assoc JAMIA"},{"key":"4247_CR3","unstructured":"Lafferty JD, McCallum A, Pereira FCN. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In: Proceedings of the eighteenth international conference on machine learning. ICML \u201901. San Francisco: Morgan Kaufmann Publishers Inc.; 2001. p. 282\u20139. http:\/\/dl.acm.org\/citation.cfm?id=645530.655813."},{"key":"4247_CR4","unstructured":"Segura-Bedmar I, Martinez P, Sanchez-Cisneros D. The 1st ddiextraction-2011 challenge task: extraction of drug-drug interactions from biomedical texts, vol. 2011; 2011. p. 1\u20139."},{"key":"4247_CR5","doi-asserted-by":"publisher","first-page":"152","DOI":"10.1016\/j.jbi.2014.05.007","volume":"51","author":"I Segura-Bedmar","year":"2014","unstructured":"Segura-Bedmar I, Mart\u00ednez P, Herrero-Zazo M. Lessons learnt from the ddiextraction-2013 shared task. J Biomed Inform. 2014;51:152\u201364. https:\/\/doi.org\/10.1016\/j.jbi.2014.05.007.","journal-title":"J Biomed Inform"},{"key":"4247_CR6","doi-asserted-by":"crossref","unstructured":"Pilevar MT, Camacho-collados J. Embeddings in natural language processing: theory and advances in vector representation of meaning. Technical report; 2020.","DOI":"10.1007\/978-3-031-02177-0"},{"key":"4247_CR7","doi-asserted-by":"publisher","first-page":"103323","DOI":"10.1016\/j.jbi.2019.103323","volume":"101","author":"KS Kalyan","year":"2020","unstructured":"Kalyan KS, Sangeetha S. SECNLP: a survey of embeddings in clinical natural language processing. J Biomed Inform. 2020;101:103323. https:\/\/doi.org\/10.1016\/j.jbi.2019.103323.","journal-title":"J Biomed Inform"},{"key":"4247_CR8","doi-asserted-by":"crossref","unstructured":"Gonzalez-Agirre A, Marimon M, Intxaurrondo A, Rabal O, Villegas M, Krallinger M. Pharmaconer: Pharmacological substances, compounds and proteins named entity recognition track. In: Proceedings of the BioNLP Open Shared Tasks (BioNLP-OST). Hong Kong: Association for Computational Linguistics; 2019. p. 1.","DOI":"10.18653\/v1\/D19-5701"},{"key":"4247_CR9","doi-asserted-by":"crossref","unstructured":"Soares F, Villegas M, Gonzalez-Agirre A, Krallinger M, Armengol-Estap\u00e9 J. Medical word embeddings for Spanish: development and evaluation. In: Proceedings of the 2nd clinical natural language processing workshop. Minneapolis: Association for Computational Linguistics; 2019. p. 124\u201333. https:\/\/www.aclweb.org\/anthology\/W19-1916.","DOI":"10.18653\/v1\/W19-1916"},{"key":"4247_CR10","doi-asserted-by":"publisher","unstructured":"Xiong Y, Shen Y, Huang Y, Chen S, Tang B, Wang X, Chen Q, Yan J, Zhou Y. A deep learning-based system for PharmaCoNER. In: Proceedings of The 5th workshop on BioNLP Open Shared Tasks. Hong Kong: Association for Computational Linguistics; 2019. p. 33\u20137. https:\/\/doi.org\/10.18653\/v1\/D19-5706. https:\/\/www.aclweb.org\/anthology\/D19-5706.","DOI":"10.18653\/v1\/D19-5706"},{"key":"4247_CR11","doi-asserted-by":"publisher","unstructured":"Stoeckel M, Hemati W, Mehler A. When specialization helps: using pooled contextualized embeddings to detect chemical and biomedical entities in Spanish. In: Proceedings of The 5th Workshop on BioNLP Open Shared Tasks. Hong Kong: Association for Computational Linguistics; 2019. p. 11\u201315. https:\/\/doi.org\/10.18653\/v1\/D19-5702. https:\/\/www.aclweb.org\/anthology\/D19-5702.","DOI":"10.18653\/v1\/D19-5702"},{"key":"4247_CR12","unstructured":"Le\u00f3n FS, Ledesma AG. Annotating and normalizing biomedical NEs with limited knowledge. 2019;1912:09152."},{"issue":"3","key":"4247_CR13","doi-asserted-by":"publisher","first-page":"324","DOI":"10.1016\/j.cmpb.2011.01.002","volume":"101","author":"TS De Silva","year":"2011","unstructured":"De Silva TS, MacDonald D, Paterson G, Sikdar KC, Cochrane B. Systematized nomenclature of medicine clinical terms (SNOMED CT) to represent computed tomography procedures. Comput Methods Prog Biomed. 2011;101(3):324\u20139. https:\/\/doi.org\/10.1016\/j.cmpb.2011.01.002.","journal-title":"Comput Methods Prog Biomed"},{"issue":"1","key":"4247_CR14","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1186\/s13321-018-0327-2","volume":"11","author":"W Hemati","year":"2019","unstructured":"Hemati W, Mehler A. LSTMVOTER: chemical named entity recognition using a conglomerate of sequence labeling tools. J Cheminform. 2019;11(1):3. https:\/\/doi.org\/10.1186\/s13321-018-0327-2.","journal-title":"J Cheminform"},{"key":"4247_CR15","unstructured":"P\u00e9rez-P\u00e9rez M, Rabal O, P\u00e9rez-Rodr\u00edguez G, Vazquez M, Fdez-Riverola F, Oyarz\u00e1bal J, Valencia A, Louren\u00e7o A, Krallinger M. Evaluation of chemical and gene\/protein entity recognition systems at biocreative v.5: the cemp and gpro patents tracks. 2017."},{"key":"4247_CR16","doi-asserted-by":"publisher","first-page":"103285","DOI":"10.1016\/j.jbi.2019.103285","volume":"99","author":"V Su\u00e1rez-Paniagua","year":"2019","unstructured":"Su\u00e1rez-Paniagua V, Zavala RMR, Segura-Bedmar I, Mart\u00ednez P. A two-stage deep learning approach for extracting entities and relationships from medical texts. J Biomed Inform. 2019;99:103285. https:\/\/doi.org\/10.1016\/j.jbi.2019.103285.","journal-title":"J Biomed Inform."},{"issue":"1","key":"4247_CR17","doi-asserted-by":"publisher","first-page":"295","DOI":"10.1093\/bioinformatics\/btz528","volume":"36","author":"L Weber","year":"2019","unstructured":"Weber L, M\u00fcnchmeyer J, Rockt\u00e4schel T, Habibi M, Leser U. HUNER: improving biomedical NER with pretraining. Bioinformatics. 2019;36(1):295\u2013302. https:\/\/doi.org\/10.1093\/bioinformatics\/btz528.","journal-title":"Bioinformatics"},{"key":"4247_CR18","doi-asserted-by":"publisher","first-page":"161","DOI":"10.1186\/1471-2105-13-161","volume":"13","author":"M Bada","year":"2012","unstructured":"Bada M, Eckert M, Evans D, Garcia K, Shipley K, Sitnikov D, Baumgartner W Jr, Cohen K, Verspoor K, Blake J, Hunter L. Concept annotation in the craft corpus. BMC Bioinform. 2012;13:161. https:\/\/doi.org\/10.1186\/1471-2105-13-161.","journal-title":"BMC Bioinform"},{"issue":"2","key":"4247_CR19","doi-asserted-by":"publisher","first-page":"15","DOI":"10.5808\/GI.2019.17.2.e15","volume":"17","author":"J Armengol-Estap\u00e9","year":"2019","unstructured":"Armengol-Estap\u00e9 J, Soares F, Marimon M, Krallinger M. Pharmaconer tagger: a deep learning-based tool for automatically finding chemicals and drugs in Spanish medical texts. Genomics Inform. 2019;17(2):15. https:\/\/doi.org\/10.5808\/GI.2019.17.2.e15.","journal-title":"Genomics Inform"},{"key":"4247_CR20","doi-asserted-by":"publisher","unstructured":"Dernoncourt F, Lee JY, Szolovits P. NeuroNER: an easy-to-use program for named-entity recognition based on neural networks. In: Proceedings of the 2017 conference on empirical methods in natural language processing: system demonstrations. Copenhagen: Association for Computational Linguistics; 2017. p. 97\u2013102. https:\/\/doi.org\/10.18653\/v1\/D17-2017. https:\/\/www.aclweb.org\/anthology\/D17-2017.","DOI":"10.18653\/v1\/D17-2017"},{"key":"4247_CR21","unstructured":"Cardellino C. Spanish billion words corpus and embeddings. http:\/\/crscardellino.me\/SBWCE\/ (2016)."},{"key":"4247_CR22","unstructured":"Trask A, Michalak P, Liu J. sense2vec: a fast and accurate method for word sense disambiguation in neural word embeddings. CoRR abs\/1511.06388. arXiv:1511.06388 (2015)"},{"key":"4247_CR23","unstructured":"PharmaCoNER Evaluation. https:\/\/temu.bsc.es\/pharmaconer\/index.php\/evaluation\/. Accessed 12 April 2021."},{"key":"4247_CR24","unstructured":"Lu Wang L, Lo K, Chandrasekhar Y, Reas R, Yang J, Eide D, Funk K, Kinney R, Liu Z, Merrill W, Mooney P, Murdick D, Rishi D, Sheehan J, Shen Z, Stilson B, Wade AD, Wang K, Wilhelm C, Xie B, Raymond D, Weld DS, Etzioni O, Kohlmeier S. CORD-19: the covid-19 open research dataset. arXiv. arXiv:2004.10706 (2020)."},{"key":"4247_CR25","doi-asserted-by":"publisher","unstructured":"Gonzalez-Agirre A, Marimon M, Intxaurrondo A, Rabal O, Villegas M, Krallinger M. PharmaCoNER: pharmacological substances, compounds and proteins named entity recognition track. In: Proceedings of The 5th workshop on BioNLP open shared tasks. Hong Kong: Association for Computational Linguistics; 2019. p. 1\u201310. https:\/\/doi.org\/10.18653\/v1\/D19-5701. https:\/\/www.aclweb.org\/anthology\/D19-5701.","DOI":"10.18653\/v1\/D19-5701"},{"key":"4247_CR26","doi-asserted-by":"publisher","unstructured":"Sun C, Yang Z. Transfer learning in biomedical named entity recognition: an evaluation of BERT in the PharmaCoNER task. In: Proceedings of The 5th workshop on BioNLP open shared tasks. Hong Kong: Association for Computational Linguistics; 2019. p. 100\u20134. https:\/\/doi.org\/10.18653\/v1\/D19-5715. https:\/\/www.aclweb.org\/anthology\/D19-5715.","DOI":"10.18653\/v1\/D19-5715"},{"key":"4247_CR27","unstructured":"ZENODO AbreMES-DB. https:\/\/zenodo.org\/record\/2207130#.XvxA7ChKg2x. Accessed 12 April 2021."},{"key":"4247_CR28","unstructured":"Diccionario de Siglas Medicas. http:\/\/www.sedom.es\/diccionario\/. Accessed 12 April 2021."},{"key":"4247_CR29","unstructured":"Miller FP, Vandome AF, McBrewster J. Levenshtein distance: information theory, computer science, string (computer science), string metric, damerau? Levenshtein distance, spell checker, hamming distance. Alpha Press; 2009."},{"key":"4247_CR30","doi-asserted-by":"crossref","unstructured":"Beam AL, Kompa B, Schmaltz A, Fried I, Weber G, Palmer NP, Shi X, Cai T, Kohane IS. Clinical concept embeddings learned from massive sources of multimodal medical data. 2018;1804:01486.","DOI":"10.1142\/9789811215636_0027"},{"key":"4247_CR31","doi-asserted-by":"crossref","unstructured":"Rasmy L, Xiang Y, Xie Z, Tao C, Zhi D. Med-BERT: pre-trained contextualized embeddings on large-scale structured electronic health records for disease prediction. arXiv:2005.12833 (2020).","DOI":"10.1038\/s41746-021-00455-y"},{"key":"4247_CR32","unstructured":"Explosion AI: spaCy - Industrial-strength Natural Language Processing in Python. https:\/\/spacy.io\/ Accessed 12 April 2021."},{"key":"4247_CR33","unstructured":"Stenetorp P, Pyysalo S, Topi\u0107 G, Ohta T, Ananiadou S, Tsujii JI. BRAT: a web-based tool for NLP-assisted text annotation. Technical report. https:\/\/dl.acm.org\/citation.cfm?id=2380942 (2012)."},{"key":"4247_CR34","unstructured":"Farkas R, Vincze V, M\u00f3ra G, Csirik J, Szarvas G. The CoNLL-2010 shared task: learning to detect hedges and their scope in natural language text. Technical Report July. http:\/\/www.aclweb.org\/anthology\/W10-3001 (2010)."},{"key":"4247_CR35","unstructured":"Borthwick A, Sterling J, Agichtein E, Grishman R. Exploiting diverse knowledge sources via maximum entropy in named entity recognition. In: Sixth workshop on very large corpora. https:\/\/www.aclweb.org\/anthology\/W98-1118 (1998)."},{"issue":"12","key":"4247_CR36","doi-asserted-by":"publisher","first-page":"18953","DOI":"10.2196\/18953","volume":"8","author":"R Rivera Zavala","year":"2020","unstructured":"Rivera Zavala R, Martinez P. The impact of pretrained language models on negation and speculation detection in cross-lingual medical text: comparative study. JMIR Med Inform. 2020;8(12):18953. https:\/\/doi.org\/10.2196\/18953.","journal-title":"JMIR Med Inform"},{"key":"4247_CR37","doi-asserted-by":"publisher","unstructured":"Sang EFTK, De\u00a0Meulder F. Introduction to the conll-2003 shared task: Language-independent named entity recognition. In: Proceedings of the seventh conference on natural language learning at HLT-NAACL 2003. CONLL \u201903, vol. 4. Association for Computational Linguistics; 2003. p. 142\u20137 https:\/\/doi.org\/10.3115\/1119176.1119195.","DOI":"10.3115\/1119176.1119195"},{"key":"4247_CR38","unstructured":"The Spanish bibliographical index in health science. http:\/\/ibecs.isciii.es. Accessed 12 April 2021."},{"key":"4247_CR39","unstructured":"Scientific electronic library online. https:\/\/scielo.org\/es\/. Accessed 12 April 2021."},{"key":"4247_CR40","unstructured":"National library of medicine. https:\/\/www.ncbi.nlm.nih.gov\/pubmed. Accessed 12 April 2021."},{"key":"4247_CR41","unstructured":"MedlinePlus. https:\/\/medlineplus.gov\/. Accessed 12 April 2021."},{"key":"4247_CR42","unstructured":"UFAL medical corpus. https:\/\/ufal.mff.cuni.cz\/ufal_medical_corpus. Accessed 12 April 2021."},{"key":"4247_CR43","unstructured":"Industrial-l. https:\/\/spacy.io\/. Accessed 12 April 2021."},{"issue":"23","key":"4247_CR44","doi-asserted-by":"publisher","first-page":"4087","DOI":"10.1093\/bioinformatics\/bty449","volume":"34","author":"JM Giorgi","year":"2018","unstructured":"Giorgi JM, Bader GD. Transfer learning for biomedical named entity recognition with neural networks. Bioinformatics (Oxford, England). 2018;34(23):4087\u201394. https:\/\/doi.org\/10.1093\/bioinformatics\/bty449.","journal-title":"Bioinformatics (Oxford, England)"},{"key":"4247_CR45","doi-asserted-by":"crossref","unstructured":"Wang D, Zheng TF. Transfer learning for speech and language processing. CoRR abs\/1511.06066. arXiv:1511.06066 (2015).","DOI":"10.1109\/APSIPA.2015.7415532"},{"key":"4247_CR46","doi-asserted-by":"crossref","unstructured":"Mou L, Meng Z, Yan R, Li G, Xu Y, Zhang L, Jin Z. How transferable are neural networks in NLP applications? CoRR abs\/1603.06111. arXiv:1603.06111 (2016).","DOI":"10.18653\/v1\/D16-1046"},{"key":"4247_CR47","unstructured":"Lee JY, Dernoncourt F, Szolovits P. Transfer learning for named-entity recognition with neural networks. In: 11th International conference on language resources and evaluation, LREC 2018. p. 4470\u20133. arXiv:1705.06273 (2019)."},{"key":"4247_CR48","doi-asserted-by":"publisher","unstructured":"Ling W, Dyer C, Black AW, Trancoso I, Fermandez R, Amir S, Marujo L, Luis T. Finding function in form: compositional character models for open vocabulary word representation. In: Proceedings of the 2015 conference on empirical methods in natural language processing. Lisbon: Association for Computational Linguistics; 2015. p. 1520\u201330. https:\/\/doi.org\/10.18653\/v1\/D15-1176.","DOI":"10.18653\/v1\/D15-1176"},{"key":"4247_CR49","unstructured":"Mart\u00ed MA, Taul\u00e9 M, Bertran M, M\u00e1rquez L. AnCora: multilingual and multilevel annotated corpora. http:\/\/clic.ub.edu\/ancora\/ancora-corpus.pdf (2007)."},{"key":"4247_CR50","unstructured":"Mikolov T, Sutskever I, Chen K, Corrado G, Dean J. Distributed representations of words and phrases and their compositionality. CoRR abs\/1310.4546. arXiv:1310.4546 (2013)."},{"issue":"2","key":"4247_CR51","doi-asserted-by":"publisher","first-page":"15","DOI":"10.5808\/GI.2019.17.2.e15","volume":"17","author":"J Armengol-Estap\u00e9","year":"2019","unstructured":"Armengol-Estap\u00e9 J, Soares F, Marimon M, Krallinger M. PharmacoNER Tagge: a deep learning-based tool for automatically finding chemicals and drugs in Spanish medical texts. Genomics Inform. 2019;17(2):15. https:\/\/doi.org\/10.5808\/GI.2019.17.2.e15.","journal-title":"Genomics Inform"},{"key":"4247_CR52","unstructured":"Mikolov T, Grave E, Bojanowski P, Puhrsch C, Joulin A. Advances in pre-training distributed word representations. In: Proceedings of the international conference on language resources and evaluation (LREC 2018); 2018."},{"key":"4247_CR53","unstructured":"Pyysalo S, Ginter F, Moen H, Salakoski T, Ananiadou S. Distributional semantics resources for biomedical text processing. Technical report. https:\/\/github.com\/spyysalo\/nxml2txt (2013)."},{"key":"4247_CR54","doi-asserted-by":"crossref","unstructured":"Bojanowski P, Grave E, Joulin A, Mikolov T. Enriching word vectors with subword information. CoRR abs\/1607.04606. arXiv:1607.04606 (2016).","DOI":"10.1162\/tacl_a_00051"},{"key":"4247_CR55","doi-asserted-by":"publisher","first-page":"924","DOI":"10.3233\/978-1-61499-512-8-924","volume":"210","author":"JB Lamy","year":"2015","unstructured":"Lamy JB, Venot A, Duclos C. PyMedTermino: an open-source generic API for advanced terminology services. Stud Health Technol Inform. 2015;210:924\u20138. https:\/\/doi.org\/10.3233\/978-1-61499-512-8-924.","journal-title":"Stud Health Technol Inform"},{"key":"4247_CR56","unstructured":"Zavala RMR. GitHub-rmriveraz\/PharmaCoNER: Biomedical Spanish Word and Concept embeddings-pretrained models. https:\/\/github.com\/rmriveraz\/PharmaCoNER. Accessed 12 April 2021."},{"key":"4247_CR57","doi-asserted-by":"publisher","unstructured":"Pennington J, Socher R, Manning C. Glove: global vectors for word representation. In: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP); 2014. p. 1532\u201343. https:\/\/doi.org\/10.3115\/v1\/D14-1162. arXiv:1504.06654.","DOI":"10.3115\/v1\/D14-1162"},{"key":"4247_CR58","doi-asserted-by":"publisher","unstructured":"Peters ME, Neumann M, Iyyer M, Gardner M, Clark C, Lee K, Zettlemoyer L. Deep contextualized word representations. In: Proceedings of the Conference of the North American chapter of the association for computational linguistics: human language technologies, NAACL HLT, vol. 1. Association for Computational Linguistics (ACL); 2018. p. 2227\u201337. https:\/\/doi.org\/10.18653\/v1\/n18-1202. arXiv:1802.05365.","DOI":"10.18653\/v1\/n18-1202"},{"key":"4247_CR59","unstructured":"McCann B, Bradbury J, Xiong C, Socher R. Learned in translation: contextualized word vectors. Technical report. arXiv:1708.00107 (2017)."},{"key":"4247_CR60","unstructured":"Devlin J, Chang MW, Lee K, Toutanova K. BERT: Pre-training of deep bidirectional transformers for language understanding. Technical report. arXiv:1810.04805 (2019)."},{"key":"4247_CR61","unstructured":"Ca\u00f1ete J, Chaperon G, Fuentes R, P\u00e9rez J. Spanish pre-trained bert model and evaluation data. In: PML4DC at ICLR 2020; 2020 (to appear)."},{"key":"4247_CR62","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/btz682","author":"J Lee","year":"2019","unstructured":"Lee J, Yoon W, Kim S, Kim D, Kim S, So CH, Kang J. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics. 2019. https:\/\/doi.org\/10.1093\/bioinformatics\/btz682.","journal-title":"Bioinformatics"},{"key":"4247_CR63","doi-asserted-by":"crossref","unstructured":"Kudo T, Richardson J. SentencePiece: a simple and language independent subword tokenizer and detokenizer for Neural Text Processing. arXiv:1808.06226 (2018).","DOI":"10.18653\/v1\/D18-2012"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-021-04247-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12859-021-04247-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-021-04247-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,12,8]],"date-time":"2023-12-08T18:03:35Z","timestamp":1702058615000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-021-04247-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,12,17]]},"references-count":63,"journal-issue":{"issue":"S1","published-online":{"date-parts":[[2021,12]]}},"alternative-id":["4247"],"URL":"https:\/\/doi.org\/10.1186\/s12859-021-04247-9","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,12,17]]},"assertion":[{"value":"27 May 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 June 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 December 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to publish"}},{"value":"The authors declare that they have no competing interests.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"601"}}