{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T14:52:51Z","timestamp":1778856771536,"version":"3.51.4"},"reference-count":33,"publisher":"Springer Science and Business Media LLC","issue":"S1","license":[{"start":{"date-parts":[[2021,12,17]],"date-time":"2021-12-17T00:00:00Z","timestamp":1639699200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,12,17]],"date-time":"2021-12-17T00:00:00Z","timestamp":1639699200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                <jats:title>Background<\/jats:title>\n                <jats:p>The recognition of pharmacological substances, compounds and proteins is essential for biomedical relation extraction, knowledge graph construction, drug discovery, as well as medical question answering. Although considerable efforts have been made to recognize biomedical entities in English texts, to date, only few limited attempts were made to recognize them from biomedical texts in other languages. PharmaCoNER is a named entity recognition challenge to recognize pharmacological entities from Spanish texts. Because there are currently abundant resources in the field of natural language processing, how to leverage these resources to the PharmaCoNER challenge is a meaningful study.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Methods<\/jats:title>\n                <jats:p>Inspired by the success of deep learning with language models, we compare and explore various representative BERT models to promote the development of the PharmaCoNER task.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Results<\/jats:title>\n                <jats:p>The experimental results show that deep learning with language models can effectively improve model performance on the PharmaCoNER dataset. Our method achieves state-of-the-art performance on the PharmaCoNER dataset, with a max F1-score of 92.01%.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Conclusion<\/jats:title>\n                <jats:p>For the BERT models on the PharmaCoNER dataset, biomedical domain knowledge has a greater impact on model performance than the native language (i.e., Spanish). The BERT models can obtain competitive performance by using WordPiece to alleviate the out of vocabulary limitation. The performance on the BERT model can be further improved by constructing a specific vocabulary based on domain knowledge. Moreover, the character case also has a certain impact on model performance.<\/jats:p>\n              <\/jats:sec>","DOI":"10.1186\/s12859-021-04260-y","type":"journal-article","created":{"date-parts":[[2021,12,17]],"date-time":"2021-12-17T16:04:31Z","timestamp":1639757071000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Deep learning with language models improves named entity recognition for PharmaCoNER"],"prefix":"10.1186","volume":"22","author":[{"given":"Cong","family":"Sun","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhihao","family":"Yang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lei","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yin","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hongfei","family":"Lin","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jian","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,12,17]]},"reference":[{"issue":"12","key":"4260_CR1","doi-asserted-by":"publisher","first-page":"7673","DOI":"10.1021\/acs.chemrev.6b00851","volume":"117","author":"M Krallinger","year":"2017","unstructured":"Krallinger M, Rabal O, Lourenco A, et al. Information retrieval and text mining technologies for chemistry. Chem Rev. 2017;117(12):7673\u2013761.","journal-title":"Chem Rev"},{"issue":"1","key":"4260_CR2","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/1758-2946-7-S1-S1","volume":"7","author":"M Krallinger","year":"2015","unstructured":"Krallinger M, Leitner F, Rabal O, et al. CHEMDNER: The drugs and chemical names extraction challenge. J Cheminform. 2015;7(1):1\u201311.","journal-title":"J Cheminform"},{"key":"4260_CR3","doi-asserted-by":"crossref","unstructured":"Elhadad N, Pradhan S, Gorman S, et al. SemEval-2015 task 14: analysis of clinical text. In: Proceedings of the 9th international workshop on semantic evaluation (SemEval 2015). Denver: Association for Computational Linguistics; 2015. p. 303\u201310.","DOI":"10.18653\/v1\/S15-2051"},{"issue":"5","key":"4260_CR4","doi-asserted-by":"publisher","first-page":"552","DOI":"10.1136\/amiajnl-2011-000203","volume":"18","author":"\u00d6 Uzuner","year":"2011","unstructured":"Uzuner \u00d6, South BR, Shen S, et al. 2010 i2b2\/VA challenge on concepts, assertions, and relations in clinical text. J Am Med Inform Assoc. 2011;18(5):552\u20136.","journal-title":"J Am Med Inform Assoc"},{"key":"4260_CR5","doi-asserted-by":"crossref","unstructured":"Agirre AG, Marimon M, Intxaurrondo A, et al. Pharmaconer: Pharmacological substances, compounds and proteins named entity recognition track. In: Proceedings of The 5th workshop on BioNLP open shared tasks. Hong Kong: Association for Computational Linguistics; 2019; p. 1\u201310.","DOI":"10.18653\/v1\/D19-5701"},{"issue":"1","key":"4260_CR6","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s13321-014-0049-z","volume":"7","author":"R Leaman","year":"2015","unstructured":"Leaman R, Wei CH, Lu Z. tmChem: a high performance approach for chemical named entity recognition and normalization. J Cheminform. 2015;7(1):1\u201310.","journal-title":"J Cheminform"},{"issue":"12","key":"4260_CR7","doi-asserted-by":"publisher","first-page":"1633","DOI":"10.1093\/bioinformatics\/bts183","volume":"28","author":"T Rockt\u00e4schel","year":"2012","unstructured":"Rockt\u00e4schel T, Weidlich M, Leser U. ChemSpot: a hybrid system for chemical named entity recognition. Bioinformatics. 2012;28(12):1633\u201340.","journal-title":"Bioinformatics"},{"key":"4260_CR8","unstructured":"Huang Z, Xu W, Yu K. Bidirectional LSTM-CRF models for sequence tagging. arXiv preprint arXiv:1508.01991 (2015)."},{"key":"4260_CR9","doi-asserted-by":"crossref","unstructured":"Lample G, Ballesteros M, Subramanian S, et al. Neural architectures for named entity recognition. In: Proceedings of the 2016 conference of the North American chapter of the association for computational linguistics: human language technologies. San Diego: Association for Computational Linguistics; 2016. p. 260\u201370.","DOI":"10.18653\/v1\/N16-1030"},{"key":"4260_CR10","unstructured":"Li L, Jin L, Jiang Z, et al. Biomedical named entity recognition based on extended recurrent neural networks. In: IEEE international conference on bioinformatics and biomedicine (BIBM). IEEE; 2015. p. 649\u201352."},{"key":"4260_CR11","doi-asserted-by":"crossref","unstructured":"Chalapathy R, Borzeshi EZ, Piccardi M. An investigation of recurrent neural architectures for drug name recognition. arXiv preprint arXiv:1609.07585 (2016).","DOI":"10.18653\/v1\/W16-6101"},{"key":"4260_CR12","unstructured":"Mikolov T, Chen K, Corrado G, et al. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781 (2013)."},{"key":"4260_CR13","unstructured":"Mikolov T, Sutskever I, Chen K, et al. Distributed representations of words and phrases and their compositionality. In: Advances in neural information processing systems; 2013. p. 3111\u20139."},{"key":"4260_CR14","doi-asserted-by":"crossref","unstructured":"Pennington J, Socher R, Manning CD. Glove: global vectors for word representation. In: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP); 2014. p. 1532\u201343.","DOI":"10.3115\/v1\/D14-1162"},{"key":"4260_CR15","doi-asserted-by":"crossref","unstructured":"Peters M, Neumann M, Iyyer M, et al. Deep contextualized word representations. In: Proceedings of the conference of the North American chapter of the association for computational linguistics; 2018. p. 2227\u201337.","DOI":"10.18653\/v1\/N18-1202"},{"key":"4260_CR16","unstructured":"Akbik A, Blythe D, Vollgraf R. Contextual string embeddings for sequence labeling. In: Proceedings of the 27th international conference on computational linguistics; 2018. p. 1638\u201349."},{"key":"4260_CR17","unstructured":"Devlin J, Chang M-W, Lee K et al. BERT: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the conference of the North American chapter of the association for computational linguistics; 2019. p. 4171\u20134186."},{"key":"4260_CR18","doi-asserted-by":"crossref","unstructured":"Xiong Y, Shen Y, Huang Y, et al. A deep learning-based system for PharmaCoNER. In: Proceedings of The 5th workshop on BioNLP open shared tasks. Hong Kong: Association for Computational Linguistics; 2019. p. 33\u20137.","DOI":"10.18653\/v1\/D19-5706"},{"key":"4260_CR19","doi-asserted-by":"crossref","unstructured":"Stoeckel M, Hemati W, Mehler A. When specialization helps: using pooled contextualized embeddings to detect chemical and biomedical entities in Spanish. In: Proceedings of the 5th workshop on BioNLP open shared tasks. Hong Kong: Association for Computational Linguistics; 2019. p. 11\u20135.","DOI":"10.18653\/v1\/D19-5702"},{"key":"4260_CR20","doi-asserted-by":"crossref","unstructured":"Akbik A, Bergmann T, Vollgraf R. Pooled contextualized embeddings for named entity recognition. In: Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, vol. 1 (Long and Short Papers); 2019. p. 724\u20138.","DOI":"10.18653\/v1\/N19-1078"},{"key":"4260_CR21","doi-asserted-by":"crossref","unstructured":"Sun C, Yang Z. Transfer learning in biomedical named entity recognition: an evaluation of BERT in the PharmaCoNER task. In: Proceedings of The 5th workshop on BioNLP open shared tasks. Hong Kong: Association for Computational Linguistics; 2019. p. 100\u20134.","DOI":"10.18653\/v1\/D19-5715"},{"issue":"4","key":"4260_CR22","doi-asserted-by":"publisher","first-page":"1234","DOI":"10.1093\/bioinformatics\/btz682","volume":"36","author":"J Lee","year":"2020","unstructured":"Lee J, Yoon W, Kim S, et al. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics. 2020;36(4):1234\u201340.","journal-title":"Bioinformatics"},{"key":"4260_CR23","unstructured":"P\u00e9rez-P\u00e9rez M, Rabal O, P\u00e9rez-Rodr\u00edguez G, et al. Evaluation of chemical and gene\/protein entity recognition systems at BioCreative V. 5: the CEMP and GPRO patents tracks; 2017. p. 1\u20138."},{"key":"4260_CR24","unstructured":"Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. In: Advances in neural information processing systems; 2017. p. 5998\u20136008."},{"key":"4260_CR25","unstructured":"Wu Y, Schuster M, Chen Z, et al. Google\u2019s neural machine translation system: bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144 (2016)."},{"issue":"4","key":"4260_CR26","first-page":"415","volume":"30","author":"WL Taylor","year":"1953","unstructured":"Taylor WL. \u201cCloze procedure\u2019\u2019: a new tool for measuring readability. J Q. 1953;30(4):415\u201333.","journal-title":"J Q"},{"key":"4260_CR27","doi-asserted-by":"crossref","unstructured":"Zhu Y, Kiros R, Zemel R, et al. Aligning books and movies: towards story-like visual explanations by watching movies and reading books. In: Proceedings of the IEEE international conference on computer vision; 2015. p. 19\u201327.","DOI":"10.1109\/ICCV.2015.11"},{"key":"4260_CR28","doi-asserted-by":"crossref","unstructured":"Peng Y, Yan S, Lu Z. Transfer learning in biomedical natural language processing: an evaluation of BERT and ELMo on ten benchmarking datasets. In: Proceedings of the 2019 workshop on biomedical natural language processing (BioNLP 2019); 2019. p. 58\u201365.","DOI":"10.18653\/v1\/W19-5006"},{"key":"4260_CR29","unstructured":"Canete J, Chaperon G, Fuentes R, et al. Spanish pre-trained bert model and evaluation data. PML4DC at ICLR, 2020."},{"key":"4260_CR30","unstructured":"Tiedemann J. Parallel data, tools and interfaces in OPUS. In: Lrec; 2012. p. 2214\u201318."},{"key":"4260_CR31","doi-asserted-by":"crossref","unstructured":"Beltagy I, Lo K, Cohan A. SciBERT: a pretrained language model for scientific text. In: Conference on empirical methods in natural language processing. Hong Kong: Association for Computational Linguistics; 2019. p. 3613\u201318.","DOI":"10.18653\/v1\/D19-1371"},{"issue":"2","key":"4260_CR32","doi-asserted-by":"publisher","first-page":"e15","DOI":"10.5808\/GI.2019.17.2.e15","volume":"17","author":"J Armengol-Estap\u00e9","year":"2019","unstructured":"Armengol-Estap\u00e9 J, Soares F, Marimon M, et al. PharmacoNER Tagger: a deep learning-based tool for automatically finding chemicals and drugs in Spanish medical texts. Genomics Inform. 2019;17(2):e15.","journal-title":"Genomics Inform"},{"key":"4260_CR33","doi-asserted-by":"crossref","unstructured":"Soares F, Villegas M, Gonzalez-Agirre A, et al. Medical word embeddings for Spanish: development and evaluation. In: Proceedings of the 2nd clinical natural language processing workshop. Minneapolis: Association for Computational Linguistics; 2019. p. 124\u201333","DOI":"10.18653\/v1\/W19-1916"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-021-04260-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12859-021-04260-y\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-021-04260-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,12,8]],"date-time":"2023-12-08T18:02:31Z","timestamp":1702058551000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-021-04260-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,12,17]]},"references-count":33,"journal-issue":{"issue":"S1","published-online":{"date-parts":[[2021,12]]}},"alternative-id":["4260"],"URL":"https:\/\/doi.org\/10.1186\/s12859-021-04260-y","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,12,17]]},"assertion":[{"value":"19 May 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"31 May 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"17 December 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}}],"article-number":"602"}}