{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,30]],"date-time":"2025-07-30T11:48:09Z","timestamp":1753876089334,"version":"3.41.2"},"reference-count":108,"publisher":"Oxford University Press (OUP)","license":[{"start":{"date-parts":[[2022,7,1]],"date-time":"2022-07-01T00:00:00Z","timestamp":1656633600000},"content-version":"vor","delay-in-days":181,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,7,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The identification of chemicals in articles has attracted a large interest in the biomedical scientific community, given its importance in drug development research. Most of previous research have focused on PubMed abstracts, and further investigation using full-text documents is required because these contain additional valuable information that must be explored. The manual expert task of indexing Medical Subject Headings (MeSH) terms to these articles later helps researchers find the most relevant publications for their ongoing work. The BioCreative\u00a0VII NLM-Chem track fostered the development of systems for chemical identification and indexing in PubMed full-text articles. Chemical identification consisted in identifying the chemical mentions and linking these to unique MeSH identifiers. This manuscript describes our participation system and the post-challenge improvements we made. We propose a three-stage pipeline that individually performs chemical mention detection, entity normalization and indexing. Regarding chemical identification, we adopted a deep-learning solution that utilizes the PubMedBERT contextualized embeddings followed by a multilayer perceptron and a conditional random field tagging layer. For the normalization approach, we use a sieve-based dictionary filtering followed by a deep-learning similarity search strategy. Finally, for the indexing we developed rules for identifying the more relevant MeSH codes for each article. During the challenge, our system obtained the best official results in the normalization and indexing tasks despite the lower performance in the chemical mention recognition task. In a post-contest phase we boosted our results by improving our named entity recognition model with additional techniques. The final system achieved 0.8731, 0.8275 and 0.4849 in the chemical identification, normalization and indexing tasks, respectively. The code to reproduce our experiments and run the pipeline is publicly available.<\/jats:p><jats:p>Database URL<\/jats:p><jats:p>https:\/\/github.com\/bioinformatics-ua\/biocreativeVII_track2<\/jats:p>","DOI":"10.1093\/database\/baac047","type":"journal-article","created":{"date-parts":[[2022,7,1]],"date-time":"2022-07-01T14:58:28Z","timestamp":1656687508000},"source":"Crossref","is-referenced-by-count":7,"title":["Chemical identification and indexing in PubMed full-text articles using deep learning and heuristics"],"prefix":"10.1093","volume":"2022","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4258-3350","authenticated-orcid":false,"given":"Tiago","family":"Almeida","sequence":"first","affiliation":[{"name":"Department of Electronics, Telecommunications and Informatics (DETI), Institute of Electronics and Informatics Engineering of Aveiro (IEETA), University of Aveiro , Aveiro, Portugal"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3533-8872","authenticated-orcid":false,"given":"Rui","family":"Antunes","sequence":"additional","affiliation":[{"name":"Department of Electronics, Telecommunications and Informatics (DETI), Institute of Electronics and Informatics Engineering of Aveiro (IEETA), University of Aveiro , Aveiro, Portugal"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5535-754X","authenticated-orcid":false,"given":"Jo\u00e3o","family":"F. Silva","sequence":"additional","affiliation":[{"name":"Department of Electronics, Telecommunications and Informatics (DETI), Institute of Electronics and Informatics Engineering of Aveiro (IEETA), University of Aveiro , Aveiro, Portugal"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0729-2264","authenticated-orcid":false,"given":"Jo\u00e3o R","family":"Almeida","sequence":"additional","affiliation":[{"name":"Department of Electronics, Telecommunications and Informatics (DETI), Institute of Electronics and Informatics Engineering of Aveiro (IEETA), University of Aveiro , Aveiro, Portugal"},{"name":"Department of Information and Communications Technologies, University of A Coru\u00f1a , A Coru\u00f1a, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1941-3983","authenticated-orcid":false,"given":"S\u00e9rgio","family":"Matos","sequence":"additional","affiliation":[{"name":"Department of Electronics, Telecommunications and Informatics (DETI), Institute of Electronics and Informatics Engineering of Aveiro (IEETA), University of Aveiro , Aveiro, Portugal"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2022,7,1]]},"reference":[{"key":"2022070114575651600_R1","doi-asserted-by":"publisher","first-page":"457","DOI":"10.1038\/nj7612-457a","article-title":"Scientific literature: information overload","volume":"535","author":"Landhuis","year":"2016","journal-title":"Nature"},{"key":"2022070114575651600_R2","doi-asserted-by":"publisher","first-page":"8","DOI":"10.1109\/MIS.2015.68","article-title":"Information extraction","volume":"30","author":"Grishman","year":"2015","journal-title":"IEEE Intell. Syst."},{"key":"2022070114575651600_R3","article-title":"Understanding PubMed user search behavior through log analysis","volume":"2009","author":"Dogan","year":"2009","journal-title":"Database"},{"key":"2022070114575651600_R4","first-page":"265","article-title":"Medical subject headings (MeSH)","volume":"88","author":"Lipscomb","year":"2000","journal-title":"Bull. Med. Libr. Assoc."},{"key":"2022070114575651600_R5","first-page":"pp108","article-title":"The overview of the NLM-Chem BioCreative VII track: full-text chemical identification and indexing in PubMed articles","author":"Leaman","year":"2021"},{"key":"2022070114575651600_R6","doi-asserted-by":"publisher","DOI":"10.1093\/bib\/6.1.57","article-title":"A survey of current work in biomedical text mining","volume":"6","author":"Cohen","journal-title":"Brief. Bioinform."},{"key":"2022070114575651600_R7","doi-asserted-by":"publisher","first-page":"132","DOI":"10.1093\/bib\/bbv024","article-title":"Community challenges in biomedical text mining over 10 years: success, failure and the future","volume":"17","author":"Huang","year":"2016","journal-title":"Brief. Bioinform."},{"key":"2022070114575651600_R8","doi-asserted-by":"publisher","first-page":"381","DOI":"10.1073\/pnas.98.2.381","article-title":"PubMed Central: the GenBank of the published literature","volume":"98","author":"Roberts","year":"2001","journal-title":"National Academy of Sciences of The United States Of America"},{"key":"2022070114575651600_R9","doi-asserted-by":"crossref","first-page":"261","DOI":"10.1561\/1900000003","article-title":"Information extraction","volume":"1","author":"Sarawagi","year":"2008","journal-title":"Found. Trends. Databases"},{"key":"2022070114575651600_R10","doi-asserted-by":"publisher","first-page":"i331","DOI":"10.1093\/bioinformatics\/btg1046","article-title":"Evaluation of text data mining for database curation: lessons learned from the KDD Challenge Cup","volume":"19","author":"Yeh","year":"2003","journal-title":"Bioinformatics"},{"key":"2022070114575651600_R11","doi-asserted-by":"publisher","first-page":"47","DOI":"10.1038\/455047a","article-title":"The future of biocuration","volume":"455","author":"Howe","year":"2008","journal-title":"Nature"},{"key":"2022070114575651600_R12","doi-asserted-by":"publisher","first-page":"2219","DOI":"10.1093\/bib\/bbaa054","article-title":"Biomedical named entity recognition and linking datasets: survey and our recent development","volume":"21","author":"Huang","year":"2020","journal-title":"Brief. Bioinform."},{"key":"2022070114575651600_R13","doi-asserted-by":"publisher","first-page":"552","DOI":"10.1136\/amiajnl-2011-000203","article-title":"2010 i2b2\/VA challenge on concepts, assertions, and relations in clinical text","volume":"18","author":"Uzuner","year":"2011","journal-title":"J. Am. Med. Inf. Assoc."},{"key":"2022070114575651600_R14","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1093\/jamia\/ocz166","article-title":"2018 n2c2 shared task on adverse drug events and medication extraction in electronic health records","volume":"27","author":"Henry","year":"2021","journal-title":"J. Am. Med. Inf. Assoc."},{"key":"2022070114575651600_R15","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-12-223","article-title":"Exploiting MeSH indexing in MEDLINE to generate a data set for word sense disambiguation","volume":"12","author":"Jimeno-Yepes","year":"2011","journal-title":"BMC Bioinform."},{"key":"2022070114575651600_R16","doi-asserted-by":"publisher","DOI":"10.1093\/nar\/gki031","article-title":"Entrez Gene: gene-centered information at NCBI","volume":"33","author":"Maglott","year":"2005","journal-title":"Nucleic Acids Res."},{"key":"2022070114575651600_R17","doi-asserted-by":"publisher","DOI":"10.1093\/nar\/gks1146","article-title":"The ChEBI reference database and ontology for biologically relevant chemistry: enhancements for 2013","volume":"41","author":"Hastings","year":"2013","journal-title":"Nucleic Acids Res."},{"key":"2022070114575651600_R18","first-page":"pp. 4","article-title":"Extraction of gene\u2013disease relations from Medline using domain dictionaries and machine learning","author":"Chun","year":"2006"},{"key":"2022070114575651600_R19","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-8-50","article-title":"BioInfer: a corpus for information extraction in the biomedical domain","volume":"8","author":"Pyysalo","year":"2007","journal-title":"BMC Bioinform."},{"key":"2022070114575651600_R20","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-9-S3-S6","article-title":"Comparative analysis of five protein\u2013protein interaction corpora","volume":"9","author":"Pyysalo","year":"2008","journal-title":"BMC Bioinform."},{"key":"2022070114575651600_R21","doi-asserted-by":"publisher","DOI":"10.1093\/database\/baw032","article-title":"Assessing the state of the art in biomedical relation extraction: overview of the BioCreative V chemical\u2013disease relation (CDR) task","volume":"2016","author":"Wei","year":"2016","journal-title":"Database"},{"key":"2022070114575651600_R22","first-page":"pp. 141","article-title":"Overview of the BioCreative VI chemical\u2013protein interaction track","author":"Krallinger","year":"2017"},{"key":"2022070114575651600_R23","first-page":"pp. 11","article-title":"Overview of DrugProt BioCreative VII track: quality evaluation and large scale text mining of drug-gene\/protein relations","author":"Miranda","year":"2021"},{"key":"2022070114575651600_R24","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3445965","article-title":"Named entity recognition and relation extraction: state-of-the-art","volume":"54","author":"Nasar","year":"2021","journal-title":"ACM Comput. Surv."},{"key":"2022070114575651600_R25","doi-asserted-by":"publisher","first-page":"143","DOI":"10.1136\/amiajnl-2013-002544","article-title":"Evaluating the state of the art in disorder recognition and normalization of the clinical narrative","volume":"22","author":"Pradhan","year":"2015","journal-title":"J. Am. Med. Inf. Assoc."},{"key":"2022070114575651600_R26","first-page":"pp. 147","article-title":"Design challenges and misconceptions in named entity recognition","author":"Ratinov","year":"2009"},{"key":"2022070114575651600_R27","doi-asserted-by":"publisher","DOI":"10.1186\/1758-2946-7-S1-S14","article-title":"Enhancing of chemical compound and drug name recognition using representative tag scheme and fine-grained tokenization","volume":"7","author":"Dai","year":"2015","journal-title":"J. Cheminf."},{"key":"2022070114575651600_R28","first-page":"pp. 260","article-title":"Neural architectures for named entity recognition","author":"Lample","year":"2016"},{"key":"2022070114575651600_R29","first-page":"pp175","article-title":"Biomedical named entity recognition: a survey of machine-learning tools","author":"Campos","year":"2012"},{"key":"2022070114575651600_R30","doi-asserted-by":"publisher","first-page":"i37","DOI":"10.1093\/bioinformatics\/btx228","article-title":"Deep learning with word embeddings improves biomedical named entity recognition","volume":"33","author":"Habibi","year":"2017","journal-title":"Bioinformatics"},{"key":"2022070114575651600_R31","article-title":"Efficient estimation of word representations in vector space","volume-title":"arXiv:1301.3781","author":"Mikolov","year":"2013"},{"key":"2022070114575651600_R32","first-page":"pp. 39","article-title":"Distributional semantics resources for biomedical text processing","author":"Pyysalo","year":"2013"},{"key":"2022070114575651600_R33","first-page":"pp. 1105","article-title":"End-to-end relation extraction using LSTMs on sequences and tree structures","author":"Miwa","year":"2016"},{"key":"2022070114575651600_R34","doi-asserted-by":"publisher","first-page":"34","DOI":"10.1016\/j.eswa.2018.07.032","article-title":"Joint entity recognition and relation extraction as a multi-head selection problem","volume":"114","author":"Bekoulis","year":"2018","journal-title":"Expert Syst. Appl."},{"key":"2022070114575651600_R35","first-page":"pp. 17","article-title":"Effective mapping of biomedical text to the UMLS Metathesaurus: the MetaMap program","author":"Aronson","year":"2001"},{"key":"2022070114575651600_R36","doi-asserted-by":"publisher","first-page":"229","DOI":"10.1136\/jamia.2009.002733","article-title":"An overview of MetaMap: historical perspective and recent advances","volume":"17","author":"Aronson","year":"2010","journal-title":"J. Am. Med. Inf. Assoc."},{"key":"2022070114575651600_R37","doi-asserted-by":"publisher","first-page":"507","DOI":"10.1136\/jamia.2009.001560","article-title":"Mayo clinical Text Analysis and Knowledge Extraction System (cTAKES): architecture, component evaluation and applications","volume":"17","author":"Savova","year":"2010","journal-title":"J. Am. Med. Inf. Assoc."},{"key":"2022070114575651600_R38","doi-asserted-by":"publisher","first-page":"2909","DOI":"10.1093\/bioinformatics\/btt474","article-title":"DNorm: disease name normalization with pairwise learning to rank","volume":"29","author":"Leaman","year":"2013","journal-title":"Bioinformatics"},{"key":"2022070114575651600_R39","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.jbi.2013.12.006","article-title":"NCBI disease corpus: a resource for disease name recognition and concept normalization","volume":"47","author":"Dogan","year":"2014","journal-title":"J. Biomed. Inform."},{"key":"2022070114575651600_R40","first-page":"pp. 303","article-title":"SemEval-2015 Task 14: analysis of clinical text","author":"Elhadad","year":"2015"},{"key":"2022070114575651600_R41","first-page":"pp. 406","article-title":"ULisboa: recognition and normalization of medical concepts","author":"Leal","year":"2015"},{"key":"2022070114575651600_R42","doi-asserted-by":"publisher","DOI":"10.1186\/1758-2946-7-S1-S3","article-title":"tmChem: a high performance approach for chemical named entity recognition and normalization","volume":"7","author":"Leaman","year":"2015","journal-title":"J. Cheminf."},{"key":"2022070114575651600_R43","doi-asserted-by":"publisher","first-page":"28","DOI":"10.1016\/j.jbi.2015.07.010","article-title":"Challenges in clinical natural language processing for automated disorder normalization","volume":"57","author":"Leaman","year":"2015","journal-title":"J. Biomed. Inform."},{"key":"2022070114575651600_R44","doi-asserted-by":"publisher","first-page":"2839","DOI":"10.1093\/bioinformatics\/btw343","article-title":"TaggerOne: joint named entity recognition and normalization with semi-Markov Models","volume":"32","author":"Leaman","year":"2016","journal-title":"Bioinformatics"},{"key":"2022070114575651600_R45","first-page":"pp. 173","article-title":"Annotating chemicals, diseases and their interactions in biomedical literature","author":"Li","year":"2015"},{"key":"2022070114575651600_R46","doi-asserted-by":"publisher","DOI":"10.1093\/database\/baw068","article-title":"BioCreative V CDR task corpus: a resource for chemical disease relation extraction","volume":"2016","author":"Li","journal-title":"Database"},{"key":"2022070114575651600_R47","first-page":"pp. 2045","article-title":"Biomedical term normalization of EHRs with UMLS","author":"P\u00e9rez-Miguel","year":"2018"},{"key":"2022070114575651600_R48","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbi.2019.103132","article-title":"MCN: a comprehensive corpus for medical concept normalization","volume":"92","author":"Luo","year":"2019","journal-title":"J. Biomed. Inform."},{"key":"2022070114575651600_R49","doi-asserted-by":"publisher","first-page":"1529","DOI":"10.1093\/jamia\/ocaa106","article-title":"The 2019 n2c2\/UMass Lowell shared task on clinical concept normalization","volume":"27","author":"Luo","year":"2020","journal-title":"J. Am. Med. Inf. Assoc."},{"key":"2022070114575651600_R50","first-page":"pp. 93","article-title":"Clinical concept normalization on medical records using word embeddings and heuristics","author":"Silva","year":"2020"},{"key":"2022070114575651600_R51","doi-asserted-by":"crossref","DOI":"10.1038\/s41597-019-0055-0","article-title":"BioWordVec, improving biomedical word embeddings with subword information and MeSH","volume":"6","author":"Zhang","year":"2019","journal-title":"Sci. Data"},{"key":"2022070114575651600_R52","first-page":"pp. 817","article-title":"A neural multi-task learning framework to jointly model medical named entity recognition and normalization","author":"Zhao","year":"2019"},{"key":"2022070114575651600_R53","doi-asserted-by":"publisher","first-page":"73729","DOI":"10.1109\/ACCESS.2019.2920708","article-title":"A neural named entity recognition and multi-type normalization tool for biomedical text mining","volume":"7","author":"Kim","year":"2019","journal-title":"IEEE Access"},{"key":"2022070114575651600_R54","doi-asserted-by":"crossref","first-page":"1234","DOI":"10.1093\/bioinformatics\/btz682","article-title":"BioBERT: a pre-trained biomedical language representation model for biomedical text mining","volume":"36","author":"Lee","year":"2020","journal-title":"Bioinformatics"},{"key":"2022070114575651600_R55","doi-asserted-by":"publisher","DOI":"10.1186\/s12859-020-03583-6","article-title":"pyMeSHSim: an integrative python package for biomedical named entity recognition, normalization, and comparison of MeSH terms","volume":"21","author":"Luo","year":"2020","journal-title":"BMC Bioinform."},{"key":"2022070114575651600_R56","doi-asserted-by":"publisher","first-page":"1510","DOI":"10.1093\/jamia\/ocaa080","article-title":"Unified Medical Language System resources improve sieve-based generation and Bidirectional Encoder Representations from Transformers (BERT)\u2013based ranking for concept normalization","volume":"27","author":"Xu","year":"2020","journal-title":"J. Am. Med. Inf. Assoc."},{"key":"2022070114575651600_R57","first-page":"pp. 422","article-title":"LasigeBioTM at CANTEMIST: named entity recognition and normalization of tumour morphology entities and clinical coding of Spanish health-related documents","author":"Ruas","year":"2020"},{"key":"2022070114575651600_R58","first-page":"pp. 303","article-title":"Named entity recognition, concept normalization and clinical coding: overview of the Cantemist track for cancer text mining in Spanish, corpus, guidelines, methods and results","author":"Miranda-Escalada","year":"2020"},{"key":"2022070114575651600_R59","doi-asserted-by":"publisher","first-page":"1576","DOI":"10.1093\/jamia\/ocaa155","article-title":"Clinical concept normalization with a hybrid natural language processing system combining multilevel matching and machine learning ranking","volume":"27","author":"Chen","year":"2020","journal-title":"J. Am. Med. Inf. Assoc."},{"key":"2022070114575651600_R60","doi-asserted-by":"publisher","DOI":"10.2196\/23104","article-title":"Clinical term normalization using learned edit patterns and subconcept matching: system development and evaluation","volume":"9","author":"Kate","year":"2021","journal-title":"JMIR Medical Informatics"},{"key":"2022070114575651600_R61","doi-asserted-by":"publisher","first-page":"516","DOI":"10.1093\/jamia\/ocaa269","article-title":"Ambiguity in medical concept normalization: an analysis of types and coverage in electronic health record datasets","volume":"28","author":"Newman-Griffis","journal-title":"J. Am. Med. Inf. Assoc."},{"key":"2022070114575651600_R62","first-page":"pp. 11","article-title":"Triplet-trained vector space and sieve-based search improve biomedical concept normalization","author":"Xu","year":"2021"},{"key":"2022070114575651600_R63","first-page":"pp. 6214","article-title":"An end-to-end progressive multi-task learning framework for medical named entity recognition and normalization","author":"Zhou","year":"2021"},{"key":"2022070114575651600_R64","doi-asserted-by":"publisher","DOI":"10.1016\/j.jbi.2021.103880","article-title":"Improving broad-coverage medical entity linking with semantic type prediction and large-scale datasets","volume":"121","author":"Vashishth","year":"2021","journal-title":"J. Biomed. Inform."},{"key":"2022070114575651600_R65","first-page":"pp. 460","article-title":"Gene indexing: characterization and analysis of NLM\u2019s GeneRIFs","author":"Mitchell","year":"2003"},{"key":"2022070114575651600_R66","first-page":"pp. 709","article-title":"Comparison and combination of several MeSH indexing approaches","author":"Yepes","year":"2013"},{"key":"2022070114575651600_R67","doi-asserted-by":"publisher","first-page":"i339","DOI":"10.1093\/bioinformatics\/btv237","article-title":"MeSHLabeler: improving the accuracy of large-scale MeSH indexing by integrating diverse evidence","volume":"31","author":"Liu","year":"2015","journal-title":"Bioinformatics"},{"key":"2022070114575651600_R68","doi-asserted-by":"publisher","first-page":"i70","DOI":"10.1093\/bioinformatics\/btw294","article-title":"DeepMeSH: deep semantic representation for improving large-scale MeSH indexing","volume":"32","author":"Peng","year":"2016","journal-title":"Bioinformatics"},{"key":"2022070114575651600_R69","doi-asserted-by":"publisher","first-page":"389","DOI":"10.1016\/j.sapharm.2016.04.006","article-title":"Comparison of the time-to-indexing in PubMed between biomedical journals according to impact factor, discipline, and focus","volume":"13","author":"Irwin","year":"2017","journal-title":"Res. Soc. Administrative Pharmacy"},{"key":"2022070114575651600_R70","doi-asserted-by":"publisher","DOI":"10.1186\/s13326-017-0123-3","article-title":"MeSH Now: automatic MeSH indexing at PubMed scale via learning to rank","volume":"8","author":"Mao","year":"2017","journal-title":"J. Biomed. Semant."},{"key":"2022070114575651600_R71","doi-asserted-by":"publisher","first-page":"1533","DOI":"10.1093\/bioinformatics\/btz756","article-title":"FullMeSH: improving large-scale MeSH indexing with full text","volume":"36","author":"Dai","journal-title":"Bioinformatics"},{"key":"2022070114575651600_R72","doi-asserted-by":"publisher","first-page":"684","DOI":"10.1093\/bioinformatics\/btaa837","article-title":"BERTMeSH: deep contextual representation learning for large-scale high-performance MeSH indexing with full text","volume":"37","author":"You","journal-title":"Bioinformatics"},{"key":"2022070114575651600_R73","doi-asserted-by":"publisher","DOI":"10.1016\/j.artmed.2021.102053","article-title":"NewsMeSH: a new classifier designed to annotate health news with MeSH headings","volume":"114","author":"Costa","year":"2021","journal-title":"Artificial Intelligence in Medicine"},{"key":"2022070114575651600_R74","first-page":"pp. 302","article-title":"A neural text ranking approach for automatic MeSH indexing","author":"Alastair","year":"2021"},{"key":"2022070114575651600_R75","first-page":"pp. 114","article-title":"The chemical corpus of the NLM-Chem BioCreative VII track: full-text chemical identification and indexing in PubMed articles","author":"Islamaj","year":"2021"},{"key":"2022070114575651600_R76","doi-asserted-by":"publisher","DOI":"10.1186\/1758-2946-7-S1-S2","article-title":"The CHEMDNER corpus of chemicals and drugs and its annotation principles","volume":"7","author":"Krallinger","year":"2015","journal-title":"J. Cheminf."},{"key":"2022070114575651600_R77","doi-asserted-by":"publisher","DOI":"10.1186\/s12859-017-1776-8","article-title":"A neural network multi-task learning approach to biomedical named entity recognition","volume":"18","author":"Crichton","year":"2017","journal-title":"BMC Bioinform."},{"key":"2022070114575651600_R78","first-page":"pp. 119","article-title":"Chemical detection and indexing in PubMed full text articles using deep learning and rule-based methods","author":"Almeida","year":"2021"},{"key":"2022070114575651600_R79","doi-asserted-by":"publisher","DOI":"10.1038\/s41597-021-00875-1","article-title":"NLM-Chem, a new resource for chemical entity recognition in PubMed full text literature","volume":"8","author":"Islamaj","year":"2021","journal-title":"Sci. Data"},{"key":"2022070114575651600_R80","first-page":"pp. 140","article-title":"Improving tagging consistency and entity coverage for chemical identification in full-text articles","author":"Kim","year":"2021"},{"key":"2022070114575651600_R81","first-page":"pp. 3861","article-title":"An analysis of simple data augmentation for named entity recognition","author":"Dai","year":"2020"},{"key":"2022070114575651600_R82","doi-asserted-by":"publisher","first-page":"D1138","DOI":"10.1093\/nar\/gkaa891","article-title":"Comparative Toxicogenomics Database (CTD): update 2021","volume":"49","author":"Davis","year":"2021","journal-title":"Nucleic Acids Res."},{"key":"2022070114575651600_R83","article-title":"Domain-specific language model pretraining for biomedical natural language processing","volume":"3","author":"Gu","year":"2021","journal-title":"ACM Trans. Comput. Healthcare"},{"article-title":"Experiment tracking with Weights and Biases","year":"2020","author":"Biewald","key":"2022070114575651600_R84"},{"key":"2022070114575651600_R85","first-page":"pp. 2024","article-title":"Masked conditional random fields for sequence labeling","author":"Wei","year":"2021"},{"key":"2022070114575651600_R86","first-page":"pp. 130","article-title":"A BERT-based hybrid system for chemical identification and indexing in full-text articles","author":"Erdengasileng","year":"2021"},{"key":"2022070114575651600_R87","first-page":"pp. 2623","article-title":"Optuna: a next-generation hyperparameter optimization framework","author":"Akiba","year":"2019"},{"key":"2022070114575651600_R88","first-page":"pp. 533","article-title":"Multiobjective tree-structured parzen estimator for computationally expensive optimization problems","author":"Ozaki","year":"2020"},{"key":"2022070114575651600_R89","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-9-402","article-title":"Abbreviation definition identification based on automatic precision estimates","volume":"9","author":"Sohn","year":"2008","journal-title":"BMC Bioinform."},{"key":"2022070114575651600_R90","first-page":"pp. 4228","article-title":"Self-alignment pretraining for biomedical entity representations","author":"Liu","year":"2021"},{"key":"2022070114575651600_R91","doi-asserted-by":"publisher","first-page":"75","DOI":"10.1002\/asi.4630230202","article-title":"A new comparison between conventional indexing (MEDLARS) and automatic text processing (SMART)","volume":"23","author":"Salton","journal-title":"J. Am. Soc. Inform. Sci."},{"key":"2022070114575651600_R92","doi-asserted-by":"publisher","first-page":"1381","DOI":"10.1093\/bioinformatics\/btx761","article-title":"An attention-based BiLSTM-CRF approach to document-level chemical named entity recognition","volume":"34","author":"Luo","year":"2018","journal-title":"Bioinformatics"},{"key":"2022070114575651600_R93","doi-asserted-by":"crossref","first-page":"291","DOI":"10.1162\/tacl_a_00461","article-title":"ByT5: towards a token-free future with pre-trained byte-to-byte models","volume":"10","author":"Xue","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"2022070114575651600_R94","first-page":"pp. 146","article-title":"Pretrained language models for biomedical and clinical tasks: understanding and extending the state-of-the-art","author":"Lewis","year":"2020"},{"key":"2022070114575651600_R95","first-page":"pp. 3641","article-title":"Biomedical entity representations with synonym marginalization","author":"Sung","year":"2020"},{"key":"2022070114575651600_R96","first-page":"pp. 148","article-title":"Chemical identification and indexing in PubMed articles via BERT and text-to-text approaches","author":"Adams","year":"2021"},{"key":"2022070114575651600_R97","first-page":"pp. 4700","article-title":"BioMegatron: larger biomedical domain language model","author":"Shin","year":"2020"},{"key":"2022070114575651600_R98","first-page":"pp. 127","article-title":"Recognizing chemical entity in biomedical literature using a BERT-based ensemble learning methods for the BioCreative 2021 NLM-Chem track","author":"Chiu","year":"2021"},{"key":"2022070114575651600_R99","first-page":"pp. 221","article-title":"BioM-Transformers: building large biomedical language models with BERT, ALBERT and ELECTRA","author":"Alrowili","year":"2021"},{"key":"2022070114575651600_R100","first-page":"pp. 144","article-title":"Fine-tuning transformers for automatic chemical entity identification in PubMed articles","author":"Bevan","year":"2021"},{"key":"2022070114575651600_R101","first-page":"pp. 156","article-title":"TTI-COIN at BioCreative VII Track 2: fully neural NER, linking, and indexing models","author":"Tsujimura","year":"2021"},{"key":"2022070114575651600_R102","first-page":"pp. 3615","article-title":"SciBERT: a pretrained language model for scientific text","author":"Beltagy","year":"2019"},{"key":"2022070114575651600_R103","first-page":"pp. 152","article-title":"Chemical entity recognition and MeSH normalization in PubMed full-text literature using BioBERT","author":"L\u00f3pez-\u00dabeda","year":"2021"},{"key":"2022070114575651600_R104","first-page":"pp. 2227","article-title":"Deep contextualized word representations","author":"Peters","year":"2018"},{"key":"2022070114575651600_R105","first-page":"pp. 124","article-title":"Rule-based enhancement of Stanza NER","author":"Mercer","year":"2021"},{"key":"2022070114575651600_R106","first-page":"pp. 101","article-title":"Stanza: A Python natural language processing toolkit for many human languages","author":"Qi","year":"2020"},{"key":"2022070114575651600_R107","doi-asserted-by":"publisher","first-page":"1892","DOI":"10.1093\/jamia\/ocab090","article-title":"Biomedical and clinical English model packages for the Stanza Python NLP library","volume":"28","author":"Zhang","year":"2021","journal-title":"J. Am. Med. Inf. Assoc."},{"key":"2022070114575651600_R108","first-page":"pp. 135","article-title":"Combining dictionary- and rule-based approximate entity linking with tuned BioBERT","author":"Mobasher","year":"2021"}],"container-title":["Database"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/database\/article-pdf\/doi\/10.1093\/database\/baac047\/44372699\/baac047.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/database\/article-pdf\/doi\/10.1093\/database\/baac047\/44372699\/baac047.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,10]],"date-time":"2023-02-10T13:59:19Z","timestamp":1676037559000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/database\/article\/doi\/10.1093\/database\/baac047\/6625810"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,1]]},"references-count":108,"URL":"https:\/\/doi.org\/10.1093\/database\/baac047","relation":{},"ISSN":["1758-0463"],"issn-type":[{"type":"electronic","value":"1758-0463"}],"subject":[],"published-other":{"date-parts":[[2022,1,1]]},"published":{"date-parts":[[2022,1,1]]},"article-number":"baac047"}}