{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T00:15:36Z","timestamp":1775002536987,"version":"3.50.1"},"reference-count":17,"publisher":"Oxford University Press (OUP)","funder":[{"name":"the NIH Intramural Research Program, National Library of Medicine"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,12,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>The automatic recognition of chemical names and their corresponding database identifiers in biomedical text is an important first step for many downstream text-mining applications. The task is even more challenging when considering the identification of these entities in the article\u2019s full text and, furthermore, the identification of candidate substances for that article\u2019s metadata [Medical Subject Heading (MeSH) article indexing]. The National Library of Medicine (NLM)-Chem track at BioCreative VII aimed to foster the development of algorithms that can predict with high quality the chemical entities in the biomedical literature and further identify the chemical substances that are candidates for article indexing. As a result of this challenge, the NLM-Chem track produced two comprehensive, manually curated corpora annotated with chemical entities and indexed with chemical substances: the chemical identification corpus and the chemical indexing corpus. The NLM-Chem BioCreative VII (NLM-Chem-BC7) Chemical Identification corpus consists of 204 full-text PubMed Central (PMC) articles, fully annotated for chemical entities by 12 NLM indexers for both span (i.e.\u00a0named entity recognition) and normalization (i.e.\u00a0entity linking) using MeSH. This resource was used for the training and testing of the Chemical Identification task to evaluate the accuracy of algorithms in predicting chemicals mentioned in recently published full-text articles. The NLM-Chem-BC7 Chemical Indexing corpus consists of 1333 recently published PMC articles, equipped with chemical substance indexing by manual experts at the NLM. This resource was used for the evaluation of the Chemical Indexing task, which evaluated the accuracy of algorithms in predicting the chemicals that should be indexed, i.e.\u00a0appear in the listing of MeSH terms for the document. This set was further enriched after the challenge in two ways: (i) 11 NLM indexers manually verified each of the candidate terms appearing in the prediction results of the challenge participants, but not in the MeSH indexing, and the chemical indexing terms appearing in the MeSH indexing list, but not in the prediction results, and (ii) the challenge organizers algorithmically merged the chemical entity annotations in the full text for all predicted chemical entities and used a statistical approach to keep those with the highest degree of confidence. As a result, the NLM-Chem-BC7 Chemical Indexing corpus is a gold-standard corpus for chemical indexing of journal articles and a silver-standard corpus for chemical entity identification in full-text journal articles. Together, these resources are currently the most comprehensive resources for chemical entity recognition, and we demonstrate improvements in the chemical entity recognition algorithms. We detail the characteristics of these novel resources and make them available for the community.<\/jats:p>\n               <jats:p>Database URL: https:\/\/ftp.ncbi.nlm.nih.gov\/pub\/lu\/NLM-Chem-BC7-corpus\/<\/jats:p>","DOI":"10.1093\/database\/baac102","type":"journal-article","created":{"date-parts":[[2022,12,2]],"date-time":"2022-12-02T09:45:10Z","timestamp":1669974310000},"source":"Crossref","is-referenced-by-count":7,"title":["NLM-Chem-BC7: manually annotated full-text resources for chemical entity annotation and indexing in biomedical articles"],"prefix":"10.1093","volume":"2022","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5651-1860","authenticated-orcid":false,"given":"Rezarta","family":"Islamaj","sequence":"first","affiliation":[{"name":"National Library of Medicine, National Institutes of Health , 8600 Rockville Pike, Bethesda, MD 20894, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Robert","family":"Leaman","sequence":"additional","affiliation":[{"name":"National Library of Medicine, National Institutes of Health , 8600 Rockville Pike, Bethesda, MD 20894, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"David","family":"Cissel","sequence":"additional","affiliation":[{"name":"National Library of Medicine, National Institutes of Health , 8600 Rockville Pike, Bethesda, MD 20894, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cathleen","family":"Coss","sequence":"additional","affiliation":[{"name":"National Library of Medicine, National Institutes of Health , 8600 Rockville Pike, Bethesda, MD 20894, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joseph","family":"Denicola","sequence":"additional","affiliation":[{"name":"National Library of Medicine, National Institutes of Health , 8600 Rockville Pike, Bethesda, MD 20894, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Carol","family":"Fisher","sequence":"additional","affiliation":[{"name":"National Library of Medicine, National Institutes of Health , 8600 Rockville Pike, Bethesda, MD 20894, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rob","family":"Guzman","sequence":"additional","affiliation":[{"name":"National Library of Medicine, National Institutes of Health , 8600 Rockville Pike, Bethesda, MD 20894, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Preeti Gokal","family":"Kochar","sequence":"additional","affiliation":[{"name":"National Library of Medicine, National Institutes of Health , 8600 Rockville Pike, Bethesda, MD 20894, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nicholas","family":"Miliaras","sequence":"additional","affiliation":[{"name":"National Library of Medicine, National Institutes of Health , 8600 Rockville Pike, Bethesda, MD 20894, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zoe","family":"Punske","sequence":"additional","affiliation":[{"name":"National Library of Medicine, National Institutes of Health , 8600 Rockville Pike, Bethesda, MD 20894, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Keiko","family":"Sekiya","sequence":"additional","affiliation":[{"name":"National Library of Medicine, National Institutes of Health , 8600 Rockville Pike, Bethesda, MD 20894, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dorothy","family":"Trinh","sequence":"additional","affiliation":[{"name":"National Library of Medicine, National Institutes of Health , 8600 Rockville Pike, Bethesda, MD 20894, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Deborah","family":"Whitman","sequence":"additional","affiliation":[{"name":"National Library of Medicine, National Institutes of Health , 8600 Rockville Pike, Bethesda, MD 20894, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Susan","family":"Schmidt","sequence":"additional","affiliation":[{"name":"National Library of Medicine, National Institutes of Health , 8600 Rockville Pike, Bethesda, MD 20894, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9998-916X","authenticated-orcid":false,"given":"Zhiyong","family":"Lu","sequence":"additional","affiliation":[{"name":"National Library of Medicine, National Institutes of Health , 8600 Rockville Pike, Bethesda, MD 20894, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2022,12,1]]},"reference":[{"key":"2022120211070785100_R1","doi-asserted-by":"crossref","DOI":"10.1093\/database\/bap018","article-title":"Understanding PubMed user search behavior through log analysis","volume":"2009","author":"Islamaj Dogan","year":"2009","journal-title":"Database (Oxford)"},{"key":"2022120211070785100_R2","doi-asserted-by":"crossref","first-page":"7673","DOI":"10.1021\/acs.chemrev.6b00851","article-title":"Information retrieval and text mining technologies for chemistry","volume":"117","author":"Krallinger","year":"2017","journal-title":"Chem. Rev."},{"key":"2022120211070785100_R3","doi-asserted-by":"crossref","DOI":"10.1186\/1758-2946-7-S1-S2","article-title":"The CHEMDNER corpus of chemicals and drugs and its annotation principles","volume":"7","author":"Krallinger","year":"2015","journal-title":"J. Cheminform."},{"key":"2022120211070785100_R4","doi-asserted-by":"crossref","DOI":"10.1038\/s41597-021-00875-1","article-title":"NLM-Chem, a new resource for chemical entity recognition in PubMed full text literature","volume":"8","author":"Islamaj","year":"2021","journal-title":"Sci. Data"},{"key":"2022120211070785100_R5","doi-asserted-by":"crossref","DOI":"10.1093\/database\/baw147","article-title":"The BioC-BioGRID corpus: full text articles annotated for curation of protein-protein and genetic interactions","volume":"2017","author":"Islamaj Dogan","year":"2017","journal-title":"Database (Oxford)"},{"key":"2022120211070785100_R6","doi-asserted-by":"crossref","DOI":"10.1186\/1471-2105-13-161","article-title":"Concept annotation in the CRAFT corpus","volume":"13","author":"Bada","year":"2012","journal-title":"BMC Bioinform."},{"key":"2022120211070785100_R7","first-page":"1400","article-title":"Biomedical text mining for research rigor and integrity: tasks, challenges, directions","volume":"19","author":"Kilicoglu","year":"2018","journal-title":"Brief. Bioinf."},{"key":"2022120211070785100_R8","article-title":"Chemical identification and indexing in full-text articles: overview of the NLM-Chem track at BioCreative VII","author":"Leaman","year":"2022","journal-title":"Database (Oxford)"},{"key":"2022120211070785100_R9","article-title":"Overview of the NLM-Chem BioCreative VII track: full-text chemical identification and indexing in PubMed articles","author":"Leaman","year":"2021"},{"key":"2022120211070785100_R10","article-title":"The chemical corpus of the NLM-Chem BioCreative VII track full-text chemical identification and indexing in PubMed articles","author":"Islamaj","year":"2021"},{"key":"2022120211070785100_R11","first-page":"268","article-title":"The NLM indexing initiative\u2019s medical text indexer","volume":"107","author":"Aronson","year":"2004","journal-title":"Stud. Health Technol. Inform."},{"key":"2022120211070785100_R12","article-title":"BioCreative V CDR task corpus: a resource for chemical disease relation extraction","volume":"2016","author":"Li","year":"2016","journal-title":"Database (Oxford)"},{"key":"2022120211070785100_R13","doi-asserted-by":"crossref","first-page":"W5","DOI":"10.1093\/nar\/gkaa333","article-title":"TeamTat: a collaborative text annotation tool","volume":"48","author":"Islamaj","year":"2020","journal-title":"Nucleic Acids Res."},{"key":"2022120211070785100_R14","doi-asserted-by":"crossref","DOI":"10.1093\/database\/bat064","article-title":"BioC: a minimalist approach to interoperability for biomedical text processing","volume":"2013","author":"Comeau","year":"2013","journal-title":"Database (Oxford)"},{"key":"2022120211070785100_R15","doi-asserted-by":"crossref","first-page":"3533","DOI":"10.1093\/bioinformatics\/btz070","article-title":"PMC text mining subset in BioC: about three million full-text articles and growing","volume":"35","author":"Comeau","year":"2019","journal-title":"Bioinformatics"},{"key":"2022120211070785100_R16","doi-asserted-by":"crossref","first-page":"W587","DOI":"10.1093\/nar\/gkz389","article-title":"PubTator Central: automated concept annotation for biomedical full text articles","volume":"47","author":"Wei","year":"2019","journal-title":"Nucleic Acids Res."},{"key":"2022120211070785100_R17","doi-asserted-by":"crossref","DOI":"10.1186\/s12859-015-0564-6","article-title":"An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competition","volume":"16","author":"Tsatsaronis","year":"2015","journal-title":"BMC Bioinform."}],"container-title":["Database"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/database\/article-pdf\/doi\/10.1093\/database\/baac102\/47513450\/baac102.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/database\/article-pdf\/doi\/10.1093\/database\/baac102\/47513450\/baac102.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,2]],"date-time":"2022-12-02T11:07:52Z","timestamp":1669979272000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/database\/article\/doi\/10.1093\/database\/baac102\/6858529"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,1]]},"references-count":17,"URL":"https:\/\/doi.org\/10.1093\/database\/baac102","relation":{},"ISSN":["1758-0463"],"issn-type":[{"value":"1758-0463","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2022,1,1]]},"published":{"date-parts":[[2022,1,1]]},"article-number":"baac102"}}