{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T22:11:45Z","timestamp":1780611105279,"version":"3.54.1"},"reference-count":29,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2023,9,20]],"date-time":"2023-09-20T00:00:00Z","timestamp":1695168000000},"content-version":"vor","delay-in-days":262,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,9,19]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>The recognition of dataset names is a critical task for automatic information extraction in scientific literature, enabling researchers to understand and identify research opportunities. However, existing corpora for dataset mention detection are limited in size and naming diversity. In this paper, we introduce the Dataset Mentions Detection Dataset (DMDD), the largest publicly available corpus for this task. DMDD consists of the DMDD main corpus, comprising 31,219 scientific articles with over 449,000 dataset mentions weakly annotated in the format of in-text spans, and an evaluation set, which comprises 450 scientific articles manually annotated for evaluation purposes. We use DMDD to establish baseline performance for dataset mention detection and linking. By analyzing the performance of various models on DMDD, we are able to identify open problems in dataset mention detection. We invite the community to use our dataset as a challenge to develop novel dataset mention detection models.<\/jats:p>","DOI":"10.1162\/tacl_a_00592","type":"journal-article","created":{"date-parts":[[2023,9,20]],"date-time":"2023-09-20T15:03:35Z","timestamp":1695222215000},"page":"1132-1146","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":9,"title":["DMDD: A Large-Scale Dataset for Dataset Mentions Detection"],"prefix":"10.1162","volume":"11","author":[{"given":"Huitong","family":"Pan","sequence":"first","affiliation":[{"name":"Temple University, Philadelphia, Pennsylvania, USA. huitong.pan@temple.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qi","family":"Zhang","sequence":"additional","affiliation":[{"name":"Temple University, Philadelphia, Pennsylvania, USA. qi.zhang@temple.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Eduard","family":"Dragut","sequence":"additional","affiliation":[{"name":"Temple University, Philadelphia, Pennsylvania, USA. edragut@temple.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Cornelia","family":"Caragea","sequence":"additional","affiliation":[{"name":"University of Illinois Chicago, Chicago, Illinois, USA. cornelia@uic.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Longin Jan","family":"Latecki","sequence":"additional","affiliation":[{"name":"Temple University, Philadelphia, Pennsylvania, USA. latecki@temple.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2023,9,19]]},"reference":[{"key":"2023092014493566900_bib1","article-title":"The ACE 2005 (ACE 05) evaluation plan evaluation of the detection and recognition of ace entities, values, temporal expressions, relations, and events 1","author":"ACE","year":"2005"},{"key":"2023092014493566900_bib2","doi-asserted-by":"publisher","first-page":"718","DOI":"10.18653\/v1\/P17-1067","article-title":"EmoNet: Fine-grained emotion detection with gated recurrent neural networks","volume-title":"ACL","author":"Abdul-Mageed","year":"2017"},{"key":"2023092014493566900_bib3","doi-asserted-by":"publisher","first-page":"546","DOI":"10.18653\/v1\/S17-2091","article-title":"SemEval 2017 task 10: ScienceIE - extracting keyphrases and relations from scientific publications","volume-title":"SemEval","author":"Augenstein","year":"2017"},{"key":"2023092014493566900_bib4","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1371","article-title":"Scibert: A pretrained language model for scientific text","volume-title":"EMNLP","author":"Iz","year":"2019"},{"key":"2023092014493566900_bib5","article-title":"Longformer: The long-document transformer","author":"Iz","year":"2020","journal-title":"arXiv preprint arXiv:2004.05150"},{"key":"2023092014493566900_bib6","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","author":"Devlin","year":"2018","journal-title":"CoRR"},{"issue":"1","key":"2023092014493566900_bib7","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-14-194","article-title":"bioNerDS: Exploring bioinformatics\u2019 database and software use through literature mining","volume":"14","author":"Duck","year":"2013","journal-title":"BMC Bioinformatics"},{"key":"2023092014493566900_bib8","article-title":"Identifying used methods and datasets in scientific publications.","volume-title":"SDU@ AAAI","author":"F\u00e4rber","year":"2021"},{"key":"2023092014493566900_bib9","doi-asserted-by":"publisher","first-page":"679","DOI":"10.18653\/v1\/S18-1111","article-title":"SemEval-2018 task 7: Semantic relation extraction and classification in scientific papers","volume-title":"Proceedings of The 12th International Workshop on Semantic Evaluation","author":"G\u00e1bor","year":"2018"},{"issue":"8","key":"2023092014493566900_bib10","doi-asserted-by":"publisher","DOI":"10.3390\/data6080084","article-title":"The automatic detection of dataset names in scientific articles","volume":"6","author":"Heddes","year":"2021","journal-title":"Data"},{"key":"2023092014493566900_bib11","doi-asserted-by":"publisher","first-page":"5203","DOI":"10.18653\/v1\/P19-1513","article-title":"Identification of tasks, datasets, evaluation metrics, and numeric scores for scientific leaderboards construction","volume-title":"ACL","author":"Hou","year":"2019"},{"key":"2023092014493566900_bib12","doi-asserted-by":"publisher","first-page":"707","DOI":"10.18653\/v1\/2021.eacl-main.59","article-title":"TDMSci: A specialized corpus for scientific literature entity tagging of tasks datasets and metrics","volume-title":"ACL","author":"Hou","year":"2021"},{"key":"2023092014493566900_bib13","doi-asserted-by":"publisher","first-page":"7506","DOI":"10.18653\/v1\/2020.acl-main.670","article-title":"SciREX: A challenge dataset for document-level information extraction","volume-title":"ACL","author":"Jain","year":"2020"},{"key":"2023092014493566900_bib14","first-page":"5203","article-title":"Rich context competition: Extracting research context and dataset usage information from scientific publications","volume-title":"ACL","author":"Kim","year":"2019"},{"key":"2023092014493566900_bib15","doi-asserted-by":"publisher","first-page":"2356","DOI":"10.1145\/3404835.3463238","article-title":"Pyserini: A Python toolkit for reproducible information retrieval research with sparse and dense representations","volume-title":"Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Lin","year":"2021"},{"key":"2023092014493566900_bib16","first-page":"4969","article-title":"S2ORC: The semantic scholar open research corpus","volume-title":"ACL","author":"Lo","year":"2020"},{"key":"2023092014493566900_bib17","article-title":"Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction","volume-title":"EMNLP","author":"Yi","year":"2018"},{"issue":"2","key":"2023092014493566900_bib18","doi-asserted-by":"publisher","first-page":"313","DOI":"10.21236\/ADA273556","article-title":"Building a large annotated corpus of english: The Penn treebank","volume":"19","author":"Marcus","year":"1993","journal-title":"Computational Linguistics"},{"key":"2023092014493566900_bib19","article-title":"Efficient estimation of word representations in vector space","author":"Mikolov","year":"2013"},{"key":"2023092014493566900_bib20","doi-asserted-by":"publisher","first-page":"1003","DOI":"10.3115\/1690219.1690287","article-title":"Distant supervision for relation extraction without labeled data","volume-title":"Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP","author":"Mintz","year":"2009"},{"key":"2023092014493566900_bib21","doi-asserted-by":"publisher","first-page":"23","DOI":"10.1080\/10489223.2012.738742","article-title":"Syntactic islands and learning biases: Combining experimental syntax and computational modeling to investigate the language acquisition problem","volume":"20","author":"Pearl","year":"2013","journal-title":"Language Acquisition"},{"key":"2023092014493566900_bib22","doi-asserted-by":"publisher","first-page":"1532","DOI":"10.3115\/v1\/D14-1162","article-title":"GloVe: Global vectors for word representation","volume-title":"EMNLP","author":"Pennington","year":"2014"},{"key":"2023092014493566900_bib23","doi-asserted-by":"publisher","first-page":"2227","DOI":"10.18653\/v1\/N18-1202","article-title":"Deep contextualized word representations","volume-title":"ACL","author":"Peters","year":"2018"},{"key":"2023092014493566900_bib24","article-title":"Data programming: Creating large training sets, quickly","volume-title":"Advances in Neural Information Processing Systems","author":"Ratner","year":"2016"},{"key":"2023092014493566900_bib25","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.272","article-title":"Colbertv2: Effective and efficient retrieval via lightweight late interaction","author":"Santhanam","year":"2021","journal-title":"arXiv preprint arXiv:2112.01488"},{"key":"2023092014493566900_bib26","article-title":"brat: A web-based tool for NLP-assisted text annotation","volume-title":"Proceedings of the Demonstrations Session at EACL 2012","author":"Stenetorp","year":"2012"},{"issue":"7","key":"2023092014493566900_bib27","doi-asserted-by":"publisher","first-page":"e0216913","DOI":"10.1371\/journal.pone.0216913","article-title":"Using distant supervision to augment manually annotated data for relation extraction","volume":"14","author":"Peng","year":"2019","journal-title":"PLOS ONE"},{"key":"2023092014493566900_bib28","doi-asserted-by":"publisher","DOI":"10.1109\/BigData47090.2019.9006262","article-title":"Method and dataset mining in scientific papers","author":"Yao","year":"2019","journal-title":"arXiv e-prints"},{"key":"2023092014493566900_bib29","doi-asserted-by":"publisher","first-page":"5206","DOI":"10.18653\/v1\/D19-1524","article-title":"A context-based framework for modeling the role and function of on-line resource citations in scientific literature","volume-title":"EMNLP","author":"He","year":"2019"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00592\/2159087\/tacl_a_00592.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00592\/2159087\/tacl_a_00592.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,20]],"date-time":"2023-09-20T15:03:43Z","timestamp":1695222223000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00592\/117582\/DMDD-A-Large-Scale-Dataset-for-Dataset-Mentions"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023]]},"references-count":29,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00592","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2023]]},"published":{"date-parts":[[2023]]}}}