{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T20:55:32Z","timestamp":1783803332214,"version":"3.55.0"},"reference-count":30,"publisher":"Oxford University Press (OUP)","issue":"23","license":[{"start":{"date-parts":[[2016,8,9]],"date-time":"2016-08-09T00:00:00Z","timestamp":1470700800000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"funder":[{"name":"AstraZeneca"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2016,12,1]]},"abstract":"<jats:p>Motivation: Biomedical researchers often search through massive catalogues of literature to look for potential relationships between genes and diseases. Given the rapid growth of biomedical literature, automatic relation extraction, a crucial technology in biomedical literature mining, has shown great potential to support research of gene-related diseases. Existing work in this field has produced datasets that are limited both in scale and accuracy.<\/jats:p>\n               <jats:p>Results: In this study, we propose a reliable and efficient framework that takes large biomedical literature repositories as inputs, identifies credible relationships between diseases and genes, and presents possible genes related to a given disease and possible diseases related to a given gene. The framework incorporates name entity recognition (NER), which identifies occurrences of genes and diseases in texts, association detection whereby we extract and evaluate features from gene\u2013disease pairs, and ranking algorithms that estimate how closely the pairs are related. The F1-score of the NER phase is 0.87, which is higher than existing studies. The association detection phase takes drastically less time than previous work while maintaining a comparable F1-score of 0.86. The end-to-end result achieves a 0.259 F1-score for the top 50 genes associated with a disease, which performs better than previous work. In addition, we released a web service for public use of the dataset.<\/jats:p>\n               <jats:p>Availability and Implementation: The implementation of the proposed algorithms is publicly available at http:\/\/gdr-web.rwebox.com\/public_html\/index.php?page=download.php. The web service is available at http:\/\/gdr-web.rwebox.com\/public_html\/index.php.<\/jats:p>\n               <jats:p>Contact: \u00a0jenny.wei@astrazeneca.com or kzhu@cs.sjtu.edu.cn<\/jats:p>\n               <jats:p>Supplementary information: \u00a0Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btw503","type":"journal-article","created":{"date-parts":[[2016,8,10]],"date-time":"2016-08-10T00:52:41Z","timestamp":1470790361000},"page":"3619-3626","source":"Crossref","is-referenced-by-count":30,"title":["DTMiner: identification of potential disease targets through biomedical literature mining"],"prefix":"10.1093","volume":"32","author":[{"given":"Dong","family":"Xu","sequence":"first","affiliation":[{"name":"1Department of CSE, Shanghai Jiao Tong University, Shanghai 200240, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Meizhuo","family":"Zhang","sequence":"additional","affiliation":[{"name":"2R&D Information, Innovation Center China, AstraZeneca, Pudong, Shanghai 201203, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanping","family":"Xie","sequence":"additional","affiliation":[{"name":"1Department of CSE, Shanghai Jiao Tong University, Shanghai 200240, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fan","family":"Wang","sequence":"additional","affiliation":[{"name":"1Department of CSE, Shanghai Jiao Tong University, Shanghai 200240, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ming","family":"Chen","sequence":"additional","affiliation":[{"name":"2R&D Information, Innovation Center China, AstraZeneca, Pudong, Shanghai 201203, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kenny Q.","family":"Zhu","sequence":"additional","affiliation":[{"name":"1Department of CSE, Shanghai Jiao Tong University, Shanghai 200240, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jia","family":"Wei","sequence":"additional","affiliation":[{"name":"2R&D Information, Innovation Center China, AstraZeneca, Pudong, Shanghai 201203, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2016,8,9]]},"reference":[{"key":"2023020114070524700_btw503-B1","volume-title":"Enriching very large ontologies using the","author":"Agirre","year":"2000"},{"key":"2023020114070524700_btw503-B2","doi-asserted-by":"crossref","first-page":"431","DOI":"10.1038\/ng0504-431","article-title":"The genetic association database","volume":"36","author":"Becker","year":"2004","journal-title":"Nat. Genet"},{"key":"2023020114070524700_btw503-B3","doi-asserted-by":"crossref","first-page":"D267","DOI":"10.1093\/nar\/gkh061","article-title":"The Unified Medical Language System (UMLS): integrating biomedical terminology","volume":"32","author":"Bodenreider","year":"2004","journal-title":"Nucleic Acids Res"},{"key":"2023020114070524700_btw503-B4","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1186\/s12859-015-0472-9","article-title":"Extraction of relations between genes and diseases from text and large-scale data analysis: implications for translational research","volume":"16","author":"Bravo","year":"2015","journal-title":"BMC Bioinf"},{"key":"2023020114070524700_btw503-B5","doi-asserted-by":"crossref","first-page":"109","DOI":"10.2165\/00002018-199920020-00002","article-title":"The medical dictionary for regulatory activities (MedDRA)","volume":"20","author":"Brown","year":"1999","journal-title":"Drug Safety"},{"key":"2023020114070524700_btw503-B6","doi-asserted-by":"crossref","first-page":"D36","DOI":"10.1093\/nar\/gku1055","article-title":"Gene: a gene-centered information resource at NCBI","volume":"43","author":"Brown","year":"2015","journal-title":"Nucleic Acids Res"},{"key":"2023020114070524700_btw503-B7","first-page":"724","article-title":"Proceedings of the conference on human language technology and empirical methods in natural language processing","author":"Bunescu","year":"2005","journal-title":"Association for Computational Linguistics"},{"key":"2023020114070524700_btw503-B8","first-page":"27","article-title":"LIBSVM: A library for support vector machines","volume":"2","author":"Chang","year":"2011","journal-title":"ACM Trans. Intell. Syst. Technol. (TIST)"},{"key":"2023020114070524700_btw503-B9","doi-asserted-by":"crossref","first-page":"S5","DOI":"10.1186\/2041-1480-3-S3-S5","article-title":"Ranking relations between diseases, drugs and genes for a curation task","author":"Clematide","year":"2012","journal-title":"J. Biomed. Seman"},{"key":"2023020114070524700_btw503-B27","author":"Collins","year":"1999"},{"key":"2023020114070524700_btw503-B10","first-page":"363","volume-title":"Incorporating Non-Local Information into Information Extraction Systems by Gibbs Sampling","author":"Finkel","year":"2005"},{"key":"2023020114070524700_btw503-B11","doi-asserted-by":"crossref","first-page":"W406","DOI":"10.1093\/nar\/gkn215","article-title":"CoPub: a literature-based keyword enrichment tool for microarray data analysis","volume":"36","author":"Frijters","year":"2008","journal-title":"Nucleic Acids Res"},{"key":"2023020114070524700_btw503-B12","doi-asserted-by":"crossref","first-page":"1012","DOI":"10.1038\/nature07634","article-title":"Detecting influenza epidemics using search engine query data","volume":"457","author":"Ginsberg","year":"2009","journal-title":"Nature"},{"key":"2023020114070524700_btw503-B13","doi-asserted-by":"crossref","first-page":"D1079","DOI":"10.1093\/nar\/gku1071","article-title":"Genenames.org: the HGNC resources in 2015","volume":"43","author":"Gray","year":"2015","journal-title":"Nucleic Acids Res"},{"key":"2023020114070524700_btw503-B14","doi-asserted-by":"crossref","first-page":"S14","DOI":"10.1186\/1471-2105-6-S1-S14","article-title":"ProMiner: rule-based protein and gene entity recognition","volume":"6","author":"Hanisch","year":"2005","journal-title":"BMC Bioinf"},{"key":"2023020114070524700_btw503-B16","first-page":"1","author":"Ju","year":"2011","journal-title":"Bioinformatics and Biomedical Engineering,(iCBBE) 2011 5th International Conference on IEEE"},{"key":"2023020114070524700_btw503-B17","doi-asserted-by":"crossref","first-page":"270","DOI":"10.1016\/j.jbi.2015.01.003","article-title":"LGscore: A method to identify disease-related genes using biological literature and Google data","volume":"54","author":"Kim","year":"2015","journal-title":"J. Biomed. Inf"},{"key":"2023020114070524700_btw503-B18","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1186\/1471-2105-11-107","article-title":"Walk-weighted subsequence kernels for protein-protein interaction extraction","volume":"11","author":"Kim","year":"2010","journal-title":"BMC Bioinf"},{"key":"2023020114070524700_btw503-B19","first-page":"55","author":"Manning","year":"2014"},{"key":"2023020114070524700_btw503-B20","author":"Mitraka","year":"2015"},{"key":"2023020114070524700_btw503-B21","doi-asserted-by":"crossref","first-page":"i277","DOI":"10.1093\/bioinformatics\/btn182","article-title":"Identifying gene\u2013disease associations using centrality on a literature mined gene-interaction network","volume":"24","author":"Ozgur","year":"2008","journal-title":"Bioinformatics"},{"key":"2023020114070524700_btw503-B22","author":"Page","year":"1999"},{"key":"2023020114070524700_btw503-B23","first-page":"410","article-title":"Discovery and explanation of drug\u2013drug interactions via text mining. Pacific Symposium on Biocomputing","author":"Percha","year":"2012","journal-title":"Pac. Symp. Biocomput"},{"key":"2023020114070524700_btw503-B24","doi-asserted-by":"crossref","first-page":"83","DOI":"10.1016\/j.ymeth.2014.11.020","article-title":"DISEASES: text mining and data integration of disease-gene associations","volume":"74","author":"Pletscher-Frankild","year":"2015","journal-title":"Methods"},{"key":"2023020114070524700_btw503-B25","doi-asserted-by":"crossref","first-page":"789","DOI":"10.1016\/j.jbi.2011.04.005","article-title":"Using a shallow linguistic kernel for drug\u2013drug interaction extraction","volume":"44","author":"Segura-Bedmar","year":"2011","journal-title":"J. Biomed. Inf"},{"key":"2023020114070524700_btw503-B26","first-page":"104","article-title":"Proceedings of the international joint workshop on natural language processing in biomedicine and its applications","author":"Settles","year":"2004","journal-title":"Association for Computational Linguistics"},{"key":"2023020114070524700_btw503-B28","doi-asserted-by":"crossref","first-page":"D204","DOI":"10.1093\/nar\/gku989","article-title":"UniProt: a hub for protein information","volume":"43","author":"Uniprot Consortium","year":"2015","journal-title":"Nucleic Acids Res"},{"key":"2023020114070524700_btw503-B29","doi-asserted-by":"crossref","first-page":"827","DOI":"10.1016\/j.jbi.2012.04.011","article-title":"A knowledge-driven conditional approach to extract pharmacogenomics specific drug-gene relationships from free text","volume":"45","author":"Xu","year":"2012","journal-title":"J. Biomed. Inf"},{"key":"2023020114070524700_btw503-B30","doi-asserted-by":"crossref","first-page":"42","DOI":"10.1145\/312624.312647","volume-title":"SIGIR'99","author":"Yang","year":"1999"},{"key":"2023020114070524700_btw503-B31","first-page":"1083","article-title":"Kernel methods for relation extraction","volume":"3","author":"Zelenko","year":"2003","journal-title":"J. Mach. Learn. Res"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/32\/23\/3619\/49026963\/bioinformatics_32_23_3619.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/32\/23\/3619\/49026963\/bioinformatics_32_23_3619.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,1]],"date-time":"2023-02-01T23:57:04Z","timestamp":1675295824000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/32\/23\/3619\/2525620"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,8,9]]},"references-count":30,"journal-issue":{"issue":"23","published-print":{"date-parts":[[2016,12,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btw503","relation":{},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2016,12,1]]},"published":{"date-parts":[[2016,8,9]]}}}