{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,9]],"date-time":"2026-01-09T12:23:23Z","timestamp":1767961403686,"version":"3.49.0"},"reference-count":27,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2005,3,14]],"date-time":"2005-03-14T00:00:00Z","timestamp":1110758400000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/2.0\/"},{"start":{"date-parts":[[2005,3,14]],"date-time":"2005-03-14T00:00:00Z","timestamp":1110758400000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/2.0\/"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                        <jats:title>Background<\/jats:title>\n                        <jats:p>Regions of interest identified through genetic linkage studies regularly exceed 30 centimorgans in size and can contain hundreds of genes. Traditionally this number is reduced by matching functional annotation to knowledge of the disease or phenotype in question. However, here we show that disease genes share patterns of sequence-based features that can provide a good basis for automatic prioritization of candidates by machine learning.<\/jats:p>\n                     <\/jats:sec><jats:sec>\n                        <jats:title>Results<\/jats:title>\n                        <jats:p>We examined a variety of sequence-based features and found that for many of them there are significant differences between the sets of genes known to be involved in human hereditary disease and those not known to be involved in disease. We have created an automatic classifier called PROSPECTR based on those features using the alternating decision tree algorithm which ranks genes in the order of likelihood of involvement in disease. On average, PROSPECTR enriches lists for disease genes two-fold 77% of the time, five-fold 37% of the time and twenty-fold 11% of the time.<\/jats:p>\n                     <\/jats:sec><jats:sec>\n                        <jats:title>Conclusion<\/jats:title>\n                        <jats:p>PROSPECTR is a simple and effective way to identify genes involved in Mendelian and oligogenic disorders. It performs markedly better than the single existing sequence-based classifier on novel data. PROSPECTR could save investigators looking at large regions of interest time and effort by prioritizing positional candidate genes for mutation detection and case-control association studies.<\/jats:p>\n                     <\/jats:sec>","DOI":"10.1186\/1471-2105-6-55","type":"journal-article","created":{"date-parts":[[2005,3,15]],"date-time":"2005-03-15T07:15:43Z","timestamp":1110870943000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":187,"title":["Speeding disease gene discovery by sequence based candidate prioritization"],"prefix":"10.1186","volume":"6","author":[{"given":"Euan A","family":"Adie","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Richard R","family":"Adams","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kathryn L","family":"Evans","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"David J","family":"Porteous","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ben S","family":"Pickard","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2005,3,14]]},"reference":[{"key":"380_CR1","doi-asserted-by":"publisher","first-page":"2345","DOI":"10.1126\/science.1076641","volume":"298","author":"AM Glazier","year":"2002","unstructured":"Glazier AM, Nadeau JH, Aitman TJ: Finding Genes That Underlie Complex Traits. Science 2002, 298: 2345\u20132349. 10.1126\/science.1076641","journal-title":"Science"},{"key":"380_CR2","doi-asserted-by":"publisher","first-page":"119","DOI":"10.1186\/gb-2003-4-10-119","volume":"4","author":"M McCarthy","year":"2003","unstructured":"McCarthy M, Smedley D, Hide W: New methods for finding disease-susceptibility genes: impact and potential. Genome Biology 2003, 4: 119. 10.1186\/gb-2003-4-10-119","journal-title":"Genome Biology"},{"key":"380_CR3","doi-asserted-by":"publisher","first-page":"429","DOI":"10.1016\/S0168-9525(01)02348-4","volume":"17","author":"D Devos","year":"2001","unstructured":"Devos D, Valencia A: Intrinsic errors in genome annotation. Trends in Genetics 2001, 17: 429\u2013431. 10.1016\/S0168-9525(01)02348-4","journal-title":"Trends in Genetics"},{"key":"380_CR4","doi-asserted-by":"publisher","first-page":"1641","DOI":"10.1093\/bioinformatics\/18.12.1641","volume":"18","author":"WR Gilks","year":"2002","unstructured":"Gilks WR, Audit B, De Angelis D, Tsoka S, Ouzounis CA: Modeling the percolation of annotation errors in a database of protein sequences. Bioinformatics 2002, 18: 1641\u20131649. 10.1093\/bioinformatics\/18.12.1641","journal-title":"Bioinformatics"},{"key":"380_CR5","doi-asserted-by":"publisher","first-page":"195","DOI":"10.1046\/j.1365-2958.1999.01561.x","volume":"34","author":"M Pallen","year":"1999","unstructured":"Pallen M, Wren B, Parkhill J: 'Going wrong with confidence': misleading sequence analyses of CiaB and ClpX. Molecular Microbiology 1999, 34: 195. 10.1046\/j.1365-2958.1999.01561.x","journal-title":"Molecular Microbiology"},{"key":"380_CR6","doi-asserted-by":"publisher","first-page":"57","DOI":"10.1038\/sj.ejhg.5200918","volume":"11","author":"MA Van Driel","year":"2003","unstructured":"Van Driel MA, Brunner HG, Leunissen JAM, Kemmeren PPCW, Cuelenaere K: A new web-based data mining tool for the identification of candidate genes for human genetic disorders. European Journal of Human Genetics 2003, 11: 57\u201363. 10.1038\/sj.ejhg.5200918","journal-title":"European Journal of Human Genetics"},{"key":"380_CR7","doi-asserted-by":"publisher","first-page":"110S","DOI":"10.1093\/bioinformatics\/18.suppl_2.S110","volume":"18","author":"J Freudenberg","year":"2002","unstructured":"Freudenberg J, Propping P: A similarity-based method for genome-wide prediction of disease-relevant human genes. Bioinformatics 2002, 18: 110S-1115.","journal-title":"Bioinformatics"},{"key":"380_CR8","doi-asserted-by":"crossref","first-page":"316","DOI":"10.1038\/ng895","volume":"31","author":"C Perez-Iratxeta","year":"2002","unstructured":"Perez-Iratxeta C, Bork P, Andrade MA: Association of genes to genetically inherited diseases using data mining. Nature Genetics 2002, 31: 316\u2013319.","journal-title":"Nature Genetics"},{"key":"380_CR9","volume-title":"Genome Biology","author":"FS Turner","year":"2003","unstructured":"Turner FS, Clutterbuck DR, Semple CAM: POCUS: mining genomic sequence annotation to predict disease genes. Genome Biology 2003., 4:"},{"key":"380_CR10","doi-asserted-by":"publisher","first-page":"315","DOI":"10.1093\/nar\/gkg046","volume":"31","author":"NJ Mulder","year":"2003","unstructured":"Mulder NJ, Apweiler R, Attwood TK, Bairoch A, Barrell D, Bateman A, Binns D, Biswas M, Bradley P, Bork P, et al.: The InterPro Database, 2003 brings increased coverage and new features. Nucl Acids Res 2003, 31: 315\u2013318. 10.1093\/nar\/gkg046","journal-title":"Nucl Acids Res"},{"key":"380_CR11","doi-asserted-by":"publisher","first-page":"169","DOI":"10.1016\/S0378-1119(03)00772-8","volume":"318","author":"NGC Smith","year":"2003","unstructured":"Smith NGC, Eyre-Walker A: Human disease genes: patterns and predictions. Gene 2003, 318: 169\u2013175. 10.1016\/S0378-1119(03)00772-8","journal-title":"Gene"},{"key":"380_CR12","doi-asserted-by":"publisher","first-page":"10","DOI":"10.1196\/annals.1310.003","volume":"1020","author":"IM Kapetanovic","year":"2004","unstructured":"Kapetanovic IM, Rosenfeld S, Izmirilan G: Overview of Commonly Used Bioinformatics Methods and Their Applications. Ann NY Acad Sci 2004, 1020: 10\u201321. 10.1196\/annals.1310.003","journal-title":"Ann NY Acad Sci"},{"key":"380_CR13","doi-asserted-by":"publisher","first-page":"3108","DOI":"10.1093\/nar\/gkh605","volume":"32","author":"N Lopez-Bigas","year":"2004","unstructured":"Lopez-Bigas N, Ouzounis CA: Genome-wide identification of genes likely to be involved in human genetic disease. Nucl Acids Res 2004, 32: 3108\u20133114. 10.1093\/nar\/gkh605","journal-title":"Nucl Acids Res"},{"key":"380_CR14","doi-asserted-by":"publisher","first-page":"268","DOI":"10.1016\/j.tig.2004.04.002","volume":"20","author":"MP Hammond","year":"2004","unstructured":"Hammond MP, Birney E: Genome information resources \u2013 developments at Ensembl. Trends in Genetics 2004, 20: 268\u2013272. 10.1016\/j.tig.2004.04.002","journal-title":"Trends in Genetics"},{"key":"380_CR15","doi-asserted-by":"publisher","first-page":"52","DOI":"10.1093\/nar\/30.1.52","volume":"30","author":"A Hamosh","year":"2002","unstructured":"Hamosh A, Scott AF, Amberger J, Bocchini C, Valle D, McKusick VA: Online Mendelian Inheritance in Man (OMIM), a knowledgebase of human genes and genetic disorders. Nucl Acids Res 2002, 30: 52\u201355. 10.1093\/nar\/30.1.52","journal-title":"Nucl Acids Res"},{"key":"380_CR16","doi-asserted-by":"publisher","first-page":"R47","DOI":"10.1186\/gb-2004-5-7-r47","volume":"5","author":"H Huang","year":"2004","unstructured":"Huang H, Winter E, Wang H, Weinstock K, Xing H, Goodstadt L, Stenson P, Cooper D, Smith D, Alba MM, et al.: Evolutionary conservation and selection of human disease gene orthologs in the rat and mouse genomes. Genome Biology 2004, 5: R47. 10.1186\/gb-2004-5-7-r47","journal-title":"Genome Biology"},{"key":"380_CR17","doi-asserted-by":"publisher","first-page":"54","DOI":"10.1101\/gr.1924004","volume":"14","author":"EE Winter","year":"2004","unstructured":"Winter EE, Goodstadt L, Ponting CP: Elevated Rates of Protein Secretion, Evolution, and Disease Among Tissue-Specific Genes. Genome Res 2004, 14: 54\u201361. 10.1101\/gr.1924004","journal-title":"Genome Res"},{"key":"380_CR18","doi-asserted-by":"publisher","first-page":"261","DOI":"10.1016\/0022-2836(87)90689-9","volume":"196","author":"M Gardiner-Garden","year":"1987","unstructured":"Gardiner-Garden M, Frommer M: CpG islands in vertebrate genomes. Journal of Molecular Biology 1987, 196: 261\u2013282. 10.1016\/0022-2836(87)90689-9","journal-title":"Journal of Molecular Biology"},{"key":"380_CR19","first-page":"261","volume-title":"Bioinformatics","author":"E Frank","year":"2004","unstructured":"Frank E, Hall M, Trigg L, Holmes G, Witten IH: Data mining in bioinformatics using Weka. Bioinformatics 2004, 261."},{"key":"380_CR20","unstructured":"Freund Y, Mason L: The Alternating Decision Tree Learning Algorithm. Proceedings of the Sixteenth International Conference on Machine Learning 124\u2013133."},{"key":"380_CR21","doi-asserted-by":"publisher","first-page":"577","DOI":"10.1002\/humu.10212","volume":"21","author":"PD Stenson","year":"2004","unstructured":"Stenson PD, Ball EV, Mort M, Philips AD, Shiel JA, Thomas NST, Abeysinghe S, Krawczak M, Cooper DN: Human Gene Mutation Database (HGMD\u00ae): 2003 update. Human Mutation 2004, 21: 577\u2013581. 10.1002\/humu.10212","journal-title":"Human Mutation"},{"key":"380_CR22","doi-asserted-by":"publisher","first-page":"431","DOI":"10.1038\/ng0504-431","volume":"36","author":"KG Becker","year":"2004","unstructured":"Becker KG, Barnes KC, Bright TJ, Wang SA: The Genetic Association Database. Nature Genetics 2004, 36: 431\u2013432. 10.1038\/ng0504-431","journal-title":"Nature Genetics"},{"key":"380_CR23","doi-asserted-by":"publisher","first-page":"189","DOI":"10.1007\/BF01617722","volume":"11","author":"AD Forbes","year":"1995","unstructured":"Forbes AD: Classification algorithm evaluation: five performance measures based on confusion matrices. Journal of Clinical Monitoring 1995, 11: 189\u2013206.","journal-title":"Journal of Clinical Monitoring"},{"key":"380_CR24","doi-asserted-by":"publisher","first-page":"146","DOI":"10.1128\/MCB.16.1.146","volume":"16","author":"RL Tanguay","year":"1996","unstructured":"Tanguay RL, Gallie DR: Translational efficiency is regulated by the length of the 3' untranslated region. Molecular Cellular Biology 1996, 16: 146\u2013156.","journal-title":"Molecular Cellular Biology"},{"key":"380_CR25","doi-asserted-by":"publisher","first-page":"2602","DOI":"10.1101\/gr.1169203","volume":"13","author":"F Chiaromonte","year":"2003","unstructured":"Chiaromonte F, Miller W, Eric E: Gene Length and Proximity to Neighbors Affect Genome-Wide Expression Levels. Genome Res 2003, 13: 2602\u20132608. 10.1101\/gr.1169203","journal-title":"Genome Res"},{"key":"380_CR26","doi-asserted-by":"publisher","first-page":"17008","DOI":"10.1073\/pnas.262658799","volume":"99","author":"S Karlin","year":"2002","unstructured":"Karlin S, Chen C, Gentles AJ, Cleary M: Associations between human disease genes and overlapping gene groups and multiple amino acid runs. PNAS 2002, 99: 17008\u201317013. 10.1073\/pnas.262658799","journal-title":"PNAS"},{"key":"380_CR27","doi-asserted-by":"publisher","first-page":"4465","DOI":"10.1073\/pnas.012025199","volume":"99","author":"AI Su","year":"2002","unstructured":"Su AI, Cooke MP, Ching KA, Hakak Y, Walker JR, Wiltshire T, Orth AP, Vega RG, Sapinoso LM, Moqrich A, et al.: Large-scale analysis of the human and mouse transcriptomes. PNAS 2002, 99: 4465\u20134470. 10.1073\/pnas.012025199","journal-title":"PNAS"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-6-55.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/1471-2105-6-55\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-6-55.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,7]],"date-time":"2024-10-07T12:07:51Z","timestamp":1728302871000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-6-55"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2005,3,14]]},"references-count":27,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2005,12]]}},"alternative-id":["380"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-6-55","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2005,3,14]]},"assertion":[{"value":"22 October 2004","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"14 March 2005","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"14 March 2005","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"55"}}