{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,12]],"date-time":"2025-09-12T18:45:28Z","timestamp":1757702728055},"reference-count":34,"publisher":"Springer Science and Business Media LLC","issue":"1","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2006,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:sec>\n            <jats:title>Background<\/jats:title>\n            <jats:p>Accuracy of document retrieval from MEDLINE for gene queries is crucially important for many applications in bioinformatics. We explore five information retrieval-based methods to rank documents retrieved by PubMed gene queries for the human genome. The aim is to rank relevant documents higher in the retrieved list. We address the special challenges faced due to ambiguity in gene nomenclature: gene terms that refer to multiple genes, gene terms that are also English words, and gene terms that have other biological meanings.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Results<\/jats:title>\n            <jats:p>Our two baseline ranking strategies are quite similar in performance. Two of our three LocusLink-based strategies offer significant improvements. These methods work very well even when there is ambiguity in the gene terms. Our best ranking strategy offers significant improvements on three different kinds of ambiguities over our two baseline strategies (improvements range from 15.9% to 17.7% and 11.7% to 13.3% depending on the baseline). For most genes the best ranking query is one that is built from the LocusLink (now Entrez Gene) summary and product information along with the gene names and aliases. For others, the gene names and aliases suffice. We also present an approach that successfully predicts, for a given gene, which of these two ranking queries is more appropriate.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Conclusion<\/jats:title>\n            <jats:p>We explore the effect of different post-retrieval strategies on the ranking of documents returned by PubMed for human gene queries. We have successfully applied some of these strategies to improve the ranking of relevant documents in the retrieved sets. This holds true even when various kinds of ambiguity are encountered. We feel that it would be very useful to apply strategies like ours on PubMed search results as these are not ordered by relevance in any way. This is especially so for queries that retrieve a large number of documents.<\/jats:p>\n          <\/jats:sec>","DOI":"10.1186\/1471-2105-7-220","type":"journal-article","created":{"date-parts":[[2006,4,21]],"date-time":"2006-04-21T18:21:30Z","timestamp":1145643690000},"update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["Retrieval with gene queries"],"prefix":"10.1186","volume":"7","author":[{"given":"Aditya K","family":"Sehgal","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Padmini","family":"Srinivasan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2006,4,21]]},"reference":[{"key":"959_CR1","doi-asserted-by":"publisher","first-page":"109","DOI":"10.1109\/CSB.2002.1039334","volume-title":"Proceedings of the 1st IEEE Computer Society Bioinformatics Conference","author":"LA Adamic","year":"2002","unstructured":"Adamic LA, Wilkinson D, Huberman BA, Adar E: A literature based method for identifying gene-disease connections. Proceedings of the 1st IEEE Computer Society Bioinformatics Conference 2002, 109\u2013117."},{"key":"959_CR2","first-page":"517","volume-title":"Proceedings of the Pacific Symposium on Biocomputing (PSB)","author":"TC Rindflesch","year":"2000","unstructured":"Rindflesch TC, Tanabe L, Weinstein JN, Hunter L: EDGAR: Extraction of drugs, genes, and relations from biomedical literature. Proceedings of the Pacific Symposium on Biocomputing (PSB) 2000, 517\u2013528."},{"key":"959_CR3","first-page":"317","volume-title":"Proceedings of the Eighth International Conference on Intelligent Systems for Molecular Biology (ISMB)","author":"H Shatkay","year":"2000","unstructured":"Shatkay H, Edwards S, Wilbur WJ, Boguski M: Genes, Themes, and Microarrays: Using Information Retrieval for Large-Scale Gene Analysis. Proceedings of the Eighth International Conference on Intelligent Systems for Molecular Biology (ISMB) 2000, 317\u2013328."},{"issue":"3","key":"959_CR4","doi-asserted-by":"publisher","first-page":"396","DOI":"10.1093\/bioinformatics\/btg002","volume":"19","author":"S Raychaudhuri","year":"2003","unstructured":"Raychaudhuri S, Altman RB: A literature-based method for assessing the functional coherence of a gene group. Bioinformatics 2003, 19(3):396\u2013401.","journal-title":"Bioinformatics"},{"key":"959_CR5","first-page":"548","volume-title":"Proceedings of the 2nd SIAM International Conference on Data Mining","author":"P Kankar","year":"2002","unstructured":"Kankar P, Adak S, Sarkar A, Murari K, Sharma G: MedMesh Summarizer: Text Mining for Gene Clusters. Proceedings of the 2nd SIAM International Conference on Data Mining 2002, 548\u2013565."},{"issue":"2","key":"959_CR6","doi-asserted-by":"publisher","first-page":"191","DOI":"10.1093\/bioinformatics\/btg390","volume":"20","author":"JD Wren","year":"2004","unstructured":"Wren JD, Garner HR: Shared relationship analysis: ranking set cohesion and commonalities within a literature-derived relationship network. Bioinformatics 2004, 20(2):191\u2013198.","journal-title":"Bioinformatics"},{"issue":"10","key":"959_CR7","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/gb-2002-3-10-research0055","volume":"3","author":"D Chaussabel","year":"2002","unstructured":"Chaussabel D, Sher A: Mining microarray expression data by literature profiling. Genome Biol 2002, 3(10):1\u20130055.","journal-title":"Genome Biol"},{"issue":"4","key":"959_CR8","doi-asserted-by":"publisher","first-page":"247","DOI":"10.1016\/S1532-0464(03)00014-5","volume":"35","author":"L Hirschman","year":"2002","unstructured":"Hirschman L, Morgan AA, Yeh AS: Rutabaga by any other name: extracting biological names. J Biomed Inform 2002, 35(4):247\u2013259.","journal-title":"J Biomed Inform"},{"key":"959_CR9","doi-asserted-by":"publisher","first-page":"9","DOI":"10.3115\/1118149.1118151","volume-title":"Proceedings of the Workshop on Natural Language Processing in the Biomedical Domain","author":"LK Tanabe","year":"2002","unstructured":"Tanabe LK, Wilbur WJ: Tagging gene and protein names in full text articles. Proceedings of the Workshop on Natural Language Processing in the Biomedical Domain 2002, 9\u201313."},{"key":"959_CR10","doi-asserted-by":"publisher","first-page":"1","DOI":"10.3115\/1118958.1118959","volume-title":"Proceedings of the ACL 2003 Workshop on Natural Language Processing in Biomedicine","author":"A Morgan","year":"2003","unstructured":"Morgan A, Hirschman L, Yeh A, Colosimo M: Gene Name Extraction Using FlyBase Resources. Proceedings of the ACL 2003 Workshop on Natural Language Processing in Biomedicine 2003, 1\u20138."},{"key":"959_CR11","first-page":"704","volume-title":"Proceedings of the AMIA Symposium","author":"M Weeber","year":"2003","unstructured":"Weeber M, Schijvenaars BJA, van Mulligen EM, Mons B, Jelier R, van der Eijk C, Kors JA: Ambiguity of Human Gene Symbols in LocusLink and MEDLINE: Creating an Inventory and a Disambiguation Test Collection. Proceedings of the AMIA Symposium 2003, 704\u2013708."},{"key":"959_CR12","first-page":"238","volume-title":"Proceedings of the Pacific Symposium on Biocomputing (PSB)","author":"O Tuason","year":"2004","unstructured":"Tuason O, Chen L, Liu H, Blake JA, Friedman C: Biological Nomenclatures: A Source of Lexical Knowledge and Ambiguity. Proceedings of the Pacific Symposium on Biocomputing (PSB) 2004, 238\u2013249."},{"issue":"2","key":"959_CR13","doi-asserted-by":"publisher","first-page":"248","DOI":"10.1093\/bioinformatics\/bth496","volume":"21","author":"L Chen","year":"2005","unstructured":"Chen L, Liu H, Friedman C: Gene Name Ambiguity of Eukaryotic Nomenclatures. Bioinformatics 2005, 21(2):248\u2013256.","journal-title":"Bioinformatics"},{"issue":"4","key":"959_CR14","doi-asserted-by":"publisher","first-page":"249","DOI":"10.1006\/jbin.2001.1023","volume":"34","author":"H Liu","year":"2001","unstructured":"Liu H, Lussier YA, Friedman C: Disambiguating ambiguous biomedical terms in bio medical narrative text: an unsupervised method. Journal of Biomedical Informatics 2001, 34(4):249\u2013261.","journal-title":"Journal of Biomedical Informatics"},{"issue":"3","key":"959_CR15","doi-asserted-by":"publisher","first-page":"743","DOI":"10.1142\/S0219720005001223","volume":"3","author":"RM Podowski","year":"2005","unstructured":"Podowski RM, Cleary JG, Goncharoff NT, Amoutzias G, Hayes WS: Suregene, a scalable system for automated term disambiguation of gene and protein names. Journal of Bioinformatics and Computational Biology 2005, 3(3):743\u2013770.","journal-title":"Journal of Bioinformatics and Computational Biology"},{"key":"959_CR16","first-page":"9","volume-title":"Proceedings of the HLT-NAACL 2004 Workshop: BioLINK Linking Biological Literature, Ontologies and Databases","author":"A Koike","year":"2004","unstructured":"Koike A, Takagi T: Gene\/Protein\/Family Name Recognition in Biomedical Literature. Proceedings of the HLT-NAACL 2004 Workshop: BioLINK Linking Biological Literature, Ontologies and Databases 2004, 9\u201316."},{"key":"959_CR17","first-page":"251","volume-title":"Proceedings of the 2nd IEEE Computer Society Bioinformatics Conference","author":"K Seki","year":"2003","unstructured":"Seki K, Mostafa J: A Probabilistic Model for Identifying Protein Names and their Name Boundaries. Proceedings of the 2nd IEEE Computer Society Bioinformatics Conference 2003, 251\u2013259."},{"key":"959_CR18","doi-asserted-by":"publisher","first-page":"149","DOI":"10.1186\/1471-2105-6-149","volume":"6","author":"B1 Schijvenaars","year":"2005","unstructured":"Schijvenaars B1, Mons B, Weeber M, Schuemie MJ, van Mulligen EM, Wain HM, Kors JA: Thesaurus-based disambiguation of gene symbols. BMC Bioinformatics 2005, 6: 149.","journal-title":"BMC Bioinformatics"},{"key":"959_CR19","unstructured":"KDD Cup 2002[http:\/\/www.biostat.wisc.edu\/~craven\/kddcup\/]"},{"issue":"Suppl 1","key":"959_CR20","doi-asserted-by":"publisher","first-page":"S1","DOI":"10.1186\/1471-2105-6-S1-S1","volume":"6","author":"L Hirschman","year":"2005","unstructured":"Hirschman L, Yeh A, Blaschke C, A V: Overview of BioCreAtIvE: critical assessment of information extraction for biology. BMC Bioinformatics 2005, 6(Suppl 1):S1.","journal-title":"BMC Bioinformatics"},{"key":"959_CR21","unstructured":"TREC Genomics Track[http:\/\/ir.ohsu.edu\/genomics\/]"},{"issue":"Suppl 1","key":"959_CR22","doi-asserted-by":"publisher","first-page":"S16","DOI":"10.1186\/1471-2105-6-S1-S16","volume":"6","author":"C Blaschke","year":"2005","unstructured":"Blaschke C, Leon EA, Krallinger M, Valencia A: Evaluation of BioCreAtIvE assessment of task 2. BMC Bioinformatics 2005, 6(Suppl 1):S16.","journal-title":"BMC Bioinformatics"},{"key":"959_CR23","first-page":"14","volume-title":"Proceedings of The 12th Text Retrieval Conference (TREC)","author":"W Hersh","year":"2003","unstructured":"Hersh W, Bhupatiraju RT: TREC Genomics Track Overview. Proceedings of The 12th Text Retrieval Conference (TREC) 2003, 14\u201323."},{"key":"959_CR24","first-page":"13","volume-title":"Proceedings of The 13th Text Retrieval Conference (TREC)","author":"W Hersh","year":"2004","unstructured":"Hersh W, Bhupatiraju RT, Ross L, Johnson P, Cohen AM, Kraemer DF: TREC 2004 Genomics Track Overview. Proceedings of The 13th Text Retrieval Conference (TREC) 2004, 13\u201331."},{"key":"959_CR25","first-page":"25","volume-title":"Proceedings of the 20th ACM SIGIR Conference","author":"A Singhal","year":"1997","unstructured":"Singhal A, Mitra M, Buckley C: Learning routing queries in a query zone. Proceedings of the 20th ACM SIGIR Conference 1997, 25\u201332."},{"key":"959_CR26","volume-title":"The NCBI Handbook, NCBI","author":"D Maglott","year":"2003","unstructured":"Maglott D: LocusLink: A Directory of Genes. The NCBI Handbook, NCBI 2003."},{"key":"959_CR27","unstructured":"WordNet \u2013 Princeton University Cognitive Science Laboratory[http:\/\/wordnet.princeton.edu]"},{"issue":"6","key":"959_CR28","doi-asserted-by":"publisher","first-page":"612","DOI":"10.1197\/jamia.M1139","volume":"9","author":"JT Chang","year":"2002","unstructured":"Chang JT, Sch\u00fctze H, Altman RB: Creating an Online Dictionary of Abbreviations from MEDLINE. J Am Med Inform Assoc 2002, 9(6):612\u2013620.","journal-title":"J Am Med Inform Assoc"},{"key":"959_CR29","first-page":"371","volume-title":"Proceedings of Medinfo","author":"J Pustejovsky","year":"2001","unstructured":"Pustejovsky J, Castano J, Cochran B, Kotechi M, Morrell M: Automatic extraction of acronym-meaning pairs from MEDLINE databases. Proceedings of Medinfo 2001, 371\u2013375."},{"key":"959_CR30","first-page":"451","volume-title":"Proceedings of the Pacific Symposium on Biocomputing (PSB)","author":"AS Schwartz","year":"2003","unstructured":"Schwartz AS, Hearst MA: A Simple Algorithm for Identifying Abbreviation Definitions in Biomedical Text. Proceedings of the Pacific Symposium on Biocomputing (PSB) 2003, 451\u2013462."},{"key":"959_CR31","unstructured":"Retrieval for Gene Queries[http:\/\/sulu.info-science.uiowa.edu\/genedocs\/]"},{"key":"959_CR32","first-page":"299","volume-title":"Proceedings of the 25th ACM SIGIR Conference","author":"S Cronen-Townsend","year":"2002","unstructured":"Cronen-Townsend S, Zhou Y, Croft WB: Predicting query performance. Proceedings of the 25th ACM SIGIR Conference 2002, 299\u2013306."},{"key":"959_CR33","unstructured":"ELink Entrez Utility[http:\/\/eutils.ncbi.nlm.nih.gov\/entrez\/query\/static\/elink_help.html]"},{"key":"959_CR34","unstructured":"Lemur Project[http:\/\/www-2.cs.cmu.edu\/~lemur\/]"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-7-220.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,1]],"date-time":"2021-09-01T10:59:07Z","timestamp":1630493947000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-7-220"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2006,4,21]]},"references-count":34,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2006,12]]}},"alternative-id":["959"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-7-220","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2006,4,21]]},"assertion":[{"value":"12 August 2005","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 April 2006","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 April 2006","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"220"}}