{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,22]],"date-time":"2026-01-22T06:05:13Z","timestamp":1769061913948,"version":"3.49.0"},"reference-count":37,"publisher":"Oxford University Press (OUP)","issue":"13","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2007,7,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Background: Identifying information that implicitly links two disparate sets of articles is a fundamental and intuitive data mining strategy that can help investigators address real scientific questions. The Arrowsmith two-node search finds title words and phrases (so-called B-terms) that are shared across two sets of articles within MEDLINE and displays them in a manner that facilitates human assessment. A serious stumbling-block has been the lack of a quantitative model for predicting which of the hundreds if not thousands of B-terms computed for a given search are most likely to be relevant to the investigator.<\/jats:p><jats:p>Methodology\/Principal Findings: Using a public two-node search interface, field testers devised a set of two-node searches under real life conditions and a certain number of B-terms were marked relevant. These were employed as \u2018gold standards;\u2019 each B-term was characterized according to eight complementary features that were strongly correlated with relevance. A logistic regression model was developed that permits one to estimate the probability of relevance for each B-term, to rank B-terms according to their likely relevance, and to estimate the overall number of relevant B-terms inherent in a given two-node search.<\/jats:p><jats:p>Conclusions\/Significance: The model greatly simplifies and streamlines the process of carrying out a two-node search, and may be applicable to a number of other literature-based discovery applications, including the so-called one-node search and related gene-centric strategies that incorporate implicit links to predict how genes may be related to each other and to human diseases. This should encourage much wider exploration of text mining for implicit information among the general scientific community.<\/jats:p><jats:p>Availability: Two-node searches can be carried out freely at http:\/\/arrowsmith.psych.uic.edu<\/jats:p><jats:p>Contact: neils@uic.edu, vtorvik@uic.edu<\/jats:p><jats:p>Supplementary information: Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btm161","type":"journal-article","created":{"date-parts":[[2007,4,27]],"date-time":"2007-04-27T00:29:34Z","timestamp":1177633774000},"page":"1658-1665","source":"Crossref","is-referenced-by-count":49,"title":["A quantitative model for linking two disparate sets of articles in MEDLINE"],"prefix":"10.1093","volume":"23","author":[{"given":"Vetle I.","family":"Torvik","sequence":"first","affiliation":[{"name":"Department of Psychiatry and Psychiatric Institute (MC912), University of Illinois-Chicago, Chicago, IL 60612, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Neil R.","family":"Smalheiser","sequence":"additional","affiliation":[{"name":"Department of Psychiatry and Psychiatric Institute (MC912), University of Illinois-Chicago, Chicago, IL 60612, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2007,4,26]]},"reference":[{"key":"2023062708443020100_B1","doi-asserted-by":"crossref","DOI":"10.1002\/0471249688","volume-title":"Categorical Data Analysis","author":"Agresti","year":"2002","edition":"2nd"},{"key":"2023062708443020100_B2","first-page":"17","article-title":"Effective mapping of biomedical text to the UMLS Metathesaurus: the MetaMap program","author":"Aronson","year":"2001","journal-title":"Proc. AMIA Symp"},{"key":"2023062708443020100_B3","first-page":"671","article-title":"Extracting noun phrases for all of MEDLINE","author":"Bennett","year":"1999","journal-title":"Proc. AMIA Symp"},{"key":"2023062708443020100_B4","doi-asserted-by":"crossref","first-page":"D267","DOI":"10.1093\/nar\/gkh061","article-title":"The Unified Medical Language System (UMLS): integrating biomedical terminology","volume":"32","author":"Bodenreider","year":"2004","journal-title":"Nucleic Acids Res"},{"key":"2023062708443020100_B5","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1385\/NI:1:1:043","article-title":"NeuroNames 2002","volume":"1","author":"Bowden","year":"2003","journal-title":"Neuroinformatics"},{"key":"2023062708443020100_B6","doi-asserted-by":"crossref","first-page":"147","DOI":"10.1186\/1471-2105-5-147","article-title":"Content-rich biological network constructed by mining PubMed abstracts","volume":"5","author":"Chen","year":"2004","journal-title":"BMC Bioinformatics"},{"key":"2023062708443020100_B7","doi-asserted-by":"crossref","first-page":"57","DOI":"10.1093\/bib\/6.1.57","article-title":"A survey of current work in biomedical text mining","volume":"6","author":"Cohen","year":"2005","journal-title":"Brief Bioinform"},{"key":"2023062708443020100_B8","doi-asserted-by":"crossref","first-page":"25","DOI":"10.6028\/NIST.SP.500-266.genomics-overview","article-title":"TREC 2005 Genomics track overview","volume-title":"NIST Special Publication 500-266: The Fourteenth Text Retrieval Conference (TREC 2005)","author":"Hersh","year":"2005"},{"key":"2023062708443020100_B9","first-page":"3","article-title":"Enhancing access to the Bibliome: the TREC 2004 Genomics Track","volume-title":"J. Biomed. Discovery Collaboration","author":"Hersh","year":"2006"},{"key":"2023062708443020100_B10","doi-asserted-by":"crossref","first-page":"104","DOI":"10.1093\/bioinformatics\/bth464","article-title":"Gene clustering by latent semantic indexing of MEDLINE abstracts","volume":"21","author":"Homayouni","year":"2005","journal-title":"Bioinformatics"},{"key":"2023062708443020100_B11","doi-asserted-by":"crossref","DOI":"10.1002\/0471722146","volume-title":"Applied Logistic Regression","author":"Hosmer","year":"2000","edition":"2nd"},{"key":"2023062708443020100_B12","doi-asserted-by":"crossref","first-page":"289","DOI":"10.1016\/j.ijmedinf.2004.04.024","article-title":"Using literature-based discovery to identify disease candidate genes","volume":"74","author":"Hristovski","year":"2005","journal-title":"Int. J. Med. Inform"},{"key":"2023062708443020100_B13","doi-asserted-by":"crossref","first-page":"589","DOI":"10.1016\/j.molcel.2006.02.012","article-title":"Biomedical language processing: what's beyond PubMed?","volume":"21","author":"Hunter","year":"2006","journal-title":"Mol. Cell"},{"key":"2023062708443020100_B14","doi-asserted-by":"crossref","first-page":"119","DOI":"10.1038\/nrg1768","article-title":"Literature mining for the biologist: from information retrieval to biological discovery","volume":"7","author":"Jensen","year":"2006","journal-title":"Nat. Rev. Genet"},{"key":"2023062708443020100_B15","doi-asserted-by":"crossref","first-page":"21","DOI":"10.1038\/ng0501-21","article-title":"A literature network of human genes for high-throughput analysis of gene expression","volume":"28","author":"Jenssen","year":"2001","journal-title":"Nat. Genet"},{"key":"2023062708443020100_B16","doi-asserted-by":"crossref","first-page":"515","DOI":"10.1016\/j.ijmedinf.2004.02.008","article-title":"Information content in Medline record fields","volume":"73","author":"Kostoff","year":"2004","journal-title":"Int. J. Med. Inform"},{"key":"2023062708443020100_B17","doi-asserted-by":"crossref","first-page":"224","DOI":"10.1186\/gb-2005-6-7-224","article-title":"Text-mining and information-retrieval services for molecular biology","volume":"6","author":"Krallinger","year":"2005","journal-title":"Genome Biol"},{"key":"2023062708443020100_B18","first-page":"218","article-title":"Learning from positive and unlabeled examples with different data distributions","author":"Li","year":"2005"},{"key":"2023062708443020100_B19","first-page":"216","article-title":"Aggregating UMLS semantic types for reducing conceptual complexity","volume":"10","author":"McCray","year":"2001","journal-title":"Medinfo"},{"key":"2023062708443020100_B20","first-page":"64","article-title":"Nested collocation and compound noun for term recognition","author":"Nakagawa","year":"1998"},{"key":"2023062708443020100_B21","doi-asserted-by":"crossref","first-page":"103","DOI":"10.1023\/A:1007692713085","article-title":"Text classification from labeled and unlabeled documents using EM","volume":"39","author":"Nigam","year":"2000","journal-title":"Mach. Learn"},{"key":"2023062708443020100_B22","first-page":"17","article-title":"An intelligent search engine and GUI-based efficient MEDLINE search tool based on deep syntactic parsing","author":"Ohta","year":"2006"},{"key":"2023062708443020100_B23","doi-asserted-by":"crossref","first-page":"2597","DOI":"10.1093\/bioinformatics\/bth291","article-title":"Distribution of information in biomedical abstracts and full-text publications","volume":"20","author":"Schuemie","year":"2004","journal-title":"Bioinformatics"},{"key":"2023062708443020100_B24","doi-asserted-by":"crossref","first-page":"20","DOI":"10.1186\/1471-2105-4-20","article-title":"Information extraction from full text scientific articles: where are the keywords?","volume":"4","author":"Shah","year":"2003","journal-title":"BMC Bioinformatics"},{"key":"2023062708443020100_B25","first-page":"26","article-title":"The Arrowsmith project: 2005 status report","volume-title":"Discovery Science 2005. Lecture Notes in Artificial Intelligence","author":"Smalheiser","year":"2005"},{"key":"2023062708443020100_B26","doi-asserted-by":"crossref","first-page":"149","DOI":"10.1016\/S0169-2607(98)00033-9","article-title":"Using ARROWSMITH: a computer-assisted approach to formulating and assessing scientific hypotheses","volume":"57","author":"Smalheiser","year":"1998","journal-title":"Comput. Methods Programs Biomed"},{"key":"2023062708443020100_B27","doi-asserted-by":"crossref","first-page":"8","DOI":"10.1186\/1747-5333-1-8","article-title":"Collaborative development of the Arrowsmith two-node search interface designed for laboratory investigators","volume":"1","author":"Smalheiser","year":"2006","journal-title":"J. Biomed. Discovery Collaboration"},{"key":"2023062708443020100_B28","doi-asserted-by":"crossref","first-page":"396","DOI":"10.1002\/asi.10389","article-title":"Text mining: generating hypotheses from MEDLINE","volume":"55","author":"Srinivasan","year":"2004","journal-title":"J. Am. Soc. Information Sci. Technol"},{"key":"2023062708443020100_B29","doi-asserted-by":"crossref","first-page":"183","DOI":"10.1016\/S0004-3702(97)00008-8","article-title":"An interactive system for finding complementary literatures: a stimulus to scientific discovery","volume":"91","author":"Swanson","year":"1997","journal-title":"Artif. Int"},{"key":"2023062708443020100_B30","doi-asserted-by":"crossref","first-page":"1427","DOI":"10.1002\/asi.20438","article-title":"Ranking indirect connections in literature-based discovery: The role of Medical Subject Headings (MeSH)","volume":"57","author":"Swanson","year":"2006","journal-title":"J. Am. Soc. Information Sci. Technol"},{"key":"2023062708443020100_B31","doi-asserted-by":"crossref","first-page":"611","DOI":"10.1142\/S0219720004000399","article-title":"Generation of a large gene\/protein lexicon by morphological pattern analysis","volume":"1","author":"Tanabe","year":"2004","journal-title":"J. Bioinform. Comput. Biol"},{"key":"2023062708443020100_B32","volume-title":"TREC: Experiment and Evaluation in Information Retrieval","author":"Voorhees","year":"2005"},{"key":"2023062708443020100_B33","doi-asserted-by":"crossref","first-page":"277","DOI":"10.1093\/bib\/6.3.277","article-title":"Online tools to support literature-based discovery in the life sciences","volume":"6","author":"Weeber","year":"2005","journal-title":"Brief Bioinform"},{"key":"2023062708443020100_B34","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1186\/1471-2105-5-145","article-title":"Extending the mutual information measure to rank inferred literature relationships","volume":"5","author":"Wren","year":"2004","journal-title":"BMC Bioinformatics"},{"key":"2023062708443020100_B35","doi-asserted-by":"crossref","first-page":"389","DOI":"10.1093\/bioinformatics\/btg421","article-title":"Knowledge discovery by automated identification and ranking of implicit relationships","volume":"20","author":"Wren","year":"2004","journal-title":"Bioinformatics"},{"key":"2023062708443020100_B36","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1016\/j.jbi.2005.11.010","article-title":"Using statistical and knowledge-based approaches for literature-based discovery","volume":"39","author":"Yetisgen-Yildiz","year":"2006","journal-title":"J. Biomed. Inform"},{"key":"2023062708443020100_B37","doi-asserted-by":"crossref","first-page":"2813","DOI":"10.1093\/bioinformatics\/btl480","article-title":"ADAM: Another database of abbreviations in MEDLINE","volume":"22","author":"Zhou","year":"2006","journal-title":"Bioinformatics"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/23\/13\/1658\/50714593\/bioinformatics_23_13_1658.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/23\/13\/1658\/50714593\/bioinformatics_23_13_1658.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,1,15]],"date-time":"2025-01-15T22:50:57Z","timestamp":1736981457000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/23\/13\/1658\/223810"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2007,4,26]]},"references-count":37,"journal-issue":{"issue":"13","published-print":{"date-parts":[[2007,7,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btm161","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2007,7]]},"published":{"date-parts":[[2007,4,26]]}}}