{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,5,17]],"date-time":"2024-05-17T22:59:58Z","timestamp":1715986798978},"reference-count":23,"publisher":"Springer Science and Business Media LLC","issue":"S9","license":[{"start":{"date-parts":[[2007,11,27]],"date-time":"2007-11-27T00:00:00Z","timestamp":1196121600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/2.0"},{"start":{"date-parts":[[2007,11,27]],"date-time":"2007-11-27T00:00:00Z","timestamp":1196121600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/2.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2007,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:sec>\n            <jats:title>Motivation<\/jats:title>\n            <jats:p>With more and more research dedicated to literature mining in the biomedical domain, more and more systems are available for people to choose from when building literature mining applications. In this study, we focus on one specific kind of literature mining task, i.e., detecting definitions of acronyms, abbreviations, and symbols in biomedical text. We denote acronyms, abbreviations, and symbols as short forms (SFs) and their corresponding definitions as long forms (LFs). The study was designed to answer the following questions; i) how well a system performs in detecting LFs from novel text, ii) what the coverage is for various terminological knowledge bases in including SFs as synonyms of their LFs, and iii) how to combine results from various SF knowledge bases.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Method<\/jats:title>\n            <jats:p>We evaluated the following three publicly available detection systems in detecting LFs for SFs: i) a handcrafted pattern\/rule based system by Ao and Takagi, ALICE, ii) a machine learning system by Chang et al., and iii) a simple alignment-based program by Schwartz and Hearst. In addition, we investigated the conceptual coverage of two terminological knowledge bases: i) the UMLS (the Unified Medical Language System), and ii) the BioThesaurus (a thesaurus of names for all UniProt protein records). We also implemented a web interface that provides a virtual integration of various SF knowledge bases.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Results<\/jats:title>\n            <jats:p>We found that detection systems agree with each other on most cases, and the existing terminological knowledge bases have a good coverage of synonymous relationship for frequently defined LFs. The web interface allows people to detect SF definitions from text and to search several SF knowledge bases.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Availability<\/jats:title>\n            <jats:p>The web site is <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"http:\/\/gauss.dbb.georgetown.edu\/liblab\/SFThesaurus\" ext-link-type=\"uri\">http:\/\/gauss.dbb.georgetown.edu\/liblab\/SFThesaurus<\/jats:ext-link>.<\/jats:p>\n          <\/jats:sec>","DOI":"10.1186\/1471-2105-8-s9-s5","type":"journal-article","created":{"date-parts":[[2007,11,27]],"date-time":"2007-11-27T19:13:44Z","timestamp":1196190824000},"update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":14,"title":["A comparison study on algorithms of detecting long forms for short forms in biomedical text"],"prefix":"10.1186","volume":"8","author":[{"given":"Manabu","family":"Torii","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhang-zhi","family":"Hu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Min","family":"Song","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cathy H","family":"Wu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hongfang","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2007,11,27]]},"reference":[{"issue":"4","key":"1976_CR1","doi-asserted-by":"publisher","first-page":"247","DOI":"10.1016\/S1532-0464(03)00014-5","volume":"35","author":"L Hirschman","year":"2002","unstructured":"Hirschman L, Morgan AA, Yeh AS: Rutabaga by any other name: extracting biological names. Journal of Biomedical Informatics 2002,35(4):247\u2013259. 10.1016\/S1532-0464(03)00014-5","journal-title":"Journal of Biomedical Informatics"},{"issue":"5","key":"1976_CR2","doi-asserted-by":"publisher","first-page":"589","DOI":"10.1016\/j.molcel.2006.02.012","volume":"21","author":"L Hunter","year":"2006","unstructured":"Hunter L, Cohen KB: Biomedical language processing: what's beyond PubMed? Mol Cell 2006,21(5):589\u2013594. 10.1016\/j.molcel.2006.02.012","journal-title":"Mol Cell"},{"issue":"5","key":"1976_CR3","doi-asserted-by":"publisher","first-page":"576","DOI":"10.1197\/jamia.M1757","volume":"12","author":"H Ao","year":"2005","unstructured":"Ao H, Takagi T: ALICE: an algorithm to extract abbreviations from MEDLINE. J Am Med Inform Assoc 2005,12(5):576\u2013586. 10.1197\/jamia.M1757","journal-title":"J Am Med Inform Assoc"},{"key":"1976_CR4","first-page":"451","volume-title":"Pac Symp Biocomput","author":"AS Schwartz","year":"2003","unstructured":"Schwartz AS, Hearst MA: A simple algorithm for identifying abbreviation definitions in biomedical text. Pac Symp Biocomput 2003, 451\u2013462."},{"issue":"6","key":"1976_CR5","doi-asserted-by":"publisher","first-page":"612","DOI":"10.1197\/jamia.M1139","volume":"9","author":"JT Chang","year":"2002","unstructured":"Chang JT, Schutze H, Altman RB: Creating an online dictionary of abbreviations from MEDLINE. J Am Med Inform Assoc 2002,9(6):612\u2013620. 10.1197\/jamia.M1139","journal-title":"J Am Med Inform Assoc"},{"key":"1976_CR6","first-page":"371","volume":"10","author":"J Pustejovsky","year":"2001","unstructured":"Pustejovsky J, Casta\u00f1o J, Cochran B, Kotecki M, Morrell M, Rumshisky A: Extraction and Disambiguation of Acronym-Meaning Pairs in Medline. Medinfo 2001, 10: 371\u2013375.","journal-title":"Medinfo"},{"key":"1976_CR7","volume-title":"Nucleic Acids Res","author":"O Bodenreider","year":"2004","unstructured":"Bodenreider O: The Unified Medical Language System (UMLS): integrating biomedical terminology. Nucleic Acids Res 2004, (32 Database):D267\u2013270. 10.1093\/nar\/gkh061"},{"key":"1976_CR8","volume-title":"Unpublished PhD thesis","author":"M Zahariev","year":"2004","unstructured":"Zahariev M: A (Acronyms). In Unpublished PhD thesis. School of Computing Science, Simon Fraser University, USA; 2004."},{"issue":"4","key":"1976_CR9","doi-asserted-by":"publisher","first-page":"249","DOI":"10.1006\/jbin.2001.1023","volume":"34","author":"H Liu","year":"2001","unstructured":"Liu H, Lussier YA, Friedman C: Disambiguating ambiguous biomedical terms in biomedical narrative text: an unsupervised method. J Biomed Inform 2001,34(4):249\u2013261. 10.1006\/jbin.2001.1023","journal-title":"J Biomed Inform"},{"issue":"11","key":"1976_CR10","doi-asserted-by":"publisher","first-page":"965","DOI":"10.1046\/j.1365-2923.2000.0818h.x","volume":"34","author":"T Luxton","year":"2000","unstructured":"Luxton T, Al-Qassab H: Better use of abbreviations - a lesson from a stroke. Medical Education 2000,34(11):965. 10.1046\/j.1365-2923.2000.0818h.x","journal-title":"Medical Education"},{"issue":"1","key":"1976_CR11","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1046\/j.1464-410x.2000.00717.x","volume":"86","author":"DA Bloom","year":"2000","unstructured":"Bloom DA: Acronyms, abbreviations and initialisms. BJU International 2000,86(1):1\u20136. 10.1046\/j.1464-410x.2000.00717.x","journal-title":"BJU International"},{"key":"1976_CR12","doi-asserted-by":"publisher","first-page":"D319","DOI":"10.1093\/nar\/gkj147","volume-title":"Nucleic Acids Res","author":"TA Eyre","year":"2006","unstructured":"Eyre TA, Ducluzeau F, Sneddon TP, Povey S, Bruford EA, Lush MJ: The HUGO Gene Nomenclature Database, 2006 updates. Nucleic Acids Res 2006, (34 Database):D319\u2013321. 10.1093\/nar\/gkj147","edition":"34 Database"},{"issue":"4","key":"1976_CR13","doi-asserted-by":"publisher","first-page":"191","DOI":"10.1007\/s100320050018","volume":"1","author":"K Taghva","year":"1999","unstructured":"Taghva K, Gilbreth J: Finding Acronyms and Their Definitions. Int Journal on Document Analysis and Recognition 1999,1(4):191\u2013198. 10.1007\/s100320050018","journal-title":"Int Journal on Document Analysis and Recognition"},{"issue":"5\u20136","key":"1976_CR14","doi-asserted-by":"publisher","first-page":"322","DOI":"10.1016\/S1532-0464(03)00032-7","volume":"35","author":"H Yu","year":"2002","unstructured":"Yu H, Hatzivassiloglou V, Rzhetsky A, Wilbur WJ: Automatically identifying gene\/protein terms in MEDLINE abstracts. J Biomed Inform 2002,35(5\u20136):322\u2013330. 10.1016\/S1532-0464(03)00032-7","journal-title":"J Biomed Inform"},{"issue":"2","key":"1976_CR15","doi-asserted-by":"publisher","first-page":"169","DOI":"10.1093\/bioinformatics\/16.2.169","volume":"16","author":"M Yoshida","year":"2000","unstructured":"Yoshida M, Fukuda K, Takagi T: PNAD-CSS: a workbench for constructing a protein name abbreviation dictionary. Bioinformatics 2000,16(2):169\u2013175. 10.1093\/bioinformatics\/16.2.169","journal-title":"Bioinformatics"},{"key":"1976_CR16","first-page":"319","volume-title":"18th Conference of the Canadian Society for Computational Studies of Intelligence: 2005; Victoria, BC, Canada","author":"D Nadeau","year":"2005","unstructured":"Nadeau D, Turney P: A Supervised Learning Approach to Acronym Identification. 18th Conference of the Canadian Society for Computational Studies of Intelligence: 2005; Victoria, BC, Canada 2005, 319\u2013329."},{"key":"1976_CR17","volume-title":"Conference on Empirical Methods in Natural Language Processing: 2001; Pittsburgh, PA","author":"Y Park","year":"2001","unstructured":"Park Y, Byrd RJ: Hybrid Text Mining for Finding Abbreviations and Their Definitions. Conference on Empirical Methods in Natural Language Processing: 2001; Pittsburgh, PA 2001."},{"issue":"5","key":"1976_CR18","doi-asserted-by":"crossref","first-page":"426","DOI":"10.1055\/s-0038-1634373","volume":"41","author":"JD Wren","year":"2002","unstructured":"Wren JD, Garner HR: Heuristics for identification of acronym-definition patterns within text: towards an automated construction of comprehensive acronym-definition dictionaries. Methods Inf Med 2002,41(5):426\u2013434.","journal-title":"Methods Inf Med"},{"issue":"4","key":"1976_CR19","doi-asserted-by":"publisher","first-page":"527","DOI":"10.1093\/bioinformatics\/btg439","volume":"20","author":"E Adar","year":"2004","unstructured":"Adar E: SaRAD: a Simple and Robust Abbreviation Dictionary. Bioinformatics 2004,20(4):527\u2013533. 10.1093\/bioinformatics\/btg439","journal-title":"Bioinformatics"},{"key":"1976_CR20","first-page":"415","volume-title":"Pac Symp Biocomput","author":"H Liu","year":"2003","unstructured":"Liu H, Friedman C: Mining terminological knowledge in large biomedical corpora. Pac Symp Biocomput 2003, 415\u2013426."},{"key":"1976_CR21","volume-title":"Bioinformatics","author":"N Okazaki","year":"2006","unstructured":"Okazaki N, Ananiadou S: Building an abbreviation dictionary using a term recognition approach. Bioinformatics 2006."},{"issue":"22","key":"1976_CR22","doi-asserted-by":"publisher","first-page":"2813","DOI":"10.1093\/bioinformatics\/btl480","volume":"22","author":"W Zhou","year":"2006","unstructured":"Zhou W, Torvik VI, Smalheiser NR: ADAM: another database of abbreviations in MEDLINE. Bioinformatics 2006,22(22):2813\u20132818. 10.1093\/bioinformatics\/btl480","journal-title":"Bioinformatics"},{"key":"1976_CR23","first-page":"D154","volume-title":"Nucleic Acids Res","author":"A Bairoch","year":"2005","unstructured":"Bairoch A, Apweiler R, Wu CH, Barker WC, Boeckmann B, Ferro S, Gasteiger E, Huang H, Lopez R, Magrane M, et al.: The Universal Protein Resource (UniProt). Nucleic Acids Res 2005, (33 Database):D154\u2013159."}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-8-S9-S5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/1471-2105-8-S9-S5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-8-S9-S5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,8,31]],"date-time":"2021-08-31T21:25:26Z","timestamp":1630445126000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-8-S9-S5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2007,11,27]]},"references-count":23,"journal-issue":{"issue":"S9","published-print":{"date-parts":[[2007,12]]}},"alternative-id":["1976"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-8-s9-s5","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2007,11,27]]},"assertion":[{"value":"27 November 2007","order":1,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"S5"}}