{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,20]],"date-time":"2026-05-20T21:20:00Z","timestamp":1779312000293,"version":"3.51.4"},"reference-count":30,"publisher":"Oxford University Press (OUP)","issue":"12","license":[{"start":{"date-parts":[[2020,10,21]],"date-time":"2020-10-21T00:00:00Z","timestamp":1603238400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/journals\/pages\/open_access\/funder_policies\/chorus\/standard_publication_model"}],"funder":[{"name":"Intramural Research Program of the National Library of Medicine"},{"DOI":"10.13039\/100000002","name":"National Institutes of Health.","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2020,12,9]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Objective<\/jats:title>\n                  <jats:p>In a biomedical literature search, the link between a query and a document is often not established, because they use different terms to refer to the same concept. Distributional word embeddings are frequently used for detecting related words by computing the cosine similarity between them. However, previous research has not established either the best embedding methods for detecting synonyms among related word pairs or how effective such methods may be.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Materials and Methods<\/jats:title>\n                  <jats:p>In this study, we first create the BioSearchSyn set, a manually annotated set of synonyms, to assess and compare 3 widely used word-embedding methods (word2vec, fastText, and GloVe) in their ability to detect synonyms among related pairs of words. We demonstrate the shortcomings of the cosine similarity score between word embeddings for this task: the same scores have very different meanings for the different methods. To address the problem, we propose utilizing pool adjacent violators (PAV), an isotonic regression algorithm, to transform a cosine similarity into a probability of 2 words being synonyms.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>Experimental results using the BioSearchSyn set as a gold standard reveal which embedding methods have the best performance in identifying synonym pairs. The BioSearchSyn set also allows converting cosine similarity scores into probabilities, which provides a uniform interpretation of the synonymy score over different methods.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Conclusions<\/jats:title>\n                  <jats:p>We introduced the BioSearchSyn corpus of 1000 term pairs, which allowed us to identify the best embedding method for detecting synonymy for biomedical search. Using the proposed method, we created PubTermVariants2.0: a large, automatically extracted set of synonym pairs that have augmented PubMed searches since the spring of 2019.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/jamia\/ocaa151","type":"journal-article","created":{"date-parts":[[2020,8,20]],"date-time":"2020-08-20T19:09:45Z","timestamp":1597950585000},"page":"1894-1902","source":"Crossref","is-referenced-by-count":12,"title":["Better synonyms for enriching biomedical search"],"prefix":"10.1093","volume":"27","author":[{"given":"Lana","family":"Yeganova","sequence":"first","affiliation":[{"name":"National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Maryland, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3072-6649","authenticated-orcid":false,"given":"Sun","family":"Kim","sequence":"additional","affiliation":[{"name":"National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Maryland, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qingyu","family":"Chen","sequence":"additional","affiliation":[{"name":"National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Maryland, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Grigory","family":"Balasanov","sequence":"additional","affiliation":[{"name":"National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Maryland, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"W John","family":"Wilbur","sequence":"additional","affiliation":[{"name":"National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Maryland, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhiyong","family":"Lu","sequence":"additional","affiliation":[{"name":"National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Bethesda, Maryland, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2020,10,21]]},"reference":[{"issue":"2","key":"2020121009244030500_ocaa151-B1","doi-asserted-by":"crossref","first-page":"390","DOI":"10.1016\/j.jbi.2009.02.002","article-title":"Empirical distributional semantics: methods and biomedical applications","volume":"42","author":"Cohen","year":"2009","journal-title":"J Biomed Inform"},{"issue":"8","key":"2020121009244030500_ocaa151-B2","doi-asserted-by":"crossref","first-page":"e2005343","DOI":"10.1371\/journal.pbio.2005343","article-title":"Best match: new relevance search for PubMed","volume":"16","author":"Fiorini","year":"2018","journal-title":"PLoS Biol"},{"issue":"10","key":"2020121009244030500_ocaa151-B3","doi-asserted-by":"crossref","first-page":"937","DOI":"10.1038\/nbt.4267","article-title":"How user intelligence is improving PubMed","volume":"36","author":"Fiorini","year":"2018","journal-title":"Nat Biotechnol"},{"key":"2020121009244030500_ocaa151-B4","doi-asserted-by":"crossref","DOI":"10.1007\/978-0-387-78703-9","volume-title":"Information Retrieval: A Health and Biomedical Perspective","author":"Hersh","year":"2009"},{"key":"2020121009244030500_ocaa151-B5","doi-asserted-by":"crossref","first-page":"122","DOI":"10.1016\/j.jbi.2017.09.014","article-title":"Bridging the gap: incorporating a semantic similarity measure for effectively mapping PubMed queries to documents","volume":"75","author":"Kim","year":"2017","journal-title":"J Biomed Inform"},{"key":"2020121009244030500_ocaa151-B6","author":"Yeganova","year":"2016"},{"key":"2020121009244030500_ocaa151-B7","author":"Yu","year":"2017"},{"issue":"3","key":"2020121009244030500_ocaa151-B8","doi-asserted-by":"crossref","first-page":"288","DOI":"10.1016\/j.jbi.2006.06.004","article-title":"Measures of semantic similarity and relatedness in the biomedical domain","volume":"40","author":"Pedersen","year":"2007","journal-title":"J Biomed Inform"},{"key":"2020121009244030500_ocaa151-B9","volume-title":"proceedings from the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","author":"Qu","year":"2017"},{"key":"2020121009244030500_ocaa151-B10","author":"Zhang"},{"key":"2020121009244030500_ocaa151-B11","author":"Pakhomov","year":"2010"},{"issue":"1","key":"2020121009244030500_ocaa151-B12","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s12911-017-0498-1","article-title":"Semantic relatedness and similarity of biomedical terms: examining the effects of recency, size, and section of biomedical publications on the performance of word2vec","volume":"17","author":"Zhu","year":"2017","journal-title":"BMC Med Inform Decis Mak"},{"issue":"1","key":"2020121009244030500_ocaa151-B13","doi-asserted-by":"crossref","first-page":"33","DOI":"10.1186\/s12859-018-2039-z","article-title":"Bio-SimVerb and Bio-SimLex: wide-coverage evaluation sets of word similarity in biomedicine","volume":"19","author":"Chiu","year":"2018","journal-title":"BMC Bioinform"},{"issue":"4","key":"2020121009244030500_ocaa151-B14","doi-asserted-by":"crossref","first-page":"e1007617","DOI":"10.1371\/journal.pcbi.1007617","article-title":"BioConceptVec: creating and evaluating literature-based biomedical concept embeddings on a large scale","volume":"16","author":"Chen","year":"2020","journal-title":"PLOS Comput Biol"},{"key":"2020121009244030500_ocaa151-B15","author":"Chen","year":"2019"},{"issue":"1","key":"2020121009244030500_ocaa151-B16","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1038\/s41597-019-0055-0","article-title":"BioWordVec, improving biomedical word embeddings with subword information and MeSH","volume":"6","author":"Zhang","year":"2019","journal-title":"Sci Data"},{"key":"2020121009244030500_ocaa151-B17","doi-asserted-by":"crossref","first-page":"103321","DOI":"10.1016\/j.jbi.2019.103321","article-title":"Quantifying semantic similarity of clinical evidence in the biomedical literature to facilitate related evidence synthesis","volume":"100","author":"Hassanzadeh","year":"2019","journal-title":"J Biomed Inform"},{"key":"2020121009244030500_ocaa151-B18","author":"Mikolov","year":"2013"},{"key":"2020121009244030500_ocaa151-B19","doi-asserted-by":"crossref","first-page":"135","DOI":"10.1162\/tacl_a_00051","article-title":"Enriching word vectors with subword information","volume":"5","author":"Bojanowski","year":"2017","journal-title":"Trans Assoc Comput Linguist"},{"key":"2020121009244030500_ocaa151-B20","volume-title":"proceedings from the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Pennington","year":"2014"},{"issue":"4","key":"2020121009244030500_ocaa151-B21","doi-asserted-by":"crossref","first-page":"641","DOI":"10.1214\/aoms\/1177728423","article-title":"An empirical distribution function for sampling with incomplete information","volume":"26","author":"Ayer","year":"1955","journal-title":"Ann Math Stat"},{"issue":"3","key":"2020121009244030500_ocaa151-B22","doi-asserted-by":"crossref","first-page":"130","DOI":"10.1108\/eb046814","article-title":"An algorithm for suffix stripping","volume":"14","author":"Porter","year":"1980","journal-title":"Program"},{"key":"2020121009244030500_ocaa151-B23","volume-title":"Introduction to Probability Theory and Statistical Inference","author":"Larson","year":"1982","edition":"3rd ed"},{"issue":"1","key":"2020121009244030500_ocaa151-B24","first-page":"1","article-title":"A study of the morpho-semantic relationship in Medline","volume":"6","author":"Wilbur","year":"2013","journal-title":"Open Inf Syst J"},{"key":"2020121009244030500_ocaa151-B25","author":"Lin","year":"1998"},{"key":"2020121009244030500_ocaa151-B26","volume-title":"Ranking and Information Retrieval, in Managing Gigabytes: Compressing and Indexing Documents and Images","author":"Witten","year":"1999"},{"key":"2020121009244030500_ocaa151-B27","volume-title":"proceedings from the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Kim","year":"2015"},{"issue":"1-3","key":"2020121009244030500_ocaa151-B28","doi-asserted-by":"crossref","first-page":"71","DOI":"10.1007\/s10994-005-1123-6","article-title":"The synergy between PAV and AdaBoost","volume":"61","author":"Wilbur","year":"2005","journal-title":"Mach Learn"},{"key":"2020121009244030500_ocaa151-B29","doi-asserted-by":"crossref","article-title":"Towards PubMed 2.0","author":"Fiorini","DOI":"10.7554\/eLife.28801"},{"key":"2020121009244030500_ocaa151-B30","author":"Fiorini"}],"container-title":["Journal of the American Medical Informatics Association"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/academic.oup.com\/jamia\/article-pdf\/27\/12\/1894\/34838591\/ocaa151.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"http:\/\/academic.oup.com\/jamia\/article-pdf\/27\/12\/1894\/34838591\/ocaa151.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2020,12,10]],"date-time":"2020-12-10T14:41:35Z","timestamp":1607611295000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/jamia\/article\/27\/12\/1894\/5933975"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,21]]},"references-count":30,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2020,10,21]]},"published-print":{"date-parts":[[2020,12,9]]}},"URL":"https:\/\/doi.org\/10.1093\/jamia\/ocaa151","relation":{},"ISSN":["1527-974X"],"issn-type":[{"value":"1527-974X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2020,12]]},"published":{"date-parts":[[2020,10,21]]}}}