{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T02:39:38Z","timestamp":1778812778783,"version":"3.51.4"},"reference-count":40,"publisher":"Springer Science and Business Media LLC","issue":"S1","license":[{"start":{"date-parts":[[2020,4,1]],"date-time":"2020-04-01T00:00:00Z","timestamp":1585699200000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"},{"start":{"date-parts":[[2020,4,30]],"date-time":"2020-04-30T00:00:00Z","timestamp":1588204800000},"content-version":"vor","delay-in-days":29,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Med Inform Decis Mak"],"published-print":{"date-parts":[[2020,4]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec><jats:title>Background<\/jats:title><jats:p>Capturing sentence semantics plays a vital role in a range of text mining applications. Despite continuous efforts on the development of related datasets and models in the general domain, both datasets and models are limited in biomedical and clinical domains. The BioCreative\/OHNLP2018 organizers have made the first attempt to annotate 1068 sentence pairs from clinical notes and have called for a community effort to tackle the Semantic Textual Similarity (BioCreative\/OHNLP STS) challenge.<\/jats:p><\/jats:sec><jats:sec><jats:title>Methods<\/jats:title><jats:p>We developed models using traditional machine learning and deep learning approaches. For the post challenge, we focused on two models: the Random Forest and the Encoder Network. We applied sentence embeddings pre-trained on PubMed abstracts and MIMIC-III clinical notes and updated the Random Forest and the Encoder Network accordingly.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>The official results demonstrated our best submission was the ensemble of eight models. It achieved a Person correlation coefficient of 0.8328 \u2013 the highest performance among 13 submissions from 4 teams. For the post challenge, the performance of both Random Forest and the Encoder Network was improved; in particular, the correlation of the Encoder Network was improved by ~\u200913%. During the challenge task, no end-to-end deep learning models had better performance than machine learning models that take manually-crafted features. In contrast, with the sentence embeddings pre-trained on biomedical corpora, the Encoder Network now achieves a correlation of ~\u20090.84, which is higher than the original best model. The ensembled model taking the improved versions of the Random Forest and Encoder Network as inputs further increased performance to 0.8528.<\/jats:p><\/jats:sec><jats:sec><jats:title>Conclusions<\/jats:title><jats:p>Deep learning models with sentence embeddings pre-trained on biomedical corpora achieve the highest performance on the test set. Through error analysis, we find that end-to-end deep learning models and traditional machine learning models with manually-crafted features complement each other by finding different types of sentences. We suggest a combination of these models can better find similar sentences in practice.<\/jats:p><\/jats:sec>","DOI":"10.1186\/s12911-020-1044-0","type":"journal-article","created":{"date-parts":[[2020,4,30]],"date-time":"2020-04-30T00:02:58Z","timestamp":1588204978000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":17,"title":["Deep learning with sentence embeddings pre-trained on biomedical corpora improves the performance of finding similar sentences in electronic medical records"],"prefix":"10.1186","volume":"20","author":[{"given":"Qingyu","family":"Chen","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jingcheng","family":"Du","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sun","family":"Kim","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"W. John","family":"Wilbur","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhiyong","family":"Lu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2020,4,30]]},"reference":[{"key":"1044_CR1","doi-asserted-by":"crossref","unstructured":"Allot A, Chen Q, Kim S, Vera Alvarez R, Comeau DC, Wilbur WJ, Lu Z. LitSense: making sense of biomedical literature at sentence level. Nucleic Acids Res. 2019;47(W1):W594-9.","DOI":"10.1093\/nar\/gkz289"},{"issue":"1","key":"1044_CR2","first-page":"baw156","volume":"2017","author":"K Ravikumar","year":"2017","unstructured":"Ravikumar K, Rastegar-Mojarad M, Liu H. BELMiner: adapting a rule-based relation extraction system to extract biological expression language statements from bio-medical literature evidence sentences. Database. 2017;2017(1):baw156.","journal-title":"Database"},{"key":"1044_CR3","doi-asserted-by":"crossref","unstructured":"Tafti AP, Behravesh E, Assefi M, LaRose E, Badger J, Mayer J, Doan A, Page D, Peissig P. bigNN: An open-source big data toolkit focused on biomedical sentence classification. In Proceedings of the 2017 IEEE International Conference on Big Data (Big Data). 2017. p. 3888\u201396.","DOI":"10.1109\/BigData.2017.8258394"},{"key":"1044_CR4","doi-asserted-by":"publisher","first-page":"96","DOI":"10.1016\/j.jbi.2017.03.001","volume":"68","author":"M Sarrouti","year":"2017","unstructured":"Sarrouti M, El Alaoui SO. A passage retrieval method based on probabilistic information retrieval model and UMLS concepts in biomedical question answering. J Biomed Inform. 2017;68:96\u2013103.","journal-title":"J Biomed Inform"},{"key":"1044_CR5","doi-asserted-by":"crossref","unstructured":"J. Du, Q. Chen, Y. Peng, Y. Xiang, C. Tao, and Z. Lu, \u201cML-net: multi-label classification of biomedical texts with deep neural networks,\u201d J Am Med Inform Assoc. 2019.","DOI":"10.1093\/jamia\/ocz085"},{"key":"1044_CR6","doi-asserted-by":"crossref","unstructured":"Cer D, Diab M, Agirre E, Lopez-Gazpio I, Specia L. SemEval-2017 Task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation. arXiv preprint arXiv. 2017;1708(00055).","DOI":"10.18653\/v1\/S17-2001"},{"key":"1044_CR7","doi-asserted-by":"crossref","unstructured":"Chen Q, Kim S, Wilbur WJ, Lu Z. Sentence similarity measures revisited: ranking sentences in PubMed documents. In Proceedings of the 2018 ACM International Conference on Bioinformatics, Computational Biology, and Health Informatics. 2018. p. 531\u20132.","DOI":"10.1145\/3233547.3233640"},{"key":"1044_CR8","doi-asserted-by":"crossref","unstructured":"Wang Y, Afzal N, Fu S, Wang L, Shen F, Rastegar-Mojarad M, Liu H. MedSTS: A Resource for Clinical Semantic Textual Similarity. arXiv preprint arXiv. 2018;1808(09397).","DOI":"10.1007\/s10579-018-9431-1"},{"key":"1044_CR9","unstructured":"Chen Q, Du J, Kim S, Wilbur WJ, Lu Z. Combining rich features and deep learning for finding similar sentences in electronic medical records. Proceedings of Biocreative\/OHNLP challenge. 2018;2018."},{"key":"1044_CR10","volume-title":"The 7th IEEE international conference on healthcare informatics","author":"Q Chen","year":"2019","unstructured":"Chen Q, Peng Y, Lu Z. BioSentVec: creating sentence embeddings for biomedical texts. In: The 7th IEEE international conference on healthcare informatics; 2019."},{"key":"1044_CR11","doi-asserted-by":"crossref","unstructured":"Chen Q, Peng Y, Lu Z. BioSentVec: creating sentence embeddings for biomedical texts. In 2019 IEEE International Conference on Healthcare Informatics (ICHI) 2019 Jun 10 (pp. 1\u20135). IEEE.","DOI":"10.1109\/ICHI.2019.8904728"},{"issue":"10","key":"1044_CR12","doi-asserted-by":"publisher","first-page":"937","DOI":"10.1038\/nbt.4267","volume":"36","author":"N Fiorini","year":"2018","unstructured":"Fiorini N, Leaman R, Lipman DJ, Lu Z. How user intelligence is improving PubMed. Nat Biotechnol. 2018;36(10):937.","journal-title":"Nat Biotechnol"},{"issue":"1","key":"1044_CR13","doi-asserted-by":"publisher","first-page":"80","DOI":"10.1093\/bioinformatics\/btx541","volume":"34","author":"C-H Wei","year":"2017","unstructured":"Wei C-H, Phan L, Feltz J, Maiti R, Hefferon T, Lu Z. tmVar 2.0: integrating genomic variant information from literature with dbSNP and ClinVar for precision medicine. Bioinformatics. 2017;34(1):80\u20137.","journal-title":"Bioinformatics"},{"issue":"8","key":"1044_CR14","doi-asserted-by":"publisher","first-page":"e0159644","DOI":"10.1371\/journal.pone.0159644","volume":"11","author":"Q Chen","year":"2016","unstructured":"Chen Q, Zobel J, Zhang X, Verspoor K. Supervised learning for detection of duplicates in genomic sequence databases. PLoS One. 2016;11(8):e0159644.","journal-title":"PLoS One"},{"key":"1044_CR15","doi-asserted-by":"crossref","unstructured":"Zobel J, Moffat A. Exploring the similarity space. In SIGIR Forum. 1998;32(1):18\u201334.","DOI":"10.1145\/281250.281256"},{"key":"1044_CR16","first-page":"69","volume":"38","author":"P Jaccard","year":"1902","unstructured":"Jaccard P. Lois de distribution florale dans la zone alpine. Bull Soc Vaud Sci Nat. 1902;38:69\u2013130.","journal-title":"Bull Soc Vaud Sci Nat"},{"issue":"406","key":"1044_CR17","doi-asserted-by":"publisher","first-page":"414","DOI":"10.1080\/01621459.1989.10478785","volume":"84","author":"MA Jaro","year":"1989","unstructured":"Jaro MA. Advances in record-linkage methodology as applied to matching the 1985 census of Tampa, Florida. J Am Stat Assoc. 1989;84(406):414\u201320.","journal-title":"J Am Stat Assoc"},{"issue":"3","key":"1044_CR18","doi-asserted-by":"publisher","first-page":"297","DOI":"10.2307\/1932409","volume":"26","author":"LR Dice","year":"1945","unstructured":"Dice LR. Measures of the amount of ecologic association between species. Ecology. 1945;26(3):297\u2013302.","journal-title":"Ecology"},{"key":"1044_CR19","doi-asserted-by":"publisher","first-page":"526","DOI":"10.2331\/suisan.22.526","volume":"22","author":"A Ochiai","year":"1957","unstructured":"Ochiai A. Zoogeographic studies on the soleoid fishes found in Japan and its neighbouring regions. Bulletin of Japanese Society of Scientific Fisheries. 1957;22:526\u201330.","journal-title":"Bulletin of Japanese Society of Scientific Fisheries"},{"issue":"1","key":"1044_CR20","doi-asserted-by":"publisher","first-page":"11","DOI":"10.1108\/eb026526","volume":"28","author":"K Sparck Jones","year":"1972","unstructured":"Sparck Jones K. A statistical interpretation of term specificity and its application in retrieval. J Doc. 1972;28(1):11\u201321.","journal-title":"J Doc"},{"issue":"1","key":"1044_CR21","doi-asserted-by":"publisher","first-page":"191","DOI":"10.1016\/0304-3975(92)90143-4","volume":"92","author":"E Ukkonen","year":"1992","unstructured":"Ukkonen E. Approximate string-matching with q-grams and maximal matches. Theor Comput Sci. 1992;92(1):191\u2013211.","journal-title":"Theor Comput Sci"},{"issue":"1","key":"1044_CR22","doi-asserted-by":"publisher","first-page":"49","DOI":"10.1197\/jamia.M3390","volume":"17","author":"JO Wrenn","year":"2010","unstructured":"Wrenn JO, Stein DM, Bakken S, Stetson PD. Quantifying clinical narrative redundancy in an electronic health record. J Am Med Inform Assoc. 2010;17(1):49\u201353.","journal-title":"J Am Med Inform Assoc"},{"key":"1044_CR23","doi-asserted-by":"crossref","unstructured":"Chen Q, Zobel J, Verspoor K. Duplicates, redundancies and inconsistencies in the primary nucleotide databases: a descriptive study. Database. 2017;2017:baw163.","DOI":"10.1093\/database\/baw163"},{"key":"1044_CR24","unstructured":"Navarro G. Multiple approximate string matching by counting. In WSP 1997, 4th South American Workshop on String Processing. 2011. p. 95\u2013111."},{"key":"1044_CR25","unstructured":"Levenshtein VI. Binary codes capable of correcting deletions, insertions and reversals In: Soviet Physics Doklady. 1966;10:707."},{"issue":"3","key":"1044_CR26","doi-asserted-by":"publisher","first-page":"443","DOI":"10.1016\/0022-2836(70)90057-4","volume":"48","author":"SB Needleman","year":"1970","unstructured":"Needleman SB, Wunsch CD. A general method applicable to the search for similarities in the amino acid sequence of two proteins. J Mol Biol. 1970;48(3):443\u201353.","journal-title":"J Mol Biol"},{"issue":"4","key":"1044_CR27","doi-asserted-by":"publisher","first-page":"482","DOI":"10.1016\/0196-8858(81)90046-4","volume":"2","author":"TF Smith","year":"1981","unstructured":"Smith TF, Waterman MS. Comparison of biosequences. Adv Appl Math. 1981;2(4):482\u20139.","journal-title":"Adv Appl Math"},{"key":"1044_CR28","doi-asserted-by":"publisher","first-page":"160035","DOI":"10.1038\/sdata.2016.35","volume":"3","author":"AE Johnson","year":"2016","unstructured":"Johnson AE, Pollard TJ, Shen L, Li-wei HL, Feng M, Ghassemi M, Moody B, Szolovits P, Celi LA, Mark RG. MIMIC-III, a freely accessible critical care database. Sci Data. 2016;3:160035.","journal-title":"Sci Data"},{"key":"1044_CR29","doi-asserted-by":"crossref","unstructured":"Wei C-H, Allot A, Leaman R, Lu Z. PubTator central: automated concept annotation for biomedical full text articles: Nucleic Acids Res. 2019:47(W1):W587\u201393.","DOI":"10.1093\/nar\/gkz389"},{"issue":"3","key":"1044_CR30","doi-asserted-by":"publisher","first-page":"331","DOI":"10.1093\/jamia\/ocx132","volume":"25","author":"E Soysal","year":"2017","unstructured":"Soysal E, Wang J, Jiang M, Wu Y, Pakhomov S, Liu H, Xu H. CLAMP\u2013a toolkit for efficiently building customized clinical natural language processing pipelines. J Am Med Inform Assoc. 2017;25(3):331\u20136.","journal-title":"J Am Med Inform Assoc"},{"key":"1044_CR31","unstructured":"Kusner M, Sun Y, Kolkin N, Weinberger K. From word embeddings to document distances. In International conference on machine learning. 2015. p. 957-66."},{"key":"1044_CR32","unstructured":"Chen Q, Peng Y, Keenan T, Dharssi S, Agro E. A multi-task deep learning model for the classification of Age-related Macular Degeneration. AMIA Summits on Translational Science Proceedings. 2019;2019:505."},{"issue":"2","key":"1044_CR33","doi-asserted-by":"publisher","first-page":"43","DOI":"10.1186\/s12911-018-0632-8","volume":"18","author":"J Du","year":"2018","unstructured":"Du J, Zhang Y, Luo J, Jia Y, Wei Q, Tao C, Xu H. Extracting psychiatric stressors for suicide from social media using deep learning. BMC Med Inform Decis Mak. 2018;18(2):43.","journal-title":"BMC Med Inform Decis Mak"},{"key":"1044_CR34","doi-asserted-by":"crossref","unstructured":"Do\u011fan RI, Kim S, Chatr-aryamontri A, Wei C-H, Comeau DC, Antunes R, Matos S, Chen Q, Elangovan A, Panyam NC. Overview of the BioCreative VI precision medicine track: mining protein interactions and mutations for precision medicine. Database. 2019;2019.","DOI":"10.1093\/database\/bay147"},{"key":"1044_CR35","doi-asserted-by":"crossref","unstructured":"Kim Y. Convolutional neural networks for sentence classification. arXiv preprint arXiv. 2014;1408(5882).","DOI":"10.3115\/v1\/D14-1181"},{"key":"1044_CR36","doi-asserted-by":"crossref","unstructured":"Mueller J, Thyagarajan A. Siamese recurrent architectures for learning sentence similarity. In thirtieth AAAI conference on artificial intelligence. 2016.","DOI":"10.1609\/aaai.v30i1.10350"},{"key":"1044_CR37","doi-asserted-by":"crossref","unstructured":"Serban IV, Sordoni A, Lowe R, Charlin L, Pineau J, Courville A, Bengio Y. A hierarchical latent variable encoder-decoder model for generating dialogues. In Thirty-First AAAI Conference on Artificial Intelligence. 2017.","DOI":"10.1609\/aaai.v31i1.10983"},{"key":"1044_CR38","doi-asserted-by":"crossref","unstructured":"Cer D, Yang Y, Kong S-y, Hua N, Limtiaco N, John RS, Constant N, Guajardo-Cespedes M, Yuan S, Tar C. Universal sentence encoder. arXiv preprint arXiv. 2018;1803(11175).","DOI":"10.18653\/v1\/D18-2029"},{"key":"1044_CR39","doi-asserted-by":"crossref","unstructured":"Conneau A, Kiela D, Schwenk H, Barrault L, Bordes A. Supervised learning of universal sentence representations from natural language inference data. arXiv preprint arXiv. 2017;1705(02364).","DOI":"10.18653\/v1\/D17-1070"},{"key":"1044_CR40","unstructured":"Zhang R, Pakhomov S, McInnes BT, Melton GB. Evaluating measures of redundancy in clinical texts. In AMIA Annual Symposium Proceedings. Am Med Inform Assoc. 2011;2011:1612."}],"container-title":["BMC Medical Informatics and Decision Making"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12911-020-1044-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12911-020-1044-0\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12911-020-1044-0.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,10,22]],"date-time":"2022-10-22T16:32:31Z","timestamp":1666456351000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcmedinformdecismak.biomedcentral.com\/articles\/10.1186\/s12911-020-1044-0"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,4]]},"references-count":40,"journal-issue":{"issue":"S1","published-print":{"date-parts":[[2020,4]]}},"alternative-id":["1044"],"URL":"https:\/\/doi.org\/10.1186\/s12911-020-1044-0","relation":{},"ISSN":["1472-6947"],"issn-type":[{"value":"1472-6947","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,4]]},"assertion":[{"value":"30 April 2020","order":1,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"N\/A","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"N\/A","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"73"}}