{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,6]],"date-time":"2026-06-06T15:58:32Z","timestamp":1780761512642,"version":"3.54.1"},"reference-count":27,"publisher":"Oxford University Press (OUP)","issue":"1","license":[{"start":{"date-parts":[[2021,11,13]],"date-time":"2021-11-13T00:00:00Z","timestamp":1636761600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/journals\/pages\/open_access\/funder_policies\/chorus\/standard_publication_model"}],"funder":[{"DOI":"10.13039\/100007277","name":"School of Medicine at Mount Sinai","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100007277","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2021,12,28]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Objective<\/jats:title>\n                    <jats:p>Clinical registries\u2014structured databases of demographic, diagnosis, and treatment information\u2014play vital roles in retrospective studies, operational planning, and assessment of patient eligibility for research, including clinical trials. Registry curation, a manual and time-intensive process, is always costly and often impossible for rare or underfunded diseases. Our goal was to evaluate the feasibility of natural language inference (NLI) as a scalable solution for registry curation.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Materials and Methods<\/jats:title>\n                    <jats:p>We applied five state-of-the-art, pretrained, deep learning-based NLI models to clinical, laboratory, and pathology notes to infer information about 43 different breast oncology registry fields. Model inferences were evaluated against a manually curated, 7439 patient breast oncology research database.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>NLI models showed considerable variation in performance, both within and across fields. One model, ALBERT, outperformed the others (BART, RoBERTa, XLNet, and ELECTRA) on 22 out of 43 fields. A detailed error analysis revealed that incorrect inferences primarily arose through models' tendency to misinterpret historical findings, as well as confusion based on abbreviations and subtle term variants common in clinical text.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Discussion and Conclusion<\/jats:title>\n                    <jats:p>Traditional natural language processing methods require specially annotated training sets or the construction of a separate model for each registry field. In contrast, a single pretrained NLI model can curate dozens of different fields simultaneously. Surprisingly, NLI methods remain unexplored in the clinical domain outside the realm of shared tasks and benchmarks. Modern NLI models could increase the efficiency of registry curation, even when applied \u201cout of the box\u201d with no additional training.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1093\/jamia\/ocab243","type":"journal-article","created":{"date-parts":[[2021,10,26]],"date-time":"2021-10-26T07:12:12Z","timestamp":1635232332000},"page":"97-108","source":"Crossref","is-referenced-by-count":23,"title":["Natural language inference for curation of structured clinical registries from unstructured text"],"prefix":"10.1093","volume":"29","author":[{"given":"Bethany","family":"Percha","sequence":"first","affiliation":[{"name":"Department of Medicine, Icahn School of Medicine at Mount Sinai, New York, New York, USA"},{"name":"Department of Genetics and Genomic Sciences, Icahn School of Medicine at Mount Sinai, New York, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kereeti","family":"Pisapati","sequence":"additional","affiliation":[{"name":"Mount Sinai Innovation Partners, Mount Sinai Health System, New York, New York, USA"},{"name":"Breast Surgical Oncology, Icahn School of Medicine at Mount Sinai, New York, New York, USA"},{"name":"Tisch Cancer Institute, Icahn School of Medicine at Mount Sinai, New York, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Cynthia","family":"Gao","sequence":"additional","affiliation":[{"name":"Department of Medicine, Icahn School of Medicine at Mount Sinai, New York, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hank","family":"Schmidt","sequence":"additional","affiliation":[{"name":"Breast Surgical Oncology, Icahn School of Medicine at Mount Sinai, New York, New York, USA"},{"name":"Tisch Cancer Institute, Icahn School of Medicine at Mount Sinai, New York, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2021,11,13]]},"reference":[{"issue":"2","key":"2021122823070427500_ocab243-B1","doi-asserted-by":"crossref","first-page":"377","DOI":"10.1016\/j.jbi.2008.08.010","article-title":"Research electronic data capture (REDCap)\u2014a metadata-driven methodology and workflow process for providing translational research informatics support","volume":"42","author":"Harris","year":"2009","journal-title":"J Biomed Inform"},{"issue":"4","key":"2021122823070427500_ocab243-B2","doi-asserted-by":"crossref","first-page":"605","DOI":"10.1093\/ejcts\/ezt018","article-title":"Clinical registries: governance, management, analysis and applications","volume":"44","author":"Hickey","year":"2013","journal-title":"Eur J Cardiothorac Surg"},{"issue":"4","key":"2021122823070427500_ocab243-B3","doi-asserted-by":"crossref","first-page":"843","DOI":"10.1002\/cpt.1658","article-title":"Examining the use of real-world evidence in the regulatory process","volume":"107","author":"Beaulieu-Jones","year":"2020","journal-title":"Clin Pharmacol Ther"},{"issue":"469","key":"2021122823070427500_ocab243-B4","doi-asserted-by":"crossref","first-page":"61","DOI":"10.1198\/016214504000001899","article-title":"Modeling reporting delays and reporting corrections in cancer registry data","volume":"100","author":"Midthune","year":"2005","journal-title":"J Am Stat Assoc"},{"issue":"5","key":"2021122823070427500_ocab243-B5","doi-asserted-by":"crossref","first-page":"747","DOI":"10.1016\/j.ejca.2008.11.032","article-title":"Evaluation of data quality in the cancer registry: principles and methods. Part I: comparability, validity and timeliness","volume":"45","author":"Bray","year":"2009","journal-title":"Eur J Cancer"},{"issue":"1","key":"2021122823070427500_ocab243-B6","doi-asserted-by":"crossref","first-page":"165","DOI":"10.1146\/annurev-biodatasci-030421-030931","article-title":"Modern clinical text mining: a guide and review","volume":"4","author":"Percha","year":"2021","journal-title":"Annu Rev Biomed Data Sci"},{"issue":"4","key":"2021122823070427500_ocab243-B7","doi-asserted-by":"crossref","first-page":"150","DOI":"10.3390\/info10040150","article-title":"Text classification algorithms: a survey","volume":"10","author":"Kowsari","year":"2019","journal-title":"Information"},{"key":"2021122823070427500_ocab243-B8","author":"Vaswani","year":"2017"},{"issue":"4","key":"2021122823070427500_ocab243-B9","doi-asserted-by":"crossref","first-page":"1","DOI":"10.2200\/S00509ED1V01Y201305HLT023","article-title":"Recognizing textual entailment: models and applications","volume":"6","author":"Dagan","year":"2013","journal-title":"Synth Lect Hum Lang Technol"},{"key":"2021122823070427500_ocab243-B10","author":"Nie","year":"2019"},{"key":"2021122823070427500_ocab243-B11","author":"Honnibal"},{"key":"2021122823070427500_ocab243-B12","author":"Lan","year":"2019"},{"key":"2021122823070427500_ocab243-B13","author":"Lewis","year":"2019"},{"key":"2021122823070427500_ocab243-B14","author":"Clark","year":"2020"},{"key":"2021122823070427500_ocab243-B15","author":"Liu","year":"2019"},{"key":"2021122823070427500_ocab243-B16","author":"Yang","year":"2019"},{"key":"2021122823070427500_ocab243-B17","author":"Wolf","year":"2019"},{"key":"2021122823070427500_ocab243-B18","author":"Bowman","year":"2015"},{"key":"2021122823070427500_ocab243-B19","author":"Williams","year":"2017"},{"key":"2021122823070427500_ocab243-B20","author":"Thorne","year":"2018"},{"key":"2021122823070427500_ocab243-B21","author":"Alsentzer","year":"2019"},{"key":"2021122823070427500_ocab243-B22","author":"Huang","year":"2019"},{"key":"2021122823070427500_ocab243-B23","author":"Romanov","year":"2018"},{"key":"2021122823070427500_ocab243-B24","author":"Devlin","year":"2018"},{"issue":"10","key":"2021122823070427500_ocab243-B25","doi-asserted-by":"crossref","first-page":"1345","DOI":"10.1109\/TKDE.2009.191","article-title":"A survey on transfer learning","volume":"22","author":"Pan","year":"2010","journal-title":"IEEE Trans Knowl Data Eng"},{"key":"2021122823070427500_ocab243-B26","year":"2019"},{"issue":"S3","key":"2021122823070427500_ocab243-B27","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1186\/s12911-018-0654-2","article-title":"Comparison of MetaMap and cTAKES for entity extraction in clinical notes","volume":"18","author":"Re\u00e1tegui","year":"2018","journal-title":"BMC Med Inform Decis Mak"}],"container-title":["Journal of the American Medical Informatics Association"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/jamia\/article-pdf\/29\/1\/97\/41955528\/ocab243.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jamia\/article-pdf\/29\/1\/97\/41955528\/ocab243.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,12,28]],"date-time":"2021-12-28T18:10:15Z","timestamp":1640715015000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/jamia\/article\/29\/1\/97\/6427462"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,11,13]]},"references-count":27,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2021,11,13]]},"published-print":{"date-parts":[[2021,12,28]]}},"URL":"https:\/\/doi.org\/10.1093\/jamia\/ocab243","relation":{"has-preprint":[{"id-type":"doi","id":"10.1101\/2021.06.14.21258493","asserted-by":"object"}]},"ISSN":["1527-974X"],"issn-type":[{"value":"1527-974X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2022,1,1]]},"published":{"date-parts":[[2021,11,13]]}}}