{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,15]],"date-time":"2025-08-15T01:03:42Z","timestamp":1755219822972,"version":"3.43.0"},"reference-count":0,"publisher":"IOS Press","isbn-type":[{"type":"electronic","value":"9781643686080"}],"license":[{"start":{"date-parts":[[2025,8,7]],"date-time":"2025-08-07T00:00:00Z","timestamp":1754524800000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,8,7]]},"abstract":"<jats:p>A critical first step in using large-scale data to study catatonia is the development of precise phenotyping algorithms that can identify instances of the condition. In this work, we present an ensemble approach that combines retrieval-augmented generation (RAG) large language models (LLMs) with boosting algorithms to phenotype catatonia from the electronic health records (EHRs) of 3.5 million individuals seen at a large academic medical center from 2006 to 2017. Although the ensemble model achieved an AUROC of 0.709, slightly lower than the boosting algorithm alone (AUROC = 0.713), the inclusion of the RAG-LLM component provides enhanced interpretability. In particular, the RAG-LLM can identify contextually complex clinical features, such as those described by the Bush\u2013Francis Catatonia Rating Scale, directly from clinical notes. These results highlight the potential of RAG-LLMs to capture nuanced contextual cues and fulfill complex catatonia phenotype definitions, even when overall classification performance is comparable to more traditional machine learning methods.<\/jats:p>","DOI":"10.3233\/shti250920","type":"book-chapter","created":{"date-parts":[[2025,8,7]],"date-time":"2025-08-07T11:35:09Z","timestamp":1754566509000},"source":"Crossref","is-referenced-by-count":0,"title":["An Ensemble Approach Integrating Retrieval-Augmented Large Language Models and Boosting Algorithms for Enhanced Catatonia Phenotyping"],"prefix":"10.3233","author":[{"given":"Yubo","family":"Feng","sequence":"first","affiliation":[{"name":"Vanderbilt University, Nashville, TN, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ruiyan","family":"Ma","sequence":"additional","affiliation":[{"name":"Vanderbilt University, Nashville, TN, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xinmeng","family":"Zhang","sequence":"additional","affiliation":[{"name":"Vanderbilt University, Nashville, TN, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"You","family":"Chen","sequence":"additional","affiliation":[{"name":"Vanderbilt University, Nashville, TN, USA"},{"name":"Vanderbilt University Medical Center, Nashville, TN, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"7437","container-title":["Studies in Health Technology and Informatics","MEDINFO 2025 \u2014 Healthcare Smart \u00d7 Medicine Deep"],"original-title":[],"link":[{"URL":"https:\/\/ebooks.iospress.nl\/pdf\/doi\/10.3233\/SHTI250920","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,7]],"date-time":"2025-08-07T11:35:09Z","timestamp":1754566509000},"score":1,"resource":{"primary":{"URL":"https:\/\/ebooks.iospress.nl\/doi\/10.3233\/SHTI250920"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,7]]},"ISBN":["9781643686080"],"references-count":0,"URL":"https:\/\/doi.org\/10.3233\/shti250920","relation":{},"ISSN":["0926-9630","1879-8365"],"issn-type":[{"type":"print","value":"0926-9630"},{"type":"electronic","value":"1879-8365"}],"subject":[],"published":{"date-parts":[[2025,8,7]]}}}