{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,10]],"date-time":"2026-08-10T15:14:46Z","timestamp":1786374886272,"version":"build-2736575974"},"reference-count":57,"publisher":"JMIR Publications Inc.","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["JMIR Form Res"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec sec-type=\"background\">\n                    <jats:title>Background<\/jats:title>\n                    <jats:p>Real-world psychiatric care is marked by wide heterogeneity in clinical presentations and outcomes, underscoring the need for systematic approaches to outcome measurement. The Clinical Global Impression\u2013Severity (CGI-S) scale is a brief, clinician-rated measure of overall illness severity that is widely used in psychiatric research, but rarely documented in routine care. Large language models (LLMs) may enable automated extraction of CGI-S scores from narrative clinical notes, thereby providing scalable outcome measures for real-world clinical care and research.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec sec-type=\"objective\">\n                    <jats:title>Objective<\/jats:title>\n                    <jats:p>The study aimed to evaluate whether LLMs can estimate CGI-S scores from psychiatric clinical notes for patients with major depressive disorder (MDD) and to compare performance across prompting strategies and model architectures.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec sec-type=\"methods\">\n                    <jats:title>Methods<\/jats:title>\n                    <jats:p>We extracted psychiatrist-authored notes from the Johns Hopkins electronic health record. Three board-certified psychiatrists independently rated 77 clinical notes using a validated depression-specific Clinical Global Impression (CGI) rubric. Weighted Cohen kappa coefficients were calculated to assess inter-rater reliability and model-human agreement. We evaluated GPT-4o under zero-shot and few-shot prompting conditions and Llama-4 under zero-shot prompting. Model performance was assessed by comparing LLM-generated scores to individual rater scores and consensus ratings. Exploratory analyses evaluated whether agreement varied by patient demographics, care setting, note length, or the percentage of copy-forwarded text within each note.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec sec-type=\"results\">\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>\n                      Interrater reliability among psychiatrists was high (\u03ba=0.77\u20100.78). GPT-4o with zero-shot prompting demonstrated the highest agreement with average human ratings (\u03ba=0.85, 95% CI 0.78-0.90), and few-shot prompting did not improve performance. In contrast, Llama-4 with zero-shot prompting demonstrated lower agreement with average human ratings (\u03ba=0.70, 95% CI 0.55\u20100.80). Model agreement did not significantly differ across age, sex, race, treatment location, or the percentage of copy-forwarded text, but it was significantly lower for notes below the median note length than for notes at or above the median length (\u03ba=0.72 vs 0.92;\n                      <jats:italic>P<\/jats:italic>\n                      =.003).\n                    <\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec sec-type=\"conclusions\">\n                    <jats:title>Conclusions<\/jats:title>\n                    <jats:p>LLMs can estimate clinician-rated CGI-S scores from psychiatric clinical notes for patients with MDD at a level of agreement comparable to that of expert interrater reliability. Performance varied by model architecture, with GPT-4o outperforming an open-source alternative. If further validated, this approach may support scalable outcome measurement in research settings and inform future efforts to implement measurement-based care in real-world psychiatric practice.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.2196\/86906","type":"journal-article","created":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T05:30:09Z","timestamp":1782538209000},"page":"e86906-e86906","source":"Crossref","is-referenced-by-count":0,"title":["Measuring Depression Severity With Clinical Global Impression\u2013Severity Scale Scores From Clinical Notes Using Large Language Models: Validation Study"],"prefix":"10.2196","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3004-2087","authenticated-orcid":false,"given":"Kevin","family":"Li","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8441-1741","authenticated-orcid":false,"given":"Ayah","family":"Zirikly","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-0554-9296","authenticated-orcid":false,"given":"Sarah C","family":"Collica","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6262-8264","authenticated-orcid":false,"given":"Fernando S","family":"Goes","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8540-6880","authenticated-orcid":false,"given":"Congwen","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1653-5491","authenticated-orcid":false,"given":"Trang","family":"Nguyen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4667-6607","authenticated-orcid":false,"given":"Jane P","family":"Gagliardi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5261-3632","authenticated-orcid":false,"given":"Benjamin A","family":"Goldstein","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3736-6327","authenticated-orcid":false,"given":"Hwanhee","family":"Hong","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9042-8611","authenticated-orcid":false,"given":"Elizabeth A","family":"Stuart","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8423-2623","authenticated-orcid":false,"given":"Peter P","family":"Zandi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1010","published-online":{"date-parts":[[2026,8,10]]},"reference":[{"key":"R1","doi-asserted-by":"publisher","DOI":"10.1016\/j.psychres.2024.115958","article-title":"Global, regional, and national temporal trend in burden of major depressive disorder from 1990 to 2019: an analysis of the Global Burden of Disease study","volume":"337","author":"Yan","journal-title":"Psychiatry Res"},{"key":"R2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jad.2025.120018","article-title":"Global, regional, and national burden of depression, 1990-2021: a decomposition and age-period-cohort analysis with projection to 2040","volume":"391","author":"Xu","journal-title":"J Affect Disord"},{"issue":"8","key":"R3","doi-asserted-by":"publisher","first-page":"671","DOI":"10.1176\/appi.ajp.2020.20060845","article-title":"The state of our understanding of the pathophysiology and optimal treatment of depression: glass half full or half empty?","volume":"177","author":"Nemeroff","journal-title":"Am J Psychiatry"},{"issue":"2","key":"R4","doi-asserted-by":"publisher","first-page":"167","DOI":"10.1001\/jamapsychiatry.2022.3860","article-title":"Association of treatment-resistant depression with patient outcomes and health care resource utilization in a population-wide study","volume":"80","author":"Lundberg","journal-title":"JAMA Psychiatry"},{"issue":"3","key":"R5","doi-asserted-by":"publisher","first-page":"269","DOI":"10.1002\/wps.20771","article-title":"The clinical characterization of the adult patient with depression aimed at personalization of management","volume":"19","author":"Maj","journal-title":"World Psychiatry"},{"issue":"3","key":"R6","doi-asserted-by":"publisher","first-page":"324","DOI":"10.1001\/jamapsychiatry.2018.3329","article-title":"Implementing measurement-based care in behavioral health: a review","volume":"76","author":"Lewis","journal-title":"JAMA Psychiatry"},{"key":"R7","doi-asserted-by":"publisher","DOI":"10.1016\/j.jad.2018.05.054","article-title":"The utility of PHQ-9 and CGI-S in measurement-based care for predicting suicidal ideation and behaviors","volume":"266","author":"Glazer","journal-title":"J Affect Disord"},{"issue":"5","key":"R8","doi-asserted-by":"publisher","first-page":"396","DOI":"10.1176\/appi.ps.201800383","article-title":"Systematic review of symptom assessment measures for use in measurement-based care of bipolar disorders","volume":"70","author":"Cerimele","journal-title":"Psychiatr Serv"},{"key":"R9","doi-asserted-by":"publisher","DOI":"10.1016\/j.jad.2022.09.055","article-title":"Barriers and facilitators to technology-enhanced measurement based care for depression among Canadian clinicians and patients: results of an online survey","volume":"320","author":"Cheung","journal-title":"J Affect Disord"},{"issue":"5","key":"R10","doi-asserted-by":"publisher","first-page":"456","DOI":"10.1176\/appi.ps.201900481","article-title":"Development of the National Network of Depression Centers Mood Outcomes Program: a multisite platform for measurement-based care","volume":"71","author":"Zandi","journal-title":"Psychiatr Serv"},{"issue":"7","key":"R11","first-page":"28","volume":"4","author":"Busner","journal-title":"Psychiatry (Edgmont)"},{"issue":"6","key":"R12","doi-asserted-by":"publisher","first-page":"979","DOI":"10.1111\/j.1365-2753.2007.00921.x","article-title":"The validity of the CGI severity and improvement scales as measures of clinical effectiveness suitable for routine clinical use","volume":"14","author":"Berk","journal-title":"J Eval Clin Pract"},{"key":"R13","doi-asserted-by":"publisher","DOI":"10.1186\/1471-244X-11-83","article-title":"The Clinical Global Impression scale and the influence of patient or staff perspective on outcome","volume":"11","author":"Forkmann","journal-title":"BMC Psychiatry"},{"issue":"8","key":"R14","doi-asserted-by":"publisher","first-page":"757","DOI":"10.1001\/jamapsychiatry.2024.0994","article-title":"Differential outcomes of placebo treatment across 9 psychiatric disorders: a systematic review and meta-analysis","volume":"81","author":"Bschor","journal-title":"JAMA Psychiatry"},{"key":"R15","doi-asserted-by":"publisher","DOI":"10.1016\/j.eurpsy.2018.05.006","article-title":"Patient centric measures for a patient centric era: agreement and convergent between ratings on the Patient Global Impression of Improvement (PGI-I) scale and the Clinical Global Impressions - Improvement (CGI-I) scale in bipolar and major depressive disorder","volume":"53","author":"Mohebbi","journal-title":"Eur Psychiatry"},{"issue":"2","key":"R16","doi-asserted-by":"publisher","first-page":"180","DOI":"10.1097\/NMD.0b013e3182439885","article-title":"A bidimensional solution for outcomes in bipolar disorder","volume":"200","author":"Magalh\u00e3es","journal-title":"J Nerv Ment Dis"},{"issue":"5","key":"R17","doi-asserted-by":"publisher","first-page":"386","DOI":"10.1097\/NMD.0000000000000136","article-title":"The reliability of self-assessment of affective state in different phases of bipolar disorder","volume":"202","author":"de Assis da Silva","journal-title":"J Nerv Ment Dis"},{"issue":"1","key":"R18","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1002\/da.1078","article-title":"Utility of the daily prospective National Institute of Mental Health Life-Chart Method (NIMH-LCM-p) ratings in clinical trials of bipolar disorder","volume":"15","author":"Denicoff","journal-title":"Depress Anxiety"},{"issue":"1","key":"R19","doi-asserted-by":"publisher","first-page":"117","DOI":"10.1016\/j.psychres.2010.06.021","article-title":"A video Clinical Global Impression (CGI) in obsessive compulsive disorder","volume":"186","author":"Bourredjem","journal-title":"Psychiatry Res"},{"issue":"4","key":"R20","doi-asserted-by":"publisher","first-page":"611","DOI":"10.1017\/s0033291703007414","article-title":"Evaluation of the Clinical Global Impression scale among individuals with social anxiety disorder","volume":"33","author":"Zaider","journal-title":"Psychol Med"},{"key":"R21","doi-asserted-by":"publisher","DOI":"10.1016\/j.schres.2022.03.007","article-title":"Pragmatic implementation of the Clinical Global Impression Scale of Severity as a tool for measurement-based care in a first-episode psychosis program","volume":"243","author":"Khau","journal-title":"Schizophr Res"},{"issue":"6","key":"R22","doi-asserted-by":"publisher","first-page":"458","DOI":"10.1016\/s0010-440x(99)90090-1","article-title":"Symptom correlates of global measures of severity in schizophrenia","volume":"40","author":"Goldman","journal-title":"Compr Psychiatry"},{"issue":"10","key":"R23","doi-asserted-by":"publisher","first-page":"2318","DOI":"10.1038\/sj.npp.1301147","article-title":"Linking the PANSS, BPRS, and CGI: clinical implications","volume":"31","author":"Leucht","journal-title":"Neuropsychopharmacology"},{"key":"R24","doi-asserted-by":"publisher","DOI":"10.1016\/j.jad.2016.12.041","article-title":"What does the MADRS mean? Equipercentile linking with the CGI using a company database of mirtazapine studies","volume":"210","author":"Leucht","journal-title":"J Affect Disord"},{"issue":"6","key":"R25","doi-asserted-by":"publisher","first-page":"281","DOI":"10.1097\/00004850-200211000-00003","article-title":"Relative sensitivity of the Montgomery-Asberg Depression Rating Scale, the Hamilton Depression rating scale and the Clinical Global Impressions rating scale in antidepressant clinical trials","volume":"17","author":"Khan","journal-title":"Int Clin Psychopharmacol"},{"issue":"11","key":"R26","doi-asserted-by":"publisher","first-page":"845","DOI":"10.1097\/01.nmd.0000244554.91259.27","article-title":"A comparative meta-analysis of Clinical Global Impressions change in antidepressant trials","volume":"194","author":"Spielmans","journal-title":"J Nerv Ment Dis"},{"key":"R27","doi-asserted-by":"publisher","DOI":"10.1186\/1471-244X-7-7","article-title":"The improved Clinical Global Impression scale (iCGI): development and validation in depression","volume":"7","author":"Kadouri","journal-title":"BMC Psychiatry"},{"issue":"416","key":"R28","doi-asserted-by":"publisher","first-page":"16","DOI":"10.1034\/j.1600-0447.107.s416.5.x","article-title":"The Clinical Global Impression-Schizophrenia scale: a simple instrument to measure the diversity of symptoms present in schizophrenia","author":"Haro","journal-title":"Acta Psychiatr Scand Suppl"},{"issue":"5","key":"R29","first-page":"327","volume":"13","author":"Leon","journal-title":"J Clin Psychopharmacol"},{"issue":"7","key":"R30","doi-asserted-by":"publisher","first-page":"629","DOI":"10.1002\/hup.966","article-title":"Targeted scoring criteria reduce variance in global impressions","volume":"23","author":"Targum","journal-title":"Hum Psychopharmacol"},{"issue":"4","key":"R31","doi-asserted-by":"publisher","first-page":"171","DOI":"10.1055\/s-2007-1014401","article-title":"\u201cClinical Global Impressions\u201d (ECDEU): some critical comments","volume":"25","author":"Beneke","journal-title":"Pharmacopsychiatry"},{"issue":"3","key":"R32","doi-asserted-by":"publisher","first-page":"257","DOI":"10.1016\/j.comppsych.2008.08.005","article-title":"The Clinical Global Impressions scale: errors in understanding and use","volume":"50","author":"Busner","journal-title":"Compr Psychiatry"},{"issue":"1","key":"R33","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1017\/S0033291711000997","article-title":"Using electronic medical records to enable large-scale studies in psychiatry: treatment resistant depression as a model","volume":"42","author":"Perlis","journal-title":"Psychol Med"},{"issue":"2","key":"R34","doi-asserted-by":"publisher","first-page":"405","DOI":"10.1093\/schbul\/sbaa126","article-title":"Using natural language processing on electronic health records to enhance detection and prediction of psychosis risk","volume":"47","author":"Irving","journal-title":"Schizophr Bull"},{"issue":"4","key":"R35","doi-asserted-by":"publisher","first-page":"363","DOI":"10.1176\/appi.ajp.2014.14030423","article-title":"Validation of electronic health record phenotyping of bipolar disorder cases and controls","volume":"172","author":"Castro","journal-title":"Am J Psychiatry"},{"issue":"12","key":"R36","doi-asserted-by":"publisher","first-page":"997","DOI":"10.1016\/j.biopsych.2018.01.011","article-title":"High throughput phenotyping for dimensional psychopathology in electronic health records","volume":"83","author":"McCoy","journal-title":"Biol Psychiatry"},{"issue":"10","key":"R37","doi-asserted-by":"publisher","first-page":"1064","DOI":"10.1001\/jamapsychiatry.2016.2172","article-title":"Improving prediction of suicide and accidental death after discharge from general hospitals with natural language processing","volume":"73","author":"McCoy","journal-title":"JAMA Psychiatry"},{"issue":"1","key":"R38","doi-asserted-by":"publisher","DOI":"10.1038\/s41398-021-01722-y","article-title":"Natural language processing markers in first episode psychosis and people at clinical high-risk","volume":"11","author":"Morgan","journal-title":"Transl Psychiatry"},{"key":"R39","doi-asserted-by":"publisher","DOI":"10.3389\/fpsyt.2024.1422807","article-title":"Applications of large language models in psychiatry: a systematic review","volume":"15","author":"Omar","journal-title":"Front Psychiatry"},{"key":"R40","doi-asserted-by":"publisher","DOI":"10.1016\/j.psychres.2024.116026","article-title":"Large language models in psychiatry: opportunities and challenges","volume":"339","author":"Volkmer","journal-title":"Psychiatry Res"},{"key":"R41","doi-asserted-by":"publisher","DOI":"10.2991\/jaims.d.210225.001","article-title":"Machine learning for violence risk assessment using Dutch clinical notes","volume":"2","author":"Mosteiro","journal-title":"J Artif Intell Med Sci"},{"issue":"7969","key":"R42","doi-asserted-by":"publisher","first-page":"357","DOI":"10.1038\/s41586-023-06160-y","article-title":"Health system-scale language models are all-purpose prediction engines","volume":"619","author":"Jiang","journal-title":"Nature"},{"key":"R43","doi-asserted-by":"publisher","DOI":"10.1016\/j.ajp.2024.104168","article-title":"Diagnostic accuracy of large language models in psychiatry","volume":"100","author":"Gargari","journal-title":"Asian J Psychiatr"},{"issue":"6","key":"R44","doi-asserted-by":"publisher","first-page":"347","DOI":"10.1111\/pcn.13656","article-title":"Comparing the performance of ChatGPT GPT-4, Bard, and Llama-2 in the Taiwan Psychiatric Licensing Examination and in differential diagnosis with multi-center psychiatrists","volume":"78","author":"Li","journal-title":"Psychiatry Clin Neurosci"},{"key":"R45","doi-asserted-by":"publisher","DOI":"10.1016\/j.jad.2025.04.014","article-title":"Estimating depression severity in narrative clinical notes using large language models","volume":"381","author":"McCoy","journal-title":"J Affect Disord"},{"issue":"6","key":"R46","doi-asserted-by":"publisher","first-page":"532","DOI":"10.1192\/bjp.2024.134","article-title":"Detection of suicidality from medical text using privacy-preserving large language models","volume":"225","author":"Wiest","journal-title":"Br J Psychiatry"},{"issue":"1","key":"R47","doi-asserted-by":"publisher","DOI":"10.1136\/bmjment-2025-301654","article-title":"Reasoning language models for more transparent prediction of suicide risk","volume":"28","author":"McCoy","journal-title":"BMJ Ment Health"},{"issue":"12","key":"R48","doi-asserted-by":"publisher","first-page":"940","DOI":"10.1016\/j.biopsych.2024.05.008","article-title":"Dimensional measures of psychopathology in children and adolescents using large language models","volume":"96","author":"McCoy","journal-title":"Biol Psychiatry"},{"key":"R49","unstructured":"Rotondi MA . KappaSize: sample size estimation functions for studies of interobserver agreement. The Comprehensive R Archive Network. 2018. URL: https:\/\/cran.r-project.org\/web\/packages\/kappaSize\/index.html [Accessed 27-07-2026]"},{"issue":"1","key":"R50","doi-asserted-by":"crossref","first-page":"159","DOI":"10.2307\/2529310","volume":"33","author":"Landis","journal-title":"Biometrics"},{"key":"R51","unstructured":"Gamer M Lemon J Fellows I Singh P . Irr: various coefficients of interrater reliability and agreement. The Comprehensive R Archive Network. 2026. URL: https:\/\/cran.r-project.org\/web\/packages\/irr\/index.html [Accessed 27-07-2026]"},{"key":"R52","doi-asserted-by":"publisher","DOI":"10.1016\/j.jclinepi.2017.01.013","article-title":"Spearman-Brown prophecy formula and Cronbach\u2019s alpha: different faces of reliability and opportunities for new applications","volume":"85","author":"de Vet","journal-title":"J Clin Epidemiol"},{"issue":"10","key":"R53","doi-asserted-by":"publisher","first-page":"867","DOI":"10.1001\/archpsyc.1995.03950220077014","article-title":"More reliable outcome measures can reduce sample size requirements","volume":"52","author":"Leon","journal-title":"Arch Gen Psychiatry"},{"issue":"8","key":"R54","doi-asserted-by":"publisher","first-page":"762","DOI":"10.1016\/s0006-3223(00)00837-4","article-title":"Penny-wise and pound-foolish: the impact of measurement error on sample size requirements in clinical trials","volume":"47","author":"Perkins","journal-title":"Biol Psychiatry"},{"issue":"7","key":"R55","doi-asserted-by":"publisher","DOI":"10.1001\/jamanetworkopen.2021.15334","article-title":"Length and redundancy of outpatient progress notes across a decade at an academic medical center","volume":"4","author":"Rule","journal-title":"JAMA Netw Open"},{"issue":"9","key":"R56","doi-asserted-by":"publisher","DOI":"10.1001\/jamanetworkopen.2022.33348","article-title":"Prevalence and sources of duplicate information in the electronic medical record","volume":"5","author":"Steinkamp","journal-title":"JAMA Netw Open"},{"issue":"5","key":"R57","doi-asserted-by":"publisher","first-page":"278","DOI":"10.1089\/cap.2021.0111","article-title":"A characterization of the Clinical Global Impressions scale thresholds in the treatment of adolescent depression across multiple rating scales","volume":"32","author":"Zhang","journal-title":"J Child Adolesc Psychopharmacol"}],"container-title":["JMIR Formative Research"],"original-title":[],"language":"en","deposited":{"date-parts":[[2026,8,10]],"date-time":"2026-08-10T14:24:01Z","timestamp":1786371841000},"score":1,"resource":{"primary":{"URL":"https:\/\/formative.jmir.org\/2026\/1\/e86906"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,8,10]]},"references-count":57,"URL":"https:\/\/doi.org\/10.2196\/86906","relation":{},"ISSN":["2561-326X"],"issn-type":[{"value":"2561-326X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,8,10]]},"article-number":"v10i8e86906"}}