{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,9,30]],"date-time":"2026-09-30T13:56:52Z","timestamp":1790776612790,"version":"4.1.0"},"reference-count":26,"publisher":"Association for Computing Machinery (ACM)","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["ACM AI Lett."],"abstract":"<jats:p>\n                    Large language models (LLMs) frequently endorse and elaborate on users\u2019 delusional beliefs, a failure mode termed\n                    <jats:italic toggle=\"yes\">psychogenicity<\/jats:italic>\n                    in the Psychosis-Bench study of Au Yeung et\u00a0al., whose framing we adopt. We present the first systematic evaluation of whether anti-sycophancy interventions transfer to psychosis-relevant contexts. Across 1,280 experiments spanning 10 conditions, 8 frontier LLMs, and 16 clinically derived Psychosis-Bench scenarios, a combined anti-sycophancy prompt reduces mean Delusion Confirmation Scores by 73.8% (paired\n                    <jats:inline-formula content-type=\"math\/tex\">\n                      <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(t(127)=11.74\\)<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    ,\n                    <jats:inline-formula content-type=\"math\/tex\">\n                      <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(p&lt;10^{-21}\\)<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    , Cohen's\n                    <jats:inline-formula content-type=\"math\/tex\">\n                      <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(d=1.04\\)<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    ). Adding a domain-general self-reflection prompt yields a 77.0% reduction and raises Safety Intervention rates by 66.9%, delivered entirely as a system prompt. Classifier-based guardrails (Llama Guard 3) flag only 5 of 3,072 evaluated turns; a reasoning guardrail (o4-mini) flags\n                    <jats:inline-formula content-type=\"math\/tex\">\n                      <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(14\\times\\)<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    more. Ablations isolating either mechanism alone plateau at\n                    <jats:inline-formula content-type=\"math\/tex\">\n                      <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\(\\approx 46\\%\\)<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    reduction, establishing anti-sycophancy prompting as a\n                    <jats:italic toggle=\"yes\">necessary foundation<\/jats:italic>\n                    that add-on mechanisms augment but cannot replace.\n                  <\/jats:p>","DOI":"10.1145\/3849709","type":"journal-article","created":{"date-parts":[[2026,9,30]],"date-time":"2026-09-30T13:07:36Z","timestamp":1790773656000},"source":"Crossref","is-referenced-by-count":0,"title":["Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-1238-1266","authenticated-orcid":false,"given":"Lorenzo","family":"de la Loza","sequence":"first","affiliation":[{"name":"College of Computing, Georgia Institute of Technology, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6539-6769","authenticated-orcid":false,"given":"Vijay K.","family":"Madisetti","sequence":"additional","affiliation":[{"name":"School of Cybersecurity and Privacy, Georgia Institute of Technology, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,9,30]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"A.\u00a0Aquilina C.\u00a0Nihalani V.\u00a0Varadarajan N.\u00a0S.\u00a0Fishbein Y.-R.\u00a0Lin and M.\u00a0Sap. 2026. Lost in Delusion: Examining LLM safety under user delusions and distress. arXiv:2606.00975."},{"key":"e_1_2_1_2_1","unstructured":"J.\u00a0Au Yeung J.\u00a0Dalmasso L.\u00a0Foschini R.\u00a0J.\u00a0B.\u00a0Dobson and Z.\u00a0Kraljevic. 2025. The Psychogenic Machine: Simulating AI Psychosis Delusion Reinforcement and Harm Enablement in Large Language Models. arXiv:2509.10970."},{"key":"e_1_2_1_3_1","unstructured":"K.\u00a0Chandra M.\u00a0Kleiman-Weiner J.\u00a0Ragan-Kelley and J.\u00a0B.\u00a0Tenenbaum. 2026. Sycophantic Chatbots Cause Delusional Spiraling Even in Ideal Bayesians. arXiv:2602.19141."},{"key":"e_1_2_1_4_1","doi-asserted-by":"crossref","unstructured":"M.\u00a0Cheng C.\u00a0Lee P.\u00a0Khadpe S.\u00a0Yu D.\u00a0Han and D.\u00a0Jurafsky. 2026. Sycophantic AI decreases prosocial intentions and promotes dependence. Science 391 6792 Article eaec8352.","DOI":"10.1126\/science.aec8352"},{"key":"e_1_2_1_5_1","volume-title":"AI-induced psychosis: Study reproduction and extensions on semantic drift, long term interactions and interventions. Course paper","author":"Chung B.","unstructured":"K.\u00a0Chung, B.\u00a0Liu, N.\u00a0Siwek, and L.\u00a0Zheng. 2025. AI-induced psychosis: Study reproduction and extensions on semantic drift, long term interactions and interventions. Course paper, Harvard University CS\u00a02881R (AI Safety), Boaz Barak, instructor."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1038\/s44220-026-00595-8"},{"key":"e_1_2_1_7_1","doi-asserted-by":"crossref","unstructured":"M.\u00a0Flathers S.\u00a0Roux and J.\u00a0Torous. 2026. Beyond artificial intelligence psychosis: a functional typology of large language model-associated psychotic phenomena. The Lancet Digital Health 100974. Online ahead of print.","DOI":"10.1016\/j.landig.2025.100974"},{"key":"e_1_2_1_8_1","volume-title":"AEGIS: Online adaptive AI content safety moderation with ensemble of LLM experts. arXiv:2404.05993.","author":"Ghosh P.","year":"2024","unstructured":"S.\u00a0Ghosh, P.\u00a0Varshney, E.\u00a0Galinkin, and C.\u00a0Parisien. 2024. AEGIS: Online adaptive AI content safety moderation with ensemble of LLM experts. arXiv:2404.05993."},{"key":"e_1_2_1_9_1","volume-title":"Proc.\u00a0ICLR.","author":"Gou Z.","year":"2024","unstructured":"Z.\u00a0Gou, Z.\u00a0Shao, Y.\u00a0Gong, Y.\u00a0Shen, Y.\u00a0Yang, N.\u00a0Duan, and W.\u00a0Chen. 2024. CRITIC: Large language models can self-correct with tool-interactive critiquing. In Proc.\u00a0ICLR."},{"key":"e_1_2_1_10_1","doi-asserted-by":"crossref","unstructured":"D.\u00a0Grabb M.\u00a0Lamparth and N.\u00a0Vasan. 2024. Risks from language models for automated mental healthcare: Ethics and structure for implementation. arXiv:2406.11852.","DOI":"10.1101\/2024.04.07.24305462"},{"key":"e_1_2_1_11_1","unstructured":"A.\u00a0Grattafiori et\u00a0al. 2024. The Llama 3 Herd of Models. arXiv:2407.21783."},{"key":"e_1_2_1_12_1","volume-title":"Deliberative Alignment: Reasoning enables safer language models. arXiv:2412.16339.","author":"Guan","year":"2024","unstructured":"M.\u00a0Y.\u00a0Guan et\u00a0al. 2024. Deliberative Alignment: Reasoning enables safer language models. arXiv:2412.16339."},{"key":"e_1_2_1_13_1","first-page":"2239","article-title":"Measuring sycophancy of language models in multi-turn dialogues","volume":"2025","author":"Hong G.","year":"2025","unstructured":"J.\u00a0Hong, G.\u00a0Byun, S.\u00a0Kim, K.\u00a0Shu, and J.\u00a0D.\u00a0Choi. 2025. Measuring sycophancy of language models in multi-turn dialogues. In Findings ACL: EMNLP 2025, 2239\u20132259.","journal-title":"Findings ACL: EMNLP"},{"key":"e_1_2_1_14_1","volume-title":"Llama Guard: LLM-based input-output safeguard for human-AI conversations. arXiv:2312.06674.","author":"Inan K.","year":"2023","unstructured":"H.\u00a0Inan, K.\u00a0Upasani, J.\u00a0Chi, R.\u00a0Rungta, K.\u00a0Iyer, Y.\u00a0Mao, M.\u00a0Tontchev, Q.\u00a0Hu, B.\u00a0Fuller, D.\u00a0Testuggine, and M.\u00a0Khabsa. 2023. Llama Guard: LLM-based input-output safeguard for human-AI conversations. arXiv:2312.06674."},{"key":"e_1_2_1_15_1","unstructured":"S.\u00a0Ji X.\u00a0Zheng J.\u00a0Sun R.\u00a0Chen W.\u00a0Gao and M.\u00a0Srivastava. 2024. MindGuard: Towards accessible and stigma-free mental health first aid via edge LLM. arXiv:2409.10064."},{"key":"e_1_2_1_16_1","doi-asserted-by":"crossref","unstructured":"E.\u00a0Kross E.\u00a0Bruehlman-Senecal J.\u00a0Park A.\u00a0Burson A.\u00a0Dougherty H.\u00a0Shablack R.\u00a0Bremner J.\u00a0Moser and O.\u00a0Ayduk. 2014. Self-talk as a regulatory mechanism: How you do it matters. J.\u00a0Personality and Social Psychology 106 2 304\u2013324.","DOI":"10.1037\/a0035173"},{"key":"e_1_2_1_17_1","unstructured":"Y.\u00a0Liu et\u00a0al. 2025. GuardReasoner: Towards reasoning-based LLM safeguards. arXiv:2501.18492."},{"key":"e_1_2_1_18_1","volume-title":"Proc.\u00a0ACM FAccT, 599\u2013627","author":"Moore D.","year":"2025","unstructured":"J.\u00a0Moore, D.\u00a0Grabb, W.\u00a0Agnew, K.\u00a0Klyman, S.\u00a0Chancellor, D.\u00a0C.\u00a0Ong, and N.\u00a0Haber. 2025. Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers. In Proc.\u00a0ACM FAccT, 599\u2013627."},{"key":"e_1_2_1_19_1","doi-asserted-by":"crossref","unstructured":"H.\u00a0Morrin L.\u00a0Nicholls M.\u00a0Levin J.\u00a0Yiend U.\u00a0Iyengar F.\u00a0DelGuidice S.\u00a0Bhattacharya S.\u00a0Tognin J.\u00a0MacCabe R.\u00a0Twumasi B.\u00a0Alderson-Day and T.\u00a0A.\u00a0Pollak. 2025. Delusions by design? How everyday AIs might be fuelling psychosis (and what can be done about it). PsyArXiv preprint.","DOI":"10.31234\/osf.io\/cmy7n_v1"},{"key":"e_1_2_1_20_1","unstructured":"M.\u00a0Y.\u00a0Guan M.\u00a0Wang M.\u00a0Carroll Z.\u00a0Dou A.\u00a0Y.\u00a0Wei J.\u00a0Huizinga I.\u00a0Kivlichan M.\u00a0Williams M.\u00a0Glaese B.\u00a0Arnav J.\u00a0Pachocki and B.\u00a0Baker. 2025. Monitoring Monitorability. arXiv:2512.18311."},{"key":"e_1_2_1_21_1","doi-asserted-by":"crossref","first-page":"1418","DOI":"10.1093\/schbul\/sbad128","article-title":"Will generative artificial intelligence chatbots generate delusions in individuals prone to psychosis","volume":"49","author":"\u00a0\u00d8stergaard","year":"2023","unstructured":"S.\u00a0D.\u00a0\u00d8stergaard. 2023. Will generative artificial intelligence chatbots generate delusions in individuals prone to psychosis? Schizophrenia Bulletin 49, 6, 1418\u20131419.","journal-title":"Schizophrenia Bulletin"},{"key":"e_1_2_1_22_1","doi-asserted-by":"crossref","first-page":"1005","DOI":"10.1093\/oxfordjournals.schbul.a007116","article-title":"Measuring delusional ideation: The 21-Item Peters et\u00a0al.\u00a0Delusions Inventory (PDI)","volume":"30","author":"Peters S.","year":"2004","unstructured":"E.\u00a0Peters, S.\u00a0Joseph, S.\u00a0Day, and P.\u00a0Garety. 2004. Measuring delusional ideation: The 21-Item Peters et\u00a0al.\u00a0Delusions Inventory (PDI). Schizophrenia Bulletin 30, 4, 1005\u20131022.","journal-title":"Schizophrenia Bulletin"},{"key":"e_1_2_1_23_1","doi-asserted-by":"crossref","unstructured":"M.\u00a0L.\u00a0Reese M.\u00a0Zeneli M.\u00a0Ng J.\u00a0Haimes A.\u00a0Damien and E.\u00a0Stade. 2026. Using LLM-as-a-judge\/jury to advance scalable clinically-validated safety evaluations of model responses to users demonstrating psychosis. arXiv:2604.02359.","DOI":"10.1609\/iaseai.v2i1.43055"},{"key":"e_1_2_1_24_1","volume-title":"Proc.\u00a0ICLR.","unstructured":"M.\u00a0Sharma et\u00a0al. 2024. Towards understanding sycophancy in language models. In Proc.\u00a0ICLR."},{"key":"e_1_2_1_25_1","doi-asserted-by":"crossref","first-page":"655","DOI":"10.1001\/jamapsychiatry.2026.0249","article-title":"Evaluation of large language model chatbot responses to psychotic prompts","volume":"83","author":"Shen F.","year":"2026","unstructured":"E.\u00a0Shen, F.\u00a0Hamati, M.\u00a0R.\u00a0Donohue, R.\u00a0R.\u00a0Girgis, J.\u00a0Veenstra-VanderWeele, and A.\u00a0Jutla. 2026. Evaluation of large language model chatbot responses to psychotic prompts. JAMA Psychiatry 83, 6, 655\u2013657.","journal-title":"JAMA Psychiatry"},{"key":"e_1_2_1_26_1","unstructured":"J.\u00a0Wei D.\u00a0Huang Y.\u00a0Lu D.\u00a0Zhou and Q.\u00a0V.\u00a0Le. 2023. Simple synthetic data reduces sycophancy in large language models. arXiv:2308.03958. Google DeepMind."}],"container-title":["ACM AI Letters"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3849709","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,9,30]],"date-time":"2026-09-30T13:07:42Z","timestamp":1790773662000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3849709"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,9,30]]},"references-count":26,"alternative-id":["10.1145\/3849709"],"URL":"https:\/\/doi.org\/10.1145\/3849709","relation":{},"ISSN":["3068-8590"],"issn-type":[{"value":"3068-8590","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,9,30]]},"article-number":"3849709"}}