{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T06:01:37Z","timestamp":1783490497555,"version":"3.55.0"},"reference-count":0,"publisher":"Association for the Advancement of Artificial Intelligence (AAAI)","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["AIES"],"abstract":"<jats:p>Despite the increasing usage of Large Language Models\n(LLMs) in answering in questions in a variety of domains,\ntheir reliability and accuracy remain unexamined for a\na plethora of domains, including the religious domains. In\nthis paper, we introduce a novel benchmark FiqhQA focused\non the LLM generated Islamic rulings explicitly categorized\nby the four major Sunni schools of thought, in both Arabic\nand English. Unlike prior work, which either overlooks the\ndistinctions between religious schools of thought or fails\nto evaluate abstention behavior, we assess LLMs not only on\ntheir accuracy but also on their ability to recognize when\nnot to answer. Our zero-shot and abstention experiments\nreveal significant variation across LLMs, languages, and\nlegal schools of thought. While GPT-4o outperforms all\nother models in accuracy, Gemini and Fanar demonstrate\nsuperior abstention behavior critical for minimizing\nconfident incorrect answers. Notably, all models exhibit a\nperformance drop in Arabic, highlighting the limitations in\nreligious reasoning for languages other than English. To\nthe best of our knowledge, this is the first study to\nbenchmark the efficacy of LLMs for fine-grained Islamic\nschool of thought specific ruling generation and to\nevaluate abstention for Islamic jurisprudence queries. Our\nfindings underscore the need for task-specific evaluation\nand cautious deployment of LLMs in religious applications.<\/jats:p>","DOI":"10.1609\/aies.v8i1.36543","type":"journal-article","created":{"date-parts":[[2025,10,15]],"date-time":"2025-10-15T13:19:04Z","timestamp":1760534344000},"page":"217-226","source":"Crossref","is-referenced-by-count":2,"title":["Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions"],"prefix":"10.1609","volume":"8","author":[{"given":"Farah","family":"Atif","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nursultan","family":"Askarbekuly","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kareem","family":"Darwish","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Monojit","family":"Choudhury","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"9382","published-online":{"date-parts":[[2025,10,15]]},"container-title":["Proceedings of the AAAI\/ACM Conference on AI, Ethics, and Society"],"original-title":[],"link":[{"URL":"https:\/\/ojs.aaai.org\/index.php\/AIES\/article\/download\/36543\/38681","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/ojs.aaai.org\/index.php\/AIES\/article\/download\/36543\/38681","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,15]],"date-time":"2025-10-15T13:19:04Z","timestamp":1760534344000},"score":1,"resource":{"primary":{"URL":"https:\/\/ojs.aaai.org\/index.php\/AIES\/article\/view\/36543"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,15]]},"references-count":0,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2025,10,15]]}},"URL":"https:\/\/doi.org\/10.1609\/aies.v8i1.36543","relation":{},"ISSN":["3065-8365"],"issn-type":[{"value":"3065-8365","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,15]]}}}