{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,9,20]],"date-time":"2026-09-20T12:21:14Z","timestamp":1789906874557,"version":"4.0.1"},"reference-count":80,"publisher":"Elsevier BV","license":[{"start":{"date-parts":[[2026,2,1]],"date-time":"2026-02-01T00:00:00Z","timestamp":1769904000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.elsevier.com\/tdm\/userlicense\/1.0\/"},{"start":{"date-parts":[[2026,2,1]],"date-time":"2026-02-01T00:00:00Z","timestamp":1769904000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.elsevier.com\/legal\/tdmrep-license"},{"start":{"date-parts":[[2025,11,3]],"date-time":"2025-11-03T00:00:00Z","timestamp":1762128000000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["elsevier.com","sciencedirect.com"],"crossmark-restriction":true},"short-container-title":["Safety Science"],"published-print":{"date-parts":[[2026,2]]},"DOI":"10.1016\/j.ssci.2025.107056","type":"journal-article","created":{"date-parts":[[2025,11,12]],"date-time":"2025-11-12T20:19:33Z","timestamp":1762978773000},"page":"107056","update-policy":"https:\/\/doi.org\/10.1016\/elsevier_cm_policy","source":"Crossref","is-referenced-by-count":24,"special_numbering":"C","title":["From hallucinations to hazards: benchmarking LLMs for hazard analysis in safety-critical systems"],"prefix":"10.1016","volume":"194","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3765-5217","authenticated-orcid":false,"given":"Ioannis M.","family":"Dokas","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"78","reference":[{"key":"10.1016\/j.ssci.2025.107056_b0005","unstructured":"Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Man\u00e9, D. (2016). Concrete problems in AI safety. arXiv preprint arXiv:1606.06565. https:\/\/arxiv.org\/abs\/1606.06565."},{"key":"10.1016\/j.ssci.2025.107056_b0015","unstructured":"Analytics Vidhya. (2025, March). 14 popular LLM benchmarks to know in 2025. https:\/\/www.analyticsvidhya.com\/blog\/2025\/03\/llm-benchmarks\/."},{"key":"10.1016\/j.ssci.2025.107056_b0010","series-title":"In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies","first-page":"5444","article-title":"RAG LLMs are not safer: a safety analysis of retrieval-augmented generation for large language models","author":"An","year":"2025"},{"key":"10.1016\/j.ssci.2025.107056_b0025","first-page":"2425","article-title":"VQA: Visual question answering","volume":"2015","author":"Antol","year":"2015","journal-title":"Proceedings of the IEEE International Conference on Computer Vision"},{"key":"10.1016\/j.ssci.2025.107056_b0030","unstructured":"Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Kernion, J., Ndousse, K., Olsson, C., Amodei, D., Brown, T., Clark, J., \u2026 Kaplan, J. (2021). A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861. https:\/\/arxiv.org\/abs\/2112.00861."},{"key":"10.1016\/j.ssci.2025.107056_b0035","doi-asserted-by":"crossref","unstructured":"Bang, Y., Cahyawijaya, S., Lee, N., Dai, W., Su, D., Weller, B., & Fung, P. (2023). A multitask, multilingual, multimodal evaluation of ChatGPT on reasoning, hallucination, and interactivity. In J. C. Park, Y. Arase, B. Hu, W. Lu, D. Wijaya, A. Purwarianti, & A. A. Krisnadhi (Eds.), Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 675\u2013718). Association for Computational Linguistics.","DOI":"10.18653\/v1\/2023.ijcnlp-main.45"},{"key":"10.1016\/j.ssci.2025.107056_b0040","unstructured":"Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Liang, P., \u2026 & others. (2022). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258v3. https:\/\/arxiv.org\/abs\/2108.07258."},{"key":"10.1016\/j.ssci.2025.107056_b0045","doi-asserted-by":"crossref","unstructured":"Borg, M., Schuberth, K., & Robertz, A. (2022). Can language models resolve requirements conflicts? In Requirements Engineering: Foundation for Software Quality (REFSQ 2022) (Lecture Notes in Computer Science, Vol. 13216, pp. 164\u2013179). Springer. Doi: 10.1007\/978-3-030-98464-9_13.","DOI":"10.1007\/978-3-030-98464-9_13"},{"issue":"2","key":"10.1016\/j.ssci.2025.107056_b0050","doi-asserted-by":"crossref","first-page":"281","DOI":"10.18280\/ijsse.150219","article-title":"Risk assessment using a structured combined method","volume":"15","author":"Bouasla","year":"2025","journal-title":"International Journal of Safety and Security Engineering"},{"key":"10.1016\/j.ssci.2025.107056_b0055","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown","year":"2020","journal-title":"Adv. Neural Inf. Proces. Syst."},{"key":"10.1016\/j.ssci.2025.107056_b0060","article-title":"Functional hazard analysis (FHA) as basis for safety requirement definition: the ETCS case study","author":"Calixto","year":"2022","journal-title":"European Federation of National Maintenance Societies."},{"key":"10.1016\/j.ssci.2025.107056_b0070","unstructured":"Chambers, L. (2005). A hazard analysis of human factors in safety-critical systems engineering. In T. Cant (Ed.), Conferences in Research and Practice in Information Technology, Vol. 55. Safety Critical Systems (SCS'05) (pp. 27-41). Australian Computer Society, Inc."},{"issue":"3","key":"10.1016\/j.ssci.2025.107056_b0065","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3641289","article-title":"A survey on evaluation of large language models","volume":"15","author":"Chang","year":"2023","journal-title":"ACM Trans. Intell. Syst. Technol."},{"key":"10.1016\/j.ssci.2025.107056_b0075","doi-asserted-by":"crossref","DOI":"10.1016\/j.ssci.2024.106608","article-title":"Hazard analysis in the era of AI: Assessing the usefulness of ChatGPT4 in STPA hazard analysis","volume":"178","author":"Charalampidou","year":"2024","journal-title":"Saf. Sci."},{"key":"10.1016\/j.ssci.2025.107056_b0080","unstructured":"Chen, M., Tworek, J., Jun, H., Yuan, Q., Ponde de Oliveira Pinto, H., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., \u2026 Zaremba, W. (2021). Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374. https:\/\/arxiv.org\/abs\/2107.03374."},{"key":"10.1016\/j.ssci.2025.107056_b0085","unstructured":"Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., Zhang, H., Zhu, B., Yang, M., Jordan, M. I., & Gonzalez, J. E. (2024). Chatbot Arena: An open platform for evaluating LLMs by human preference. In Proceedings of the Twelfth International Conference on Learning Representations (ICLR 2024)."},{"key":"10.1016\/j.ssci.2025.107056_b0090","unstructured":"Chollet, F. (2019). On the measure of intelligence. arXiv preprint arXiv:1911.01547. https:\/\/arxiv.org\/abs\/1911.01547."},{"key":"10.1016\/j.ssci.2025.107056_b0100","unstructured":"Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., & Schulman, J. (2021). Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168. https:\/\/arxiv.org\/abs\/2110.14168."},{"key":"10.1016\/j.ssci.2025.107056_b0105","doi-asserted-by":"crossref","unstructured":"Collier, Z. A., Gruss, R. J., & Abrahams, A. S. (2024). How good are large language models at product risk assessment? Risk Analysis: An Official Publication of the Society for Risk Analysis. Advance online publication. Doi: 10.1111\/risa.14351.","DOI":"10.1111\/risa.14351"},{"key":"10.1016\/j.ssci.2025.107056_b0115","unstructured":"Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., & Kaiser, \u0141. (2019). Universal transformers. In International Conference on Learning Representations (ICLR 2019)."},{"key":"10.1016\/j.ssci.2025.107056_b0120","series-title":"The cybernetic teammate: a field experiment on generative AI reshaping teamwork and expertise","first-page":"25","author":"Dell\u2019Acqua","year":"2025"},{"key":"10.1016\/j.ssci.2025.107056_b0125","series-title":"International Conference on Computer Safety, Reliability, and Security","first-page":"410","article-title":"Can large language models assist in hazard analysis?","author":"Diemert","year":"2023"},{"key":"10.1016\/j.ssci.2025.107056_b0130","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1016\/j.ssci.2013.03.013","article-title":"EWaSAP: an early warning sign identification approach based on a systemic hazard analysis","volume":"58","author":"Dokas","year":"2013","journal-title":"Saf. Sci."},{"key":"10.1016\/j.ssci.2025.107056_b0140","series-title":"Hazard analysis techniques for system safety","author":"Ericson","year":"2015"},{"key":"10.1016\/j.ssci.2025.107056_b0150","doi-asserted-by":"crossref","unstructured":"Fang, M., Liu, X., Wang, Z., Wang, Z., Wu, J., & Zhu, Y. (2024). MathOdyssey: Benchmarking mathematical problem-solving skills in large language models using Odyssey Math data. arXiv preprint arXiv:2406.18321. https:\/\/arxiv.org\/abs\/2406.18321.","DOI":"10.1038\/s41597-025-05283-3"},{"key":"10.1016\/j.ssci.2025.107056_b0155","unstructured":"Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., ... & Clark, J. (2022). Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv preprint arXiv:2209.07858."},{"key":"10.1016\/j.ssci.2025.107056_b0160","doi-asserted-by":"crossref","unstructured":"Gehman, S., Gururangan, S., Sap, M., Choi, Y., & Smith, N. A. (2020). RealToxicityPrompts: Evaluating neural toxic degeneration in language models. In Findings of the Association for Computational Linguistics: EMNLP 2020 (pp. 3356\u20133369). Association for Computational Linguistics. Doi: 10.18653\/v1\/2020.findings-emnlp.301.","DOI":"10.18653\/v1\/2020.findings-emnlp.301"},{"key":"10.1016\/j.ssci.2025.107056_b0165","series-title":"Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency","first-page":"2570","article-title":"June). understanding and mitigating risks of generative AI in financial services","author":"Gehrmann","year":"2025"},{"key":"10.1016\/j.ssci.2025.107056_b0170","doi-asserted-by":"crossref","first-page":"346","DOI":"10.1162\/tacl_a_00370","article-title":"Did Aristotle use a laptop? a question answering benchmark with implicit reasoning strategies","volume":"9","author":"Geva","year":"2021","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"10.1016\/j.ssci.2025.107056_b0175","series-title":"2025 IEEE International Conference on Software Quality, Reliability, and Security (QRS)","article-title":"Automated generation of defeaters for assurance cases","author":"Gohar","year":"2025"},{"key":"10.1016\/j.ssci.2025.107056_b0180","unstructured":"Golovneva, O., et al. (2023). ROSCOE: A suite of metrics for scoring step-by-step reasoning. In The Eleventh International Conference on Learning Representations, ICLR 2023."},{"issue":"2","key":"10.1016\/j.ssci.2025.107056_b0185","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1111\/j.1471-1842.2009.00848.x","article-title":"A typology of reviews: an analysis of 14 review types and associated methodologies","volume":"26","author":"Grant","year":"2009","journal-title":"Health Information and Libraries Journal"},{"key":"10.1016\/j.ssci.2025.107056_b0190","unstructured":"Graydon, M. S., & Lehman, S. M. (2025). Examining proposed uses of LLMs to produce or assess assurance arguments. NASA Technical Memorandum TM-20250001849, NASA Langley Research Center."},{"key":"10.1016\/j.ssci.2025.107056_b0195","doi-asserted-by":"crossref","unstructured":"Guha, N., Nyarko, J., Ho, D. E., R\u00e9, C., Chilton, A., Narayana, A., Chohlas-Wood, A., Peters, A., Waldon, B., Rockmore, D. N., Zambrano, D. A., Talisman, D., Hoque, E., Surani, F., Fagan, F., Sarfaty, G., Dickinson, G. M., Porat, H., Hegland, J., \u2026 Li, Z. (2023). LegalBench: A collaboratively built benchmark for measuring legal reasoning in large language models. In Advances in Neural Information Processing Systems 36 (NeurIPS 2023) (pp. 44123\u201344279).","DOI":"10.2139\/ssrn.4583531"},{"key":"10.1016\/j.ssci.2025.107056_b0200","doi-asserted-by":"crossref","unstructured":"Hartvigsen, T., Gabriel, S., Palangi, H., Sap, M., Shahid, A., & Kamar, E. (2022). ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL 2022) (Vol. 1, pp. 3309\u20133326). Association for Computational Linguistics. Doi: 10.18653\/v1\/2022.acl-long.234.","DOI":"10.18653\/v1\/2022.acl-long.234"},{"key":"10.1016\/j.ssci.2025.107056_b0205","series-title":"Proceedings of the International Conference on Learning Representations (ICLR). Arxiv Preprint arXiv:2009.03300","article-title":"Measuring massive multitask language understanding","author":"Hendrycks","year":"2021"},{"key":"10.1016\/j.ssci.2025.107056_b0210","article-title":"ChatGPT is bullshit","volume":"26","author":"Hicks","year":"2024","journal-title":"Ethics and Information Technology"},{"key":"10.1016\/j.ssci.2025.107056_b0215","unstructured":"Hu, J., Ruder, S., Siddhant, A., Neubig, G., Firat, O., & Johnson, M. (2020, July 13\u201318). XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation. In H. Daum\u00e9 III & A. Singh (Eds.), Proceedings of the 37th International Conference on Machine Learning (Vol. 119, pp. 4411\u20134421). Proceedings of Machine Learning Research. https:\/\/proceedings.mlr.press\/v119\/hu20b.htm."},{"key":"10.1016\/j.ssci.2025.107056_b0220","unstructured":"Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2024). SWE-bench: Can language models resolve real-world GitHub issues? In Proceedings of the Twelfth International Conference on Learning Representations (ICLR 2024). https:\/\/openreview.net\/forum?id=VTF8yNQM66."},{"key":"10.1016\/j.ssci.2025.107056_b0225","doi-asserted-by":"crossref","unstructured":"Jin, Q., Dhingra, B., Liu, Z., Cohen, W. W., & Lu, X. (2019). PubMedQA: A dataset for biomedical research question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 2567\u20132577). Association for Computational Linguistics. Doi: 10.18653\/v1\/D19-1259.","DOI":"10.18653\/v1\/D19-1259"},{"key":"10.1016\/j.ssci.2025.107056_b0230","unstructured":"Joshi, A., Heimdahl, M. P. E., Miller, S. P., & Whalen, M. W. (2006). Model-based safety analysis (NASA\/CR-2006-213953). National Aeronautics and Space Administration, Langley Research Center."},{"key":"10.1016\/j.ssci.2025.107056_b0235","doi-asserted-by":"crossref","unstructured":"Khastgir, S., Birrell, S., Dhadyalla, G., Sivencrona, H., & Jennings, P. (2017). Towards increased reliability by objectification of Hazard Analysis and Risk Assessment (HARA) of automated automotive systems. Safety Science, 99(Part B), 166\u2013177. Doi: 10.1016\/j.ssci.2017.03.02.","DOI":"10.1016\/j.ssci.2017.03.024"},{"key":"10.1016\/j.ssci.2025.107056_b0240","series-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT)","first-page":"4110","article-title":"Dynabench: Rethinking benchmarking in NLP","author":"Kiela","year":"2021"},{"key":"10.1016\/j.ssci.2025.107056_b0245","unstructured":"Lad, J. (2025). Beyond the scoreboard: The past, present, and future of LLM benchmarks. Medium. https:\/\/medium.com\/@lad.jai\/beyond-the-scoreboard-the-past-present-and-future-of-llm-benchmarks-df4e2edb2bdf."},{"key":"10.1016\/j.ssci.2025.107056_b0250","series-title":"Engineering a safer world: Systems thinking applied to safety","author":"Leveson","year":"2012"},{"key":"10.1016\/j.ssci.2025.107056_b0255","series-title":"Partnership for Systems Approaches to Safety and Security","article-title":"STPA handbook","author":"Leveson","year":"2018"},{"key":"10.1016\/j.ssci.2025.107056_b0265","series-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","article-title":"API-Bank: a comprehensive benchmark for tool-augmented LLMs","author":"Li","year":"2023"},{"key":"10.1016\/j.ssci.2025.107056_b0270","doi-asserted-by":"crossref","unstructured":"Lin, S., Hilton, J., & Evans, O. (2022). TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 3214\u20133252). Association for Computational Linguistics. Doi: 10.18653\/v1\/2022.acl-long.229.","DOI":"10.18653\/v1\/2022.acl-long.229"},{"key":"10.1016\/j.ssci.2025.107056_b0275","unstructured":"Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.-W., Zhu, S.-C., Tafjord, O., Clark, P., & Kalyan, A. (2022). Learn to explain: Multimodal reasoning via thought chains for science question answering. In Advances in Neural Information Processing Systems 35 (NeurIPS 2022) (pp. 2507\u20132521)."},{"key":"10.1016\/j.ssci.2025.107056_b0280","series-title":"Introduction to Information Retrieval","author":"Manning","year":"2008"},{"key":"10.1016\/j.ssci.2025.107056_b0285","unstructured":"McDermid, J. A., Jia, Y., & Habli, I. (2019). Towards a framework for safety assurance of autonomous systems. In H. Espinoza, H. Yu, X. Huang, F. Lecue, C. Chen, J. Hernandez-Orallo, S. \u00d3 h\u00c9igeartaigh, & R. Mallah (Eds.), Artificial Intelligence Safety 2019 (Vol. 2419, pp. 1\u20137). CEUR Workshop Proceedings."},{"key":"10.1016\/j.ssci.2025.107056_b0290","series-title":"In the Thirteenth International Conference on Learning Representations","article-title":"GSM-Symbolic: Understanding the limitations of mathematical reasoning in large language models","author":"Mirzadeh","year":"2025"},{"key":"10.1016\/j.ssci.2025.107056_b0295","unstructured":"Mitchell, M. S., & Winner, D. R. (2010). Subsystem hazard analysis methodology for the Ares I upper stage source controlled items (Report No. 20100022159). National Aeronautics and Space Administration."},{"key":"10.1016\/j.ssci.2025.107056_b0300","unstructured":"NASA. (2014). NASA system safety handbook: Volume 2: System safety concepts, guidelines, and implementation examples (NASA\/SP-2014-612). National Aeronautics and Space Administration."},{"key":"10.1016\/j.ssci.2025.107056_b0305","doi-asserted-by":"crossref","unstructured":"Nouri, A., Cabrero-Daniel, B., T\u00f6rner, F., Sivencrona, H., & Berger, C. (2024). Engineering safety requirements for autonomous driving with large language models. Accepted by 32nd IEEE International Requirements Engineering Conference, 2024.","DOI":"10.1109\/RE59067.2024.00029"},{"key":"10.1016\/j.ssci.2025.107056_b0310","doi-asserted-by":"crossref","unstructured":"Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002). BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (ACL 2002) (pp. 311\u2013318). Association for Computational Linguistics. Doi: 10.3115\/1073083.1073135.","DOI":"10.3115\/1073083.1073135"},{"issue":"2","key":"10.1016\/j.ssci.2025.107056_b0315","doi-asserted-by":"crossref","first-page":"183","DOI":"10.1016\/j.im.2014.08.008","article-title":"Synthesizing information systems knowledge: a typology of literature reviews","volume":"52","author":"Par\u00e9","year":"2015","journal-title":"Information & Management"},{"key":"10.1016\/j.ssci.2025.107056_b0320","unstructured":"Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., & Bowman, S. R. (2023). GPQA: A graduate\u2011level Google\u2011proof Q&A benchmark. In Proceedings of the First Conference on Language Modeling (COLM 2024)."},{"key":"10.1016\/j.ssci.2025.107056_b0325","article-title":"Systematic Derivation of Defeaters for Security Arguments","author":"Shahandashti","year":"2019","journal-title":"YorkSpace"},{"key":"10.1016\/j.ssci.2025.107056_b0330","doi-asserted-by":"crossref","first-page":"333","DOI":"10.1016\/j.jbusres.2019.07.039","article-title":"Literature review as a research methodology: an overview and guidelines","volume":"104","author":"Snyder","year":"2019","journal-title":"Journal of Business Research"},{"key":"10.1016\/j.ssci.2025.107056_b0335","first-page":"34188","article-title":"Llm-check: investigating detection of hallucinations in large language models","volume":"37","author":"Sriramanan","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"10.1016\/j.ssci.2025.107056_b0340","unstructured":"Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., Kluska, A., Lewkowycz, A., Agarwal, A., Wu, A., & Bowman, S. R. (2023). Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research."},{"key":"10.1016\/j.ssci.2025.107056_b0345","doi-asserted-by":"crossref","unstructured":"Su, W., Wang, C., Ai, Q., Hu, Y., Wu, Z., Zhou, Y., & Liu, Y. (2024). Unsupervised real-time hallucination detection based on the internal states of large language models.arXiv preprint arXiv:2403.06448.","DOI":"10.18653\/v1\/2024.findings-acl.854"},{"key":"10.1016\/j.ssci.2025.107056_b0350","doi-asserted-by":"crossref","unstructured":"Suzgun, M., Scales, N., Sch\u00e4rli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., & Wei, J. (2022). Challenging BIG\u2011Bench tasks and whether chain\u2011of\u2011thought can solve them. In Findings of the Association for Computational Linguistics: ACL 2023. Association for Computational Linguistics. Doi: 10.18653\/v1\/2023.findings-acl.824.","DOI":"10.18653\/v1\/2023.findings-acl.824"},{"issue":"5","key":"10.1016\/j.ssci.2025.107056_b0355","article-title":"Safety-critical failure analysis of industrial automotive airbag system using FMEA and FTA techniques","volume":"5","author":"Swarup","year":"2014","journal-title":"Int. J. Adv. Res. Comput. Sci."},{"key":"10.1016\/j.ssci.2025.107056_b0360","unstructured":"Unite.AI. (2025, April). 5 best large language models (LLMs) in April 2025. https:\/\/www.unite.ai\/best-large-language-models-llms\/."},{"key":"10.1016\/j.ssci.2025.107056_b0365","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., & Polosukhin, I. (2023). Attention is all you need. Attention is all you need. In Advances in Neural Information Processing Systems 30 (NIPS 2017)."},{"key":"10.1016\/j.ssci.2025.107056_b0375","series-title":"2024 IEEE International Conference on Assured Autonomy (ICAA)","article-title":"Are we there yet? Assessing the maturity of defeater generation from assurance cases","author":"Viger","year":"2024"},{"key":"10.1016\/j.ssci.2025.107056_b0380","first-page":"3266","article-title":"SuperGLUE: a stickier benchmark for general-purpose language understanding systems","volume":"32","author":"Wang","year":"2019","journal-title":"Advances in Neural Information Processing Systems"},{"key":"10.1016\/j.ssci.2025.107056_b0385","article-title":"GLUE: a multi-task benchmark and analysis platform for natural language understanding","author":"Wang","year":"2019","journal-title":"In International Conference on Learning Representations."},{"key":"10.1016\/j.ssci.2025.107056_b0390","unstructured":"Wang, B., Chen, W., Pei, H., Xie, C., Kang, M., Zhang, C., ... & Li, B. (2023). Decodingtrust: A comprehensive assessment of trustworthiness in gpt models, 2024.URL https:\/\/arxiv. org\/abs\/2306.11698."},{"key":"10.1016\/j.ssci.2025.107056_b0395","unstructured":"Wasson, K., Neogi, N., Graydon, M., Maddalon, J., Miner, P., & McCormick, G. F. (2022). Functional hazard assessment for the eVTOL aircraft supporting urban air mobility (UAM) applications: Exploratory demonstrations (NASA\/TM\u201320210024234). National Aeronautics and Space Administration."},{"key":"10.1016\/j.ssci.2025.107056_b0400","article-title":"Emergent abilities of large language models","author":"Wei","year":"2022","journal-title":"Transactions on Machine Learning Research."},{"key":"10.1016\/j.ssci.2025.107056_b0405","doi-asserted-by":"crossref","unstructured":"Wu, Z., Qiu, L., Ross, A., Aky\u00fcrek, E., Chen, B., Wang, B., Kim, N., Andreas, J., & Kim, Y. (2024). Reasoning or reciting? Exploring the capabilities and limitations of language models through counterfactual tasks. Reasoning or reciting? Exploring the capabilities and limitations of language models through counterfactual tasks. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL 2024), Volume 1: Long Papers (pp. 1819\u20131862).","DOI":"10.18653\/v1\/2024.naacl-long.102"},{"key":"10.1016\/j.ssci.2025.107056_b0410","unstructured":"Yan, F., Mao, H., Ji, C. C.-J., Zhang, T., Patil, S. G., Stoica, I., & Gonzalez, J. E. (2024). Berkeley function calling leaderboard. https:\/\/gorilla.cs.berkeley.edu."},{"key":"10.1016\/j.ssci.2025.107056_b0415","unstructured":"Yin, S., Pang, X., Ding, Y., Chen, M., Bi, Y., Xiong, Y., Huang, W., Xiang, Z., Shao, J., & Chen, S. (2025). SafeAgentBench: A benchmark for safe task planning of embodied LLM agents. arXiv preprint arXiv:2412.13178. https:\/\/arxiv.org\/abs\/2412.13178."},{"key":"10.1016\/j.ssci.2025.107056_b0420","unstructured":"Yue, Z., Wang, S., Zhou, Z., Narasimhan, K., & Chang, K.-W. (2024). MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert knowledge. arXiv preprint arXiv:2311.16502v4. https:\/\/arxiv.org\/abs\/2311.16502."},{"issue":"4","key":"10.1016\/j.ssci.2025.107056_b0425","doi-asserted-by":"crossref","first-page":"85","DOI":"10.3390\/safety10040085","article-title":"Reducing Data Uncertainties: Fuzzy Real-Time Safety Level Methodology for Socio-Technical Systems","volume":"10","author":"Zeleskidis","year":"2024","journal-title":"Safety"},{"key":"10.1016\/j.ssci.2025.107056_b0430","doi-asserted-by":"crossref","unstructured":"Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., & Choi, Y. (2019, July). HellaSwag: Can a machine really finish your sentence? In A. Korhonen, D. Traum, & L. M\u00e0rquez (Eds.), Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 4791\u20134800). Association for Computational Linguistics. Doi: 10.18653\/v1\/P19-1472.","DOI":"10.18653\/v1\/P19-1472"},{"key":"10.1016\/j.ssci.2025.107056_b0435","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Lei, L., Wu, L., Sun, R., Huang, Y., Long, C., ... & Huang, M. (2024). SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 9546-9562). Association for Computational Linguistics.","DOI":"10.18653\/v1\/2024.acl-long.830"}],"container-title":["Safety Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/api.elsevier.com\/content\/article\/PII:S0925753525002814?httpAccept=text\/xml","content-type":"text\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/api.elsevier.com\/content\/article\/PII:S0925753525002814?httpAccept=text\/plain","content-type":"text\/plain","content-version":"vor","intended-application":"text-mining"}],"deposited":{"date-parts":[[2025,11,22]],"date-time":"2025-11-22T11:23:31Z","timestamp":1763810611000},"score":1,"resource":{"primary":{"URL":"https:\/\/linkinghub.elsevier.com\/retrieve\/pii\/S0925753525002814"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2]]},"references-count":80,"alternative-id":["S0925753525002814"],"URL":"https:\/\/doi.org\/10.1016\/j.ssci.2025.107056","relation":{},"ISSN":["0925-7535"],"issn-type":[{"value":"0925-7535","type":"print"}],"subject":[],"published":{"date-parts":[[2026,2]]},"assertion":[{"value":"Elsevier","name":"publisher","label":"This article is maintained by"},{"value":"From hallucinations to hazards: benchmarking LLMs for hazard analysis in safety-critical systems","name":"articletitle","label":"Article Title"},{"value":"Safety Science","name":"journaltitle","label":"Journal Title"},{"value":"https:\/\/doi.org\/10.1016\/j.ssci.2025.107056","name":"articlelink","label":"CrossRef DOI link to publisher maintained version"},{"value":"article","name":"content_type","label":"Content Type"},{"value":"\u00a9 2025 The Author(s). Published by Elsevier Ltd.","name":"copyright","label":"Copyright"}],"article-number":"107056"}}