{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T13:41:32Z","timestamp":1776087692060,"version":"3.50.1"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"5","funder":[{"name":"Shanghai Pujiang Talent Program","award":["21PJD026"],"award-info":[{"award-number":["21PJD026"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2026,9,30]]},"abstract":"<jats:p>Large Language Models (LLMs) are being increasingly applied in various natural language processing tasks including safety-critical systems (e.g., medical diagnosis querying, and code generation for self-driving), where resilience to hardware transient faults is essential for guaranteed safety. Traditional Fault Injection (FI) approaches are time consuming due to a large number of repeated executions, which limits their scalability to fast evaluate large-scale resilience. To address these challenges, we propose LLM-IARE, a novel Input-Aware Resilience Estimation Model for LLMs under hardware transient faults. It takes advantage of twice static analysis and once dynamic execution to extract critical parameters to compute general resilience metrics such as Silent Data Corruption (SDC) rates. Fast static analysis can obtain the primary LLM parameters, while dynamic execution can provide input-sensitive attention profiling to characterize how input variations influence internal attention patterns dynamically. More importantly, our proposed LLM-IARE uses the obtained parameters for modeling at three levels (operation, module, and layer) so that the SDC rates of transient fault impacts on LLMs can be calculated quickly and accurately. Additionally, LLM-IARE is further extended to estimate the LLM application-level resilience metric, the cosine similarity reflecting the bit-upset induced semantic fault impacts on final output quality. Comprehensive experiments on six representative LLMs (for example, GPT-2, T5 and RoBERTa), and 30 BERT variants demonstrate that LLM-IARE achieves a fast and accurate LLMs resilience evaluation, with up to 7335\u00d7 (average 4500\u00d7) speedup and an average logarithmic relative error of 3.94% compared with advanced LLVM-based fault injection methods. We further extend the evaluation to four larger Qwen2.5 models (0.5B\u20137B), where LLM-IARE maintains stable accuracy with logarithmic relative errors between 1.96% and 4.21% (average 3.39%).<\/jats:p>","DOI":"10.1145\/3796531","type":"journal-article","created":{"date-parts":[[2026,2,9]],"date-time":"2026-02-09T21:09:38Z","timestamp":1770671378000},"page":"1-41","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["LLM-IARE: An Input-Aware Resilience Estimation Methodology for LLMs under Hardware Transient Faults"],"prefix":"10.1145","volume":"31","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3680-787X","authenticated-orcid":false,"given":"Jiajia","family":"Jiao","sequence":"first","affiliation":[{"name":"College of Information Engineering, Shanghai Maritime University","place":["Shanghai, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-4256-3134","authenticated-orcid":false,"given":"Tainian","family":"Zhou","sequence":"additional","affiliation":[{"name":"College of Information Engineering, Shanghai Maritime University","place":["Shanghai, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-0618-2427","authenticated-orcid":false,"given":"Ran","family":"Wen","sequence":"additional","affiliation":[{"name":"College of Information Engineering, Shanghai Maritime University","place":["Shanghai, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-3240-0259","authenticated-orcid":false,"given":"Yulian","family":"Li","sequence":"additional","affiliation":[{"name":"College of Information Engineering, Shanghai Maritime University","place":["Shanghai, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-2139-4461","authenticated-orcid":false,"given":"Jin","family":"Liu","sequence":"additional","affiliation":[{"name":"College of Information Engineering, Shanghai Marine University","place":["Shanghai, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,13]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSRE55969.2022.00036"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSRE59848.2023.00052"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ETS56758.2023.10174133"},{"key":"e_1_3_2_5_2","doi-asserted-by":"crossref","unstructured":"Ahmed Alajrami and Nikolaos Aletras. 2022. How does the pre-training objective affect what large language models learn about linguistic properties?arXiv:2203.10415. Retrieved from https:\/\/arxiv.org\/abs\/2203.10415","DOI":"10.18653\/v1\/2022.acl-short.16"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/MDT.2005.69"},{"key":"e_1_3_2_7_2","unstructured":"Iz Beltagy Matthew E. Peters and Arman Cohan. 2020. Longformer: The Long-Document Transformer. arXiv:2004.05150. Retrieved from https:\/\/arxiv.org\/abs\/2004.05150"},{"key":"e_1_3_2_8_2","first-page":"1877","volume-title":"Advances in Neural Information Processing Systems","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, et\u00a0al. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems. H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 1877\u20131901. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2020\/file\/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSRE5003.2020.00047"},{"key":"e_1_3_2_10_2","doi-asserted-by":"crossref","unstructured":"Kevin Clark Urvashi Khandelwal Omer Levy and Christopher D. Manning. 2019. What Does BERT Look At? An Analysis of BERT\u2019s Attention. arXiv:1906.04341. Retrieved from https:\/\/arxiv.org\/abs\/1906.04341","DOI":"10.18653\/v1\/W19-4828"},{"key":"e_1_3_2_11_2","unstructured":"Huangliang Dai Shixun Wu Jiajun Huang Zizhe Jian Yue Zhu Haiyang Hu and Zizhong Chen. 2025. FT-Transformer: Resilient and Reliable Transformer with End-to-End Fault Tolerant Attention. arxiv:2504.02211 [cs.DC] https:\/\/arxiv.org\/abs\/2504.02211"},{"key":"e_1_3_2_12_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arxiv:1810.04805 [cs.CL] https:\/\/arxiv.org\/abs\/1810.04805"},{"key":"e_1_3_2_13_2","unstructured":"Zhangyin Feng Daya Guo Duyu Tang Nan Duan Xiaocheng Feng Ming Gong Linjun Shou Bing Qin Ting Liu Daxin Jiang et\u00a0al. 2020. CodeBERT: A Pre-Trained Model for Programming and Natural Languages. arxiv:2002.08155 [cs.CL] https:\/\/arxiv.org\/abs\/2002.08155"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSN-S58398.2023.00025"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458754"},{"key":"e_1_3_2_16_2","unstructured":"Albert Q. Jiang Alexandre Sablayrolles Antoine Roux Arthur Mensch Blanche Savary Chris Bamford Devendra Singh Chaplot Diego de las Casas Emma Bou Hanna Florian Bressand et\u00a0al. 2024. Mixtral of Experts. arxiv:2401.04088 [cs.LG] https:\/\/arxiv.org\/abs\/2401.04088"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1631\/FITEE.2400547"},{"key":"e_1_3_2_18_2","unstructured":"Mike Lewis Yinhan Liu Naman Goyal Marjan Ghazvininejad Abdelrahman Mohamed Omer Levy Ves Stoyanov and Luke Zettlemoyer. 2019. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation Translation and Comprehension. arxiv:1910.13461 [cs.CL] https:\/\/arxiv.org\/abs\/1910.13461"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","unstructured":"Guanpeng Li Siva Kumar Sastry Hari Michael Sullivan Timothy Tsai Karthik Pattabiraman Joel Emer and Stephen W. Keckler. 2017. Understanding error propagation in deep learning neural network (DNN) accelerators and applications. In Proceedings of the International Conference for High Performance Computing Networking Storage and Analysis. Association for Computing Machinery Denver Colorado. DOI:10.1145\/3126908.3126964","DOI":"10.1145\/3126908.3126964"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSN.2018.00038"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSN.2018.00016"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3710848.3710870"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/OJCS.2024.3400696"},{"key":"e_1_3_2_24_2","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arxiv:1907.11692 [cs.CL] https:\/\/arxiv.org\/abs\/1907.11692"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSN-W50199.2020.00014"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSN53405.2022.00031"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2022.3226333"},{"key":"e_1_3_2_28_2","unstructured":"Paul Michel Omer Levy and Graham Neubig. 2019. Are Sixteen Heads Really Better than One?arxiv:1905.10650 [cs.CL] https:\/\/arxiv.org\/abs\/1905.10650"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2003.1253181"},{"key":"e_1_3_2_30_2","unstructured":"Guilherme Penedo Quentin Malartic Daniel Hesslow Ruxandra Cojocaru Alessandro Cappelli Hamza Alobeidli Baptiste Pannier Ebtesam Almazrouei and Julien Launay. 2023. The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data and Web Data Only. arxiv:2306.01116 [cs.CL] https:\/\/arxiv.org\/abs\/2306.01116"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3386263.3406938"},{"issue":"8","key":"e_1_3_2_32_2","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et\u00a0al. 2019. Language models are unsupervised multitask learners. OpenAI Blog 1, 8 (2019), 9.","journal-title":"OpenAI Blog"},{"key":"e_1_3_2_33_2","unstructured":"Colin Raffel Noam Shazeer Adam Roberts Katherine Lee Sharan Narang Michael Matena Yanqi Zhou Wei Li and Peter J. Liu. 2023. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arxiv:1910.10683 [cs.LG] https:\/\/arxiv.org\/abs\/1910.10683"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.32"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/1993316.1993518"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","unstructured":"Yu Sun Zachary Coalson Shiyang Chen Hang Liu Zhao Zhang Sanghyun Hong Bo Fang and Lishan Yang. 2025. Demystifying the resilience of large language model inference: An end-to-end perspective. In Proceedings of the International Conference for High Performance Computing Networking Storage and Analysis (SC \u201925) Association for Computing Machinery 1127\u20131144. DOI:10.1145\/3712285.3759803","DOI":"10.1145\/3712285.3759803"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/DDECS57882.2023.10139468"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2023.3283685"},{"key":"e_1_3_2_39_2","unstructured":"Qwen Team. 2024. Qwen2.5: A Party of Foundation Models. Retrieved fromhttps:\/\/qwenlm.github.io\/blog\/qwen2.5\/"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.microrel.2025.115929"},{"key":"e_1_3_2_41_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard et\u00a0al. 2023. LLaMA: Open and Efficient Foundation Language Models. arxiv:2302.13971 [cs.CL] https:\/\/arxiv.org\/abs\/2302.13971"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_3_2_43_2","unstructured":"Tong Xie Jiawang Zhao Zishen Wan Zuodong Zhang Yuan Wang Runsheng Wang Ru Huang and Meng Li. 2025. ReaLM: Reliable and Efficient Large Language Model Inference with Statistical Algorithm-Based Fault Tolerance. arxiv:2503.24053 [cs.AR] https:\/\/arxiv.org\/abs\/2503.24053"},{"key":"e_1_3_2_44_2","unstructured":"Atsuki Yamaguchi George Chrysostomou Katerina Margatina and Nikolaos Aletras. 2021. Frustratingly Simple Pretraining Alternatives to Masked Language Modeling. arxiv:2109.01819 [cs.CL] https:\/\/arxiv.org\/abs\/2109.01819"},{"key":"e_1_3_2_45_2","article-title":"Qwen2 technical report","author":"Yang An","year":"2024","unstructured":"An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et\u00a0al. 2024. Qwen2 technical report. arXiv preprint arXiv:2407.10671 (2024).","journal-title":"arXiv preprint"}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3796531","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T12:41:35Z","timestamp":1776084095000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3796531"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,13]]},"references-count":44,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,9,30]]}},"alternative-id":["10.1145\/3796531"],"URL":"https:\/\/doi.org\/10.1145\/3796531","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"value":"1084-4309","type":"print"},{"value":"1557-7309","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,13]]},"assertion":[{"value":"2025-08-08","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-02","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-04-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}