{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,20]],"date-time":"2026-03-20T11:18:36Z","timestamp":1774005516905,"version":"3.50.1"},"reference-count":51,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"DOI":"10.13039\/100020595","name":"National Science and Technology Council of Taiwan","doi-asserted-by":"crossref","award":["NSTC 114-2218-E-110-008, NSTC 113-2221-E-110 -040 -MY3, NSTC 114-2634-F-110-001-MBK, NSTC 114-2640- E-110-004"],"award-info":[{"award-number":["NSTC 114-2218-E-110-008, NSTC 113-2221-E-110 -040 -MY3, NSTC 114-2634-F-110-001-MBK, NSTC 114-2640- E-110-004"]}],"id":[{"id":"10.13039\/100020595","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2026,7,31]]},"abstract":"<jats:p>Recent advancements in artificial intelligence have led to the widespread use of convolutional neural networks (CNNs) in various fields. To improve hardware performance, numerous hardware accelerator circuits have been developed. However, the extensive use of processing elements (PEs) in these accelerators raises potential reliability concerns. Traditional modular redundancy techniques, while effective in error mitigation, come with substantial hardware costs. In this study, we introduce a cost-effective scheme designed for efficient on-line error detection and mitigation in CNNs. This scheme is based on innovative designs of approximate PEs. Specifically, we propose two potential designs for these approximate PEs and evaluate their cost-effectiveness. These approximate PEs, used in conjunction with the original PEs, form an Approximate Dual PE Redundancy (ADPR) and Approximate Triple PE Redundancy (ATPR) structure, which is essential for verifying the quality of PE computations. Additionally, we propose a novel error mitigation technique derived from our ADPR and ATPR structure, significantly enhancing the error tolerance of PEs. Notably, our approximate PE designs require only 64% of the area compared with previous designs, while maintaining similar levels of approximation errors.<\/jats:p>","DOI":"10.1145\/3793551","type":"journal-article","created":{"date-parts":[[2026,2,10]],"date-time":"2026-02-10T21:27:43Z","timestamp":1770758863000},"page":"1-34","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["A Highly Cost-Effective Online Error Detection and Mitigation Scheme for CNN Hardware Accelerators Based on Approximate PE"],"prefix":"10.1145","volume":"31","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9495-2907","authenticated-orcid":false,"given":"Wei-Ji","family":"Chao","sequence":"first","affiliation":[{"name":"National Sun Yat-sen University Department of Electrical Engineering","place":["Kaohsiung, Taiwan"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-9593-8139","authenticated-orcid":false,"given":"Yen-Chieh","family":"Tseng","sequence":"additional","affiliation":[{"name":"National Sun Yat-sen University Department of Electrical Engineering","place":["Kaohsiung, Taiwan"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7954-5569","authenticated-orcid":false,"given":"Tong-Yu","family":"Hsieh","sequence":"additional","affiliation":[{"name":"National Sun Yat-sen University Institute of Integrated Circuit Design","place":["Kaohsiung, Taiwan"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,3,20]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2016.2616357"},{"key":"e_1_3_1_3_2","article-title":"In-datacenter performance analysis of a tensor processing unit","author":"Jouppi N. P.","unstructured":"N. P. Jouppi et al. 2017. In-datacenter performance analysis of a tensor processing unit. Int. Symp. on Computer Architecture. 1--12.","journal-title":"Int. Symp. on Computer Architecture"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2020.3012753"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/VTS.2018.8368656"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/DFT52944.2021.9568340"},{"key":"e_1_3_1_7_2","first-page":"1","article-title":"FAT: Training neural networks for reliable inference under hardware faults","author":"Zahid U.","unstructured":"U. Zahid, G. Gambardella, N. J. Fraser, M. Blott and K. Vissers2020. FAT: Training neural networks for reliable inference under hardware faults. IEEE Int'l. Test Conf. 1\u201310.","journal-title":"IEEE Int'l. Test Conf."},{"key":"e_1_3_1_8_2","first-page":"211","article-title":"FTT-NAS: Discovering fault-tolerant neural architecture","author":"Li W.","unstructured":"W. Li, X. Ning, G. Ge, X. Chen, Y. Wang and H. Yang. 2020. FTT-NAS: Discovering fault-tolerant neural architecture. Asia and South Pacific Design Automation Conf. 211\u2013216.","journal-title":"Asia and South Pacific Design Automation Conf"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126964"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE48585.2020.9116571"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSN48987.2021.00018"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE54114.2022.9774635"},{"issue":"2164","key":"e_1_3_1_13_2","first-page":"1","article-title":"Salvagednn: Salvaging deep neural network accelerators with permanent faults through saliency-driven fault-aware mapping","volume":"378","author":"Abdullah Hanif M.","unstructured":"M. Abdullah Hanif and M. Shafique. 2020. Salvagednn: Salvaging deep neural network accelerators with permanent faults through saliency-driven fault-aware mapping. Philos. Trans. Roy. Soc. A 378, 2164 (2020), 1\u201323.","journal-title":"Philos. Trans. Roy. Soc. A"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CAHPC.2018.8645906"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TDMR.2022.3159089"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-56258-2_18"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNS.2018.2884460"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/MDAT.2022.3174181"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11265-021-01642-6"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCD56317.2022.00086"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/PRDC.2017.24"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCKE54056.2021.9721442"},{"key":"e_1_3_1_23_2","first-page":"1","volume-title":"Symposium on Integrated Circuits and Systems Design","author":"Gomes I. A. C.","unstructured":"I. A. C. Gomes and F. G. L. Kastensmidt. 2013. Reducing TMR overhead by combining approximate circuit, transistor topology and input permutation approaches. Symposium on Integrated Circuits and Systems Design. 1\u20136."},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSD.2016.57"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TDMR.2017.2781186"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.7873\/DATE.2015.0024"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2017.2653780"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2017.2761740"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1049\/cdt2.12016"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/LSSC.2019.2954780"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2021.3138491"},{"key":"e_1_3_1_32_2","first-page":"512","article-title":"Computer multiplication and division using binary logarithms","author":"Mitchell J. N.","unstructured":"J. N. Mitchell et al. 1962. Computer multiplication and division using binary logarithms. IRE Transactions on Electronic Computers. 512\u2013517.","journal-title":"IRE Transactions on Electronic Computers"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCSS51193.2021.9464180"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE.2019.8714868"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ARITH.2016.25"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSRE52982.2021.00025"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3065386"},{"key":"e_1_3_1_39_2","unstructured":"K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556. Retrieved from https:\/\/arxiv.org\/abs\/1409.1556"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_41_2","unstructured":"The MNIST DATABAS Networks. [Online]. Retrieved from http:\/\/yann.lecun.com\/exdb\/mnist\/"},{"key":"e_1_3_1_42_2","unstructured":"Alex Krizhevsky Networks. [Online]. Retrieved from https:\/\/www.cs.toronto.edu\/\u223ckriz\/cifar.html"},{"key":"e_1_3_1_43_2","unstructured":"Stanford Vision Lab Networks. [Online]. Retrieved from https:\/\/image-net.org\/index.php"},{"key":"e_1_3_1_44_2","first-page":"1","article-title":"Cost-effective error-mitigation for high memory error rate of DNN: A case study on YOLOv4","author":"Chao W.-J.","unstructured":"W.-J. Chao and T.-Y. Hsieh. 2023. Cost-effective error-mitigation for high memory error rate of DNN: A case study on YOLOv4. IEEE International Test Conference in Asia. 1\u20136.","journal-title":"IEEE International Test Conference in Asia"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/DSN-W50199.2020.00014"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2020.3032495"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3445814.3446747"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/LATS53581.2021.9651807"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/DFT50435.2020.9250866"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01318"},{"key":"e_1_3_1_51_2","unstructured":"Alom Md Zahangir et al. The history began from AlexNet: A comprehensive survey on deep learning approaches. arXiv:1803.01164. Retrieved from https:\/\/arxiv.org\/abs\/1803.01164"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISOCC66390.2025.11329963"}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3793551","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,20]],"date-time":"2026-03-20T09:12:11Z","timestamp":1773997931000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3793551"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,20]]},"references-count":51,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2026,7,31]]}},"alternative-id":["10.1145\/3793551"],"URL":"https:\/\/doi.org\/10.1145\/3793551","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"value":"1084-4309","type":"print"},{"value":"1557-7309","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,20]]},"assertion":[{"value":"2025-06-12","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-08","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-20","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}