{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,2]],"date-time":"2026-08-02T10:42:30Z","timestamp":1785667350058,"version":"3.56.0"},"reference-count":80,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2022,6,26]],"date-time":"2022-06-26T00:00:00Z","timestamp":1656201600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,6,26]],"date-time":"2022-06-26T00:00:00Z","timestamp":1656201600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001691","name":"japan society for the promotion of science","doi-asserted-by":"publisher","award":["JP20H00590"],"award-info":[{"award-number":["JP20H00590"]}],"id":[{"id":"10.13039\/501100001691","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Artif Intell Rev"],"published-print":{"date-parts":[[2023,3]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Deep neural networks (DNNs) have made significant achievements in a wide variety of domains. For the deep learning tasks, multiple excellent hardware platforms provide efficient solutions, including graphics processing units (GPUs), central processing units (CPUs), field programmable gate arrays (FPGAs), and application-specific integrated circuit (ASIC). Nonetheless, CPUs outperform other solutions including GPUs in many cases for the inference workload of DNNs with the support of various techniques, such as the high-performance libraries being the basic building blocks for DNNs. Thus, CPUs have been a preferred choice for DNN inference applications, particularly in the low-latency demand scenarios. However, the DNN inference efficiency remains a critical issue, especially when low latency is required under conditions with limited hardware resources, such as embedded systems. At the same time, the hardware features have not been fully exploited for DNNs and there is much room for improvement. To this end, this paper conducts a series of experiments to make a thorough study for the inference workload of prominent state-of-the-art DNN architectures on a single-instruction-multiple-data (SIMD) CPU platform, as well as with widely applicable scopes for multiple hardware platforms. The study goes into depth in DNNs: the CPU kernel-instruction level performance characteristics of DNNs including branches, branch prediction misses, cache misses, etc, and the underlying convolutional computing mechanism at the SIMD level; The thorough layer-wise time consumption details with potential time-cost bottlenecks; And the exhaustive dynamic activation sparsity with exact details on the redundancy of DNNs. The research provides researchers with comprehensive and insightful details, as well as crucial target areas for optimising and improving the efficiency of DNNs at both the hardware and software levels.<\/jats:p>","DOI":"10.1007\/s10462-022-10221-5","type":"journal-article","created":{"date-parts":[[2022,6,26]],"date-time":"2022-06-26T11:02:41Z","timestamp":1656241361000},"page":"1971-2010","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":65,"title":["An architecture-level analysis on deep learning models for low-impact computations"],"prefix":"10.1007","volume":"56","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4112-7297","authenticated-orcid":false,"given":"Hengyi","family":"Li","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhichen","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xuebin","family":"Yue","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenwen","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hiroyuki","family":"Tomiyama","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lin","family":"Meng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,6,26]]},"reference":[{"key":"10221_CR1","unstructured":"Agarap Abien\u00a0Fred (2018) Deep learning using rectified linear units (relu). CoRR, abs\/1803.08375"},{"issue":"3","key":"10221_CR2","doi-asserted-by":"publisher","first-page":"656","DOI":"10.1109\/TCAD.2021.3064914","volume":"41","author":"B Ahn","year":"2022","unstructured":"Ahn B, Kim T (2022) Deeper weight pruning without accuracy loss in deep neural networks: Signed-digit representation-based approach. IEEE Trans Comput Aided Des Integr Circuits Syst 41(3):656\u2013668","journal-title":"IEEE Trans Comput Aided Des Integr Circuits Syst"},{"key":"10221_CR3","unstructured":"Bochkovskiy Alexey, Wang Chien-Yao, Liao Hong-Yuan\u00a0Mark (2020) Yolov4: Optimal speed and accuracy of object detection,"},{"key":"10221_CR4","doi-asserted-by":"crossref","unstructured":"Camci E, Gupta M, Wu M, Lin J (2022) Qlp: Deep q-learning for pruning deep neural networks. IEEE Transactions on Circuits and Systems for Video Technology","DOI":"10.1109\/TCSVT.2022.3167951"},{"key":"10221_CR5","doi-asserted-by":"crossref","unstructured":"Cardoso VB, Oliveira AS, Forechi A, Azevedo P, Mutz F, Oliveira-Santos T, Badue C, De Souza AF (2020) A large-scale mapping method based on deep neural networks applied to self-driving car localization, pp 1\u20138","DOI":"10.1109\/IJCNN48605.2020.9207449"},{"key":"10221_CR6","doi-asserted-by":"crossref","unstructured":"Cardoso Jo\u00e3o\u00a0M.P, Coutinho Jos\u00e9 Gabriel\u00a0F, Diniz Pedro\u00a0C (2017) Chapter 2 - high-performance embedded computing. In Embedded Computing for High Performance, pages 17\u201356","DOI":"10.1016\/B978-0-12-804189-5.00002-8"},{"key":"10221_CR7","doi-asserted-by":"crossref","unstructured":"Carion Nicolas, Massa Francisco, Synnaeve Gabriel, Usunier Nicolas, Kirillov Alexander, Zagoruyko Sergey (2020) End-to-end object detection with transformers. In: Computer Vision \u2013 ECCV 2020, pages 213\u2013229, Cham","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"10221_CR8","doi-asserted-by":"publisher","first-page":"106","DOI":"10.1109\/TASLP.2020.3036783","volume":"29","author":"L Chai","year":"2021","unstructured":"Chai L, Jun D, Liu Q-F, Lee C-H (2021) A cross-entropy-guided measure (cegm) for assessing speech recognition performance and optimizing dnn-based speech enhancement. IEEE\/ACM Trans Audio Speech Lang Process 29:106\u2013117","journal-title":"IEEE\/ACM Trans Audio Speech Lang Process"},{"issue":"2","key":"10221_CR9","doi-asserted-by":"publisher","first-page":"799","DOI":"10.1109\/TNNLS.2020.2979517","volume":"32","author":"Z Chen","year":"2021","unstructured":"Chen Z, Ting-Bing X, Changde D, Liu C-L, He H (2021) Dynamical channel pruning by conditional accuracy change for deep neural networks. IEEE Trans Neural Networks Learn Syst 32(2):799\u2013813","journal-title":"IEEE Trans Neural Networks Learn Syst"},{"issue":"5","key":"10221_CR10","doi-asserted-by":"publisher","first-page":"1747","DOI":"10.1109\/TNNLS.2019.2927224","volume":"31","author":"K Chen","year":"2020","unstructured":"Chen K, Yao L, Zhang D, Wang X, Chang X, Nie F (2020) A semisupervised recurrent convolutional attention model for human activity recognition. IEEE Trans Neural Networks Learn Syst 31(5):1747\u20131756","journal-title":"IEEE Trans Neural Networks Learn Syst"},{"key":"10221_CR11","doi-asserted-by":"crossref","unstructured":"Chen Z (2021) Application of artificial intelligence in the inheritance and innovation of excellent traditional chinese culture. In: 2021 International Conference on Computer Information Science and Artificial Intelligence (CISAI) pp\u00a0665\u2013668, Kunming,China, 09","DOI":"10.1109\/CISAI54367.2021.00134"},{"key":"10221_CR12","unstructured":"Chen T, Ji B, Ding T, Fang B, Wang G, Zhu Z, Liang L, Shi Y, Yi S, Tu X (2021) Only train once: A one-shot neural network training and pruning framework. In: Advances in Neural Information Processing Systems, volume\u00a034, pages 19637\u201319651. Curran Associates, Inc.,"},{"key":"10221_CR13","doi-asserted-by":"crossref","unstructured":"Chernikova A, Oprea A, Nita-Rotaru C, Kim BG (2019) Are self-driving cars secure? evasion attacks against deep neural networks for steering angle prediction. In: Security and Privacy Workshops (SPW), pp 132\u2013137","DOI":"10.1109\/SPW.2019.00033"},{"key":"10221_CR14","doi-asserted-by":"crossref","unstructured":"Chung S-H, Abdelrahman TS (2020) Optimizing opencl kernels and runtime for dnn inference on fpgas. In: 2020 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), pp\u00a0151\u2013154, New Orleans, LA, USA","DOI":"10.1109\/IPDPSW50202.2020.00034"},{"key":"10221_CR15","unstructured":"Facebook\u00a0Investor Relations. \u201cfacebook reports first quarter 2018 results. 2018"},{"key":"10221_CR16","doi-asserted-by":"crossref","unstructured":"Flautner K, Uhlig R, Reinhardt S, Mudge T (2000) Thread-level parallelism and interactive performance of desktop applications. In: ACM SIGARCH Computer Architecture News, vol\u00a028, pp 129\u2013138, Online, December","DOI":"10.1145\/378995.379233"},{"key":"10221_CR17","first-page":"1","volume":"16","author":"D Fong","year":"2017","unstructured":"Fong D, Motamedi M, Ghiasi S (2017) Machine intelligence on resource-constrained iot devices: The case of thread granularity optimization for cnn inference. ACM Transactions on Embedded Computing Systems 16:1\u201319","journal-title":"ACM Transactions on Embedded Computing Systems"},{"key":"10221_CR18","doi-asserted-by":"crossref","unstructured":"Fujikawa Y, Li H, Yue X, Aravinda CV, Amar PG, Meng L (2022) Recognition of oracle bone inscriptions by using two deep learning models. Int J Digit Hum","DOI":"10.1007\/s42803-022-00044-9"},{"issue":"2","key":"10221_CR19","doi-asserted-by":"publisher","first-page":"652","DOI":"10.1109\/TPAMI.2019.2938758","volume":"43","author":"S-H Gao","year":"2021","unstructured":"Gao S-H, Cheng M-M, Zhao K, Zhang X-Y, Yang M-H, Torr P (2021) Res2net: a new multi-scale backbone architecture. IEEE Trans Pattern Anal Mach Intell 43(2):652\u2013662","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"10221_CR20","doi-asserted-by":"crossref","unstructured":"Goel A, Tung C, Lu Y-H, Thiruvathukal GK (2020) A survey of methods for low-power deep learning and computer vision. In: 2020 IEEE 6th World Forum on Internet of Things (WF-IoT), pp\u00a01\u20136","DOI":"10.1109\/WF-IoT48130.2020.9221198"},{"issue":"5","key":"10221_CR21","doi-asserted-by":"publisher","first-page":"696","DOI":"10.1109\/TC.2020.2995593","volume":"70","author":"C Gong","year":"2021","unstructured":"Gong C, Chen Y, Ye L, Li T, Hao C, Chen D (2021) Vecq: minimal loss dnn model compression with vectorized weight quantization. IEEE Trans Comput 70(5):696\u2013710","journal-title":"IEEE Trans Comput"},{"key":"10221_CR22","doi-asserted-by":"crossref","unstructured":"Gong Zhangxiaowen, Ji Houxiang, Fletcher Christopher\u00a0W, Hughes Christopher\u00a0J, Baghsorkhi Sara, Torrellas Josep (2020) Save: Sparsity-aware vector engine for accelerating dnn training and inference on cpus. In: 2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO), pages 796\u2013810","DOI":"10.1109\/MICRO50266.2020.00070"},{"issue":"7","key":"10221_CR23","doi-asserted-by":"publisher","first-page":"931","DOI":"10.1109\/TC.2020.2981080","volume":"69","author":"Y Guan","year":"2020","unstructured":"Guan Y, Sun G, Yuan Z, Li X, Ningyi X, Chen S, Cong J, Xie Y (2020) Crane: Mitigating accelerator under-utilization caused by sparsity irregularities in cnns. IEEE Trans Comput 69(7):931\u2013943","journal-title":"IEEE Trans Comput"},{"key":"10221_CR24","doi-asserted-by":"crossref","unstructured":"Guo Cong, Hsueh Bo\u00a0Yang, Leng Jingwen, Qiu Yuxian, Guan Yue, Wang Zehuan, Jia Xiaoying, Li Xipeng, Guo Minyi, Zhu Yuhao (2020) Accelerating sparse dnn models without hardware-support via tile-wise sparsity. In: SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1\u201315,","DOI":"10.1109\/SC41405.2020.00020"},{"key":"10221_CR25","doi-asserted-by":"crossref","unstructured":"Han K, Wang Y, Tian Q, Guo J, Xu C, Xu C (2020) Ghostnet: More features from cheap operations. In: 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1577\u20131586, Seattle, WA, USA, jun","DOI":"10.1109\/CVPR42600.2020.00165"},{"key":"10221_CR26","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 770\u2013778, Los Alamitos, CA, USA, June","DOI":"10.1109\/CVPR.2016.90"},{"key":"10221_CR27","unstructured":"Howard AG, Zhu M, Chen B, Kalenichenko D, Wang W, Weyand T, Andreetto M, Adam H (2017) Mobilenets: Efficient convolutional neural networks for mobile vision applications. In: CoRR, abs\/1704.04861"},{"key":"10221_CR28","doi-asserted-by":"crossref","unstructured":"Huang G, Liu Z, Van\u00a0Der Maaten L, Weinberger KQ (2017) Densely connected convolutional networks. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 2261\u20132269, Los Alamitos, CA, USA, July","DOI":"10.1109\/CVPR.2017.243"},{"key":"10221_CR29","unstructured":"Intel. Intel oneAPI Programming Guide, 2020"},{"key":"10221_CR30","unstructured":"Intel. Intel 64 and IA-32 ArchitecturesSoftware Developer\u2019s Manual, (2021)"},{"key":"10221_CR31","unstructured":"Ioffe S, Szegedy C (2015) Batch normalization: Accelerating deep network training by reducing internal covariate shift. The 32nd International Conference on Machine Learning (ICML). volume 37. Lille, France, pp 448\u2013456"},{"issue":"1","key":"10221_CR32","doi-asserted-by":"publisher","first-page":"15","DOI":"10.1109\/LES.2021.3087707","volume":"14","author":"EJ Jeong","year":"2022","unstructured":"Jeong EJ, Kim J, Tan S, Lee J, Ha S (2022) Deep learning inference parallelization on heterogeneous processors with tensorrt. IEEE Embed Syst Lett 14(1):15\u201318","journal-title":"IEEE Embed Syst Lett"},{"issue":"2","key":"10221_CR33","doi-asserted-by":"publisher","first-page":"397","DOI":"10.1109\/TVLSI.2020.3041517","volume":"29","author":"D Ji","year":"2021","unstructured":"Ji D, Shin D, Park J (2021) An error compensation technique for low-voltage dnn accelerators. IEEE Trans Very Large Scale Integr Syst 29(2):397\u2013408","journal-title":"IEEE Trans Very Large Scale Integr Syst"},{"key":"10221_CR34","unstructured":"Jongsoo Park, Maxim Naumov, Protonu Basu, Summer Deng, Aravind Kalaiah,Daya\u00a0Shanker Khudia, James Law, Parth Malani, Andrey Malevich, NadathurSatish, Juan\u00a0Miguel Pino, Martin Schatz, Alexander Sidorov, ViswanathSivakumar, Andrew Tulloch, Xiaodong Wang, Yiming Wu, Hector Yuen, Utku Diril,Dmytro Dzhulgakov, Kim\u00a0M. Hazelwood, Bill Jia, Yangqing Jia, Lin Qiao, VijayRao, Nadav Rotem, Sungjoo Yoo, and Mikhail Smelyanskiy.Deep learning inference in facebook data centers: Characterization,performance optimizations and hardware implications.CoRR, 2018"},{"key":"10221_CR35","doi-asserted-by":"publisher","first-page":"211422","DOI":"10.1109\/ACCESS.2020.3039278","volume":"8","author":"Y Kim","year":"2020","unstructured":"Kim Y, Kong J, Munir A (2020) Cpu-accelerator co-scheduling for cnn acceleration at the edge. IEEE Access 8:211422\u2013211433","journal-title":"IEEE Access"},{"key":"10221_CR36","doi-asserted-by":"crossref","unstructured":"Kim H, Nam H, Jung W, Lee J (2017) Performance analysis of cnn frameworks for gpus. In: The 2017 IEEE International Symposium on Performance Analysis of Systems and Software, pp\u00a055\u201364, Santa Rosa, CA, USA, April","DOI":"10.1109\/ISPASS.2017.7975270"},{"key":"10221_CR37","doi-asserted-by":"crossref","unstructured":"Kim Youngsok, Kim Joonsung, Chae Dongju, Kim Daehyun, Kim Jangwoo (2019) layer: Low latency on-device inference using cooperative single-layer acceleration and processor-friendly quantization. pages 1\u201315, 03","DOI":"10.1145\/3302424.3303950"},{"key":"10221_CR38","unstructured":"Krizhevsky Alex , Sutskever Ilya, Hinton Geoffrey\u00a0E (2012) Imagenet classification with deep convolutional neural networks. In: Advances in Neural Information Processing Systems, volume\u00a025, Lake Tahoe, Nevada, USA, December . Curran Associates, Inc"},{"key":"10221_CR39","doi-asserted-by":"crossref","unstructured":"Kwon H, Lai L, Pellauer M, Krishna T, Chen Y-H, Chandra V (2021) Heterogeneous dataflow accelerators for multi-dnn workloads. In: 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), pp\u00a071\u201383","DOI":"10.1109\/HPCA51647.2021.00016"},{"key":"10221_CR40","doi-asserted-by":"crossref","unstructured":"Lee G, Park H, Ryu S, Lee H-J (2020) Acceleration of dnn training regularization: Dropout accelerator. In: ICEIC2020, pp 1\u20132","DOI":"10.1109\/ICEIC49074.2020.9051194"},{"key":"10221_CR41","doi-asserted-by":"crossref","unstructured":"Li J, Liang W, Li Y, Xu Z, Jia X, Guo S (2021) Throughput maximization of delay-aware dnn inference in edge computing by exploring dnn model partitioning and inference parallelism. IEEE Transactions on Mobile Computing, pp 1","DOI":"10.1109\/TMC.2021.3125949"},{"key":"10221_CR42","doi-asserted-by":"crossref","unstructured":"Li Hengyi, Wang Zhichen, Yue Xuebin, Wang Wenwen, Hiroyuki Tomiyama, Meng Lin (2021) A comprehensive analysis of low-impact computations in deep learning workloads. In: Proceedings of the 2021 on Great Lakes Symposium on VLSI, pages 385\u2013390, New York, USA, June","DOI":"10.1145\/3453688.3461747"},{"key":"10221_CR43","doi-asserted-by":"publisher","first-page":"1016","DOI":"10.1109\/TASLP.2021.3133209","volume":"30","author":"Y-C Lin","year":"2022","unstructured":"Lin Y-C, Cheng Yu, Hsu Y-T, Szu-Wei F, Tsao Yu, Kuo T-W (2022) Seofp-net: Compression and acceleration of deep neural networks for speech enhancement using sign-exponent-only floating-points. IEEE\/ACM Tran Audio Speech Lang Process 30:1016\u20131031","journal-title":"IEEE\/ACM Tran Audio Speech Lang Process"},{"key":"10221_CR44","unstructured":"Liu Y, Wang Y, Yu R, Li M, Sharma V, Wang Y (2019) Optimizing CNN model inference on cpus. In Proceedings of the 2019 USENIX Conference on Usenix Annual Technical Conference, pages 1025\u20131040, Berkeley, CA, USA, July"},{"key":"10221_CR45","doi-asserted-by":"crossref","unstructured":"Low Cheng-Yaw, Park Jaewoo, Teoh Andrew Beng-Jin (2020) Stacking-based deep neural network: Deep analytic network for pattern classification. IEEE Transactions on Cybernetics, 50(12):5021\u20135034","DOI":"10.1109\/TCYB.2019.2908387"},{"key":"10221_CR46","doi-asserted-by":"crossref","unstructured":"Ma N et al (2018) Shufflenet V2: practical guidelines for efficient CNN architecture design. The 15th European Conference on Computer Vision (ECCV). volume 11218. Munich, Germany, pp 122\u2013138","DOI":"10.1007\/978-3-030-01264-9_8"},{"key":"10221_CR47","doi-asserted-by":"crossref","unstructured":"Ma Yun, Xiang Dongwei, Zheng Shuyu, Tian Deyu, Liu Xuanzhe (2019) Moving deep learning into web browser: How far can we go? pages 1234\u20131244, 05","DOI":"10.1145\/3308558.3313639"},{"issue":"4","key":"10221_CR48","doi-asserted-by":"publisher","first-page":"532","DOI":"10.1109\/JETCAS.2021.3129415","volume":"11","author":"AN Mazumder","year":"2021","unstructured":"Mazumder AN, Meng J, Rashid H-A, Kallakuri U, Zhang X, Seo J-S, Mohsenin T (2021) A survey on the optimization of neural network accelerators for micro-ai on-device inference. IEEE J Emerg Sel Top Circ Syst 11(4):532\u2013547","journal-title":"IEEE J Emerg Sel Top Circ Syst"},{"key":"10221_CR49","doi-asserted-by":"publisher","first-page":"9102","DOI":"10.1109\/ACCESS.2020.2964608","volume":"8","author":"V Mazzia","year":"2020","unstructured":"Mazzia V, Khaliq A, Salvetti F, Chiaberge M (2020) Real-time apple detection system using embedded systems with hardware accelerators: An edge ai application. IEEE Access 8:9102\u20139114","journal-title":"IEEE Access"},{"key":"10221_CR50","doi-asserted-by":"publisher","first-page":"17880","DOI":"10.1109\/ACCESS.2018.2820326","volume":"6","author":"L Meng","year":"2018","unstructured":"Meng L, Hirayama T, Oyanagi S (2018) Underwater-drone with panoramic camera for automatic fish recognition based on deep learning. IEEE Access 6:17880\u201317886","journal-title":"IEEE Access"},{"key":"10221_CR51","unstructured":"Meng Lin et\u00a0al. (2012) A novel branch predictor using local history for miss-prediction. The 2012 International Conference on Computer Design, pages 77\u201383"},{"key":"10221_CR52","doi-asserted-by":"publisher","DOI":"10.1016\/j.sysarc.2019.101635","volume":"99","author":"S Mittal","year":"2019","unstructured":"Mittal S, Vaishay S (2019) A survey of techniques for optimizing deep learning on gpus. J Syst Architect 99:101635","journal-title":"J Syst Architect"},{"key":"10221_CR53","unstructured":"Mittal Sparsh, Rajput Poonam, Subramoney Sreenivas (2021) A survey of deep learning on cpus: Opportunities and co-optimizations. IEEE Transactions on Neural Networks and Learning Systems, pages 1\u201321"},{"key":"10221_CR54","doi-asserted-by":"crossref","unstructured":"Nasiri Nasibeh, Segal Oren, Margala Martin (2014) Modified fused multiply-accumulate chained unit. In: 2014 IEEE 57th International Midwest Symposium on Circuits and Systems (MWSCAS), pages 889\u2013892,","DOI":"10.1109\/MWSCAS.2014.6908558"},{"issue":"2","key":"10221_CR55","doi-asserted-by":"publisher","first-page":"604","DOI":"10.1109\/TNNLS.2020.2979670","volume":"32","author":"DW Otter","year":"2021","unstructured":"Otter DW, Medina JR, Kalita JK (2021) A survey of the usages of deep learning for natural language processing. IEEE Trans Neural Networks Learn Syst 32(2):604\u2013624","journal-title":"IEEE Trans Neural Networks Learn Syst"},{"issue":"2","key":"10221_CR56","doi-asserted-by":"publisher","first-page":"1737","DOI":"10.1109\/LRA.2021.3060442","volume":"6","author":"R Ozaki","year":"2021","unstructured":"Ozaki R, Kuroda Y (2021) Ekf-based real-time self-attitude estimation with camera dnn learning landscape regularities. IEEE Robot Autom Lett 6(2):1737\u20131744","journal-title":"IEEE Robot Autom Lett"},{"key":"10221_CR57","doi-asserted-by":"crossref","unstructured":"Patel K, Mistry C, Mehta D, Thakker U, Tanwar S, Gupta R, Kumar N (2022) A survey on artificial intelligence techniques for chronic diseases: Open issues and challenges. Artif Intell Rev, pp 1\u201344, 06","DOI":"10.1007\/s10462-021-10084-2"},{"key":"10221_CR58","unstructured":"Patel Keyur, Mistry Chinmay, Mehta Dev, Thakker Urvish, Tanwar Sudeep, Gupta Rajesh, Kumar Neeraj (2022) Ai on the edge: a comprehensive review. Artificial Intelligence Review, 03"},{"issue":"2","key":"10221_CR59","doi-asserted-by":"publisher","first-page":"572","DOI":"10.1109\/JBHI.2021.3098662","volume":"26","author":"T Pokaprakarn","year":"2022","unstructured":"Pokaprakarn T, Kitzmiller RR, Moorman R, Lake DE, Krishnamurthy AK, Kosorok MR (2022) Sequence to sequence ecg cardiac rhythm classification using convolutional recurrent neural networks. IEEE J Biomed Health Inform 26(2):572\u2013580","journal-title":"IEEE J Biomed Health Inform"},{"issue":"11","key":"10221_CR60","doi-asserted-by":"publisher","first-page":"2293","DOI":"10.1109\/TCAD.2020.3046568","volume":"40","author":"M de Prado","year":"2021","unstructured":"de Prado M, Mundy A, Saeed R, Denna M, Pazos N, Benini L (2021) Automated design space exploration for optimized deployment of dnn on arm cortex-a cpus. IEEE Trans Comput Aided Des Integr Circuits Syst 40(11):2293\u20132305","journal-title":"IEEE Trans Comput Aided Des Integr Circuits Syst"},{"key":"10221_CR61","doi-asserted-by":"crossref","unstructured":"Putro Muhamad\u00a0Dwisnanto, Kurnianggoro Laksono, Jo Kang-Hyun (2021) High performance and efficient real-time face detector on central processing unit based on convolutional neural network. IEEE Transactions on Industrial Informatics, 17(7):4449\u20134457","DOI":"10.1109\/TII.2020.3022501"},{"key":"10221_CR62","unstructured":"Ren Shaoqing, He Kaiming, Girshick Ross, Sun Jian (2015) Faster r-cnn: Towards real-time object detection with region proposal networks. In C.\u00a0Cortes, N.\u00a0Lawrence, D.\u00a0Lee, M.\u00a0Sugiyama, and R.\u00a0Garnett, editors, Advances in Neural Information Processing Systems, volume\u00a028. Curran Associates, Inc.,"},{"key":"10221_CR63","doi-asserted-by":"crossref","unstructured":"Sebastian A, Boybat I, Dazzi M, Giannopoulos I, Jonnalagadda V, Joshi V, Karunaratne G, Kersting B, Khaddam-Aljameh R, Nandakumar SR, Petropoulos A, Piveteau C, Antonakopoulos T, Rajendran B, Le Gallo M, Eleftheriou E (2019) Computational memory-based inference and training of deep neural networks. In: 2019 Symposium on VLSI Technology, pages T168\u2013T169, Kyoto, Japan, June","DOI":"10.23919\/VLSIT.2019.8776518"},{"key":"10221_CR64","unstructured":"Simonyan K, Zisserman A (2015) Very deep convolutional networks for large-scale image recognition. In: The 3rd International Conference on Learning Representations(ICLR), San Diego, CA, USA, May"},{"key":"10221_CR65","doi-asserted-by":"crossref","unstructured":"Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, Erhan D, Vanhoucke V, Rabinovich A (2015) Going deeper with convolutions. In: 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1\u20139, Boston, MA, USA, June","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"10221_CR66","doi-asserted-by":"crossref","unstructured":"Szegedy C, Vanhoucke V, Ioffe S, Shlens J, Wojna Z (2016) Rethinking the inception architecture for computer vision. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 2818\u20132826, Las Vegas, NV, USA, June","DOI":"10.1109\/CVPR.2016.308"},{"key":"10221_CR67","doi-asserted-by":"crossref","unstructured":"Szegedy C, Ioffe S, Vanhoucke V (2016) Inception-v4, inception-resnet and the impact of residual connections on learning. In: Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, pages 4278\u20134284, San Francisco California USA, Feburary","DOI":"10.1609\/aaai.v31i1.11231"},{"key":"10221_CR68","doi-asserted-by":"crossref","unstructured":"Szegedy Christian, Ioffe Sergey, Vanhoucke Vincent, Alemi Alexander\u00a0A (2017) Inception-v4, inception-resnet and the impact of residual connections on learning. In: Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, AAAI\u201917, page 4278-4284. AAAI Press,","DOI":"10.1609\/aaai.v31i1.11231"},{"key":"10221_CR69","doi-asserted-by":"crossref","unstructured":"Tan M, Chen B, Pang R, Vasudevan V, Le QV (2019) Mnasnet: Platform-aware neural architecture search for mobile. In: 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 2815\u20132823, Long Beach, CA, USA, June","DOI":"10.1109\/CVPR.2019.00293"},{"key":"10221_CR70","unstructured":"Wu Carole-Jean et al. (2019) Machine learning at facebook: Understanding inference at the edge. In: 2019 IEEE International Symposium on High Performance Computer Architecture (HPCA), pages 331\u2013344, Washington, DC, USA, February"},{"issue":"1","key":"10221_CR71","doi-asserted-by":"publisher","first-page":"94","DOI":"10.1109\/TCBB.2020.2986544","volume":"18","author":"Yu Xiang","year":"2021","unstructured":"Xiang Yu, Kang C, Guttery DS, Kadry S, Chen Y, Zhang Y-D (2021) Resnet-scda-50 for breast abnormality classification. IEEE\/ACM Trans Comput Biol Bioinf 18(1):94\u2013102","journal-title":"IEEE\/ACM Trans Comput Biol Bioinf"},{"issue":"1","key":"10221_CR72","doi-asserted-by":"publisher","first-page":"246","DOI":"10.1109\/TCSVT.2020.2975566","volume":"31","author":"J Xie","year":"2021","unstructured":"Xie J, He N, Fang L, Ghamisi P (2021) Multiscale densely-connected fusion networks for hyperspectral images classification. IEEE Trans Circuits Syst Video Technol 31(1):246\u2013259","journal-title":"IEEE Trans Circuits Syst Video Technol"},{"key":"10221_CR73","doi-asserted-by":"crossref","unstructured":"Yang L, Jiang H, Cai R, Wang Y, Song S, Huang G, Tian Q (2021) Condensenet v2: Sparse feature reactivation for deep networks. In: 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 3568\u20133577","DOI":"10.1109\/CVPR46437.2021.00357"},{"key":"10221_CR74","doi-asserted-by":"crossref","unstructured":"Yu Y, Beyls K, D\u2019Hollander EH (2001) Visualizing the impact of the cache on program execution. In: Proceedings Fifth International Conference on Information Visualisation, pages 336\u2013341, London, UK, July","DOI":"10.1109\/IV.2001.942079"},{"key":"10221_CR75","doi-asserted-by":"crossref","unstructured":"Yue X, Li H, Shimizu M, Kawamura S (2022) and Lin Meng. A deep learning-based object detection algorithm for empty-dish recycling robots. Machines, Yolo-gd","DOI":"10.23919\/ASCC56756.2022.9828060"},{"key":"10221_CR76","doi-asserted-by":"crossref","unstructured":"Yunyang Xiong, Hanxiao Liu, Suyog Gupta, Berkin Akin, Gabriel Bender, Yongzhe Wang, Pieter-Jan Kindermans, Mingxing Tan, Vikas Singh, and Bo\u00a0Chen. Mobiledets: Searching for object detection architectures for mobile accelerators. In 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3824\u20133833, Nashville, TN, USA, 2021","DOI":"10.1109\/CVPR46437.2021.00382"},{"issue":"12","key":"10221_CR77","doi-asserted-by":"publisher","first-page":"3901","DOI":"10.1109\/TMI.2021.3101616","volume":"40","author":"X Zhang","year":"2021","unstructured":"Zhang X, Han Z, Shangguan H, Han X, Cui X, Wang A (2021) Artifact and detail attention generative adversarial networks for low-dose ct denoising. IEEE Trans Med Imaging 40(12):3901\u20133918","journal-title":"IEEE Trans Med Imaging"},{"issue":"7","key":"10221_CR78","doi-asserted-by":"publisher","first-page":"3033","DOI":"10.1109\/TCYB.2019.2905157","volume":"50","author":"D Zhang","year":"2020","unstructured":"Zhang D, Yao L, Chen K, Wang S, Chang X, Liu Y (2020) Making sense of spatio-temporal preserving representations for eeg-based human intention recognition. IEEE Trans Cybern 50(7):3033\u20133044","journal-title":"IEEE Trans Cybern"},{"key":"10221_CR79","doi-asserted-by":"crossref","unstructured":"Zhang , ZX, Lin M, Sun J (2018) Shufflenet: an extremely efficient convolutional neural network for mobile devices. In: 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 6848\u20136856, Salt Lake City","DOI":"10.1109\/CVPR.2018.00716"},{"key":"10221_CR80","doi-asserted-by":"crossref","unstructured":"Zhou X et al. (2018)Cambricon-s: Addressing irregularity in sparse neural networks through a cooperative software\/hardware approach. In: Proc. 51st IEEE\/ACM Int. Symp. Microarchitecture, pages 15\u201328","DOI":"10.1109\/MICRO.2018.00011"}],"container-title":["Artificial Intelligence Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-022-10221-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10462-022-10221-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-022-10221-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,9,27]],"date-time":"2024-09-27T20:41:19Z","timestamp":1727469679000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10462-022-10221-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,26]]},"references-count":80,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,3]]}},"alternative-id":["10221"],"URL":"https:\/\/doi.org\/10.1007\/s10462-022-10221-5","relation":{},"ISSN":["0269-2821","1573-7462"],"issn-type":[{"value":"0269-2821","type":"print"},{"value":"1573-7462","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,6,26]]},"assertion":[{"value":"26 June 2022","order":1,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}