{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T04:42:05Z","timestamp":1780548125565,"version":"3.54.1"},"reference-count":31,"publisher":"Springer Science and Business Media LLC","issue":"9","license":[{"start":{"date-parts":[[2023,1,21]],"date-time":"2023-01-21T00:00:00Z","timestamp":1674259200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,1,21]],"date-time":"2023-01-21T00:00:00Z","timestamp":1674259200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Agencia Estatal de Investigaci\u00f3n,Spain","award":["PID2020-113656RB-C21\/C22"],"award-info":[{"award-number":["PID2020-113656RB-C21\/C22"]}]},{"name":"Agencia Estatal de Investigaci\u00f3n,Spain","award":["PID2020-113656RB-C21\/C22"],"award-info":[{"award-number":["PID2020-113656RB-C21\/C22"]}]},{"name":"Agencia Estatal de Investigaci\u00f3n,Spain","award":["PID2020-113656RB-C21\/C22"],"award-info":[{"award-number":["PID2020-113656RB-C21\/C22"]}]},{"name":"Agencia Estatal de Investigaci\u00f3n,Spain","award":["PID2020-113656RB-C21\/C22"],"award-info":[{"award-number":["PID2020-113656RB-C21\/C22"]}]},{"DOI":"10.13039\/501100016386","name":"Conselleria de Innovaci\u00f3n, Universidades, Ciencia y Sociedad Digital, Generalitat Valenciana","doi-asserted-by":"publisher","award":["CDEIGENT\/2018\/014"],"award-info":[{"award-number":["CDEIGENT\/2018\/014"]}],"id":[{"id":"10.13039\/501100016386","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100002878","name":"Consejer\u00eda de Econom\u00eda, Innovaci\u00f3n, Ciencia y Empleo, Junta de Andaluc\u00eda","doi-asserted-by":"publisher","award":["POSTDOC_21_00025"],"award-info":[{"award-number":["POSTDOC_21_00025"]}],"id":[{"id":"10.13039\/501100002878","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100011033","name":"Agencia Estatal de Investigaci\u00f3n","doi-asserted-by":"publisher","award":["FJC2019-039222-I"],"award-info":[{"award-number":["FJC2019-039222-I"]}],"id":[{"id":"10.13039\/501100011033","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100011033","name":"Agencia Estatal de Investigaci\u00f3n","doi-asserted-by":"publisher","award":["PRE2021-099284"],"award-info":[{"award-number":["PRE2021-099284"]}],"id":[{"id":"10.13039\/501100011033","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004834","name":"Universitat Jaume I","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004834","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2023,6]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>In this work, we assess the performance and energy efficiency of high-performance codes for the convolution operator, based on the direct, explicit\/implicit lowering and Winograd algorithms used for deep learning (DL) inference on a series of ARM-based processor architectures. Specifically, we evaluate the NVIDIA Denver2 and Carmel processors, as well as the ARM Cortex-A57 and Cortex-A78AE CPUs as part of a recent set of NVIDIA Jetson platforms. The performance\u2013energy evaluation is carried out using the ResNet-50 v1.5 convolutional neural network (CNN) on varying configurations of convolution algorithms, number of threads\/cores, and operating frequencies on the tested processor cores. The results demonstrate that the best throughput is obtained on all platforms with the Winograd convolution operator running on all the cores at their highest frequency. However, if the goal is to reduce the energy footprint, there is no rule of thumb for the optimal configuration.<\/jats:p>","DOI":"10.1007\/s11227-023-05050-4","type":"journal-article","created":{"date-parts":[[2023,1,21]],"date-time":"2023-01-21T07:02:40Z","timestamp":1674284560000},"page":"9819-9836","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Performance\u2013energy trade-offs of deep learning convolution algorithms on ARM processors"],"prefix":"10.1007","volume":"79","author":[{"given":"Manuel F.","family":"Dolz","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sergio","family":"Barrachina","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"H\u00e9ctor","family":"Mart\u00ednez","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Adri\u00e1n","family":"Castell\u00f3","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Antonio","family":"Maci\u00e1","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Germ\u00e1n","family":"Fabregat","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andr\u00e9s E.","family":"Tom\u00e1s","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,1,21]]},"reference":[{"issue":"5","key":"5050_CR1","doi-asserted-by":"publisher","first-page":"92:1","DOI":"10.1145\/3234150","volume":"51","author":"S Pouyanfar","year":"2018","unstructured":"Pouyanfar S, Sadiq S, Yan Y, Tian H, Tao Y, Reyes MP, Shyu M-L, Chen S-C, Iyengar SS (2018) A survey on deep learning: Algorithms, techniques, and applications. ACM Comput Surv 51(5):92:1-92:36. https:\/\/doi.org\/10.1145\/3234150. ([Online])","journal-title":"ACM Comput Surv"},{"issue":"12","key":"5050_CR2","doi-asserted-by":"publisher","first-page":"2295","DOI":"10.1109\/JPROC.2017.2761740","volume":"105","author":"V Sze","year":"2017","unstructured":"Sze V, Chen Y-H, Yang T-J, Emer JS (2017) Efficient processing of deep neural networks: a tutorial and survey. Proc IEEE 105(12):2295\u20132329","journal-title":"Proc IEEE"},{"key":"5050_CR3","doi-asserted-by":"crossref","unstructured":"San\u00a0Juan P, Castell\u00f3 A, Dolz MF, Alonso-Jord\u00e1 P, Quintana-Ort\u00ed ES (2020) High performance and portable convolution operators for multicore processors. In: 2020 IEEE 32nd International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD), pp 91\u201398","DOI":"10.1109\/SBAC-PAD49847.2020.00023"},{"key":"5050_CR4","unstructured":"Zhang J, Franchetti F, Low TM (2018) High performance zero-memory overhead direct convolutions. In: Proceedings of the 35th International Conference on Machine Learning \u2013 ICML, Vol 80, pp 5776\u20135785"},{"issue":"6","key":"5050_CR5","doi-asserted-by":"publisher","first-page":"1955","DOI":"10.3390\/s21061955","volume":"21","author":"MJH Pantho","year":"2021","unstructured":"Pantho MJH, Bhowmik P, Bobda C (2021) Towards an efficient CNN inference architecture enabling in-sensor processing. Sensors 21(6):1955","journal-title":"Sensors"},{"key":"5050_CR6","doi-asserted-by":"publisher","first-page":"09","DOI":"10.1007\/s11227-021-03673-z","volume":"77","author":"S Barrachina","year":"2021","unstructured":"Barrachina S, Castell\u00f3 A, Catalan M, Dolz MF, Mestre J (2021) PyDTNN: a user-friendly and extensible framework for distributed deep learning. J Supercomput 77:09","journal-title":"J Supercomput"},{"key":"5050_CR7","unstructured":"Chellapilla K, Puri S, Simard P (2006) High performance convolutional neural networks for document processing. In: International Workshop on Frontiers in Handwriting Recognition"},{"key":"5050_CR8","doi-asserted-by":"crossref","unstructured":"Georganas E, Avancha S, Banerjee K, Kalamkar D, Henry G, Pabst H, Heinecke A (2018) Anatomy of high-performance deep learning convolutions on SIMD architectures. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis, ser. SC \u201918. IEEE Press","DOI":"10.1109\/SC.2018.00069"},{"key":"5050_CR9","doi-asserted-by":"publisher","unstructured":"Zlateski A, Jia Z, Li K, Durand F (2019) The anatomy of efficient FFT and Winograd convolutions on modern CPUs. In: Proceedings of the ACM International Conference on Supercomputing, ser. ICS \u201919. New York, NY, USA: Association for Computing Machinery, p 414-424. Available: https:\/\/doi.org\/10.1145\/3330345.3330382","DOI":"10.1145\/3330345.3330382"},{"key":"5050_CR10","doi-asserted-by":"publisher","first-page":"248","DOI":"10.1007\/978-3-030-57675-2_16","volume-title":"Euro-Par 2020: Parallel Processing","author":"Q Wang","year":"2020","unstructured":"Wang Q, Li D, Huang X, Shen S, Mei S, Liu J (2020) Optimizing FFT-based convolution on ARMv8 multi-core CPUs. In: Malawski M, Rzadca K (eds) Euro-Par 2020: Parallel Processing. Springer International Publishing, Cham, pp 248\u2013262"},{"key":"5050_CR11","unstructured":"Zlateski A, Jia Z, Li K, Durand F (2018) FFT convolutions are faster than Winograd on modern CPUs, here is why"},{"key":"5050_CR12","doi-asserted-by":"publisher","unstructured":"Lavin A, Gray S (2016) Fast algorithms for convolutional neural networks. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 4013\u20134021. Available: https:\/\/doi.org\/10.1109\/CVPR.2016.435","DOI":"10.1109\/CVPR.2016.435"},{"key":"5050_CR13","unstructured":"Lai L, Suda N, Chandra V (2018) Cmsis-nn: Efficient neural network kernels for arm cortex-m cpus. arXiv preprint arXiv:1801.06601"},{"key":"5050_CR14","unstructured":"Sun D, Liu S, Gaudiot J-L (2017) Enabling embedded inference engine with arm compute library: A case study. arXiv preprint arXiv:1704.03751"},{"key":"5050_CR15","unstructured":"Dukhan M Nnpack: an acceleration package for neural network computations. Available: https:\/\/github.com\/Maratyszcza\/NNPACK"},{"key":"5050_CR16","doi-asserted-by":"crossref","unstructured":"Castell\u00f3 A, Barrachina S, Dolz MF, Quintana-Ort\u00ed ES, Juan PS, Tom\u00e1s AE (2022) High performance and energy efficient inference for deep learning on multicore arm processors using general optimization techniques and blis. Journal of Systems Architecture, vol 125, p 102459. Available: https:\/\/www.sciencedirect.com\/science\/article\/pii\/S1383762122000509","DOI":"10.1016\/j.sysarc.2022.102459"},{"key":"5050_CR17","doi-asserted-by":"crossref","unstructured":"Bhattacharya S, Lane ND (2016) Sparsification and separation of deep learning layers for constrained resource inference on wearables. In: Proceedings of the 14th ACM Conference on Embedded Network Sensor Systems CD-ROM, pp 176\u2013189","DOI":"10.1145\/2994551.2994564"},{"key":"5050_CR18","doi-asserted-by":"crossref","unstructured":"Lane ND, Bhattacharya S, Georgiev P, Forlivesi C, Jiao L, Qendro L, Kawsar F (2016) DeepX: A software accelerator for low-power deep learning inference on mobile devices. In: 2016 15th ACM\/IEEE International Conference on Information Processing in Sensor Networks (IPSN). IEEE, pp 1\u201312","DOI":"10.1109\/IPSN.2016.7460664"},{"key":"5050_CR19","unstructured":"Barrachina S, Castello A, Dolz MF, Low TM, Martinez H, Quintana-Orti ES, Sridhar U, Tomas AE Convdirect: A library with different implementations of the direct convolution operation. Available: https:\/\/github.com\/hpca-uji\/convDirect.git"},{"issue":"3","key":"5050_CR20","doi-asserted-by":"crossref","first-page":"14:1","DOI":"10.1145\/2764454","volume":"41","author":"FG Van Zee","year":"2015","unstructured":"Van Zee FG, van de Geijn RA (2015) BLIS: a framework for rapidly instantiating BLAS functionality. ACM Trans on Math Soft 41(3):14:1-14:33","journal-title":"ACM Trans on Math Soft"},{"key":"5050_CR21","unstructured":"Dolz MF, Barrachina S, Castello A, Quintana-Orti ES, Tomas AE Convwinograd: An implementation of the winograd-based convolution transform. Available: https:\/\/github.com\/hpca-uji\/convWinograd.git"},{"key":"5050_CR22","doi-asserted-by":"crossref","unstructured":"Winograd S (1980) Arithmetic Complexity of Computations. Society for Industrial and Applied Mathematics","DOI":"10.1137\/1.9781611970364"},{"key":"5050_CR23","doi-asserted-by":"publisher","unstructured":"Barabasz B, Anderson A, Soodhalter KM, Gregg D (2020) Error analysis and improving the accuracy of Winograd convolution for deep neural networks. ACM Trans. Math. Softw. 46(4):1. Available: https:\/\/doi.org\/10.1145\/3412380","DOI":"10.1145\/3412380"},{"key":"5050_CR24","doi-asserted-by":"crossref","unstructured":"Dolz MF, Castell\u00f3 A, Quintana-Ort\u00ed ES (2022) Towards portable realizations of Winograd-based convolution with vector intrinsics and OpenMP. In: 2022 30th Euromicro International Conference on Parallel, Distributed and Network-based Processing (PDP), pp 39\u201346","DOI":"10.1109\/PDP55904.2022.00015"},{"key":"5050_CR25","doi-asserted-by":"crossref","unstructured":"Masliah I, Abdelfattah A, Haidar A, Tomov S, Falcou J, Dongarra J (2016) High-performance matrix-matrix multiplications of very small matrices. In: 22nd International European Conference on Parallel and Distributed Computing (Euro-Par\u201916). Grenoble, France: Springer International Publishing, 2016-08","DOI":"10.1007\/978-3-319-43659-3_48"},{"key":"5050_CR26","unstructured":"Tegra hardware information. Available: https:\/\/developer.nvidia.com\/tegra-hardware-sales-inquiries"},{"key":"5050_CR27","unstructured":"Jetson-stats is a package for monitoring and control the nvidia jetson. Available: https:\/\/github.com\/rbonghi\/jetson_stats"},{"key":"5050_CR28","unstructured":"INA3221 Triple-Channel, High-Side Measurement, Shunt and Bus Voltage Monitor with I2C- and SMBUS-Compatible Interface, https:\/\/www.ti.com\/product\/INA3221#tech-docs"},{"key":"5050_CR29","unstructured":"Barrachina S, Barreda M, Catal\u00e1n S, Dolz MF, Fabregat G, Mayo R, Quintana-Ort\u00ed E (2013) An integrated framework for power-performance analysis of parallel scientific workloads. Energy, pp 114\u2013119"},{"key":"5050_CR30","doi-asserted-by":"crossref","unstructured":"Barrachina S, Castell\u00f3 A, Catal\u00e1n M, Dolz MF, Mestre JI (2021) A flexible research-oriented framework for distributed training of deep neural networks. In: 2021 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), pp 730\u2013739","DOI":"10.1109\/IPDPSW52791.2021.00110"},{"key":"5050_CR31","unstructured":"Krizhevsky A, Sutskever I, Hinton GE (2012) Imagenet classification with deep convolutional neural networks. In: Proceedings of the 25th International Conference on Neural Information Processing Systems - Volume 1, ser. NIPS\u201912. USA: Curran Associates Inc., pp 1097\u20131105. Available: http:\/\/dl.acm.org\/citation.cfm?id=2999134.2999257"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-023-05050-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-023-05050-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-023-05050-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,4,24]],"date-time":"2023-04-24T22:06:29Z","timestamp":1682373989000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-023-05050-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,21]]},"references-count":31,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2023,6]]}},"alternative-id":["5050"],"URL":"https:\/\/doi.org\/10.1007\/s11227-023-05050-4","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"value":"0920-8542","type":"print"},{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1,21]]},"assertion":[{"value":"9 January 2023","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 January 2023","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}