{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,12]],"date-time":"2026-08-12T17:34:21Z","timestamp":1786556061386,"version":"3.56.0"},"reference-count":19,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2025,3,14]],"date-time":"2025-03-14T00:00:00Z","timestamp":1741910400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,3,14]],"date-time":"2025-03-14T00:00:00Z","timestamp":1741910400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100003359","name":"Generalitat Valenciana","doi-asserted-by":"publisher","award":["ACIF\/2021\/281"],"award-info":[{"award-number":["ACIF\/2021\/281"]}],"id":[{"id":"10.13039\/501100003359","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003359","name":"Generalitat Valenciana","doi-asserted-by":"publisher","award":["CIDEXG\/2022\/013"],"award-info":[{"award-number":["CIDEXG\/2022\/013"]}],"id":[{"id":"10.13039\/501100003359","id-type":"DOI","asserted-by":"publisher"}]},{"name":"European Union NextGenerationEU\/PRTR","award":["TED2021-129334B- I00"],"award-info":[{"award-number":["TED2021-129334B- I00"]}]},{"DOI":"10.13039\/501100004834","name":"Universitat Jaume I","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004834","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>The deployment of deep learning models on resource-constrained devices requires the development of new optimisation techniques to effectively exploit the computational and storage capacities of these devices. Thus, the primary objective of this research is to introduce an innovative and efficient approach for fusing convolution (or fully connected), ReLU, and batch normalisation neural network layers into a unified, single-layer structure, alongside a quantisation method for this new fused layer. This approach has been evaluated using the Arduino BLE Sense ARM Cortex-M4 and the Arduino Portenta H7 Lite ARM Cortex-M4 and M7 processors, known for their widespread adoption in various Internet of Things devices. Depending on the microcontroller unit and compilation flag used, the fused layers can reduce the overall execution time by up to 1.53<jats:inline-formula>\n              <jats:alternatives>\n                <jats:tex-math>$$\\times$$<\/jats:tex-math>\n                <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mo>\u00d7<\/mml:mo>\n                <\/mml:math>\n              <\/jats:alternatives>\n            <\/jats:inline-formula>, and on individual layers it can reach a speedup of 2.95<jats:inline-formula>\n              <jats:alternatives>\n                <jats:tex-math>$$\\times$$<\/jats:tex-math>\n                <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mo>\u00d7<\/mml:mo>\n                <\/mml:math>\n              <\/jats:alternatives>\n            <\/jats:inline-formula>.<\/jats:p>","DOI":"10.1007\/s11227-025-07107-y","type":"journal-article","created":{"date-parts":[[2025,3,14]],"date-time":"2025-03-14T14:48:33Z","timestamp":1741963713000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Deep learning inference optimisation for IoT: Conv2D-ReLU-BN layer fusion and quantisation"],"prefix":"10.1007","volume":"81","author":[{"given":"Jose I.","family":"Mestre","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sergio","family":"Barrachina","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Darwin","family":"Quezada","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Manuel F.","family":"Dolz","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,3,14]]},"reference":[{"key":"7107_CR1","doi-asserted-by":"publisher","first-page":"100164","DOI":"10.1016\/j.crbiot.2023.100164","volume":"7","author":"C Chakraborty","year":"2024","unstructured":"Chakraborty C, Bhattacharya M, Pal S, Lee S-S (2024) From machine learning to deep learning: advances of the recent data-driven paradigm shift in medicine and healthcare. Curr Res Biotechnol 7:100164","journal-title":"Curr Res Biotechnol"},{"key":"7107_CR2","doi-asserted-by":"publisher","first-page":"107271","DOI":"10.1016\/j.engappai.2023.107271","volume":"127","author":"I-I Prado-Rujas","year":"2024","unstructured":"Prado-Rujas I-I, Garc\u00eda-Dopico A, Serrano E, C\u00f3rdoba ML, P\u00e9rez MS (2024) A multivariable sensor-agnostic framework for spatio-temporal air quality forecasting based on deep learning. Eng Appl Artif Intell 127:107271","journal-title":"Eng Appl Artif Intell"},{"key":"7107_CR3","unstructured":"Tavakoli M, Baldi P, Carlton AM, Chiu YT, Shmakov A, Van\u00a0Vranken D (2024) AI for interpretable chemistry: predicting radical mechanistic pathways via contrastive learning. Adv Neural Inf Process Syst 36"},{"issue":"15","key":"7107_CR4","doi-asserted-by":"publisher","first-page":"26203","DOI":"10.1109\/JIOT.2024.3395335","volume":"11","author":"A Maci\u00e1-Lillo","year":"2024","unstructured":"Maci\u00e1-Lillo A, Barrachina S, Fabregat G, Dolz MF (2024) Optimizing convolutions for deep learning inference on ARM Cortex-M processors. IEEE Internet Things J 11(15):26203\u201326219. https:\/\/doi.org\/10.1109\/JIOT.2024.3395335","journal-title":"IEEE Internet Things J"},{"key":"7107_CR5","doi-asserted-by":"publisher","unstructured":"Cai X, Wang Y, Zhang L (2021) Optimus: towards optimal layer-fusion on deep learning processors. In: Proceedings of the 22nd ACM SIGPLAN\/SIGBED International Conference on Languages, Compilers, and Tools for Embedded Systems. LCTES 2021, pp. 67\u201379. Association for Computing Machinery, New York, NY, USA https:\/\/doi.org\/10.1145\/3461648.3463848","DOI":"10.1145\/3461648.3463848"},{"key":"7107_CR6","unstructured":"Chen T, Moreau T, Jiang Z, Zheng L, Yan E, Cowan M, Shen H, Wang L, Hu Y, Ceze L, Guestrin C, Krishnamurthy A (2018) TVM: an automated end-to-end optimizing compiler for deep learning. In: Proceedings of the 13th USENIX Conference on Operating Systems Design and Implementation. OSDI\u201918, pp. 579\u2013594. USENIX Association, USA"},{"issue":"3","key":"7107_CR7","doi-asserted-by":"publisher","first-page":"708","DOI":"10.1109\/TPDS.2020.3030548","volume":"32","author":"M Li","year":"2021","unstructured":"Li M, Liu Y, Liu X, Sun Q, You X, Yang H, Luan Z, Gan L, Yang G, Qian D (2021) The deep learning compiler: a comprehensive survey. IEEE Trans Parallel Distrib Syst 32(3):708\u2013727. https:\/\/doi.org\/10.1109\/TPDS.2020.3030548","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"7107_CR8","unstructured":"Ioffe S, Szegedy C (2015) Batch normalization: accelerating deep network training by reducing internal covariate shift"},{"key":"7107_CR9","unstructured":"Ducha A. CaffeNet benchmark - understanding batch normalization. https:\/\/github.com\/ducha-aiki\/caffenet-enchmark\/blob\/master\/batchnorm.md. Accessed 21 Jan 2025"},{"key":"7107_CR10","unstructured":"Batch Normalization Before or After ReLU?. https:\/\/www.reddit.com\/r\/MachineLearning\/comments\/67gonq\/d_batch_normalization_before_or_after_relu\/?rdt=65436. Accessed 21 Jan 2025"},{"key":"7107_CR11","doi-asserted-by":"publisher","first-page":"240","DOI":"10.1016\/j.jpdc.2022.05.009","volume":"167","author":"S Barrachina","year":"2022","unstructured":"Barrachina S, Dolz MF, San Juan P, Quintana-Ort\u00ed ES (2022) Efficient and portable gemm-based convolution operators for deep neural network training on multicore processors. J Parallel Distrib Comput 167:240\u2013254. https:\/\/doi.org\/10.1016\/j.jpdc.2022.05.009","journal-title":"J Parallel Distrib Comput"},{"key":"7107_CR12","doi-asserted-by":"publisher","unstructured":"Jacob B, Kligys S, Chen B, Zhu M, Tang M, Howard A, Adam H, Kalenichenko D (2018) Quantization and training of neural networks for efficient integer-arithmetic-only inference. In: 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp 2704\u20132713 https:\/\/doi.org\/10.1109\/CVPR.2018.00286","DOI":"10.1109\/CVPR.2018.00286"},{"key":"7107_CR13","unstructured":"Arduino Nano 33 BLE Sense. https:\/\/store.arduino.cc\/arduino-nano-33-ble-sense. Accessed 21 Jan 2025"},{"key":"7107_CR14","unstructured":"STMicroelectronics: STM32H747XI Datasheet. Datasheet. https:\/\/www.st.com\/en\/microcontrollers-microprocessors\/stm32h747xi.html. Accessed 21 Jan 2025"},{"key":"7107_CR15","unstructured":"Arduino Portenta H7 Lite. https:\/\/www.arduino.cc\/pro\/hardware\/product\/portenta-h7-lite. Accessed 21 Jan 2025"},{"key":"7107_CR16","unstructured":"Simonyan K, Zisserman A (2015) Very deep convolutional networks for large-scale image recognition"},{"issue":"12","key":"7107_CR17","doi-asserted-by":"publisher","first-page":"2734","DOI":"10.14778\/3407790.3407857","volume":"13","author":"J Fang","year":"2020","unstructured":"Fang J, Shen Y, Wang Y, Chen L (2020) Optimizing dnn computation graph using graph substitutions. Proc VLDB Endow 13(12):2734\u20132746. https:\/\/doi.org\/10.14778\/3407790.3407857","journal-title":"Proc VLDB Endow"},{"key":"7107_CR18","unstructured":"O\u2019Neill J, Steeg GV, Galstyan A (2021) Layer-wise neural network compression via layer fusion. In: Balasubramanian, V.N., Tsang, I. (eds.) Proceedings of The 13th Asian Conference on Machine Learning. Proceedings of Machine Learning Research, vol 157, pp 1381\u20131396. PMLR, Online https:\/\/proceedings.mlr.press\/v157\/o-neill21a.html"},{"key":"7107_CR19","doi-asserted-by":"publisher","unstructured":"Niu W, Guan J, Wang Y, Agrawal G, Ren B (2021) Dnnfusion: accelerating deep neural networks execution with advanced operator fusion. In: Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation. PLDI 2021, pp. 883\u2013898. Association for Computing Machinery, New York, NY, USA https:\/\/doi.org\/10.1145\/3453483.3454083","DOI":"10.1145\/3453483.3454083"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-025-07107-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-025-07107-y\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-025-07107-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,3,14]],"date-time":"2025-03-14T14:48:39Z","timestamp":1741963719000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-025-07107-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,14]]},"references-count":19,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2025,3]]}},"alternative-id":["7107"],"URL":"https:\/\/doi.org\/10.1007\/s11227-025-07107-y","relation":{},"ISSN":["1573-0484"],"issn-type":[{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,3,14]]},"assertion":[{"value":"21 February 2025","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"14 March 2025","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"621"}}