{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T15:21:01Z","timestamp":1781104861502,"version":"3.54.1"},"reference-count":62,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2022,12,21]],"date-time":"2022-12-21T00:00:00Z","timestamp":1671580800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100000266","name":"Engineering and Physical Sciences Research Council","doi-asserted-by":"publisher","award":["EP\/M50659X\/1,EP\/R018677\/1,EP\/S001530\/1"],"award-info":[{"award-number":["EP\/M50659X\/1,EP\/R018677\/1,EP\/S001530\/1"]}],"id":[{"id":"10.13039\/501100000266","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Interact. Mob. Wearable Ubiquitous Technol."],"published-print":{"date-parts":[[2022,12,21]]},"abstract":"<jats:p>Wearable, embedded, and IoT devices are a centrepiece of many ubiquitous computing applications, such as fitness tracking, health monitoring, home security and voice assistants. By gathering user data through a variety of sensors and leveraging machine learning (ML), applications can adapt their behaviour: in other words, devices become \"smart\". Such devices are typically powered by microcontroller units (MCUs). As MCUs continue to improve, smart devices become capable of performing a non-trivial amount of sensing and data processing, including machine learning inference, which results in a greater degree of user data privacy and autonomy, compared to offloading the execution of ML models to another device.<\/jats:p>\n          <jats:p>Advanced predictive capabilities across many tasks make neural networks an attractive ML model for ubiquitous computing applications; however, on-device inference on MCUs remains extremely challenging. Orders of magnitude less storage, memory and computational ability, compared to what is typically required to execute neural networks, impose strict structural constraints on the network architecture and call for specialist model compression methodology. In this work, we present a differentiable structured pruning method for convolutional neural networks, which integrates a model's MCU-specific resource usage and parameter importance feedback to obtain highly compressed yet accurate models. Compared to related network pruning work, compressed models are more accurate due to better use of MCU resource budget, and compared to MCU specialist work, compressed models are produced faster. The user only needs to specify the amount of available computational resources and the pruning algorithm will automatically compress the network during training to satisfy them.<\/jats:p>\n          <jats:p>We evaluate our methodology using benchmark image and audio classification tasks and find that it (a) improves key resource usage of neural networks up to 80x; (b) has little to no overhead or even improves model training time; (c) produces compressed models with matching or improved resource usage up to 1.4x in less time compared to prior MCU-specific model compression methods.<\/jats:p>","DOI":"10.1145\/3569468","type":"journal-article","created":{"date-parts":[[2023,1,11]],"date-time":"2023-01-11T15:34:01Z","timestamp":1673451241000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":27,"title":["Differentiable Neural Network Pruning to Enable Smart Applications on Microcontrollers"],"prefix":"10.1145","volume":"6","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2621-2285","authenticated-orcid":false,"given":"Edgar","family":"Liberis","sequence":"first","affiliation":[{"name":"University of Cambridge, Cambridge, UK and Samsung AI Centre Cambridge, Cambridge, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2728-8273","authenticated-orcid":false,"given":"Nicholas D.","family":"Lane","sequence":"additional","affiliation":[{"name":"University of Cambridge, Cambridge, UK and Samsung AI Centre Cambridge, Cambridge, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,1,11]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448083"},{"key":"e_1_2_2_2_1","volume-title":"DynO: Dynamic Onloading of Deep Neural Networks from Cloud to Device. arXiv preprint arXiv:2104.09949","author":"Almeida Mario","year":"2021","unstructured":"Mario Almeida, Stefanos Laskaridis, Stylianos I Venieris, Ilias Leontiadis, and Nicholas D Lane. 2021. DynO: Dynamic Onloading of Deep Neural Networks from Cloud to Device. arXiv preprint arXiv:2104.09949 (2021)."},{"key":"e_1_2_2_3_1","volume-title":"gen. 4. Retrieved","author":"Dot Echo","year":"2022","unstructured":"Amazon. 2022. Echo Dot, gen. 4. Retrieved February 1, 2022 from https:\/\/www.amazon.co.uk\/all-new-echo-dot-4th-generation-smart-speaker-with-alexa-charcoal\/dp\/B084DWCZXZ"},{"key":"e_1_2_2_4_1","volume-title":"Structured Pruning of Deep Convolutional Neural Networks. CoRR abs\/1512.08571","author":"Anwar Sajid","year":"2015","unstructured":"Sajid Anwar, Kyuyeon Hwang, and Wonyong Sung. 2015. Structured Pruning of Deep Convolutional Neural Networks. CoRR abs\/1512.08571 (2015). arXiv preprint arXiv:1512.08571 (2015)."},{"key":"e_1_2_2_5_1","unstructured":"ARM mbed. 2022. NUCLEO-F446RE. Retrieved February 1 2022 from https:\/\/os.mbed.com\/platforms\/ST-Nucleo-F446RE\/"},{"key":"e_1_2_2_6_1","unstructured":"ARM mbed. 2022. NUCLEO-F767ZI. Retrieved February 1 2022 from https:\/\/os.mbed.com\/platforms\/ST-Nucleo-F767ZI\/"},{"key":"e_1_2_2_7_1","unstructured":"ARM mbed. 2022. NUCLEO-H743ZI2. Retrieved February 1 2022 from https:\/\/os.mbed.com\/platforms\/ST-Nucleo-H743ZI2\/"},{"key":"e_1_2_2_8_1","volume-title":"Urmish Thakkar, Dibakar Gope, Vijay Janapa Reddi, Matthew Mattina, and Paul N Whatmough.","author":"Banbury Colby","year":"2020","unstructured":"Colby Banbury, Chuteng Zhou, Igor Fedorov, Ramon Matas Navarro, Urmish Thakkar, Dibakar Gope, Vijay Janapa Reddi, Matthew Mattina, and Paul N Whatmough. 2020. MicroNets: Neural Network Architectures for Deploying TinyML Applications on Commodity Microcontrollers. arXiv preprint arXiv:2010.11267 (2020)."},{"key":"e_1_2_2_9_1","volume-title":"Jonathan Frankle, and John Guttag.","author":"Blalock Davis","year":"2020","unstructured":"Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag. 2020. What is the state of neural network pruning? arXiv preprint arXiv:2003.03033 (2020)."},{"key":"e_1_2_2_10_1","volume-title":"Megan K O'Brien, Nicholas Shawen, John A Rogers, Richard L Lieber, Kathryn J Reid, Phyllis C Zee, and Arun Jayaraman.","author":"Boe Alexander J","year":"2019","unstructured":"Alexander J Boe, Lori L McGee Koch, Megan K O'Brien, Nicholas Shawen, John A Rogers, Richard L Lieber, Kathryn J Reid, Phyllis C Zee, and Arun Jayaraman. 2019. Automating sleep stage classification using wireless, wearable sensors. NPJ digital medicine 2, 1 (2019), 1--9."},{"key":"e_1_2_2_11_1","volume-title":"BRP-NAS: Prediction-based NAS using GCNs. arXiv preprint arXiv:2007.08668","author":"Chau Thomas","year":"2020","unstructured":"Thomas Chau, \u0141ukasz Dudziak, Mohamed S Abdelfattah, Royson Lee, Hyeji Kim, and Nicholas D Lane. 2020. BRP-NAS: Prediction-based NAS using GCNs. arXiv preprint arXiv:2007.08668 (2020)."},{"key":"e_1_2_2_12_1","volume-title":"SwiftNet: Using Graph Propagation as Meta-knowledge to Search Highly Representative Neural Architectures. arXiv preprint arXiv:1906.08305","author":"Cheng Hsin-Pai","year":"2019","unstructured":"Hsin-Pai Cheng, Tunhou Zhang, Yukun Yang, Feng Yan, Shiyu Li, Harris Teague, Hai Li, and Yiran Chen. 2019. SwiftNet: Using Graph Propagation as Meta-knowledge to Search Highly Representative Neural Architectures. arXiv preprint arXiv:1906.08305 (2019)."},{"key":"e_1_2_2_13_1","volume-title":"Visual wake words dataset. arXiv preprint arXiv:1906.05721","author":"Chowdhery Aakanksha","year":"2019","unstructured":"Aakanksha Chowdhery, Pete Warden, Jonathon Shlens, Andrew Howard, and Rocky Rhodes. 2019. Visual wake words dataset. arXiv preprint arXiv:1906.05721 (2019)."},{"key":"e_1_2_2_14_1","volume-title":"Nat Jeffries, Jian Li, Nick Kreeger, Ian Nappier, Meghna Natraj, Shlomi Regev, et al.","author":"David Robert","year":"2020","unstructured":"Robert David, Jared Duke, Advait Jain, Vijay Janapa Reddi, Nat Jeffries, Jian Li, Nick Kreeger, Ian Nappier, Meghna Natraj, Shlomi Regev, et al. 2020. TensorFlow Lite Micro: Embedded Machine Learning on TinyML Systems. arXiv preprint arXiv:2010.08678 (2020)."},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3161190"},{"key":"e_1_2_2_16_1","unstructured":"Igor Fedorov Ryan P Adams Matthew Mattina and Paul Whatmough. 2019. SpArSe: Sparse architecture search for CNNs on resource-constrained microcontrollers. In Advances in Neural Information Processing Systems. 4977--4989."},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3131895"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3351240"},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00171"},{"key":"e_1_2_2_20_1","volume-title":"Share & Trends Report. Retrieved","author":"Research Grand View","year":"2022","unstructured":"Grand View Research. 2022. Microcontroller Market Size, Share & Trends Report. Retrieved February 1, 2022 from https:\/\/www.grandviewresearch.com\/industry-analysis\/microcontroller-market"},{"key":"e_1_2_2_21_1","volume-title":"Learning both weights and connections for efficient neural networks. arXiv preprint arXiv:1506.02626","author":"Han Song","year":"2015","unstructured":"Song Han, Jeff Pool, John Tran, and William J Dally. 2015. Learning both weights and connections for efficient neural networks. arXiv preprint arXiv:1506.02626 (2015)."},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3214269"},{"key":"e_1_2_2_24_1","volume-title":"SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and &lt","author":"Iandola Forrest N","year":"2016","unstructured":"Forrest N Iandola, Song Han, Matthew W Moskewicz, Khalid Ashraf, William J Dally, and Kurt Keutzer. 2016. SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and &lt; 0.5 MB model size. arXiv preprint arXiv:1602.07360 (2016)."},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00286"},{"key":"e_1_2_2_26_1","unstructured":"Alex Krizhevsky. 2009. Learning multiple layers of features from tiny images. Technical Report."},{"key":"e_1_2_2_27_1","volume-title":"International Conference on Machine Learning. 1935--1944","author":"Kumar Ashish","year":"2017","unstructured":"Ashish Kumar, Saurabh Goyal, and Manik Varma. 2017. Resource-efficient machine learning in 2KB RAM for the internet of things. In International Conference on Machine Learning. 1935--1944."},{"key":"e_1_2_2_28_1","volume-title":"CMSIS-NN: Efficient neural network kernels for ARM Cortex-M CPUs. arXiv preprint arXiv:1801.06601","author":"Lai Liangzhen","year":"2018","unstructured":"Liangzhen Lai, Naveen Suda, and Vikas Chandra. 2018. CMSIS-NN: Efficient neural network kernels for ARM Cortex-M CPUs. arXiv preprint arXiv:1801.06601 (2018)."},{"key":"e_1_2_2_29_1","volume-title":"SNIP: Single-shot network pruning based on connection sensitivity. arXiv preprint arXiv:1810.02340","author":"Lee Namhoon","year":"2018","unstructured":"Namhoon Lee, Thalaiyasingam Ajanthan, and Philip HS Torr. 2018. SNIP: Single-shot network pruning based on connection sensitivity. arXiv preprint arXiv:1810.02340 (2018)."},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00932"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3437984.3458836"},{"key":"e_1_2_2_32_1","volume-title":"Neural networks on microcontrollers: saving memory at inference via operator reordering. arXiv preprint arXiv:1910.05110","author":"Liberis Edgar","year":"2019","unstructured":"Edgar Liberis and Nicholas D Lane. 2019. Neural networks on microcontrollers: saving memory at inference via operator reordering. arXiv preprint arXiv:1910.05110 (2019)."},{"key":"e_1_2_2_33_1","volume-title":"MCUNet: Tiny deep learning on IoT devices. arXiv preprint arXiv:2007.10319","author":"Lin Ji","year":"2020","unstructured":"Ji Lin, Wei-Ming Chen, Yujun Lin, John Cohn, Chuang Gan, and Song Han. 2020. MCUNet: Tiny deep learning on IoT devices. arXiv preprint arXiv:2007.10319 (2020)."},{"key":"e_1_2_2_34_1","volume-title":"Dynamic model pruning with feedback. arXiv preprint arXiv:2006.07253","author":"Lin Tao","year":"2020","unstructured":"Tao Lin, Sebastian U Stich, Luis Barba, Daniil Dmitriev, and Martin Jaggi. 2020. Dynamic model pruning with feedback. arXiv preprint arXiv:2006.07253 (2020)."},{"key":"e_1_2_2_35_1","volume-title":"AutoCompress: An automatic DNN structured pruning framework for ultra-high compression rates. arXiv preprint arXiv:1907.03141","author":"Liu Ning","year":"2019","unstructured":"Ning Liu, Xiaolong Ma, Zhiyuan Xu, Yanzhi Wang, Jian Tang, and Jieping Ye. 2019. AutoCompress: An automatic DNN structured pruning framework for ultra-high compression rates. arXiv preprint arXiv:1907.03141 (2019)."},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3494971"},{"key":"e_1_2_2_37_1","first-page":"1","article-title":"WR-Hand: Wearable Armband Can Track User's Hand","volume":"5","author":"Liu Yang","year":"2021","unstructured":"Yang Liu, Chengdong Lin, and Zhenjiang Li. 2021. WR-Hand: Wearable Armband Can Track User's Hand. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 5, 3 (2021), 1--27.","journal-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"},{"key":"e_1_2_2_38_1","unstructured":"Zhuang Liu Jianguo Li Zhiqiang Shen Gao Huang Shoumeng Yan and Changshui Zhang. 2017. Learning Efficient Convolutional Networks through Network Slimming. In ICCV."},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICECS46596.2019.8964993"},{"key":"e_1_2_2_40_1","volume-title":"DSA: More efficient budgeted pruning via differentiable sparsity allocation. arXiv preprint arXiv:2004.02164","author":"Ning Xuefei","year":"2020","unstructured":"Xuefei Ning, Tianchen Zhao, Wenshuo Li, Peng Lei, Yu Wang, and Huazhong Yang. 2020. DSA: More efficient budgeted pruning via differentiable sparsity allocation. arXiv preprint arXiv:2004.02164 (2020)."},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/EMBC.2019.8856722"},{"key":"e_1_2_2_42_1","unstructured":"Biswajit Paria Kirthevasan Kandasamy and Barnab\u00e1s P\u00f3czos. 2020. A flexible framework for multi-objective bayesian optimization using random scalarizations. In Uncertainty in Artificial Intelligence. PMLR 766--776."},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3161174"},{"key":"e_1_2_2_44_1","volume-title":"Video doorbell, gen. 2. Retrieved","year":"2022","unstructured":"Ring. 2022. Video doorbell, gen. 2. Retrieved February 1, 2022 from https:\/\/en-uk.ring.com\/products\/video-doorbell-gen-2"},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.3022818"},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_2_2_47_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3432701"},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neuri.2021.100028"},{"key":"e_1_2_2_50_1","volume-title":"International Conference on Machine Learning. PMLR, 6105--6114","author":"Tan Mingxing","year":"2019","unstructured":"Mingxing Tan and Quoc Le. 2019. Efficientnet: Rethinking model scaling for convolutional neural networks. In International Conference on Machine Learning. PMLR, 6105--6114."},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8462688"},{"key":"e_1_2_2_52_1","volume-title":"Faster gaze prediction with dense networks and fisher pruning. arXiv preprint arXiv:1801.05787","author":"Theis Lucas","year":"2018","unstructured":"Lucas Theis, Iryna Korshunova, Alykhan Tejani, and Ferenc Husz\u00e1r. 2018. Faster gaze prediction with dense networks and fisher pruning. arXiv preprint arXiv:1801.05787 (2018)."},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3494995"},{"key":"e_1_2_2_54_1","volume-title":"Single Shot Structured Pruning Before Training. arXiv preprint arXiv:2007.00389","author":"van Amersfoort Joost","year":"2020","unstructured":"Joost van Amersfoort, Milad Alizadeh, Sebastian Farquhar, Nicholas Lane, and Yarin Gal. 2020. Single Shot Structured Pruning Before Training. arXiv preprint arXiv:2007.00389 (2020)."},{"key":"e_1_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240925.3240937"},{"key":"e_1_2_2_56_1","volume-title":"Picking winning tickets before training by preserving gradient flow. arXiv preprint arXiv:2002.07376","author":"Wang Chaoqi","year":"2020","unstructured":"Chaoqi Wang, Guodong Zhang, and Roger Grosse. 2020. Picking winning tickets before training by preserving gradient flow. arXiv preprint arXiv:2002.07376 (2020)."},{"key":"e_1_2_2_57_1","volume-title":"Speech commands: A dataset for limited-vocabulary speech recognition. arXiv preprint arXiv:1804.03209","author":"Warden Pete","year":"2018","unstructured":"Pete Warden. 2018. Speech commands: A dataset for limited-vocabulary speech recognition. arXiv preprint arXiv:1804.03209 (2018)."},{"key":"e_1_2_2_58_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i4.20387"},{"key":"e_1_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/3380980"},{"key":"e_1_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/3397334"},{"key":"e_1_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/3494990"},{"key":"e_1_2_2_62_1","volume-title":"Hello Edge: Keyword spotting on microcontrollers. arXiv preprint arXiv:1711.07128","author":"Zhang Yundong","year":"2017","unstructured":"Yundong Zhang, Naveen Suda, Liangzhen Lai, and Vikas Chandra. 2017. Hello Edge: Keyword spotting on microcontrollers. arXiv preprint arXiv:1711.07128 (2017)."}],"container-title":["Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3569468","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3569468","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,15]],"date-time":"2025-07-15T20:52:12Z","timestamp":1752612732000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3569468"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,21]]},"references-count":62,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2022,12,21]]}},"alternative-id":["10.1145\/3569468"],"URL":"https:\/\/doi.org\/10.1145\/3569468","relation":{},"ISSN":["2474-9567"],"issn-type":[{"value":"2474-9567","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12,21]]},"assertion":[{"value":"2023-01-11","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}