{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T14:21:46Z","timestamp":1778682106515,"version":"3.51.4"},"reference-count":38,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T00:00:00Z","timestamp":1778630400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"Strategic Priority Research Program of the Chinese Academy of Sciences","award":["XDB0660102"],"award-info":[{"award-number":["XDB0660102"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2026,5,31]]},"abstract":"<jats:p>\n                    Mixed-precision neural network (MPNN) that utilizes just enough data width for the neural network processing is an effective approach to meet the stringent resources constraints including memory and computing of MCUs. Nevertheless, there is still a lack of sub-byte and mixed-precision SIMD operations in MCU-class ISA and the limited computing capability of MCUs remains underutilized, which further aggravates the computing bound encountered in neural network processing. As a result, the benefits of MPNNs cannot be fully unleashed. In this work, we propose to pack multiple low-bitwidth arithmetic operations within a single instruction multiple data (SIMD) instructions in typical MCUs, and then develop an efficient convolution operator by exploring both the data parallelism and computing parallelism in convolution along with the proposed SIMD packing. Finally, we further leverage Neural Architecture Search (NAS) to build a HW\/SW co-designed MPNN design framework, namely MCU-MixQ. This framework can optimize both the MPNN quantization and MPNN implementation efficiency, striking an optimized balance between neural network performance and accuracy. According to our experiment results, MCU-MixQ achieves 2.1\u00d7 and 1.4\u00d7 speedup over CMix-NN and MCUNet respectively under the same resource constraints. MCU-MixQ is also open sourced on GitHub.\n                    <jats:xref ref-type=\"fn\">\n                      <jats:sup>1<\/jats:sup>\n                    <\/jats:xref>\n                  <\/jats:p>","DOI":"10.1145\/3803553","type":"journal-article","created":{"date-parts":[[2026,3,19]],"date-time":"2026-03-19T20:56:25Z","timestamp":1773953785000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["MCU-MixQ: A HW\/SW Co-optimized Mixed-precision Neural Network Design Framework for MCUs"],"prefix":"10.1145","volume":"25","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-8367-5846","authenticated-orcid":false,"given":"Junfeng","family":"Gong","sequence":"first","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences","place":["Beijing, China"]},{"name":"University of the Chinese Academy of Sciences","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1638-059X","authenticated-orcid":false,"given":"Long","family":"Cheng","sequence":"additional","affiliation":[{"name":"North China Electric Power University","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9693-3467","authenticated-orcid":false,"given":"Jiawei","family":"Nian","sequence":"additional","affiliation":[{"name":"North China Electric Power University","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5542-7306","authenticated-orcid":false,"given":"Cheng","family":"Liu","sequence":"additional","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology Chinese Academy of Sciences","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8082-4218","authenticated-orcid":false,"given":"Huawei","family":"Li","sequence":"additional","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,5,13]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00540"},{"key":"e_1_3_2_3_2","unstructured":"Colby Banbury Chuteng Zhou Igor Fedorov Ramon Matas Navarro Urmish Thakker Dibakar Gope Vijay Janapa Reddi Matthew Mattina and Paul N. Whatmough. 2021. MicroNets: Neural network architectures for deploying tiny ML applications on commodity microcontrollers. In Proceedings of Machine Learning and Systems (MLSys) 3 (2021) 517\u2013532."},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/tc.2021.3066883"},{"key":"e_1_3_2_5_2","unstructured":"Yaohui Cai Zhewei Yao Zhen Dong Amir Gholami Michael W. Mahoney and Kurt Keutzer. 2020. ZeroQ: A novel zero shot quantization framework. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 13169\u201313178."},{"key":"e_1_3_2_6_2","unstructured":"Zhaowei Cai and Nuno Vasconcelos. 2020. Rethinking differentiable search for mixed-precision neural networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSII.2020.2983648"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/tnnls.2020.2980041"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3368826.3377912"},{"key":"e_1_3_2_10_2","doi-asserted-by":"crossref","unstructured":"Zhen Dong Zhewei Yao Amir Gholami Michael W. Mahoney and Kurt Keutzer. 2019. HAWQ: Hessian aware quantization of neural networks with mixed-precision. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV).","DOI":"10.1109\/ICCV.2019.00038"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/ESSCIRC53450.2021.9567767"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE48585.2020.9116529"},{"key":"e_1_3_2_13_2","unstructured":"Qing Jin Jian Ren Richard Zhuang Sumant Hanumante Zhengang Li Zhiyu Chen Yanzhi Wang Kaiyuan Yang and Sergey Tulyakov. 2022. F8Net: Fixed-point 8-bit only multiplication for network quantization. In The Tenth International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00767"},{"key":"e_1_3_2_15_2","unstructured":"Andrey Kuzmin Mart van Baalen Yuwei Ren Markus Nagel Jorn Peters and Tijmen Blankevoort. 2022. FP8 quantization: The power of the exponent. In Advances in Neural Information Processing Systems (NeurIPS 2022). 35."},{"key":"e_1_3_2_16_2","unstructured":"Liangzhen Lai Naveen Suda and Vikas Chandra. 2018. CMSIS-NN: Efficient Neural Network Kernels for Arm Cortex-M CPUs. arxiv:1801.06601. Retrieved from https:\/\/arxiv.org\/abs\/1801.06601"},{"key":"e_1_3_2_17_2","first-page":"289","volume-title":"Proceedings of the 12th Asian Conference on Machine Learning.","author":"Li Bowen","year":"2020","unstructured":"Bowen Li, Kai Huang, Siang Chen, Dongliang Xiong, Haitian Jiang, and Luc Claesen. 2020. DFQF: Data free quantization-aware fine-tuning. In Proceedings of the 12th Asian Conference on Machine Learning.Sinno Jialin Pan and Masashi Sugiyama (Eds.), PMLR, 289\u2013304. Retrieved from http:\/\/proceedings.mlr.press\/v129\/li20a.html"},{"key":"e_1_3_2_18_2","unstructured":"Lianqiang Li Chenqian Yan and Yefei Chen. 2024. Differentiable Search for Finding Optimal Quantization Strategy. arxiv:2404.08010. Retrieved from https:\/\/arxiv.org\/abs\/2404.08010"},{"key":"e_1_3_2_19_2","unstructured":"Yuhong Li Cong Hao Xiaofan Zhang Xinheng Liu Yao Chen Jinjun Xiong Wen mei Hwu and Deming Chen. 2020. EDD: Efficient Differentiable DNN Architecture and Implementation Co-search for Embedded AI Solutions. arxiv:2005.02563. Retrieved from https:\/\/arxiv.org\/abs\/2005.02563"},{"key":"e_1_3_2_20_2","unstructured":"Edgar Liberis and Nicholas D. Lane. 2023. Pex: Memory-efficient Microcontroller Deep Learning through Partial Execution. arxiv:2211.17246. Retrieved from https:\/\/arxiv.org\/abs\/2211.17246"},{"key":"e_1_3_2_21_2","unstructured":"Ji Lin Wei-Ming Chen Han Cai Chuang Gan and Song Han. 2021. Memory-efficient patch-based inference for tiny deep learning. In Advances in Neural Information Processing Systems (NeurIPS 2021). 34."},{"key":"e_1_3_2_22_2","unstructured":"Ji Lin Wei-Ming Chen Yujun Lin John Cohn Chuang Gan and Song Han. 2020. MCUNet: Tiny deep learning on IoT devices. In Advances in Neural Information Processing Systems (NeurIPS 2020). 33."},{"key":"e_1_3_2_23_2","unstructured":"Hanxiao Liu Karen Simonyan and Yiming Yang. 2019. DARTS: Differentiable architecture search. In The Seventh International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3582016.3582062"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASP-DAC52403.2022.9712553"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3608447"},{"key":"e_1_3_2_27_2","unstructured":"Paulius Micikevicius Dusan Stosic Neil Burgess Marius Cornea Pradeep Dubey Richard Grisenthwaite Sangwon Ha Alexander Heinecke Patrick Judd John Kamalu et\u00a0al. 2022. FP8 Formats for Deep Learning. arxiv:2209.05433. Retrieved from https:\/\/arxiv.org\/abs\/2209.05433"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSII.2022.3205029"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.3390\/s21092984"},{"key":"e_1_3_2_30_2","unstructured":"Nilesh Prasad Pandey Markus Nagel Mart van Baalen Yin Huang Chirag Patel and Tijmen Blankevoort. 2023. A Practical Mixed Precision Algorithm for Post-Training Quantization. arxiv:2302.05397. Retrieved from https:\/\/arxiv.org\/abs\/2302.05397"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3394390"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/fpl57034.2022.00035"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3517207.3526978"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3125501.3125528"},{"key":"e_1_3_2_35_2","unstructured":"Colin White Mahmoud Safari Rhea Sukthanker Binxin Ru Thomas Elsken Arber Zela Debadeepta Dey and Frank Hutter. 2023. Neural Architecture Search: Insights from 1000 Papers. arxiv:2301.08727."},{"key":"e_1_3_2_36_2","first-page":"52","volume-title":"Proceedings of the Machine Learning and Systems.","volume":"4","author":"Won Jaeyeon","year":"2022","unstructured":"Jaeyeon Won, Jeyeon Si, Sam Son, Tae Jun Ham, and Jae W. Lee. 2022. ULPPACK: Fast sub-8-bit matrix multiply on commodity SIMD hardware. In Proceedings of the Machine Learning and Systems.D. Marculescu, Y. Chi, and C. Wu (Eds.), Vol. 4, 52\u201363. Retrieved from https:\/\/arxiv.org\/abs\/1812.00090https:\/\/proceedings.mlsys.org\/paper_files\/paper\/2022\/file\/e09d45e14e9ece7142217550ddd3c4d0-Paper.pdf"},{"key":"e_1_3_2_37_2","unstructured":"Bichen Wu Yanghan Wang Peizhao Zhang Yuandong Tian Peter Vajda and Kurt Keutzer. 2019. Mixed precision quantization of convnets via differentiable neural architecture search. In The Seventh International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_38_2","unstructured":"Xilinx. 2017. Deep Learning with INT8 Optimization on Xilinx Devices. Retrieved April 20 2025 from https:\/\/docs.xilinx.com\/v\/u\/en-US\/wp486-deep-learning-int8"},{"key":"e_1_3_2_39_2","unstructured":"Xilinx. 2020. Convolutional Neural Network with INT4 Optimization on Xilinx Devices. Retrieved April 20 2025 from https:\/\/docs.xilinx.com\/v\/u\/en-US\/wp521-4bit-optimization"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3803553","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T13:57:26Z","timestamp":1778680646000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3803553"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,13]]},"references-count":38,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,5,31]]}},"alternative-id":["10.1145\/3803553"],"URL":"https:\/\/doi.org\/10.1145\/3803553","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,13]]},"assertion":[{"value":"2025-06-03","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-22","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}