{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,1]],"date-time":"2025-10-01T15:14:33Z","timestamp":1759331673658,"version":"build-2065373602"},"reference-count":42,"publisher":"Association for Computing Machinery (ACM)","issue":"5s","funder":[{"name":"China Higher Education Institution Industry-University-Research Innovation Fund","award":["2024HY001"],"award-info":[{"award-number":["2024HY001"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2025,11,30]]},"abstract":"<jats:p>In power-constrained and real-time-demanding embedded scenarios, Field-Programmable Gate Arrays (FPGAs) emerge as ideal options for accelerating neural network inference, owing to the reconfigurability, high reliability, and flexibility of FPGAs. Mixed precision quantization technology significantly reduces computational complexity and bandwidth requirements while preserving model accuracy. However, existing FPGA accelerators fail to fully leverage the parallel advantages of mixed precision, which leads to the actual inference speedup being markedly lower than the theoretical prediction. Reviewing existing methods, we found three main drawbacks in enhancing practical computational performance. Firstly, FPGAs primarily rely on DSP slices to achieve high-performance parallel multiplication and accumulation (MAC) in neural network inference. However, the current DSP PE design and data packing methods are not compatible with mixed-precision models. Secondly, the remaining logic resources are not fully utilized to accelerate computations. Thirdly, the mixed-precision quantization bit-width selection method without hardware-guided guidance leads to additional model accuracy loss. To address these challenges, we propose a high-performance heterogeneous mixed-precision systolic array (SA) accelerator, HMSA. It aims to leverage mixed-precision quantization fully, enhancing the practical inference efficiency of neural networks on embedded FPGAs. We propose an optimized DSP data packing method guided by a resource-performance cost model, which enhances the parallel computing performance of accelerators. We propose a heterogeneous convolutional acceleration architecture based on SA architecture with high scalability, enabling efficient utilization of FPGA\u2019s heterogeneous computing resources. In terms of the algorithm, we propose an optimization method for bit-width selection based on FPGA hardware architecture, aiming to avoid accuracy degradation without improving inference speed. Experiments confirm that HMSA on the Xilinx XC7VX690T FPGA reaches a peak throughput of 6.385 TOP\/s at W1A8 precision. When inferring mixed precision neural networks, HMSA achieves 3.53\u00d7, 5.46\u00d7, and 1.58\u00d7 improvements in actual throughput\/DSP compared to state-of-the-art MPA, MSD, and MP-OPU. Compared with the state-of-the-art MBFQuant and Edge-MPQ mixed quantization algorithms, the proposed optimized bit-width selection method effectively reduces the model accuracy loss.<\/jats:p>","DOI":"10.1145\/3759458","type":"journal-article","created":{"date-parts":[[2025,8,7]],"date-time":"2025-08-07T11:13:16Z","timestamp":1754565196000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["HMSA: High-Performance Heterogeneous Mixed-Precision CNN Systolic Array Accelerator on FPGA"],"prefix":"10.1145","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-2020-5116","authenticated-orcid":false,"given":"Yongxiang","family":"Cao","sequence":"first","affiliation":[{"name":"School of computer science and engineering, Beihang University","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6633-6605","authenticated-orcid":false,"given":"Hongxu","family":"Jiang","sequence":"additional","affiliation":[{"name":"School of computer science and engineering, Beihang University","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0518-7313","authenticated-orcid":false,"given":"Huiyong","family":"Li","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, Beihang University","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-8186-4567","authenticated-orcid":false,"given":"Yu","family":"Tang","sequence":"additional","affiliation":[{"name":"School of computer science and engineering, Beihang University","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-1905-3870","authenticated-orcid":false,"given":"Dongcheng","family":"Shi","sequence":"additional","affiliation":[{"name":"School of computer science and engineering, Beihang University","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-5477-3020","authenticated-orcid":false,"given":"Guocheng","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of computer science and engineering, Beihang University","place":["Beijing, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,9,26]]},"reference":[{"key":"e_1_3_1_2_2","article-title":"Fpga-based real-time object detection and classification system using yolo for edge computing","author":"Amin Rashed Al","year":"2024","unstructured":"Rashed Al Amin, Mehrab Hasan, Veit Wiese, and Roman Obermaisser. 2024. Fpga-based real-time object detection and classification system using yolo for edge computing. IEEE Access (2024).","journal-title":"IEEE Access"},{"key":"e_1_3_1_3_2","volume-title":"7 Series DSP48E1 Slice User Guide (UG479)","author":"Xilinx AMD","year":"2024","unstructured":"AMD Xilinx. 2024. 7 Series DSP48E1 Slice User Guide (UG479). AMD Corporation, Santa Clara, CA. Retrieved from https:\/\/docs.amd.com\/v\/u\/en-US\/ug479_7Series_DSP48E1"},{"key":"e_1_3_1_4_2","first-page":"35","volume-title":"Proceedings of the 2018 28th International Conference on Field Programmable Logic and Applications (FPL)","author":"Boutros Andrew","year":"2018","unstructured":"Andrew Boutros, Sadegh Yazdanshenas, and Vaughn Betz. 2018. Embracing diversity: Enhanced DSP blocks for low-precision deep learning on FPGAs. In Proceedings of the 2018 28th International Conference on Field Programmable Logic and Applications (FPL). IEEE, 35\u2013357."},{"key":"e_1_3_1_5_2","first-page":"57","volume-title":"Proceedings of the IEEE 1st International Conference on Artificial Intelligence Circuits and Systems (AICAS)","author":"Camus V.","year":"2019","unstructured":"V. Camus, C. Enz, and M. Verhelst. 2019. Survey of precision-scalable multiply-accumulate units for neural-network processing. In Proceedings of the IEEE 1st International Conference on Artificial Intelligence Circuits and Systems (AICAS). 57\u201361."},{"issue":"4","key":"e_1_3_1_6_2","first-page":"697","article-title":"Arithmetic core generation using bit heaps","volume":"9","author":"Camus Vincent","year":"2019","unstructured":"Vincent Camus, Linyan Mei, Christian Enz, and Marian Verhelst. 2019. Arithmetic core generation using bit heaps. IEEE Journal on Emerging and Selected Topics in Circuits and Systems 9, 4 (Dec.2019), 697\u2013711.","journal-title":"IEEE Journal on Emerging and Selected Topics in Circuits and Systems"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2019.2950386"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2019.2950386"},{"issue":"4","key":"e_1_3_1_9_2","first-page":"697","article-title":"Small logic-based multipliers with incomplete sub-multipliers for FPGAs","volume":"9","author":"Camus Vincent","year":"2019","unstructured":"Vincent Camus, Linyan Mei, Christian Enz, and Marian Verhelst. 2019. Small logic-based multipliers with incomplete sub-multipliers for FPGAs. IEEE Journal on Emerging and Selected Topics in Circuits and Systems 9, 4 (Dec.2019), 697\u2013711.","journal-title":"IEEE Journal on Emerging and Selected Topics in Circuits and Systems"},{"key":"e_1_3_1_10_2","doi-asserted-by":"crossref","first-page":"208","DOI":"10.1109\/HPCA51647.2021.00027","volume-title":"Proceedings of the 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA)","author":"Chang Sung-En","year":"2021","unstructured":"Sung-En Chang, Yanyu Li, Mengshu Sun, Runbin Shi, Hayden K-H So, Xuehai Qian, Yanzhi Wang, and Xue Lin. 2021. Mix and match: A novel fpga-centric deep neural network quantization framework. In Proceedings of the 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 208\u2013220."},{"key":"e_1_3_1_11_2","doi-asserted-by":"crossref","first-page":"69","DOI":"10.1109\/ICFPT59805.2023.00013","volume-title":"Proceedings of the 2023 International Conference on Field Programmable Technology (ICFPT)","author":"Chen Yuzong","year":"2023","unstructured":"Yuzong Chen, Jordan Dotzel, and Mohamed S Abdelfattah. 2023. M4bram: Mixed-precision matrix-matrix multiplication in fpga block rams. In Proceedings of the 2023 International Conference on Field Programmable Technology (ICFPT). IEEE, 69\u201378."},{"key":"e_1_3_1_12_2","doi-asserted-by":"crossref","first-page":"222","DOI":"10.1109\/ISVLSI61997.2024.00049","volume-title":"Proceedings of the 2024 IEEE Computer Society Annual Symposium on VLSI (ISVLSI)","author":"Ghavami Behnam","year":"2024","unstructured":"Behnam Ghavami, Mahdi Sajadi, Lesley Shannon, and Steve Wilton. 2024. Boosting multiple multipliers packing on FPGA DSP blocks via truncation and compensation-based approximation. In Proceedings of the 2024 IEEE Computer Society Annual Symposium on VLSI (ISVLSI). IEEE, 222\u2013227."},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","first-page":"112","DOI":"10.1145\/3490422.3502367","volume-title":"Proceedings of the 2022 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays","author":"Gong Yu","year":"2022","unstructured":"Yu Gong, Zhihan Xu, Zhezhi He, Weifeng Zhang, Xiaobing Tu, Xiaoyao Liang, and Li Jiang. 2022. N3H-core: Neuron-designed neural network accelerator via FPGA-based heterogeneous computing cores. In Proceedings of the 2022 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays. 112\u2013122."},{"issue":"3","key":"e_1_3_1_14_2","doi-asserted-by":"crossref","first-page":"64","DOI":"10.1007\/s11554-024-01442-8","article-title":"Survey of convolutional neural network accelerators on field-programmable gate array platforms: Architectures and optimization techniques","volume":"21","author":"Hong Hyeonseok","year":"2024","unstructured":"Hyeonseok Hong, Dahun Choi, Namjoon Kim, Haein Lee, Beomjin Kang, Huibeom Kang, and Hyun Kim. 2024. Survey of convolutional neural network accelerators on field-programmable gate array platforms: Architectures and optimization techniques. Journal of Real-Time Image Processing 21, 3 (2024), 64.","journal-title":"Journal of Real-Time Image Processing"},{"key":"e_1_3_1_15_2","volume-title":"Intel\u00ae Stratix\u00ae 10 Variable Precision DSP Blocks User Guide","author":"Intel FPGA","year":"2017","unstructured":"FPGA Intel. 2017. Intel\u00ae Stratix\u00ae 10 Variable Precision DSP Blocks User Guide."},{"key":"e_1_3_1_16_2","first-page":"62414","article-title":"Pruning vs quantization: Which is better?","volume":"36","author":"Kuzmin Andrey","year":"2023","unstructured":"Andrey Kuzmin, Markus Nagel, Mart Van Baalen, Arash Behboodi, and Tijmen Blankevoort. 2023. Pruning vs quantization: Which is better? Advances in Neural inFormation Processing Systems 36 (2023), 62414\u201362427.","journal-title":"Advances in Neural inFormation Processing Systems"},{"issue":"4","key":"e_1_3_1_17_2","first-page":"697","article-title":"Design of high-throughput mixed-precision CNN accelerators on FPGA","volume":"9","author":"Latotzke Cecilia","year":"2019","unstructured":"Cecilia Latotzke, Tim Ciesielski, and Tobias Gemmeke. 2019. Design of high-throughput mixed-precision CNN accelerators on FPGA. IEEE Journal on Emerging and Selected Topics in Circuits and Systems 9, 4 (Dec.2019), 697\u2013711.","journal-title":"IEEE Journal on Emerging and Selected Topics in Circuits and Systems"},{"key":"e_1_3_1_18_2","first-page":"218","volume-title":"Proceedings of the IEEE International Solid-State Circuits Conference (ISSCC) Digest of Technical Papers","author":"Lee J.","year":"2018","unstructured":"J. Lee, C. Kim, S. Kang, D. Shin, S. Kim, and H.-J. Yoo. 2018. UNPU: A 50.6TOPS\/W unified deep neural network accelerator with 1b-to-16b fully-variable weight bit-precision. In Proceedings of the IEEE International Solid-State Circuits Conference (ISSCC) Digest of Technical Papers. 218\u2013220."},{"key":"e_1_3_1_19_2","article-title":"A mixed-precision transformer accelerator with vector tiling systolic array for license plate recognition in unconstrained scenarios","author":"Li Jie","year":"2024","unstructured":"Jie Li, Dingjiang Yan, Fangzhou He, Zhicheng Dong, and Mingfei Jiang. 2024. A mixed-precision transformer accelerator with vector tiling systolic array for license plate recognition in unconstrained scenarios. IEEE Transactions on Intelligent Transportation Systems (2024).","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"key":"e_1_3_1_20_2","first-page":"578","volume-title":"Proceedings of the 6th International Conference on Computer Information Science and Application Technology (CISAT 2023)","volume":"12800","author":"Li Runzhou","year":"2023","unstructured":"Runzhou Li, Bo Jiang, and Hong Xu. 2023. Mixed DSP packing method for convolutional neural network on FPGA. In Proceedings of the 6th International Conference on Computer Information Science and Application Technology (CISAT 2023), Vol. 12800. SPIE, 578\u2013585."},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2023.127210"},{"key":"e_1_3_1_22_2","first-page":"1","volume-title":"Proceedings of the 2024 IEEE International Symposium on Circuits and Systems (ISCAS)","author":"Liu Xinyan","year":"2024","unstructured":"Xinyan Liu, Xiao Wu, Haikuo Shao, and Zhongfeng Wang. 2024. A flexible FPGA-based accelerator for efficient inference of multi-precision CNNs. In Proceedings of the 2024 IEEE International Symposium on Circuits and Systems (ISCAS). IEEE, 1\u20135."},{"issue":"6","key":"e_1_3_1_23_2","doi-asserted-by":"crossref","first-page":"1051","DOI":"10.1109\/TCSVT.2014.2360030","article-title":"Evaluation and acceleration of high-throughput fixed-point object detection on FPGAs","volume":"25","author":"Ma Xiaoyin","year":"2014","unstructured":"Xiaoyin Ma, Walid A Najjar, and Amit K Roy-Chowdhury. 2014. Evaluation and acceleration of high-throughput fixed-point object detection on FPGAs. IEEE Transactions on Circuits and Systems for Video Technology 25, 6 (2014), 1051\u20131062.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_24_2","first-page":"6","volume-title":"Proceedings of the IEEE 1stInternational Conference on Artificial Intelligence Circuits and Systems (AICAS)","author":"Mei L.","year":"2019","unstructured":"L. Mei, M. Dandekar, D. Rodopoulos, J. Constantin; P. Debacker, and R. Lauwereins2019. Sub-word parallel precision-scalable MAC engines for efficient embedded DNN inference. In Proceedings of the IEEE 1stInternational Conference on Artificial Intelligence Circuits and Systems (AICAS). 6\u201310."},{"key":"e_1_3_1_25_2","first-page":"1","volume-title":"Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV)","author":"Moons B.","year":"2016","unstructured":"B. Moons, B. De Brabandere, L. Van Gool, and M. Verhelst. 2016. Energy-efficient ConvNets through approximate computing. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV). 1\u20138."},{"key":"e_1_3_1_26_2","first-page":"488","volume-title":"Proceedings of the Design, Automation and Test in Europe Conference and Exhibition (DATE)","author":"Moons B.","year":"2017","unstructured":"B. Moons, R. Uytterhoeven, W. Dehaene, and M. Verhelst. 2017. DVAFS: Trading computational accuracy for energy through dynamic-voltage-accuracy-frequency-scaling. In Proceedings of the Design, Automation and Test in Europe Conference and Exhibition (DATE). 488\u2013493."},{"key":"e_1_3_1_27_2","article-title":"A 119.64 GOPs\/W FPGA-Based ResNet50 mixed-precision accelerator using the dynamic DSP packing","author":"Ou Yaozhong","year":"2024","unstructured":"Yaozhong Ou, Wei-Han Yu, Ka-Fai Un, Chi-Hang Chan, and Yan Zhu. 2024. A 119.64 GOPs\/W FPGA-Based ResNet50 mixed-precision accelerator using the dynamic DSP packing. IEEE Transactions on Circuits and Systems II: Express Briefs (2024).","journal-title":"IEEE Transactions on Circuits and Systems II: Express Briefs"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.sysarc.2022.102561"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3268562"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/2847263.2847265"},{"key":"e_1_3_1_31_2","first-page":"20:1\u201320:6","volume-title":"Proceedings of the ACM\/EDAC\/IEEE Design Automation Conference (DAC)","author":"Sharify S.","year":"2018","unstructured":"S. Sharify, A. D. Lascorz, K. Siu, P. Judd, and A. Moshovos. 2018. Loom: Exploiting weight and activation precisions to accelerate convolutional neural networks. In Proceedings of the ACM\/EDAC\/IEEE Design Automation Conference (DAC). 20:1\u201320:6."},{"key":"e_1_3_1_32_2","first-page":"764","volume-title":"Proceedings of the 45th IEEE International Symposium on Computer Architecture (ISCA)","author":"Sharma H.","year":"2018","unstructured":"H. Sharma, J. Park, N. Suda, L. Lai, B. Chau, J. K. Kim, V. Chandra, and H. Esmaeilzadeh. 2018. Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural networks. In Proceedings of the 45th IEEE International Symposium on Computer Architecture (ISCA). 764\u2013775."},{"key":"e_1_3_1_33_2","doi-asserted-by":"crossref","first-page":"764","DOI":"10.1109\/ISCA.2018.00069","volume-title":"Proceedings of the 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA)","author":"Sharma Hardik","year":"2018","unstructured":"Hardik Sharma, Jongse Park, Naveen Suda, Liangzhen Lai, Benson Chau, Joon Kyung Kim, Vikas Chandra, and Hadi Esmaeilzadeh. 2018. Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural network. In Proceedings of the 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). IEEE, 764\u2013775."},{"key":"e_1_3_1_34_2","first-page":"240","volume-title":"Proceedings of the IEEE International Solid-State Circuits Conference (ISSCC) Digest of Technical Papers","author":"Shin D.","year":"2017","unstructured":"D. Shin, J. Lee, J. Lee, and H.-J. Yoo. 2017. DNPU: An 8.1TOPS\/W reconfigurable CNN-RNN processor for general-purpose deep neural networks. In Proceedings of the IEEE International Solid-State Circuits Conference (ISSCC) Digest of Technical Papers. 240\u2013241."},{"key":"e_1_3_1_35_2","first-page":"160","volume-title":"Proceedings of the 2022 32nd International Conference on Field-Programmable Logic and Applications (FPL)","author":"Sommer Jan","year":"2022","unstructured":"Jan Sommer, M Akif \u00d6zkan, Oliver Keszocze, and J\u00fcrgen Teich. 2022. Dsp-packing: Squeezing low-precision arithmetic into fpga dsp blocks. In Proceedings of the 2022 32nd International Conference on Field-Programmable Logic and Applications (FPL). IEEE, 160\u2013166."},{"key":"e_1_3_1_36_2","doi-asserted-by":"crossref","first-page":"66","DOI":"10.1109\/HPCA.2018.00016","volume-title":"Proceedings of the 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA)","author":"Song Mingcong","year":"2018","unstructured":"Mingcong Song, Jiaqi Zhang, Huixiang Chen, and Tao Li. 2018. Towards efficient microarchitectural design for accelerating unsupervised gan-based deep learning. In Proceedings of the 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 66\u201377."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3490422.3502364"},{"key":"e_1_3_1_38_2","article-title":"Low-bit mixed-precision quantization and acceleration of CNN for FPGA deployment","author":"Wang JianRong","year":"2024","unstructured":"JianRong Wang, Zhijun He, Hongbo Zhao, and Rongke Liu. 2024. Low-bit mixed-precision quantization and acceleration of CNN for FPGA deployment. IEEE Transactions on Emerging Topics in Computational Intelligence (2024).","journal-title":"IEEE Transactions on Emerging Topics in Computational Intelligence"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00881"},{"key":"e_1_3_1_40_2","first-page":"33","volume-title":"Proceedings of the 2021 31st International Conference on Field-Programmable Logic and Applications (FPL)","author":"Wu Chen","year":"2021","unstructured":"Chen Wu, Jinming Zhuang, Kun Wang, and Lei He. 2021. MP-OPU: A mixed precision FPGA-based overlay processor for convolutional neural networks. In Proceedings of the 2021 31st International Conference on Field-Programmable Logic and Applications (FPL). IEEE, 33\u201337."},{"key":"e_1_3_1_41_2","first-page":"94","volume-title":"Proceedings of the 2023 IEEE 31st Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM)","author":"Wu Jiajun","year":"2023","unstructured":"Jiajun Wu, Jiajun Zhou, Yizhao Gao, Yuhao Ding, Ngai Wong, and Hayden Kwok-Hay So. 2023. Msd: Mixing signed digit representations for hardware-efficient dnn acceleration on fpga with heterogeneous resources. In Proceedings of the 2023 IEEE 31st Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM). IEEE, 94\u2013104."},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.3390\/electronics14071345"},{"key":"e_1_3_1_43_2","article-title":"Edge-MPQ: Layer-wise mixed-precision quantization with tightly integrated versatile inference units for edge computing","author":"Zhao Xiaotian","year":"2024","unstructured":"Xiaotian Zhao, Ruge Xu, Yimin Gao, Vaibhav Verma, Mircea R Stan, and Xinfei Guo. 2024. Edge-MPQ: Layer-wise mixed-precision quantization with tightly integrated versatile inference units for edge computing. IEEE Transactions on Computers (2024).","journal-title":"IEEE Transactions on Computers"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3759458","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,30]],"date-time":"2025-09-30T13:44:24Z","timestamp":1759239864000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3759458"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,26]]},"references-count":42,"journal-issue":{"issue":"5s","published-print":{"date-parts":[[2025,11,30]]}},"alternative-id":["10.1145\/3759458"],"URL":"https:\/\/doi.org\/10.1145\/3759458","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"type":"print","value":"1539-9087"},{"type":"electronic","value":"1558-3465"}],"subject":[],"published":{"date-parts":[[2025,9,26]]},"assertion":[{"value":"2025-07-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-29","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-26","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}