{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T22:40:58Z","timestamp":1779144058753,"version":"3.51.4"},"reference-count":49,"publisher":"Association for Computing Machinery (ACM)","issue":"6","funder":[{"name":"Chips Joint Undertaking and its members","award":["101112338"],"award-info":[{"award-number":["101112338"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2025,11,30]]},"abstract":"<jats:p>Deep neural networks (DNNs) have become indispensable in many real-life applications like natural language processing, and autonomous systems. However, deploying DNNs on resource-constrained devices, e.g., in RISC-V platforms, remains challenging due to the high computational and memory demands of fully connected (FC) layers, which dominate resource consumption. Low-rank factorization (LRF) offers an effective approach to compressing FC layers, but the vast design space of LRF solutions involves complex tradeoffs among FLOPs, memory size, inference time, and accuracy, making the LRF process complex and time-consuming. This article introduces an end-to-end LRF design space exploration methodology and a specialized design tool for optimizing FC layers on RISC-V processors. Using Tensor Train Decomposition (TTD) offered by TensorFlow T3F library, the proposed work prunes the LRF design space by excluding first, inefficient decomposition shapes and second, solutions with poor inference performance on RISC-V architectures. Compiler optimizations are then applied to enhance custom T3F layer performance, minimizing inference time and boosting computational efficiency. On average, our TT-decomposed layers run 3\u00d7 faster than IREE and 8\u00d7 faster than Pluto on the same compressed model. This work provides an efficient solution for deploying DNNs on edge and embedded devices powered by RISC-V architectures.<\/jats:p>\n                  <jats:p\/>","DOI":"10.1145\/3768624","type":"journal-article","created":{"date-parts":[[2025,9,18]],"date-time":"2025-09-18T11:36:36Z","timestamp":1758195396000},"page":"1-34","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Optimizing Tensor Train Decomposition in DNNs for RISC-V Architectures Using Design Space Exploration and Compiler Optimizations"],"prefix":"10.1145","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-3620-4637","authenticated-orcid":false,"given":"Theologos","family":"Anthimopoulos","sequence":"first","affiliation":[{"name":"School of Informatics, Aristotle University of Thessaloniki","place":["Thessaloniki, Greece"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-7430-3691","authenticated-orcid":false,"given":"Milad","family":"Kokhazadeh","sequence":"additional","affiliation":[{"name":"School of Informatics, Aristotle University of Thessaloniki","place":["Thessaloniki, Greece"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9591-913X","authenticated-orcid":false,"given":"Vasilios","family":"Kelefouras","sequence":"additional","affiliation":[{"name":"School of Engineering, Computing and Mathematics, University of Plymouth","place":["Plymouth, United Kingdom of Great Britain and Northern Ireland"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3737-8379","authenticated-orcid":false,"given":"Benjamin","family":"Himpel","sequence":"additional","affiliation":[{"name":"Informatics, Hochschule Reutlingen","place":["Reutlingen, Germany"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0460-6061","authenticated-orcid":false,"given":"Georgios","family":"Keramidas","sequence":"additional","affiliation":[{"name":"School of Informatics","place":["Thessaloniki, Greece"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,10,24]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2956508"},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2021.05.103"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2896880"},{"key":"e_1_3_3_5_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.mlwa.2021.100164"},{"key":"e_1_3_3_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2018.2876865"},{"issue":"16","key":"e_1_3_3_7_2","doi-asserted-by":"crossref","first-page":"7509","DOI":"10.1007\/s00500-021-06480-z","article-title":"Research on intelligent language translation system based on deep learning algorithm","volume":"26","author":"Shi Chunliu","year":"2022","unstructured":"Chunliu Shi. 2022. Research on intelligent language translation system based on deep learning algorithm. Soft Computing 26, 16 (2022), 7509\u20137518.","journal-title":"Soft Computing"},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-75855-4"},{"key":"e_1_3_3_9_2","first-page":"1","article-title":"Design possibilities and challenges of DNN models: A review on the perspective of end devices","author":"Hussain Hanan","year":"2022","unstructured":"Hanan Hussain, P. S. Tamizharasan, and C. S. Rahul. 2022. Design possibilities and challenges of DNN models: A review on the perspective of end devices. Artificial Intelligence Review 55, 7 (2022), 1\u201359.","journal-title":"Artificial Intelligence Review"},{"key":"e_1_3_3_10_2","first-page":"1","volume-title":"2024 IEEE International Conference on Omni-Layer Intelligent Systems (COINS)","author":"Liu Qiankun","year":"2024","unstructured":"Qiankun Liu, Sam Amiri, and Luciano Ost. 2024. Exploring RISC-V based DNN accelerators. In 2024 IEEE International Conference on Omni-Layer Intelligent Systems (COINS). IEEE, 1\u20136."},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2022.11.072"},{"issue":"1","key":"e_1_3_3_12_2","doi-asserted-by":"crossref","first-page":"20","DOI":"10.1007\/s10766-024-00762-3","article-title":"A practical approach for employing tensor train decomposition in edge devices","volume":"52","author":"Kokhazadeh Milad","year":"2024","unstructured":"Milad Kokhazadeh, Georgios Keramidas, Vasilios Kelefouras, and Iakovos Stamoulis. 2024. A practical approach for employing tensor train decomposition in edge devices. International Journal of Parallel Programming 52, 1 (2024), 20\u201339.","journal-title":"International Journal of Parallel Programming"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1137\/090752286"},{"key":"e_1_3_3_14_2","doi-asserted-by":"crossref","unstructured":"Ga\u00ebtan Frusque Gabriel Michau and Olga Fink. 2021. Canonical polyadic decomposition and deep learning for machine fault detection. arXiv:2107.09519. Retrieved from https:\/\/arxiv.org\/abs\/\/2107.09519.","DOI":"10.36001\/phme.2021.v6i1.2881"},{"key":"e_1_3_3_15_2","first-page":"5806","volume-title":"International Conference on Machine Learning","author":"Zhang Jiong","year":"2018","unstructured":"Jiong Zhang, Qi Lei, and Inderjit Dhillon. 2018. Stabilizing gradients for deep neural networks via efficient svd parameterization. In International Conference on Machine Learning. PMLR, 5806\u20135814."},{"issue":"30","key":"e_1_3_3_16_2","first-page":"1","article-title":"Tensor train decomposition on tensorflow (t3f)","volume":"21","author":"Novikov Alexander","year":"2020","unstructured":"Alexander Novikov, Pavel Izmailov, Valentin Khrulkov, Michael Figurnov, and Ivan Oseledets. 2020. Tensor train decomposition on tensorflow (t3f). Journal of Machine Learning Research 21, 30 (2020), 1\u20137.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2205.14479"},{"issue":"3","key":"e_1_3_3_18_2","first-page":"1","article-title":"PLUTO: An automatic parallelizer and locality optimizer for affine loop nests","volume":"30","author":"Bondhugula Uday","year":"2008","unstructured":"Uday Bondhugula et\u00a0al. 2008. PLUTO: An automatic parallelizer and locality optimizer for affine loop nests. ACM Transactions on Programming Languages and Systems (TOPLAS) 30, 3 (2008), 1\u201354.","journal-title":"ACM Transactions on Programming Languages and Systems (TOPLAS)"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","unstructured":"Alexander Novikov Dmitry Podoprikhin Anton Osokin and Dmitry Vetrov. 2015. Tensorizing Neural Networks. (2015). DOI: 10.48550\/ARXIV.1509.06569","DOI":"10.48550\/ARXIV.1509.06569"},{"issue":"26","key":"e_1_3_3_20_2","doi-asserted-by":"crossref","first-page":"753","DOI":"10.21105\/joss.00753","article-title":"Opt \\(\\backslash\\) _einsum-a python package for optimizing contraction order for einsum-like expressions","volume":"3","author":"Daniel G","year":"2018","unstructured":"G Daniel, Johnnie Gray, et\u00a0al. 2018. Opt \\(\\backslash\\) _einsum-a python package for optimizing contraction order for einsum-like expressions. Journal of Open Source Software 3, 26 (2018), 753.","journal-title":"Journal of Open Source Software"},{"key":"e_1_3_3_21_2","unstructured":"Arsenia Chorti and David Picard. 2021. Rate analysis and deep neural network detectors for SEFDM FTN systems. arXiv:2103.02306. Retrieved from https:\/\/arxiv.org\/abs\/\/2103.02306."},{"key":"e_1_3_3_22_2","first-page":"38060","article-title":"Model preserving compression for neural networks","volume":"35","author":"Chee Jerry","year":"2022","unstructured":"Jerry Chee, Anil Damle, Christopher M. De Sa, et\u00a0al. 2022. Model preserving compression for neural networks. Advances in Neural Information Processing Systems 35, 10 (2022), 38060\u201338074.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_3_23_2","unstructured":"Teck Kai Chan Cheng Siong Chin and Ye Li. 2020. Non-negative matrix factorization-convolutional neural network (NMF-CNN) for sound event detection. arXiv:2001.07874. Retrieved from https:\/\/arxiv.org\/abs\/\/2001.07874."},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2020.107538"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00807"},{"key":"e_1_3_3_26_2","unstructured":"Vadim Lebedev Yaroslav Ganin Maksim Rakhuba Ivan Oseledets and Victor Lempitsky. 2014. Speeding-up convolutional neural networks using fine-tuned cp-decomposition. arXiv:1412.6553. Retrieved from https:\/\/arxiv.org\/abs\/\/1412.6553."},{"key":"e_1_3_3_27_2","doi-asserted-by":"crossref","first-page":"118","DOI":"10.1109\/NICS51282.2020.9335842","volume-title":"2020 7th NAFOSTED Conference on Information and Computer Science (NICS)","author":"Mai An","year":"2020","unstructured":"An Mai, Loc Tran, Linh Tran, and Nguyen Trinh. 2020. VGG deep neural network compression via SVD and CUR decomposition techniques. In 2020 7th NAFOSTED Conference on Information and Computer Science (NICS). IEEE, 118\u2013123."},{"key":"e_1_3_3_28_2","first-page":"507","volume-title":"2020 IEEE 22nd International Conference on High Performance Computing and Communications; IEEE 18th International Conference on Smart City; IEEE 6th International Conference on Data Science and Systems (HPCC\/SmartCity\/DSS)","author":"Dai Cheng","year":"2020","unstructured":"Cheng Dai, Hongqiang Cheng, and Xingang Liu. 2020. A tucker decomposition based on adaptive genetic algorithm for efficient deep model compression. In 2020 IEEE 22nd International Conference on High Performance Computing and Communications; IEEE 18th International Conference on Smart City; IEEE 6th International Conference on Data Science and Systems (HPCC\/SmartCity\/DSS). IEEE, 507\u2013512."},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/BIGCOMP.2017.7881725"},{"key":"e_1_3_3_30_2","doi-asserted-by":"crossref","first-page":"522","DOI":"10.1007\/978-3-030-58526-6_31","volume-title":"Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, August 23\u201328, 2020, Proceedings, Part XXIX 16","author":"Phan Anh-Huy","year":"2020","unstructured":"Anh-Huy Phan, Konstantin Sobolev, Konstantin Sozykin, Dmitry Ermilov, Julia Gusak, Petr Tichavsk\u1ef3, Valeriy Glukhov, Ivan Oseledets, and Andrzej Cichocki. 2020. Stable low-rank tensor decomposition for compression of convolutional neural network. In Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, August 23\u201328, 2020, Proceedings, Part XXIX 16. Springer, 522\u2013539."},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","DOI":"10.5555\/2567709.2502582"},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.2968357"},{"key":"e_1_3_3_33_2","unstructured":"Walid Ahmed Habib Hajimolahoseini Austin Wen and Yang Liu. 2023. Speeding up resnet architecture with layers targeted low rank decomposition. arXiv:2309.12412. Retrieved from https:\/\/arxiv.org\/abs\/\/2309.12412."},{"key":"e_1_3_3_34_2","doi-asserted-by":"crossref","first-page":"347","DOI":"10.23919\/ICACT.2018.8323750","volume-title":"2018 20th International Conference on Advanced Communication Technology (ICACT)","author":"Astrid Marcella","year":"2018","unstructured":"Marcella Astrid, Seung-Ik Lee, and Beom-Su Seo. 2018. Rank selection of CP-decomposed convolutional layers with variational Bayesian matrix factorization. In 2018 20th International Conference on Advanced Communication Technology (ICACT). IEEE, 347\u2013350."},{"key":"e_1_3_3_35_2","first-page":"173","volume-title":"International Conference on Embedded Computer Systems","author":"Kokhazadeh Milad","year":"2022","unstructured":"Milad Kokhazadeh, Georgios Keramidas, Vasilios Kelefouras, and Iakovos Stamoulis. 2022. A design space exploration methodology for enabling tensor train decomposition in edge devices. In International Conference on Embedded Computer Systems. Springer, 173\u2013186."},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3649153.3649183"},{"key":"e_1_3_3_37_2","first-page":"1","volume-title":"2024 IEEE International Symposium on Circuits and Systems (ISCAS)","author":"Luo Yuan-June","year":"2024","unstructured":"Yuan-June Luo, Yu-Shan Tai, Ming-Guang Lin, and An-Yeu Andy Wu. 2024. Similarity-aware fast low-rank decomposition framework for vision transformers. In 2024 IEEE International Symposium on Circuits and Systems (ISCAS). IEEE, 1\u20135."},{"key":"e_1_3_3_38_2","article-title":"OneDNN: Open-source cross-platform performance library of basic building blocks for deep learning applications","author":"Corporation Intel","year":"2024","unstructured":"Intel Corporation. 2024. OneDNN: Open-source cross-platform performance library of basic building blocks for deep learning applications. Retrieved from https:\/\/github.com\/oneapi-src\/oneDNN. (2024). Accessed: 2024-11-24.","journal-title":"https:\/\/github.com\/oneapi-src\/oneDNN"},{"key":"e_1_3_3_39_2","unstructured":"NVIDIA Corporation. 2024. cuDNN: NVIDIA CUDA Deep Neural Network Library. NVIDIA Developer Santa Clara CA USA. Available at: https:\/\/developer.nvidia.com\/cudnn. (Accessed November 24 2025)."},{"key":"e_1_3_3_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/2764454"},{"key":"e_1_3_3_41_2","first-page":"137","volume-title":"Proceedings of the International Conference on Parallel Architectures and Compilation Techniques","author":"Le Thien T.","year":"2010","unstructured":"Thien T. Le et\u00a0al. 2010. PoCC: A polyhedral compiler collection for affine computations. In Proceedings of the International Conference on Parallel Architectures and Compilation Techniques. 137\u2013146."},{"key":"e_1_3_3_42_2","volume-title":"Proceedings of the 2011 International Conference on High Performance Computing","author":"Pupapetis K.","year":"2011","unstructured":"K. Pupapetis et\u00a0al. 2011. PPCG: Polyhedral parallel code generation. In Proceedings of the 2011 International Conference on High Performance Computing."},{"key":"e_1_3_3_43_2","first-page":"67","volume-title":"Proceedings of the 2017 International Conference on Compiler Construction","author":"Boulanger Gr\u00e9gory","year":"2017","unstructured":"Gr\u00e9gory Boulanger et\u00a0al. 2017. Tiramisu: A high-level language and code generation framework for high-performance stencil computations. In Proceedings of the 2017 International Conference on Compiler Construction. 67\u201378."},{"key":"e_1_3_3_44_2","unstructured":"Nadav Rotem Jerry Fix Asher Abdulrasool et\u00a0al. 2018. Glow: Graph lowering compiler techniques for neural networks. arXiv:1805.00907. Retrieved from https:\/\/arxiv.org\/abs\/\/1805.00907."},{"key":"e_1_3_3_45_2","unstructured":"2024. TensorFlow Lite for Microcontrollers. Retrieved from https:\/\/github.com\/tensorflow\/tflite-micro. (2024). Accessed: 2024-11-24."},{"key":"e_1_3_3_46_2","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1007\/978-3-031-42785-5_2","volume-title":"Architecture of Computing Systems","author":"Anthimopoulos Theologos","year":"2023","unstructured":"Theologos Anthimopoulos, Georgios Keramidas, Vasilios Kelefouras, and Iakovos Stamoulis. 2023. A comparative study of neural network compilers on ARMv8 architecture. In Architecture of Computing Systems. 18\u201322."},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3747183"},{"key":"e_1_3_3_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3649153.3649194"},{"key":"e_1_3_3_49_2","unstructured":"2025. TFLite Model Benchmark Tool. Retrieved from https:\/\/github.com\/sourcecode369\/tensorflow-1\/blob\/master\/tensorflow\/lite\/tools\/benchmark. (2025). Accessed: 2025-08-05."},{"key":"e_1_3_3_50_2","unstructured":"Zack Smith. 2024. Bandwidth a benchmark to estimate memory transfer bandwidth. (July2024.). Retrieved from https:\/\/zsmith.co\/bandwidth.php"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3768624","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,24]],"date-time":"2025-10-24T14:01:16Z","timestamp":1761314476000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3768624"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,24]]},"references-count":49,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,11,30]]}},"alternative-id":["10.1145\/3768624"],"URL":"https:\/\/doi.org\/10.1145\/3768624","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,24]]},"assertion":[{"value":"2025-02-17","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-08-31","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-10-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}