{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,24]],"date-time":"2025-10-24T16:50:28Z","timestamp":1761324628392,"version":"3.44.0"},"reference-count":34,"publisher":"Association for Computing Machinery (ACM)","issue":"5","funder":[{"name":"National Authorities of Italy, Turkey, Portugal, The Netherlands, Czech Republic, Latvia, Greece, and Romania","award":["101112338"],"award-info":[{"award-number":["101112338"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2025,9,30]]},"abstract":"<jats:p>Register Blocking (RB), also known as \u2018Register-level Tiling\u2019 or \u2018unroll-and-jam,\u2019 is a key compiler optimization for developing efficient micro-kernels. However, applying RB effectively is a complex task due to several challenges. First, the exploration space of possible RB configurations is vast. Second, RB and loop permutation are interdependent; therefore, addressing both optimizations simultaneously further inflates the exploration space. Third, the effectiveness of RB is highly dependent on the target hardware platform and the specific loop kernel being optimized. As a result, an extensive and time-consuming fine-tuning process is necessary for achieving an efficient implementation.<\/jats:p>\n          <jats:p>To address these challenges, a source-to-source analytical modelling approach is proposed. The RB factors, the loops to apply RB, the number of allocated variables\/registers per array reference, and the loops\u2019 ordering are generated by an analytical model, leveraging the target hardware architecture details and loop kernel characteristics. The proposed methodology has been evaluated on both embedded and general-purpose CPUs, using seven well-known loop kernels and three machine learning applications. The results show significant speedups over the GCC compiler, the Pluto tool, and related work.<\/jats:p>\n          <jats:p\/>","DOI":"10.1145\/3747183","type":"journal-article","created":{"date-parts":[[2025,7,5]],"date-time":"2025-07-05T06:54:32Z","timestamp":1751698472000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Register Blocking: A Source-to-Source Analytical Modelling Approach for Affine Loop Kernels"],"prefix":"10.1145","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-3620-4637","authenticated-orcid":false,"given":"Theologos","family":"Anthimopoulos","sequence":"first","affiliation":[{"name":"School of Informatics, Aristotle University of Thessaloniki","place":["Thessaloniki, Greece"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0460-6061","authenticated-orcid":false,"given":"Georgios","family":"Keramidas","sequence":"additional","affiliation":[{"name":"School of Informatics, Aristotle University of Thessaloniki","place":["Thessalonike, Greece"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9591-913X","authenticated-orcid":false,"given":"Vasilios","family":"Kelefouras","sequence":"additional","affiliation":[{"name":"School of Engineering, Computing and Mathematics, University of Plymouth","place":["Plymouth, United Kingdom of Great Britain and Northern Ireland"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-4624-6089","authenticated-orcid":false,"given":"Iakovos","family":"Stamoulis","sequence":"additional","affiliation":[{"name":"Think Silicon S.A., An Applied Materials Company","place":["Patras, Greece"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,9,13]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3416510"},{"key":"e_1_3_2_3_2","unstructured":"AlexNet. https:\/\/cvml.ista.ac.at\/courses\/DLWT_W17\/material\/AlexNet.pdf. Accessed on August 5 2025."},{"volume-title":"Proceedings of the Computing Frontiers Conference.","author":"Anthimopoulos T.","key":"e_1_3_2_4_2","unstructured":"T. Anthimopoulos, G. Keramidas, V. I. Kelefouras, and I. Stamoulis. 2024. Register blocking: An analytical modelling approach for affine loop kernels. In Proceedings of the Computing Frontiers Conference."},{"volume-title":"Proceedings of the International Conference on Programming Language Design and Implementation.","author":"Bondhugula U.","key":"e_1_3_2_5_2","unstructured":"U. Bondhugula, A. Hartono, J. Ramanujam, and P. Sadayappan. 2008. A practical automatic polyhedral parallelizer and locality optimizer. In Proceedings of the International Conference on Programming Language Design and Implementation."},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/197320.197366"},{"volume-title":"Proceedings of the International Symposium on Microarchitecture.","author":"Carr S.","key":"e_1_3_2_7_2","unstructured":"S. Carr and Y. Guan. 1997. Unroll-and-jam using uniformly generated sets. In Proceedings of the International Symposium on Microarchitecture."},{"volume-title":"Proceedings of the International Symposium on Code Generation and Optimization.","author":"Chen C.","key":"e_1_3_2_8_2","unstructured":"C. Chen, J. Chame, and M. Hall. 2005. Combining models and guided empirical search to optimize for multiple levels of the memory hierarchy. In Proceedings of the International Symposium on Code Generation and Optimization."},{"volume-title":"Proceedings of the International Conference ParCo.","author":"Herruzo E.","key":"e_1_3_2_9_2","unstructured":"E. Herruzo, G. Bandera, E. L. Zapata, and O. Plata. 2006. Reducing cache misses by loop reordering. In Proceedings of the International Conference ParCo."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-019-02880-z"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2023.3322037"},{"volume-title":"Proceedings of the International Conference on Programming Language Design and Implementation.","author":"Kong M.","key":"e_1_3_2_12_2","unstructured":"M. Kong and L. N. Pouchet. 2019. Model-driven transformations for multi- and many-core CPUs. In Proceedings of the International Conference on Programming Language Design and Implementation."},{"key":"e_1_3_2_13_2","unstructured":"M. Kong and L. N. Pouchet. 2018. A performance vocabulary for affine loop transformations. Retrieved August 2025 from https:\/\/arxiv.org\/abs\/1811.06043"},{"key":"e_1_3_2_14_2","unstructured":"L. Lai N. Suda and V. Chandra. 2018. CMSIS-NN: Efficient neural network kernels for arm cortex-M CPUs. Retrieved August 2025 from https:\/\/arxiv.org\/abs\/1801.06601"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_3_2_16_2","unstructured":"J. Li Z. Qin Y. Mei J. Cui Y. Song C. Chen Y. Zhang L. Du X. Cheng B. Jin J. Ye E. Lin and D. Lavery. 2023. oneDNN graph compiler: A hybrid approach for high-performance deep learning compilation. Retrieved August 2025 from https:\/\/arxiv.org\/abs\/2301.01333"},{"volume-title":"Proceedings of the MATEC Web of Conferences.","author":"Liu X.","key":"e_1_3_2_17_2","unstructured":"X. Liu, L. Ding, Y. Li, G. Chen, and J. Du. 2018. Research of register pressure aware loop unrolling optimizations for compiler. In Proceedings of the MATEC Web of Conferences."},{"key":"e_1_3_2_18_2","unstructured":"LLVM Compiler: https:\/\/github.com\/LLVM\/LLVM-project\/issues\/38004. Accessed on August 5 2025."},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2020.01.044"},{"volume-title":"Proceedings of the Encyclopedia of Parallel Computing.","author":"Meister B.","key":"e_1_3_2_20_2","unstructured":"B. Meister, N. Vasilache, D. Wohlford, M. Baskaran, A. Leung, and R. Lethin. 2011. R-stream compiler. In Proceedings of the Encyclopedia of Parallel Computing."},{"key":"e_1_3_2_21_2","unstructured":"Orio Tool: https:\/\/github.com\/brnorris03\/Orio. Accessed on August 5 2025."},{"key":"e_1_3_2_22_2","unstructured":"Pluto Tool: https:\/\/Pluto-compiler.sourceforge.net\/. Accessed on August 5 2025."},{"key":"e_1_3_2_23_2","unstructured":"Pmbw tool: https:\/\/panthema.net\/2013\/pmbw\/. Accessed on August 5 2025."},{"volume-title":"Proceedings of the International Symposium on Principles of Programming Languages.","author":"Pouchet L. N.","key":"e_1_3_2_24_2","unstructured":"L. N. Pouchet, U. Bondhugula, C. Bastoul, A. Cohen, J. Ramanujam, P. Sadayappan, and N. Vasilache. 2011. Loop transformations: Convexity, pruning and optimization. In Proceedings of the International Symposium on Principles of Programming Languages."},{"key":"e_1_3_2_25_2","unstructured":"Polly Tool: https:\/\/polly.LLVM.org\/docs\/UsingPollyWithClang.html. Accessed on August 5 2025."},{"key":"e_1_3_2_26_2","unstructured":"Github url Register Blocking Source-to-Source:https:\/\/github.com\/Theoo1997\/RB_s2s. Accessed on August 5 2025."},{"key":"e_1_3_2_27_2","first-page":"397","article-title":"Optimized unrolling of nested loops","volume":"29","author":"Sarkar V.","year":"2001","unstructured":"V. Sarkar. 2001. Optimized unrolling of nested loops. Journal of Parallel Programming 29, 5 (2001), 397\u2013423.","journal-title":"Journal of Parallel Programming"},{"key":"e_1_3_2_28_2","volume-title":"Fast transformer decoding: One write-head is all you need. Retrieved","author":"Shazeer N.","year":"2025","unstructured":"N. Shazeer. 2019. Fast transformer decoding: One write-head is all you need. Retrieved August 2025 from https:\/\/arxiv.org\/abs\/1911.02150"},{"key":"e_1_3_2_29_2","unstructured":"Valgrind Tool: https:\/\/valgrind.org\/. Accessed on August 5 2025."},{"key":"e_1_3_2_30_2","unstructured":"TensorFlow: https:\/\/www.tensorflow.org\/api_docs\/python\/tf\/einsum. Accessed on August 5 2025."},{"volume-title":"Proceedings of the International Workshop on Polyhedral Compilation Techniques.","author":"Vasilache N.","key":"e_1_3_2_31_2","unstructured":"N. Vasilache, B. Meister, M. Baskaran, and R. Lethin. 2012. Joint scheduling and layout optimization to enable multi-level vectorization. In Proceedings of the International Workshop on Polyhedral Compilation Techniques."},{"volume-title":"International Conference on Programming Languages.","author":"Wilkinson L.","key":"e_1_3_2_32_2","unstructured":"L. Wilkinson, K. Cheshmi, and M. M. Dehnavi. 2023. Register tiling for unstructured sparsity in neural network inference. International Conference on Programming Languages."},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1145\/3570641"},{"key":"e_1_3_2_34_2","unstructured":"Zack Smith tool: https:\/\/zs3.me\/bandwidth.php. Accessed on August 5 2025."},{"volume-title":"International Conference on Compiler Construction.","author":"Zinenko O.","key":"e_1_3_2_35_2","unstructured":"O. Zinenko, S. Verdoolaege, C. Reddy, J. Shirako, T. Grosser, V. Sarkar, and A. Cohen. 2018. Modeling the conflicting demands of parallelism and temporal\/spatial locality in affine scheduling. International Conference on Compiler Construction."}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3747183","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,13]],"date-time":"2025-09-13T13:44:31Z","timestamp":1757771071000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3747183"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,13]]},"references-count":34,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2025,9,30]]}},"alternative-id":["10.1145\/3747183"],"URL":"https:\/\/doi.org\/10.1145\/3747183","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"type":"print","value":"1539-9087"},{"type":"electronic","value":"1558-3465"}],"subject":[],"published":{"date-parts":[[2025,9,13]]},"assertion":[{"value":"2024-09-24","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-19","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}