{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T16:11:31Z","timestamp":1777651891863,"version":"3.51.4"},"reference-count":27,"publisher":"Association for Computing Machinery (ACM)","issue":"6","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2025,11,30]]},"abstract":"<jats:p>Predictable wavefront splitting (PWS) is an optimization technique for graphics processing units (GPUs) to address the performance and worst-case execution time (WCET) impacts of branch divergence. PWS relies on manual annotation by the GPU programmer; these choices affect the resulting WCET. This work automates this process with two key approaches. First, we formulate the optimal annotation as an integer quadratic programming (IQP) problem such that the solution guarantees the lowest WCET. Second, we show that the problem can be solved with an optimal polynomial-time dynamic programming algorithm that achieves the same solutions as the IQP. We implement our algorithm in a compiler flow for an AMD GPU, and we deploy the annotated executable on a gem5 micro-architectural implementation of the AMD GCN3 GPU. We evaluate our implementation on a benchmark suite provided by AMD and supplement it with an extensive set of synthetic benchmarks. Our evaluation shows that these two approaches are able to reduce the WCET by between 13% and 31% compared to five baseline algorithms.<\/jats:p>","DOI":"10.1145\/3769118","type":"journal-article","created":{"date-parts":[[2025,9,23]],"date-time":"2025-09-23T11:36:39Z","timestamp":1758627399000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Optimal Split Point Placement for Predictable GPU Wavefront Splitting"],"prefix":"10.1145","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-8458-8512","authenticated-orcid":false,"given":"Artem","family":"Klashtorny","sequence":"first","affiliation":[{"name":"Electrical and Computer Engineering, University of Waterloo","place":["Waterloo, Canada"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3615-9393","authenticated-orcid":false,"given":"Mahesh","family":"Tripunitara","sequence":"additional","affiliation":[{"name":"Electrical and Computer Engineering, University of Waterloo","place":["Waterloo, Canada"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2750-4471","authenticated-orcid":false,"given":"Hiren","family":"Patel","sequence":"additional","affiliation":[{"name":"Electrical and Computer Engineering, University of Waterloo","place":["Waterloo, Canada"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,10,24]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"AMD. 2024. ROCm Examples. Retrieved May 22 2024 from https:\/\/github.com\/ROCm\/rocm-examples. (2024)."},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTSS.2017.00017"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2009.4919648"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/ECRTS.2013.29"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2012.6237005"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2009.5306797"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2012.18"},{"key":"e_1_3_1_9_2","first-page":"1184","volume-title":"Proceedings of the International Symposium on High-Performance Computer Architecture","author":"Damani Sana","year":"2022","unstructured":"Sana Damani, Mark Stephenson, Ram Rangan, Daniel Johnson, Rishkul Kulkarni, and Stephen W. Keckler. 2022. GPU subwarp interleaving. In Proceedings of the International Symposium on High-Performance Computer Architecture. IEEE, New Jersey, USA, 1184\u20131197."},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2007.30"},{"key":"e_1_3_1_11_2","unstructured":"Gurobi Optimization LLC. 2024. Gurobi Optimizer Reference Manual. (2024). https:\/\/www.gurobi.com"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.4230\/OASIcs.WCET.2014.43"},{"key":"e_1_3_1_13_2","first-page":"1","volume-title":"Proceedings of the 2024 Design, Automation & Test in Europe Conference & Exhibition","author":"Klashtorny Artem","year":"2024","unstructured":"Artem Klashtorny, Mahesh Tripunitara, and Hiren Patel. 2024. A compiler phase to optimally split GPU wavefronts for safety-critical systems. In Proceedings of the 2024 Design, Automation & Test in Europe Conference & Exhibition. IEEE, New Jersey, USA, 1\u20136."},{"issue":"5","key":"e_1_3_1_14_2","first-page":"25","article-title":"Predictable GPU wavefront splitting for safety-critical systems","volume":"22","author":"Klashtorny Artem","year":"2023","unstructured":"Artem Klashtorny, Zhuanhao Wu, Anirudh Mohan Kaushik, and Hiren Patel. 2023. Predictable GPU wavefront splitting for safety-critical systems. ACM Transactions on Embedded Computing Systems 22, 5s (2023), 25 pages.","journal-title":"ACM Transactions on Embedded Computing Systems"},{"key":"e_1_3_1_15_2","unstructured":"Artem Klashtorny Zhuanhao Wu Anirudh Mohan Kaushik and Hiren Patel. 2023. PWS GPU. Retrieved June 6 2024 from https:\/\/github.com\/caesr-uwaterloo\/gem5-pws. (2023)."},{"key":"e_1_3_1_16_2","unstructured":"Artem Klashtorny Zhuanhao Wu Anirudh Mohan Kaushik and Hiren Patel. 2024. Split Point Placement Solver. Retrieved June 6 2024 from https:\/\/github.com\/caesr-uwaterloo\/pws-solver. (2024)."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2020.02.003"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3242089"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/1815961.1815992"},{"key":"e_1_3_1_20_2","first-page":"308","volume-title":"Proceedings of the 2011 44th Annual IEEE\/ACM International Symposium on Microarchitecture","author":"Narasiman Veynu","year":"2011","unstructured":"Veynu Narasiman, Michael Shebanow, Chang Joo Lee, Rustam Miftakhutdinov, Onur Mutlu, and Yale N. Patt. 2011. Improving GPU performance via large warps and two-level warp scheduling. In Proceedings of the 2011 44th Annual IEEE\/ACM International Symposium on Microarchitecture. IEEE, New Jersey, USA, 308\u2013317."},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.4230\/LIPIcs.ECRTS.2020.10"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3453417.3453432"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.4230\/LIPIcs.ECRTS.2021.1"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2013.6522352"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750410"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2021.3064290"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.4230\/LIPIcs.ECRTS.2018.20"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/1810085.1810104"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3769118","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,24]],"date-time":"2025-10-24T14:01:10Z","timestamp":1761314470000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3769118"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,24]]},"references-count":27,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,11,30]]}},"alternative-id":["10.1145\/3769118"],"URL":"https:\/\/doi.org\/10.1145\/3769118","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,24]]},"assertion":[{"value":"2025-01-13","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-15","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-10-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}