{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T06:29:48Z","timestamp":1781677788001,"version":"3.54.5"},"reference-count":30,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2020,8,3]],"date-time":"2020-08-03T00:00:00Z","timestamp":1596412800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2020,9,30]]},"abstract":"<jats:p>The Halide DSL and compiler have enabled high-performance code generation for image processing pipelines targeting heterogeneous architectures through the separation of algorithmic description and optimization schedule. However, automatic schedule generation is currently only possible for multi-core CPU architectures. As a result, expert knowledge is still required when optimizing for platforms with GPU capabilities. In this work, we extend the current Halide Autoscheduler with novel optimization passes to efficiently generate schedules for CUDA-based GPU architectures. We evaluate our proposed method across a variety of applications and show that it can achieve performance competitive with that of manually tuned Halide schedules, or in many cases even better performance. Experimental results show that our schedules are on average 10% faster than manual schedules and over 2\u00d7 faster than previous autoscheduling attempts.<\/jats:p>","DOI":"10.1145\/3406117","type":"journal-article","created":{"date-parts":[[2020,8,3]],"date-time":"2020-08-03T23:17:08Z","timestamp":1596496628000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":14,"title":["Schedule Synthesis for Halide Pipelines on GPUs"],"prefix":"10.1145","volume":"17","author":[{"given":"Savvas","family":"Sioutas","sequence":"first","affiliation":[{"name":"Eindhoven University of Technology, Eindhoven, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sander","family":"Stuijk","sequence":"additional","affiliation":[{"name":"Eindhoven University of Technology, Eindhoven, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Twan","family":"Basten","sequence":"additional","affiliation":[{"name":"Eindhoven University of Technology and ESI, TNO, Eindhoven, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Henk","family":"Corporaal","sequence":"additional","affiliation":[{"name":"Eindhoven University of Technology, Eindhoven, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lou","family":"Somers","sequence":"additional","affiliation":[{"name":"Canon Production Printing and Eindhoven University of Technology, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,8,3]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3306346.3322967"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2628071.2628092"},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the International Conference on Field Programmable Logic and Applications. 126--131","author":"Asano S.","year":"2009","unstructured":"S. Asano , T. Maruyama , and Y. Yamaguchi . 2009. Performance comparison of FPGA, GPU, and CPU in image processing . In Proceedings of the International Conference on Field Programmable Logic and Applications. 126--131 . DOI:https:\/\/doi.org\/10.1109\/FPL. 2009 .5272532 S. Asano, T. Maruyama, and Y. Yamaguchi. 2009. Performance comparison of FPGA, GPU, and CPU in image processing. In Proceedings of the International Conference on Field Programmable Logic and Applications. 126--131. DOI:https:\/\/doi.org\/10.1109\/FPL.2009.5272532"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2019.8661197"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2018.2872064"},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201918)","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen , Thierry Moreau , Ziheng Jiang , Lianmin Zheng , Eddie Yan , Meghan Cowan , Haichen Shen , Leyuan Wang , Yuwei Hu , Luis Ceze , Carlos Guestrin , and Arvind Krishnamurthy . 2018 . TVM: An automated end-to-end optimizing compiler for deep learning . In Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201918) . USENIX Association, Berkeley, CA, 579--594. Retrieved from http:\/\/dl.acm.org\/citation.cfm?id=3291168.3291211. Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Meghan Cowan, Haichen Shen, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: An automated end-to-end optimizing compiler for deep learning. In Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation (OSDI\u201918). USENIX Association, Berkeley, CA, 579--594. Retrieved from http:\/\/dl.acm.org\/citation.cfm?id=3291168.3291211."},{"key":"e_1_2_1_7_1","volume-title":"NVIDIA Nsight Compute.","author":"NVIDIA Corporation","year":"2019","unstructured":"NVIDIA Corporation . 2019. NVIDIA Nsight Compute. Retrieved from https:\/\/developer.nvidia.com\/nsight-compute-2019_5 version 2019 .5.0. NVIDIA Corporation. 2019. NVIDIA Nsight Compute. Retrieved from https:\/\/developer.nvidia.com\/nsight-compute-2019_5 version 2019.5.0."},{"key":"e_1_2_1_8_1","unstructured":"Halide. 2018. Halide GitHub Repository (MIT License). Retrieved from https:\/\/github.com\/halide\/Halide.  Halide. 2018. Halide GitHub Repository (MIT License). Retrieved from https:\/\/github.com\/halide\/Halide."},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the 26th ACM International Conference on Supercomputing (ICS\u201912)","author":"Holewinski Justin","unstructured":"Justin Holewinski , Louis-No\u00ebl Pouchet , and P. Sadayappan . 2012. High-performance code generation for stencil computations on GPU architectures . In Proceedings of the 26th ACM International Conference on Supercomputing (ICS\u201912) . Association for Computing Machinery, New York, NY, 311--320. DOI:https:\/\/doi.org\/10.1145\/2304576.2304619 Justin Holewinski, Louis-No\u00ebl Pouchet, and P. Sadayappan. 2012. High-performance code generation for stencil computations on GPU architectures. In Proceedings of the 26th ACM International Conference on Supercomputing (ICS\u201912). Association for Computing Machinery, New York, NY, 311--320. DOI:https:\/\/doi.org\/10.1145\/2304576.2304619"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201383"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/233561.233564"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2015.2394802"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925952"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2786763.2694364"},{"key":"e_1_2_1_15_1","unstructured":"Nvidia. 2019. Cuda occupancy calculator. Retrieved from https:\/\/docs.nvidia.com\/cuda\/cuda-occupancy-calculator\/index.html.  Nvidia. 2019. Cuda occupancy calculator. Retrieved from https:\/\/docs.nvidia.com\/cuda\/cuda-occupancy-calculator\/index.html."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3018743.3018744"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2184319.2184337"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3207719.3207723"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2019.8661176"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2491956.2462176"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2716282.2716290"},{"key":"e_1_2_1_22_1","volume-title":"Proceedings of the International Conference on Parallel Architectures and Compilation (PACT\u201916)","author":"Rawat Prashant Singh","unstructured":"Prashant Singh Rawat , Changwan Hong , Mahesh Ravishankar , Vinod Grover , Louis-Noel Pouchet , Atanas Rountev , and P. Sadayappan . 2016. Resource conscious reuse-driven tiling for GPUs . In Proceedings of the International Conference on Parallel Architectures and Compilation (PACT\u201916) . Association for Computing Machinery, New York, NY, 99--111. DOI:https:\/\/doi.org\/10.1145\/2967938.2967967 Prashant Singh Rawat, Changwan Hong, Mahesh Ravishankar, Vinod Grover, Louis-Noel Pouchet, Atanas Rountev, and P. Sadayappan. 2016. Resource conscious reuse-driven tiling for GPUs. In Proceedings of the International Conference on Parallel Architectures and Compilation (PACT\u201916). Association for Computing Machinery, New York, NY, 99--111. DOI:https:\/\/doi.org\/10.1145\/2967938.2967967"},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of the International Symposium on Code Generation and Optimization (CGO\u201918)","author":"Sioutas Savvas","year":"2018","unstructured":"Savvas Sioutas , Sander Stuijk , Henk Corp oraal, Twan Basten , and Lou Somers . 2018 . Loop transformations leveraging hardware prefetching . In Proceedings of the International Symposium on Code Generation and Optimization (CGO\u201918) . ACM, New York, NY, 254--264. DOI:https:\/\/doi.org\/10.1145\/3168823 Savvas Sioutas, Sander Stuijk, Henk Corporaal, Twan Basten, and Lou Somers. 2018. Loop transformations leveraging hardware prefetching. In Proceedings of the International Symposium on Code Generation and Optimization (CGO\u201918). ACM, New York, NY, 254--264. DOI:https:\/\/doi.org\/10.1145\/3168823"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3310248"},{"key":"e_1_2_1_25_1","volume-title":"the Regents of the University of California","author":"Lawrence Berkeley National Laboratory","year":"2019","unstructured":"Lawrence Berkeley National Laboratory , the Regents of the University of California . 2019 . Empirical Roofline Tool (ERT). Retrieved from https:\/\/bitbucket.org\/berkeleylab\/cs-roofline-toolkit\/src\/master\/. Lawrence Berkeley National Laboratory, the Regents of the University of California. 2019. Empirical Roofline Tool (ERT). Retrieved from https:\/\/bitbucket.org\/berkeleylab\/cs-roofline-toolkit\/src\/master\/."},{"key":"e_1_2_1_26_1","volume-title":"Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions. CoRR abs\/1802.04730","author":"Vasilache Nicolas","year":"2018","unstructured":"Nicolas Vasilache , Oleksandr Zinenko , Theodoros Theodoridis , Priya Goyal , Zachary DeVito , William S. Moses , Sven Verdoolaege , Andrew Adams , and Albert Cohen . 2018. Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions. CoRR abs\/1802.04730 ( 2018 ). Nicolas Vasilache, Oleksandr Zinenko, Theodoros Theodoridis, Priya Goyal, Zachary DeVito, William S. Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen. 2018. Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions. CoRR abs\/1802.04730 (2018)."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2400682.2400713"},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC\u201914)","author":"Wahib M.","year":"2014","unstructured":"M. Wahib and N. Maruyama . 2014. Scalable kernel fusion for memory-bound GPU applications . In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC\u201914) . 191--202. DOI:https:\/\/doi.org\/10.1109\/SC. 2014 .21 M. Wahib and N. Maruyama. 2014. Scalable kernel fusion for memory-bound GPU applications. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC\u201914). 191--202. DOI:https:\/\/doi.org\/10.1109\/SC.2014.21"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/GreenCom-CPSCom.2010.102"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/1498765.1498785"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3406117","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3406117","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:31:52Z","timestamp":1750195912000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3406117"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,8,3]]},"references-count":30,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2020,9,30]]}},"alternative-id":["10.1145\/3406117"],"URL":"https:\/\/doi.org\/10.1145\/3406117","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,8,3]]},"assertion":[{"value":"2019-11-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-06-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-08-03","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}