{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T18:11:58Z","timestamp":1783534318201,"version":"3.55.0"},"reference-count":58,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2022,9,16]],"date-time":"2022-09-16T00:00:00Z","timestamp":1663286400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Booz Allen Hamilton Inc."}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2022,12,31]]},"abstract":"<jats:p>\n            As CUDA becomes the de facto programming language among data parallel applications such as high-performance computing or machine learning applications, running CUDA on other platforms becomes a compelling option. Although several efforts have attempted to support CUDA on devices other than NVIDIA GPUs, due to extra steps in the translation, the support is always a few years behind CUDA\u2019s latest features. In particular, the new CUDA programming model exposes the warp concept in the programming language, which greatly changes the way the CUDA code should be mapped to CPU programs. In this article, hierarchical collapsing that\n            <jats:italic>correctly<\/jats:italic>\n            supports CUDA warp-level functions on CPUs is proposed. To verify hierarchical collapsing , we build a framework, COX , that supports executing CUDA source code on the CPU backend. With hierarchical collapsing , 90% of kernels in CUDA SDK samples can be executed on CPUs, much higher than previous works (68%). We also evaluate the performance with benchmarks for real applications and show that hierarchical collapsing can generate CPU programs with comparable or even higher performance than previous projects in general.\n          <\/jats:p>","DOI":"10.1145\/3554736","type":"journal-article","created":{"date-parts":[[2022,8,2]],"date-time":"2022-08-02T11:04:27Z","timestamp":1659438267000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["COX : Exposing CUDA Warp-level Functions to CPUs"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3090-3951","authenticated-orcid":false,"given":"Ruobing","family":"Han","sequence":"first","affiliation":[{"name":"Georgia Institute of Technology, North Avenue Atlanta, GA , USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0768-384X","authenticated-orcid":false,"given":"Jaewon","family":"Lee","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, North Avenue Atlanta, GA , USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0403-9928","authenticated-orcid":false,"given":"Jaewoong","family":"Sim","sequence":"additional","affiliation":[{"name":"Seoul National University, Gwanak-gu, Seoul, South Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6061-7825","authenticated-orcid":false,"given":"Hyesoon","family":"Kim","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, North Avenue Atlanta, GA , USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,9,16]]},"reference":[{"key":"e_1_3_3_2_2","unstructured":"Compiling and Executing CUDA Programs in Emulation Mode. Retrieved from https:\/\/developer.nvidia.com\/cuda-toolkit."},{"key":"e_1_3_3_3_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jsb.2014.11.009"},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3388333.3388658"},{"key":"e_1_3_3_5_2","unstructured":"AMD. 2021. HIP. Retrieved from https:\/\/github.com\/ROCm-Developer-Tools\/HIP."},{"key":"e_1_3_3_6_2","unstructured":"AMD. 2021. HIP-CPU. Retrieved from https:\/\/github.com\/ROCm-Developer-Tools\/HIP-CPU."},{"key":"e_1_3_3_7_2","unstructured":"AMD. 2021. HIPIFY. Retrieved from https:\/\/github.com\/ROCm-Developer-Tools\/HIPIFY."},{"key":"e_1_3_3_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICEEE52452.2021.9415927"},{"key":"e_1_3_3_9_2","unstructured":"Vera Blomkvist Karlsson. 2021. Cumulus - translating CUDA to sequential C++: Simplifying the process of debugging CUDA programs. Dissertation. KTH ROYAL INSTITUTE OF TECHNOLOGY."},{"key":"e_1_3_3_10_2","unstructured":"Berenger Bramas. 2017. Fast sorting algorithms using AVX-512 on intel knights landing. arXiv preprint arXiv:1704.08579 305 (2017) 315."},{"key":"e_1_3_3_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2009.5306797"},{"key":"e_1_3_3_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3177960"},{"key":"e_1_3_3_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/1854273.1854318"},{"key":"e_1_3_3_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/1854273.1854318"},{"key":"e_1_3_3_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS48437.2020.00020"},{"key":"e_1_3_3_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/1854273.1854302"},{"key":"e_1_3_3_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2011.62"},{"key":"e_1_3_3_18_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cpc.2010.12.052"},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.1145\/1854273.1854303"},{"key":"e_1_3_3_20_2","unstructured":"Intel. 2021. DPCT. Retrieved from https:\/\/software.intel.com\/content\/www\/cn\/zh\/develop\/tools\/oneapi\/components\/dpc-compatibility-tool.html."},{"key":"e_1_3_3_21_2","unstructured":"Intel. 2021. Intel OneAPI Community. Retrieved from https:\/\/community.intel.com\/t5\/Intel-oneAPI-Data-Parallel-C\/How-oneAPI-support-parallel-tasks-on-CPU\/td-p\/1365943."},{"key":"e_1_3_3_22_2","unstructured":"Intel. 2021. Intel-OpenCL. Retrieved from https:\/\/software.intel.com\/content\/www\/us\/en\/develop\/tools\/opencl-sdk.html."},{"key":"e_1_3_3_23_2","unstructured":"Intel. 2021. IntelSDE. Retrieved from https:\/\/www.intel.com\/content\/www\/us\/en\/developer\/articles\/tool\/software-development-emulator.html."},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10766-014-0320-y"},{"key":"e_1_3_3_25_2","article-title":"Performance of SSE and AVX instruction sets","author":"Jeong Hwancheol","year":"2012","unstructured":"Hwancheol Jeong, Sunghoon Kim, Weonjong Lee, and Seok-Ho Myung. 2012. Performance of SSE and AVX instruction sets. In Proceedings of the 30th International Symposium on Lattice Field Theory.","journal-title":"Proceedings of the 30th International Symposium on Lattice Field Theory"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2011.5764682"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-28652-0_1"},{"key":"e_1_3_3_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/2304576.2304623"},{"key":"e_1_3_3_29_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10766-013-0252-y"},{"key":"e_1_3_3_30_2","unstructured":"LLVM. 2013. LLVM Loop Terminology. Retrieved from https:\/\/llvm.org\/docs\/LoopTerminology.html."},{"key":"e_1_3_3_31_2","unstructured":"MacFinder. 2020. 2020GPGPU Roundup. Retrieved from https:\/\/macfinder.co.uk\/blog\/2020-gpgpu-roundup-metal-vs-cuda-vs-opencl-amd-vs-nvidia\/."},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2011.68"},{"key":"e_1_3_3_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/HOTCHIPS.2009.7478342"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/2807591.2807626"},{"key":"e_1_3_3_35_2","volume-title":"International Symposium on Code Generation and Optimization (CGO\u201906)","author":"Nuzman Dorit","year":"2006","unstructured":"Dorit Nuzman and Richard Henderson. 2006. Multi-platform auto-vectorization. In International Symposium on Code Generation and Optimization (CGO\u201906). IEEE."},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/1133255.1133997"},{"key":"e_1_3_3_37_2","first-page":"145","volume-title":"Proceedings of the GCC Developers Summit","author":"Nuzman Dorit","year":"2006","unstructured":"Dorit Nuzman and Ayal Zaks. 2006. Autovectorization in GCC\u2014Two years later. In Proceedings of the GCC Developers Summit. 145\u2013158."},{"key":"e_1_3_3_38_2","unstructured":"NVIDIA. 2011. PGI. Retrieved from https:\/\/developer.nvidia.com\/pgi-cuda-cc-x86."},{"key":"e_1_3_3_39_2","unstructured":"NVIDIA. 2019. CUDA Aligned Barrier. Retrieved from https:\/\/docs.nvidia.com\/cuda\/parallel-thread-execution\/index.html#parallel-synchronization-and-communication-instructions-bar."},{"key":"e_1_3_3_40_2","unstructured":"NVIDIA. 2019. CUDA Warp Shuffle Functions. Retrieved from https:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide\/index.html#warp-shuffle-functions."},{"key":"e_1_3_3_41_2","unstructured":"NVIDIA. 2021. CUB. Retrieved from https:\/\/nvlabs.github.io\/cub\/."},{"key":"e_1_3_3_42_2","unstructured":"NVIDIA. 2021. CUDA-SYNC. Retrieved from https:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide\/index.html#synchronization-functions."},{"key":"e_1_3_3_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458744.3473356"},{"key":"e_1_3_3_44_2","unstructured":"Hugh Perkins. 2016. Cltorch: A hardware-agnostic backend for the torch deep neural network library based on opencl. arXiv preprint arXiv:1606.04884 (2016)."},{"key":"e_1_3_3_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3078155.3078156"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/InPar.2012.6339601"},{"key":"e_1_3_3_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2017.21"},{"key":"e_1_3_3_48_2","volume-title":"GCC Developers Summit","author":"Rosen Ira","year":"2007","unstructured":"Ira Rosen, Dorit Nuzman, and Ayal Zaks. 2007. Loop-aware SLP in GCC. In GCC Developers Summit. Citeseer."},{"key":"e_1_3_3_49_2","article-title":"Supporting CUDA for an extended RISC-V GPU architecture","author":"Tine Jaewon Lee Jaewoong Sim Hyesoon Kim Ruobing Han, Blaise","year":"2021","unstructured":"Jaewon Lee Jaewoong Sim Hyesoon Kim Ruobing Han, Blaise Tine. 2021. Supporting CUDA for an extended RISC-V GPU architecture. In Proceedings of the 5th Workshop on Computer Architecture Research with RISC-V (CARRV\u201921).","journal-title":"Proceedings of the 5th Workshop on Computer Architecture Research with RISC-V (CARRV\u201921)"},{"key":"e_1_3_3_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3293320.3293338"},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3380595"},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/1542275.1542304"},{"key":"e_1_3_3_53_2","volume-title":"Proceedings of the International Conference on Parallel and Distributed Processing Techniques and Applications (PDPTA\u201900)","author":"Silveira Andr\u00e9","year":"2000","unstructured":"Andr\u00e9 Silveira, Rafael Bohrer Avila, Marcos E. Barreto, and Philippe Olivier Alexandre Navaux. 2000. DPC++: Object-oriented programming applied to cluster computing. In Proceedings of the International Conference on Parallel and Distributed Processing Techniques and Applications (PDPTA\u201900)."},{"key":"e_1_3_3_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/1772954.1772971"},{"key":"e_1_3_3_55_2","article-title":"Performance Portability in Accelerated Parallel Kernels","author":"Stratton John A.","year":"2013","unstructured":"John A. Stratton, Hee-Seok Kim, Thoman B. Jablin, and Wen-Mei W. Hwu. 2013. Performance Portability in Accelerated Parallel Kernels. Center for Reliable and High-Performance Computing.","journal-title":"Center for Reliable and High-Performance Computing"},{"key":"e_1_3_3_56_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-89740-8_2"},{"key":"e_1_3_3_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2016.7581262"},{"key":"e_1_3_3_58_2","unstructured":"Yuhsiang M. Tsai Terry Cojean and Hartwig Anzt. 2021. Porting a sparse linear algebra math library to Intel GPUs. arXiv:2103.10116 [cs.DC]."},{"key":"e_1_3_3_59_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-38750-0_11"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3554736","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3554736","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:29Z","timestamp":1750182569000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3554736"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,9,16]]},"references-count":58,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2022,12,31]]}},"alternative-id":["10.1145\/3554736"],"URL":"https:\/\/doi.org\/10.1145\/3554736","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,9,16]]},"assertion":[{"value":"2021-12-18","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-07-21","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-09-16","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}