{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:24:01Z","timestamp":1750307041784,"version":"3.41.0"},"reference-count":23,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2012,6,1]],"date-time":"2012-06-01T00:00:00Z","timestamp":1338508800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000143","name":"Division of Computing and Communication Foundations","doi-asserted-by":"publisher","award":["CCF-0968667"],"award-info":[{"award-number":["CCF-0968667"]}],"id":[{"id":"10.13039\/100000143","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2012,6]]},"abstract":"<jats:p>This article presents a novel optimizing compiler for general purpose computation on graphics processing units (GPGPU). It addresses two major challenges of developing high performance GPGPU programs: effective utilization of GPU memory hierarchy and judicious management of parallelism. The input to our compiler is a na\u00efve GPU kernel function, which is functionally correct but without any consideration for performance optimization. The compiler generates two kernels, one optimized for global memories and the other for texture memories. The proposed compilation process is effective for both AMD\/ATI and NVIDIA GPUs. The experiments show that our optimized code achieves very high performance, either superior or very close to highly fine-tuned libraries.<\/jats:p>","DOI":"10.1145\/2207222.2207225","type":"journal-article","created":{"date-parts":[[2012,6,15]],"date-time":"2012-06-15T15:31:37Z","timestamp":1339774297000},"page":"1-33","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":28,"title":["A unified optimizing compiler framework for different GPGPU architectures"],"prefix":"10.1145","volume":"9","author":[{"given":"Yi","family":"Yang","sequence":"first","affiliation":[{"name":"North Carolina State University, Raleigh, NC"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ping","family":"Xiang","sequence":"additional","affiliation":[{"name":"North Carolina State University, Raleigh, NC"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jingfei","family":"Kong","sequence":"additional","affiliation":[{"name":"Advanced Micro Devices"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mike","family":"Mantor","sequence":"additional","affiliation":[{"name":"Advanced Micro Devices"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Huiyang","family":"Zhou","sequence":"additional","affiliation":[{"name":"North Carolina State University, Raleigh, NC"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2012,6,15]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.5555\/1177220"},{"key":"e_1_2_1_2_1","unstructured":"AMD INC. 2011. AMD Accelerated Parallel Processing OpenCL Programming Guide 2.4.  AMD INC. 2011. AMD Accelerated Parallel Processing OpenCL Programming Guide 2.4."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1693453.1693470"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/1375527.1375562"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1090\/S0025-5718-1965-0178586-1"},{"volume-title":"Proceedings of the IEEE International Parallel & Distributed Processing Symposium (IPDPS'08)","author":"Fujimoto N.","key":"e_1_2_1_6_1","unstructured":"Fujimoto , N. Faster matrix-vector multiplication on GeForce 8800 GTX. 2008 . In Proceedings of the IEEE International Parallel & Distributed Processing Symposium (IPDPS'08) . IEEE, 1--8. Fujimoto, N. Faster matrix-vector multiplication on GeForce 8800 GTX. 2008. In Proceedings of the IEEE International Parallel & Distributed Processing Symposium (IPDPS'08). IEEE, 1--8."},{"volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC'08)","author":"Govindaraju N.","key":"e_1_2_1_7_1","unstructured":"Govindaraju , N. , Lloyd , B. , Dotsenko , Y. , Smith , B. , and Manferdelli , J . 2008. High performance discrete Fourier transforms on graphics processors . In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC'08) . IEEE. 1--12. Govindaraju, N., Lloyd, B., Dotsenko, Y., Smith, B., and Manferdelli, J. 2008. High performance discrete Fourier transforms on graphics processors. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC'08). IEEE. 1--12."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1555754.1555775"},{"volume-title":"Proceedings of Workshops on Languages and Compilers for Parallel Computing (LCPC'03)","author":"Lee S.-I.","key":"e_1_2_1_9_1","unstructured":"Lee , S.-I. , Johnson , T. , and Eigenmann , R . 2003. Cetus\u2014An extensible compiler infrastructure for source-to-source transformation . In Proceedings of Workshops on Languages and Compilers for Parallel Computing (LCPC'03) . 539--553. Lee, S.-I., Johnson, T., and Eigenmann, R. 2003. Cetus\u2014An extensible compiler infrastructure for source-to-source transformation. In Proceedings of Workshops on Languages and Compilers for Parallel Computing (LCPC'03). 539--553."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1504176.1504194"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2009.5160988"},{"key":"e_1_2_1_12_1","unstructured":"Nath R. Tomov S. and Dongarra J. 2010. An improved MAGMA GEMM for Fermi GPUs. Tech. rep. UT-CS-10-655. University of Tennessee Computer Science.  Nath R. Tomov S. and Dongarra J. 2010. An improved MAGMA GEMM for Fermi GPUs. Tech. rep. UT-CS-10-655. University of Tennessee Computer Science."},{"key":"e_1_2_1_13_1","unstructured":"NVIDIA Inc. 2010. NVIDIA CUDA C Programming Guide 3.2.  NVIDIA Inc. 2010. NVIDIA CUDA C Programming Guide 3.2."},{"key":"e_1_2_1_14_1","unstructured":"OpenCL. http:\/\/www.khronos.org\/opencl\/.  OpenCL. http:\/\/www.khronos.org\/opencl\/."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CGO.2007.21"},{"key":"e_1_2_1_16_1","unstructured":"Ruetsch G. and Micikevicius P. 2009. Optimize matrix transpose in CUDA. http:\/\/developer.download.nvidia.com\/compute\/cuda\/sdk\/website\/C\/src\/transpose\/doc\/MatrixTranspose.pdf.  Ruetsch G. and Micikevicius P. 2009. Optimize matrix transpose in CUDA. http:\/\/developer.download.nvidia.com\/compute\/cuda\/sdk\/website\/C\/src\/transpose\/doc\/MatrixTranspose.pdf."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1345206.1345220"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/1356058.1356084"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-89740-8_2"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-89740-8_1"},{"volume-title":"Proceedings of the International Conference for High Performance Computing (SC'08)","author":"Volkov V.","key":"e_1_2_1_21_1","unstructured":"Volkov , V. and Demmel , J. W . Benchmarking GPUs to tune dense linear algebra. 2008 . In Proceedings of the International Conference for High Performance Computing (SC'08) , ACM. 1--11. Volkov, V. and Demmel, J. W. Benchmarking GPUs to tune dense linear algebra. 2008. In Proceedings of the International Conference for High Performance Computing (SC'08), ACM. 1--11."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1806596.1806606"},{"key":"e_1_2_1_23_1","unstructured":"Yang Y. and Zhou H. 2010. GPGPU compiler. http:\/\/code.google.com\/p\/gpgpucompiler\/.  Yang Y. and Zhou H. 2010. GPGPU compiler. http:\/\/code.google.com\/p\/gpgpucompiler\/."}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2207222.2207225","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2207222.2207225","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T09:20:55Z","timestamp":1750238455000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2207222.2207225"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,6]]},"references-count":23,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2012,6]]}},"alternative-id":["10.1145\/2207222.2207225"],"URL":"https:\/\/doi.org\/10.1145\/2207222.2207225","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"type":"print","value":"1544-3566"},{"type":"electronic","value":"1544-3973"}],"subject":[],"published":{"date-parts":[[2012,6]]},"assertion":[{"value":"2011-03-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2011-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2012-06-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}