{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T23:41:10Z","timestamp":1783035670829,"version":"3.54.6"},"publisher-location":"New York, NY, USA","reference-count":44,"publisher":"ACM","license":[{"start":{"date-parts":[[2019,11,17]],"date-time":"2019-11-17T00:00:00Z","timestamp":1573948800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000015","name":"U.S. Department of Energy","doi-asserted-by":"publisher","award":["DE-AC02-05CH11231"],"award-info":[{"award-number":["DE-AC02-05CH11231"]}],"id":[{"id":"10.13039\/100000015","id-type":"DOI","asserted-by":"publisher"}]},{"name":"U.S. Department of Energy and National Nuclear Security Administration","award":["17-SC-20-SC"],"award-info":[{"award-number":["17-SC-20-SC"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2019,11,17]]},"DOI":"10.1145\/3295500.3356210","type":"proceedings-article","created":{"date-parts":[[2019,11,7]],"date-time":"2019-11-07T19:43:22Z","timestamp":1573155802000},"page":"1-44","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":37,"title":["Exploiting reuse and vectorization in blocked stencil computations on CPUs and GPUs"],"prefix":"10.1145","author":[{"given":"Tuowen","family":"Zhao","sequence":"first","affiliation":[{"name":"University of Utah"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Protonu","family":"Basu","sequence":"additional","affiliation":[{"name":"Facebook"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Samuel","family":"Williams","sequence":"additional","affiliation":[{"name":"Lawrence Berkeley National Lab"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mary","family":"Hall","sequence":"additional","affiliation":[{"name":"University of Utah"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hans","family":"Johansen","sequence":"additional","affiliation":[{"name":"Lawrence Berkeley National Lab"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,11,17]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"2016. High-Performance Geometric Multigrid. http:\/\/hpgmg.org  2016. High-Performance Geometric Multigrid. http:\/\/hpgmg.org"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1155\/2009\/382638"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2015.103"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2011.70"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1137\/070693199"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"crossref","unstructured":"Kaushik Datta Mark Murphy Vasily Volkov Samuel Williams Jonathan Carter Leonid Oliker David Patterson John Shalf and Katherine Yelick. 2008. Stencil Computation Optimization and Auto-Tuning on State-of-the-art Multicore Architectures. In Supercomputing (SC).  Kaushik Datta Mark Murphy Vasily Volkov Samuel Williams Jonathan Carter Leonid Oliker David Patterson John Shalf and Katherine Yelick. 2008. Stencil Computation Optimization and Auto-Tuning on State-of-the-art Multicore Architectures. In Supercomputing (SC).","DOI":"10.1109\/SC.2008.5222004"},{"key":"e_1_3_2_1_8_1","volume-title":"Introducing the Semi-stencil Algorithm. In International Conference on Parallel Processing and Applied Mathematics: Part I (PPAM). 11","author":"La Cruz Ra\u00fal De","year":"2010","unstructured":"Ra\u00fal De La Cruz , Mauricio Araya-Polo , and Jos\u00e9 Mar\u00eda Cela . 2010 . Introducing the Semi-stencil Algorithm. In International Conference on Parallel Processing and Applied Mathematics: Part I (PPAM). 11 . Ra\u00fal De La Cruz, Mauricio Araya-Polo, and Jos\u00e9 Mar\u00eda Cela. 2010. Introducing the Semi-stencil Algorithm. In International Conference on Parallel Processing and Applied Mathematics: Part I (PPAM). 11."},{"key":"e_1_3_2_1_9_1","volume-title":"International Conference on High Performance Computing. Springer, 489--507","author":"Deakin Tom","year":"2016","unstructured":"Tom Deakin , James Price , Matt Martineau , and Simon McIntosh-Smith . 2016 . GPU-STREAM v2. 0: benchmarking the achievable memory bandwidth of many-core processors across diverse parallel programming models . In International Conference on High Performance Computing. Springer, 489--507 . Tom Deakin, James Price, Matt Martineau, and Simon McIntosh-Smith. 2016. GPU-STREAM v2. 0: benchmarking the achievable memory bandwidth of many-core processors across diverse parallel programming models. In International Conference on High Performance Computing. Springer, 489--507."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/377792.377807"},{"key":"e_1_3_2_1_11_1","first-page":"21","article-title":"Cache Optimization for Structured and Unstructured Grid Multigrid","volume":"10","author":"Douglas Craig C.","year":"2000","unstructured":"Craig C. Douglas , Jonathan Hu , Markus Kowarschik , Ulrich R\u00fcde , and Christian Weiss . 2000 . Cache Optimization for Structured and Unstructured Grid Multigrid . Elect. Trans. Numer. Anal 10 (2000), 21 -- 40 . Craig C. Douglas, Jonathan Hu, Markus Kowarschik, Ulrich R\u00fcde, and Christian Weiss. 2000. Cache Optimization for Structured and Unstructured Grid Multigrid. Elect. Trans. Numer. Anal 10 (2000), 21--40.","journal-title":"Elect. Trans. Numer. Anal"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1080\/13647830.2014.919410"},{"key":"e_1_3_2_1_13_1","volume-title":"Proc. ACM International Conference on Supercomputing (ICS).","author":"Frigo M.","unstructured":"M. Frigo and V. Strumpen . 2005. Evaluation of cache-based superscalar and cache-less vector architectures for scientific computations . In Proc. ACM International Conference on Supercomputing (ICS). M. Frigo and V. Strumpen. 2005. Evaluation of cache-based superscalar and cache-less vector architectures for scientific computations. In Proc. ACM International Conference on Supercomputing (ICS)."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1002\/nla.1808"},{"key":"e_1_3_2_1_15_1","volume-title":"Compiler Construction","author":"Henretty Tom","unstructured":"Tom Henretty , Kevin Stock , Louis-No\u00ebl Pouchet , Franz Franchetti , J Ramanujam , and P Sadayappan . 2011. Data layout transformation for stencil computations on short-vector simd architectures . In Compiler Construction . Springer , 225--245. Tom Henretty, Kevin Stock, Louis-No\u00ebl Pouchet, Franz Franchetti, J Ramanujam, and P Sadayappan. 2011. Data layout transformation for stencil computations on short-vector simd architectures. In Compiler Construction. Springer, 225--245."},{"key":"e_1_3_2_1_16_1","volume-title":"International Conference on Supercomputing (ICS).","author":"Holewinski Justin","unstructured":"Justin Holewinski , Louis-No\u00ebl Pouchet , and P. Sadayappan . 2012. High-performance code generation for stencil computations on GPU architectures . In International Conference on Supercomputing (ICS). Justin Holewinski, Louis-No\u00ebl Pouchet, and P. Sadayappan. 2012. High-performance code generation for stencil computations on GPU architectures. In International Conference on Supercomputing (ICS)."},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW.2014.117"},{"key":"e_1_3_2_1_19_1","volume-title":"DiMEPACK - A Cache-Optimized Multigrid Library. In International Conference on Parallel and Distributed Processing Techniques and Applications (PDPTA)","author":"Kowarschik Markus","year":"2001","unstructured":"Markus Kowarschik and Christian WeiSS. 2001 . DiMEPACK - A Cache-Optimized Multigrid Library. In International Conference on Parallel and Distributed Processing Techniques and Applications (PDPTA) , volume I . Markus Kowarschik and Christian WeiSS. 2001. DiMEPACK - A Cache-Optimized Multigrid Library. In International Conference on Parallel and Distributed Processing Techniques and Applications (PDPTA), volume I."},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/1250734.1250761"},{"key":"e_1_3_2_1_21_1","volume-title":"Technical Report DCS-TR-379. Department of Computer Science","author":"McCalpin J.","year":"1999","unstructured":"J. McCalpin and D. Wonnacott . 1999 . Time skewing: A value-based approach to optimizing for memory locality. Technical Report DCS-TR-379. Department of Computer Science , Rutgers University. J. McCalpin and D. Wonnacott. 1999. Time skewing: A value-based approach to optimizing for memory locality. Technical Report DCS-TR-379. Department of Computer Science, Rutgers University."},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1513895.1513905"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2010.2"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3178487.3178500"},{"key":"e_1_3_2_1_25_1","volume-title":"Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis (SC '18)","author":"Rawat Prashant Singh","unstructured":"Prashant Singh Rawat , Aravind Sukumaran-Rajam , Atanas Rountev , Fabrice Rastello , Louis-No\u00ebl Pouchet , and P. Sadayappan . 2018. Associative Instruction Reordering to Alleviate Register Pressure . In Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis (SC '18) . IEEE Press, Piscataway, NJ, USA, Article 46, 13 pages. http:\/\/dl.acm.org\/citation.cfm?id=3291656.3291718 Prashant Singh Rawat, Aravind Sukumaran-Rajam, Atanas Rountev, Fabrice Rastello, Louis-No\u00ebl Pouchet, and P. Sadayappan. 2018. Associative Instruction Reordering to Alleviate Register Pressure. In Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis (SC '18). IEEE Press, Piscataway, NJ, USA, Article 46, 13 pages. http:\/\/dl.acm.org\/citation.cfm?id=3291656.3291718"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"crossref","unstructured":"G. Rivera and C. Tseng. 2000. Tiling Optimizations for 3D Scientific Computations. In Supercomputing (SC).  G. Rivera and C. Tseng. 2000. Tiling Optimizations for 3D Scientific Computations. In Supercomputing (SC).","DOI":"10.1109\/SC.2000.10015"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1177\/1094342004041295"},{"key":"e_1_3_2_1_28_1","volume-title":"Proc. ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI).","author":"Song Y.","unstructured":"Y. Song and Z. Li . 1999. New tiling techniques to improve cache temporal locality . In Proc. ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI). Y. Song and Z. Li. 1999. New tiling techniques to improve cache temporal locality. In Proc. ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI)."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2666356.2594342"},{"key":"e_1_3_2_1_30_1","volume-title":"ACM symposium on Parallelism in algorithms and architectures.","author":"Tang Yuan","unstructured":"Yuan Tang , Rezaul Alam Chowdhury , Bradley C. Kuszmaul , Chi-Keung Luk , and Charles E. Leiserson . 2011. The pochoir stencil compiler . In ACM symposium on Parallelism in algorithms and architectures. Yuan Tang, Rezaul Alam Chowdhury, Bradley C. Kuszmaul, Chi-Keung Luk, and Charles E. Leiserson. 2011. The pochoir stencil compiler. In ACM symposium on Parallelism in algorithms and architectures."},{"key":"e_1_3_2_1_31_1","volume-title":"Burak Bastem, George Michelogiannakis, Ann Almgren, and John Shalf.","author":"Unat Didem","year":"2016","unstructured":"Didem Unat , Tan Nguyen , Weiqun Zhang , Muhammed Nufail Farooqi , Burak Bastem, George Michelogiannakis, Ann Almgren, and John Shalf. 2016 . TiDA: High- Level Programming Abstractions for Data Locality Management. Springer International Publishing , Cham, 116--135. Didem Unat, Tan Nguyen, Weiqun Zhang, Muhammed Nufail Farooqi, Burak Bastem, George Michelogiannakis, Ann Almgren, and John Shalf. 2016. TiDA: High-Level Programming Abstractions for Data Locality Management. Springer International Publishing, Cham, 116--135."},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.2514\/6.2016-3965"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/COMPSAC.2009.82"},{"key":"e_1_3_2_1_34_1","volume-title":"Lattice Boltzmann Simulation Optimization on Leading Multicore Platforms. In Interational Conference on Parallel and Distributed Computing Systems (IPDPS).","author":"Williams S.","unstructured":"S. Williams , J. Carter , L. Oliker , J. Shalf , and K. Yelick . 2008 . Lattice Boltzmann Simulation Optimization on Leading Multicore Platforms. In Interational Conference on Parallel and Distributed Computing Systems (IPDPS). S. Williams, J. Carter, L. Oliker, J. Shalf, and K. Yelick. 2008. Lattice Boltzmann Simulation Optimization on Leading Multicore Platforms. In Interational Conference on Parallel and Distributed Computing Systems (IPDPS)."},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"crossref","unstructured":"Samuel Williams Leonid Oliker Jonathan Carter and John Shalf. 2011. Extracting ultra-scale Lattice Boltzmann performance via hierarchical and distributed auto-tuning. In Supercomputing (SC).  Samuel Williams Leonid Oliker Jonathan Carter and John Shalf. 2011. Extracting ultra-scale Lattice Boltzmann performance via hierarchical and distributed auto-tuning. In Supercomputing (SC).","DOI":"10.1145\/2063384.2063458"},{"key":"e_1_3_2_1_36_1","volume-title":"Proc. Conference on Computing Frontiers.","author":"Williams S.","unstructured":"S. Williams , J. Shalf , L. Oliker , S. Kamil , P. Husbands , and K. Yelick . 2006. The potential of the Cell processor for scientific computing . In Proc. Conference on Computing Frontiers. S. Williams, J. Shalf, L. Oliker, S. Kamil, P. Husbands, and K. Yelick. 2006. The potential of the Cell processor for scientific computing. In Proc. Conference on Computing Frontiers."},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/1498765.1498785"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2000.845979"},{"key":"e_1_3_2_1_39_1","volume-title":"Vector Folding: Improving Stencil Performance via Multidimensional SIMD-vector Representation. In 2015 IEEE 17th International Conference on High Performance Computing and Communications","author":"Yount C.","year":"2015","unstructured":"C. Yount . 2015. Vector Folding: Improving Stencil Performance via Multidimensional SIMD-vector Representation. In 2015 IEEE 17th International Conference on High Performance Computing and Communications , 2015 IEEE 7th International Symposium on Cyberspace Safety and Security, and 2015 IEEE 12th International Conference on Embedded Software and Systems . 865--870. C. Yount. 2015. Vector Folding: Improving Stencil Performance via Multidimensional SIMD-vector Representation. In 2015 IEEE 17th International Conference on High Performance Computing and Communications, 2015 IEEE 7th International Symposium on Cyberspace Safety and Security, and 2015 IEEE 12th International Conference on Embedded Software and Systems. 865--870."},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/WOLFHPC.2016.08"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"crossref","unstructured":"T. Zeiser G. Wellein A. Nitsure K. Iglberger U. Rude and G. Hager. 2008. Introducing a parallel cache oblivious blocking approach for the lattice Boltzmann method. Progress in Computational Fluid Dynamics 8 (2008).  T. Zeiser G. Wellein A. Nitsure K. Iglberger U. Rude and G. Hager. 2008. Introducing a parallel cache oblivious blocking approach for the lattice Boltzmann method. Progress in Computational Fluid Dynamics 8 (2008).","DOI":"10.1504\/PCFD.2008.018088"},{"key":"e_1_3_2_1_42_1","volume-title":"Snowflake: A Lightweight Portable Stencil DSL. In 2017IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). 795--804","author":"Zhang N.","unstructured":"N. Zhang , M. Driscoll , C. Markley , S. Williams , P. Basu , and A. Fox . 2017 . Snowflake: A Lightweight Portable Stencil DSL. In 2017IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). 795--804 . N. Zhang, M. Driscoll, C. Markley, S. Williams, P. Basu, and A. Fox. 2017. Snowflake: A Lightweight Portable Stencil DSL. In 2017IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). 795--804."},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/2259016.2259037"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/P3HPC.2018.00009"},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/2259016.2259044"},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-36803-5_25"}],"event":{"name":"SC '19: The International Conference for High Performance Computing, Networking, Storage, and Analysis","location":"Denver Colorado","acronym":"SC '19","sponsor":["SIGHPC ACM Special Interest Group on High Performance Computing, Special Interest Group on High Performance Computing","IEEE CS"]},"container-title":["Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3295500.3356210","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3295500.3356210","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3295500.3356210","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:12:49Z","timestamp":1750201969000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3295500.3356210"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,11,17]]},"references-count":44,"alternative-id":["10.1145\/3295500.3356210","10.1145\/3295500"],"URL":"https:\/\/doi.org\/10.1145\/3295500.3356210","relation":{},"subject":[],"published":{"date-parts":[[2019,11,17]]},"assertion":[{"value":"2019-11-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}