{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:10:55Z","timestamp":1750219855587,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":42,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,2,25]],"date-time":"2023-02-25T00:00:00Z","timestamp":1677283200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["1833332, 2015254"],"award-info":[{"award-number":["1833332, 2015254"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,2,25]]},"DOI":"10.1145\/3582514.3582523","type":"proceedings-article","created":{"date-parts":[[2023,2,24]],"date-time":"2023-02-24T17:27:41Z","timestamp":1677259661000},"page":"60-69","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Exploring OpenMP GPU Offloading for Implementing Convolutional Neural Networks"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0406-1276","authenticated-orcid":false,"given":"Kewei","family":"Yan","sequence":"first","affiliation":[{"name":"University of North Carolina at Charlotte, Charlotte, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8721-8509","authenticated-orcid":false,"given":"Yaying","family":"Shi","sequence":"additional","affiliation":[{"name":"University of North Carolina at Charlotte, Charlotte, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5274-8526","authenticated-orcid":false,"given":"Yonghong","family":"Yan","sequence":"additional","affiliation":[{"name":"University of North Carolina at Charlotte, Charlotte, United States of America"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,2,25]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"12th USENIX symposium on operating systems design and implementation (OSDI 16)","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi , Paul Barham , Jianmin Chen , Zhifeng Chen , Andy Davis , Jeffrey Dean , Matthieu Devin , Sanjay Ghemawat , Geoffrey Irving , Michael Isard , 2016 . {TensorFlow}: A System for {Large-Scale} Machine Learning . In 12th USENIX symposium on operating systems design and implementation (OSDI 16) . 265--283. Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. 2016. {TensorFlow}: A System for {Large-Scale} Machine Learning. In 12th USENIX symposium on operating systems design and implementation (OSDI 16). 265--283."},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0731-7085(99)00272-1"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICEngTechnol.2017.8308186"},{"key":"e_1_3_2_1_4_1","volume-title":"Low-memory gemm-based convolution algorithms for deep neural networks. arXiv preprint arXiv:1709.03395","author":"Anderson Andrew","year":"2017","unstructured":"Andrew Anderson , Aravind Vasudevan , Cormac Keane , and David Gregg . 2017. Low-memory gemm-based convolution algorithms for deep neural networks. arXiv preprint arXiv:1709.03395 ( 2017 ). Andrew Anderson, Aravind Vasudevan, Cormac Keane, and David Gregg. 2017. Low-memory gemm-based convolution algorithms for deep neural networks. arXiv preprint arXiv:1709.03395 (2017)."},{"key":"e_1_3_2_1_5_1","volume-title":"2016 Third Workshop on the LLVM Compiler Infrastructure in HPC (LLVM-HPC). IEEE.","author":"Antao Samuel F","year":"2016","unstructured":"Samuel F Antao , Alexey Bataev , Arpith C Jacob , Gheorghe-Teodor Bercea , Alexandre E Eichenberger , Georgios Rokos , Matt Martineau , Tian Jin , Guray Ozen , Zehra Sura , 2016 . Offloading support for OpenMP in Clang and LLVM . In 2016 Third Workshop on the LLVM Compiler Infrastructure in HPC (LLVM-HPC). IEEE. Samuel F Antao, Alexey Bataev, Arpith C Jacob, Gheorghe-Teodor Bercea, Alexandre E Eichenberger, Georgios Rokos, Matt Martineau, Tian Jin, Guray Ozen, Zehra Sura, et al. 2016. Offloading support for OpenMP in Clang and LLVM. In 2016 Third Workshop on the LLVM Compiler Infrastructure in HPC (LLVM-HPC). IEEE."},{"key":"e_1_3_2_1_6_1","volume-title":"Raja: Portable performance for large-scale scientific applications. In 2019 ieee\/acm international workshop on performance, portability and productivity in hpc (p3hpc)","author":"Beckingsale David A","year":"2019","unstructured":"David A Beckingsale , Jason Burmark , Rich Hornung , Holger Jones , William Killian , Adam J Kunen , Olga Pearce , Peter Robinson , Brian S Ryujin , and Thomas RW Scogland . 2019 . Raja: Portable performance for large-scale scientific applications. In 2019 ieee\/acm international workshop on performance, portability and productivity in hpc (p3hpc) . IEEE , 71--81. David A Beckingsale, Jason Burmark, Rich Hornung, Holger Jones, William Killian, Adam J Kunen, Olga Pearce, Peter Robinson, Brian S Ryujin, and Thomas RW Scogland. 2019. Raja: Portable performance for large-scale scientific applications. In 2019 ieee\/acm international workshop on performance, portability and productivity in hpc (p3hpc). IEEE, 71--81."},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2833157.2833161"},{"key":"e_1_3_2_1_8_1","volume-title":"Arpith C Jacob, Tong Chen, and Olivier Sallenave.","author":"Bertolli Carlo","year":"2014","unstructured":"Carlo Bertolli , Samuel F Antao , Alexandre E Eichenberger , Kevin OBrien Zehra Sura , Arpith C Jacob, Tong Chen, and Olivier Sallenave. 2014 . Coordinating GPU threads for OpenMP 4.0 in LLVM. In 2014 LLVM Compiler Infrastructure in HPC. IEEE , 12--21. Carlo Bertolli, Samuel F Antao, Alexandre E Eichenberger, Kevin OBrien Zehra Sura, Arpith C Jacob, Tong Chen, and Olivier Sallenave. 2014. Coordinating GPU threads for OpenMP 4.0 in LLVM. In 2014 LLVM Compiler Infrastructure in HPC. IEEE, 12--21."},{"key":"e_1_3_2_1_9_1","volume-title":"International Workshop on OpenMP. Springer, 108--121","author":"Beyer James C","year":"2011","unstructured":"James C Beyer , Eric J Stotzer , Alistair Hart , and Bronis R de Supinski . 2011 . OpenMP for accelerators . In International Workshop on OpenMP. Springer, 108--121 . James C Beyer, Eric J Stotzer, Alistair Hart, and Bronis R de Supinski. 2011. OpenMP for accelerators. In International Workshop on OpenMP. Springer, 108--121."},{"key":"e_1_3_2_1_10_1","volume-title":"International Workshop on OpenMP. Springer.","author":"Chapman Barbara","year":"2021","unstructured":"Barbara Chapman , Buu Pham , Charlene Yang , Christopher Daley , Colleen Bertoni , Dhruva Kulkarni , Dossay Oryspayev , Ed D'Azevedo , Johannes Doerfert , Keren Zhou , 2021 . Outcomes of OpenMP Hackathon: OpenMP Application Experiences with the Offloading Model (Part II) . In International Workshop on OpenMP. Springer. Barbara Chapman, Buu Pham, Charlene Yang, Christopher Daley, Colleen Bertoni, Dhruva Kulkarni, Dossay Oryspayev, Ed D'Azevedo, Johannes Doerfert, Keren Zhou, et al. 2021. Outcomes of OpenMP Hackathon: OpenMP Application Experiences with the Offloading Model (Part II). In International Workshop on OpenMP. Springer."},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2009.5306797"},{"key":"e_1_3_2_1_12_1","volume-title":"cudnn: Efficient primitives for deep learning. arXiv preprint arXiv:1410.0759","author":"Chetlur Sharan","year":"2014","unstructured":"Sharan Chetlur , Cliff Woolley , Philippe Vandermersch , Jonathan Cohen , John Tran , Bryan Catanzaro , and Evan Shelhamer . 2014. cudnn: Efficient primitives for deep learning. arXiv preprint arXiv:1410.0759 ( 2014 ). Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. 2014. cudnn: Efficient primitives for deep learning. arXiv preprint arXiv:1410.0759 (2014)."},{"key":"e_1_3_2_1_13_1","volume-title":"International Workshop on Accelerator Programming Using Directives. Springer, 25--44","author":"Davis Joshua Hoke","year":"2021","unstructured":"Joshua Hoke Davis , Christopher Daley , Swaroop Pophale , Thomas Huber , Sunita Chandrasekaran , and Nicholas J Wright . 2021 . Performance assessment of OpenMP compilers targeting NVIDIA V100 GPUs . In International Workshop on Accelerator Programming Using Directives. Springer, 25--44 . Joshua Hoke Davis, Christopher Daley, Swaroop Pophale, Thomas Huber, Sunita Chandrasekaran, and Nicholas J Wright. 2021. Performance assessment of OpenMP compilers targeting NVIDIA V100 GPUs. In International Workshop on Accelerator Programming Using Directives. Springer, 25--44."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2012.2211477"},{"key":"e_1_3_2_1_16_1","volume-title":"Proceedings of the 19th international conference on Parallel architectures and compilation techniques. 353--364","author":"Diamos Gregory Frederick","year":"2010","unstructured":"Gregory Frederick Diamos , Andrew Robert Kerr , Sudhakar Yalamanchili , and Nathan Clark . 2010 . Ocelot: a dynamic optimization framework for bulk-synchronous applications in heterogeneous systems . In Proceedings of the 19th international conference on Parallel architectures and compilation techniques. 353--364 . Gregory Frederick Diamos, Andrew Robert Kerr, Sudhakar Yalamanchili, and Nathan Clark. 2010. Ocelot: a dynamic optimization framework for bulk-synchronous applications in heterogeneous systems. In Proceedings of the 19th international conference on Parallel architectures and compilation techniques. 353--364."},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"crossref","first-page":"102546","DOI":"10.1016\/j.parco.2019.102546","article-title":"Analysis of OpenMP 4.5 offloading in implementations: correctness and overhead","volume":"89","author":"Diaz Jose Monsalve","year":"2019","unstructured":"Jose Monsalve Diaz , Kyle Friedline , Swaroop Pophale , Oscar Hernandez , David E Bernholdt , and Sunita Chandrasekaran . 2019 . Analysis of OpenMP 4.5 offloading in implementations: correctness and overhead . Parallel Comput. 89 (2019), 102546 . Jose Monsalve Diaz, Kyle Friedline, Swaroop Pophale, Oscar Hernandez, David E Bernholdt, and Sunita Chandrasekaran. 2019. Analysis of OpenMP 4.5 offloading in implementations: correctness and overhead. Parallel Comput. 89 (2019), 102546.","journal-title":"Parallel Comput."},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3229710.3229717"},{"key":"e_1_3_2_1_19_1","volume-title":"International Workshop on OpenMP. Springer, 153--167","author":"Doerfert Johannes","year":"2019","unstructured":"Johannes Doerfert , Jose Manuel Monsalve Diaz , and Hal Finkel . 2019 . The TRegion interface and compiler optimizations for OpenMP target regions . In International Workshop on OpenMP. Springer, 153--167 . Johannes Doerfert, Jose Manuel Monsalve Diaz, and Hal Finkel. 2019. The TRegion interface and compiler optimizations for OpenMP target regions. In International Workshop on OpenMP. Springer, 153--167."},{"key":"e_1_3_2_1_20_1","volume-title":"International Workshop on Languages and Compilers for Parallel Computing. Springer, 112--119","author":"Doerfert Johannes","year":"2018","unstructured":"Johannes Doerfert and Hal Finkel . 2018 . Compiler optimizations for parallel programs . In International Workshop on Languages and Compilers for Parallel Computing. Springer, 112--119 . Johannes Doerfert and Hal Finkel. 2018. Compiler optimizations for parallel programs. In International Workshop on Languages and Compilers for Parallel Computing. Springer, 112--119."},{"key":"e_1_3_2_1_21_1","volume-title":"2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, 504--514","author":"Doerfert Johannes","year":"2022","unstructured":"Johannes Doerfert , Atemn Patel , Joseph Huber , Shilei Tian , Jose M Monsalve Diaz , Barbara Chapman , and Giorgis Georgakoudis . 2022 . Co-Designing an OpenMP GPU runtime and optimizations for near-zero overhead execution . In 2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, 504--514 . Johannes Doerfert, Atemn Patel, Joseph Huber, Shilei Tian, Jose M Monsalve Diaz, Barbara Chapman, and Giorgis Georgakoudis. 2022. Co-Designing an OpenMP GPU runtime and optimizations for near-zero overhead execution. In 2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, 504--514."},{"key":"e_1_3_2_1_22_1","volume-title":"Neocognitron: A self-organizing neural network model for a mechanism of visual pattern recognition. In Competition and cooperation in neural nets","author":"Fukushima Kunihiko","year":"1982","unstructured":"Kunihiko Fukushima and Sei Miyake . 1982 . Neocognitron: A self-organizing neural network model for a mechanism of visual pattern recognition. In Competition and cooperation in neural nets . Springer . Kunihiko Fukushima and Sei Miyake. 1982. Neocognitron: A self-organizing neural network model for a mechanism of visual pattern recognition. In Competition and cooperation in neural nets. Springer."},{"key":"e_1_3_2_1_23_1","volume-title":"International Workshop on Accelerator Programming Using Directives. Springer, 75--95","author":"Gayatri Rahulkumar","year":"2019","unstructured":"Rahulkumar Gayatri , Charlene Yang , Thorsten Kurth , and Jack Deslippe . 2019 . A case study for performance portability using OpenMP 4.5 . In International Workshop on Accelerator Programming Using Directives. Springer, 75--95 . Rahulkumar Gayatri, Charlene Yang, Thorsten Kurth, and Jack Deslippe. 2019. A case study for performance portability using OpenMP 4.5. In International Workshop on Accelerator Programming Using Directives. Springer, 75--95."},{"key":"e_1_3_2_1_24_1","volume-title":"2022 IEEE\/ACM International Symposium on Code Generation and Optimization (CGO). IEEE, 41--52","author":"Huber Joseph","year":"2022","unstructured":"Joseph Huber , Melanie Cornelius , Giorgis Georgakoudis , Shilei Tian , Jose M Monsalve Diaz , Kuter Dinel , Barbara Chapman , and Johannes Doerfert . 2022 . Efficient execution of OpenMP on GPUs . In 2022 IEEE\/ACM International Symposium on Code Generation and Optimization (CGO). IEEE, 41--52 . Joseph Huber, Melanie Cornelius, Giorgis Georgakoudis, Shilei Tian, Jose M Monsalve Diaz, Kuter Dinel, Barbara Chapman, and Johannes Doerfert. 2022. Efficient execution of OpenMP on GPUs. In 2022 IEEE\/ACM International Symposium on Code Generation and Optimization (CGO). IEEE, 41--52."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654889"},{"key":"e_1_3_2_1_26_1","first-page":"1","article-title":"Beyond Data and Model Parallelism for Deep Neural Networks","volume":"1","author":"Jia Zhihao","year":"2019","unstructured":"Zhihao Jia , Matei Zaharia , and Alex Aiken . 2019 . Beyond Data and Model Parallelism for Deep Neural Networks . Proceedings of Machine Learning and Systems 1 (2019), 1 -- 13 . Zhihao Jia, Matei Zaharia, and Alex Aiken. 2019. Beyond Data and Model Parallelism for Deep Neural Networks. Proceedings of Machine Learning and Systems 1 (2019), 1--13.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_3_2_1_27_1","volume-title":"International Workshop on OpenMP. Springer, 281--292","author":"Karlin Ian","year":"2016","unstructured":"Ian Karlin , Tom Scogland , Arpith C Jacob , Samuel F Antao , Gheorghe-Teodor Bercea , Carlo Bertolli , Bronis R de Supinski , Erik W Draeger , Alexandre E Eichenberger , Jim Glosli , 2016 . Early experiences porting three applications to OpenMP 4.5 . In International Workshop on OpenMP. Springer, 281--292 . Ian Karlin, Tom Scogland, Arpith C Jacob, Samuel F Antao, Gheorghe-Teodor Bercea, Carlo Bertolli, Bronis R de Supinski, Erik W Draeger, Alexandre E Eichenberger, Jim Glosli, et al. 2016. Early experiences porting three applications to OpenMP 4.5. In International Workshop on OpenMP. Springer, 281--292."},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/2.214439"},{"key":"e_1_3_2_1_29_1","volume-title":"SC20: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 1--14","author":"Lambert Jacob","year":"2020","unstructured":"Jacob Lambert , Seyong Lee , Jeffrey S Vetter , and Allen D Malony . 2020 . CCAMP: an integrated translation and optimization framework for OpenACC and OpenMP . In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 1--14 . Jacob Lambert, Seyong Lee, Jeffrey S Vetter, and Allen D Malony. 2020. CCAMP: an integrated translation and optimization framework for OpenACC and OpenMP. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 1--14."},{"key":"e_1_3_2_1_30_1","volume-title":"2008 5th IEEE international symposium on biomedical imaging: from nano to macro. IEEE, 836--838","author":"Luebke David","year":"2008","unstructured":"David Luebke . 2008 . CUDA: Scalable parallel programming for high-performance scientific computing . In 2008 5th IEEE international symposium on biomedical imaging: from nano to macro. IEEE, 836--838 . David Luebke. 2008. CUDA: Scalable parallel programming for high-performance scientific computing. In 2008 5th IEEE international symposium on biomedical imaging: from nano to macro. IEEE, 836--838."},{"key":"e_1_3_2_1_31_1","volume-title":"International Workshop on OpenMP. Springer, 185--200","author":"Martineau Matt","year":"2017","unstructured":"Matt Martineau and Simon McIntosh-Smith . 2017 . The productivity, portability and performance of OpenMP 4.5 for scientific applications targeting Intel CPUs, IBM CPUs, and NVIDIA GPUs . In International Workshop on OpenMP. Springer, 185--200 . Matt Martineau and Simon McIntosh-Smith. 2017. The productivity, portability and performance of OpenMP 4.5 for scientific applications targeting Intel CPUs, IBM CPUs, and NVIDIA GPUs. In International Workshop on OpenMP. Springer, 185--200."},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3148173.3148184"},{"key":"e_1_3_2_1_33_1","volume-title":"Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke , Sam Gross , Francisco Massa , Adam Lerer , James Bradbury , Gregory Chanan , Trevor Killeen , Zeming Lin , Natalia Gimelshein , Luca Antiga , 2019 . Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32 (2019). Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems 32 (2019)."},{"key":"e_1_3_2_1_34_1","volume-title":"2018 IEEE\/ACM International Workshop on Performance, Portability and Productivity in HPC (P3HPC). IEEE.","author":"Pennycook Simon J","year":"2018","unstructured":"Simon J Pennycook , Jason D Sewall , and Jeff R Hammond . 2018 . Evaluating the impact of proposed openmp 5.0 features on performance, portability and productivity . In 2018 IEEE\/ACM International Workshop on Performance, Portability and Productivity in HPC (P3HPC). IEEE. Simon J Pennycook, Jason D Sewall, and Jeff R Hammond. 2018. Evaluating the impact of proposed openmp 5.0 features on performance, portability and productivity. In 2018 IEEE\/ACM International Workshop on Performance, Portability and Productivity in HPC (P3HPC). IEEE."},{"key":"e_1_3_2_1_35_1","unstructured":"Joseph Redmon. 2013--2016. Darknet: Open Source Neural Networks in C. http:\/\/pjreddie.com\/darknet\/.  Joseph Redmon. 2013--2016. Darknet: Open Source Neural Networks in C. http:\/\/pjreddie.com\/darknet\/."},{"key":"e_1_3_2_1_36_1","volume-title":"2017 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). IEEE.","author":"Salehian Solmaz","year":"2017","unstructured":"Solmaz Salehian , Jiawen Liu , and Yonghong Yan . 2017 . Comparison of threading programming models . In 2017 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). IEEE. Solmaz Salehian, Jiawen Liu, and Yonghong Yan. 2017. Comparison of threading programming models. In 2017 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). IEEE."},{"key":"e_1_3_2_1_37_1","volume-title":"Measuring the effects of data parallelism on neural network training. arXiv preprint arXiv:1811.03600","author":"Shallue Christopher J","year":"2018","unstructured":"Christopher J Shallue , Jaehoon Lee , Joseph Antognini , Jascha Sohl-Dickstein , Roy Frostig , and George E Dahl . 2018. Measuring the effects of data parallelism on neural network training. arXiv preprint arXiv:1811.03600 ( 2018 ). Christopher J Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E Dahl. 2018. Measuring the effects of data parallelism on neural network training. arXiv preprint arXiv:1811.03600 (2018)."},{"key":"e_1_3_2_1_38_1","volume-title":"International Workshop on OpenMP. Springer, 159--169","author":"Tian Shilei","year":"2021","unstructured":"Shilei Tian , Jon Chesterfield , Johannes Doerfert , and Barbara Chapman . 2021 . Experience Report: Writing a Portable GPU Runtime with OpenMP 5.1 . In International Workshop on OpenMP. Springer, 159--169 . Shilei Tian, Jon Chesterfield, Johannes Doerfert, and Barbara Chapman. 2021. Experience Report: Writing a Portable GPU Runtime with OpenMP 5.1. In International Workshop on OpenMP. Springer, 159--169."},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.3648"},{"key":"e_1_3_2_1_40_1","first-page":"14","article-title":"OpenMP 4.5 compiler optimization for GPU offloading","volume":"64","author":"Tiotto Ettore","year":"2019","unstructured":"Ettore Tiotto , Bardia Mahjour , Whitney Tsang , Xing Xue , Tarique Islam , and Wang Chen . 2019 . OpenMP 4.5 compiler optimization for GPU offloading . IBM Journal of Research and Development 64 , 3\/4 (2019), 14 -- 11 . Ettore Tiotto, Bardia Mahjour, Whitney Tsang, Xing Xue, Tarique Islam, and Wang Chen. 2019. OpenMP 4.5 compiler optimization for GPU offloading. IBM Journal of Research and Development 64, 3\/4 (2019), 14--1.","journal-title":"IBM Journal of Research and Development"},{"key":"e_1_3_2_1_41_1","volume-title":"Proceedings of the 6th ACM\/SPEC International Conference on Performance Engineering. 253--264","author":"Ukidave Yash","year":"2015","unstructured":"Yash Ukidave , Fanny Nina Paravecino , Leiming Yu , Charu Kalra , Amir Momeni , Zhongliang Chen , Nick Materise , Brett Daley , Perhaad Mistry , and David Kaeli . 2015 . Nupar: A benchmark suite for modern gpu architectures . In Proceedings of the 6th ACM\/SPEC International Conference on Performance Engineering. 253--264 . Yash Ukidave, Fanny Nina Paravecino, Leiming Yu, Charu Kalra, Amir Momeni, Zhongliang Chen, Nick Materise, Brett Daley, Perhaad Mistry, and David Kaeli. 2015. Nupar: A benchmark suite for modern gpu architectures. In Proceedings of the 6th ACM\/SPEC International Conference on Performance Engineering. 253--264."},{"volume-title":"C++ parallel programming with threading building blocks","author":"Voss Michael","key":"e_1_3_2_1_42_1","unstructured":"Michael Voss , Rafael Asenjo , and James Reinders . 2019. Pro TBB : C++ parallel programming with threading building blocks . Springer . Michael Voss, Rafael Asenjo, and James Reinders. 2019. Pro TBB: C++ parallel programming with threading building blocks. Springer."}],"event":{"name":"PMAM'23: 14th International Workshop on Programming Models and Applications for Multicores and Manycores","sponsor":["SIGHPC ACM Special Interest Group on High Performance Computing","SIGPLAN ACM Special Interest Group on Programming Languages"],"location":"Montreal QC Canada","acronym":"PMAM'23"},"container-title":["Proceedings of the 14th International Workshop on Programming Models and Applications for Multicores and Manycores"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3582514.3582523","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:47:14Z","timestamp":1750178834000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3582514.3582523"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,25]]},"references-count":42,"alternative-id":["10.1145\/3582514.3582523","10.1145\/3582514"],"URL":"https:\/\/doi.org\/10.1145\/3582514.3582523","relation":{},"subject":[],"published":{"date-parts":[[2023,2,25]]},"assertion":[{"value":"2023-02-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}