{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,1]],"date-time":"2025-11-01T16:30:00Z","timestamp":1762014600970,"version":"build-2065373602"},"publisher-location":"New York, NY, USA","reference-count":39,"publisher":"ACM","license":[{"start":{"date-parts":[[2019,5,13]],"date-time":"2019-05-13T00:00:00Z","timestamp":1557705600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2019,5,13]]},"DOI":"10.1145\/3318170.3318192","type":"proceedings-article","created":{"date-parts":[[2019,7,1]],"date-time":"2019-07-01T19:23:35Z","timestamp":1562009015000},"page":"1-6","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Evaluating data parallelism in C++ using the Parallel Research Kernels"],"prefix":"10.1145","author":[{"given":"Jeff R.","family":"Hammond","sequence":"first","affiliation":[{"name":"Data Center Group, Intel Corporation, Hillsboro, Oregon"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Timothy G.","family":"Mattson","sequence":"additional","affiliation":[{"name":"Parallel Computing Lab, Intel Corporation, Hillsboro, Oregon"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,5,13]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"2019. Parallel Research Kernels. https:\/\/github.com\/ParRes\/Kernels"},{"key":"e_1_3_2_1_2_1","unstructured":"2019. Travis CI -- ParRes\/Kernels. https:\/\/travis-ci.org\/ParRes\/Kernels"},{"volume-title":"Evolving OpenMP for Evolving Architectures, Bronis R","author":"Agrawal Vishakha","key":"e_1_3_2_1_3_1","unstructured":"Vishakha Agrawal, Michael J. Voss, Pablo Reble, Vasanth Tovinkere, Jeff Hammond, and Michael Klemm. 2018. Visualization of OpenMP* Task Dependencies Using Intel\u00ae Advisor -- Flow Graph Analyzer. In Evolving OpenMP for Evolving Architectures, Bronis R. de Supinski, Pedro Valero-Lara, Xavier Martorell, Sergi Mateo Bellido, and Jesus Labarta (Eds.). Springer International Publishing, Cham, 175--188."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2015.30"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-05215-1_12"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/2141702.2141703"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1155\/2012\/917630"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/XSW.2013.7"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2014.07.003"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2966884.2966916"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3127024.3127026"},{"volume-title":"2016 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 73--82","author":"Georganas Evangelos","key":"e_1_3_2_1_12_1","unstructured":"Evangelos Georganas, Rob F. Van der Wijngaart, and Timothy G. Mattson. 2016. Design and Implementation of a Parallel Research Kernel for Assessing Dynamic Load-Balancing Capabilities. In 2016 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 73--82."},{"key":"e_1_3_2_1_13_1","unstructured":"The Khronos 'Group. {n. d.}. Open Source Parallel STL implementation. https:\/\/github.com\/KhronosGroup\/SyclParallelSTL"},{"key":"e_1_3_2_1_14_1","unstructured":"Georg Hager. 2019. The McCalpin STREAM benchmark: How do do it right and interpret the results. https:\/\/blogs.fau.de\/hager\/archives\/8263"},{"key":"e_1_3_2_1_15_1","volume-title":"Keasler","author":"Hornung Richard D.","year":"2014","unstructured":"Richard D. Hornung and Jeffrey A. Keasler. 2014. The RAJA Portability Layer: Overview and Status. (9 2014)."},{"key":"e_1_3_2_1_16_1","unstructured":"Intel Corporation. {n. d.}. Threading Building Blocks (TBB). https:\/\/github.com\/01org\/tbb. https:\/\/www.threadingbuildingblocks.org\/"},{"volume-title":"ISO. 2017. ISO\/IEC 14882:2017 Information technology --- Programming languages --- C++","key":"e_1_3_2_1_17_1","unstructured":"ISO. 2017. ISO\/IEC 14882:2017 Information technology --- Programming languages --- C++ (fifth ed.). International Organization for Standardization, Geneva, Switzerland. 1605 pages. https:\/\/www.iso.org\/standard\/68564.html"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2832241.2832244"},{"volume-title":"Comparative Performance and Optimization of Chapel in Modern Manycore Architectures. In 2017 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). 1105--1114","author":"Kayraklioglu E.","key":"e_1_3_2_1_19_1","unstructured":"E. Kayraklioglu, W. Chang, and T. El-Ghazawi. 2017. Comparative Performance and Optimization of Chapel in Modern Manycore Architectures. In 2017 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). 1105--1114."},{"key":"e_1_3_2_1_20_1","unstructured":"Khronos OpenCL Working Group. 2012. The OpenCL Specification Version 1.2 Aaftab Munshi (Ed.). https:\/\/www.khronos.org\/registry\/cl\/specs\/opencl-1.2.pdf"},{"key":"e_1_3_2_1_21_1","unstructured":"Lawrence Livermore National Laboratory. {n. d.}. RAJA Performance Portability Layer. https:\/\/github.com\/LLNL\/RAJA"},{"key":"e_1_3_2_1_22_1","unstructured":"Sandia National Laboratory. {n. d.}. Kokkos C++ Performance Portability Programming EcoSystem: The Programming Model -- Parallel Execution and Memory Abstraction. https:\/\/github.com\/Kokkos\/kokkos"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1188455.1188677"},{"key":"e_1_3_2_1_24_1","unstructured":"Devin Matthews. {n. d.}. TBLIS (Tensor BLIS). https:\/\/github.com\/devinamatthews\/tblis"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.5555\/1406956"},{"key":"e_1_3_2_1_26_1","volume-title":"Memory bandwidth and machine balance in high performance computers","author":"McCalpin John","year":"1995","unstructured":"John McCalpin. 1995. Memory bandwidth and machine balance in high performance computers. IEEE Technical Committee on Computer Architecture Newsletter (12 1995), 19--25."},{"key":"e_1_3_2_1_27_1","volume-title":"STREAM: Sustainable Memory Bandwidth in High Performance Computers. https:\/\/www.cs.virginia.edu\/stream\/","author":"McCalpin John D.","year":"2015","unstructured":"John D. McCalpin. 2015. STREAM: Sustainable Memory Bandwidth in High Performance Computers. https:\/\/www.cs.virginia.edu\/stream\/"},{"volume-title":"Symmetric Memory Partitions in OpenSHMEM: A Case Study with Intel KNL","author":"Namashivayam Naveen","key":"e_1_3_2_1_28_1","unstructured":"Naveen Namashivayam, Bob Cernohous, Krishna Kandalla, Dan Pou, Joseph Robichaux, James Dinan, and Mark Pagel. 2018. Symmetric Memory Partitions in OpenSHMEM: A Case Study with Intel KNL. In OpenSHMEM and Related Technologies. Big Compute and Big Data Convergence, Manjunath Gorentla Venkata, Neena Imam, and Swaroop Pophale (Eds.). Springer International Publishing, Cham, 3--18."},{"key":"e_1_3_2_1_29_1","unstructured":"OpenMP Architecture Review Board. 2015. OpenMP Aplication Program Interface -- Version 4.5. https:\/\/www.openmp.org\/wp-content\/uploads\/openmp-4.5.pdf."},{"key":"e_1_3_2_1_30_1","unstructured":"OpenMP Architecture Review Board. 2018. OpenMP Aplication Program Interface - Version 5.0. https:\/\/www.openmp.org\/wp-content\/uploads\/OpenMP-API-Specification-5.0.pdf."},{"volume-title":"2006 IEEE International Conference on Cluster Computing. 1--7.","author":"Plimpton S.J.","key":"e_1_3_2_1_31_1","unstructured":"S.J. Plimpton, R. Brightwell, C. Vaughan, K. Underwood, and M. Davis. 2006. A Simple Synchronous Distributed-Memory Algorithm for the HPCC RandomAccess Benchmark. In 2006 IEEE International Conference on Cluster Computing. 1--7."},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.5555\/2396095.2396141"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2014.110"},{"key":"e_1_3_2_1_34_1","unstructured":"Khronos\u00ae OpenCL\u2122 Working Group SYCL\u2122 subgroup. 2019. SYCL\u2122 Specification. https:\/\/www.khronos.org\/registry\/SYCL\/specs\/sycl-1.2.1.pdf Ronan Keryell Maria Rovatsou and Lee Howes (Eds.)."},{"volume-title":"A New Parallel Research Kernel to Expand Research on Dynamic Load-Balancing Capabilities","author":"Van der Wijngaart Rob F.","key":"e_1_3_2_1_35_1","unstructured":"Rob F. Van der Wijngaart, Evangelos Georganas, Timothy G. Mattson, and Andrew Wissink. 2017. A New Parallel Research Kernel to Expand Research on Dynamic Load-Balancing Capabilities. In High Performance Computing, Julian M. Kunkel, Rio Yokota, Pavan Balaji, and David Keyes (Eds.). Springer International Publishing, Cham, 256--274."},{"volume-title":"Comparing Runtime Systems with Exascale Ambitions Using the Parallel Research Kernels","author":"Van der Wijngaart Rob F.","key":"e_1_3_2_1_36_1","unstructured":"Rob F. Van der Wijngaart, Abdullah Kayi, Jeff R. Hammond, Gabriele Jost, Tom St. John, Srinivas Sridharan, Timothy G. Mattson, John Abercrombie, and Jacob Nelson. 2016. Comparing Runtime Systems with Exascale Ambitions Using the Parallel Research Kernels. In High Performance Computing, Julian M. Kunkel, Pavan Balaji, and Jack Dongarra (Eds.). Springer International Publishing, Cham, 321--339."},{"volume-title":"Proceedings of the IEEE High Performance Extreme Computing Conference. IEEE.","author":"Rob","key":"e_1_3_2_1_37_1","unstructured":"Rob F. Van der Wijngaart and Timothy G. Mattson. 2014. The Parallel Research Kernels: A tool for architecture and programming system investigation. In Proceedings of the IEEE High Performance Extreme Computing Conference. IEEE."},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/PGAS.2015.24"},{"key":"e_1_3_2_1_39_1","unstructured":"Field van Zee et al. {n. d.}. BLIS. https:\/\/github.com\/flame\/blis"}],"event":{"name":"IWOCL'19: International Workshop on OpenCL","sponsor":["Khronos Khronos Group","Northeastern University","Codeplay Codeplay Software Ltd.","Intel Intel","The University of Bristol The University of Bristol"],"location":"Boston MA USA","acronym":"IWOCL'19"},"container-title":["Proceedings of the International Workshop on OpenCL"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3318170.3318192","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3318170.3318192","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T17:49:34Z","timestamp":1750268974000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3318170.3318192"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,5,13]]},"references-count":39,"alternative-id":["10.1145\/3318170.3318192","10.1145\/3318170"],"URL":"https:\/\/doi.org\/10.1145\/3318170.3318192","relation":{},"subject":[],"published":{"date-parts":[[2019,5,13]]},"assertion":[{"value":"2019-05-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}