{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,12]],"date-time":"2026-07-12T04:16:38Z","timestamp":1783829798329,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":101,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,5,31]],"date-time":"2020-05-31T00:00:00Z","timestamp":1590883200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"publisher","award":["MA4662-5"],"award-info":[{"award-number":["MA4662-5"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100007601","name":"Horizon 2020","doi-asserted-by":"publisher","award":["780245"],"award-info":[{"award-number":["780245"]}],"id":[{"id":"10.13039\/501100007601","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Bundesministerium f\u00fcr Bildung und Forschung","award":["01IS14013A, 01IS18025A and 01IS18037A"],"award-info":[{"award-number":["01IS14013A, 01IS18025A and 01IS18037A"]}]},{"name":"Bundesministerium f\u00fcr Wirtschaft und Energie","award":["01MD19002B"],"award-info":[{"award-number":["01MD19002B"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,6,11]]},"DOI":"10.1145\/3318464.3389705","type":"proceedings-article","created":{"date-parts":[[2020,5,29]],"date-time":"2020-05-29T17:12:33Z","timestamp":1590772353000},"page":"1633-1649","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":90,"title":["Pump Up the Volume"],"prefix":"10.1145","author":[{"given":"Clemens","family":"Lutz","sequence":"first","affiliation":[{"name":"DFKI GmbH, Berlin, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sebastian","family":"Bre\u00df","sequence":"additional","affiliation":[{"name":"TU Berlin, Berlin, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Steffen","family":"Zeuch","sequence":"additional","affiliation":[{"name":"DFKI GmbH, Berlin, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tilmann","family":"Rabl","sequence":"additional","affiliation":[{"name":"HPI, University of Potsdam, Potsdam, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Volker","family":"Markl","sequence":"additional","affiliation":[{"name":"DFKI GmbH, TU Berlin, Berlin, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,5,31]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/HOTCHIPS.2009.7478337"},{"key":"e_1_3_2_2_2_1","unstructured":"Brian Allison. 2018. Introduction to the OpenCAPI Interface. https:\/\/openpowerfoundation.org\/wp-content\/uploads\/2018\/10\/ Brian-Allison.OPF_OpenCAPI_FPGA_Overview_V1--1.pdf. In Open- POWER Summit Europe.  Brian Allison. 2018. Introduction to the OpenCAPI Interface. https:\/\/openpowerfoundation.org\/wp-content\/uploads\/2018\/10\/ Brian-Allison.OPF_OpenCAPI_FPGA_Overview_V1--1.pdf. In Open- POWER Summit Europe."},{"key":"e_1_3_2_2_3_1","volume-title":"AMD Radeon Instinct GPUs and ROCm Open Source Software to Power World's Fastest Supercomputer at Oak Ridge National Laboratory. Retrieved","author":"AMD.","year":"2019","unstructured":"AMD. 2019. AMD EPYC CPUs , AMD Radeon Instinct GPUs and ROCm Open Source Software to Power World's Fastest Supercomputer at Oak Ridge National Laboratory. Retrieved July 5, 2019 from https:\/\/www.amd.com\/en\/press-releases\/2019-05-07-amd-epyccpus- radeon-instinct-gpus-and-rocm-open-source-software-topower AMD. 2019. AMD EPYC CPUs, AMD Radeon Instinct GPUs and ROCm Open Source Software to Power World's Fastest Supercomputer at Oak Ridge National Laboratory. Retrieved July 5, 2019 from https:\/\/www.amd.com\/en\/press-releases\/2019-05-07-amd-epyccpus- radeon-instinct-gpus-and-rocm-open-source-software-topower"},{"key":"e_1_3_2_2_4_1","unstructured":"Raja Appuswamy Manos Karpathiotakis Danica Porobic and Anastasia Ailamaki. 2017. The Case For Heterogeneous HTAP. In CIDR.  Raja Appuswamy Manos Karpathiotakis Danica Porobic and Anastasia Ailamaki. 2017. The Case For Heterogeneous HTAP. In CIDR."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"crossref","unstructured":"Arash Ashari Shirish Tatikonda Matthias Boehm Berthold Reinwald Keith Campbell John Keenleyside and P. Sadayappan. 2015. On optimizing machine learning workloads via kernel fusion. In PPoPP. 173--182. https:\/\/doi.org\/10.1145\/2688500.2688521  Arash Ashari Shirish Tatikonda Matthias Boehm Berthold Reinwald Keith Campbell John Keenleyside and P. Sadayappan. 2015. On optimizing machine learning workloads via kernel fusion. In PPoPP. 173--182. https:\/\/doi.org\/10.1145\/2688500.2688521","DOI":"10.1145\/2858788.2688521"},{"key":"e_1_3_2_2_6_1","volume-title":"Owens","author":"Ashkiani Saman","year":"2018","unstructured":"Saman Ashkiani , Shengren Li , Martin Farach-Colton , Nina Amenta , and John D . Owens . 2018 . GPU LSM : A Dynamic Dictionary Data Structure for the GPU. In IPDPS. 430--440. https:\/\/doi.org\/10.1109\/ IPDPS. 2018.00053 Saman Ashkiani, Shengren Li, Martin Farach-Colton, Nina Amenta, and John D. Owens. 2018. GPU LSM: A Dynamic Dictionary Data Structure for the GPU. In IPDPS. 430--440. https:\/\/doi.org\/10.1109\/ IPDPS.2018.00053"},{"key":"e_1_3_2_2_7_1","volume-title":"Martin Farach- Colton, and John D. Owens","author":"Awad Muhammad A.","year":"2019","unstructured":"Muhammad A. Awad , Saman Ashkiani , Rob Johnson , Martin Farach- Colton, and John D. Owens . 2019 . Engineering a high-performance GPU B-Tree. In PPoPP. 145--157. https:\/\/doi.org\/10.1145\/3293883. 3295706 Muhammad A. Awad, Saman Ashkiani, Rob Johnson, Martin Farach- Colton, and John D. Owens. 2019. Engineering a high-performance GPU B-Tree. In PPoPP. 145--157. https:\/\/doi.org\/10.1145\/3293883. 3295706"},{"key":"e_1_3_2_2_8_1","volume-title":"Main-memory hash joins on multi-core CPUs: Tuning to the underlying hardware","author":"Balkesen Cagri","year":"2013","unstructured":"Cagri Balkesen , Jens Teubner , Gustavo Alonso , and M. Tamer \u00d6zsu . 2013. Main-memory hash joins on multi-core CPUs: Tuning to the underlying hardware . In ICDE. IEEE , New York, NY, USA , 362--373. https:\/\/doi.org\/10.1109\/ICDE. 2013 .6544839 Cagri Balkesen, Jens Teubner, Gustavo Alonso, and M. Tamer \u00d6zsu. 2013. Main-memory hash joins on multi-core CPUs: Tuning to the underlying hardware. In ICDE. IEEE, New York, NY, USA, 362--373. https:\/\/doi.org\/10.1109\/ICDE.2013.6544839"},{"key":"e_1_3_2_2_9_1","volume-title":"SIGMOD","author":"Barthels Claude","unstructured":"Claude Barthels , Simon Loesing , Gustavo Alonso , and Donald Kossmann . 2015. Rack-Scale In-Memory Join Processing using RDMA . In SIGMOD . ACM , New York, NY, USA , 1463--1475. https:\/\/doi.org\/10. 1145\/2723372.2750547 Claude Barthels, Simon Loesing, Gustavo Alonso, and Donald Kossmann. 2015. Rack-Scale In-Memory Join Processing using RDMA. In SIGMOD. ACM, New York, NY, USA, 1463--1475. https:\/\/doi.org\/10. 1145\/2723372.2750547"},{"key":"e_1_3_2_2_10_1","volume-title":"Patel","author":"Blanas Spyros","year":"2011","unstructured":"Spyros Blanas , Yinan Li , and Jignesh M . Patel . 2011 . Design and evaluation of main memory hash join algorithms for multi-core CPUs. In SIGMOD. ACM, New York, NY, USA , 37--48. https:\/\/doi.org\/10. 1145\/1989323.1989328 Spyros Blanas, Yinan Li, and Jignesh M. Patel. 2011. Design and evaluation of main memory hash join algorithms for multi-core CPUs. In SIGMOD. ACM, New York, NY, USA, 37--48. https:\/\/doi.org\/10. 1145\/1989323.1989328"},{"key":"e_1_3_2_2_11_1","unstructured":"Rajesh Bordawekar and Pidad Gasfer D'Souza. 2018. Evaluation of hybrid cache-coherent concurrent hash table on IBM POWER9 AC922 system with NVLink2. http:\/\/on-demand.gputechconf.com\/gtc\/2018\/ video\/S8172\/. In GTC. Nvidia.  Rajesh Bordawekar and Pidad Gasfer D'Souza. 2018. Evaluation of hybrid cache-coherent concurrent hash table on IBM POWER9 AC922 system with NVLink2. http:\/\/on-demand.gputechconf.com\/gtc\/2018\/ video\/S8172\/. In GTC. Nvidia."},{"key":"e_1_3_2_2_12_1","volume-title":"Applying AMD's Kaveri APU for heterogeneous computing","author":"Bouvier Dan","year":"2014","unstructured":"Dan Bouvier and Ben Sander . 2014. Applying AMD's Kaveri APU for heterogeneous computing . In HCS. IEEE , New York, NY, USA , 1--42. https:\/\/doi.org\/10.1109\/HOTCHIPS. 2014 .7478810 Dan Bouvier and Ben Sander. 2014. Applying AMD's Kaveri APU for heterogeneous computing. In HCS. IEEE, New York, NY, USA, 1--42. https:\/\/doi.org\/10.1109\/HOTCHIPS.2014.7478810"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2882936"},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-018-0512-y"},{"key":"e_1_3_2_2_15_1","volume-title":"IBM power system AC922 introduction and technical overview. IBM","author":"Caldeira Alexandre Bicas","unstructured":"Alexandre Bicas Caldeira . 2018. IBM power system AC922 introduction and technical overview. IBM , International Technical Support Organization . Alexandre Bicas Caldeira. 2018. IBM power system AC922 introduction and technical overview. IBM, International Technical Support Organization."},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.14778\/3303753.3303760"},{"key":"e_1_3_2_2_17_1","unstructured":"Periklis Chrysogelos Panagiotis Sioulas and Anastasia Ailamaki. 2019. Hardware-conscious Query Processing in GPU-accelerated Analytical Engines. In CIDR.  Periklis Chrysogelos Panagiotis Sioulas and Anastasia Ailamaki. 2019. Hardware-conscious Query Processing in GPU-accelerated Analytical Engines. In CIDR."},{"key":"e_1_3_2_2_18_1","volume-title":"Making kernel","author":"Corbet Jonathan","year":"2015","unstructured":"Jonathan Corbet . 2015. Making kernel pages movable. LWN. net ( July 2015 ). https:\/\/lwn.net\/Articles\/650917\/ Jonathan Corbet. 2015. Making kernel pages movable. LWN.net (July 2015). https:\/\/lwn.net\/Articles\/650917\/"},{"key":"e_1_3_2_2_19_1","unstructured":"CXL 2019. Compute Express Link Specification Revision 1.1. CXL. https:\/\/www.computeexpresslink.org  CXL 2019. Compute Express Link Specification Revision 1.1. CXL. https:\/\/www.computeexpresslink.org"},{"key":"e_1_3_2_2_20_1","article-title":"From a Comprehensive Experimental Survey to a Cost-based Selection Strategy for Lightweight Integer Compression","volume":"44","author":"Damme Patrick","year":"2019","unstructured":"Patrick Damme , Annett Ungeth\u00fcm , Juliana Hildebrandt , Dirk Habich , and Wolfgang Lehner . 2019 . From a Comprehensive Experimental Survey to a Cost-based Selection Strategy for Lightweight Integer Compression Algorithms. Trans. Database Syst. 44 , 3 (2019), 9:1--9:46. https:\/\/doi.org\/10.1145\/3323991 Patrick Damme, Annett Ungeth\u00fcm, Juliana Hildebrandt, Dirk Habich, and Wolfgang Lehner. 2019. From a Comprehensive Experimental Survey to a Cost-based Selection Strategy for Lightweight Integer Compression Algorithms. Trans. Database Syst. 44, 3 (2019), 9:1--9:46. https:\/\/doi.org\/10.1145\/3323991","journal-title":"Algorithms. Trans. Database Syst."},{"key":"e_1_3_2_2_21_1","volume-title":"Hitting the accelerator: the next generation of machine-learning chips. Retrieved","year":"2019","unstructured":"Deloitte. 2017. Hitting the accelerator: the next generation of machine-learning chips. Retrieved Oct 1, 2019 from https:\/\/www2.deloitte.com\/content\/dam\/Deloitte\/global\/Images\/ infographics\/technologymediatelecommunications\/gx-deloittetmt- 2018-nextgen-machine-learning-report.pdf Deloitte. 2017. Hitting the accelerator: the next generation of machine-learning chips. Retrieved Oct 1, 2019 from https:\/\/www2.deloitte.com\/content\/dam\/Deloitte\/global\/Images\/ infographics\/technologymediatelecommunications\/gx-deloittetmt- 2018-nextgen-machine-learning-report.pdf"},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.14778\/3352063.3352137"},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.14778\/2994509.2994515"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.14778\/1920841.1920927"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDCS.2009.32"},{"key":"e_1_3_2_2_26_1","volume-title":"SIGMOD","author":"Funke Henning","year":"1837","unstructured":"Henning Funke , Sebastian Bre\u00df , Stefan Noll , Volker Markl , and Jens Teubner . 2018. Pipelined Query Processing in Coprocessor Environments . In SIGMOD . ACM , New York, NY, USA , 1603--1618. https:\/\/doi.org\/10.1145\/3 1837 13.3183734 Henning Funke, Sebastian Bre\u00df, Stefan Noll, Volker Markl, and Jens Teubner. 2018. Pipelined Query Processing in Coprocessor Environments. In SIGMOD. ACM, New York, NY, USA, 1603--1618. https:\/\/doi.org\/10.1145\/3183713.3183734"},{"key":"e_1_3_2_2_27_1","volume-title":"Pe\u00f1a","author":"Garcia-Flores Victor","year":"2017","unstructured":"Victor Garcia-Flores , Eduard Ayguad\u00e9 , and Antonio J . Pe\u00f1a . 2017 . Efficient Data Sharing on Heterogeneous Systems. In ICPP. 121--130. https:\/\/doi.org\/10.1109\/ICPP.2017.21 Victor Garcia-Flores, Eduard Ayguad\u00e9, and Antonio J. Pe\u00f1a. 2017. Efficient Data Sharing on Heterogeneous Systems. In ICPP. 121--130. https:\/\/doi.org\/10.1109\/ICPP.2017.21"},{"key":"e_1_3_2_2_28_1","volume-title":"Gartner Says the Future of the Database Market Is the Cloud. Retrieved","year":"2019","unstructured":"Gartner. 2019. Gartner Says the Future of the Database Market Is the Cloud. Retrieved Oct 1, 2019 from https:\/\/www.gartner.com\/en\/newsroom\/press-releases\/2019- 07-01-gartner-says-the-future-of-the-database-market-is-the Gartner. 2019. Gartner Says the Future of the Database Market Is the Cloud. Retrieved Oct 1, 2019 from https:\/\/www.gartner.com\/en\/newsroom\/press-releases\/2019- 07-01-gartner-says-the-future-of-the-database-market-is-the"},{"key":"e_1_3_2_2_29_1","volume-title":"SIGMOD","author":"Govindaraju Naga K.","unstructured":"Naga K. Govindaraju , Brandon Lloyd , Wei Wang , Ming C. Lin , and Dinesh Manocha . 2004. Fast Computation of Database Operations using Graphics Processors . In SIGMOD . ACM , New York, NY, USA , 215--226. https:\/\/doi.org\/10.1145\/1007568.1007594 Naga K. Govindaraju, Brandon Lloyd, Wei Wang, Ming C. Lin, and Dinesh Manocha. 2004. Fast Computation of Database Operations using Graphics Processors. In SIGMOD. ACM, New York, NY, USA, 215--226. https:\/\/doi.org\/10.1145\/1007568.1007594"},{"key":"e_1_3_2_2_30_1","first-page":"1","article-title":"Accelerating the Unacceleratable: Hybrid CPU\/GPU Algorithms for Memory-Bound Database Primitives. In DaMoN. ACM, New York","volume":"7","author":"Gowanlock Michael","year":"2019","unstructured":"Michael Gowanlock , Ben Karsin , Zane Fink , and Jordan Wright . 2019 . Accelerating the Unacceleratable: Hybrid CPU\/GPU Algorithms for Memory-Bound Database Primitives. In DaMoN. ACM, New York , NY, USA , 7 : 1 -- 7 :11. https:\/\/doi.org\/10.1145\/3329785.3329926 Michael Gowanlock, Ben Karsin, Zane Fink, and Jordan Wright. 2019. Accelerating the Unacceleratable: Hybrid CPU\/GPU Algorithms for Memory-Bound Database Primitives. In DaMoN. ACM, New York, NY, USA, 7:1--7:11. https:\/\/doi.org\/10.1145\/3329785.3329926","journal-title":"NY, USA"},{"key":"e_1_3_2_2_31_1","volume-title":"Hazelwood","author":"Gregg Chris","year":"2011","unstructured":"Chris Gregg and Kim M . Hazelwood . 2011 . Where is the data? Why you cannot debate CPU vs. GPU performance without the answer. In ISPASS. IEEE, New York, NY, USA, 134--144. https:\/\/doi.org\/10.1109\/ ISPASS. 2011.5762730 Chris Gregg and Kim M. Hazelwood. 2011. Where is the data? Why you cannot debate CPU vs. GPU performance without the answer. In ISPASS. IEEE, New York, NY, USA, 134--144. https:\/\/doi.org\/10.1109\/ ISPASS.2011.5762730"},{"key":"e_1_3_2_2_32_1","first-page":"1","article-title":"Fluid Co-processing: GPU Bloom-filters for CPU Joins. In DaMoN. ACM, New York","volume":"9","author":"Gubner Tim","year":"2019","unstructured":"Tim Gubner , Diego G. Tom\u00e9 , Harald Lang , and Peter A. Boncz . 2019 . Fluid Co-processing: GPU Bloom-filters for CPU Joins. In DaMoN. ACM, New York , NY, USA , 9 : 1 -- 9 :10. https:\/\/doi.org\/10.1145\/3329785. 3329934 Tim Gubner, Diego G. Tom\u00e9, Harald Lang, and Peter A. Boncz. 2019. Fluid Co-processing: GPU Bloom-filters for CPU Joins. In DaMoN. ACM, New York, NY, USA, 9:1--9:10. https:\/\/doi.org\/10.1145\/3329785. 3329934","journal-title":"NY, USA"},{"key":"e_1_3_2_2_33_1","unstructured":"Prabhat K. Gupta. 2016. Accelerating Datacenter Workloads. In FPL. 1--27.  Prabhat K. Gupta. 2016. Accelerating Datacenter Workloads. In FPL. 1--27."},{"key":"e_1_3_2_2_34_1","volume-title":"Research 18: Main Memory Databases and Modern Hardware SIGMOD '20, June 14--19","author":"Bingsheng","year":"2020","unstructured":"Bingsheng He et al. 2009. Relational query coprocessing on graphics processors. TODS 34, 4 (2009) . Research 18: Main Memory Databases and Modern Hardware SIGMOD '20, June 14--19 , 2020 , Portland, OR, USA 1647 Bingsheng He et al. 2009. Relational query coprocessing on graphics processors. TODS 34, 4 (2009). Research 18: Main Memory Databases and Modern Hardware SIGMOD '20, June 14--19, 2020, Portland, OR, USA 1647"},{"key":"e_1_3_2_2_35_1","volume-title":"Sander","author":"He Bingsheng","year":"2008","unstructured":"Bingsheng He , Ke Yang , Rui Fang , Mian Lu , Naga K. Govindaraju , Qiong Luo , and Pedro V . Sander . 2008 . Relational joins on graphics processors. In SIGMOD. ACM, New York, NY, USA , 511--524. https: \/\/doi.org\/10.1145\/1376616.1376670 Bingsheng He, Ke Yang, Rui Fang, Mian Lu, Naga K. Govindaraju, Qiong Luo, and Pedro V. Sander. 2008. Relational joins on graphics processors. In SIGMOD. ACM, New York, NY, USA, 511--524. https: \/\/doi.org\/10.1145\/1376616.1376670"},{"key":"e_1_3_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.14778\/2536206.2536216"},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.14778\/2735496.2735497"},{"key":"e_1_3_2_2_38_1","first-page":"709","article-title":"Hardware-oblivious parallelism for inmemory column-stores","volume":"6","author":"Max Heimel","year":"2013","unstructured":"Max Heimel et al. 2013 . Hardware-oblivious parallelism for inmemory column-stores . PVLDB 6 , 9 (2013), 709 -- 720 . Max Heimel et al. 2013. Hardware-oblivious parallelism for inmemory column-stores. PVLDB 6, 9 (2013), 709--720.","journal-title":"PVLDB"},{"key":"e_1_3_2_2_39_1","volume-title":"Wood","author":"Hestness Joel","year":"2014","unstructured":"Joel Hestness , Stephen W. Keckler , and David A . Wood . 2014 . A comparative analysis of microarchitecture effects on CPU and GPU memory system behavior. In IISWC. IEEE, New York, NY, USA , 150-- 160. https:\/\/doi.org\/10.1109\/IISWC.2014.6983054 Joel Hestness, Stephen W. Keckler, and David A. Wood. 2014. A comparative analysis of microarchitecture effects on CPU and GPU memory system behavior. In IISWC. IEEE, New York, NY, USA, 150-- 160. https:\/\/doi.org\/10.1109\/IISWC.2014.6983054"},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"crossref","unstructured":"IBM 2018. POWER9 Processor User's Manual Version 2.0. IBM.  IBM 2018. POWER9 Processor User's Manual Version 2.0. IBM.","DOI":"10.1147\/JRD.2018.2854039"},{"key":"e_1_3_2_2_41_1","first-page":"9","article-title":"Functionality and performance of NVLink with IBM POWER9 processors","volume":"62","author":"IBM","year":"2018","unstructured":"IBM POWER9 NPU team. 2018 . Functionality and performance of NVLink with IBM POWER9 processors . IBM Journal of Research and Development 62 , 4\/5 (2018), 9 . IBM POWER9 NPU team. 2018. Functionality and performance of NVLink with IBM POWER9 processors. IBM Journal of Research and Development 62, 4\/5 (2018), 9.","journal-title":"IBM Journal of Research and Development"},{"key":"e_1_3_2_2_42_1","unstructured":"Intel 2018. Intel 64 and IA-32 Architectures Software Developer's Manual. Intel.  Intel 2018. Intel 64 and IA-32 Architectures Software Developer's Manual. Intel."},{"key":"e_1_3_2_2_43_1","volume-title":"Intel Stratix 10 DX FPGA Product Brief. Retrieved Accessed","year":"2019","unstructured":"Intel. 2019. Intel Stratix 10 DX FPGA Product Brief. Retrieved Accessed : Oct 2, 2019 from https:\/\/www.intel.com\/content\/dam\/ www\/programmable\/us\/en\/pdfs\/literature\/solution-sheets\/stratix- 10-dx-product-brief.pdf Intel. 2019. Intel Stratix 10 DX FPGA Product Brief. Retrieved Accessed: Oct 2, 2019 from https:\/\/www.intel.com\/content\/dam\/ www\/programmable\/us\/en\/pdfs\/literature\/solution-sheets\/stratix- 10-dx-product-brief.pdf"},{"key":"e_1_3_2_2_44_1","volume-title":"Retrieved","year":"2019","unstructured":"Intel. 2019 . Intel Unveils New GPU Architecture with High- Performance Computing and AI Acceleration, and oneAPI Software Stack with Unified and Scalable Abstraction for Heterogeneous Architectures . Retrieved Jan 29, 2020 from https:\/\/newsroom.intel.com\/news-releases\/intel-unveils-new-gpuarchitecture- optimized-for-hpc-ai-oneapi Intel. 2019. Intel Unveils New GPU Architecture with High- Performance Computing and AI Acceleration, and oneAPI Software Stack with Unified and Scalable Abstraction for Heterogeneous Architectures. Retrieved Jan 29, 2020 from https:\/\/newsroom.intel.com\/news-releases\/intel-unveils-new-gpuarchitecture- optimized-for-hpc-ai-oneapi"},{"key":"e_1_3_2_2_45_1","volume-title":"Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking. CoRR abs\/1804.06826","author":"Jia Zhe","year":"2018","unstructured":"Zhe Jia , Marco Maggioni , Benjamin Staiger , and Daniele Paolo Scarpazza . 2018. Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking. CoRR abs\/1804.06826 ( 2018 ). arXiv:1804.06826 Zhe Jia, Marco Maggioni, Benjamin Staiger, and Daniele Paolo Scarpazza. 2018. Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking. CoRR abs\/1804.06826 (2018). arXiv:1804.06826"},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"crossref","unstructured":"Krzysztof Kaczmarski. 2012. B + -Tree Optimized for GPGPU. In OTM. 843--854. https:\/\/doi.org\/10.1007\/978--3--642--33615--7_27  Krzysztof Kaczmarski. 2012. B + -Tree Optimized for GPGPU. In OTM. 843--854. https:\/\/doi.org\/10.1007\/978--3--642--33615--7_27","DOI":"10.1007\/978-3-642-33615-7_27"},{"key":"e_1_3_2_2_47_1","volume-title":"DaMoN","author":"Kaldewey Tim","unstructured":"Tim Kaldewey , Guy M. Lohman , Ren\u00e9 M\u00fcller , and Peter Benjamin Volk . 2012. GPU join processing revisited . In DaMoN . ACM , New York, NY, USA , 55--62. https:\/\/doi.org\/10.1145\/2236584.2236592 Tim Kaldewey, Guy M. Lohman, Ren\u00e9 M\u00fcller, and Peter Benjamin Volk. 2012. GPU join processing revisited. In DaMoN. ACM, New York, NY, USA, 55--62. https:\/\/doi.org\/10.1145\/2236584.2236592"},{"key":"e_1_3_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.14778\/3297753.3297756"},{"key":"e_1_3_2_2_49_1","first-page":"1","article-title":"Big data causing big (TLB) problems: taming random memory accesses on the GPU. In DaMoN. ACM, New York","volume":"6","author":"Karnagel Tomas","year":"2017","unstructured":"Tomas Karnagel , Tal Ben-Nun , Matthias Werner , Dirk Habich , and Wolfgang Lehner . 2017 . Big data causing big (TLB) problems: taming random memory accesses on the GPU. In DaMoN. ACM, New York , NY, USA , 6 : 1 -- 6 :10. https:\/\/doi.org\/10.1145\/3076113.3076115 Tomas Karnagel, Tal Ben-Nun, Matthias Werner, Dirk Habich, and Wolfgang Lehner. 2017. Big data causing big (TLB) problems: taming random memory accesses on the GPU. In DaMoN. ACM, New York, NY, USA, 6:1--6:10. https:\/\/doi.org\/10.1145\/3076113.3076115","journal-title":"NY, USA"},{"key":"e_1_3_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.14778\/3067421.3067423"},{"key":"e_1_3_2_2_51_1","volume-title":"Lohman","author":"Karnagel Tomas","year":"2015","unstructured":"Tomas Karnagel , Ren\u00e9 M\u00fcller , and Guy M . Lohman . 2015 . Optimizing GPU-accelerated group-by and aggregation. In ADMS. ACM, New York, NY, USA , 13--24. Tomas Karnagel, Ren\u00e9 M\u00fcller, and Guy M. Lohman. 2015. Optimizing GPU-accelerated group-by and aggregation. In ADMS. ACM, New York, NY, USA, 13--24."},{"key":"e_1_3_2_2_52_1","volume-title":"Bhuyan","author":"Khorasani Farzad","year":"2015","unstructured":"Farzad Khorasani , Mehmet E. Belviranli , Rajiv Gupta , and Laxmi N . Bhuyan . 2015 . Stadium Hashing : Scalable and Flexible Hashing on GPUs. In PACT. IEEE, New York, NY, USA , 63--74. https:\/\/doi.org\/10. 1109\/PACT.2015.13 Farzad Khorasani, Mehmet E. Belviranli, Rajiv Gupta, and Laxmi N. Bhuyan. 2015. Stadium Hashing: Scalable and Flexible Hashing on GPUs. In PACT. IEEE, New York, NY, USA, 63--74. https:\/\/doi.org\/10. 1109\/PACT.2015.13"},{"key":"e_1_3_2_2_53_1","unstructured":"Changkyu Kim Jatin Chhugani Nadathur Satish Eric Sedlar Anthony D. Nguyen Tim Kaldewey Victor W. Lee Scott A. Brandt and Pradeep Dubey. 2010. FAST: fast architecture sensitive tree search on modern CPUs and GPUs. In SIGMOD. 339--350. https: \/\/doi.org\/10.1145\/1807167.1807206  Changkyu Kim Jatin Chhugani Nadathur Satish Eric Sedlar Anthony D. Nguyen Tim Kaldewey Victor W. Lee Scott A. Brandt and Pradeep Dubey. 2010. FAST: fast architecture sensitive tree search on modern CPUs and GPUs. In SIGMOD. 339--350. https: \/\/doi.org\/10.1145\/1807167.1807206"},{"key":"e_1_3_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.14778\/1687553.1687564"},{"key":"e_1_3_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.14778\/3342263.3342276"},{"key":"e_1_3_2_2_56_1","volume-title":"Alexander L. Wolf, Paolo Costa, and Peter R. Pietzuch.","author":"Koliousis Alexandros","year":"2016","unstructured":"Alexandros Koliousis , Matthias Weidlich , Raul Castro Fernandez , Alexander L. Wolf, Paolo Costa, and Peter R. Pietzuch. 2016 . SABER : Window-Based Hybrid Stream Processing for Heterogeneous Architectures. In SIGMOD. ACM, New York, NY, USA , 555--569. https: \/\/doi.org\/10.1145\/2882903.2882906 Alexandros Koliousis, Matthias Weidlich, Raul Castro Fernandez, Alexander L. Wolf, Paolo Costa, and Peter R. Pietzuch. 2016. SABER: Window-Based Hybrid Stream Processing for Heterogeneous Architectures. In SIGMOD. ACM, New York, NY, USA, 555--569. https: \/\/doi.org\/10.1145\/2882903.2882906"},{"key":"e_1_3_2_2_57_1","volume-title":"SIGMOD","author":"Leis Viktor","unstructured":"Viktor Leis , Peter A. Boncz , Alfons Kemper , and Thomas Neumann . 2014. Morsel-driven parallelism: a NUMA-aware query evaluation framework for the many-core age . In SIGMOD . ACM , New York, NY, USA , 743--754. https:\/\/doi.org\/10.1145\/2588555.2610507 Viktor Leis, Peter A. Boncz, Alfons Kemper, and Thomas Neumann. 2014. Morsel-driven parallelism: a NUMA-aware query evaluation framework for the many-core age. In SIGMOD. ACM, New York, NY, USA, 743--754. https:\/\/doi.org\/10.1145\/2588555.2610507"},{"key":"e_1_3_2_2_58_1","volume-title":"Jieyang Chen, Jiajia Li, Xu Liu, Nathan R. Tallent, and Kevin J. Barker.","author":"Li Ang","year":"2019","unstructured":"Ang Li , Shuaiwen Leon Song , Jieyang Chen, Jiajia Li, Xu Liu, Nathan R. Tallent, and Kevin J. Barker. 2019 . Evaluating Modern GPU Interconnect: PC Ie , NVLink, NV-SLI, NVSwitch and GPUDirect. CoRR abs\/1903.04611 (2019). arXiv:1903.04611 Ang Li, Shuaiwen Leon Song, Jieyang Chen, Jiajia Li, Xu Liu, Nathan R. Tallent, and Kevin J. Barker. 2019. Evaluating Modern GPU Interconnect: PCIe, NVLink, NV-SLI, NVSwitch and GPUDirect. CoRR abs\/1903.04611 (2019). arXiv:1903.04611"},{"key":"e_1_3_2_2_59_1","volume-title":"Jieyang Chen, Xu Liu, Nathan R. Tallent, and Kevin J. Barker.","author":"Li Ang","year":"2018","unstructured":"Ang Li , Shuaiwen Leon Song , Jieyang Chen, Xu Liu, Nathan R. Tallent, and Kevin J. Barker. 2018 . Tartan : Evaluating Modern GPU Interconnect via a Multi-GPU Benchmark Suite. In IISWC. IEEE, New York, NY, USA , 191--202. https:\/\/doi.org\/10.1109\/IISWC.2018.8573483 Ang Li, Shuaiwen Leon Song, Jieyang Chen, Xu Liu, Nathan R. Tallent, and Kevin J. Barker. 2018. Tartan: Evaluating Modern GPU Interconnect via a Multi-GPU Benchmark Suite. In IISWC. IEEE, New York, NY, USA, 191--202. https:\/\/doi.org\/10.1109\/IISWC.2018.8573483"},{"key":"e_1_3_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.14778\/3007328.3007331"},{"key":"e_1_3_2_2_61_1","volume-title":"Lohman","author":"Li Yinan","year":"2013","unstructured":"Yinan Li , Ippokratis Pandis , Ren\u00e9 M\u00fcller , Vijayshankar Raman , and Guy M . Lohman . 2013 . NUMA-aware algorithms: the case of data shuffling. In CIDR. Yinan Li, Ippokratis Pandis, Ren\u00e9 M\u00fcller, Vijayshankar Raman, and Guy M. Lohman. 2013. NUMA-aware algorithms: the case of data shuffling. In CIDR."},{"key":"e_1_3_2_2_62_1","volume-title":"ADMS","author":"Lisa Nusrat Jahan","unstructured":"Nusrat Jahan Lisa , Annett Ungeth\u00fcm , Dirk Habich , Wolfgang Lehner , Tuan D. A. Nguyen , and Akash Kumar . 2018. Column Scan Acceleration in Hybrid CPU-FPGA Systems . In ADMS . ACM , New York, NY, USA , 22--33. Nusrat Jahan Lisa, Annett Ungeth\u00fcm, Dirk Habich,Wolfgang Lehner, Tuan D. A. Nguyen, and Akash Kumar. 2018. Column Scan Acceleration in Hybrid CPU-FPGA Systems. In ADMS. ACM, New York, NY, USA, 22--33."},{"key":"e_1_3_2_2_63_1","volume-title":"ASPLOS","author":"Lustig Daniel","unstructured":"Daniel Lustig , Sameer Sahasrabuddhe , and Olivier Giroux . 2019. A Formal Analysis of the NVIDIA PTX Memory Consistency Model . In ASPLOS . ACM , New York, NY, USA , 257--270. https:\/\/doi.org\/10. 1145\/3297858.3304043 Daniel Lustig, Sameer Sahasrabuddhe, and Olivier Giroux. 2019. A Formal Analysis of the NVIDIA PTX Memory Consistency Model. In ASPLOS. ACM, New York, NY, USA, 257--270. https:\/\/doi.org\/10. 1145\/3297858.3304043"},{"key":"e_1_3_2_2_64_1","first-page":"1","article-title":"Efficient k-means on GPUs. In DaMoN. ACM, New York","volume":"3","author":"Clemens Lutz","year":"2018","unstructured":"Clemens Lutz et al. 2018 . Efficient k-means on GPUs. In DaMoN. ACM, New York , NY, USA , 3 : 1 -- 3 :3. https:\/\/doi.org\/10.1145\/3211922.3211925 Clemens Lutz et al. 2018. Efficient k-means on GPUs. In DaMoN. ACM, New York, NY, USA, 3:1--3:3. https:\/\/doi.org\/10.1145\/3211922.3211925","journal-title":"NY, USA"},{"key":"e_1_3_2_2_65_1","doi-asserted-by":"publisher","DOI":"10.1007\/s13222-018-0293-x"},{"key":"e_1_3_2_2_66_1","doi-asserted-by":"publisher","DOI":"10.14778\/3236187.3236188"},{"key":"e_1_3_2_2_67_1","volume-title":"GPU Database Market. Retrieved","author":"Research MarketsandMarkets","year":"2019","unstructured":"MarketsandMarkets Research . 2018. GPU Database Market. Retrieved Oct 1, 2019 from https:\/\/www.marketsandmarkets.com\/ Market-Reports\/ gpu-database-market-259046335.html MarketsandMarkets Research. 2018. GPU Database Market. Retrieved Oct 1, 2019 from https:\/\/www.marketsandmarkets.com\/ Market-Reports\/gpu-database-market-259046335.html"},{"key":"e_1_3_2_2_68_1","volume-title":"andWolfgang Rehm","author":"Mietke Frank","year":"2006","unstructured":"Frank Mietke , Robert Rex , Robert Baumgartl , Torsten Mehlan , Torsten Hoefler , andWolfgang Rehm . 2006 . Analysis of the Memory Registration Process in the Mellanox InfiniBand Software Stack. In Euro-Par . 124--133. https:\/\/doi.org\/10.1007\/11823285_13 Frank Mietke, Robert Rex, Robert Baumgartl, Torsten Mehlan, Torsten Hoefler, andWolfgang Rehm. 2006. Analysis of the Memory Registration Process in the Mellanox InfiniBand Software Stack. In Euro-Par. 124--133. https:\/\/doi.org\/10.1007\/11823285_13"},{"key":"e_1_3_2_2_69_1","volume-title":"Research 18: Main Memory Databases and Modern Hardware SIGMOD '20, June 14--19","author":"Negrut Dan","year":"2014","unstructured":"Dan Negrut , Radu Serban , Ang Li , and Andrew Seidl . 2014 . Unified memory in CUDA 6.0: a brief overview of related data access and transfer issues. University of Wisconsin-Madison. TR-2014--09 . Research 18: Main Memory Databases and Modern Hardware SIGMOD '20, June 14--19 , 2020, Portland, OR, USA 1648 Dan Negrut, Radu Serban, Ang Li, and Andrew Seidl. 2014. Unified memory in CUDA 6.0: a brief overview of related data access and transfer issues. University of Wisconsin-Madison. TR-2014--09. Research 18: Main Memory Databases and Modern Hardware SIGMOD '20, June 14--19, 2020, Portland, OR, USA 1648"},{"key":"e_1_3_2_2_70_1","volume-title":"Yury Audzevich, Sergio L\u00f3pez-Buedo, and Andrew W. Moore.","author":"Neugebauer Rolf","year":"2018","unstructured":"Rolf Neugebauer , Gianni Antichi , Jos\u00e9 Fernando Zazo , Yury Audzevich, Sergio L\u00f3pez-Buedo, and Andrew W. Moore. 2018 . Understanding PCIe performance for end host networking. In SIGCOMM. ACM, New York, NY, USA , 327--341. https:\/\/doi.org\/10.1145\/3230543. 3230560 Rolf Neugebauer, Gianni Antichi, Jos\u00e9 Fernando Zazo, Yury Audzevich, Sergio L\u00f3pez-Buedo, and Andrew W. Moore. 2018. Understanding PCIe performance for end host networking. In SIGCOMM. ACM, New York, NY, USA, 327--341. https:\/\/doi.org\/10.1145\/3230543. 3230560"},{"key":"e_1_3_2_2_71_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-018-09679-z"},{"key":"e_1_3_2_2_72_1","unstructured":"Nvidia 2016. Nvidia Tesla P100. Nvidia. https:\/\/images.nvidia.com\/ content\/pdf\/tesla\/whitepaper\/pascal-architecture-whitepaper.pdf WP-08019-001_v01.1.  Nvidia 2016. Nvidia Tesla P100. Nvidia. https:\/\/images.nvidia.com\/ content\/pdf\/tesla\/whitepaper\/pascal-architecture-whitepaper.pdf WP-08019-001_v01.1."},{"key":"e_1_3_2_2_73_1","unstructured":"Nvidia 2017. Nvidia Tesla V100 GPU Architecture. Nvidia. https:\/\/images.nvidia.com\/content\/volta-architecture\/pdf\/voltaarchitecture- whitepaper.pdf WP-08608-001_v1.1.  Nvidia 2017. Nvidia Tesla V100 GPU Architecture. Nvidia. https:\/\/images.nvidia.com\/content\/volta-architecture\/pdf\/voltaarchitecture- whitepaper.pdf WP-08608-001_v1.1."},{"key":"e_1_3_2_2_74_1","unstructured":"Nvidia 2018. CUDA C Programming Guide. Nvidia. http:\/\/docs.nvidia. com\/pdf\/CUDA_C_Programming_Guide.pdf PG-02829-001_v10.0.  Nvidia 2018. CUDA C Programming Guide. Nvidia. http:\/\/docs.nvidia. com\/pdf\/CUDA_C_Programming_Guide.pdf PG-02829-001_v10.0."},{"key":"e_1_3_2_2_75_1","unstructured":"Nvidia 2019. CUDA C Best Practices Guide. Nvidia. https:\/\/docs. nvidia.com\/cuda\/pdf\/CUDA_C_Best_Practices_Guide.pdf DG-05603- 001_v10.1.  Nvidia 2019. CUDA C Best Practices Guide. Nvidia. https:\/\/docs. nvidia.com\/cuda\/pdf\/CUDA_C_Best_Practices_Guide.pdf DG-05603- 001_v10.1."},{"key":"e_1_3_2_2_76_1","unstructured":"Nvidia 2019. Tuning CUDA Applications for Pascal. Nvidia. https:\/\/ docs.nvidia.com\/cuda\/pdf\/Pascal_Tuning_Guide.pdf DA-08134-001_- v10.1.  Nvidia 2019. Tuning CUDA Applications for Pascal. Nvidia. https:\/\/ docs.nvidia.com\/cuda\/pdf\/Pascal_Tuning_Guide.pdf DA-08134-001_- v10.1."},{"key":"e_1_3_2_2_77_1","doi-asserted-by":"publisher","DOI":"10.1109\/ReConFig.2011.4"},{"key":"e_1_3_2_2_78_1","doi-asserted-by":"publisher","DOI":"10.14778\/3357377.3357383"},{"key":"e_1_3_2_2_79_1","doi-asserted-by":"crossref","unstructured":"C. Pearson I. Chung Z. Sura W. Hwu and J. Xiong. 2018. NUMAaware Data-transfer Measurements for Power\/NVLink Multi-GPU Systems. In IWOPH. Springer Heidelberg Germany.  C. Pearson I. Chung Z. Sura W. Hwu and J. Xiong. 2018. NUMAaware Data-transfer Measurements for Power\/NVLink Multi-GPU Systems. In IWOPH. Springer Heidelberg Germany.","DOI":"10.1007\/978-3-030-02465-9_32"},{"key":"e_1_3_2_2_80_1","volume-title":"ICPE","author":"Pearson Carl","unstructured":"Carl Pearson , Abdul Dakkak , Sarah Hashash , Cheng Li , I- Hsin Chung , Jinjun Xiong , and Wen-Mei Hwu . 2019. Evaluating Characteristics of CUDA Communication Primitives on High-Bandwidth Interconnects . In ICPE . ACM , New York, NY, USA , 209--218. https:\/\/doi.org\/10.1145\/ 3297663.3310299 Carl Pearson, Abdul Dakkak, Sarah Hashash, Cheng Li, I-Hsin Chung, Jinjun Xiong, and Wen-Mei Hwu. 2019. Evaluating Characteristics of CUDA Communication Primitives on High-Bandwidth Interconnects. In ICPE. ACM, New York, NY, USA, 209--218. https:\/\/doi.org\/10.1145\/ 3297663.3310299"},{"key":"e_1_3_2_2_81_1","volume-title":"Kersten","author":"Pirk Holger","year":"2014","unstructured":"Holger Pirk , Stefan Manegold , and Martin L . Kersten . 2014 . Waste not. . . Efficient co-processing of relational data. In ICDE. IEEE, New York, NY, USA , 508--519. Holger Pirk, Stefan Manegold, and Martin L. Kersten. 2014. Waste not. . . Efficient co-processing of relational data. In ICDE. IEEE, New York, NY, USA, 508--519."},{"key":"e_1_3_2_2_82_1","unstructured":"Aunn Raza Periklis Chrysogelos Panagiotis Sioulas Vladimir Indjic Angelos-Christos G. Anadiotis and Anastasia Ailamaki. 2020. GPUaccelerated data management under the test of time. In CIDR.  Aunn Raza Periklis Chrysogelos Panagiotis Sioulas Vladimir Indjic Angelos-Christos G. Anadiotis and Anastasia Ailamaki. 2020. GPUaccelerated data management under the test of time. In CIDR."},{"key":"e_1_3_2_2_83_1","first-page":"1","article-title":"MapD: a GPU-powered big data analytics and visualization platform. In SIGGRAPH. ACM, New York","volume":"73","author":"Root Christopher","year":"2016","unstructured":"Christopher Root and Todd Mostak . 2016 . MapD: a GPU-powered big data analytics and visualization platform. In SIGGRAPH. ACM, New York , NY, USA , 73 : 1 -- 73 :2. https:\/\/doi.org\/10.1145\/2897839.2927468 Christopher Root and Todd Mostak. 2016. MapD: a GPU-powered big data analytics and visualization platform. In SIGGRAPH. ACM, New York, NY, USA, 73:1--73:2. https:\/\/doi.org\/10.1145\/2897839.2927468","journal-title":"NY, USA"},{"key":"e_1_3_2_2_84_1","volume-title":"DaMoN","author":"Rosenfeld Viktor","unstructured":"Viktor Rosenfeld , Sebastian Bre\u00df , Steffen Zeuch , Tilmann Rabl , and Volker Markl . 2019. Performance Analysis and Automatic Tuning of Hash Aggregation on GPUs . In DaMoN . ACM , New York, NY, USA , 8. https:\/\/doi.org\/10.1145\/3329785.3329922 Viktor Rosenfeld, Sebastian Bre\u00df, Steffen Zeuch, Tilmann Rabl, and Volker Markl. 2019. Performance Analysis and Automatic Tuning of Hash Aggregation on GPUs. In DaMoN. ACM, New York, NY, USA, 8. https:\/\/doi.org\/10.1145\/3329785.3329922"},{"key":"e_1_3_2_2_85_1","first-page":"1","article-title":"Faster across the PCIe bus: a GPU library for lightweight decompression: including support for patched compression schemes. In DaMoN. ACM, New York","volume":"8","author":"Rozenberg Eyal","year":"2017","unstructured":"Eyal Rozenberg and Peter A. Boncz . 2017 . Faster across the PCIe bus: a GPU library for lightweight decompression: including support for patched compression schemes. In DaMoN. ACM, New York , NY, USA , 8 : 1 -- 8 :5. https:\/\/doi.org\/10.1145\/3076113.3076122 Eyal Rozenberg and Peter A. Boncz. 2017. Faster across the PCIe bus: a GPU library for lightweight decompression: including support for patched compression schemes. In DaMoN. ACM, New York, NY, USA, 8:1--8:5. https:\/\/doi.org\/10.1145\/3076113.3076122","journal-title":"NY, USA"},{"key":"e_1_3_2_2_86_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2882917"},{"key":"e_1_3_2_2_87_1","unstructured":"Amirhesam Shahvarani and Hans-Arno Jacobsen. 2016. A Hybrid B+- tree as Solution for In-Memory Indexing on CPU-GPU Heterogeneous Computing Platforms. In SIGMOD. 1523--1538. https:\/\/doi.org\/10. 1145\/2882903.2882918  Amirhesam Shahvarani and Hans-Arno Jacobsen. 2016. A Hybrid B+- tree as Solution for In-Memory Indexing on CPU-GPU Heterogeneous Computing Platforms. In SIGMOD. 1523--1538. https:\/\/doi.org\/10. 1145\/2882903.2882918"},{"key":"e_1_3_2_2_88_1","volume-title":"Which GPU database is right for me? Retrieved","author":"Shimoni Arnon","year":"2019","unstructured":"Arnon Shimoni . 2017. Which GPU database is right for me? Retrieved Oct 1, 2019 from https:\/\/hackernoon.com\/which-gpu-database-isright- for-me-6ceef6a17505 Arnon Shimoni. 2017. Which GPU database is right for me? Retrieved Oct 1, 2019 from https:\/\/hackernoon.com\/which-gpu-database-isright- for-me-6ceef6a17505"},{"key":"e_1_3_2_2_89_1","volume-title":"Hardware-conscious Hash-Joins on GPUs","author":"Sioulas Panagiotis","unstructured":"Panagiotis Sioulas , Periklis Chrysogelos , Manos Karpathiotakis , Raja Appuswamy , and Anastasia Ailamaki . 2019. Hardware-conscious Hash-Joins on GPUs . In ICDE. IEEE , New York, NY, USA . Panagiotis Sioulas, Periklis Chrysogelos, Manos Karpathiotakis, Raja Appuswamy, and Anastasia Ailamaki. 2019. Hardware-conscious Hash-Joins on GPUs. In ICDE. IEEE, New York, NY, USA."},{"key":"e_1_3_2_2_90_1","volume-title":"SIGMOD","author":"Stehle Elias","unstructured":"Elias Stehle and Hans-Arno Jacobsen . 2017. A Memory Bandwidth- Efficient Hybrid Radix Sort on GPUs . In SIGMOD . ACM , New York, NY, USA , 417--432. Elias Stehle and Hans-Arno Jacobsen. 2017. A Memory Bandwidth- Efficient Hybrid Radix Sort on GPUs. In SIGMOD. ACM, New York, NY, USA, 417--432."},{"key":"e_1_3_2_2_91_1","unstructured":"Nathan R. Tallent Nitin A. Gawande Charles Siegel Abhinav Vishnu and Adolfy Hoisie. 2017. Evaluating On-Node GPU Interconnects for Deep Learning Workloads. In PMBS@SC. 3--21. https:\/\/doi.org\/10. 1007\/978--3--319--72971--8_1  Nathan R. Tallent Nitin A. Gawande Charles Siegel Abhinav Vishnu and Adolfy Hoisie. 2017. Evaluating On-Node GPU Interconnects for Deep Learning Workloads. In PMBS@SC. 3--21. https:\/\/doi.org\/10. 1007\/978--3--319--72971--8_1"},{"key":"e_1_3_2_2_92_1","volume-title":"Retrieved","year":"2019","unstructured":"Top500. 2019 . Top500 Highlights . Retrieved Mar 16, 2020 from https:\/\/www.top500.org\/lists\/2019\/11\/highs\/ Top500. 2019. Top500 Highlights. Retrieved Mar 16, 2020 from https:\/\/www.top500.org\/lists\/2019\/11\/highs\/"},{"key":"e_1_3_2_2_93_1","volume-title":"Gross","author":"Trivedi Animesh","year":"2015","unstructured":"Animesh Trivedi , Patrick Stuedi , Bernard Metzler , Clemens Lutz , Martin Schmatz , and Thomas R . Gross . 2015 . RStore: A Direct-Access DRAM-based Data Store. In ICDCS. 674--685. https:\/\/doi.org\/10.1109\/ ICDCS. 2015.74 Animesh Trivedi, Patrick Stuedi, Bernard Metzler, Clemens Lutz, Martin Schmatz, and Thomas R. Gross. 2015. RStore: A Direct-Access DRAM-based Data Store. In ICDCS. 674--685. https:\/\/doi.org\/10.1109\/ ICDCS.2015.74"},{"key":"e_1_3_2_2_95_1","volume-title":"Srihari Cadambi, and Sudhakar Yalamanchili.","author":"Wu Haicheng","year":"2012","unstructured":"Haicheng Wu , Gregory Frederick Diamos , Srihari Cadambi, and Sudhakar Yalamanchili. 2012 . KernelWeaver: Automatically Fusing Database Primitives for Efficient GPU Computation. In MICRO. IEEE\/ACM, New York, NY, USA , 107--118. https:\/\/doi.org\/10.1109\/MICRO.2012.19 Haicheng Wu, Gregory Frederick Diamos, Srihari Cadambi, and Sudhakar Yalamanchili. 2012. KernelWeaver: Automatically Fusing Database Primitives for Efficient GPU Computation. In MICRO. IEEE\/ACM, New York, NY, USA, 107--118. https:\/\/doi.org\/10.1109\/MICRO.2012.19"},{"key":"e_1_3_2_2_96_1","unstructured":"Xilinx 2017. Vivado Design Suite: AXI Reference Guide. Xilinx. https: \/\/www.xilinx.com\/support\/documentation\/ip_documentation\/axi_ ref_guide\/latest\/ug1037-vivado-axi-reference-guide.pdf UG1037 (v4.0).  Xilinx 2017. Vivado Design Suite: AXI Reference Guide. Xilinx. https: \/\/www.xilinx.com\/support\/documentation\/ip_documentation\/axi_ ref_guide\/latest\/ug1037-vivado-axi-reference-guide.pdf UG1037 (v4.0)."},{"key":"e_1_3_2_2_97_1","unstructured":"Rengan Xu Frank Han and Quy Ta. 2018. Deep Learning at Scale on Nvidia V100 Accelerators. In PMBS@SC. 23--32. https:\/\/doi.org\/10. 1109\/PMBS.2018.8641600  Rengan Xu Frank Han and Quy Ta. 2018. Deep Learning at Scale on Nvidia V100 Accelerators. In PMBS@SC. 23--32. https:\/\/doi.org\/10. 1109\/PMBS.2018.8641600"},{"key":"e_1_3_2_2_98_1","volume-title":"Harmonia: A high throughput B+tree for GPUs. In PPoPP. 133--144. https:\/\/doi.org\/10.1145\/3293883.3295704","author":"Yan Zhaofeng","year":"2019","unstructured":"Zhaofeng Yan , Yuzhe Lin , Lu Peng , and Weihua Zhang . 2019 . Harmonia: A high throughput B+tree for GPUs. In PPoPP. 133--144. https:\/\/doi.org\/10.1145\/3293883.3295704 Zhaofeng Yan, Yuzhe Lin, Lu Peng, and Weihua Zhang. 2019. Harmonia: A high throughput B+tree for GPUs. In PPoPP. 133--144. https:\/\/doi.org\/10.1145\/3293883.3295704"},{"key":"e_1_3_2_2_99_1","volume-title":"DaMoN","author":"Yang Ke","unstructured":"Ke Yang , Bingsheng He , Rui Fang , Mian Lu , Naga K. Govindaraju , Qiong Luo , Pedro V. Sander , and Jiaoying Shi . 2007. In-memory grid files on graphics processors . In DaMoN . ACM , New York, NY, USA , 5. https:\/\/doi.org\/10.1145\/1363189.1363196 Ke Yang, Bingsheng He, Rui Fang, Mian Lu, Naga K. Govindaraju, Qiong Luo, Pedro V. Sander, and Jiaoying Shi. 2007. In-memory grid files on graphics processors. In DaMoN. ACM, New York, NY, USA, 5. https:\/\/doi.org\/10.1145\/1363189.1363196"},{"key":"e_1_3_2_2_100_1","doi-asserted-by":"publisher","DOI":"10.14778\/2536206.2536210"},{"key":"e_1_3_2_2_101_1","volume-title":"ISCA","author":"Zhao Xia","unstructured":"Xia Zhao , Almutaz Adileh , Zhibin Yu , Zhiying Wang , Aamer Jaleel , and Lieven Eeckhout . 2019. Adaptive memory-side last-level GPU caching . In ISCA . ISCA , Winona, MN, USA , 411--423. https:\/\/doi.org\/ 10.1145\/3307650.3322235 Xia Zhao, Almutaz Adileh, Zhibin Yu, Zhiying Wang, Aamer Jaleel, and Lieven Eeckhout. 2019. Adaptive memory-side last-level GPU caching. In ISCA. ISCA, Winona, MN, USA, 411--423. https:\/\/doi.org\/ 10.1145\/3307650.3322235"},{"key":"e_1_3_2_2_102_1","volume-title":"Keckler","author":"Zheng Tianhao","year":"2016","unstructured":"Tianhao Zheng , David W. Nellans , Arslan Zulfiqar , Mark Stephenson , and Stephen W . Keckler . 2016 . Towards high performance paged memory for GPUs. In HPCA. IEEE, New York, NY, USA , 345--357. https:\/\/doi.org\/10.1109\/HPCA.2016.7446077 Tianhao Zheng, DavidW. Nellans, Arslan Zulfiqar, Mark Stephenson, and Stephen W. Keckler. 2016. Towards high performance paged memory for GPUs. In HPCA. IEEE, New York, NY, USA, 345--357. https:\/\/doi.org\/10.1109\/HPCA.2016.7446077"}],"event":{"name":"SIGMOD\/PODS '20: International Conference on Management of Data","location":"Portland OR USA","acronym":"SIGMOD\/PODS '20","sponsor":["SIGMOD ACM Special Interest Group on Management of Data"]},"container-title":["Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3318464.3389705","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3318464.3389705","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:38:44Z","timestamp":1750199924000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3318464.3389705"}},"subtitle":["Processing Large Data on GPUs with Fast Interconnects"],"short-title":[],"issued":{"date-parts":[[2020,5,31]]},"references-count":101,"alternative-id":["10.1145\/3318464.3389705","10.1145\/3318464"],"URL":"https:\/\/doi.org\/10.1145\/3318464.3389705","relation":{},"subject":[],"published":{"date-parts":[[2020,5,31]]},"assertion":[{"value":"2020-05-31","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}