{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T02:48:33Z","timestamp":1783738113655,"version":"3.55.0"},"reference-count":62,"publisher":"Association for Computing Machinery (ACM)","issue":"13","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2024,9]]},"abstract":"<jats:p>GPU-accelerated databases have been gaining popularity in recent years due to their massive parallelism and high memory bandwidth. The limited GPU memory capacity, however, is still a major bottleneck for GPU databases.<\/jats:p>\n          <jats:p>\n            Existing approaches have attempted to address this limitation by using (1) hybrid CPU-GPU DBMS or (2) multi-GPU DBMS. We aim to improve prior solutions further by leveraging both hybrid CPU-GPU DBMS and multi-GPU DBMS at the same time. In particular, we explore the design space and optimize the\n            <jats:italic>data placement<\/jats:italic>\n            and\n            <jats:italic>query execution<\/jats:italic>\n            in hybrid CPU and multi-GPU DBMS. To improve data placement, we introduce the\n            <jats:italic>cache-aware replication policy<\/jats:italic>\n            which takes into account the cost of shuffle when replicating data and could coordinate both caching and replication decisions for the best performance. To improve query execution, we extend the existing hybrid CPU-GPU query execution strategy with distributed query processing techniques to support multiple GPUs. We build a system called\n            <jats:italic>Lancelot<\/jats:italic>\n            , a hybrid CPU and Multi-GPU data analytics engine with all the optimizations integrated.\n          <\/jats:p>\n          <jats:p>\n            Our evaluation shows that the\n            <jats:italic>cache-aware replication<\/jats:italic>\n            outperforms other policies by up to 2.5\u00d7 and Lancelot outperforms existing GPU DBMSes by at least 2\u00d7 on Star Schema Benchmark and 12\u00d7 on TPC-H Benchmark.\n          <\/jats:p>","DOI":"10.14778\/3704965.3704977","type":"journal-article","created":{"date-parts":[[2025,2,18]],"date-time":"2025-02-18T17:22:57Z","timestamp":1739899377000},"page":"4709-4722","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Scaling your Hybrid CPU-GPU DBMS to Multiple GPUs"],"prefix":"10.14778","volume":"17","author":[{"given":"Bobbi","family":"Yogatama","sequence":"first","affiliation":[{"name":"University of Wisconsin-Madison"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Weiwei","family":"Gong","sequence":"additional","affiliation":[{"name":"Oracle Corporation"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiangyao","family":"Yu","sequence":"additional","affiliation":[{"name":"University of Wisconsin-Madison"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,2,18]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2024. AMD Instinct MI300 Series Accelerators. https:\/\/www.amd.com\/en\/products\/accelerators\/instinct\/mi300.html."},{"key":"e_1_2_1_2_1","unstructured":"2024. BlazingSQL. https:\/\/blazingsql.com."},{"key":"e_1_2_1_3_1","unstructured":"2024. CUDA C Programming Guide. http:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide\/index.html."},{"key":"e_1_2_1_4_1","unstructured":"2024. cuDF-Performance Comparison. https:\/\/github.com\/rapidsai\/cudf\/blob\/branch-23.04\/docs\/cudf\/source\/user_guide\/performance_comparisons.ipynb."},{"key":"e_1_2_1_5_1","unstructured":"2024. Dask-CUDA. https:\/\/docs.rapids.ai\/api\/dask-cuda\/nightly\/."},{"key":"e_1_2_1_6_1","unstructured":"2024. HeavyAI. https:\/\/www.heavy.ai\/."},{"key":"e_1_2_1_7_1","unstructured":"2024. HIP Programming Guide. https:\/\/github.com\/ROCm-Developer-Tools\/HIP."},{"key":"e_1_2_1_8_1","unstructured":"2024. Kinetica. https:\/\/kinetica.com\/."},{"key":"e_1_2_1_9_1","unstructured":"2024. NVIDIA Collective Communication Library. https:\/\/developer.nvidia.com\/nccl."},{"key":"e_1_2_1_10_1","unstructured":"2024. NVIDIA H100 Tensor Core GPU. https:\/\/www.nvidia.com\/en-us\/data-center\/h100\/."},{"key":"e_1_2_1_11_1","unstructured":"2024. NVLINK. https:\/\/www.nvidia.com\/en-us\/data-center\/nvlink\/."},{"key":"e_1_2_1_12_1","unstructured":"2024. Opencl. https:\/\/www.khronos.org\/opencl\/."},{"key":"e_1_2_1_13_1","unstructured":"2024. PG-Storm. https:\/\/github.com\/heterodb\/pg-strom."},{"key":"e_1_2_1_14_1","unstructured":"2024. RAPIDS. https:\/\/rapids.ai."},{"key":"e_1_2_1_15_1","unstructured":"2024. The RAPIDS Accelerator for Apache Spark. https:\/\/nvidia.github.io\/spark-rapids\/."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1007\/s13222-014-0164-z"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.is.2013.05.004"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2882936"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.14778\/2536274.2536325"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.datak.2014.07.003"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33074-2_5"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.14778\/3632093.3632107"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.14778\/3303753.3303760"},{"key":"e_1_2_1_24_1","volume-title":"9th Biennial Conference on Innovative Data Systems Research, CIDR","author":"Chrysogelos Periklis","year":"2019","unstructured":"Periklis Chrysogelos, Panagiotis Sioulas, and Anastasia Ailamaki. 2019. Hardware-conscious Query Processing in GPU-accelerated Analytical Engines. In 9th Biennial Conference on Innovative Data Systems Research, CIDR 2019, Asilomar, CA, USA, January 13-16, 2019, Online Proceedings. www.cidrdb.org. http:\/\/cidrdb.org\/cidr2019\/papers\/p127-chrysogelos-cidr19.pdf"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3183734"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.14778\/3380750.3380758"},{"key":"e_1_2_1_27_1","unstructured":"Hao Gao. 2021. Scaling Joins to a Thousand GPUs. In ADMS@VLDB. https:\/\/api.semanticscholar.org\/CorpusID:237250537"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/191839.191886"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1620585.1620588"},{"key":"e_1_2_1_30_1","unstructured":"Bingsheng He Ke Yang Rui Fang Mian Lu Naga Govindaraju Qiong Luo and Pedro Sander. 2008. Relational joins on graphics processors. In SIGMOD."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.14778\/3551793.3551833"},{"key":"e_1_2_1_32_1","volume-title":"Revisiting co-processing for hash joins on the coupled cpu-gpu architecture. PVLDB","author":"He Jiong","year":"2013","unstructured":"Jiong He, Mian Lu, and Bingsheng He. 2013. Revisiting co-processing for hash joins on the coupled cpu-gpu architecture. PVLDB (2013)."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.14778\/2735496.2735497"},{"key":"e_1_2_1_34_1","volume-title":"Hardware-oblivious parallelism for in-memory column-stores. PVLDB","author":"Heimel Max","year":"2013","unstructured":"Max Heimel, Michael Saecker, Holger Pirk, Stefan Manegold, and Volker Markl. 2013. Hardware-oblivious parallelism for in-memory column-stores. PVLDB (2013)."},{"key":"e_1_2_1_35_1","doi-asserted-by":"crossref","unstructured":"Tim Kaldewey Guy Lohman Rene Mueller and Peter Volk. 2012. GPU join processing revisited. In DaMoN.","DOI":"10.1145\/2236584.2236592"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.14778\/3067421.3067423"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.14778\/3007328.3007331"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3389705"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3517911"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2903735"},{"key":"e_1_2_1_41_1","unstructured":"Hamish Nicholson Aunn Raza Periklis Chrysogelos and Anastasia Ailamaki. 2023. HetCache: Synergising NVMe Storage and GPU acceleration for Memory-Efficient Analytics. (2023). https:\/\/www.cidrdb.org\/cidr2023\/papers\/p84-nicholson.pdf"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-10424-4_17"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.14778\/3425879.3425890"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3457254"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3320212"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.14778\/3436905.3436927"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3085504.3085521"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3526132"},{"key":"e_1_2_1_49_1","volume-title":"Proceedings of the 2020 International Conference on Management of Data. ACM.","author":"Shanbhag Anil","year":"2020","unstructured":"Anil Shanbhag, Xiangyao Yu, and Samuel Madden. 2020. A Study of the Fundamental Performance Charecteristics of GPUs and CPUs for Database Analytics. In Proceedings of the 2020 International Conference on Management of Data. ACM."},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485278.2485282"},{"key":"e_1_2_1_52_1","doi-asserted-by":"crossref","unstructured":"Elias Stehle and Hans-Arno Jacobsen. 2017. A Memory Bandwidth-Efficient Hybrid Radix Sort on GPUs. In SIGMOD. ACM.","DOI":"10.1145\/3035918.3064043"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3588709"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3315508.3329973"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/319732.319734"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/MASCOTS.2006.16"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.19"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2017.2677451"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/3592980.3595307"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.14778\/3551793.3551809"},{"key":"e_1_2_1_61_1","volume-title":"The Yin and Yang of processing data warehousing queries on GPU devices. PVLDB","author":"Yuan Yuan","year":"2013","unstructured":"Yuan Yuan, Rubao Lee, and Xiaodong Zhang. 2013. The Yin and Yang of processing data warehousing queries on GPU devices. PVLDB (2013)."},{"key":"e_1_2_1_62_1","volume-title":"Hetero-DB: Next Generation High-Performance Database Systems by Best Utilizing Heterogeneous Computing and Storage Resources. Journal of Computer Science and Technology 30","author":"Zhang Kai","year":"2015","unstructured":"Kai Zhang, Feng Chen, Xiaoning Ding, Yin Huai, Rubao Lee, Tian Luo, Kaibo Wang, Yuan Yuan, and Xiaodong Zhang. 2015. Hetero-DB: Next Generation High-Performance Database Systems by Best Utilizing Heterogeneous Computing and Storage Resources. Journal of Computer Science and Technology 30 (2015)."},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.14778\/2536274.2536319"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3704965.3704977","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,2,18]],"date-time":"2025-02-18T17:30:54Z","timestamp":1739899854000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3704965.3704977"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,9]]},"references-count":62,"journal-issue":{"issue":"13","published-print":{"date-parts":[[2024,9]]}},"alternative-id":["10.14778\/3704965.3704977"],"URL":"https:\/\/doi.org\/10.14778\/3704965.3704977","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2024,9]]},"assertion":[{"value":"2025-02-18","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}