{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,17]],"date-time":"2025-09-17T03:18:47Z","timestamp":1758079127024,"version":"3.44.0"},"reference-count":39,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2025,8]]},"abstract":"<jats:p>Buffer management is critical for DBMSs but often suffers from scalability bottlenecks and poor cache locality, which stems from centralized reference counting in page access and intensive locking in page-to-buffer translation. However, prior radical approaches like pointer swizzling or optimistic lock can hardly be adopted in production-grade DBMSs due to its inherent complexity and incompatibility.<\/jats:p>\n          <jats:p>\n            This paper proposes\n            <jats:italic toggle=\"yes\">ScaleCache<\/jats:italic>\n            , a scalable, highly-efficient and production-grade buffer management system with three key designs.\n            <jats:italic toggle=\"yes\">ScaleCache<\/jats:italic>\n            first incorporates a novel compact per-group buffer reference counting technique, which enables scalable buffer pinning and unpinning by concurrent threads on many-core servers. It then devised a novel read-write lock based on copy-on-write and per-group reference counting, which is suitable for B-link tree. At last, we propose an optimistic, CPU-cache friendly and SIMD-accelerated hash table for fast and scalable page-to-buffer translation, which eliminates most contention on modern many-core hardware. ScaleCache has been adopted in\n            <jats:italic toggle=\"yes\">Huawei GaussDB<\/jats:italic>\n            , a commercial high-performance DBMS. Evaluation on a 128-core server demonstrates that\n            <jats:italic toggle=\"yes\">ScaleCache<\/jats:italic>\n            exhibits near-linear scalability and can significantly improve index query throughput of both classic B-link tree index and complex graph-based vector index.\n          <\/jats:p>","DOI":"10.14778\/3750601.3750628","type":"journal-article","created":{"date-parts":[[2025,9,16]],"date-time":"2025-09-16T13:38:05Z","timestamp":1758029885000},"page":"5073-5085","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["<i>ScaleCache<\/i>\n            : Scalable and Production-Grade Buffer Management for Disk-Based Database Systems"],"prefix":"10.14778","volume":"18","author":[{"given":"Mingyu","family":"Liu","sequence":"first","affiliation":[{"name":"Huawei Company, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Junbin","family":"Kang","sequence":"additional","affiliation":[{"name":"Huawei Company, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kai","family":"Wang","sequence":"additional","affiliation":[{"name":"Huawei Company, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lu","family":"Zhang","sequence":"additional","affiliation":[{"name":"Huawei Company, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haibo","family":"Chen","sequence":"additional","affiliation":[{"name":"Huawei Company, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiuchang","family":"Li","sequence":"additional","affiliation":[{"name":"Huawei Company, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tianhong","family":"Ding","sequence":"additional","affiliation":[{"name":"Huawei Company, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,9,16]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.14778\/3611479.3611485"},{"key":"e_1_2_1_2_1","volume-title":"9th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2010, October 4\u20136, 2010, Vancouver, BC, Canada, Proceedings. 1\u201316","author":"Boyd-Wickizer Silas","year":"2010","unstructured":"Silas Boyd-Wickizer, Austin T. Clements, Yandong Mao, Aleksey Pesterev, M. Frans Kaashoek, Robert Tappan Morris, and Nickolai Zeldovich. 2010. An Analysis of Linux Scalability to Many Cores. In 9th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2010, October 4\u20136, 2010, Vancouver, BC, Canada, Proceedings. 1\u201316."},{"key":"e_1_2_1_3_1","volume-title":"Cache-Conscious Concurrency Control of Main-Memory Indexes on Shared-Memory Multiprocessor Systems. In VLDB 2001, Proceedings of 27th International Conference on Very Large Data Bases, September 11\u201314","author":"Cha Sang Kyun","year":"2001","unstructured":"Sang Kyun Cha, Sangyong Hwang, Kihong Kim, and Keunjoo Kwon. [n.d.]. Cache-Conscious Concurrency Control of Main-Memory Indexes on Shared-Memory Multiprocessor Systems. In VLDB 2001, Proceedings of 27th International Conference on Very Large Data Bases, September 11\u201314, 2001, Roma, Italy. 181\u2013190."},{"key":"e_1_2_1_4_1","volume-title":"Eighth Eurosys Conference 2013","author":"Clements Austin T.","year":"2013","unstructured":"Austin T. Clements, M. Frans Kaashoek, and Nickolai Zeldovich. 2013. RadixVM: scalable address spaces for multithreaded applications. In Eighth Eurosys Conference 2013, EuroSys '13, Prague, Czech Republic, April 14\u201317, 2013. 211\u2013224."},{"key":"e_1_2_1_5_1","unstructured":"CMU. [n.d.]. BenchBase (formerly OLTPBench) is a Multi-DBMS SQL Benchmarking Framework via JDBC. https:\/\/github.com\/cmu-db\/benchbase"},{"key":"e_1_2_1_6_1","unstructured":"Jonathan Corbet. 2006. The search for fast scalable counters. https:\/\/lwn.net\/Articles\/170003\/"},{"key":"e_1_2_1_7_1","unstructured":"Jonathan Corbet. 2013. Per-CPU reference counts. https:\/\/lwn.net\/Articles\/557478\/"},{"key":"e_1_2_1_8_1","volume-title":"CLoF: A Compositional Lock Framework for Multi-level NUMA Systems. In SOSP '21: ACM SIGOPS 28th Symposium on Operating Systems Principles, Virtual Event \/ Koblenz, Germany, October 26\u201329","author":"de Lima Chehab Rafael Lourenco","year":"2021","unstructured":"Rafael Lourenco de Lima Chehab, Antonio Paolillo, Diogo Behrens, Ming Fu, Hermann H\u00e4rtig, and Haibo Chen. 2021. CLoF: A Compositional Lock Framework for Multi-level NUMA Systems. In SOSP '21: ACM SIGOPS 28th Symposium on Operating Systems Principles, Virtual Event \/ Koblenz, Germany, October 26\u201329, 2021. ACM, 851\u2013865."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.14778\/2556549.2556575"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2463676.2463710"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.14778\/2732240.2732246"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.14778\/2732967.2732968"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.14778\/2735461.2735465"},{"key":"e_1_2_1_14_1","volume-title":"Adaptive Factorization Using Linear-Chained Hash Tables. In CIDR","author":"Gro\u00df Paul","year":"2025","unstructured":"Paul Gro\u00df, Daniel ten Wolde, and Peter Boncz. 2025. Adaptive Factorization Using Linear-Chained Hash Tables. In CIDR 2025."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1376616.1376713"},{"key":"e_1_2_1_16_1","unstructured":"HUAWEI. [n.d.]. Kunpeng 920-6426 - HiSilicon. https:\/\/en.wikichip.org\/wiki\/hisilicon\/kunpeng\/920-6426"},{"key":"e_1_2_1_17_1","volume-title":"17th USENIX Conference on File and Storage Technologies, FAST 2019","author":"Jung Seokyong","year":"2019","unstructured":"Seokyong Jung, Jong-Bin Kim, Minsoo Ryu, Sooyong Kang, and Hyungsoo Jung. 2019. Pay Migration Tax to Homeland: Anchor-based Scalable Reference Counting for Multicores. In 17th USENIX Conference on File and Storage Technologies, FAST 2019, Boston, MA, February 25\u201328, 2019. USENIX Association, 79\u201391."},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the 27th International Conference on Data Engineering, ICDE 2011, April 11\u201316","author":"Kemper Alfons","year":"2011","unstructured":"Alfons Kemper and Thomas Neumann. 2011. HyPer: A hybrid OLTP&OLAP main memory database system based on virtual memory snapshots. In Proceedings of the 27th International Conference on Data Engineering, ICDE 2011, April 11\u201316, 2011, Hannover, Germany. IEEE Computer Society, 195\u2013206."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/319628.319663"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3588687"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2018.00026"},{"key":"e_1_2_1_22_1","first-page":"73","article-title":"Optimistic Lock Coupling: A Scalable and Efficient General-Purpose Synchronization Method","volume":"42","author":"Leis Viktor","year":"2019","unstructured":"Viktor Leis, Michael Haubenschild, and Thomas Neumann. 2019. Optimistic Lock Coupling: A Scalable and Efficient General-Purpose Synchronization Method. IEEE Data Eng. Bull. 42, 1 (2019), 73\u201384.","journal-title":"IEEE Data Eng. Bull."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2933349.2933352"},{"key":"e_1_2_1_24_1","volume-title":"The RCU API","author":"McKenney Paul","year":"2019","unstructured":"Paul McKenney. 2019. The RCU API, 2019 edition. https:\/\/lwn.net\/Articles\/777036\/"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3421473.3421481"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447786.3456242"},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of the 2016 USENIX Annual Technical Conference, USENIX ATC 2016","author":"Papagiannis Anastasios","year":"2016","unstructured":"Anastasios Papagiannis, Giorgos Saloustros, Pilar Gonz\u00e1lez-F\u00e9rez, and Angelos Bilas. 2016. Tucana: Design and Implementation of a Fast and Efficient Scale-up Key-value Store. In Proceedings of the 2016 USENIX Annual Technical Conference, USENIX ATC 2016, Denver, CO, USA, June 22\u201324, 2016. USENIX Association, 537\u2013550."},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of the ACM Symposium on Cloud Computing, SoCC 2018","author":"Papagiannis Anastasios","year":"2018","unstructured":"Anastasios Papagiannis, Giorgos Saloustros, Pilar Gonz\u00e1lez-F\u00e9rez, and Angelos Bilas. 2018. An Efficient Memory-Mapped Key-Value Store for Flash Storage. In Proceedings of the ACM Symposium on Cloud Computing, SoCC 2018, Carlsbad, CA, USA, October 11\u201313, 2018. ACM, 490\u2013502."},{"key":"e_1_2_1_29_1","article-title":"Kreon: An Efficient Memory-Mapped Key-Value Store for Flash Storage","volume":"17","author":"Papagiannis Anastasios","year":"2021","unstructured":"Anastasios Papagiannis, Giorgos Saloustros, Giorgos Xanthakis, Giorgos Kalaentzis, Pilar Gonz\u00e1lez-F\u00e9rez, and Angelos Bilas. 2021. Kreon: An Efficient Memory-Mapped Key-Value Store for Flash Storage. ACM Trans. Storage 17, 1 (2021), 7:1\u20137:32.","journal-title":"ACM Trans. Storage"},{"key":"e_1_2_1_30_1","volume-title":"Proceedings of the 2020 USENIX Annual Technical Conference, USENIX ATC 2020, July 15\u201317","author":"Papagiannis Anastasios","year":"2020","unstructured":"Anastasios Papagiannis, Giorgos Xanthakis, Giorgos Saloustros, Manolis Marazakis, and Angelos Bilas. 2020. Optimizing Memory-mapped I\/O for Fast Storage Devices. In Proceedings of the 2020 USENIX Annual Technical Conference, USENIX ATC 2020, July 15\u201317, 2020. USENIX Association, 813\u2013827."},{"key":"e_1_2_1_31_1","unstructured":"Andrew Pavlo. 2014. On Scalable Transaction Execution in Partitioned Main Memory Database Management Systems. Ph.D. Dissertation. Brown University USA."},{"key":"e_1_2_1_32_1","volume-title":"11th USENIX Symposium on Operating Systems Design and Implementation, OSDI '14","author":"Peter Simon","year":"2014","unstructured":"Simon Peter, Jialin Li, Irene Zhang, Dan R. K. Ports, Doug Woos, Arvind Krishnamurthy, Thomas E. Anderson, and Timothy Roscoe. 2014. Arrakis: The Operating System is the Control Plane. In 11th USENIX Symposium on Operating Systems Design and Implementation, OSDI '14, Broomfield, CO, USA, October 6\u20138, 2014. USENIX Association, 1\u201316."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/2213836.2213946"},{"key":"e_1_2_1_34_1","volume-title":"9th USENIX Symposium on Operating Systems Design and Implementation, OSDI","author":"Soares Livio","year":"2010","unstructured":"Livio Soares and Michael Stumm. 2010. FlexSC:Flexible System Call Scheduling with Exception-Less System Calls. In 9th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2010, October 4\u20136, 2010, Vancouver, BC, Canada, Proceedings. USENIX Association, 33\u201346."},{"key":"e_1_2_1_35_1","volume-title":"NeurIPS","author":"Subramanya Suhas Jayaram","year":"2019","unstructured":"Suhas Jayaram Subramanya, Devvrit, Rohan Kadekodi, Ravishankar Krishnaswamy, and Harsha Simhadri. 2019. DiskANN: Fast Accurate Billion-point Nearest Neighbor Search on a Single Node. In NeurIPS 2019."},{"key":"e_1_2_1_36_1","unstructured":"TPCH. [n.d.]. TPC-H is a Decision Support Benchmark. http:\/\/www.tpc.org\/tpch\/"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2517349.2522713"},{"key":"e_1_2_1_38_1","volume-title":"Fast Databases with Fast Durability and Recovery Through Multicore Parallelism. In 11th USENIX Symposium on Operating Systems Design and Implementation, OSDI '14","author":"Zheng Wenting","year":"2014","unstructured":"Wenting Zheng, Stephen Tu, Eddie Kohler, and Barbara Liskov. 2014. Fast Databases with Fast Durability and Recovery Through Multicore Parallelism. In 11th USENIX Symposium on Operating Systems Design and Implementation, OSDI '14, Broomfield, CO, USA, October 6\u20138, 2014. USENIX Association, 465\u2013477."},{"key":"e_1_2_1_39_1","volume-title":"Userspace Bypass: Accelerating Syscall-intensive Applications. In 17th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2023","author":"Zhou Zhe","year":"2023","unstructured":"Zhe Zhou, Yanxiang Bi, Junpeng Wan, Yangfan Zhou, and Zhou Li. 2023. Userspace Bypass: Accelerating Syscall-intensive Applications. In 17th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2023, Boston, MA, USA, July 10\u201312, 2023. USENIX Association, 33\u201349."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3750601.3750628","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,16]],"date-time":"2025-09-16T13:42:22Z","timestamp":1758030142000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3750601.3750628"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8]]},"references-count":39,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2025,8]]}},"alternative-id":["10.14778\/3750601.3750628"],"URL":"https:\/\/doi.org\/10.14778\/3750601.3750628","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2025,8]]},"assertion":[{"value":"2025-09-16","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}