{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,20]],"date-time":"2026-02-20T19:39:26Z","timestamp":1771616366845,"version":"3.50.1"},"reference-count":31,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2023,6,13]],"date-time":"2023-06-13T00:00:00Z","timestamp":1686614400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100000923","name":"Australian Research Council","doi-asserted-by":"publisher","award":["DP220101434; DE230100366"],"award-info":[{"award-number":["DP220101434; DE230100366"]}],"id":[{"id":"10.13039\/501100000923","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2023,6,13]]},"abstract":"<jats:p>Although many updatable learned indexes have been proposed in recent years, whether they can outperform traditional approaches on disk remains unknown. In this study, we revisit and implement four state-of-the-art updatable learned indexes on disk, and compare them against the B+-tree under a wide range of settings. Through our evaluation, we make some key observations: 1) Overall, the B+-tree performs well across a range of workload types and datasets. 2) A learned index could outperform B+-tree or other learned indexes on disk for a specific workload. For example, PGM achieves the best performance in write-only workloads while LIPP significantly outperforms others in lookup-only workloads. We further conduct a detailed performance analysis to reveal the strengths and weaknesses of these learned indexes on disk. Moreover, we summarize the observed common shortcomings in five categories and propose four design principles to guide future design of on-disk, updatable learned indexes: (1) reducing the index's tree height, (2) better data structures to lower operation overheads, (3) improving the efficiency of scan operations, and (4) more efficient storage layout.<\/jats:p>","DOI":"10.1145\/3589284","type":"journal-article","created":{"date-parts":[[2023,6,20]],"date-time":"2023-06-20T20:26:45Z","timestamp":1687292805000},"page":"1-22","source":"Crossref","is-referenced-by-count":17,"title":["Updatable Learned Indexes Meet Disk-Resident DBMS - From Evaluations to Design Choices"],"prefix":"10.1145","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-4433-9232","authenticated-orcid":false,"given":"Hai","family":"Lan","sequence":"first","affiliation":[{"name":"RMIT University, Melbourne, VIC, Australia"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2477-381X","authenticated-orcid":false,"given":"Zhifeng","family":"Bao","sequence":"additional","affiliation":[{"name":"RMIT University, Melbourne, VIC, Australia"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1902-9087","authenticated-orcid":false,"given":"J. Shane","family":"Culpepper","sequence":"additional","affiliation":[{"name":"RMIT University, Melbourne, VIC, Australia"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3503-4123","authenticated-orcid":false,"given":"Renata","family":"Borovica-Gajic","sequence":"additional","affiliation":[{"name":"University of Melbourne, Melbourne, VIC, Australia"}]}],"member":"320","published-online":{"date-parts":[[2023,6,20]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"[n. d.]. Source Code. https:\/\/github.com\/rmitbggroup\/LearnedIndexDiskExp."},{"key":"e_1_2_2_2_1","unstructured":"[n. d.]. STX B. https:\/\/panthema.net\/2007\/stx-btree\/."},{"key":"e_1_2_2_4_1","volume-title":"Lyric Doshi, Tim Kraska, Xiaozhou Li, Andy Ly, and Christopher Olston.","author":"Abu-Libdeh Hussam","year":"2020","unstructured":"Hussam Abu-Libdeh, Deniz Altinb\u00fcken, Alex Beutel, Ed H. Chi, Lyric Doshi, Tim Kraska, Xiaozhou Li, Andy Ly, and Christopher Olston. 2020. Learned Indexes for a Google-scale Disk-based Database. CoRR abs\/2012.12501 (2020)."},{"key":"e_1_2_2_5_1","volume-title":"Non-blocking interpolation search trees with doublylogarithmic running time. PPoPP","author":"Brown Trevor","year":"2020","unstructured":"Trevor Brown, Aleksandar Prokopec, and Dan Alistarh. 2020. Non-blocking interpolation search trees with doublylogarithmic running time. PPoPP (2020), 276--291."},{"key":"e_1_2_2_6_1","volume-title":"Arpaci-Dusseau","author":"Dai Yifan","year":"2020","unstructured":"Yifan Dai, Yien Xu, Aishwarya Ganesan, Ramnatthan Alagappan, Brian Kroth, Andrea C. Arpaci-Dusseau, and Remzi H. Arpaci-Dusseau. 2020. From WiscKey to Bourbon: A Learned Index for Log-Structured Merge Trees. In OSDI. 155--171."},{"key":"e_1_2_2_7_1","volume-title":"Jia Yu, Chi Wang, Jaeyoung Do, Yinan Li, Hantian Zhang, Badrish Chandramouli, Johannes Gehrke, Donald Kossmann, David B. Lomet, and Tim Kraska.","author":"Ding Jialin","year":"2020","unstructured":"Jialin Ding, Umar Farooq Minhas, Jia Yu, Chi Wang, Jaeyoung Do, Yinan Li, Hantian Zhang, Badrish Chandramouli, Johannes Gehrke, Donald Kossmann, David B. Lomet, and Tim Kraska. 2020. ALEX: An Updatable Adaptive Learned Index. In SIGMOD. 969--984."},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.14778\/3425879.3425880"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.14778\/3389133.3389135"},{"key":"e_1_2_2_10_1","doi-asserted-by":"crossref","unstructured":"Alex Galakatos Michael Markovitch Carsten Binnig Rodrigo Fonseca and Tim Kraska. 2019. FITing-Tree: A Dataaware Index Structure. In SIGMOD. 1189--1206.","DOI":"10.1145\/3299869.3319860"},{"key":"e_1_2_2_11_1","first-page":"1","article-title":"RadixSpline: a single-pass learned index. In aiDM@SIGMOD","volume":"5","author":"Kipf Andreas","year":"2020","unstructured":"Andreas Kipf, Ryan Marcus, Alexander van Renen, Mihail Stoian, Alfons Kemper, Tim Kraska, and Thomas Neumann. 2020. RadixSpline: a single-pass learned index. In aiDM@SIGMOD. ACM, 5:1--5:5.","journal-title":"ACM"},{"key":"e_1_2_2_12_1","volume-title":"Jeffrey Dean, and Neoklis Polyzotis.","author":"Kraska Tim","year":"2018","unstructured":"Tim Kraska, Alex Beutel, Ed H. Chi, Jeffrey Dean, and Neoklis Polyzotis. 2018. The Case for Learned Index Structures. In SIGMOD. 489--504."},{"key":"e_1_2_2_13_1","first-page":"3","article-title":"ML-In-Databases","volume":"44","author":"Kraska Tim","year":"2021","unstructured":"Tim Kraska, Umar Farooq Minhas, Thomas Neumann, Olga Papaemmanouil, Jignesh M. Patel, Christopher R\u00e9, and Michael Stonebraker. 2021. ML-In-Databases: Assessment and Prognosis. IEEE Data Eng. Bull. 44, 1 (2021), 3--10.","journal-title":"Assessment and Prognosis. IEEE Data Eng. Bull."},{"key":"e_1_2_2_14_1","volume-title":"The adaptive radix tree: ARTful indexing for main-memory databases","author":"Leis Viktor","unstructured":"Viktor Leis, Alfons Kemper, and Thomas Neumann. 2013. The adaptive radix tree: ARTful indexing for main-memory databases. In ICDE. IEEE Computer Society, 38--49."},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.14778\/3489496.3489512"},{"key":"e_1_2_2_16_1","volume-title":"LISA: A Learned Index Structure for Spatial Data. In SIMGOD. ACM, 2119--2133.","author":"Li Pengfei","year":"2020","unstructured":"Pengfei Li, Hua Lu, Qian Zheng, Long Yang, and Gang Pan. 2020. LISA: A Learned Index Structure for Spatial Data. In SIMGOD. ACM, 2119--2133."},{"key":"e_1_2_2_17_1","volume-title":"Proceedings of the 1st International Workshop on Applied AI for Database Systems and Applications.","author":"Llaveshi Anisa","year":"2019","unstructured":"Anisa Llaveshi, Utku Sirin, Anastasia Ailamaki, and Robert West. 2019. Accelerating B tree search by using simple machine learning techniques. In Proceedings of the 1st International Workshop on Applied AI for Database Systems and Applications."},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.14778\/3421424.3421425"},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/174130.174139"},{"key":"e_1_2_2_20_1","volume-title":"RUSLI: Real-time Updatable Spline Learned Index. In aiDM@SIGMOD. ACM, 1--8.","author":"Mishra Mayank","year":"2021","unstructured":"Mayank Mishra and Rekha Singhal. 2021. RUSLI: Real-time Updatable Spline Learned Index. In aiDM@SIGMOD. ACM, 1--8."},{"key":"e_1_2_2_21_1","doi-asserted-by":"crossref","unstructured":"Vikram Nathan Jialin Ding Mohammad Alizadeh and Tim Kraska. 2020. Learning Multi-Dimensional Indexes. In SIGMOD. ACM 985--1000.","DOI":"10.1145\/3318464.3380579"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1007\/s002360050048"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/358746.358758"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.14778\/3407790.3407829"},{"key":"e_1_2_2_25_1","volume-title":"ICML","volume":"37","author":"Rezende Danilo Jimenez","year":"2015","unstructured":"Danilo Jimenez Rezende and Shakir Mohamed. 2015. Variational Inference with Normalizing Flows. In ICML, Vol. 37. JMLR.org, 1530--1538."},{"key":"e_1_2_2_26_1","volume-title":"Umar Farooq Minhas, and Tim Kraska","author":"Spector Benjamin","year":"2021","unstructured":"Benjamin Spector, Andreas Kipf, Kapil Vaidya, Chi Wang, Umar Farooq Minhas, and Tim Kraska. 2021. Bounding the Last Mile: Efficient Learned String Indexing. CoRR abs\/2111.14905 (2021)."},{"key":"e_1_2_2_27_1","doi-asserted-by":"crossref","unstructured":"Chuzhe Tang YouyunWang Zhiyuan Dong Gansen Hu ZhaoguoWang MinjieWang and Haibo Chen. 2020. XIndex: a scalable learned index for multicore data storage. In PPoPP. ACM 308--320.","DOI":"10.1145\/3332466.3374547"},{"key":"e_1_2_2_28_1","doi-asserted-by":"crossref","unstructured":"Youyun Wang Chuzhe Tang Zhaoguo Wang and Haibo Chen. 2020. SIndex: a scalable learned index for string keys. In APSys. ACM 17--24.","DOI":"10.1145\/3409963.3410496"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.14778\/3551793.3551848"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.14778\/3457390.3457393"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.14778\/3547305.3547322"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.14778\/3551793.3551823"}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589284","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3589284","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:48:54Z","timestamp":1750182534000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589284"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,13]]},"references-count":31,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2023,6,13]]}},"alternative-id":["10.1145\/3589284"],"URL":"https:\/\/doi.org\/10.1145\/3589284","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,13]]}}}