{"status":"ok","message-type":"work-list","message-version":"1.0.0","message":{"facets":{},"total-results":1160,"items":[{"indexed":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T21:47:44Z","timestamp":1775598464166,"version":"3.50.1"},"reference-count":33,"publisher":"Association for Computing Machinery (ACM)","issue":"1","funder":[{"name":"Natural Sciences and Engineering Research Council of Canada","award":["RGPIN-2020-05408"],"award-info":[{"award-number":["RGPIN-2020-05408"]}]},{"name":"European Union's Horizon Europe Research and Innovation programme","award":["101188416"],"award-info":[{"award-number":["101188416"]}]},{"name":"Research Grant Council of Hong Kong","award":["17202325"],"award-info":[{"award-number":["17202325"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2026,4,2]]},"abstract":"<jats:p>\n                    We study Aggregation Queries over Nearest Neighbors (AQNN), which compute aggregates over the learned representations of the neighborhood of a designated query object. For example, a medical professional may be interested in\n                    <jats:italic toggle=\"yes\">the average heart rate of patients whose representations are similar to that of an insomnia patient.<\/jats:italic>\n                    Answering AQNNs accurately and efficiently is challenging due to the high cost of generating high-quality representations (e.g., via a deep learning model trained on human expert annotations) and the different sensitivities of different aggregation functions to neighbor selection errors. We address these challenges by combining high-quality and low-cost representations to approximate the aggregate. We characterize value- and count-sensitive AQNNs and propose the\n                    <jats:italic toggle=\"yes\">Sampler with Precision-Recall in Target<\/jats:italic>\n                    (\n                    <jats:sc>SPRinT<\/jats:sc>\n                    ), a query answering framework that works in three steps: (1) sampling, (2) nearest neighbor selection, and (3) aggregation. We further establish theoretical bounds on sample sizes and aggregation errors. Extensive experiments on five datasets from three domains (medical, social media, and e-commerce) demonstrate that\n                    <jats:sc>SPRinT<\/jats:sc>\n                    achieves the lowest aggregation error with minimal computation cost in most cases compared to existing solutions.\n                    <jats:sc>SPRinT<\/jats:sc>\n                    's performance remains stable as dataset size grows, confirming its scalability for large-scale applications requiring both accuracy and efficiency.\n                  <\/jats:p>","DOI":"10.1145\/3786672","type":"journal-article","created":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T17:54:13Z","timestamp":1775584453000},"page":"1-26","source":"Crossref","is-referenced-by-count":0,"title":["On Efficient Approximate Aggregate Nearest Neighbor Queries over Learned Representations"],"prefix":"10.1145","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-7638-1448","authenticated-orcid":false,"given":"Carrie","family":"Wang","sequence":"first","affiliation":[{"name":"The University of Hong Kong, Hong Kong, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6194-4502","authenticated-orcid":false,"given":"Sihem","family":"Amer-Yahia","sequence":"additional","affiliation":[{"name":"CNRS, Univ. Grenoble Alpes, Grenoble, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9775-4241","authenticated-orcid":false,"given":"Laks V.S.","family":"Lakshmanan","sequence":"additional","affiliation":[{"name":"The University of British Columbia, Vancouver, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9480-9809","authenticated-orcid":false,"given":"Reynold","family":"Cheng","sequence":"additional","affiliation":[{"name":"The University of Hong Kong, Hong Kong, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,7]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2015. Yelp Dataset. https:\/\/www.yelp.com\/dataset."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2465351.2465355"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2019.00132"},{"key":"e_1_2_1_4_1","volume-title":"Hd-index: Pushing the scalability-accuracy boundary for approximate knn search in high-dimensional spaces. arXiv preprint arXiv:1804.06829","author":"Arora Akhil","year":"2018","unstructured":"Akhil Arora, Sakshi Sinha, Piyush Kumar, and Arnab Bhattacharya. 2018. Hd-index: Pushing the scalability-accuracy boundary for approximate knn search in high-dimensional spaces. arXiv preprint arXiv:1804.06829 (2018)."},{"key":"e_1_2_1_5_1","doi-asserted-by":"crossref","unstructured":"Jon L Bentley. 1975. A survey of techniques for fixed radius near neighbor searching. Technical Report. Stanford University Stanford CA USA.","DOI":"10.2172\/1453938"},{"key":"e_1_2_1_6_1","volume-title":"Approximate Query Processing Using Wavelets. In VLDB 2000, Proceedings of 26th International Conference on Very Large Data Bases, September 10--14","author":"Chakrabarti Kaushik","year":"2000","unstructured":"Kaushik Chakrabarti, Minos N. Garofalakis, Rajeev Rastogi, and Kyuseok Shim. 2000. Approximate Query Processing Using Wavelets. In VLDB 2000, Proceedings of 26th International Conference on Very Large Data Bases, September 10--14, 2000, Cairo, Egypt. Morgan Kaufmann, 111--122."},{"key":"e_1_2_1_7_1","volume-title":"nithum, and Will Cukierski","author":"Sorensen Jeffrey","year":"2017","unstructured":"cjadams, Jeffrey Sorensen, Julia Elliott, Lucas Dixon, Mark McDonald, nithum, and Will Cukierski. 2017. Toxic Comment Classification Challenge. https:\/\/kaggle.com\/competitions\/jigsaw-toxic-comment-classification-challenge."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.14778\/3574245.3574273"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/355744.355745"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4612-0865-5_26"},{"key":"e_1_2_1_11_1","volume-title":"Bridging Language and Items for Retrieval and Recommendation. arXiv preprint arXiv:2403.03952","author":"Hou Yupeng","year":"2024","unstructured":"Yupeng Hou, Jiacheng Li, Zhankui He, An Yan, Xiusi Chen, and Julian McAuley. 2024. Bridging Language and Items for Retrieval and Recommendation. arXiv preprint arXiv:2403.03952 (2024)."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.14778\/2850469.2850470"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/276698.276876"},{"key":"e_1_2_1_14_1","volume-title":"Conference On Learning Theory. PMLR","author":"Indyk Piotr","year":"2018","unstructured":"Piotr Indyk and Tal Wagner. 2018. Approximate nearest neighbors in limited space. In Conference On Learning Theory. PMLR, 2012--2036."},{"key":"e_1_2_1_15_1","volume-title":"Leo Anthony Celi, and Roger G Mark","author":"Johnson Alistair EW","year":"2016","unstructured":"Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. 2016. MIMIC-III, a freely accessible critical care database. Scientific data 3, 1 (2016), 1--9."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1016\/J.INS.2020.09.024"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.14778\/3137628.3137664"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.14778\/3407790.3407804"},{"key":"e_1_2_1_19_1","volume-title":"Central limit theorem: the cornerstone of modern statistics. Korean journal of anesthesiology 70, 2","author":"Kwak Sang Gyu","year":"2017","unstructured":"Sang Gyu Kwak and Jong Hae Kim. 2017. Central limit theorem: the cornerstone of modern statistics. Korean journal of anesthesiology 70, 2 (2017), 144--156."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3452786"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3380600"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.1982.1056489"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3183751"},{"key":"e_1_2_1_24_1","volume-title":"Proc. of 5th Berkeley Symposium on Math. Stat. and Prob. 281--297","author":"McQueen James B","year":"1967","unstructured":"James B McQueen. 1967. Some methods of classification and analysis of multivariate observations. In Proc. of 5th Berkeley Symposium on Math. Stat. and Prob. 281--297."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2014.2321376"},{"key":"e_1_2_1_26_1","volume-title":"Jesse D Raffa, Leo A Celi, Roger G Mark, and Omar Badawi.","author":"Pollard Tom J","year":"2018","unstructured":"Tom J Pollard, Alistair EW Johnson, Jesse D Raffa, Leo A Celi, Roger G Mark, and Omar Badawi. 2018. The eICU Collaborative Research Database, a freely available multi-center database for critical care research. Scientific data 5, 1 (2018), 1--13."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00045"},{"key":"e_1_2_1_28_1","volume-title":"Probability inequalities for the sum in sampling without replacement. The Annals of Statistics","author":"Serfling Robert J","year":"1974","unstructured":"Robert J Serfling. 1974. Probability inequalities for the sum in sampling without replacement. The Annals of Statistics (1974), 39--48."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2008.4587638"},{"key":"e_1_2_1_30_1","volume-title":"Mpnet: Masked and permuted pre-training for language understanding. Advances in neural information processing systems 33","author":"Song Kaitao","year":"2020","unstructured":"Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2020. Mpnet: Masked and permuted pre-training for language understanding. Advances in neural information processing systems 33 (2020), 16857--16867."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2017.2761740"},{"key":"e_1_2_1_32_1","volume-title":"Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems","author":"Wang Wenhui","year":"2020","unstructured":"Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020. MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6--12, 2020, virtual."},{"key":"e_1_2_1_33_1","volume-title":"Gliner: Generalist model for named entity recognition using bidirectional transformer. arXiv preprint arXiv:2311.08526","author":"Zaratiana Urchade","year":"2023","unstructured":"Urchade Zaratiana, Nadi Tomeh, Pierre Holat, and Thierry Charnois. 2023. Gliner: Generalist model for named entity recognition using bidirectional transformer. arXiv preprint arXiv:2311.08526 (2023)."}],"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3786672","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T21:01:01Z","timestamp":1775595661000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3786672"}},"issued":{"date-parts":[[2026,4,2]]},"references-count":33,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,4,2]]}},"alternative-id":["10.1145\/3786672"],"URL":"https:\/\/doi.org\/10.1145\/3786672","ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"published":{"date-parts":[[2026,4,2]]}},{"indexed":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T19:00:48Z","timestamp":1774983648138,"version":"3.50.1"},"reference-count":66,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2025,2,10]],"date-time":"2025-02-10T00:00:00Z","timestamp":1739145600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100006374","name":"Japan Society for the Promotion of Science","doi-asserted-by":"publisher","award":["Kakenhi JP23K17456, JP23K25157, JP23K28096"],"award-info":[{"award-number":["Kakenhi JP23K17456, JP23K25157, JP23K28096"]}],"id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100006374","name":"Japan Science and Technology Agency","doi-asserted-by":"publisher","award":["CREST JPMJCR22M2"],"award-info":[{"award-number":["CREST JPMJCR22M2"]}],"id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2025,2,10]]},"abstract":"<jats:p>Existing what-if analysis systems are predominantly tailored to operate on either only the application layer or only the database layer of software. This isolated approach limits their effectiveness in scenarios where intensive interaction between applications and database systems occurs. To address this gap, we introduce Ultraverse, a what-if analysis framework that seamlessly integrates both application and database layers. Ultraverse employs dynamic symbolic execution to effectively translate application code into compact SQL procedure representations, thereby synchronizing application semantics at both SQL and application levels during what-if replays. A novel aspect of Ultraverse is its use of advanced query dependency analysis, which serves two key purposes: (1) it eliminates the need to replay irrelevant transactions that do not influence the outcome, and (2) it facilitates parallel replay of mutually independent transactions, significantly enhancing the analysis efficiency. Ultraverse is applicable to existing unmodified database systems and legacy application codes. Our extensive evaluations of the framework have demonstrated remarkable improvements in what-if analysis speed, achieving performance gains ranging from 7.7x to 291x across diverse benchmarks.<\/jats:p>","DOI":"10.1145\/3709734","type":"journal-article","created":{"date-parts":[[2025,2,11]],"date-time":"2025-02-11T15:45:06Z","timestamp":1739288706000},"page":"1-27","source":"Crossref","is-referenced-by-count":1,"title":["Ultraverse: An Efficient What-if Analysis Framework for Software Applications Interacting with Database Systems"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1282-5208","authenticated-orcid":false,"given":"Ronny","family":"Ko","sequence":"first","affiliation":[{"name":"Osaka University, Osaka, Kansai, Japan, Ohio State University, Columbus, USA, and Sensor Tech Inc., Daejeon, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7239-5134","authenticated-orcid":false,"given":"Chuan","family":"Xiao","sequence":"additional","affiliation":[{"name":"Osaka University, Osaka, Japan and Nagoya University, Nagoya, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5559-8300","authenticated-orcid":false,"given":"Makoto","family":"Onizuka","sequence":"additional","affiliation":[{"name":"Osaka University, Osaka, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6527-5994","authenticated-orcid":false,"given":"Zhiqiang","family":"Lin","sequence":"additional","affiliation":[{"name":"The Ohio State University, Columbus, OH, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8717-8343","authenticated-orcid":false,"given":"Yihe","family":"Huang","sequence":"additional","affiliation":[{"name":"Databricks, San Francisco, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,2,11]]},"reference":[{"key":"e_1_2_1_1_1","first-page":"1179","volume-title":"VLDB","author":"Alexe B.","year":"2006","unstructured":"B. Alexe, L. Chiticariu, andW.-C. Tan. Spider: A schema mapping debugger. In VLDB, pages 1179--1182, 2006."},{"key":"e_1_2_1_2_1","volume-title":"Astore: ecommerce web","author":"Pham Anh","year":"2023","unstructured":"Anh Pham. Astore: ecommerce web, 2023. https:\/\/github.com\/anhpham1509\/eCommerceWeb."},{"key":"e_1_2_1_3_1","unstructured":"Apache Iceberg 2024. https:\/\/iceberg.apache.org\/."},{"key":"e_1_2_1_4_1","volume-title":"Supported Sources and More","author":"Pte Atlan","year":"2024","unstructured":"Atlan Pte. Ltd. DataHub Column-Level Lineage: Features, Supported Sources and More, 2024. https:\/\/atlan.com\/know\/ data-catalog\/datahub\/column-level-lineage\/."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP40001.2021.00077"},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the Sixth Conference on Uncertainty in Artificial Intelligence","author":"T. BALL, J.","unstructured":"T. BALL, J. DANIEL, and T. Ball. Deconstructing dynamic symbolic execution. Technical Report MSR-TR-2015--95, January 2015. Proceedings of the Sixth Conference on Uncertainty in Artificial Intelligence, Boston, MA."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.5555\/645926.672016"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-005-0156-6"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/377674.377665"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2016.7498289"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF01231700"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.5555\/645504.656274"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3526138"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2043556.2043567"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.5555\/2685048.2685092"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2499370.2462180"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1066157.1066296"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/357775.357777"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.5555\/862872"},{"key":"e_1_2_1_20_1","first-page":"715","volume-title":"VLDB","author":"Daudjee K.","year":"2006","unstructured":"K. Daudjee and K. Salem. Lazy database replication with snapshot isolation. In VLDB, pages 715--726, 2006."},{"key":"e_1_2_1_21_1","volume-title":"CIDR","author":"Deutch D.","year":"2013","unstructured":"D. Deutch, Z. Ives, T. Milo, and V. Tannen. Caravan: Provisioning for what-if analysis. In CIDR, 2013."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.14778\/2732240.2732246"},{"key":"e_1_2_1_23_1","unstructured":"DPDK Project. Ring Library 2019. https:\/\/doc.dpdk.org\/guides\/prog_guide\/ring_lib.html."},{"key":"e_1_2_1_24_1","volume-title":"Tableau 201: How to Make a What-If Analysis Using Parameters","year":"2024","unstructured":"Evolytics. Tableau 201: How to Make a What-If Analysis Using Parameters, 2024. https:\/\/evolytics.com\/blog\/tableau-201-how-to-make-a-what-if-analysis-using-parameters\/."},{"issue":"11","key":"e_1_2_1_25_1","first-page":"1190","article-title":"Rethinking serializable multiversion concurrency control","volume":"8","author":"Abadi J.M.","year":"2015","unstructured":"J.M.FaleiroandD. J. Abadi. Rethinking serializable multiversion concurrency control. Proc.VLDBEndow., 8(11):1190--1201, jul 2015.","journal-title":"Proc.VLDBEndow."},{"key":"e_1_2_1_26_1","unstructured":"Felix Campbell. What-if Reproducibility 2023. https:\/\/github.com\/fsalc\/whatif-reproducibility."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3520175"},{"key":"e_1_2_1_28_1","volume-title":"Built-in functions for atomic memory access","author":"GCC","year":"2019","unstructured":"GCC GNU. Built-in functions for atomic memory access, 2019. https:\/\/gcc.gnu.org\/onlinedocs\/gcc-4.1.0\/gcc\/Atomic- Builtins.html."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2001420.2001424"},{"key":"e_1_2_1_30_1","unstructured":"Google Inc. Puppeteer 2015. https:\/\/pptr.dev\/."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/1247480.1247631"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3034786.3056125"},{"key":"e_1_2_1_33_1","first-page":"1186","volume-title":"ICDE","author":"Hong C.","year":"2013","unstructured":"C. Hong, D. Zhou, M. Yang, C. Kuo, L. Zhang, and L. Zhou. KuaFu: Closing the parallelism gap in database replication. In ICDE, pages 1186--1195, 2013."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-019-00594-5"},{"key":"e_1_2_1_35_1","first-page":"1","volume-title":"Introduction to temporal database research. Temporal database management","author":"Jensen C.","year":"2000","unstructured":"C. Jensen. Introduction to temporal database research. Temporal database management, pages 1--29, 2000."},{"key":"e_1_2_1_36_1","volume-title":"The Right AndWrongWays To Use 'What-If' Analysis","author":"Shapiro Joel","year":"2020","unstructured":"Joel Shapiro. The Right AndWrongWays To Use 'What-If' Analysis, 2020. https:\/\/www.forbes.com\/sites\/joelshapiro\/ 2020\/07\/28\/the-right-and-wrong-ways-to-use-what-if-analysis\/'sh=3f3617eb2b98."},{"key":"e_1_2_1_37_1","first-page":"34","article-title":"Temporal features","author":"Kulkarni K.","year":"2011","unstructured":"K. Kulkarni and J.-E. Michels. Temporal features in SQL:2011. SIGMOD Rec., 41(3):34--43, 2012.","journal-title":"SQL"},{"key":"e_1_2_1_38_1","volume-title":"Proceedings of the 40th International Conference on Machine Learning, ICML'23. JMLR.org","author":"Longpre S.","year":"2023","unstructured":"S. Longpre, L. Hou, T. Vu, A. Webson, H. W. Chung, Y. Tay, D. Zhou, Q. V. Le, B. Zoph, J. Wei, and A. Roberts. The flan collection: Designing data and methods for effective instruction tuning. In Proceedings of the 40th International Conference on Machine Learning, ICML'23. JMLR.org, 2023."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3092282.3092295"},{"key":"e_1_2_1_40_1","first-page":"2","author":"Sridharan Manu","year":"2023","unstructured":"Manu Sridharan. Jalangi 2, 2023. https:\/\/github.com\/Samsung\/jalangi2.","journal-title":"Jalangi"},{"key":"e_1_2_1_41_1","volume-title":"MariaDB Source Code","author":"DB.","year":"2019","unstructured":"MariaDB. MariaDB Source Code, 2019. https:\/\/mariadb.com\/kb\/en\/library\/getting-the-mariadb-source-code\/."},{"key":"e_1_2_1_42_1","volume-title":"Temporal Data Tables","author":"DB.","year":"2019","unstructured":"MariaDB. Temporal Data Tables, 2019. https:\/\/mariadb.com\/kb\/en\/library\/temporal-data-tables\/."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/2213836.2213875"},{"key":"e_1_2_1_44_1","volume-title":"Introduction to What-If Analysis","year":"2024","unstructured":"Microsoft. Introduction to What-If Analysis, 2024. https:\/\/support.microsoft.com\/en-gb\/office\/introduction-to-what-ifanalysis- 22bffa5f-e891--4acc-bf7a-e4645c446fb4."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-012-0294-6"},{"key":"e_1_2_1_46_1","volume-title":"The Binary Log","year":"2023","unstructured":"Oracle. The Binary Log, 2023. https:\/\/dev.mysql.com\/doc\/refman\/8.0\/en\/binary-log.html\/."},{"key":"e_1_2_1_47_1","volume-title":"Tableau 201: Create What-If Analysis Using Parameters in Oracle Analytics","year":"2024","unstructured":"Oracle. Tableau 201: Create What-If Analysis Using Parameters in Oracle Analytics, 2024. https:\/\/docs.oracle.com\/en\/ cloud\/paas\/analytics-cloud\/tutorial-what-if-analysis\/index.html."},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/69.404027"},{"key":"e_1_2_1_49_1","volume-title":"Thick Client Proxying","author":"Den Parsia's","year":"2015","unstructured":"Parsia's Den. Thick Client Proxying, 2015. https:\/\/parsiya.net\/blog\/2015--10--19-proxying-hipchat-part-3\\-ssl-addedand- removed-here\/."},{"key":"e_1_2_1_50_1","first-page":"155","volume-title":"Middleware","author":"Plattner C.","year":"2004","unstructured":"C. Plattner and G. Alonso. Ganymed: Scalable replication for transactional web applications. In Middleware, pages 155--174, 2004."},{"key":"e_1_2_1_51_1","volume-title":"A step-by-step guide on creating What-if analysis in Power BI","author":"Tubsamon Ploii","year":"2023","unstructured":"Ploii Tubsamon. A step-by-step guide on creating What-if analysis in Power BI, 2023. https:\/\/ploiitubsamon.medium. com\/a-step-by-step-guide-on-creating-what-if-analysis-in-power-bi-c5c09def5cf5."},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3357223.3362702"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/1806596.1806598"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1016\/0097-3165(77)90082-6"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/69.50911"},{"key":"e_1_2_1_56_1","volume-title":"Conflict serializability in dbms","author":"Srivastava Sonal","year":"2019","unstructured":"Sonal Srivastava. Conflict serializability in dbms, 2019. https:\/\/www.geeksforgeeks.org\/conflict-serializability-in-dbms\/."},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.14778\/1920841.1920855"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/2213836.2213838"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.14722\/ndss.2017.23100"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3035925"},{"key":"e_1_2_1_61_1","first-page":"519","volume-title":"VLDB","author":"Wang Y. R.","year":"1990","unstructured":"Y. R.Wang and S. E. Madnick. A polygen model for heterogeneous database systems: The source tagging perspective. In VLDB, pages 519--538, 1990."},{"key":"e_1_2_1_62_1","first-page":"1199","volume-title":"27th USENIX Security Symposium (USENIX Security 18)","author":"Webster A.","year":"2018","unstructured":"A.Webster, R. Eckenrod, and J. Purtilo. Fast and service-preserving recovery from malware infections using CRIU. In 27th USENIX Security Symposium (USENIX Security 18), pages 1199--1211, Baltimore,MD, Aug. 2018. USENIX Association."},{"key":"e_1_2_1_63_1","volume-title":"Transactional Information Systems: Theory, Algorithms, and the Practice of Concurrency Control and Recovery","author":"Weikum G.","year":"2001","unstructured":"G.Weikum and G. Vossen. Transactional Information Systems: Theory, Algorithms, and the Practice of Concurrency Control and Recovery. Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 2001."},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2017.2788018"},{"key":"e_1_2_1_65_1","first-page":"465","volume-title":"OSDI","author":"Zheng W.","year":"2014","unstructured":"W. Zheng, S. Tu, E. Kohler, and B. Liskov. Fast databases with fast durability and recovery through multicore parallelism. In OSDI, pages 465--477, 2014."},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/568271.223848"}],"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3709734","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3709734","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T18:17:08Z","timestamp":1774981028000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3709734"}},"issued":{"date-parts":[[2025,2,10]]},"references-count":66,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,2,10]]}},"alternative-id":["10.1145\/3709734"],"URL":"https:\/\/doi.org\/10.1145\/3709734","ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"published":{"date-parts":[[2025,2,10]]}},{"indexed":{"date-parts":[[2026,3,14]],"date-time":"2026-03-14T09:53:09Z","timestamp":1773481989478,"version":"3.50.1"},"reference-count":54,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,3,12]],"date-time":"2024-03-12T00:00:00Z","timestamp":1710201600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2024,3,12]]},"abstract":"<jats:p>Substantial research efforts have been devoted to studying the performance optimality problem for distributed database transactions. However, they focus just on optimizing transactional reads, and thus overlook crucial factors, such as the efficiency of writes, which also impact the overall system performance. Motivated by a recent study on Twitter's workloads showing the prominence of write-heavy workloads in practice, we make a substantial step towards performance-optimal distributed transactions by also aiming to optimize writes, a fundamentally new dimension to this problem. We propose a new design objective and establish impossibility results with respect to the achievable isolation levels. Guided by these results, we present two new transaction algorithms with different isolation guarantees that fulfill this design objective. Our evaluation demonstrates that these algorithms outperform the state of the art.<\/jats:p>","DOI":"10.1145\/3639264","type":"journal-article","created":{"date-parts":[[2024,3,26]],"date-time":"2024-03-26T18:51:32Z","timestamp":1711479092000},"page":"1-25","source":"Crossref","is-referenced-by-count":6,"title":["NOC-NOC: Towards Performance-optimal Distributed Transactions"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3578-7432","authenticated-orcid":false,"given":"Si","family":"Liu","sequence":"first","affiliation":[{"name":"ETH Zurich, Zurich, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-1393-4732","authenticated-orcid":false,"given":"Luca","family":"Multazzu","sequence":"additional","affiliation":[{"name":"ETH Zurich, Zurich, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0427-9710","authenticated-orcid":false,"given":"Hengfeng","family":"Wei","sequence":"additional","affiliation":[{"name":"Nanjing University, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2952-939X","authenticated-orcid":false,"given":"David A.","family":"Basin","sequence":"additional","affiliation":[{"name":"ETH Zurich, Zurich, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,3,26]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF01784241"},{"key":"e_1_2_2_2_1","volume-title":"Preguicc a, and Marc Shapiro","author":"Akkoorath Deepthi Devaki","year":"2016","unstructured":"Deepthi Devaki Akkoorath, Alejandro Z. Tomsic, Manuel Bravo, Zhongmiao Li, Tyler Crain, Annette Bieniusa, Nuno M. Preguicc a, and Marc Shapiro. 2016. Cure: Strong Semantics Meets High Availability and Low Latency. In ICDCS 2016. IEEE Computer Society, 405--414."},{"key":"e_1_2_2_3_1","volume-title":"The Impossibility of Fast Transactions. In IPDPS'20","author":"Antoniadis Karolos","year":"2020","unstructured":"Karolos Antoniadis, Diego Didona, Rachid Guerraoui, and Willy Zwaenepoel. 2020. The Impossibility of Fast Transactions. In IPDPS'20. IEEE, 1143--1154."},{"key":"e_1_2_2_4_1","unstructured":"Masoud Saeida Ardekani Pierre Sutra Nuno Pregui\u00e7a and Marc Shapiro. 2013. Non-Monotonic Snapshot Isolation. arxiv: 1306.3906 [cs.DC]"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2909870"},{"key":"e_1_2_2_6_1","volume-title":"O'Neil","author":"Berenson Hal","year":"1995","unstructured":"Hal Berenson, Philip A. Bernstein, Jim Gray, Jim Melton, Elizabeth J. O'Neil, and Patrick E. O'Neil. 1995. A Critique of ANSI SQL Isolation Levels. In SIGMOD'95. ACM Press, 1--10."},{"key":"e_1_2_2_7_1","volume-title":"Concurrency Control and Recovery in Database Systems","author":"Bernstein Philip A.","unstructured":"Philip A. Bernstein, Vassos Hadzilacos, and Nathan Goodman. 1987. Concurrency Control and Recovery in Database Systems. Addison-Wesley."},{"key":"e_1_2_2_8_1","doi-asserted-by":"crossref","unstructured":"Eric A. Brewer. 2000. Towards robust distributed systems (abstract). In PODC. 7.","DOI":"10.1145\/343477.343502"},{"key":"e_1_2_2_9_1","volume-title":"ATC'13","author":"Bronson Nathan","year":"2013","unstructured":"Nathan Bronson, Zach Amsden, George Cabrera, Prasad Chakka, Peter Dimov, Hui Ding, Jack Ferris, Anthony Giardullo, Sachin Kulkarni, Harry C. Li, Mark Marchukov, Dmitri Petrov, Lovro Puzar, Yee Jiun Song, and Venkateshwaran Venkataramani. 2013. TAO: Facebook's Distributed Data Store for the Social Graph. In ATC'13. USENIX Association, 49--60."},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.14778\/3236187.3236209"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.14778\/3538598.3538616"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476311.3476379"},{"key":"e_1_2_2_13_1","volume-title":"Lazy Database Replication with Ordering Guarantees. In ICDE","author":"Daudjee Khuzaima","year":"2004","unstructured":"Khuzaima Daudjee and Kenneth Salem. 2004. Lazy Database Replication with Ordering Guarantees. In ICDE 2004. IEEE Computer Society, 424--435."},{"key":"e_1_2_2_14_1","volume-title":"Distributed Transactional Systems Cannot Be Fast. In SPAA'19","author":"Didona Diego","year":"2019","unstructured":"Diego Didona, Panagiota Fatourou, Rachid Guerraoui, Jingjing Wang, and Willy Zwaenepoel. 2019. Distributed Transactional Systems Cannot Be Fast. In SPAA'19. ACM, 369--380."},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.14778\/3236187.3236210"},{"key":"e_1_2_2_16_1","first-page":"1","article-title":"GentleRain: Cheap and Scalable Causal Consistency with Physical Clocks. In SoCC 2014","volume":"4","author":"Du Jiaqing","year":"2014","unstructured":"Jiaqing Du, Calin Iorgulescu, Amitabha Roy, and Willy Zwaenepoel. 2014. GentleRain: Cheap and Scalable Causal Consistency with Physical Clocks. In SoCC 2014. ACM, 4:1--4:13.","journal-title":"ACM"},{"key":"e_1_2_2_17_1","volume-title":"The Design and Operation of CloudLab. In ATC'19","author":"Duplyakin Dmitry","year":"2019","unstructured":"Dmitry Duplyakin, Robert Ricci, Aleksander Maricq, Gary Wong, Jonathon Duerig, Eric Eide, Leigh Stoller, Mike Hibler, David Johnson, Kirk Webb, Aditya Akella, Kuangching Wang, Glenn Ricart, Larry Landweber, Chip Elliott, Michael Zink, Emmanuel Cecchet, Snigdhaswin Kar, and Prabodh Mishra. 2019. The Design and Operation of CloudLab. In ATC'19. USENIX Association, 1--14."},{"key":"e_1_2_2_18_1","volume-title":"Accessed","author":"SQL.","year":"2023","unstructured":"ElectricSQL. Accessed in July, 2023. https:\/\/electric-sql.com\/."},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","unstructured":"Shabnam Ghasemirad. 2022. Mechanized Data Consistency Models for Distributed Database Transactions. Master's thesis. ETH Zurich. https:\/\/doi.org\/10.3929\/ethz-b-000581334.","DOI":"10.3929\/ethz-b-000581334"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/564585.564601"},{"key":"e_1_2_2_21_1","volume-title":"PODC'11","author":"Golab Wojciech","unstructured":"Wojciech Golab, Xiaozhou Li, and Mehul A. Shah. 2011. Analyzing consistency properties for fun and profit. In PODC'11. ACM, 197--206."},{"key":"e_1_2_2_22_1","volume-title":"Regular Sequential Serializability and Regular Sequential Consistency. In SOSP'21","author":"Helt Jeffrey","year":"2021","unstructured":"Jeffrey Helt, Matthew Burke, Amit Levy, and Wyatt Lloyd. 2021. Regular Sequential Serializability and Regular Sequential Consistency. In SOSP'21. ACM, 163--179."},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.14778\/1454159.1454211"},{"key":"e_1_2_2_24_1","volume-title":"Lynch","author":"Konwar Kishori M.","year":"2021","unstructured":"Kishori M. Konwar, Wyatt Lloyd, Haonan Lu, and Nancy A. Lynch. 2021. SNOW Revisited: Understanding When Ideal READ Transactions Are Possible. In IPDPS'21. IEEE, 922--931."},{"key":"e_1_2_2_25_1","volume-title":"Sturgis","author":"Lampson Butler","year":"1979","unstructured":"Butler Lampson and Howard E. Sturgis. 1979. Crash recovery in a distributed storage system. Xerox Palo Alto Research Center."},{"key":"e_1_2_2_26_1","volume-title":"Andersen","author":"Lim Hyeontaek","year":"2017","unstructured":"Hyeontaek Lim, Michael Kaminsky, and David G. Andersen. 2017. Cicada: Dependably Fast Multi-Core In-Memory Transactions. In SIGMOD '17. ACM, 21--35."},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3494517"},{"key":"e_1_2_2_28_1","doi-asserted-by":"crossref","unstructured":"Si Liu Luca Multazzu Hengfeng Wei and David Basin. 2023. Artifact and technical report for \u201cNOC-NOC: Towards Performance-optimal Distributed Transactions\u201d. https:\/\/github.com\/siliunobi\/NOC-NOC.","DOI":"10.1145\/3639264"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2851613.2851838"},{"key":"e_1_2_2_30_1","volume-title":"SOSP","author":"Lloyd Wyatt","year":"2011","unstructured":"Wyatt Lloyd, Michael J. Freedman, Michael Kaminsky, and David G. Andersen. 2011. Don't settle for eventual: scalable causal consistency for wide-area storage with COPS. In SOSP 2011. ACM, 401--416."},{"key":"e_1_2_2_31_1","volume-title":"Andersen","author":"Lloyd Wyatt","year":"2013","unstructured":"Wyatt Lloyd, Michael J. Freedman, Michael Kaminsky, and David G. Andersen. 2013. Stronger Semantics for Low-Latency Geo-Replicated Storage. In NSDI'13. USENIX Association, 313--328."},{"key":"e_1_2_2_32_1","volume-title":"The SNOW Theorem and Latency-Optimal Read-Only Transactions. In OSDI'16","author":"Lu Haonan","year":"2016","unstructured":"Haonan Lu, Christopher Hodsdon, Khiem Ngo, Shuai Mu, and Wyatt Lloyd. 2016. The SNOW Theorem and Latency-Optimal Read-Only Transactions. In OSDI'16. USENIX Association, 135--150."},{"key":"e_1_2_2_33_1","volume-title":"OSDI'23","author":"Lu Haonan","year":"2023","unstructured":"Haonan Lu, Shuai Mu, Siddhartha Sen, and Wyatt Lloyd. 2023. NCC: Natural Concurrency Control for Strictly Serializable Datastores by Avoiding the Timestamp-Inversion Pitfall. In OSDI'23. USENIX Association, 305--323."},{"key":"e_1_2_2_34_1","volume-title":"Performance-Optimal Read-Only Transactions. In OSDI'20","author":"Lu Haonan","year":"2020","unstructured":"Haonan Lu, Siddhartha Sen, and Wyatt Lloyd. 2020a. Performance-Optimal Read-Only Transactions. In OSDI'20. USENIX Association, 333--349."},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2815400.2815426"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.14778\/3407790.3407808"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.14778\/3342263.3342270"},{"key":"e_1_2_2_38_1","volume-title":"Accessed","year":"2023","unstructured":"Microsoft. Accessed in July, 2023. Azure Cosmos DB. https:\/\/learn.microsoft.com\/en-us\/azure\/cosmos-db\/consistency-levels."},{"key":"e_1_2_2_39_1","volume-title":"Accessed","author":"DB.","year":"2023","unstructured":"MongoDB. Accessed in July, 2023. Read Isolation, Consistency, and Recency. https:\/\/docs.mongodb.com\/manual\/core\/read-isolation-consistency-recency\/."},{"key":"e_1_2_2_40_1","volume-title":"Accessed","author":"SQL.","year":"2023","unstructured":"MySQL. Accessed in July, 2023. MySQL Cluster CGE. https:\/\/www.mysql.com\/products\/cluster\/."},{"key":"e_1_2_2_41_1","volume-title":"Accessed","year":"2023","unstructured":"Neo4j. Accessed in July, 2023. https:\/\/neo4j.com\/docs\/operations-manual\/current\/clustering\/introduction\/."},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/322154.322158"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/2851141.2851170"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.14778\/2536222.2536232"},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/2043556.2043592"},{"key":"e_1_2_2_46_1","volume-title":"PaRiS: Causally Consistent Transactions with Non-blocking Reads and Partial Replication. In ICDCS","author":"Spirovska Kristina","year":"2019","unstructured":"Kristina Spirovska, Diego Didona, and Willy Zwaenepoel. 2019. PaRiS: Causally Consistent Transactions with Non-blocking Reads and Partial Replication. In ICDCS 2019. IEEE, 304--316."},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2020.3026778"},{"key":"e_1_2_2_48_1","volume-title":"Welch","author":"Terry Douglas B.","year":"1994","unstructured":"Douglas B. Terry, Alan J. Demers, Karin Petersen, Mike Spreitzer, Marvin Theimer, and Brent B. Welch. 1994. Session Guarantees for Weakly Consistent Replicated Data. In PDIS. IEEE Computer Society, 140--149."},{"key":"e_1_2_2_49_1","doi-asserted-by":"crossref","unstructured":"Alejandro Z. Tomsic Manuel Bravo and Marc Shapiro. 2018. Distributed transactional reads: the strong the quick the fresh & the impossible. In Middleware'18. ACM 120--133.","DOI":"10.1145\/3274808.3274818"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.14778\/3523210.3523221"},{"key":"e_1_2_2_51_1","volume-title":"An Integrated Experimental Environment for Distributed Systems and Networks","author":"White Brian","unstructured":"Brian White, Jay Lepreau, Leigh Stoller, Robert Ricci, Shashi Guruprasad, Mac Newbold, Mike Hibler, Chad Barb, and Abhijeet Joglekar. 2002. An Integrated Experimental Environment for Distributed Systems and Networks. In OSDI. USENIX Association, 255--270."},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3468521"},{"key":"e_1_2_2_53_1","volume-title":"SOSP","author":"Zhang Irene","year":"2015","unstructured":"Irene Zhang, Naveen Kr. Sharma, Adriana Szekeres, Arvind Krishnamurthy, and Dan R. K. Ports. 2015. Building consistent transactions with inconsistent replication. In SOSP 2015. ACM, 263--278."},{"key":"e_1_2_2_54_1","volume-title":"OSDI'19","author":"Zhou Fang","year":"2018","unstructured":"Fang Zhou, Yifan Gan, Sixiang Ma, and Yang Wang. 2018. wPerf: Generic Off-CPU Analysis to Identify Bottleneck Waiting Events. In OSDI'19. USENIX Association, 527--543."}],"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3639264","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3639264","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T15:12:09Z","timestamp":1755789129000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3639264"}},"issued":{"date-parts":[[2024,3,12]]},"references-count":54,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,3,12]]}},"alternative-id":["10.1145\/3639264"],"URL":"https:\/\/doi.org\/10.1145\/3639264","ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"published":{"date-parts":[[2024,3,12]]}},{"indexed":{"date-parts":[[2025,9,20]],"date-time":"2025-09-20T18:45:35Z","timestamp":1758393935527,"version":"3.41.0"},"reference-count":68,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2023,6,13]],"date-time":"2023-06-13T00:00:00Z","timestamp":1686614400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2023,6,13]]},"abstract":"<jats:p>Modern data analytics and AI jobs become increasingly complex and involve multiple tasks performed on specialized systems. Sharing of intermediate data between different systems is often a significant bottleneck in such jobs. When the intermediate data is large, it is mostly exchanged through files in standard formats (e.g., CSV and ORC), causing high I\/O and (de)serialization overheads. To solve these problems, we develop Vineyard, a high-performance, extensible, and cloud-native object store, trying to provide an intuitive experience for users to share data across systems in complex real-life workflows. Since different systems usually work on data structures (e.g., dataframes, graphs, hashmaps) with similar interfaces, and their computation logic is often loosely-coupled with how such interfaces are implemented over specific memory layouts, it enables Vineyard to conduct data sharing efficiently at a high level via memory mapping and method sharing. Vineyard provides an IDL named VCDL to facilitate users to register their own intermediate data types into Vineyard such that objects of the registered types can then be efficiently shared across systems in a polyglot workflow. As a cloud-native system, Vineyard is designed to work closely with Kubernetes, as well as achieve fault-tolerance and high performance in production environments. Evaluations on real-life datasets and data analytics jobs show that the above optimizations of Vineyard can significantly improve the end-to-end performance of data analytics jobs, by reducing their data-sharing time up to 68.4x.<\/jats:p>","DOI":"10.1145\/3589780","type":"journal-article","created":{"date-parts":[[2023,6,20]],"date-time":"2023-06-20T20:26:45Z","timestamp":1687292805000},"page":"1-27","source":"Crossref","is-referenced-by-count":5,"title":["Vineyard: Optimizing Data Sharing in Data-Intensive Analytics"],"prefix":"10.1145","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-5641-2452","authenticated-orcid":false,"given":"Wenyuan","family":"Yu","sequence":"first","affiliation":[{"name":"Alibaba Group, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-7687-7342","authenticated-orcid":false,"given":"Tao","family":"He","sequence":"additional","affiliation":[{"name":"Alibaba Group, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-7535-452X","authenticated-orcid":false,"given":"Lei","family":"Wang","sequence":"additional","affiliation":[{"name":"Alibaba Group, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3629-7892","authenticated-orcid":false,"given":"Ke","family":"Meng","sequence":"additional","affiliation":[{"name":"Alibaba Group, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-7273-0575","authenticated-orcid":false,"given":"Ye","family":"Cao","sequence":"additional","affiliation":[{"name":"Alibaba Group, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-7175-0784","authenticated-orcid":false,"given":"Diwen","family":"Zhu","sequence":"additional","affiliation":[{"name":"Alibaba Group, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-2869-7944","authenticated-orcid":false,"given":"Sanhong","family":"Li","sequence":"additional","affiliation":[{"name":"Alibaba Group, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6851-1366","authenticated-orcid":false,"given":"Jingren","family":"Zhou","sequence":"additional","affiliation":[{"name":"Alibaba Group, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,6,20]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"2019. Google Analytics Customer Revenue Prediction. https:\/\/www.kaggle.com\/c\/ga-customer-revenue-prediction."},{"key":"e_1_2_2_2_1","unstructured":"2023. ioctl(2) - Linux manual page. https:\/\/man7.org\/linux\/man-pages\/man2\/ioctl.2.html."},{"key":"e_1_2_2_3_1","unstructured":"2023. LD_PRELOAD - Linux manual page. https:\/\/man7.org\/linux\/man-pages\/man8\/ld.so.8.html."},{"key":"e_1_2_2_4_1","unstructured":"2023. Data-intensive computing. https:\/\/en.wikipedia.org\/wiki\/Data-intensive_computing."},{"key":"e_1_2_2_5_1","unstructured":"2023. Kubernets Scheduling Framework. https:\/\/kubernetes.io\/docs\/concepts\/scheduling-eviction\/scheduling-framework."},{"key":"e_1_2_2_6_1","unstructured":"2023. Node Property Prediction. https:\/\/ogb.stanford.edu\/docs\/nodeprop\/."},{"key":"e_1_2_2_7_1","unstructured":"2023. Production-Grade Container Orchestration. https:\/\/kubernetes.io."},{"key":"e_1_2_2_8_1","unstructured":"Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dandelion Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. https:\/\/www.tensorflow.org\/ Software available from tensorflow.org."},{"key":"e_1_2_2_9_1","unstructured":"Sajid Alam Nok Lam Chan Gabriel Comym Yetunde Dada Ivan Danov Deepyaman Datta Tynan DeBold Jannic Holzer Rashida Kanchwala Ankita Katiyar Amanda Koh Andrew Mackay Ahdra Merali Antony Milne Huong Nguyen Nero Okwa Juan Luis Cano Rodr\u00edguez Joel Schwarzmann Jo Stichbury and Merel Theisen. 2023. Kedro. https:\/\/github.com\/kedro-org\/kedro"},{"key":"e_1_2_2_10_1","volume-title":"Amazon Web Service","author":"Inc.","year":"2022","unstructured":"Inc. Amazon Web Service. 2022. Amazon Simple Storage Service: Object Storage built to retrieve any amount of data from anywhere. https:\/\/aws.amazon.com\/s3\/."},{"key":"e_1_2_2_11_1","volume-title":"9th USENIX Symposium on Networked Systems Design and Implementation (NSDI 12)","author":"Ananthanarayanan Ganesh","year":"2012","unstructured":"Ganesh Ananthanarayanan, Ali Ghodsi, Andrew Warfield, Dhruba Borthakur, Srikanth Kandula, Scott Shenker, and Ion Stoica. 2012. Pacman: Coordinated memory caching for parallel jobs. In 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI 12). 267--280."},{"key":"e_1_2_2_12_1","unstructured":"Fluid Authors. 2021. Fluid: elastic data abstraction and acceleration for BigData\/AI applications in cloud. https:\/\/fluid-cloudnative.github.io."},{"key":"e_1_2_2_13_1","unstructured":"Kubernetes Authors. 2022. Kubernets Custom Resources. https:\/\/kubernetes.io\/docs\/concepts\/extend-kubernetes\/api-extension\/custom-resources."},{"key":"e_1_2_2_14_1","unstructured":"Kubernetes Authors. 2022. Kubernets Operator Pattern. https:\/\/kubernetes.io\/docs\/concepts\/extend-kubernetes\/operator\/."},{"key":"e_1_2_2_15_1","unstructured":"NumPy Authors. 2022. NumPy: The fundamental package for scientific computing with Python. https:\/\/www.numpy.org\/."},{"key":"e_1_2_2_16_1","volume-title":"Pandas: Python Data Analysis Library. https:\/\/pandas.pydata.org\/.","author":"Pandas","year":"2022","unstructured":"Pandas authors. 2022. Pandas: Python Data Analysis Library. https:\/\/pandas.pydata.org\/."},{"key":"e_1_2_2_17_1","volume-title":"Polars: Fast multi-threaded, hybrid-streaming DataFrame library. https:\/\/www.pola.rs.","author":"Authors Polars","year":"2022","unstructured":"Polars Authors. 2022. Polars: Fast multi-threaded, hybrid-streaming DataFrame library. https:\/\/www.pola.rs."},{"key":"e_1_2_2_18_1","volume-title":"SWIG: Simplified Wrapper and Interface Generator. https:\/\/github.com\/swig\/swig.","author":"Authors SWIG","year":"2019","unstructured":"SWIG Authors. 2019. SWIG: Simplified Wrapper and Interface Generator. https:\/\/github.com\/swig\/swig."},{"key":"e_1_2_2_19_1","unstructured":"Inc. ClickHouse. 2022. ClickHouse: Fast Open-Source OLAP DBMS. https:\/\/clickhouse.com\/."},{"key":"e_1_2_2_20_1","unstructured":"Dormando. 2022. memcached: a distributed memory object caching system. https:\/\/memcached.org\/."},{"key":"e_1_2_2_21_1","volume-title":"Dagster: An orchestration platform for the development, production, and observation of data assets. https:\/\/github.com\/dagster-io\/dagster.","author":"Inc. Elementl.","year":"2023","unstructured":"Inc. Elementl. 2023. Dagster: An orchestration platform for the development, production, and observation of data assets. https:\/\/github.com\/dagster-io\/dagster."},{"key":"e_1_2_2_22_1","unstructured":"etcd Authors. 2022. etcd: A distributed reliable key-value store for the most critical data of a distributed system. https:\/\/etcd.io\/."},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476311.3476369"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3282488"},{"key":"e_1_2_2_25_1","volume-title":"Scaling Large Production Clusters with Partitioned Synchronization. In 2021 USENIX Annual Technical Conference (USENIX ATC 21)","author":"Feng Yihui","year":"2021","unstructured":"Yihui Feng, Zhi Liu, Yunjian Zhao, Tatiana Jin, Yidi Wu, Yang Zhang, James Cheng, Chao Li, and Tao Guan. 2021. Scaling Large Production Clusters with Partitioned Synchronization. In 2021 USENIX Annual Technical Conference (USENIX ATC 21). 81--97."},{"key":"e_1_2_2_26_1","unstructured":"Linux Foundation. 2015. Data Plane Development Kit (DPDK). http:\/\/www.dpdk.org."},{"key":"e_1_2_2_27_1","volume-title":"Apache Airflow: A platform to programmatically author, schedule, and monitor workflows. https:\/\/airflow.apache.org\/.","author":"Software Foundation The Apache","year":"2022","unstructured":"The Apache Software Foundation. 2022. Apache Airflow: A platform to programmatically author, schedule, and monitor workflows. https:\/\/airflow.apache.org\/."},{"key":"e_1_2_2_28_1","unstructured":"The Apache Software Foundation. 2022. Apache Data Fusion SQL Query Engine. https:\/\/arrow.apache.org\/datafusion\/."},{"key":"e_1_2_2_29_1","volume-title":"Apache Doris: An easy-to-use, high-performance and unified analytical database. https:\/\/doris.apache.org\/.","author":"Software Foundation The Apache","year":"2022","unstructured":"The Apache Software Foundation. 2022. Apache Doris: An easy-to-use, high-performance and unified analytical database. https:\/\/doris.apache.org\/."},{"key":"e_1_2_2_30_1","volume-title":"Apache Dremio: The Easy and Open Data Lakehouse. https:\/\/www.dremio.com\/.","author":"Software Foundation The Apache","year":"2022","unstructured":"The Apache Software Foundation. 2022. Apache Dremio: The Easy and Open Data Lakehouse. https:\/\/www.dremio.com\/."},{"key":"e_1_2_2_31_1","volume-title":"Arrow: A cross-language development platform for in-memory analytics. https:\/\/github.com\/apache\/arrow.","author":"Software Foundation The Apache","year":"2022","unstructured":"The Apache Software Foundation. 2022. Arrow: A cross-language development platform for in-memory analytics. https:\/\/github.com\/apache\/arrow."},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/945445.945450"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/2741948.2741968"},{"key":"e_1_2_2_34_1","volume-title":"10th USENIX Symposium on Operating Systems Design and Implementation (OSDI 12)","author":"Gonzalez Joseph E","year":"2012","unstructured":"Joseph E Gonzalez, Yucheng Low, Haijie Gu, Danny Bickson, and Carlos Guestrin. 2012. Powergraph: Distributed graph-parallel computation on natural graphs. In 10th USENIX Symposium on Operating Systems Design and Implementation (OSDI 12). 17--30."},{"key":"e_1_2_2_35_1","volume-title":"Protocol Buffers: A language-neutral, platform-neutral extensible mechanism for serializing structured data. https:\/\/developers.google.com\/protocol-buffers.","author":"Inc. Google.","year":"2022","unstructured":"Inc. Google. 2022. Protocol Buffers: A language-neutral, platform-neutral extensible mechanism for serializing structured data. https:\/\/developers.google.com\/protocol-buffers."},{"key":"e_1_2_2_36_1","volume-title":"Whiz: Data-Driven Analytics Execution. In 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21)","author":"Grandl Robert","year":"2021","unstructured":"Robert Grandl, Arjun Singhvi, Raajay Viswanathan, and Aditya Akella. 2021. Whiz: Data-Driven Analytics Execution. In 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21)."},{"key":"e_1_2_2_37_1","unstructured":"gRPC Authors. 2022. gRPC: A high performance open source universal RPC framework. https:\/\/grpc.io."},{"key":"e_1_2_2_38_1","volume-title":"Ogb-lsc: A large-scale challenge for machine learning on graphs. arXiv preprint arXiv:2103.09430","author":"Hu Weihua","year":"2021","unstructured":"Weihua Hu, Matthias Fey, Hongyu Ren, Maho Nakata, Yuxiao Dong, and Jure Leskovec. 2021. Ogb-lsc: A large-scale challenge for machine learning on graphs. arXiv preprint arXiv:2103.09430 (2021)."},{"key":"e_1_2_2_39_1","unstructured":"Inc. Juicedata. 2022. JuiceFS: A POSIX HDFS and S3 compatible distributed file system for cloud. https:\/\/juicefs.com\/en\/."},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2670979.2670985"},{"key":"e_1_2_2_41_1","unstructured":"Jingdong Li Zhao Li Jiaming Huang Ji Zhang Xiaoling Wang Xingjian Lu and Jingren Zhou. 2021. Large-scale Fake Click Detection for E-commerce Recommendation Systems. In ICDE."},{"key":"e_1_2_2_42_1","unstructured":"libclang Authors. 2022. libclang: C interface to Clang. https:\/\/clang.llvm.org\/doxygen\/group__CINDEX.html."},{"key":"e_1_2_2_43_1","unstructured":"libfuse authors. 2022. libfuse: The reference implementation of the Linux FUSE (Filesystem in Userspace) interface. https:\/\/github.com\/libfuse\/libfuse."},{"key":"e_1_2_2_44_1","volume-title":"Redis: The open source, in-memory data store. https:\/\/redis.io\/.","author":"Ltd Redis","year":"2022","unstructured":"Redis Ltd. 2022. Redis: The open source, in-memory data store. https:\/\/redis.io\/."},{"key":"e_1_2_2_45_1","unstructured":"The Alibaba Group Holding Ltd. 2022. Mars: a tensor-based unified framework for large-scale data computation. https:\/\/github.com\/mars-project\/mars."},{"key":"e_1_2_2_46_1","unstructured":"Ruotian Luo. 2017. An Image Captioning codebase in PyTorch. https:\/\/github.com\/ruotianluo\/ImageCaptioning.pytorch."},{"key":"e_1_2_2_47_1","volume-title":"15th Workshop on Hot Topics in Operating Systems (HotOS XV ).","author":"McSherry Frank","year":"2015","unstructured":"Frank McSherry, Michael Isard, and Derek G Murray. 2015. Scalability! But at what COST?. In 15th Workshop on Hot Topics in Operating Systems (HotOS XV )."},{"volume-title":"Handbook of cloud computing","author":"Middleton Anthony M","key":"e_1_2_2_48_1","unstructured":"Anthony M Middleton. 2010. Data-intensive technologies for cloud computing. In Handbook of cloud computing. Springer, 83--136."},{"key":"e_1_2_2_49_1","volume-title":"13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)","author":"Moritz Philipp","year":"2018","unstructured":"Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I Jordan, et al . 2018. Ray: A distributed framework for emerging AI applications. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). 561--577."},{"volume-title":"Cocomo ii forum","author":"Nguyen Vu","key":"e_1_2_2_50_1","unstructured":"Vu Nguyen, Sophia Deeds-Rubin, Thomas Tan, and Barry Boehm. 2007. A SLOC counting standard. In Cocomo ii forum, Vol. 2007. Citeseer, 1--16."},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359652"},{"volume-title":"PyTorch: An Imperative Style","author":"Paszke Adam","key":"e_1_2_2_52_1","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems 32. Curran Associates, Inc., 8024--8035. http:\/\/papers.neurips.cc\/paper\/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf"},{"key":"e_1_2_2_53_1","volume-title":"18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21)","author":"Qian Zhengping","year":"2021","unstructured":"Zhengping Qian, Chenqiang Min, Longbin Lai, Yong Fang, Gaofeng Li, Youyang Yao, Bingqing Lyu, Xiaoli Zhou, Zhimin Chen, and Jingren Zhou. 2021. GAIA: A System for Interactive Analysis on Distributed Graphs Using a High-Level Language. In 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21)."},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3012408.3012416"},{"key":"e_1_2_2_55_1","unstructured":"scikit-learn Authors. 2022. scikit-learn: Machine-Learning in Python. https:\/\/scikit-learn.org\/."},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2019.00196"},{"key":"e_1_2_2_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSST.2010.5496972"},{"key":"e_1_2_2_58_1","volume-title":"Thrift: Scalable cross-language services implementation. Facebook white paper 5, 8","author":"Slee Mark","year":"2007","unstructured":"Mark Slee, Aditya Agarwal, and Marc Kwiatkowski. 2007. Thrift: Scalable cross-language services implementation. Facebook white paper 5, 8 (2007), 127."},{"key":"e_1_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/1064979.1064997"},{"key":"e_1_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2010.5447738"},{"key":"e_1_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/3308558.3313411"},{"key":"e_1_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/2509578.2509581"},{"key":"e_1_2_2_63_1","volume-title":"9th USENIX Symposium on Networked Systems Design and Implementation (NSDI 12)","author":"Zaharia Matei","year":"2012","unstructured":"Matei Zaharia, Mosharaf Chowdhury, Tathagata Das, Ankur Dave, Justin Ma, Murphy McCauly, Michael J Franklin, Scott Shenker, and Ion Stoica. 2012. Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing. In 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI 12). 15--28."},{"key":"e_1_2_2_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/2934664"},{"key":"e_1_2_2_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/2934664"},{"key":"e_1_2_2_66_1","doi-asserted-by":"publisher","DOI":"10.14778\/3352063.3352127"},{"key":"e_1_2_2_67_1","volume-title":"12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16)","author":"Zhu Xiaowei","year":"2016","unstructured":"Xiaowei Zhu, Wenguang Chen, Weimin Zheng, and Xiaosong Ma. 2016. Gemini: A computation-centric distributed graph processing system. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16). 301--316."},{"key":"e_1_2_2_68_1","doi-asserted-by":"publisher","DOI":"10.14778\/3384345.3384351"}],"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589780","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3589780","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:22Z","timestamp":1750182562000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589780"}},"issued":{"date-parts":[[2023,6,13]]},"references-count":68,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2023,6,13]]}},"alternative-id":["10.1145\/3589780"],"URL":"https:\/\/doi.org\/10.1145\/3589780","ISSN":["2836-6573"],"issn-type":[{"type":"electronic","value":"2836-6573"}],"published":{"date-parts":[[2023,6,13]]}},{"indexed":{"date-parts":[[2025,8,22]],"date-time":"2025-08-22T02:40:58Z","timestamp":1755830458376,"version":"3.44.0"},"reference-count":58,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2024,5,29]],"date-time":"2024-05-29T00:00:00Z","timestamp":1716940800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2024,5,29]]},"abstract":"<jats:p>Extracting interesting patterns from data is the main objective of Data Mining. In this context, Frequent Itemset Mining has shown its usefulness in providing insights from transactional databases, which, in turn, can be used to gain insights about the structure of Knowledge Graphs. While there have been a lot of advances in the field, due to the NP-hard nature of the problem, the main approaches still struggle when they are faced with large databases with large and sparse vocabularies, such as the ones obtained from graph propositionalizations. There have been efforts to propose parallel algorithms, but, so far, the goal has not been to tackle this source of complexity (i.e., vocabulary size), thus, in this paper, we propose to parallelize frequent itemset mining algorithms by partitioning the database horizontally (i.e., transaction-wise) while not neglecting all the possible vertical information (i.e., item-wise). Instead of relying on pure item co-appearance metrics, we advocate for the adoption of a different approach: modeling databases as documents, where each transaction is a sentence, and each item a word. In this way, we can apply recent language modeling techniques (i.e., word embeddings) to obtain a continuous representation of the database, clusterize it in different partitions, and apply any mining algorithm to them. We show how our proposal leads to informed partitions with a reduced vocabulary size and a reduced entropy (i.e., disorder). This enhances the scalability, allowing us to speed up mining even in very large databases with sparse vocabularies. We have carried out a thorough experimental evaluation over both synthetic and real datasets showing the benefits of our proposal.<\/jats:p>","DOI":"10.1145\/3654987","type":"journal-article","created":{"date-parts":[[2024,5,30]],"date-time":"2024-05-30T09:44:53Z","timestamp":1717062293000},"page":"1-27","source":"Crossref","is-referenced-by-count":0,"title":["Language-Model Based Informed Partition of Databases to Speed Up Pattern Mining"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4239-8785","authenticated-orcid":false,"given":"Carlos","family":"Bobed Lisbona","sequence":"first","affiliation":[{"name":"Aragon Institute of Engineering Research (I3A), University of Zaragoza, Zaragoza, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8531-353X","authenticated-orcid":false,"given":"Jordi","family":"Bernad","sequence":"additional","affiliation":[{"name":"Aragon Institute of Engineering Research (I3A), University of Zaragoza, Zaragoza, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9814-439X","authenticated-orcid":false,"given":"Pierre","family":"Maillot","sequence":"additional","affiliation":[{"name":"University Cote d'Azur Inria, CNRS, I3S, Nice, France"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,5,30]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.5555\/645920.672836"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2003.1211462"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/584792.584888"},{"key":"e_1_2_1_4_1","volume-title":"A Neural Probabilistic Language Model. J. Mach. Learn. Res. 3 (mar","author":"Bengio Yoshua","year":"2003","unstructured":"Yoshua Bengio, R\u00e9jean Ducharme, Pascal Vincent, and Christian Janvin. 2003. A Neural Probabilistic Language Model. J. Mach. Learn. Res. 3 (mar 2003), 1137--1155."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10618-021-00749-5"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.3233\/SW-200368"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00051"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1002\/widm.1074"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/2033831.2033854"},{"key":"e_1_2_1_10_1","unstructured":"Tom B. Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell Sandhini Agarwal Ariel Herbert-Voss Gretchen Krueger Tom Henighan Rewon Child Aditya Ramesh Daniel M. Ziegler Jeffrey Wu Clemens Winter Christopher Hesse Mark Chen Eric Sigler Mateusz Litwin Scott Gray Benjamin Chess Jack Clark Christopher Berner Sam McCandlish Alec Radford Ilya Sutskever and Dario Amodei. 2020. Language Models are Few-Shot Learners. http:\/\/arxiv.org\/abs\/2005.14165"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2007.190649"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2003.1250893"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/978--1--4419-"},{"volume-title":"The LUCS--KDD discretised\/normalised ARM and CARM data library. Department of Computer Science","author":"Coenen F","key":"e_1_2_1_14_1","unstructured":"F Coenen. 2003. The LUCS--KDD discretised\/normalised ARM and CARM data library. Department of Computer Science, The University of Liverpool, UK. https:\/\/cgi.csc.liv.ac.uk\/~frans\/KDD\/Software\/LUCS-KDD-DN\/DataSets\/dataSets.html"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19--1423"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.3836"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2018.06.060"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3439771"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.5555\/3001460.3001507"},{"key":"e_1_2_1_20_1","volume-title":"Studies in Linguistic Analysis","author":"Firth J.","year":"1930","unstructured":"J. Firth. 1957. Studies in Linguistic Analysis. Oxford: Blackwell, Chapter A Synopsis of Linguistic Theory, 1930--1955, 1--32."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447548.3467348"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.2307\/1403797"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1002\/widm.1207"},{"volume-title":"Principles of Data Mining and Knowledge Discovery","author":"Giannotti Fosca","key":"e_1_2_1_24_1","unstructured":"Fosca Giannotti, Cristian Gozzi, and Giuseppe Manco. 2002. Clustering Transactional Data. In Principles of Data Mining and Knowledge Discovery, Tapio Elomaa, Heikki Mannila, and Hannu Toivonen (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 175--187."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3097983.3098034"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1080\/00437956"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","unstructured":"Z. R. Hesabi Z. Tari A. Goscinski A. Fahad I. Khalil and C. Queiroz. 2015. Data Summarization Techniques for Big Data-A Survey. Springer New York New York NY 1109--1152. https:\/\/doi.org\/10.1007\/978--1--4939--2092--1_38","DOI":"10.1007\/978--1--4939--2092--1_38"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447772"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-41398-8_20"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-011-0672--7"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/TBDATA.2019.2921572"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/978--1--4684--2001--2_9"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1007\/978--3--662-04599--2_11"},{"key":"e_1_2_1_34_1","volume-title":"S\u00f6ren Auer, et al.","author":"Lehmann Jens","year":"2015","unstructured":"Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, Dimitris Kontokostas, Pablo N Mendes, Sebastian Hellmann, Mohamed Morsey, Patrick Van Kleef, S\u00f6ren Auer, et al. 2015. Dbpedia--a large-scale, multilingual knowledge base extracted from wikipedia. Semantic web 6, 2 (2015), 167--195."},{"key":"e_1_2_1_35_1","volume-title":"Proceedings of the 27th International Conference on Neural Information Processing Systems -","volume":"2","author":"Levy Omer","year":"2014","unstructured":"Omer Levy and Yoav Goldberg. 2014. Neural word embedding as implicit matrix factorization. In Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2 (Montreal, Canada) (NIPS'14). MIT Press, Cambridge, MA, USA, 2177--2185."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.3904"},{"key":"e_1_2_1_37_1","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. http:\/\/arxiv.org\/abs\/1907. 11692"},{"key":"e_1_2_1_38_1","volume-title":"Proceedings of the fifth Berkeley symposium on mathematical statistics and probability","volume":"1","author":"James","unstructured":"James MacQueen et al. 1967. Some methods for classification and analysis of multivariate observations. In Proceedings of the fifth Berkeley symposium on mathematical statistics and probability, Vol. 1. Oakland, CA, USA, 281--297."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3167132.3167342"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2017.106"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1162\/153244303322533223"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/BigData.2013.6691742"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.is.2021.101754"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_2_1_45_1","volume-title":"Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks. ELRA, Valletta, Malta, 45--50","author":"Petr Sojka Radim","year":"2010","unstructured":"Radim ?eh??ek and Petr Sojka. 2010. Software Framework for Topic Modelling with Large Corpora. In Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks. ELRA, Valletta, Malta, 45--50. http:\/\/is.muni.cz\/ publication\/884893\/en."},{"volume-title":"Minimum Description Length Principle","author":"Rissanen Jorma","key":"e_1_2_1_46_1","unstructured":"Jorma Rissanen. 2014. Minimum Description Length Principle. In Wiley StatsRef: Statistics Reference Online. John Wiley & Sons, Inc."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.3233\/SW-180317"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1214\/aos\/1176344136"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611972825.21"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10618-010-0202-x"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/319950.320054"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1007\/s40745-015-0040--1"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSMC.2015"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2016.2560176"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/775047.775149"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.is.2004.11.008"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/69.846291"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER.2011.61"}],"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3654987","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3654987","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T14:39:26Z","timestamp":1755787166000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3654987"}},"issued":{"date-parts":[[2024,5,29]]},"references-count":58,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,5,29]]}},"alternative-id":["10.1145\/3654987"],"URL":"https:\/\/doi.org\/10.1145\/3654987","ISSN":["2836-6573"],"issn-type":[{"type":"electronic","value":"2836-6573"}],"published":{"date-parts":[[2024,5,29]]}},{"indexed":{"date-parts":[[2025,8,24]],"date-time":"2025-08-24T01:42:38Z","timestamp":1755999758758,"version":"3.41.0"},"reference-count":70,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2023,5,26]],"date-time":"2023-05-26T00:00:00Z","timestamp":1685059200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["IIS-1910014"],"award-info":[{"award-number":["IIS-1910014"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2023,5,26]]},"abstract":"<jats:p>Most data analytical pipelines often encounter the problem of querying inconsistent data that violate pre-determined integrity constraints. Data cleaning is an extensively studied paradigm that singles out a consistent repair of the inconsistent data. Consistent query answering (CQA) is an alternative approach to data cleaning that asks for all tuples guaranteed to be returned by a given query on all (in most cases, exponentially many) repairs of the inconsistent data. In this paper, we identify a class of acyclic select-project-join (SPJ) queries for which CQA can be solved via SQL rewriting with a linear time guarantee. Our rewriting method can be viewed as a generalization of Yannakakis' algorithm for acyclic joins to the inconsistent setting. We present LinCQA, a system that takes as input any query in our class and outputs rewritings in both SQL and non-recursive Datalog with negation. We show that LinCQA often outperforms the existing CQA systems on both synthetic and real-world workloads, and in some cases, by orders of magnitude.<\/jats:p>","DOI":"10.1145\/3588718","type":"journal-article","created":{"date-parts":[[2023,5,30]],"date-time":"2023-05-30T17:42:05Z","timestamp":1685468525000},"page":"1-25","source":"Crossref","is-referenced-by-count":2,"title":["LinCQA: Faster Consistent Query Answering with Linear Time Guarantees"],"prefix":"10.1145","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-1131-3761","authenticated-orcid":false,"given":"Zhiwei","family":"Fan","sequence":"first","affiliation":[{"name":"University of Wisconsin-Madison, Madison, WI, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6309-1702","authenticated-orcid":false,"given":"Paraschos","family":"Koutris","sequence":"additional","affiliation":[{"name":"University of Wisconsin-Madison, Madison, WI, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5214-4688","authenticated-orcid":false,"given":"Xiating","family":"Ouyang","sequence":"additional","affiliation":[{"name":"University of Wisconsin-Madison, Madison, WI, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8216-273X","authenticated-orcid":false,"given":"Jef","family":"Wijsen","sequence":"additional","affiliation":[{"name":"University of Mons, Mons, Belgium"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,5,30]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0022-0000(05)80071-6"},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2008.4497507"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1559845.1559871"},{"volume-title":"Consistent Query Answers in Inconsistent Databases","author":"Arenas Marcelo","key":"e_1_2_2_4_1","unstructured":"Marcelo Arenas, Leopoldo E. Bertossi, and Jan Chomicki. 1999. Consistent Query Answers in Inconsistent Databases. In PODS. ACM Press, 68--79."},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1017\/S1471068403001832"},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.14778\/2850578.2850579"},{"key":"e_1_2_2_7_1","volume-title":"ICDT (LIPIcs","volume":"397","author":"Barcel\u00f3 Pablo","year":"2015","unstructured":"Pablo Barcel\u00f3 and Ga\u00eblle Fontaine. 2015. On the Data Complexity of Consistent Query Answering over Graph Databases. In ICDT (LIPIcs, Vol. 31). Schloss Dagstuhl - Leibniz-Zentrum f\u00fcr Informatik, 380--397."},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcss.2017.03.015"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/2402.322389"},{"key":"e_1_2_2_10_1","volume-title":"Query-Oriented Data Cleaning with Oracles. In SIGMOD Conference. ACM, 1199--1214","author":"Bergman Moria","year":"2015","unstructured":"Moria Bergman, Tova Milo, Slava Novgorodov, and Wang Chiew Tan. 2015. Query-Oriented Data Cleaning with Oracles. In SIGMOD Conference. ACM, 1199--1214."},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00224-012-9402-7"},{"key":"e_1_2_2_12_1","doi-asserted-by":"crossref","unstructured":"Leopoldo E. Bertossi. 2019. Database Repairs and Consistent Query Answering: Origins and Further Developments. In PODS. ACM 48--58.","DOI":"10.1145\/3294052.3322190"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00224-012-9402-7"},{"key":"e_1_2_2_14_1","volume-title":"Certifiable Robustness for Naive Bayes Classifiers. CoRR abs\/2303.04811","author":"Bian Song","year":"2023","unstructured":"Song Bian, Xiating Ouyang, Zhiwei Fan, and Paraschos Koutris. 2023. Certifiable Robustness for Naive Bayes Classifiers. CoRR abs\/2303.04811 (2023)."},{"volume-title":"Conditional Functional Dependencies for Data Cleaning","author":"Bohannon Philip","key":"e_1_2_2_15_1","unstructured":"Philip Bohannon, Wenfei Fan, Floris Geerts, Xibei Jia, and Anastasios Kementsietsidis. 2007. Conditional Functional Dependencies for Data Cleaning. In ICDE. IEEE Computer Society, 746--755."},{"key":"e_1_2_2_16_1","doi-asserted-by":"crossref","unstructured":"Marco Calautti Marco Console and Andreas Pieris. 2021. Benchmarking Approximate Consistent Query Answering. In PODS. ACM 233--246.","DOI":"10.1145\/3452021.3458309"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.14778\/1453856.1453935"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ic.2004.04.007"},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-24741-8_53"},{"key":"e_1_2_2_20_1","volume-title":"Data Cleaning: Overview and Emerging Challenges. In SIGMOD Conference. ACM, 2201--2206","author":"Chu Xu","year":"2016","unstructured":"Xu Chu, Ihab F. Ilyas, Sanjay Krishnan, and Jiannan Wang. 2016. Data Cleaning: Overview and Emerging Challenges. In SIGMOD Conference. ACM, 2201--2206."},{"volume-title":"Holistic data cleaning: Putting violations into context","author":"Chu Xu","key":"e_1_2_2_21_1","unstructured":"Xu Chu, Ihab F. Ilyas, and Paolo Papotti. 2013. Holistic data cleaning: Putting violations into context. In ICDE. IEEE Computer Society, 458--469."},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2749431"},{"key":"e_1_2_2_23_1","unstructured":"CloudLab 2018. https:\/\/www.cloudlab.us\/."},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-23786-7_19"},{"key":"e_1_2_2_25_1","unstructured":"Akhil Anand Dixit. 2021. Answering Queries Over Inconsistent Databases Using SAT Solvers. Ph. D. Dissertation. UC Santa Cruz."},{"key":"e_1_2_2_26_1","volume-title":"Kolaitis","author":"Dixit Akhil A.","year":"2019","unstructured":"Akhil A. Dixit and Phokion G. Kolaitis. 2019. A SAT-Based System for Consistent Query Answering. In SAT (Lecture Notes in Computer Science, Vol. 11628). Springer, 117--135."},{"key":"e_1_2_2_27_1","volume-title":"Kolaitis","author":"Dixit Akhil A.","year":"2021","unstructured":"Akhil A. Dixit and Phokion G. Kolaitis. 2021. Consistent Answers of Aggregation Queries using SAT Solvers. CoRR abs\/2103.03314 (2021)."},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.14778\/2536274.2536280"},{"key":"e_1_2_2_29_1","first-page":"1","article-title":"Certifiable Robustness for Nearest Neighbor Classifiers. In ICDT (LIPIcs, Vol. 220)","volume":"6","author":"Fan Austen Z.","year":"2022","unstructured":"Austen Z. Fan and Paraschos Koutris. 2022. Certifiable Robustness for Nearest Neighbor Classifiers. In ICDT (LIPIcs, Vol. 220). Schloss Dagstuhl - Leibniz-Zentrum f\u00fcr Informatik, 6:1--6:20.","journal-title":"Schloss Dagstuhl - Leibniz-Zentrum f\u00fcr Informatik"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2208.12339"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/1066157.1066176"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcss.2006.10.013"},{"volume-title":"A Hybrid Data Cleaning Framework Using Markov Logic Networks (Extended Abstract)","author":"Ge Congcong","key":"e_1_2_2_33_1","unstructured":"Congcong Ge, Yunjun Gao, Xiaoye Miao, Bin Yao, and Haobo Wang. 2021. A Hybrid Data Cleaning Framework Using Markov Logic Networks (Extended Abstract). In ICDE. IEEE, 2344--2345."},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.14778\/2536360.2536363"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2003.1245280"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.14778\/3231751.3231766"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.2147\/CLEP.S242080"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.14778\/3430915.3430917"},{"volume-title":"Inconsistency resolution in online databases","author":"Katsis Yannis","key":"e_1_2_2_39_1","unstructured":"Yannis Katsis, Alin Deutsch, Yannis Papakonstantinou, and Vasilis Vassalos. 2010. Inconsistency resolution in online databases. In ICDE. IEEE Computer Society, 1205--1208."},{"key":"e_1_2_2_40_1","doi-asserted-by":"crossref","unstructured":"Aziz Amezian El Khalfioui Jonathan Joertz Dorian Labeeuw Ga\u00ebtan Staquet and Jef Wijsen. 2020. Optimization of Answer Set Programs for Consistent Query Answering by Means of First-Order Rewriting. In CIKM. ACM 25--34.","DOI":"10.1145\/3340531.3411911"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2747646"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2021.3062318"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipl.2011.10.018"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.14778\/2536336.2536341"},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.14778\/2536336.2536341"},{"key":"e_1_2_2_46_1","doi-asserted-by":"crossref","unstructured":"Paraschos Koutris Xiating Ouyang and Jef Wijsen. 2021. Consistent Query Answering for Primary Keys on Path Queries. In PODS. ACM 215--232.","DOI":"10.1145\/3452021.3458334"},{"key":"e_1_2_2_47_1","unstructured":"Paraschos Koutris and Dan Suciu. 2014. A Dichotomy on the Complexity of Consistent Query Answering for Atoms with Simple Keys. In ICDT. OpenProceedings.org 165--176."},{"key":"e_1_2_2_48_1","doi-asserted-by":"crossref","unstructured":"Paraschos Koutris and Jef Wijsen. 2015. The Data Complexity of Consistent Query Answering for Self-Join-Free Conjunctive Queries Under Primary Key Constraints. In PODS. ACM 17--29.","DOI":"10.1145\/2745754.2745769"},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3068334"},{"key":"e_1_2_2_50_1","doi-asserted-by":"crossref","unstructured":"Paraschos Koutris and Jef Wijsen. 2018. Consistent Query Answering for Primary Keys and Conjunctive Queries with Negated Atoms. In PODS. ACM 209--224.","DOI":"10.1145\/3196959.3196982"},{"key":"e_1_2_2_51_1","first-page":"1","article-title":"Consistent Query Answering for Primary Keys in Logspace. In ICDT (LIPIcs, Vol. 127)","volume":"23","author":"Koutris Paraschos","year":"2019","unstructured":"Paraschos Koutris and Jef Wijsen. 2019. Consistent Query Answering for Primary Keys in Logspace. In ICDT (LIPIcs, Vol. 127). Schloss Dagstuhl - Leibniz-Zentrum f\u00fcr Informatik, 23:1--23:19.","journal-title":"Schloss Dagstuhl - Leibniz-Zentrum f\u00fcr Informatik"},{"key":"e_1_2_2_52_1","doi-asserted-by":"crossref","unstructured":"Paraschos Koutris and Jef Wijsen. 2020. First-Order Rewritability in Consistent Query Answering with Respect to Multiple Keys. In PODS. ACM 113--129.","DOI":"10.1145\/3375395.3387654"},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00224-020-09985-6"},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.14778\/2994509.2994514"},{"volume-title":"CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification Tasks","author":"Li Peng","key":"e_1_2_2_55_1","unstructured":"Peng Li, Xi Rao, Jennifer Blase, Yue Zhang, Xu Chu, and Ce Zhang. 2021. CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification Tasks. In ICDE. IEEE, 13--24."},{"key":"e_1_2_2_56_1","volume-title":"Bertossi","author":"Lopatenko Andrei","year":"2007","unstructured":"Andrei Lopatenko and Leopoldo E. Bertossi. 2007. Complexity of Consistent Query Answering in Databases Under Cardinality-Based and Incremental Repair Semantics. In ICDT (Lecture Notes in Computer Science, Vol. 4353). Springer, 179--193."},{"key":"e_1_2_2_57_1","doi-asserted-by":"publisher","DOI":"10.1017\/S1471068415000320"},{"key":"e_1_2_2_58_1","volume-title":"Bertossi","author":"Marileo M\u00f3nica Caniup\u00e1n","year":"2005","unstructured":"M\u00f3nica Caniup\u00e1n Marileo and Leopoldo E. Bertossi. 2005. Optimizing repair programs for consistent query answering. In SCCC. IEEE Computer Society, 3--12."},{"key":"e_1_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.1007\/978--3--642--10424--4_17"},{"key":"e_1_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/369275.369291"},{"key":"e_1_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.14778\/2856318.2856325"},{"key":"e_1_2_2_62_1","first-page":"3","article-title":"Data cleaning: Problems and current approaches","volume":"23","author":"Rahm Erhard","year":"2000","unstructured":"Erhard Rahm and Hong Hai Do. 2000. Data cleaning: Problems and current approaches. IEEE Data Eng. Bull. 23, 4 (2000), 3--13.","journal-title":"IEEE Data Eng. Bull."},{"key":"e_1_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.14778\/3137628.3137631"},{"key":"e_1_2_2_64_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476249.3476301"},{"key":"e_1_2_2_65_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.is.2012.08.005"},{"key":"e_1_2_2_66_1","volume-title":"Chen Jason Zhang, Yatao Li, and Lei Chen.","author":"Tong Yongxin","year":"2014","unstructured":"Yongxin Tong, Caleb Chen Cao, Chen Jason Zhang, Yatao Li, and Lei Chen. 2014. CrowdCleaner: Data cleaning for multi-version data on the web via crowdsourcing. In ICDE. IEEE Computer Society, 1182--1185."},{"key":"e_1_2_2_67_1","doi-asserted-by":"crossref","unstructured":"Jef Wijsen. 2010. On the first-order expressibility of computing certain answers to conjunctive queries over uncertain databases. In PODS. ACM 179--190.","DOI":"10.1145\/1807085.1807111"},{"key":"e_1_2_2_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/2188349.2188351"},{"key":"e_1_2_2_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/3377391.3377393"},{"key":"e_1_2_2_70_1","volume-title":"7th International Conference, September 9--11","author":"Yannakakis Mihalis","year":"1981","unstructured":"Mihalis Yannakakis. 1981. Algorithms for Acyclic Database Schemes. In Very Large Data Bases, 7th International Conference, September 9--11, 1981, Cannes, France, Proceedings. IEEE Computer Society, 82--94."}],"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3588718","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3588718","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3588718","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:47:35Z","timestamp":1750178855000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3588718"}},"issued":{"date-parts":[[2023,5,26]]},"references-count":70,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2023,5,26]]}},"alternative-id":["10.1145\/3588718"],"URL":"https:\/\/doi.org\/10.1145\/3588718","ISSN":["2836-6573"],"issn-type":[{"type":"electronic","value":"2836-6573"}],"published":{"date-parts":[[2023,5,26]]}},{"indexed":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T15:47:44Z","timestamp":1778255264549,"version":"3.51.4"},"reference-count":42,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"DOI":"10.13039\/501100006374","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62372122, 92270123"],"award-info":[{"award-number":["62372122, 92270123"]}],"id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Research Grants Council, Hong Kong SAR, China","award":["15225921, 15208923, 25207224, and C2003-23Y"],"award-info":[{"award-number":["15225921, 15208923, 25207224, and C2003-23Y"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2025,6,17]]},"abstract":"<jats:p>\n                    The increasing collection and analysis of personal data driven by digital technologies has raised concerns about individual privacy. Local Differential Privacy (LDP) has emerged as a promising solution to provide rigorous privacy guarantee for users, without relying on a trusted data collector. In the context of LDP, range mean estimation over numerical values is an important yet challenging problem. Simply applying existing work may introduce overly large noise sensitivity, since all of them focus on statistical tasks (e.g., mean or distribution) across the entire domain. In this paper, we propose a novel framework for &lt;u&gt;Priv&lt;\/u&gt;ate &lt;u&gt;R&lt;\/u&gt;ange &lt;u&gt;M&lt;\/u&gt;ean (\n                    <jats:italic toggle=\"yes\">PrivRM<\/jats:italic>\n                    ) estimation under LDP. Two implementations of the framework, namely\n                    <jats:italic toggle=\"yes\">PrivRM<\/jats:italic>\n                    <jats:sup>I<\/jats:sup>\n                    and\n                    <jats:italic toggle=\"yes\">PrivRM<\/jats:italic>\n                    <jats:sup>*<\/jats:sup>\n                    , are developed, which are adaptable to all existing numerical value perturbation mechanisms. As an optimization of the framework, we also propose a distribution-aware Adaptive Adjustment (AA) strategy to dynamically confine the perturbation space for skewed data distributions. Extensive experimental results show that under the same privacy guarantee and query range, our framework\n                    <jats:italic toggle=\"yes\">PrivRM<\/jats:italic>\n                    significantly improve over existing solutions.\n                  <\/jats:p>","DOI":"10.1145\/3725414","type":"journal-article","created":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T21:23:29Z","timestamp":1750281809000},"page":"1-26","source":"Crossref","is-referenced-by-count":3,"title":["<i>PrivRM<\/i>\n                    : A Framework for Range Mean Estimation under Local Differential Privacy"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-7279-4352","authenticated-orcid":false,"given":"Liantong","family":"Yu","sequence":"first","affiliation":[{"name":"The Hong Kong Polytechnic University, Hong Kong SAR, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1547-2847","authenticated-orcid":false,"given":"Qingqing","family":"Ye","sequence":"additional","affiliation":[{"name":"The Hong Kong Polytechnic University, Hong Kong SAR, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5478-896X","authenticated-orcid":false,"given":"Rong","family":"Du","sequence":"additional","affiliation":[{"name":"The Hong Kong Polytechnic University, Hong Kong SAR, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,6,18]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2019. New york taxi trip record data. https:\/\/www.nyc.gov\/site\/tlc\/about\/tlc-trip-record-data.page"},{"key":"e_1_2_1_2_1","unstructured":"2020. Frequent itemset mining dataset repository. http:\/\/fimi.ua.ac.be\/data\/"},{"key":"e_1_2_1_3_1","unstructured":"2024. PrivRM: A Framework for Range Mean Estimation under Local Differential Privacy (Full version). https:\/\/github.com\/Yunicom-YLT\/Range-mean"},{"key":"e_1_2_1_4_1","unstructured":"2024. USA Real Estate Dataset. https:\/\/www.kaggle.com\/datasets\/ahmedshahriarsakib\/usa-real-estate-dataset"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476249.3476277"},{"key":"e_1_2_1_6_1","volume-title":"Practical locally private heavy hitters. Advances in Neural Information Processing Systems 30","author":"Bassily Raef","year":"2017","unstructured":"Raef Bassily, Kobbi Nissim, Uri Stemmer, and Abhradeep Guha Thakurta. 2017. Practical locally private heavy hitters. Advances in Neural Information Processing Systems 30 (2017)."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2746539.2746632"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3344722"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3197390"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3196906"},{"key":"e_1_2_1_11_1","volume-title":"Collecting telemetry data privately. Advances in Neural Information Processing Systems 30","author":"Ding Bolin","year":"2017","unstructured":"Bolin Ding, Janardhan Kulkarni, and Sergey Yekhanin. 2017. Collecting telemetry data privately. Advances in Neural Information Processing Systems 30 (2017)."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3460120.3485668"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE55515.2023.00169"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.14778\/3594512.3594520"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE53745.2022.00035"},{"key":"e_1_2_1_16_1","volume-title":"Local privacy and statistical minimax rates. In 2013 IEEE 54th annual symposium on foundations of computer science","author":"Duchi John C","unstructured":"John C Duchi, Michael I Jordan, and Martin J Wainwright. 2013. Local privacy and statistical minimax rates. In 2013 IEEE 54th annual symposium on foundations of computer science. IEEE, 429--438."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.2017.1389735"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2666468"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/11681878_14"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2660267.2660348"},{"key":"e_1_2_1_21_1","volume-title":"29th USENIX security symposium (USENIX security 20). 967--984.","author":"Gu Xiaolan","unstructured":"Xiaolan Gu, Ming Li, Yueqiang Cheng, Li Xiong, and Yang Cao. 2020. {PCKV}: Locally differentially private correlated {Key-Value} data collection with optimized utility. In 29th USENIX security symposium (USENIX security 20). 967--984."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.14778\/1920841.1920970"},{"key":"e_1_2_1_23_1","volume-title":"Extremal mechanisms for local differential privacy. Advances in neural information processing systems 27","author":"Kairouz Peter","year":"2014","unstructured":"Peter Kairouz, Sewoong Oh, and Pramod Viswanath. 2014. Extremal mechanisms for local differential privacy. Advances in neural information processing systems 27 (2014)."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3300102"},{"key":"e_1_2_1_25_1","volume-title":"32nd USENIX Security Symposium (USENIX Security 23)","author":"Li Xiaoguang","year":"2023","unstructured":"Xiaoguang Li, Ninghui Li, Wenhai Sun, Neil Zhenqiang Gong, and Hui Li. 2023. Fine-grained poisoning attack to local differential privacy protocols for mean and variance estimation. In 32nd USENIX Security Symposium (USENIX Security 23). 1739--1756."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3389700"},{"key":"e_1_2_1_27_1","unstructured":"Alexander McFarlane Mood. 1950. Introduction to the Theory of Statistics. (1950)."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133956.3134086"},{"key":"e_1_2_1_29_1","first-page":"705","article-title":"Emoji frequency detection and deep link frequency","volume":"9","author":"Thakurta Abhradeep Guha","year":"2017","unstructured":"Abhradeep Guha Thakurta, Andrew H Vyrros, Umesh S Vaishampayan, Gaurav Kapoor, Julien Freudinger, Vipul Ved Prakash, Arnaud Legendre, and Steven Duplinsky. 2017. Emoji frequency detection and deep link frequency. US Patent 9,705,908.","journal-title":"US Patent"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2019.00063"},{"key":"e_1_2_1_31_1","volume-title":"26th USENIX Security Symposium (USENIX Security 17)","author":"Wang Tianhao","year":"2017","unstructured":"Tianhao Wang, Jeremiah Blocki, Ninghui Li, and Somesh Jha. 2017. Locally differentially private protocols for frequency estimation. In 26th USENIX Security Symposium (USENIX Security 17). 729--745."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3319891"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.1965.10480775"},{"key":"e_1_2_1_34_1","volume-title":"31st USENIX Security Symposium (USENIX Security 22)","author":"Wu Yongji","year":"2022","unstructured":"Yongji Wu, Xiaoyu Cao, Jinyuan Jia, and Neil Zhenqiang Gong. 2022. Poisoning Attacks to Local Differential Privacy Protocols for {Key-Value} Data. In 31st USENIX Security Symposium (USENIX Security 22). 519--536."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2010.247"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.2018.2809790"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2020.3047124"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM53939.2023.10229063"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM42981.2021.9488899"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP.2019.00018"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.14778\/3603581.3603597"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3243734.3243742"}],"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3725414","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T18:58:43Z","timestamp":1774983523000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3725414"}},"issued":{"date-parts":[[2025,6,17]]},"references-count":42,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,6,17]]}},"alternative-id":["10.1145\/3725414"],"URL":"https:\/\/doi.org\/10.1145\/3725414","ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"published":{"date-parts":[[2025,6,17]]}},{"indexed":{"date-parts":[[2026,2,8]],"date-time":"2026-02-08T09:24:21Z","timestamp":1770542661613,"version":"3.49.0"},"reference-count":52,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2024,5,29]],"date-time":"2024-05-29T00:00:00Z","timestamp":1716940800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100006374","name":"NSF","doi-asserted-by":"publisher","award":["2337806"],"award-info":[{"award-number":["2337806"]}],"id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2024,5,29]]},"abstract":"<jats:p>B-trees are widely recognized as one of the most important index structures in database systems, providing efficient query processing capabilities. Over the past few decades, many techniques have been developed to enhance the efficiency of B-trees from various perspectives. Among them, B-tree compression is an important technique introduced as early as the 1970s to improve both space efficiency and query performance. Since then, several B-tree compression techniques have been developed. However, to our surprise, we have found that these B-tree compression techniques were never compared against each other in prior works. Consequently, many important questions remain unanswered, such as whether B-tree compression is truly effective or not. If it is effective, under what scenarios and which B-tree compression methods should be employed? In this paper, we conduct the first experimental evaluation of seven widely used B-tree compression techniques using both synthetic and real datasets. Based on our evaluation, we present lessons and insights that can be leveraged to guide system design decisions in modern databases regarding the use of B-tree compression.<\/jats:p>","DOI":"10.1145\/3654972","type":"journal-article","created":{"date-parts":[[2024,5,30]],"date-time":"2024-05-30T09:44:53Z","timestamp":1717062293000},"page":"1-25","source":"Crossref","is-referenced-by-count":5,"title":["Revisiting B-tree Compression: An Experimental Study"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-2171-5036","authenticated-orcid":false,"given":"Chuqing","family":"Gao","sequence":"first","affiliation":[{"name":"Purdue University, West Lafayette, IN, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3932-6011","authenticated-orcid":false,"given":"Shreya","family":"Ballijepalli","sequence":"additional","affiliation":[{"name":"Purdue University, West Lafayette, IN, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3039-1175","authenticated-orcid":false,"given":"Jianguo","family":"Wang","sequence":"additional","affiliation":[{"name":"Purdue University, West Lafayette, IN, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,5,30]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"2007. WEBSPAM-UK2007 Dataset. https:\/\/chato.cl\/webspam\/datasets\/uk2007\/"},{"key":"e_1_2_2_2_1","unstructured":"2008. SNAP Memetracker Dataset. https:\/\/www.kaggle.com\/datasets\/snap\/snap-memetracker"},{"key":"e_1_2_2_3_1","unstructured":"2022. The Default Page Size Change of SQLite 3.12.0. https:\/\/www.sqlite.org\/pgszchng2016.html"},{"key":"e_1_2_2_4_1","unstructured":"2022. Source Code of WiredTiger's B-Tree Implementation. https:\/\/github.com\/wiredtiger\/wiredtiger\/tree\/develop\/ src\/btree"},{"key":"e_1_2_2_5_1","unstructured":"2023. CREATE INDEX Statement in SAP HANA (https:\/\/help.sap.com\/docs\/SAP_HANA_PLATFORM\/ 4fe29514fd584807ac9f2a04f6754767\/20d44b4175191014a940afff4b47c7ea.html)."},{"key":"e_1_2_2_6_1","unstructured":"2023. Database Page Layout in PostgreSQL 16. https:\/\/www.postgresql.org\/docs\/current\/storage-page-layout.html"},{"key":"e_1_2_2_7_1","unstructured":"2023. MyISAM Source Code in MySQL. https:\/\/github.com\/mysql\/mysql-server\/tree\/ a246bad76b9271cb4333634e954040a970222e0a\/storage\/myisam"},{"key":"e_1_2_2_8_1","unstructured":"2023. MySQL Reference Manual. https:\/\/dev.mysql.com\/doc\/refman\/8.0\/en\/key-space.html#: :text=Prefix% 20compression%20is%20used%20on when%20you%20create%20the%20table."},{"key":"e_1_2_2_9_1","unstructured":"2023. TPC-H Benchmark. https:\/\/www.tpc.org\/tpch\/"},{"key":"e_1_2_2_10_1","unstructured":"2023. WiredTiger Documentation. https:\/\/source.wiredtiger.com\/11.1.0\/file_formats.html#file_formats_compression"},{"key":"e_1_2_2_11_1","doi-asserted-by":"crossref","unstructured":"Daniel J. Abadi Samuel Madden and Miguel Ferreira. 2006. Integrating compression and execution in column-oriented database systems. In SIGMOD. 671--682.","DOI":"10.1145\/1142473.1142548"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1007\/s007780050031"},{"key":"e_1_2_2_13_1","doi-asserted-by":"crossref","unstructured":"G. Antoshenkov D. Lomet and J. Murray. 1996. Order Preserving String Compression. In ICDE. 655--663.","DOI":"10.1109\/ICDE.1996.492216"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF00288683"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/320521.320530"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.14778\/1687553.1687573"},{"key":"e_1_2_2_17_1","doi-asserted-by":"crossref","unstructured":"Carsten Binnig Stefan Hildenbrand and Franz F\u00e4rber. 2009. Dictionary-based Order-preserving String Compression for Main Memory Column Stores. In SIGMOD. 283--296.","DOI":"10.1145\/1559845.1559877"},{"key":"e_1_2_2_18_1","doi-asserted-by":"crossref","unstructured":"Philip Bohannon Peter McIlroy and Rajeev Rastogi. 2001. Main-Memory Index Structures with Fixed-Size Partial Keys. In SIGMOD. 163--174.","DOI":"10.1145\/375663.375681"},{"key":"e_1_2_2_19_1","unstructured":"Lars Breddemann. 2020. What is CPB-Tree in SAP HANA? https:\/\/www.lbreddemann.org\/what-is-cpb-tree-in-saphana\/"},{"key":"e_1_2_2_20_1","volume-title":"Bowman","author":"Bumbulis Peter","year":"2002","unstructured":"Peter Bumbulis and Ivan T. Bowman. 2002. A Compact B-tree. In SIGMOD. 533--541."},{"key":"e_1_2_2_21_1","volume-title":"Converged Index: The Secret Sauce Behind Rockset's Fast Queries. https:\/\/rockset.com\/blog\/ converged-indexing-the-secret-sauce-behind-rocksets-fast-queries\/","author":"Canadi Igor","year":"2019","unstructured":"Igor Canadi. 2019. Converged Index: The Secret Sauce Behind Rockset's Fast Queries. https:\/\/rockset.com\/blog\/ converged-indexing-the-secret-sauce-behind-rocksets-fast-queries\/"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.14778\/3648160.3648180"},{"key":"e_1_2_2_23_1","doi-asserted-by":"crossref","unstructured":"Zhiyuan Chen Johannes Gehrke and Flip Korn. 2001. Query Optimization In Compressed Database Systems. In SIGMOD. 271--282.","DOI":"10.1145\/375663.375692"},{"key":"e_1_2_2_24_1","unstructured":"Gregg Christman. 2022. Compression Features Included with Oracle Database Enterprise Edition. https:\/\/blogs.oracle. com\/dbstorage\/post\/compression-features-included-with-oracle-database-enterprise-edition"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/356770.356776"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2463676.2463710"},{"key":"e_1_2_2_27_1","volume-title":"Jia Yu, Chi Wang, Jaeyoung Do, Yinan Li, Hantian Zhang, Badrish Chandramouli, Johannes Gehrke, Donald Kossmann, David B. Lomet, and Tim Kraska.","author":"Ding Jialin","year":"2020","unstructured":"Jialin Ding, Umar Farooq Minhas, Jia Yu, Chi Wang, Jaeyoung Do, Yinan Li, Hantian Zhang, Badrish Chandramouli, Johannes Gehrke, Donald Kossmann, David B. Lomet, and Tim Kraska. 2020. ALEX: An Updatable Adaptive Learned Index. In SIGMOD. 969--984."},{"key":"e_1_2_2_28_1","volume-title":"A Survey of B-tree Locking Techniques. TODS 35, 3","author":"Graefe Goetz","year":"2010","unstructured":"Goetz Graefe. 2010. A Survey of B-tree Locking Techniques. TODS 35, 3 (2010), 16:1--16:26."},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1561\/1900000028"},{"key":"e_1_2_2_30_1","doi-asserted-by":"crossref","unstructured":"Goetz Graefe and Per-\u00c5ke Larson. 2001. B-Tree Indexes and CPU Caches. In ICDE. 349--358.","DOI":"10.1109\/ICDE.2001.914847"},{"key":"e_1_2_2_31_1","volume-title":"Patel","author":"Hankins Richard A.","year":"2003","unstructured":"Richard A. Hankins and Jignesh M. Patel. 2003. Effect of Node Size on the Performance of Cache-conscious B-trees. In SIGMETRICS. 283--294."},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.14778\/1454159.1454211"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1807167.1807206"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.is.2016.07.004"},{"key":"e_1_2_2_35_1","volume-title":"Jeffrey Dean, and Neoklis Polyzotis.","author":"Kraska Tim","year":"2018","unstructured":"Tim Kraska, Alex Beutel, Ed H. Chi, Jeffrey Dean, and Neoklis Polyzotis. 2018. The Case for Learned Index Structures. In SIGMOD. 489--504."},{"key":"e_1_2_2_36_1","doi-asserted-by":"crossref","unstructured":"Harald Lang Alexander Beischl Viktor Leis Peter A. Boncz Thomas Neumann and Alfons Kemper. 2020. Tree-Encoded Bitmaps. In SIGMOD. 937--967.","DOI":"10.1145\/3318464.3380588"},{"key":"e_1_2_2_37_1","doi-asserted-by":"crossref","unstructured":"Viktor Leis Alfons Kemper and Thomas Neumann. 2013. The Adaptive Radix Tree: ARTful Indexing for Main-memory Databases. In ICDE. 38--49.","DOI":"10.1109\/ICDE.2013.6544812"},{"key":"e_1_2_2_38_1","doi-asserted-by":"crossref","unstructured":"Justin J. Levandoski David B. Lomet and Sudipta Sengupta. 2013. The Bw-Tree: A B-tree for New Hardware Platforms. In ICDE. 302--313.","DOI":"10.1109\/ICDE.2013.6544834"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476249.3476305"},{"key":"e_1_2_2_40_1","volume-title":"Elmore","author":"Liu Chunwei","year":"2019","unstructured":"Chunwei Liu, McKade Umbenhower, Hao Jiang, Pranav Subramaniam, Jihong Ma, and Aaron J. Elmore. 2019. Mostly Order Preserving Dictionaries. In ICDE. 1214--1225."},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/603867.603878"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.14778\/2824032.2824078"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.14778\/3090163.3090164"},{"key":"e_1_2_2_44_1","doi-asserted-by":"crossref","unstructured":"Jianguo Wang Chunbin Lin Yannis Papakonstantinou and Steven Swanson. 2017. An Experimental Study of Bitmap Compression vs. Inverted List Compression. In SIGMOD. 993--1008.","DOI":"10.1145\/3035918.3064007"},{"key":"e_1_2_2_45_1","volume-title":"Andersen","author":"Wang Ziqi","year":"2018","unstructured":"Ziqi Wang, Andrew Pavlo, Hyeontaek Lim, Viktor Leis, Huanchen Zhang, Michael Kaminsky, and David G. Andersen. 2018. Building a Bw-Tree Takes More Than Just Buzz Words. In SIGMOD. 473--488."},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/1132863.1132864"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.14778\/3611479.3611502"},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.14778\/3352063.3352124"},{"key":"e_1_2_2_49_1","doi-asserted-by":"crossref","unstructured":"Feng Zhang Weitao Wan Chenyang Zhang Jidong Zhai Yunpeng Chai Haixiang Li and Xiaoyong Du. 2022. CompressDB: Enabling Efficient Compressed Data Direct Processing for Various Databases. In SIGMOD. 1655--1669.","DOI":"10.1145\/3514221.3526130"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-020-00636-3"},{"key":"e_1_2_2_51_1","doi-asserted-by":"crossref","unstructured":"Huanchen Zhang Xiaoxuan Liu David G. Andersen Michael Kaminsky Kimberly Keeton and Andrew Pavlo. 2020. Order-Preserving Key Compression for In-Memory Search Trees. In SIGMOD. 1601--1615.","DOI":"10.1145\/3318464.3380583"},{"key":"e_1_2_2_52_1","volume-title":"Carsten Binnig, Rodrigo Fonseca, and Tim Kraska.","author":"Ziegler Tobias","year":"2019","unstructured":"Tobias Ziegler, Sumukha Tumkur Vani, Carsten Binnig, Rodrigo Fonseca, and Tim Kraska. 2019. Designing Distributed Tree-based Index Structures for Fast RDMA-capable Networks. In SIGMOD. 741--758."}],"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3654972","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3654972","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T14:37:18Z","timestamp":1755787038000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3654972"}},"issued":{"date-parts":[[2024,5,29]]},"references-count":52,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,5,29]]}},"alternative-id":["10.1145\/3654972"],"URL":"https:\/\/doi.org\/10.1145\/3654972","ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"published":{"date-parts":[[2024,5,29]]}},{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T15:44:23Z","timestamp":1783784663590,"version":"3.55.0"},"reference-count":96,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2025,2,10]],"date-time":"2025-02-10T00:00:00Z","timestamp":1739145600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100006374","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62472400 and 62072428"],"award-info":[{"award-number":["62472400 and 62072428"]}],"id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2025,2,10]]},"abstract":"<jats:p>Cutting-edge platforms of graph neural networks (GNNs), such as DGL and PyG, harness the parallel processing power of GPUs to extract structural information from graph data, achieving state-of-the-art (SOTA) performance in fields such as recommendation systems, knowledge graphs, and bioinformatics. Despite the computational advantages provided by GPUs, these GNN platforms struggle with scalability challenges due to the colossal graphical structures processed and the limited memory capacities of GPUs. In response, this work introduces Capsule, a new out-of-core mechanism for large-scale GNN training. Unlike existing out-of-core GNN systems, which use main or secondary memory as operative memory and use CPU kernels during non-backpropagation computation, Capsule uses GPU memory and GPU kernels. By substantially leveraging the parallelization capabilities of GPUs, Capsule significantly enhances GNN training efficiency. In addition, Capsule can be smoothly integrated to mainstream open-source GNN frameworks, DGL and PyG, in a play-and-plug manner. Through a prototype implementation and comprehensive experiments on real datasets, we demonstrate that Capsule can achieve up to a 12.02\u00d7 improvement in runtime efficiency, while using only 22.24% of the main memory, compared to SOTA out-of-core GNN systems.<\/jats:p>","DOI":"10.1145\/3709669","type":"journal-article","created":{"date-parts":[[2025,2,11]],"date-time":"2025-02-11T15:45:06Z","timestamp":1739288706000},"page":"1-30","source":"Crossref","is-referenced-by-count":7,"title":["Capsule: An Out-of-Core Training Mechanism for Colossal GNNs"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-0572-5479","authenticated-orcid":false,"given":"Yongan","family":"Xiang","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, University of Science and Technology of China (USTC), Hefei, Anhui, China, &amp; Data Darkness Lab, Suzhou Institute for Advanced Research, USTC, Suzhou, Jiangsu, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6286-8679","authenticated-orcid":false,"given":"Zezhong","family":"Ding","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence and Data Science, USTC, Hefei, Anhui, China, &amp; Data Darkness Lab, Suzhou Institute for Advanced Research, USTC, Suzhou, Jiangsu, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-4033-3816","authenticated-orcid":false,"given":"Rui","family":"Guo","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence and Data Science, USTC, Hefei, Anhui, China, &amp; Data Darkness Lab, Suzhou Institute for Advanced Research, USTC, Suzhou, Jiangsu, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-5738-985X","authenticated-orcid":false,"given":"Shangyou","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence and Data Science, USTC, Hefei, Anhui, China, &amp; Data Darkness Lab, Suzhou Institute for Advanced Research, USTC, Suzhou, Jiangsu, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5290-5408","authenticated-orcid":false,"given":"Xike","family":"Xie","sequence":"additional","affiliation":[{"name":"School of Biomedical Engineering, USTC, Suzhou, Jiangsu, China, &amp; Data Darkness Lab, Suzhou Institute for Advanced Research, USTC, Suzhou, Jiangsu, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6881-4444","authenticated-orcid":false,"given":"S. Kevin","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Biomedical Engineering, USTC, Suzhou, Jiangsu, China, &amp; MIRACLE Center, Suzhou Institute for Advanced Research, USTC, Suzhou, Jiangsu, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,2,11]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i10.28948"},{"key":"e_1_2_1_2_1","unstructured":"Amazon. 2024. Best Sellers in Computer Graphics Cards. https:\/\/www.amazon.com\/Best-Sellers-Computer-Graphics-Cards\/zgbs\/pc\/284822"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2021.3065737"},{"key":"e_1_2_1_4_1","volume-title":"Layer-Neighbor Sampling - Defusing Neighborhood Explosion in GNNs. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems, NeurIPS 2023","author":"Fatih Muhammed","year":"2023","unstructured":"Muhammed Fatih Balin and \u00dcmit V. \u00c7ataly\u00fcrek. 2023. Layer-Neighbor Sampling - Defusing Neighborhood Explosion in GNNs. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023. https:\/\/papers.nips.cc\/paper_files\/paper\/2023\/hash\/51f9036d5e7ae822da8f6d4adda1fb39-Abstract-Conference.html"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3160017"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1963405.1963488"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/988672.988752"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.socnet.2007.04.002"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/2623330.2623660"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1080\/0022250X.2001.9990249"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0169--7552(98)00110-X"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3572848.3577528"},{"key":"e_1_2_1_13_1","volume-title":"International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=rytstxWAW","author":"Chen Jie","year":"2018","unstructured":"Jie Chen, Tengfei Ma, and Cao Xiao. 2018a. FastGCN: Fast Learning with Graph Convolutional Networks via Importance Sampling. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=rytstxWAW"},{"key":"e_1_2_1_14_1","volume-title":"Proceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research","volume":"950","author":"Chen Jianfei","year":"2018","unstructured":"Jianfei Chen, Jun Zhu, and Le Song. 2018b. Stochastic Training of Graph Convolutional Networks with Variance Reduction. In Proceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 80). PMLR, 942--950. https:\/\/proceedings.mlr.press\/v80\/chen18p.html"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330925"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403192"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403192"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.5555\/3157382.3157527"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2024.3475568"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3654965"},{"key":"e_1_2_1_21_1","article-title":"Improving Lipschitz-Constrained Neural Networks by Learning Activation Functions","volume":"25","author":"Ducotterd Stanislas","year":"2024","unstructured":"Stanislas Ducotterd, Alexis Goujon, Pakshal Bohra, Dimitris Perdios, Sebastian Neumayer, and Michael Unser. 2024. Improving Lipschitz-Constrained Neural Networks by Learning Activation Functions. J. Mach. Learn. Res., Vol. 25 (2024), 65:1--65:30. https:\/\/jmlr.org\/papers\/v25\/22--1347.html","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_2_1_22_1","unstructured":"Leo Egghe et al. 2006. An improvement of the h-index: The g-index. ISSI newsletter Vol. 2 1 (2006) 8--9. http:\/\/www2.stat-athens.aueb.gr\/ jpan\/Egghe-ISSI-2006.pdf"},{"key":"e_1_2_1_23_1","volume-title":"Fast graph representation learning with PyTorch Geometric. (2019). arxiv","author":"Fey Matthias","year":"1903","unstructured":"Matthias Fey and Jan Eric Lenssen. 2019. Fast graph representation learning with PyTorch Geometric. (2019). arxiv: 1903.02428 [cs.LG] http:\/\/arxiv.org\/abs\/1903.02428"},{"key":"e_1_2_1_24_1","volume-title":"15th USENIX Symposium on Operating Systems Design and Implementation (OSDI 21)","author":"Gandhi Swapnil","year":"2021","unstructured":"Swapnil Gandhi and Anand Padmanabha Iyer. 2021. P3: Distributed Deep Graph Learning at Scale. In 15th USENIX Symposium on Operating Systems Design and Implementation (OSDI 21). USENIX Association, 551--568. https:\/\/www.usenix.org\/conference\/osdi21\/presentation\/gandhi"},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics (Proceedings of Machine Learning Research","volume":"323","author":"Glorot Xavier","year":"2011","unstructured":"Xavier Glorot, Antoine Bordes, and Yoshua Bengio. 2011. Deep Sparse Rectifier Neural Networks. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics (Proceedings of Machine Learning Research, Vol. 15). PMLR, Fort Lauderdale, FL, USA, 315--323. https:\/\/proceedings.mlr.press\/v15\/glorot11a.html"},{"key":"e_1_2_1_26_1","volume-title":"PowerGraph: Distributed Graph-Parallel Computation on Natural Graphs. In 10th USENIX Symposium on Operating Systems Design and Implementation (OSDI 12)","author":"Gonzalez Joseph E.","year":"2012","unstructured":"Joseph E. Gonzalez, Yucheng Low, Haijie Gu, Danny Bickson, and Carlos Guestrin. 2012. PowerGraph: Distributed Graph-Parallel Computation on Natural Graphs. In 10th USENIX Symposium on Operating Systems Design and Implementation (OSDI 12). USENIX Association, Hollywood, CA, 17--30. https:\/\/www.usenix.org\/conference\/osdi12\/technical-sessions\/presentation\/gonzalez"},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research","volume":"5209","author":"Gower Robert Mansel","year":"2019","unstructured":"Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richt\u00e1rik. 2019. SGD: General Analysis and Improved Rates. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 97). PMLR, 5200--5209. https:\/\/proceedings.mlr.press\/v97\/qian19b.html"},{"key":"e_1_2_1_28_1","volume-title":"Inductive Representation Learning on Large Graphs. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems, December 4--9, 2017","author":"Hamilton William L.","year":"2017","unstructured":"William L. Hamilton, Zhitao Ying, and Jure Leskovec. 2017. Inductive Representation Learning on Large Graphs. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems, December 4--9, 2017, Long Beach, CA, USA. 1024--1034. https:\/\/proceedings.neurips.cc\/paper\/2017\/hash\/5dd9db5e033da9c6fb5ba83c7a7ebea9-Abstract.html"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.14778\/3358701.3358706"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3591647"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401063"},{"key":"e_1_2_1_32_1","volume-title":"Open Graph Benchmark: Datasets for Machine Learning on Graphs. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems, NeurIPS 2020","author":"Hu Weihua","year":"2020","unstructured":"Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. 2020. Open Graph Benchmark: Datasets for Machine Learning on Graphs. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems, NeurIPS 2020, December 6--12, 2020, virtual. https:\/\/proceedings.neurips.cc\/paper\/2020\/hash\/fb60d411a5c5b72b2e7d3527cfc84fd0-Abstract.html"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2018.2890515"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3591720"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3534540.3534710"},{"key":"e_1_2_1_36_1","volume-title":"Adaptive Sampling Towards Fast Graph Representation Learning. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems, NeurIPS 2018","author":"Zhang Tong","year":"2018","unstructured":"Wen-bing Huang, Tong Zhang, Yu Rong, and Junzhou Huang. 2018. Adaptive Sampling Towards Fast Graph Representation Learning. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems, NeurIPS 2018, December 3--8, 2018, Montr\u00e9al, Canada. 4563--4572. https:\/\/proceedings.neurips.cc\/paper\/2018\/hash\/01eee509ee2f68dc6014898c309e86bf-Abstract.html"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-022--12201--9"},{"key":"e_1_2_1_38_1","volume-title":"Proceedings of Machine Learning and Systems 2020","author":"Jia Zhihao","year":"2020","unstructured":"Zhihao Jia, Sina Lin, Mingyu Gao, Matei Zaharia, and Alex Aiken. 2020a. Improving the Accuracy, Scalability, and Performance of Graph Neural Networks with Roc. In Proceedings of Machine Learning and Systems 2020, MLSys 2020, Austin, TX, USA, March 2--4, 2020. mlsys.org. https:\/\/proceedings.mlsys.org\/paper_files\/paper\/2020\/hash\/91fc23ceccb664ebb0cf4257e1ba9c51-Abstract.html"},{"key":"e_1_2_1_39_1","first-page":"187","article-title":"Improving the accuracy, scalability, and performance of graph neural networks with roc","volume":"2","author":"Jia Zhihao","year":"2020","unstructured":"Zhihao Jia, Sina Lin, Mingyu Gao, Matei Zaharia, and Alex Aiken. 2020b. Improving the accuracy, scalability, and performance of graph neural networks with roc. Proceedings of Machine Learning and Systems , Vol. 2 (2020), 187--198. https:\/\/proceedings.mlsys.org\/paper_files\/paper\/2020\/hash\/91fc23ceccb664ebb0cf4257e1ba9c51-Abstract.html","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1016\/J.JCSS.2015.06.003"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1137\/S1064827595287997"},{"key":"e_1_2_1_42_1","volume-title":"Kipf and Max Welling","author":"Thomas","year":"2017","unstructured":"Thomas N. Kipf and Max Welling. 2017. Semi-Supervised Classification with Graph Convolutional Networks. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24--26, 2017, Conference Track Proceedings. OpenReview.net. https:\/\/openreview.net\/forum?id=SJU4ayYgl"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE53745.2022.00049"},{"key":"e_1_2_1_44_1","unstructured":"Jure Leskovec and Andrej Krevl. 2014. SNAP Datasets: Stanford Large Network Dataset Collection. http:\/\/snap.stanford.edu\/data."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE55515.2023.00379"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3419111.3421281"},{"key":"e_1_2_1_47_1","unstructured":"Renjie Liu Yichuan Wang Xiao Yan Zhenkun Cai Minjie Wang Haitian Jiang Bo Tang and Jinyang Li. 2024. DiskGNN: Bridging I\/O Efficiency and Model Accuracy for Out-of-Core GNN Training. (2024). arxiv: 2405.05231 [cs.LG] https:\/\/arxiv.org\/abs\/2405.05231"},{"key":"e_1_2_1_48_1","volume-title":"20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23)","author":"Liu Tianfeng","year":"2023","unstructured":"Tianfeng Liu, Yangrui Chen, Dan Li, Chuan Wu, Yibo Zhu, Jun He, Yanghua Peng, Hongzheng Chen, Hongzhi Chen, and Chuanxiong Guo. 2023. BGL: GPU-Efficient GNN Training by Optimizing Graph Data I\/O and Preprocessing. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). USENIX Association, Boston, MA, 103--118. https:\/\/www.usenix.org\/conference\/nsdi23\/presentation\/liu-tianfeng"},{"key":"e_1_2_1_49_1","volume-title":"NeuGraph: Parallel Deep Neural Network Computation on Large Graphs. In 2019 USENIX Annual Technical Conference (USENIX ATC 19)","author":"Ma Lingxiao","year":"2019","unstructured":"Lingxiao Ma, Zhi Yang, Youshan Miao, Jilong Xue, Ming Wu, Lidong Zhou, and Yafei Dai. 2019. NeuGraph: Parallel Deep Neural Network Computation on Large Graphs. In 2019 USENIX Annual Technical Conference (USENIX ATC 19). USENIX Association, Renton, WA, 443--458. https:\/\/www.usenix.org\/conference\/atc19\/presentation\/ma"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDCS.2018.00072"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3457300"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE53745.2022.00242"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3480856"},{"key":"e_1_2_1_54_1","volume-title":"Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems 2013. Proceedings of a meeting held December 5--8, 2013","author":"Mikolov Tom\u00e1s","year":"2013","unstructured":"Tom\u00e1s Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013. Distributed Representations of Words and Phrases and their Compositionality. In Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems 2013. Proceedings of a meeting held December 5--8, 2013, Lake Tahoe, Nevada, United States,, Christopher J. C. Burges, L\u00e9on Bottou, Zoubin Ghahramani, and Kilian Q. Weinberger (Eds.). 3111--3119. https:\/\/proceedings.neurips.cc\/paper\/2013\/hash\/9aa42b31882ec039965f3c4923ce901b-Abstract.html"},{"key":"e_1_2_1_55_1","volume-title":"15th USENIX Symposium on Operating Systems Design and Implementation (OSDI 21)","author":"Mohoney Jason","year":"2021","unstructured":"Jason Mohoney, Roger Waleffe, Henry Xu, Theodoros Rekatsinas, and Shivaram Venkataraman. 2021. Marius: Learning Massive Graph Embeddings on a Single Machine. In 15th USENIX Symposium on Operating Systems Design and Implementation (OSDI 21). USENIX Association, 533--549. https:\/\/www.usenix.org\/conference\/osdi21\/presentation\/mohoney"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/2739008"},{"key":"e_1_2_1_57_1","unstructured":"NVIDIA. 2016. GeForce GTX 1050 Ti. https:\/\/www.nvidia.com\/en-gb\/geforce\/graphics-cards\/geforce-gtx-1050-ti\/specifications\/. Accessed: 2024-07--12."},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2012.01.004"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1007\/978--3--540--69311--621"},{"key":"e_1_2_1_60_1","volume-title":"Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021","author":"Paetzold Johannes C.","year":"2021","unstructured":"Johannes C. Paetzold, Julian McGinnis, Suprosanna Shit, Ivan Ezhov, Paul B\u00fcschl, Chinmay Prabhakar, Anjany Sekuboyina, Mihail I. Todorov, Georgios Kaissis, Ali Ert\u00fcrk, Stephan G\u00fcnnemann, and Bjoern H. Menze. 2021. Whole Brain Vessel Graphs: A Dataset and Benchmark for Graph Learning and Neuroscience. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual, Joaquin Vanschoren and Sai-Kit Yeung (Eds.). https:\/\/datasets-benchmarks-proceedings.neurips.cc\/paper\/2021\/hash\/c9f0f895fb98ab9159f51fd0297e236d-Abstract-round2.html"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403280"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.14778\/3551793.3551819"},{"key":"e_1_2_1_63_1","volume-title":"High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems, NeurIPS 2019","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas K\u00f6pf, Edward Z. Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems, NeurIPS 2019, December 8--14, 2019, Vancouver, BC, Canada. 8024--8035. https:\/\/proceedings.neurips.cc\/paper\/2019\/hash\/bdbca288fee7f92f2bfa9f7012727740-Abstract.html"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/2806416.2806424"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE55515.2023.00083"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.14778\/2536258.2536264"},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1038\/323533a0"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1109\/FOCS.2007.56"},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2310.00837"},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1186\/s40537-019-0257--5"},{"key":"e_1_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41592-020-0792--1"},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1016\/J.ARTINT.2016.08.001"},{"key":"e_1_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.hm.2020.04.003"},{"key":"e_1_2_1_74_1","volume-title":"6th International Conference on Learning Representations, ICLR","author":"Velickovic Petar","year":"2018","unstructured":"Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Li\u00f2, and Yoshua Bengio. 2018. Graph Attention Networks. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net. https:\/\/openreview.net\/forum?id=rJXMpikCZ"},{"key":"e_1_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1145\/3552326.3567501"},{"key":"e_1_2_1_76_1","volume-title":"Proceedings of Machine Learning and Systems, MLSys 2022","author":"Wan Cheng","year":"2022","unstructured":"Cheng Wan, Youjie Li, Ang Li, Nam Sung Kim, and Yingyan Lin. 2022. BNS-GCN: Efficient Full-Graph Training of Graph Convolutional Networks with Partition-Parallelism and Random Boundary Node Sampling. In Proceedings of Machine Learning and Systems, MLSys 2022, Santa Clara, CA, USA, August 29 - September 1, 2022. mlsys.org. https:\/\/proceedings.mlsys.org\/paper_files\/paper\/2022\/hash\/676638b91bc90529e09b22e58abb01d6-Abstract.html"},{"key":"e_1_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1145\/3589288"},{"key":"e_1_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1162\/QSS_A_00021"},{"key":"e_1_2_1_79_1","volume-title":"Highly-Performant Package for Graph Neural Networks.","author":"Wang Minjie","year":"2019","unstructured":"Minjie Wang, Da Zheng, Zihao Ye, Quan Gan, Mufei Li, Xiang Song, Jinjing Zhou, Chao Ma, Lingfan Yu, Yu Gai, Tianjun Xiao, Tong He, George Karypis, Jinyang Li, and Zheng Zhang. 2019. Deep Graph Library: A Graph-Centric, Highly-Performant Package for Graph Neural Networks. (2019). arxiv: 1909.01315 [cs.LG] http:\/\/arxiv.org\/abs\/1909.01315"},{"key":"e_1_2_1_80_1","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3591634"},{"key":"e_1_2_1_81_1","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.1992.287150"},{"key":"e_1_2_1_82_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2023.3257507"},{"key":"e_1_2_1_83_1","volume-title":"Capsule: An Out-of-Core Training Mechanism for Colossal GNNs (Supplementary Materials). Technical Report. GitHub Repository. https:\/\/github.com\/USTC-DataDarknessLab\/Capsule\/blob\/master\/supp.pdf","author":"Xiang Yongan","year":"2025","unstructured":"Yongan Xiang, Zezhong Ding, Rui Guo, Shangyou Wang, Xike Xie, and S. Kevin Zhou. 2025. Capsule: An Out-of-Core Training Mechanism for Colossal GNNs (Supplementary Materials). Technical Report. GitHub Repository. https:\/\/github.com\/USTC-DataDarknessLab\/Capsule\/blob\/master\/supp.pdf"},{"key":"e_1_2_1_84_1","volume-title":"Proceedings of the 27th International Conference on Neural Information Processing Systems -","volume":"1","author":"Xie Cong","year":"2014","unstructured":"Cong Xie, Ling Yan, Wu-Jun Li, and Zhihua Zhang. 2014. Distributed power-law graph computing: theoretical and empirical analysis. In Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 1 (Montreal, Canada) (NIPS'14). MIT Press, Cambridge, MA, USA, 1673--1681."},{"key":"e_1_2_1_85_1","volume-title":"7th International Conference on Learning Representations, ICLR 2019","author":"Xu Keyulu","year":"2019","unstructured":"Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. 2019. How Powerful are Graph Neural Networks?. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6--9, 2019. OpenReview.net. https:\/\/openreview.net\/forum?id=ryGs6iA5Km"},{"key":"e_1_2_1_86_1","doi-asserted-by":"publisher","DOI":"10.1145\/2350190.2350193"},{"key":"e_1_2_1_87_1","doi-asserted-by":"publisher","DOI":"10.1007\/978--3--319--49178--336"},{"key":"e_1_2_1_88_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219890"},{"key":"e_1_2_1_89_1","volume-title":"The Eleventh International Conference on Learning Representations, ICLR 2023","author":"Zaidi Sheheryar","year":"2023","unstructured":"Sheheryar Zaidi, Michael Schaarschmidt, James Martens, Hyunjik Kim, Yee Whye Teh, Alvaro Sanchez-Gonzalez, Peter W. Battaglia, Razvan Pascanu, and Jonathan Godwin. 2023. Pre-training via Denoising for Molecular Property Prediction. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1--5, 2023. OpenReview.net. https:\/\/openreview.net\/forum?id=tYIMtogyee"},{"key":"e_1_2_1_90_1","volume-title":"GraphSAINT: Graph Sampling Based Inductive Learning Method. In 8th International Conference on Learning Representations, ICLR 2020","author":"Zeng Hanqing","year":"2020","unstructured":"Hanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan, and Viktor K. Prasanna. 2020. GraphSAINT: Graph Sampling Based Inductive Learning Method. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26--30, 2020. OpenReview.net. https:\/\/openreview.net\/forum?id=BJe8pkHFwS"},{"key":"e_1_2_1_91_1","doi-asserted-by":"publisher","DOI":"10.1145\/3097983.3098033"},{"key":"e_1_2_1_92_1","doi-asserted-by":"publisher","DOI":"10.1145\/3580305.3599404"},{"key":"e_1_2_1_93_1","doi-asserted-by":"publisher","DOI":"10.1109\/IA351965.2020.00011"},{"key":"e_1_2_1_94_1","doi-asserted-by":"publisher","DOI":"10.14778\/3352063.3352127"},{"key":"e_1_2_1_95_1","volume-title":"Layer-Dependent Importance Sampling for Training Deep and Large Graph Convolutional Networks. In Annual Conference on Neural Information Processing Systems (NeurIPS 2019","author":"Zou Difan","year":"2019","unstructured":"Difan Zou, Ziniu Hu, Yewen Wang, Song Jiang, Yizhou Sun, and Quanquan Gu. 2019. Layer-Dependent Importance Sampling for Training Deep and Large Graph Convolutional Networks. In Annual Conference on Neural Information Processing Systems (NeurIPS 2019). Vancouver, BC, Canada, 11247--11256. https:\/\/proceedings.neurips.cc\/paper\/2019\/hash\/91ba4a4478a66bee9812b0804b6f9d1b-Abstract.html"},{"key":"e_1_2_1_96_1","doi-asserted-by":"publisher","DOI":"10.1145\/3533702.3534920"}],"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3709669","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3709669","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T18:19:10Z","timestamp":1774981150000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3709669"}},"issued":{"date-parts":[[2025,2,10]]},"references-count":96,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,2,10]]}},"alternative-id":["10.1145\/3709669"],"URL":"https:\/\/doi.org\/10.1145\/3709669","ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"published":{"date-parts":[[2025,2,10]]}},{"indexed":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T20:58:41Z","timestamp":1775595521352,"version":"3.50.1"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"1","funder":[{"name":"New Generation Artificial Intelligence-National Science and Technology Major Project","award":["2025ZD0123304"],"award-info":[{"award-number":["2025ZD0123304"]}]},{"DOI":"10.13039\/501100001809","name":"NSFC","doi-asserted-by":"crossref","award":["62532001"],"award-info":[{"award-number":["62532001"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"ARC","award":["DP230101445"],"award-info":[{"award-number":["DP230101445"]}]},{"name":"ARC","award":["FT210100303"],"award-info":[{"award-number":["FT210100303"]}]},{"name":"Australian Research Council Centre of Excellence for Mathematical Modelling of Cellular Systems","award":["CE230100001"],"award-info":[{"award-number":["CE230100001"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2026,4,2]]},"abstract":"<jats:p>Continuous subgraph matching (CSM), which finds incremental matches of a query graph for each update in a dynamic graph, has gained significant research attention. Most CSM algorithms follow a common indexing-enumeration paradigm: they first use indexes to identify candidate vertices and edges for the query graph, and then enumerate matches based on these candidates. Although there have been several comprehensive experimental analyses of CSM algorithms, they tend to evaluate CSM algorithms holistically, obscuring the distinct contributions of the indexing and enumeration methods to overall performance. In this paper, we decouple the indexing method and enumeration method of existing CSM algorithms, and focus on the comparison of indexing methods. Our experimental results offer guidance on index selection across different scenarios, serving as a reference for future research and industrial applications. They further reveal the relative importance of different index components, informing strategies to discard less essential parts when memory is limited. Additionally, we show that the commonly used candidate count metric may underestimate the filtering effectiveness of certain indexes, suggesting that future research should adopt more reliable evaluation metrics.<\/jats:p>","DOI":"10.1145\/3786623","type":"journal-article","created":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T17:54:13Z","timestamp":1775584453000},"page":"1-26","source":"Crossref","is-referenced-by-count":0,"title":["An Extensive Experimental Study of Indexes in Continuous Subgraph Matching: [Experiments &amp; Analysis]"],"prefix":"10.1145","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-6657-3781","authenticated-orcid":false,"given":"Xiangyang","family":"Gou","sequence":"first","affiliation":[{"name":"University of New South Wales, Sydney, NSW, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8586-4400","authenticated-orcid":false,"given":"Lei","family":"Zou","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9738-827X","authenticated-orcid":false,"given":"Jeffrey Xu","family":"Yu","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6572-2600","authenticated-orcid":false,"given":"Wenjie","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of New South Wales, Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,7]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2009. LiveJournal dataset on SNAP website. https:\/\/snap.stanford.edu\/data\/socLiveJournal1.html"},{"key":"e_1_2_1_2_1","unstructured":"2017. Lsbench codes. https:\/\/code.google.com\/archive\/p\/lsbench\/"},{"key":"e_1_2_1_3_1","unstructured":"2025. Source code datasets and queries for this paper. https:\/\/github.com\/XiangyangGouUNSW\/CSM-Experiments"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3129246"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19433-7_41"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2742796"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3300086"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2915236"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-019-00558-9"},{"key":"e_1_2_1_10_1","volume-title":"A subgraph isomorphism algorithm and its application to biochemical data. BMC bioinformatics 14","author":"Bonnici Vincenzo","year":"2013","unstructured":"Vincenzo Bonnici, Rosalba Giugno, Alfredo Pulvirenti, Dennis Shasha, and Alfredo Ferro. 2013. A subgraph isomorphism algorithm and its application to biochemical data. BMC bioinformatics 14 (2013), 1-13."},{"key":"e_1_2_1_11_1","volume-title":"A selectivity based approach to continuous pattern detection in streaming graphs. arXiv preprint arXiv:1503.00849","author":"Choudhury Sutanay","year":"2015","unstructured":"Sutanay Choudhury, Lawrence Holder, George Chin, Khushbu Agarwal, and John Feo. 2015. A selectivity based approach to continuous pattern detection in streaming graphs. arXiv preprint arXiv:1503.00849 (2015)."},{"key":"e_1_2_1_12_1","volume-title":"Carlo Sansone, and Mario Vento","author":"Cordella Luigi P","year":"2004","unstructured":"Luigi P Cordella, Pasquale Foggia, Carlo Sansone, and Mario Vento. 2004. A (sub) graph isomorphism algorithm for matching large graphs. IEEE transactions on pattern analysis and machine intelligence 26, 10 (2004), 1367-1372."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2489791"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.14778\/2733004.2733010"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3319880"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2463676.2465300"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/SWAT.1971.1"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3064027"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.14778\/3587136.3587144"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3056445"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3457265"},{"key":"e_1_2_1_22_1","volume-title":"Seo,Wook-Shin Han, Jeong-Hoon Lee, Sungpack Hong, Hassan Chafi, Hyungyu Shin, and Geonhwa Jeong.","author":"Kim Kyoungmin","year":"2018","unstructured":"Kyoungmin Kim, In Seo,Wook-Shin Han, Jeong-Hoon Lee, Sungpack Hong, Hassan Chafi, Hyungyu Shin, and Geonhwa Jeong. 2018. Turboflux: A fast continuous subgraph matching system for streaming graph data. In Proceedings of the 2018 international conference on management of data. 411-426."},{"key":"e_1_2_1_23_1","volume-title":"Quantum query complexity of subgraph isomorphism and homomorphism. arXiv preprint arXiv:1509.06361","author":"Kulkarni Raghav","year":"2015","unstructured":"Raghav Kulkarni and Supartha Podder. 2015. Quantum query complexity of subgraph isomorphism and homomorphism. arXiv preprint arXiv:1509.06361 (2015)."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-35173-0_20"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3654950"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1080\/15427951.2009.10129177"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2019.00100"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE60146.2024.00257"},{"key":"e_1_2_1_29_1","volume-title":"Efficient Multi-Query Oriented Continuous Subgraph Matching. In 2024 IEEE 40th International Conference on Data Engineering (ICDE). IEEE, 3230-3243","author":"Ma Ziyi","year":"2024","unstructured":"Ziyi Ma, Jianye Yang, Xu Zhou, Guoqing Xiao, Jianhua Wang, Liang Yang, Kenli Li, and Xuemin Lin. 2024. Efficient Multi-Query Oriented Continuous Subgraph Matching. In 2024 IEEE 40th International Conference on Data Engineering (ICDE). IEEE, 3230-3243."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.14778\/3342263.3342643"},{"key":"e_1_2_1_31_1","volume-title":"Time-Constrained Continuous Subgraph Matching Using Temporal Information for Filtering and Backtracking. In 2024 IEEE 40th International Conference on Data Engineering (ICDE). IEEE, 3257-3269","author":"Min Seunghwan","year":"2024","unstructured":"Seunghwan Min, Jihoon Jang, Kunsoo Park, Dora Giammarresi, Giuseppe F Italiano, and Wook-Shin Han. 2024. Time-Constrained Continuous Subgraph Matching Using Temporal Information for Filtering and Backtracking. In 2024 IEEE 40th International Conference on Data Engineering (ICDE). IEEE, 3257-3269."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.14778\/3457390.3457395"},{"key":"e_1_2_1_33_1","volume-title":"GPU-Accelerated Batch-Dynamic Subgraph Matching. In 2024 IEEE 40th International Conference on Data Engineering (ICDE). IEEE, 3204-3216","author":"Qiu Linshan","year":"2024","unstructured":"Linshan Qiu, Lu Chen, Hailiang Jie, Xiangyu Ke, Yunjun Gao, Yang Liu, and Zetao Zhang. 2024. GPU-Accelerated Batch-Dynamic Subgraph Matching. In 2024 IEEE 40th International Conference on Data Engineering (ICDE). IEEE, 3204-3216."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.14778\/2735479.2735493"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.14778\/3551793.3551803"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.14778\/3523210.3523218"},{"key":"e_1_2_1_37_1","volume-title":"Efficient subgraph matching on billion node graphs. arXiv preprint arXiv:1205.6691","author":"Sun Zhao","year":"2012","unstructured":"Zhao Sun, Hongzhi Wang, Haixun Wang, Bin Shao, and Jianzhong Li. 2012. Efficient subgraph matching on billion node graphs. arXiv preprint arXiv:1205.6691 (2012)."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/321921.321925"},{"key":"e_1_2_1_39_1","volume-title":"Proc. International Conference on Database Theory.","author":"Veldhuizen Todd L","year":"2014","unstructured":"Todd L Veldhuizen. 2014. Leapfrog triejoin: A simple,worst-case optimal join algorithm. In Proc. International Conference on Database Theory."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.3233\/SW-212864"},{"key":"e_1_2_1_41_1","volume-title":"GCSM: GPU-Accelerated Continuous Subgraph Matching for Large Graphs. In 2024 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, 1046-1057","author":"Wei Yihua","year":"2024","unstructured":"Yihua Wei and Peng Jiang. 2024. GCSM: GPU-Accelerated Continuous Subgraph Matching for Large Graphs. In 2024 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, 1046-1057."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.14778\/3681954.3681963"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3588695"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.14778\/2002974.2002976"}],"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3786623","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T19:58:48Z","timestamp":1775591928000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3786623"}},"issued":{"date-parts":[[2026,4,2]]},"references-count":44,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,4,2]]}},"alternative-id":["10.1145\/3786623"],"URL":"https:\/\/doi.org\/10.1145\/3786623","ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"published":{"date-parts":[[2026,4,2]]}},{"indexed":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T19:06:50Z","timestamp":1779131210249,"version":"3.51.4"},"reference-count":78,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["Grant 62472119"],"award-info":[{"award-number":["Grant 62472119"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Cyber Security-National Science and Technology Major Project","award":["Grant 2025ZD1500204"],"award-info":[{"award-number":["Grant 2025ZD1500204"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2026,5,18]]},"abstract":"<jats:p>Audit logs serve as a fundamental data source for system security. However, extracting high-value threat information from massive log data poses significant data management challenges: traditional unsupervised models produce high false positive rates due to their inherent assumption of equating statistical anomalies with malicious activities; Emerging Large Language Model (LLM) based solutions, despite their powerful semantic understanding capabilities, are constrained by high computational costs and context window lengths, making it difficult to detect attacks within massive logs. Furthermore, the outputs of existing methods differ significantly from the practical attack reports required by security analysts.<\/jats:p>\n                  <jats:p>To overcome these limitations, this paper proposes ANTEATER, an innovative end-to-end attack investigation framework based on raw logs that features a cascading ''filter-then-scrutinize'' architecture. The ''Filter'' stage is a lightweight, flow-based anomaly detection model that efficiently filters massive logs and reduces the data scale for investigation. Subsequently, the ''Scrutinize'' stage is an attack investigation model with a three-agent LLM collaboration. It operates on a provenance graph constructed from the filtered anomalous logs. The agents collaboratively and autonomously explore and reconstruct the attack subgraph, then generate a structured natural-language report. ANTEATER not only effectively mitigates the LLM bottleneck from cost and context window, enabling it to tackle long-term, stealthy attacks, but also bridges the critical gap between raw data detection and the generation of readable attack reports.<\/jats:p>","DOI":"10.1145\/3802012","type":"journal-article","created":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T18:19:16Z","timestamp":1779128356000},"page":"1-27","source":"Crossref","is-referenced-by-count":0,"title":["ANTEATER: A Filter-then-Scrutinize Architecture for End-to-End Attack Investigation"],"prefix":"10.1145","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-0072-163X","authenticated-orcid":false,"given":"Yiming","family":"Ren","sequence":"first","affiliation":[{"name":"Institute of Information Engineering\uff0cChinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-9516-7994","authenticated-orcid":false,"given":"Haoqiang","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering\uff0cChinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-6967-4128","authenticated-orcid":false,"given":"Linghao","family":"Li","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering\uff0cChinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-9034-7598","authenticated-orcid":false,"given":"Haoyang","family":"Chen","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering\uff0cChinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2646-6100","authenticated-orcid":false,"given":"Chengxiang","family":"Si","sequence":"additional","affiliation":[{"name":"National Computer Network Emergency Response Technical Team\/Coordination Center of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6924-5848","authenticated-orcid":false,"given":"Zhou","family":"Zhou","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4815-3463","authenticated-orcid":false,"given":"Qingyun","family":"Liu","sequence":"additional","affiliation":[{"name":"Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,5,18]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al.","author":"Achiam Josh","year":"2023","unstructured":"Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al., 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)."},{"key":"e_1_2_1_2_1","volume-title":"30th USENIX security symposium (USENIX security 21). 3005-3022.","author":"Alsaheel Abdulellah","unstructured":"Abdulellah Alsaheel, Yuhong Nan, Shiqing Ma, Le Yu, Gregory Walkup, Z Berkay Celik, Xiangyu Zhang, and Dongyan Xu. 2021. : A sequence-based learning approach for attack investigation. In 30th USENIX security symposium (USENIX security 21). 3005-3022."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/COMST.2019.2891891"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/comst.2019.2891891"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/milcom.2018.8599708"},{"key":"e_1_2_1_6_1","unstructured":"Authors. 2025. The paper website. https:\/\/sites.google.com\/view\/anteatermaterials"},{"key":"e_1_2_1_7_1","unstructured":"Shuai Bai Keqin Chen Xuejing Liu Jialin Wang Wenbin Ge Sibo Song Kai Dang Peng Wang Shijie Wang Jun Tang et al. 2025. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923 (2025)."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/IKT.2013.6620049"},{"key":"e_1_2_1_9_1","unstructured":"Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared D Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et al. 2020. Language models are few-shot learners. Advances in neural information processing systems Vol. 33 (2020) 1877-1901."},{"key":"e_1_2_1_10_1","volume-title":"2025 IEEE 41st International Conference on Data Engineering (ICDE). IEEE, 4025-4037","author":"Chen Xiaolei","year":"2025","unstructured":"Xiaolei Chen, Jia Chen, Jie Shi, Peng Wang, and Wei Wang. 2025. EPAS: Efficient Online Log Parsing via Asynchronous Scheduling of LLM Queries. In 2025 IEEE 41st International Conference on Data Engineering (ICDE). IEEE, 4025-4037."},{"key":"e_1_2_1_11_1","unstructured":"CrowdStrike. 2024. Living off the Land (LotL) Attack. https:\/\/www.crowdstrike.com\/en-us\/cybersecurity-101\/cyberattacks\/living-off-the-land-attack\/ Accessed on: 2024-03-08."},{"key":"e_1_2_1_12_1","unstructured":"DARPA i2o. 2024. Transparent Computing. https:\/\/github.com\/darpa-i2o\/Transparent-Computing Accessed: 2024-12-03."},{"key":"e_1_2_1_13_1","unstructured":"DepImpact. 2023. DepImpact Source. https:\/\/github.com\/depimpact\/depimpact\/tree\/master\/depimpact-source-master Accessed: 2025-06-06."},{"key":"e_1_2_1_14_1","first-page":"3277","volume-title":"32nd USENIX Security Symposium (USENIX Security 23)","author":"Ding Hailun","year":"2023","unstructured":"Hailun Ding, Juan Zhai, Dong Deng, and Shiqing Ma. 2023a. The case for learned provenance graph storage systems. In 32nd USENIX Security Symposium (USENIX Security 23). 3277-3294."},{"key":"e_1_2_1_15_1","first-page":"373","volume-title":"32nd USENIX Security Symposium (USENIX Security 23)","author":"Ding Hailun","year":"2023","unstructured":"Hailun Ding, Juan Zhai, Yuhong Nan, and Shiqing Ma. 2023b. AIRTAG: Towards Automated Attack Investigation by Unsupervised Learning with Log Texts. In 32nd USENIX Security Symposium (USENIX Security 23). 373-390."},{"key":"e_1_2_1_16_1","volume-title":"Nice: Non-linear independent components estimation. arXiv preprint arXiv:1410.8516","author":"Dinh Laurent","year":"2014","unstructured":"Laurent Dinh, David Krueger, and Yoshua Bengio. 2014. Nice: Non-linear independent components estimation. arXiv preprint arXiv:1410.8516 (2014)."},{"key":"e_1_2_1_17_1","volume-title":"Density estimation using real nvp. arXiv preprint arXiv:1605.08803","author":"Dinh Laurent","year":"2016","unstructured":"Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. 2016. Density estimation using real nvp. arXiv preprint arXiv:1605.08803 (2016)."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133956.3134015"},{"key":"e_1_2_1_19_1","unstructured":"Brown Farinholt Mohammad Rezaeirad Paul Pearce Hitesh Dharmdasani Haikuo Yin StevensLe Blond Damon Mccoy and Kirill Levchenko. [n.d.]. To Catch a Ratter: Monitoring the Behavior of Amateur DarkComet RAT Operators in the Wild. ([n.d.])."},{"key":"e_1_2_1_20_1","first-page":"113","volume-title":"2018 USENIX Annual Technical Conference (USENIX ATC 18)","author":"Gao Peng","year":"2018","unstructured":"Peng Gao, Xusheng Xiao, Zhichun Li, Fengyuan Xu, Sanjeev R Kulkarni, and Prateek Mittal. 2018. : Enabling efficient attack investigation from system monitoring data. In 2018 USENIX Annual Technical Conference (USENIX ATC 18). 113-126."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2018.06.055"},{"key":"e_1_2_1_22_1","unstructured":"Google and Hugging Face. 2018. BERT base uncased. https:\/\/huggingface.co\/google-bert\/bert-base-uncased. Accessed: 2025-06-06."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.14722\/ndss.2023.24207"},{"key":"e_1_2_1_24_1","unstructured":"Daya Guo Dejian Yang Haowei Zhang Junxiao Song Ruoyu Zhang Runxin Xu Qihao Zhu Shirong Ma Peiyi Wang Xiao Bi et al. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948 (2025)."},{"key":"e_1_2_1_25_1","volume-title":"Unicorn: Runtime provenance-based detector for advanced persistent threats. arXiv preprint arXiv:2001.01525","author":"Han Xueyuan","year":"2020","unstructured":"Xueyuan Han, Thomas Pasquier, Adam Bates, James Mickens, and Margo Seltzer. 2020. Unicorn: Runtime provenance-based detector for advanced persistent threats. arXiv preprint arXiv:2001.01525 (2020)."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP40000.2020.00096"},{"key":"e_1_2_1_27_1","volume-title":"NoDoze: Combatting Threat Alert Fatigue with Automated Provenance Triage. In Network and Distributed System Security Symposium.","author":"Hassan Wajih Ul","year":"2019","unstructured":"Wajih Ul Hassan, Shengjian Guo, Ding Li, Zhengzhang Chen, Kangkook Jee, Zhichun Li, and Adam Bates. 2019. NoDoze: Combatting Threat Alert Fatigue with Automated Provenance Triage. In Network and Distributed System Security Symposium."},{"key":"e_1_2_1_28_1","volume-title":"A survey on automated log analysis for reliability engineering. ACM computing surveys (CSUR)","author":"He Shilin","year":"2021","unstructured":"Shilin He, Pinjia He, Zhuangbin Chen, Tianyi Yang, Yuxin Su, and Michael R Lyu. 2021. A survey on automated log analysis for reliability engineering. ACM computing surveys (CSUR), Vol. 54, 6 (2021), 1-37."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3534678.3539321"},{"key":"e_1_2_1_30_1","unstructured":"Peiwei Hu Ruigang Liang and Kai Chen. [n.d.]. DeGPT: Optimizing Decompiler Output with LLM. ([n.d.])."},{"key":"e_1_2_1_31_1","first-page":"10331","article-title":"Joint Multi-Facts Reasoning Network For Complex Temporal Question Answering Over Knowledge Graph. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Huang Rikui","year":"2024","unstructured":"Rikui Huang, Wei Wei, Xiaoye Qu, Wenfeng Xie, Xianling Mao, and Dangyang Chen. 2024. Joint Multi-Facts Reasoning Network For Complex Temporal Question Answering Over Knowledge Graph. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 10331-10335.","journal-title":"IEEE"},{"key":"e_1_2_1_32_1","unstructured":"Shaohan Huang Yi Liu Carol Fung He Wang Hailong Yang and Zhongzhi Luan. [n.d.]. Improving Log-Based Anomaly Detection by Pre-Training Hierarchical Transformers. ([n.d.])."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP46215.2023.10179405"},{"key":"e_1_2_1_34_1","volume-title":"Proceedings of the 2017 ACM SIGSAC conference on computer and communications security. 377-390","author":"Ji Yang","year":"2017","unstructured":"Yang Ji, Sangho Lee, Evan Downing, Weiren Wang, Mattia Fazzini, Taesoo Kim, Alessandro Orso, and Wenke Lee. 2017. Rain: Refinable attack investigation with on-demand inter-process information flow tracking. In Proceedings of the 2017 ACM SIGSAC conference on computer and communications security. 377-390."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3588918"},{"key":"e_1_2_1_36_1","first-page":"5197","volume-title":"33rd USENIX Security Symposium (USENIX Security 24)","author":"Jia Zian","year":"2024","unstructured":"Zian Jia, Yun Xiong, Yuhong Nan, Yao Zhang, Jinjing Zhao, and Mi Wen. 2024. MAGIC: Detecting Advanced Persistent Threats via Masked Graph Representation Learning. In 33rd USENIX Security Symposium (USENIX Security 24). 5197-5214."},{"key":"e_1_2_1_37_1","volume-title":"Glow: Generative flow with invertible 1x1 convolutions. Advances in neural information processing systems","author":"Kingma Durk P","year":"2018","unstructured":"Durk P Kingma and Prafulla Dhariwal. 2018. Glow: Generative flow with invertible 1x1 convolutions. Advances in neural information processing systems, Vol. 31 (2018)."},{"key":"e_1_2_1_38_1","volume-title":"Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114","author":"Kingma Diederik P","year":"2013","unstructured":"Diederik P Kingma and Max Welling. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114 (2013)."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.5555\/1036285"},{"key":"e_1_2_1_40_1","volume-title":"International Conference on Detection of Intrusions and Malware, and Vulnerability Assessment. Springer, 366-387","author":"Lamprakis Pavlos","year":"2017","unstructured":"Pavlos Lamprakis, Ruggiero Dargenio, David Gugelmann, Vincent Lenders, Markus Happe, and Laurent Vanbever. 2017. Unsupervised detection of APT C&C channels using web request graphs. In International Conference on Detection of Intrusions and Malware, and Vulnerability Assessment. Springer, 366-387."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3654966"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3722212.3724427"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3319535.3363224"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3639478.3643108"},{"key":"e_1_2_1_45_1","doi-asserted-by":"crossref","unstructured":"Yushan Liu Mu Zhang Ding Li Kangkook Jee Zhichun Li Zhenyu Wu Junghwan Rhee and Prateek Mittal. 2018. Towards a Timely Causality Analysis for Enterprise Security.. In NDSS.","DOI":"10.14722\/ndss.2018.23254"},{"key":"e_1_2_1_46_1","volume-title":"Reasoning on graphs: Faithful and interpretable large language model reasoning. arXiv preprint arXiv:2310.01061","author":"Luo Linhao","year":"2023","unstructured":"Linhao Luo, Yuan-Fang Li, Gholamreza Haffari, and Shirui Pan. 2023. Reasoning on graphs: Faithful and interpretable large language model reasoning. arXiv preprint arXiv:2310.01061 (2023)."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3677139"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3319535.3363217"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP.2019.00026"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01367"},{"key":"e_1_2_1_51_1","first-page":"1199","volume-title":"32nd USENIX Security Symposium (USENIX Security 23)","author":"Mukherjee Kunal","year":"2023","unstructured":"Kunal Mukherjee, Joshua Wiedemeier, Tianhao Wang, James Wei, Feng Chen, Muhyun Kim, Murat Kantarcioglu, and Kangkook Jee. 2023. Evading detectors with adversarial system actions. In 32nd USENIX Security Symposium (USENIX Security 23). 1199-1216."},{"key":"e_1_2_1_52_1","first-page":"575","article-title":"What supercomputers say: A study of five system logs. In 37th annual IEEE\/IFIP international conference on dependable systems and networks (DSN'07)","author":"Oliner Adam","year":"2007","unstructured":"Adam Oliner and Jon Stearley. 2007. What supercomputers say: A study of five system logs. In 37th annual IEEE\/IFIP international conference on dependable systems and networks (DSN'07). IEEE, 575-584.","journal-title":"IEEE"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/2991079.2991122"},{"key":"e_1_2_1_54_1","volume-title":"James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, et al.","author":"Petroni Fabio","year":"2020","unstructured":"Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, et al., 2020. KILT: a benchmark for knowledge intensive language tasks. arXiv preprint arXiv:2009.02252 (2020)."},{"key":"e_1_2_1_55_1","unstructured":"Purseclab. 2023. ATLAS. https:\/\/github.com\/purseclab\/ATLAS. Accessed: 2025-06-06."},{"key":"e_1_2_1_56_1","first-page":"273","article-title":"Loggpt: Exploring chatgpt for log-based anomaly detection. In 2023 IEEE International Conference on High Performance Computing & Communications, Data Science & Systems, Smart City & Dependability in Sensor, Cloud & Big Data Systems & Application (HPCC\/DSS\/SmartCity\/DependSys)","author":"Qi Jiaxing","year":"2023","unstructured":"Jiaxing Qi, Shaohan Huang, Zhongzhi Luan, Shu Yang, Carol Fung, Hailong Yang, Depei Qian, Jing Shang, Zhiwen Xiao, and Zhihui Wu. 2023. Loggpt: Exploring chatgpt for log-based anomaly detection. In 2023 IEEE International Conference on High Performance Computing & Communications, Data Science & Systems, Smart City & Dependability in Sensor, Cloud & Big Data Systems & Application (HPCC\/DSS\/SmartCity\/DependSys). IEEE, 273-280.","journal-title":"IEEE"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.3390\/app10113874"},{"key":"e_1_2_1_58_1","unstructured":"Alec Radford Karthik Narasimhan Tim Salimans Ilya Sutskever et al. 2018. Improving language understanding by generative pre-training. (2018)."},{"key":"e_1_2_1_59_1","unstructured":"Alec Radford Jeffrey Wu Rewon Child David Luan Dario Amodei Ilya Sutskever et al. 2019. Language models are unsupervised multitask learners. OpenAI blog Vol. 1 8 (2019) 9."},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/319709.319712"},{"key":"e_1_2_1_61_1","volume-title":"Audit-llm: Multi-agent collaboration for log-based insider threat detection. arXiv preprint arXiv:2408.08902","author":"Song Chengyu","year":"2024","unstructured":"Chengyu Song, Linru Ma, Jianming Zheng, Jinzhi Liao, Hongyu Kuang, and Lin Yang. 2024. Audit-llm: Multi-agent collaboration for log-based insider threat detection. arXiv preprint arXiv:2408.08902 (2024)."},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.2991\/nceece-15.2016.187"},{"key":"e_1_2_1_63_1","volume-title":"Think-on-graph: Deep and responsible reasoning of large language model on knowledge graph. arXiv preprint arXiv:2307.07697","author":"Sun Jiashuo","year":"2023","unstructured":"Jiashuo Sun, Chengjin Xu, Lumingyuan Tang, Saizhuo Wang, Chen Lin, Yeyun Gong, Lionel M Ni, Heung-Yeung Shum, and Jian Guo. 2023. Think-on-graph: Deep and responsible reasoning of large language model on knowledge graph. arXiv preprint arXiv:2307.07697 (2023)."},{"key":"e_1_2_1_64_1","volume-title":"The Web as a Knowledge-base for Answering Complex Questions. arXiv: Computation and Language,arXiv: Computation and Language (Mar","author":"Talmor Alon","year":"2018","unstructured":"Alon Talmor and Jonathan Berant. 2018. The Web as a Knowledge-base for Answering Complex Questions. arXiv: Computation and Language,arXiv: Computation and Language (Mar 2018)."},{"key":"e_1_2_1_65_1","volume-title":"CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge. arXiv: Computation and Language,arXiv: Computation and Language (Nov","author":"Talmor Alon","year":"2018","unstructured":"Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2018. CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge. arXiv: Computation and Language,arXiv: Computation and Language (Nov 2018)."},{"key":"e_1_2_1_66_1","volume-title":"A comprehensive survey of hallucination mitigation techniques in large language models. arXiv preprint arXiv:2401.01313","author":"Tonmoy SM","year":"2024","unstructured":"SM Tonmoy, SM Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das. 2024. A comprehensive survey of hallucination mitigation techniques in large language models. arXiv preprint arXiv:2401.01313, Vol. 6 (2024)."},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.14722\/ndss.2025.240266"},{"key":"e_1_2_1_68_1","volume-title":"Ding Li, Kangkook Jee, Xiao Yu, Kexuan Zou, Junghwan Rhee, Zhengzhang Chen, Wei Cheng, Carl A Gunter, et al.","author":"Wang Qi","year":"2020","unstructured":"Qi Wang, Wajih Ul Hassan, Ding Li, Kangkook Jee, Xiao Yu, Kexuan Zou, Junghwan Rhee, Zhengzhang Chen, Wei Cheng, Carl A Gunter, et al., 2020. You Are What You Do: Hunting Stealthy Malware via Data Provenance Analysis.. In NDSS."},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIFS.2022.3208815"},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1109\/icc.2016.7511197"},{"key":"e_1_2_1_71_1","volume-title":"Beyond chain-of-thought: A survey of chain-of-x paradigms for llms. arXiv preprint arXiv:2404.15676","author":"Xia Yu","year":"2024","unstructured":"Yu Xia, Rui Wang, Xu Liu, Mingyan Li, Tong Yu, Xiang Chen, Julian McAuley, and Shuai Li. 2024. Beyond chain-of-thought: A survey of chain-of-x paradigms for llms. arXiv preprint arXiv:2404.15676 (2024)."},{"key":"e_1_2_1_72_1","first-page":"4355","volume-title":"32nd USENIX Security Symposium (USENIX Security 23)","author":"Yang Fan","year":"2023","unstructured":"Fan Yang, Jiacen Xu, Chunlin Xiong, Zhou Li, and Kehuan Zhang. 2023. PROGRAPHER: An Anomaly Detection System based on Provenance Graph Embedding. In 32nd USENIX Security Symposium (USENIX Security 23). 4355-4372."},{"key":"e_1_2_1_73_1","first-page":"1525","volume-title":"33rd USENIX Security Symposium (USENIX Security 24)","author":"Yang Limin","year":"2024","unstructured":"Limin Yang, Zhi Chen, Chenkai Wang, Zhenning Zhang, Sushruth Booma, Phuong Cao, Constantin Adam, Alexander Withers, Zbigniew Kalbarczyk, Ravishankar K Iyer, et al., 2024. True attacks, attack attempts, or benign triggers? an empirical measurement of network alerts in a security operations center. In 33rd USENIX Security Symposium (USENIX Security 24). 1525-1542."},{"key":"e_1_2_1_74_1","series-title":"Lecture Notes in Computer Science,Lecture Notes in Computer Science (Jan","doi-asserted-by":"crossref","DOI":"10.1007\/11890881","volume-title":"Integrating IDS alert correlation and OS-level dependency tracking","author":"Zhai Yan","year":"2006","unstructured":"Yan Zhai, Peng Ning, and Jun Xu. 2006. Integrating IDS alert correlation and OS-level dependency tracking. Lecture Notes in Computer Science,Lecture Notes in Computer Science (Jan 2006)."},{"key":"e_1_2_1_75_1","volume-title":"Noisy pair corrector for dense retrieval. arXiv preprint arXiv:2311.03798","author":"Zhang Hang","year":"2023","unstructured":"Hang Zhang, Yeyun Gong, Xingwei He, Dayiheng Liu, Daya Guo, Jiancheng Lv, and Jian Guo. 2023. Noisy pair corrector for dense retrieval. arXiv preprint arXiv:2311.03798 (2023)."},{"key":"e_1_2_1_76_1","doi-asserted-by":"publisher","unstructured":"G. Zhao K. Xu L. Xu and B. Wu. 2015. Detecting APT Malware Infections Based on Malicious DNS and Traffic Analysis. IEEE Access (Jan 2015) 1132\u20131142. doi:10.1109\/access.2015.2458581","DOI":"10.1109\/access.2015.2458581"},{"key":"e_1_2_1_77_1","volume-title":"Loghub: A large collection of system log datasets towards automated log analytics. arXiv preprint arXiv:2008.06448","author":"Zhu Jinyang","year":"2020","unstructured":"Jinyang Zhu, Shilin He, Jinyu Liu, Pinjia He, Qi Zheng, and Michael R Lyu. 2020. Loghub: A large collection of system log datasets towards automated log analytics. arXiv preprint arXiv:2008.06448 (2020)."},{"key":"e_1_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.14722\/ndss.2019.23492"}],"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3802012","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T18:28:38Z","timestamp":1779128918000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3802012"}},"issued":{"date-parts":[[2026,5,18]]},"references-count":78,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,5,18]]}},"alternative-id":["10.1145\/3802012"],"URL":"https:\/\/doi.org\/10.1145\/3802012","ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"published":{"date-parts":[[2026,5,18]]}},{"indexed":{"date-parts":[[2026,4,23]],"date-time":"2026-04-23T15:47:46Z","timestamp":1776959266216,"version":"3.51.4"},"reference-count":0,"publisher":"Association for Computing Machinery (ACM)","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2026,4,2]]},"abstract":"<jats:p>This is a corrigendum for the article ''A Comprehensive Benchmark on Spectral GNNs: The Impact on Efficiency, Memory, and Effectiveness: [Experiments &amp; Analysis]'' published in Proc. ACM Manag. Data 3, 4 (SIGMOD), Article 238 (September 2025), 29 pages.<\/jats:p>","DOI":"10.1145\/3803526","type":"journal-article","created":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T17:54:13Z","timestamp":1775584453000},"page":"1-1","source":"Crossref","is-referenced-by-count":0,"title":["Corrigendum: A Comprehensive Benchmark on Spectral GNNs: The Impact on Efficiency, Memory, and Effectiveness: [Experiments &amp; Analysis]"],"prefix":"10.1145","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3176-4401","authenticated-orcid":false,"given":"Ningyi","family":"Liao","sequence":"first","affiliation":[{"name":"College of Computing and Data Science, Nanyang Technological University, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0839-5460","authenticated-orcid":false,"given":"Haoyu","family":"Liu","sequence":"additional","affiliation":[{"name":"College of Computing and Data Science, Nanyang Technological University, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5176-6378","authenticated-orcid":false,"given":"Zulun","family":"Zhu","sequence":"additional","affiliation":[{"name":"College of Computing and Data Science, Nanyang Technological University, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8197-0903","authenticated-orcid":false,"given":"Siqiang","family":"Luo","sequence":"additional","affiliation":[{"name":"College of Computing and Data Science, Nanyang Technological University, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9775-4241","authenticated-orcid":false,"given":"Laks V.S.","family":"Lakshmanan","sequence":"additional","affiliation":[{"name":"Department of Computer Science, The University of British Columbia, Vancouver, BC, Canada"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,7]]},"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3803526","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,23]],"date-time":"2026-04-23T15:30:04Z","timestamp":1776958204000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3803526"}},"issued":{"date-parts":[[2026,4,2]]},"references-count":0,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,4,2]]}},"alternative-id":["10.1145\/3803526"],"URL":"https:\/\/doi.org\/10.1145\/3803526","ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"published":{"date-parts":[[2026,4,2]]}},{"indexed":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T19:07:39Z","timestamp":1779131259294,"version":"3.51.4"},"reference-count":0,"publisher":"Association for Computing Machinery (ACM)","issue":"3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2026,5,18]]},"abstract":"<jats:p>The Proceedings of the ACM on Management of Data (PACMMOD) is concerned with the principles, algorithms, techniques, systems, and applications of database management systems, data management technology, and science and engineering of data. It includes articles reporting cutting-edge data management, data engineering, and data science research. We are pleased to present the 3rd issue of Volume 4 of PACMMOD. This issue contains papers that were submitted to the SIGMOD research track in October 2025 and is the last issue of the research track of SIGMOD 2026.<\/jats:p>","DOI":"10.1145\/3802001","type":"journal-article","created":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T18:19:16Z","timestamp":1779128356000},"page":"1-2","source":"Crossref","is-referenced-by-count":0,"title":["PACMMOD V4, N3 (SIGMOD), June 2026 Editorial"],"prefix":"10.1145","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2744-7836","authenticated-orcid":false,"given":"Carsten","family":"Binnig","sequence":"first","affiliation":[{"name":"Technical University of Darmstadt, Darmstadt, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-8300-7891","authenticated-orcid":false,"given":"Sudeepa","family":"Roy","sequence":"additional","affiliation":[{"name":"Duke University, Durham, North Carolina, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0215-9539","authenticated-orcid":false,"given":"Divy","family":"Agrawal","sequence":"additional","affiliation":[{"name":"University of California, Santa Barbara, Santa Barbara, California, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9582-869X","authenticated-orcid":false,"given":"Angela","family":"Bonifati","sequence":"additional","affiliation":[{"name":"Lyon 1 University, Lyon, France"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,5,18]]},"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3802001","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T18:33:02Z","timestamp":1779129182000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3802001"}},"issued":{"date-parts":[[2026,5,18]]},"references-count":0,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,5,18]]}},"alternative-id":["10.1145\/3802001"],"URL":"https:\/\/doi.org\/10.1145\/3802001","ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"published":{"date-parts":[[2026,5,18]]}},{"indexed":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T07:15:47Z","timestamp":1779174947345,"version":"3.51.4"},"reference-count":124,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"DOI":"10.13039\/501100006374","name":"Cisco Systems","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100006374","name":"Meta","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2025,6,17]]},"abstract":"<jats:p>\n                    <jats:italic toggle=\"yes\">Symbolic approximations<\/jats:italic>\n                    are dimensionality reduction techniques that convert time series into sequences of discrete symbols, enhancing interpretability while reducing computational and storage costs. To construct symbolic representations, first numeric representations approximate and capture properties of raw time series, followed by a discretization step that converts these numeric dimensions into symbols. Despite decades of development, existing approaches have several key limitations that often result in unsatisfactory performance: they (i) rely on data-agnostic numeric approximations, disregarding intrinsic properties of the time series; (ii) decompose dimensions into equal-sized subspaces, assuming independence among dimensions; and (iii) allocate a uniform encoding budget for discretizing each dimension or subspace, assuming balanced importance. To address these shortcomings, we propose SPARTAN, a novel data-adaptive symbolic approximation method that intelligently allocates the encoding budget according to the importance of the constructed uncorrelated dimensions. Specifically, SPARTAN (i) leverages intrinsic dimensionality reduction properties to derive non-overlapping, uncorrelated latent dimensions; (ii) adaptively distributes the budget based on the importance of each dimension by solving a constrained optimization problem; and (iii) prevents false dismissals in similarity search by ensuring a lower bound on the true distance in the original space. To demonstrate SPARTAN's robustness, we conduct the most comprehensive study to date, comparing SPARTAN with seven state-of-the-art symbolic methods across four tasks: classification, clustering, indexing, and anomaly detection. Rigorous statistical analysis across hundreds of datasets shows that SPARTAN outperforms competing methods significantly on\n                    <jats:italic toggle=\"yes\">all<\/jats:italic>\n                    tasks in terms of downstream accuracy, given the same budget. Notably, SPARTAN achieves up to a 2x speedup compared to the most accurate rival. Overall, SPARTAN effectively improves the symbolic representation quality without storage or runtime overheads, paving the way for future advancements.\n                  <\/jats:p>","DOI":"10.1145\/3725357","type":"journal-article","created":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T21:22:29Z","timestamp":1750281749000},"page":"1-30","source":"Crossref","is-referenced-by-count":12,"title":["SPARTAN: Data-Adaptive Symbolic Time-Series Approximation"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-8289-3462","authenticated-orcid":false,"given":"Fan","family":"Yang","sequence":"first","affiliation":[{"name":"The Ohio State University, Columbus, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7592-748X","authenticated-orcid":false,"given":"John","family":"Paparrizos","sequence":"additional","affiliation":[{"name":"The Ohio State University, Columbus, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,6,18]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"2025. SPARTAN: Data-Adaptive Symbolic Time-Series Approximation. https:\/\/github.com\/TheDatumOrg\/SPARTAN. Accessed: 2025-04-07."},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/645415.652239"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1088\/0067-0049\/219\/1\/12"},{"key":"e_1_2_2_4_1","volume-title":"Badal","author":"Andr\u00e9-J\u00f6nsson Henrik","year":"1997","unstructured":"Henrik Andr\u00e9-J\u00f6nsson and Dushan Z. Badal. 1997. Using signature files for querying time-series data. In Principles of Data Mining and Knowledge Discovery, Jan Komorowski and Jan Zytkow, (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 211-220."},{"key":"e_1_2_2_5_1","volume-title":"Sebastian Pineda Arango, Shubham Kapoor, Jasper Zschiegner, Danielle C. Maddix, Hao Wang, Michael W. Mahoney, Kari Torkkola, Andrew Gordon Wilson, Michael Bohlke-Schneider, and Yuyang Wang.","author":"Ansari Abdul Fatir","year":"2024","unstructured":"Abdul Fatir Ansari, Lorenzo Stella, Caner Turkmen, Xiyuan Zhang, Pedro Mercado, Huibin Shen, Oleksandr Shchur, Syama Syndar Rangapuram, Sebastian Pineda Arango, Shubham Kapoor, Jasper Zschiegner, Danielle C. Maddix, Hao Wang, Michael W. Mahoney, Kari Torkkola, Andrew Gordon Wilson, Michael Bohlke-Schneider, and Yuyang Wang. 2024. Chronos: Learning the Language of Time Series. arXiv preprint arXiv:2403.07815, (2024)."},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/565196.565200"},{"key":"e_1_2_2_7_1","volume-title":"James Large, and Jason Lines","author":"Bagnall Anthony","year":"2016","unstructured":"Anthony Bagnall, Aaron Bostrom, James Large, and Jason Lines. 2016. The Great Time Series Classification Bake Off: An Experimental Evaluation of Recently Proposed Algorithms. Extended Version. arxiv:1602.01711 [cs.LG]"},{"key":"e_1_2_2_8_1","volume-title":"Aaron Bostrom, James Large, and Eamonn Keogh.","author":"Bagnall Anthony","year":"2017","unstructured":"Anthony Bagnall, Jason Lines, Aaron Bostrom, James Large, and Eamonn Keogh. 2017. The great time series classification bake off: a review and experimental evaluation of recent algorithmic advances. Data mining and knowledge discovery, Vol. 31 (2017), 606-660."},{"key":"e_1_2_2_9_1","volume-title":"k-shapestream: Probabilistic streaming clustering for electric grid events. In 2021 IEEE Madrid PowerTech","author":"Bariya Mohini","year":"2021","unstructured":"Mohini Bariya, Alexandra von Meier, John Paparrizos, and Michael J Franklin. 2021. k-shapestream: Probabilistic streaming clustering for electric grid events. In 2021 IEEE Madrid PowerTech. IEEE, 1's6, (2021)."},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-025-00907-x"},{"key":"e_1_2_2_11_1","doi-asserted-by":"crossref","unstructured":"Paul Boniol Qinghua Liu Mingyi Huang Themis Palpanas and John Paparrizos. 2024a. Dive into Time-Series Anomaly Detection: A Decade Review. arXiv preprint arXiv:2412.20512 (2024).","DOI":"10.1109\/ICDE60146.2024.00409"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.14778\/3554821.3554879"},{"key":"e_1_2_2_13_1","first-page":"847","article-title":"New Trends in Time Series Anomaly Detection","author":"Boniol Paul","year":"2023","unstructured":"Paul Boniol, John Paparrizos, and Themis Palpanas. 2023. New Trends in Time Series Anomaly Detection. In EDBT. 847-850.","journal-title":"EDBT."},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE60146.2024.00409"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476311.3476365"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.14778\/3467861.3467863"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE60146.2024.00423"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.23919\/Eusipco47968.2020.9287474"},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.1999.754915"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3532622"},{"key":"e_1_2_2_21_1","volume-title":"Yan Zhu, Shaghayegh Gharghabi, Chotirat Ann Ratanamahatana, Yanping, Bing Hu, Nurjahan Begum, Anthony Bagnall, Abdullah Mueen, Gustavo Batista, and Hexagon-ML.","author":"Dau Hoang Anh","year":"2018","unstructured":"Hoang Anh Dau, Eamonn Keogh, Kaveh Kamgar, Chin-Chia Michael Yeh, Yan Zhu, Shaghayegh Gharghabi, Chotirat Ann Ratanamahatana, Yanping, Bing Hu, Nurjahan Begum, Anthony Bagnall, Abdullah Mueen, Gustavo Batista, and Hexagon-ML. 2018. The UCR Time Series Classification Archive. https:\/\/www.cs.ucr.edu\/ eamonn\/time_series_data_2018\/."},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1063\/1.1531823"},{"key":"e_1_2_2_23_1","volume-title":"Mach. Learn. Res.","volume":"7","author":"Dem\u0161ar Janez","year":"2006","unstructured":"Janez Dem\u0161ar. 2006. Statistical Comparisons of Classifiers over Multiple Data Sets. J. Mach. Learn. Res., Vol. 7 (dec 2006), 1--30."},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3725258"},{"key":"e_1_2_2_25_1","first-page":"107","article-title":"Beyond the Dimensions: A Structured Evaluation of Multivariate Time Series Distance Measures. In 2024 IEEE 40th International Conference on Data Engineering Workshops (ICDEW)","author":"Hondt Jens E","year":"2024","unstructured":"Jens E d'Hondt, Odysseas Papapetrou, and John Paparrizos. 2024. Beyond the Dimensions: A Structured Evaluation of Multivariate Time Series Distance Measures. In 2024 IEEE 40th International Conference on Data Engineering Workshops (ICDEW). IEEE, 107-112.","journal-title":"IEEE"},{"key":"e_1_2_2_26_1","unstructured":"Karima Echihabi Kostas Zoumpatianos Themis Palpanas and Houda Benbrahim. 2020. The Lernaean Hydra of Data Series Similarity Search: An Experimental Evaluation of the State of the Art. arxiv:2006.11454 [cs.DB]"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10618-020-00689-6"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/191839.191925"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.1937.10503522"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.379"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/3-540-44794-6_10"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-47880-7_3"},{"key":"e_1_2_2_33_1","volume-title":"MOMENT: A Family of Open Time-series Foundation Models. In International Conference on Machine Learning.","author":"Goswami Mononito","year":"2024","unstructured":"Mononito Goswami, Konrad Szafer, Arjun Choudhry, Yifu Cai, Shuo Li, and Artur Dubrawski. 2024. MOMENT: A Family of Open Time-series Foundation Models. In International Conference on Machine Learning."},{"key":"e_1_2_2_34_1","volume-title":"Advances in Neural Information Processing Systems","author":"Gruver Nate","year":"1962","unstructured":"Nate Gruver, Marc Finzi, Shikai Qiu, and Andrew G Wilson. 2023. Large Language Models Are Zero-Shot Time Series Forecasters. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, (Eds.), Vol. 36. Curran Associates, Inc., 19622-19635. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2023\/file\/3eb7ca52e8207697361b2c0fb3926511-Paper-Conference.pdf"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3722212.3725135"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3725420"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1137\/090771806"},{"key":"e_1_2_2_38_1","volume-title":"Srikanta Patnaik","author":"Xiaoxu He.","unstructured":"Xiaoxu He. 2023. A Survey on Time Series Forecasting. In 3D Imaging--Multidimensional Signal Processing and Deep Learning, Srikanta Patnaik, Roumen Kountchev, Yonghang Tai, and Roumiana Kountcheva, (Eds.). Springer Nature Singapore, Singapore, 13-23."},{"key":"e_1_2_2_39_1","volume-title":"AutoML: A survey of the state-of-the-art. Knowledge-based systems","author":"He Xin","year":"2021","unstructured":"Xin He, Kaiyong Zhao, and Xiaowen Chu. 2021. AutoML: A survey of the state-of-the-art. Knowledge-based systems, Vol. 212 (2021), 106622."},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-023-01952-0"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/312129.318357"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.3390\/en6020579"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/SUTC.2010.29"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.14778\/3380750.3380761"},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3457283"},{"key":"e_1_2_2_46_1","volume-title":"The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=Unb5CVPtae","author":"Jin Ming","year":"2024","unstructured":"Ming Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu, James Y. Zhang, Xiaoming Shi, Pin-Yu Chen, Yuxuan Liang, Yuan-Fang Li, Shirui Pan, and Qingsong Wen. 2024. Time-LLM: Time Series Forecasting by Reprogramming Large Language Models. In The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=Unb5CVPtae"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1007\/b98835"},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1098\/rsta.2015.0202"},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3470918"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/376284.375680"},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1007\/978--1--4899--7687--1_192"},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1007\/PL00011669"},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3352020.3352022"},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177729694"},{"key":"e_1_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.20965\/jaciii.2013.p0263"},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.140"},{"key":"e_1_2_2_57_1","doi-asserted-by":"publisher","unstructured":"Yuan Li Jessica Lin and Tim Oates. [n.d.]. Visualizing Variable-Length Time Series Motifs. 895-906. doi:10.1137\/1.9781611972825.77 arXiv:https:\/\/epubs.siam.org\/doi\/pdf\/10.1137\/1.9781611972825.77","DOI":"10.1137\/1.9781611972825.77"},{"key":"e_1_2_2_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/882082.882086"},{"key":"e_1_2_2_59_1","volume-title":"Experiencing SAX: a novel symbolic representation of time series. Data Mining and knowledge discovery","author":"Lin Jessica","year":"2007","unstructured":"Jessica Lin, Eamonn Keogh, Li Wei, and Stefano Lonardi. 2007. Experiencing SAX: a novel symbolic representation of time series. Data Mining and knowledge discovery, Vol. 15 (2007), 107-144."},{"key":"e_1_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10844-012-0196-5"},{"key":"e_1_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2016.0133"},{"key":"e_1_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476249.3476305"},{"key":"e_1_2_2_63_1","volume-title":"AdaEdge: A Dynamic Compression Selection Framework for Resource Constrained Devices. In 2024 IEEE 40th International Conference on Data Engineering (ICDE). IEEE, 1506-1519","author":"Liu Chunwei","year":"2024","unstructured":"Chunwei Liu, John Paparrizos, and Aaron J Elmore. 2024d. AdaEdge: A Dynamic Compression Selection Framework for Resource Constrained Devices. In 2024 IEEE 40th International Conference on Data Engineering (ICDE). IEEE, 1506-1519."},{"key":"e_1_2_2_64_1","volume-title":"Proceedings of the 41st International Conference on Machine Learning, (Proceedings of Machine Learning Research","volume":"31325","author":"Liu Haoxin","unstructured":"Haoxin Liu, Harshavardhan Kamarthi, Lingkai Kong, Zhiyuan Zhao, Chao Zhang, and B. Aditya Prakash. 2024b. Time-Series Forecasting for Out-of-Distribution Generalization Using Invariant Learning. In Proceedings of the 41st International Conference on Machine Learning, (Proceedings of Machine Learning Research, Vol. 235), Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, (Eds.). PMLR, 31312-31325. https:\/\/proceedings.mlr.press\/v235\/liu24ae.html"},{"key":"e_1_2_2_65_1","unstructured":"Haoxin Liu Chenghao Liu and B Aditya Prakash. 2024c. A picture is worth a thousand numbers: Enabling llms reason about time series via visualization. arXiv preprint arXiv:2411.06018 (2024)."},{"key":"e_1_2_2_66_1","first-page":"77888","article-title":"e. Time-mmd: Multi-domain multimodal dataset for time series analysis","volume":"37","author":"Liu Haoxin","year":"2024","unstructured":"Haoxin Liu, Shangqing Xu, Zhiyuan Zhao, Lingkai Kong, Harshavardhan Prabhakar Kamarthi, Aditya Sasanur, Megha Sharma, Jiaming Cui, Qingsong Wen, Chao Zhang, et al., 2024 e. Time-mmd: Multi-domain multimodal dataset for time series analysis. Advances in Neural Information Processing Systems, Vol. 37 (2024), 77888-77933.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_2_67_1","unstructured":"Haoxin Liu Zhiyuan Zhao Shiduo Li and B Aditya Prakash. 2025. Evaluating System 1 vs. 2 Reasoning Approaches for Zero-Shot Time-Series Forecasting: A Benchmark and Insights. arXiv preprint arXiv:2503.01895 (2025)."},{"key":"e_1_2_2_68_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.466"},{"key":"e_1_2_2_69_1","doi-asserted-by":"publisher","DOI":"10.14778\/3685800.3685842"},{"key":"e_1_2_2_70_1","volume-title":"The Elephant in the Room: Towards A Reliable Time-Series Anomaly Detection Benchmark. In The Thirty-eight Conference on Neural Information Processing Systems.","author":"Liu Qinghua","year":"2024","unstructured":"Qinghua Liu and John Paparrizos. 2024. The Elephant in the Room: Towards A Reliable Time-Series Anomaly Detection Benchmark. In The Thirty-eight Conference on Neural Information Processing Systems."},{"key":"e_1_2_2_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/3631429"},{"key":"e_1_2_2_72_1","volume-title":"DEWS2006 4A-i8","volume":"7","author":"Lkhagva Battuguldur","year":"2006","unstructured":"Battuguldur Lkhagva, Yu Suzuki, and Kyoji Kawagoe. 2006. Extended SAX: Extension of symbolic aggregate approximation for financial time series data representation. DEWS2006 4A-i8, Vol. 7 (2006)."},{"key":"e_1_2_2_73_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.dcan.2017.10.002"},{"key":"e_1_2_2_74_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-41398-8_24"},{"key":"e_1_2_2_75_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.acha.2010.02.003"},{"key":"e_1_2_2_76_1","doi-asserted-by":"publisher","DOI":"10.1002\/asi.23612"},{"key":"e_1_2_2_77_1","doi-asserted-by":"publisher","DOI":"10.1007\/978--3-030--67658--2_38"},{"key":"e_1_2_2_78_1","doi-asserted-by":"publisher","DOI":"10.1007\/978--3-030--33607--3_2"},{"key":"e_1_2_2_79_1","volume-title":"Distribution-free multiple comparisons","author":"Nemenyi Peter Bjorn","unstructured":"Peter Bjorn Nemenyi. 1963. Distribution-free multiple comparisons., Princeton University."},{"key":"e_1_2_2_80_1","unstructured":"Thach Le Nguyen and Georgiana Ifrim. 2022. MrSQM: Fast Time Series Classification with Symbolic Representations. arxiv:2109.01036 [cs.LG]"},{"key":"e_1_2_2_81_1","volume-title":"International Conference on Learning Representations.","author":"Nie Yuqi","year":"2023","unstructured":"Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam. 2023. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers. In International Conference on Learning Representations."},{"key":"e_1_2_2_82_1","unstructured":"Ioannis Paparrizos. 2018a. Fast scalable and accurate algorithms for time-series analysis. Ph.D. Dissertation. Columbia University USA."},{"key":"e_1_2_2_83_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2011.5767943"},{"key":"e_1_2_2_84_1","unstructured":"John Paparrizos. 2018b. ucr time-series archive: Backward compatibility missing values and varying lengths. URL: https:\/\/github. com\/johnpaparrizos\/UCRArchiveFixes (2018)."},{"key":"e_1_2_2_85_1","doi-asserted-by":"publisher","DOI":"10.14778\/3551793.3551830"},{"key":"e_1_2_2_86_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE53745.2022.00268"},{"key":"e_1_2_2_87_1","doi-asserted-by":"publisher","DOI":"10.14778\/3342263.3342648"},{"key":"e_1_2_2_88_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2737793"},{"key":"e_1_2_2_89_1","doi-asserted-by":"publisher","DOI":"10.1145\/3044711"},{"key":"e_1_2_2_90_1","doi-asserted-by":"publisher","DOI":"10.14778\/3529337.3529354"},{"key":"e_1_2_2_91_1","unstructured":"John Paparrizos Haojun Li Fan Yang Kaize Wu Jens E d'Hondt and Odysseas Papapetrou. 2024a. A survey on time-series distance measures. arXiv preprint arXiv:2412.20574 (2024)."},{"key":"e_1_2_2_92_1","unstructured":"John Paparrizos Chunwei Liu Bruno Barbarioli Johnny Hwang Ikraduya Edian Aaron J Elmore Michael J Franklin and Sanjay Krishnan. 2021. VergeDB: A Database for IoT Analytics on Edge Devices. In CIDR."},{"key":"e_1_2_2_93_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3389760"},{"key":"e_1_2_2_94_1","volume-title":"Data Engineering","author":"Paparrizos John","year":"2023","unstructured":"John Paparrizos, Chunwei Liu, Aaron J Elmore, and Michael J Franklin. 2023a. Querying Time-Series Data: A Comprehensive Comparison of Distance Measures. Data Engineering, (2023), 69."},{"key":"e_1_2_2_95_1","doi-asserted-by":"publisher","DOI":"10.14778\/3611540.3611622"},{"key":"e_1_2_2_96_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939722"},{"key":"e_1_2_2_97_1","doi-asserted-by":"publisher","DOI":"10.1200\/JOP.2015.010504"},{"key":"e_1_2_2_98_1","doi-asserted-by":"publisher","DOI":"10.14778\/3594512.3594530"},{"key":"e_1_2_2_99_1","unstructured":"John Paparrizos Fan Yang and Haojun Li. 2024b. Bridging the gap: A decade review of time-series clustering methods. arXiv preprint arXiv:2412.20582 (2024)."},{"key":"e_1_2_2_100_1","doi-asserted-by":"publisher","DOI":"10.1145\/342009.335437"},{"key":"e_1_2_2_101_1","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.1971.10482356"},{"key":"e_1_2_2_102_1","doi-asserted-by":"publisher","DOI":"10.1142\/9789812813305_0005"},{"key":"e_1_2_2_103_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10618-014-0377--7"},{"key":"e_1_2_2_104_1","doi-asserted-by":"publisher","DOI":"10.1145\/2247596.2247656"},{"key":"e_1_2_2_105_1","volume-title":"Multivariate Time Series Classification with WEASELMUSE. ArXiv","author":"Sch\u00e4fer Patrick","year":"2017","unstructured":"Patrick Sch\u00e4fer and Ulf Leser. 2017. Multivariate Time Series Classification with WEASELMUSE. ArXiv, Vol. abs\/1711.11343 (2017). https:\/\/api.semanticscholar.org\/CorpusID:8727328"},{"key":"e_1_2_2_106_1","doi-asserted-by":"publisher","DOI":"10.1145\/3132847.3132980"},{"key":"e_1_2_2_107_1","first-page":"481","article-title":"Time series anomaly discovery with grammar-based compression","author":"Senin Pavel","year":"2015","unstructured":"Pavel Senin, Jessica Lin, Xing Wang, Tim Oates, Sunil Gandhi, Arnold P Boedihardjo, Crystal Chen, and Susan Frankenstein. 2015. Time series anomaly discovery with grammar-based compression. In Edbt. 481-492.","journal-title":"Edbt."},{"key":"e_1_2_2_108_1","volume-title":"Interactive discovery of variable-length time series patterns. ACM Transactions on Knowledge Discovery from Data (TKDD)","author":"Senin Pavel","year":"2018","unstructured":"Pavel Senin, Jessica Lin, Xing Wang, Tim Oates, Sunil Gandhi, Arnold P Boedihardjo, Crystal Chen, and Susan Frankenstein. 2018. Grammarviz 3.0: Interactive discovery of variable-length time series patterns. ACM Transactions on Knowledge Discovery from Data (TKDD), Vol. 12, 1 (2018), 1-28."},{"key":"e_1_2_2_109_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-662-44845-8_37"},{"key":"e_1_2_2_110_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2013.52"},{"key":"e_1_2_2_111_1","volume-title":"Mehmet Ugur Gudelek, and Ahmet Murat Ozbayoglu","author":"Sezer Omer Berat","year":"2020","unstructured":"Omer Berat Sezer, Mehmet Ugur Gudelek, and Ahmet Murat Ozbayoglu. 2020. Financial time series forecasting with deep learning: A systematic literature review: 2005-2019. Applied soft computing, Vol. 90 (2020), 106181."},{"key":"e_1_2_2_112_1","doi-asserted-by":"publisher","DOI":"10.1145\/1401890.1401966"},{"key":"e_1_2_2_113_1","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/btq422"},{"key":"e_1_2_2_114_1","doi-asserted-by":"publisher","DOI":"10.14778\/3611479.3611536"},{"key":"e_1_2_2_115_1","volume-title":"International Conference on Learning Representations.","author":"Wu Haixu","year":"2023","unstructured":"Haixu Wu, Tengge Hu, Yong Liu, Hang Zhou, Jianmin Wang, and Mingsheng Long. 2023. TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis. In International Conference on Learning Representations."},{"key":"e_1_2_2_116_1","doi-asserted-by":"publisher","DOI":"10.4236\/jcc.2022.106005"},{"key":"e_1_2_2_117_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2024.101641"},{"key":"e_1_2_2_118_1","volume-title":"International Conference on Machine Learning. PMLR, 25038-25054","author":"Yang Ling","year":"2022","unstructured":"Ling Yang and Shenda Hong. 2022. Unsupervised time-series representation learning with iterative bilinear temporal-spectral fusion. In International Conference on Machine Learning. PMLR, 25038-25054."},{"key":"e_1_2_2_119_1","doi-asserted-by":"crossref","unstructured":"Yufeng Yu Yuelong Zhu Dingsheng Wan Qun Zhao and Huan Liu. 2019. A novel trend symbolic aggregate approximation for time series. arXiv preprint arXiv:1905.00421 (2019).","DOI":"10.1007\/978-3-030-19063-7_65"},{"key":"e_1_2_2_120_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i8.20881"},{"key":"e_1_2_2_121_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447548.3467401"},{"key":"e_1_2_2_122_1","first-page":"3988","article-title":"Self-supervised contrastive pre-training for time series via time-frequency consistency","volume":"35","author":"Zhang Xiang","year":"2022","unstructured":"Xiang Zhang, Ziyuan Zhao, Theodoros Tsiligkaridis, and Marinka Zitnik. 2022. Self-supervised contrastive pre-training for time series via time-frequency consistency. Advances in Neural Information Processing Systems, Vol. 35 (2022), 3988-4003.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_2_123_1","unstructured":"Zhiyuan Zhao Alexander Rodriguez and B Aditya Prakash. 2023. Performative time-series forecasting. arXiv preprint arXiv:2310.06077 (2023)."},{"key":"e_1_2_2_124_1","doi-asserted-by":"crossref","unstructured":"Tian Zhou Peisong Niu Liang Sun Rong Jin et al. 2023. One fits all: Power general time series analysis by pretrained lm. Advances in neural information processing systems Vol. 36 (2023) 43322-43355.","DOI":"10.52202\/075280-1877"}],"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3725357","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T19:00:17Z","timestamp":1774983617000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3725357"}},"issued":{"date-parts":[[2025,6,17]]},"references-count":124,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,6,17]]}},"alternative-id":["10.1145\/3725357"],"URL":"https:\/\/doi.org\/10.1145\/3725357","ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"published":{"date-parts":[[2025,6,17]]}},{"indexed":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T07:17:12Z","timestamp":1779175032145,"version":"3.51.4"},"reference-count":69,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,3,26]],"date-time":"2024-03-26T00:00:00Z","timestamp":1711411200000},"content-version":"vor","delay-in-days":14,"URL":"http:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100006374","name":"National Science Foundation","doi-asserted-by":"publisher","award":["1122374, CCF-1845763"],"award-info":[{"award-number":["1122374, CCF-1845763"]}],"id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000015","name":"Department of Energy","doi-asserted-by":"crossref","award":["DE-SC0018947"],"award-info":[{"award-number":["DE-SC0018947"]}],"id":[{"id":"10.13039\/100000015","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100000185","name":"DARPA","doi-asserted-by":"crossref","award":["HR0011-18-3-0007"],"award-info":[{"award-number":["HR0011-18-3-0007"]}],"id":[{"id":"10.13039\/100000185","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Jump Center - SRC and DARPA"},{"DOI":"10.13039\/501100006374","name":"Google","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2024,3,12]]},"abstract":"<jats:p>Nucleus decompositions have been shown to be a useful tool for finding dense subgraphs. The coreness value of a clique represents its density based on the number of other cliques it is adjacent to. One useful output of nucleus decomposition is to generate a hierarchy among dense subgraphs at different resolutions. However, existing parallel algorithms for nucleus decomposition do not generate this hierarchy, and only compute the coreness values. This paper presents a scalable parallel algorithm for hierarchy construction, with practical optimizations, such as interleaving the coreness computation with hierarchy construction and using a concurrent union-find data structure in an innovative way to generate the hierarchy. We also introduce a parallel approximation algorithm for nucleus decomposition, which achieves much lower span in theory and better performance in practice. We prove strong theoretical bounds on the work and span (parallel time) of our algorithms.<\/jats:p>\n          <jats:p>On a 30-core machine with two-way hyper-threading, our parallel hierarchy construction algorithm achieves up to a 58.84x speedup over the state-of-the-art sequential hierarchy construction algorithm by Sariyuce et al. and up to a 30.96x self-relative parallel speedup. On the same machine, our approximation algorithm achieves a 3.3x speedup over our exact algorithm, while generating coreness estimates with a multiplicative error of 1.33x on average.<\/jats:p>","DOI":"10.1145\/3639287","type":"journal-article","created":{"date-parts":[[2024,3,26]],"date-time":"2024-03-26T18:51:32Z","timestamp":1711479092000},"page":"1-27","source":"Crossref","is-referenced-by-count":2,"title":["Parallel Algorithms for Hierarchical Nucleus Decomposition"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4485-5492","authenticated-orcid":false,"given":"Jessica","family":"Shi","sequence":"first","affiliation":[{"name":"MIT CSAIL, Cambridge, MA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0685-064X","authenticated-orcid":false,"given":"Laxman","family":"Dhulipala","sequence":"additional","affiliation":[{"name":"University of Maryland, College Park, College Park, MA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6163-6625","authenticated-orcid":false,"given":"Julian","family":"Shun","sequence":"additional","affiliation":[{"name":"MIT CSAIL, Cambridge, MA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,3,26]]},"reference":[{"key":"e_1_2_2_1_1","first-page":"11","article-title":"Truss-Based Community Search","volume":"10","author":"Akbas Esra","year":"2017","unstructured":"Esra Akbas and Peixiang Zhao. 2017. Truss-Based Community Search: A Truss-Equivalence Based Indexing Approach. Proc. VLDB Endow., Vol. 10, 11 (Aug. 2017), 1298--1309.","journal-title":"A Truss-Equivalence Based Indexing Approach. Proc. VLDB Endow."},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2933267.2933299"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-4-2"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.5555\/3433701.3433833"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPEC.2019.8916473"},{"key":"e_1_2_2_6_1","volume-title":"ACM Symposium on Parallelism in Algorithms and Architectures (SPAA).","author":"Blelloch Guy E.","year":"2020","unstructured":"Guy E. Blelloch, Daniel Anderson, and Laxman Dhulipala. 2020. Brief Announcement: ParlayLib -- A Toolkit for Parallel Algorithms on Shared-Memory Multicore Machines. In ACM Symposium on Parallelism in Algorithms and Architectures (SPAA)."},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/324133.324234"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.14778\/3401960.3401971"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/BigData.2014.7004264"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1137\/0214017"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE53745.2022.00090"},{"key":"e_1_2_2_12_1","unstructured":"Jonathan Cohen. 2008. Trusses: Cohesive Subgraphs for Social Network Analysis. National Security Agency Technical Report Vol. 16 3.1 (2008)."},{"key":"e_1_2_2_13_1","volume-title":"Dense Subgroup Identifying in Social Network. In International Conference on Advances in Social Networks Analysis and Mining. 555--556","author":"Conghuan Ye","year":"2011","unstructured":"Ye Conghuan. 2011. Dense Subgroup Identifying in Social Network. In International Conference on Advances in Social Networks Analysis and Mining. 555--556."},{"key":"e_1_2_2_14_1","volume-title":"IEEE High Performance Extreme Computing Conference (HPEC). 1--6.","author":"Conte Alessio","year":"2018","unstructured":"Alessio Conte, Daniele De Sensi, Roberto Grossi, Andrea Marino, and Luca Versari. 2018. Discovering k-Trusses in Large-Scale Networks. In IEEE High Performance Extreme Computing Conference (HPEC). 1--6."},{"key":"e_1_2_2_15_1","volume-title":"Introduction to Algorithms (3. ed.)","author":"Cormen Thomas H.","unstructured":"Thomas H. Cormen, Charles E. Leiserson, Ronald L. Rivest, and Clifford Stein. 2009. Introduction to Algorithms (3. ed.). MIT Press."},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3087556.3087580"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/FOCS54457.2022.00077"},{"key":"e_1_2_2_18_1","volume-title":"Nucleus Decomposition in Probabilistic Graphs: Hardness and Algorithms. In IEEE International Conference on Data Engineering (ICDE). 218--231","author":"Esfahani Fatemeh","year":"2022","unstructured":"Fatemeh Esfahani, Venkatesh Srinivasan, Alex Thomo, and Kui Wu. 2022. Nucleus Decomposition in Probabilistic Graphs: Hardness and Algorithms. In IEEE International Conference on Data Engineering (ICDE). 218--231."},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.14778\/3342263.3342645"},{"key":"e_1_2_2_20_1","volume-title":"Computing the Degeneracy of Large Graphs. In Latin American Symposium on Theoretical Informatics. 250--260","author":"Farach-Colton Martin","year":"2014","unstructured":"Martin Farach-Colton and Meng-Tsung Tsai. 2014. Computing the Degeneracy of Large Graphs. In Latin American Symposium on Theoretical Informatics. 250--260."},{"key":"e_1_2_2_21_1","first-page":"e150","article-title":"MotifCut","volume":"22","author":"Fratkin Eugene","year":"2006","unstructured":"Eugene Fratkin, Brian T Naughton, Douglas L Brutlag, and Serafim Batzoglou. 2006. MotifCut: Regulatory Motifs Finding with Maximum Density Subgraphs. Bioinformatics, Vol. 22, 14 (2006), e150--e157.","journal-title":"Regulatory Motifs Finding with Maximum Density Subgraphs. Bioinformatics"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1137\/0220066"},{"key":"e_1_2_2_23_1","volume-title":"Proceedings of the 36th International Conference on Machine Learning. 2201--2210","author":"Ghaffari Mohsen","year":"2019","unstructured":"Mohsen Ghaffari, Silvio Lattanzi, and Slobodan Mitrovi\u0107. 2019. Improved Parallel Algorithms for Density-Based Network Clustering. In Proceedings of the 36th International Conference on Machine Learning. 2201--2210."},{"key":"e_1_2_2_24_1","volume-title":"Proc. VLDB Endow. 721--732","author":"Gibson David","year":"2005","unstructured":"David Gibson, Ravi Kumar, and Andrew Tomkins. 2005. Discovering Large Dense Subgraphs in Massive Graphs. In Proc. VLDB Endow. 721--732."},{"key":"e_1_2_2_25_1","volume-title":"IEEE Symposium on Foundations of Computer Science (FOCS). 698--710","author":"Gil J.","unstructured":"J. Gil, Y. Matias, and U. Vishkin. 1991. Towards a Theory of Nearly Constant Time Parallel Algorithms. In IEEE Symposium on Foundations of Computer Science (FOCS). 698--710."},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2019.2960226"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2588555.2610495"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.14778\/2856318.2856323"},{"key":"e_1_2_2_29_1","volume-title":"Efficient Algorithms for Parallel Bi-core Decomposition. In Symposium on Algorithmic Principles of Computer Systems (APOCS). 17--32","author":"Huang Yihao","year":"2023","unstructured":"Yihao Huang, Claire Wang, Jessica Shi, and Julian Shun. 2023. Efficient Algorithms for Parallel Bi-core Decomposition. In Symposium on Algorithmic Principles of Computer Systems (APOCS). 17--32."},{"key":"e_1_2_2_30_1","volume-title":"Introduction to Parallel Algorithms","author":"Jaja J.","unstructured":"J. Jaja. 1992. Introduction to Parallel Algorithms. Addison-Wesley Professional."},{"key":"e_1_2_2_31_1","volume-title":"ACM Symposium on Principles of Distributed Computing (PODC). 75--82","author":"Siddhartha","unstructured":"Siddhartha V. Jayanti and Robert E. Tarjan. 2016. A Randomized Concurrent Algorithm for Disjoint Set Union. In ACM Symposium on Principles of Distributed Computing (PODC). 75--82."},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2018.2835441"},{"key":"e_1_2_2_33_1","volume-title":"IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). 1482--1491","author":"Kabir H.","unstructured":"H. Kabir and K. Madduri. 2017a. Parallel k-Core Decomposition on Multicore Platforms. In IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). 1482--1491."},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPEC.2017.8091052"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.14778\/2850469.2850471"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.14778\/3430915.3430929"},{"key":"e_1_2_2_37_1","unstructured":"Jure Leskovec and Andrej Krevl. 2019. SNAP Datasets: Stanford Large Network Dataset Collection. http:\/\/snap.stanford.edu\/data."},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2013.158"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.14778\/3446095.3446099"},{"key":"e_1_2_2_40_1","first-page":"1075","article-title":"Efficient ((\u03b1), (\u03b2))-Core Computation in Bipartite Graphs","volume":"29","author":"Liu Boge","year":"2020","unstructured":"Boge Liu, Long Yuan, Xuemin Lin, Lu Qin, Wenjie Zhang, and Jingren Zhou. 2020. Efficient ((\u03b1), (\u03b2))-Core Computation in Bipartite Graphs. Proc. VLDB Endow., Vol. 29, 5 (2020), 1075--1099.","journal-title":"Proc. VLDB Endow."},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3490148.3538569"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPEC.2018.8547718"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSS.2020.3026574"},{"key":"e_1_2_2_44_1","doi-asserted-by":"crossref","unstructured":"Qi Luo Dongxiao Yu Hao Sheng Jiguo Yu and Xiuzhen Cheng. 2021. Distributed Algorithm for Truss Maintenance in Dynamic Graphs. In Parallel and Distributed Computing Applications and Technologies (PDCAT). 104--115.","DOI":"10.1007\/978-3-030-69244-5_9"},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/2402.322385"},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2012.124"},{"key":"e_1_2_2_47_1","volume-title":"Motif-Driven Dense Subgraph Discovery in Directed and Labeled Networks. In The Web Conference (WWW). 379--390","author":"Ahmet Erdem","year":"2021","unstructured":"Ahmet Erdem Sariy\u00fc ce. 2021. Motif-Driven Dense Subgraph Discovery in Directed and Labeled Networks. In The Web Conference (WWW). 379--390."},{"key":"e_1_2_2_48_1","first-page":"425","article-title":"Incremental k-Core Decomposition","volume":"25","author":"Sariy\u00fcce Ahmet Erdem","year":"2016","unstructured":"Ahmet Erdem Sariy\u00fcce, Buug ra Gedik, Gabriela Jacques-Silva, Kun-Lung Wu, and \u00dcmit V cC ataly\u00fcrek. 2016. Incremental k-Core Decomposition: Algorithms and Evaluation. Proc. VLDB Endow., Vol. 25, 3 (2016), 425--447.","journal-title":"Algorithms and Evaluation. Proc. VLDB Endow."},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.14778\/3021924.3021927"},{"key":"e_1_2_2_50_1","volume-title":"Peeling Bipartite Networks for Dense Subgraph Discovery. In ACM International Conference on Web Search and Data Mining (WSDM). 504--512","author":"Erdem Ahmet","unstructured":"Ahmet Erdem Sariy\u00fc ce and Ali Pinar. 2018. Peeling Bipartite Networks for Dense Subgraph Discovery. In ACM International Conference on Web Search and Data Mining (WSDM). 504--512."},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.14778\/3275536.3275540"},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3057742"},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1016\/0378-8733(83)90028-X"},{"key":"e_1_2_2_54_1","volume-title":"Parallel Clique Counting and Peeling Algorithms. In SIAM Conference on Applied and Computational Discrete Algorithms (ACDA). 135--146","author":"Shi Jessica","year":"2021","unstructured":"Jessica Shi, Laxman Dhulipala, and Julian Shun. 2021. Parallel Clique Counting and Peeling Algorithms. In SIAM Conference on Applied and Computational Discrete Algorithms (ACDA). 135--146."},{"key":"e_1_2_2_55_1","first-page":"3","article-title":"Theoretically and Practically Efficient Parallel Nucleus Decomposition","volume":"15","author":"Shi Jessica","year":"2022","unstructured":"Jessica Shi, Laxman Dhulipala, and Julian Shun. 2022. Theoretically and Practically Efficient Parallel Nucleus Decomposition. Proc. VLDB Endow., Vol. 15, 3 (feb 2022), 583--596.","journal-title":"Proc. VLDB Endow."},{"key":"e_1_2_2_56_1","unstructured":"Jessica Shi Laxman Dhulipala and Julian Shun. 2023. Parallel Algorithms for Hierarchical Nucleus Decomposition. arxiv: 2306.08623 [cs.DC]"},{"key":"e_1_2_2_57_1","volume-title":"Parallel Algorithms for Butterfly Computations. In SIAM Symposium on Algorithmic Principles of Computer Systems (APoCS). 16--30","author":"Shi Jessica","year":"2020","unstructured":"Jessica Shi and Julian Shun. 2020. Parallel Algorithms for Butterfly Computations. In SIAM Symposium on Algorithmic Principles of Computer Systems (APoCS). 16--30."},{"key":"e_1_2_2_58_1","volume-title":"Truss Decomposition on Shared-Memory Parallel Systems. In IEEE High Performance Extreme Computing Conference (HPEC). 1--6.","author":"Smith Shaden","year":"2017","unstructured":"Shaden Smith, Xing Liu, Nesreen K Ahmed, Ancy Sarah Tom, Fabrizio Petrini, and George Karypis. 2017. Truss Decomposition on Shared-Memory Parallel Systems. In IEEE High Performance Extreme Computing Conference (HPEC). 1--6."},{"key":"e_1_2_2_59_1","article-title":"Fully Dynamic Approximate k-Core Decomposition in Hypergraphs","volume":"14","author":"Sun Bintao","year":"2020","unstructured":"Bintao Sun, T.-H. Hubert Chan, and Mauro Sozio. 2020. Fully Dynamic Approximate k-Core Decomposition in Hypergraphs. ACM Trans. Knowl. Discov. Data (TKDD), Vol. 14, 4, Article 39 (May 2020).","journal-title":"ACM Trans. Knowl. Discov. Data (TKDD)"},{"key":"e_1_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/2736277.2741098"},{"key":"e_1_2_2_61_1","volume-title":"Finding Dense Subgraph for Community Detection on Social Network Based on Information Diffusion. In International Conference on Data and Software Engineering (ICoDSE). 1--6.","author":"Venica Liptia","year":"2021","unstructured":"Liptia Venica and Gusti Ayu Putri Saptawati. 2021. Finding Dense Subgraph for Community Detection on Social Network Based on Information Diffusion. In International Conference on Data and Software Engineering (ICoDSE). 1--6."},{"key":"e_1_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.14778\/2311906.2311909"},{"key":"e_1_2_2_63_1","volume-title":"Efficient Bitruss Decomposition for Large-Scale Bipartite Graphs. In IEEE International Conference on Data Engineering (ICDE). 661--672","author":"Wang Kai","year":"2020","unstructured":"Kai Wang, Xuemin Lin, Lu Qin, Wenjie Zhang, and Ying Zhang. 2020. Efficient Bitruss Decomposition for Large-Scale Bipartite Graphs. In IEEE International Conference on Data Engineering (ICDE). 661--672."},{"key":"e_1_2_2_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2018.2833070"},{"key":"e_1_2_2_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2012.35"},{"key":"e_1_2_2_66_1","volume-title":"Unboundedness and Efficiency of Truss Maintenance in Evolving Graphs. In ACM SIGMOD International Conference on Management of Data. 1024--1041","author":"Zhang Yikai","year":"2019","unstructured":"Yikai Zhang and Jeffrey Xu Yu. 2019. Unboundedness and Efficiency of Truss Maintenance in Evolving Graphs. In ACM SIGMOD International Conference on Management of Data. 1024--1041."},{"key":"e_1_2_2_67_1","volume-title":"IEEE International Conference on Data Engineering (ICDE). 337--348","author":"Zhang Y.","unstructured":"Y. Zhang, J. X. Yu, Y. Zhang, and L. Qin. 2017. A Fast Order-Based Approach for Core Maintenance. In IEEE International Conference on Data Engineering (ICDE). 337--348."},{"key":"e_1_2_2_68_1","doi-asserted-by":"publisher","DOI":"10.14778\/2535568.2448942"},{"key":"e_1_2_2_69_1","doi-asserted-by":"crossref","unstructured":"Zhaonian Zou. 2016. Bitruss Decomposition of Bipartite Graphs. In Database Systems for Advanced Applications (DASFAA). 218--233.","DOI":"10.1007\/978-3-319-32049-6_14"}],"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3639287","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3639287","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3639287","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T15:16:28Z","timestamp":1755789388000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3639287"}},"issued":{"date-parts":[[2024,3,12]]},"references-count":69,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,3,12]]}},"alternative-id":["10.1145\/3639287"],"URL":"https:\/\/doi.org\/10.1145\/3639287","ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"published":{"date-parts":[[2024,3,12]]}},{"indexed":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T21:10:14Z","timestamp":1775596214426,"version":"3.50.1"},"reference-count":63,"publisher":"Association for Computing Machinery (ACM)","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2026,4,2]]},"abstract":"<jats:p>Log-structured merge-trees (LSM-trees) are widely used in modern key-value stores, but their multi-level structure reduces lookup efficiency, especially for range scans. Existing caching solutions, like block caches or full query caches, are memory-inefficient because they fail to exploit a critical asymmetry: eliminating an I\/O from upper LSM-tree levels requires caching far fewer key-value pairs (KVs) than from lower levels. To address this, we introduce Group Cache, which uses KV Groups, the minimal set of KVs within a block for a specific query, as its fundamental caching unit. By employing a size-aware policy that prioritizes small, high-utility KV Groups, Group Cache maximizes I\/O savings per unit of memory. We also address practical challenges like compaction management, intra-group hotness difference and scalability. Our theoretical analysis and extensive experiments in RocksDB demonstrate that Group Cache significantly outperforms traditional caching methods, achieving up to 3\u00d7 faster query performance with the same memory budget, or achieving similar performance while using 75% less space.<\/jats:p>","DOI":"10.1145\/3786661","type":"journal-article","created":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T17:54:13Z","timestamp":1775584453000},"page":"1-26","source":"Crossref","is-referenced-by-count":0,"title":["Improving Range Scan Performance in LSM-trees with Group Caching"],"prefix":"10.1145","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8006-4220","authenticated-orcid":false,"given":"Hengrui","family":"Wang","sequence":"first","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-3075-6147","authenticated-orcid":false,"given":"Jiaoyi","family":"Zhang","sequence":"additional","affiliation":[{"name":"East China Normal University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-8463-1768","authenticated-orcid":false,"given":"Jiansheng","family":"Qiu","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-1720-8258","authenticated-orcid":false,"given":"Fangzhou","family":"Yuan","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-4821-1558","authenticated-orcid":false,"given":"Huanchen","family":"Zhang","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,7]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"[n.d.]. Asynchronous IO in RocksDB. https:\/\/github.com\/facebook\/rocksdb\/wiki\/Asynchronous-IO."},{"key":"e_1_2_1_2_1","unstructured":"2012. Sanjay Ghemawat and Jeff Dean. LevelDB. https:\/\/github.com\/google\/leveldb."},{"key":"e_1_2_1_3_1","unstructured":"2024. Apache. Cassandra. http:\/\/cassandra.apache.org."},{"key":"e_1_2_1_4_1","unstructured":"2024. Apache. HBase. http:\/\/hbase.apache.org\/."},{"key":"e_1_2_1_5_1","unstructured":"2024. CockroachDB. https:\/\/github.com\/cockroachdb\/cockroach.."},{"key":"e_1_2_1_6_1","unstructured":"2024. Facebook. RocksDB. https:\/\/github.com\/facebook\/rocksdb.."},{"key":"e_1_2_1_7_1","unstructured":"2024. LinkedIn. Voldemort. http:\/\/www.project-voldemort.com."},{"key":"e_1_2_1_8_1","unstructured":"2024. WiredTiger. https:\/\/github.com\/wiredtiger\/wiredtiger.."},{"key":"e_1_2_1_9_1","volume-title":"ACM SIGMOD Conference. https:\/\/api.semanticscholar.org\/CorpusID:11759711","author":"Armstrong Timothy G.","unstructured":"Timothy G. Armstrong, Vamsi Ponnekanti, Dhruba Borthakur, and Mark D. Callaghan. 2013. LinkBench: a database benchmark based on the Facebook social graph. In ACM SIGMOD Conference. https:\/\/api.semanticscholar.org\/CorpusID:11759711"},{"key":"e_1_2_1_10_1","volume-title":"14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17)","author":"Berger Daniel S.","year":"2017","unstructured":"Daniel S. Berger, Ramesh K. Sitaraman, and Mor Harchol-Balter. 2017. AdaptSize: Orchestrating the Hot Object Memory Cache in a Content Delivery Network. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17). USENIX Association, Boston, MA, 483--498. https:\/\/www.usenix.org\/conference\/nsdi17\/technical-sessions\/presentation\/berger"},{"key":"e_1_2_1_11_1","volume-title":"USENIX Conference on File and Storage Technologies. https:\/\/api.semanticscholar.org\/CorpusID:211137004","author":"Cao Zhichao","year":"2020","unstructured":"Zhichao Cao, Siying Dong, Sagar Vemuri, and David Hung-Chang Du. 2020. Characterizing, Modeling, and Benchmarking RocksDB Key-Value Workloads at Facebook. In USENIX Conference on File and Storage Technologies. https:\/\/api.semanticscholar.org\/CorpusID:211137004"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.14778\/3659437.3659447"},{"key":"e_1_2_1_13_1","volume-title":"Proceedings of the 13th Usenix Conference on Networked Systems Design and Implementation","author":"Cidon Asaf","year":"2016","unstructured":"Asaf Cidon, Assaf Eisenman, Mohammad Alizadeh, and Sachin Katti. 2016. Cliffhanger: scaling performance cliffs in web memory caches. In Proceedings of the 13th Usenix Conference on Networked Systems Design and Implementation (Santa Clara, CA) (NSDI'16). USENIX Association, USA, 379--392."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3588726"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1807128.1807152"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3639258"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3064054"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2915219"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3196927"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3319903"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3457273"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3064033"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3698820"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2741948.2741973"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.14778\/3681954.3682012"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3314041"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3314041"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1002\/rsa.20008"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.orl.2009.03.011"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1017\/S096354830400625X"},{"key":"e_1_2_1_31_1","volume-title":"2021 USENIX Annual Technical Conference (USENIX ATC 21)","author":"Kassa Hiwot Tadese","year":"2021","unstructured":"Hiwot Tadese Kassa, Jason Akers, Mrinmoy Ghosh, Zhichao Cao, Vaibhav Gogte, and Ronald Dreslinski. 2021. Improving Performance of Flash Based Key-Value Stores Using Storage Class Memory as a Volatile Memory Extension. In 2021 USENIX Annual Technical Conference (USENIX ATC 21). USENIX Association, 821--837. https:\/\/www.usenix.org\/conference\/atc21\/presentation\/kassa"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3526167"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.14778\/3199517.3199519"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3654978"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3389731"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3187009.3177736"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.14778\/3421424.3421425"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2020.3027191"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3617333"},{"key":"e_1_2_1_40_1","unstructured":"Jiansheng Qiu Fangzhou Yuan Mingyu Gao and Huanchen Zhang. 2024. HotRAP: Hot Record Retention and Promotion for LSM-trees with Tiered Storage. arXiv:2402.02070 [cs.DB] https:\/\/arxiv.org\/abs\/2402.02070"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3582016.3582052"},{"key":"e_1_2_1_42_1","volume-title":"USENIX Workshop on Hot Topics in Storage and File Systems. https:\/\/api.semanticscholar.org\/CorpusID:46996220","author":"Raju Pandian","year":"2018","unstructured":"Pandian Raju, Soujanya Ponnapalli, Evan Kaminsky, Gilad Oved, Zachary Keener, Vijay Chidambaram, and Ittai Abraham. 2018. mLSM: Making Authenticated Storage Faster in Ethereum. In USENIX Workshop on Hot Topics in Storage and File Systems. https:\/\/api.semanticscholar.org\/CorpusID:46996220"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3056102"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476249.3476274"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/2213836.2213862"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0166--5316(01)00045--1"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.14778\/3529337.3529347"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3654944"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3725344"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE60146.2024.00044"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3626736"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE55515.2023.00158"},{"key":"e_1_2_1_53_1","volume-title":"AC-Key: Adaptive Caching for LSM-based Key-Value Stores. In 2020 USENIX Annual Technical Conference (USENIX ATC 20)","author":"Wu Fenggang","unstructured":"Fenggang Wu, Ming-Hong Yang, Baoquan Zhang, and David H.C. Du. 2020. AC-Key: Adaptive Caching for LSM-based Key-Value Stores. In 2020 USENIX Annual Technical Conference (USENIX ATC 20). USENIX Association, 603--615. https:\/\/www.usenix.org\/conference\/atc20\/presentation\/wu-fenggang"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3593856.3595887"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3600006.3613147"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.14778\/3407790.3407803"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3267809.3267846"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.14778\/3561261.3561270"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/3677138"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3196931"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.14778\/3547305.3547320"},{"key":"e_1_2_1_62_1","volume-title":"REMIX: Efficient Range Query for LSM-trees. In 19th USENIX Conference on File and Storage Technologies (FAST 21)","author":"Zhong Wenshao","year":"2021","unstructured":"Wenshao Zhong, Chen Chen, Xingbo Wu, and Song Jiang. 2021. REMIX: Efficient Range Query for LSM-trees. In 19th USENIX Conference on File and Storage Technologies (FAST 21). USENIX Association, 51--64. https:\/\/www.usenix.org\/conference\/fast21\/presentation\/zhong"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/3725327"}],"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3786661","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T20:03:12Z","timestamp":1775592192000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3786661"}},"issued":{"date-parts":[[2026,4,2]]},"references-count":63,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,4,2]]}},"alternative-id":["10.1145\/3786661"],"URL":"https:\/\/doi.org\/10.1145\/3786661","ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"published":{"date-parts":[[2026,4,2]]}},{"indexed":{"date-parts":[[2026,3,8]],"date-time":"2026-03-08T00:35:41Z","timestamp":1772930141555,"version":"3.50.1"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2024,5,29]],"date-time":"2024-05-29T00:00:00Z","timestamp":1716940800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2024,5,29]]},"abstract":"<jats:p>The study of uncertain graphs is crucial in diverse fields, including but not limited to protein interaction analysis, viral marketing, and network reliability. Processing queries on uncertain graphs presents formidable challenges due to the vast probabilistic space they encapsulate. While existing systems employ batch processing to address these challenges, their performance is often compromised by the suboptimal selection of parallel graph traversal methods, the excessive costs in random number generation, and additional workloads intrinsic to batch processing. In this paper, we introduce uBlade, an efficient batch-processing framework for uncertain graph queries on multi-core CPUs. uBlade utilizes the work-efficient graph traversal, achieving superior parallelism in the batch processing model. Additionally, our Quasi-Sampling technique reduces the random number generation cost by a factor of B, with O(B) denoting the batch size. We further examine the extra workload resulting from batch processing and introduce an efficient strategy to reorder possible worlds, minimizing this associated overhead. Through comprehensive evaluations, we showcase that uBlade achieves up to two orders of magnitude speedups against the state-of-the-art CPU and GPU-based solutions.<\/jats:p>","DOI":"10.1145\/3654982","type":"journal-article","created":{"date-parts":[[2024,5,30]],"date-time":"2024-05-30T09:44:53Z","timestamp":1717062293000},"page":"1-24","source":"Crossref","is-referenced-by-count":4,"title":["uBlade: Efficient Batch Processing for Uncertainty Graph Queries"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-8243-3947","authenticated-orcid":false,"given":"Siyuan","family":"Yao","sequence":"first","affiliation":[{"name":"Singapore Management University &amp; National University of Singapore, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9646-291X","authenticated-orcid":false,"given":"Yuchen","family":"Li","sequence":"additional","affiliation":[{"name":"Singapore Management University, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4060-9438","authenticated-orcid":false,"given":"Shixuan","family":"Sun","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8748-3225","authenticated-orcid":false,"given":"Jiaxin","family":"Jiang","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8618-4581","authenticated-orcid":false,"given":"Bingsheng","family":"He","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,5,30]]},"reference":[{"issue":"2","key":"e_1_2_1_1_1","first-page":"15","article-title":"Managing uncertainty in social networks","volume":"30","author":"Adar E.","year":"2007","unstructured":"E. Adar and C. Re. Managing uncertainty in social networks. IEEE Data Eng. Bull., 30(2):15--22, 2007.","journal-title":"IEEE Data Eng. Bull."},{"key":"e_1_2_1_2_1","volume-title":"On a routing problem. Quarterly of applied mathematics, 16(1):87--90","author":"Bellman R.","year":"1958","unstructured":"R. Bellman. On a routing problem. Quarterly of applied mathematics, 16(1):87--90, 1958."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.14778\/2350229.2350254"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/1557019.1557047"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2015.7113362"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1516360.1516438"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDCS.2015.109"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2016.2535444"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3490422.3502358"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2014.45"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3087556.3087580"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF01386390"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFCOMW.2019.8845088"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/TR.1986.4335388"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1142\/S0218488513500074"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFCOM.2007.201"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1963192.1963217"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3078447.3078452"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.14778\/2002938.2002941"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1006\/jpdc.1997.1404"},{"key":"e_1_2_1_21_1","first-page":"87","volume-title":"International Workshop of Algorithmic Aspects of Cloud Computing","author":"Kassiano V.","year":"2016","unstructured":"V. Kassiano, A. Gounaris, A. N. Papadopoulos, and K. Tsichlas. Mining uncertain graphs: An overview. In International Workshop of Algorithmic Aspects of Cloud Computing, pages 87--116. Springer, 2016."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.14778\/3324301.3324304"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/956750.956769"},{"key":"e_1_2_1_24_1","first-page":"535","volume-title":"Proceedings of the 17th International Conference on Extending Database Technology","author":"Khan A.","year":"2014","unstructured":"A. Khan, F. Bonchi, A. Gionis, and F. Gullo. Fast reliability search in uncertain graphs. In Proceedings of the 17th International Conference on Extending Database Technology, pages 535--546, 2014."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.14778\/2824032.2824133"},{"key":"e_1_2_1_26_1","volume-title":"Morgan & Claypool","author":"Khan A.","year":"2018","unstructured":"A. Khan, Y. Ye, and L. Chen. On uncertain graphs. synthesis lectures on data management. Morgan & Claypool, 2018."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2011.243"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1038\/nature04670"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.14778\/3565838.3565844"},{"key":"e_1_2_1_30_1","unstructured":"J. Leskovec and A. Krevl. SNAP Datasets: Stanford large network dataset collection. http:\/\/snap.stanford.edu\/data June 2014."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-011-0220-3"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2015.2485212"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2019.00047"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2018.2807843"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.14778\/2794367.2794376"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2882959"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3457253"},{"key":"e_1_2_1_38_1","volume-title":"An indexing framework for queries on probabilistic graphs. ACM Transactions on Database Systems (TODS), 42(2):1--34","author":"Maniu S.","year":"2017","unstructured":"S. Maniu, R. Cheng, and P. Senellart. An indexing framework for queries on probabilistic graphs. ACM Transactions on Database Systems (TODS), 42(2):1--34, 2017."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0196-6774(03)00076-2"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2588555.2593668"},{"key":"e_1_2_1_41_1","first-page":"06","article-title":"Interactive exploration of heterogeneous biological networks with biomine explorer","author":"V.","year":"2019","unstructured":"V. Podpe?an, ?. Ram?ak, K. Gruden, H. Toivonen, and N. Lavra?. Interactive exploration of heterogeneous biological networks with biomine explorer. Bioinformatics, 06 2019.","journal-title":"Bioinformatics"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.14778\/1920841.1920967"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v29i1.9277"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.14778\/3450980.3450988"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/2442516.2442530"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476249.3476257"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.14778\/2735496.2735507"},{"key":"e_1_2_1_48_1","volume-title":"The complexity of enumeration and reliability problems. siam Journal on Computing, 8(3):410--421","author":"Valiant L. G.","year":"1979","unstructured":"L. G. Valiant. The complexity of enumeration and reliability problems. siam Journal on Computing, 8(3):410--421, 1979."},{"issue":"2","key":"e_1_2_1_49_1","first-page":"2019","article-title":"Scaleg: A distributed disk-based system for vertex-centric graph processing","volume":"35","author":"Wang X.","year":"2023","unstructured":"X. Wang, D. Wen, L. Qin, L. Chang, Y. Zhang, and W. Zhang. Scaleg: A distributed disk-based system for vertex-centric graph processing. IEEE Transactions on Knowledge and Data Engineering, 35(2):2019--2033, 2023.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3639288"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3524059.3532379"},{"key":"e_1_2_1_52_1","first-page":"441","volume-title":"2018 USENIX Annual Technical Conference (USENIX ATC 18)","author":"Zhang Y.","year":"2018","unstructured":"Y. Zhang, X. Liao, H. Jin, L. Gu, L. He, B. He, and H. Liu. {CGraph}: A correlations-aware approach for efficient concurrent iterative graph processing. In 2018 USENIX Annual Technical Conference (USENIX ATC 18), pages 441--452, 2018."},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3295500.3356143"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.4304\/jnw.9.9.2353-2359"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2015.64"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2017.70"}],"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3654982","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3654982","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T14:39:11Z","timestamp":1755787151000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3654982"}},"issued":{"date-parts":[[2024,5,29]]},"references-count":56,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,5,29]]}},"alternative-id":["10.1145\/3654982"],"URL":"https:\/\/doi.org\/10.1145\/3654982","ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"published":{"date-parts":[[2024,5,29]]}},{"indexed":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T19:01:58Z","timestamp":1774983718511,"version":"3.50.1"},"reference-count":57,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2025,2,10]],"date-time":"2025-02-10T00:00:00Z","timestamp":1739145600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100006374","name":"HORIZON EUROPE Framework Programme","doi-asserted-by":"publisher","award":["101070122"],"award-info":[{"award-number":["101070122"]}],"id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2025,2,10]]},"abstract":"<jats:p>Entity Resolution (ER) is typically implemented as a batch task that processes all available data before identifying duplicate records. However, applications with time or computational constraints, e.g., those running in the cloud, require a progressive approach that produces results in a pay-as-you-go fashion. Numerous algorithms have been proposed for Progressive ER in the literature. In this work, we propose a novel framework for Progressive Entity Matching that organizes relevant techniques into four consecutive steps: (i) filtering, which reduces the search space to the most likely candidate matches, (ii) weighting, which associates every pair of candidate matches with a similarity score, (iii) scheduling, which prioritizes the execution of the candidate matches so that the real duplicates precede the non-matching pairs, and (iv) matching, which applies a complex, matching function to the pairs in the order defined by the previous step. We associate each step with existing and novel techniques, illustrating that our framework overall generates a superset of the main existing works in the field. We select the most representative combinations resulting from our framework and fine-tune them over 10 established datasets for Record Linkage and 8 for Deduplication, with our results indicating that our taxonomy yields a wide range of high performing progressive techniques both in terms of effectiveness and time efficiency.<\/jats:p>","DOI":"10.1145\/3709715","type":"journal-article","created":{"date-parts":[[2025,2,11]],"date-time":"2025-02-11T15:45:06Z","timestamp":1739288706000},"page":"1-25","source":"Crossref","is-referenced-by-count":6,"title":["Progressive Entity Matching: A Design Space Exploration"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-8307-8843","authenticated-orcid":false,"given":"Jakub","family":"Maciejewski","sequence":"first","affiliation":[{"name":"National and Kapodistrian University of Athens, Athens, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3465-1197","authenticated-orcid":false,"given":"Konstantinos","family":"Nikoletos","sequence":"additional","affiliation":[{"name":"National and Kapodistrian University of Athens, Athens, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7298-9431","authenticated-orcid":false,"given":"George","family":"Papadakis","sequence":"additional","affiliation":[{"name":"National and Kapodistrian University of Athens, Athens, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6332-0296","authenticated-orcid":false,"given":"Yannis","family":"Velegrakis","sequence":"additional","affiliation":[{"name":"University of Trento and Utrecht University, Utrecht, Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,2,11]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.is.2019.02.006"},{"key":"e_1_2_2_2_1","volume-title":"Enriching Word Vectors with Subword Information. CoRR","author":"Bojanowski Piotr","year":"2016","unstructured":"Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tom\u00e1s Mikolov. 2016. Enriching Word Vectors with Subword Information. CoRR, Vol. abs\/1607.04606 (2016)."},{"key":"e_1_2_2_3_1","volume-title":"Entity Resolution, and Duplicate Detection","author":"Christen Peter","unstructured":"Peter Christen. 2012. Data Matching - Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection. Springer."},{"key":"e_1_2_2_4_1","doi-asserted-by":"crossref","unstructured":"Peter Christen Ross W. Gayler and David Hawking. 2009. Similarity-aware indexing for real-time entity resolution. In CIKM. 1565--1568.","DOI":"10.1145\/1645953.1646173"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3418896"},{"key":"e_1_2_2_6_1","volume-title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT. 4171--4186.","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT. 4171--4186."},{"key":"e_1_2_2_7_1","volume-title":"Big Data Integration","author":"Dong Xin Luna","unstructured":"Xin Luna Dong and Divesh Srivastava. 2015. Big Data Integration. Morgan & Claypool Publishers."},{"key":"e_1_2_2_8_1","first-page":"44","article-title":"Unicorn","volume":"53","author":"Fan Ju","year":"2024","unstructured":"Ju Fan, Jianhong Tu, Guoliang Li, Peng Wang, Xiaoyong Du, Xiaofeng Jia, Song Gao, and Nan Tang. 2024. Unicorn: A Unified Multi-Tasking Matching Model. SIGMOD Rec., Vol. 53, 1 (2024), 44--53.","journal-title":"A Unified Multi-Tasking Matching Model. SIGMOD Rec."},{"key":"e_1_2_2_9_1","volume-title":"BEER: Blocking for Effective Entity Resolution. In SIGMOD. 2711--2715.","author":"Galhotra Sainyam","year":"2021","unstructured":"Sainyam Galhotra, Donatella Firmani, Barna Saha, and Divesh Srivastava. 2021a. BEER: Blocking for Effective Entity Resolution. In SIGMOD. 2711--2715."},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-021-00656-7"},{"key":"e_1_2_2_11_1","unstructured":"Leonardo Gazzarri and Melanie Herschel. 2023. Progressive Entity Resolution over Incremental Data. In EDBT. 80--91."},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3626711"},{"key":"e_1_2_2_13_1","first-page":"2018","article-title":"Entity Resolution","volume":"5","author":"Getoor Lise","year":"2012","unstructured":"Lise Getoor and Ashwin Machanavajjhala. 2012. Entity Resolution: Theory, Practice & Open Challenges. PVLDB, Vol. 5, 12 (2012), 2018--2019.","journal-title":"Theory, Practice & Open Challenges. PVLDB"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.14778\/1687627.1687771"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.14778\/3485450.3485455"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/TBDATA.2019.2921572"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.14778\/1920841.1920904"},{"key":"e_1_2_2_18_1","volume-title":"ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. In ICLR.","author":"Lan Zhenzhong","year":"2020","unstructured":"Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. In ICLR."},{"key":"e_1_2_2_19_1","volume-title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR, Vol. abs\/1907.11692 (2019)."},{"key":"e_1_2_2_20_1","unstructured":"Jakub Maciejewski Konstantinos Nikoletos George Papadakis and Yannis Velegrakis. 2022. Progressive Entity Matching: A Design Space Exploration. https:\/\/github.com\/JacobMaciejewski\/PER-Design-Space-Exploration\/paper\/PEMextended.pdf. In SIGMOD (extended version)."},{"key":"e_1_2_2_21_1","doi-asserted-by":"crossref","unstructured":"Venkata Vamsikrishna Meduri Lucian Popa Prithviraj Sen and Mohamed Sarwat. 2020. A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching. In SIGMOD. 1133--1147.","DOI":"10.1145\/3318464.3380597"},{"key":"e_1_2_2_22_1","unstructured":"Tom\u00e1s Mikolov Kai Chen Greg Corrado and Jeffrey Dean. 2013a. Efficient Estimation of Word Representations in Vector Space. In ICLR."},{"key":"e_1_2_2_23_1","unstructured":"Tom\u00e1s Mikolov Ilya Sutskever Kai Chen Gregory S. Corrado and Jeffrey Dean. 2013b. Distributed Representations of Words and Phrases and their Compositionality. In NIPS. 3111--3119."},{"key":"e_1_2_2_24_1","doi-asserted-by":"crossref","unstructured":"Sidharth Mudgal Han Li Theodoros Rekatsinas AnHai Doan Youngchoon Park Ganesh Krishnan Rohit Deep Esteban Arcaute and Vijay Raghavendra. 2018. Deep Learning for Entity Matching: A Design Space Exploration. In SIGMOD. ACM 19--34.","DOI":"10.1145\/3183713.3196926"},{"key":"e_1_2_2_25_1","volume-title":"Open benchmark for filtering techniques in entity resolution. The VLDB Journal","author":"Neuhof Franziska","year":"2024","unstructured":"Franziska Neuhof, Marco Fisichella, George Papadakis, Konstantinos Nikoletos, Nikolaus Augsten, Wolfgang Nejdl, and Manolis Koubarakis. 2024. Open benchmark for filtering techniques in entity resolution. The VLDB Journal (2024), 1--26."},{"key":"e_1_2_2_26_1","unstructured":"Daniel Obraczka Jonathan Schuchart and Erhard Rahm. 2021. Embedding-Assisted Entity Resolution for Knowledge Graphs. In KGCW@ESWC."},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-023-00791-3"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE55515.2023.00389"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-01878-7"},{"key":"e_1_2_2_30_1","doi-asserted-by":"crossref","unstructured":"George Papadakis Nishadi Kirielle Peter Christen and Themis Palpanas. 2024. A Critical Re-evaluation of Record Linkage Benchmarks for Learning-Based Matching Algorithms. In ICDE. 3435--3448.","DOI":"10.1109\/ICDE60146.2024.00265"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.is.2020.101565"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3377455"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.14778\/2947618.2947624"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2014.2359666"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.14778\/3583140.3583163"},{"key":"e_1_2_2_36_1","volume-title":"Manning","author":"Pennington Jeffrey","year":"2014","unstructured":"Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014. Glove: Global Vectors for Word Representation. In EMNLP. 1532--1543."},{"key":"e_1_2_2_37_1","article-title":"Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. J. Mach. Learn. Res., Vol. 21 (2020), 140:1--140:67.","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_2_2_38_1","doi-asserted-by":"crossref","unstructured":"Banda Ramadan and Peter Christen. 2015. Unsupervised Blocking Key Selection for Real-Time Entity Resolution. In PAKDD. 574--585.","DOI":"10.1007\/978-3-319-18032-8_45"},{"key":"e_1_2_2_39_1","article-title":"Dynamic Sorted Neighborhood Indexing for Real-Time Entity Resolution","volume":"6","author":"Ramadan Banda","year":"2015","unstructured":"Banda Ramadan, Peter Christen, Huizhi Liang, and Ross W. Gayler. 2015. Dynamic Sorted Neighborhood Indexing for Real-Time Entity Resolution. ACM J. Data Inf. Qual., Vol. 6, 4 (2015), 15:1--15:29.","journal-title":"ACM J. Data Inf. Qual."},{"key":"e_1_2_2_40_1","doi-asserted-by":"crossref","unstructured":"Banda Ramadan Peter Christen Huizhi Liang Ross W. Gayler and David Hawking. 2013. Dynamic Similarity-Aware Inverted Indexing for Real-Time Entity Resolution. In PAKDD. 47--58.","DOI":"10.1007\/978-3-642-40319-4_5"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.7250\/csimq.2018-16.04"},{"key":"e_1_2_2_42_1","volume-title":"a distilled version of BERT: smaller, faster, cheaper and lighter. CoRR","author":"Sanh Victor","year":"2019","unstructured":"Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. CoRR, Vol. abs\/1910.01108 (2019)."},{"key":"e_1_2_2_43_1","doi-asserted-by":"crossref","unstructured":"Sunita Sarawagi and Anuradha Bhamidipaty. 2002. Interactive deduplication using active learning. In KDD. ACM 269--278.","DOI":"10.1145\/775047.775087"},{"key":"e_1_2_2_44_1","volume-title":"Proceedings of the 5th International Workshop on Ontology Matching (OM-2010)","volume":"689","author":"Shvaiko Pavel","year":"2010","unstructured":"Pavel Shvaiko, J\u00e9r\u00f4me Euzenat, Fausto Giunchiglia, Heiner Stuckenschmidt, Ming Mao, and Isabel F. Cruz (Eds.). 2010. Proceedings of the 5th International Workshop on Ontology Matching (OM-2010), Shanghai, China, November 7, 2010. Vol. 689."},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2018.2852763"},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.14778\/3523210.3523226"},{"key":"e_1_2_2_47_1","unstructured":"Kaitao Song Xu Tan Tao Qin Jianfeng Lu and Tie-Yan Liu. 2020. MPNet: Masked and Permuted Pre-training for Language Understanding. In NIPS."},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3139987"},{"key":"e_1_2_2_49_1","first-page":"228","article-title":"Exploring the Design Space of Unsupervised Blocking with Pre-trained Language Models in Entity Resolution","volume":"14176","author":"Sun Chenchen","year":"2023","unstructured":"Chenchen Sun, Yuyuan Jin, Yang Xu, Derong Shen, Tiezheng Nie, and Xite Wang. 2023. Exploring the Design Space of Unsupervised Blocking with Pre-trained Language Models in Entity Resolution. In ADMA, Vol. 14176. 228--244.","journal-title":"ADMA"},{"key":"e_1_2_2_50_1","first-page":"2459","article-title":"Deep Learning for Blocking in Entity Matching","volume":"14","author":"Thirumuruganathan Saravanan","year":"2021","unstructured":"Saravanan Thirumuruganathan, Han Li, Nan Tang, Mourad Ouzzani, Yash Govind, Derek Paulsen, Glenn Fung, and AnHai Doan. 2021. Deep Learning for Blocking in Entity Matching: A Design Space Exploration. PVLDB, Vol. 14, 11 (2021), 2459--2472.","journal-title":"A Design Space Exploration. PVLDB"},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.14778\/2732977.2732982"},{"key":"e_1_2_2_52_1","unstructured":"Runhui Wang and Yongfeng Zhang. 2024. Pre-trained Language Models for Entity Blocking: A Reproducibility Study. In NAACL. 8712--8722."},{"key":"e_1_2_2_53_1","unstructured":"Wenhui Wang Furu Wei Li Dong Hangbo Bao Nan Yang and Ming Zhou. 2020. MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers. In NIPS."},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2012.43"},{"key":"e_1_2_2_55_1","volume-title":"Le","author":"Yang Zhilin","year":"2019","unstructured":"Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019. XLNet: Generalized Autoregressive Pretraining for Language Understanding. In NeurIPS. 5754--5764."},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.14778\/3598581.3598594"},{"key":"e_1_2_2_57_1","first-page":"4026","article-title":"BrewER","volume":"16","author":"Zecchini Luca","year":"2023","unstructured":"Luca Zecchini, Giovanni Simonini, Sonia Bergamaschi, and Felix Naumann. 2023. BrewER: Entity Resolution On-Demand. PVLDB, Vol. 16, 12 (2023), 4026--4029.","journal-title":"Entity Resolution On-Demand. PVLDB"}],"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3709715","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3709715","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T18:20:55Z","timestamp":1774981255000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3709715"}},"issued":{"date-parts":[[2025,2,10]]},"references-count":57,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,2,10]]}},"alternative-id":["10.1145\/3709715"],"URL":"https:\/\/doi.org\/10.1145\/3709715","ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"published":{"date-parts":[[2025,2,10]]}},{"indexed":{"date-parts":[[2026,6,1]],"date-time":"2026-06-01T17:38:41Z","timestamp":1780335521750,"version":"3.54.1"},"reference-count":42,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,3,12]],"date-time":"2024-03-12T00:00:00Z","timestamp":1710201600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2024,3,12]]},"abstract":"<jats:p>Sparse operators, i.e., operators that take sparse tensors as input, are of great importance in deep learning models. Due to the diverse sparsity patterns in different sparse tensors, it is challenging to optimize sparse operators by seeking an optimal sparse format, i.e., leading to the lowest operator latency. Existing works propose to decompose a sparse tensor into several parts and search for a hybrid of sparse formats to handle diverse sparse patterns. However, they often make a trade-off between search space and search time: their search spaces are limited in some cases, resulting in limited operator running efficiency they can achieve. In this paper, we try to extend the search space in its breadth (by doing flexible sparse tensor transformations) and depth (by enabling multi-level decomposition). We formally define the multi-level sparse format decomposition problem, which is NP-hard, and we propose a framework STile for it. To search efficiently, a greedy algorithm is used, which is guided by a cost model about the latency of computing a sub-task of the original operator after decomposing the sparse tensor. Experiments of two common kinds of sparse operators, SpMM and SDDMM, are conducted on various sparsity patterns, and we achieve 2.1-18.0\u00d7 speedup against cuSPARSE on SpMMs and 1.5 - 6.9\u00d7 speedup against DGL on SDDMM. The search time is less than one hour for any tested sparse operator, which can be amortized.<\/jats:p>","DOI":"10.1145\/3639323","type":"journal-article","created":{"date-parts":[[2024,3,26]],"date-time":"2024-03-26T18:51:32Z","timestamp":1711479092000},"page":"1-26","source":"Crossref","is-referenced-by-count":6,"title":["STile: Searching Hybrid Sparse Formats for Sparse Deep Learning Operators Automatically"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6462-5825","authenticated-orcid":false,"given":"Jingzhi","family":"Fang","sequence":"first","affiliation":[{"name":"The Hong Kong University of Science and Technology, Hong Kong SAR, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8364-3674","authenticated-orcid":false,"given":"Yanyan","family":"Shen","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8618-9806","authenticated-orcid":false,"given":"Yue","family":"Wang","sequence":"additional","affiliation":[{"name":"Shenzhen Institute of Computing Sciences, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8257-5806","authenticated-orcid":false,"given":"Lei","family":"Chen","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology &amp; The Hong Kong University of Science and Technology (Guangzhou), Hong Kong SAR, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,3,26]]},"reference":[{"key":"e_1_2_2_1_1","volume-title":"Boman","author":"Ahrens Peter","year":"2020","unstructured":"Peter Ahrens and Erik G. Boman. 2020. On Optimal Partitioning For Sparse Matrices In Variable Block Row Format. CoRR, Vol. abs\/2005.12414 (2020). showeprint[arXiv]2005.12414 https:\/\/arxiv.org\/abs\/2005.12414"},{"key":"e_1_2_2_2_1","volume-title":"Diameter of the world-wide web. nature","author":"Albert R\u00e9ka","year":"1999","unstructured":"R\u00e9ka Albert, Hawoong Jeong, and Albert-L\u00e1szl\u00f3 Barab\u00e1si. 1999. Diameter of the world-wide web. nature, Vol. 401, 6749 (1999), 130--131."},{"key":"e_1_2_2_3_1","volume-title":"Longformer: The Long-Document Transformer. CoRR","author":"Beltagy Iz","year":"2020","unstructured":"Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The Long-Document Transformer. CoRR, Vol. abs\/2004.05150 (2020). showeprint[arXiv]2004.05150 https:\/\/arxiv.org\/abs\/2004.05150"},{"key":"e_1_2_2_4_1","volume-title":"The Tenth International Conference on Learning Representations, ICLR 2022","author":"Chen Beidi","year":"2022","unstructured":"Beidi Chen, Tri Dao, Kaizhao Liang, Jiaming Yang, Zhao Song, Atri Rudra, and Christopher R\u00e9. 2022. Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25--29, 2022. OpenReview.net. https:\/\/openreview.net\/forum?id=Nfl-iXa-y7R"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476182"},{"key":"e_1_2_2_6_1","volume-title":"Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509","author":"Child Rewon","year":"2019","unstructured":"Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019a. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509 (2019)."},{"key":"e_1_2_2_7_1","volume-title":"Generating Long Sequences with Sparse Transformers. CoRR","author":"Child Rewon","year":"2019","unstructured":"Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019b. Generating Long Sequences with Sparse Transformers. CoRR, Vol. abs\/1904.10509 (2019). showeprint[arXiv]1904.10509 http:\/\/arxiv.org\/abs\/1904.10509"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1693453.1693471"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3276493"},{"key":"e_1_2_2_10_1","volume-title":"cuSPARSE :: CUDA Toolkit Documentation v11.7.1. https:\/\/docs.nvidia.com\/cuda\/cusparse\/index.html Retrieved","author":"NVIDIA Corporation","year":"2023","unstructured":"NVIDIA Corporation. 2022. cuSPARSE :: CUDA Toolkit Documentation v11.7.1. https:\/\/docs.nvidia.com\/cuda\/cusparse\/index.html Retrieved July 15, 2023 from"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC41404.2022.00071"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00021"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.5555\/3433701.3433723"},{"key":"e_1_2_2_14_1","volume-title":"Inductive representation learning on large graphs. Advances in neural information processing systems","author":"Hamilton Will","year":"2017","unstructured":"Will Hamilton, Zhitao Ying, and Jure Leskovec. 2017. Inductive representation learning on large graphs. Advances in neural information processing systems, Vol. 30 (2017)."},{"key":"e_1_2_2_15_1","volume-title":"Dally","author":"Han Song","year":"2015","unstructured":"Song Han, Jeff Pool, John Tran, and William J. Dally. 2015. Learning both Weights and Connections for Efficient Neural Networks. CoRR, Vol. abs\/1506.02626 (2015). showeprint[arXiv]1506.02626 http:\/\/arxiv.org\/abs\/1506.02626"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1137\/S0097539704444750"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447818.3461703"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3293883.3295712"},{"key":"e_1_2_2_19_1","volume-title":"Open graph benchmark: Datasets for machine learning on graphs. Advances in neural information processing systems","author":"Hu Weihua","year":"2020","unstructured":"Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. 2020a. Open graph benchmark: Datasets for machine learning on graphs. Advances in neural information processing systems, Vol. 33 (2020), 22118--22133."},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00075"},{"key":"e_1_2_2_21_1","volume-title":"Blocking Techniques for Sparse Matrix Multiplication on Tensor Accelerators. CoRR","author":"Labini Paolo Sylos","year":"2022","unstructured":"Paolo Sylos Labini, Massimo Bernaschi, Francesco Silvestri, and Flavio Vella. 2022. Blocking Techniques for Sparse Matrix Multiplication on Tensor Accelerators. CoRR, Vol. abs\/2202.05868 (2022). showeprint[arXiv]2202.05868 https:\/\/arxiv.org\/abs\/2202.05868"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.829"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/CGO53902.2022.9741270"},{"key":"e_1_2_2_24_1","volume-title":"Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting. Advances in neural information processing systems","author":"Li Shiyang","year":"2019","unstructured":"Shiyang Li, Xiaoyong Jin, Yao Xuan, Xiyou Zhou, Wenhu Chen, Yu-Xiang Wang, and Xifeng Yan. 2019. Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting. Advances in neural information processing systems, Vol. 32 (2019)."},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC41404.2022.00042"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2751205.2751209"},{"key":"e_1_2_2_27_1","volume-title":"Structured pruning of a bert-based question answering model. arXiv preprint arXiv:1910.06360","author":"McCarley JS","year":"2019","unstructured":"JS McCarley, Rishav Chakravarti, and Avirup Sil. 2019. Structured pruning of a bert-based question answering model. arXiv preprint arXiv:1910.06360 (2019)."},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS51385.2021.00016"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS49936.2021.00016"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503221.3508431"},{"key":"e_1_2_2_31_1","volume-title":"Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020","author":"Sanh Victor","year":"2020","unstructured":"Victor Sanh, Thomas Wolf, and Alexander M. Rush. 2020. Movement Pruning: Adaptive Sparsity by Fine-Tuning. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6--12, 2020, virtual, Hugo Larochelle, Marc'Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (Eds.). https:\/\/proceedings.neurips.cc\/paper\/2020\/hash\/eae15aabaa768ae4a5993a8a4f4fa6e4-Abstract.html"},{"key":"e_1_2_2_32_1","volume-title":"Collective classification in network data. AI magazine","author":"Sen Prithviraj","year":"2008","unstructured":"Prithviraj Sen, Galileo Namata, Mustafa Bilgic, Lise Getoor, Brian Galligher, and Tina Eliassi-Rad. 2008. Collective classification in network data. AI magazine, Vol. 29, 3 (2008), 93--93."},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3315508.3329973"},{"key":"e_1_2_2_34_1","unstructured":"Minjie Wang Da Zheng Zihao Ye Quan Gan Mufei Li Xiang Song Jinjing Zhou Chao Ma Lingfan Yu Yu Gai et al. 2019. Deep graph library: A graph-centric highly-performant package for graph neural networks. arXiv preprint arXiv:1909.01315 (2019)."},{"key":"e_1_2_2_35_1","volume-title":"TC-GNN: Accelerating Sparse Graph Neural Network Computation Via Dense Tensor Core on GPUs. CoRR","author":"Wang Yuke","year":"2052","unstructured":"Yuke Wang, Boyuan Feng, and Yufei Ding. 2021a. TC-GNN: Accelerating Sparse Graph Neural Network Computation Via Dense Tensor Core on GPUs. CoRR, Vol. abs\/2112.02052 (2021). showeprint[arXiv]2112.02052 https:\/\/arxiv.org\/abs\/2112.02052"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00088"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/1498765.1498785"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3575693.3575742"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3582016.3582047"},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCD53106.2021.00092"},{"key":"e_1_2_2_41_1","volume-title":"16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22)","author":"Zheng Ningxin","year":"2022","unstructured":"Ningxin Zheng, Bin Lin, Quanlu Zhang, Lingxiao Ma, Yuqing Yang, Fan Yang, Yang Wang, Mao Yang, and Lidong Zhou. 2022. SparTA: Deep-Learning Model Sparsity via $$Tensor-with-Sparsity-Attribute$$. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22). 213--232."},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3575693.3575723"}],"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3639323","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3639323","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,21]],"date-time":"2025-08-21T15:16:59Z","timestamp":1755789419000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3639323"}},"issued":{"date-parts":[[2024,3,12]]},"references-count":42,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,3,12]]}},"alternative-id":["10.1145\/3639323"],"URL":"https:\/\/doi.org\/10.1145\/3639323","ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"published":{"date-parts":[[2024,3,12]]}},{"indexed":{"date-parts":[[2026,8,11]],"date-time":"2026-08-11T15:46:55Z","timestamp":1786463215143,"version":"build-2736575974"},"reference-count":60,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2023,6,13]],"date-time":"2023-06-13T00:00:00Z","timestamp":1686614400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"publisher","award":["BI2011\/1,BI2011\/2,SFB1053"],"award-info":[{"award-number":["BI2011\/1,BI2011\/2,SFB1053"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2023,6,13]]},"abstract":"<jats:p>Remote data structures built with one-sided Remote Direct Memory Access (RDMA) are at the heart of many disaggregated database management systems today. Concurrent access to these data structures by thousands of remote workers necessitates a highly efficient synchronization scheme. Remarkably, our investigation reveals that existing synchronization schemes display substantial variations in performance and scalability. Even worse, some schemes do not correctly synchronize, resulting in rare and hard-to-detect data corruption. Motivated by these observations, we conduct the first comprehensive analysis of one-sided synchronization techniques and provide general principles for correct synchronization using one-sided RDMA. Our research demonstrates that adherence to these principles not only guarantees correctness but also results in substantial performance enhancements.<\/jats:p>","DOI":"10.1145\/3589276","type":"journal-article","created":{"date-parts":[[2023,6,20]],"date-time":"2023-06-20T20:26:45Z","timestamp":1687292805000},"page":"1-26","source":"Crossref","is-referenced-by-count":32,"title":["Design Guidelines for Correct, Efficient, and Scalable Synchronization using One-Sided RDMA"],"prefix":"10.1145","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1602-4512","authenticated-orcid":false,"given":"Tobias","family":"Ziegler","sequence":"first","affiliation":[{"name":"Technische Universit\u00e4t Darmstadt, Darmstadt, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-1968-8863","authenticated-orcid":false,"given":"Jacob","family":"Nelson-Slivon","sequence":"additional","affiliation":[{"name":"Lehigh University, Bethlehem, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5676-8017","authenticated-orcid":false,"given":"Viktor","family":"Leis","sequence":"additional","affiliation":[{"name":"Technische Universit\u00e4t M\u00fcnchen, M\u00fcnchen, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2744-7836","authenticated-orcid":false,"given":"Carsten","family":"Binnig","sequence":"additional","affiliation":[{"name":"Technische Universit\u00e4t Darmstadt, Darmstadt, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,6,20]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"ARM. 2018. Arm CoreLink CCI-550 Cache Coherent Interconnect Technical Reference Manual. https:\/\/developer.arm. com\/documentation\/100282\/0100\/?lang=en. https:\/\/developer.arm.com\/documentation\/100282\/0100\/?lang=en"},{"key":"e_1_2_2_2_1","unstructured":"ARM. 2021. Introducing the AMBA Coherent Hub Interface. https:\/\/developer.arm.com\/documentation\/102407\/0100"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3457560"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2017.2729545"},{"key":"e_1_2_2_5_1","volume-title":"Using RDMA for Lock Management. CoRR abs\/1507.03274","author":"Chung Yeounoh","year":"2015","unstructured":"Yeounoh Chung and Erfan Zamanian. 2015. Using RDMA for Lock Management. CoRR abs\/1507.03274 (2015). arXiv:1507.03274 http:\/\/arxiv.org\/abs\/1507.03274"},{"key":"e_1_2_2_6_1","unstructured":"NVIDIA Coporation. 2021. NVIDIA InfiniBand Adaptive Routing Technology. Whitepaper WP-10326-001_v01."},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2983990.2984033"},{"key":"e_1_2_2_8_1","unstructured":"Aleksandar Dragojevic Dushyanth Narayanan Miguel Castro and Orion Hodson. 2014. FaRM: Fast Remote Memory. In NSDI."},{"key":"e_1_2_2_9_1","doi-asserted-by":"crossref","unstructured":"Aleksandar Dragojevic Dushyanth Narayanan Edmund B. Nightingale Matthew Renzelmann Alex Shamis Anirudh Badam and Miguel Castro. 2015. No compromises: distributed transactions with consistency availability and performance. In SOSP.","DOI":"10.1145\/2815400.2815425"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3587096"},{"key":"e_1_2_2_11_1","doi-asserted-by":"crossref","unstructured":"Philipp Fent Alexander van Renen Andreas Kipf Viktor Leis Thomas Neumann and Alfons Kemper. 2020. Low- Latency Communication for Fast DBMS Using RDMA and Shared Memory. In ICDE.","DOI":"10.1109\/ICDE48307.2020.00131"},{"key":"e_1_2_2_12_1","unstructured":"Torsten Hoefler Duncan Roweth Keith Underwood Bob Alverson Mark Griswold Vahid Tabatabaee Mohan Kalkunte Surendra Anubolu Siyan Shen Abdul Kabbani Moray McLaren and Steve Scott. 2023. Datacenter Ethernet and RDMA: Issues at Hyperscale. arXiv:2302.03337 [cs.NI]"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11704-"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2005.23"},{"key":"e_1_2_2_15_1","volume-title":"InfiniBand Architecture Specification","author":"InfiniBand Trade Association 2007.","unstructured":"InfiniBand Trade Association 2007. InfiniBand Architecture Specification Volume 1. InfiniBand Trade Association. Release 1.2.1."},{"key":"e_1_2_2_16_1","unstructured":"InfiniBand Trade Association. 2010. RDMA Over Converged Ethernet (RoCE). https:\/\/cw.infinibandta.org\/document\/ dl\/7148."},{"key":"e_1_2_2_17_1","unstructured":"Intel. 2012. Intel Data Direct I\/O Technology (Intel DDIO): A P rimer. https:\/\/www.intel.com\/content\/dam\/www\/ public\/us\/en\/documents\/technology-briefs\/data-direct-i-o-technology-brief.pdf"},{"key":"e_1_2_2_18_1","volume-title":"Andersen","author":"Kalia Anuj","year":"2014","unstructured":"Anuj Kalia, Michael Kaminsky, and David G. Andersen. 2014. Using RDMA efficiently for key-value services. In SIGCOMM."},{"key":"e_1_2_2_19_1","volume-title":"Andersen","author":"Kalia Anuj","year":"2016","unstructured":"Anuj Kalia, Michael Kaminsky, and David G. Andersen. 2016. Design Guidelines for High Performance RDMA Systems. login Usenix Mag. 41, 3 (2016)."},{"key":"e_1_2_2_20_1","volume-title":"Andersen","author":"Kalia Anuj","year":"2016","unstructured":"Anuj Kalia, Michael Kaminsky, and David G. Andersen. 2016. FaSST: Fast, Scalable and Simple Distributed Transactions with Two-Sided (RDMA) Datagram RPCs. In OSDI."},{"key":"e_1_2_2_21_1","unstructured":"Tejas Karmarkar. 2015. Availability of Linux RDMA on Microsoft Azure. Online. https:\/\/azure.microsoft.com\/enus\/ blog\/azure-linux-rdma-hpc-available\/"},{"key":"e_1_2_2_22_1","volume-title":"Farview: Disaggregated Memory with Operator Off-loading for Database Engines. In 12th Conference on Innovative Data Systems Research, CIDR 2022","author":"Korolija Dario","year":"2022","unstructured":"Dario Korolija, Dimitrios Koutsoukos, Kimberly Keeton, Konstantin Taranov, Dejan S. Milojicic, and Gustavo Alonso. 2022. Farview: Disaggregated Memory with Operator Off-loading for Database Engines. In 12th Conference on Innovative Data Systems Research, CIDR 2022 , Chaminade, CA, USA, January 9--12, 2022. www.cidrdb.org. https:\/\/www.cidrdb.org\/cidr2022\/papers\/p11-korolija.pdf"},{"key":"e_1_2_2_23_1","first-page":"73","article-title":"Optimistic Lock Coupling: A Scalable and Efficient General-Purpose Synchronization Method","volume":"42","author":"Leis Viktor","year":"2019","unstructured":"Viktor Leis, Michael Haubenschild, and Thomas Neumann. 2019. Optimistic Lock Coupling: A Scalable and Efficient General-Purpose Synchronization Method. IEEE Data Eng. Bull. 42 (2019), 73--84.","journal-title":"IEEE Data Eng. Bull."},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2933349.2933352"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/HOTI.2007.8"},{"key":"e_1_2_2_26_1","volume-title":"Panda","author":"Liu Jiuxing","year":"2003","unstructured":"Jiuxing Liu, Jiesheng Wu, Sushmitha P. Kini, Pete Wyckoff, and Dhabaleswar K. Panda. 2003. High performance RDMA-based MPI implementation over InfiniBand. In ICS."},{"key":"e_1_2_2_27_1","doi-asserted-by":"crossref","unstructured":"Simon Loesing Markus Pilman Thomas Etter and Donald Kossmann. 2015. On the Design and Scalability of Distributed Shared-Data Databases. In SIGMOD.","DOI":"10.1145\/2723372.2751519"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/Cluster48925.2021.00033"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/COMPSAC51774.2021.00021"},{"key":"e_1_2_2_30_1","volume-title":"CPU-Efficient Key-Value Store. In 2013 USENIX Annual Technical Conference","author":"Mitchell Christopher","year":"2013","unstructured":"Christopher Mitchell, Yifeng Geng, and Jinyang Li. 2013. Using One-Sided RDMA Reads to Build a Fast, CPU-Efficient Key-Value Store. In 2013 USENIX Annual Technical Conference, San Jose, CA, USA, June 26--28, 2013, Andrew Birrell and Emin G\u00fcn Sirer (Eds.). USENIX Association, 103--114. https:\/\/www.usenix.org\/conference\/atc13\/technicalsessions\/ presentation\/mitchell"},{"key":"e_1_2_2_31_1","unstructured":"Christopher Mitchell Yifeng Geng and Jinyang Li. 2013. Using One-Sided RDMA Reads to Build a Fast CPU-Efficient Key-Value Store. In USENIX ATC."},{"key":"e_1_2_2_32_1","unstructured":"Christopher Mitchell Kate Montgomery Lamont Nelson Siddhartha Sen and Jinyang Li. 2016. Balancing CPU and Network in the Cell Distributed B-Tree Store. In USENIX ATC."},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/CCGRID.2007.58"},{"key":"e_1_2_2_34_1","doi-asserted-by":"crossref","unstructured":"Jacob Nelson and Roberto Palmieri. 2020. Performance Evaluation of the Impact of NUMA on One-sided RDMA Interactions. In SRDS.","DOI":"10.1109\/SRDS51746.2020.00036"},{"key":"e_1_2_2_35_1","unstructured":"PCI-SIG. 2014. PCI Express Base Specification Revision 4.0. (2014)."},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","unstructured":"R. Recio B. Metzler P. Culley J. Hilland and D. Garcia. 2007. A Remote Direct Memory Access Protocol Specification. Technical Report. https:\/\/doi.org\/10.17487\/rfc5040","DOI":"10.17487\/rfc5040"},{"key":"e_1_2_2_37_1","volume-title":"Robertazzi","author":"Ren Yufei","year":"2013","unstructured":"Yufei Ren, Tan Li, Dantong Yu, Shudong Jin, and Thomas G. Robertazzi. 2013. Design and performance evaluation of NUMA-aware RDMA-based end-to-end data transfer systems. In HiPC."},{"key":"e_1_2_2_38_1","volume-title":"Application- Integrated Far Memory. In 14th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2020","author":"Ruan Zhenyuan","year":"2020","unstructured":"Zhenyuan Ruan, Malte Schwarzkopf, Marcos K. Aguilera, and Adam Belay. 2020. AIFM: High-Performance, Application- Integrated Far Memory. In 14th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2020, Virtual Event, November 4--6, 2020. USENIX Association, 315--332. https:\/\/www.usenix.org\/conference\/osdi20\/presentation\/ruan"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","unstructured":"H. Shah F. Marti W. Noureddine A. Eiriksson and R. Sharp. 2014. Remote Direct Memory Access (RDMA) Protocol Extensions. Technical Report. https:\/\/doi.org\/10.17487\/rfc7306","DOI":"10.17487\/rfc7306"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3452296.3472934"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2010.5416638"},{"key":"e_1_2_2_43_1","unstructured":"Konstantin Taranov Fabian Fischer and Torsten Hoefler. 2022. Efficient RDMA Communication Protocols. arXiv:2212.09134 [cs.NI]"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3452817"},{"key":"e_1_2_2_45_1","volume-title":"Disaggregating Persistent Memory and Controlling Them Remotely: An Exploration of Passive Disaggregated Key-Value Stores. In 2020 USENIX Annual Technical Conference, USENIX ATC 2020","author":"Tsai Shin-Yeh","year":"2020","unstructured":"Shin-Yeh Tsai, Yizhou Shan, and Yiying Zhang. 2020. Disaggregating Persistent Memory and Controlling Them Remotely: An Exploration of Passive Disaggregated Key-Value Stores. In 2020 USENIX Annual Technical Conference, USENIX ATC 2020, July 15--17, 2020, Ada Gavrilovska and Erez Zadok (Eds.). USENIX Association, 33--48. https:\/\/www.usenix.org\/conference\/atc20\/presentation\/tsai"},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/tcc.2021.3116516"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3517824"},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2207.03027"},{"key":"e_1_2_2_49_1","volume-title":"Efficient Usage of One-Sided RDMA for Linear Probing. In International Workshop on Accelerating Analytics and Data Management Systems Using Modern Processor and Storage Architectures, ADMS@VLDB 2020","author":"Wang Tinggang","year":"2020","unstructured":"Tinggang Wang, Shuo Yang, Hideaki Kimura, Garret Swart, and Spyros Blanas. 2020. Efficient Usage of One-Sided RDMA for Linear Probing. In International Workshop on Accelerating Analytics and Data Management Systems Using Modern Processor and Storage Architectures, ADMS@VLDB 2020, Tokyo, Japan, August 31, 2020, Rajesh Bordawekar and Tirthankar Lahiri (Eds.). 1--13. http:\/\/www.adms-conf.org\/2020-camera-ready\/ADMS20_06.pdf"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/2807591.2807614"},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3468520"},{"key":"e_1_2_2_52_1","unstructured":"Xingda Wei Zhiyuan Dong Rong Chen and Haibo Chen. 2018. Deconstructing RDMA-enabled Distributed Transactions: Hybrid is Better!. In OSDI."},{"key":"e_1_2_2_53_1","doi-asserted-by":"crossref","unstructured":"Xingda Wei Jiaxin Shi Yanzhe Chen Rong Chen and Haibo Chen. 2015. Fast in-memory transaction processing using RDMA and HTM. In SOSP.","DOI":"10.1145\/2815400.2815419"},{"key":"e_1_2_2_54_1","volume-title":"The End of a Myth: Distributed Transactions Can Scale. CoRR abs\/1607.00655","author":"Zamanian Erfan","year":"2016","unstructured":"Erfan Zamanian, Carsten Binnig, Tim Kraska, and Tim Harris. 2016. The End of a Myth: Distributed Transactions Can Scale. CoRR abs\/1607.00655 (2016)."},{"key":"e_1_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3471485.3471490"},{"key":"e_1_2_2_56_1","volume-title":"FORD: Fast One-sided RDMA-based Distributed Transactions for Disaggregated Persistent Memory. In 20th USENIX Conference on File and Storage Technologies, FAST 2022","author":"Zhang Ming","year":"2022","unstructured":"Ming Zhang, Yu Hua, Pengfei Zuo, and Lurong Liu. 2022. FORD: Fast One-sided RDMA-based Distributed Transactions for Disaggregated Persistent Memory. In 20th USENIX Conference on File and Storage Technologies, FAST 2022, Santa Clara, CA, USA, February 22--24, 2022, Dean Hildebrand and Donald E. Porter (Eds.). USENIX Association, 51--68. https:\/\/www.usenix.org\/conference\/fast22\/presentation\/zhang-ming"},{"key":"e_1_2_2_57_1","volume-title":"FORD: Fast One-sided RDMA-based Distributed Transactions for Disaggregated Persistent Memory. In 20th USENIX Conference on File and Storage Technologies (FAST 22)","author":"Zhang Ming","year":"2022","unstructured":"Ming Zhang, Yu Hua, Pengfei Zuo, and Lurong Liu. 2022. FORD: Fast One-sided RDMA-based Distributed Transactions for Disaggregated Persistent Memory. In 20th USENIX Conference on File and Storage Technologies (FAST 22). USENIX Association, Santa Clara, CA, 51--68. https:\/\/www.usenix.org\/conference\/fast22\/presentation\/zhang-ming"},{"key":"e_1_2_2_58_1","doi-asserted-by":"publisher","DOI":"10.14778\/3467861.3467877"},{"key":"e_1_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3526187"},{"key":"e_1_2_2_60_1","volume-title":"Carsten Binnig, Rodrigo Fonseca, and Tim Kraska.","author":"Ziegler Tobias","year":"2019","unstructured":"Tobias Ziegler, Sumukha Tumkur Vani, Carsten Binnig, Rodrigo Fonseca, and Tim Kraska. 2019. Designing Distributed Tree-based Index Structures for Fast RDMA-capable Networks. In SIGMOD."},{"key":"e_1_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/3511895"}],"container-title":["Proceedings of the ACM on Management of Data"],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589276","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3589276","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:48:54Z","timestamp":1750182534000},"score":0.0,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589276"}},"issued":{"date-parts":[[2023,6,13]]},"references-count":60,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2023,6,13]]}},"alternative-id":["10.1145\/3589276"],"URL":"https:\/\/doi.org\/10.1145\/3589276","ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"published":{"date-parts":[[2023,6,13]]}}],"items-per-page":20,"query":{"start-index":0,"search-terms":null}}}