{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T04:58:24Z","timestamp":1781326704496,"version":"3.54.1"},"reference-count":54,"publisher":"Association for Computing Machinery (ACM)","issue":"6","funder":[{"name":"NUS Faculty Development Fund"},{"name":"Emerging Areas Research Projects (EARP) Funding Initiative, Singapore Blockchain Innovation Programme"},{"name":"Alibaba Group, Alibaba Innovative Research Program"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2025,12,4]]},"abstract":"<jats:p>With the prevalence of in-database AI-powered analytics, there is an increasing demand for database systems to efficiently manage the ever-expanding number and size of deep learning models. However, existing database systems typically store entire models as monolithic files or apply compression techniques that overlook the structural characteristics of deep learning models, resulting in suboptimal model storage overhead. This paper presents NeurStore, a novel in-database model management system that enables efficient storage and utilization of deep learning models. First, NeurStore employs a tensor-based model storage engine to enable fine-grained model storage within databases. In particular, we enhance the hierarchical navigable small world (HNSW) graph to index tensors, and only store additional deltas for tensors within a predefined similarity threshold to ensure tensor-level deduplication. Second, we propose a delta quantization algorithm that effectively compresses delta tensors, thus achieving a superior compression ratio with controllable model accuracy loss. Finally, we devise a compression-aware model loading mechanism, which improves model utilization performance by enabling direct computation on compressed tensors. Experimental evaluations demonstrate that NeurStore achieves superior compression ratios and competitive model loading throughput compared to state-of-the-art approaches.<\/jats:p>","DOI":"10.1145\/3769809","type":"journal-article","created":{"date-parts":[[2025,12,6]],"date-time":"2025-12-06T04:32:13Z","timestamp":1764995533000},"page":"1-26","source":"Crossref","is-referenced-by-count":0,"title":["NeurStore: Efficient In-database Deep Learning Model Management System"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-5669-5776","authenticated-orcid":false,"given":"Siqi","family":"Xiang","sequence":"first","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-1582-3316","authenticated-orcid":false,"given":"Sheng","family":"Wang","sequence":"additional","affiliation":[{"name":"Alibaba Group, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0914-4580","authenticated-orcid":false,"given":"Xiaokui","family":"Xiao","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0263-8879","authenticated-orcid":false,"given":"Cong","family":"Yue","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4044-7742","authenticated-orcid":false,"given":"Zhanhao","family":"Zhao","sequence":"additional","affiliation":[{"name":"National University of Singapore, Singapore, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4446-1100","authenticated-orcid":false,"given":"Beng Chin","family":"Ooi","sequence":"additional","affiliation":[{"name":"Zhejiang University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,12,5]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"2014. Avazu Dataset. https:\/\/www.kaggle.com\/c\/avazu-ctr-prediction."},{"key":"e_1_2_1_2_1","unstructured":"2020. Beans Dataset. https:\/\/github.com\/AI-Lab-Makerere\/ibean."},{"key":"e_1_2_1_3_1","unstructured":"2024. zlib. https:\/\/zlib.net."},{"key":"e_1_2_1_4_1","unstructured":"2025. Azure SQL. https:\/\/azure.microsoft.com."},{"key":"e_1_2_1_5_1","unstructured":"2025. Hugging Face. https:\/\/huggingface.co."},{"key":"e_1_2_1_6_1","unstructured":"2025. NeurDB. https:\/\/github.com\/neurdb\/neurdb."},{"key":"e_1_2_1_7_1","unstructured":"2025. NeurStore. https:\/\/github.com\/neurdb\/neurstore."},{"key":"e_1_2_1_8_1","unstructured":"2025. Oracle Machine Learning. https:\/\/docs.oracle.com\/en\/database\/oracle\/machine-learning."},{"key":"e_1_2_1_9_1","unstructured":"2025. PostgresML. https:\/\/postgresml.org."},{"key":"e_1_2_1_10_1","unstructured":"2025. Zstandard. https:\/\/github.com\/facebook\/zstd."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.14778\/3611540.3611554"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2019.2953897"},{"key":"e_1_2_1_13_1","volume-title":"EfficientQAT: Efficient Quantization-Aware Training for Large Language Models. CoRR abs\/2407.11062","author":"Chen Mengzhao","year":"2024","unstructured":"Mengzhao Chen, Wenqi Shao, Peng Xu, Jiahao Wang, Peng Gao, Kaipeng Zhang, Yu Qiao, and Ping Luo. 2024. EfficientQAT: Efficient Quantization-Aware Training for Large Language Models. CoRR abs\/2407.11062 (2024)."},{"key":"e_1_2_1_14_1","volume-title":"The Faiss library. CoRR abs\/2401.08281","author":"Douze Matthijs","year":"2024","unstructured":"Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazar\u00e9, Maria Lomeli, Lucas Hosseini, and Herv\u00e9 J\u00e9gou. 2024. The Faiss library. CoRR abs\/2401.08281 (2024)."},{"key":"e_1_2_1_15_1","volume-title":"Vertica-ML: Distributed Machine Learning in Vertica Database. In SIGMOD Conference. 755-768","author":"Fard Arash","year":"2020","unstructured":"Arash Fard, Anh Le, George Larionov,Waqas Dhillon, and Chuck Bear. 2020. Vertica-ML: Distributed Machine Learning in Vertica Database. In SIGMOD Conference. 755-768."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2213836.2213874"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.14778\/3303753.3303754"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.14778\/3603581.3603589"},{"key":"e_1_2_1_19_1","volume-title":"A Survey of Quantization Methods for Efficient Neural Network Inference. CoRR abs\/2103.13630","author":"Gholami Amir","year":"2021","unstructured":"Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer. 2021. A Survey of Quantization Methods for Efficient Neural Network Inference. CoRR abs\/2103.13630 (2021)."},{"key":"e_1_2_1_20_1","volume-title":"Dally","author":"Han Song","year":"2016","unstructured":"Song Han, Huizi Mao, and William J. Dally. 2016. Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding. In ICLR."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.14778\/2367502.2367510"},{"key":"e_1_2_1_22_1","unstructured":"Edward J. Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang Lu Wang and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In ICLR."},{"key":"e_1_2_1_23_1","first-page":"604","article-title":"Approximate Nearest Neighbors: Towards Removing the Curse of Dimensionality","author":"Indyk Piotr","year":"1998","unstructured":"Piotr Indyk and Rajeev Motwani. 1998. Approximate Nearest Neighbors: Towards Removing the Curse of Dimensionality. In STOC. 604-613.","journal-title":"STOC."},{"key":"e_1_2_1_24_1","first-page":"2704","article-title":"Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference","author":"Jacob Benoit","year":"2018","unstructured":"Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew G. Howard, Hartwig Adam, and Dmitry Kalenichenko. 2018. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference. In CVPR. 2704-2713.","journal-title":"CVPR."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2010.57"},{"key":"e_1_2_1_26_1","volume-title":"Quantizing deep convolutional networks for efficient inference: A whitepaper. CoRR abs\/1806.08342","author":"Krishnamoorthi Raghuraman","year":"2018","unstructured":"Raghuraman Krishnamoorthi. 2018. Quantizing deep convolutional networks for efficient inference: A whitepaper. CoRR abs\/1806.08342 (2018)."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2014.2346458"},{"key":"e_1_2_1_28_1","first-page":"28092","article-title":"Post-Training Quantization for Vision Transformer","author":"Liu Zhenhua","year":"2021","unstructured":"Zhenhua Liu, Yunhe Wang, Kai Han, Wei Zhang, Siwei Ma, and Wen Gao. 2021. Post-Training Quantization for Vision Transformer. In NeurIPS. 28092-28103.","journal-title":"NeurIPS."},{"key":"e_1_2_1_29_1","first-page":"1655","article-title":"MLCask: Efficient management of component evolution in collaborative data analytics pipelines","author":"Luo Zhaojing","year":"2021","unstructured":"Zhaojing Luo, Sai Ho Yeung, Meihui Zhang, Kaiping Zheng, Lei Zhu, Gang Chen, Feiyi Fan, Qian Lin, Kee Yuan Ngiam, and Beng Chin Ooi. 2021. MLCask: Efficient management of component evolution in collaborative data analytics pipelines. In ICDE. 1655-1666.","journal-title":"ICDE."},{"key":"e_1_2_1_30_1","first-page":"142","article-title":"Learning Word Vectors for Sentiment Analysis","author":"Maas Andrew L.","year":"2011","unstructured":"Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning Word Vectors for Sentiment Analysis. In ACL-HLT. 142-150.","journal-title":"ACL-HLT."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2889473"},{"key":"e_1_2_1_32_1","first-page":"571","article-title":"Towards Unified Data and Lifecycle Management for Deep Learning","author":"Miao Hui","year":"2017","unstructured":"Hui Miao, Ang Li, Larry S. Davis, and Amol Deshpande. 2017. Towards Unified Data and Lifecycle Management for Deep Learning. In ICDE. 571-582.","journal-title":"ICDE."},{"key":"e_1_2_1_33_1","first-page":"7197","article-title":"Up or Down? Adaptive Rounding for Post-Training Quantization","volume":"119","author":"Nagel Markus","year":"2020","unstructured":"Markus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos, and Tijmen Blankevoort. 2020. Up or Down? Adaptive Rounding for Post-Training Quantization. In ICML, Vol. 119. 7197-7206.","journal-title":"ICML"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-024-4125-9"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2733373.2807410"},{"key":"e_1_2_1_36_1","volume-title":"End-to-end Optimization of Machine Learning Prediction Queries. In SIGMOD Conference. 587-601","author":"Park Kwanghyun","year":"2022","unstructured":"Kwanghyun Park, Karla Saur, Dalitso Banda, Rathijit Sen, Matteo Interlandi, and Konstantinos Karanasos. 2022. End-to-end Optimization of Machine Learning Prediction Queries. In SIGMOD Conference. 587-601."},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3320212"},{"key":"e_1_2_1_38_1","first-page":"1378","article-title":"Revisiting kd-tree for Nearest Neighbor Search","author":"Ram Parikshit","year":"2019","unstructured":"Parikshit Ram and Kaushik Sinha. 2019. Revisiting kd-tree for Nearest Neighbor Search. In KDD. 1378-1388.","journal-title":"KDD."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.14778\/3570690.3570695"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.14778\/3659437.3659441"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.14778\/3685800.3685802"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.14778\/3659437.3659456"},{"key":"e_1_2_1_43_1","volume-title":"MODELDB: A System for Machine Learning Model Management. In CIDR.","author":"Vartak Manasi","year":"2017","unstructured":"Manasi Vartak. 2017. MODELDB: A System for Machine Learning Model Management. In CIDR."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476249.3476255"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.14778\/3641204.3641212"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3526150"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3526142"},{"key":"e_1_2_1_48_1","volume-title":"Omar Mohamed Awad, and Andreas Moshovos","author":"Zadeh Ali Hadi","year":"2020","unstructured":"Ali Hadi Zadeh, Isak Edo, Omar Mohamed Awad, and Andreas Moshovos. 2020. GOBO: Quantizing Attention-Based NLP Models for Low Latency and Energy Efficient Inference. In MICRO. 811-824."},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.14778\/3704965.3704985"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3626246.3654754"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3725326"},{"key":"e_1_2_1_52_1","first-page":"838","volume-title":"ICML","volume":"32","author":"Zhang Ting","year":"2014","unstructured":"Ting Zhang, Chao Du, and Jingdong Wang. 2014. Composite Quantization for Approximate Nearest Neighbor Search. In ICML, Vol. 32. 838-846."},{"key":"e_1_2_1_53_1","volume-title":"Yanyan Shen, Yuncheng Wu, and Meihui Zhang.","author":"Zhao Zhanhao","year":"2025","unstructured":"Zhanhao Zhao, Shaofeng Cai, Haotian Gao, Hexiang Pan, Siqi Xiang, Naili Xing, Gang Chen, Beng Chin Ooi, Yanyan Shen, Yuncheng Wu, and Meihui Zhang. 2025. NeurDB: On the Design and Implementation of an AI-powered Autonomous Database. In CIDR."},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.14778\/3547305.3547325"}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3769809","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T04:44:56Z","timestamp":1781325896000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3769809"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,4]]},"references-count":54,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,12,4]]}},"alternative-id":["10.1145\/3769809"],"URL":"https:\/\/doi.org\/10.1145\/3769809","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,12,4]]}}}