{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T07:11:58Z","timestamp":1779174718993,"version":"3.51.4"},"reference-count":53,"publisher":"Association for Computing Machinery (ACM)","issue":"8","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2024,4]]},"abstract":"<jats:p>The performance of inference with machine learning (ML) models and its integration with analytical query processing have become critical bottlenecks for data analysis in many organizations. An ML inference pipeline typically consists of a preprocessing workflow followed by prediction with an ML model. Current approaches for in-database inference implement preprocessing operators and ML algorithms in the database either natively, by transpiling code to SQL, or by executing user-defined functions in guest languages such as Python. In this work, we present a radically different approach that approximates an end-to-end inference pipeline (preprocessing plus prediction) using a light-weight embedding that discretizes a carefully selected subset of the input features and an index that maps data points in the embedding space to aggregated predictions of an ML model. We replace a complex preprocessing workflow and model-based inference with a simple feature transformation and an index lookup. Our framework improves inference latency by several orders of magnitude while maintaining similar prediction accuracy compared to the pipeline it approximates.<\/jats:p>","DOI":"10.14778\/3659437.3659441","type":"journal-article","created":{"date-parts":[[2024,5,31]],"date-time":"2024-05-31T16:22:27Z","timestamp":1717172547000},"page":"1830-1842","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["InferDB: In-Database Machine Learning Inference Using Indexes"],"prefix":"10.14778","volume":"17","author":[{"given":"Ricardo","family":"Salazar-D\u00edaz","sequence":"first","affiliation":[{"name":"Hasso Plattner Institute, University of Potsdam"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Boris","family":"Glavic","sequence":"additional","affiliation":[{"name":"University of Illinois Chicago"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tilmann","family":"Rabl","sequence":"additional","affiliation":[{"name":"Hasso Plattner Institute, University of Potsdam"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,5,31]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Rakesh Agrawal and Kyuseok Shim. 1996. Developing Tightly-Coupled Data Mining Applications on a Relational Database System. In SIGKDD. 287--290."},{"key":"e_1_2_1_2_1","volume-title":"Retrieved","year":"2020","unstructured":"Amazon. 2020. Create, train, and deploy machine learning models in Amazon Redshift using SQL with Amazon Redshift ML. Retrieved April 11, 2024 from https:\/\/aws.amazon.com\/de\/blogs\/big-data\/create-train-and-deploy-machine-learning-models-in-amazon-redshift-using-sql-with-amazon-redshift-ml\/"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/361002.361007"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.14778\/3007263.3007279"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.15832\/ankutbd.862482"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939785"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.1967.1053964"},{"key":"e_1_2_1_8_1","doi-asserted-by":"crossref","unstructured":"Anirban Dasgupta Ravi Kumar and Tam\u00e1s Sarl\u00f3s. 2011. Fast locality-sensitive hashing. In SIGKDD. 1073--1081.","DOI":"10.1145\/2020408.2020578"},{"key":"e_1_2_1_9_1","doi-asserted-by":"crossref","unstructured":"James Dougherty Ron Kohavi and Mehran Sahami. 1995. Supervised and Unsupervised Discretization of Continuous Features. In ICML. 194--202.","DOI":"10.1016\/B978-1-55860-377-6.50032-3"},{"key":"e_1_2_1_10_1","volume-title":"Irani","author":"Fayyad Usama M.","year":"1993","unstructured":"Usama M. Fayyad and Keki B. Irani. 1993. Multi-Interval Discretization of Continuous-Valued Attributes for Classification Learning. In IJCAI. 1022--1029."},{"key":"e_1_2_1_11_1","doi-asserted-by":"crossref","unstructured":"Xixuan Feng Arun Kumar Benjamin Recht and Christopher R\u00e9. 2012. Towards a Unified Architecture for In-RDBMS Analytics. In SIGMOD. 325--336.","DOI":"10.1145\/2213836.2213874"},{"key":"e_1_2_1_12_1","volume-title":"Retrieved","author":"Centers for Disease Control and Prevention.","year":"2023","unstructured":"Centers for Disease Control and Prevention. 2023. Daily Census Tract-level PM2.5 concentrations, 2016-2020. Retrieved April 11, 2024 from https:\/\/healthdata.gov\/dataset\/Daily-Census-Tract-Level-PM2-5-Concentrations-2016\/k9st-jhz8\/data"},{"key":"e_1_2_1_13_1","article-title":"SP-GIST: An extensible database index for supporting Space Partitioning Trees","author":"Ilyas Walid","year":"2001","unstructured":"Walid G. and Ihab F. Ilyas. 2001. SP-GIST: An extensible database index for supporting Space Partitioning Trees. Journal of Intelligent Information Systems, 215--240.","journal-title":"Journal of Intelligent Information Systems, 215--240."},{"key":"e_1_2_1_14_1","volume-title":"Jesus Camacho-Rodriguez, and Matteo Interlandi.","author":"Gandhi Apurva","year":"2023","unstructured":"Apurva Gandhi, Yuki Asada, Victor Fu, Advitya Gemawat, Lihao Zhang, Rathijit Sen, Carlo Curino, Jesus Camacho-Rodriguez, and Matteo Interlandi. 2023. The Tensor Data Platform: Towards an AI-centric Database System. CIDR (2023)."},{"key":"e_1_2_1_15_1","volume-title":"Retrieved","year":"2023","unstructured":"Google. 2023. Make predictions with imported TensorFlow models. Retrieved April 11, 2024 from https:\/\/cloud.google.com\/bigquery\/docs\/making-predictions-with-imported-tensorflow-models?hl=de"},{"key":"e_1_2_1_16_1","volume-title":"On The Move to Meaningful Internet Systems 2003: CoopIS, DOA, and ODBASE","author":"Guo Gongde","unstructured":"Gongde Guo, Hui Wang, David Bell, Yaxin Bi, and Kieran Greer. 2003. KNN model-based approach in classification. In On The Move to Meaningful Internet Systems 2003: CoopIS, DOA, and ODBASE. Springer Berlin Heidelberg, 986--996. https:\/\/link.springer.com\/chapter\/10.1007\/978-3-540-39964-3_62#citeas"},{"key":"e_1_2_1_17_1","volume-title":"An Introduction to Variable and Feature Selection. J. Mach. Learn. Res. 3, null","author":"Guyon Isabelle","year":"2003","unstructured":"Isabelle Guyon and Andr\u00e9 Elisseeff. 2003. An Introduction to Variable and Feature Selection. J. Mach. Learn. Res. 3, null (2003), 1157--1182."},{"key":"e_1_2_1_18_1","first-page":"62","article-title":"Experimental Analysis of Locality Sensitive Hashing Techniques for High-Dimensional Approximate Nearest Neighbor Searches","volume":"12610","author":"Jafari Omid","year":"2021","unstructured":"Omid Jafari and Parth Nagarkar. 2021. Experimental Analysis of Locality Sensitive Hashing Techniques for High-Dimensional Approximate Nearest Neighbor Searches. In ADC, Vol. 12610. 62--73.","journal-title":"ADC"},{"key":"e_1_2_1_19_1","volume-title":"Video. Retrieved","author":"Jassy Andy","year":"2018","unstructured":"Andy Jassy. 2018. AWS re:Invent 2018 keynote. Video. Retrieved April 11, 2024 from https:\/\/www.youtube.com\/watch?v=ZOIkOnW640A"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1098\/rspa.1946.0056"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCCNT56998.2023.10308091"},{"key":"e_1_2_1_22_1","volume-title":"Retrieved","year":"2017","unstructured":"Kaggle. 2017. New York City taxi trip duration. Retrieved April 11, 2024 from https:\/\/www.kaggle.com\/c\/nyc-taxi-trip-duration"},{"key":"e_1_2_1_23_1","volume-title":"Retrieved","year":"2017","unstructured":"Kaggle. 2017. New York City taxi trip duration evaluation. Retrieved April 11, 2024 from https:\/\/www.kaggle.com\/competitions\/nyc-taxi-trip-duration\/overview\/evaluation"},{"key":"e_1_2_1_24_1","volume-title":"Eui Chul Richard Shin, and Dawn Song","author":"Katz Gilad","year":"2016","unstructured":"Gilad Katz, Eui Chul Richard Shin, and Dawn Song. 2016. ExploreKit: Automatic Feature Generation and Selection. In ICDM. 979--984."},{"key":"e_1_2_1_25_1","volume-title":"NIPS (Long Beach, California, USA) (NIPS'17)","author":"Ke Guolin","unstructured":"Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. 2017. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In NIPS (Long Beach, California, USA) (NIPS'17). Curran Associates Inc., Red Hook, NY, USA, 3149--3157."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.18420\/btw2021-10"},{"key":"e_1_2_1_27_1","unstructured":"Steffen Kl\u00e4be Stefan Hagedorn and Kai-Uwe Sattler. 2022. Exploration of Approaches for In-Database ML. In EDBT. OpenProceedings.org 311--323. https:\/\/openproceedings.org\/2023\/conf\/edbt\/paper-7.pdf"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.14778\/3342263.3342633"},{"key":"e_1_2_1_29_1","volume-title":"MNIST handwritten digit database. ATT Labs [Online]. Available: http:\/\/yann.lecun.com\/exdb\/mnist 2","author":"LeCun Yann","year":"2010","unstructured":"Yann LeCun, Corinna Cortes, and CJ Burges. 2010. MNIST handwritten digit database. ATT Labs [Online]. Available: http:\/\/yann.lecun.com\/exdb\/mnist 2 (2010)."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-019-09709-4"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1016304305535"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.18420\/BTW2023-26"},{"key":"e_1_2_1_33_1","volume-title":"Retrieved","year":"2023","unstructured":"Microsoft. 2023. Predict (transact-SQL) - SQL machine learning. Retrieved April 11, 2024 from https:\/\/learn.microsoft.com\/en-us\/sql\/t-sql\/queries\/predict-transact-sql?view=sql-server-ver15"},{"key":"e_1_2_1_34_1","volume-title":"Optimal binning: mathematical programming formulation. abs\/2001.08025","author":"Navas-Palencia Guillermo","year":"2020","unstructured":"Guillermo Navas-Palencia. 2020. Optimal binning: mathematical programming formulation. abs\/2001.08025 (2020). arXiv:2001.08025 [cs.LG]"},{"key":"e_1_2_1_35_1","volume-title":"Five balltree construction algorithms","author":"Omohundro Stephen M","year":"2009","unstructured":"Stephen M Omohundro. 1989. Five balltree construction algorithms. International Computer Science Institute Berkeley. https:\/\/omohundro.files.wordpress.com\/2009\/03\/omohundro89_five_balltree_construction_algorithms.pdf"},{"key":"e_1_2_1_36_1","volume-title":"Retrieved","author":"Oneapi-Src Intel","year":"2020","unstructured":"Intel Oneapi-Src. 2020. Oneapi-src\/onedal: Oneapi data analytics library (onedal). Retrieved April 11, 2024 from https:\/\/github.com\/oneapi-src\/oneDAL?tab=readme-ov-file"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2006.31"},{"key":"e_1_2_1_38_1","doi-asserted-by":"crossref","unstructured":"Jia Pan and Dinesh Manocha. 2011. Fast GPU-based locality sensitive hashing for k-nearest neighbor computation. In SIGSPATIAL. 211--220.","DOI":"10.1145\/2093973.2094002"},{"key":"e_1_2_1_39_1","doi-asserted-by":"crossref","unstructured":"Kwanghyun Park Karla Saur Dalitso Banda Rathijit Sen Matteo Interlandi and Konstantinos Karanasos. 2022. End-to-End Optimization of Machine Learning Prediction Queries. In SIGMOD. 587--601.","DOI":"10.1145\/3514221.3526141"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.14778\/2850583.2850589"},{"key":"e_1_2_1_41_1","unstructured":"Postgresml. 2022. Postgresml\/postgresml: PostgresML is an AI application database. Download open source models from Huggingface or train your own to create and index LLM embeddings generate text or make online predictions using only SQL. Retrieved April 11 2024 from https:\/\/github.com\/postgresml\/postgresml"},{"key":"e_1_2_1_42_1","doi-asserted-by":"crossref","unstructured":"Andrea Dal Pozzolo Olivier Caelen Reid A. Johnson and Gianluca Bontempi. 2015. Calibrating Probability with Undersampling for Unbalanced Classification. In SSCI. 159--166.","DOI":"10.1109\/SSCI.2015.33"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3552490.3552496"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","unstructured":"Mark Raasveldt Pedro Holanda Hannes M\u00fchleisen and Stefan Manegold. 2018. Deep Integration of Machine Learning Into Column Stores. In EDBT. OpenProceedings.org 473--476. 10.5441\/002\/EDBT.2018.50","DOI":"10.5441\/002\/EDBT.2018.50"},{"key":"e_1_2_1_45_1","unstructured":"Maximilian Rieger Moritz Sichert and Thomas Neumann. 2022. Integrating deep learning frameworks into main-memory databases. In Proceedings of the VLDB 2022 Applied AI for Database Systems and Applications Workshop co-located with (VLDB 2022) (AIDB Workshop Proceedings). https:\/\/drive.google.com\/file\/d\/1GfZH3Y1sQKgplnnpTEM_E4skWdhmyrfe\/edit"},{"key":"e_1_2_1_46_1","doi-asserted-by":"crossref","unstructured":"Sunita Sarawagi Shiby Thomas and Rakesh Agrawal. 1998. Integrating Association Rule Mining with Relational Database Systems: Alternatives and Implications. In SIGMOD. 343--354.","DOI":"10.1145\/276305.276335"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1007\/s42979-021-00592-x"},{"key":"e_1_2_1_48_1","doi-asserted-by":"crossref","unstructured":"Kai-Uwe Sattler and Oliver Dunemann. 2001. SQL Database Primitives for Decision Tree Classifiers. In CIKM. 379--386.","DOI":"10.1145\/502585.502650"},{"key":"e_1_2_1_49_1","volume-title":"BTW","author":"Sch\u00fcle Maximilian Emanuel","unstructured":"Maximilian Emanuel Sch\u00fcle, Alfons Kemper, and Thomas Neumann. 2023. NN2SQL: Let SQL Think for Neural Networks. In BTW, Vol. P-331. 183--194."},{"key":"e_1_2_1_50_1","unstructured":"Maximilian E. Sch\u00fcle Luca Scalerandi Alfons Kemper and Thomas Neumann. 2023. Blue Elephants Inspecting Pandas: Inspection and Execution of Machine Learning Pipelines in SQL. In EDBT. 40--52."},{"key":"e_1_2_1_51_1","volume-title":"Retrieved","year":"2024","unstructured":"Scikit-learn. [n.d.]. 8.2. computational performance. Retrieved April 11, 2024 from https:\/\/scikit-learn.org\/stable\/computing\/computational_performance.html"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/2990508"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.14778\/3467861.3467867"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3659437.3659441","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,20]],"date-time":"2024-11-20T21:35:12Z","timestamp":1732138512000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3659437.3659441"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4]]},"references-count":53,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2024,4]]}},"alternative-id":["10.14778\/3659437.3659441"],"URL":"https:\/\/doi.org\/10.14778\/3659437.3659441","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2024,4]]},"assertion":[{"value":"2024-05-31","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}