{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,20]],"date-time":"2026-06-20T01:11:19Z","timestamp":1781917879056,"version":"3.54.5"},"reference-count":70,"publisher":"Association for Computing Machinery (ACM)","issue":"11","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2023,7]]},"abstract":"<jats:p>Although dominant for tabular data, ML libraries that train tree models over normalized databases (e.g., LightGBM, XGBoost) require the data to be denormalized as a single table, materialized, and exported. This process is not scalable, slow, and poses security risks. In-DB ML aims to train models within DBMSes to avoid data movement and provide data governance. Rather than modify a DBMS to support In-DB ML, is it possible to offer competitive tree training performance to specialized ML libraries...with only SQL?<\/jats:p>\n          <jats:p>\n            We present JoinBoost, a Python library that rewrites tree training algorithms over normalized databases into pure SQL. It is portable to any DBMS, offers performance competitive with specialized ML libraries, and scales with the underlying DBMS capabilities. JoinBoost extends prior work from both algorithmic and systems perspectives. Algorithmically, we support factorized gradient boosting, by updating the\n            <jats:italic>Y<\/jats:italic>\n            variable to the residual in the\n            <jats:italic>non-materialized join result.<\/jats:italic>\n            Although this view update problem is generally ambiguous, we identify\n            <jats:italic>addition-to-multiplication preserving<\/jats:italic>\n            , the key property of variance semi-ring to support\n            <jats:italic>rmse<\/jats:italic>\n            the most widely used criterion. System-wise, we identify residual updates as a performance bottleneck. Such overhead can be natively minimized on columnar DBMSes by creating a new column of residual values and adding it as a projection. We validate this with two implementations on DuckDB, with no or minimal modifications to its internals for portability. Our experiment shows that JoinBoost is 3\u00d7 (1.1\u00d7) faster for random forests (gradient boosting) compared to LightGBM, and over an order of magnitude faster than state-of-the-art In-DB ML systems. Further, JoinBoost scales well beyond LightGBM in terms of the # features, DB size (TPC-DS SF=1000), and join graph complexity (galaxy schemas).\n          <\/jats:p>","DOI":"10.14778\/3611479.3611509","type":"journal-article","created":{"date-parts":[[2023,8,25]],"date-time":"2023-08-25T02:08:08Z","timestamp":1692929288000},"page":"3071-3084","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["JoinBoost: Grow Trees over Normalized Data Using Only SQL"],"prefix":"10.14778","volume":"16","author":[{"given":"Zezhou","family":"Huang","sequence":"first","affiliation":[{"name":"Columbia University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rathijit","family":"Sen","sequence":"additional","affiliation":[{"name":"Microsoft"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiaxiang","family":"Liu","sequence":"additional","affiliation":[{"name":"Columbia University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Eugene","family":"Wu","sequence":"additional","affiliation":[{"name":"DSI, Columbia University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,8,24]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"22nd International Conference on Data Engineering (ICDE'06)","unstructured":"2006. Updates through views: A new hope . In 22nd International Conference on Data Engineering (ICDE'06) . IEEE, 2--2. 2006. Updates through views: A new hope. In 22nd International Conference on Data Engineering (ICDE'06). IEEE, 2--2."},{"key":"e_1_2_1_2_1","unstructured":"2013. IMDB. https:\/\/www.imdb.com\/interfaces\/.  2013. IMDB. https:\/\/www.imdb.com\/interfaces\/."},{"key":"e_1_2_1_3_1","unstructured":"2017. Corporaci\u00f3n Favorita Grocery Sales Forecasting. https:\/\/www.kaggle.com\/c\/favorita-grocery-sales-forecasting.  2017. Corporaci\u00f3n Favorita Grocery Sales Forecasting. https:\/\/www.kaggle.com\/c\/favorita-grocery-sales-forecasting."},{"key":"e_1_2_1_4_1","unstructured":"2017. Lightgbm memory explodes in start train. https:\/\/github.com\/microsoft\/LightGBM\/issues\/1032.  2017. Lightgbm memory explodes in start train. https:\/\/github.com\/microsoft\/LightGBM\/issues\/1032."},{"key":"e_1_2_1_5_1","unstructured":"2020. Looker data modeling. https:\/\/www.looker.com\/platform\/data-modeling\/.  2020. Looker data modeling. https:\/\/www.looker.com\/platform\/data-modeling\/."},{"key":"e_1_2_1_6_1","unstructured":"2020. The Tableau Data Model. https:\/\/help.tableau.com\/current\/online\/en-us\/datasource_datamodel.htm.  2020. The Tableau Data Model. https:\/\/help.tableau.com\/current\/online\/en-us\/datasource_datamodel.htm."},{"key":"e_1_2_1_7_1","unstructured":"2021. Client APIs Overview. https:\/\/duckdb.org\/docs\/api\/overview.  2021. Client APIs Overview. https:\/\/duckdb.org\/docs\/api\/overview."},{"key":"e_1_2_1_8_1","unstructured":"2021. Kaggle Data Science and Machine Learning Survey. https:\/\/www.kaggle.com\/code\/paultimothymooney\/2021-kaggle-data-science-machine-learning-survey\/notebook.  2021. Kaggle Data Science and Machine Learning Survey. https:\/\/www.kaggle.com\/code\/paultimothymooney\/2021-kaggle-data-science-machine-learning-survey\/notebook."},{"key":"e_1_2_1_9_1","unstructured":"2022. ClickBench: a Benchmark For Analytical Databases. https:\/\/benchmark.clickhouse.com\/.  2022. ClickBench: a Benchmark For Analytical Databases. https:\/\/benchmark.clickhouse.com\/."},{"key":"e_1_2_1_10_1","unstructured":"2022. Personal-Data-Protection-Act. https:\/\/www.pdpc.gov.sg\/Overview-of-PDPA\/The-Legislation\/Personal-Data-Protection-Act.  2022. Personal-Data-Protection-Act. https:\/\/www.pdpc.gov.sg\/Overview-of-PDPA\/The-Legislation\/Personal-Data-Protection-Act."},{"key":"e_1_2_1_11_1","unstructured":"2023. Azure Machine Learning documentation. https:\/\/learn.microsoft.com\/en-us\/azure\/machine-learning\/.  2023. Azure Machine Learning documentation. https:\/\/learn.microsoft.com\/en-us\/azure\/machine-learning\/."},{"key":"e_1_2_1_12_1","unstructured":"2023. The dbt Semantic Layer. https:\/\/www.getdbt.com\/product\/semantic-layer\/.  2023. The dbt Semantic Layer. https:\/\/www.getdbt.com\/product\/semantic-layer\/."},{"key":"e_1_2_1_13_1","unstructured":"2023. Snowflake Machine Learning Platforms. https:\/\/www.snowflake.com\/guides\/machine-learning-platforms.  2023. Snowflake Machine Learning Platforms. https:\/\/www.snowflake.com\/guides\/machine-learning-platforms."},{"key":"e_1_2_1_14_1","unstructured":"2023. Using machine learning in Amazon Redshift. https:\/\/docs.amazonaws.cn\/en_us\/redshift\/latest\/dg\/machine_learning.html.  2023. Using machine learning in Amazon Redshift. https:\/\/docs.amazonaws.cn\/en_us\/redshift\/latest\/dg\/machine_learning.html."},{"key":"e_1_2_1_15_1","unstructured":"Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dandelion Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. https:\/\/www.tensorflow.org\/ Software available from tensorflow.org.  Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dandelion Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. https:\/\/www.tensorflow.org\/ Software available from tensorflow.org."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3129246"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2902251.2902280"},{"key":"e_1_2_1_18_1","volume-title":"Alessandro De Palma, and Haralampos Pozidis","author":"Anghel Andreea","year":"2018","unstructured":"Andreea Anghel , Nikolaos Papandreou , Thomas Parnell , Alessandro De Palma, and Haralampos Pozidis . 2018 . Benchmarking and optimization of gradient boosting decision tree algorithms. arXiv preprint arXiv:1809.04559 (2018). Andreea Anghel, Nikolaos Papandreou, Thomas Parnell, Alessandro De Palma, and Haralampos Pozidis. 2018. Benchmarking and optimization of gradient boosting decision tree algorithms. arXiv preprint arXiv:1809.04559 (2018)."},{"key":"e_1_2_1_19_1","doi-asserted-by":"crossref","unstructured":"Yuki Asada Victor Fu Apurva Gandhi Advitya Gemawat Lihao Zhang Dong He Vivek Gupta Ehi Nosakhare Dalitso Banda Rathijit Sen etal 2022. Share the tensor tea: how databases can leverage the machine learning ecosystem. arXiv preprint arXiv:2209.04579 (2022).  Yuki Asada Victor Fu Apurva Gandhi Advitya Gemawat Lihao Zhang Dong He Vivek Gupta Ehi Nosakhare Dalitso Banda Rathijit Sen et al. 2022. Share the tensor tea: how databases can leverage the machine learning ecosystem. arXiv preprint arXiv:2209.04579 (2022).","DOI":"10.14778\/3554821.3554853"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.14778\/3007263.3007279"},{"key":"e_1_2_1_21_1","volume-title":"Random forests. Machine learning 45, 1","author":"Breiman Leo","year":"2001","unstructured":"Leo Breiman . 2001. Random forests. Machine learning 45, 1 ( 2001 ), 5--32. Leo Breiman. 2001. Random forests. Machine learning 45, 1 (2001), 5--32."},{"key":"e_1_2_1_22_1","volume-title":"Classification and regression trees","author":"Breiman Leo","unstructured":"Leo Breiman , Jerome H Friedman , Richard A Olshen , and Charles J Stone . 2017. Classification and regression trees . Routledge . Leo Breiman, Jerome H Friedman, Richard A Olshen, and Charles J Stone. 2017. Classification and regression trees. Routledge."},{"key":"e_1_2_1_23_1","unstructured":"Bishop PRML Ch and Alireza Ghane. 1993. Sampling Methods. (1993).  Bishop PRML Ch and Alireza Ghane. 1993. Sampling Methods. (1993)."},{"key":"e_1_2_1_24_1","volume-title":"EDBT\/ICDT Workshops.","author":"Chatziantoniou Damianos","year":"2020","unstructured":"Damianos Chatziantoniou and Verena Kantere . 2020 . Data Virtual Machines: Data-Driven Conceptual Modeling of Big Data Infrastructures .. In EDBT\/ICDT Workshops. Damianos Chatziantoniou and Verena Kantere. 2020. Data Virtual Machines: Data-Driven Conceptual Modeling of Big Data Infrastructures.. In EDBT\/ICDT Workshops."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939785"},{"key":"e_1_2_1_26_1","unstructured":"Francois Chollet et al. 2015. Keras. https:\/\/github.com\/fchollet\/keras  Francois Chollet et al. 2015. Keras. https:\/\/github.com\/fchollet\/keras"},{"key":"e_1_2_1_27_1","volume-title":"International Conference on Artificial Intelligence and Statistics. PMLR, 2742--2752","author":"Curtin Ryan","year":"2020","unstructured":"Ryan Curtin , Benjamin Moseley , Hung Ngo , XuanLong Nguyen , Dan Olteanu , and Maximilian Schleich . 2020 . Rk-means: Fast clustering for relational data . In International Conference on Artificial Intelligence and Statistics. PMLR, 2742--2752 . Ryan Curtin, Benjamin Moseley, Hung Ngo, XuanLong Nguyen, Dan Olteanu, and Maximilian Schleich. 2020. Rk-means: Fast clustering for relational data. In International Conference on Artificial Intelligence and Statistics. PMLR, 2742--2752."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/2213836.2213874"},{"key":"e_1_2_1_29_1","volume-title":"Stochastic gradient boosting. Computational statistics & data analysis 38, 4","author":"Friedman Jerome H","year":"2002","unstructured":"Jerome H Friedman . 2002. Stochastic gradient boosting. Computational statistics & data analysis 38, 4 ( 2002 ), 367--378. Jerome H Friedman. 2002. Stochastic gradient boosting. Computational statistics & data analysis 38, 4 (2002), 367--378."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/1265530.1265535"},{"key":"e_1_2_1_31_1","volume-title":"Why do tree-based models still outperform deep learning on tabular data? arXiv preprint arXiv:2207.08815","author":"Grinsztajn L\u00e9o","year":"2022","unstructured":"L\u00e9o Grinsztajn , Edouard Oyallon , and Ga\u00ebl Varoquaux . 2022. Why do tree-based models still outperform deep learning on tabular data? arXiv preprint arXiv:2207.08815 ( 2022 ). L\u00e9o Grinsztajn, Edouard Oyallon, and Ga\u00ebl Varoquaux. 2022. Why do tree-based models still outperform deep learning on tabular data? arXiv preprint arXiv:2207.08815 (2022)."},{"key":"e_1_2_1_32_1","volume-title":"Eugene Fratkin, Aleksander Gorajek, Kee Siong Ng, Caleb Welton, Xixuan Feng, Kun Li, et al.","author":"Hellerstein Joe","year":"2012","unstructured":"Joe Hellerstein , Christopher R\u00e9 , Florian Schoppmann , Daisy Zhe Wang , Eugene Fratkin, Aleksander Gorajek, Kee Siong Ng, Caleb Welton, Xixuan Feng, Kun Li, et al. 2012 . The MADlib analytics library or MAD skills, the SQL. arXiv preprint arXiv:1208.4165 (2012). Joe Hellerstein, Christopher R\u00e9, Florian Schoppmann, Daisy Zhe Wang, Eugene Fratkin, Aleksander Gorajek, Kee Siong Ng, Caleb Welton, Xixuan Feng, Kun Li, et al. 2012. The MADlib analytics library or MAD skills, the SQL. arXiv preprint arXiv:1208.4165 (2012)."},{"key":"e_1_2_1_33_1","volume-title":"TCUDB: Accelerating Database with Tensor Processors. arXiv preprint arXiv:2112.07552","author":"Hu Yu-Ching","year":"2021","unstructured":"Yu-Ching Hu , Yuliang Li , and Hung-Wei Tseng . 2021 . TCUDB: Accelerating Database with Tensor Processors. arXiv preprint arXiv:2112.07552 (2021). Yu-Ching Hu, Yuliang Li, and Hung-Wei Tseng. 2021. TCUDB: Accelerating Database with Tensor Processors. arXiv preprint arXiv:2112.07552 (2021)."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3597465.3605224"},{"key":"e_1_2_1_35_1","volume-title":"Joinboost: Grow trees over normalized data using only SQL. arXiv preprint arXiv:2307.00422","author":"Huang Zezhou","year":"2023","unstructured":"Zezhou Huang , Rathijit Sen , Jiaxiang Liu , and Eugene Wu . 2023 . Joinboost: Grow trees over normalized data using only SQL. arXiv preprint arXiv:2307.00422 (2023). Zezhou Huang, Rathijit Sen, Jiaxiang Liu, and Eugene Wu. 2023. Joinboost: Grow trees over normalized data using only SQL. arXiv preprint arXiv:2307.00422 (2023)."},{"key":"e_1_2_1_36_1","volume-title":"Calibration: A Simple Trick for Wide-table Delta Analytics. arXiv:2210.03851 [cs.DB]","author":"Huang Zezhou","year":"2022","unstructured":"Zezhou Huang and Eugene Wu . 2022 . Calibration: A Simple Trick for Wide-table Delta Analytics. arXiv:2210.03851 [cs.DB] Zezhou Huang and Eugene Wu. 2022. Calibration: A Simple Trick for Wide-table Delta Analytics. arXiv:2210.03851 [cs.DB]"},{"key":"e_1_2_1_37_1","volume-title":"Pass-glm: polynomial approximate sufficient statistics for scalable bayesian glm inference. Advances in Neural Information Processing Systems 30","author":"Huggins Jonathan","year":"2017","unstructured":"Jonathan Huggins , Ryan P Adams , and Tamara Broderick . 2017. Pass-glm: polynomial approximate sufficient statistics for scalable bayesian glm inference. Advances in Neural Information Processing Systems 30 ( 2017 ). Jonathan Huggins, Ryan P Adams, and Tamara Broderick. 2017. Pass-glm: polynomial approximate sufficient statistics for scalable bayesian glm inference. Advances in Neural Information Processing Systems 30 (2017)."},{"key":"e_1_2_1_38_1","volume-title":"Declarative recursive computation on an rdbms, or, why you should use a database for distributed machine learning. arXiv preprint arXiv:1904.11121","author":"Jankov Dimitrije","year":"2019","unstructured":"Dimitrije Jankov , Shangyu Luo , Binhang Yuan , Zhuhua Cai , Jia Zou , Chris Jermaine , and Zekai J Gao . 2019. Declarative recursive computation on an rdbms, or, why you should use a database for distributed machine learning. arXiv preprint arXiv:1904.11121 ( 2019 ). Dimitrije Jankov, Shangyu Luo, Binhang Yuan, Zhuhua Cai, Jia Zou, Chris Jermaine, and Zekai J Gao. 2019. Declarative recursive computation on an rdbms, or, why you should use a database for distributed machine learning. arXiv preprint arXiv:1904.11121 (2019)."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.14778\/3450980.3450991"},{"key":"e_1_2_1_40_1","volume-title":"Aggregations over generalized hypertree decompositions. arXiv preprint arXiv:1508.07532","author":"Joglekar Manas","year":"2015","unstructured":"Manas Joglekar , Rohan Puttagunta , and Christopher R\u00e9. 2015. Aggregations over generalized hypertree decompositions. arXiv preprint arXiv:1508.07532 ( 2015 ). Manas Joglekar, Rohan Puttagunta, and Christopher R\u00e9. 2015. Aggregations over generalized hypertree decompositions. arXiv preprint arXiv:1508.07532 (2015)."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2902251.2902293"},{"key":"e_1_2_1_42_1","volume-title":"Lightgbm: A highly efficient gradient boosting decision tree. Advances in neural information processing systems 30","author":"Ke Guolin","year":"2017","unstructured":"Guolin Ke , Qi Meng , Thomas Finley , Taifeng Wang , Wei Chen , Weidong Ma , Qiwei Ye , and Tie-Yan Liu . 2017 . Lightgbm: A highly efficient gradient boosting decision tree. Advances in neural information processing systems 30 (2017). Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. 2017. Lightgbm: A highly efficient gradient boosting decision tree. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_2_1_43_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3426865","article-title":"Functional Aggregate Queries with Additive Inequalities","volume":"45","author":"Khamis Mahmoud Abo","year":"2020","unstructured":"Mahmoud Abo Khamis , Ryan R Curtin , Benjamin Moseley , Hung Q Ngo , XuanLong Nguyen , Dan Olteanu , and Maximilian Schleich . 2020 . Functional Aggregate Queries with Additive Inequalities . ACM Transactions on Database Systems (TODS) 45 , 4 (2020), 1 -- 41 . Mahmoud Abo Khamis, Ryan R Curtin, Benjamin Moseley, Hung Q Ngo, XuanLong Nguyen, Dan Olteanu, and Maximilian Schleich. 2020. Functional Aggregate Queries with Additive Inequalities. ACM Transactions on Database Systems (TODS) 45, 4 (2020), 1--41.","journal-title":"ACM Transactions on Database Systems (TODS)"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3209889.3209896"},{"key":"e_1_2_1_45_1","first-page":"2","article-title":"MLbase: A Distributed Machine-learning System","volume":"1","author":"Kraska Tim","year":"2013","unstructured":"Tim Kraska , Ameet Talwalkar , John C Duchi , Rean Griffith , Michael J Franklin , and Michael I Jordan . 2013 . MLbase: A Distributed Machine-learning System .. In Cidr , Vol. 1. 2 -- 1 . Tim Kraska, Ameet Talwalkar, John C Duchi, Rean Griffith, Michael J Franklin, and Michael I Jordan. 2013. MLbase: A Distributed Machine-learning System.. In Cidr, Vol. 1. 2--1.","journal-title":"Cidr"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.14778\/3137765.3137812"},{"key":"e_1_2_1_47_1","volume-title":"BigQuery for Data Warehousing","author":"Mucchetti Mark","unstructured":"Mark Mucchetti . 2020. BigQuery ML . In BigQuery for Data Warehousing . Springer , 419--468. Mark Mucchetti. 2020. BigQuery ML. In BigQuery for Data Warehousing. Springer, 419--468."},{"key":"e_1_2_1_48_1","volume-title":"Machine learning: a probabilistic perspective","author":"Murphy Kevin P","unstructured":"Kevin P Murphy . 2012. Machine learning: a probabilistic perspective . MIT press . Kevin P Murphy. 2012. Machine learning: a probabilistic perspective. MIT press."},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2749436"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3183758"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3003665.3003667"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/2656335"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.5281\/zenodo.3509134"},{"key":"e_1_2_1_54_1","doi-asserted-by":"crossref","unstructured":"Johns Paul Shengliang Lu Bingsheng He etal 2021. Database Systems on GPUs. Foundations and Trends\u00ae in Databases 11 1 (2021) 1--108.  Johns Paul Shengliang Lu Bingsheng He et al. 2021. Database Systems on GPUs. Foundations and Trends \u00ae in Databases 11 1 (2021) 1--108.","DOI":"10.1561\/1900000076"},{"key":"e_1_2_1_55_1","unstructured":"Judea Pearl. 1982. Reverend Bayes on inference engines: A distributed hierarchical approach. Cognitive Systems Laboratory School of Engineering and Applied Science . . . .  Judea Pearl. 1982. Reverend Bayes on inference engines: A distributed hierarchical approach. Cognitive Systems Laboratory School of Engineering and Applied Science . . . ."},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2078195"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3320212"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.14778\/2536222.2536233"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.25080\/Majora-7b98e3ed-013"},{"key":"e_1_2_1_60_1","first-page":"28","article-title":"Data Warehouse Designing: Dimensional Modelling and ER Modelling","volume":"3","author":"Saxena Geetika","year":"2014","unstructured":"Geetika Saxena and Bharat Bhushan Agarwal . 2014 . Data Warehouse Designing: Dimensional Modelling and ER Modelling . International Journal of Engineering Inventions 3 , 9 (2014), 28 -- 34 . Geetika Saxena and Bharat Bhushan Agarwal. 2014. Data Warehouse Designing: Dimensional Modelling and ER Modelling. International Journal of Engineering Inventions 3, 9 (2014), 28--34.","journal-title":"International Journal of Engineering Inventions"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3324961"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2882939"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3457302"},{"key":"e_1_2_1_64_1","volume-title":"Operationalizing Machine Learning: An Interview Study. arXiv preprint arXiv:2209.09125","author":"Shankar Shreya","year":"2022","unstructured":"Shreya Shankar , Rolando Garcia , Joseph M Hellerstein , and Aditya G Parameswaran . 2022. Operationalizing Machine Learning: An Interview Study. arXiv preprint arXiv:2209.09125 ( 2022 ). Shreya Shankar, Rolando Garcia, Joseph M Hellerstein, and Aditya G Parameswaran. 2022. Operationalizing Machine Learning: An Interview Study. arXiv preprint arXiv:2209.09125 (2022)."},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1006\/ijhc.1995.1017"},{"key":"e_1_2_1_67_1","doi-asserted-by":"crossref","unstructured":"Mike Stonebraker Daniel J Abadi Adam Batkin Xuedong Chen Mitch Cherniack Miguel Ferreira Edmond Lau Amerson Lin Sam Madden Elizabeth O'Neil etal 2018. C-store: a column-oriented DBMS. In Making Databases Work: the Pragmatic Wisdom of Michael Stonebraker. 491--518.  Mike Stonebraker Daniel J Abadi Adam Batkin Xuedong Chen Mitch Cherniack Miguel Ferreira Edmond Lau Amerson Lin Sam Madden Elizabeth O'Neil et al. 2018. C-store: a column-oriented DBMS. In Making Databases Work: the Pragmatic Wisdom of Michael Stonebraker. 491--518.","DOI":"10.1145\/3226595.3226638"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-22351-8_1"},{"key":"e_1_2_1_69_1","doi-asserted-by":"crossref","unstructured":"Swarup Acharya Phillip B Gibbons Viswanath and Poosala Sridhar Ramaswamy. 1998. Join Synopses for Approximate Query Answering. (1998).  Swarup Acharya Phillip B Gibbons Viswanath and Poosala Sridhar Ramaswamy. 1998. Join Synopses for Approximate Query Answering. (1998).","DOI":"10.1145\/304182.304207"},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1145\/2463676.2465288"},{"key":"e_1_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3183739"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3611479.3611509","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,23]],"date-time":"2023-09-23T22:24:28Z","timestamp":1695507868000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3611479.3611509"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7]]},"references-count":70,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2023,7]]}},"alternative-id":["10.14778\/3611479.3611509"],"URL":"https:\/\/doi.org\/10.14778\/3611479.3611509","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2023,7]]},"assertion":[{"value":"2023-08-24","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}