{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,8]],"date-time":"2026-02-08T07:54:11Z","timestamp":1770537251421,"version":"3.49.0"},"reference-count":78,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2023,12,8]],"date-time":"2023-12-08T00:00:00Z","timestamp":1701993600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100006374","name":"NSF","doi-asserted-by":"publisher","award":["1845638,2008295,2106197,2103794,2312991"],"award-info":[{"award-number":["1845638,2008295,2106197,2103794,2312991"]}],"id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Amazon"},{"DOI":"10.13039\/501100006374","name":"Columbia Data Science Institute?s Avanessian PhD fellowship","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100006374","name":"Adobe Systems","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2023,12,8]]},"abstract":"<jats:p>Dashboards are vital in modern business intelligence tools, providing non-technical users with an interface to access comprehensive business data. With the rise of cloud technology, there is an increased number of data sources to provide enriched contexts for various analytical tasks, leading to a demand for interactive dashboards over a large number of joins. Nevertheless, joins are among the most expensive operations in DBMSes, making the support of interactive dashboards over joins challenging.<\/jats:p>\n          <jats:p>In this paper, we present Treant, a dashboard accelerator for queries over large joins. Treant uses factorized query execution to handle aggregation queries over large joins, which alone is still insufficient for interactive speeds. To address this, we exploit the incremental nature of user interactions using Calibrated Junction Hypertree (CJT), a novel data structure that applies lightweight materialization of the intermediates during factorized execution. CJT ensures that the work needed to compute a query is proportional to how different it is from the previous query, rather than the overall complexity. Treant manages CJTs to share work between queries and performs materialization offline or during user \"think-times.\" Implemented as a middleware that rewrites SQL, Treant is portable to any SQL-based DBMS. Our experiments on single node and cloud DBMSes show that Treant improves dashboard interactions by two orders of magnitude, and provides 10x improvement for ML augmentation compared to SOTA factorized ML system.<\/jats:p>","DOI":"10.1145\/3626735","type":"journal-article","created":{"date-parts":[[2023,12,12]],"date-time":"2023-12-12T14:01:21Z","timestamp":1702389681000},"page":"1-27","source":"Crossref","is-referenced-by-count":1,"title":["Lightweight Materialization for Fast Dashboards Over Joins"],"prefix":"10.1145","volume":"1","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-6613-0337","authenticated-orcid":false,"given":"Zezhou","family":"Huang","sequence":"first","affiliation":[{"name":"Columbia University, NYC, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4254-6688","authenticated-orcid":false,"given":"Eugene","family":"Wu","sequence":"additional","affiliation":[{"name":"Columbia University, NYC, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,12,12]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"2016. Looker. https:\/\/www.looker.com\/."},{"key":"e_1_2_2_2_1","unstructured":"2017. Corporaci\u00f3n Favorita Grocery Sales Forecasting. https:\/\/www.kaggle.com\/c\/favorita-grocery-sales-forecasting."},{"key":"e_1_2_2_3_1","unstructured":"2018. Modin: Scale your pandas workflows by changing one line of code. https:\/\/github.com\/modin-project\/modin."},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3129246"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2902251.2902280"},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/2902251.2902289"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.14569\/IJACSA.2016.070124"},{"key":"e_1_2_2_8_1","volume-title":"Atta Rahman, Rachid Zagrouba, Fahd Alhaidari, Tariq Ali, and Farzana Zahid.","author":"Ahmad Munir","year":"2020","unstructured":"Munir Ahmad, Muhammad Abdul Qadir, Atta Rahman, Rachid Zagrouba, Fahd Alhaidari, Tariq Ali, and Farzana Zahid. 2020. Enhanced query processing over semantic cache for cloud based relational databases. Journal of Ambient Intelligence and Humanized Computing (2020), 1--19."},{"key":"e_1_2_2_9_1","volume-title":"Massively parallel sort-merge joins in main memory multi-core database systems. arXiv preprint arXiv:1207.0145","author":"Albutiu Martina-Cezara","year":"2012","unstructured":"Martina-Cezara Albutiu, Alfons Kemper, and Thomas Neumann. 2012. Massively parallel sort-merge joins in main memory multi-core database systems. arXiv preprint arXiv:1207.0145 (2012)."},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1504\/IJBIDM.2009.029086"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/FOCS.2008.43"},{"key":"e_1_2_2_12_1","unstructured":"Oscar Bashaw. 2022. Drive Revenue by Using Sigma with Salesforce. https:\/\/www.sigmacomputing.com\/blog\/drive-revenue-by-using-sigma-with-salesforce."},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46073-4_3"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.1995.380392"},{"key":"e_1_2_2_15_1","volume-title":"Tim Kraska, and David Karger.","author":"Chepurko Nadiia","year":"2020","unstructured":"Nadiia Chepurko, Ryan Marcus, Emanuel Zgraggen, Raul Castro Fernandez, Tim Kraska, and David Karger. 2020. ARDA: automatic relational data augmentation for machine learning. arXiv preprint arXiv:2003.09758 (2020)."},{"key":"e_1_2_2_16_1","volume-title":"Approximation complexity of maximum a posteriori inference in sum-product networks. arXiv preprint arXiv:1703.06045","author":"Conaty Diarmaid","year":"2017","unstructured":"Diarmaid Conaty, Denis D Mau\u00e1, and Cassio P De Campos. 2017. Approximation complexity of maximum a posteriori inference in sum-product networks. arXiv preprint arXiv:1703.06045 (2017)."},{"key":"e_1_2_2_17_1","volume-title":"Sketch techniques for approximate query processing. Foundations and Trends in Databases","author":"Cormode Graham","year":"2011","unstructured":"Graham Cormode. 2011. Sketch techniques for approximate query processing. Foundations and Trends in Databases. NOW publishers (2011), 15."},{"key":"e_1_2_2_18_1","volume-title":"Workshop on probabilistic reasoning in artificial intelligence. Citeseer, 27--32","author":"Gagliardi Fabio","unstructured":"Fabio Gagliardi Cozman et al. 2000. Generalizing variable elimination in Bayesian networks. In Workshop on probabilistic reasoning in artificial intelligence. Citeseer, 27--32."},{"key":"e_1_2_2_19_1","volume-title":"International Conference on Artificial Intelligence and Statistics. PMLR, 2742--2752","author":"Curtin Ryan","year":"2020","unstructured":"Ryan Curtin, Benjamin Moseley, Hung Ngo, XuanLong Nguyen, Dan Olteanu, and Maximilian Schleich. 2020. Rk-means: Fast clustering for relational data. In International Conference on Artificial Intelligence and Statistics. PMLR, 2742--2752."},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10590-1_4"},{"key":"e_1_2_2_21_1","unstructured":"Donko Donjerkovic and Raghu Ramakrishnan. 1999. Probabilistic optimization of top N queries. Technical Report. University of Wisconsin-Madison Department of Computer Sciences."},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3380574"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2018.00094"},{"key":"e_1_2_2_24_1","volume-title":"Data market platforms: Trading data assets to solve data problems. arXiv preprint arXiv:2002.01047","author":"Fernandez Raul Castro","year":"2020","unstructured":"Raul Castro Fernandez, Pranav Subramaniam, and Michael J Franklin. 2020. Data market platforms: Trading data assets to solve data problems. arXiv preprint arXiv:2002.01047 (2020)."},{"key":"e_1_2_2_25_1","volume-title":"Introducing Microsoft Power BI","author":"Ferrari Alberto","unstructured":"Alberto Ferrari and Marco Russo. 2016. Introducing Microsoft Power BI. Microsoft Press."},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3196959.3196962"},{"key":"e_1_2_2_27_1","volume-title":"Sigma Worksheet: Interactive Construction of OLAP Queries. arXiv preprint arXiv:2012.00697","author":"Gale James","year":"2020","unstructured":"James Gale, Max Seiden, Gretchen Atwood, Jason Frantz, Rob Woollen, and \u00c7agatay Demiralp. 2020. Sigma Worksheet: Interactive Construction of OLAP Queries. arXiv preprint arXiv:2012.00697 (2020)."},{"key":"e_1_2_2_28_1","volume-title":"Data cube: A relational aggregation operator generalizing group-by, cross-tab, and sub-totals. Data mining and knowledge discovery 1, 1","author":"Gray Jim","year":"1997","unstructured":"Jim Gray, Surajit Chaudhuri, Adam Bosworth, Andrew Layman, Don Reichart, Murali Venkatrao, Frank Pellow, and Hamid Pirahesh. 1997. Data cube: A relational aggregation operator generalizing group-by, cross-tab, and sub-totals. Data mining and knowledge discovery 1, 1 (1997), 29--53."},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1265530.1265535"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.1998.655768"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/69.536253"},{"key":"e_1_2_2_32_1","unstructured":"Ashish Gupta Venky Harinarayan and Dallan Quass. 1995. Aggregate-query processing in data warehousing environments. (1995)."},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/223784.223817"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/1516360.1516376"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3064027"},{"key":"e_1_2_2_36_1","volume-title":"Aggregations over generalized hypertree decompositions. arXiv preprint arXiv:1508.07532","author":"Joglekar Manas","year":"2015","unstructured":"Manas Joglekar, Rohan Puttagunta, and Christopher R\u00e9. 2015. Aggregations over generalized hypertree decompositions. arXiv preprint arXiv:1508.07532 (2015)."},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2902251.2902293"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF02289635"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3426865"},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3209889.3209896"},{"key":"e_1_2_2_41_1","volume-title":"Probabilistic graphical models: principles and techniques","author":"Koller Daphne","unstructured":"Daphne Koller and Nir Friedman. 2009. Probabilistic graphical models: principles and techniques. MIT press."},{"key":"e_1_2_2_42_1","doi-asserted-by":"crossref","unstructured":"Steffen L Lauritzen and Nuala A Sheehan. 2003. Graphical models for genetic analyses. Statist. Sci. (2003) 489--514.","DOI":"10.1214\/ss\/1081443232"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-017-0480-7"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1007\/11547686_10"},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2013.179"},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2014.2346452"},{"key":"e_1_2_2_47_1","volume-title":"Computer Graphics Forum","author":"Liu Zhicheng","unstructured":"Zhicheng Liu, Biye Jiang, and Jeffrey Heer. 2013. imMens: Real-time visual querying of big data. In Computer Graphics Forum, Vol. 32. Wiley Online Library, 421--430."},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3290605.3300924"},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.6017\/ital.v31i3.1919"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3183733"},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3180143"},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3183758"},{"key":"e_1_2_2_53_1","volume-title":"of Transportation Statisticsw","author":"B.","year":"2017","unstructured":"B. of Transportation Statisticsw. 2017. Bureau of transportation statistics. http:\/\/www.transtats.bts.gov."},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/2656335"},{"key":"e_1_2_2_55_1","volume-title":"Hashedcubes: Simple, low memory, real-time visual exploration of big data","author":"Pahins Cicero AL","year":"2016","unstructured":"Cicero AL Pahins, Sean A Stephens, Carlos Scheidegger, and Joao LD Comba. 2016. Hashedcubes: Simple, low memory, real-time visual exploration of big data. IEEE transactions on visualization and computer graphics 23, 1 (2016), 671--680."},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","unstructured":"The pandas development team. 2020. pandas-dev\/pandas: Pandas. https:\/\/doi.org\/10.5281\/zenodo.3509134","DOI":"10.5281\/zenodo.3509134"},{"key":"e_1_2_2_57_1","unstructured":"Judea Pearl. 1982. Reverend Bayes on inference engines: A distributed hierarchical approach. Cognitive Systems Laboratory School of Engineering and Applied Science . . . ."},{"key":"e_1_2_2_58_1","volume-title":"AMIA Summits on Translational Science Proceedings","author":"Pineda Arturo Lopez","year":"2015","unstructured":"Arturo Lopez Pineda and Vanathi Gopalakrishnan. 2015. Novel application of junction trees to the interpretation of epigenetic differences among lung cancer subtypes. AMIA Summits on Translational Science Proceedings (2015), 31."},{"key":"e_1_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3320212"},{"key":"e_1_2_2_60_1","volume-title":"Robotics and Automotive Mechanics Conference (CERMA). IEEE, 301--306","author":"Ramirez Julio C","year":"2009","unstructured":"Julio C Ramirez, Guillermina Munoz, and Ludivina Gutierrez. 2009. Fault diagnosis in an industrial process using Bayesian Networks: Application of the junction tree algorithm. In 2009 Electronics, Robotics and Automotive Mechanics Conference (CERMA). IEEE, 301--306."},{"key":"e_1_2_2_61_1","volume-title":"Semantic caching and query processing","author":"Ren Qun","year":"2003","unstructured":"Qun Ren, Margaret H Dunham, and Vijay Kumar. 2003. Semantic caching and query processing. IEEE transactions on knowledge and data engineering 15, 1 (2003), 192--210."},{"key":"e_1_2_2_62_1","volume-title":"Graph databases: new opportunities for connected data. \" O'Reilly Media","author":"Robinson Ian","unstructured":"Ian Robinson, Jim Webber, and Emil Eifrem. 2015. Graph databases: new opportunities for connected data. \" O'Reilly Media, Inc.\"."},{"key":"e_1_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/342009.335419"},{"key":"e_1_2_2_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3461670"},{"key":"e_1_2_2_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3324961"},{"key":"e_1_2_2_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2882939"},{"key":"e_1_2_2_67_1","doi-asserted-by":"publisher","DOI":"10.1145\/582095.582099"},{"key":"e_1_2_2_68_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF01531015"},{"key":"e_1_2_2_69_1","doi-asserted-by":"publisher","DOI":"10.1007\/s007780050040"},{"key":"e_1_2_2_70_1","doi-asserted-by":"publisher","DOI":"10.1145\/1031763.1031783"},{"key":"e_1_2_2_71_1","volume-title":"Proc. International Conference on Database Theory.","author":"Veldhuizen Todd L","year":"2014","unstructured":"Todd L Veldhuizen. 2014. Leapfrog triejoin: A simple, worst-case optimal join algorithm. In Proc. International Conference on Database Theory."},{"key":"e_1_2_2_72_1","volume-title":"MAP estimation via agreement on trees: message-passing and linear programming","author":"Wainwright Martin J","year":"2005","unstructured":"Martin J Wainwright, Tommi S Jaakkola, and Alan S Willsky. 2005. MAP estimation via agreement on trees: message-passing and linear programming. IEEE transactions on information theory 51, 11 (2005), 3697--3717."},{"key":"e_1_2_2_73_1","doi-asserted-by":"publisher","DOI":"10.1109\/SNPD.2007.461"},{"key":"e_1_2_2_74_1","volume-title":"Memory-efficient group-by aggregates over multi-way joins. arXiv preprint arXiv:1906.05745","author":"Xirogiannopoulos Konstantinos","year":"2019","unstructured":"Konstantinos Xirogiannopoulos and Amol Deshpande. 2019. Memory-efficient group-by aggregates over multi-way joins. arXiv preprint arXiv:1906.05745 (2019)."},{"key":"e_1_2_2_75_1","first-page":"82","article-title":"Algorithms for acyclic database schemes","volume":"81","author":"Yannakakis Mihalis","year":"1981","unstructured":"Mihalis Yannakakis. 1981. Algorithms for acyclic database schemes. In VLDB, Vol. 81. 82--94.","journal-title":"VLDB"},{"key":"e_1_2_2_76_1","volume-title":"Proceedings. Computer Software and The IEEE Computer Society's Third International Applications Conference","author":"Yu Clement Tak","year":"1979","unstructured":"Clement Tak Yu and Meral Z Ozsoyoglu. 1979. An algorithm for tree-query membership of a distributed query. In COMPSAC 79. Proceedings. Computer Software and The IEEE Computer Society's Third International Applications Conference, 1979. IEEE, 306--312."},{"key":"e_1_2_2_77_1","doi-asserted-by":"publisher","DOI":"10.14778\/3342263.3342267"},{"key":"e_1_2_2_78_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.trc.2014.12.009"}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3626735","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3626735","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,22]],"date-time":"2025-08-22T12:59:29Z","timestamp":1755867569000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3626735"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,12,8]]},"references-count":78,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2023,12,8]]}},"alternative-id":["10.1145\/3626735"],"URL":"https:\/\/doi.org\/10.1145\/3626735","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,12,8]]}}}