{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T10:15:21Z","timestamp":1781518521680,"version":"3.54.1"},"reference-count":47,"publisher":"Association for Computing Machinery (ACM)","issue":"11","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2023,7]]},"abstract":"<jats:p>Analysts and scientists are interested in querying streams of video, audio, and text to extract quantitative insights. For example, an urban planner may wish to measure congestion by querying the live feed from a traffic camera. Prior work has used deep neural networks (DNNs) to answer such queries in the batch setting. However, much of this work is not suited for the streaming setting because it requires access to the entire dataset before a query can be submitted or is specific to video. Thus, to the best of our knowledge, no prior work addresses the problem of efficiently answering queries over multiple modalities of streams.<\/jats:p>\n          <jats:p>In this work we propose InQuest, a system for accelerating aggregation queries on unstructured streams of data with statistical guarantees on query accuracy. InQuest leverages inexpensive approximation models (\"proxies\") and sampling techniques to limit the execution of an expensive high-precision model (an \"oracle\") to a subset of the stream. It then uses the oracle predictions to compute an approximate query answer in real-time. We theoretically analyzed InQuest and show that the expected error of its query estimates converges on stationary streams at a rate inversely proportional to the oracle budget. We evaluated our algorithm on six real-world video and text datasets and show that InQuest achieves the same root mean squared error (RMSE) as two streaming baselines with up to 5.0x fewer oracle invocations. We further show that InQuest can achieve up to 1.9x lower RMSE at a fixed number of oracle invocations than a state-of-the-art batch setting algorithm.<\/jats:p>","DOI":"10.14778\/3611479.3611496","type":"journal-article","created":{"date-parts":[[2023,8,25]],"date-time":"2023-08-25T02:08:08Z","timestamp":1692929288000},"page":"2897-2910","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Accelerating Aggregation Queries on Unstructured Streams of Data"],"prefix":"10.14778","volume":"16","author":[{"given":"Matthew","family":"Russo","sequence":"first","affiliation":[{"name":"Stanford University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tatsunori","family":"Hashimoto","sequence":"additional","affiliation":[{"name":"Stanford University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Daniel","family":"Kang","sequence":"additional","affiliation":[{"name":"University of Illinois Urbana-Champaign"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yi","family":"Sun","sequence":"additional","affiliation":[{"name":"University of Chicago"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Matei","family":"Zaharia","sequence":"additional","affiliation":[{"name":"Stanford University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,8,24]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.5555\/1182635.1164180"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/SSDBM.2007.29"},{"key":"e_1_2_1_3_1","volume-title":"Retrieved","year":"2022","unstructured":"ALERTWildfire. 2022 . AlertWildfire . Retrieved Dec. 28, 2022 from https:\/\/www.alertwildfire.org\/ ALERTWildfire. 2022. AlertWildfire. Retrieved Dec. 28, 2022 from https:\/\/www.alertwildfire.org\/"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/icde.2019.00132"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/872757.872854"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/603867.603884"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3389692"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00051"},{"key":"e_1_2_1_9_1","volume-title":"Approximation, Randomization, and Combinatorial Optimization. Algorithms and Techniques","author":"Braverman Vladimir","unstructured":"Vladimir Braverman and Rafail Ostrovsky . 2013. Generalizing the layering method of indyk and woodruff: Recursive sketches for frequency-based vectors on streams . In Approximation, Randomization, and Combinatorial Optimization. Algorithms and Techniques . Springer , 58--70. Vladimir Braverman and Rafail Ostrovsky. 2013. Generalizing the layering method of indyk and woodruff: Recursive sketches for frequency-based vectors on streams. In Approximation, Randomization, and Combinatorial Optimization. Algorithms and Techniques. Springer, 58--70."},{"key":"e_1_2_1_10_1","volume-title":"On the use of a pilot sample for sample size determination. Statistics in medicine 14, 17","author":"Browne Richard H","year":"1995","unstructured":"Richard H Browne . 1995. On the use of a pilot sample for sample size determination. Statistics in medicine 14, 17 ( 1995 ), 1933--1940. Richard H Browne. 1995. On the use of a pilot sample for sample size determination. Statistics in medicine 14, 17 (1995), 1933--1940."},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the 2nd SysML Conference","author":"Canel Christopher","year":"1905","unstructured":"Christopher Canel , Thomas Kim , Giulio Zhou , Conglong Li , Hyeontaek Lim , David G. Andersen , Michael Kaminsky , and Subramanya R. Dulloor . 2019. Scaling Video Analytics on Constrained Edge Nodes . In Proceedings of the 2nd SysML Conference . Palo Alto, CA, USA, 12 pages. arXiv : 1905 .13536 http:\/\/arxiv.org\/abs\/1905.13536 Christopher Canel, Thomas Kim, Giulio Zhou, Conglong Li, Hyeontaek Lim, David G. Andersen, Michael Kaminsky, and Subramanya R. Dulloor. 2019. Scaling Video Analytics on Constrained Edge Nodes. In Proceedings of the 2nd SysML Conference. Palo Alto, CA, USA, 12 pages. arXiv:1905.13536 http:\/\/arxiv.org\/abs\/1905.13536"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1016\/B978-155860869-6\/50026-3"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448016.3452803"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/n19-1423"},{"key":"e_1_2_1_15_1","volume-title":"Retrieved","author":"AWS","year":"2023","unstructured":"AWS EC2. 2023 . G4 On-Demand Pricing . Retrieved Apr. 11, 2023 from https:\/\/aws.amazon.com\/ec2\/instance-types\/g4\/ AWS EC2. 2023. G4 On-Demand Pricing. Retrieved Apr. 11, 2023 from https:\/\/aws.amazon.com\/ec2\/instance-types\/g4\/"},{"key":"e_1_2_1_16_1","volume-title":"Retrieved","author":"Flink Apache","year":"2023","unstructured":"Apache Flink . 2023 . Tumbling Windows . Retrieved May 13, 2023 from https:\/\/nightlies.apache.org\/flink\/flink-docs-master\/docs\/dev\/datastream\/operators\/windows\/#tumbling-windows Apache Flink. 2023. Tumbling Windows. Retrieved May 13, 2023 from https:\/\/nightlies.apache.org\/flink\/flink-docs-master\/docs\/dev\/datastream\/operators\/windows\/#tumbling-windows"},{"key":"e_1_2_1_17_1","volume-title":"SOSP 2019 Workshop on AI Systems. 16 pages. arXiv:1910","author":"Fu Daniel Y.","year":"2019","unstructured":"Daniel Y. Fu , Will Crichton , James Hong , Xinwei Yao , Haotian Zhang , Anh Truong , Avanika Narayan , Maneesh Agrawala , Christopher R\u00e9 , and Kayvon Fatahalian . 2019 . Rekall: Specifying Video Events using Compositions of Spatiotemporal Labels . In SOSP 2019 Workshop on AI Systems. 16 pages. arXiv:1910 .02993 http:\/\/arxiv.org\/abs\/1910.02993 Daniel Y. Fu, Will Crichton, James Hong, Xinwei Yao, Haotian Zhang, Anh Truong, Avanika Narayan, Maneesh Agrawala, Christopher R\u00e9, and Kayvon Fatahalian. 2019. Rekall: Specifying Video Events using Compositions of Spatiotemporal Labels. In SOSP 2019 Workshop on AI Systems. 16 pages. arXiv:1910.02993 http:\/\/arxiv.org\/abs\/1910.02993"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.322"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447548.3467134"},{"key":"e_1_2_1_21_1","volume-title":"Focus: Querying Large Video Datasets with Low Latency and Low Cost. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)","author":"Hsieh Kevin","year":"2018","unstructured":"Kevin Hsieh , Ganesh Ananthanarayanan , Peter Bodik , Shivaram Venkataraman , Paramvir Bahl , Matthai Philipose , Phillip B. Gibbons , and Onur Mutlu . 2018 . Focus: Querying Large Video Datasets with Low Latency and Low Cost. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) . USENIX Association, Carlsbad, CA, 269--286. https:\/\/www.usenix.org\/conference\/osdi18\/presentation\/hsieh Kevin Hsieh, Ganesh Ananthanarayanan, Peter Bodik, Shivaram Venkataraman, Paramvir Bahl, Matthai Philipose, Phillip B. Gibbons, and Onur Mutlu. 2018. Focus: Querying Large Video Datasets with Low Latency and Low Cost. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). USENIX Association, Carlsbad, CA, 269--286. https:\/\/www.usenix.org\/conference\/osdi18\/presentation\/hsieh"},{"key":"e_1_2_1_22_1","volume-title":"Retrieved","year":"2022","unstructured":"HuggingFace. 2022 . Twitter-roBERTa-base for Sentiment Analysis . Retrieved Dec. 29, 2022 from https:\/\/huggingface.co\/cardiffnlp\/twitter-roberta-base-sentiment-latest HuggingFace. 2022. Twitter-roBERTa-base for Sentiment Analysis. Retrieved Dec. 29, 2022 from https:\/\/huggingface.co\/cardiffnlp\/twitter-roberta-base-sentiment-latest"},{"key":"e_1_2_1_23_1","volume-title":"Retrieved","year":"2022","unstructured":"Kaggle. 2022 . Customer Support on Twitter . Retrieved Dec. 29, 2022 from https:\/\/www.kaggle.com\/datasets\/thoughtvector\/customer-support-on-twitter Kaggle. 2022. Customer Support on Twitter. Retrieved Dec. 29, 2022 from https:\/\/www.kaggle.com\/datasets\/thoughtvector\/customer-support-on-twitter"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.14778\/3372716.3372725"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.14778\/3137628.3137664"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.14778\/3407790.3407804"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476249.3476285"},{"key":"e_1_2_1_28_1","volume-title":"Retrieved","author":"Kang Daniel","year":"2021","unstructured":"Daniel Kang , John Guibas , Peter Bailis , Tatsunori Hashimoto , Yi Sun , and Matei Zaharia . 2021 . Proof: Accelerating Approximate Aggregation Queries with Expensive Predicates . Retrieved Dec. 30, 2022 from https:\/\/ddkang.github.io\/papers\/2021\/abae-tech-report.pdf Daniel Kang, John Guibas, Peter Bailis, Tatsunori Hashimoto, Yi Sun, and Matei Zaharia. 2021. Proof: Accelerating Approximate Aggregation Queries with Expensive Predicates. Retrieved Dec. 30, 2022 from https:\/\/ddkang.github.io\/papers\/2021\/abae-tech-report.pdf"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3517897"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.14778\/3425879.3425881"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/IC2EW.2016.56"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2020.3048606"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-demo.25"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2002.994774"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE53745.2022.00266"},{"key":"e_1_2_1_36_1","volume-title":"Retrieved","author":"Penn State Eberly College of Science","year":"2023","unstructured":"Penn State Eberly College of Science : Department of Statistics. 2023. Lesson 6: Stratified Sampling . Retrieved Jul. 17, 2023 from https:\/\/online.stat.psu.edu\/stat506\/book\/export\/html\/655 Penn State Eberly College of Science: Department of Statistics. 2023. Lesson 6: Stratified Sampling. Retrieved Jul. 17, 2023 from https:\/\/online.stat.psu.edu\/stat506\/book\/export\/html\/655"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.2000.10473909"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1002\/9781118445112.stat05999"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/971697.602294"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201394"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1561\/2200000070"},{"key":"e_1_2_1_43_1","volume-title":"Retrieved","author":"Stats Internet Live","year":"2013","unstructured":"Internet Live Stats . 2013 . Twitter Usage Statistics . Retrieved Nov. 27, 2022 from https:\/\/www.internetlivestats.com\/twitter-statistics\/ Internet Live Stats. 2013. Twitter Usage Statistics. Retrieved Nov. 27, 2022 from https:\/\/www.internetlivestats.com\/twitter-statistics\/"},{"key":"e_1_2_1_44_1","volume-title":"Retrieved","author":"Tracker Twitch","year":"2022","unstructured":"Twitch Tracker . 2022 . Twitch broadcast time for all channels by month . Retrieved Nov. 27, 2022 from https:\/\/twitchtracker.com\/statistics\/stream-time Twitch Tracker. 2022. Twitch broadcast time for all channels by month. Retrieved Nov. 27, 2022 from https:\/\/twitchtracker.com\/statistics\/stream-time"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2021.3094997"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/2517349.2522737"},{"key":"e_1_2_1_47_1","volume-title":"14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17)","author":"Zhang Haoyu","unstructured":"Haoyu Zhang , Ganesh Ananthanarayanan , Peter Bodik , Matthai Philipose , Paramvir Bahl , and Michael J. Freedman . 2017. Live Video Analytics at Scale with Approximation and Delay-Tolerance . In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17) . USENIX Association, Boston, MA, 377--392. https:\/\/www.usenix.org\/conference\/nsdi17\/technical-sessions\/presentation\/zhang Haoyu Zhang, Ganesh Ananthanarayanan, Peter Bodik, Matthai Philipose, Paramvir Bahl, and Michael J. Freedman. 2017. Live Video Analytics at Scale with Approximation and Delay-Tolerance. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17). USENIX Association, Boston, MA, 377--392. https:\/\/www.usenix.org\/conference\/nsdi17\/technical-sessions\/presentation\/zhang"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.14778\/3372716.3372721"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3611479.3611496","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,23]],"date-time":"2023-09-23T22:14:16Z","timestamp":1695507256000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3611479.3611496"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7]]},"references-count":47,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2023,7]]}},"alternative-id":["10.14778\/3611479.3611496"],"URL":"https:\/\/doi.org\/10.14778\/3611479.3611496","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2023,7]]},"assertion":[{"value":"2023-08-24","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}