{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T19:07:23Z","timestamp":1779131243527,"version":"3.51.4"},"reference-count":38,"publisher":"Association for Computing Machinery (ACM)","issue":"3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2026,5,18]]},"abstract":"<jats:p>Several data warehouse and database providers have recently introduced extensions to SQL called AI Queries, enabling users to specify functions and conditions in SQL that are evaluated by LLMs, thereby broadening significantly the kinds of queries one can express over the combination of structured and unstructured data. LLMs offer remarkable semantic reasoning capabilities, making them an essential tool for complex and nuanced queries that blend structured and unstructured data. While extremely powerful, these AI queries can become prohibitively costly when invoked thousands of times.<\/jats:p>\n                  <jats:p>\n                    This paper provides an extensive evaluation of a recent AI query approximation approach that enables low cost analytics and database applications to benefit from AI queries. The approach delivers &gt;100x cost and latency reduction for the semantic filter (AI.IF) operator and also important gains for semantic ranking (AI.RANK). The cost and performance gains come from utilizing cheap and accurate proxy models over embedding vectors. We show that despite the massive gains in latency and cost, these proxy models preserve accuracy and occasionally improve accuracy across various benchmark datasets, including the extended Amazon reviews benchmark that has 10M rows. We present an OLAP-friendly architecture within Google\n                    <jats:italic toggle=\"yes\">BigQuery<\/jats:italic>\n                    for this approach for purely online (ad hoc) queries, and a low-latency HTAP database-friendly architecture in\n                    <jats:italic toggle=\"yes\">AlloyDB<\/jats:italic>\n                    that could further improve the latency by moving the proxy model training offline. We present techniques that accelerate the proxy model training.\n                  <\/jats:p>","DOI":"10.1145\/3802002","type":"journal-article","created":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T18:19:16Z","timestamp":1779128356000},"page":"1-23","source":"Crossref","is-referenced-by-count":0,"title":["100x Cost &amp; Latency Reduction: Performance Analysis of AI Query Approximation using Lightweight Proxy Models: [Experiments &amp; Analysis]"],"prefix":"10.1145","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-3947-2782","authenticated-orcid":false,"given":"Yeounoh","family":"Chung","sequence":"first","affiliation":[{"name":"Google Cloud, Sunnyvale, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-7406-8433","authenticated-orcid":false,"given":"Rushabh","family":"Desai","sequence":"additional","affiliation":[{"name":"Google Cloud, Seattle, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-5088-4308","authenticated-orcid":false,"given":"Jian","family":"He","sequence":"additional","affiliation":[{"name":"Google Cloud, Seattle, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-4837-151X","authenticated-orcid":false,"given":"Yu","family":"Xiao","sequence":"additional","affiliation":[{"name":"Google Cloud, Seattle, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-4491-2227","authenticated-orcid":false,"given":"Thibaud","family":"Hottelier","sequence":"additional","affiliation":[{"name":"Google Cloud, Seattle, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2901-6930","authenticated-orcid":false,"given":"Yves-Laurent","family":"Kom Samo","sequence":"additional","affiliation":[{"name":"Google Cloud, Seattle, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-6625-1423","authenticated-orcid":false,"given":"Pushkar","family":"Khadilkar","sequence":"additional","affiliation":[{"name":"Google Cloud, Pune, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-7659-7152","authenticated-orcid":false,"given":"Xianshun","family":"Chen","sequence":"additional","affiliation":[{"name":"Google Cloud, Seattle, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-4562-0557","authenticated-orcid":false,"given":"Sam","family":"Idicula","sequence":"additional","affiliation":[{"name":"Google Cloud, Sunnyvale, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4418-4724","authenticated-orcid":false,"given":"Fatma","family":"Ozcan","sequence":"additional","affiliation":[{"name":"Google Cloud, Sunnyvale, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8717-7356","authenticated-orcid":false,"given":"Alon","family":"Halevy","sequence":"additional","affiliation":[{"name":"Google Cloud, Sunnyvale, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-6360-9496","authenticated-orcid":false,"given":"Yannis","family":"Papakonstantinou","sequence":"additional","affiliation":[{"name":"Google Cloud, Sunnyvale, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,5,18]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Boost your Search and RAG agents with Vertex AI's new state-of-the-art Ranking API. Google Cloud Blog (30","author":"Bruderer Lukas","year":"2025","unstructured":"Lukas Bruderer and Mihai Ciorobea. 2025. Boost your Search and RAG agents with Vertex AI's new state-of-the-art Ranking API. Google Cloud Blog (30 May 2025). https:\/\/cloud.google.com\/blog\/products\/ai-machine-learning\/launching-our-new-state-of-the-art-vertex-ai-ranking-api"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/1622407.1622416"},{"key":"e_1_2_1_3_1","volume-title":"Wayne Xin Zhao, et al","author":"Chern Ethan","year":"2023","unstructured":"Ethan Chern, Steffi Freihat, Yangni Shieh, Stephen Wan, Junjie Zhao, Wayne Xin Zhao, et al., 2023. FacTool: Factuality Detection in Generative AI - A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios. arXiv preprint arXiv:2307.13528 (2023). https:\/\/arxiv.org\/abs\/2307.13528"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2019.2916074"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2019.00138"},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the Thirty-First Text REtrieval Conference (TREC 2022)","author":"Craswell Nick","year":"2022","unstructured":"Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, Jimmy Lin, Ellen M. Voorhees, and Ian Soboroff. 2022. Overview of the TREC 2022 Deep Learning Track. In Proceedings of the Thirty-First Text REtrieval Conference (TREC 2022) (NIST Special Publication 500-338). https:\/\/trec.nist.gov\/pubs\/trec31\/papers\/Overview_deep.pdf"},{"key":"e_1_2_1_7_1","first-page":"29807","article-title":"UQE: A Query Engine for Unstructured Databases","volume":"37","author":"Dai Hanjun","year":"2024","unstructured":"Hanjun Dai, Bethany Wang, Xingchen Wan, Bo Dai, Sherry Yang, Azade Nova, Pengcheng Yin, Mangpo Phothilimthana, Charles Sutton, and Dale Schuurmans. 2024. UQE: A Query Engine for Unstructured Databases. Advances in Neural Information Processing Systems, Vol. 37 (2024), 29807-29838.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_8_1","unstructured":"Databricks. 2025. AI Functions on Databricks. https:\/\/docs.databricks.com\/aws\/en\/large-language-models\/ai-functions. Accessed: 2025-07-31."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.14778\/3750601.3750685"},{"key":"e_1_2_1_10_1","volume-title":"A Comprehensive Survey on Imbalanced Data Learning. arXiv preprint arXiv:2502.08960","author":"Gao Xinyi","year":"2025","unstructured":"Xinyi Gao, Dongting Xie, Yihang Zhang, Zhengren Wang, Chong Chen, Conghui He, Hongzhi Yin, and Wentao Zhang. 2025. A Comprehensive Survey on Imbalanced Data Learning. arXiv preprint arXiv:2502.08960 (2025)."},{"key":"e_1_2_1_11_1","unstructured":"Google. 2025. google\/embedding-gemma-300m. https:\/\/huggingface.co\/google\/embeddinggemma-300m"},{"key":"e_1_2_1_12_1","unstructured":"Google Cloud. 2025a. AlloyDB AI. https:\/\/cloud.google.com\/alloydb\/ai?e=48754805. Accessed: 2025-07-31."},{"key":"e_1_2_1_13_1","volume-title":"Accessed","author":"Cloud Google","year":"2025","unstructured":"Google Cloud. 2025b. BigQuery ML overview. https:\/\/cloud.google.com\/bigquery\/docs\/bqml-introduction. Accessed: July 31, 2025."},{"key":"e_1_2_1_14_1","unstructured":"Google Cloud. 2025c. Generative AI pricing. Google. https:\/\/cloud.google.com\/vertex-ai\/generative-ai\/pricing"},{"key":"e_1_2_1_15_1","volume-title":"Gemini: Model Thinking Updates. Google DeepMind Blog Post. https:\/\/blog.google\/technology\/google-deepmind\/gemini-model-thinking-updates-march-2025\/ Accessed","author":"DeepMind Google","year":"2025","unstructured":"Google DeepMind. 2025. Gemini: Model Thinking Updates. Google DeepMind Blog Post. https:\/\/blog.google\/technology\/google-deepmind\/gemini-model-thinking-updates-march-2025\/ Accessed: October 2025."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.100"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3654989"},{"key":"e_1_2_1_18_1","volume-title":"Dspy: Compiling declarative language model calls into self-improving pipelines. arXiv preprint arXiv:2310.03714","author":"Khattab Omar","year":"2023","unstructured":"Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T Joshi, Hanna Moazam, et al., 2023. Dspy: Compiling declarative language model calls into self-improving pipelines. arXiv preprint arXiv:2310.03714 (2023)."},{"key":"e_1_2_1_19_1","unstructured":"Jinhyuk Lee Feiyang Chen Sahil Dua Daniel Cer Madhuri Shanbhogue Iftekhar Naim Gustavo Hernandez Abrego Zhe Li Kaifeng Chen Henrique Schechter Vera Xiaoqi Ren Shanfeng Zhang Daniel Salz Michael Boratko Jay Han Blair Chen Shuo Huang Vikram Rao Paul Suganthan Feng Han Andreas Doumanoglou Nithi Gupta Fedor Moiseev Cathy Yip Aashi Jain Simon Baumgartner Shahrokh Shahi Frank Palma Gomez Sandeep Mariserla Min Choi Parashar Shah Sonam Goenka Ke Chen Ye Xia Koert Chen Sai Meher Karthik Duddu Yichang Chen Trevor Walker Wenlei Zhou Rakesh Ghiya Zach Gleicher Karan Gill Zhe Dong Mojtaba Seyedhosseini Yunhsuan Sung Raphael Hoffmann and Tom Duerig. 2025. Gemini Embedding: Generalizable Embeddings from Gemini. arXiv:2503.07891 [cs.CL]"},{"key":"e_1_2_1_20_1","volume-title":"Gustavo Hernandez Abrego, Weiqiang Shi, Nithi Gupta, Aditya Kusupati, Prateek Jain, Siddhartha Reddy Jonnalagadda, Ming-Wei Chang, and Iftekhar Naim.","author":"Lee Jinhyuk","year":"2024","unstructured":"Jinhyuk Lee, Zhuyun Dai, Xiaoqi Ren, Blair Chen, Daniel Cer, Jeremy R. Cole, Kai Hui, Michael Boratko, Rajvi Kapadia, Wen Ding, Yi Luan, Sai Meher Karthik Duddu, Gustavo Hernandez Abrego, Weiqiang Shi, Nithi Gupta, Aditya Kusupati, Prateek Jain, Siddhartha Reddy Jonnalagadda, Ming-Wei Chang, and Iftekhar Naim. 2024. Gecko: Versatile Text Embeddings Distilled from Large Language Models. arXiv:2403.20327 [cs.CL] https:\/\/arxiv.org\/abs\/2403.20327"},{"key":"e_1_2_1_21_1","volume-title":"Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Cand\u00e8s, and Tatsunori Hashimoto.","author":"Muennighoff Niklas","year":"2025","unstructured":"Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Cand\u00e8s, and Tatsunori Hashimoto. 2025. s1: Simple test-time scaling. arXiv:2501.19393 [cs.CL]"},{"key":"e_1_2_1_22_1","volume-title":"Re-Ranking Step by Step: Investigating Pre-Filtering for Re-Ranking with Large Language Models. arXiv preprint arXiv:2406.18740","author":"Nouriinanloo Baharan","year":"2024","unstructured":"Baharan Nouriinanloo and Maxime Lamothe. 2024. Re-Ranking Step by Step: Investigating Pre-Filtering for Re-Ranking with Large Language Models. arXiv preprint arXiv:2406.18740 (2024)."},{"key":"e_1_2_1_23_1","unstructured":"OpenAI. 2022. Classification using embeddings. https:\/\/cookbook.openai.com\/examples\/classification_using_embeddings. Accessed: 2025-07-29."},{"key":"e_1_2_1_24_1","volume-title":"Palimpzest: Optimizing AI-Powered Analytics with Declarative Query Processing. In Proceedings of the 11th International Conference on Very Large Databases (CIDR","author":"Patel Liana","year":"2025","unstructured":"Liana Patel, Siddharth Jha, Carlos Guestrin, and Matei Zaharia. [n.d.]. Palimpzest: Optimizing AI-Powered Analytics with Declarative Query Processing. In Proceedings of the 11th International Conference on Very Large Databases (CIDR 2025). 12."},{"key":"e_1_2_1_25_1","volume-title":"Lotus: Enabling semantic queries with llms over tables of unstructured and structured data. arXiv e-prints","author":"Patel Liana","year":"2024","unstructured":"Liana Patel, Siddharth Jha, Carlos Guestrin, and Matei Zaharia. 2024. Lotus: Enabling semantic queries with llms over tables of unstructured and structured data. arXiv e-prints (2024), arXiv-2407."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2078195"},{"key":"e_1_2_1_27_1","volume-title":"Abacus: A Cost-Based Optimizer for Semantic Operator Systems. arXiv preprint arXiv:2505.14661","author":"Russo Matthew","year":"2025","unstructured":"Matthew Russo, Sivaprasad Sudhir, Gerardo Vitagliano, Chunwei Liu, Tim Kraska, Samuel Madden, and Michael Cafarella. 2025. Abacus: A Cost-Based Optimizer for Semantic Operator Systems. arXiv preprint arXiv:2505.14661 (2025)."},{"key":"e_1_2_1_28_1","first-page":"30811","volume-title":"Matryoshka Representation Learning. International Conference on Machine Learning (ICML)","volume":"202","author":"Shaham Uri","year":"2023","unstructured":"Uri Shaham, Mor Geva, Roi Levi, and Yoav Shoham. 2023. Matryoshka Representation Learning. International Conference on Machine Learning (ICML), Vol. 202 (2023), 30811-30829."},{"key":"e_1_2_1_29_1","volume-title":"Docetl: Agentic query rewriting and evaluation for complex document processing. arXiv preprint arXiv:2410.12189","author":"Shankar Shreya","year":"2024","unstructured":"Shreya Shankar, Tristan Chambers, Tarak Shah, Aditya G Parameswaran, and Eugene Wu. 2024. Docetl: Agentic query rewriting and evaluation for complex document processing. arXiv preprint arXiv:2410.12189 (2024)."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3726302.3730331"},{"key":"e_1_2_1_31_1","volume-title":"Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. arXiv preprint arXiv:2408.03314","author":"Snell Charlie Victor","year":"2024","unstructured":"Charlie Victor Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024. Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters. arXiv preprint arXiv:2408.03314 (2024)."},{"key":"e_1_2_1_32_1","unstructured":"Snowflake Inc. 2025. AI SQL. https:\/\/docs.snowflake.com\/en\/user-guide\/snowflake-cortex\/aisql. Accessed: 2025-07-31."},{"key":"e_1_2_1_33_1","volume-title":"Is ChatGPT good at search? investigating large language models as re-ranking agents. arXiv preprint arXiv:2304.09542","author":"Sun Weiwei","year":"2023","unstructured":"Weiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang, Pengjie Ren, Zhumin Chen, Dawei Yin, and Zhaochun Ren. 2023. Is ChatGPT good at search? investigating large language models as re-ranking agents. arXiv preprint arXiv:2304.09542 (2023)."},{"key":"e_1_2_1_34_1","unstructured":"TensorFlow Tutorial. 2023. Word embeddings. https:\/\/www.tensorflow.org\/text\/guide\/word_embeddings. Accessed: 2025-07-29."},{"key":"e_1_2_1_35_1","unstructured":"text2vec.org. 2018. Vectorization. https:\/\/text2vec.org\/vectorization.html. Accessed: 2025-07-29."},{"key":"e_1_2_1_36_1","unstructured":"The Devastator. 2022. DBpedia Ontology: Text Classification Dataset with 14 Classes. https:\/\/www.kaggle.com\/datasets\/thedevastator\/dbpedia-ontology-dataset Kaggle Dataset."},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.14778\/3704965.3704989"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3725411"}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3802002","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T18:31:54Z","timestamp":1779129114000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3802002"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,18]]},"references-count":38,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,5,18]]}},"alternative-id":["10.1145\/3802002"],"URL":"https:\/\/doi.org\/10.1145\/3802002","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,18]]}}}