{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T04:12:53Z","timestamp":1784261573685,"version":"3.55.0"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"9","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2025,5]]},"abstract":"<jats:p>\n            Analyzing unstructured data has been a persistent challenge in data processing. Recent proposals offer declarative frameworks for LLM-powered processing of unstructured data, but they typically execute user-specified operations as-is in a single LLM call\u2014focusing on cost rather than accuracy. This is problematic for complex tasks, where even well-prompted LLMs can miss relevant information. For instance, reliably extracting\n            <jats:italic toggle=\"yes\">all<\/jats:italic>\n            instances of a specific clause from legal documents often requires decomposing the task, the data, or both.\n          <\/jats:p>\n          <jats:p>\n            We present DocETL, a system that optimizes complex document processing pipelines, while accounting for LLM shortcomings. DocETL offers a declarative interface for users to deine such pipelines and uses an agent-based approach to automatically optimize them, leveraging novel agent-based rewrites (that we call\n            <jats:italic toggle=\"yes\">rewrite directives<\/jats:italic>\n            ), as well as an optimization and evaluation framework. We introduce\n            <jats:italic toggle=\"yes\">(i)<\/jats:italic>\n            logical rewriting of pipelines, tailored for LLM-based tasks,\n            <jats:italic toggle=\"yes\">(ii)<\/jats:italic>\n            an agent-guided plan evaluation mechanism, and\n            <jats:italic toggle=\"yes\">(iii)<\/jats:italic>\n            an optimization algorithm that efficiently finds promising plans, considering the latencies of LLM execution. Across four real-world document processing tasks, DocETL improves accuracy by 21\u201380% over strong baselines. DocETL is open-source at docetl.org and, as of March 2025, has over 1.7k GitHub stars across diverse domains.\n          <\/jats:p>","DOI":"10.14778\/3746405.3746426","type":"journal-article","created":{"date-parts":[[2025,9,3]],"date-time":"2025-09-03T17:06:20Z","timestamp":1756919180000},"page":"3035-3048","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":20,"title":["DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing"],"prefix":"10.14778","volume":"18","author":[{"given":"Shreya","family":"Shankar","sequence":"first","affiliation":[{"name":"UC Berkeley EECS"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tristan","family":"Chambers","sequence":"additional","affiliation":[{"name":"BIDS Police Records Access Project"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tarak","family":"Shah","sequence":"additional","affiliation":[{"name":"BIDS Police Records Access Project"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Aditya G.","family":"Parameswaran","sequence":"additional","affiliation":[{"name":"UC Berkeley EECS"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Eugene","family":"Wu","sequence":"additional","affiliation":[{"name":"Columbia University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,9,3]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Bohou Li, Mark Lindblad, Henry Lindeman, Alex Meyer, Parth Parmar, Tanvi Ranade, Mehul A. Shah, Benjamin Sowell, Dan Tecuci, Vinayak Thapliyal, and Matt Welsh.","author":"Anderson Eric","year":"2024","unstructured":"Eric Anderson, Jonathan Fritz, Austin Lee, Bohou Li, Mark Lindblad, Henry Lindeman, Alex Meyer, Parth Parmar, Tanvi Ranade, Mehul A. Shah, Benjamin Sowell, Dan Tecuci, Vinayak Thapliyal, and Matt Welsh. 2024. The Design of an LLM-powered Unstructured Analytics System. arXiv:2409.00847 [cs.DB] https:\/\/arxiv.org\/abs\/2409.00847"},{"key":"e_1_2_1_2_1","volume-title":"Welcome to LLMflation - LLM inference cost is going down fast. a16z Blog, https:\/\/a16z.com\/llmflation-llm-inference-cost\/","author":"Appenzeller Guido","year":"2024","unstructured":"Guido Appenzeller. 2024. Welcome to LLMflation - LLM inference cost is going down fast. a16z Blog, https:\/\/a16z.com\/llmflation-llm-inference-cost\/ (2024)."},{"key":"e_1_2_1_3_1","volume-title":"Language models enable simple systems for generating structured views of heterogeneous data lakes. arXiv preprint arXiv:2304.09433","author":"Arora Simran","year":"2023","unstructured":"Simran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan, Andrew Hojel, Immanuel Trummer, and Christopher R\u00e9. 2023. Language models enable simple systems for generating structured views of heterogeneous data lakes. arXiv preprint arXiv:2304.09433 (2023)."},{"key":"e_1_2_1_4_1","volume-title":"Longbench: A bilingual, multitask benchmark for long context understanding. arXiv preprint arXiv:2308.14508","author":"Bai Yushi","year":"2023","unstructured":"Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al. 2023. Longbench: A bilingual, multitask benchmark for long context understanding. arXiv preprint arXiv:2308.14508 (2023)."},{"key":"e_1_2_1_5_1","volume-title":"Natural language processing with Python: analyzing text with the natural language toolkit. \"O'Reilly Media","author":"Bird Steven","unstructured":"Steven Bird, Ewan Klein, and Edward Loper. 2009. Natural language processing with Python: analyzing text with the natural language toolkit. \"O'Reilly Media, Inc.\"."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/275487.275492"},{"key":"e_1_2_1_7_1","first-page":"20","article-title":"MapReduce online","volume":"10","author":"Condie Tyson","year":"2010","unstructured":"Tyson Condie, Neil Conway, Peter Alvaro, Joseph M Hellerstein, Khaled Elmeleegy, and Russell Sears. 2010. MapReduce online.. In Nsdi, Vol. 10. 20.","journal-title":"Nsdi"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.14778\/3636218.3636237"},{"key":"e_1_2_1_9_1","volume-title":"DeepJoin: Joinable Table Discovery with Pre-trained Language Models. arXiv preprint arXiv:2212.07588","author":"Dong Yuyang","year":"2022","unstructured":"Yuyang Dong, Chuan Xiao, Takuma Nozawa, Masafumi Enomoto, and Masafumi Oyamada. 2022. DeepJoin: Joinable Table Discovery with Pre-trained Language Models. arXiv preprint arXiv:2212.07588 (2022)."},{"key":"e_1_2_1_10_1","volume-title":"From local to global: A graph rag approach to query-focused summarization. arXiv preprint arXiv:2404.16130","author":"Edge Darren","year":"2024","unstructured":"Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, and Jonathan Larson. 2024. From local to global: A graph rag approach to query-focused summarization. arXiv preprint arXiv:2404.16130 (2024)."},{"key":"e_1_2_1_11_1","volume-title":"Jiani Zhang, Ziqing Hu, Yanjun Qi, Scott Nickleach, Diego Socolinsky, Srinivasan Sengamedu, and Christos Faloutsos.","author":"Fang Xi","year":"2024","unstructured":"Xi Fang, Weijie Xu, Fiona Anting Tan, Jiani Zhang, Ziqing Hu, Yanjun Qi, Scott Nickleach, Diego Socolinsky, Srinivasan Sengamedu, and Christos Faloutsos. 2024. Large Language Models on Tabular Data-A Survey. arXiv preprint arXiv:2402.17944 (2024)."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.14778\/3611479.3611527"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.14778\/3611479.3611527"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/1989323.1989331"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/163090.163100"},{"key":"e_1_2_1_16_1","volume-title":"The Cascades Framework for Query Optimization","author":"Graefe Goetz","year":"1995","unstructured":"Goetz Graefe. 1995. The Cascades Framework for Query Optimization. IEEE Data(base) Engineering Bulletin 18 (1995), 19\u201329. https:\/\/api.semanticscholar.org\/CorpusID:260706023"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/170036.170066"},{"key":"e_1_2_1_18_1","volume-title":"Readings in Database Systems","author":"Hellerstein Joseph M","year":"2005","unstructured":"Joseph M Hellerstein and Michael Stonebraker. 2005. Anatomy of a database system. Readings in Database Systems, (2005)."},{"key":"e_1_2_1_19_1","volume-title":"CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review. NeurIPS","author":"Hendrycks Dan","year":"2021","unstructured":"Dan Hendrycks, Collin Burns, Anya Chen, and Spencer Ball. 2021. CUAD: An Expert-Annotated NLP Dataset for Legal Contract Review. NeurIPS (2021)."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.5281\/zenodo.1212303"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1609\/icwsm.v8i1.14550"},{"key":"e_1_2_1_22_1","volume-title":"Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression. arXiv preprint arXiv:2310.06839","author":"Jiang Huiqiang","year":"2023","unstructured":"Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023. Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression. arXiv preprint arXiv:2310.06839 (2023)."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3618260.3649777"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.14778\/3137628.3137664"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.14778\/3659437.3659461"},{"key":"e_1_2_1_26_1","volume-title":"The Twelfth International Conference on Learning Representations.","author":"Khattab Omar","year":"2024","unstructured":"Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Saiful Haq, Ashutosh Sharma, Thomas T Joshi, Hanna Moazam, Heather Miller, et al. 2024. DSPy: Compiling Declarative Language Model Calls into State-of-the-Art Pipelines. In The Twelfth International Conference on Learning Representations."},{"key":"e_1_2_1_27_1","volume-title":"Same task, more tokens: the impact of input length on the reasoning performance of large language models. arXiv preprint arXiv:2402.14848","author":"Levy Mosh","year":"2024","unstructured":"Mosh Levy, Alon Jacoby, and Yoav Goldberg. 2024. Same task, more tokens: the impact of input length on the reasoning performance of large language models. arXiv preprint arXiv:2402.14848 (2024)."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.14778\/3229863.3236226"},{"key":"e_1_2_1_29_1","volume-title":"Towards Accurate and Efficient Document Analytics with Large Language Models. arXiv preprint arXiv:2405.04674","author":"Lin Yiming","year":"2024","unstructured":"Yiming Lin, Madelon Hulsebos, Ruiying Ma, Shreya Shankar, Sepanta Zeigham, Aditya G Parameswaran, and Eugene Wu. 2024. Towards Accurate and Efficient Document Analytics with Large Language Models. arXiv preprint arXiv:2405.04674 (2024)."},{"key":"e_1_2_1_30_1","volume-title":"Proceedings of the Conference on Innovative Database Research (CIDR).","author":"Liu Chunwei","year":"2025","unstructured":"Chunwei Liu, Matthew Russo, Michael Cafarella, Lei Cao, Peter Baile Chen, Zui Chen, Michael Franklin, Tim Kraska, Samuel Madden, Rana Shahout, et al. 2025. Palimpzest: Optimizing ai-powered analytics with declarative query processing. In Proceedings of the Conference on Innovative Database Research (CIDR)."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00638"},{"key":"e_1_2_1_32_1","volume-title":"Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators. In First Conference on Language Modeling. https:\/\/openreview.net\/forum?id=9gdZI7c6yr","author":"Liu Yinhong","year":"2024","unstructured":"Yinhong Liu, Han Zhou, Zhijiang Guo, Ehsan Shareghi, Ivan Vuli\u0107, Anna Korhonen, and Nigel Collier. 2024. Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators. In First Conference on Language Modeling. https:\/\/openreview.net\/forum?id=9gdZI7c6yr"},{"key":"e_1_2_1_33_1","unstructured":"Adam Marcus Eugene Wu David R Karger Samuel Madden and Robert C Miller. 2011. Crowdsourced databases: Query processing with people. Cidr."},{"key":"e_1_2_1_34_1","volume-title":"Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al.","author":"Nye Maxwell","year":"2021","unstructured":"Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al. 2021. Show your work: Scratchpads for intermediate computation with language models. arXiv preprint arXiv:2112.00114 (2021)."},{"key":"e_1_2_1_35_1","unstructured":"Pallets. 2024. Jinja. https:\/\/github.com\/pallets\/jinja\/. Version 3.1.x."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/2396761.2398421"},{"key":"e_1_2_1_37_1","volume-title":"Revisiting Prompt Engineering via Declarative Crowdsourcing. Cidr","author":"Parameswaran Aditya G","year":"2024","unstructured":"Aditya G Parameswaran, Shreya Shankar, Parth Asawa, Naman Jain, and Yujie Wang. 2024. Revisiting Prompt Engineering via Declarative Crowdsourcing. Cidr (2024)."},{"key":"e_1_2_1_38_1","volume-title":"Semantic Operators: A Declarative Model for Rich, AI-based Analytics Over Text Data. arXiv preprint arXiv:2407.11418","author":"Patel Liana","year":"2024","unstructured":"Liana Patel, Siddharth Jha, Parth Asawa, Melissa Pan, Carlos Guestrin, and Matei Zaharia. 2024. Semantic Operators: A Declarative Model for Rich, AI-based Analytics Over Text Data. arXiv preprint arXiv:2407.11418 (2024)."},{"key":"e_1_2_1_39_1","volume-title":"On limitations of the transformer architecture. arXiv preprint arXiv:2402.08164","author":"Peng Binghui","year":"2024","unstructured":"Binghui Peng, Srini Narayanan, and Christos Papadimitriou. 2024. On limitations of the transformer architecture. arXiv preprint arXiv:2402.08164 (2024)."},{"key":"e_1_2_1_40_1","volume-title":"Yu Gan, Amin Saberi, Fatma Ozcan, and Sercan O Arik.","author":"Pourreza Mohammadreza","year":"2024","unstructured":"Mohammadreza Pourreza, Hailong Li, Ruoxi Sun, Yeounoh Chung, Shayan Talaei, Gaurav Tarlok Kakkar, Yu Gan, Amin Saberi, Fatma Ozcan, and Sercan O Arik. 2024. Chase-sql: Multi-path reasoning and preference optimized candidate selection in text-to-sql. arXiv preprint arXiv:2410.01943 (2024)."},{"key":"e_1_2_1_41_1","volume-title":"CleanAgent: Automating Data Standardization with LLM-based Agents. arXiv preprint arXiv:2403.08291","author":"Qi Danrui","year":"2024","unstructured":"Danrui Qi and Jiannan Wang. 2024. CleanAgent: Automating Data Standardization with LLM-based Agents. arXiv preprint arXiv:2403.08291 (2024)."},{"key":"e_1_2_1_42_1","volume-title":"Abacus: A Cost-Based Optimizer for Semantic Operator Systems. arXiv preprint arXiv:2505.14661","author":"Russo Matthew","year":"2025","unstructured":"Matthew Russo, Sivaprasad Sudhir, Gerardo Vitagliano, Chunwei Liu, Tim Kraska, Samuel Madden, and Michael Cafarella. 2025. Abacus: A Cost-Based Optimizer for Semantic Operator Systems. arXiv preprint arXiv:2505.14661 (2025)."},{"key":"e_1_2_1_43_1","volume-title":"DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing. arXiv preprint arXiv:2410.12189","author":"Shankar Shreya","year":"2024","unstructured":"Shreya Shankar, Tristan Chambers, Tarak Shah, Aditya G Parameswaran, and Eugene Wu. 2024. DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing. arXiv preprint arXiv:2410.12189 (2024)."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.14778\/3685800.3685835"},{"key":"e_1_2_1_45_1","volume-title":"Building Reactive Large Language Model Pipelines with Motion. In Companion of the 2024 International Conference on Management of Data. 520\u2013523","author":"Shankar Shreya","year":"2024","unstructured":"Shreya Shankar and Aditya G Parameswaran. 2024. Building Reactive Large Language Model Pipelines with Motion. In Companion of the 2024 International Conference on Management of Data. 520\u2013523."},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3654777.3676450"},{"key":"e_1_2_1_47_1","volume-title":"International Conference on Machine Learning. PMLR, 31210\u201331227","author":"Shi Freda","year":"2023","unstructured":"Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Scharli, and Denny Zhou. 2023. Large language models can be easily distracted by irrelevant context. In International Conference on Machine Learning. PMLR, 31210\u201331227."},{"key":"e_1_2_1_48_1","volume-title":"Confabulation: The Surprising Value of Large Language Model Hallucinations. arXiv preprint arXiv:2406.04175","author":"Sui Peiqi","year":"2024","unstructured":"Peiqi Sui, Eamon Duede, Sophie Wu, and Richard Jean So. 2024. Confabulation: The Surprising Value of Large Language Model Hallucinations. arXiv preprint arXiv:2406.04175 (2024)."},{"key":"e_1_2_1_49_1","volume-title":"Found in the middle: Permutation self-consistency improves listwise ranking in large language models. arXiv preprint arXiv:2310.07712","author":"Tang Raphael","year":"2023","unstructured":"Raphael Tang, Xinyu Zhang, Xueguang Ma, Jimmy Lin, and Ferhan Ture. 2023. Found in the middle: Permutation self-consistency improves listwise ranking in large language models. arXiv preprint arXiv:2310.07712 (2023)."},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3517843"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3626246.3654732"},{"key":"e_1_2_1_52_1","volume-title":"Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. https:\/\/api.semanticscholar.org\/CorpusID:271114432","author":"Tempest","unstructured":"Tempest A. van Schaik and Brittany Pugh. 2024. A Field Guide to Automatic Evaluation of LLM-Generated Summaries. In Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. https:\/\/api.semanticscholar.org\/CorpusID:271114432"},{"key":"e_1_2_1_53_1","volume-title":"Idk cascades: Fast deep learning by learning not to overthink. arXiv preprint arXiv:1706.00885","author":"Wang Xin","year":"2017","unstructured":"Xin Wang, Yujia Luo, Daniel Crankshaw, Alexey Tumanov, Fisher Yu, and Joseph E Gonzalez. 2017. Idk cascades: Fast deep learning by learning not to overthink. arXiv preprint arXiv:1706.00885 (2017)."},{"key":"e_1_2_1_54_1","volume-title":"Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery. Advances in Neural Information Processing Systems 36","author":"Wen Yuxin","year":"2024","unstructured":"Yuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum, Jonas Geiping, and Tom Goldstein. 2024. Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery. Advances in Neural Information Processing Systems 36 (2024)."},{"key":"e_1_2_1_55_1","volume-title":"A prompt pattern catalog to enhance prompt engineering with chatgpt. arXiv preprint arXiv:2302.11382","author":"White Jules","year":"2023","unstructured":"Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry Gilbert, Ashraf Elnashar, Jesse Spencer-Smith, and Douglas C Schmidt. 2023. A prompt pattern catalog to enhance prompt engineering with chatgpt. arXiv preprint arXiv:2302.11382 (2023)."},{"key":"e_1_2_1_56_1","volume-title":"LongAgent: Scaling Language Models to 128k Context through Multi-Agent Collaboration. arXiv preprint arXiv:2402.11550","author":"Zhao Jun","year":"2024","unstructured":"Jun Zhao, Can Zu, Hao Xu, Yi Lu, Wei He, Yiwen Ding, Tao Gui, Qi Zhang, and Xuanjing Huang. 2024. LongAgent: Scaling Language Models to 128k Context through Multi-Agent Collaboration. arXiv preprint arXiv:2402.11550 (2024)."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3746405.3746426","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,4]],"date-time":"2025-09-04T19:50:20Z","timestamp":1757015420000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3746405.3746426"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5]]},"references-count":56,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2025,5]]}},"alternative-id":["10.14778\/3746405.3746426"],"URL":"https:\/\/doi.org\/10.14778\/3746405.3746426","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2025,5]]},"assertion":[{"value":"2025-09-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}