{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T19:07:31Z","timestamp":1779131251129,"version":"3.51.4"},"reference-count":104,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"DOI":"10.13039\/100006785","name":"Google","doi-asserted-by":"publisher","award":["junior ML faculty award"],"award-info":[{"award-number":["junior ML faculty award"]}],"id":[{"id":"10.13039\/100006785","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2026,5,18]]},"abstract":"<jats:p>Analytical join queries over unstructured data are increasingly prevalent in data analytics. Applying machine learning (ML) models to label every pair in the cross product of tables can achieve state-of-the-art accuracy, but the cost of pairwise execution of ML models is prohibitive. Existing algorithms, such as embedding-based blocking and sampling, aim to reduce this cost. However, they either fail to provide statistical guarantees (leading to errors up to 79% higher than expected) or become as inefficient as uniform sampling.<\/jats:p>\n                  <jats:p>We propose blocking-augmented sampling (BaS), which simultaneously achieves statistical guarantees and high efficiency. BaS optimally orchestrates embedding-based blocking and sampling to mitigate their respective limitations. Specifically, BaS allocates data tuples in the cross product into two regimes based on the failure modes of embeddings. In the regime of false negatives, BaS uses sampling to estimate the result. In the regime of false positives, BaS applies embedding-based blocking to improve efficiency. To minimize the estimation error given a budget for ML executions, we design a novel two-stage algorithm that adaptively allocates the budget between blocking and sampling. Theoretically, we prove that BaS asymptotically outperforms or matches standalone sampling. On real-world datasets across different modalities, we show that BaS provides valid confidence intervals and reduces estimation errors by up to 19\u00d7 compared to state-of-the-art baselines.<\/jats:p>","DOI":"10.1145\/3802004","type":"journal-article","created":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T18:19:16Z","timestamp":1779128356000},"page":"1-28","source":"Crossref","is-referenced-by-count":0,"title":["Accelerating Approximate Analytical Join Queries over Unstructured Data with Statistical Guarantees"],"prefix":"10.1145","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-6738-8930","authenticated-orcid":false,"given":"Yuxuan","family":"Zhu","sequence":"first","affiliation":[{"name":"University of Illinois Urbana-Champaign, Urbana, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-0353-4184","authenticated-orcid":false,"given":"Tengjun","family":"Jin","sequence":"additional","affiliation":[{"name":"University of Illinois Urbana-Champaign, Urbana, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-9860-6325","authenticated-orcid":false,"given":"Chenghao","family":"Mo","sequence":"additional","affiliation":[{"name":"University of Illinois Urbana-Champaign, Urbana, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9860-9938","authenticated-orcid":false,"given":"Daniel","family":"Kang","sequence":"additional","affiliation":[{"name":"University of Illinois Urbana-Champaign, Urbana, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,5,18]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"[n. d.]. First Quora Dataset Release: Question Pairs. https:\/\/quoradata.quora.com\/First-Quora-Dataset-Release-Question-Pairs. Accessed: 2024-01-08."},{"key":"e_1_2_1_2_1","volume-title":"Foundations of databases","author":"Abiteboul Serge","unstructured":"Serge Abiteboul, Richard Hull, and Victor Vianu. 1995. Foundations of databases. Addison-Wesley Reading."},{"key":"e_1_2_1_3_1","doi-asserted-by":"crossref","unstructured":"Swarup Acharya Phillip B Gibbons Viswanath Poosala and Sridhar Ramaswamy. 1999. Join synopses for approximate query answering. In SIGMOD.","DOI":"10.1145\/304182.304207"},{"key":"e_1_2_1_4_1","doi-asserted-by":"crossref","unstructured":"Mehdi Akbarian Rastaghi Ehsan Kamalloo and Davood Rafiei. 2022. Probing the Robustness of Pre-trained Language Models for Entity Matching. In CIKM.","DOI":"10.1145\/3511808.3557673"},{"key":"e_1_2_1_5_1","unstructured":"Alberto Bacchelli. 2013. Mining Challenge 2013: Stack Overflow. In MSR."},{"key":"e_1_2_1_6_1","volume-title":"Plagiarism meets paraphrasing: Insights for the next generation in automatic plagiarism detection. Computational Linguistics","author":"Alberto","year":"2013","unstructured":"Alberto Barr\u00f3n-Cede no, Marta Vila, M Ant\u00f2nia Mart\u00ed, and Paolo Rosso. 2013. Plagiarism meets paraphrasing: Insights for the next generation in automatic plagiarism detection. Computational Linguistics (2013)."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2006.13"},{"key":"e_1_2_1_8_1","doi-asserted-by":"crossref","unstructured":"Surajit Chaudhuri Venkatesh Ganti and Raghav Kaushik. 2006. A primitive operator for similarity joins in data cleaning. In ICDE.","DOI":"10.1109\/ICDE.2006.9"},{"key":"e_1_2_1_9_1","volume-title":"On random sampling over joins. SIGMOD","author":"Chaudhuri Surajit","year":"1999","unstructured":"Surajit Chaudhuri, Rajeev Motwani, and Vivek Narasayya. 1999. On random sampling over joins. SIGMOD (1999)."},{"key":"e_1_2_1_10_1","volume-title":"A survey of indexing techniques for scalable record linkage and deduplication","author":"Christen Peter","year":"2011","unstructured":"Peter Christen. 2011. A survey of indexing techniques for scalable record linkage and deduplication. IEEE transactions on knowledge and data engineering, Vol. 24, 9 (2011), 1537-1555."},{"key":"e_1_2_1_11_1","volume-title":"Entity resolution in the web of data","author":"Christophides Vassilis","unstructured":"Vassilis Christophides, Vasilis Efthymiou, and Kostas Stefanidis. 2015. Entity resolution in the web of data. Vol. 5. Springer."},{"key":"e_1_2_1_12_1","unstructured":"William Gemmell Cochran. 1977. Sampling techniques. john wiley & sons."},{"key":"e_1_2_1_13_1","volume-title":"Vehicle re-identification and travel time measurement in real-time on freeways using existing loop detector infrastructure. Transportation Research Record","author":"Coifman Benjamin","year":"1998","unstructured":"Benjamin Coifman. 1998. Vehicle re-identification and travel time measurement in real-time on freeways using existing loop detector infrastructure. Transportation Research Record (1998)."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/1631862.1631865"},{"key":"e_1_2_1_15_1","volume-title":"Proceedings 2003 VLDB Conference. Elsevier, 464-475","author":"Cormode Graham","year":"2003","unstructured":"Graham Cormode, Flip Korn, Shanmugavelayutham Muthukrishnan, and Divesh Srivastava. 2003. Finding hierarchical heavy hitters in data streams. In Proceedings 2003 VLDB Conference. Elsevier, 464-475."},{"key":"e_1_2_1_16_1","volume-title":"Extreme value theory: an introduction","author":"Haan Laurens De","unstructured":"Laurens De Haan and Ana Ferreira. 2006. Extreme value theory: an introduction. Springer."},{"key":"e_1_2_1_17_1","doi-asserted-by":"crossref","unstructured":"Timothy De Vries Hui Ke Sanjay Chawla and Peter Christen. 2009. Robust record linkage blocking using suffix arrays. In CIKM.","DOI":"10.1145\/1645953.1645994"},{"key":"e_1_2_1_18_1","volume-title":"Learning to recognize dialect features. arXiv preprint arXiv:2010.12707","author":"Demszky Dorottya","year":"2020","unstructured":"Dorottya Demszky, Devyani Sharma, Jonathan H Clark, Vinodkumar Prabhakaran, and Jacob Eisenstein. 2020. Learning to recognize dialect features. arXiv preprint arXiv:2010.12707 (2020)."},{"key":"e_1_2_1_19_1","volume-title":"International Conference on Data and Knowledge Engineering.","author":"Draisbach Uwe","year":"2011","unstructured":"Uwe Draisbach and Felix Naumann. 2011. A generalization of blocking and windowing algorithms for duplicate detection. In International Conference on Data and Knowledge Engineering."},{"key":"e_1_2_1_20_1","unstructured":"DunZhang. 2024. dunzhang\/stella_en_1.5B_v5. https:\/\/huggingface.co\/dunzhang\/stella_en_1.5B_v5 Accessed: 2024-11-20."},{"key":"e_1_2_1_21_1","volume-title":"Proceedings of the USENIX Annual Technical Conference (ATC). 1-14","author":"Duplyakin Dmitry","year":"2019","unstructured":"Dmitry Duplyakin, Robert Ricci, Aleksander Maricq, Gary Wong, Jonathon Duerig, Eric Eide, Leigh Stoller, Mike Hibler, David Johnson, Kirk Webb, Aditya Akella, Kuangching Wang, Glenn Ricart, Larry Landweber, Chip Elliott, Michael Zink, Emmanuel Cecchet, Snigdhaswin Kar, and Prabodh Mishra. 2019. The Design and Operation of CloudLab. In Proceedings of the USENIX Annual Technical Conference (ATC). 1-14. https:\/\/www.flux.utah.edu\/paper\/duplyakin-atc19"},{"key":"e_1_2_1_22_1","volume-title":"Duplicate record detection: A survey","author":"Elmagarmid Ahmed K","year":"2006","unstructured":"Ahmed K Elmagarmid, Panagiotis G Ipeirotis, and Vassilios S Verykios. 2006. Duplicate record detection: A survey. IEEE TKDE (2006)."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.1969.10501049"},{"key":"e_1_2_1_24_1","volume-title":"Academic plagiarism detection: a systematic literature review. Comput. Surveys","author":"Folt\u1ef3nek Tom\u00e1\u0161","year":"2019","unstructured":"Tom\u00e1\u0161 Folt\u1ef3nek, Norman Meuschke, and Bela Gipp. 2019. Academic plagiarism detection: a systematic literature review. Comput. Surveys (2019)."},{"key":"e_1_2_1_25_1","volume-title":"Ripple joins for online aggregation. ACM SIGMOD Record","author":"Haas Peter J","year":"1999","unstructured":"Peter J Haas and Joseph M Hellerstein. 1999. Ripple joins for online aggregation. ACM SIGMOD Record (1999)."},{"key":"e_1_2_1_26_1","volume-title":"Theoretical comparison of bootstrap confidence intervals. The Annals of Statistics","author":"Hall Peter","year":"1988","unstructured":"Peter Hall. 1988. Theoretical comparison of bootstrap confidence intervals. The Annals of Statistics (1988)."},{"key":"e_1_2_1_27_1","volume-title":"Transreid: Transformer-based object re-identification. In CVPR.","author":"He Shuting","year":"2021","unstructured":"Shuting He, Hao Luo, Pichao Wang, Fan Wang, Hao Li, and Wei Jiang. 2021. Transreid: Transformer-based object re-identification. In CVPR."},{"key":"e_1_2_1_28_1","volume-title":"Focus: Querying large video datasets with low latency and low cost. In OSDI.","author":"Hsieh Kevin","year":"2018","unstructured":"Kevin Hsieh, Ganesh Ananthanarayanan, Peter Bodik, Shivaram Venkataraman, Paramvir Bahl, Matthai Philipose, Phillip B Gibbons, and Onur Mutlu. 2018. Focus: Querying large video datasets with low latency and low cost. In OSDI."},{"key":"e_1_2_1_29_1","volume-title":"Seth Pettie, and Barzan Mozafari.","author":"Huang Dawei","year":"2019","unstructured":"Dawei Huang, Dong Young Yoon, Seth Pettie, and Barzan Mozafari. 2019. Joins on samples: a theoretical guide for practitioners. PVLDB (2019)."},{"key":"e_1_2_1_30_1","unstructured":"Lei Huang Weijiang Yu Weitao Ma Weihong Zhong Zhangyin Feng Haotian Wang Qianglong Chen Weihua Peng Xiaocheng Feng Bing Qin et al. 2023. A survey on hallucination in large language models: Principles taxonomy challenges and open questions. ACM Transactions on Information Systems (2023)."},{"key":"e_1_2_1_31_1","volume-title":"Smart: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization. In ACL.","author":"Jiang Haoming","year":"2020","unstructured":"Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Tuo Zhao. 2020. Smart: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization. In ACL."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3654989"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.14778\/3372716.3372725"},{"key":"e_1_2_1_34_1","volume-title":"NoScope: Optimizing Neural Network Queries over Video at Scale. PVLDB","author":"Kang Daniel","year":"2017","unstructured":"Daniel Kang, John Emmons, Firas Abuzaid, Peter Bailis, and Matei Zaharia. 2017. NoScope: Optimizing Neural Network Queries over Video at Scale. PVLDB (2017)."},{"key":"e_1_2_1_35_1","volume-title":"Approximate selection with guarantees using proxies. PVLDB","author":"Kang Daniel","year":"2020","unstructured":"Daniel Kang, Edward Gan, Peter Bailis, Tatsunori Hashimoto, and Matei Zaharia. 2020. Approximate selection with guarantees using proxies. PVLDB (2020)."},{"key":"e_1_2_1_36_1","volume-title":"Accelerating approximate aggregation queries with expensive predicates. PVLDB","author":"Kang Daniel","year":"2021","unstructured":"Daniel Kang, John Guibas, Peter Bailis, Tatsunori Hashimoto, Yi Sun, and Matei Zaharia. 2021. Accelerating approximate aggregation queries with expensive predicates. PVLDB (2021)."},{"key":"e_1_2_1_37_1","volume-title":"Colbert: Efficient and effective passage search via contextualized late interaction over bert. In SIGIR.","author":"Khattab Omar","year":"2020","unstructured":"Omar Khattab and Matei Zaharia. 2020. Colbert: Efficient and effective passage search via contextualized late interaction over bert. In SIGIR."},{"key":"e_1_2_1_38_1","unstructured":"Leslie Kish. 2011. Survey sampling."},{"key":"e_1_2_1_39_1","volume-title":"Bayesian estimates of equation system parameters: an application of integration by Monte Carlo. Econometrica: Journal of the Econometric Society","author":"Kloek Teun","year":"1978","unstructured":"Teun Kloek and Herman K Van Dijk. 1978. Bayesian estimates of equation system parameters: an application of integration by Monte Carlo. Econometrica: Journal of the Econometric Society (1978), 1-19."},{"key":"e_1_2_1_40_1","doi-asserted-by":"crossref","unstructured":"Pradap Konda Sanjib Das AnHai Doan Adel Ardalan Jeffrey R Ballard Han Li Fatemah Panahi Haojun Zhang Jeff Naughton Shishir Prasad et al. 2016. Magellan: toward building entity matching management systems over data science stacks. PVLDB (2016).","DOI":"10.14778\/3007263.3007314"},{"key":"e_1_2_1_41_1","unstructured":"Jiale Lao Andreas Zimmerer Olga Ovcharenko Tianji Cong Matthew Russo Gerardo Vitagliano Michael Cochez Fatma \u00d6zcan Gautam Gupta Thibaud Hottelier et al. 2025. SemBench: A Benchmark for Semantic Query Processing Engines. arXiv preprint arXiv:2511.01716 (2025)."},{"key":"e_1_2_1_42_1","volume-title":"NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models. arXiv preprint arXiv:2405.17428","author":"Lee Chankyu","year":"2024","unstructured":"Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2024. NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models. arXiv preprint arXiv:2405.17428 (2024)."},{"key":"e_1_2_1_43_1","volume-title":"DR-GAT: Dynamic routing graph attention network for stock recommendation. Information Sciences","author":"Lei Zengyu","year":"2024","unstructured":"Zengyu Lei, Caiming Zhang, Yunyang Xu, and Xuemei Li. 2024. DR-GAT: Dynamic routing graph attention network for stock recommendation. Information Sciences (2024)."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11390-020-0350-4"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3539597.3570379"},{"key":"e_1_2_1_46_1","unstructured":"Feifei Li Bin Wu Ke Yi and Zhuoyue Zhao. 2016. Wander join: Online aggregation for joins. In SIGMOD."},{"key":"e_1_2_1_47_1","volume-title":"Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML.","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML."},{"key":"e_1_2_1_48_1","volume-title":"Align before fuse: Vision and language representation learning with momentum distillation. Advances in neural information processing systems","author":"Li Junnan","year":"2021","unstructured":"Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021b. Align before fuse: Vision and language representation learning with momentum distillation. Advances in neural information processing systems, Vol. 34 (2021), 9694-9705."},{"key":"e_1_2_1_49_1","volume-title":"Deep entity matching with pre-trained language models. PVLDB","author":"Li Yuliang","year":"2020","unstructured":"Yuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan, and Wang-Chiew Tan. 2020a. Deep entity matching with pre-trained language models. PVLDB (2020)."},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3431816"},{"key":"e_1_2_1_51_1","volume-title":"Towards general text embeddings with multi-stage contrastive learning. arXiv preprint arXiv:2308.03281","author":"Li Zehan","year":"2023","unstructured":"Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. 2023b. Towards general text embeddings with multi-stage contrastive learning. arXiv preprint arXiv:2308.03281 (2023)."},{"key":"e_1_2_1_52_1","volume-title":"Optimizing llm queries in relational workloads. arXiv preprint arXiv:2403.05821","author":"Liu Shu","year":"2024","unstructured":"Shu Liu, Asim Biswal, Audrey Cheng, Xiangxi Mo, Shiyi Cao, Joseph E Gonzalez, Ion Stoica, and Matei Zaharia. 2024. Optimizing llm queries in relational workloads. arXiv preprint arXiv:2403.05821 (2024)."},{"key":"e_1_2_1_53_1","unstructured":"Xinchen Liu Wu Liu Tao Mei and Huadong Ma. 2016. A deep learning-based approach to progressive vehicle re-identification for urban surveillance. In ECCV."},{"key":"e_1_2_1_54_1","volume-title":"Provid: Progressive and multimodal vehicle reidentification for large-scale urban surveillance","author":"Liu Xinchen","year":"2017","unstructured":"Xinchen Liu, Wu Liu, Tao Mei, and Huadong Ma. 2017. Provid: Progressive and multimodal vehicle reidentification for large-scale urban surveillance. IEEE Transactions on Multimedia (2017)."},{"key":"e_1_2_1_55_1","doi-asserted-by":"crossref","unstructured":"Yao Lu Aakanksha Chowdhery Srikanth Kandula and Surajit Chaudhuri. 2018. Accelerating machine learning inference with probabilistic predicates. In SIGMOD.","DOI":"10.1145\/3183713.3183751"},{"key":"e_1_2_1_56_1","unstructured":"David Maier. 1983. The theory of relational databases. Computer science press Rockville."},{"key":"e_1_2_1_57_1","first-page":"440","article-title":"Learning blocking schemes for record linkage","volume":"6","author":"Michelson Matthew","year":"2006","unstructured":"Matthew Michelson and Craig A Knoblock. 2006. Learning blocking schemes for record linkage. In AAAI, Vol. 6. 440-445.","journal-title":"AAAI"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.5555\/1182635.1164207"},{"key":"e_1_2_1_59_1","doi-asserted-by":"crossref","unstructured":"Sidharth Mudgal Han Li Theodoros Rekatsinas AnHai Doan Youngchoon Park Ganesh Krishnan Rohit Deep Esteban Arcaute and Vijay Raghavendra. 2018. Deep learning for entity matching: A design space exploration. In SIGMOD.","DOI":"10.1145\/3183713.3196926"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.eacl-main.148"},{"key":"e_1_2_1_61_1","volume-title":"Can Foundation Models Wrangle Your Data? PVLDB","author":"Narayan Avanika","year":"2022","unstructured":"Avanika Narayan, Ines Chami, Laurel Orr, and Christopher R\u00e9. 2022. Can Foundation Models Wrangle Your Data? PVLDB (2022)."},{"key":"e_1_2_1_62_1","volume-title":"Passage Re-ranking with BERT. arXiv preprint arXiv:1901.04085","author":"Nogueira Rodrigo","year":"2019","unstructured":"Rodrigo Nogueira and Kyunghyun Cho. 2019. Passage Re-ranking with BERT. arXiv preprint arXiv:1901.04085 (2019)."},{"key":"e_1_2_1_63_1","unstructured":"OpenAI. [n.d.]. Pricing. https:\/\/openai.com\/api\/pricing\/ Accessed: 2024-10-22."},{"key":"e_1_2_1_64_1","unstructured":"OpenAI. 2024. New embedding models and API updates. https:\/\/openai.com\/index\/new-embedding-models-and-api-updates\/"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.14778\/3529337.3529356"},{"key":"e_1_2_1_66_1","volume-title":"Schema-agnostic vs schema-based configurations for blocking methods on homogeneous data. PVLDB","author":"Papadakis George","year":"2015","unstructured":"George Papadakis, George Alexiou, George Papastefanatos, and Georgia Koutrika. 2015. Schema-agnostic vs schema-based configurations for blocking methods on homogeneous data. PVLDB (2015)."},{"key":"e_1_2_1_67_1","doi-asserted-by":"crossref","first-page":"2665","DOI":"10.1109\/TKDE.2012.150","article-title":"A blocking framework for entity resolution in highly heterogeneous information spaces","volume":"25","author":"Papadakis George","year":"2012","unstructured":"George Papadakis, Ekaterini Ioannou, Themis Palpanas, Claudia Nieder\u00e9e, and Wolfgang Nejdl. 2012. A blocking framework for entity resolution in highly heterogeneous information spaces. IEEE Transactions on Knowledge and Data Engineering, Vol. 25, 12 (2012), 2665-2682.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"e_1_2_1_68_1","volume-title":"Blocking and filtering techniques for entity resolution: A survey. Comput. Surveys","author":"Papadakis George","year":"2020","unstructured":"George Papadakis, Dimitrios Skoutas, Emmanouil Thanos, and Themis Palpanas. 2020. Blocking and filtering techniques for entity resolution: A survey. Comput. Surveys (2020)."},{"key":"e_1_2_1_69_1","volume-title":"Semantic Operators: A Declarative Model for Rich, AI-based Data Processing. arXiv preprint arXiv:2407.11418","author":"Patel Liana","year":"2025","unstructured":"Liana Patel, Siddharth Jha, Carlos Guestrin, and Matei Zaharia. 2025. Semantic Operators: A Declarative Model for Rich, AI-based Data Processing. arXiv preprint arXiv:2407.11418 (2025)."},{"key":"e_1_2_1_70_1","volume-title":"28th International Conference on Extending Database Technology.","author":"Peeters Ralph","year":"2025","unstructured":"Ralph Peeters, Aaron Steiner, and Christian Bizer. 2025. Entity matching using large language models. In 28th International Conference on Extending Database Technology."},{"key":"e_1_2_1_71_1","volume-title":"WiC: the word-in-context dataset for evaluating context-sensitive meaning representations. arXiv preprint arXiv:1808.09121","author":"Pilehvar Mohammad Taher","year":"2018","unstructured":"Mohammad Taher Pilehvar and Jose Camacho-Collados. 2018. WiC: the word-in-context dataset for evaluating context-sensitive meaning representations. arXiv preprint arXiv:1808.09121 (2018)."},{"key":"e_1_2_1_72_1","volume-title":"Flickr30K Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models. ICCV","author":"Plummer Bryan A.","year":"2017","unstructured":"Bryan A. Plummer, Liwei Wang, Christopher M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2017. Flickr30K Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models. ICCV (2017)."},{"key":"e_1_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.1145\/1066157.1066224"},{"key":"e_1_2_1_74_1","volume-title":"International Conference on Extending Database Technology.","author":"Puhlmann Sven","year":"2006","unstructured":"Sven Puhlmann, Melanie Weis, and Felix Naumann. 2006. XML duplicate detection using sorted neighborhoods. In International Conference on Extending Database Technology."},{"key":"e_1_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00598"},{"key":"e_1_2_1_76_1","volume-title":"Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al.","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., 2021. Learning transferable visual models from natural language supervision. In ICML."},{"key":"e_1_2_1_77_1","unstructured":"Juan Ramos et al. 2003. Using tf-idf to determine word relevance in document queries. In ICML."},{"key":"e_1_2_1_78_1","doi-asserted-by":"crossref","unstructured":"Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In EMNLP.","DOI":"10.18653\/v1\/D19-1410"},{"key":"e_1_2_1_79_1","volume-title":"Caiming Xiong, Yingbo Zhou, and Semih Yavuz.","author":"Meng","year":"2024","unstructured":"Meng Rui*, Ye* Liu, Shafiq Rayhan Joty, Caiming Xiong, Yingbo Zhou, and Semih Yavuz. 2024. SFR-Embedding-2: Advanced Text Embedding with Multi-stage Training. https:\/\/huggingface.co\/Salesforce\/SFR-Embedding-2_R"},{"key":"e_1_2_1_80_1","volume-title":"Entity linking with a knowledge base: Issues, techniques, and solutions","author":"Shen Wei","year":"2014","unstructured":"Wei Shen, Jianyong Wang, and Jiawei Han. 2014. Entity linking with a knowledge base: Issues, techniques, and solutions. IEEE TKDE (2014)."},{"key":"e_1_2_1_81_1","volume-title":"Abdol Hossein Ahmadi, Reza Kafipour, and Kyle Albert Beattie.","author":"Shoyukhi Moohebat","year":"2023","unstructured":"Moohebat Shoyukhi, Paul Hubert Vossen, Abdol Hossein Ahmadi, Reza Kafipour, and Kyle Albert Beattie. 2023. Developing a comprehensive plagiarism assessment rubric. Education and Information Technologies (2023)."},{"key":"e_1_2_1_82_1","volume-title":"Improving Temporal Joins Using Histograms. In International Conference on Database and Expert Systems Applications.","author":"Sitzmann Inga","year":"2000","unstructured":"Inga Sitzmann and Peter J Stuckey. 2000. Improving Temporal Joins Using Histograms. In International Conference on Database and Expert Systems Applications."},{"key":"e_1_2_1_83_1","volume-title":"Proceedings of the 29th ACM International Conference on Information & Knowledge Management. 2241-2244","author":"Teong Kai-Sheng","year":"2020","unstructured":"Kai-Sheng Teong, Lay-Ki Soon, and Tin Tin Su. 2020. Schema-agnostic entity matching using pre-trained language models. In Proceedings of the 29th ACM International Conference on Information & Knowledge Management. 2241-2244."},{"key":"e_1_2_1_84_1","unstructured":"Nandan Thakur Nils Reimers Andreas R\u00fcckl\u00e9 Abhishek Srivastava and Iryna Gurevych. [n.d.]. BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. ([n.d.])."},{"key":"e_1_2_1_85_1","volume-title":"Deep learning for blocking in entity matching: a design space exploration. PVLDB","author":"Thirumuruganathan Saravanan","year":"2021","unstructured":"Saravanan Thirumuruganathan, Han Li, Nan Tang, Mourad Ouzzani, Yash Govind, Derek Paulsen, Glenn Fung, and AnHai Doan. 2021. Deep learning for blocking in entity matching: a design space exploration. PVLDB (2021)."},{"key":"e_1_2_1_86_1","volume-title":"Implementing Semantic Join Operators Efficiently. arXiv preprint arXiv:2510.08489","author":"Trummer Immanuel","year":"2025","unstructured":"Immanuel Trummer. 2025. Implementing Semantic Join Operators Efficiently. arXiv preprint arXiv:2510.08489 (2025)."},{"key":"e_1_2_1_87_1","volume-title":"Asymptotic statistics","author":"Van der Vaart Aad W","unstructured":"Aad W Van der Vaart. 2000. Asymptotic statistics. Vol. 3. Cambridge university press."},{"key":"e_1_2_1_88_1","first-page":"126500","article-title":"INQUIRE: A Natural World Text-to-Image Retrieval Benchmark","volume":"37","author":"Vendrow Edward","year":"2025","unstructured":"Edward Vendrow, Omiros Pantazis, Alexander Shepard, Gabriel Brostow, Kate Jones, Oisin Mac Aodha, Sara Beery, and Grant Van Horn. 2025. INQUIRE: A Natural World Text-to-Image Retrieval Benchmark. Advances in Neural Information Processing Systems, Vol. 37 (2025), 126500-126514.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_89_1","volume-title":"Mohamed Zait, and Sunil P Chakkappen.","author":"Vengerov David","year":"2015","unstructured":"David Vengerov, Andre Cavalheiro Menck, Mohamed Zait, and Sunil P Chakkappen. 2015. Join size estimation subject to filter conditions. PVLDB (2015)."},{"key":"e_1_2_1_90_1","volume-title":"Jeffrey Xu Yu, and Jianhua Feng","author":"Wang Jiannan","year":"2011","unstructured":"Jiannan Wang, Guoliang Li, Jeffrey Xu Yu, and Jianhua Feng. 2011. Entity matching: How similar is similar. PVLDB (2011)."},{"key":"e_1_2_1_91_1","volume-title":"Proceedings of the 2018 International Conference on Management of Data. 1113-1127","author":"Wang Sibo","year":"2018","unstructured":"Sibo Wang and Yufei Tao. 2018. Efficient algorithms for finding approximate heavy hitters in personalized pageranks. In Proceedings of the 2018 International Conference on Management of Data. 1113-1127."},{"key":"e_1_2_1_92_1","volume-title":"An Investigation of Large Language Models for Entity Matching. In International Conference on Computational Linguistics.","author":"Wang Tianshu","year":"2025","unstructured":"Tianshu Wang, Xiaoyang Chen, Hongyu Lin, Xuanang Chen, Xianpei Han, Hao Wang, Zhenyu Zeng, and Le Sun. 2025. Match, Compare, or Select? An Investigation of Large Language Models for Entity Matching. In International Conference on Computational Linguistics."},{"key":"e_1_2_1_93_1","doi-asserted-by":"publisher","DOI":"10.1145\/1086339.1086341"},{"key":"e_1_2_1_94_1","unstructured":"Shitao Xiao Zheng Liu Peitian Zhang and Niklas Muennighoff. 2023. C-Pack: Packaged Resources To Advance General Chinese Embedding. arXiv:2309.07597 [cs.CL]"},{"key":"e_1_2_1_95_1","volume-title":"Adaptive sorted neighborhood methods for efficient record linkage","author":"Yan Su","unstructured":"Su Yan, Dongwon Lee, Min-Yen Kan, and Lee C Giles. 2007. Adaptive sorted neighborhood methods for efficient record linkage. In ACM\/IEEE-CS joint conference on Digital libraries."},{"key":"e_1_2_1_96_1","doi-asserted-by":"publisher","DOI":"10.1145\/1559795.1559820"},{"key":"e_1_2_1_97_1","volume-title":"CVPR workshops.","author":"Zapletal Dominik","year":"2016","unstructured":"Dominik Zapletal and Adam Herout. 2016. Vehicle re-identification for automatic video traffic surveillance. In CVPR workshops."},{"key":"e_1_2_1_98_1","volume-title":"Similarity Analysis of Knowledge Graph-based Company Embedding for Stocks Portfolio. In IEEE International Conference on Smart Cloud.","author":"Zhang Boyao","year":"2021","unstructured":"Boyao Zhang, Zhongrui Li, Chao Yang, Zongguo Wang, Yonghua Zhao, Jingqi Sun, and Lihua Wang. 2021. Similarity Analysis of Knowledge Graph-based Company Embedding for Stocks Portfolio. In IEEE International Conference on Smart Cloud."},{"key":"e_1_2_1_99_1","volume-title":"Multi-factor duplicate question detection in stack overflow. Journal of Computer Science and Technology","author":"Zhang Yun","year":"2015","unstructured":"Yun Zhang, David Lo, Xin Xia, and Jian-Ling Sun. 2015. Multi-factor duplicate question detection in stack overflow. Journal of Computer Science and Technology (2015)."},{"key":"e_1_2_1_100_1","doi-asserted-by":"crossref","unstructured":"Zhuoyue Zhao Robert Christensen Feifei Li Xiao Hu and Ke Yi. 2018. Random sampling over joins revisited. In SIGMOD.","DOI":"10.1145\/3183713.3183739"},{"key":"e_1_2_1_101_1","doi-asserted-by":"crossref","unstructured":"Kaiyang Zhou Yongxin Yang Andrea Cavallaro and Tao Xiang. 2019. Omni-scale feature learning for person re-identification. In ICCV.","DOI":"10.1109\/ICCV.2019.00380"},{"key":"e_1_2_1_102_1","volume-title":"Accelerating Approximate Analytical Join Queries over Unstructured Data with Error Guarantees. arXiv preprint","author":"Zhu Yuxuan","year":"2026","unstructured":"Yuxuan Zhu, Tengjun Jin, Chenghao Mo, and Daniel Kang. 2026. Accelerating Approximate Analytical Join Queries over Unstructured Data with Error Guarantees. arXiv preprint (2026)."},{"key":"e_1_2_1_103_1","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3592047"},{"key":"e_1_2_1_104_1","doi-asserted-by":"publisher","DOI":"10.1162\/coli_a_00502"}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3802004","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T18:32:32Z","timestamp":1779129152000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3802004"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,18]]},"references-count":104,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,5,18]]}},"alternative-id":["10.1145\/3802004"],"URL":"https:\/\/doi.org\/10.1145\/3802004","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,18]]}}}