{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T21:08:58Z","timestamp":1775596138063,"version":"3.50.1"},"reference-count":114,"publisher":"Association for Computing Machinery (ACM)","issue":"1","funder":[{"name":"National Natural Science Foundation of China &#x28;NSFC&#x29;","award":["62232005"],"award-info":[{"award-number":["62232005"]}]},{"name":"National Natural Science Foundation of China &#x28;NSFC&#x29;","award":["62202126"],"award-info":[{"award-number":["62202126"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2026,4,2]]},"abstract":"<jats:p>\n                    Modern databases often use the\n                    <jats:italic toggle=\"yes\">LIKE<\/jats:italic>\n                    predicate to search text data. However, when the search condition is interrupted by wildcards, the existing search structure can degrade to a worst-case complexity linear to full table scale, resulting in poor performance. Traditional methods, such as B+-trees, fail to handle wildcards at both ends efficiently. Recent advances in language models offer a promising solution. These models can decode complex\n                    <jats:italic toggle=\"yes\">LIKE<\/jats:italic>\n                    patterns into a small set of candidate values, which are then verified in dataset-size-invariant time via hash table lookups, greatly improving efficiency. However, integrating LLMs into databases faces challenges such as high latency, large storage requirements, and sensitivity to data distribution drifts.\n                  <\/jats:p>\n                  <jats:p>\n                    To address these issues, we propose SMILE, a\n                    <jats:underline>S<\/jats:underline>\n                    mall language\n                    <jats:underline>M<\/jats:underline>\n                    odel\n                    <jats:underline>I<\/jats:underline>\n                    ntegrated\n                    <jats:underline>L<\/jats:underline>\n                    IKE\n                    <jats:underline>E<\/jats:underline>\n                    ngine that learns column-local character distributions through small but exquisite parameters. Our SMILE acts as a neural translator that converts complex LIKE patterns into their corresponding result sets. Our approach achieves asymptotic complexity improvements while preserving SQL LIKE logic. We conduct comprehensive evaluation across diverse datasets to validate the efficacy of our approach. Our compact SMILE, with a parameter size 5 orders of magnitude smaller than large language models, achieves strong LIKE decoding efficiency and quality. Specifically, SMILE obtains high recall ability while accelerating LIKE by 3 orders of magnitude compared to large language models and sequential scans, 1.8-41.6 times faster than trigram indexes, and 2 orders of magnitude faster than B+-trees. Moreover, our model demonstrates robustness against potential data and query distribution drifts.\n                  <\/jats:p>","DOI":"10.1145\/3786703","type":"journal-article","created":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T17:54:13Z","timestamp":1775584453000},"page":"1-30","source":"Crossref","is-referenced-by-count":0,"title":["The Case For Language Model Approximated LIKE Predicate"],"prefix":"10.1145","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-9035-7218","authenticated-orcid":false,"given":"Yingze","family":"Li","sequence":"first","affiliation":[{"name":"Harbin Institute of Technology, Harbin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-3755-6412","authenticated-orcid":false,"given":"Dong","family":"Wang","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Harbin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-3896-9847","authenticated-orcid":false,"given":"Zixuan","family":"Wang","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Harbin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-5630-6822","authenticated-orcid":false,"given":"Yingli","family":"Zhou","sequence":"additional","affiliation":[{"name":"The Chinese University of Hong Kong, Shenzhen, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2078-3343","authenticated-orcid":false,"given":"Yu","family":"Yan","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Harbin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-0964-585X","authenticated-orcid":false,"given":"Jian","family":"Geng","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Harbin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-9792-4309","authenticated-orcid":false,"given":"Xinyue","family":"Wang","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Harbin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-3930-0675","authenticated-orcid":false,"given":"Ziqing","family":"Zeng","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Harbin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7521-2871","authenticated-orcid":false,"given":"Hongzhi","family":"Wang","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Harbin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,7]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Information technology -- Database languages -- SQL -- Part 2: Foundation","unstructured":"2023. Information technology -- Database languages -- SQL -- Part 2: Foundation (SQL\/Foundation). https:\/\/webstore.iec.ch\/en\/publication\/86046"},{"key":"e_1_2_1_2_1","unstructured":"2025. morfessor. https:\/\/morfessor.readthedocs.io\/en\/latest\/ Accessed: 2025--10--31."},{"key":"e_1_2_1_3_1","unstructured":"2025. Pyphen. https:\/\/pyphen.org\/ Accessed: 2025--10--31."},{"key":"e_1_2_1_4_1","unstructured":"2025. SMILE: Source Code and Appendix. https:\/\/github.com\/LiYingZe\/SMILE. Accessed: 2025--12--29."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1465482.1465560"},{"key":"e_1_2_1_6_1","unstructured":"AutoDL. 2024. AutoDL GPU Marketplace. https:\/\/www.autodl.com\/market\/list Accessed: 2024-06--15."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3639309"},{"key":"e_1_2_1_8_1","unstructured":"Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2016. Neural Machine Translation by Jointly Learning to Align and Translate. (2016). arXiv:1409.0473 [cs.CL] https:\/\/arxiv.org\/abs\/1409.0473"},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the 32nd International Conference on Very Large Data Bases","author":"Bast Holger","year":"2006","unstructured":"Holger Bast, Debapriyo Majumdar, Ralf Schenkel, Martin Theobald, and Gerhard Weikum. 2006. IO-Top-k: index-access optimized top-k query processing. In Proceedings of the 32nd International Conference on Very Large Data Bases (Seoul, Korea) (VLDB '06). VLDB Endowment, 475--486."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2009.32"},{"key":"e_1_2_1_11_1","unstructured":"Rishi Bommasani Drew A. Hudson Ehsan Adeli Russ Altman Simran Arora Sydney von Arx Michael S. Bernstein Jeannette Bohg Antoine Bosselut Emma Brunskill Erik Brynjolfsson Shyamal Buch Dallas Card Rodrigo Castellon Niladri Chatterji Annie Chen Kathleen Creel Jared Quincy Davis Dora Demszky Chris Donahue Moussa Doumbouya Esin Durmus Stefano Ermon John Etchemendy Kawin Ethayarajh Li Fei-Fei Chelsea Finn Trevor Gale Lauren Gillespie Karan Goel Noah Goodman Shelby Grossman Neel Guha Tatsunori Hashimoto Peter Henderson John Hewitt Daniel E. Ho Jenny Hong Kyle Hsu Jing Huang Thomas Icard Saahil Jain Dan Jurafsky Pratyusha Kalluri Siddharth Karamcheti Geoff Keeling Fereshte Khani Omar Khattab Pang Wei Koh Mark Krass Ranjay Krishna Rohith Kuditipudi Ananya Kumar Faisal Ladhak Mina Lee Tony Lee Jure Leskovec Isabelle Levent Xiang Lisa Li Xuechen Li Tengyu Ma Ali Malik Christopher D. Manning Suvir Mirchandani Eric Mitchell Zanele Munyikwa Suraj Nair Avanika Narayan Deepak Narayanan Ben Newman Allen Nie Juan Carlos Niebles Hamed Nilforoshan Julian Nyarko Giray Ogut Laurel Orr Isabel Papadimitriou Joon Sung Park Chris Piech Eva Portelance Christopher Potts Aditi Raghunathan Rob Reich Hongyu Ren Frieda Rong Yusuf Roohani Camilo Ruiz Jack Ryan Christopher R\u00e9 Dorsa Sadigh Shiori Sagawa Keshav Santhanam Andy Shih Krishnan Srinivasan Alex Tamkin Rohan Taori Armin W. Thomas Florian Tram\u00e8r Rose E. Wang William Wang Bohan Wu Jiajun Wu Yuhuai Wu Sang Michael Xie Michihiro Yasunaga Jiaxuan You Matei Zaharia Michael Zhang Tianyi Zhang Xikun Zhang Yuhui Zhang Lucia Zheng Kaitlyn Zhou and Percy Liang. 2022. On the Opportunities and Risks of Foundation Models. (2022). arXiv:2108.07258 [cs.LG] https:\/\/arxiv.org\/abs\/2108.07258"},{"key":"e_1_2_1_12_1","doi-asserted-by":"crossref","unstructured":"Felix Brauer Ralf Rieger Adrian Mocan Craig A. Knoblock and Pedro Szekely. 2011. Enabling Information Extraction by Inference of Regular Expressions from Sample Entities. In CIKM. 1285--1290.","DOI":"10.1145\/2063576.2063763"},{"key":"e_1_2_1_13_1","unstructured":"Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et al. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165 (2020)."},{"key":"e_1_2_1_14_1","doi-asserted-by":"crossref","unstructured":"Chen Chai Lei Cao Guoliang Li Samuel Madden and Nan Tang. 2020. Human-in-the-loop Outlier Detection. In SIGMOD. 19--33.","DOI":"10.1145\/3318464.3389772"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.5555\/645926.671851"},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of the 2003 ACM SIGMOD International Conference on Management of Data (SIGMOD). 313--324","author":"Chaudhuri S.","unstructured":"S. Chaudhuri, K. Ganjam, V. Ganti, and R. Motwani. 2003. Robust and Efficient Fuzzy Match for Online Data Cleaning. In Proceedings of the 2003 ACM SIGMOD International Conference on Management of Data (SIGMOD). 313--324."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.5555\/645925.671359"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/1559845.1559868"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3166054.3166058"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3622863"},{"key":"e_1_2_1_21_1","unstructured":"Cherry Servers. 2024. AMD EPYC 9654 Dedicated Server Pricing. https:\/\/www.cherryservers.com\/pricing\/dedicated-servers\/amd-epyc-9654?billing=3\u00aeion=lt-siauliai"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2463676.2465327"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1457838.1457895"},{"key":"e_1_2_1_24_1","unstructured":"DeepSeek-AI. 2024. DeepSeek-V2: A Strong Economical and Efficient Mixture-of-Experts Language Model. arXiv:2405.04434 [cs.CL]"},{"key":"e_1_2_1_25_1","unstructured":"DeepSeek-AI. 2025. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv:2501.12948 [cs.CL] https:\/\/arxiv.org\/abs\/2501.12948"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2012.29"},{"key":"e_1_2_1_27_1","volume-title":"BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin, Michael-W Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.14778\/3565838.3565857"},{"key":"e_1_2_1_29_1","unstructured":"PostgreSQL Documentation. [n.d.]. F.33. pg_trgm \u2014 support for similarity of text using trigram matching. https:\/\/www.postgresql.org\/docs\/current\/pgtrgm.html. Accessed: 2025-03-07."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.emnlp-main.64"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.14778\/1687627.1687667"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/s40474-018-0153--2"},{"key":"e_1_2_1_33_1","unstructured":"DBA Stack Exchange. [n.d.]. LIKE Query Optimization. https:\/\/dba.stackexchange.com\/questions\/203748\/like-query-optimization. Accessed: 2025-03-07."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.14778\/3681954.3681960"},{"key":"e_1_2_1_35_1","first-page":"104","article-title":"Human-in-the-loop Rule Learning for Data Integration","volume":"41","author":"Fan Ju","year":"2018","unstructured":"Ju Fan and Guoliang Li. 2018. Human-in-the-loop Rule Learning for Data Integration. IEEE Data Eng. Bull. 41, 2 (2018), 104--115.","journal-title":"IEEE Data Eng. Bull."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","unstructured":"Wenfei Fan Jianzhong Li Shuai Ma Nan Tang and Wenyuan Yu. 2011. Interaction between record matching and data repairing. (2011) 469--480. doi:10.1145\/1989323.1989373","DOI":"10.1145\/1989323.1989373"},{"key":"e_1_2_1_37_1","doi-asserted-by":"crossref","unstructured":"Zhangyin Feng Daya Guo Duyu Tang Nan Duan Xiaocheng Feng Ming Gong Linjun Shou Bing Qin Ting Liu Daxin Jiang and Ming Zhou. 2020. CodeBERT: A Pre-Trained Model for Programming and Natural Languages. (2020). arXiv:2002.08155 [cs.CL] https:\/\/arxiv.org\/abs\/2002.08155","DOI":"10.18653\/v1\/2020.findings-emnlp.139"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/301970.301973"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1016\/S1364-6613(99)01294-2"},{"key":"e_1_2_1_40_1","volume-title":"An Introduction to Language. Wadsworth Cengage Learning","author":"Fromkin Victoria","unstructured":"Victoria Fromkin, Robert Rodman, and Nina Hyams. [n.d.]. An Introduction to Language. Wadsworth Cengage Learning, Boston."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3319860"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.99"},{"key":"e_1_2_1_43_1","volume-title":"Proceedings of the 27th International Conference on Very Large Data Bases (VLDB). 491--500","author":"Gravano L.","unstructured":"L. Gravano, P. G. Ipeirotis, H. V. Jagadish, N. Koudas, S. Muthukrishnan, and D. Srivastava. 2001. Approximate String Joins in a Database (Almost) for Free. In Proceedings of the 27th International Conference on Very Large Data Bases (VLDB). 491--500."},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18--1065"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.14778\/3503585.3503586"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3543873.3584633"},{"key":"e_1_2_1_47_1","volume-title":"IBM Db2 13 for z\/OS SQL Reference","author":"IBM Corporation 2023.","unstructured":"IBM Corporation 2023. IBM Db2 13 for z\/OS SQL Reference. IBM Corporation. Sec. ''LIKE predicate'', https:\/\/www.ibm.com\/docs\/en\/SSEPEK_13.0.0\/pdf\/db2z_13_sqlrefbook.pdf."},{"key":"e_1_2_1_48_1","unstructured":"IMDB. [n.d.]. Latest IMDB-Name(2025). https:\/\/datasets.imdbws.com\/. Accessed: 2025--10--21."},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/1526709.1526718"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.5555\/1083592.1083640"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00065"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2013.04.037"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/2463676.2465324"},{"key":"e_1_2_1_54_1","volume-title":"Learned Cardinalities: Estimating Correlated Joins with Deep Learning.","author":"Kipf Andreas","year":"2018","unstructured":"Andreas Kipf, Thomas Kipf, Bernhard Radke, Viktor Leis, Peter Boncz, and Alfons Kemper. 2018. Learned Cardinalities: Estimating Correlated Joins with Deep Learning. (2018). arXiv:1809.00677 [cs.DB] https:\/\/arxiv.org\/abs\/1809.00677"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.1611835114"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3196909"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3389752"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","unstructured":"Toshitaka Kuwa Shigehiko Schamoni and Stefan Riezler. 2020. Embedding Meta-Textual Information for Improved Learning to Rank. In Proceedings of the 28th International Conference on Computational Linguistics Donia Scott Nuria Bel and Chengqing Zong (Eds.). International Committee on Computational Linguistics Barcelona Spain (Online) 5558--5568. doi:10.18653\/v1\/2020.coling-main.487","DOI":"10.18653\/v1\/2020.coling-main.487"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/3709670"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/1516360.1516455"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/1066157.1066173"},{"key":"e_1_2_1_62_1","volume-title":"Proceedings of the 33rd International Conference on Very Large Data Bases","author":"Li Chen","year":"2007","unstructured":"Chen Li, Bin Wang, and Xiaochun Yang. 2007. VGRAM: improving performance of approximate queries on string collections using variable-length grams. In Proceedings of the 33rd International Conference on Very Large Data Bases (Vienna, Austria) (VLDB '07). VLDB Endowment, 303--314."},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.14778\/3137765.3137833"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/1989323.1989379"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00778-011-0218-x"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.14778\/3626292.3626302"},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1145\/3639270"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.14778\/2876473.2876478"},{"key":"e_1_2_1_69_1","volume-title":"Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692 (2019)."},{"key":"e_1_2_1_70_1","volume-title":"Magneto: Combining Small and Large Language Models for Schema Matching.","author":"Liu Yurong","year":"2025","unstructured":"Yurong Liu, Eduardo Pena, Aecio Santos, Eden Wu, and Juliana Freire. 2025. Magneto: Combining Small and Large Language Models for Schema Matching. (2025). arXiv:2412.08194 [cs.DB] https:\/\/arxiv.org\/abs\/2412.08194"},{"key":"e_1_2_1_71_1","unstructured":"Microsoft 2023. LIKE (Transact-SQL). Microsoft. https:\/\/learn.microsoft.com\/en-us\/sql\/t-sql\/language-elements\/like-transact-sql?view=sql-server-ver17."},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.201"},{"key":"e_1_2_1_73_1","unstructured":"Colin Morris. 2017. Reddit Usernames. https:\/\/www.kaggle.com\/datasets\/colinmorris\/reddit-usernames. Accessed: 2025-03--18."},{"key":"e_1_2_1_74_1","unstructured":"John X. Morris Chawin Sitawarin Chuan Guo Narine Kokhlikyan G. Edward Suh Alexander M. Rush et al. 2025. How much do language models memorize? arXiv preprint arXiv:2505.24832 (2025). https:\/\/arxiv.org\/abs\/2505.24832"},{"key":"e_1_2_1_75_1","unstructured":"Brent Ozar. [n.d.]. Sargability: Why %string% Is Slow. https:\/\/www.brentozar.com\/archive\/2010\/06\/sargable-why-string-is-slow\/. Accessed: 2025-03-07."},{"key":"e_1_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1145\/3183713.3183747"},{"key":"e_1_2_1_77_1","unstructured":"PostgreSQL Core Team. 2025. GIN Indexes. https:\/\/www.postgresql.org\/docs\/current\/gin.html. Accessed: 2025-03--18."},{"key":"e_1_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.14778\/3377369.3377377"},{"key":"e_1_2_1_79_1","unstructured":"Alec Radford Karthik Narasimhan Tim Salimans and Ilya Sutskever. 2018. Improving language understanding by generative pre-training. (2018)."},{"key":"e_1_2_1_80_1","volume-title":"Language models are unsupervised multitask learners. OpenAI blog 1, 8","author":"Radford Alec","year":"2019","unstructured":"Alec Radford, Jeffrey Wu, Rewon Child, David Chen, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019)."},{"key":"e_1_2_1_81_1","volume-title":"Hellerstein","author":"Raman Vijayshankar","year":"2001","unstructured":"Vijayshankar Raman and Joseph M. Hellerstein. 2001. Potter's Wheel: An Interactive Data Cleaning System. In VLDB. 381--390."},{"key":"e_1_2_1_82_1","volume-title":"Experience Replay for Continual Learning. NeurIPS 32","author":"Rolnick David","year":"2019","unstructured":"David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy P. Lillicrap, and Gregory Wayne. 2019. Experience Replay for Continual Learning. NeurIPS 32 (2019)."},{"key":"e_1_2_1_83_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cognition.2017.11.003"},{"key":"e_1_2_1_84_1","doi-asserted-by":"publisher","DOI":"10.14778\/3570690.3570702"},{"key":"e_1_2_1_85_1","volume-title":"SE 2023. SAP IQ 16","author":"SAP","unstructured":"SAP SE 2023. SAP IQ 16.0 SP04 Reference: Building Blocks. SAP SE. PDF p. 80, https:\/\/infocenter.sybase.com\/help\/topic\/com.sybase.infocenter.dc38151.1604\/doc\/pdf\/iqrefbb.pdf."},{"key":"e_1_2_1_86_1","doi-asserted-by":"publisher","DOI":"10.1145\/3555041.3589677"},{"key":"e_1_2_1_87_1","doi-asserted-by":"publisher","DOI":"10.1145\/3318464.3380595"},{"key":"e_1_2_1_88_1","doi-asserted-by":"publisher","DOI":"10.14778\/3436905.3436907"},{"key":"e_1_2_1_89_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.230"},{"key":"e_1_2_1_90_1","unstructured":"Charlie Snell Jaehoon Lee Kelvin Xu and Aviral Kumar. 2024. Scaling LLM Test-Time Compute Optimally Can Be More Effective Than Scaling Model Parameters. arXiv:2408.03314 [cs.CL]"},{"key":"e_1_2_1_91_1","volume-title":"Umar Farooq Minhas, and Tim Kraska","author":"Spector Benjamin","year":"2021","unstructured":"Benjamin Spector, Andreas Kipf, Kapil Vaidya, Chi Wang, Umar Farooq Minhas, and Tim Kraska. 2021. Bounding the Last Mile: Efficient Learned String Indexing. In AIDB."},{"key":"e_1_2_1_92_1","volume-title":"Proceedings of the 28th International Conference on Neural Information Processing Systems -","volume":"2","author":"Sutskever Ilya","unstructured":"Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to sequence learning with neural networks. In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 2 (Montreal, Canada) (NIPS'14). MIT Press, Cambridge, MA, USA, 3104--3112."},{"key":"e_1_2_1_93_1","unstructured":"DeepSeek Team. 2024. DeepSeek-V3 Technical Report. https:\/\/arxiv.org\/html\/2412.19437v1 arXiv preprint arXiv:2412.19437v1."},{"key":"e_1_2_1_94_1","unstructured":"Qwen Team. 2025. Qwen2.5 Technical Report. Technical Report 2412.15115. arXiv. https:\/\/arxiv.org\/pdf\/2412.15115"},{"key":"e_1_2_1_95_1","unstructured":"Together Computer. 2023. RedPajama-Data-1T-Sample: A 1.2 B-token sample of the RedPajama pre-training corpus. https:\/\/huggingface.co\/datasets\/togethercomputer\/RedPajama-Data-1T-Sample Hugging Face dataset snapshot accessed 2025--10--24."},{"key":"e_1_2_1_96_1","unstructured":"Ask TOM. [n.d.]. Using the LIKE predicate in a WHERE clause. https:\/\/asktom.oracle.com\/ords\/asktom.search?tag=using-the-like-predicate-in-a-where-clause. Accessed: 2025-03-07."},{"key":"e_1_2_1_97_1","unstructured":"TPC. [n.d.]. TPC-H Homepage. https:\/\/www.tpc.org\/tpch\/. Accessed: 2025-03--18."},{"key":"e_1_2_1_98_1","doi-asserted-by":"publisher","DOI":"10.14778\/3529337.3529347"},{"key":"e_1_2_1_99_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_2_1_100_1","doi-asserted-by":"publisher","DOI":"10.1145\/3409963.3410496"},{"key":"e_1_2_1_101_1","unstructured":"WIKI. [n.d.]. Latest WIKI-dump(2025). https:\/\/dumps.wikimedia.org\/enwiki\/. Accessed: 2025--10--21."},{"key":"e_1_2_1_102_1","doi-asserted-by":"publisher","DOI":"10.14778\/3457390.3457393"},{"key":"e_1_2_1_103_1","volume-title":"The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=iLUcsecZJp","author":"Wu Shiguang","year":"2025","unstructured":"Shiguang Wu, Yaqing Wang, and Quanming Yao. 2025. Why In-Context Learning Models are Good Few-Shot Learners?. In The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=iLUcsecZJp"},{"key":"e_1_2_1_104_1","doi-asserted-by":"publisher","DOI":"10.1145\/3626246.3653391"},{"key":"e_1_2_1_105_1","doi-asserted-by":"publisher","DOI":"10.1145\/2000824.2000825"},{"key":"e_1_2_1_106_1","doi-asserted-by":"publisher","DOI":"10.1145\/1376616.1376655"},{"key":"e_1_2_1_107_1","doi-asserted-by":"publisher","DOI":"10.14778\/3681954.3682010"},{"key":"e_1_2_1_108_1","doi-asserted-by":"publisher","DOI":"10.14778\/3421424.3421432"},{"key":"e_1_2_1_109_1","doi-asserted-by":"publisher","DOI":"10.14778\/3368289.3368294"},{"key":"e_1_2_1_110_1","doi-asserted-by":"crossref","unstructured":"Shaoxiong Yu Li Han Marta Indulska and Hongzhi Liu. 2023. Human-in-the-loop Regular Expression Extraction for Single Column Format Inconsistency. In WWW. 3859--3867.","DOI":"10.1145\/3543507.3583515"},{"key":"e_1_2_1_111_1","unstructured":"Yicheng Zhang Yuhao Wang Ziyu Zhang Yiming Zhang Yizhe Zhang Haoyang Zhang Shengyu Zhang Yifan Zhang Yuchen Zhang Yuxuan Zhang et al. 2023. Math Problem Solving with Large Language Models: A Survey. arXiv preprint arXiv:2305.00772 (2023)."},{"key":"e_1_2_1_112_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i5.16602"},{"key":"e_1_2_1_113_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i1.27798"},{"key":"e_1_2_1_114_1","doi-asserted-by":"publisher","DOI":"10.1145\/1132956.1132959"}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3786703","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T20:02:40Z","timestamp":1775592160000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3786703"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,2]]},"references-count":114,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,4,2]]}},"alternative-id":["10.1145\/3786703"],"URL":"https:\/\/doi.org\/10.1145\/3786703","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,2]]}}}