{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T19:48:30Z","timestamp":1774986510683,"version":"3.50.1"},"reference-count":100,"publisher":"Association for Computing Machinery (ACM)","issue":"3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2025,6,17]]},"abstract":"<jats:p>\n                    Number of Distinct Values (NDV) estimation of a multiset\/column is a basis for many data management tasks, especially within databases. Despite decades of research, most existing methods require either a significant amount of samples through uniform random sampling or access to the entire column to produce estimates, leading to substantial data access costs and potentially ineffective estimations in scenarios with limited data access. In this paper, we propose leveraging semantic information, i.e., schema, to address these challenges. The schema contains rich semantic information that can benefit the NDV estimation. To this end, we propose\n                    <jats:italic toggle=\"yes\">PLM4NDV,<\/jats:italic>\n                    a learned method incorporating Pre-trained Language Models (PLMs) to extract semantic schema information for NDV estimation. Specifically,\n                    <jats:italic toggle=\"yes\">PLM4NDV<\/jats:italic>\n                    leverages the semantics of the target column and the corresponding table to gain a comprehensive understanding of the column's meaning. By using the semantics,\n                    <jats:italic toggle=\"yes\">PLM4NDV<\/jats:italic>\n                    reduces data access costs, provides accurate NDV estimation, and can even operate effectively without any data access. Extensive experiments on a large-scale real-world dataset demonstrate the superiority of\n                    <jats:italic toggle=\"yes\">PLM4NDV<\/jats:italic>\n                    over baseline methods. Our code is available at https:\/\/github.com\/bytedance\/plm4ndv.\n                  <\/jats:p>","DOI":"10.1145\/3725336","type":"journal-article","created":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T21:23:29Z","timestamp":1750281809000},"page":"1-28","source":"Crossref","is-referenced-by-count":1,"title":["PLM4NDV: Minimizing Data Access for Number of Distinct Values Estimation with Pre-trained Language Models"],"prefix":"10.1145","volume":"3","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2447-4107","authenticated-orcid":false,"given":"Xianghong","family":"Xu","sequence":"first","affiliation":[{"name":"ByteDance, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7959-2157","authenticated-orcid":false,"given":"Xiao","family":"He","sequence":"additional","affiliation":[{"name":"ByteDance, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-2250-5528","authenticated-orcid":false,"given":"Tieying","family":"Zhang","sequence":"additional","affiliation":[{"name":"Bytedance, San Jose, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-1681-1956","authenticated-orcid":false,"given":"Lei","family":"Zhang","sequence":"additional","affiliation":[{"name":"ByteDance, San Jose, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-9122-4703","authenticated-orcid":false,"given":"Rui","family":"Shi","sequence":"additional","affiliation":[{"name":"ByteDance, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3734-892X","authenticated-orcid":false,"given":"Jianjun","family":"Chen","sequence":"additional","affiliation":[{"name":"ByteDance, San Jose, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,6,18]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"2024. Llama 3.1 model. https:\/\/llama.meta.com\/"},{"key":"e_1_2_2_2_1","unstructured":"2024. OpenAI ChatGPT. https:\/\/openai.com\/chatgpt\/"},{"key":"e_1_2_2_3_1","unstructured":"2024. OpenAI GPT-4o. https:\/\/openai.com\/index\/hello-gpt-4o\/"},{"key":"e_1_2_2_4_1","unstructured":"2024. Pydistinct - Population Distinct Value Estimators. https:\/\/pydistinct.readthedocs.io\/"},{"key":"e_1_2_2_5_1","unstructured":"2024. Source Code of MySQL. https:\/\/github.com\/mysql\/mysql-server\/blob\/ 824e2b4064053f7daf17d7f3f84b7a3ed92e5fb4\/sql\/join_optimizer\/cost_model.cc"},{"key":"e_1_2_2_6_1","unstructured":"2024. Source Code of PostgreSQL. https:\/\/github.com\/postgres\/postgres\/blob\/master\/src\/backend\/optimizer\/plan\/analyzejoins.c"},{"key":"e_1_2_2_7_1","unstructured":"2024. Source Code of Spark. https:\/\/github.com\/apache\/spark\/blob\/master\/sql\/catalyst\/src\/main\/scala\/org\/apache\/spark\/sql\/catalyst\/plans\/logical\/statsEstimation\/JoinEstimation.scala"},{"key":"e_1_2_2_8_1","unstructured":"2024. tablib-v1-sample dataset. https:\/\/huggingface.co\/datasets\/approximatelabs\/tablib-v1-sample"},{"key":"e_1_2_2_9_1","volume-title":"International Conference on Machine Learning. PMLR, 11--21","author":"Acharya Jayadev","year":"2017","unstructured":"Jayadev Acharya, Hirakendu Das, Alon Orlitsky, and Ananda Theertha Suresh. 2017. A unified maximum likelihood approach for estimating symmetric properties of discrete distributions. In International Conference on Machine Learning. PMLR, 11--21."},{"key":"e_1_2_2_10_1","volume-title":"Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al.","author":"Achiam Josh","year":"2023","unstructured":"Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)."},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/237814.237823"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1007\/3-540-45726-7_1"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1247480.1247504"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-76153-9_28"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1198\/106186002760180572"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.1993.10594330"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/65.3.625"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.2307\/1936861"},{"key":"e_1_2_2_19_1","volume-title":"Nonparametric estimation of the number of classes in a population. Scandinavian Journal of statistics","author":"Chao Anne","year":"1984","unstructured":"Anne Chao. 1984. Nonparametric estimation of the number of classes in a population. Scandinavian Journal of statistics (1984), 265--270."},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.1992.10475194"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/335168.335230"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1007568.1007602"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/276305.276343"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3534678.3539377"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-21042-1_29"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1006\/jcss.1997.1534"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNET.2019.2940705"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3542700.3542709"},{"key":"e_1_2_2_29_1","unstructured":"Ning Ding Yujia Qin Guang Yang Fuchao Wei Zonghan Yang Yusheng Su Shengding Hu Yulin Chen Chi-Min Chan Weize Chen et al. 2022. Delta tuning: A comprehensive study of parameter efficient methods for pre-trained language models. arXiv preprint arXiv:2203.06904 (2022)."},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1198\/106186006X142744"},{"key":"e_1_2_2_31_1","volume-title":"Proceedings 11","author":"Durand Marianne","year":"2003","unstructured":"Marianne Durand and Philippe Flajolet. 2003. Loglog counting of large cardinalities. In Algorithms-ESA 2003: 11th Annual European Symposium, Budapest, Hungary, September 16--19, 2003. Proceedings 11. Springer, 605--617."},{"key":"e_1_2_2_32_1","unstructured":"Gus Eggert Kevin Huo Mike Biven and Justin Waugh. 2023. TabLib: A Dataset of 627M Tables with Context. arXiv:2310.07875 [cs.CL]"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.14778\/3654621.3654632"},{"key":"e_1_2_2_34_1","volume-title":"Hyperloglog: the analysis of a near-optimal cardinality estimation algorithm. Discrete mathematics & theoretical computer science Proceedings","author":"Flajolet Philippe","year":"2007","unstructured":"Philippe Flajolet, \u00c9ric Fusy, Olivier Gandouet, and Fr\u00e9d\u00e9ric Meunier. 2007. Hyperloglog: the analysis of a near-optimal cardinality estimation algorithm. Discrete mathematics & theoretical computer science Proceedings (2007)."},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1016\/0022-0000(85)90041-8"},{"key":"e_1_2_2_36_1","volume-title":"Text-to-sql empowered by large language models: A benchmark evaluation. arXiv preprint arXiv:2308.15363","author":"Gao Dawei","year":"2023","unstructured":"Dawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun, Yichen Qian, Bolin Ding, and Jingren Zhou. 2023. Text-to-sql empowered by large language models: A benchmark evaluation. arXiv preprint arXiv:2308.15363 (2023)."},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.dam.2008.06.020"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177729949"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3589292"},{"key":"e_1_2_2_40_1","volume-title":"Interleaving pre-trained language models and large language models for zero-shot nl2sql generation. arXiv preprint arXiv:2306.08891","author":"Gu Zihui","year":"2023","unstructured":"Zihui Gu, Ju Fan, Nan Tang, Songyue Zhang, Yuxin Zhang, Zui Chen, Lei Cao, Guoliang Li, Sam Madden, and Xiaoyong Du. 2023. Interleaving pre-trained language models and large language models for zero-shot nl2sql generation. arXiv preprint arXiv:2306.08891 (2023)."},{"key":"e_1_2_2_41_1","first-page":"311","article-title":"Sampling-based estimation of the number of distinct values of an attribute","volume":"95","author":"Haas Peter J","year":"1995","unstructured":"Peter J Haas, Jeffrey F Naughton, S Seshadri, and Lynne Stokes. 1995. Sampling-based estimation of the number of distinct values of an attribute. In VLDB, Vol. 95. 311--322.","journal-title":"VLDB"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.aiopen.2021.08.002"},{"key":"e_1_2_2_43_1","volume-title":"ByteCard: Enhancing Data Warehousing with Learned Cardinality Estimation. arXiv preprint arXiv:2403.16110","author":"Han Yuxing","year":"2024","unstructured":"Yuxing Han, Haoyu Wang, Lixiang Chen, Yifeng Dong, Xing Chen, Benquan Yu, Chengcheng Yang, and Weining Qian. 2024. ByteCard: Enhancing Data Warehousing with Learned Cardinality Estimation. arXiv preprint arXiv:2403.16110 (2024)."},{"key":"e_1_2_2_44_1","volume-title":"Unified sample-optimal property estimation in near-linear time. Advances in Neural Information Processing Systems 32","author":"Hao Yi","year":"2019","unstructured":"Yi Hao and Alon Orlitsky. 2019. Unified sample-optimal property estimation in near-linear time. Advances in Neural Information Processing Systems 32 (2019)."},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3186728.3164145"},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE55515.2023.00068"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/2452376.2452456"},{"key":"e_1_2_2_48_1","volume-title":"LLMTune: Accelerate Database Knob Tuning with Large Language Models. arXiv preprint arXiv:2404.11581","author":"Huang Xinmei","year":"2024","unstructured":"Xinmei Huang, Haoyang Li, Jing Zhang, Xinxin Zhao, Zhiming Yao, Yiyan Li, Zhuohao Yu, Tieying Zhang, Hong Chen, and Cuiping Li. 2024. LLMTune: Accelerate Database Knob Tuning with Large Language Models. arXiv preprint arXiv:2404.11581 (2024)."},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3588710"},{"key":"e_1_2_2_50_1","volume-title":"Proceedings of NAACL-HLT. 4171--4186","author":"Ming-Wei Chang Jacob Devlin","year":"2019","unstructured":"Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL-HLT. 4171--4186."},{"key":"e_1_2_2_51_1","volume-title":"Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7--9, 2015, Conference Track Proceedings.","author":"Diederik","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7--9, 2015, Conference Track Proceedings."},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3654994"},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.14778\/3626292.3626302"},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.14778\/3421424.3421431"},{"key":"e_1_2_2_55_1","volume-title":"Query Rewriting via Large Language Models. arXiv preprint arXiv:2403.09060","author":"Liu Jie","year":"2024","unstructured":"Jie Liu and Barzan Mozafari. 2024. Query Rewriting via Large Language Models. arXiv preprint arXiv:2403.09060 (2024)."},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3560815"},{"key":"e_1_2_2_57_1","volume-title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR abs\/1907.11692","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR abs\/1907.11692 (2019). arXiv:1907.11692 http:\/\/arxiv.org\/abs\/1907.11692"},{"key":"e_1_2_2_58_1","doi-asserted-by":"publisher","DOI":"10.14778\/1687627.1687738"},{"key":"e_1_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611972962.7"},{"key":"e_1_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/1340771.1340773"},{"key":"e_1_2_2_61_1","volume-title":"Noah Constant, Ji Ma, Keith B Hall, Daniel Cer, and Yinfei Yang.","author":"Ni Jianmo","year":"2021","unstructured":"Jianmo Ni, Gustavo Hernandez Abrego, Noah Constant, Ji Ma, Keith B Hall, Daniel Cer, and Yinfei Yang. 2021. Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models. arXiv preprint arXiv:2108.08877 (2021)."},{"key":"e_1_2_2_62_1","unstructured":"Frank Olken. 1993. Random sampling from databases. Ph. D. Dissertation. Citeseer."},{"key":"e_1_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.5555\/1036843.1036895"},{"key":"e_1_2_2_64_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-7091-7555-2_68"},{"key":"e_1_2_2_65_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10619-010-7067-2"},{"key":"e_1_2_2_66_1","first-page":"1","article-title":"Approximate profile maximum likelihood","volume":"20","author":"Pavlichin Dmitri S","year":"2019","unstructured":"Dmitri S Pavlichin, Jiantao Jiao, and Tsachy Weissman. 2019. Approximate profile maximum likelihood. Journal of Machine Learning Research 20, 122 (2019), 1--55.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_2_2_67_1","doi-asserted-by":"publisher","DOI":"10.1145\/3583780.3614731"},{"key":"e_1_2_2_68_1","volume-title":"Online Index Recommendation for Slow Queries. In 2024 IEEE 40th International Conference on Data Engineering (ICDE). IEEE, 5294--5306","author":"Peng Gan","year":"2024","unstructured":"Gan Peng, Peng Cai, Kaikai Ye, Kai Li, Jinlong Cai, Yufeng Shen, Han Su, and Weiyuan Xu. 2024. Online Index Recommendation for Slow Queries. In 2024 IEEE 40th International Conference on Data Engineering (ICDE). IEEE, 5294--5306."},{"key":"e_1_2_2_69_1","volume-title":"Din-sql: Decomposed in-context learning of text-to-sql with self-correction. Advances in Neural Information Processing Systems 36","author":"Pourreza Mohammadreza","year":"2024","unstructured":"Mohammadreza Pourreza and Davood Rafiei. 2024. Din-sql: Decomposed in-context learning of text-to-sql with self-correction. Advances in Neural Information Processing Systems 36 (2024)."},{"key":"e_1_2_2_70_1","unstructured":"Bowen Qin Binyuan Hui Lihan Wang Min Yang Jinyang Li Binhua Li Ruiying Geng Rongyu Cao Jian Sun Luo Si et al. 2022. A survey on text-to-sql parsing: Concepts methods and future directions. arXiv preprint arXiv:2208.13629 (2022)."},{"key":"e_1_2_2_71_1","unstructured":"Alec Radford Karthik Narasimhan Tim Salimans Ilya Sutskever et al. 2018. Improving language understanding by generative pre-training. (2018)."},{"key":"e_1_2_2_72_1","unstructured":"Alec Radford Jeffrey Wu Rewon Child David Luan Dario Amodei Ilya Sutskever et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1 8 (2019) 9."},{"key":"e_1_2_2_73_1","first-page":"1","article-title":"Exploring the limits of transfer learning with a unified text-to-text transformer","volume":"21","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research 21, 140 (2020), 1--67.","journal-title":"Journal of machine learning research"},{"key":"e_1_2_2_74_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1410"},{"key":"e_1_2_2_75_1","doi-asserted-by":"publisher","DOI":"10.1145\/3589319"},{"key":"e_1_2_2_76_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4612-4378-6"},{"key":"e_1_2_2_77_1","doi-asserted-by":"publisher","DOI":"10.1145\/582095.582099"},{"key":"e_1_2_2_78_1","doi-asserted-by":"publisher","DOI":"10.14778\/3565838.3565848"},{"key":"e_1_2_2_79_1","first-page":"97","article-title":"On estimation of the size of the dictionary of a long text on the basis of a sample","volume":"19","author":"Shlosser A","year":"1981","unstructured":"A Shlosser. 1981. On estimation of the size of the dictionary of a long text on the basis of a sample. Engineering Cybernetics 19, 1 (1981), 97--102.","journal-title":"Engineering Cybernetics"},{"key":"e_1_2_2_80_1","doi-asserted-by":"publisher","DOI":"10.1080\/03610928608829161"},{"key":"e_1_2_2_81_1","first-page":"45","article-title":"Word frequency distributions and type-token characteristics","volume":"11","author":"Sichel Herbert S","year":"1986","unstructured":"Herbert S Sichel. 1986. Word frequency distributions and type-token characteristics. Math. Scientist 11 (1986), 45--72.","journal-title":"Math. Scientist"},{"key":"e_1_2_2_82_1","doi-asserted-by":"publisher","DOI":"10.1016\/0306-4573(92)90088-H"},{"key":"e_1_2_2_83_1","doi-asserted-by":"publisher","DOI":"10.14778\/3632093.3632111"},{"key":"e_1_2_2_84_1","volume-title":"Nonparametric estimation of species richness. Biometrics","author":"Smith Eric P","year":"1984","unstructured":"Eric P Smith and Gerald van Belle. 1984. Nonparametric estimation of species richness. Biometrics (1984), 119--129."},{"key":"e_1_2_2_85_1","volume-title":"Stanford Workshop on Buffer Sizing.","author":"Spang Bruce","year":"2019","unstructured":"Bruce Spang and Nick McKeown. 2019. On estimating the number of flows. In Stanford Workshop on Buffer Sizing."},{"key":"e_1_2_2_86_1","doi-asserted-by":"publisher","DOI":"10.1145\/3514221.3517906"},{"key":"e_1_2_2_87_1","volume-title":"Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971","author":"Touvron Hugo","year":"2023","unstructured":"Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth\u00e9e Lacroix, Baptiste Rozi\u00e8re, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)."},{"key":"e_1_2_2_88_1","doi-asserted-by":"publisher","DOI":"10.1145\/3125643"},{"key":"e_1_2_2_89_1","volume-title":"Estimating the Unseen: Improved Estimators for Entropy and other Properties. Advances in Neural Information Processing Systems 26","author":"Valiant Paul","year":"2013","unstructured":"Paul Valiant and Gregory Valiant. 2013. Estimating the Unseen: Improved Estimators for Entropy and other Properties. Advances in Neural Information Processing Systems 26 (2013)."},{"key":"e_1_2_2_90_1","volume-title":"Attention is all you need. Advances in neural information processing systems 30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_2_2_91_1","doi-asserted-by":"publisher","DOI":"10.1145\/78922.78925"},{"key":"e_1_2_2_92_1","doi-asserted-by":"publisher","DOI":"10.14778\/3489496.3489508"},{"key":"e_1_2_2_93_1","doi-asserted-by":"publisher","DOI":"10.1214\/17-AOS1665"},{"key":"e_1_2_2_94_1","unstructured":"Xianghong Xu Tieying Zhang Xiao He Haoyang Li Rong Kang Shuai Wang Linhui Xu Zhimin Liang Shangyu Luo Lei Zhang et al. 2025. AdaNDV: Adaptive Number of Distinct Value Estimation via Learning to Select and Fuse Estimators. arXiv preprint arXiv:2502.16190 (2025)."},{"key":"e_1_2_2_95_1","unstructured":"An Yang Baosong Yang Binyuan Hui Bo Zheng Bowen Yu Chang Zhou Chengpeng Li Chengyuan Li Dayiheng Liu Fei Huang et al. 2024. Qwen2 technical report. arXiv preprint arXiv:2407.10671 (2024)."},{"key":"e_1_2_2_96_1","doi-asserted-by":"publisher","DOI":"10.14778\/3654621.3654622"},{"key":"e_1_2_2_97_1","unstructured":"Wayne Xin Zhao Kun Zhou Junyi Li Tianyi Tang Xiaolei Wang Yupeng Hou Yingqian Min Beichen Zhang Junjie Zhang Zican Dong et al. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223 (2023)."},{"key":"e_1_2_2_98_1","volume-title":"Llm as dba. arXiv preprint arXiv:2308.05481","author":"Zhou Xuanhe","year":"2023","unstructured":"Xuanhe Zhou, Guoliang Li, and Zhiyuan Liu. 2023. Llm as dba. arXiv preprint arXiv:2308.05481 (2023)."},{"key":"e_1_2_2_99_1","doi-asserted-by":"publisher","DOI":"10.14778\/3675034.3675043"},{"key":"e_1_2_2_100_1","volume-title":"LLM-Enhanced Data Management. arXiv preprint arXiv:2402.02643","author":"Zhou Xuanhe","year":"2024","unstructured":"Xuanhe Zhou, Xinyang Zhao, and Guoliang Li. 2024. LLM-Enhanced Data Management. arXiv preprint arXiv:2402.02643 (2024)."}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3725336","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,31]],"date-time":"2026-03-31T18:52:19Z","timestamp":1774983139000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3725336"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,17]]},"references-count":100,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,6,17]]}},"alternative-id":["10.1145\/3725336"],"URL":"https:\/\/doi.org\/10.1145\/3725336","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,17]]}}}