{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,7]],"date-time":"2026-03-07T00:12:40Z","timestamp":1772842360102,"version":"3.50.1"},"reference-count":62,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"name":"Outstanding Innovative Talents Cultivation Funded Programs 2023 of Renmin University of China"},{"name":"National Social Science Fund of China","award":["25BTQ030"],"award-info":[{"award-number":["25BTQ030"]}]},{"name":"Major Project of Humanities and Social Sciences Key Research Base of the Ministry of Education","award":["22JJD870001"],"award-info":[{"award-number":["22JJD870001"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2026,4,30]]},"abstract":"<jats:p>Large language models (LLMs) significantly advance data augmentation research. However, existing approaches largely overlook two fundamental issues. First, empirical evidence demonstrating the alignment between LLM natural language (LLMNL) and human natural language (HNL) remains limited. Second, current methodologies often neglect the variability of LLMNL in comparison with HNL. To address the gap, we introduce a comprehensive scaling-law-based framework for examining the congruence between LLMNL and HNL. Through extensive experiments, we uncover a progression of findings: LLMNL fails to achieve congruence with HNL; there is a consistent discrepancy, with Mandelbrot exponents for LLMNL being approximately 0.2 lower than those of HNL; LLMNL exhibits reduced fractal complexity, corroborated our analysis to stylistic factors such as readability, sentiment, and semantics. Furthermore, we propose a new data augmentation approach for text classification, which leverages scaling laws to make decisions on LLM-generated texts. Extensive experiments under real-world scenarios demonstrate that our approach is competitive and robust, outperforming recent methods and consistently maintaining performance advantages across different LLMs and prompts.<\/jats:p>","DOI":"10.1145\/3787100","type":"journal-article","created":{"date-parts":[[2026,1,5]],"date-time":"2026-01-05T11:23:22Z","timestamp":1767612202000},"page":"1-37","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Data Augmentation with Large Language Models: A Scaling Law-Guided Approach"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0369-2765","authenticated-orcid":false,"given":"Zhenhua","family":"Wang","sequence":"first","affiliation":[{"name":"School of Information Resource Management, Renmin University of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7348-1875","authenticated-orcid":false,"given":"Guang","family":"Xu","sequence":"additional","affiliation":[{"name":"School of Information and Communication, Nankai University, Tianjin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2819-6653","authenticated-orcid":false,"given":"Ming","family":"Ren","sequence":"additional","affiliation":[{"name":"School of Information Resource Management, Renmin University of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,2,19]]},"reference":[{"issue":"5","key":"e_1_3_2_2_2","first-page":"1","article-title":"Improving short text classification with augmented data using GPT-3","volume":"30","author":"Balkus S. V.","year":"2022","unstructured":"S. V. Balkus and D. Yan. 2022. Improving short text classification with augmented data using GPT-3. Natural Language Engineering 30, 5 (2022), 1\u201330.","journal-title":"Natural Language Engineering"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2024.104035"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/TBDATA.2025.3536934"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3654674"},{"issue":"2","key":"e_1_3_2_6_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3686807","article-title":"Knowledge-tuning large language models with structured medical knowledge bases for trustworthy response generation in Chinese","volume":"19","author":"Wang H.","year":"2024","unstructured":"H. Wang, S. Zhao, Z. Qiang, Z. Li, C. Liu, N. Xi, Y. Du, B. Qin, and T. Liu. 2024. Knowledge-tuning large language models with structured medical knowledge bases for trustworthy response generation in Chinese. ACM Transactions on Knowledge Discovery from Data 19, 2 (2024), 1\u201317.","journal-title":"ACM Transactions on Knowledge Discovery from Data"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3649506"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.compind.2024.104082"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-61057-8_31"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.eacl-long.39"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.inffus.2023.101861"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.emnlp-main.73"},{"key":"e_1_3_2_13_2","first-page":"24457","volume-title":"International Conference on Machine Learning","author":"Meng Y.","year":"2023","unstructured":"Y. Meng, M. Michalski, J. Huang, Y. Zhang, T. Abdelzaher, and J. Han. 2023. Tuning language models as training data generators for augmentation-enhanced few-shot learning. In International Conference on Machine Learning. PMLR, 24457\u201324477."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.eacl-short.17"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1162\/coli_a_00355"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1017\/S0272263121000553"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.physrep.2023.12.002"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.chaos.2021.111489"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2015.10.023"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.compind.2023.103875"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2023.103561"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.joi.2023.101453"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3748305"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0256133"},{"key":"e_1_3_2_25_2","unstructured":"\u0141. D\u0119bowski. 2023. A simplistic model of neural scaling laws: Multiperiodic Santa Fe processes. arXiv:2302.09049. Retrieved from https:\/\/arxiv.org\/abs\/2302.09049"},{"key":"e_1_3_2_26_2","first-page":"27730","article-title":"Training language models to follow instructions with human feedback","volume":"35","author":"Ouyang L.","year":"2022","unstructured":"L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al. 2022. Training language models to follow instructions with human feedback. In Proceedings of the Advances in Neural Information Processing Systems, Vol. 35, 27730\u201327744.","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3358168"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.compind.2022.103647"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3638781"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3638057"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3608953"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/3673232"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1002\/asi.24872"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1670"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2019.2890970"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2025.104139"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/JBHI.2024.3435085"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.dash-1.4"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3736418"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2023.119809"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.psep.2022.11.005"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2024.120653"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2022.103073"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.dss.2025.114578"},{"key":"e_1_3_2_45_2","unstructured":"B. Guo X. Zhang Z. Wang M. Jiang J. Nie Y. Ding J. Yue and Y. Wu. 2023. How close is ChatGPT to human experts? Comparison corpus evaluation and detection. arXiv:2301.07597. Retrieved from https:\/\/arxiv.org\/abs\/2301.07597"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2021.101227"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10936-019-09673-8"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1177\/2329488416675456"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.586"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2024.104054"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10339-010-0368-6"},{"issue":"2","key":"e_1_3_2_52_2","first-page":"145","article-title":"The problems of LLM-generated data in social science research","volume":"18","author":"Rossi L.","year":"2024","unstructured":"L. Rossi, K. Harrison, and I. Shklovski. 2024. The problems of LLM-generated data in social science research. Sociologica 18, 2 (2024), 145\u2013168.","journal-title":"Sociologica"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevE.98.052139"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-023-45644-9"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3571736"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.elerap.2025.101521"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-024-10772-9"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2024.120234"},{"key":"e_1_3_2_59_2","unstructured":"P. Sahoo A. K. Singh S. Saha V. Jain S. Mondal and A. Chadha. 2024. A systematic survey of prompt engineering in large language models: Techniques and applications. arXiv:2402.07927. Retrieved from https:\/\/arxiv.org\/abs\/2402.07927"},{"key":"e_1_3_2_60_2","first-page":"8048","volume-title":"Proceedings of the 33rd International Joint Conference on Artificial Intelligence","author":"Guo T.","year":"2024","unstructured":"T. Guo, X. Chen, Y. Wang, R. Chang, S. Pei, N. V. Chawla, O. Wiest, and X. Zhang. 2024. Large language model based multi-agents: A survey of progress and challenges. In Proceedings of the 33rd International Joint Conference on Artificial Intelligence, 8048\u20138057."},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41562-024-01815-w"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41562-017-0184-4"},{"key":"e_1_3_2_63_2","unstructured":"K. Wang J. Zhu M. Ren Z. Liu S. Li Z. Zhang C. Zhang X. Wu Q. Zhan Q. Liu et al. 2024. A survey on data synthesis and augmentation for large language models. arXiv:2410.12896. Retrieved from https:\/\/arxiv.org\/abs\/2410.12896"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3787100","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,6]],"date-time":"2026-03-06T14:35:11Z","timestamp":1772807711000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3787100"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,19]]},"references-count":62,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,4,30]]}},"alternative-id":["10.1145\/3787100"],"URL":"https:\/\/doi.org\/10.1145\/3787100","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"value":"1556-4681","type":"print"},{"value":"1556-472X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,19]]},"assertion":[{"value":"2024-06-25","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-12-27","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-19","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}