{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T03:12:42Z","timestamp":1784776362129,"version":"3.55.0"},"reference-count":60,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2025,2,10]],"date-time":"2025-02-10T00:00:00Z","timestamp":1739145600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"},{"start":{"date-parts":[[2025,2,10]],"date-time":"2025-02-10T00:00:00Z","timestamp":1739145600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Nat Mach Intell"],"DOI":"10.1038\/s42256-025-00984-1","type":"journal-article","created":{"date-parts":[[2025,2,10]],"date-time":"2025-02-10T10:03:15Z","timestamp":1739181795000},"page":"172-180","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":18,"title":["On the caveats of AI autophagy"],"prefix":"10.1038","volume":"7","author":[{"given":"Xiaodan","family":"Xing","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fadong","family":"Shi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8032-9097","authenticated-orcid":false,"given":"Jiahao","family":"Huang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3857-6112","authenticated-orcid":false,"given":"Yinzhe","family":"Wu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4542-3336","authenticated-orcid":false,"given":"Yang","family":"Nan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sheng","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yingying","family":"Fang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3484-5031","authenticated-orcid":false,"given":"Michael","family":"Roberts","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Carola-Bibiane","family":"Sch\u00f6nlieb","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Javier","family":"Del Ser","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7344-7733","authenticated-orcid":false,"given":"Guang","family":"Yang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,2,10]]},"reference":[{"key":"984_CR1","unstructured":"Villalobos, P. et al. Will we run out of data? an analysis of the limits of scaling datasets in machine learning. Preprint at https:\/\/arxiv.org\/abs\/2211.04325 (2022)."},{"key":"984_CR2","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s40537-019-0197-0","volume":"6","author":"C Shorten","year":"2019","unstructured":"Shorten, C. & Khoshgoftaar, T. M. A survey on image data augmentation for deep learning. J. Big Data 6, 1\u201348 (2019).","journal-title":"J. Big Data"},{"key":"984_CR3","doi-asserted-by":"publisher","first-page":"755","DOI":"10.1038\/s41586-024-07566-y","volume":"631","author":"I Shumailov","year":"2024","unstructured":"Shumailov, I. et al. AI models collapse when trained on recursively generated data. Nature 631, 755\u2013759 (2024).","journal-title":"Nature"},{"key":"984_CR4","unstructured":"Alemohammad, S. et al. Self-consuming generative models go mad. In The Twelfth International Conference on Learning Representations (ICLR, 2024). A highly relevant theoretical analysis and empirical finding on AI autophagy, introducing the term \u2018autophagy\u2019 using image synthesis models."},{"key":"984_CR5","unstructured":"Yamaguchi, S. & Fukuda, T. On the limitation of diffusion models for synthesizing training datasets. In NeurIPS Workshop on Synthetic Data Generation with Generative AI (NeurIPS, 2023)."},{"key":"984_CR6","doi-asserted-by":"crossref","unstructured":"Hataya, R., Bao, H. & Arai, H. Will large-scale generative models corrupt future datasets? In Proc. IEEE\/CVF International Conference on Computer Vision 20555\u201320565 (IEEE, 2023).","DOI":"10.1109\/ICCV51070.2023.01879"},{"key":"984_CR7","doi-asserted-by":"publisher","first-page":"101896","DOI":"10.1016\/j.inffus.2023.101896","volume":"99","author":"N D\u00edaz-Rodr\u00edguez","year":"2023","unstructured":"D\u00edaz-Rodr\u00edguez, N. et al. Connecting the dots in trustworthy artificial intelligence: from AI principles, ethics, and key requirements to responsible AI systems and regulation. Inf. Fusion 99, 101896 (2023).","journal-title":"Inf. Fusion"},{"key":"984_CR8","unstructured":"Bertrand, Q., Bose, A. J., Duplessis, A., Jiralerspong, M. & Gidel, G. On the stability of iterative retraining of generative models on their own data. In The Twelfth International Conference on Learning Representations (ICLR, 2024). A theoretical analysis and potential solution, proposing the stability of training loops that incorporate fixed real data."},{"key":"984_CR9","unstructured":"Gerstgrasser, M. et al. Is model collapse inevitable? Breaking the curse of recursion by accumulating real and synthetic data. In First Conference on Language Modeling (COLM, 2024). A potential solution, demonstrating that, using a linear regression model, if a fixed set of real data is maintained, the test error stabilizes with a finite upper bound, independent of the number of iterations."},{"key":"984_CR10","unstructured":"Mart\u00ednez, G. et al. Combining generative artificial intelligence (AI) and the internet: heading towards evolution or degradation? Preprint at https:\/\/arxiv.org\/abs\/2303.01255 (2023). A highly relevant empirical finding on AI autophagy, demonstrated through image synthesis models."},{"key":"984_CR11","doi-asserted-by":"crossref","unstructured":"Mart\u00ednez, G. et al. Towards understanding the interplay of generative artificial intelligence and the internet. In International Workshop on Epistemic Uncertainty in Artificial Intelligence (Epi UAI 223) 59\u201373 (Springer Nature, 2023). A highly relevant empirical finding on AI autophagy, demonstrated through image synthesis models.","DOI":"10.1007\/978-3-031-57963-9_5"},{"key":"984_CR12","unstructured":"Briesch, M., Sobania, D. & Rothlauf, F. Large language models suffer from their own output: an analysis of the self-consuming training loop. Preprint at https:\/\/arxiv.org\/abs\/2311.16822 (2023). A highly relevant theoretical analysis."},{"key":"984_CR13","unstructured":"Feng, Y., Dohmatob, E., Yang, P., Charton, F. & Kempe, J. Beyond model collapse: scaling up with synthesized data requires reinforcement. Preprint at https:\/\/arxiv.org\/abs\/2406.07515 (2024). A potential solution, proposing pruning incorrect predictions and selecting optimal guesses from multiple outputs, to counteract model collapse when scaling with synthesized data."},{"key":"984_CR14","unstructured":"Ferbach, D., Bertrand, Q., Bose, A. J. & Gidel, G. Self-consuming generative models with curated data provably optimize human preferences. In Advances in Neural Information Processing Systems 1\u201327 (2024). A potential solution to AI autophagy. This paper demonstrates that when data are curated according to a reward model, the expected reward of the iterative retraining process is maximized."},{"key":"984_CR15","unstructured":"Bohacek, M. & Farid, H. Nepotistically trained generative-AI models collapse. Preprint at https:\/\/arxiv.org\/abs\/2311.12202 (2023). A highly relevant empirical finding on AI autophagy, demonstrated through image synthesis models."},{"key":"984_CR16","unstructured":"Dohmatob, E., Feng, Y. & Kempe, J. Model collapse demystified: the case of regression. In Advances in Neural Information Processing Systems (NeurIPS, 2024). A highly relevant theoretical analysis of AI autophagy using a linear regression model, analysing the accumulation of error in fully synthetic loops."},{"key":"984_CR17","doi-asserted-by":"crossref","unstructured":"Guo, Y., Shang, G., Vazirgiannis, M. & Clavel, C. The curious decline of linguistic diversity: Training language models on synthetic text. In Findings of the Association for Computational Linguistics: NAACL 2024 3589\u20133604 (ACL, 2024). A highly relevant empirical finding on AI autophagy, demonstrating the decline of diversity across multiple aspects.","DOI":"10.18653\/v1\/2024.findings-naacl.228"},{"key":"984_CR18","unstructured":"Fu, S., Zhang, S., Wang, Y., Tian, X. & Tao, D. Towards theoretical understandings of self-consuming generative models. In Proc. 41st International Conference on Machine Learning Vol. 235, 14228\u201314255 (PMLR, 2025). A theoretical support for fixed real dataset loops, providing a theoretical guarantee on stability and error bounds when training models with a combination of real and synthetic data."},{"key":"984_CR19","unstructured":"Dohmatob, E., Feng, Y., Yang, P., Charton, F. & Kempe, J. A tale of tails: model collapse as a change of scaling laws. In Proc. 41st International Conference on Machine Learning Vol. 235, 11165\u201311197 (PMLR, 2025). This study reveals how scaling with synthetic data can disrupt the scaling laws of large models, leading to model collapse."},{"key":"984_CR20","unstructured":"Azizi, S., Kornblith, S., Saharia, C., Norouzi, M. & Fleet, D. J. Synthetic data from diffusion models improves imagenet classification. In Transactions on Machine Learning Research (TMLR, 2023)."},{"key":"984_CR21","unstructured":"Trabucco, B., Doherty, K., Gurinas, M. & Salakhutdinov, R. Effective data augmentation with diffusion models. In Twelfth International Conference on Learning Representations (ICLR, 204)."},{"key":"984_CR22","unstructured":"Ravuri, S. & Vinyals, O. Classification accuracy score for conditional generative models. In Proc. 33rd International Conference on Neural Information Processing Systems 12268\u201312279 (ACM, 2019)."},{"key":"984_CR23","doi-asserted-by":"crossref","unstructured":"Chen, T., Hirota, Y., Otani, M., Garcia, N. & Nakashima, Y. Would deep generative models amplify bias in future models? In Proc. IEEE\/CVF Conference on Computer Vision and Pattern Recognition 10833\u201310843 (IEEE, 2024).","DOI":"10.1109\/CVPR52733.2024.01030"},{"key":"984_CR24","unstructured":"Brock, A., Donahue, J. & Simonyan, K. Large scale GAN training for high fidelity natural image synthesis. In International Conference on Learning Representations 2019 (ICLR, 2018)."},{"key":"984_CR25","unstructured":"Gillman, N., Freeman, M., Aggarwal, D., Hsu, C.-H., Luo, C., Tian, Y. & Sun, C. Self-correcting self-consuming loops for generative model training. In Proc. 41st International Conference on Machine Learning Vol. 235, 15646\u201315677 (PMLR, 2025\uff09A potential solution, proposing using a predefined criteria to filter out incorrect synthetic contents."},{"key":"984_CR26","unstructured":"Alemohammad, S., Humayun, A. I., Agarwal, S., Collomosse, J. & Baraniuk, R. Self-improving diffusion models with synthetic data. Preprint at https:\/\/arxiv.org\/abs\/2408.16333 (2024). A solution for AI autophagy by using synthetic datasets to optimize the diffusion model\u2019s score function through a fine-tuning step."},{"key":"984_CR27","unstructured":"Wen, Y., Kirchenbauer, J., Geiping, J. & Goldstein, T. Tree-ring watermarks: fingerprints for diffusion images that are invisible and robust. In Advances in Neural Information Processing Systems (NeurIPS, 2023)."},{"key":"984_CR28","doi-asserted-by":"crossref","unstructured":"Fernandez, P., Couairon, G., J\u00e9gou, H., Douze, M. & Furon, T. The stable signature: rooting watermarks in latent diffusion models. In Proc. IEEE\/CVF International Conference on Computer Vision 22466\u201322477 (IEEE, 2023).","DOI":"10.1109\/ICCV51070.2023.02053"},{"key":"984_CR29","unstructured":"Bui, T., Agarwal, S. & Collomosse, J. TrustMark: universal watermarking for arbitrary resolution images. Preprint at https:\/\/arxiv.org\/abs\/2311.18297 (2023)."},{"key":"984_CR30","doi-asserted-by":"crossref","unstructured":"Tancik, M., Mildenhall, B. & Ng, R. StegaStamp: invisible hyperlinks in physical photographs. In Proc. IEEE\/CVF Conference on Computer Vision and Pattern Recognition 2117\u20132126 (IEEE, 2020).","DOI":"10.1109\/CVPR42600.2020.00219"},{"key":"984_CR31","doi-asserted-by":"crossref","unstructured":"Abdelnabi, S. & Fritz, M. Adversarial watermarking transformer: towards tracing text provenance with data hiding. In 2021 IEEE Symposium on Security and Privacy 121\u2013140 (IEEE, 2021).","DOI":"10.1109\/SP40001.2021.00083"},{"key":"984_CR32","doi-asserted-by":"crossref","unstructured":"Yoo, K., Ahn, W., Jang, J. & Kwak, N. Robust multi-bit natural language watermarking through invariant features. In Proc. 61st Annual Meeting of the Association for Computational Linguistics Vol. 1, 2092\u20132115 (ACL, 2023).","DOI":"10.18653\/v1\/2023.acl-long.117"},{"key":"984_CR33","unstructured":"Yang, X. et al. Watermarking text generated by black-box language models. Preprint at https:\/\/arxiv.org\/abs\/2305.08883 (2023)."},{"key":"984_CR34","unstructured":"Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I. & Goldstein, T. A watermark for large language models. In Proc. 40th International Conference on Machine Learning Vol. 202, 17061\u201317084 (PMLR, 2023)."},{"key":"984_CR35","unstructured":"Wang, L. et al. Towards codable text watermarking for large language models. In The 12th International Conference on Learning Representations (ICLR, 2024)."},{"key":"984_CR36","unstructured":"Sadasivan, V. S., Kumar, A., Balasubramanian, S., Wang, W. & Feizi, S. Can AI-generated text be reliably detected? In The 12th International Conference on Learning Representations (ICLR, 2024)."},{"key":"984_CR37","doi-asserted-by":"publisher","first-page":"95","DOI":"10.1016\/j.eswa.2017.12.003","volume":"97","author":"Y Liu","year":"2018","unstructured":"Liu, Y., Tang, S., Liu, R., Zhang, L. & Ma, Z. Secure and robust digital image watermarking scheme using logistic and RSA encryption. Expert Syst. Appl. 97, 95\u2013105 (2018).","journal-title":"Expert Syst. Appl."},{"key":"984_CR38","doi-asserted-by":"publisher","first-page":"11852","DOI":"10.3390\/app132111852","volume":"13","author":"X Zhong","year":"2023","unstructured":"Zhong, X., Das, A., Alrasheedi, F. & Tanvir, A. A brief, in-depth survey of deep learning-based image watermarking. Appl. Sci. 13, 11852 (2023).","journal-title":"Appl. Sci."},{"key":"984_CR39","doi-asserted-by":"publisher","first-page":"e13570","DOI":"10.1111\/exsy.13570","volume":"41","author":"LA Passos","year":"2024","unstructured":"Passos, L. A. et al. A review of deep learning-based approaches for deepfake content detection. Expert Syst. 41, e13570 (2024).","journal-title":"Expert Syst."},{"key":"984_CR40","doi-asserted-by":"crossref","unstructured":"Wang, S. Y., Wang, O., Zhang, R., Owens, A. & Efros, A. A. CNN-generated images are surprisingly easy to spot\u2026 for now. In Proc. IEEE\/CVF Conference on Computer Vision and Pattern Recognition 8695\u20138704 (IEEE, 2020).","DOI":"10.1109\/CVPR42600.2020.00872"},{"key":"984_CR41","doi-asserted-by":"crossref","unstructured":"Ju, Y. et al. Fusing global and local features for generalized AI-synthesized image detection. In 2022 IEEE International Conference on Image Processing 3465\u20133469 (IEEE, 2022).","DOI":"10.1109\/ICIP46576.2022.9897820"},{"key":"984_CR42","doi-asserted-by":"crossref","unstructured":"Mandelli, S., Bonettini, N., Bestagini, P. & Tubaro, S. Detecting GAN-generated images by orthogonal training of multiple CNNs. In 2022 IEEE International Conference on Image Processing 3091\u20133095 (IEEE, 2022).","DOI":"10.1109\/ICIP46576.2022.9897310"},{"key":"984_CR43","unstructured":"New AI classifier for indicating AI-written text. OpenAI https:\/\/openai.com\/blog\/new-ai-classifier-for-indicating-ai-written-text (accessed 17 March 2024)."},{"key":"984_CR44","unstructured":"Yu, X. et al. GPT paternity test: GPT generated text detection with GPT genetic inheritance. Preprint at https:\/\/arxiv.org\/abs\/2305.12519 (2023)."},{"key":"984_CR45","unstructured":"Hu, X., Chen, P.-Y. & Ho, T.-Y. RADAR: Robust AI-text detection via adversarial learning. In Advances in Neural Information Processing Systems (NeurIPS, 2023)."},{"key":"984_CR46","unstructured":"Tian, Y., Chen, H., Wang, X., Bai, Z., Zhang, Q., Li, R., Xu, C. & Wang, Y. Multiscale positive-unlabeled detection of AI-generated texts. In International Conference on Learning Representations (ICLR, 2024)."},{"key":"984_CR47","doi-asserted-by":"publisher","first-page":"100779","DOI":"10.1016\/j.patter.2023.100779","volume":"4","author":"W Liang","year":"2023","unstructured":"Liang, W., Yuksekgonul, M., Mao, Y., Wu, E. & Zou, J. GPT detectors are biased against non-native English writers. Patterns 4, 100779 (2023).","journal-title":"Patterns"},{"key":"984_CR48","unstructured":"Coley, M. Guidance on AI detection and why we\u2019re disabling Turnitin\u2019s AI detector. Vanderbilt University (2023)."},{"key":"984_CR49","unstructured":"OpenAI. DALL\u00b7E 2 pre-training mitigations; https:\/\/openai.com\/index\/dall-e-2-pre-training-mitigations\/ (accessed 17 March 2024)."},{"key":"984_CR50","unstructured":"OpenAI. An update on our safety & security practices; https:\/\/openai.com\/index\/update-on-safety-and-security-practices (accessed 25 September 2024)."},{"key":"984_CR51","unstructured":"Stability AI. Artificial intelligence and the content ecosystem; https:\/\/stability.ai\/artificial-intelligence-and-the-content-ecosystem (accessed 25 September 2024)."},{"key":"984_CR52","unstructured":"Google AI. Google AI model documentation and responsible AI practices; https:\/\/ai.google\/discover\/palm2 (accessed 25 September 2024)."},{"key":"984_CR53","unstructured":"Cyberspace Administration of China. Notice on the issuance of the 2023 regulations on internet security; http:\/\/www.cac.gov.cn\/2023-07\/13\/c_1690898327029107.htm (accessed 17 March 2024)."},{"key":"984_CR54","unstructured":"Madiega, T. Artificial intelligence act; 'EU Legislation in Progress' briefings (European Parliament, 2021)."},{"key":"984_CR55","unstructured":"European Parliament and Council of the European Union. Regulation (EU) 2024\/1689 of the European Parliament and of the Council (European Union, 2024)."},{"key":"984_CR56","unstructured":"National Institute of Standards and Technology (NIST). Safe, secure, and trustworthy development and use of artificial intelligence; Executive Order 14110, 75191\u201375226 (2023)."},{"key":"984_CR57","doi-asserted-by":"publisher","first-page":"55","DOI":"10.1007\/s10676-023-09728-4","volume":"25","author":"A Knott","year":"2023","unstructured":"Knott, A. et al. Generative AI models should include detection mechanisms as a condition for public release. Ethics Inf. Technol. 25, 55 (2023).","journal-title":"Ethics Inf. Technol."},{"key":"984_CR58","doi-asserted-by":"publisher","first-page":"141","DOI":"10.1109\/MSP.2012.2211477","volume":"29","author":"L Deng","year":"2012","unstructured":"Deng, L. The MNIST database of handwritten digit images for machine learning research [best of the web]. IEEE Signal Process. Mag. 29, 141\u2013142 (2012).","journal-title":"IEEE Signal Process. Mag."},{"key":"984_CR59","doi-asserted-by":"crossref","unstructured":"Karras, T. et al. Analyzing and improving the image quality of stylegan. In Proc. IEEE\/CVF Conference on Computer Vision and Pattern Recognition 8110\u20138119 (IEEE, 2020).","DOI":"10.1109\/CVPR42600.2020.00813"},{"key":"984_CR60","doi-asserted-by":"crossref","unstructured":"Choi, Y., Uh, Y., Yoo, J. & Ha, J. W. StarGAN v2: diverse image synthesis for multiple domains. In Proc. IEEE\/CVF Conference on Computer Vision and Pattern Recognition 8188\u20138197 (IEEE, 2020).","DOI":"10.1109\/CVPR42600.2020.00821"}],"container-title":["Nature Machine Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.nature.com\/articles\/s42256-025-00984-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s42256-025-00984-1","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s42256-025-00984-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,2,23]],"date-time":"2025-02-23T23:03:14Z","timestamp":1740351794000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.nature.com\/articles\/s42256-025-00984-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,10]]},"references-count":60,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2025,2]]}},"alternative-id":["984"],"URL":"https:\/\/doi.org\/10.1038\/s42256-025-00984-1","relation":{},"ISSN":["2522-5839"],"issn-type":[{"value":"2522-5839","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,2,10]]},"assertion":[{"value":"24 June 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 December 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 February 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare no competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}]}}