{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,14]],"date-time":"2026-05-14T16:41:58Z","timestamp":1778776918086,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":96,"publisher":"ACM","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,11,19]]},"DOI":"10.1145\/3719027.3765063","type":"proceedings-article","created":{"date-parts":[[2025,11,22]],"date-time":"2025-11-22T23:37:25Z","timestamp":1763854645000},"page":"21-35","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["The Odyssey of robots.txt Governance: Measuring Convention Implications of Web Bots in Large Language Model Services"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5212-0501","authenticated-orcid":false,"given":"Jian","family":"Cui","sequence":"first","affiliation":[{"name":"University of Illinois Urbana-Champaign, Champaign, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7827-9369","authenticated-orcid":false,"given":"Mingming","family":"Zha","sequence":"additional","affiliation":[{"name":"Indiana University Bloomington, Bloomington, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0607-4946","authenticated-orcid":false,"given":"XiaoFeng","family":"Wang","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7555-1673","authenticated-orcid":false,"given":"Xiaojing","family":"Liao","sequence":"additional","affiliation":[{"name":"University of Illinois Urbana-Champaign, Champaign, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,11,22]]},"reference":[{"key":"e_1_3_2_2_1_1","unstructured":"Bots - openai platform documentation. https:\/\/platform.openai.com\/docs\/bots."},{"key":"e_1_3_2_2_2_1","unstructured":"Which crawlers does bing use? https:\/\/www.bing.com\/webmasters\/help\/whichcrawlers-does-bing-use-8c184ec0."},{"key":"e_1_3_2_2_3_1","volume-title":"an open dataset for training large language models","author":"Redpajama","year":"2023","unstructured":"Redpajama: an open dataset for training large language models, 2023."},{"key":"e_1_3_2_2_4_1","volume-title":"The odyssey of robots.txt governance: Measuring compliance implications of web crawling bots in large language model services. https:\/\/sites.google.com\/view\/botcompliance\/home","author":"Artifact","year":"2024","unstructured":"Artifact: The odyssey of robots.txt governance: Measuring compliance implications of web crawling bots in large language model services. https:\/\/sites.google.com\/view\/botcompliance\/home, 2024."},{"key":"e_1_3_2_2_5_1","volume-title":"Robots exclusion protocol. https:\/\/www.rfc-editor","author":"RFC","year":"2024","unstructured":"RFC 9309. Robots exclusion protocol. https:\/\/www.rfc-editor.org\/rfc\/rfc9309.html, 2024."},{"key":"e_1_3_2_2_6_1","volume-title":"Bots guide. https:\/\/docs.perplexity.ai\/guides\/bots","author":"Perplexity","year":"2024","unstructured":"Perplexity AI. Bots guide. https:\/\/docs.perplexity.ai\/guides\/bots, 2024."},{"key":"e_1_3_2_2_7_1","unstructured":"Amazon. Amazonbot. https:\/\/developer.amazon.com\/amazonbot."},{"key":"e_1_3_2_2_8_1","unstructured":"Anthropic. Does Anthropic crawl data from the web and how can site owners block the crawler? https:\/\/support.anthropic.com\/en\/articles\/8896518-doesanthropic-crawl-data-from-the-web-and-how-can-site-owners-block-thecrawler."},{"key":"e_1_3_2_2_9_1","volume-title":"Does anthropic crawl data from the web, and how can site owners block the crawler?","year":"2024","unstructured":"Anthropic. Does anthropic crawl data from the web, and how can site owners block the crawler?, 2024."},{"key":"e_1_3_2_2_10_1","volume-title":"Wayback machine. https:\/\/wayback-api.archive.org\/","author":"Archive Internet","year":"2024","unstructured":"Internet Archive. Wayback machine. https:\/\/wayback-api.archive.org\/, 2024."},{"key":"e_1_3_2_2_11_1","unstructured":"Baidu. Baiduspider Help Center - How to Block the Crawling. https:\/\/www.baidu.com\/search\/robots_english.html."},{"key":"e_1_3_2_2_12_1","volume-title":"Genai blocking robots.txt. https:\/\/twitter.com\/AndyBeard\/status\/1740647491027267946","author":"Beard Andy","year":"2024","unstructured":"Andy Beard. Genai blocking robots.txt. https:\/\/twitter.com\/AndyBeard\/status\/1740647491027267946, 2024."},{"key":"e_1_3_2_2_13_1","volume-title":"Goose3: A python html content\/article extractor","author":"Bertram Corey","year":"2023","unstructured":"Corey Bertram and Contributors. Goose3: A python html content\/article extractor, 2023. Version 3.1.11."},{"key":"e_1_3_2_2_14_1","first-page":"36","article-title":"Emergent and predictable memorization in large language models","author":"Biderman Stella","year":"2024","unstructured":"Stella Biderman, USVSN PRASHANTH, Lintang Sutawika, Hailey Schoelkopf, Quentin Anthony, Shivanshu Purohit, and Edward Raff. Emergent and predictable memorization in large language models. Advances in Neural Information Processing Systems, 36, 2024.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_2_15_1","volume-title":"Language models are few-shot learners. arXiv preprint arXiv:2005.14165","author":"Brown Tom B","year":"2020","unstructured":"Tom B Brown. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020."},{"key":"e_1_3_2_2_16_1","volume-title":"Quantifying memorization across neural language models. arXiv preprint arXiv:2202.07646","author":"Carlini Nicholas","year":"2022","unstructured":"Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. Quantifying memorization across neural language models. arXiv preprint arXiv:2202.07646, 2022."},{"key":"e_1_3_2_2_17_1","first-page":"2633","volume-title":"30th USENIX Security Symposium (USENIX Security 21)","author":"Carlini Nicholas","year":"2021","unstructured":"Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert- Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21), pages 2633-2650, 2021."},{"key":"e_1_3_2_2_18_1","volume-title":"March","author":"Kim Chan Woo","year":"2022","unstructured":"Chan Woo Kim. Guiding text generation with constrained beam search in transformers, March 2022."},{"key":"e_1_3_2_2_19_1","volume-title":"User experience report. https:\/\/developer.chrome.com\/docs\/crux\/","year":"2024","unstructured":"Chrome. User experience report. https:\/\/developer.chrome.com\/docs\/crux\/, 2024."},{"key":"e_1_3_2_2_20_1","volume-title":"Umbrella popularity list. https:\/\/umbrella-static.s3-us-west-1.amazonaws.com\/index.html","year":"2024","unstructured":"Cisco. Umbrella popularity list. https:\/\/umbrella-static.s3-us-west-1.amazonaws.com\/index.html, 2024."},{"key":"e_1_3_2_2_21_1","volume-title":"Natural language understanding. https:\/\/cloud.ibm.com\/apidocs\/natural-language-understanding#categories","author":"Cloud IBM","year":"2024","unstructured":"IBM Cloud. Natural language understanding. https:\/\/cloud.ibm.com\/apidocs\/natural-language-understanding#categories, 2024."},{"key":"e_1_3_2_2_22_1","volume-title":"Domain ranking. https:\/\/radar.cloudflare.com\/domains","year":"2024","unstructured":"Cloudflare. Domain ranking. https:\/\/radar.cloudflare.com\/domains, 2024."},{"key":"e_1_3_2_2_23_1","unstructured":"Common Crawl. CCBot. https:\/\/commoncrawl.org\/ccbot."},{"key":"e_1_3_2_2_24_1","unstructured":"Summa contributors. Summa. PyPI 2023."},{"key":"e_1_3_2_2_25_1","unstructured":"Common Crawl. Common crawl: Open web data. https:\/\/commoncrawl.org."},{"key":"e_1_3_2_2_26_1","volume-title":"Anthropic ai crawler","year":"2024","unstructured":"Darkvisitors.com. Anthropic ai crawler, 2024."},{"key":"e_1_3_2_2_27_1","volume-title":"Cohere ai crawler","year":"2024","unstructured":"Darkvisitors.com. Cohere ai crawler, 2024."},{"key":"e_1_3_2_2_28_1","volume-title":"July","year":"2024","unstructured":"Davis, Wes. Anthropic's crawler is ignoring websites' anti-ai scraping policies, July 2024."},{"key":"e_1_3_2_2_29_1","volume-title":"Overviewof google crawlers (user agents). https:\/\/developers.google.com\/search\/docs\/crawling-indexing\/overview-google-crawlers","author":"Developers Google","year":"2024","unstructured":"Google Developers. Overviewof google crawlers (user agents). https:\/\/developers.google.com\/search\/docs\/crawling-indexing\/overview-google-crawlers, 2024."},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3589335.3651893"},{"key":"e_1_3_2_2_31_1","unstructured":"DuckDuckGo. Is DuckDuckBot related to DuckDuckGo? https:\/\/duckduckgo.com\/duckduckgo-help-pages\/results\/duckduckbot\/."},{"key":"e_1_3_2_2_32_1","volume-title":"Recurrent ventures - terms and conditions. https:\/\/recurrent.io\/termsand-conditions\/","year":"2024","unstructured":"Dwell. Recurrent ventures - terms and conditions. https:\/\/recurrent.io\/termsand-conditions\/, 2024."},{"key":"e_1_3_2_2_33_1","volume-title":"Domain tools. https:\/\/www.domaintools.com\/resources\/blog\/mirrormirror-on-the-wall-whos-the-fairest-website-of-them-all\/","year":"2024","unstructured":"Farsight. Domain tools. https:\/\/www.domaintools.com\/resources\/blog\/mirrormirror-on-the-wall-whos-the-fairest-website-of-them-all\/, 2024."},{"key":"e_1_3_2_2_34_1","volume-title":"The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027","author":"Gao Leo","year":"2020","unstructured":"Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020."},{"key":"e_1_3_2_2_35_1","volume-title":"Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997","author":"Gao Yunfan","year":"2023","unstructured":"Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997, 2023."},{"key":"e_1_3_2_2_36_1","volume-title":"Google deepmind: Advancing gemini with deep research. https:\/\/blog.google\/products\/gemini\/google-gemini-deep-research\/","year":"2024","unstructured":"Google. Google deepmind: Advancing gemini with deep research. https:\/\/blog.google\/products\/gemini\/google-gemini-deep-research\/, 2024."},{"key":"e_1_3_2_2_37_1","volume-title":"Iab tech lab content taxonomy","author":"Tech Lab IAB","year":"2024","unstructured":"IAB Tech Lab. Iab tech lab content taxonomy, 2024."},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3589334.3645699"},{"key":"e_1_3_2_2_39_1","volume-title":"Alpaca against vicuna: Using llms to uncover memorization of llms. arXiv preprint arXiv:2403.04801","author":"Kassem Aly M","year":"2024","unstructured":"Aly M Kassem, Omar Mahmoud, Niloofar Mireshghallah, Hyunwoo Kim, Yulia Tsvetkov, Yejin Choi, Sherif Saad, and Santu Rana. Alpaca against vicuna: Using llms to uncover memorization of llms. arXiv preprint arXiv:2403.04801, 2024."},{"key":"e_1_3_2_2_40_1","volume-title":"Open Future","author":"Keller Paul","year":"2023","unstructured":"Paul Keller and Zuzanna Warso. Defining best practices for opting out of ml training. Open Future, 2023."},{"key":"e_1_3_2_2_41_1","volume-title":"Ai web crawler bots gone wild! e.g. claudebot, dotbot, petalbot. https:\/\/dev.lucee.org\/t\/ai-web-crawler-bots-gone-wild-e-g-claudebotdotbot-petalbot\/13832","year":"2024","unstructured":"kenricashe. Ai web crawler bots gone wild! e.g. claudebot, dotbot, petalbot. https:\/\/dev.lucee.org\/t\/ai-web-crawler-bots-gone-wild-e-g-claudebotdotbot-petalbot\/13832, 2024. Accessed on [insert your access date]."},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/1367497.1367711"},{"key":"e_1_3_2_2_43_1","unstructured":"Hugo Lauren\u00e7on Lucile Saulnier Thomas Wang Christopher Akiki Albert Villanova del Moral Teven Le Scao Leandro VonWerra Chenghao Mou Eduardo Gonz\u00e1lez Ponferrada Huu Nguyen et al. The bigscience roots corpus: A 1.6 tb composite multilingual dataset. Advances in Neural Information Processing Systems 35:31809-31826 2022."},{"key":"e_1_3_2_2_44_1","first-page":"9459","article-title":"Retrieval-augmented generation for knowledge-intensive nlp tasks","volume":"33","author":"Lewis Patrick","year":"2020","unstructured":"Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K\u00fcttler, Mike Lewis, Wen-tau Yih, Tim Rockt\u00e4schel, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459-9474, 2020.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASID.2019.8925189"},{"key":"e_1_3_2_2_46_1","first-page":"74","volume-title":"Text Summarization Branches Out","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74-81, Barcelona, Spain, July 2004. Association for Computational Linguistics."},{"key":"e_1_3_2_2_47_1","first-page":"25","volume-title":"Proceedings of the ACL 2012 system demonstrations","author":"Lui Marco","year":"2012","unstructured":"Marco Lui and Timothy Baldwin. langid. py: An off-the-shelf language identification tool. In Proceedings of the ACL 2012 system demonstrations, pages 25-30, 2012."},{"key":"e_1_3_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/SP46215.2023.10179300"},{"key":"e_1_3_2_2_49_1","volume-title":"Wayback cdx server api. https:\/\/archive.org\/developers\/wayback-cdx-server.html","author":"Wayback Machine Internet Archive","year":"2024","unstructured":"Internet Archive Wayback Machine. Wayback cdx server api. https:\/\/archive.org\/developers\/wayback-cdx-server.html, 2024."},{"key":"e_1_3_2_2_50_1","volume-title":"Majestic million. https:\/\/majestic.com\/reports\/majestic-million","year":"2024","unstructured":"Majestic. Majestic million. https:\/\/majestic.com\/reports\/majestic-million, 2024."},{"key":"e_1_3_2_2_51_1","volume-title":"Membership inference attacks against language models via neighbourhood comparison. arXiv preprint arXiv:2305.18462","author":"Mattern Justus","year":"2023","unstructured":"Justus Mattern, Fatemehsadat Mireshghallah, Zhijing Jin, Bernhard Sch\u00f6lkopf, Mrinmaya Sachan, and Taylor Berg-Kirkpatrick. Membership inference attacks against language models via neighbourhood comparison. arXiv preprint arXiv:2305.18462, 2023."},{"key":"e_1_3_2_2_52_1","first-page":"2369","volume-title":"33rd USENIX Security Symposium (USENIX Security 24)","author":"Meeus Matthieu","year":"2024","unstructured":"Matthieu Meeus, Shubham Jain, Marek Rei, and Yves-Alexandre de Montjoye. Did the neurons read your book? document-level membership inference for large language models. In 33rd USENIX Security Symposium (USENIX Security 24), pages 2369-2385, 2024."},{"key":"e_1_3_2_2_53_1","volume-title":"June","author":"Marchman Dhruv","year":"2024","unstructured":"Mehrotra, Dhruv and Marchman, Tim. Perplexity is a bullshit machine, June 2024."},{"key":"e_1_3_2_2_54_1","unstructured":"Meta for Developers. MetaWeb Crawlers. https:\/\/developers.facebook.com\/docs\/sharing\/webmasters\/web-crawlers\/."},{"key":"e_1_3_2_2_55_1","volume-title":"Scalable extraction of training data from (production) language models. arXiv preprint arXiv:2311.17035","author":"Nasr Milad","year":"2023","unstructured":"Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A Feder Cooper, Daphne Ippolito, Christopher A Choquette-Choo, Eric Wallace, Florian Tram\u00e8r, and Katherine Lee. Scalable extraction of training data from (production) language models. arXiv preprint arXiv:2311.17035, 2023."},{"key":"e_1_3_2_2_56_1","unstructured":"Naver. Web Document Crawling and Removal Policy. https:\/\/help.naver.com\/service\/5626\/contents\/8026?lang=ko."},{"key":"e_1_3_2_2_57_1","unstructured":"OpenAI. Gptbot: Openai's web crawler. https:\/\/platform.openai.com\/docs\/gptbot."},{"key":"e_1_3_2_2_58_1","volume-title":"Controlling the extraction of memorized data from large language models via prompt-tuning. arXiv preprint arXiv:2305.11759","author":"Ozdayi Mustafa Safa","year":"2023","unstructured":"Mustafa Safa Ozdayi, Charith Peris, Jack FitzGerald, Christophe Dupuy, Jimit Majmudar, Haidar Khan, Rahil Parikh, and Rahul Gupta. Controlling the extraction of memorized data from large language models via prompt-tuning. arXiv preprint arXiv:2305.11759, 2023."},{"key":"e_1_3_2_2_59_1","volume-title":"robots.txt checker. https:\/\/pagedart.com\/tools\/robots-txt-file-checker\/","year":"2024","unstructured":"PageDart. robots.txt checker. https:\/\/pagedart.com\/tools\/robots-txt-file-checker\/, 2024."},{"key":"e_1_3_2_2_60_1","unstructured":"Guilherme Penedo Hynek Kydl\u00edek Leandro von Werra and Thomas Wolf. Fineweb 2024."},{"key":"e_1_3_2_2_61_1","volume-title":"Samaneh Tajalizadehkhoob, Maciej Korczyski, and Wouter Joosen. Tranco: A research-oriented top sites ranking hardened against manipulation. arXiv preprint arXiv:1806.01156","author":"Pochat Victor Le","year":"2018","unstructured":"Victor Le Pochat, Tom Van Goethem, Samaneh Tajalizadehkhoob, Maciej Korczyski, and Wouter Joosen. Tranco: A research-oriented top sites ranking hardened against manipulation. arXiv preprint arXiv:1806.01156, 2018."},{"key":"e_1_3_2_2_62_1","volume-title":"robotsparser - a parser for robots.txt files. https:\/\/docs.python.org\/3\/library\/urllib.robotparser.html","author":"Foundation Python Software","year":"2023","unstructured":"Python Software Foundation. robotsparser - a parser for robots.txt files. https:\/\/docs.python.org\/3\/library\/urllib.robotparser.html, 2023."},{"key":"e_1_3_2_2_63_1","volume-title":"Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1-67","author":"Raffel Colin","year":"2020","unstructured":"Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1-67, 2020."},{"key":"e_1_3_2_2_64_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-2099"},{"key":"e_1_3_2_2_65_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1410"},{"key":"e_1_3_2_2_66_1","volume-title":"An update on web publisher controls. https:\/\/blog.google\/technology\/ai\/an-update-on-web-publisher-controls\/","author":"Romain Danielle","year":"2023","unstructured":"Danielle Romain. An update on web publisher controls. https:\/\/blog.google\/technology\/ai\/an-update-on-web-publisher-controls\/, 2023."},{"key":"e_1_3_2_2_67_1","volume-title":"Yossi Adi, Jingyu Liu, Tal Remez, J\u00e9r\u00e9my Rapin, et al. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950","author":"Roziere Baptiste","year":"2023","unstructured":"Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, J\u00e9r\u00e9my Rapin, et al. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950, 2023."},{"key":"e_1_3_2_2_68_1","volume-title":"Undesirable memorization in large language models: A survey. arXiv preprint arXiv:2410.02650","author":"Satvaty Ali","year":"2024","unstructured":"Ali Satvaty, Suzan Verberne, and Fatih Turkmen. Undesirable memorization in large language models: A survey. arXiv preprint arXiv:2410.02650, 2024."},{"key":"e_1_3_2_2_69_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.clsr.2013.09.003"},{"key":"e_1_3_2_2_70_1","first-page":"25278","article-title":"Laion-5b: An open large-scale dataset for training next generation image-text models","volume":"35","author":"Schuhmann Christoph","year":"2022","unstructured":"Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35:25278-25294, 2022.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_2_71_1","volume-title":"Rethinking llm memorization through the lens of adversarial compression. arXiv preprint arXiv:2404.15146","author":"Schwarzschild Avi","year":"2024","unstructured":"Avi Schwarzschild, Zhili Feng, Pratyush Maini, Zachary C Lipton, and J Zico Kolter. Rethinking llm memorization through the lens of adversarial compression. arXiv preprint arXiv:2404.15146, 2024."},{"key":"e_1_3_2_2_72_1","volume-title":"Detecting pretraining data from large language models. arXiv preprint arXiv:2310.16789","author":"Shi Weijia","year":"2023","unstructured":"Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer. Detecting pretraining data from large language models. arXiv preprint arXiv:2310.16789, 2023."},{"key":"e_1_3_2_2_73_1","volume-title":"Privacy at scale: Introducing the privaseer corpus of web privacy policies. arXiv preprint arXiv:2004.11131","author":"Srinath Mukund","year":"2020","unstructured":"Mukund Srinath, Shomir Wilson, and C Lee Giles. Privacy at scale: Introducing the privaseer corpus of web privacy policies. arXiv preprint arXiv:2004.11131, 2020."},{"key":"e_1_3_2_2_74_1","doi-asserted-by":"publisher","DOI":"10.4324\/9780203856949-27"},{"key":"e_1_3_2_2_75_1","volume-title":"Extracting memorized training data via decomposition. arXiv preprint arXiv:2409.12367","author":"Su Ellen","year":"2024","unstructured":"Ellen Su, Anu Vellore, Amy Chang, Raffaele Mura, Blaine Nelson, Paul Kassianik, and Amin Karbasi. Extracting memorized training data via decomposition. arXiv preprint arXiv:2409.12367, 2024."},{"key":"e_1_3_2_2_76_1","doi-asserted-by":"publisher","DOI":"10.1109\/WI.2007.98"},{"key":"e_1_3_2_2_77_1","doi-asserted-by":"publisher","DOI":"10.1145\/1242572.1242726"},{"key":"e_1_3_2_2_78_1","doi-asserted-by":"publisher","DOI":"10.1145\/3469096.3474940"},{"key":"e_1_3_2_2_79_1","volume-title":"Ai crawlers and how to block them with robots.txt. https: \/\/blog.openreplay.com\/ai-crawlers-block-robots-txt\/","author":"Team OpenReplay","year":"2025","unstructured":"OpenReplay Team. Ai crawlers and how to block them with robots.txt. https: \/\/blog.openreplay.com\/ai-crawlers-block-robots-txt\/, 2025. Accessed on [insert your access date]."},{"key":"e_1_3_2_2_80_1","doi-asserted-by":"publisher","DOI":"10.1145\/3589334.3645719"},{"key":"e_1_3_2_2_81_1","volume-title":"Terms of service. https:\/\/www.thirteen.org\/about\/terms-of-service\/","author":"THIRTEEN.","year":"2024","unstructured":"THIRTEEN. Terms of service. https:\/\/www.thirteen.org\/about\/terms-of-service\/, 2024."},{"key":"e_1_3_2_2_82_1","volume-title":"Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239","author":"Thoppilan Romal","year":"2022","unstructured":"Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239, 2022."},{"key":"e_1_3_2_2_83_1","first-page":"38274","article-title":"Memorization without overfitting: Analyzing the training dynamics of large language models","volume":"35","author":"Tirumala Kushal","year":"2022","unstructured":"Kushal Tirumala, Aram Markosyan, Luke Zettlemoyer, and Armen Aghajanyan. Memorization without overfitting: Analyzing the training dynamics of large language models. Advances in Neural Information Processing Systems, 35:38274-38290, 2022.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_2_84_1","volume-title":"et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971","author":"Touvron Hugo","year":"2023","unstructured":"Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth\u00e9e Lacroix, Baptiste Rozi\u00e8re, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023."},{"key":"e_1_3_2_2_85_1","volume-title":"Database includes detailed information about every single user agent and operating system. https:\/\/udger.com\/resources\/ua-list\/crawlers","year":"2024","unstructured":"Udger. Database includes detailed information about every single user agent and operating system. https:\/\/udger.com\/resources\/ua-list\/crawlers, 2024."},{"key":"e_1_3_2_2_86_1","volume-title":"A list of known ai agents on the internet. https:\/\/darkvisitors.com","author":"Visitors Dark","year":"2024","unstructured":"Dark Visitors. A list of known ai agents on the internet. https:\/\/darkvisitors.com, 2024."},{"key":"e_1_3_2_2_87_1","volume-title":"Unlocking memorization in large language models with dynamic soft prompting. arXiv preprint arXiv:2409.13853","author":"Wang Zhepeng","year":"2024","unstructured":"Zhepeng Wang, Runxue Bao, Yawen Wu, Jackson Taylor, Cao Xiao, Feng Zheng, Weiwen Jiang, Shangqian Gao, and Yanfu Zhang. Unlocking memorization in large language models with dynamic soft prompting. arXiv preprint arXiv:2409.13853, 2024."},{"key":"e_1_3_2_2_88_1","volume-title":"Ccnet: Extracting high quality monolingual datasets from web crawl data. arXiv preprint arXiv:1911.00359","author":"Wenzek Guillaume","year":"2019","unstructured":"Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzm\u00e1n, Armand Joulin, and Edouard Grave. Ccnet: Extracting high quality monolingual datasets from web crawl data. arXiv preprint arXiv:1911.00359, 2019."},{"key":"e_1_3_2_2_89_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-demos.6"},{"key":"e_1_3_2_2_90_1","volume-title":"et al. Open-finllms: Open multimodal large language models for financial applications. arXiv preprint arXiv:2408.11878","author":"Xie Qianqian","year":"2024","unstructured":"Qianqian Xie, Dong Li, Mengxi Xiao, Zihao Jiang, Ruoyu Xiang, Xiao Zhang, Zhengyu Chen, Yueru He, Weiguang Han, Yuzhe Yang, et al. Open-finllms: Open multimodal large language models for financial applications. arXiv preprint arXiv:2408.11878, 2024."},{"key":"e_1_3_2_2_91_1","unstructured":"Yandex Support. How to Make Sure That a Robot Belongs to Yandex. https:\/\/yandex.com\/support\/webmaster\/robot-workings\/check-yandex-robots.html."},{"key":"e_1_3_2_2_92_1","unstructured":"You.com. YouBot. https:\/\/web.archive.org\/web\/20240423043032\/https:\/\/about.you.com\/youbot\/."},{"key":"e_1_3_2_2_93_1","first-page":"39321","article-title":"Counterfactual memorization in neural language models","volume":"36","author":"Zhang Chiyuan","year":"2023","unstructured":"Chiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski, Florian Tram\u00e8r, and Nicholas Carlini. Counterfactual memorization in neural language models. Advances in Neural Information Processing Systems, 36:39321-39362, 2023.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_2_94_1","volume-title":"Instruction tuning for large language models: A survey. arXiv preprint arXiv:2308.10792","author":"Zhang Shengyu","year":"2023","unstructured":"Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, et al. Instruction tuning for large language models: A survey. arXiv preprint arXiv:2308.10792, 2023."},{"key":"e_1_3_2_2_95_1","volume-title":"Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. Opt: Open pre-trained transformer language models","author":"Zhang Susan","year":"2022","unstructured":"Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. Opt: Open pre-trained transformer language models, 2022."},{"key":"e_1_3_2_2_96_1","volume-title":"Ilyas Chahed, Younes Belkada, Guillaume Kunsch, and Hakim Hacid. Falcon mamba: The first competitive attention-free 7b language model. arXiv preprint arXiv:2410.05355","author":"Zuo Jingwei","year":"2024","unstructured":"Jingwei Zuo, Maksim Velikanov, Dhia Eddine Rhaiem, Ilyas Chahed, Younes Belkada, Guillaume Kunsch, and Hakim Hacid. Falcon mamba: The first competitive attention-free 7b language model. arXiv preprint arXiv:2410.05355, 2024."}],"event":{"name":"CCS '25: ACM SIGSAC Conference on Computer and Communications Security","location":"Taipei Taiwan","acronym":"CCS '25","sponsor":["SIGSAC ACM Special Interest Group on Security, Audit, and Control"]},"container-title":["Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3719027.3765063","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,22]],"date-time":"2025-12-22T22:23:41Z","timestamp":1766442221000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3719027.3765063"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,19]]},"references-count":96,"alternative-id":["10.1145\/3719027.3765063","10.1145\/3719027"],"URL":"https:\/\/doi.org\/10.1145\/3719027.3765063","relation":{},"subject":[],"published":{"date-parts":[[2025,11,19]]},"assertion":[{"value":"2025-11-22","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}