{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T09:46:16Z","timestamp":1782380776942,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":138,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T00:00:00Z","timestamp":1776038400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,4,13]]},"DOI":"10.1145\/3772318.3791344","type":"proceedings-article","created":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T04:12:26Z","timestamp":1776053546000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["From Reflection to Repair: A Scoping Review of Dataset Documentation Tools"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5690-1850","authenticated-orcid":false,"given":"Pedro","family":"Reynolds-Cu\u00e9llar","sequence":"first","affiliation":[{"name":"Robotics, Ethics and Society, Robotics and AI Institute, Cambridge, Massachusetts, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8071-6716","authenticated-orcid":false,"given":"Marisol","family":"Wong-Villacres","sequence":"additional","affiliation":[{"name":"Facultad de Ingenier\u00eda en Electricidad y Computaci\u00f3n, Escuela Superior Polit\u00e9cnica del Litoral, Guayaquil, Ecuador"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4230-3777","authenticated-orcid":false,"given":"Adriana","family":"Alvarado Garcia","sequence":"additional","affiliation":[{"name":"IBM Research, Yorktown Heights, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-2580-7040","authenticated-orcid":false,"given":"Heila","family":"Precel","sequence":"additional","affiliation":[{"name":"Faculty of Computing &amp; Data Sciences, Boston University, Boston, Massachusetts, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,4,13]]},"reference":[{"key":"e_1_3_3_2_2_2","unstructured":"[n. d.]. Andersen v. Stability AI Ltd. 3:23-Cv-00201 (N.D. Cal.). https:\/\/www.courtlistener.com\/docket\/66732129\/andersen-v-stability-ai-ltd\/"},{"key":"e_1_3_3_2_3_2","unstructured":"[n. d.]. The Common Crawl. https:\/\/commoncrawl.org\/overview"},{"key":"e_1_3_3_2_4_2","unstructured":"2024. Concord Music Group Inc. v. Anthropic PBC 3:23-Cv-01092 (M.D. Tenn.). https:\/\/www.courtlistener.com\/docket\/67894459\/concord-music-group-inc-v-anthropic-pbc\/"},{"key":"e_1_3_3_2_5_2","unstructured":"[n. d.]. Dow Jones & Company Inc. v. Perplexity AI Inc. 1:24-Cv-07984 (S.D.N.Y.). https:\/\/www.courtlistener.com\/docket\/69280523\/dow-jones-company-inc-v-perplexity-ai-inc\/"},{"key":"e_1_3_3_2_6_2","volume-title":"Hugging Face Dataset Cards","unstructured":"Hugging Face Dataset Cards [n. d.]. Hugging Face Dataset Cards. Hugging Face Dataset Cards. https:\/\/huggingface.co\/docs\/hub\/datasets-cards"},{"key":"e_1_3_3_2_7_2","unstructured":"[n. d.]. In Re: OpenAI Inc. Copyright Infringement Litigation 1:25-Md-03143 (S.D.N.Y.). https:\/\/www.courtlistener.com\/docket\/69879510\/in-re-openai-inc-copyright-infringement-litigation\/"},{"key":"e_1_3_3_2_8_2","volume-title":"NeurIPS Code of Ethics","unstructured":"NeurIPS Code of Ethics [n. d.]. NeurIPS Code of Ethics. NeurIPS Code of Ethics. https:\/\/nips.cc\/public\/EthicsGuidelines"},{"key":"e_1_3_3_2_9_2","unstructured":"2024. . Number AB 2013. https:\/\/leginfo.legislature.ca.gov\/faces\/billNavClient.xhtml?bill_id=202320240AB2013"},{"key":"e_1_3_3_2_10_2","unstructured":"2024. . Number SB24-205. https:\/\/leg.colorado.gov\/sites\/default\/files\/2024a_205_signed.pdf"},{"key":"e_1_3_3_2_11_2","doi-asserted-by":"publisher","unstructured":"Mohsen Abbasi Sorelle\u00a0A. Friedler Carlos Scheidegger and Suresh Venkatasubramanian. 2019. Fairness in representation: 19th SIAM International Conference on Data Mining SDM 2019. SIAM International Conference on Data Mining SDM 2019 (2019) 801\u2013809. 10.1137\/1.9781611975673.90","DOI":"10.1137\/1.9781611975673.90"},{"key":"e_1_3_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3522664.3528600"},{"key":"e_1_3_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/SMDS53860.2021.00016"},{"key":"e_1_3_3_2_14_2","doi-asserted-by":"publisher","unstructured":"Raia\u00a0Abu Ahmad Jennifer D\u2019Souza Matth\u00e4us Zloch Wolfgang Otto Georg Rehm Allard Oelen Stefan Dietze and S\u00f6ren Auer. 2024. Toward FAIR Semantic Publishing of Research Dataset Metadata in the Open Research Knowledge Graph. 10.48550\/arXiv.2404.08443arXiv:https:\/\/arXiv.org\/abs\/2404.08443 [cs].","DOI":"10.48550\/arXiv.2404.08443"},{"key":"e_1_3_3_2_15_2","unstructured":"Nur Ahmed and Neil\u00a0C. Thompson. 2023. What should be done about the growing influence of industry in AI research?https:\/\/www.brookings.edu\/articles\/what-should-be-done-about-the-growing-influence-of-industry-in-ai-research"},{"key":"e_1_3_3_2_16_2","doi-asserted-by":"publisher","unstructured":"Nur Ahmed Muntasir Wahed and Neil\u00a0C. Thompson. 2023. The growing influence of industry in AI research. Science 379 6635 (2023) 884\u2013886. arXiv:https:\/\/www.science.org\/doi\/pdf\/10.1126\/science.ade242010.1126\/science.ade2420","DOI":"10.1126\/science.ade2420"},{"key":"e_1_3_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3650203.3663326"},{"key":"e_1_3_3_2_18_2","doi-asserted-by":"publisher","unstructured":"Joseph\u00a0E Alderman Joanne Palmer Elinor Laws Melissa\u00a0D McCradden Johan Ordish Marzyeh Ghassemi Stephen\u00a0R Pfohl Negar Rostamzadeh Heather Cole-Lewis Ben Glocker Melanie Calvert Tom\u00a0J Pollard Jaspret Gill Jacqui Gath Adewale Adebajo Jude Beng Cassandra\u00a0H Leung Stephanie Kuku Lesley-Anne Farmer Rubeta\u00a0N Matin Bilal\u00a0A Mateen Francis McKay Katherine Heller Alan Karthikesalingam Darren Treanor Maxine Mackintosh Lauren Oakden-Rayner Russell Pearson Arjun\u00a0K Manrai Puja Myles Judit Kumuthini Zoher Kapacee Neil\u00a0J Sebire Lama\u00a0H Nazer Jarrel Seah Ashley Akbari Lew Berman Judy\u00a0W Gichoya Lorenzo Righetto Diana Samuel William Wasswa Maria Charalambides Anmol Arora Sameer Pujari Charlotte Summers Elizabeth Sapey Sharon Wilkinson Vishal Thakker Alastair Denniston and Xiaoxuan Liu. 2025. Tackling algorithmic bias and promoting transparency in health datasets: the STANDING Together consensus recommendations. The Lancet Digital Health 7 1 (2025) e64\u2013e88. 10.1016\/S2589-7500(24)00224-3","DOI":"10.1016\/S2589-7500(24)00224-3"},{"key":"e_1_3_3_2_19_2","doi-asserted-by":"publisher","unstructured":"Victor Alencar Troy Kohwalter Vanessa Braganholo Jos\u00e9\u00a0Ricardo da Silva and Leonardo Murta. 2024. Prov-Dominoes: An approach for knowledge discovery from provenance data. Expert Systems with Applications 245 (2024) 123030. 10.1016\/j.eswa.2023.123030","DOI":"10.1016\/j.eswa.2023.123030"},{"key":"e_1_3_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3593013.3593990"},{"key":"e_1_3_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3706598.3714069"},{"key":"e_1_3_3_2_22_2","doi-asserted-by":"publisher","unstructured":"Hilary Arksey and Lisa O\u2019Malley. 2005. Scoping studies: towards a methodological framework. International Journal of Social Research Methodology 8 1 (Feb. 2005) 19\u201332. 10.1080\/1364557032000119616","DOI":"10.1080\/1364557032000119616"},{"key":"e_1_3_3_2_23_2","doi-asserted-by":"publisher","unstructured":"M. Arnold R.\u00a0K.\u00a0E. Bellamy M. Hind S. Houde S. Mehta A. Mojsilovi\u0107 R. Nair K.\u00a0Natesan Ramamurthy A. Olteanu D. Piorkowski D. Reimer J. Richards J. Tsay and K.\u00a0R. Varshney. 2019. FactSheets: Increasing trust in AI services through supplier\u2019s declarations of conformity. IBM Journal of Research and Development 63 4\/5 (july 2019) 6:1\u20136:13. 10.1147\/JRD.2019.2942288","DOI":"10.1147\/JRD.2019.2942288"},{"key":"e_1_3_3_2_24_2","doi-asserted-by":"publisher","unstructured":"Ruben\u00a0C Arslan. 2019. How to automatically document data with the codebook package to facilitate data reuse. Advances in Methods and Practices in Psychological Science 2 2 (2019) 169\u2013187. 10.1177\/2515245919838783","DOI":"10.1177\/2515245919838783"},{"key":"e_1_3_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510548.3519376"},{"key":"e_1_3_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3600211.3604674"},{"key":"e_1_3_3_2_27_2","doi-asserted-by":"publisher","unstructured":"Jack Bandy and Nicholas Vincent. 2021. Addressing \u201cDocumentation Debt\u201d in Machine Learning Research: A Retrospective Datasheet for BookCorpus. arXiv:2105.05241 (May 2021). 10.48550\/arXiv.2105.05241","DOI":"10.48550\/arXiv.2105.05241"},{"key":"e_1_3_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/QoMEX58391.2023.10178546"},{"key":"e_1_3_3_2_29_2","doi-asserted-by":"publisher","unstructured":"Emily\u00a0M. Bender and Batya Friedman. 2018. Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science. Transactions of the Association for Computational Linguistics 6 (Dec. 2018) 587\u2013604. 10.1162\/tacl_a_00041","DOI":"10.1162\/tacl_a_00041"},{"key":"e_1_3_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1145\/3442188.3445922"},{"key":"e_1_3_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3630106.3658955"},{"key":"e_1_3_3_2_32_2","doi-asserted-by":"crossref","unstructured":"Eshta Bhardwaj Harshit Gujral Siyi Wu Ciara Zogheib Tegan Maharaj and Christoph Becker. 2024. The State of Data Curation at NeurIPS: An Assessment of Dataset Development Practices in the Datasets and Benchmarks Track. Advances in Neural Information Processing Systems 37 (Dec. 2024) 53626\u201353648.","DOI":"10.52202\/079017-1698"},{"key":"e_1_3_3_2_33_2","volume-title":"Advances in Neural Information Processing Systems","author":"Bolukbasi Tolga","year":"2016","unstructured":"Tolga Bolukbasi, Kai-Wei Chang, James\u00a0Y Zou, Venkatesh Saligrama, and Adam\u00a0T Kalai. 2016. Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings. In Advances in Neural Information Processing Systems , Vol.\u00a029. Curran Associates, Inc.https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2016\/hash\/a486cd07e4ac3d270571622f4f316ec5-Abstract.html"},{"key":"e_1_3_3_2_34_2","unstructured":"Rishi Bommasani Kevin Klyman Shayne Longpre Sayash Kapoor Nestor Maslej Betty Xiong Daniel Zhang and Percy Liang. 2023. The foundation model transparency index. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2310.12941 (2023)."},{"key":"e_1_3_3_2_35_2","doi-asserted-by":"publisher","unstructured":"Karen\u00a0L. Boyd. 2021. Datasheets for Datasets help ML Engineers Notice and Understand Ethical Issues in Training Data. Proc. ACM Hum.-Comput. Interact. 5 CSCW2 (Oct. 2021) 438:1\u2013438:27. 10.1145\/3479582","DOI":"10.1145\/3479582"},{"key":"e_1_3_3_2_36_2","doi-asserted-by":"publisher","unstructured":"Virginia Braun and Victoria Clarke. 2021. One size fits all? What counts as quality practice in (reflexive) thematic analysis? Qualitative Research in Psychology 18 3 (July 2021) 328\u2013352. 10.1080\/14780887.2020.1769238","DOI":"10.1080\/14780887.2020.1769238"},{"key":"e_1_3_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-69909-7_3470-2"},{"key":"e_1_3_3_2_38_2","unstructured":"Joy\u00a0Adowaa Buolamwini. 2017. Gender shades: intersectional phenotypic and demographic evaluation of face datasets and gender classifiers. Thesis. Massachusetts Institute of Technology. https:\/\/dspace.mit.edu\/handle\/1721.1\/114068 Accepted: 2018-03-12T19:28:30Z."},{"key":"e_1_3_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-43823-4_1"},{"key":"e_1_3_3_2_40_2","unstructured":"Kasia Chmielinski Sarah Newman Chris\u00a0N. Kranzinger Michael Hind Jennifer\u00a0Wortman Vaughan Margaret Mitchell Julia Stoyanovich Angelina McMillan-Major Emily McReynolds Kathleen Esfahany Mary\u00a0L. Gray Maui Hudson and Audrey Chang. 2024."},{"key":"e_1_3_3_2_41_2","volume-title":"The Dataset Nutrition Label (2nd Gen): Leveraging Context to Mitigate Harms in Artificial Intelligence","author":"Chmielinski Kasia\u00a0S.","year":"2022","unstructured":"Kasia\u00a0S. Chmielinski, Sarah Newman, Matt Taylor, Josh Joseph, Kemi Thomas, Jessica Yurkofsky, and Yue\u00a0Chelsea Qiu. 2022. The Dataset Nutrition Label (2nd Gen): Leveraging Context to Mitigate Harms in Artificial Intelligence. arXiv:https:\/\/arXiv.org\/abs\/2201.03954\u00a0[cs] http:\/\/arxiv.org\/abs\/2201.03954"},{"key":"e_1_3_3_2_42_2","doi-asserted-by":"publisher","unstructured":"Robert Cinca Enrico Costanza and Mirco Musolesi. 2025. Practitioners and Bias in Machine Learning: A Study. ACM Trans. Interact. Intell. Syst. 15 2 (June 2025) 12:1\u201312:28. 10.1145\/3733838","DOI":"10.1145\/3733838"},{"key":"e_1_3_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533143"},{"key":"e_1_3_3_2_44_2","unstructured":"Kate Crawford. 2017. The Trouble with Bias. https:\/\/www.youtube.com\/watch?v=fMym_BKWQzk"},{"key":"e_1_3_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533108"},{"key":"e_1_3_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3463274.3463362"},{"key":"e_1_3_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3544548.3581026"},{"key":"e_1_3_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533113"},{"key":"e_1_3_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3593013.3594037"},{"key":"e_1_3_3_2_50_2","doi-asserted-by":"publisher","unstructured":"Nouha Dziri Sivan Milton Mo Yu Osmar Zaiane and Siva Reddy. 2022. On the Origin of Hallucinations in Conversational Models: Is it the Datasets or the Models?arXiv:2204.07931 (April 2022). 10.48550\/arXiv.2204.07931","DOI":"10.48550\/arXiv.2204.07931"},{"key":"e_1_3_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3534647"},{"key":"e_1_3_3_2_52_2","unstructured":"Brian Eastwood. 2023. Study: Industry now dominates AI Research. https:\/\/mitsloan.mit.edu\/ideas-made-to-matter\/study-industry-now-dominates-ai-research"},{"key":"e_1_3_3_2_53_2","doi-asserted-by":"publisher","unstructured":"Tyna Eloundou Sam Manning Pamela Mishkin and Daniel Rock. 2023. GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models. arXiv:2303.10130 (Aug. 2023). 10.48550\/arXiv.2303.10130","DOI":"10.48550\/arXiv.2303.10130"},{"key":"e_1_3_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/3551624.3555286"},{"key":"e_1_3_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3627673.3679157"},{"key":"e_1_3_3_2_56_2","unstructured":"Timnit Gebru Jamie Morgenstern Briana Vecchione Jennifer\u00a0Wortman Vaughan Hanna Wallach Hal Daum\u00e9\u00a0III and Kate Crawford. 2021. Datasheets for Datasets. http:\/\/arxiv.org\/abs\/1803.09010 arXiv:https:\/\/arXiv.org\/abs\/1803.09010 [cs]."},{"key":"e_1_3_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.12987\/9780300235029"},{"key":"e_1_3_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3550356.3559087"},{"key":"e_1_3_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1145\/3583780.3614737"},{"key":"e_1_3_3_2_60_2","doi-asserted-by":"publisher","unstructured":"Joan Giner-Miguelez Abel G\u00f3mez and Jordi Cabot. 2023. A domain-specific language for describing machine learning datasets. Journal of Computer Languages 76 (2023) 101209. 10.1016\/j.cola.2023.101209","DOI":"10.1016\/j.cola.2023.101209"},{"key":"e_1_3_3_2_61_2","doi-asserted-by":"publisher","unstructured":"Joan Giner-Miguelez Abel G\u00f3mez and Jordi Cabot. 2024. Using Large Language Models to Enrich the Documentation of Datasets for Machine Learning. 10.48550\/arXiv.2404.15320arXiv:https:\/\/arXiv.org\/abs\/2404.15320 [cs].","DOI":"10.48550\/arXiv.2404.15320"},{"key":"e_1_3_3_2_62_2","volume-title":"Ghost Work: How to Stop Silicon Valley from Building a New Global Underclass","author":"Gray Mary\u00a0L.","year":"2019","unstructured":"Mary\u00a0L. Gray and Siddharth Suri. 2019. Ghost Work: How to Stop Silicon Valley from Building a New Global Underclass. Houghton Mifflin Harcourt."},{"key":"e_1_3_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2903730"},{"key":"e_1_3_3_2_64_2","doi-asserted-by":"publisher","unstructured":"Amy\u00a0K. Heger Liz\u00a0B. Marquis Mihaela Vorvoreanu Hanna Wallach and Jennifer Wortman\u00a0Vaughan. 2022. Understanding Machine Learning Practitioners\u2019 Data Documentation Perceptions Needs Challenges and Desiderata. Proc. ACM Hum.-Comput. Interact. 6 CSCW2 (Nov. 2022) 340:1\u2013340:29. 10.1145\/3555760","DOI":"10.1145\/3555760"},{"key":"e_1_3_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/3599733.3600249"},{"key":"e_1_3_3_2_66_2","doi-asserted-by":"publisher","unstructured":"Sarah Holland Ahmed Hosny Sarah Newman Joshua Joseph and Kasia Chmielinski. 2018. The Dataset Nutrition Label: A Framework To Drive Higher Data Quality Standards. 10.48550\/arXiv.1805.03677arXiv:https:\/\/arXiv.org\/abs\/1805.03677 [cs].","DOI":"10.48550\/arXiv.1805.03677"},{"key":"e_1_3_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1145\/3290605.3300830"},{"key":"e_1_3_3_2_68_2","doi-asserted-by":"publisher","unstructured":"Rachel Hong Jevan Hutson William Agnew Imaad Huda Tadayoshi Kohno and Jamie Morgenstern. 2025. A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset. arXiv:2506.17185 (june 2025). 10.48550\/arXiv.2506.17185arXiv:https:\/\/arXiv.org\/abs\/2506.17185 [cs].","DOI":"10.48550\/arXiv.2506.17185"},{"key":"e_1_3_3_2_69_2","doi-asserted-by":"publisher","unstructured":"Graeme Horsman and James\u00a0R. Lyle. 2021. Dataset construction challenges for digital forensics. Forensic Science International: Digital Investigation 38 (2021) 301264. 10.1016\/j.fsidi.2021.301264","DOI":"10.1016\/j.fsidi.2021.301264"},{"key":"e_1_3_3_2_70_2","unstructured":"The\u00a0White House. 2025. America\u2019s AI Action Plan. (july 2025). https:\/\/www.whitehouse.gov\/wp-content\/uploads\/2025\/07\/Americas-AI-Action-Plan.pdf"},{"key":"e_1_3_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1145\/3442188.3445918"},{"key":"e_1_3_3_2_72_2","doi-asserted-by":"publisher","unstructured":"Nitisha Jain Mubashara Akhtar Joan Giner-Miguelez Rajat Shinde Joaquin Vanschoren Steffen Vogler Sujata Goswami Yuhan Rao Tim Santos Luis Oala Michalis Karamousadakis Manil Maskey Pierre Marcenac Costanza Conforti Michael Kuchnik Lora Aroyo Omar Benjelloun and Elena Simperl. 2024. A Standardized Machine-readable Dataset Documentation Format for Responsible AI. arxiv:https:\/\/arXiv.org\/abs\/2407.16883\u00a0[cs] 10.48550\/arXiv.2407.16883","DOI":"10.48550\/arXiv.2407.16883"},{"key":"e_1_3_3_2_73_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642111"},{"key":"e_1_3_3_2_74_2","doi-asserted-by":"publisher","unstructured":"Nikhil Kandpal and Colin Raffel. 2025. Position: The Most Expensive Part of an LLM should be its Training Data. arXiv:2504.12427 (April 2025). 10.48550\/arXiv.2504.12427","DOI":"10.48550\/arXiv.2504.12427"},{"key":"e_1_3_3_2_75_2","doi-asserted-by":"publisher","DOI":"10.1145\/3322276.3323691"},{"key":"e_1_3_3_2_76_2","doi-asserted-by":"publisher","DOI":"10.1145\/3411764.3445261"},{"key":"e_1_3_3_2_77_2","doi-asserted-by":"publisher","DOI":"10.1145\/3411764.3445261"},{"key":"e_1_3_3_2_78_2","doi-asserted-by":"publisher","unstructured":"Shayne Longpre Robert Mahari Anthony Chen Naana Obeng-Marnu Damien Sileo William Brannon Niklas Muennighoff Nathan Khazam Jad Kabbara Kartik Perisetla Xinyi\u00a0(Alexis) Wu Enrico Shippole Kurt Bollacker Tongshuang Wu Luis Villa Sandy Pentland and Sara Hooker. 2024. A large-scale audit of dataset licensing and attribution in AI. Nature Machine Intelligence 6 8 (Aug. 2024) 975\u2013987. 10.1038\/s42256-024-00878-8","DOI":"10.1038\/s42256-024-00878-8"},{"key":"e_1_3_3_2_79_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-short.24"},{"key":"e_1_3_3_2_80_2","doi-asserted-by":"publisher","DOI":"10.63317\/3w9yw2grx32r"},{"key":"e_1_3_3_2_81_2","doi-asserted-by":"publisher","unstructured":"Michael Madaio Lisa Egede Hariharan Subramonyam Jennifer Wortman\u00a0Vaughan and Hanna Wallach. 2022. Assessing the Fairness of AI Systems: AI Practitioners\u2019 Processes Challenges and Needs for Support. Proc. ACM Hum.-Comput. Interact. 6 CSCW1 Article 52 (April 2022) 26\u00a0pages. 10.1145\/3512899","DOI":"10.1145\/3512899"},{"key":"e_1_3_3_2_82_2","doi-asserted-by":"publisher","unstructured":"Michael\u00a0A. Madaio Jingya Chen Hanna Wallach and Jennifer Wortman\u00a0Vaughan. 2024. Tinker Tailor Configure Customize: The Articulation Work of Contextualizing an AI Fairness Checklist. Proceedings of the ACM on Human-Computer Interaction 8 CSCW1 (April 2024) 1\u201320. 10.1145\/3653705","DOI":"10.1145\/3653705"},{"key":"e_1_3_3_2_83_2","doi-asserted-by":"publisher","DOI":"10.1145\/3313831.3376445"},{"key":"e_1_3_3_2_84_2","unstructured":"Laura Manduchi Clara Meister Kushagra Pandey Robert Bamler Ryan Cotterell Sina D\u00e4ubener Sophie Fellenz Asja Fischer Thomas G\u00e4rtner Matthias Kirchler et\u00a0al. 2024. On the challenges and opportunities in generative ai. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2403.00025 (2024)."},{"key":"e_1_3_3_2_85_2","doi-asserted-by":"publisher","unstructured":"Ramtin\u00a0Zargari Marandi Anne\u00a0Svane Frahm and Maja Milojevic. 2025. Datasheets for AI and medical datasets (DAIMS): a data validation and documentation framework before machine learning analysis in medical research. 10.48550\/arXiv.2501.14094arXiv:https:\/\/arXiv.org\/abs\/2501.14094 [cs].","DOI":"10.48550\/arXiv.2501.14094"},{"key":"e_1_3_3_2_86_2","volume-title":"Data Statements | Tech Policy Lab","author":"McMillan-Major Angelina","year":"2023","unstructured":"Angelina McMillan-Major and Emily\u00a0M. Bender. 2023. Data Statements | Tech Policy Lab. Technical Report. University of Washington. https:\/\/techpolicylab.uw.edu\/data-statements\/"},{"key":"e_1_3_3_2_87_2","doi-asserted-by":"publisher","unstructured":"Angelina McMillan-Major Emily\u00a0M. Bender and Batya Friedman. 2024. Data Statements: From Technical Concept to Community Practice. ACM J. Responsib. Comput. 1 1 (March 2024) 1:1\u20131:17. 10.1145\/3594737","DOI":"10.1145\/3594737"},{"key":"e_1_3_3_2_88_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.gem-1.11"},{"key":"e_1_3_3_2_89_2","unstructured":"Angelina\u00a0Yvonne McMillan-Major. 2023. Language Dataset Documentation Design: Learning from Deaf and Indigenous Communities. hdl:1773\/50854http:\/\/hdl.handle.net\/1773\/50854"},{"key":"e_1_3_3_2_90_2","doi-asserted-by":"publisher","unstructured":"Jacob Metcalf Emanuel Moss and Danah Boyd. 2019. Owning ethics: Corporate logics Silicon Valley and the institutionalization of ethics. Social Research: An International Quarterly 86 2 (2019) 449\u2013476. 10.1353\/sor.2019.0022","DOI":"10.1353\/sor.2019.0022"},{"key":"e_1_3_3_2_91_2","unstructured":"Microsoft. 2022. Aether Data Documentation Template. https:\/\/www.microsoft.com\/en-us\/research\/wp-content\/uploads\/2022\/07\/aether-datadoc-082522.pdf"},{"key":"e_1_3_3_2_92_2","unstructured":"Microsoft. 2022. Microsoft RAI Impact Assessment Template. https:\/\/msblogs.thesourcemediaassets.com\/sites\/5\/2022\/06\/Microsoft-RAI-Impact-Assessment-Template.pdf"},{"key":"e_1_3_3_2_93_2","doi-asserted-by":"publisher","unstructured":"Surbhi Mittal Kartik Thakral Richa Singh Mayank Vatsa Tamar Glaser Cristian\u00a0Canton Ferrer and Tal Hassner. 2024. On Responsible Machine Learning Datasets with Fairness Privacy and Regulatory Norms. 10.48550\/arXiv.2310.15848arXiv:https:\/\/arXiv.org\/abs\/2310.15848 [cs].","DOI":"10.48550\/arXiv.2310.15848"},{"key":"e_1_3_3_2_94_2","doi-asserted-by":"publisher","DOI":"10.1109\/INDIN58382.2024.10774524"},{"key":"e_1_3_3_2_95_2","unstructured":"Global Future\u00a0Council on Human\u00a0Rights. 2018. How to Prevent Discriminatory Outcomes in Machine Learning. https:\/\/www3.weforum.org\/docs\/WEF_40065_White_Paper_How_to_Prevent_Discriminatory_Outcomes_in_Machine_Learning.pdf"},{"key":"e_1_3_3_2_96_2","doi-asserted-by":"publisher","unstructured":"Will Orr and Kate Crawford. 2024. The social construction of datasets: On the practices processes and challenges of dataset creation for machine learning. 26 (sept 2024) 4955\u20134972. 10.1177\/14614448241251797","DOI":"10.1177\/14614448241251797"},{"key":"e_1_3_3_2_97_2","doi-asserted-by":"publisher","unstructured":"Leon\u00a0J. Osterweil Lori\u00a0A. Clarke Aaron\u00a0M. Ellison Emery Boose Rodion Podorozhny and Alexander Wise. 2010. Clear and precise specification of ecological data management processes and dataset provenance. IEEE Transactions on Automation Science and Engineering 7 1 (Jan. 2010) 189\u2013195. 10.1109\/TASE.2009.2021774","DOI":"10.1109\/TASE.2009.2021774"},{"key":"e_1_3_3_2_98_2","doi-asserted-by":"publisher","DOI":"10.1145\/3593013.3594049"},{"key":"e_1_3_3_2_99_2","doi-asserted-by":"publisher","unstructured":"Amandalynne Paullada Inioluwa\u00a0Deborah Raji Emily\u00a0M. Bender Emily Denton and Alex Hanna. 2021. Data and its (dis)contents: A survey of dataset development and use in machine learning research. Patterns 2 11 (Nov. 2021) 100336. 10.1016\/j.patter.2021.100336","DOI":"10.1016\/j.patter.2021.100336"},{"key":"e_1_3_3_2_100_2","volume-title":"Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks","volume":"1","author":"Peng Kenneth","year":"2021","unstructured":"Kenneth Peng, Arunesh Mathur, and Arvind Narayanan. 2021. Mitigating dataset harms requires stewardship: Lessons from 1000 papers. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks , J.\u00a0Vanschoren and S.\u00a0Yeung (Eds.), Vol.\u00a01. https:\/\/datasets-benchmarks-proceedings.neurips.cc\/paper_files\/paper\/2021\/file\/077e29b11be80ab57e1a2ecabb7da330-Paper-round2.pdf"},{"key":"e_1_3_3_2_101_2","doi-asserted-by":"publisher","unstructured":"Anne\u00a0Helby Petersen and Claus\u00a0Thorn Ekstr\u00f8m. 2019. dataMaid: Your Assistant for Documenting Supervised Data Quality Screening in R. Journal of Statistical Software 90 (July 2019) 1\u201338. 10.18637\/jss.v090.i06","DOI":"10.18637\/jss.v090.i06"},{"key":"e_1_3_3_2_102_2","series-title":"Theory on Demand","volume-title":"Economies of Virtue: The Circulation of \u2018Ethics\u2019 in AI","author":"Phan Thao","year":"2022","unstructured":"Thao Phan, Jake Goldenfein, Declan Kuch, and Monique Mann (Eds.). 2022. Economies of Virtue: The Circulation of \u2018Ethics\u2019 in AI. Theory on Demand, Vol.\u00a046. Institute of Network Cultures, Amsterdam. https:\/\/networkcultures.org\/blog\/publication\/economies-of-virtue-the-circulation-of-ethics-in-ai\/"},{"key":"e_1_3_3_2_103_2","doi-asserted-by":"publisher","unstructured":"Thao Phan Jake Goldenfein Monique Mann and Declan Kuch. 2021. Economies of Virtue: The Circulation of \u2018Ethics\u2019 Big Tech. Science as Culture 31 1 (2021) 121\u2013135. 10.1080\/09505431.2021.1990875","DOI":"10.1080\/09505431.2021.1990875"},{"key":"e_1_3_3_2_104_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSREW51248.2020.00085"},{"key":"e_1_3_3_2_105_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2201.13224"},{"key":"e_1_3_3_2_106_2","unstructured":"Vyacheslav Polonski. 2018. The hard problem of AI ethics\u2014Three guidelines for building morality into machines. https:\/\/www.oecd-forum.org\/posts\/30743-the-hard-problem-of-ai-ethics-three-guidelines-for-building-morality-into-machines. OECD Forum."},{"key":"e_1_3_3_2_107_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642749"},{"key":"e_1_3_3_2_108_2","doi-asserted-by":"publisher","unstructured":"Erich Prem. 2023. From ethical AI frameworks to tools: a review of approaches. 3 (Aug. 2023) 699\u2013716. 10.1007\/s43681-023-00258-9","DOI":"10.1007\/s43681-023-00258-9"},{"key":"e_1_3_3_2_109_2","unstructured":"Executive Office of\u00a0the President. 2023. Safe Secure and Trustworthy Development and Use of Artificial Intelligence. https:\/\/www.federalregister.gov\/documents\/2023\/11\/01\/2023-24283\/safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence (2023)."},{"key":"e_1_3_3_2_110_2","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533231"},{"key":"e_1_3_3_2_111_2","doi-asserted-by":"publisher","unstructured":"Xiangyu Qi Yi Zeng Tinghao Xie Pin-Yu Chen Ruoxi Jia Prateek Mittal and Peter Henderson. 2023. Fine-tuning Aligned Language Models Compromises Safety Even When Users Do Not Intend To!arXiv:2310.03693 (Oct. 2023). 10.48550\/arXiv.2310.03693","DOI":"10.48550\/arXiv.2310.03693"},{"key":"e_1_3_3_2_112_2","doi-asserted-by":"publisher","unstructured":"Bogdana Rakova Jingying Yang Henriette Cramer and Rumman Chowdhury. 2021. Where Responsible AI meets Reality: Practitioner Perspectives on Enablers for Shifting Organizational Practices. Proc. ACM Hum.-Comput. Interact. 5 CSCW1 Article 7 (April 2021) 23\u00a0pages. 10.1145\/3449081","DOI":"10.1145\/3449081"},{"key":"e_1_3_3_2_113_2","first-page":"51","volume-title":"Proceedings of the 21st annual workshop of the australasian language technology association","author":"Reid Kathy","year":"2023","unstructured":"Kathy Reid and Elizabeth\u00a0T. Williams. 2023. Right the docs: Characterising voice dataset documentation practices used in machine learning. In Proceedings of the 21st annual workshop of the australasian language technology association, Smaranda Muresan, Vivian Chen, Kennington Casey, Vandyke David, Dethlefs Nina, Inoue Koji, Ekstedt Erik, and Ultes Stefan (Eds.). Association for Computational Linguistics, Melbourne, Australia, 51\u201366. https:\/\/aclanthology.org\/2023.alta-1.6\/"},{"key":"e_1_3_3_2_114_2","doi-asserted-by":"publisher","unstructured":"John Richards David Piorkowski Michael Hind Stephanie Houde and Aleksandra Mojsilovi\u0107. 2020. A Methodology for Creating AI FactSheets. 10.48550\/arXiv.2006.13796arXiv:https:\/\/arXiv.org\/abs\/2006.13796 [cs].","DOI":"10.48550\/arXiv.2006.13796"},{"key":"e_1_3_3_2_115_2","doi-asserted-by":"publisher","DOI":"10.1145\/3411764.3445604"},{"key":"e_1_3_3_2_116_2","doi-asserted-by":"publisher","unstructured":"Jasper Roe and Mike Perkins. 2023. \u2018What they\u2019re not telling you about ChatGPT\u2019: exploring the discourse of AI in UK news media headlines. Humanities and Social Sciences Communications 10 1 (Oct. 2023) 753. 10.1057\/s41599-023-02282-w","DOI":"10.1057\/s41599-023-02282-w"},{"key":"e_1_3_3_2_117_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2312.06153"},{"key":"e_1_3_3_2_118_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-49008-8_7"},{"key":"e_1_3_3_2_119_2","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533239"},{"key":"e_1_3_3_2_120_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642810"},{"key":"e_1_3_3_2_121_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1163"},{"key":"e_1_3_3_2_122_2","doi-asserted-by":"publisher","unstructured":"Jodi Schneider Di Ye Alison\u00a0M. Hill and Ashley\u00a0S. Whitehorn. 2020. Continued post-retraction citation of a fraudulent clinical trial report 11 years after it was retracted for falsifying data. Scientometrics 125 3 (Dec. 2020) 2877\u20132913. 10.1007\/s11192-020-03631-1","DOI":"10.1007\/s11192-020-03631-1"},{"key":"e_1_3_3_2_123_2","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533192"},{"key":"e_1_3_3_2_124_2","doi-asserted-by":"publisher","unstructured":"Shreya Shankar Yoni Halpern Eric Breck James Atwood Jimbo Wilson and D. Sculley. 2017. No Classification without Representation: Assessing Geodiversity Issues in Open Data Sets for the Developing World. (Nov. 2017). 10.48550\/arXiv.1711.08536ADS Bibcode: 2017arXiv171108536S.","DOI":"10.48550\/arXiv.1711.08536"},{"key":"e_1_3_3_2_125_2","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533110"},{"key":"e_1_3_3_2_126_2","doi-asserted-by":"publisher","unstructured":"Marjia Siddik and Harshvardhan\u00a0J. Pandit. 2025. Datasheets for Healthcare AI: A Framework for Transparency and Bias Mitigation. 10.48550\/arXiv.2501.05617arXiv:https:\/\/arXiv.org\/abs\/2501.05617 [cs].","DOI":"10.48550\/arXiv.2501.05617"},{"key":"e_1_3_3_2_127_2","unstructured":"Ramya Srinivasan Emily Denton Jordan Famularo Negar Rostamzadeh Fernando Diaz and Beth Coleman. 2021. Artsheets for Art Datasets. Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1 (Dec. 2021). https:\/\/datasets-benchmarks-proceedings.neurips.cc\/paper_files\/paper\/2021\/hash\/9b8619251a19057cff70779273e95aa6-Abstract-round2.html"},{"key":"e_1_3_3_2_128_2","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3357853"},{"key":"e_1_3_3_2_129_2","doi-asserted-by":"publisher","DOI":"10.6028\/NIST.AI.100-1"},{"key":"e_1_3_3_2_130_2","unstructured":"David Thiel. 2023. Identifying and eliminating csam in generative ml training data and models. Stanford Internet Observatory Cyber Policy Center December 23 (2023) 3."},{"key":"e_1_3_3_2_131_2","doi-asserted-by":"publisher","unstructured":"Andrea\u00a0C. Tricco Erin Lillie Wasifa Zarin Kelly\u00a0K. O\u2019Brien Heather Colquhoun Danielle Levac David Moher Micah\u00a0D.J. Peters Tanya Horsley Laura Weeks Susanne Hempel Elie\u00a0A. Akl Christine Chang Jessie McGowan Lesley Stewart Lisa Hartling Adrian Aldcroft Michael\u00a0G. Wilson Chantelle Garritty Simon Lewin Christina\u00a0M. Godfrey Marilyn\u00a0T. Macdonald Etienne\u00a0V. Langlois Karla Soares-Weiser Jo Moriarty Tammy Clifford \u00d6zge Tun\u00e7alp and Sharon\u00a0E. Straus. 2018. PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. Annals of Internal Medicine 169 7 (Oct. 2018) 467\u2013473. 10.7326\/M18-0850","DOI":"10.7326\/M18-0850"},{"key":"e_1_3_3_2_132_2","doi-asserted-by":"publisher","DOI":"10.6028\/NIST.AI.100-2e2025"},{"key":"e_1_3_3_2_133_2","doi-asserted-by":"publisher","DOI":"10.1145\/3544548.3581278"},{"key":"e_1_3_3_2_134_2","doi-asserted-by":"publisher","DOI":"10.1145\/3706598.3713814"},{"key":"e_1_3_3_2_135_2","doi-asserted-by":"publisher","unstructured":"Richmond\u00a0Y. Wong. 2021. Tactics of Soft Resistance in User Experience Professionals\u2019 Values Work. Proc. ACM Hum.-Comput. Interact. 5 CSCW2 Article 355 (Oct. 2021) 28\u00a0pages. 10.1145\/3479499","DOI":"10.1145\/3479499"},{"key":"e_1_3_3_2_136_2","doi-asserted-by":"publisher","unstructured":"Richmond\u00a0Y. Wong Michael\u00a0A. Madaio and Nick Merrill. 2023. Seeing Like a Toolkit: How Toolkits Envision the Work of AI Ethics. Proceedings of the ACM on Human-Computer Interaction 7 CSCW1 (April 2023) 1\u201327. 10.1145\/3579621","DOI":"10.1145\/3579621"},{"key":"e_1_3_3_2_137_2","unstructured":"Xinyu Yang Weixin Liang and James Zou. 2023. Navigating Dataset Documentations in AI: A Large-Scale Analysis of Dataset Cards on HuggingFace. https:\/\/openreview.net\/forum?id=xC8xh2RSs2"},{"key":"e_1_3_3_2_138_2","doi-asserted-by":"publisher","DOI":"10.1145\/3511808.3557115"},{"key":"e_1_3_3_2_139_2","doi-asserted-by":"publisher","unstructured":"Tanja \u0160ar\u010devi\u0107 Alicja Karlowicz Rudolf Mayer Ricardo Baeza-Yates and Andreas Rauber. 2024. U Can\u2019t Gen This? A Survey of Intellectual Property Protection Methods for Data in Generative AI. arXiv:2406.15386 (April 2024). 10.48550\/arXiv.2406.15386","DOI":"10.48550\/arXiv.2406.15386"}],"event":{"name":"CHI 2026: CHI Conference on Human Factors in Computing Systems","location":"Barcelona Spain","acronym":"CHI '26","sponsor":["SIGCHI ACM Special Interest Group on Computer-Human Interaction"]},"container-title":["Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3772318.3791344","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T08:49:16Z","timestamp":1782377356000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3772318.3791344"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,13]]},"references-count":138,"alternative-id":["10.1145\/3772318.3791344","10.1145\/3772318"],"URL":"https:\/\/doi.org\/10.1145\/3772318.3791344","relation":{},"subject":[],"published":{"date-parts":[[2026,4,13]]},"assertion":[{"value":"2026-04-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}