{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,29]],"date-time":"2026-06-29T19:48:50Z","timestamp":1782762530429,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":46,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,6,25]]},"DOI":"10.1145\/3805689.3812403","type":"proceedings-article","created":{"date-parts":[[2026,6,29]],"date-time":"2026-06-29T17:52:08Z","timestamp":1782755528000},"page":"6554-6586","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Validating and Refining Measurements for Generative AI Evaluation Via Stakeholder Engagement"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7262-0247","authenticated-orcid":false,"given":"Tonya","family":"Nguyen","sequence":"first","affiliation":[{"name":"Information Science, University of California, Berkeley, Berkeley, California, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-5550-640X","authenticated-orcid":false,"given":"Jean","family":"Garcia-Gathright","sequence":"additional","affiliation":[{"name":"Microsoft Research, New York City, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-8124-0147","authenticated-orcid":false,"given":"Hannah","family":"Washington","sequence":"additional","affiliation":[{"name":"Microsoft Research, New York City, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2337-9610","authenticated-orcid":false,"given":"Alexandra","family":"Chouldechova","sequence":"additional","affiliation":[{"name":"Microsoft Research, New York City, New York, USA, and Abridge AI Inc., New York City, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3395-7186","authenticated-orcid":false,"given":"Hanna","family":"Wallach","sequence":"additional","affiliation":[{"name":"Microsoft Research, New York City, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7807-2018","authenticated-orcid":false,"given":"Jennifer","family":"Wortman Vaughan","sequence":"additional","affiliation":[{"name":"Microsoft Research, New York City, New York, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,25]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3351095.3372871"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1017\/S0003055401003100"},{"key":"e_1_3_2_1_3_1","volume-title":"How to Do Things with Words","author":"Austin J.L.","unstructured":"J.L. Austin. 1962. How to Do Things with Words. Harvard University Press."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-024-56648-4"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3551624.3555290"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.81"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","unstructured":"David Cash William C. Clark Frank Alcock Nancy M. Dickson Noelle Eckley and Jill J\u00e4ger. 2002. Salience Credibility Legitimacy and Boundaries: Linking Research Assessment and Decision Making. doi:10.2139\/ssrn.372280","DOI":"10.2139\/ssrn.372280"},{"key":"e_1_3_2_1_8_1","unstructured":"Kathy Charmaz. 2006. Constructing grounded theory: A practical guide through qualitative analysis. sage."},{"key":"e_1_3_2_1_9_1","unstructured":"Alexandra Chouldechova Chad Atalla Solon Barocas A. Feder Cooper Emily Corvi P. Alex Dow Jean Garcia-Gathright Nicholas Pangakis Stefanie Reed Emily Sheng Dan Vann Matthew Vogel Hannah Washington and Hanna Wallach. 2024. A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities Risks and Impacts. arXiv:2412.01934 [cs.CY] https:\/\/arxiv.org\/abs\/2412.01934"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3544548.3581134"},{"key":"e_1_3_2_1_11_1","unstructured":"A. Feder Cooper and James Grimmelmann. 2025. The Files are in the Computer: On Copyright Memorization and Generative AI. arXiv:2404.12590 [cs.CY] https:\/\/arxiv.org\/abs\/2404.12590"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2311.06477"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","unstructured":"Emily Corvi Hannah Washington Stefanie Reed Chad Atalla Alexandra Chouldechova P. Alex Dow Jean Garcia-Gathright Nicholas J Pangakis Emily Sheng Dan Vann Matthew Vogel and Hanna Wallach. 2025. Taxonomizing Representational Harms using Speech Act Theory. 3907-3932 pages. doi:10.18653\/v1\/2025.findings-acl.202","DOI":"10.18653\/v1\/2025.findings-acl.202"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"crossref","unstructured":"Amanda Coston Anna Kawakami Haiyi Zhu Ken Holstein and Hoda Heidari. 2023. A Validity Perspective on Evaluating the Justified Use of Data-driven Decision-making Algorithms. arXiv:2206.14983 [cs.LG] https:\/\/arxiv.org\/abs\/2206.14983","DOI":"10.1109\/SaTML54575.2023.00050"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3617694.3623261"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1561\/1100000015"},{"key":"e_1_3_2_1_17_1","volume-title":"Kahn Jr","author":"Friedman Batya","year":"2003","unstructured":"Batya Friedman and Peter H. Kahn Jr. 2003. Human Values, Ethics, and Design. In The Human-Computer Interaction Handbook: Fundamentals, Evolving Technologies, and Emerging Applications, Jacko JA Sears A (Ed.). L. Erlbaum Associates Inc., USA, 1177\u20131201."},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-94-007-7844-3_4"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1038\/s42256-020-00257-z"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.2307\/3235246"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1177\/20539517241252131"},{"key":"e_1_3_2_1_22_1","volume-title":"MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming. arXiv:2505.17147 [cs.CR] https:\/\/arxiv.org\/abs\/2505.17147","author":"Guo Weiyang","year":"2025","unstructured":"Weiyang Guo, Jing Li, Wenya Wang, YU LI, Daojing He, Jun Yu, and Min Zhang. 2025. MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming. arXiv:2505.17147 [cs.CR] https:\/\/arxiv.org\/abs\/2505.17147"},{"key":"e_1_3_2_1_23_1","volume-title":"Alexandra Chouldechova, Jean Garcia-Gathright, Alexandra Olteanu, and Hanna Wallach.","author":"Harvey Emma","year":"2024","unstructured":"Emma Harvey, Emily Sheng, Su Lin Blodgett, Alexandra Chouldechova, Jean Garcia-Gathright, Alexandra Olteanu, and Hanna Wallach. 2024. Gaps Between Research and Practice When Measuring Representational Harms Caused by LLM-Based Systems. arXiv:2411.15662 [cs.CY] https:\/\/arxiv.org\/abs\/2411.15662"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/1993060.1993065"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3442188.3445901"},{"key":"e_1_3_2_1_26_1","volume-title":"Proceedings of the 9th ACM Conference on Fairness, Accountability, and Transparency.","author":"Johnson Nari","year":"2026","unstructured":"Nari Johnson, Deepthi Sudharsan, Hamna, Samantha Dalal, Theo Holroyd, Anja Thieme, Hoda Heidari, Daniela Massiceti, Jennifer Wortman Vaughan, and Cecily Morrison. 2026. Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics. In Proceedings of the 9th ACM Conference on Fairness, Accountability, and Transparency."},{"key":"e_1_3_2_1_27_1","volume-title":"AI Safety on whose terms? Science 381, 6654","author":"Lazar Seth","year":"2023","unstructured":"Seth Lazar and Alondra Nelson. 2023. AI Safety on whose terms? Science 381, 6654 (2023), 138\u2013138."},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3664616"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/153571.255960"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.6028\/NIST.AI.600-1"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3687058"},{"key":"e_1_3_2_1_32_1","unstructured":"Shiwen Ni Guhong Chen Shuaimin Li Xuanang Chen Siyi Li Bingli Wang Qiyao Wang Xingjian Wang Yifan Zhang Liyang Fan Chengming Li Ruifeng Xu Le Sun and Min Yang. 2025. A Survey on Large Language Model Benchmarks. arXiv:2508.15361 [cs.CL] https:\/\/arxiv.org\/abs\/2508.15361"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"crossref","unstructured":"Ethan Perez Saffron Huang Francis Song Trevor Cai Roman Ring John Aslanides Amelia Glaese Nat McAleese and Geoffrey Irving. 2022. Red Teaming Language Models with Language Models. arXiv:2202.03286 [cs.CL] https:\/\/arxiv.org\/abs\/2202.03286","DOI":"10.18653\/v1\/2022.emnlp-main.225"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2503.19075"},{"key":"e_1_3_2_1_35_1","unstructured":"Inioluwa Deborah Raji Emily M. Bender Amandalynne Paullada Emily Denton and Alex Hanna. 2021. AI and the Everything in the Whole Wide World Benchmark. arXiv:2111.15366 [cs.LG] https:\/\/arxiv.org\/abs\/2111.15366"},{"key":"e_1_3_2_1_36_1","unstructured":"Olawale Salaudeen Anka Reuel Ahmed Ahmed Suhana Bedi Zachary Robertson Sudharsan Sundar Ben Domingue Angelina Wang and Sanmi Koyejo. 2025. Measurement to Meaning: A Validity-Centered Framework for AI Evaluation. arXiv:2505.10573 [cs.CY] https:\/\/arxiv.org\/abs\/2505.10573"},{"key":"e_1_3_2_1_37_1","volume-title":"Seeing Like a State: How Certain Schemes to Improve the Human Condition Have Failed","author":"Scott James C.","unstructured":"James C. Scott. 1998. Seeing Like a State: How Certain Schemes to Improve the Human Condition Have Failed. Yale University Press, New Haven."},{"key":"e_1_3_2_1_38_1","unstructured":"Mrinank Sharma Meg Tong Tomasz Korbak David Duvenaud Amanda Askell Samuel R. Bowman Newton Cheng Esin Durmus Zac Hatfield-Dodds Scott R. Johnston Shauna Kravec Timothy Maxwell Sam McCandlish Kamal Ndousse Oliver Rausch Nicholas Schiefer Da Yan Miranda Zhang and Ethan Perez. 2025. Towards Understanding Sycophancy in Language Models. arXiv:2310.13548 [cs.CL] https:\/\/arxiv.org\/abs\/2310.13548"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3600211.3604673"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2531602.2531625"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3551624.3555285"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3290605.3300791"},{"key":"e_1_3_2_1_43_1","volume-title":"Forty-second International Conference on Machine Learning Position Paper Track. https:\/\/openreview.net\/forum?id=1ZC4RNjqzU","author":"Wallach Hanna","unstructured":"Hanna Wallach, Meera Desai, A. Feder Cooper, Angelina Wang, Chad Atalla, Solon Barocas, Su Lin Blodgett, Alexandra Chouldechova, Emily Corvi, P. Alex Dow, Jean Garcia-Gathright, Alexandra Olteanu, Nicholas J Pangakis, Stefanie Reed, Emily Sheng, Dan Vann, Jennifer Wortman Vaughan, Matthew Vogel, Hannah Washington, and Abigail Z. Jacobs. 2025. Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge. In Forty-second International Conference on Machine Learning Position Paper Track. https:\/\/openreview.net\/forum?id=1ZC4RNjqzU"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"crossref","unstructured":"Jennifer Wang Andrew Selbst Solon Barocas and Suresh Venkatasubramanian. 2026. Distinguishing Task-Specific and General-Purpose AI in Regulation. arXiv:2506.17347 [cs.CY] https:\/\/arxiv.org\/abs\/2506.17347","DOI":"10.1145\/3788646.3789523"},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3290605.3300428"},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.5210\/fm.v29i4.13642"}],"event":{"name":"FAccT '26: The 2026 ACM Conference on Fairness, Accountability, and Transparency","location":"Montreal QC Canada","acronym":"FAccT '26","sponsor":["ACM\/SIG"]},"container-title":["Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3805689.3812403","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,29]],"date-time":"2026-06-29T19:20:47Z","timestamp":1782760847000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3805689.3812403"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,25]]},"references-count":46,"alternative-id":["10.1145\/3805689.3812403","10.1145\/3805689"],"URL":"https:\/\/doi.org\/10.1145\/3805689.3812403","relation":{},"subject":[],"published":{"date-parts":[[2026,6,25]]},"assertion":[{"value":"2026-06-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}