{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T06:48:36Z","timestamp":1782370116119,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":74,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T00:00:00Z","timestamp":1776038400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"Danish Novo Nordisk Foundation","award":["NNF20OC0066119"],"award-info":[{"award-number":["NNF20OC0066119"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,4,13]]},"DOI":"10.1145\/3772318.3791069","type":"proceedings-article","created":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T04:12:26Z","timestamp":1776053546000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the Wild"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0245-1633","authenticated-orcid":false,"given":"Willem","family":"van der Maden","sequence":"first","affiliation":[{"name":"HCI &amp; Design Section, IT University of Copenhagen, Copenhagen, Denmark"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8284-890X","authenticated-orcid":false,"given":"Malak","family":"Sadek","sequence":"additional","affiliation":[{"name":"Centre for Human-Inspired Artificial Intelligence (CHIA), Cambridge University, Cambridge, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3368-0180","authenticated-orcid":false,"given":"Ziang","family":"Xiao","sequence":"additional","affiliation":[{"name":"Computer Science, Johns Hopkins University, Baltimore, Maryland, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1827-8513","authenticated-orcid":false,"given":"Aske","family":"Mottelson","sequence":"additional","affiliation":[{"name":"HCI &amp; Design Section, IT University of Copenhagen, Copenhagen, Denmark"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4543-7196","authenticated-orcid":false,"given":"Q. Vera","family":"Liao","sequence":"additional","affiliation":[{"name":"Computer Science and Engineering, University of Michigan, Ann Arbor, Ann Arbor, Michigan, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6740-4550","authenticated-orcid":false,"given":"Jichen","family":"Zhu","sequence":"additional","affiliation":[{"name":"HCI &amp; Design Section, IT University of Copenhagen, Copenhagen, Denmark"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,4,13]]},"reference":[{"key":"e_1_3_3_2_2_2","volume-title":"Introduction to measurement theory","author":"Allen Mary\u00a0J","year":"2001","unstructured":"Mary\u00a0J Allen and Wendy\u00a0M Yen. 2001. Introduction to measurement theory. Waveland Press, Long Grove, IL."},{"key":"e_1_3_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-SEIP.2019.00042"},{"key":"e_1_3_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.acl-long.776"},{"key":"e_1_3_3_2_5_2","doi-asserted-by":"crossref","unstructured":"Lora Aroyo and Chris Welty. 2015. Truth is a lie: Crowd truth and the seven myths of human annotation. AI Magazine 36 1 (2015) 15\u201324.","DOI":"10.1609\/aimag.v36i1.2564"},{"key":"e_1_3_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.acl-short.20"},{"key":"e_1_3_3_2_7_2","first-page":"313","volume-title":"11th Conference of the European Chapter of the Association for Computational Linguistics","author":"Belz Anja","year":"2006","unstructured":"Anja Belz and Ehud Reiter. 2006. Comparing Automatic and Human Evaluation of NLG Systems. In 11th Conference of the European Chapter of the Association for Computational Linguistics. Association for Computational Linguistics, Trento, Italy, 313\u2013320."},{"key":"e_1_3_3_2_8_2","unstructured":"Stella Biderman Hailey Schoelkopf Lintang Sutawika Leo Gao Jonathan Tow Baber Abbasi Alham\u00a0Fikri Aji Pawan\u00a0Sasanka Ammanamanchi Sidney Black Jordan Clive et\u00a0al. 2024. Lessons from the trenches on reproducible evaluation of language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2405.14782 (2024)."},{"key":"e_1_3_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-69909-7_3470-2"},{"key":"e_1_3_3_2_10_2","doi-asserted-by":"publisher","unstructured":"Yupeng Chang Xu Wang Jindong Wang Yuan Wu Linyi Yang Kaijie Zhu Hao Chen Xiaoyuan Yi Cunxiang Wang Yidong Wang Wei Ye Yue Zhang Yi Chang Philip\u00a0S. Yu Qiang Yang and Xing Xie. 2024. A Survey on Evaluation of Large Language Models. ACM Trans. Intell. Syst. Technol. 15 3 Article 39 (March 2024) 45\u00a0pages. 10.1145\/3641289","DOI":"10.1145\/3641289"},{"key":"e_1_3_3_2_11_2","unstructured":"Elizabeth Clark Tal August Sofia Serrano Nikita Haduong Suchin Gururangan and Noah\u00a0A. Smith. 2021. All That\u2019s \u2019Human\u2019 Is Not Gold: Evaluating Human Evaluation of Generated Text. arXiv:https:\/\/arXiv.org\/abs\/2107.00061 [cs]."},{"key":"e_1_3_3_2_12_2","doi-asserted-by":"publisher","unstructured":"Katherine\u00a0M. Collins Albert\u00a0Q. Jiang Simon Frieder Lionel Wong Miri Zilka Umang Bhatt Thomas Lukasiewicz Yuhuai Wu Joshua\u00a0B. Tenenbaum William Hart Timothy Gowers Wenda Li Adrian Weller and Mateja Jamnik. 2024. Evaluating language models for mathematics through interactions. Proceedings of the National Academy of Sciences 121 24 (June 2024) e2318124121. 10.1073\/pnas.2318124121Publisher: Proceedings of the National Academy of Sciences.","DOI":"10.1073\/pnas.2318124121"},{"key":"e_1_3_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3064663.3064667"},{"key":"e_1_3_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1145\/157709.157715"},{"key":"e_1_3_3_2_15_2","doi-asserted-by":"crossref","unstructured":"Aida\u00a0Mostafazadeh Davani Mark D\u00edaz and Vinodkumar Prabhakaran. 2022. Dealing with disagreements: Looking beyond the majority vote in subjective annotations. Transactions of the Association for Computational Linguistics 10 (2022) 92\u2013110.","DOI":"10.1162\/tacl_a_00449"},{"key":"e_1_3_3_2_16_2","unstructured":"Deloitte AI Institute. 2021. Women in AI: Infographic. https:\/\/www2.deloitte.com\/content\/dam\/Deloitte\/us\/Documents\/deloitte-analytics\/us-consulting-ai-institute-women-in-ai-infographic.pdf. [Accessed: 2025-09-01]."},{"key":"e_1_3_3_2_17_2","unstructured":"Finale Doshi-Velez and Been Kim. 2017. Towards a Rigorous Science of Interpretable Machine Learning. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/1702.08608 (2017)."},{"key":"e_1_3_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.acl-long.63"},{"key":"e_1_3_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-06516-3_10"},{"key":"e_1_3_3_2_20_2","unstructured":"Steven Fokkinga Pieter Desmet and Paul Hekkert. 2020. Impact-centered design: Introducing an integrated framework of the psychological and behavioral effects of design. International Journal of Design 14 3 (2020) 97."},{"key":"e_1_3_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3639477.3639720"},{"key":"e_1_3_3_2_22_2","doi-asserted-by":"publisher","unstructured":"Sebastian Gehrmann Elizabeth Clark and Thibault Sellam. 2023. Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text. Journal of Artificial Intelligence Research 77 (May 2023) 103\u2013166. 10.1613\/jair.1.13715","DOI":"10.1613\/jair.1.13715"},{"key":"e_1_3_3_2_23_2","doi-asserted-by":"publisher","unstructured":"Zishan Guo Renren Jin Chuang Liu Yufei Huang Dan Shi Supryadi Linhao Yu Yan Liu Jiaxuan Li Bojian Xiong and Deyi Xiong. 2023. Evaluating Large Language Models: A Comprehensive Survey. 10.48550\/arXiv.2310.19736arXiv:https:\/\/arXiv.org\/abs\/2310.19736.","DOI":"10.48550\/arXiv.2310.19736"},{"key":"e_1_3_3_2_24_2","unstructured":"Steve Harrison Deborah Tatar and Phoebe Sengers. 2007. The three paradigms of HCI(alt.CHI \u201907). Association for Computing Machinery New York NY USA 1\u201318."},{"key":"e_1_3_3_2_25_2","doi-asserted-by":"publisher","unstructured":"Dan Hendrycks Collin Burns Steven Basart Andy Zou Mantas Mazeika Dawn Song and Jacob Steinhardt. 2021. Measuring Massive Multitask Language Understanding. 10.48550\/arXiv.2009.03300arXiv:https:\/\/arXiv.org\/abs\/2009.03300 [cs].","DOI":"10.48550\/arXiv.2009.03300"},{"key":"e_1_3_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3290605.3300830"},{"key":"e_1_3_3_2_27_2","doi-asserted-by":"publisher","unstructured":"Lujain Ibrahim Saffron Huang Lama Ahmad and Markus Anderljung. 2024. Beyond static AI evaluations: advancing human interaction evaluations for LLM harms and risks. 10.48550\/arXiv.2405.10632arXiv:https:\/\/arXiv.org\/abs\/2405.10632 [cs].","DOI":"10.48550\/arXiv.2405.10632"},{"key":"e_1_3_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3442188.3445901"},{"key":"e_1_3_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.emnlp-main.248"},{"key":"e_1_3_3_2_30_2","unstructured":"Mina Lee Megha Srivastava Amelia Hardy John Thickstun Esin Durmus Ashwin Paranjape Ines Gerard-Ursin Xiang\u00a0Lisa Li Faisal Ladhak Frieda Rong Rose\u00a0E. Wang Minae Kwon Joon\u00a0Sung Park Hancheng Cao Tony Lee Rishi Bommasani Michael Bernstein and Percy Liang. 2023. Evaluating Human-Language Model Interaction. arXiv:https:\/\/arXiv.org\/abs\/2212.09746 [cs]."},{"key":"e_1_3_3_2_31_2","volume-title":"UNESCO Science Report: The Race Against Time for Smarter Development","author":"Leibbrandt Alexia","year":"2021","unstructured":"Alexia Leibbrandt. 2021. Women and the Digital Revolution. In UNESCO Science Report: The Race Against Time for Smarter Development. UNESCO, Chapter\u00a03. [Accessed: 2025-09-01]."},{"key":"e_1_3_3_2_32_2","doi-asserted-by":"publisher","unstructured":"Percy Liang Rishi Bommasani Tony Lee Dimitris Tsipras Dilara Soylu Michihiro Yasunaga Yian Zhang Deepak Narayanan Yuhuai Wu Ananya Kumar Benjamin Newman Binhang Yuan Bobby Yan Ce Zhang Christian Cosgrove Christopher\u00a0D. Manning Christopher R\u00e9 Diana Acosta-Navas Drew\u00a0A. Hudson Eric Zelikman Esin Durmus Faisal Ladhak Frieda Rong Hongyu Ren Huaxiu Yao Jue Wang Keshav Santhanam Laurel Orr Lucia Zheng Mert Yuksekgonul Mirac Suzgun Nathan Kim Neel Guha Niladri Chatterji Omar Khattab Peter Henderson Qian Huang Ryan Chi Sang\u00a0Michael Xie Shibani Santurkar Surya Ganguli Tatsunori Hashimoto Thomas Icard Tianyi Zhang Vishrav Chaudhary William Wang Xuechen Li Yifan Mai Yuhui Zhang and Yuta Koreeda. 2023. Holistic Evaluation of Language Models. 10.48550\/arXiv.2211.09110arXiv:https:\/\/arXiv.org\/abs\/2211.09110 [cs].","DOI":"10.48550\/arXiv.2211.09110"},{"key":"e_1_3_3_2_33_2","unstructured":"Q.\u00a0Vera Liao and Ziang Xiao. 2023. Rethinking Model Evaluation as Narrowing the Socio-Technical Gap. arXiv:https:\/\/arXiv.org\/abs\/2306.03100 [cs]."},{"key":"e_1_3_3_2_34_2","volume-title":"Proceedings of the Twelfth International Conference on Learning Representations","author":"Liu Xiao","year":"2024","unstructured":"Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et\u00a0al. 2024. AgentBench: Evaluating LLMs as Agents. In Proceedings of the Twelfth International Conference on Learning Representations."},{"key":"e_1_3_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-naacl.280"},{"key":"e_1_3_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.153"},{"key":"e_1_3_3_2_37_2","unstructured":"Yinhong Liu Han Zhou Zhijiang Guo Ehsan Shareghi Ivan Vuli\u0107 Anna Korhonen and Nigel Collier. 2024. Aligning with human judgement: The role of pairwise preference in large language model evaluators. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2403.16950 (2024)."},{"key":"e_1_3_3_2_38_2","doi-asserted-by":"crossref","unstructured":"Yu\u00a0Lu Liu Su\u00a0Lin Blodgett Jackie Chi\u00a0Kit Cheung Q.\u00a0Vera Liao Alexandra Olteanu and Ziang Xiao. 2024. ECBD: Evidence-Centered Benchmark Design for NLP.","DOI":"10.18653\/v1\/2024.acl-long.861"},{"key":"e_1_3_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3644815.3644950"},{"key":"e_1_3_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/CAIN66642.2025.00011"},{"key":"e_1_3_3_2_41_2","unstructured":"Raiza Martin and Usama\u00a0Bin Shafqat. 2024. How NotebookLM Was Made. Latent Space podcast. https:\/\/www.latent.space\/p\/notebooklm"},{"key":"e_1_3_3_2_42_2","doi-asserted-by":"publisher","unstructured":"Timothy\u00a0R McIntosh Teo Susnjak Nalin Arachchilage Tong Liu Dan Xu Paul Watters and Malka\u00a0N Halgamuge. 2025. Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence. IEEE Transactions on Artificial Intelligence (2025) 1\u201318. 10.1109\/TAI.2025.3569516","DOI":"10.1109\/TAI.2025.3569516"},{"key":"e_1_3_3_2_43_2","doi-asserted-by":"crossref","unstructured":"Samuel Messick. 1995. Validity of psychological assessment: Validation of inferences from persons\u2019 responses and performances as scientific inquiry into score meaning. American Psychologist 50 9 (1995) 741\u2013749.","DOI":"10.1037\/0003-066X.50.9.741"},{"key":"e_1_3_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.741"},{"key":"e_1_3_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-SEIP66354.2025.00051"},{"key":"e_1_3_3_2_46_2","doi-asserted-by":"crossref","unstructured":"David Nigenda Zohar Karnin Muhammad\u00a0Bilal Zafar Raghu Ramesha Alan Tan Michele Donini and Krishnaram Kenthapadi. 2022. Amazon SageMaker Model Monitor: A System for Real-Time Insights into Deployed Machine Learning Models.","DOI":"10.1145\/3534678.3539145"},{"key":"e_1_3_3_2_47_2","doi-asserted-by":"crossref","unstructured":"Donald\u00a0A Norman and Pieter\u00a0Jan Stappers. 2016. DesignX: complex sociotechnical systems. She Ji: The Journal of Design Economics and Innovation 1 2 (2016) 83\u2013106.","DOI":"10.1016\/j.sheji.2016.01.002"},{"key":"e_1_3_3_2_48_2","doi-asserted-by":"crossref","unstructured":"Qian Pan Zahra Ashktorab Michael Desmond Martin\u00a0Santillan Cooper James Johnson Rahul Nair Elizabeth Daly and Werner Geyer. 2024. Human-Centered Design Recommendations for LLM-as-a-Judge. arXiv:https:\/\/arXiv.org\/abs\/2407.03479 [cs].","DOI":"10.18653\/v1\/2024.hucllm-1.2"},{"key":"e_1_3_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/SANER64311.2025.00039"},{"key":"e_1_3_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3351095.3372873"},{"key":"e_1_3_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/2470654.2466257"},{"key":"e_1_3_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/3571884.3597143"},{"key":"e_1_3_3_2_53_2","doi-asserted-by":"publisher","unstructured":"Malak Sadek and Celine Mougenot. 2025. Challenges in Value-Sensitive AI Design: Insights from AI Practitioner Interviews. International Journal of Human\u2013Computer Interaction 41 17 (2025) 10877\u201310894. 10.1080\/10447318.2024.2439021","DOI":"10.1080\/10447318.2024.2439021"},{"key":"e_1_3_3_2_54_2","doi-asserted-by":"crossref","unstructured":"Shreya Shankar J.\u00a0D. Zamfirescu-Pereira Bj\u00f6rn Hartmann Aditya\u00a0G. Parameswaran and Ian Arawjo. 2024. Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences. arXiv:https:\/\/arXiv.org\/abs\/2404.12272 [cs].","DOI":"10.1145\/3654777.3676450"},{"key":"e_1_3_3_2_55_2","doi-asserted-by":"crossref","unstructured":"Murtuza\u00a0N. Shergadwala Himabindu Lakkaraju and Krishnaram Kenthapadi. 2022. A Human-Centric Take on Model Monitoring.","DOI":"10.1609\/hcomp.v10i1.21997"},{"key":"e_1_3_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.543"},{"key":"e_1_3_3_2_57_2","doi-asserted-by":"publisher","unstructured":"Aarohi et\u00a0al. Srivastava. 2023. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models. 10.48550\/arXiv.2206.04615arXiv:https:\/\/arXiv.org\/abs\/2206.04615 [cs stat].","DOI":"10.48550\/arXiv.2206.04615"},{"key":"e_1_3_3_2_58_2","doi-asserted-by":"publisher","unstructured":"Tammy Y.\u00a0C. Tam Sumathy Sivarajkumar Shauna Kapoor et\u00a0al. 2024. A framework for human evaluation of large language models in healthcare derived from literature review. npj Digital Medicine 7 1 (2024) 258. 10.1038\/s41746-024-01258-7","DOI":"10.1038\/s41746-024-01258-7"},{"key":"e_1_3_3_2_59_2","first-page":"404","volume-title":"Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM\u00b2)","author":"Thakur Aman\u00a0Singh","year":"2025","unstructured":"Aman\u00a0Singh Thakur, Kartik Choudhary, Venkat\u00a0Srinik Ramayapally, Sankaran Vaidyanathan, and Dieuwke Hupkes. 2025. Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges. In Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM\u00b2), Ofir Arviv, Miruna Clinciu, Kaustubh Dhole, Rotem Dror, Sebastian Gehrmann, Eliya Habba, Itay Itzhak, Simon Mille, Yotam Perlitz, Enrico Santus, Jo\u00e3o Sedoc, Michal Shmueli\u00a0Scheuer, Gabriel Stanovsky, and Oyvind Tafjord (Eds.). Association for Computational Linguistics, Vienna, Austria and virtual meeting, 404\u2013430."},{"key":"e_1_3_3_2_60_2","first-page":"187","volume-title":"Proceedings of the 31st International Conference on Computational Linguistics: Industry Track","author":"Urlana Ashok","year":"2025","unstructured":"Ashok Urlana, Charaka Vinayak\u00a0Kumar, Bala\u00a0Mallikarjunarao Garlapati, Ajeet\u00a0Kumar Singh, and Rahul Mishra. 2025. No Size Fits All: The Perils and Pitfalls of Leveraging LLMs Vary with Company Size. In Proceedings of the 31st International Conference on Computational Linguistics: Industry Track, Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara\u00a0Di Eugenio, Steven Schockaert, Kareem Darwish, and Apoorv Agarwal (Eds.). Association for Computational Linguistics, Abu Dhabi, UAE, 187\u2013203."},{"key":"e_1_3_3_2_61_2","doi-asserted-by":"publisher","unstructured":"Willem van\u00a0der Maden Derek Lomas and Paul Hekkert. 2024. Developing and evaluating a design method for positive artificial intelligence. Artificial Intelligence for Engineering Design Analysis and Manufacturing 38 (2024) e14. 10.1017\/S0890060424000155","DOI":"10.1017\/S0890060424000155"},{"key":"e_1_3_3_2_62_2","volume-title":"SuperGLUE: a stickier benchmark for general-purpose language understanding systems","author":"Wang Alex","year":"2019","unstructured":"Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel\u00a0R. Bowman. 2019. SuperGLUE: a stickier benchmark for general-purpose language understanding systems. Curran Associates Inc., Red Hook, NY, USA."},{"key":"e_1_3_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-5446"},{"key":"e_1_3_3_2_64_2","doi-asserted-by":"publisher","unstructured":"Chenyu Wang Zhou Yang Zewei Li Daniela\u00a0E. Damian and David Lo. 2024. Quality Assurance for Artificial Intelligence: A Study of Industrial Concerns Challenges and Best Practices. ArXiv abs\/2402.16391 (2024). 10.48550\/arXiv.2402.16391","DOI":"10.48550\/arXiv.2402.16391"},{"key":"e_1_3_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3641917"},{"key":"e_1_3_3_2_66_2","doi-asserted-by":"crossref","unstructured":"Jiayin Wang Fengran Mo Weizhi Ma Peijie Sun Min Zhang and Jian-Yun Nie. 2024. A User-Centric Benchmark for Evaluating Large Language Models. arXiv:https:\/\/arXiv.org\/abs\/2404.13940 [cs].","DOI":"10.18653\/v1\/2024.emnlp-main.210"},{"key":"e_1_3_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.52202\/079017-3018"},{"key":"e_1_3_3_2_68_2","unstructured":"Laura Weidinger Maribeth Rauh Nahema Marchal Arianna Manzini Lisa\u00a0Anne Hendricks Juan Mateos-Garcia Stevie Bergman Jackie Kay Conor Griffin Ben Bariach Iason Gabriel Verena Rieser and William Isaac. 2023. Sociotechnical Safety Evaluation of Generative AI Systems. arXiv:https:\/\/arXiv.org\/abs\/2310.11986 [cs]."},{"key":"e_1_3_3_2_69_2","unstructured":"World Economic Forum and LinkedIn. 2025. Gender Parity in the Intelligent Age. [Accessed: 2025-09-01]."},{"key":"e_1_3_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.676"},{"key":"e_1_3_3_2_71_2","doi-asserted-by":"publisher","unstructured":"Tianyi Zhang Faisal Ladhak Esin Durmus Percy Liang Kathleen McKeown and Tatsunori\u00a0B. Hashimoto. 2024. Benchmarking Large Language Models for News Summarization. Transactions of the Association for Computational Linguistics 12 (2024) 39\u201357. 10.1162\/tacl_a_00632","DOI":"10.1162\/tacl_a_00632"},{"key":"e_1_3_3_2_72_2","doi-asserted-by":"crossref","unstructured":"Lianmin Zheng Wei-Lin Chiang Ying Sheng Siyuan Zhuang Zhanghao Wu Yonghao Zhuang Zi Lin Zhuohan Li Dacheng Li Eric Xing et\u00a0al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems 36 (2023) 46595\u201346623.","DOI":"10.52202\/075280-2020"},{"key":"e_1_3_3_2_73_2","doi-asserted-by":"publisher","unstructured":"Lianmin Zheng Wei-Lin Chiang Ying Sheng Siyuan Zhuang Zhanghao Wu Yonghao Zhuang Zi Lin Zhuohan Li Dacheng Li Eric\u00a0P. Xing Hao Zhang Joseph\u00a0E. Gonzalez and Ion Stoica. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. 10.48550\/arXiv.2306.05685arXiv:https:\/\/arXiv.org\/abs\/2306.05685 [cs].","DOI":"10.48550\/arXiv.2306.05685"},{"key":"e_1_3_3_2_74_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.24"},{"key":"e_1_3_3_2_75_2","volume-title":"Proceedings of the Thirteenth International Conference on Learning Representations","author":"Zhuo Terry\u00a0Yue","year":"2024","unstructured":"Terry\u00a0Yue Zhuo, Minh\u00a0Chien Vu, Jenny Chim, Han Hu, Wenhao Yu, Ratnadira Widyasari, Imam Nur\u00a0Bani Yusuf, Haolan Zhan, Junda He, Indraneil Paul, et\u00a0al. 2024. BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions. In Proceedings of the Thirteenth International Conference on Learning Representations."}],"event":{"name":"CHI 2026: CHI Conference on Human Factors in Computing Systems","location":"Barcelona Spain","acronym":"CHI '26","sponsor":["SIGCHI ACM Special Interest Group on Computer-Human Interaction"]},"container-title":["Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3772318.3791069","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T06:18:23Z","timestamp":1782368303000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3772318.3791069"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,13]]},"references-count":74,"alternative-id":["10.1145\/3772318.3791069","10.1145\/3772318"],"URL":"https:\/\/doi.org\/10.1145\/3772318.3791069","relation":{},"subject":[],"published":{"date-parts":[[2026,4,13]]},"assertion":[{"value":"2026-04-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}