{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,29]],"date-time":"2026-06-29T18:45:31Z","timestamp":1782758731069,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":55,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,6,25]]},"DOI":"10.1145\/3805689.3806721","type":"proceedings-article","created":{"date-parts":[[2026,6,23]],"date-time":"2026-06-23T16:20:39Z","timestamp":1782231639000},"page":"7565-7586","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["LLM Evaluation as Sociotechnical Practice: Experiences from a Research-Practice Partnership in Community Health"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-9606-2613","authenticated-orcid":false,"given":"Roshini","family":"Deva","sequence":"first","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, Georgia, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8352-6929","authenticated-orcid":false,"given":"Aradhana","family":"Thapa","sequence":"additional","affiliation":[{"name":"Emory University, Atlanta, Georgia, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-3037-8249","authenticated-orcid":false,"given":"Zeel","family":"Mehta","sequence":"additional","affiliation":[{"name":"Myna Mahila Foundation, Mumbai, Maharashtra, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-2495-8279","authenticated-orcid":false,"given":"Shraddha Kale","family":"Kapile","sequence":"additional","affiliation":[{"name":"Myna Mahila Foundation, Mumbai, Maharashtra, India"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3770-9369","authenticated-orcid":false,"given":"Suhani","family":"Jalota","sequence":"additional","affiliation":[{"name":"Stanford University, Hoover Institution, Stanford, California, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0355-3041","authenticated-orcid":false,"given":"Naveena","family":"Karusala","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology, Atlanta, Georgia, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7570-9474","authenticated-orcid":false,"given":"Azra","family":"Ismail","sequence":"additional","affiliation":[{"name":"Emory University, Atlanta, Georgia, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,25]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Overview of the MEDIQA 2019 shared task on textual inference, question entailment and question answering. In Proceedings of the 18th bioNLP workshop and shared task. 370\u2013379","author":"Abacha Asma Ben","year":"2019","unstructured":"Asma Ben Abacha, Chaitanya Shivade, and Dina Demner-Fushman. 2019. Overview of the MEDIQA 2019 shared task on textual inference, question entailment and question answering. In Proceedings of the 18th bioNLP workshop and shared task. 370\u2013379."},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41746-024-01074-z"},{"key":"e_1_3_2_1_3_1","volume-title":"The challenges of evaluating llm applications: An analysis of automated, human, and llm-based approaches. arXiv preprint arXiv:2406.03339","author":"Abeysinghe Bhashithe","year":"2024","unstructured":"Bhashithe Abeysinghe and Ruhan Circi. 2024. The challenges of evaluating llm applications: An analysis of automated, human, and llm-based approaches. arXiv preprint arXiv:2406.03339 (2024)."},{"key":"e_1_3_2_1_4_1","volume-title":"Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al.","author":"Achiam Josh","year":"2023","unstructured":"Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)."},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.258"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.naacl-long.143"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1101\/2023.12.22.23300458"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"crossref","unstructured":"John W Ayers Adam Poliak Mark Dredze Eric C Leas Zechariah Zhu Jessica B Kelley Dennis J Faix Aaron M Goodman Christopher A Longhurst Michael Hogarth et al. 2023. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA internal medicine 183 6 (2023) 589\u2013596.","DOI":"10.1001\/jamainternmed.2023.1838"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3442188.3445922"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1080\/08839514.2010.492259"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.2196\/11510"},{"key":"e_1_3_2_1_12_1","volume-title":"What is participatory design? In Participatory design","author":"B\u00f8dker Susanne","unstructured":"Susanne B\u00f8dker, Christian Dindler, Ole S Iversen, and Rachel C Smith. 2022. What is participatory design? In Participatory design. Springer, 5\u201313."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.5555\/380681"},{"key":"e_1_3_2_1_14_1","volume-title":"One size fits all? What counts as quality practice in (reflexive) thematic analysis? Qualitative research in psychology 18, 3","author":"Braun Virginia","year":"2021","unstructured":"Virginia Braun and Victoria Clarke. 2021. One size fits all? What counts as quality practice in (reflexive) thematic analysis? Qualitative research in psychology 18, 3 (2021), 328\u2013352."},{"key":"e_1_3_2_1_15_1","volume-title":"Who's thinking? A push for human-centered evaluation of LLMs using the XAI playbook. arXiv preprint arXiv:2303.06223","author":"Datta Teresa","year":"2023","unstructured":"Teresa Datta and John P Dickerson. 2023. Who's thinking? A push for human-centered evaluation of LLMs using the XAI playbook. arXiv preprint arXiv:2303.06223 (2023)."},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.5749\/9781452953465"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3617694.3623261"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3706598.3713362"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"crossref","unstructured":"Kathleen Kara Fitzpatrick Alison Darcy and Molly Vierhile. 2017. Delivering cognitive behavior therapy to young adults with symptoms of depression and anxiety using a fully automated conversational agent (Woebot): a randomized controlled trial. JMIR mental health 4 2 (2017) e7785.","DOI":"10.2196\/mental.7785"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.7551\/mitpress\/7585.001.0001"},{"key":"e_1_3_2_1_21_1","volume-title":"Proceedings of the ACM on Human-Computer Interaction 7, CSCW","author":"Ismail A.","year":"2023","unstructured":"A. Ismail, D. H. Thakkar, N. Madhiwalla, and N. Kumar. 2023. Public Health calls for\/with AI: An ethnographic perspective. Proceedings of the ACM on Human-Computer Interaction 7, CSCW (2023), 1\u201324."},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3442188.3445901"},{"key":"e_1_3_2_1_23_1","volume-title":"Andrea Madotto, and Pascale Fung.","author":"Ji Ziwei","year":"2023","unstructured":"Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM computing surveys 55, 12 (2023), 1\u201338."},{"key":"e_1_3_2_1_24_1","volume-title":"The state and fate of linguistic diversity and inclusion in the NLP world. arXiv preprint arXiv:2004.09095","author":"Joshi Pratik","year":"2020","unstructured":"Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. The state and fate of linguistic diversity and inclusion in the NLP world. arXiv preprint arXiv:2004.09095 (2020)."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3359286"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3610210"},{"key":"e_1_3_2_1_27_1","volume-title":"Lived data: tinkering with bodies, code, and care work. Human\u2013Computer Interaction 33, 1","author":"Kaziunas Elizabeth","year":"2018","unstructured":"Elizabeth Kaziunas, Silvia Lindtner, Mark S Ackerman, and Joyce M Lee. 2018. Lived data: tinkering with bodies, code, and care work. Human\u2013Computer Interaction 33, 1 (2018), 49\u201392."},{"key":"e_1_3_2_1_28_1","volume-title":"GLUECoS: An evaluation benchmark for code-switched NLP. arXiv preprint arXiv:2004.12376","author":"Khanuja Simran","year":"2020","unstructured":"Simran Khanuja, Sandipan Dandapat, Anirudh Srinivasan, Sunayana Sitaram, and Monojit Choudhury. 2020. GLUECoS: An evaluation benchmark for code-switched NLP. arXiv preprint arXiv:2004.12376 (2020)."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.2196\/72153"},{"key":"e_1_3_2_1_30_1","volume-title":"New directions in eHealth communication: opportunities and challenges. Patient education and counseling 78, 3","author":"Kreps Gary L","year":"2010","unstructured":"Gary L Kreps and Linda Neuhauser. 2010. New directions in eHealth communication: opportunities and challenges. Patient education and counseling 78, 3 (2010), 329\u2013336."},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3274368"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1093\/jamia\/ocy072"},{"key":"e_1_3_2_1_33_1","unstructured":"Percy Liang Rishi Bommasani Tony Lee Dimitris Tsipras Dilara Soylu Michihiro Yasunaga Yian Zhang Deepak Narayanan Yuhuai Wu Ananya Kumar et al. 2022. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110 (2022)."},{"key":"e_1_3_2_1_34_1","volume-title":"Rethinking model evaluation as narrowing the socio-technical gap. arXiv preprint arXiv:2306.03100","author":"Vera Liao Q","year":"2023","unstructured":"Q Vera Liao and Ziang Xiao. 2023. Rethinking model evaluation as narrowing the socio-technical gap. arXiv preprint arXiv:2306.03100 (2023)."},{"key":"e_1_3_2_1_35_1","volume-title":"Proceedings of the ACM on Human-Computer Interaction 8, CSCW2","author":"Lin Hongjin","year":"2024","unstructured":"Hongjin Lin, Naveena Karusala, Chinasa T Okolo, Catherine D'Ignazio, and Krzysztof Z Gajos. 2024. \u201cCome to us first\u201d: Centering Community Organizations in Artificial Intelligence for Social Good Partnerships. Proceedings of the ACM on Human-Computer Interaction 8, CSCW2 (2024), 1\u201328."},{"key":"e_1_3_2_1_36_1","volume-title":"On faithfulness and factuality in abstractive summarization. arXiv preprint arXiv:2005.00661","author":"Maynez Joshua","year":"2020","unstructured":"Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. arXiv preprint arXiv:2005.00661 (2020)."},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3287560.3287596"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.4324\/9780203927076"},{"key":"e_1_3_2_1_39_1","unstructured":"Annemarie Mol Ingunn Moser and Jeannette Pols. 2010. Care in practice: On tinkering in clinics homes and farms. transcript Verlag."},{"key":"e_1_3_2_1_40_1","volume-title":"Ethics and governance of artificial intelligence for health: large multi-modal models","author":"World Health Organization","unstructured":"World Health Organization. 2024. Ethics and governance of artificial intelligence for health: large multi-modal models. WHO guidance. World Health Organization."},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3287560.3287567"},{"key":"e_1_3_2_1_42_1","volume-title":"Proceedings of the 17th ACM conference on Computer supported cooperative work & social computing. 283\u2013294","author":"Pine Kathleen H","year":"2014","unstructured":"Kathleen H Pine and Melissa Mazmanian. 2014. Institutional logics of the EMR and the problem of \u2018perfect\u2019 but inaccurate accounts. In Proceedings of the 17th ACM conference on Computer supported cooperative work & social computing. 283\u2013294."},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3351095.3372873"},{"key":"e_1_3_2_1_44_1","volume-title":"Balancing the local and the global in infrastructural information systems. The information society 18, 2","author":"Rolland Knut H","year":"2002","unstructured":"Knut H Rolland and Eric Monteiro. 2002. Balancing the local and the global in infrastructural information systems. The information society 18, 2 (2002), 87\u2013100."},{"key":"e_1_3_2_1_45_1","volume-title":"Algorithms as culture: Some tactics for the ethnography of algorithmic systems. Big data & society 4, 2","author":"Seaver Nick","year":"2017","unstructured":"Nick Seaver. 2017. Algorithms as culture: Some tactics for the ethnography of algorithmic systems. Big data & society 4, 2 (2017), 2053951717738104."},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3287560.3287598"},{"key":"e_1_3_2_1_47_1","volume-title":"Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al.","author":"Singhal Karan","year":"2023","unstructured":"Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. 2023. Large language models encode clinical knowledge. Nature 620, 7972 (2023), 172\u2013180."},{"key":"e_1_3_2_1_48_1","volume-title":"Institutional ecology,translations' and boundary objects: Amateurs and professionals in Berkeley's Museum of Vertebrate Zoology","author":"Star Susan Leigh","year":"1907","unstructured":"Susan Leigh Star and James R Griesemer. 1989. Institutional ecology,translations' and boundary objects: Amateurs and professionals in Berkeley's Museum of Vertebrate Zoology, 1907\u201339. Social studies of science 19, 3 (1989), 387\u2013420."},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.5555\/1212951"},{"key":"e_1_3_2_1_50_1","doi-asserted-by":"crossref","unstructured":"Thomas Yu Chow Tam Sonish Sivarajkumar Sumit Kapoor Alisa V Stolyar Katelyn Polanska Karleigh R McCarthy Hunter Osterhoudt Xizhi Wu Shyam Visweswaran Sunyang Fu et al. 2024. A framework for human evaluation of large language models in healthcare derived from literature review. NPJ digital medicine 7 1 (2024) 258.","DOI":"10.1038\/s41746-024-01258-7"},{"key":"e_1_3_2_1_51_1","unstructured":"Ting Fang Tan Kabilan Elangovan Jasmine Chiat Ling Ong Aaron Lee Nigam H Shah Joseph JY Sung Tien Yin Wong Xue Lan Nan Liu Haibo Wang et al. [n. d.]. A Proposed SCORE Evaluation Framework for Large Language Models\u2013Safety Consensus & Context Objectivity Reproducibility and Explainability. Consensus & Context Objectivity Reproducibility and Explainability ([n. d.])."},{"key":"e_1_3_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3555532"},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1177\/0706743719828977"},{"key":"e_1_3_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533088"},{"key":"e_1_3_2_1_55_1","volume-title":"Care and disability. Practices of experimenting, tinkering with, and arranging people and technical aids. Care in practice. On tinkering in clinics, homes and farms","author":"Winance Myriam","year":"2010","unstructured":"Myriam Winance. 2010. Care and disability. Practices of experimenting, tinkering with, and arranging people and technical aids. Care in practice. On tinkering in clinics, homes and farms (2010), 93\u2013117."}],"event":{"name":"FAccT '26: The 2026 ACM Conference on Fairness, Accountability, and Transparency","location":"Montreal QC Canada","acronym":"FAccT '26","sponsor":["ACM\/SIG"]},"container-title":["Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3805689.3806721","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,29]],"date-time":"2026-06-29T18:09:19Z","timestamp":1782756559000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3805689.3806721"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,25]]},"references-count":55,"alternative-id":["10.1145\/3805689.3806721","10.1145\/3805689"],"URL":"https:\/\/doi.org\/10.1145\/3805689.3806721","relation":{},"subject":[],"published":{"date-parts":[[2026,6,25]]},"assertion":[{"value":"2026-06-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}