{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T20:14:49Z","timestamp":1776111289006,"version":"3.50.1"},"publisher-location":"New York, NY, USA","reference-count":23,"publisher":"ACM","license":[{"start":{"date-parts":[[2025,3,24]],"date-time":"2025-03-24T00:00:00Z","timestamp":1742774400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"U.S. Department of Agriculture, National Institute of Food and Agriculture","award":["2021-67022-33447"],"award-info":[{"award-number":["2021-67022-33447"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,3,24]]},"DOI":"10.1145\/3708557.3716146","type":"proceedings-article","created":{"date-parts":[[2025,3,18]],"date-time":"2025-03-18T07:08:31Z","timestamp":1742281711000},"page":"215-217","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Evaluation Workflows for Large Language Models (LLMs) that Integrate Domain Expertise for Complex Knowledge Tasks"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-5472-282X","authenticated-orcid":false,"given":"Annalisa","family":"Szymanski","sequence":"first","affiliation":[{"name":"Computer Science and Engineering, University of Notre Dame, Notre Dame, Indiana, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,3,24]]},"reference":[{"key":"e_1_3_3_1_2_2","doi-asserted-by":"crossref","unstructured":"Muhammad\u00a0Aurangzeb Ahmad Ilker Yaramis and Taposh\u00a0Dutta Roy. 2023. Creating trustworthy llms: Dealing with hallucinations in healthcare ai. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2311.01463 (2023).","DOI":"10.20944\/preprints202310.1662.v1"},{"key":"e_1_3_3_1_3_2","unstructured":"Konstantinos Andriopoulos and Johan Pouwelse. 2023. Augmenting LLMs with Knowledge: A survey on hallucination prevention. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2309.16459 (2023)."},{"key":"e_1_3_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642016"},{"key":"e_1_3_3_1_5_2","doi-asserted-by":"crossref","unstructured":"John\u00a0W Ayers Adam Poliak Mark Dredze Eric\u00a0C Leas Zechariah Zhu Jessica\u00a0B Kelley Dennis\u00a0J Faix Aaron\u00a0M Goodman Christopher\u00a0A Longhurst Michael Hogarth et\u00a0al. 2023. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA internal medicine 183 6 (2023) 589\u2013596.","DOI":"10.1001\/jamainternmed.2023.1838"},{"key":"e_1_3_3_1_6_2","doi-asserted-by":"crossref","unstructured":"Angeline Chatelan Aur\u00e9lien Clerc and Pierre-Alexandre Fonta. 2023. ChatGPT and future artificial intelligence chatbots: what may be the influence on credentialed nutrition and dietetics practitioners? Journal of the Academy of Nutrition and Dietetics 123 11 (2023) 1525\u20131531.","DOI":"10.1016\/j.jand.2023.08.001"},{"key":"e_1_3_3_1_7_2","unstructured":"Yida Chen Aoyu Wu Trevor DePodesta Catherine Yeh Kenneth Li Nicholas\u00a0Castillo Marin Oam Patel Jan Riecke Shivam Raval Olivia Seow et\u00a0al. 2024. Designing a Dashboard for Transparency and Control of Conversational AI. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2406.07882 (2024)."},{"key":"e_1_3_3_1_8_2","doi-asserted-by":"crossref","unstructured":"Sebastian Gehrmann Elizabeth Clark and Thibault Sellam. 2023. Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text. Journal of Artificial Intelligence Research 77 (2023) 103\u2013166.","DOI":"10.1613\/jair.1.13715"},{"key":"e_1_3_3_1_9_2","doi-asserted-by":"crossref","unstructured":"Suchin Gururangan Ana Marasovi\u0107 Swabha Swayamdipta Kyle Lo Iz Beltagy Doug Downey and Noah\u00a0A Smith. 2020. Don\u2019t stop pretraining: Adapt language models to domains and tasks. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2004.10964 (2020).","DOI":"10.18653\/v1\/2020.acl-main.740"},{"key":"e_1_3_3_1_10_2","unstructured":"Junfeng Jiao Saleh Afroogh Yiming Xu and Connor Phillips. 2024. Navigating llm ethics: Advancements challenges and future directions. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2406.18841 (2024)."},{"key":"e_1_3_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613905.3650755"},{"key":"e_1_3_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3618260.3649777"},{"key":"e_1_3_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642216"},{"key":"e_1_3_3_1_14_2","unstructured":"Fang Liu Yang Liu Lin Shi Houkun Huang Ruifeng Wang Zhen Yang and Li Zhang. 2024. Exploring and evaluating hallucinations in llm-powered code generation. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2404.00971 (2024)."},{"key":"e_1_3_3_1_15_2","doi-asserted-by":"crossref","unstructured":"Adian Liusie Vatsal Raina Yassir Fathullah and Mark Gales. 2024. Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2405.05894 (2024).","DOI":"10.18653\/v1\/2024.emnlp-main.389"},{"key":"e_1_3_3_1_16_2","doi-asserted-by":"crossref","unstructured":"Qian Pan Zahra Ashktorab Michael Desmond Martin\u00a0Santillan Cooper James Johnson Rahul Nair Elizabeth Daly and Werner Geyer. 2024. Human-Centered Design Recommendations for LLM-as-a-Judge. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2407.03479 (2024).","DOI":"10.18653\/v1\/2024.hucllm-1.2"},{"key":"e_1_3_3_1_17_2","first-page":"887","volume-title":"Healthcare","author":"Sallam Malik","year":"2023","unstructured":"Malik Sallam. 2023. ChatGPT utility in healthcare education, research, and practice: systematic review on the promising perspectives and valid concerns. In Healthcare, Vol.\u00a011. MDPI, 887."},{"key":"e_1_3_3_1_18_2","unstructured":"Souvika Sarkar Mohammad\u00a0Fakhruddin Babar Monowar Hasan and Shubhra\u00a0Kanti Karmaker. 2024. LLMs as On-demand Customizable Service. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2401.16577 (2024)."},{"key":"e_1_3_3_1_19_2","doi-asserted-by":"crossref","unstructured":"Shreya Shankar JD Zamfirescu-Pereira Bj\u00f6rn Hartmann Aditya\u00a0G Parameswaran and Ian Arawjo. 2024. Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2404.12272 (2024).","DOI":"10.1145\/3654777.3676450"},{"key":"e_1_3_3_1_20_2","unstructured":"Annalisa Szymanski Simret\u00a0Araya Gebreegziabher Oghenemaro Anuyah Ronald\u00a0A Metoyer and Toby Jia-Jun Li. 2024. Comparing Criteria Development Across Domain Experts Lay Users and Models in Large Language Model Evaluation. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2410.02054 (2024)."},{"key":"e_1_3_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3641924"},{"key":"e_1_3_3_1_22_2","unstructured":"Annalisa Szymanski Noah Ziems Heather\u00a0A Eicher-Miller Toby Jia-Jun Li Meng Jiang and Ronald\u00a0A Metoyer. 2024. Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2410.20266 (2024)."},{"key":"e_1_3_3_1_23_2","unstructured":"Weixuan Wang Barry Haddow Alexandra Birch and Wei Peng. 2023. Assessing the reliability of large language model knowledge. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2310.09820 (2023)."},{"key":"e_1_3_3_1_24_2","unstructured":"Ian Webster. 2023. promptfoo: Test your prompts. https:\/\/www.promptfoo.dev\/."}],"event":{"name":"IUI '25: 30th International Conference on Intelligent User Interfaces Companion","location":"Cagliari Italy","acronym":"IUI '25","sponsor":["SIGAI ACM Special Interest Group on Artificial Intelligence","SIGCHI ACM Special Interest Group on Computer-Human Interaction"]},"container-title":["Companion Proceedings of the 30th International Conference on Intelligent User Interfaces"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3708557.3716146","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3708557.3716146","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3708557.3716146","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:17:46Z","timestamp":1750295866000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3708557.3716146"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,24]]},"references-count":23,"alternative-id":["10.1145\/3708557.3716146","10.1145\/3708557"],"URL":"https:\/\/doi.org\/10.1145\/3708557.3716146","relation":{},"subject":[],"published":{"date-parts":[[2025,3,24]]},"assertion":[{"value":"2025-03-24","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}