{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T07:02:02Z","timestamp":1782802922462,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":28,"publisher":"ACM","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,6,23]]},"DOI":"10.1145\/3715275.3732204","type":"proceedings-article","created":{"date-parts":[[2025,6,23]],"date-time":"2025-06-23T17:01:18Z","timestamp":1750698078000},"page":"3196-3206","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Detecting Prefix Bias in LLM-based Reward Models"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-8782-1388","authenticated-orcid":false,"given":"Ashwin","family":"Kumar","sequence":"first","affiliation":[{"name":"Washington University in St Louis, St Louis, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2083-1349","authenticated-orcid":false,"given":"Yuzi","family":"He","sequence":"additional","affiliation":[{"name":"Meta Platforms, Inc., Menlo Park, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3213-8333","authenticated-orcid":false,"given":"Aram H","family":"Markosyan","sequence":"additional","affiliation":[{"name":"Meta Platforms, Inc., Menlo Park, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-2313-4390","authenticated-orcid":false,"given":"Bobbie","family":"Chern","sequence":"additional","affiliation":[{"name":"Meta Platforms, Inc, Sunnyvale, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4402-2618","authenticated-orcid":false,"given":"Imanol","family":"Arrieta-Ibarra","sequence":"additional","affiliation":[{"name":"Independent, San Mateo, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,6,23]]},"reference":[{"key":"e_1_3_3_2_2_2","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia\u00a0Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat et\u00a0al. 2023. Gpt-4 technical report. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2303.08774 (2023)."},{"key":"e_1_3_3_2_3_2","unstructured":"Yuntao Bai Andy Jones Kamal Ndousse Amanda Askell Anna Chen Nova DasSarma Dawn Drain Stanislav Fort Deep Ganguli Tom Henighan et\u00a0al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2204.05862 (2022)."},{"key":"e_1_3_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/3442188.3445922"},{"key":"e_1_3_3_2_5_2","unstructured":"Stephen Casper Xander Davies Claudia Shi Thomas\u00a0Krendl Gilbert J\u00e9r\u00e9my Scheurer Javier Rando Rachel Freedman Tomasz Korbak David Lindner Pedro Freire et\u00a0al. 2023. Open problems and fundamental limitations of reinforcement learning from human feedback. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2307.15217 (2023)."},{"key":"e_1_3_3_2_6_2","unstructured":"Isha Chaudhary Qian Hu Manoj Kumar Morteza Ziyadi Rahul Gupta and Gagandeep Singh. 2024. Quantitative Certification of Bias in Large Language Models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2405.18780 (2024)."},{"key":"e_1_3_3_2_7_2","unstructured":"Thomas Coste Usman Anwar Robert Kirk and David Krueger. 2023. Reward model ensembles help mitigate overoptimization. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2310.02743 (2023)."},{"key":"e_1_3_3_2_8_2","doi-asserted-by":"crossref","unstructured":"Jesse Dodge Maarten Sap Ana Marasovi\u0107 William Agnew Gabriel Ilharco Dirk Groeneveld Margaret Mitchell and Matt Gardner. 2021. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2104.08758 (2021).","DOI":"10.18653\/v1\/2021.emnlp-main.98"},{"key":"e_1_3_3_2_9_2","unstructured":"Tyna Eloundou Alex Beutel David\u00a0G Robinson Keren Gu-Lemberg Anna-Luisa Brakman Pamela Mishkin Meghan Shah Johannes Heidecke Lilian Weng and Adam\u00a0Tauman Kalai. 2024. First-person fairness in chatbots. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2410.19803 (2024)."},{"key":"e_1_3_3_2_10_2","unstructured":"Kawin Ethayarajh Yejin Choi and Swabha Swayamdipta. 2022. Understanding Dataset Difficulty with \\(\\mathcal {V}\\) Chaudhuri Stefanie Jegelka Le\u00a0Song Csaba Szepesvari Gang Niu and Sivan Sabato (Eds.). PMLR 5988\u20136008."},{"key":"e_1_3_3_2_11_2","unstructured":"Isabel\u00a0O Gallegos Ryan\u00a0A Rossi Joe Barrow Md\u00a0Mehrab Tanjim Sungchul Kim Franck Dernoncourt Tong Yu Ruiyi Zhang and Nesreen\u00a0K Ahmed. 2023. Bias and fairness in large language models: A survey. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2309.00770 (2023)."},{"key":"e_1_3_3_2_12_2","first-page":"10835","volume-title":"International Conference on Machine Learning","author":"Gao Leo","year":"2023","unstructured":"Leo Gao, John Schulman, and Jacob Hilton. 2023. Scaling laws for reward model overoptimization. In International Conference on Machine Learning. PMLR, 10835\u201310866."},{"key":"e_1_3_3_2_13_2","unstructured":"Amit Haim Alejandro Salinas and Julian Nyarko. 2024. What\u2019s in a Name? Auditing Large Language Models for Race and Gender Bias. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2402.14875 (2024)."},{"key":"e_1_3_3_2_14_2","unstructured":"Anjali Kantharuban Jeremiah Milbauer Emma Strubell and Graham Neubig. 2024. Stereotype or Personalization? User Identity Biases Chatbot Recommendations. ArXiv abs\/2410.05613 (2024). https:\/\/api.semanticscholar.org\/CorpusID:273228304"},{"key":"e_1_3_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3582269.3615599"},{"key":"e_1_3_3_2_16_2","first-page":"22631","volume-title":"International Conference on Machine Learning","author":"Longpre Shayne","year":"2023","unstructured":"Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung\u00a0Won Chung, Yi Tay, Denny Zhou, Quoc\u00a0V Le, Barret Zoph, Jason Wei, et\u00a0al. 2023. The flan collection: Designing data and methods for effective instruction tuning. In International Conference on Machine Learning. PMLR, 22631\u201322648."},{"key":"e_1_3_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.findings-naacl.417"},{"key":"e_1_3_3_2_18_2","doi-asserted-by":"crossref","unstructured":"Jesutofunmi\u00a0A Omiye Jenna\u00a0C Lester Simon Spichak Veronica Rotemberg and Roxana Daneshjou. 2023. Large language models propagate race-based medicine. NPJ Digital Medicine 6 1 (2023) 195.","DOI":"10.1038\/s41746-023-00939-z"},{"key":"e_1_3_3_2_19_2","unstructured":"Long Ouyang Jeffrey Wu Xu Jiang Diogo Almeida Carroll Wainwright Pamela Mishkin Chong Zhang Sandhini Agarwal Katarina Slama Alex Ray et\u00a0al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems 35 (2022) 27730\u201327744."},{"key":"e_1_3_3_2_20_2","unstructured":"Rafael Rafailov Archit Sharma Eric Mitchell Christopher\u00a0D Manning Stefano Ermon and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems 36 (2024)."},{"key":"e_1_3_3_2_21_2","unstructured":"Baptiste Roziere Jonas Gehring Fabian Gloeckle Sten Sootla Itai Gat Xiaoqing\u00a0Ellen Tan Yossi Adi Jingyu Liu Tal Remez J\u00e9r\u00e9my Rapin et\u00a0al. 2023. Code llama: Open foundation models for code. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2308.12950 (2023)."},{"key":"e_1_3_3_2_22_2","unstructured":"John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/1707.06347 (2017)."},{"key":"e_1_3_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.330"},{"key":"e_1_3_3_2_24_2","unstructured":"Karan Singhal Tao Tu Juraj Gottweis Rory Sayres Ellery Wulczyn Le Hou Kevin Clark Stephen Pfohl Heather Cole-Lewis Darlene Neal et\u00a0al. 2023. Towards expert-level medical question answering with large language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2305.09617 (2023)."},{"key":"e_1_3_3_2_25_2","unstructured":"Nisan Stiennon Long Ouyang Jeffrey Wu Daniel Ziegler Ryan Lowe Chelsea Voss Alec Radford Dario Amodei and Paul\u00a0F Christiano. 2020. Learning to summarize with human feedback. Advances in Neural Information Processing Systems 33 (2020) 3008\u20133021."},{"key":"e_1_3_3_2_26_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et\u00a0al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2302.13971 (2023)."},{"key":"e_1_3_3_2_27_2","unstructured":"Hugo Touvron Louis Martin Kevin Stone Peter Albert Amjad Almahairi Yasmine Babaei Nikolay Bashlykov Soumya Batra Prajjwal Bhargava Shruti Bhosale et\u00a0al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2307.09288 (2023)."},{"key":"e_1_3_3_2_28_2","unstructured":"Susan Zhang Stephen Roller Naman Goyal Mikel Artetxe Moya Chen Shuohui Chen Christopher Dewan Mona Diab Xian Li Xi\u00a0Victoria Lin et\u00a0al. 2022. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2205.01068 (2022)."},{"key":"e_1_3_3_2_29_2","unstructured":"Yoshua\u00a0X ZXhang Yann\u00a0M Haxo and Ying\u00a0X Mat. 2023. Falcon llm: A new frontier in natural language processing. AC Investment Research Journal 220 44 (2023)."}],"event":{"name":"FAccT '25: The 2025 ACM Conference on Fairness, Accountability, and Transparency","location":"Athens Greece","acronym":"FAccT '25"},"container-title":["Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3715275.3732204","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,24]],"date-time":"2025-06-24T11:03:52Z","timestamp":1750763032000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3715275.3732204"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,23]]},"references-count":28,"alternative-id":["10.1145\/3715275.3732204","10.1145\/3715275"],"URL":"https:\/\/doi.org\/10.1145\/3715275.3732204","relation":{},"subject":[],"published":{"date-parts":[[2025,6,23]]},"assertion":[{"value":"2025-06-23","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}