{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,24]],"date-time":"2025-09-24T00:14:36Z","timestamp":1758672876852,"version":"3.44.0"},"publisher-location":"California","reference-count":0,"publisher":"International Joint Conferences on Artificial Intelligence Organization","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,9]]},"abstract":"<jats:p>Human preference alignment (HPA) aims to ensure Large Language Models (LLMs) responding appropriately to meet human moral and ethical requirements. Existing methods, such as RLHF and DPO, rely heavily on high-quality human annotation, which restrict the efficiency of iterative online model refinement.\n\n\t\tTo address the inefficiencies of human annotation acquisition, iterated online strategy advocates the use of fine-tuned LLMs to self-generate preference data. However, this approach is prone to distribution bias, because of differences between human and model annotations, as well as modeling errors between simulators and real-world contexts. To mitigate the impact of distribution bias, we adopt the principles of adversarial training, framing a zero-sum two-player game with a protagonist agent and an adversarial agent. With the adversarial agent challenging the alignment of protagonist agent, we continuously refine the protagonist\u2019s performance. By utilizing min-max equilibrium and Nash equilibrium strategies, we propose Indirect Online Preference Optimization (IOPO) mechanism that enables the protagonist agent to converge without bias while maintaining linear computational complexity. Extensive experiments across three real-world datasets demonstrate that IOPO outperforms state-of-the-art alignment methods in both offline and online scenarios, evidenced by standard alignment metrics and human evaluations. This innovation reduces the time required for model iterations from months to one week, alleviates distribution shifts, and significantly cuts annotation costs.<\/jats:p>","DOI":"10.24963\/ijcai.2025\/61","type":"proceedings-article","created":{"date-parts":[[2025,9,19]],"date-time":"2025-09-19T08:10:40Z","timestamp":1758269440000},"page":"538-546","source":"Crossref","is-referenced-by-count":0,"title":["Indirect Online Preference Optimization via Reinforcement Learning"],"prefix":"10.24963","author":[{"given":"En","family":"Wang","sequence":"first","affiliation":[{"name":"Jilin University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xingyu","family":"Lin","sequence":"additional","affiliation":[{"name":"Jilin University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Du","family":"Su","sequence":"additional","affiliation":[{"name":"Institute of Computing Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chenfu","family":"Bao","sequence":"additional","affiliation":[{"name":"Baidu Inc."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhonghou","family":"Lv","sequence":"additional","affiliation":[{"name":"Baidu Inc."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Funing","family":"Yang","sequence":"additional","affiliation":[{"name":"Jilin University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuanbo","family":"Xu","sequence":"additional","affiliation":[{"name":"Jilin University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenbin","family":"Liu","sequence":"additional","affiliation":[{"name":"Jilin University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"10584","event":{"number":"34","sponsor":["International Joint Conferences on Artificial Intelligence Organization (IJCAI)"],"acronym":"IJCAI-2025","name":"Thirty-Fourth International Joint Conference on Artificial Intelligence {IJCAI-25}","start":{"date-parts":[[2025,8,16]]},"theme":"Artificial Intelligence","location":"Montreal, Canada","end":{"date-parts":[[2025,8,22]]}},"container-title":["Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence"],"original-title":[],"deposited":{"date-parts":[[2025,9,23]],"date-time":"2025-09-23T11:32:47Z","timestamp":1758627167000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.ijcai.org\/proceedings\/2025\/61"}},"subtitle":[],"proceedings-subject":"Artificial Intelligence Research Articles","short-title":[],"issued":{"date-parts":[[2025,9]]},"references-count":0,"URL":"https:\/\/doi.org\/10.24963\/ijcai.2025\/61","relation":{},"subject":[],"published":{"date-parts":[[2025,9]]}}}