{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,16]],"date-time":"2025-10-16T00:47:32Z","timestamp":1760575652692,"version":"build-2065373602"},"reference-count":0,"publisher":"Association for the Advancement of Artificial Intelligence (AAAI)","issue":"2","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["AIES"],"abstract":"<jats:p>Offline reinforcement learning (RL) offers a promising path\nfor training domain-specific conversational agents (CAs)\nusing large-scale historical dialogue data, without the\nneed for costly online interactions or human annotations.\nIn the legal domain, vast amounts of publicly available\ncourtroom transcripts provide a rich and underutilized\nresource for developing intelligent legal CAs. However,\noffline training suffers from distribution shift between\nthe learned policy and the behavior policy embedded in the\ntraining data, which can degrade agent performance at\ndeployment. We address this challenge with a novel offline\nRL method, Reward-on-the-Line (ROL), which calibrates\nrewards based on action-selection agreement among an\nensemble of CAs. We apply ROL to the U.S. Supreme Court\ndataset to demonstrate its effectiveness in learning\nproactive, legally-informed dialogue strategies from\nhistorical court proceedings. To show the broader\napplicability of our approach, we also evaluate ROL on the\nCraigslistBargain negotiation dataset. Results in both\ndomains confirm that ROL reduces distribution shift and\nimproves agent performance in unseen dialogue scenarios.<\/jats:p>","DOI":"10.1609\/aies.v8i2.36657","type":"journal-article","created":{"date-parts":[[2025,10,15]],"date-time":"2025-10-15T13:20:43Z","timestamp":1760534443000},"page":"1575-1584","source":"Crossref","is-referenced-by-count":0,"title":["Reward-on-the-Line: A Novel Offline Reinforcement Learning Method for Building Legal Conversational Agents"],"prefix":"10.1609","volume":"8","author":[{"given":"Xubo","family":"Lin","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mingze","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Grace Hui","family":"Yang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Daniel","family":"Chen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"9382","published-online":{"date-parts":[[2025,10,15]]},"container-title":["Proceedings of the AAAI\/ACM Conference on AI, Ethics, and Society"],"original-title":[],"link":[{"URL":"https:\/\/ojs.aaai.org\/index.php\/AIES\/article\/download\/36657\/38795","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/ojs.aaai.org\/index.php\/AIES\/article\/download\/36657\/38795","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,15]],"date-time":"2025-10-15T13:20:44Z","timestamp":1760534444000},"score":1,"resource":{"primary":{"URL":"https:\/\/ojs.aaai.org\/index.php\/AIES\/article\/view\/36657"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,15]]},"references-count":0,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2025,10,15]]}},"URL":"https:\/\/doi.org\/10.1609\/aies.v8i2.36657","relation":{},"ISSN":["3065-8365"],"issn-type":[{"value":"3065-8365","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,15]]}}}