{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,22]],"date-time":"2025-10-22T23:37:05Z","timestamp":1761176225198,"version":"build-2065373602"},"reference-count":0,"publisher":"IOS Press","isbn-type":[{"value":"9781643686318","type":"electronic"}],"license":[{"start":{"date-parts":[[2025,10,21]],"date-time":"2025-10-21T00:00:00Z","timestamp":1761004800000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,10,21]]},"abstract":"<jats:p>Inverse reinforcement learning (IRL) aims to infer the reward function from expert demonstrations. However, as IRL techniques are increasingly applied in high-stakes domains such as autonomous driving and military decision-making, reward function leakage has emerged as a critical risk, potentially leading to severe security threats and unintended consequences. To address this challenge, we propose Safe Accelerated Policy Gradient (Safe APG), a method designed to enhance learning security of the demonstrating agent by preventing observers from inferring its reward function. The core idea behind Safe APG is to incorporate a delicately constructed and theoretically guaranteed structural noise into Nesterov\u2019s Accelerated Gradient (NAG) for policy updating, with the goal of concealing critical gradient information from the learning agent as well as keeping the geometric convergence property of NAG. The results from numerical experiments and simulations in reinforcement learning environments demonstrate that the proposed method not only significantly mitigates reward function leakage, but also achieves superior convergence rates even under the perturbation of the introduced structural noise.<\/jats:p>","DOI":"10.3233\/faia251154","type":"book-chapter","created":{"date-parts":[[2025,10,22]],"date-time":"2025-10-22T09:53:05Z","timestamp":1761126785000},"source":"Crossref","is-referenced-by-count":0,"title":["Safe APG: Accelerated Policy Gradient Algorithm for Secure Policy Updating in Reinforcement Learning"],"prefix":"10.3233","author":[{"given":"Jianan","family":"Lin","sequence":"first","affiliation":[{"name":"Southwestern University of Finance and Economics, Chengdu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yao","family":"Chen","sequence":"additional","affiliation":[{"name":"Southwestern University of Finance and Economics, Chengdu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhengyang","family":"Ji","sequence":"additional","affiliation":[{"name":"Zhejiang University of Technology, Hangzhou, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Meng","family":"Yuan","sequence":"additional","affiliation":[{"name":"Southwestern University of Finance and Economics, Chengdu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bo","family":"Hou","sequence":"additional","affiliation":[{"name":"Rocket Force Engineering University, Xi\u2019an, China"},{"name":"Northwestern Polytechnical University, Xi\u2019an, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shaolin","family":"Tan","sequence":"additional","affiliation":[{"name":"The Zhongguancun Laboratory, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"7437","container-title":["Frontiers in Artificial Intelligence and Applications","ECAI 2025"],"original-title":[],"link":[{"URL":"https:\/\/ebooks.iospress.nl\/pdf\/doi\/10.3233\/FAIA251154","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,22]],"date-time":"2025-10-22T09:53:15Z","timestamp":1761126795000},"score":1,"resource":{"primary":{"URL":"https:\/\/ebooks.iospress.nl\/doi\/10.3233\/FAIA251154"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,21]]},"ISBN":["9781643686318"],"references-count":0,"URL":"https:\/\/doi.org\/10.3233\/faia251154","relation":{},"ISSN":["0922-6389","1879-8314"],"issn-type":[{"value":"0922-6389","type":"print"},{"value":"1879-8314","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,21]]}}}