{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,25]],"date-time":"2026-03-25T16:15:39Z","timestamp":1774455339064,"version":"3.50.1"},"publisher-location":"California","reference-count":0,"publisher":"International Joint Conferences on Artificial Intelligence Organization","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2021,8]]},"abstract":"<jats:p>Reinforcement Learning (RL) with sparse rewards is a major challenge. We pro-\n\npose Hindsight Trust Region Policy Optimization (HTRPO), a new RL algorithm\n\nthat extends the highly successful TRPO algorithm with hindsight to tackle the\n\nchallenge of sparse rewards. Hindsight refers to the algorithm\u2019s ability to learn\n\nfrom information across goals, including past goals not intended for the current\n\ntask. We derive the hindsight form of TRPO, together with QKL, a quadratic\n\napproximation to the KL divergence constraint on the trust region. QKL reduces\n\nvariance in KL divergence estimation and improves stability in policy updates. We\n\nshow that HTRPO has similar convergence property as TRPO. We also present\n\nHindsight Goal Filtering (HGF), which further improves the learning performance\n\nfor suitable tasks. HTRPO has been evaluated on various sparse-reward tasks,\n\nincluding Atari games and simulated robot control. Experimental results show that\n\nHTRPO consistently outperforms TRPO, as well as HPG, a state-of-the-art policy\n\n14 gradient algorithm for RL with sparse rewards.<\/jats:p>","DOI":"10.24963\/ijcai.2021\/459","type":"proceedings-article","created":{"date-parts":[[2021,8,11]],"date-time":"2021-08-11T07:00:49Z","timestamp":1628665249000},"page":"3335-3341","source":"Crossref","is-referenced-by-count":5,"title":["Hindsight Trust Region Policy Optimization"],"prefix":"10.24963","author":[{"given":"Hanbo","family":"Zhang","sequence":"first","affiliation":[{"name":"Xi'an Jiaotong University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Site","family":"Bai","sequence":"additional","affiliation":[{"name":"Xi'an Jiaotong University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xuguang","family":"Lan","sequence":"additional","affiliation":[{"name":"Xi'an Jiaotong University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"David","family":"Hsu","sequence":"additional","affiliation":[{"name":"National University of Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nanning","family":"Zheng","sequence":"additional","affiliation":[{"name":"Xi'an Jiaotong University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"10584","event":{"name":"Thirtieth International Joint Conference on Artificial Intelligence {IJCAI-21}","theme":"Artificial Intelligence","location":"Montreal, Canada","acronym":"IJCAI-2021","number":"30","sponsor":["International Joint Conferences on Artificial Intelligence Organization (IJCAI)"],"start":{"date-parts":[[2021,8,19]]},"end":{"date-parts":[[2021,8,27]]}},"container-title":["Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence"],"original-title":[],"deposited":{"date-parts":[[2021,8,11]],"date-time":"2021-08-11T07:03:30Z","timestamp":1628665410000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.ijcai.org\/proceedings\/2021\/459"}},"subtitle":[],"proceedings-subject":"Artificial Intelligence Research Articles","short-title":[],"issued":{"date-parts":[[2021,8]]},"references-count":0,"URL":"https:\/\/doi.org\/10.24963\/ijcai.2021\/459","relation":{},"subject":[],"published":{"date-parts":[[2021,8]]}}}