{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,8,7]],"date-time":"2024-08-07T07:35:59Z","timestamp":1723016159535},"publisher-location":"California","reference-count":0,"publisher":"International Joint Conferences on Artificial Intelligence Organization","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2020,7]]},"abstract":"<jats:p>A major challenge in reinforcement learning is exploration, when local dithering methods such as epsilon-greedy sampling are insufficient to solve a given task. Many recent methods have proposed to intrinsically motivate an agent to seek novel states, driving the agent to discover improved reward. However, while state-novelty exploration methods are suitable for tasks where novel observations correlate well with improved reward, they may not explore more efficiently than epsilon-greedy approaches in environments where the two are not well-correlated. In this paper, we distinguish between exploration tasks in which seeking novel states aids in finding new reward, and those where it does not, such as goal-conditioned tasks and escaping local reward maxima. We propose a new exploration objective, maximizing the reward prediction error (RPE) of a value function trained to predict extrinsic reward. We then propose a deep reinforcement learning method, QXplore, which exploits the temporal difference error of a Q-function to solve hard exploration tasks in high-dimensional MDPs. We demonstrate the exploration behavior of QXplore on several OpenAI Gym MuJoCo tasks and Atari games and observe that QXplore is comparable to or better than a baseline state-novelty method in all cases, outperforming the baseline on tasks where state novelty is not well-correlated with improved reward.<\/jats:p>","DOI":"10.24963\/ijcai.2020\/390","type":"proceedings-article","created":{"date-parts":[[2020,7,8]],"date-time":"2020-07-08T12:12:10Z","timestamp":1594210330000},"page":"2816-2823","source":"Crossref","is-referenced-by-count":1,"title":["Reward Prediction Error as an Exploration Objective in Deep RL"],"prefix":"10.24963","author":[{"given":"Riley","family":"Simmons-Edler","sequence":"first","affiliation":[{"name":"Princeton University"},{"name":"Samsung AI Center NYC"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ben","family":"Eisner","sequence":"additional","affiliation":[{"name":"Samsung AI Center NYC"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Daniel","family":"Yang","sequence":"additional","affiliation":[{"name":"Samsung AI Center NYC"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anthony","family":"Bisulco","sequence":"additional","affiliation":[{"name":"Samsung AI Center NYC"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eric","family":"Mitchell","sequence":"additional","affiliation":[{"name":"Samsung AI Center NYC"},{"name":"Stanford University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sebastian","family":"Seung","sequence":"additional","affiliation":[{"name":"Samsung AI Center NYC"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Daniel","family":"Lee","sequence":"additional","affiliation":[{"name":"Samsung AI Center NYC"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"10584","event":{"number":"28","sponsor":["International Joint Conferences on Artificial Intelligence Organization (IJCAI)"],"acronym":"IJCAI-PRICAI-2020","name":"Twenty-Ninth International Joint Conference on Artificial Intelligence and Seventeenth Pacific Rim International Conference on Artificial Intelligence {IJCAI-PRICAI-20}","start":{"date-parts":[[2020,7,11]]},"theme":"Artificial Intelligence","location":"Yokohama, Japan","end":{"date-parts":[[2020,7,17]]}},"container-title":["Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence"],"original-title":[],"deposited":{"date-parts":[[2020,7,9]],"date-time":"2020-07-09T02:14:52Z","timestamp":1594260892000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.ijcai.org\/proceedings\/2020\/390"}},"subtitle":[],"proceedings-subject":"Artificial Intelligence Research Articles","short-title":[],"issued":{"date-parts":[[2020,7]]},"references-count":0,"URL":"https:\/\/doi.org\/10.24963\/ijcai.2020\/390","relation":{},"subject":[],"published":{"date-parts":[[2020,7]]}}}