{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,8,7]],"date-time":"2024-08-07T07:37:42Z","timestamp":1723016262989},"publisher-location":"California","reference-count":0,"publisher":"International Joint Conferences on Artificial Intelligence Organization","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2018,7]]},"abstract":"<jats:p>We investigate the task of learning to interpret natural language instructions by jointly reasoning with visual observations and language inputs. Unlike current methods which start with learning from demonstrations (LfD) and then use reinforcement learning (RL) to fine-tune the model parameters, we propose a novel policy optimization algorithm which can dynamically schedule demonstration learning and RL. The proposed training paradigm provides efficient exploration and generalization beyond existing methods. Comparing to existing ensemble models, the best single model based on our proposed method tremendously decreases the execution error by 55% on a block-world environment. To further illustrate the exploration strategy of our RL algorithm, our paper includes systematic studies on the evolution of policy entropy during training.<\/jats:p>","DOI":"10.24963\/ijcai.2018\/626","type":"proceedings-article","created":{"date-parts":[[2018,7,5]],"date-time":"2018-07-05T05:49:10Z","timestamp":1530769750000},"page":"4503-4509","source":"Crossref","is-referenced-by-count":2,"title":["Scheduled Policy Optimization for Natural Language Communication with Intelligent Agents"],"prefix":"10.24963","author":[{"given":"Wenhan","family":"Xiong","sequence":"first","affiliation":[{"name":"University of California, Santa Barbara"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaoxiao","family":"Guo","sequence":"additional","affiliation":[{"name":"IBM Research"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mo","family":"Yu","sequence":"additional","affiliation":[{"name":"IBM Research"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shiyu","family":"Chang","sequence":"additional","affiliation":[{"name":"IBM Research"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bowen","family":"Zhou","sequence":"additional","affiliation":[{"name":"JD AI Research"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"William Yang","family":"Wang","sequence":"additional","affiliation":[{"name":"University of California, Santa Barbara"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"10584","event":{"number":"27","sponsor":["International Joint Conferences on Artificial Intelligence Organization (IJCAI)"],"acronym":"IJCAI-2018","name":"Twenty-Seventh International Joint Conference on Artificial Intelligence {IJCAI-18}","start":{"date-parts":[[2018,7,13]]},"theme":"Artificial Intelligence","location":"Stockholm, Sweden","end":{"date-parts":[[2018,7,19]]}},"container-title":["Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence"],"original-title":[],"deposited":{"date-parts":[[2018,7,5]],"date-time":"2018-07-05T05:54:38Z","timestamp":1530770078000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.ijcai.org\/proceedings\/2018\/626"}},"subtitle":[],"proceedings-subject":"Artificial Intelligence Research Articles","short-title":[],"issued":{"date-parts":[[2018,7]]},"references-count":0,"URL":"https:\/\/doi.org\/10.24963\/ijcai.2018\/626","relation":{},"subject":[],"published":{"date-parts":[[2018,7]]}}}