{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,5]],"date-time":"2026-06-05T03:08:12Z","timestamp":1780628892481,"version":"3.54.1"},"reference-count":17,"publisher":"MIT Press - Journals","issue":"6","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Neural Computation"],"published-print":{"date-parts":[[2002,6,1]]},"abstract":"<jats:p> We propose a modular reinforcement learning architecture for nonlinear, nonstationary control tasks, which we call multiple model-based reinforcement learning (MMRL). The basic idea is to decompose a complex task into multiple domains in space and time based on the predictability of the environmental dynamics. The system is composed of multiple modules, each of which consists of a state prediction model and a reinforcement learning controller. The \u201cresponsibility signal,\u201d which is given by the softmax function of the prediction errors, is used to weight the outputs of multiple modules, as well as to gate the learning of the prediction models and the reinforcement learning controllers. We formulate MMRL for both discrete-time, finite-state case and continuous-time, continuous-state case. The performance of MMRL was demonstrated for discrete case in a nonstationary hunting task in a grid world and for continuous case in a nonlinear, nonstationary control task of swinging up a pendulum with variable physical parameters. <\/jats:p>","DOI":"10.1162\/089976602753712972","type":"journal-article","created":{"date-parts":[[2002,7,27]],"date-time":"2002-07-27T11:58:45Z","timestamp":1027771125000},"page":"1347-1369","source":"Crossref","is-referenced-by-count":349,"title":["Multiple Model-Based Reinforcement Learning"],"prefix":"10.1162","volume":"14","author":[{"given":"Kenji","family":"Doya","sequence":"first","affiliation":[{"name":"Human Information Science Laboratories, ATR International, Seika, Soraku, Kyoto 619-0288, Japan; CREST, Japan Science and Technology Corporation, Seika, Soraku, Kyoto 619-0288, Japan; Kawato Dynamic Brain Project, ERATO, Japan Science and Technology Corporation, Seika, Soraku, Kyoto 619-0288, Japan; and Nara Institute of Science and Technology, Ikoma, Nara 630-0101, Japan,"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kazuyuki","family":"Samejima","sequence":"additional","affiliation":[{"name":"Human Information Science Laboratories, ATR International, Seika, Soraku, Kyoto 619-0288, Japan, and Kawato Dynamic Brain Project, ERATO, Japan Science and Technology Corporation, Seika, Soraku, Kyoto 619-0288, Japan,"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ken-ichi","family":"Katagiri","sequence":"additional","affiliation":[{"name":"ATR Human Information Processing Research Laboratories, Seika, Soraku, Kyoto 619-0288, Japan, and Nara Institute of Science and Technology, Ikoma, Nara 630-0101, Japan,"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mitsuo","family":"Kawato","sequence":"additional","affiliation":[{"name":"Human Information Science Laboratories, ATR International, Seika, Soraku, Kyoto 619-0288, Japan; Kawato Dynamic Brain Project, ERATO, Japan Science and Technology Corporation, Seika, Soraku, Kyoto 619-0288, Japan; and Nara Institute of Science and Technology, Ikoma, Nara 630-0101, Japan,"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","reference":[{"key":"p_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSMC.1983.6313077"},{"key":"p_5","doi-asserted-by":"publisher","DOI":"10.1162\/089976600300015961"},{"key":"p_6","doi-asserted-by":"publisher","DOI":"10.1016\/S0893-6080(05)80053-X"},{"key":"p_8","doi-asserted-by":"publisher","DOI":"10.1162\/089976601750541778"},{"key":"p_10","doi-asserted-by":"publisher","DOI":"10.1038\/35003194"},{"key":"p_11","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1991.3.1.79"},{"key":"p_13","doi-asserted-by":"publisher","DOI":"10.1016\/S0921-8890(01)00113-0"},{"key":"p_14","doi-asserted-by":"publisher","DOI":"10.1109\/37.387616"},{"key":"p_16","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1996.8.2.340"},{"key":"p_18","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992700"},{"key":"p_19","doi-asserted-by":"publisher","DOI":"10.1007\/BF00115009"},{"key":"p_21","doi-asserted-by":"publisher","DOI":"10.1016\/S0004-3702(99)00052-1"},{"key":"p_22","doi-asserted-by":"publisher","DOI":"10.1016\/S0893-6080(99)00060-X"},{"key":"p_23","doi-asserted-by":"publisher","DOI":"10.1177\/105971239700600202"},{"key":"p_24","doi-asserted-by":"publisher","DOI":"10.1038\/81497"},{"key":"p_25","doi-asserted-by":"publisher","DOI":"10.1016\/S0893-6080(98)00066-5"},{"key":"p_26","doi-asserted-by":"publisher","DOI":"10.1016\/S1364-6613(98)01221-2"}],"container-title":["Neural Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mitpressjournals.org\/doi\/pdf\/10.1162\/089976602753712972","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,3,12]],"date-time":"2021-03-12T21:49:39Z","timestamp":1615585779000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/neco\/article\/14\/6\/1347-1369\/6614"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2002,6,1]]},"references-count":17,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2002,6,1]]}},"alternative-id":["10.1162\/089976602753712972"],"URL":"https:\/\/doi.org\/10.1162\/089976602753712972","relation":{},"ISSN":["0899-7667","1530-888X"],"issn-type":[{"value":"0899-7667","type":"print"},{"value":"1530-888X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2002,6,1]]}}}