{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,5,18]],"date-time":"2025-05-18T05:04:13Z","timestamp":1747544653207},"reference-count":5,"publisher":"World Scientific Pub Co Pte Lt","issue":"03","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Int. J. Neur. Syst."],"published-print":{"date-parts":[[1999,6]]},"abstract":"<jats:p> Autonomous learning techniques are based on experience acquisition. In most realistic applications, experience is time-consuming: it implies sensor reading, actuator control and algorithmic update, constrained by the learning system dynamics. The information crudeness upon which classical learning algorithms operate make such problems too difficult and unrealistic, Nonetheless, additional information for facilitating the learning process ideally should be embedded in such a way that the structural, well-studied characteristics of these fundamental algorithms are maintained. We investigate in this article a more general formulation of the Q-learning method that allows for a spreading of information derived from single updates towards a neighbourhood of the instantly visited state and converges to optimality. We show how this new formulation can be used as a mechanism to safely embed prior knowledge about the structure of the state space, and demonstrate it in a modified implementation of a reinforcement learning algorithm in a real robot navigation task. <\/jats:p>","DOI":"10.1142\/s0129065799000241","type":"journal-article","created":{"date-parts":[[2003,5,14]],"date-time":"2003-05-14T07:53:01Z","timestamp":1052898781000},"page":"243-249","source":"Crossref","is-referenced-by-count":2,"title":["Autonomous Learning Based on Cost Assumptions: Theoretical Studies And Experiments in Robot Control"],"prefix":"10.1142","volume":"09","author":[{"given":"CARLOS H.C.","family":"RIBEIRO","sequence":"first","affiliation":[{"name":"Technological Institute of Aeronautics, Pra\u00e7a Mal. Eduardo Gomes, 50,  12228\u2013900 S\u00e3o Jos\u00e9 dos Campos \u2013 SP, Brazil"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"ELDER M.","family":"HEMERLY","sequence":"additional","affiliation":[{"name":"Technological Institute of Aeronautics, Pra\u00e7a Mal. Eduardo Gomes, 50,  12228\u2013900 S\u00e3o Jos\u00e9 dos Campos \u2013 SP, Brazil"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2011,11,21]]},"reference":[{"key":"p_4","doi-asserted-by":"publisher","DOI":"10.1109\/3477.499792"},{"key":"p_5","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1994.6.6.1185"},{"key":"p_9","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(92)90058-6"},{"key":"p_17","first-page":"02912","author":"Szepesvari C.","year":"1996","journal-title":"Rhode Island"},{"key":"p_18","first-page":"59","volume":"22","author":"Tsitsiklis J. N.","year":"1996","journal-title":"Machine Learning"}],"container-title":["International Journal of Neural Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S0129065799000241","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,8,6]],"date-time":"2019-08-06T22:08:47Z","timestamp":1565129327000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/abs\/10.1142\/S0129065799000241"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[1999,6]]},"references-count":5,"journal-issue":{"issue":"03","published-online":{"date-parts":[[2011,11,21]]},"published-print":{"date-parts":[[1999,6]]}},"alternative-id":["10.1142\/S0129065799000241"],"URL":"https:\/\/doi.org\/10.1142\/s0129065799000241","relation":{},"ISSN":["0129-0657","1793-6462"],"issn-type":[{"value":"0129-0657","type":"print"},{"value":"1793-6462","type":"electronic"}],"subject":[],"published":{"date-parts":[[1999,6]]}}}