{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T04:07:13Z","timestamp":1777522033689,"version":"3.51.4"},"reference-count":26,"publisher":"SAGE Publications","issue":"3","license":[{"start":{"date-parts":[[1994,1,1]],"date-time":"1994-01-01T00:00:00Z","timestamp":757382400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Adaptive Behavior"],"published-print":{"date-parts":[[1994,1]]},"abstract":"<jats:p>This article is concerned with training an agent to perform sequential behavior. In previous work, we have been applying reinforcement learning techniques to control a reactive agent. Obviously, a purely reactive system is limited in the kind of interactions it can learn. In particular, it can learn what we call pseudosequences\u2014that is, sequences of actions in which each action is selected on the basis of current sensory stimuli. It cannot learn proper sequences, in which actions must be selected also on the basis of some internal state. Moreover, it is a result of our research that effective learning of proper sequences is improved by letting the agent and the trainer communicate. First, we consider trainer-to-agent communication, introducing the concept of reinforcement sensor, which lets the learning robot explicitly know whether the last reinforcement was a reward or a punishment. We also show how the use of this sensor makes error recovery rules emerge. Then we introduce agent-to-trainer communication, which is used to disambiguate ambiguous training situations\u2014that is, situations in which the observation of the agent's behavior does not provide the trainer with enough information to decide whether the agent's move is right or wrong. We also show an alternative solution to the problem of ambiguous situations, which involves learning to coordinate behavior in a simpler, unambiguous setting and then transferring what has been learned to a more complex situation. All the design choices we make are discussed and compared by means of experiments in a simulated world.<\/jats:p>","DOI":"10.1177\/105971239400200302","type":"journal-article","created":{"date-parts":[[2007,3,11]],"date-time":"2007-03-11T01:45:25Z","timestamp":1173577525000},"page":"247-275","source":"Crossref","is-referenced-by-count":36,"title":["Training Agents to Perform Sequential Behavior"],"prefix":"10.1177","volume":"2","author":[{"given":"Marco","family":"Colombetti","sequence":"first","affiliation":[{"name":"Politecnico di Milano"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Marco","family":"Dorigo","sequence":"additional","affiliation":[{"name":"Universit\u00e9 Libre de Bruxelles"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[1994,1,1]]},"reference":[{"key":"atypb1","author":"Beer, R.D.","year":"1994","journal-title":"Artificial Intelligence"},{"key":"atypb2","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(93)90071-I"},{"key":"atypb3","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(89)90050-7"},{"key":"atypb4","volume-title":"Proceedings of \"From Animals to Animats,\" Second International Conference on Simulation of Adaptive Behavior (SAB92)","author":"Colombetti, M."},{"key":"atypb5","volume-title":"ALECSYS and the AutonoMouse: Learning to control a real robot by distributed classifier systems (Tech. Rep. No. 92-011)","author":"Dorigo, M.","year":"1992"},{"key":"atypb6","doi-asserted-by":"publisher","DOI":"10.1162\/evco.1993.1.2.151"},{"key":"atypb7","author":"Dorigo, M.","year":"1994","journal-title":"Artificial Intelligence"},{"key":"atypb8","doi-asserted-by":"publisher","DOI":"10.1109\/21.214773"},{"key":"atypb9","volume-title":"Proceedings of the Fourth International Conference on Genetic Algorithms","author":"Dorigo, M."},{"key":"atypb10","doi-asserted-by":"publisher","DOI":"10.1007\/BF00113893"},{"key":"atypb11","volume-title":"Genetic algorithms in search, optimization and machine learning","author":"Goldberg, D.E.","year":"1989"},{"key":"atypb12","volume-title":"Adaptation in natural and artificial systems","author":"Holland, J.H.","year":"1975"},{"key":"atypb13","volume-title":"Escaping brittleness: The possibilities of general purpose learning algorithms applied to parallel rule-based systems","author":"Holland, J.H.","year":"1986"},{"key":"atypb14","volume-title":"Memory approaches to reinforcement learning in non-Markovian domains (Tech. Rep. CMU-CS-92-138)","author":"Lin, L.-J.","year":"1992"},{"key":"atypb15","volume-title":"Proceedings of \"From Animals to Animats,\" Second International Conference on Simulation of Adaptive Behavior (SAB92)","author":"Littman, M.L."},{"key":"atypb16","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(92)90058-6"},{"key":"atypb17","volume-title":"Logic design principles","author":"McCluskey, E.J.","year":"1986"},{"key":"atypb18","volume-title":"Proceedings of the Third International Conference on Genetic Algorithms","author":"Riolo, R.L."},{"key":"atypb19","volume-title":"Proceedings of the 1986 Conference on Theoretical Aspects of Reasoning About Knowledge","author":"Rosenschein, S.J."},{"key":"atypb20","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992700"},{"key":"atypb21","volume-title":"In Proceedings of the Fourth International Conference on Genetic Algorithms","author":"Spiessens, P."},{"key":"atypb22","volume-title":"Learning with delayed rewards","author":"Watkins, C.J.C.H.","year":"1989"},{"key":"atypb23","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992698"},{"key":"atypb24","author":"Whitehead, S.D.","year":"1994","journal-title":"Artificial Intelligence"},{"key":"atypb25","doi-asserted-by":"publisher","DOI":"10.1007\/BF00058679"},{"key":"atypb26","volume-title":"First International Conference on the Simulation of Adaptive Behavior (SAB90)","author":"Wilson, S."}],"container-title":["Adaptive Behavior"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/105971239400200302","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/105971239400200302","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T16:16:08Z","timestamp":1777392968000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/105971239400200302"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[1994,1]]},"references-count":26,"journal-issue":{"issue":"3","published-print":{"date-parts":[[1994,1]]}},"alternative-id":["10.1177\/105971239400200302"],"URL":"https:\/\/doi.org\/10.1177\/105971239400200302","relation":{},"ISSN":["1059-7123","1741-2633"],"issn-type":[{"value":"1059-7123","type":"print"},{"value":"1741-2633","type":"electronic"}],"subject":[],"published":{"date-parts":[[1994,1]]}}}