{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T04:17:43Z","timestamp":1777522663600,"version":"3.51.4"},"reference-count":43,"publisher":"SAGE Publications","issue":"3-4","license":[{"start":{"date-parts":[[1997,1,1]],"date-time":"1997-01-01T00:00:00Z","timestamp":852076800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Adaptive Behavior"],"published-print":{"date-parts":[[1997,1]]},"abstract":"<jats:p>We present an approach to support effective learning and adaptation of behaviors for autonomous agents with reinforcement learning algorithms. These methods can identify control systems that optimize a reinforcement program, which is, usually, a straightforward representation of the designer's goals. Reinforcement learning algorithms usually are too slow to be applied in real time on embodied agents, although they provide a suitable way to represent the desired behavior. We have tackled three aspects of this problem: the speed of the algorithm, the learning procedure, and the control system architecture. The learning algorithm we have developed includes features to speed up learning, such as niche-based learning, and a representation of the control modules in terms of fuzzy rules that reduces the search space and improves robustness to noisy data. Our learning procedure exploits methodologies such as learning from easy missions and transfer of policy from simpler environments to the more complex. The architecture of our control system is layered and modular, so that each module has a low complexity and can be learned in a short time. The composition of the actions proposed by the modules is either learned or predefined. Finally, we adopt an anytime learning approach to improve the quality of the control system on-line and to adapt it to dynamic environments.<\/jats:p>\n                  <jats:p>The experiments we present in this article concern learning to reach another moving agent in a real, dynamic environment that includes nontrivial situations such as that in which the moving target is faster than the agent and that in which the target is hidden by obstacles.<\/jats:p>","DOI":"10.1177\/105971239700500304","type":"journal-article","created":{"date-parts":[[2007,3,17]],"date-time":"2007-03-17T21:21:19Z","timestamp":1174166479000},"page":"281-315","source":"Crossref","is-referenced-by-count":25,"title":["Anytime Learning and Adaptation of Structured Fuzzy Behaviors"],"prefix":"10.1177","volume":"5","author":[{"given":"Andrea","family":"Bonarini","sequence":"first","affiliation":[{"name":"Politecnico di Milano Artificial Intelligence and Robotics Project, Dipartimento di Elettronica e Informazione-Politecnico di Milano, Piazza Leonardo da Vinci, 32-20133 Milano, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[1997,1,1]]},"reference":[{"key":"atypb1","doi-asserted-by":"publisher","DOI":"10.1177\/105971239200100204"},{"key":"atypb2","volume-title":"Purposive behavior acquisition on a real robot by vision-based reinforcement learning. Proceedings of the Machine Learning Conference\u2014Conference on Learning Theory Workshop on Robot Learning","author":"Asada, M.","year":"1994"},{"key":"atypb3","volume-title":"Proceedings of the European Congress of Intelligent Techniques and Soft Computing (EUFIT '93)","author":"Bonarini, A."},{"key":"atypb4","volume-title":"Proceedings of the IEEE World Conference on Computational Intelligence\u2014Evolutionary Computation","author":"Bonarini, A."},{"key":"atypb5","volume-title":"Proceedings of the European Congress of Intelligent Techniques and Soft Computing (EUFIT '93)","author":"Bonarini, A."},{"key":"atypb6","volume-title":"Evolutionary learning of fuzzy rules: Competition and cooperation","author":"Bonarini, A.","year":"1996"},{"key":"atypb7","volume-title":"Delayed reinforcement, fuzzy Q-learning and fuzzy logic controllers","author":"Bonarini, A.","year":"1996"},{"key":"atypb8","volume-title":"Proceedings of the International Conference on Machine Learning (ICML '96) Pre-Conference Workshop on Evolutionary Computing and Machine Learning","author":"Bonarini, A."},{"key":"atypb9","volume-title":"Symbol grounding and a neuro-fuzzy architecture for multisensor fusion. Proceedings of the World Automation Congress '96","author":"Bonarini, A."},{"key":"atypb10","doi-asserted-by":"publisher","DOI":"10.1007\/BF00113896"},{"key":"atypb11","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(89)90050-7"},{"key":"atypb12","doi-asserted-by":"publisher","DOI":"10.1109\/JRA.1986.1087032"},{"key":"atypb13","volume-title":"The behavior language: User's guide. AI Memo 1227","author":"Brooks, R.A.","year":"1990"},{"key":"atypb14","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4615-3184-5_8"},{"key":"atypb15","doi-asserted-by":"publisher","DOI":"10.1109\/72.159059"},{"key":"atypb16","volume-title":"Proceedings of the Seventh National Conference on AI (AAAI-88)","author":"Dean, T."},{"key":"atypb17","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(94)90047-7"},{"key":"atypb18","volume-title":"Fuzzy sets and systems: Theory and applications","author":"Dubois, D.","year":"1980"},{"key":"atypb19","volume-title":"Proceedings of the Third Fuzz-IEEE","author":"Glorennec, P.Y."},{"key":"atypb20","volume-title":"Genetic algorithms in search, optimization and machine learning","author":"Goldberg, D.E.","year":"1989"},{"key":"atypb21","volume-title":"An approach to anytime learning. Proceedings of the Ninth International Workshop on Machine Learning (ML92)","author":"Grefenstette, J.J.","year":"1992"},{"key":"atypb22","volume-title":"Genetic learning for adaptation in autonomous robots. Proceedings of the World Automation Congress '96","author":"Grefenstette, J.J."},{"key":"atypb23","doi-asserted-by":"publisher","DOI":"10.1007\/BF00344744"},{"key":"atypb24","volume-title":"Theories of learning","author":"Hilgard, E.L.","year":"1975"},{"key":"atypb25","volume-title":"Properties of the bucket brigade algorithm. Proceedings of the First International Conference on Genetic Algorithms","author":"Holland, J.H.","year":"1985"},{"key":"atypb26","doi-asserted-by":"publisher","DOI":"10.1613\/jair.301"},{"key":"atypb27","volume-title":"Optimal control theory: An introduction","author":"Kirk, D.E.","year":"1970"},{"key":"atypb28","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(92)90058-6"},{"key":"atypb29","volume-title":"A distributed model for mobile robot environment-learning and navigation. Technical Report 1228","author":"Matari\u0107, M.J.","year":"1989"},{"key":"atypb30","doi-asserted-by":"publisher","DOI":"10.1177\/105971239500400104"},{"key":"atypb31","doi-asserted-by":"publisher","DOI":"10.1023\/A:1008819414322"},{"key":"atypb32","volume-title":"Fuzzy control and fuzzy systems","author":"Pedrycz, W.","year":"1992"},{"key":"atypb33","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(94)00088-I"},{"key":"atypb34","volume-title":"Adaptive control: Stability, convergence, and robustness","author":"Sastry, S.","year":"1989"},{"key":"atypb35","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992700"},{"key":"atypb36","volume-title":"Temporal credit assignment in reinforcement learning. Unpublished doctoral thesis","author":"Sutton, R.S.","year":"1984"},{"key":"atypb37","doi-asserted-by":"publisher","DOI":"10.1007\/BF00115009"},{"key":"atypb38","volume-title":"The fuzzy classifier system: A classifier system for continuously varying variables. Proceedings of the Fourth International Conference on Genetic Algorithms","author":"Valenzuela-Rend\u00f3n, M.","year":"1991"},{"key":"atypb39","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992698"},{"key":"atypb40","doi-asserted-by":"publisher","DOI":"10.1007\/BF00058926"},{"key":"atypb41","volume-title":"Knowledge growth in an artificial animal. Proceedings of the Fitst International Conference on Genetic Algorithms and their Applications","author":"Wilson, S.W.","year":"1985"},{"key":"atypb42","doi-asserted-by":"publisher","DOI":"10.1162\/evco.1994.2.1.1"},{"key":"atypb43","doi-asserted-by":"publisher","DOI":"10.1016\/S0019-9958(65)90241-X"}],"container-title":["Adaptive Behavior"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/105971239700500304","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/105971239700500304","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T16:18:20Z","timestamp":1777393100000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/105971239700500304"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[1997,1]]},"references-count":43,"journal-issue":{"issue":"3-4","published-print":{"date-parts":[[1997,1]]}},"alternative-id":["10.1177\/105971239700500304"],"URL":"https:\/\/doi.org\/10.1177\/105971239700500304","relation":{},"ISSN":["1059-7123","1741-2633"],"issn-type":[{"value":"1059-7123","type":"print"},{"value":"1741-2633","type":"electronic"}],"subject":[],"published":{"date-parts":[[1997,1]]}}}