{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T04:04:00Z","timestamp":1777521840947,"version":"3.51.4"},"reference-count":24,"publisher":"SAGE Publications","issue":"4","license":[{"start":{"date-parts":[[2006,12,1]],"date-time":"2006-12-01T00:00:00Z","timestamp":1164931200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Adaptive Behavior"],"published-print":{"date-parts":[[2006,12]]},"abstract":"<jats:p>Multi-agent reinforcement learning (MRL) is a growing area of research. What makes it particularly challenging is that multiple learners render each other's environments non-stationary. In addition to adapting their behaviors to other learning agents, online learners must also provide assurances about their online performance in order to promote user trust of adaptive agent systems deployed in real world applications. In this article, instead of developing new algorithms with such assurances, we study the question of safety in online performance of some existing MRL algorithms. We identify the key notion of reactivity of a learner by analyzing how an algorithm (PHC-Exploiter), designed to exploit some simpler opponents, can itself be exploited by them. We quantify and analyze this concept of reac tivity in the context of these algorithms to explain their experimental behaviors. We argue that no learner can be designed that can deliberately avoid exploitation. We also show that any attempt to opti mize reactivity must take into account a tradeoff with sensitivity to noise, and devise an adaptive method (based on environmental feedback) designed to maximize the learner's safety and minimize its sensitivity to noise.<\/jats:p>","DOI":"10.1177\/1059712306072334","type":"journal-article","created":{"date-parts":[[2006,11,8]],"date-time":"2006-11-08T05:57:46Z","timestamp":1162965466000},"page":"339-356","source":"Crossref","is-referenced-by-count":0,"title":["Reactivity and Safe Learning in Multi-Agent Systems"],"prefix":"10.1177","volume":"14","author":[{"given":"Bikramjit","family":"Banerjee","sequence":"first","affiliation":[{"name":"Department of Electrical Engineering and Computer Science, Tulane                         University, New Orleans, LA 70118, USA,"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jing","family":"Peng","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering and Computer Science, Tulane                         University, New Orleans, LA 70118, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2006,12,1]]},"reference":[{"key":"atypb1","first-page":"322","volume-title":"Proceedings of the 36th Annual Symposium on Foundations of Computer Science","author":"Auer, P."},{"key":"atypb2","volume-title":"Proceedings of the Sixth International Workshop on Trust, Privacy, Deception, and Fraud in Agent Societies","author":"Banerjee, B."},{"key":"atypb3","first-page":"2","volume-title":"Proceedings of the 19th National Conference on Artificial Intelligence (AAAI-04)","author":"Banerjee, B."},{"key":"atypb4","first-page":"825","volume-title":"Proceedings of the Seventeenth International Joint Conference on Artificial Intelligence (IJCAI)","author":"Banerjee, B."},{"key":"atypb5","first-page":"209","volume-title":"Advances in Neural Information Processing Systems 17","author":"Bowling, M.","year":"2005"},{"key":"atypb6","doi-asserted-by":"publisher","DOI":"10.1016\/S0004-3702(02)00121-2"},{"key":"atypb7","volume-title":"AAAI Workshop Proceedings on Game Theoretic and Decision Theoretic Agents","author":"Bowling, M."},{"key":"atypb8","first-page":"1483","volume-title":"Advances in Neural Information Processing Systems, 14","author":"Chang, Y. H.","year":"2001"},{"key":"atypb9","unstructured":"Claus, C. &                      Boutilier, C.                  (1998). The dynamics of reinforcement learning in cooperative                     multiagent systems. In Proceedings of the 15th National                     Conference on Artificial Intelligence (pp.                      746\u2013752                 ). Menlo Park, CA: AAAI Press\/ MIT Press."},{"key":"atypb10","doi-asserted-by":"publisher","DOI":"10.1006\/game.1999.0738"},{"key":"atypb11","doi-asserted-by":"publisher","DOI":"10.1016\/0165-1889(94)00819-4"},{"key":"atypb12","doi-asserted-by":"publisher","DOI":"10.1613\/jair.720"},{"key":"atypb13","first-page":"242","volume-title":"Proceedings of the 15th International Conference on Machine Learning (ML'98)","author":"Hu, J."},{"key":"atypb14","first-page":"522","volume-title":"Proceedings of the 3rd International Joint Conference on Autonomous Agents and Multi-Agent Systems","author":"Lee, C."},{"key":"atypb15","first-page":"157","volume-title":"Proceedings of the 11th International Conference on Machine Learning","author":"Littman, M. L."},{"key":"atypb16","doi-asserted-by":"publisher","DOI":"10.2307\/2171894"},{"key":"atypb17","doi-asserted-by":"publisher","DOI":"10.2307\/1969529"},{"key":"atypb18","doi-asserted-by":"publisher","DOI":"10.1038\/364056a0"},{"key":"atypb19","volume-title":"Game theory","author":"Owen, G.","year":"1995"},{"key":"atypb20","first-page":"1089","volume-title":"Advances in Neural Information Processing Systems 17","author":"Powers, R.","year":"2005"},{"key":"atypb21","volume-title":"The online proceedings of the 1st and the 2nd international workshop on safety and security in multiagent systems","author":"SASEMAS"},{"key":"atypb22","first-page":"541","volume-title":"Proceedings of the Sixteenth Conference on Uncertainty in Artificial Intelligence","author":"Singh, S."},{"key":"atypb23","first-page":"1057","volume-title":"Advances of Neural Information Processing Systems, 12","author":"Sutton, R. S.","year":"2000"},{"key":"atypb24","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4419-8909-3_2"}],"container-title":["Adaptive Behavior"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1059712306072334","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1059712306072334","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T16:15:33Z","timestamp":1777392933000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1059712306072334"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2006,12]]},"references-count":24,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2006,12]]}},"alternative-id":["10.1177\/1059712306072334"],"URL":"https:\/\/doi.org\/10.1177\/1059712306072334","relation":{},"ISSN":["1059-7123","1741-2633"],"issn-type":[{"value":"1059-7123","type":"print"},{"value":"1741-2633","type":"electronic"}],"subject":[],"published":{"date-parts":[[2006,12]]}}}