{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T01:40:50Z","timestamp":1787017250411,"version":"build-2736575974"},"reference-count":19,"publisher":"World Scientific Pub Co Pte Ltd","issue":"02","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Advs. Complex Syst."],"published-print":{"date-parts":[[2011,4]]},"abstract":"<jats:p>This paper investigates the impact of reward shaping in multi-agent reinforcement learning as a way to incorporate domain knowledge about good strategies. In theory, potential-based reward shaping does not alter the Nash Equilibria of a stochastic game, only the exploration of the shaped agent. We demonstrate empirically the performance of reward shaping in two problem domains within the context of RoboCup KeepAway by designing three reward shaping schemes, encouraging specific behaviour such as keeping a minimum distance from other players on the same team and taking on specific roles. The results illustrate that reward shaping with multiple, simultaneous learning agents can reduce the time needed to learn a suitable policy and can alter the final group performance.<\/jats:p>","DOI":"10.1142\/s0219525911002998","type":"journal-article","created":{"date-parts":[[2011,4,20]],"date-time":"2011-04-20T07:53:22Z","timestamp":1303286002000},"page":"251-278","source":"Crossref","is-referenced-by-count":38,"title":["AN EMPIRICAL STUDY OF POTENTIAL-BASED REWARD SHAPING AND ADVICE IN COMPLEX, MULTI-AGENT SYSTEMS"],"prefix":"10.1142","volume":"14","author":[{"given":"SAM","family":"DEVLIN","sequence":"first","affiliation":[{"name":"University of York, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"DANIEL","family":"KUDENKO","sequence":"additional","affiliation":[{"name":"University of York, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"MAREK","family":"GRZE\u015a","sequence":"additional","affiliation":[{"name":"University of Waterloo, CA, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"219","published-online":{"date-parts":[[2011,11,20]]},"reference":[{"key":"rf3","volume-title":"Dynamic Programming and Optimal Control (2 Vol Set)","author":"Bertsekas D. P.","year":"2007"},{"key":"rf4","volume-title":"Fun and Games \u2014 A Text on Game Theory","author":"Binmore K.","year":"1991"},{"key":"rf5","doi-asserted-by":"publisher","DOI":"10.1109\/TSMCC.2007.913919"},{"key":"rf10","volume-title":"Game Theory","author":"Fudenberg D.","year":"1991"},{"key":"rf13","first-page":"1039","volume":"4","author":"Hu J.","journal-title":"J. Mach. Learn. Res."},{"key":"rf15","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-11876-0_14"},{"key":"rf16","first-page":"251","author":"Maclin R.","journal-title":"Lect. Notes Artif. Int."},{"key":"rf18","unstructured":"M.\u00a0Mihaylov, K.\u00a0Tuyls and A.\u00a0Now\u00e9, Adaptive and Learning Agents (2009)\u00a0pp. 60\u201373."},{"key":"rf20","doi-asserted-by":"publisher","DOI":"10.2307\/1969529"},{"key":"rf23","doi-asserted-by":"publisher","DOI":"10.1002\/9780470316887"},{"key":"rf25","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2006.02.006"},{"key":"rf26","doi-asserted-by":"publisher","DOI":"10.1007\/11780519_9"},{"key":"rf27","doi-asserted-by":"publisher","DOI":"10.1177\/105971230501300301"},{"key":"rf28","first-page":"1038","author":"Sutton R.","journal-title":"Adv. Neur. In."},{"key":"rf30","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton R. S.","year":"1998"},{"key":"rf32","doi-asserted-by":"publisher","DOI":"10.1142\/S0219525909002301"},{"key":"rf34","first-page":"1603","author":"Wang X.","journal-title":"Adv. Neur. In."},{"key":"rf35","doi-asserted-by":"crossref","first-page":"205","DOI":"10.1613\/jair.1190","volume":"19","author":"Wiewiora E.","journal-title":"J. Artif. Intell. Res."},{"key":"rf37","volume-title":"An Introduction to MultiAgent Systems","author":"Wooldridge M.","year":"2002"}],"container-title":["Advances in Complex Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S0219525911002998","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,8,6]],"date-time":"2019-08-06T16:21:50Z","timestamp":1565108510000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/abs\/10.1142\/S0219525911002998"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2011,4]]},"references-count":19,"journal-issue":{"issue":"02","published-online":{"date-parts":[[2011,11,20]]},"published-print":{"date-parts":[[2011,4]]}},"alternative-id":["10.1142\/S0219525911002998"],"URL":"https:\/\/doi.org\/10.1142\/s0219525911002998","relation":{},"ISSN":["0219-5259","1793-6802"],"issn-type":[{"value":"0219-5259","type":"print"},{"value":"1793-6802","type":"electronic"}],"subject":[],"published":{"date-parts":[[2011,4]]}}}