{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2022,12,25]],"date-time":"2022-12-25T05:30:42Z","timestamp":1671946242351},"reference-count":21,"publisher":"World Scientific Pub Co Pte Lt","issue":"04n05","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Advs. Complex Syst."],"published-print":{"date-parts":[[2009,8]]},"abstract":"<jats:p> In large, distributed systems composed of adaptive and interactive components (agents), ensuring the coordination among the agents so that the system achieves certain performance objectives is a challenging proposition. The key difficulty to overcome in such systems is one of credit assignment: How to apportion credit (or blame) to a particular agent based on the performance of the entire system. In this paper, we show how this problem can be solved in general for a large class of reward functions whose analytical form may be unknown (hence \"black box\" reward). This method combines the salient features of global solutions (e.g. \"team games\") which are broadly applicable but provide poor solutions in large problems with those of local solutions (e.g. \"difference rewards\") which learn quickly, but can be computationally burdensome. We introduce two estimates for local rewards for a class of problems where the mapping from the agent actions to system reward functions can be decomposed into a linear combination of nonlinear functions of the agents' actions. We test our method's performance on a distributed marketing problem and an air traffic flow management problem and show a 44% performance improvement over team games and a speedup of order n for difference rewards (for an n agent system). <\/jats:p>","DOI":"10.1142\/s0219525909002295","type":"journal-article","created":{"date-parts":[[2009,9,25]],"date-time":"2009-09-25T17:44:46Z","timestamp":1253900686000},"page":"475-492","source":"Crossref","is-referenced-by-count":5,"title":["MULTIAGENT LEARNING FOR BLACK BOX SYSTEM REWARD FUNCTIONS"],"prefix":"10.1142","volume":"12","author":[{"given":"KAGAN","family":"TUMER","sequence":"first","affiliation":[{"name":"Oregon State University, 204 Rogers Hall, Corvallis, Oregon 97331, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"ADRIAN","family":"AGOGINO","sequence":"additional","affiliation":[{"name":"UCSC, NASA Ames Research Center, Mailstop 269-3, Moffett Field, California 94035, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"219","published-online":{"date-parts":[[2011,11,20]]},"reference":[{"key":"rf4","doi-asserted-by":"publisher","DOI":"10.1007\/s10458-006-6105-y"},{"key":"rf5","doi-asserted-by":"publisher","DOI":"10.1162\/evco.2008.16.2.257"},{"key":"rf6","volume":"18","author":"Bagnell J. A.","journal-title":"News Physiol. Sci."},{"key":"rf7","doi-asserted-by":"publisher","DOI":"10.2514\/1.15242"},{"key":"rf8","volume":"9","author":"Bilimoria K. D.","journal-title":"Air Traffic Cont. Q."},{"key":"rf9","doi-asserted-by":"publisher","DOI":"10.1016\/S0378-4371(98)00260-X"},{"key":"rf10","unstructured":"R. H.\u00a0Crites and A. G.\u00a0Barto, Advances in Neural Information Processing Systems\u00a08, eds. D. S.\u00a0Touretzky, M. C.\u00a0Mozer and M. E.\u00a0Hasselmo (MIT Press, 1996)\u00a0pp. 1017\u20131023."},{"key":"rf12","doi-asserted-by":"crossref","first-page":"227","DOI":"10.1613\/jair.639","volume":"13","author":"Dietterich T. G.","journal-title":"J. Artif. Intell."},{"key":"rf16","volume":"65","author":"Jefferies P.","journal-title":"Phys. Rev. E"},{"key":"rf17","first-page":"277","volume":"177","author":"Jennings N. R.","journal-title":"Artif. Intell."},{"key":"rf20","doi-asserted-by":"publisher","DOI":"10.1613\/jair.301"},{"key":"rf22","first-page":"58","volume":"1","author":"McGlohon M.","journal-title":"Int. J. Lateral Comput."},{"key":"rf23","doi-asserted-by":"publisher","DOI":"10.2514\/2.4384"},{"key":"rf25","first-page":"791","volume":"16","author":"Parkes D.","journal-title":"News Physiol. Sci."},{"key":"rf27","author":"Stone P.","journal-title":"Adapt. Behav."},{"key":"rf29","doi-asserted-by":"publisher","DOI":"10.1109\/9.664154"},{"key":"rf33","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4419-8909-3"},{"key":"rf35","volume":"171","author":"Tuyls K.","journal-title":"Artif. Intell."},{"key":"rf37","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-32274-0_18"},{"key":"rf40","volume":"15","author":"Whiteson S.","journal-title":"Adapt. Behav."},{"key":"rf41","doi-asserted-by":"publisher","DOI":"10.1142\/S0219525901000188"}],"container-title":["Advances in Complex Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.worldscientific.com\/doi\/pdf\/10.1142\/S0219525909002295","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,8,6]],"date-time":"2019-08-06T20:26:32Z","timestamp":1565123192000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.worldscientific.com\/doi\/abs\/10.1142\/S0219525909002295"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,8]]},"references-count":21,"journal-issue":{"issue":"04n05","published-online":{"date-parts":[[2011,11,20]]},"published-print":{"date-parts":[[2009,8]]}},"alternative-id":["10.1142\/S0219525909002295"],"URL":"https:\/\/doi.org\/10.1142\/s0219525909002295","relation":{},"ISSN":["0219-5259","1793-6802"],"issn-type":[{"value":"0219-5259","type":"print"},{"value":"1793-6802","type":"electronic"}],"subject":[],"published":{"date-parts":[[2009,8]]}}}