{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,16]],"date-time":"2026-05-16T15:55:15Z","timestamp":1778946915028,"version":"3.51.4"},"reference-count":44,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2024,10,18]],"date-time":"2024-10-18T00:00:00Z","timestamp":1729209600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,10,18]],"date-time":"2024-10-18T00:00:00Z","timestamp":1729209600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Effective Giving"},{"DOI":"10.13039\/100006919","name":"Massachusetts Institute of Technology","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100006919","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Auton Agent Multi-Agent Syst"],"published-print":{"date-parts":[[2024,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Multi-agent Reinforcement Learning (MARL) is a powerful tool for training autonomous agents acting independently in a common environment. However, it can lead to sub-optimal behavior when individual incentives and group incentives diverge. Humans are remarkably capable at solving these social dilemmas. It is an open problem in MARL to replicate such cooperative behaviors in selfish agents. In this work, we draw upon the idea of formal contracting from economics to overcome diverging incentives between agents in MARL. We propose an augmentation to a Markov game where agents voluntarily agree to binding transfers of reward, under pre-specified conditions. Our contributions are theoretical and empirical. First, we show that this augmentation makes all subgame-perfect equilibria of all Fully Observable Markov Games exhibit socially optimal behavior, given a sufficiently rich space of contracts. Next, we show that for general contract spaces, and even under partial observability, richer contract spaces lead to higher welfare. Hence, contract space design solves an exploration-exploitation tradeoff, sidestepping incentive issues. We complement our theoretical analysis with experiments. Issues of exploration in the contracting augmentation are mitigated using a training methodology inspired by multi-objective reinforcement learning: Multi-Objective Contract Augmentation Learning. We test our methodology in static, single-move games, as well as dynamic domains that simulate traffic, pollution management, and common pool resource management.<\/jats:p>","DOI":"10.1007\/s10458-024-09682-5","type":"journal-article","created":{"date-parts":[[2024,10,18]],"date-time":"2024-10-18T02:02:07Z","timestamp":1729216927000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Formal contracts mitigate social dilemmas in multi-agent reinforcement learning"],"prefix":"10.1007","volume":"38","author":[{"given":"Andreas","family":"Haupt","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Phillip","family":"Christoffersen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mehul","family":"Damani","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dylan","family":"Hadfield-Menell","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,10,18]]},"reference":[{"key":"9682_CR1","unstructured":"Andrychowicz, M., Wolski, F., Ray, A., et\u00a0al. (2017). Hindsight experience replay. Advances in neural information processing systems. 30"},{"issue":"4","key":"9682_CR2","doi-asserted-by":"publisher","first-page":"819","DOI":"10.1287\/moor.27.4.819.297","volume":"27","author":"DS Bernstein","year":"2002","unstructured":"Bernstein, D. S., Givan, R., Immerman, N., et al. (2002). The complexity of decentralized control of markov decision processes. Mathematics of Operations Research, 27(4), 819\u2013840.","journal-title":"Mathematics of Operations Research"},{"key":"9682_CR3","doi-asserted-by":"crossref","unstructured":"Curry, M., Sandholm, T., & Dickerson, J. (2022). Differentiable economics for randomized affine maximizer auctions. arXiv preprint arXiv:2202.02872","DOI":"10.24963\/ijcai.2023\/293"},{"issue":"1","key":"9682_CR4","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/s10458-019-09424-y","volume":"34","author":"D De Jonge","year":"2020","unstructured":"De Jonge, D., & Zhang, D. (2020). Strategic negotiations for extensive-form games. Autonomous Agents and Multi-Agent Systems, 34(1), 1\u201341.","journal-title":"Autonomous Agents and Multi-Agent Systems"},{"key":"9682_CR5","doi-asserted-by":"crossref","unstructured":"Devin, C., Gupta, A., Darrell, T., et\u00a0al. (2017). Learning modular neural network policies for multi-task and multi-robot transfer. In: 2017 IEEE international conference on robotics and automation (ICRA), IEEE, pp 2169\u20132176.","DOI":"10.1109\/ICRA.2017.7989250"},{"key":"9682_CR6","doi-asserted-by":"publisher","DOI":"10.1145\/3630749","author":"P D\u00fctting","year":"2023","unstructured":"D\u00fctting, P., Feng, Z., Narasimhan, H., et al. (2023). Optimal auctions through deep learning: Advances in differentiable economics. Journal of the ACM. https:\/\/doi.org\/10.1145\/3630749","journal-title":"Journal of the ACM"},{"issue":"3","key":"9682_CR7","doi-asserted-by":"publisher","first-page":"533","DOI":"10.2307\/1911307","volume":"54","author":"D Fudenberg","year":"1986","unstructured":"Fudenberg, D., & Maskin, E. (1986). The folk theorem in repeated games with discounting or with incomplete information. Econometrica, 54(3), 533\u2013554.","journal-title":"Econometrica"},{"issue":"2","key":"9682_CR8","doi-asserted-by":"publisher","first-page":"200","DOI":"10.1016\/j.jebo.2004.09.010","volume":"58","author":"R Gibbons","year":"2005","unstructured":"Gibbons, R. (2005). Four formal (izable) theories of the firm? Journal of Economic Behavior & Organization, 58(2), 200\u2013245.","journal-title":"Journal of Economic Behavior & Organization"},{"key":"9682_CR9","doi-asserted-by":"publisher","DOI":"10.2307\/j.ctvcmxrzd","volume-title":"Game theory for applied economists","author":"RS Gibbons","year":"1992","unstructured":"Gibbons, R. S. (1992). Game theory for applied economists. Princeton University Press."},{"key":"9682_CR10","doi-asserted-by":"publisher","first-page":"141","DOI":"10.1146\/annurev-lawsocsci-110316-113413","volume":"13","author":"R Gil","year":"2017","unstructured":"Gil, R., & Zanarone, G. (2017). Formal and informal contracting: Theory and evidence. Annual Review of Law and Social Science, 13, 141\u2013159.","journal-title":"Annual Review of Law and Social Science"},{"issue":"3","key":"9682_CR11","doi-asserted-by":"publisher","first-page":"561","DOI":"10.1007\/s10458-016-9338-4","volume":"31","author":"TA Han","year":"2017","unstructured":"Han, T. A., Pereira, L. M., & Lenaerts, T. (2017). Evolution of commitment and level of participation in public goods games. Autonomous Agents and Multi-Agent Systems, 31(3), 561\u2013583. https:\/\/doi.org\/10.1007\/s10458-016-9338-4","journal-title":"Autonomous Agents and Multi-Agent Systems"},{"key":"9682_CR12","doi-asserted-by":"crossref","unstructured":"Hardt, M., Megiddo, N., Papadimitriou, C., et\u00a0al. (2016). Strategic classification. In: Proceedings of the 2016 ACM conference on innovations in theoretical computer science, pp 111\u2013122.","DOI":"10.1145\/2840728.2840730"},{"key":"9682_CR13","doi-asserted-by":"publisher","first-page":"74","DOI":"10.2307\/3003320","volume":"13","author":"B Holmstr\u00f6m","year":"1979","unstructured":"Holmstr\u00f6m, B. (1979). Moral hazard and observability. The Bell Journal of Economics, 13, 74\u201391.","journal-title":"The Bell Journal of Economics"},{"key":"9682_CR14","unstructured":"Hughes, E., Leibo, J.Z., Phillips, M.G., et\u00a0al. (2018). Inequity aversion improves cooperation in intertemporal social dilemmas. arXiv preprint arXiv:1803.08884."},{"key":"9682_CR15","unstructured":"Hughes, E., Anthony, T.W., Eccles, T., et\u00a0al. (2020). Learning to resolve alliance dilemmas in many-player zero-sum games. In: Proceedings of the 19th International Conference on Autonomous Agents and MultiAgent Systems, pp 538\u2013547."},{"issue":"2","key":"9682_CR16","first-page":"1","volume":"63","author":"L Hurwicz","year":"1973","unstructured":"Hurwicz, L. (1973). The design of mechanisms for resource allocation. The American Economic Review, 63(2), 1\u201330.","journal-title":"The American Economic Review"},{"key":"9682_CR17","unstructured":"Janssen, M., & Ahn, T. (2003). Adaptation versus anticipation in public-good games. In: Annual meeting of the American Political Science Association, Philadelphia, PA."},{"key":"9682_CR18","unstructured":"K\u00f6ster, R., McKee, K.R., Everett, R., et\u00a0al. (2020). Model-free conventions in multi-agent reinforcement learning with heterogeneous preferences. arXiv preprint arXiv:2010.09054."},{"key":"9682_CR19","doi-asserted-by":"crossref","unstructured":"K\u00f6ster, R., Hadfield-Menell, D., Everett, R., et\u00a0al. (2022). Spurious normativity enhances learning of compliance and enforcement behavior in artificial agents. In Proceedings of the National Academy of Sciences 119(3)","DOI":"10.1073\/pnas.2106028118"},{"issue":"1","key":"9682_CR20","doi-asserted-by":"publisher","first-page":"7214","DOI":"10.1038\/s41467-022-34473-5","volume":"13","author":"J Kram\u00e1r","year":"2022","unstructured":"Kram\u00e1r, J., Eccles, T., Gemp, I., et al. (2022). Negotiation and honesty in artificial intelligence methods for the board game of diplomacy. Nature Communications, 13(1), 7214.","journal-title":"Nature Communications"},{"key":"9682_CR21","unstructured":"Leibo, J.Z., Zambaldi, V., Lanctot, M., et\u00a0al. (2017). Multi-agent reinforcement learning in sequential social dilemmas. arXiv preprint arXiv:1702.03037."},{"key":"9682_CR22","unstructured":"Liang, E., Liaw, R., Nishihara, R., et\u00a0al. (2018). Rllib: Abstractions for distributed reinforcement learning. In: International Conference on Machine Learning, PMLR, pp 3053\u20133062."},{"key":"9682_CR23","unstructured":"Lupu, A., & Precup, D. (2020). Gifting in multi-agent reinforcement learning. In: Proceedings of the 19th International Conference on Autonomous Agents and MultiAgent Systems, pp 789\u2013797"},{"key":"9682_CR24","unstructured":"McAleer, S., Lanier, J., Dennis, M., et\u00a0al. (2021). Improving social welfare while preserving autonomy via a pareto mediator. arXiv preprint arXiv:2106.03927."},{"key":"9682_CR25","doi-asserted-by":"crossref","unstructured":"Milli, S., Miller, J., Dragan, A.D., et\u00a0al. (2019). The social cost of strategic classification. In: Proceedings of the Conference on Fairness, Accountability, and Transparency, pp 230\u2013239","DOI":"10.1145\/3287560.3287576"},{"issue":"1","key":"9682_CR26","doi-asserted-by":"publisher","first-page":"37","DOI":"10.1613\/jair.1231","volume":"21","author":"D Monderer","year":"2004","unstructured":"Monderer, D., & Tennenholtz, M. (2004). K-implementation. Journal of Artificial Intelligence Research, 21(1), 37\u201362.","journal-title":"Journal of Artificial Intelligence Research"},{"issue":"2","key":"9682_CR27","doi-asserted-by":"publisher","first-page":"301","DOI":"10.1016\/0022-0531(78)90085-6","volume":"18","author":"M Mussa","year":"1978","unstructured":"Mussa, M., & Rosen, S. (1978). Monopoly and product quality. Journal of Economic theory, 18(2), 301\u2013317.","journal-title":"Journal of Economic theory"},{"issue":"1","key":"9682_CR28","doi-asserted-by":"publisher","first-page":"58","DOI":"10.1287\/moor.6.1.58","volume":"6","author":"RB Myerson","year":"1981","unstructured":"Myerson, R. B. (1981). Optimal auction design. Mathematics of Operations Research, 6(1), 58\u201373.","journal-title":"Mathematics of Operations Research"},{"key":"9682_CR29","volume-title":"A course in game theory","author":"MJ Osborne","year":"1994","unstructured":"Osborne, M. J., & Rubinstein, A. (1994). A course in game theory. MIT press."},{"key":"9682_CR30","unstructured":"Parisotto, E., Ba, J.L., & Salakhutdinov, R. (2015). Actor-mimic: Deep multitask and transfer reinforcement learning. arXiv preprint arXiv:1511.06342."},{"key":"9682_CR31","unstructured":"Radke, D., Larson, K., & Brecht, T. (2022). The importance of credo in multiagent learning. arXiv preprint arXiv:2204.07471."},{"key":"9682_CR32","unstructured":"Rusu, A.A., Colmenarejo, S.G., & Gulcehre, C., et\u00a0al. (2015). Policy distillation. arXiv preprint arXiv:1511.06295."},{"key":"9682_CR33","unstructured":"Samvelyan, M., Rashid, T., De\u00a0Witt, C.S., et\u00a0al. (2019). The starcraft multi-agent challenge. arXiv preprint arXiv:1902.04043"},{"key":"9682_CR34","unstructured":"Sandholm, T.W., & Lesser, V.R. (1996). Advantages of a leveled commitment contracting protocol. In: AAAI\/IAAI, Vol. 1, Citeseer, pp 126\u2013133"},{"key":"9682_CR35","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., et\u00a0al. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347"},{"issue":"12","key":"9682_CR36","doi-asserted-by":"publisher","first-page":"1104","DOI":"10.1109\/TC.1980.1675516","volume":"29","author":"RG Smith","year":"1980","unstructured":"Smith, R. G. (1980). The contract net protocol: High-level communication and control in a distributed problem solver. IEEE Transactions on Computers, 29(12), 1104\u20131113.","journal-title":"IEEE Transactions on Computers"},{"key":"9682_CR37","unstructured":"Sodomka, E., Hilliard, E., Littman, M., et\u00a0al. (2013). Coco-q: Learning in stochastic games with side payments. In: International Conference on Machine Learning, PMLR, pp 1471\u20131479."},{"key":"9682_CR38","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1707.04175","author":"Y Teh","year":"2017","unstructured":"Teh, Y., Bapst, V., Czarnecki, W. M., et al. (2017). Distral: Robust multitask reinforcement learning. Advances in Neural Information Processing Systems. https:\/\/doi.org\/10.48550\/arXiv.1707.04175","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"3","key":"9682_CR39","doi-asserted-by":"publisher","first-page":"228","DOI":"10.2307\/3027092","volume":"14","author":"AW Tucker","year":"1983","unstructured":"Tucker, A. W., & Straffin, P. D., Jr. (1983). The mathematics of tucker: A sampler. The Two-Year College Mathematics Journal, 14(3), 228\u2013232.","journal-title":"The Two-Year College Mathematics Journal"},{"issue":"1","key":"9682_CR40","doi-asserted-by":"publisher","first-page":"8","DOI":"10.1111\/j.1540-6261.1961.tb02789.x","volume":"16","author":"W Vickrey","year":"1961","unstructured":"Vickrey, W. (1961). Counterspeculation, auctions, and competitive sealed tenders. The Journal of Finance, 16(1), 8\u201337.","journal-title":"The Journal of Finance"},{"key":"9682_CR41","unstructured":"Vinitsky, E., K\u00f6ster, R., Agapiou, J.P., et\u00a0al. (2021). A learning agent that acquires social norms from public sanctions in decentralized multi-agent settings. arXiv preprint arXiv:2106.09012."},{"key":"9682_CR42","doi-asserted-by":"crossref","unstructured":"Wang, W.Z., Beliaev, M., B\u0131y\u0131k, E., et\u00a0al. (2021). Emergent prosociality in multi-agent games through gifting. arXiv preprint arXiv:2105.06593.","DOI":"10.24963\/ijcai.2021\/61"},{"key":"9682_CR43","first-page":"15208","volume":"33","author":"J Yang","year":"2020","unstructured":"Yang, J., Li, A., Farajtabar, M., et al. (2020). Learning to incentivize other learning agents. Advances in Neural Information Processing Systems, 33, 15208\u201315219.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"9682_CR44","first-page":"15257","volume":"34","author":"T Zrnic","year":"2021","unstructured":"Zrnic, T., Mazumdar, E., Sastry, S., et al. (2021). Who leads and who follows in strategic classification? Advances in Neural Information Processing Systems, 34, 15257\u201315269.","journal-title":"Advances in Neural Information Processing Systems"}],"container-title":["Autonomous Agents and Multi-Agent Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10458-024-09682-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10458-024-09682-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10458-024-09682-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,13]],"date-time":"2024-11-13T15:25:38Z","timestamp":1731511538000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10458-024-09682-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,18]]},"references-count":44,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,12]]}},"alternative-id":["9682"],"URL":"https:\/\/doi.org\/10.1007\/s10458-024-09682-5","relation":{},"ISSN":["1387-2532","1573-7454"],"issn-type":[{"value":"1387-2532","type":"print"},{"value":"1573-7454","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,10,18]]},"assertion":[{"value":"3 October 2024","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"18 October 2024","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"51"}}