{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,28]],"date-time":"2026-01-28T14:56:17Z","timestamp":1769612177263,"version":"3.49.0"},"reference-count":33,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2026,1,21]],"date-time":"2026-01-21T00:00:00Z","timestamp":1768953600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["2144646"],"award-info":[{"award-number":["2144646"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000183","name":"Army Research Office through the Cooperative Agreement","doi-asserted-by":"publisher","award":["W911NF-24-2-0133"],"award-info":[{"award-number":["W911NF-24-2-0133"]}],"id":[{"id":"10.13039\/100000183","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Robotics"],"abstract":"<jats:p>Value factorization methods have become a standard tool for cooperative multi-agent reinforcement learning (MARL) in the centralized-training, decentralized-execution (CTDE) setting. QMIX (a monotonic mixing network for value factorization), in particular, constrains the joint action\u2013value function to be a monotonic mixing of per-agent utilities, which guarantees consistency with individual greedy policies but can severely limit expressiveness on tasks with non-monotonic agent interactions. This work revisits this design choice and proposes Relaxed Monotonic QMIX (R-QMIX), a simple regularized variant of QMIX that encourages but does not strictly enforce the monotonicity constraint. R-QMIX removes the sign constraints on the mixing network weights and introduces a differentiable penalty on negative partial derivatives of the joint value with respect to each agent\u2019s utility. This preserves the computational benefits of value factorization while allowing the joint value to deviate from strict monotonicity when beneficial. R-QMIX is implemented in a standard PyMARL (an open-source MARL codebase) and evaluated on the StarCraft Multi-Agent Challenge (SMAC). On a simple map (3m), R-QMIX matches the asymptotic performance of QMIX while learning substantially faster. On more challenging maps (MMM2, 6h vs. 8z, and 27m vs. 30m), R-QMIX significantly improves both sample efficiency and final win rate (WR), for example increasing the final-quarter mean win rate from 42.3% to 97.1% on MMM2, from 0.0% to 57.5% on 6h vs. 8z, and from 58.0% to 96.6% on 27m vs. 30m. These results suggest that soft monotonicity regularization is a practical way to bridge the gap between strictly monotonic value factorization and fully unconstrained joint value functions. A further comparison against QTRAN (Q-value transformation), a more expressive value factorization method, shows that R-QMIX achieves higher and more reliably convergent win rates on the challenging SMAC maps considered.<\/jats:p>","DOI":"10.3390\/robotics15010028","type":"journal-article","created":{"date-parts":[[2026,1,21]],"date-time":"2026-01-21T15:31:46Z","timestamp":1769009506000},"page":"28","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Relaxed Monotonic QMIX (R-QMIX): A Regularized Value Factorization Approach to Decentralized Multi-Agent Reinforcement Learning"],"prefix":"10.3390","volume":"15","author":[{"given":"Liam","family":"O\u2019Brien","sequence":"first","affiliation":[{"name":"Department of Electrical and Biomedical Engineering, University of Nevada, Reno, NV 89557, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hao","family":"Xu","sequence":"additional","affiliation":[{"name":"Department of Electrical and Biomedical Engineering, University of Nevada, Reno, NV 89557, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,1,21]]},"reference":[{"key":"ref_1","unstructured":"Sutton, R.S., and Barto, A.G. (2018). Reinforcement Learning: An Introduction, MIT Press. [2nd ed.]."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"ref_3","unstructured":"Lillicrap, T.P., Hunt, J.J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2016, January 2\u20134). Continuous control with deep reinforcement learning. Proceedings of the International Conference on Learning Representations (ICLR), San Juan, Puerto Rico."},{"key":"ref_4","unstructured":"Mnih, V., Badia, A.P., Mirza, M., Graves, A., Lillicrap, T.P., Harley, T., Silver, D., and Kavukcuoglu, K. (2016, January 19\u201324). Asynchronous methods for deep reinforcement learning. Proceedings of the 33rd International Conference on Machine Learning (ICML), New York, NY, USA."},{"key":"ref_5","unstructured":"Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015, January 6\u201311). Trust region policy optimization. Proceedings of the 32nd International Conference on Machine Learning (ICML), Lille, France."},{"key":"ref_6","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017). Proximal policy optimization algorithms. arXiv."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Oliehoek, F.A., and Amato, C. (2016). A Concise Introduction to Decentralized POMDPs, Springer.","DOI":"10.1007\/978-3-319-28929-8"},{"key":"ref_8","unstructured":"Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, P., and Mordatch, I. (2017, January 4\u20139). Multi-agent actor-critic for mixed cooperative-competitive environments. Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS), Long Beach, CA, USA."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W.M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J.Z., and Tuyls, K. (2018, January 10\u201315). Value-decomposition networks for cooperative multi-agent learning. Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems (AAMAS), Stockholm, Sweden.","DOI":"10.65109\/JSRC7365"},{"key":"ref_10","unstructured":"Rashid, T., Samvelyan, M., de Witt, C.S., Farquhar, G., Foerster, J., and Whiteson, S. (2018, January 10\u201315). QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning. Proceedings of the 35th International Conference on Machine Learning (ICML), Stockholm, Sweden."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Samvelyan, M., Rashid, T., de Witt, C.S., Farquhar, G., Nardelli, N., Rudner, T.G., Hung, C., Torr, P.H.S., Foerster, J., and Whiteson, S. (2019, January 13\u201317). The StarCraft Multi-Agent Challenge. Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems (AAMAS), Montreal, QC, Canada.","DOI":"10.65109\/LVZZ5205"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"387","DOI":"10.1007\/s10458-005-2631-2","article-title":"Cooperative multi-agent learning: The state of the art","volume":"11","author":"Panait","year":"2005","journal-title":"Auton. Agents Multi-Agent Syst."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"156","DOI":"10.1109\/TSMCC.2007.913919","article-title":"A comprehensive survey of multiagent reinforcement learning","volume":"38","author":"Busoniu","year":"2008","journal-title":"IEEE Trans. Syst. Man Cybern. Part C"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"750","DOI":"10.1007\/s10458-019-09421-1","article-title":"A survey and critique of multiagent deep reinforcement learning","volume":"33","author":"Kartal","year":"2019","journal-title":"Auton. Agents Multi-Agent Syst."},{"key":"ref_15","unstructured":"Oroojlooyjadid, A., and Hajinezhad, D. (2019). A review of cooperative multi-agent deep reinforcement learning. arXiv."},{"key":"ref_16","unstructured":"Papoudakis, G., Christianos, F., Sch\u00e4fer, L., and Albrecht, S.V. (2019). Dealing with non-stationarity in multi-agent deep reinforcement learning. arXiv."},{"key":"ref_17","unstructured":"Bellman, R. (1957). Dynamic Programming, Princeton University Press."},{"key":"ref_18","first-page":"279","article-title":"Q-learning","volume":"8","author":"Watkins","year":"1992","journal-title":"Mach. Learn."},{"key":"ref_19","unstructured":"Hausknecht, M., and Stone, P. (2015, January 12\u201314). Deep recurrent Q-learning for partially observable MDPs. Proceedings of the AAAI Fall Symposium on Sequential Decision Making for Intelligent Agents, Arlington, VA, USA."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Cho, K., Merri\u00ebnboer, B.V., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014). Learning Phrase Representations using RNN Encoder\u2013Decoder for Statistical Machine Translation. arXiv.","DOI":"10.3115\/v1\/D14-1179"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long Short-Term Memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Tampuu, A., Matiisen, T., Kodelja, D., Kuzovkin, I., Korjus, K., Aru, J., Aru, J., and Vicente, R. (2017). Multiagent cooperation and competition with deep reinforcement learning. PLoS ONE, 12.","DOI":"10.1371\/journal.pone.0172395"},{"key":"ref_23","unstructured":"Tan, M. (1993, January 27\u201329). Multi-agent reinforcement learning: Independent vs. cooperative agents. Proceedings of the 10th International Conference on Machine Learning (ICML), Amherst, MA, USA."},{"key":"ref_24","unstructured":"Rashid, T., de Witt, C.S., Farquhar, G., and Whiteson, S. (2020). Weighted QMIX: Expanding monotonic value function factorisation for deep multi-agent reinforcement learning. arXiv."},{"key":"ref_25","unstructured":"Wang, T., Han, B., Xu, H., Wang, X., Dong, H., and Zhang, C. (2021, January 3\u20137). QPLEX: Duplex dueling multi-agent Q-learning. Proceedings of the International Conference on Learning Representations (ICLR), Virtual Event."},{"key":"ref_26","unstructured":"Yang, Y., Meng, Z., Hao, J., Zhang, H., and Wang, Z. (2020). Qatten: A general framework for cooperative multiagent reinforcement learning. arXiv."},{"key":"ref_27","unstructured":"B\u00f6hmer, W., Rashid, T., Thoma, J., Oliehoek, F.A., and Whiteson, S. (2020, January 12\u201318). Deep coordination graphs. Proceedings of the 37th International Conference on Machine Learning (ICML), Virtual Event."},{"key":"ref_28","unstructured":"Son, K., Kim, D., Kang, W., Hostallero, D.E., and Yi, Y. (2019, January 9\u201315). QTRAN: Learning to factorize with transformation for cooperative multi-agent reinforcement learning. Proceedings of the 36th International Conference on Machine Learning (ICML), Long Beach, CA, USA."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Foerster, J.N., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S. (2018, January 2\u20137). Counterfactual multi-agent policy gradients. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11794"},{"key":"ref_30","unstructured":"Mahajan, A., Samvelyan, M., Rashid, T., de Witt, C.S., and Whiteson, S. (2019, January 8\u201314). MAVEN: Multi-agent variational exploration. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada."},{"key":"ref_31","unstructured":"Peng, P., Wen, Y., Yang, Y., Yuan, Q., Tang, Z., Long, H., and Wang, J. (2017). Multiagent bidirectionally-coordinated nets: Emergence of human-level coordination in learning to play StarCraft combat games. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1007\/s10458-023-09633-6","article-title":"A survey of multi-agent deep reinforcement learning with communication","volume":"38","author":"Zhu","year":"2024","journal-title":"Auton. Agents Multi-Agent Syst."},{"key":"ref_33","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A Method for Stochastic Optimization. arXiv."}],"container-title":["Robotics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2218-6581\/15\/1\/28\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,27]],"date-time":"2026-01-27T09:38:21Z","timestamp":1769506701000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2218-6581\/15\/1\/28"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,21]]},"references-count":33,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,1]]}},"alternative-id":["robotics15010028"],"URL":"https:\/\/doi.org\/10.3390\/robotics15010028","relation":{},"ISSN":["2218-6581"],"issn-type":[{"value":"2218-6581","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,21]]}}}