{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T10:26:49Z","timestamp":1775039209409,"version":"3.50.1"},"reference-count":34,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T00:00:00Z","timestamp":1775001600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>In cooperative multi-agent reinforcement learning (MARL), communication can address the challenges of partial observability and environmental non-stationarity by conveying environmental features and decision intents, respectively. However, existing methods either focus on only one type of information\u2014failing to tackle both challenges simultaneously\u2014or conflate these signals, causing agents to confuse environmental context with decision intents. This paper introduces Disentangling Environment and Decision messages for Multi-Agent Communication (DEDMAC), a framework that explicitly separates these two information types into two distinct message streams and processes them independently. Specifically, environment messages are integrated into long-term memory to resolve partial observability, while decision messages provide instantaneous intent signals to mitigate non-stationarity and facilitate coordination. To prevent semantic confusion between the two message streams, we employ mutual information constraints to ensure semantic disentanglement. Furthermore, we design a mechanism that leverages global information to correct intent biases in decision messages resulting from limited local perspectives during generation. Evaluations across complex multi-agent benchmarks demonstrate that DEDMAC significantly outperforms state-of-the-art communication-based methods. These findings indicate that the explicit separation and specialized processing of environment and decision semantics are critical for achieving optimal performance in dynamic, collaborative multi-agent systems.<\/jats:p>","DOI":"10.3390\/info17040332","type":"journal-article","created":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T08:08:34Z","timestamp":1775030914000},"page":"332","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["DEDMAC: Disentangling Environment and Decision Messages for Multi-Agent Communication"],"prefix":"10.3390","volume":"17","author":[{"given":"Yihan","family":"Liang","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, University of Science and Technology of China, Hefei 230026, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jinlong","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, University of Science and Technology of China, Hefei 230026, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,4,1]]},"reference":[{"key":"ref_1","unstructured":"H\u00fcttenrauch, M., \u0160o\u0161i\u0107, A., and Neumann, G. (2017). Guided deep reinforcement learning for swarm systems. arXiv."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"427","DOI":"10.1109\/TII.2012.2219061","article-title":"An overview of recent progress in the study of distributed multi-agent coordination","volume":"9","author":"Cao","year":"2012","journal-title":"IEEE Trans. Ind. Inform."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Du, X., Wang, J., Chen, S., and Liu, Z. (2021). Multi-agent deep reinforcement learning with spatio-temporal feature fusion for traffic signal control. Proceedings of the Machine Learning and Knowledge Discovery in Databases, Applied Data Science Track: European Conference, ECML PKDD 2021, Bilbao, Spain, 13\u201317 September 2021, Springer. Proceedings, Part IV 21.","DOI":"10.1007\/978-3-030-86514-6_29"},{"key":"ref_4","first-page":"20147","article-title":"Multi-agent dynamic algorithm configuration","volume":"Volume 35","author":"Xue","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"ref_5","first-page":"4026","article-title":"Deep exploration via bootstrapped DQN","volume":"Volume 29","author":"Osband","year":"2016","journal-title":"Advances in Neural Information Processing Systems"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Mao, W., Zhang, K., Miehling, E., and Ba\u015far, T. (2020). Information state embedding in partially observable cooperative multi-agent reinforcement learning. Proceedings of the 2020 59th IEEE Conference on Decision and Control (CDC), IEEE.","DOI":"10.1109\/CDC42340.2020.9303801"},{"key":"ref_7","unstructured":"Papoudakis, G., Christianos, F., Rahman, A., and Albrecht, S.V. (2019). Dealing with non-stationarity in multi-agent deep reinforcement learning. arXiv."},{"key":"ref_8","unstructured":"Lowe, R., Wu, Y.I., Tamar, A., Harb, J., Pieter Abbeel, O., and Mordatch, I. (2017). Multi-agent actor-critic for mixed cooperative-competitive environments. Advances in Neural Information Processing Systems, ACM Digital Library."},{"key":"ref_9","first-page":"7234","article-title":"Monotonic value function factorisation for deep multi-agent reinforcement learning","volume":"21","author":"Rashid","year":"2020","journal-title":"J. Mach. Learn. Res."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"66","DOI":"10.1016\/j.artint.2018.01.002","article-title":"Autonomous agents modelling other agents: A comprehensive survey and open problems","volume":"258","author":"Albrecht","year":"2018","journal-title":"Artif. Intell."},{"key":"ref_11","unstructured":"He, H., Boyd-Graber, J., Kwok, K., and Daum\u00e9, H. (2016). Opponent modeling in deep reinforcement learning. Proceedings of the International Conference on Machine Learning, PMLR, New York, NY, USA, 20\u201322 June 2016, Association for Computing Machinery."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Zhang, K., Yang, Z., and Ba\u015far, T. (2021). Multi-agent reinforcement learning: A selective overview of theories and algorithms. Handbook of Reinforcement Learning and Control, Springer.","DOI":"10.1007\/978-3-030-60990-0_12"},{"key":"ref_13","first-page":"2137","article-title":"Learning to communicate with deep multi-agent reinforcement learning","volume":"Volume 29","author":"Foerster","year":"2016","journal-title":"Advances in Neural Information Processing Systems"},{"key":"ref_14","first-page":"2244","article-title":"Learning multiagent communication with backpropagation","volume":"Volume 29","author":"Sukhbaatar","year":"2016","journal-title":"Advances in Neural Information Processing Systems"},{"key":"ref_15","first-page":"1020","article-title":"Efficient multi-agent communication via self-supervised information aggregation","volume":"Volume 35","author":"Guan","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"9466","DOI":"10.1609\/aaai.v36i9.21179","article-title":"Multi-agent incentive communication via decentralized teammate modeling","volume":"Volume 36","author":"Yuan","year":"2022","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"ref_17","unstructured":"Das, A., Gervet, T., Romoff, J., Batra, D., Parikh, D., Rabbat, M., and Pineau, J. (2019). Tarmac: Targeted multi-agent communication. Proceedings of the International Conference on Machine Learning, PMLR, Long Beach, CA, USA, 9\u201315 June 2019, Association for Computing Machinery."},{"key":"ref_18","unstructured":"Chung, J., Gulcehre, C., Cho, K., and Bengio, Y. (2014). Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv."},{"key":"ref_19","unstructured":"Cheng, P., Hao, W., Dai, S., Liu, J., Gan, Z., and Carin, L. (2020). Club: A contrastive log-ratio upper bound of mutual information. Proceedings of the International Conference on Machine Learning, PMLR, Virtual, 13\u201318 July 2020, Association for Computing Machinery."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Oliehoek, F.A., and Amato, C. (2016). A Concise Introduction to Decentralized POMDPs, Springer.","DOI":"10.1007\/978-3-319-28929-8"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Sunehag, P., Lever, G., Gruslys, A., Czarnecki, W.M., Zambaldi, V., Jaderberg, M., Lanctot, M., Sonnerat, N., Leibo, J.Z., and Tuyls, K. (2017). Value-decomposition networks for cooperative multi-agent learning. arXiv.","DOI":"10.65109\/JSRC7365"},{"key":"ref_23","unstructured":"Wang, J., Ren, Z., Liu, T., Yu, Y., and Zhang, C. (2020). Qplex: Duplex dueling multi-agent q-learning. arXiv."},{"key":"ref_24","unstructured":"Bahdanau, D., Cho, K., and Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv."},{"key":"ref_25","unstructured":"Ye, T., Dong, L., Xia, Y., Sun, Y., Zhu, Y., Huang, G., and Wei, F. (2024). Differential transformer. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1561\/2200000001","article-title":"Graphical models, exponential families, and variational inference","volume":"Volume 1","author":"Wainwright","year":"2008","journal-title":"Foundations and Trends\u00ae in Machine Learning"},{"key":"ref_27","unstructured":"Alemi, A.A., Fischer, I., Dillon, J.V., and Murphy, K. (2016). Deep variational information bottleneck. arXiv."},{"key":"ref_28","first-page":"15154","article-title":"T2mac: Targeted and trusted multi-agent communication through selective engagement and evidence-driven integration","volume":"Volume 38","author":"Sun","year":"2024","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 20\u201327 February 2024"},{"key":"ref_29","unstructured":"Wang, T., Wang, J., Zheng, C., and Zhang, C. (2019). Learning nearly decomposable value functions via communication minimization. arXiv."},{"key":"ref_30","unstructured":"Hu, J., Jiang, S., Harding, S.A., Wu, H., and Liao, S.-w. (2021). Rethinking the implementation tricks and monotonicity constraint in cooperative multi-agent reinforcement learning. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Samvelyan, M., Rashid, T., de Witt, C.S., Farquhar, G., Nardelli, N., Rudner, T.G.J., Hung, C.M., Torr, P.H.S., Foerster, J., and Whiteson, S. (2019). The StarCraft Multi-Agent Challenge. arXiv.","DOI":"10.65109\/LVZZ5205"},{"key":"ref_32","unstructured":"Papoudakis, G., Christianos, F., Sch\u00e4fer, L., and Albrecht, S.V. (2020). Benchmarking multi-agent deep reinforcement learning algorithms in cooperative tasks. arXiv."},{"key":"ref_33","unstructured":"Singh, A., Jain, T., and Sukhbaatar, S. (2018). Learning when to Communicate at Scale in Multiagent Cooperative and Competitive Tasks. arXiv."},{"key":"ref_34","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"Maaten","year":"2008","journal-title":"J. Mach. Learn. Res."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/17\/4\/332\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,1]],"date-time":"2026-04-01T08:21:41Z","timestamp":1775031701000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/17\/4\/332"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,1]]},"references-count":34,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2026,4]]}},"alternative-id":["info17040332"],"URL":"https:\/\/doi.org\/10.3390\/info17040332","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,1]]}}}