{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T16:35:26Z","timestamp":1783701326479,"version":"3.55.0"},"reference-count":52,"publisher":"ASME International","issue":"2","license":[{"start":{"date-parts":[[2021,12,9]],"date-time":"2021-12-09T00:00:00Z","timestamp":1639008000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.asme.org\/publications-submissions\/publishing-information\/legal-policies"}],"content-domain":{"domain":["asmedigitalcollection.asme.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,4,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Self-organizing systems (SOS) can perform complex tasks in unforeseen situations with adaptability. Previous work has introduced field-based approaches and rule-based social structuring for individual agents to not only comprehend the task situations but also take advantage of the social rule-based agent relations to accomplish their tasks without a centralized controller. Although the task fields and social rules can be predefined for relatively simple task situations, when the task complexity increases and the task environment changes, having a priori knowledge about these fields and the rules may not be feasible. In this paper, a multiagent reinforcement learning (RL) based model is proposed as a design approach to solving the rule generation problem with complex SOS tasks. A deep multiagent reinforcement learning algorithm was devised as a mechanism to train SOS agents for knowledge acquisition of the task field and social rules. Learning stability, functional differentiation, and robustness properties of this learning approach were investigated with respect to the changing team sizes and task variations. Through computer simulation studies of a box-pushing problem, the results have shown that there is an optimal range of the number of agents that achieves good learning stability; agents in a team learn to differentiate from other agents with changing team sizes and box dimensions; the robustness of the learned knowledge shows to be stronger to the external noises than with changing task constraints.<\/jats:p>","DOI":"10.1115\/1.4052800","type":"journal-article","created":{"date-parts":[[2021,10,22]],"date-time":"2021-10-22T08:09:50Z","timestamp":1634890190000},"update-policy":"https:\/\/doi.org\/10.1115\/crossmarkpolicy-asme","source":"Crossref","is-referenced-by-count":13,"title":["Knowledge Acquisition of Self-Organizing Systems With Deep Multiagent Reinforcement Learning"],"prefix":"10.1115","volume":"22","author":[{"given":"Hao","family":"Ji","sequence":"first","affiliation":[{"name":"Department of Aerospace and Mechanical Engineering, University of Southern California, 3650 McClintock Avenue, OHE 400, Los Angeles, CA 90089-1453"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yan","family":"Jin","sequence":"additional","affiliation":[{"name":"Department of Aerospace and Mechanical Engineering, University of Southern California, 3650 McClintock Avenue, OHE 400, Los Angeles, CA 90089-1453"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"33","published-online":{"date-parts":[[2021,12,9]]},"reference":[{"key":"2021120911152022800_CIT0001","first-page":"25","article-title":"Flocks, Herds and Schools: A Distributed Behavioral Model","author":"Reynolds","year":"1987"},{"key":"2021120911152022800_CIT0002","doi-asserted-by":"crossref","first-page":"405","DOI":"10.1007\/978-1-4899-0718-9_28","volume-title":"Facets of Systems Science","author":"Ashby","year":"1991"},{"key":"2021120911152022800_CIT0003","first-page":"511","article-title":"Design of Cellular Self-Organizing Systems","author":"Chiang","year":"2012"},{"issue":"3","key":"2021120911152022800_CIT0004","doi-asserted-by":"publisher","first-page":"259","DOI":"10.1017\/s0890060414000213","article-title":"Evolutionary Computational Synthesis of Self-Organizing Systems","volume":"28","author":"Humann","year":"2014","journal-title":"AI EDAM"},{"issue":"4","key":"2021120911152022800_CIT0005","doi-asserted-by":"publisher","first-page":"041101","DOI":"10.1115\/1.4032265","article-title":"Effect of Social Structuring in Self-Organizing Systems","volume":"138","author":"Khani","year":"2016","journal-title":"ASME J. Mech. Des."},{"key":"2021120911152022800_CIT0006","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1007\/978-3-319-14956-1_1","volume-title":"Design Computing and Cognition\u201914","author":"Khani","year":"2015"},{"key":"2021120911152022800_CIT0007","doi-asserted-by":"crossref","DOI":"10.1115\/DETC2018-86006","article-title":"Modeling Trust in Self-Organizing Systems With Heterogeneity","author":"Ji","year":"2018"},{"key":"2021120911152022800_CIT0008","first-page":"95","article-title":"A Behavior Based Approach to Cellular Self-Organizing Systems Design","author":"Chen","year":"2011"},{"key":"2021120911152022800_CIT0009","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton","year":"2018"},{"key":"2021120911152022800_CIT0010","first-page":"4295","article-title":"Qmix: Monotonic Value Function Factorisation for Deep Multiagent Reinforcement Learning","author":"Rashid","year":"2018"},{"key":"2021120911152022800_CIT0011","volume-title":"General Features of Complex Systems. Encyclopedia of Life Support Systems (EOLSS)","author":"Bar-Yam","year":"2002"},{"issue":"9","key":"2021120911152022800_CIT0012","doi-asserted-by":"publisher","first-page":"091101","DOI":"10.1115\/1.4040317","article-title":"Exploring Natural Strategies for Bio-Inspired Fault Adaptive Systems Design","volume":"140","author":"Arroyo","year":"2018","journal-title":"ASME J. Mech. Des."},{"issue":"1","key":"2021120911152022800_CIT0013","doi-asserted-by":"publisher","first-page":"011102","DOI":"10.1115\/1.4031714","article-title":"Comparing Strategies for Topologic and Parametric Rule Application in Automated Computational Design Synthesis","volume":"138","author":"K\u00f6nigseder","year":"2016","journal-title":"ASME J. Mech. Des."},{"issue":"12","key":"2021120911152022800_CIT0014","doi-asserted-by":"publisher","first-page":"121101","DOI":"10.1115\/1.4039494","article-title":"Gaming the System: An Agent-Based Model of Estimation Strategies and Their Effects on System Performance","volume":"140","author":"Meluso","year":"2018","journal-title":"ASME J. Mech. Des."},{"issue":"4","key":"2021120911152022800_CIT0015","doi-asserted-by":"publisher","DOI":"10.1115\/1.4035793","article-title":"Optimizing Design Teams Based on Problem Properties: Computational Team Simulations and an Applied Empirical Test","volume":"139","author":"McComb","year":"2017","journal-title":"ASME J. Mech. Des."},{"issue":"2","key":"2021120911152022800_CIT0016","doi-asserted-by":"publisher","first-page":"021102","DOI":"10.1115\/1.4032091","article-title":"System Architecture, Level of Decomposition, and Structural Complexity: Analysis and Observations","volume":"138","author":"Min","year":"2016","journal-title":"ASME J. Mech. Des."},{"issue":"4","key":"2021120911152022800_CIT0017","doi-asserted-by":"publisher","first-page":"868","DOI":"10.2514\/1.17147","article-title":"Effective Development of Reconfigurable Systems Using Linear State-Feedback Control","volume":"44","author":"Ferguson","year":"2006","journal-title":"AIAA J."},{"key":"2021120911152022800_CIT0018","doi-asserted-by":"crossref","DOI":"10.1115\/DETC97\/DFM-4359","article-title":"Design for Variety: Development of Complexity Indices and Design Charts","author":"Martin","year":"1997"},{"key":"2021120911152022800_CIT0019","doi-asserted-by":"crossref","first-page":"115","DOI":"10.1007\/978-3-642-33902-8_5","volume-title":"Morphogenetic Engineering","author":"Werfel","year":"2012"},{"key":"2021120911152022800_CIT0020","doi-asserted-by":"crossref","first-page":"1008","DOI":"10.1007\/978-94-010-0870-9_63","volume-title":"Prerational Intelligence: Adaptive Behavior and Intelligent Systems Without Symbols and Logic, Volume 1, Volume 2 Prerational Intelligence: Interdisciplinary Perspectives on the Behavior of Natural and Artificial Systems, Volume 3","author":"Beckers","year":"2000"},{"issue":"3","key":"2021120911152022800_CIT0021","doi-asserted-by":"publisher","first-page":"549","DOI":"10.1109\/TSMCA.2008.918619","article-title":"A Multiagent Swarming System for Distributed Automatic Target Recognition Using Unmanned Aerial Vehicles","volume":"38","author":"Dasgupta","year":"2008","journal-title":"IEEE Trans. Syst. Man Cybern. Part A Syst. Humans"},{"issue":"5\u20136","key":"2021120911152022800_CIT0022","doi-asserted-by":"publisher","first-page":"812","DOI":"10.1016\/j.neunet.2009.06.032","article-title":"Extending the Evolutionary Robotics Approach to Flying Machines: An Application to MAV Teams","volume":"22","author":"Ruini","year":"2009","journal-title":"Neural Networks"},{"key":"2021120911152022800_CIT0023","first-page":"10","article-title":"UAV Swarm Mission Planning and Routing Using Multi-Objective Evolutionary Algorithms","author":"Lamont","year":"2007"},{"key":"2021120911152022800_CIT0024","first-page":"1","article-title":"Agent-Based Simulation for UAV Swarm Mission Planning and Execution","author":"Wei","year":"2013"},{"key":"2021120911152022800_CIT0025","first-page":"1307","article-title":"GA Directed Self-Organized Search and Attack UAV Swarms","author":"Price","year":"2006"},{"issue":"2","key":"2021120911152022800_CIT0026","doi-asserted-by":"publisher","first-page":"156","DOI":"10.1109\/TSMCC.2007.913919","article-title":"A Comprehensive Survey of Multiagent Reinforcement Learning","volume":"38","author":"Busoniu","year":"2008","journal-title":"IEEE Trans. Syst. Man Cybern. Part C Appl. Rev."},{"issue":"4","key":"2021120911152022800_CIT0027","doi-asserted-by":"publisher","first-page":"e0172395","DOI":"10.1371\/journal.pone.0172395","article-title":"Multiagent Cooperation and Competition With Deep Reinforcement Learning","volume":"12","author":"Tampuu","year":"2017","journal-title":"PLoS One"},{"key":"2021120911152022800_CIT0028","article-title":"Counterfactual Multiagent Policy Gradients","author":"Foerster","year":"2018"},{"issue":"4","key":"2021120911152022800_CIT0029","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3072959.3073602","article-title":"Deeploco: Dynamic Locomotion Skills Using Hierarchical Deep Reinforcement Learning","volume":"36","author":"Peng","year":"2017","journal-title":"ACM Trans. Graph."},{"key":"2021120911152022800_CIT0030","first-page":"330","article-title":"Multiagent Reinforcement Learning: Independent vs. Cooperative Agents","author":"Tan","year":"1993"},{"key":"2021120911152022800_CIT0031","unstructured":"Watkins, C. J. C. H. , 1989, \u201cLearning From Delayed Rewards,\u201d Ph.D. dissertation, Cambridge University, Cambridge, UK."},{"issue":"7540","key":"2021120911152022800_CIT0032","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level Control Through Deep Reinforcement Learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"2021120911152022800_CIT0033","first-page":"1146","article-title":"Stabilising Experience Replay for Deep Multiagent Reinforcement Learning","author":"Foerster","year":"2017"},{"key":"2021120911152022800_CIT0034","article-title":"Deep Recurrent Q-Learning for Partially Observable MDPs","author":"Hausknecht","year":"2015"},{"issue":"8","key":"2021120911152022800_CIT0035","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long Short-Term Memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"2021120911152022800_CIT0036","article-title":"Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling","author":"Chung","year":"2014"},{"key":"2021120911152022800_CIT0037","article-title":"Multiagent Actor-Critic for Mixed Cooperative-Competitive Environments","author":"Lowe","year":"2017"},{"issue":"6456","key":"2021120911152022800_CIT0038","doi-asserted-by":"publisher","first-page":"885","DOI":"10.1126\/science.aay2400","article-title":"Superhuman AI for Multiplayer Poker","volume":"365","author":"Brown","year":"2019","journal-title":"Science"},{"key":"2021120911152022800_CIT0039","article-title":"Emergent Tool Use From Multiagent Autocurricula","author":"Baker","year":"2019"},{"issue":"2","key":"2021120911152022800_CIT0040","doi-asserted-by":"publisher","first-page":"414","DOI":"10.1111\/tops.12525","article-title":"Too Many Cooks: Bayesian Inference for Coordinating Multi-Agent Collaboration","volume":"13","author":"Wu","year":"2021","journal-title":"Top. Cogn. Sci."},{"key":"2021120911152022800_CIT0041","first-page":"3694","article-title":"Multi-Robot Box-Pushing: Single-Agent Q-Learning vs. Team Q-Learning","author":"Wang","year":"2006"},{"key":"2021120911152022800_CIT0042","first-page":"16","article-title":"A Comparison of Various Approaches to Reinforcement Learning Algorithms for Multi-Robot Box Pushing","author":"Rahimi","year":"2018"},{"key":"2021120911152022800_CIT0043","article-title":"Playing Atari With Deep Reinforcement Learning","author":"Mnih","year":"2013"},{"key":"2021120911152022800_CIT0044","first-page":"1995","article-title":"Dueling Network Architectures for Deep Reinforcement Learning","author":"Wang","year":"2016"},{"key":"2021120911152022800_CIT0045","article-title":"Learning to Communicate to Solve Riddles With Deep Distributed Recurrent Q-Networks","author":"Foerster","year":"2016"},{"key":"2021120911152022800_CIT0046","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511546877","volume-title":"Planning Algorithms","author":"LaValle","year":"2006"},{"key":"2021120911152022800_CIT0047","first-page":"1969","article-title":"Adaptive Division of Labor in Large-Scale Minimalist Multi-Robot Systems","author":"Jones","year":"2003"},{"issue":"6","key":"2021120911152022800_CIT0048","doi-asserted-by":"publisher","first-page":"1115","DOI":"10.1109\/TRO.2006.882919","article-title":"Autonomous Self-Assembly in Swarm-Bots","volume":"22","author":"Gro\u00df","year":"2006","journal-title":"IEEE Trans. Rob."},{"key":"2021120911152022800_CIT0049","doi-asserted-by":"crossref","DOI":"10.1115\/DETC2016-60053","article-title":"Adaptability Tradeoffs in the Design of Self-Organizing Systems","author":"Humann","year":"2016"},{"key":"2021120911152022800_CIT0050","first-page":"303","article-title":"Design of Transfer Reinforcement Learning Mechanisms for Autonomous Collision Avoidance","author":"Liu","year":"2018"},{"key":"2021120911152022800_CIT0051","volume-title":"An Introduction to Cybernetics","author":"Ashby","year":"1961"},{"key":"2021120911152022800_CIT0052","first-page":"246","article-title":"Hierarchical Multiagent Reinforcement Learning","author":"Makar","year":"2001"}],"container-title":["Journal of Computing and Information Science in Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/asmedigitalcollection.asme.org\/computingengineering\/article-pdf\/22\/2\/021010\/6808765\/jcise_22_2_021010.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/asmedigitalcollection.asme.org\/computingengineering\/article-pdf\/22\/2\/021010\/6808765\/jcise_22_2_021010.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,11]],"date-time":"2023-11-11T14:58:09Z","timestamp":1699714689000},"score":1,"resource":{"primary":{"URL":"https:\/\/asmedigitalcollection.asme.org\/computingengineering\/article\/22\/2\/021010\/1122800\/Knowledge-Acquisition-of-Self-Organizing-Systems"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,12,9]]},"references-count":52,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2022,4,1]]}},"URL":"https:\/\/doi.org\/10.1115\/1.4052800","relation":{},"ISSN":["1530-9827","1944-7078"],"issn-type":[{"value":"1530-9827","type":"print"},{"value":"1944-7078","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,12,9]]},"article-number":"021010"}}