{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,26]],"date-time":"2025-10-26T15:09:01Z","timestamp":1761491341769,"version":"3.40.5"},"reference-count":55,"publisher":"Cambridge University Press (CUP)","issue":"4","license":[{"start":{"date-parts":[[2021,12,27]],"date-time":"2021-12-27T00:00:00Z","timestamp":1640563200000},"content-version":"unspecified","delay-in-days":56,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":["cambridge.org"],"crossmark-restriction":true},"short-container-title":["AIEDAM"],"published-print":{"date-parts":[[2021,11]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Self-organizing systems (SOS) are developed to perform complex tasks in unforeseen situations with adaptability. Predefining rules for self-organizing agents can be challenging, especially in tasks with high complexity and changing environments. Our previous work has introduced a multiagent reinforcement learning (RL) model as a design approach to solving the rule generation problem of SOS. A deep multiagent RL algorithm was devised to train agents to acquire the task and self-organizing knowledge. However, the simulation was based on one specific task environment. Sensitivity of SOS to reward functions and systematic evaluation of SOS designed with multiagent RL remain an issue. In this paper, we introduced a rotation reward function to regulate agent behaviors during training and tested different weights of such reward on SOS performance in two case studies: box-pushing and T-shape assembly. Additionally, we proposed three metrics to evaluate the SOS: learning stability, quality of learned knowledge, and scalability. Results show that depending on the type of tasks; designers may choose appropriate weights of rotation reward to obtain the full potential of agents\u2019 learning capability. Good learning stability and quality of knowledge can be achieved with an optimal range of team sizes. Scaling up to larger team sizes has better performance than scaling downwards.<\/jats:p>","DOI":"10.1017\/s089006042100024x","type":"journal-article","created":{"date-parts":[[2021,12,27]],"date-time":"2021-12-27T10:47:12Z","timestamp":1640602032000},"page":"404-422","update-policy":"https:\/\/doi.org\/10.1017\/policypage","source":"Crossref","is-referenced-by-count":1,"title":["Evaluating the learning and performance characteristics of self-organizing systems with different task features"],"prefix":"10.1017","volume":"35","author":[{"given":"Hao","family":"Ji","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6502-5837","authenticated-orcid":false,"given":"Yan","family":"Jin","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2021,12,27]]},"reference":[{"key":"S089006042100024X_ref26","doi-asserted-by":"publisher","DOI":"10.1007\/0-387-27705-6_6"},{"key":"S089006042100024X_ref4","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4899-0718-9_28"},{"key":"S089006042100024X_ref31","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511546877"},{"key":"S089006042100024X_ref37","doi-asserted-by":"publisher","DOI":"10.1115\/1.4032091"},{"key":"S089006042100024X_ref25","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2003.1248936"},{"key":"S089006042100024X_ref1","unstructured":"Abramson, J , Ahuja, A , Barr, I , Brussee, A , Carnevale, F , Cassin, M , Chhaparia, R , Clark, S , Damoc, B , Dudzik, A , Georgiev, P , Guy, A , Harley, T , Hill, F , Hung, A , Kenton, Z , Landon, J , Lillicrap, T , Mathewson, K , Mokr\u00e1, S , Muldal, A , Santoro, A , Savinov, N , Varma, V , Wayne, G , Williams, D , Wong, N , Yan, C and Zhu, R (2020) Imitating interactive intelligence. arXiv preprint arXiv:2012.05672."},{"key":"S089006042100024X_ref50","doi-asserted-by":"crossref","unstructured":"Tan, M (1993) Multiagent reinforcement learning: Independent vs. cooperative agents. Proceedings of the Tenth International Conference on Machine Learning. pp. 330\u2013337.","DOI":"10.1016\/B978-1-55860-307-3.50049-6"},{"key":"S089006042100024X_ref7","doi-asserted-by":"publisher","DOI":"10.1109\/TSMCC.2007.913919"},{"key":"S089006042100024X_ref39","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"S089006042100024X_ref24","unstructured":"Ji, H and Jin, Y (2020) Designing self-assembly systems with deep multiagent reinforcement learning. Design Computing and Cognition\u201914. Springer, Cham, pp. xx\u2013xx."},{"key":"S089006042100024X_ref33","unstructured":"Lowe, R , Wu, Y , Tamar, A , Harb, J , Abbeel, P and Mordatch, I (2017) Multiagent actor-critic for mixed cooperative-competitive environments. arXiv preprint arXiv:1706.02275."},{"key":"S089006042100024X_ref13","unstructured":"Drogoul, A and Zucker, JD (1998) Methodological Issues for Designing Multiagent Systems with Machine Learning Techniques: Capitalizing Experiences from the Robocup Challenge (Doctoral dissertation, LIP6)."},{"key":"S089006042100024X_ref18","unstructured":"Hausknecht, M and Stone, P (2015) Deep recurrent q-learning for partially observable mdps. arXiv preprint arXiv:1507.06527."},{"key":"S089006042100024X_ref42","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2014.6906590"},{"key":"S089006042100024X_ref53","unstructured":"Watkins, CJCH (1989) Learning from delayed rewards."},{"key":"S089006042100024X_ref17","doi-asserted-by":"crossref","first-page":"1115","DOI":"10.1109\/TRO.2006.882919","article-title":"Autonomous self-assembly in swarm-bots","volume":"22","author":"Gro\u00df","year":"2006","journal-title":"IEEE Transactions on Robotics"},{"volume-title":"Reinforcement Learning: An Introduction","year":"2018","author":"Sutton","key":"S089006042100024X_ref48"},{"key":"S089006042100024X_ref38","unstructured":"Mnih, V , Kavukcuoglu, K , Silver, D , Graves, A , Antonoglou, I , Wierstra, D and Riedmiller, M (2013) Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602."},{"key":"S089006042100024X_ref30","doi-asserted-by":"publisher","DOI":"10.1109\/MCDM.2007.369410"},{"key":"S089006042100024X_ref8","doi-asserted-by":"publisher","DOI":"10.1115\/DETC2011-48833"},{"key":"S089006042100024X_ref28","doi-asserted-by":"publisher","DOI":"10.1115\/1.4032265"},{"key":"S089006042100024X_ref16","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11794"},{"key":"S089006042100024X_ref54","unstructured":"Wei, Y , Madey, GR and Blake, MB (2013) Agent-based simulation for uav swarm mission planning and execution. Proceedings of the Agent-Directed Simulation Symposium, pp. 1\u20138."},{"key":"S089006042100024X_ref41","unstructured":"Pippin, CE (2013) Trust and Reputation for Formation and Evolution of Multi-robot Teams (Doctoral dissertation). Georgia Institute of Technology."},{"volume-title":"Encyclopedia of Life Support Systems (EOLSS)","year":"2002","author":"Bar-Yam","key":"S089006042100024X_ref5"},{"key":"S089006042100024X_ref9","doi-asserted-by":"publisher","DOI":"10.1115\/DETC2012-71216"},{"key":"S089006042100024X_ref6","doi-asserted-by":"publisher","DOI":"10.1007\/978-94-010-0870-9_63"},{"key":"S089006042100024X_ref21","doi-asserted-by":"publisher","DOI":"10.1115\/DETC2016-60053"},{"key":"S089006042100024X_ref12","doi-asserted-by":"publisher","DOI":"10.1109\/TSMCA.2008.918619"},{"key":"S089006042100024X_ref45","unstructured":"Rashid, T , Samvelyan, M , Schroeder, C , Farquhar, G , Foerster, J and Whiteson, S (2018) Qmix: monotonic value function factorisation for deep multiagent reinforcement learning. International Conference on Machine Learning. PMLR, pp. 4295\u20134304."},{"key":"S089006042100024X_ref49","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0172395"},{"key":"S089006042100024X_ref55","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33902-8_5"},{"key":"S089006042100024X_ref2","doi-asserted-by":"publisher","DOI":"10.1115\/1.4040317"},{"key":"S089006042100024X_ref46","doi-asserted-by":"crossref","unstructured":"Reynolds, CW (1987) Flocks, herds and schools: a distributed behavioral model. Proceedings of the 14th Annual Conference on Computer Graphics and Interactive Techniques. pp. 25\u201334.","DOI":"10.1145\/37402.37406"},{"key":"S089006042100024X_ref19","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"S089006042100024X_ref11","doi-asserted-by":"publisher","DOI":"10.1080\/088395198117794"},{"key":"S089006042100024X_ref44","doi-asserted-by":"crossref","unstructured":"Rahimi, M , Gibb, S , Shen, Y and La, HM (2018) A comparison of various approaches to reinforcement learning algorithms for multi-robot box pushing. International Conference on Engineering Research and Applications. Cham: Springer, pp. 16\u201330.","DOI":"10.1007\/978-3-030-04792-4_6"},{"key":"S089006042100024X_ref52","unstructured":"Wang, Z , Schaul, T , Hessel, M , Hasselt, H , Lanctot, M and Freitas, N (2016) Dueling network architectures for deep reinforcement learning. International Conference on Machine Learning. PMLR, pp. 1995\u20132003."},{"key":"S089006042100024X_ref20","first-page":"259","article-title":"Evolutionary computational synthesis of self-organizing systems","volume":"28","author":"Humann","year":"2014","journal-title":"AI EDAM"},{"key":"S089006042100024X_ref29","doi-asserted-by":"publisher","DOI":"10.1115\/1.4031714"},{"key":"S089006042100024X_ref34","first-page":"DFM-4359-1","article-title":"Design for variety: development of complexity indices and design charts","author":"Martin","year":"1997","journal-title":"Proceedings of ASME 1997 Design Engineering Technical Conferences, September 14\u201317, 1997, Sacramento, CA"},{"key":"S089006042100024X_ref22","doi-asserted-by":"publisher","DOI":"10.1115\/DETC2018-86006"},{"key":"S089006042100024X_ref43","doi-asserted-by":"crossref","unstructured":"Price, IC and Lamont, GB (2006) GA directed self-organized search and attack UAV swarms. Proceedings of the 2006 Winter Simulation Conference. IEEE, pp. 1307\u20131315.","DOI":"10.1109\/WSC.2006.323229"},{"key":"S089006042100024X_ref36","doi-asserted-by":"publisher","DOI":"10.1115\/1.4039494"},{"volume-title":"An Introduction to Cybernetics","year":"1961","author":"Ashby","key":"S089006042100024X_ref3"},{"key":"S089006042100024X_ref40","first-page":"1","article-title":"Deeploco: dynamic locomotion skills using hierarchical deep reinforcement learning","volume":"36","author":"Peng","year":"2017","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"S089006042100024X_ref47","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2009.06.032"},{"key":"S089006042100024X_ref35","doi-asserted-by":"publisher","DOI":"10.1115\/1.4035793"},{"key":"S089006042100024X_ref10","unstructured":"Chung, J , Gulcehre, C , Cho, K and Bengio, Y (2014) Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555."},{"key":"S089006042100024X_ref15","unstructured":"Foerster, J , Nardelli, N , Farquhar, G , Afouras, T , Torr, PH , Kohli, P and Whiteson, S (2017) Stabilising experience replay for deep multiagent reinforcement learning. International Conference on Machine Learning. PMLR, pp. 1146\u20131155."},{"key":"S089006042100024X_ref27","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1007\/978-3-319-14956-1_1","volume-title":"Design Computing and Cognition\u201914","author":"Khani","year":"2015"},{"key":"S089006042100024X_ref14","doi-asserted-by":"publisher","DOI":"10.2514\/1.17147"},{"key":"S089006042100024X_ref23","doi-asserted-by":"publisher","DOI":"10.1115\/DETC2019-98268"},{"key":"S089006042100024X_ref32","doi-asserted-by":"crossref","unstructured":"Liu, X and Jin, Y (2018) Design of transfer reinforcement learning mechanisms for autonomous collision avoidance. International Conference on-Design Computing and Cognition. Cham: Springer, pp. 303\u2013319.","DOI":"10.1007\/978-3-030-05363-5_17"},{"key":"S089006042100024X_ref51","doi-asserted-by":"crossref","unstructured":"Wang, Y and De Silva, CW (2006) Multi-robot box-pushing: single-agent q-learning vs. team q-learning.2006 IEEE\/RSJ International Conference on Intelligent Robots and Systems. IEEE, pp. 3694\u20133699.","DOI":"10.1109\/IROS.2006.281729"}],"container-title":["Artificial Intelligence for Engineering Design, Analysis and Manufacturing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S089006042100024X","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,11,14]],"date-time":"2023-11-14T21:54:08Z","timestamp":1699998848000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S089006042100024X\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,11]]},"references-count":55,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2021,11]]}},"alternative-id":["S089006042100024X"],"URL":"https:\/\/doi.org\/10.1017\/s089006042100024x","relation":{},"ISSN":["0890-0604","1469-1760"],"issn-type":[{"type":"print","value":"0890-0604"},{"type":"electronic","value":"1469-1760"}],"subject":[],"published":{"date-parts":[[2021,11]]},"assertion":[{"value":"Copyright \u00a9 The Author(s), 2021. Published by Cambridge University Press","name":"copyright","label":"Copyright","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}}]}}