{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,12]],"date-time":"2026-05-12T20:19:07Z","timestamp":1778617147833,"version":"3.51.4"},"reference-count":30,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2023,6,27]],"date-time":"2023-06-27T00:00:00Z","timestamp":1687824000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"European Research Council (ERC)","award":["833168"],"award-info":[{"award-number":["833168"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>We studied the ability of deep reinforcement learning and self-organizing approaches to adapt to dynamic complex systems, using the applied example of traffic signal control in a simulated urban environment. We highlight the general limitations of deep learning for control in complex systems, even when employing state-of-the-art meta-learning methods, and contrast it with self-organization-based methods. Accordingly, we argue that complex systems are a good and challenging study environment for developing and improving meta-learning approaches. At the same time, we point to the importance of baselines to which meta-learning methods can be compared and present a self-organizing analytic traffic signal control that outperforms state-of-the-art meta-learning in some scenarios. We also show that meta-learning methods outperform classical learning methods in our simulated environment (around 1.5\u20132\u00d7 improvement, in most scenarios). Our conclusions are that, in order to develop effective meta-learning methods that are able to adapt to a variety of conditions, it is necessary to test them in demanding, complex settings (such as, for example, urban traffic control) and compare them against established methods.<\/jats:p>","DOI":"10.3390\/e25070982","type":"journal-article","created":{"date-parts":[[2023,6,28]],"date-time":"2023-06-28T00:50:52Z","timestamp":1687913452000},"page":"982","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Deep Reinforcement Meta-Learning and Self-Organization in Complex Systems: Applications to Traffic Signal Control"],"prefix":"10.3390","volume":"25","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2784-5172","authenticated-orcid":false,"given":"Marcin","family":"Korecki","sequence":"first","affiliation":[{"name":"ETH Zurich, Computational Social Science, 8092 Zurich, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,6,27]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Van den Berg, J., Miller, S., Duckworth, D., Hu, H., Wan, A., Fu, X.Y., Goldberg, K., and Abbeel, P. (2010, January 3\u20138). Superhuman performance of surgical tasks by robots using iterative learning from human-guided demonstrations. Proceedings of the 2010 IEEE International Conference on Robotics and Automation, Anchorage, AK, USA.","DOI":"10.1109\/ROBOT.2010.5509621"},{"key":"ref_2","unstructured":"Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., and Graepel, T. (2017). Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm. arXiv."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"26339137221114874","DOI":"10.1177\/26339137221114874","article-title":"Collective Intelligence for Deep Learning: A Survey of Recent Developments","volume":"1","author":"Ha","year":"2021","journal-title":"Collect. Intell."},{"key":"ref_4","unstructured":"Finn, C., Abbeel, P., and Levine, S. (2017, January 6\u201311). Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks. Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Wei, H., Xu, N., Zhang, H., Zheng, G., Zang, X., Chen, C., Zhang, W., Zhu, Y., Xu, K., and Li, Z. (2019, January 3\u20137). Colight: Learning network-level cooperation for traffic signal control. Proceedings of the 28th ACM International Conference on Information and Knowledge Management, Beijing, China.","DOI":"10.1145\/3357384.3357902"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Wei, H., Zheng, G., Yao, H., and Li, Z. (2018, January 19\u201323). IntelliLight. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, London, UK.","DOI":"10.1145\/3219819.3220096"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Wei, H., Chen, C., Zheng, G., Wu, K., Gayah, V., Xu, K., and Li, Z. (2019, January 4\u20138). Presslight: Learning Max pressure control to coordinate traffic signals in arterial network. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Anchorage, AK, USA.","DOI":"10.1145\/3292500.3330949"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Zang, X., Yao, H., Zheng, G., Xu, N., Xu, K., and Li, Z. (2020, January 7\u201312). MetaLight: Value-Based Meta-Reinforcement Learning for Traffic Signal Control. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i01.5467"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Zhang, H., Liu, C., Zhang, W., Zheng, G., and Yu, Y. (2020, January 19\u201323). GeneraLight: Improving Environment Generalization of Traffic Signal Control via Meta Reinforcement Learning. Proceedings of the 29th ACM International Conference on Information & Knowledge Management, Online.","DOI":"10.1145\/3340531.3411859"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"686","DOI":"10.1016\/j.jcp.2018.10.045","article-title":"Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations","volume":"378","author":"Raissi","year":"2019","journal-title":"J. Comput. Phys."},{"key":"ref_11","first-page":"1","article-title":"AI Pontryagin or how artificial neural networks learn to control dynamical systems","volume":"13","author":"Asikis","year":"2022","journal-title":"Nat. Commun."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"96348","DOI":"10.1109\/ACCESS.2022.3204057","article-title":"Analytically guided reinforcement learning for green it and fluent traffic","volume":"10","author":"Korecki","year":"2022","journal-title":"IEEE Access"},{"key":"ref_13","first-page":"cnz018","article-title":"Deep learning systems as complex networks","volume":"8","author":"Testolin","year":"2018","journal-title":"J. Complex Netw."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3457607","article-title":"A Survey on Bias and Fairness in Machine Learning","volume":"54","author":"Mehrabi","year":"2021","journal-title":"ACM Comput. Surv."},{"key":"ref_15","unstructured":"Tommasi, T., Patricia, N., Caputo, B., and Tuytelaars, T. (2017). Domain Adaptation in Computer Vision Applications, Springer."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Mannion, P., Duggan, J., and Howley, E. (2016). An Experimental Review of Reinforcement Learning Algorithms for Adaptive Traffic Signal Control. Autonomic Road Transport Support Systems, Springer.","DOI":"10.1007\/978-3-319-25808-9_4"},{"key":"ref_17","unstructured":"Hospedales, T., Antoniou, A., Micaelli, P., and Storkey, A. (2020). IEEE Transactions on Pattern Analysis and Machine Intelligence, IEEE."},{"key":"ref_18","unstructured":"Huang, X., Wu, D., Jenkin, M., and Boulet, B. (2021). ModelLight: Model-Based Meta-Reinforcement Learning for Traffic Signal Control. arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Yang, S., and Yang, B. (2021, January 25\u201328). A Meta Multi-agent Reinforcement Learning Algorithm for Multi-intersection Traffic Signal Control. Proceedings of the 2021 IEEE International Symposium on Dependable, Autonomic and Secure Computing (DASC), Online.","DOI":"10.1109\/DASC-PICom-CBDCom-CyberSciTech52372.2021.00019"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"33","DOI":"10.1007\/s13194-012-0056-8","article-title":"What is a complex system?","volume":"3","author":"Ladyman","year":"2013","journal-title":"Eur. J. Philos. Sci."},{"key":"ref_21","unstructured":"Wolf, T.D., and Holvoet, T. Emergence versus self-organisation: Different concepts but promising when combined. Proceedings of the International Workshop on Engineering Self-Organising Applications."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1067","DOI":"10.1103\/RevModPhys.73.1067","article-title":"Traffic and related self-driven many-particle systems","volume":"73","author":"Helbing","year":"2001","journal-title":"Rev. Mod. Phys."},{"key":"ref_23","first-page":"253","article-title":"The science of self-organization and adaptivity","volume":"5","author":"Heylighen","year":"2001","journal-title":"Encycl. Life Support Syst."},{"key":"ref_24","unstructured":"Hayek, F.A. (1973). Law, Legislation and Liberty, Volume 1: Rules and Order, University of Chicago Press."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Gershenson, C., and Heylighen, F. When Can we Call a System Self-organizing? In Proceedings of the Advances in Artificial Life: 7th European Conference, ECAL 2003, Dortmund, Germany, 14\u201317 September 2003.","DOI":"10.1007\/978-3-540-39432-7_65"},{"key":"ref_26","unstructured":"Gershenson, C. (2004). Self-organizing traffic lights. arXiv."},{"key":"ref_27","first-page":"1","article-title":"Self-control of traffic lights and vehicle flows in urban road networks","volume":"2008","author":"Helbing","year":"2008","journal-title":"J. Stat. Mech. Theory Exp."},{"key":"ref_28","unstructured":"Zhang, H., Ding, Y., Zhang, W., Feng, S., Zhu, Y., Yu, Y., Li, Z., Liu, C., Zhou, Z., and Jin, H. (2019). The Web Conference 2019\u2014Proceedings of the World Wide Web Conference, WWW 2019, Association for Computing Machinery."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"16681","DOI":"10.1038\/s41598-022-21125-3","article-title":"Adaptability and sustainability of machine learning approaches to traffic signal control","volume":"12","author":"Korecki","year":"2022","journal-title":"Sci. Rep."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"36504","DOI":"10.1109\/ACCESS.2023.3266644","article-title":"How Well Do Reinforcement Learning Approaches Cope With Disruptions? The Case of Traffic Signal Control","volume":"11","author":"Korecki","year":"2023","journal-title":"IEEE Access"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/25\/7\/982\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:01:49Z","timestamp":1760126509000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/25\/7\/982"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,27]]},"references-count":30,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2023,7]]}},"alternative-id":["e25070982"],"URL":"https:\/\/doi.org\/10.3390\/e25070982","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,27]]}}}