{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,5]],"date-time":"2026-02-05T04:35:05Z","timestamp":1770266105477,"version":"3.49.0"},"reference-count":111,"publisher":"Association for Computing Machinery (ACM)","issue":"7","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2026,5,31]]},"abstract":"<jats:p>The Digital Twins (DT) paradigm has emerged as a powerful tool for simulating and analyzing complex systems in various domains. A DT is a virtual representation of a real-world object(s) whose goal is to accurately emulate real systems, optimize processes, minimize synchronization delays, cut down on overhead, and automate decision-making. DT technology is moving at a faster than expected pace with advances in Artificial Intelligence (AI), Internet of Things (IoT), Distributed Computing, and 5\/6G. Being a highly beneficial technology, DT still faces issues of - (1) limited adaptability, (2) incomplete model representation, (3) suboptimal decision making, (4) limited generalization, and (5) scalability and computational efficiency. Reinforcement Learning (RL) offers unsupervised decision-making and intelligence, which can be immensely beneficial in addressing the current challenges faced by DT. This study offers a thorough analysis of the DT paradigm from the standpoint of RL. The survey compares and contrasts existing reinforcement learning-based Digital Twin frameworks, assessing their advantages and disadvantages. Moreover, discussions of approaches highlighting the tradeoffs between simulation fidelity and computing complexity is also studied. Additionally, a thorough understanding of the Digital Twins paradigm from a reinforcement learning perspective, is presented as a helpful resource for academics and industry professionals in the field. Finally, future research directions in this developing field at the nexus of digital modeling, simulation, and artificial intelligence is discussed.<\/jats:p>","DOI":"10.1145\/3777367","type":"journal-article","created":{"date-parts":[[2025,12,2]],"date-time":"2025-12-02T11:03:37Z","timestamp":1764673417000},"page":"1-33","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Digital Twins Paradigm: A Systematic Review from the Reinforcement Learning Perspective"],"prefix":"10.1145","volume":"58","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0740-5257","authenticated-orcid":false,"given":"Shahmir Khan","family":"Mohammed","sequence":"first","affiliation":[{"name":"Khalifa University of Science and Technology","place":["Abu Dhabi, United Arab Emirates"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7977-7178","authenticated-orcid":false,"given":"Shakti","family":"Singh","sequence":"additional","affiliation":[{"name":"Khalifa University of Science and Technology","place":["Abu Dhabi, United Arab Emirates"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6915-3759","authenticated-orcid":false,"given":"Rabeb","family":"Mizouni","sequence":"additional","affiliation":[{"name":"Khalifa University of Science and Technology","place":["Abu Dhabi, United Arab Emirates"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9574-5384","authenticated-orcid":false,"given":"Hadi","family":"Otrok","sequence":"additional","affiliation":[{"name":"Khalifa University of Science and Technology","place":["Abu Dhabi, United Arab Emirates"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9557-6496","authenticated-orcid":false,"given":"Ernesto","family":"Damiani","sequence":"additional","affiliation":[{"name":"Khalifa University of Science and Technology","place":["Abu Dhabi, United Arab Emirates"]},{"name":"Computer Science Department, Universit\u00e0 degli Studi di Milano","place":["Abu Dhabi, United Arab Emirates"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,2,3]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"2023. Digital twin market by application. Retrieved October 20 2023 from https:\/\/www.marketsandmarkets.com\/Market-Reports\/digital-twin-market-225269522.html (2023)."},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2023.3296809"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.vehcom.2025.100874"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/COMST.2022.3171465"},{"key":"e_1_3_2_6_2","volume-title":"Proceedings of the Digital Twin Summit","author":"Allen B. Danette","year":"2021","unstructured":"B. Danette Allen. 2021. Digital twins and living models at NASA. In Proceedings of the Digital Twin Summit."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/MCE.2022.3212570"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2953499"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.rser.2021.110801"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSIPN.2022.3171336"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.compind.2019.103130"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/COASE.2019.8842888"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/LCOMM.2020.3022737"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2020.3016320"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.23919\/JCIN.2022.9745481"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00422-022-00940-x"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.conengprac.2021.104790"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.3389\/frobt.2021.758099"},{"key":"e_1_3_2_19_2","unstructured":"Jeff Druce Michael Harradon and James Tittle. 2021. Explainable artificial intelligence (XAI) for increasing user trust in deep reinforcement learning driven autonomous systems. arXiv:2106.03775. Retrieved from https:\/\/arxiv.org\/abs\/2106.03775"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2021.3051158"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/mnet.201.2000768"},{"issue":"1","key":"e_1_3_2_22_2","first-page":"8","article-title":"Applying digital twins in metaverse: User interface, security and privacy challenges","volume":"2","author":"Far Saeed Banaeian","year":"2022","unstructured":"Saeed Banaeian Far and Azadeh Imani Rad. 2022. Applying digital twins in metaverse: User interface, security and privacy challenges. Journal of Metaverse 2, 1 (2022), 8\u201315.","journal-title":"Journal of Metaverse"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-38756-7_4"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1504\/IJPD.2005.006669"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSMC.2022.3231299"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2022.118302"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.3390\/app14030977"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2022.3227655"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2023.3313887"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2022.3205778"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jmsy.2020.02.004"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.3390\/s19204410"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cirpj.2020.02.002"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.5555\/1622737.1622748"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCTURKEY53027.2021.9654360"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.3390\/a14080240"},{"key":"e_1_3_2_37_2","article-title":"Actor-critic algorithms","volume":"12","author":"Konda Vijay","year":"1999","unstructured":"Vijay Konda and John Tsitsiklis. 1999. Actor-critic algorithms. Advances in Neural Information Processing Systems 12 (1999).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/BigData.2018.8622160"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ifacol.2018.08.474"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.aei.2022.101710"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.autcon.2022.104716"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2022.3182647"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.rcim.2022.102321"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.rcim.2022.102471"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1038\/s43017-023-00409-w"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/VTC2022-Fall57202.2022.10012857"},{"key":"e_1_3_2_47_2","first-page":"357","volume-title":"Proceedings of the Federated and Transfer Learning","author":"Liang Xinle","year":"2022","unstructured":"Xinle Liang, Yang Liu, Tianjian Chen, Ming Liu, and Qiang Yang. 2022. Federated transfer reinforcement learning for autonomous driving. In Proceedings of the Federated and Transfer Learning. Springer, 357\u2013371."},{"key":"e_1_3_2_48_2","unstructured":"Timothy P. Lillicrap Jonathan J. Hunt Alexander Pritzel Nicolas Heess Tom Erez Yuval Tassa David Silver and Daan Wierstra. 2015. Continuous control with deep reinforcement learning. arXiv:1509.02971. Retrieved from https:\/\/arxiv.org\/abs\/1509.02971"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jmsy.2020.04.014"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.rcim.2022.102365"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2909828"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1063\/1.5031520"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1061\/(ASCE)ME.1943-5479.0000763"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.autcon.2020.103277"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2020.3015772"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2020.3017668"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2021.3098508"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cag.2021.01.011"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/COMST.2022.3208773"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2020.2998530"},{"key":"e_1_3_2_61_2","unstructured":"Volodymyr Mnih Koray Kavukcuoglu David Silver Alex Graves Ioannis Antonoglou Daan Wierstra and Martin Riedmiller. 2013. Playing atari with deep reinforcement learning. arXiv:1312.5602. Retrieved from https:\/\/arxiv.org\/abs\/1312.5602"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature14236"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.iot.2023.100713"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jnca.2023.103793"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.apmt.2018.11.003"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/ETFA46521.2020.9211946"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA48506.2021.9561138"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/IRC.2019.00120"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/MCOM.001.2000343"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jmsy.2020.06.018"},{"key":"e_1_3_2_71_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jobe.2021.102726"},{"key":"e_1_3_2_72_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.oceaneng.2021.109433"},{"key":"e_1_3_2_73_2","article-title":"Adaptive batch size for safe policy gradients","volume":"30","author":"Papini Matteo","year":"2017","unstructured":"Matteo Papini, Matteo Pirotta, and Marcello Restelli. 2017. Adaptive batch size for safe policy gradients. Advances in Neural Information Processing Systems 30 (2017).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_74_2","doi-asserted-by":"publisher","DOI":"10.1080\/00207543.2021.1884309"},{"key":"e_1_3_2_75_2","doi-asserted-by":"publisher","DOI":"10.3390\/app11072977"},{"key":"e_1_3_2_76_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSS.2023.3262958"},{"key":"e_1_3_2_77_2","doi-asserted-by":"publisher","DOI":"10.1109\/INDIN45523.2021.9557372"},{"key":"e_1_3_2_78_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2022.3160709"},{"key":"e_1_3_2_79_2","doi-asserted-by":"publisher","DOI":"10.1109\/TEM.2022.3201434"},{"key":"e_1_3_2_80_2","doi-asserted-by":"publisher","DOI":"10.1109\/IROS45743.2020.9340848"},{"key":"e_1_3_2_81_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-020-03051-4"},{"key":"e_1_3_2_82_2","unstructured":"John Schulman Filip Wolski Prafulla Dhariwal Alec Radford and Oleg Klimov. 2017. Proximal policy optimization algorithms. arXiv:1707.06347. Retrieved from https:\/\/arxiv.org\/abs\/1707.06347"},{"key":"e_1_3_2_83_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2021.3127873"},{"key":"e_1_3_2_84_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.iot.2023.100867"},{"key":"e_1_3_2_85_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2022.107145"},{"key":"e_1_3_2_86_2","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2020.3034674"},{"key":"e_1_3_2_87_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2020.3018817"},{"key":"e_1_3_2_88_2","volume-title":"Reinforcement Learning: An Introduction","year":"1998","unstructured":"Richard S. Sutton and Andrew G. Barto. 1998. Reinforcement Learning: An Introduction. Vol. 1. MIT press Cambridge."},{"key":"e_1_3_2_89_2","first-page":"1","volume-title":"Proceedings of the 2020 IEEE 31st Annual International Symposium on Personal, Indoor and Mobile Radio Communications","author":"Szab\u00f3 G\u00e9za","year":"2020","unstructured":"G\u00e9za Szab\u00f3, J\u00f3zsef Pet\u0151, Levente N\u00e9meth, and Attila Vid\u00e1cs. 2020. Information gain regulation in reinforcement learning with the digital twins\u2019 level of realism. In Proceedings of the 2020 IEEE 31st Annual International Symposium on Personal, Indoor and Mobile Radio Communications. IEEE, 1\u20137."},{"key":"e_1_3_2_90_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00170-022-09700-4"},{"key":"e_1_3_2_91_2","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2018.2873186"},{"key":"e_1_3_2_92_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ymssp.2021.108284"},{"key":"e_1_3_2_93_2","doi-asserted-by":"publisher","DOI":"10.1109\/MECOM61498.2024.10881535"},{"key":"e_1_3_2_94_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2022.3162714"},{"key":"e_1_3_2_95_2","article-title":"Bibliometric analysis of digital twin literature: A review of influencing factors and conceptual structure","author":"Wang Jing","year":"2024","unstructured":"Jing Wang, Xinchun Li, Peng Wang, and Quanlong Liu. 2024. Bibliometric analysis of digital twin literature: A review of influencing factors and conceptual structure. Technology Analysis and Strategic Management 36, 1 (2024), 166\u2013180.","journal-title":"Technology Analysis and Strategic Management"},{"key":"e_1_3_2_96_2","doi-asserted-by":"publisher","DOI":"10.1109\/VTC2022-Spring54318.2022.9860495"},{"key":"e_1_3_2_97_2","first-page":"1995","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Wang Ziyu","year":"2016","unstructured":"Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Hasselt, Marc Lanctot, and Nando Freitas. 2016. Dueling network architectures for deep reinforcement learning. In Proceedings of the International Conference on Machine Learning. PMLR, 1995\u20132003."},{"key":"e_1_3_2_98_2","unstructured":"Christopher John Cornish Hellaby Watkins et\u00a0al. 1989. Learning from delayed rewards. (1989)."},{"key":"e_1_3_2_99_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cities.2020.103064"},{"key":"e_1_3_2_100_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jmsy.2020.06.012"},{"key":"e_1_3_2_101_2","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2020.3040180"},{"key":"e_1_3_2_102_2","doi-asserted-by":"publisher","DOI":"10.3390\/info13060286"},{"key":"e_1_3_2_103_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cor.2022.105823"},{"key":"e_1_3_2_104_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.dcan.2022.05.005"},{"key":"e_1_3_2_105_2","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2022.3183465"},{"key":"e_1_3_2_106_2","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2022.3204585"},{"key":"e_1_3_2_107_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i01.5467"},{"key":"e_1_3_2_108_2","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2021.3088407"},{"key":"e_1_3_2_109_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSS.2021.3096824"},{"key":"e_1_3_2_110_2","unstructured":"Oleksii Zhelo Jingwei Zhang Lei Tai Ming Liu and Wolfram Burgard. 2018. Curiosity-driven exploration for mapless navigation with deep reinforcement learning. arXiv:1804.00456. Retrieved from https:\/\/arxiv.org\/abs\/1804.00456"},{"key":"e_1_3_2_111_2","first-page":"3757","article-title":"Episodic multi-agent reinforcement learning with curiosity-driven exploration","volume":"34","author":"Zheng Lulu","year":"2021","unstructured":"Lulu Zheng, Jiarui Chen, Jianhao Wang, Jiamin He, Yujing Hu, Yingfeng Chen, Changjie Fan, Yang Gao, and Chongjie Zhang. 2021. Episodic multi-agent reinforcement learning with curiosity-driven exploration. Advances in Neural Information Processing Systems 34 (2021), 3757\u20133769.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_112_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCCN.2022.3204572"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3777367","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,3]],"date-time":"2026-02-03T14:19:13Z","timestamp":1770128353000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3777367"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,3]]},"references-count":111,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,5,31]]}},"alternative-id":["10.1145\/3777367"],"URL":"https:\/\/doi.org\/10.1145\/3777367","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,3]]},"assertion":[{"value":"2023-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-07","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}