{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,28]],"date-time":"2026-07-28T22:15:19Z","timestamp":1785276919378,"version":"3.55.0"},"reference-count":59,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2023,7,26]],"date-time":"2023-07-26T00:00:00Z","timestamp":1690329600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100004837","name":"Spanish Ministry of Science and Innovation","doi-asserted-by":"crossref","award":["PID2021- 122136OB-C21"],"award-info":[{"award-number":["PID2021- 122136OB-C21"]}],"id":[{"id":"10.13039\/501100004837","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100010661","name":"Horizon 2020 Framework Programme","doi-asserted-by":"publisher","award":["739578"],"award-info":[{"award-number":["739578"]}],"id":[{"id":"10.13039\/100010661","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2023,8]]},"abstract":"<jats:p>Simulating crowds with realistic behaviors is a difficult but very important task for a variety of applications. Quantifying how a person balances between different conflicting criteria such as goal seeking, collision avoidance and moving within a group is not intuitive, especially if we consider that behaviors differ largely between people. Inspired by recent advances in Deep Reinforcement Learning, we propose Guided REinforcement Learning (GREIL) Crowds, a method that learns a model for pedestrian behaviors which is guided by reference crowd data. The model successfully captures behaviors such as goal seeking, being part of consistent groups without the need to define explicit relationships and wandering around seemingly without a specific purpose. Two fundamental concepts are important in achieving these results: (a) the per agent state representation and (b) the reward function. The agent state is a temporal representation of the situation around each agent. The reward function is based on the idea that people try to move in situations\/states in which they feel comfortable in. Therefore, in order for agents to stay in a comfortable state space, we first obtain a distribution of states extracted from real crowd data; then we evaluate states based on how much of an outlier they are compared to such a distribution. We demonstrate that our system can capture and simulate many complex and subtle crowd interactions in varied scenarios. Additionally, the proposed method generalizes to unseen situations, generates consistent behaviors and does not suffer from the limitations of other data-driven and reinforcement learning approaches.<\/jats:p>","DOI":"10.1145\/3592459","type":"journal-article","created":{"date-parts":[[2023,7,26]],"date-time":"2023-07-26T14:29:21Z","timestamp":1690381761000},"page":"1-15","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":46,"title":["GREIL-Crowds: Crowd Simulation with Deep Reinforcement Learning and Examples"],"prefix":"10.1145","volume":"42","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7230-5132","authenticated-orcid":false,"given":"Panayiotis","family":"Charalambous","sequence":"first","affiliation":[{"name":"CYENS - Centre of Excellence, Nicosia, Cyprus"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1812-1436","authenticated-orcid":false,"given":"Julien","family":"Pettre","sequence":"additional","affiliation":[{"name":"Univ Rennes, Inria, CNRS, IRISA, Rennes, France"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1336-5629","authenticated-orcid":false,"given":"Vassilis","family":"Vassiliades","sequence":"additional","affiliation":[{"name":"CYENS - Centre of Excellence, Nicosia, Cyprus"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5136-8890","authenticated-orcid":false,"given":"Yiorgos","family":"Chrysanthou","sequence":"additional","affiliation":[{"name":"CYENS - Centre of Excellence, Nicosia, Cyprus"},{"name":"University of Cyprus, Nicosia, Cyprus"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1437-245X","authenticated-orcid":false,"given":"Nuria","family":"Pelechano","sequence":"additional","affiliation":[{"name":"Universitat Politecnica de Catalunya (UPC), Barcelona, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,7,26]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1015330.1015430"},{"key":"e_1_2_2_2_1","doi-asserted-by":"crossref","unstructured":"Alexandre Alahi Kratarth Goel Vignesh Ramanathan Alexandre Robicquet Li Fei-Fei and Silvio Savarese. 2016. Social lstm: Human trajectory prediction in crowded spaces. (2016) 961--971.","DOI":"10.1109\/CVPR.2016.110"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2019.00359"},{"key":"e_1_2_2_4_1","unstructured":"Marc G Bellemare Will Dabney and R\u00e9mi Munos. 2017. A distributional perspective on reinforcement learning. (2017) 449--458."},{"key":"e_1_2_2_5_1","series-title":"SIAM review 53, 3","volume-title":"On the modeling of traffic and crowds: A survey of models, speculations, and perspectives","author":"Bellomo Nicola","year":"2011","unstructured":"Nicola Bellomo and Christian Dogbe. 2011. On the modeling of traffic and crowds: A survey of models, speculations, and perspectives. SIAM review 53, 3 (2011), 409--463."},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.12403"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.12472"},{"key":"e_1_2_2_8_1","first-page":"4","article-title":"Crowd motion capture","volume":"18","author":"N. Courty and T. Corp","year":"2007","unstructured":"N. Courty and T. Corpetti. 2007. Crowd motion capture. Computer Animation and Virtual Worlds 18, 4--5 (2007), 361--370.","journal-title":"Computer Animation and Virtual Worlds"},{"key":"e_1_2_2_9_1","volume-title":"International Conference on Machine Learning. 49--58","author":"Finn Chelsea","year":"2016","unstructured":"Chelsea Finn, Sergey Levine, and Pieter Abbeel. 2016. Guided cost learning: Deep inverse optimal control via policy optimization. In International Conference on Machine Learning. 49--58."},{"key":"e_1_2_2_10_1","volume-title":"Proceedings of the 2015 International Conference on Autonomous Agents and Multiagent Systems. International Foundation for Autonomous Agents and Multiagent Systems, 1577--1585","author":"Godoy Julio E","year":"2015","unstructured":"Julio E Godoy, Ioannis Karamouzas, Stephen J Guy, and Maria Gini. 2015. Adaptive learning for multi-agent navigation. In Proceedings of the 2015 International Conference on Autonomous Agents and Multiagent Systems. International Foundation for Autonomous Agents and Multiagent Systems, 1577--1585."},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00240"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2366145.2366209"},{"key":"e_1_2_2_13_1","volume-title":"International conference on machine learning. PMLR, 1352--1361","author":"Haarnoja Tuomas","year":"2017","unstructured":"Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine. 2017. Reinforcement learning with deep energy-based policies. In International conference on machine learning. PMLR, 1352--1361."},{"key":"e_1_2_2_14_1","volume-title":"A system for the notation of proxemic behavior. American anthropologist 65, 5","author":"Hall Edward T","year":"1963","unstructured":"Edward T Hall. 1963. A system for the notation of proxemic behavior. American anthropologist 65, 5 (1963), 1003--1026."},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevE.51.4282"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/ROBOT.2010.5509772"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11796"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1093\/beheco\/arq149"},{"key":"e_1_2_2_19_1","volume-title":"Heterogeneous crowd simulation using parametric reinforcement learning","author":"Hu Kaidong","year":"2021","unstructured":"Kaidong Hu, Brandon Haworth, Glen Berseth, Vladimir Pavlovic, Petros Faloutsos, and Mubbasir Kapadia. 2021. Heterogeneous crowd simulation using parametric reinforcement learning. IEEE Transactions on Visualization and Computer Graphics (2021)."},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/1882261.1866162"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2019406.2019414"},{"key":"e_1_2_2_22_1","volume-title":"ACM Transactions on Graphics (TOG)","volume":"27","author":"Kwon T.","unstructured":"T. Kwon, K.H. Lee, J. Lee, and S. Takahashi. 2008. Group motion editing. In ACM Transactions on Graphics (TOG), Vol. 27. ACM, 80."},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1073368.1073409"},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3274247.3274510"},{"key":"e_1_2_2_25_1","volume-title":"Proceedings of the 2007 ACM SIGGRAPH\/Eurographics Symposium on Computer Animation","author":"Lee Kang Hoon","year":"2007","unstructured":"Kang Hoon Lee, Myung Geol Choi, Qyoun Hong, and Jehee Lee. 2007. Group Behavior from Video: A Data-driven Approach to Crowd Simulation. In Proceedings of the 2007 ACM SIGGRAPH\/Eurographics Symposium on Computer Animation (San Diego, California) (SCA '07). Eurographics Association, Aire-la-Ville, Switzerland, Switzerland, 109--118. http:\/\/dl.acm.org\/citation.cfm?id=1272690.1272706"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.1467-8659.2007.01089.x"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.1467-8659.2010.01808.x"},{"key":"e_1_2_2_28_1","volume-title":"Proceedings of the 30th International Conference on Machine Learning (ICML-13)","author":"Levine Sergey","year":"2013","unstructured":"Sergey Levine and Vladlen Koltun. 2013. Guided policy search. In Proceedings of the 30th International Conference on Machine Learning (ICML-13). 1--9."},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/2422356.2422385"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3083723"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2018.8461113"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01495"},{"key":"e_1_2_2_33_1","volume-title":"Multi-Agent Reinforcement Learning for Simulating Pedestrian Navigation. In International Workshop on Adaptive and Learning Agents. 53","author":"Martinez-Gil Francisco","year":"2011","unstructured":"Francisco Martinez-Gil, Miguel Lozano, and Fernando Fern\u00e1ndez. 2011. Multi-Agent Reinforcement Learning for Simulating Pedestrian Navigation. In International Workshop on Adaptive and Learning Agents. 53."},{"key":"e_1_2_2_34_1","volume-title":"CASA '03: Proceedings of the 16th International Conference on Computer Animation and Social Agents (CASA","author":"Ronald","year":"2003","unstructured":"Ronald A. Metoyer and Jessica K. Hodgins. 2003. Reactive Pedestrian Path Following from Examples. In CASA '03: Proceedings of the 16th International Conference on Computer Animation and Social Agents (CASA 2003). IEEE Computer Society, Washington, DC, USA, 149."},{"key":"e_1_2_2_35_1","volume-title":"International Conference on Machine Learning. 1928--1937","author":"Mnih Volodymyr","year":"2016","unstructured":"Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016. Asynchronous methods for deep reinforcement learning. In International Conference on Machine Learning. 1928--1937."},{"key":"e_1_2_2_36_1","doi-asserted-by":"crossref","unstructured":"Volodymyr Mnih Koray Kavukcuoglu David Silver Andrei A Rusu Joel Veness Marc G Bellemare Alex Graves Martin Riedmiller Andreas K Fidjeland Georg Ostrovski et al. 2015. Human-level control through deep reinforcement learning. Nature 518 7540 (2015) 529--533.","DOI":"10.1038\/nature14236"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0010047"},{"key":"e_1_2_2_38_1","volume-title":"Simulating the Motion of Virtual Agents Based on Examples. In ACM\/EG Symposium on Computer Animation, Short Papers","author":"Musse S. R.","unstructured":"S. R. Musse, C. R. Jung, A. Braun, and J. J. Junior. 2006. Simulating the Motion of Virtual Agents Based on Examples. In ACM\/EG Symposium on Computer Animation, Short Papers. Vienna, Austria."},{"key":"e_1_2_2_39_1","volume-title":"CCP: Configurable Crowd Profiles. In ACM SIGGRAPH 2022 Conference Proceedings. 1--10","author":"Panayiotou Andreas","year":"2022","unstructured":"Andreas Panayiotou, Theodoros Kyriakou, Marilena Lemonari, Yiorgos Chrysanthou, and Panayiotis Charalambous. 2022. CCP: Configurable Crowd Profiles. In ACM SIGGRAPH 2022 Conference Proceedings. 1--10."},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.1467-8659.2007.01090.x"},{"key":"e_1_2_2_41_1","volume-title":"Simulating heterogeneous crowds with interactive behaviors","author":"Pelechano Nuria","unstructured":"Nuria Pelechano, Jan M Allbeck, Mubbasir Kapadia, and Norman I Badler. 2016. Simulating heterogeneous crowds with interactive behaviors. CRC Press."},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201311"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073602"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/1599470.1599495"},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58523-5_40"},{"key":"e_1_2_2_46_1","volume-title":"Prioritized experience replay. arXiv preprint arXiv:1511.05952","author":"Schaul Tom","year":"2015","unstructured":"Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver. 2015. Prioritized experience replay. arXiv preprint arXiv:1511.05952 (2015)."},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1137\/050627113"},{"key":"e_1_2_2_48_1","volume-title":"Reinforcement learning: An introduction","author":"Sutton Richard S","unstructured":"Richard S Sutton and Andrew G Barto. 2018. Reinforcement learning: An introduction. MIT press."},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/1276377.1276386"},{"key":"e_1_2_2_50_1","doi-asserted-by":"crossref","unstructured":"B. van Basten S. Jansen and I. Karamouzas. 2009. Exploiting motion capture to enhance avoidance behaviour in games. Motion in Games (2009) 29--40.","DOI":"10.1007\/978-3-642-10347-6_3"},{"key":"e_1_2_2_51_1","doi-asserted-by":"crossref","unstructured":"Hado van Hasselt Arthur Guez and David Silver. 2016. Deep Reinforcement Learning with Double Q-Learning.","DOI":"10.1609\/aaai.v30i1.10295"},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2016.2642963"},{"key":"e_1_2_2_53_1","volume-title":"International conference on machine learning. PMLR","author":"Wang Ziyu","year":"2016","unstructured":"Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Hasselt, Marc Lanctot, and Nando Freitas. 2016. Dueling network architectures for deep reinforcement learning. In International conference on machine learning. PMLR, 1995--2003."},{"key":"e_1_2_2_54_1","volume-title":"Machine learning 8, 3--4","author":"Watkins Christopher JCH","year":"1992","unstructured":"Christopher JCH Watkins and Peter Dayan. 1992. Q-learning. Machine learning 8, 3--4 (1992), 279--292."},{"key":"e_1_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.12328"},{"key":"e_1_2_2_56_1","volume-title":"Maximum entropy deep inverse reinforcement learning. arXiv preprint arXiv:1507.04888","author":"Wulfmeier Markus","year":"2015","unstructured":"Markus Wulfmeier, Peter Ondruska, and Ingmar Posner. 2015. Maximum entropy deep inverse reinforcement learning. arXiv preprint arXiv:1507.04888 (2015)."},{"key":"e_1_2_2_57_1","volume-title":"CLUST: Simulating Realistic Crowd Behaviour by Mining Pattern from Crowd Videos. In Computer Graphics Forum","author":"Zhao M","year":"2017","unstructured":"M Zhao, W Cai, and SJ Turner. 2017. CLUST: Simulating Realistic Crowd Behaviour by Mining Pattern from Crowd Videos. In Computer Graphics Forum. Wiley Online Library."},{"key":"e_1_2_2_58_1","unstructured":"M. Zhao and V. Saligrama. 2009. Anomaly detection with score functions based on nearest neighbor graphs. In Advances in Neural Information Processing Systems."},{"key":"e_1_2_2_59_1","volume-title":"Dey","author":"Ziebart Brian D.","year":"2008","unstructured":"Brian D. Ziebart, Andrew Maas, J. Andrew Bagnell, and Anind K. Dey. 2008. Maximum Entropy Inverse Reinforcement Learning. In Proc. AAAI. 1433--1438."}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3592459","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3592459","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:48:59Z","timestamp":1750182539000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3592459"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,26]]},"references-count":59,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2023,8]]}},"alternative-id":["10.1145\/3592459"],"URL":"https:\/\/doi.org\/10.1145\/3592459","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,26]]},"assertion":[{"value":"2023-07-26","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}