{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,23]],"date-time":"2026-04-23T10:05:11Z","timestamp":1776938711554,"version":"3.51.4"},"reference-count":60,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2014,2,20]],"date-time":"2014-02-20T00:00:00Z","timestamp":1392854400000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Auton Agent Multi-Agent Syst"],"published-print":{"date-parts":[[2015,1]]},"DOI":"10.1007\/s10458-014-9252-6","type":"journal-article","created":{"date-parts":[[2014,2,19]],"date-time":"2014-02-19T07:50:13Z","timestamp":1392796213000},"page":"98-130","source":"Crossref","is-referenced-by-count":26,"title":["Strategies for simulating pedestrian navigation with multiple reinforcement learning agents"],"prefix":"10.1007","volume":"29","author":[{"given":"Francisco","family":"Martinez-Gil","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Miguel","family":"Lozano","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fernando","family":"Fern\u00e1ndez","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2014,2,20]]},"reference":[{"key":"9252_CR1","unstructured":"Agre, P. & Chapman, D. (1987). Pengi: An implementation of a theory of activity. In: Proceedings of the Sixth National Conference on Artificial Intelligence, (pp. 268\u2013272). Burlington: Morgan Kaufmann"},{"key":"9252_CR2","doi-asserted-by":"crossref","first-page":"621","DOI":"10.1177\/0037549709340659","volume":"85","author":"B Banerjee","year":"2009","unstructured":"Banerjee, B., Abukmail, A., & Kraemer, L. (2009). Layered intelligence for agent-based crowd simulation. Simulation, 85, 621\u2013632.","journal-title":"Simulation"},{"key":"9252_CR3","unstructured":"van den Berg, J., Lin, M. & Manocha, D. (2008). Reciprocal velocity obstales for real-time multi-agent navigator. In: Proceedings of the IEEE International Conference on Robotics and Automation (pp. 1928\u20131935)."},{"key":"9252_CR4","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1108\/9781848557512-001","volume-title":"Pedestrian Behavior","author":"M Bierlaire","year":"2009","unstructured":"Bierlaire, M., & Robin, T. (2009). Pedestrians choices. In H. Timmermans (Ed.), Pedestrian Behavior (pp. 1\u201326). Bradford: Emerald."},{"issue":"1","key":"9252_CR5","doi-asserted-by":"crossref","first-page":"52","DOI":"10.1007\/s10458-012-9201-1","volume":"27","author":"T Bosse","year":"2013","unstructured":"Bosse, T., Hoogendoorn, M., Klein, M. C. A., Treur, J., van der Wal, C. N., & van Wissen, A. (2013). Modelling collective decision making in groups and crowds: Integrating social contagion and interacting emotions, beliefs and intentions. Autonomous Agents and Multi-Agent Systems, 27(1), 52\u201384.","journal-title":"Autonomous Agents and Multi-Agent Systems"},{"key":"9252_CR6","unstructured":"Campanella, M., Hoogendoorn, S., Daamen, W. (2010). Calibrating walker models: A methodology and applications. In: Proceedings of the 12th World Conference on Transport Research WCTR 2010. Lisbon: 12th WCTR Comitee."},{"key":"9252_CR7","unstructured":"Claus, C. & Boutilier, C. (1998). The dynamics of reinforcement learning in cooperative multiagent systems. In: Proceedings of the Fifteenth National Conference on Artificial Intelligence (pp. 746\u2013752). Menlo Park: AAAI Press."},{"key":"9252_CR8","unstructured":"Daamen, W. & Hoogendoorn, S. (2003). Experimental research of pedestrian walking behavior. In: Transportation Research Board Annual Meeting 2003, (pp. 1\u201316). Washington: National Academy Press."},{"issue":"2","key":"9252_CR9","doi-asserted-by":"crossref","first-page":"213","DOI":"10.1002\/int.20255","volume":"23","author":"F Fern\u00e1ndez","year":"2008","unstructured":"Fern\u00e1ndez, F., & Borrajo, D. (2008). Two steps reinforcement learning. International Journal of Intelligent Systems, 23(2), 213\u2013245.","journal-title":"International Journal of Intelligent Systems"},{"issue":"2\u20134","key":"9252_CR10","doi-asserted-by":"crossref","first-page":"161","DOI":"10.1007\/s10846-005-5137-x","volume":"43","author":"F Fern\u00e1ndez","year":"2005","unstructured":"Fern\u00e1ndez, F., Borrajo, D., & Parker, L. (2005). A reinforcement learning algorithm in cooperative multi-robot domains. Journal of Intelligent Robotics Systems, 43(2\u20134), 161\u2013174.","journal-title":"Journal of Intelligent Robotics Systems"},{"key":"9252_CR11","doi-asserted-by":"crossref","unstructured":"Fern\u00e1ndez, F., Garc\u00eda, J., & Veloso, M. (2010). Probabilistic policy reuse for inter-task transfer learning. Robotics and Autonomous Systems, 58(7), 866\u2013871.","DOI":"10.1016\/j.robot.2010.03.007"},{"issue":"7","key":"9252_CR12","first-page":"866","volume":"58","author":"JG Fernando Fern\u00e1ndez","year":"2010","unstructured":"Fernando Fern\u00e1ndez, J. G., & Veloso, M. (2010). Probabilistic policy reuse for inter-task transfer learning. Robotics and Autonomous Systems. Special Issue on Advances in Autonomous Robots for Service and Entertainment, 58(7), 866\u2013871.","journal-title":"Special Issue on Advances in Autonomous Robots for Service and Entertainment"},{"key":"9252_CR13","unstructured":"Fruin, J. (1971). Pedestrian and planning design. Tech. rep., Metropolitan Association of Urban Designers and Environmental Planners. New York, Library of congress catalogue number 70\u2013159312."},{"key":"9252_CR14","unstructured":"Garc\u00eda, J., L\u00f3pez-Bueno, I., Fern\u00e1ndez, F. & Borrajo, D. (2010). A Comparative Study of Discretization Approaches for State Space Generalization in the Keepaway Soccer Task. In: Reinforcement Learning: Algorithms, Implementations and Aplications. Hauppauge: Nova Science Publishers."},{"key":"9252_CR15","doi-asserted-by":"crossref","first-page":"95","DOI":"10.1016\/0378-4754(85)90027-8","volume":"27","author":"P Gipps","year":"1985","unstructured":"Gipps, P., & Marsjo, B. (1985). A microsimulation model for pedestrian flows. Mathematics and Computers in Simulation, 27, 95\u2013105.","journal-title":"Mathematics and Computers in Simulation"},{"issue":"2","key":"9252_CR16","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1109\/MASSP.1984.1162229","volume":"1","author":"RM Gray","year":"1984","unstructured":"Gray, R. M. (1984). Vector quantization. IEEE ASSP Magazine, 1(2), 4\u201329.","journal-title":"IEEE ASSP Magazine"},{"key":"9252_CR17","doi-asserted-by":"crossref","first-page":"180","DOI":"10.1016\/j.commatsci.2004.01.026","volume":"30","author":"D Helbing","year":"2004","unstructured":"Helbing, D. (2004). Collective phenomena and states in traffic and self-driven many-particle systems. Computational Materials Science, 30, 180\u2013187.","journal-title":"Computational Materials Science"},{"issue":"1","key":"9252_CR18","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1287\/trsc.1040.0108","volume":"39","author":"D Helbing","year":"2005","unstructured":"Helbing, D., Buzna, L., Johansson, A., & Werner, T. (2005). Self-organized pedestrian crowd dynamics: Experiments, simulations, and design solutions. Transportation Science, 39(1), 1\u201324.","journal-title":"Transportation Science"},{"key":"9252_CR19","doi-asserted-by":"crossref","first-page":"487","DOI":"10.1038\/35035023","volume":"407","author":"D Helbing","year":"2000","unstructured":"Helbing, D., Farkas, I., & Vicsek, T. (2000). Simulating dynamical features of escape panic. Nature, 407, 487.","journal-title":"Nature"},{"key":"9252_CR20","unstructured":"Helbing, D. & Johansson, A. (2009). Pedestrian, Crowd and Evacuation Dynamics. Encyclopedia of Complexity and Systems Science, Part 16. (pp. 6476\u20136495). New York: Springer. ."},{"key":"9252_CR21","doi-asserted-by":"crossref","first-page":"046109","DOI":"10.1103\/PhysRevE.75.046109","volume":"75","author":"D Helbing","year":"2007","unstructured":"Helbing, D., Johansson, A., & Al-Abideen, H. Z. (2007). Dynamics of crowd disasters: An empirical study. Physical Review E, 75, 046109.","journal-title":"Physical Review E"},{"key":"9252_CR22","doi-asserted-by":"crossref","first-page":"4282","DOI":"10.1103\/PhysRevE.51.4282","volume":"51","author":"D Helbing","year":"1995","unstructured":"Helbing, D., & Moln\u00e1r, P. (1995). Social force model for pedestrian dynamics. Physics Review E, 51, 4282\u20134286.","journal-title":"Physics Review E"},{"key":"9252_CR23","doi-asserted-by":"crossref","first-page":"361","DOI":"10.1068\/b2697","volume":"28","author":"D Helbing","year":"2001","unstructured":"Helbing, D., Moln\u00e1r, P., Farkas, I., & Bolay, K. (2001). Self-organizing pedestrian movement. Environment and Planning. Part B. Planning and Design, 28, 361\u2013383.","journal-title":"Planning and Design"},{"issue":"5","key":"9252_CR24","doi-asserted-by":"crossref","first-page":"935","DOI":"10.1142\/S0219622012500277","volume":"11","author":"FB Javier Garc\u00eda","year":"2012","unstructured":"Javier Garc\u00eda, F. B., & Fern\u00e1ndez, F. (2012). Reinforcement learning for decision-making in a business simulator. International Journal of Information Technology & Decision Making, 11(5), 935\u2013960.","journal-title":"International Journal of Information Technology & Decision Making"},{"key":"9252_CR25","doi-asserted-by":"crossref","first-page":"394","DOI":"10.1109\/TVCG.2011.133","volume":"18","author":"I Karamouzas","year":"2012","unstructured":"Karamouzas, I., & Overmars, M. (2012). Simulating and evaluating the local behavior of small pedestrian groups. IEEE Transactions on Visualization and Computer Graphics, 18, 394\u2013406.","journal-title":"IEEE Transactions on Visualization and Computer Graphics"},{"key":"9252_CR26","unstructured":"Klein, F., Bourjot, C. & Chevrier, V. (2009). Application of reinforcement learning to control a multiagent system. In: International Conference on Agents and Artificial Intelligence. Berlin: Springer."},{"key":"9252_CR27","doi-asserted-by":"crossref","unstructured":"Lane, T., Ridens, M., Stevens, S. (2007). Reinforcement learning in nonstationary environment navigation tasks. In: Advances in Artificial Intelligence (LNCS 4509), pp. 429\u2013440. Berlin: Springer.","DOI":"10.1007\/978-3-540-72665-4_37"},{"issue":"1","key":"9252_CR28","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1109\/TCOM.1980.1094577","volume":"28","author":"Y Linde","year":"1980","unstructured":"Linde, Y., Buzo, A., & Gray, R. (1980). An algorithm for vector quantizer design. IEEE Transactions on Communications, 28(1), 84\u201395.","journal-title":"IEEE Transactions on Communications"},{"key":"9252_CR29","unstructured":"Littman, M.L. (2005). Markov games as a framework for multi-agent reinforcement learning. In: Proceedings of the Eleventh International Conference on Machine Learning (pp. 157\u2013163). New Brunswick: Morgan Kaufmann."},{"key":"9252_CR30","doi-asserted-by":"crossref","first-page":"429","DOI":"10.1016\/0191-2615(94)90013-2","volume":"28B","author":"G Lovas","year":"1994","unstructured":"Lovas, G. (1994). Modelling and simulation of pedestrian traffic flow. Transportation Research, 28B, 429\u2013443.","journal-title":"Transportation Research"},{"key":"9252_CR31","unstructured":"Martinez-Gil, F., Barber, F., Lozano, M., Grimaldo, F., Fern\u00e1ndez, F. (2010). A reinforcement learning approach for multiagent navigation. In: ICAART 2010\u2014Proceedings of the International Conferencenon Agents and Artificial Intelligence, Volume 1 (pp. 607\u2013610). Artificial Intelligence: Valencia, January 22\u201324, 2010."},{"key":"9252_CR32","doi-asserted-by":"crossref","unstructured":"Martinez-Gil, F., Lozano, M. & Fern\u00e1ndez, F. (2012). Calibrating a motion model based on reinforcement learning for pedestrian simulation. In: Motion in Games - 5th International Conference, MIG 2012, Rennes, France, November 15\u201317, 2012. Proceedings, Lecture Notes in Computer Science, vol. 7660, pp. 302\u2013313. Springer.","DOI":"10.1007\/978-3-642-34710-8_28"},{"key":"9252_CR33","doi-asserted-by":"crossref","unstructured":"Martinez-Gil, F., Lozano, M. & Fern\u00e1ndez, F. (2012). Multi-agent reinforcement learning for simulating pedestrian navigation. In: Adaptive and Learning Agents - International Workshop, ALA 2011, Held at AAMAS 2011, Taipei, Taiwan, May 2, 2011, Revised Selected Papers, Lecture Notes in Computer Science, vol. 7113, pp. 54\u201369. Springer.","DOI":"10.1007\/978-3-642-28499-1_4"},{"key":"9252_CR34","unstructured":"Mataric, M. J. (1994). Learning to behave socially. In: From Animals to Animats: International Conference on Simulation of Adaptive Behavior (pp. 453\u2013462). Cambridge: MIT Press."},{"key":"9252_CR35","unstructured":"Pelechano, N., Allbeck, J. & Badler, N. (2007). Controlling individual agents in high-density crowd simulation. In: Proc. ACM\/SIGGRAPH\/Eurographycs Symp. Computer Animation, pp. 99\u2013108."},{"key":"9252_CR36","unstructured":"Pettr\u00e9, J., Ondrej, J., Olivier, A., Cr\u00e9tual, A., Donikian, S. (2009). Experiment-based modeling.simulation and validation of interactions between virtual walkers. In: Proceedings of the Symposium on Computer, Animation SCA\u201909 (pp. 189\u2013198)."},{"key":"9252_CR37","unstructured":"Reynolds, C. (2003). Evolution of corridor following behavior in a noisy world. In: From animals to animats. Proceedings of the third international conference on simulation of adaptive behavior. Cambridge: MIT Press."},{"key":"9252_CR38","first-page":"9","volume":"3","author":"G Rindsf\u00fcser","year":"2007","unstructured":"Rindsf\u00fcser, G., & Kl\u00fcgl, F. (2007). Agent-based pedestrian simulation: A case study of the Bern railway station. disP, 3, 9\u201318.","journal-title":"disP"},{"key":"9252_CR39","doi-asserted-by":"crossref","first-page":"36","DOI":"10.1016\/j.trb.2008.06.010","volume":"43","author":"T Robin","year":"2009","unstructured":"Robin, T., Antonioni, G., Bierlaire, M., & Cruz, J. (2009). Specification, estimation and validation of a pedestrian walking behavior model. Transportation Research, 43, 36\u201356.","journal-title":"Transportation Research"},{"key":"9252_CR40","first-page":"3142","volume-title":"Encyclopedia of Complexity and Systems Science","author":"A Schadschneider","year":"2008","unstructured":"Schadschneider, A., Klingsch, W., Kluepfel, H., Kretz, T., Rogsch, C., & Seyfried, A. (2008). Evacuation dynamics: empirical results, modelling and applications. In R. A. Meyers (Ed.), Encyclopedia of Complexity and Systems Science (pp. 3142\u20133176). Heidelberg: Springer."},{"key":"9252_CR41","doi-asserted-by":"crossref","first-page":"545","DOI":"10.3934\/nhm.2011.6.545","volume":"6","author":"A Schadschneider","year":"2011","unstructured":"Schadschneider, A., & Syfried, A. (2011). Empirical results for pedestrian dynamics and their implications for modeling. Networks and Heterogeneous Media, 6, 545\u2013560.","journal-title":"Networks and Heterogeneous Media"},{"key":"9252_CR42","unstructured":"Sen, S. & Sekaran, M. (1996). Multiagent coordination with learning classifier systems. In: IJCAI95 Workshop on Adaptation and Learning in Multiagent Systems (pp. 218\u2013233). Berlin: Springer."},{"key":"9252_CR43","unstructured":"Seyfried, A., Steffen, B., Klingsch, W. & Boltes, M. (2005). The fundamental diagram of pedestrian movement revisited. Journal of Statistical Mechanics: Theory and Experiment, p. P10002."},{"key":"9252_CR44","doi-asserted-by":"crossref","DOI":"10.4271\/2005-01-2699","volume-title":"Autonomous pedestrians. In: Proceedings of the 2005 ACM SIGGRAPH symposium on Computer animation","author":"W Shao","year":"2005","unstructured":"Shao, W., & Terzopoulos, D. (2005). Autonomous pedestrians. In: Proceedings of the 2005 ACM SIGGRAPH symposium on Computer animation. New York: ACM Press."},{"key":"9252_CR45","unstructured":"Steiner, A., Philipp, M. & Schmid, A. (2007). Parameter stimation for pedestrian simulation model. In: Proc. 7th Swiss Transport Research Conference (pp. 1\u201329)."},{"key":"9252_CR46","unstructured":"Still, K. (2000). Crowd dynamics. Ph.D. thesis, Department of Mathematics. Warwick University, UK."},{"issue":"3","key":"9252_CR47","doi-asserted-by":"crossref","first-page":"165","DOI":"10.1177\/105971230501300301","volume":"13","author":"P Stone","year":"2005","unstructured":"Stone, P., Sutton, R. S., & Kuhlmann, G. (2005). Reinforcement learning for RoboCup-soccer keepaway. Adaptive Behavior, 13(3), 165\u2013188.","journal-title":"Adaptive Behavior"},{"key":"9252_CR48","volume-title":"Reinforcement Learning: An Introduction","author":"RS Sutton","year":"1998","unstructured":"Sutton, R. S., & Barto, A. G. (1998). Reinforcement Learning: An Introduction. Cambridge: MIT Press."},{"key":"9252_CR49","doi-asserted-by":"crossref","first-page":"343","DOI":"10.1002\/cav.105","volume":"16","author":"T Sakuma","year":"2005","unstructured":"Sakuma, T., & Mukai, S. K. (2005). Psychological model for animating crowded pedestrians: virtual humans and social agents. Computer animation virtual worlds, 16, 343\u2013351.","journal-title":"Computer animation virtual worlds"},{"key":"9252_CR50","unstructured":"Taylor, M. & Stone, P. (2007). Representation transfer in reinforcement learning. In: AAAI 2007 Fall Symposium on Computational Approacher to Representation Change during Learning and Development."},{"key":"9252_CR51","first-page":"1633","volume":"10","author":"M Taylor","year":"2009","unstructured":"Taylor, M., & Stone, P. (2009). Transfer learning for reinforcement learning domains: a survey. Journal of Machine Learning Research, 10, 1633\u20131685.","journal-title":"Journal of Machine Learning Research"},{"key":"9252_CR52","unstructured":"Taylor, M.E., Suay, H.B. & Chernova, S. (2011). Integrating reinforcement learning with human demonstrations of varying ability. In: Proceedings International Conference on Autonomous Agents and Multiagent Systems."},{"key":"9252_CR53","unstructured":"Teknomo, K. (2002). Microscopic pedestrian flow characteristics: Development of an image processing data collection and simulation model. Ph.D. thesis, Department of Human Social Information Sciencies. Tohoku University, Japan."},{"key":"9252_CR54","unstructured":"Thesauro, G. & Kephart, J. (2002). Pricing in agent economies using multi-agent q-learning. In: International Conference on Autonomous Agents and Multiagents Systems (AAMAS\u201902)."},{"key":"9252_CR55","unstructured":"Torrey, L. (2010). Crowd simulation via multi-agent reinforcement learning. In: Proceedings of the Sixth AAAI Conference On Artificial Intelligence and Interactive Digital Entertainment. Menlo Park: AAAI Press."},{"key":"9252_CR56","unstructured":"Torrey, L. & Taylor, M.E. (2012). Help an agent out: Student\/teacher learning in sequential decision tasks. In: Proceedings of the Adaptive and Learning Agents workshop (at AAMAS-12)."},{"issue":"1","key":"9252_CR57","doi-asserted-by":"crossref","first-page":"225","DOI":"10.1016\/j.asoc.2009.07.004","volume":"10","author":"G Vigueras","year":"2010","unstructured":"Vigueras, G., Lozano, M., Ordu\u00f1a, J. M., & Grimaldo, F. (2010). A comparative study of partitioning methods for crowd simulations. Applied Soft Computing, 10(1), 225\u2013235.","journal-title":"Applied Soft Computing"},{"key":"9252_CR58","first-page":"279","volume":"8","author":"C Watkins","year":"1992","unstructured":"Watkins, C., & Dayan, P. (1992). Q-learning. Machine Learning, 8, 279\u2013292.","journal-title":"Machine Learning"},{"key":"9252_CR59","unstructured":"Weidmann, U. (1993). Transporttechnik der fussg\u00e4nger - transporttechnische eigenschaften des fussgngerverkehrs (literaturstudie). Literature Research 90, Institut f\u00fcer Verkehrsplanung, Transporttechnik, Strassen- undEisenbahnbau IVT an der ETH Z\u00fcrich, ETH-H\u00f6nggerberg, CH-8093 Z\u00fcrich."},{"key":"9252_CR60","doi-asserted-by":"crossref","unstructured":"Whitehead, S.D. & Ballard, D.H. (1991). Learning to perceive and act by trial and error. Machine Learning pp. 45\u201383.","DOI":"10.1007\/BF00058926"}],"container-title":["Autonomous Agents and Multi-Agent Systems"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s10458-014-9252-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s10458-014-9252-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s10458-014-9252-6","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,8,7]],"date-time":"2019-08-07T19:20:02Z","timestamp":1565205602000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s10458-014-9252-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,2,20]]},"references-count":60,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2015,1]]}},"alternative-id":["9252"],"URL":"https:\/\/doi.org\/10.1007\/s10458-014-9252-6","relation":{},"ISSN":["1387-2532","1573-7454"],"issn-type":[{"value":"1387-2532","type":"print"},{"value":"1573-7454","type":"electronic"}],"subject":[],"published":{"date-parts":[[2014,2,20]]}}}