{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,7,30]],"date-time":"2025-07-30T16:30:56Z","timestamp":1753893056432,"version":"3.41.2"},"reference-count":48,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2023,1,25]],"date-time":"2023-01-25T00:00:00Z","timestamp":1674604800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100007601","name":"Horizon 2020","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100007601","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Artif. Intell."],"abstract":"<jats:p>Learning from only real-world collected data can be unrealistic and time consuming in many scenario. One alternative is to use synthetic data as learning environments to learn rare situations and replay buffers to speed up the learning. In this work, we examine the hypothesis of how the creation of the environment affects the training of reinforcement learning agent through auto-generated environment mechanisms. We take the autonomous vehicle as an application. We compare the effect of two approaches to generate training data for artificial cognitive agents. We consider the added value of curriculum learning\u2014just as in human learning\u2014as a way to structure novel training data that the agent has not seen before as well as that of using a replay buffer to train further on data the agent has seen before. In other words, the focus of this paper is on characteristics of the training data rather than on learning algorithms. We therefore use two tasks that are commonly trained early on in autonomous vehicle research: lane keeping and pedestrian avoidance. Our main results show that curriculum learning indeed offers an additional benefit over a vanilla reinforcement learning approach (using Deep-Q Learning), but the replay buffer actually has a detrimental effect in most (but not all) combinations of data generation approaches we considered here. The benefit of curriculum learning does depend on the existence of a well-defined difficulty metric with which various training scenarios can be ordered. In the lane-keeping task, we can define it as a function of the curvature of the road, in which the steeper and more occurring curves on the road, the more difficult it gets. Defining such a difficulty metric in other scenarios is not always trivial. In general, the results of this paper emphasize both the importance of considering data characterization, such as curriculum learning, and the importance of defining an appropriate metric for the task.<\/jats:p>","DOI":"10.3389\/frai.2023.1098982","type":"journal-article","created":{"date-parts":[[2023,1,25]],"date-time":"2023-01-25T09:27:56Z","timestamp":1674638876000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["How to train a self-driving vehicle: On the added value (or lack thereof) of curriculum learning and replay buffers"],"prefix":"10.3389","volume":"6","author":[{"given":"Sara","family":"Mahmoud","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Erik","family":"Billing","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Henrik","family":"Svensson","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Serge","family":"Thill","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2023,1,25]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","first-page":"19817","DOI":"10.1109\/TITS.2022.3160673","article-title":"An end-to-end curriculum learning approach for autonomous driving scenarios","volume":"23","author":"Anzalone","year":"2022","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"B2","doi-asserted-by":"publisher","first-page":"89249","DOI":"10.1109\/ACCESS.2021.3090907","article-title":"Curriculum learning for vehicle lateral stability estimations","volume":"9","author":"Bae","year":"2021","journal-title":"IEEE Access"},{"key":"B3","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1145\/1553374.1553380","article-title":"\u201cCurriculum learning,\u201d","author":"Bengio","year":"2009","journal-title":"Proceedings of the 26th Annual International Conference on Machine Learning"},{"key":"B4","article-title":"\u201cProgressive reinforcement learning with distillation for multi-skilled motion control,\u201d","author":"Berseth","year":"2018","journal-title":"International Conference on Learning Representations"},{"key":"B5","doi-asserted-by":"publisher","first-page":"9","DOI":"10.3389\/frobt.2016.00009","article-title":"Finding your way from the bed to the kitchen: reenacting and recombining sensorimotor episodes learned from human demonstration","volume":"3","author":"Billing","year":"2016","journal-title":"Front. Robot. AI"},{"key":"B6","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1604.07316","article-title":"End to end learning for self-driving cars","author":"Bojarski","year":"2016","journal-title":"arXiv"},{"key":"B7","doi-asserted-by":"publisher","first-page":"1","DOI":"10.4114\/intartif.vol24iss68pp1-20","article-title":"Evaluating the impact of curriculum learning on the training process for an intelligent agent in a video game","volume":"24","author":"Camargo","year":"2021","journal-title":"Intel. Artif"},{"key":"B8","doi-asserted-by":"publisher","first-page":"1856","DOI":"10.1109\/IVS.2017.7995975","article-title":"\u201cEnd-to-end learning for lane keeping of self-driving cars,\u201d","author":"Chen","year":"2017","journal-title":"2017 IEEE Intelligent Vehicles Symposium (IV)"},{"key":"B9","doi-asserted-by":"publisher","first-page":"9329","DOI":"10.1109\/ICCV.2019.00942","article-title":"\u201cExploring the limitations of behavior cloning for autonomous driving,\u201d","author":"Codevilla","year":"2019","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision"},{"year":"2017","author":"Da Lio","first-page":"1","key":"B10"},{"key":"B11","doi-asserted-by":"publisher","first-page":"71","DOI":"10.1016\/0010-0277(93)90058-4","article-title":"Learning and development in neural networks: the importance of starting small","volume":"48","author":"Elman","year":"1993","journal-title":"Cognition"},{"key":"B12","article-title":"\u201cCurriculum-guided hindsight experience replay,\u201d","author":"Fang","year":"2019","journal-title":"Advances in Neural Information Processing Systems, Vol. 32"},{"key":"B13","doi-asserted-by":"publisher","first-page":"1231","DOI":"10.1177\/0278364913491297","article-title":"Vision meets robotics: the KITTI dataset","volume":"32","author":"Geiger","year":"2013","journal-title":"Int. J. Robot. Res"},{"key":"B14","doi-asserted-by":"publisher","first-page":"101","DOI":"10.1146\/annurev-psych-122414-033625","article-title":"Reinforcement learning and episodic memory in humans and animals: an integrative framework","volume":"68","author":"Gershman","year":"2017","journal-title":"Ann. Rev. Psychol"},{"key":"B15","first-page":"2672","article-title":"\u201cGenerative adversarial nets,\u201d","author":"Goodfellow","year":"2014","journal-title":"Advances in Neural Information Processing Systems, Vol. 63"},{"key":"B16","doi-asserted-by":"publisher","first-page":"362","DOI":"10.1002\/rob.21918","article-title":"A survey of deep learning techniques for autonomous driving","volume":"37","author":"Grigorescu","year":"2020","journal-title":"J. Field Robot"},{"key":"B17","doi-asserted-by":"crossref","first-page":"3389","DOI":"10.1109\/ICRA.2017.7989385","article-title":"\u201cDeep reinforcement learning for robotic manipulation with asynchronous off-policy updates,\u201d","author":"Gu","year":"2017","journal-title":"Robotics and Automation (ICRA), 2017 IEEE International Conference"},{"key":"B18","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1803.10122","article-title":"World models","author":"Ha","year":"2018","journal-title":"arXiv"},{"key":"B19","doi-asserted-by":"publisher","first-page":"40","DOI":"10.1007\/978-3-030-35664-4_4","article-title":"\u201cAutonomous vehicle control: End-to-end learning in simulated urban environments,\u201d","author":"Haavaldsen","year":"2019","journal-title":"Symposium of the Norwegian AI Society"},{"key":"B20","first-page":"2535","article-title":"\u201cOn the power of curriculum learning in training deep networks,\u201d","author":"Hacohen","year":"2019","journal-title":"Proceedings of Machine Learning Research, Vol. 97"},{"key":"B21","doi-asserted-by":"publisher","first-page":"22","DOI":"10.1016\/j.neunet.2006.07.003","article-title":"Perception through visuomotor anticipation in a mobile robot","volume":"20","author":"Hoffmann","year":"2007","journal-title":"Neural Netw"},{"key":"B22","article-title":"\u201cDistributed prioritized experience replay,\u201d","author":"Horgan","year":"2018","journal-title":"International Conference on Learning Representations"},{"key":"B23","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1801.00904","article-title":"Screenernet: learning self-paced curriculum for deep neural networks","author":"Kim","year":"2018","journal-title":"arXiv"},{"key":"B24","doi-asserted-by":"publisher","first-page":"380","DOI":"10.1016\/j.cognition.2008.11.014","article-title":"Flexible shaping: how learning in small steps helps","volume":"110","author":"Krueger","year":"2009","journal-title":"Cognition"},{"key":"B25","first-page":"3675","article-title":"\u201cHierarchical deep reinforcement learning: integrating temporal abstraction and intrinsic motivation,\u201d","author":"Kulkarni","year":"2016","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B26","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1509.02971","article-title":"Continuous control with deep reinforcement learning","author":"Lillicrap","year":"2015","journal-title":"arXiv"},{"year":"1993","author":"Lin","journal-title":"Reinforcement Learning for Robots Using Neural Networks","key":"B27"},{"key":"B28","doi-asserted-by":"publisher","first-page":"63","DOI":"10.1016\/j.cogsys.2022.09.005","article-title":"Where to from here? On the future development of autonomous vehicles from a cognitive systems perspective","volume":"76","author":"Mahmoud","year":"2022","journal-title":"Cogn. Syst. Res"},{"key":"B29","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1312.5602","article-title":"Playing atari with deep reinforcement learning","author":"Mnih","year":"2013","journal-title":"arXiv"},{"key":"B30","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"B31","first-page":"1","article-title":"Curriculum learning for reinforcement learning domains: a framework and survey","volume":"21","author":"Narvekar","year":"2022","journal-title":"J. Mach. Learn. Res"},{"key":"B32","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2017\/353","article-title":"\u201cAutonomous task sequencing for customized curriculum design in reinforcement learning,\u201d","author":"Narvekar","year":"2017","journal-title":"The Twenty-Sixth International Joint Conference on Artificial Intelligence (IJCAI)"},{"key":"B33","first-page":"25","article-title":"\u201cLearning curriculum policies for reinforcement learning,\u201d","author":"Narvekar","year":"2019","journal-title":"Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems"},{"key":"B34","doi-asserted-by":"publisher","first-page":"54","DOI":"10.1016\/j.neunet.2019.01.012","article-title":"Continual lifelong learning with neural networks: a review","volume":"113","author":"Parisi","year":"2019","journal-title":"Neural Netw"},{"key":"B35","doi-asserted-by":"publisher","first-page":"877","DOI":"10.1017\/S0140525X00004015","article-title":"The reinterpretation of dreams: an evolutionary hypothesis of the function of dreaming","volume":"23","author":"Revonsuo","year":"2000","journal-title":"Behav. Brain Sci"},{"key":"B36","doi-asserted-by":"publisher","first-page":"70","DOI":"10.2352\/ISSN.2470-1173.2017.19.AVM-023","article-title":"Deep reinforcement learning framework for autonomous driving","volume":"2017","author":"Sallab","year":"2017","journal-title":"Electro. Imaging"},{"key":"B37","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1608.01230","article-title":"Learning a driving simulator","author":"Santana","year":"2016","journal-title":"arXiv"},{"key":"B38","article-title":"\u201cPrioritized experience replay,\u201d","author":"Schaul","year":"2016","journal-title":"International Conference on Learning Representations"},{"key":"B39","doi-asserted-by":"publisher","first-page":"160","DOI":"10.1145\/122344.122377","article-title":"Dyna, an integrated architecture for learning, planning, and reacting","volume":"2","author":"Sutton","year":"1991","journal-title":"ACM SIGART Bull"},{"year":"2018","author":"Sutton","journal-title":"Reinforcement Learning: An Introduction","key":"B40"},{"key":"B41","doi-asserted-by":"publisher","first-page":"222","DOI":"10.1177\/1059712313491295","article-title":"Dreaming of electric sheep? Exploring the functions of dream-like mechanisms in the development of mental imagery simulations","volume":"21","author":"Svensson","year":"2013","journal-title":"Adapt. Behav"},{"key":"B42","doi-asserted-by":"publisher","first-page":"1131","DOI":"10.1016\/S0893-6080(99)00060-X","article-title":"Learning to perceive the world as articulated: an approach for hierarchical learning in sensory-motor systems","volume":"12","author":"Tani","year":"1999","journal-title":"Neural Netw"},{"key":"B43","first-page":"2314","article-title":"\u201cA deeper look at planning as learning from replay,\u201d","author":"Vanseijen","year":"2015","journal-title":"International Conference on Machine Learning"},{"key":"B44","doi-asserted-by":"publisher","DOI":"10.1155\/2015\/250461","article-title":"An automatic traffic sign detection and recognition system based on colour segmentation, shape matching, and svm","author":"Wali","year":"2015","journal-title":"Math. Prob. Eng"},{"key":"B45","doi-asserted-by":"publisher","first-page":"267","DOI":"10.1177\/1059712319896489","article-title":"On the utility of dreaming: a general model for how learning in artificial agents can benefit from data hallucination","volume":"29","author":"Windridge","year":"2020","journal-title":"Adapt. Behav"},{"key":"B46","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1712.01275","article-title":"A deepe look at experience replay","author":"Zhang","year":"2017","journal-title":"arXiv"},{"key":"B47","doi-asserted-by":"publisher","first-page":"737","DOI":"10.1109\/SSCI47803.2020.9308468","article-title":"\u201cSim-to-real transfer in deep reinforcement learning for robotics: a survey,\u201d","author":"Zhao","year":"2020","journal-title":"2020 IEEE Symposium Series on Computational Intelligence (SSCI)"},{"key":"B48","doi-asserted-by":"publisher","first-page":"85","DOI":"10.1016\/j.neucom.2004.12.005","article-title":"Internal simulation of perception: a minimal neuro-robotic model","volume":"68","author":"Ziemke","year":"2005","journal-title":"Neurocomputing"}],"container-title":["Frontiers in Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2023.1098982\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,25]],"date-time":"2023-01-25T09:28:18Z","timestamp":1674638898000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/frai.2023.1098982\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,25]]},"references-count":48,"alternative-id":["10.3389\/frai.2023.1098982"],"URL":"https:\/\/doi.org\/10.3389\/frai.2023.1098982","relation":{},"ISSN":["2624-8212"],"issn-type":[{"type":"electronic","value":"2624-8212"}],"subject":[],"published":{"date-parts":[[2023,1,25]]},"article-number":"1098982"}}