{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T20:31:16Z","timestamp":1780086676949,"version":"3.54.0"},"reference-count":53,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2023,12,28]],"date-time":"2023-12-28T00:00:00Z","timestamp":1703721600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Direct policy learning (DPL) is a widely used approach in imitation learning for time-efficient and effective convergence when training mobile robots. However, using DPL in real-world applications is not sufficiently explored due to the inherent challenges of mobilizing direct human expertise and the difficulty of measuring comparative performance. Furthermore, autonomous systems are often resource-constrained, thereby limiting the potential application and implementation of highly effective deep learning models. In this work, we present a lightweight DPL-based approach to train mobile robots in navigational tasks. We integrated a safety policy alongside the navigational policy to safeguard the robot and the environment. The approach was evaluated in simulations and real-world settings and compared with recent work in this space. The results of these experiments and the efficient transfer from simulations to real-world settings demonstrate that our approach has improved performance compared to its hardware-intensive counterparts. We show that using the proposed methodology, the training agent achieves closer performance to the expert within the first 15 training iterations in simulation and real-world settings.<\/jats:p>","DOI":"10.3390\/s24010185","type":"journal-article","created":{"date-parts":[[2023,12,28]],"date-time":"2023-12-28T09:35:21Z","timestamp":1703756121000},"page":"185","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Hardware Efficient Direct Policy Imitation Learning for Robotic Navigation in Resource-Constrained Settings"],"prefix":"10.3390","volume":"24","author":[{"given":"Vidura","family":"Sumanasena","sequence":"first","affiliation":[{"name":"Research Centre for Data Analytics and Cognition, La Trobe University, Bundoora, VIC 3083, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Heshan","family":"Fernando","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering, Rensselaer Polytechnic Institute, New York, NY 12180, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3878-5969","authenticated-orcid":false,"given":"Daswin","family":"De Silva","sequence":"additional","affiliation":[{"name":"Research Centre for Data Analytics and Cognition, La Trobe University, Bundoora, VIC 3083, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-6891-3553","authenticated-orcid":false,"given":"Beniel","family":"Thileepan","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Warwick, Coventry CV4 7AL, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Amila","family":"Pasan","sequence":"additional","affiliation":[{"name":"Centre for Wireless Communications, University of Oulu, 90570 Oulu, Finland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jayathu","family":"Samarawickrama","sequence":"additional","affiliation":[{"name":"Department of Electronic and Telecom Engineering, University of Moratuwa, Moratuwa 10400, Sri Lanka"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Evgeny","family":"Osipov","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Electrical and Space Engineering, Lule\u00e5 University of Technology, 97187 Lule\u00e5, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Damminda","family":"Alahakoon","sequence":"additional","affiliation":[{"name":"Research Centre for Data Analytics and Cognition, La Trobe University, Bundoora, VIC 3083, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,12,28]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1109\/MRA.2010.936952","article-title":"Imitation and reinforcement learning","volume":"17","author":"Kober","year":"2010","journal-title":"IEEE Robot. Autom. Mag."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"13753","DOI":"10.1109\/ACCESS.2022.3146518","article-title":"A survey of domain-specific architectures for reinforcement learning","volume":"10","author":"Rothmann","year":"2022","journal-title":"IEEE Access"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"27091","DOI":"10.1109\/ACCESS.2017.2777827","article-title":"System design perspective for human-level agents using deep reinforcement learning: A survey","volume":"5","author":"Nguyen","year":"2017","journal-title":"IEEE Access"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Yang, C., Liang, P., Ajoudani, A., Li, Z., and Bicchi, A. (2016, January 9\u201314). Development of a robotic teaching interface for human to human skill transfer. Proceedings of the 2016 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Daejeon, Republic of Korea.","DOI":"10.1109\/IROS.2016.7759130"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Luo, J., Dong, X., and Yang, H. (2015, January 27\u201330). Session search by direct policy learning. Proceedings of the 2015 International Conference on the Theory of Information Retrieval, New York, NY, USA.","DOI":"10.1145\/2808194.2809461"},{"key":"ref_6","first-page":"14128","article-title":"A survey on imitation learning techniques for end-to-end autonomous vehicles","volume":"5","author":"Yi","year":"2022","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Pan, Y., Cheng, C.A., Saigol, K., Lee, K., Yan, X., Theodorou, E., and Boots, B. (2017). Agile autonomous driving using end-to-end deep imitation learning. arXiv.","DOI":"10.15607\/RSS.2018.XIV.056"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Kelly, M., Sidrane, C., Driggs-Campbell, K., and Kochenderfer, M.J. (2019, January 20\u201324). Hg-dagger: Interactive imitation learning with human experts. Proceedings of the 2019 International Conference on Robotics and Automation (ICRA), IEEE, Montreal, QC, Canada.","DOI":"10.1109\/ICRA.2019.8793698"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"373","DOI":"10.1017\/S0140525X0000546X","article-title":"Against direct perception","volume":"3","author":"Ullman","year":"1980","journal-title":"Behav. Brain Sci."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1231","DOI":"10.1177\/0278364913491297","article-title":"Vision meets robotics: The kitti dataset","volume":"32","author":"Geiger","year":"2013","journal-title":"Int. J. Robot. Res."},{"key":"ref_11","unstructured":"\u00d6lsner, F., and Milz, S. (2020, January 8\u201314). Catch me, if you can! A mediated perception approach towards fully autonomous drone racing. Proceedings of the NeurIPS 2019 Competition and Demonstration Track, PMLR, Vancouver, BC, Canada."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Chen, C., Seff, A., Kornhauser, A., and Xiao, J. (2015, January 7\u201313). DeepDriving: Learning affordance for direct perception in autonomous driving. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.312"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1627","DOI":"10.1109\/TPAMI.2009.167","article-title":"Object detection with discriminatively trained part-based models","volume":"32","author":"Felzenszwalb","year":"2009","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Lenz, P., Ziegler, J., Geiger, A., and Roser, M. (2011, January 5\u20139). Sparse scene flow segmentation for moving object detection in urban environments. Proceedings of the 2011 IEEE Intelligent Vehicles Symposium (IV), IEEE, Baden-Baden, Germany.","DOI":"10.1109\/IVS.2011.5940558"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Aly, M. (2008, January 4\u20136). Real time detection of lane markers in urban streets. Proceedings of the 2008 IEEE Intelligent Vehicles Symposium, IEEE, Eindhoven, The Netherlands.","DOI":"10.1109\/IVS.2008.4621152"},{"key":"ref_16","unstructured":"Franke, U., and Kutzbach, I. (1996, January 19\u201320). Fast stereo based object detection for stop&go traffic. Proceedings of the Conference on Intelligent Vehicles, Tokyo, Japan."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"40","DOI":"10.1109\/5254.736001","article-title":"Autonomous driving goes downtown","volume":"13","author":"Franke","year":"1998","journal-title":"IEEE Intell. Syst. Their Appl."},{"key":"ref_18","unstructured":"Bertozzi, M., Broggi, A., Conte, G., and Fascioli, A. (1997, January 8\u201312). Obstacle and lane detection on ARGO. Proceedings of the Conference on Intelligent Transportation Systems, Macau, China."},{"key":"ref_19","unstructured":"Bertozzi, M., and Broggi, A. (1996, January 11\u201317). Real-time lane and obstacle detection on the GOLD system. Proceedings of the Conference on Intelligent Vehicles, Nagoya, Japan."},{"key":"ref_20","first-page":"1","article-title":"Alvinn: An autonomous land vehicle in a neural network","volume":"1","author":"Pomerleau","year":"1988","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_21","unstructured":"Pomerleau, D.A. (2012). Neural Network Perception for Mobile Robot Guidance, Springer Science & Business Media."},{"key":"ref_22","unstructured":"Bojarski, M., Del Testa, D., Dworakowski, D., Firner, B., Flepp, B., Goyal, P., Jackel, L.D., Monfort, M., Muller, U., and Zhang, J. (2016). End to end learning for self-driving cars. arXiv."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Cultrera, L., Seidenari, L., Becattini, F., Pala, P., and Del Bimbo, A. (2020, January 20\u201325). Explaining autonomous driving by learning end-to-end visual attention. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, Nashville, TN, USA.","DOI":"10.1109\/CVPRW50498.2020.00178"},{"key":"ref_24","unstructured":"Ruiz-del Solar, J., Loncomilla, P., and Soto, N. (2018). A survey on deep learning methods for robot vision. arXiv."},{"key":"ref_25","unstructured":"Zhou, B., Lapedriza, A., Xiao, J., Torralba, A., and Oliva, A. (2014). Learning deep features for scene recognition using places database. Adv. Neural Inf. Process. Syst., 27."},{"key":"ref_26","unstructured":"Gomez-Ojeda, R., Lopez-Antequera, M., Petkov, N., and Gonzalez-Jimenez, J. (2015). Training a convolutional neural network for appearance-invariant place recognition. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Hou, Y., Zhang, H., and Zhou, S. (2015, January 8\u201310). Convolutional neural network-based image representation for visual loop closure detection. Proceedings of the 2015 IEEE International Conference on Information and Automation, IEEE, Lijiang, China.","DOI":"10.1109\/ICInfA.2015.7279659"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"S\u00fcnderhauf, N., Dayoub, F., McMahon, S., Talbot, B., Schulz, R., Corke, P., Wyeth, G., Upcroft, B., and Milford, M. (2016, January 16\u201321). Place categorization and semantic mapping on a mobile robot. Proceedings of the 2016 IEEE International Conference on Robotics and Automation (ICRA), Stockholm, Sweden.","DOI":"10.1109\/ICRA.2016.7487796"},{"key":"ref_29","unstructured":"Liao, Y., Kodagoda, S., Wang, Y., Shi, L., and Liu, Y. (2016, January 16\u201321). Understand scene categories by objects: A semantic regularized scene classifier using convolutional neural networks. Proceedings of the 2016 IEEE International Conference on Robotics and Automation (ICRA), Stockholm, Sweden."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 7\u201312). Going deeper with convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_31","unstructured":"(2023, December 21). GPU-Based Deep Learning Inference: A Performance and Power Analysis\u2014Nvidia, 2015. Available online: https:\/\/www.nvidia.com\/content\/tegra\/embedded-systems\/pdf\/jetson_tx1_whitepaper.pdf."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Haavaldsen, H., Aasboe, M., and Lindseth, F. (2019, January 27\u201328). Autonomous vehicle control: End-to-end learning in simulated urban environments. Proceedings of the Nordic Artificial Intelligence Research and Development: Third Symposium of the Norwegian AI Society, NAIS 2019, Trondheim, Norway.","DOI":"10.1007\/978-3-030-35664-4_4"},{"key":"ref_33","unstructured":"Codevilla, F., Santana, E., L\u00f3pez, A.M., and Gaidon, A. (November, January 27). Exploring the limitations of behavior cloning for autonomous driving. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Repluc of Korea."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Russell, S. (1998, January 24\u201326). Learning agents for uncertain environments. Proceedings of the Eleventh Annual Conference on Computational Learning Theory, Madison, WI, USA.","DOI":"10.1145\/279943.279964"},{"key":"ref_35","unstructured":"Ng, A.Y., and Russell, S. (July, January 29). Algorithms for inverse reinforcement learning. Proceedings of the ICML, Stanford, CA, USA."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Abbeel, P., and Ng, A.Y. (2004, January 4\u20138). Apprenticeship learning via inverse reinforcement learning. Proceedings of the Twenty-First International Conference on Machine Learning, New Yor, NY, USA.","DOI":"10.1145\/1015330.1015430"},{"key":"ref_37","unstructured":"Sadigh, D., Sastry, S., Seshia, S.A., and Dragan, A.D. (2016, January 18\u201322). Planning for autonomous cars that leverage effects on human actions. Proceedings of the Robotics: Science and Systems, Ann Arbor, MI, USA."},{"key":"ref_38","unstructured":"Ziebart, B.D., Maas, A., Bagnell, J.A., and Dey, A.K. (2018, January 13\u201317). Maximum entropy inverse reinforcement learning. Proceedings of the 23rd National Conference on Artificial Intelligence-Volume 3, Chicago, IL, USA."},{"key":"ref_39","unstructured":"Wulfmeier, M., Ondruska, P., and Posner, I. (2015). Maximum entropy deep inverse reinforcement learning. arXiv."},{"key":"ref_40","first-page":"4572","article-title":"Generative adversarial imitation learning","volume":"29","author":"Ho","year":"2016","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_41","unstructured":"Ross, S., Gordon, G., and Bagnell, D. (2011, January 10\u201314). A reduction of imitation learning and structured prediction to no-regret online learning. Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, JMLR Workshop and Conference Proceedings, Rome, Italy."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Zhang, J., and Cho, K. (2016). Query-efficient imitation learning for end-to-end autonomous driving. arXiv.","DOI":"10.1609\/aaai.v31i1.10857"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Li, G., Mueller, M., Casser, V., Smith, N., Michels, D.L., and Ghanem, B. (2018). Oil: Observational imitation learning. arXiv.","DOI":"10.15607\/RSS.2019.XV.005"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"902","DOI":"10.1007\/s11263-018-1073-7","article-title":"Sim4cv: A photo-realistic simulator for computer vision applications","volume":"126","author":"Casser","year":"2018","journal-title":"Int. J. Comput. Vis."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Jayaratne, M., Alahakoon, D., De Silva, D., and Yu, X. (2018, January 21\u201323). Bio-inspired multisensory Fusion for autonomous robots. Proceedings of the IECON 2018\u201444th Annual Conference of the IEEE Industrial Electronics Society, Washington, DC, USA.","DOI":"10.1109\/IECON.2018.8592809"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Durrant-Whyte, H., and Henderson, T.C. (2016). Multisensor Data Fusion, Springer Handbook of Robotics.","DOI":"10.1007\/978-3-319-32552-1_35"},{"key":"ref_47","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). ImageNet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_49","unstructured":"Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., and Koltun, V. (2017, January 13\u201315). CARLA: An Open Urban Driving Simulator. Proceedings of the Conference on Robot Learning, Mountain View, CA, USA."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Nawaratne, R., Bandaragoda, T., Adikari, A., Alahakoon, D., De Silva, D., and Yu, X. (November, January 29). Incremental knowledge acquisition and self-learning for autonomous video surveillance. Proceedings of the IECON 2017\u201443rd Annual Conference of the IEEE Industrial Electronics Society, Beijing, China.","DOI":"10.1109\/IECON.2017.8216826"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016, January 30\u201327). Rethinking the inception architecture for computer vision. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.308"},{"key":"ref_52","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv."},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Ketkar, N., and Ketkar, N. (2017). Introduction to keras. Deep Learning with Python: A Hands-On Introduction, Springer.","DOI":"10.1007\/978-1-4842-2766-4"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/1\/185\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T21:43:31Z","timestamp":1760132611000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/1\/185"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,12,28]]},"references-count":53,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,1]]}},"alternative-id":["s24010185"],"URL":"https:\/\/doi.org\/10.3390\/s24010185","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,12,28]]}}}