{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T02:33:56Z","timestamp":1760150036077,"version":"build-2065373602"},"reference-count":50,"publisher":"MDPI AG","issue":"20","license":[{"start":{"date-parts":[[2023,10,15]],"date-time":"2023-10-15T00:00:00Z","timestamp":1697328000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Autonomous driving is a complex task that requires high-level hierarchical reasoning. Various solutions based on hand-crafted rules, multi-modal systems, or end-to-end learning have been proposed over time but are not quite ready to deliver the accuracy and safety necessary for real-world urban autonomous driving. Those methods require expensive hardware for data collection or environmental perception and are sensitive to distribution shifts, making large-scale adoption impractical. We present an approach that solely uses monocular camera inputs to generate valuable data without any supervision. Our main contributions involve a mechanism that can provide steering data annotations starting from unlabeled data alongside a different pipeline that generates path labels in a completely self-supervised manner. Thus, our method represents a natural step towards leveraging the large amounts of available online data ensuring the complexity and the diversity required to learn a robust autonomous driving policy.<\/jats:p>","DOI":"10.3390\/s23208473","type":"journal-article","created":{"date-parts":[[2023,10,15]],"date-time":"2023-10-15T10:47:32Z","timestamp":1697366852000},"page":"8473","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Self-Supervised Steering and Path Labeling for Autonomous Driving"],"prefix":"10.3390","volume":"23","author":[{"given":"Andrei","family":"Mihalea","sequence":"first","affiliation":[{"name":"Department of Computer Science, Faculty of Automatic Control and Computers, University Politehnica of Bucharest, 060042 Bucharest, Romania"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Robert-Florian","family":"Samoilescu","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Faculty of Automatic Control and Computers, University Politehnica of Bucharest, 060042 Bucharest, Romania"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7249-1871","authenticated-orcid":false,"given":"Adina Magda","family":"Florea","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Faculty of Automatic Control and Computers, University Politehnica of Bucharest, 060042 Bucharest, Romania"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,10,15]]},"reference":[{"key":"ref_1","unstructured":"Pomerleau, D.A. (1989, January 27\u201330). Alvinn: An autonomous land vehicle in a neural network. Proceedings of the Advances in Neural Information Processing Systems, Denver, CO, USA."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Hawke, J., Shen, R., Gurau, C., Sharma, S., Reda, D., Nikolov, N., Mazur, P., Micklethwaite, S., Griffiths, N., and Shah, A. (2019). Urban Driving with Conditional Imitation Learning. arXiv.","DOI":"10.1109\/ICRA40945.2020.9197408"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"Segnet: A deep convolutional encoder-decoder architecture for image segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Zhou, T., Brown, M., Snavely, N., and Lowe, D.G. (2017, January 21\u201326). Unsupervised learning of depth and ego-motion from video. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.700"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Yin, Z., and Shi, J. (2018, January 18\u201323). Geonet: Unsupervised learning of dense depth, optical flow and camera pose. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00212"},{"key":"ref_7","unstructured":"Bojarski, M., Del Testa, D., Dworakowski, D., Firner, B., Flepp, B., Goyal, P., Jackel, L.D., Monfort, M., Muller, U., and Zhang, J. (2016). End to end learning for self-driving cars. arXiv."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Drews, P., Williams, G., Goldfain, B., Theodorou, E.A., and Rehg, J.M. (2017). Aggressive Deep Driving: Model Predictive Control with a CNN Cost Model. arXiv, Available online: http:\/\/xxx.lanl.gov\/abs\/1707.05303.","DOI":"10.1109\/ICRA.2016.7487277"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Barnes, D., Maddern, W.P., and Posner, I. (2016). Find Your Own Way: Weakly-Supervised Segmentation of Path Proposals for Urban Autonomy. arXiv, Available online: http:\/\/xxx.lanl.gov\/abs\/1610.01238.","DOI":"10.1109\/ICRA.2017.7989025"},{"key":"ref_10","unstructured":"Bian, J.W., Li, Z., Wang, N., Zhan, H., Shen, C., Cheng, M.M., and Reid, I. (2019, January 8\u201314). Unsupervised Scale-consistent Depth and Ego-motion Learning from Monocular Video. Proceedings of the Thirty-third Conference on Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Xu, H., Gao, Y., Yu, F., and Darrell, T. (2017, January 21\u201326). End-to-end learning of driving models from large-scale video datasets. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.376"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Mehta, A., Subramanian, A., and Subramanian, A. (2018). Learning end-to-end autonomous driving using guided auxiliary supervision. arXiv.","DOI":"10.1145\/3293353.3293364"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Chen, Y., Praveen, P., Priyantha, M., Muelling, K., and Dolan, J. (2019, January 7\u201311). Learning on-road visual control for self-driving vehicles with auxiliary tasks. Proceedings of the 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), Waikoloa Village, HI, USA.","DOI":"10.1109\/WACV.2019.00041"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Bansal, M., Krizhevsky, A., and Ogale, A. (2018). Chauffeurnet: Learning to drive by imitating the best and synthesizing the worst. arXiv.","DOI":"10.15607\/RSS.2019.XV.031"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Kendall, A., Hawke, J., Janz, D., Mazur, P., Reda, D., Allen, J.M., Lam, V.D., Bewley, A., and Shah, A. (2019, January 20\u201324). Learning to drive in a day. Proceedings of the 2019 International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada.","DOI":"10.1109\/ICRA.2019.8793742"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1143","DOI":"10.1109\/LRA.2020.2966414","article-title":"Learning robust control policies for end-to-end autonomous driving from data-driven simulation","volume":"5","author":"Amini","year":"2020","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3467978","article-title":"An Uncertainty-Based Neural Network for Explainable Trajectory Segmentation","volume":"13","author":"Bi","year":"2021","journal-title":"ACM Trans. Intell. Syst. Technol."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"167855","DOI":"10.1109\/ACCESS.2021.3134889","article-title":"Pixel Level Segmentation Based Drivable Road Region Detection and Steering Angle Estimation Method for Autonomous Driving on Unstructured Roads","volume":"9","author":"Rasib","year":"2021","journal-title":"IEEE Access"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"3066","DOI":"10.1109\/LRA.2020.2975414","article-title":"See the future: A semantic segmentation network predicting ego-vehicle trajectory with a single monocular camera","volume":"5","author":"Sun","year":"2020","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Tang, L., Ding, X., Yin, H., Wang, Y., and Xiong, R. (2017, January 5\u20138). From one to many: Unsupervised traversable area segmentation in off-road environment. Proceedings of the 2017 IEEE International Conference on Robotics and Biomimetics (ROBIO), Macau, China.","DOI":"10.1109\/ROBIO.2017.8324513"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1509","DOI":"10.1109\/LRA.2019.2895390","article-title":"Where Should I Walk? Predicting Terrain Properties From Images Via Self-Supervised Learning","volume":"4","author":"Wellhausen","year":"2019","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_22","unstructured":"Faust, A., Hsu, D., and Neumann, G. (2022, January 8\u201311). Semantic Terrain Classification for Off-Road Autonomous Driving. Proceedings of the 5th Conference on Robot Learning, London, UK."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Ha, H., Im, S., Park, J., Jeon, H.G., and So Kweon, I. (2016, January 27\u201330). High-quality depth from uncalibrated small motion clip. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.584"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Kong, N., and Black, M.J. (2015, January 7\u201313). Intrinsic depth: Improving depth transfer with intrinsic images. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.401"},{"key":"ref_25","unstructured":"Eigen, D., Puhrsch, C., and Fergus, R. (2014, January 8\u201313). Depth map prediction from a single image using a multi-scale deep network. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_26","first-page":"1426","article-title":"Monocular depth estimation using multi-scale continuous CRFs as sequential deep networks","volume":"41","author":"Ricci","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Xu, D., Wang, W., Tang, H., Liu, H., Sebe, N., and Ricci, E. (2018, January 18\u201323). Structured attention guided convolutional neural fields for monocular depth estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00412"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Ummenhofer, B., Zhou, H., Uhrig, J., Mayer, N., Ilg, E., Dosovitskiy, A., and Brox, T. (2016). DeMoN: Depth and Motion Network for Learning Monocular Stereo. arXiv, Available online: http:\/\/xxx.lanl.gov\/abs\/1612.02401.","DOI":"10.1109\/CVPR.2017.596"},{"key":"ref_29","unstructured":"Wei, X., Zhang, Y., Li, Z., Fu, Y., and Xue, X. (2019). DeepSFM: Structure From Motion Via Deep Bundle Adjustment. arXiv, Available online: http:\/\/xxx.lanl.gov\/abs\/1912.09697."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Guizilini, V., Ambrus, R., Pillai, S., and Gaidon, A. (2019). PackNet-SfM: 3D Packing for Self-Supervised Monocular Depth Estimation. arXiv, Available online: http:\/\/xxx.lanl.gov\/abs\/1905.02693.","DOI":"10.1109\/CVPR42600.2020.00256"},{"key":"ref_31","unstructured":"Casser, V., Pirk, S., Mahjourian, R., and Angelova, A. (February, January 27). Depth prediction without the sensors: Leveraging structure for unsupervised learning from monocular videos. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_32","unstructured":"Godard, C., Mac Aodha, O., Firman, M., and Brostow, G.J. (November, January 27). Digging into self-supervised monocular depth estimation. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Godard, C., Mac Aodha, O., and Brostow, G.J. (2017, January 21\u201326). Unsupervised monocular depth estimation with left-right consistency. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.699"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Chen, Y., Schmid, C., and Sminchisescu, C. (2019). Self-supervised Learning with Geometric Constraints in Monocular Video: Connecting Flow, Depth, and Camera. arXiv, Available online: http:\/\/xxx.lanl.gov\/abs\/1907.05820.","DOI":"10.1109\/ICCV.2019.00716"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Ranjan, A., Jampani, V., Kim, K., Sun, D., Wulff, J., and Black, M.J. (2018). Adversarial Collaboration: Joint Unsupervised Learning of Depth, Camera Motion, Optical Flow and Motion Segmentation. arXiv, Available online: http:\/\/xxx.lanl.gov\/abs\/1805.09806.","DOI":"10.1109\/CVPR.2019.01252"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Mahjourian, R., Wicke, M., and Angelova, A. (2018). Unsupervised Learning of Depth and Ego-Motion from Monocular Video Using 3D Geometric Constraints. arXiv, Available online: http:\/\/xxx.lanl.gov\/abs\/1802.05522.","DOI":"10.1109\/CVPR.2018.00594"},{"key":"ref_37","unstructured":"Rusinkiewicz, S., and Levoy, M. (June, January 28). Efficient variants of the ICP algorithm. Proceedings of the Third International Conference on 3-D Digital Imaging and Modeling, Quebec City, QC, Canada."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"239","DOI":"10.1109\/34.121791","article-title":"A method for registration of 3-D shapes","volume":"14","author":"Besl","year":"1992","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_39","unstructured":"Chen, Y., and Medioni, G. (1991, January 9\u201311). Object modeling by registration of multiple range images. Proceedings of the 1991 IEEE International Conference on Robotics and Automation, Sacramento, CA, USA."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"1751","DOI":"10.1109\/TCSVT.2021.3080928","article-title":"Depth Estimation Using a Self-Supervised Network Based on Cross-Layer Feature Fusion and the Quadtree Constraint","volume":"32","author":"Tian","year":"2022","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Gordon, A., Li, H., Jonschkowski, R., and Angelova, A. (2019). Depth from Videos in the Wild: Unsupervised Monocular Depth Learning from Unknown Cameras. arXiv, Available online: http:\/\/xxx.lanl.gov\/abs\/1904.04998.","DOI":"10.1109\/ICCV.2019.00907"},{"key":"ref_42","unstructured":"Mertan, A., Duff, D.J., and Unal, G. (2021). Single Image Depth Estimation: An Overview. arXiv, Available online: http:\/\/xxx.lanl.gov\/abs\/2104.06456."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Nica, A., Tr\u0103sc\u0103u, M., Rotaru, A.A., Andreescu, C., Sorici, A., Florea, A.M., and Bacue, V. (2019, January 28\u201330). Collecting and Processing a Self-Driving Dataset in the UPB Campus. Proceedings of the The 22nd International Conference on Control Systems and Computer Science, Bucharest, Romania.","DOI":"10.1109\/CSCS.2019.00041"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Geiger, A., Lenz, P., and Urtasun, R. (2012, January 16\u201321). Are we ready for Autonomous Driving? The KITTI Vision Benchmark Suite. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Mayer, N., Ilg, E., H\u00e4usser, P., Fischer, P., Cremers, D., Dosovitskiy, A., and Brox, T. (2015). A Large Dataset to Train Convolutional Networks for Disparity, Optical Flow, and Scene Flow Estimation. arXiv, Available online: http:\/\/xxx.lanl.gov\/abs\/1512.02134.","DOI":"10.1109\/CVPR.2016.438"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2015). Deep Residual Learning for Image Recognition. arXiv, Available online: http:\/\/xxx.lanl.gov\/abs\/1512.03385.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_47","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_48","unstructured":"Chen, L., Papandreou, G., Schroff, F., and Adam, H. (2017). Rethinking Atrous Convolution for Semantic Image Segmentation. arXiv, Available online: http:\/\/xxx.lanl.gov\/abs\/1706.05587."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). ImageNet: A Large-Scale Hierarchical Image Database. Proceedings of the CVPR09, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_50","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., and Antiga, L. (2019). Advances in Neural Information Processing Systems 32, Curran Associates, Inc."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/20\/8473\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T21:07:06Z","timestamp":1760130426000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/20\/8473"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,15]]},"references-count":50,"journal-issue":{"issue":"20","published-online":{"date-parts":[[2023,10]]}},"alternative-id":["s23208473"],"URL":"https:\/\/doi.org\/10.3390\/s23208473","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2023,10,15]]}}}