{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T17:45:31Z","timestamp":1777657531631,"version":"3.51.4"},"reference-count":37,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2024,1,8]],"date-time":"2024-01-08T00:00:00Z","timestamp":1704672000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["42074039"],"award-info":[{"award-number":["42074039"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>This work presents a novel RGB-D dynamic Simultaneous Localisation and Mapping (SLAM) method that improves the precision, stability, and efficiency of localisation while relying on lightweight deep learning in a dynamic environment compared to the traditional static feature-based visual SLAM algorithm. Based on ORB-SLAM3, the GCNv2-tiny network instead of the ORB method, improves the reliability of feature extraction and matching and the accuracy of position estimation; then, the semantic segmentation thread employs the lightweight YOLOv5s object detection algorithm based on the GSConv network combined with a depth image to determine potentially dynamic regions of the image. Finally, to guarantee that the static feature points are used for position estimation, dynamic probability is employed to determine the true dynamic feature points based on the optical flow, semantic labels, and the state in last frame. We have performed experiments on the TUM datasets to verify the feasibility of the algorithm. Compared with the classical dynamic visual SLAM algorithm, the experimental results demonstrate that the absolute trajectory error is greatly reduced in dynamic environments, and that the computing efficiency is improved by 31.54% compared with the real-time dynamic visual SLAM algorithm with close accuracy, demonstrating the superiority of DLD-SLAM in accuracy, stability, and efficiency.<\/jats:p>","DOI":"10.3390\/rs16020246","type":"journal-article","created":{"date-parts":[[2024,1,8]],"date-time":"2024-01-08T07:59:20Z","timestamp":1704700760000},"page":"246","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":23,"title":["DLD-SLAM: RGB-D Visual Simultaneous Localisation and Mapping in Indoor Dynamic Environments Based on Deep Learning"],"prefix":"10.3390","volume":"16","author":[{"given":"Han","family":"Yu","sequence":"first","affiliation":[{"name":"School of Instrument Science and Engineering, Southeast University, Nanjing 210096, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qing","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Instrument Science and Engineering, Southeast University, Nanjing 210096, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chao","family":"Yan","sequence":"additional","affiliation":[{"name":"School of Instrument Science and Engineering, Southeast University, Nanjing 210096, China"},{"name":"School of Electrical Engineering and Automation, Changshu Institute of Technology, Changshu 215500, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5980-4164","authenticated-orcid":false,"given":"Youyang","family":"Feng","sequence":"additional","affiliation":[{"name":"School of Instrument Science and Engineering, Southeast University, Nanjing 210096, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yang","family":"Sun","sequence":"additional","affiliation":[{"name":"School of Instrument Science and Engineering, Southeast University, Nanjing 210096, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lu","family":"Li","sequence":"additional","affiliation":[{"name":"School of Instrument Science and Engineering, Southeast University, Nanjing 210096, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2024,1,8]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"117734","DOI":"10.1016\/j.eswa.2022.117734","article-title":"A Survey of State-of-the-Art on Visual SLAM","volume":"205","author":"Fitzgerald","year":"2022","journal-title":"Expert Syst. Appl."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1255","DOI":"10.1109\/TRO.2017.2705103","article-title":"ORB-SLAM2: An open-source slam system for monocular, stereo, and rgb-d cameras","volume":"33","year":"2017","journal-title":"IEEE Trans. Robot."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1874","DOI":"10.1109\/TRO.2021.3075644","article-title":"ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual\u2013Inertial, and Multimap SLAM","volume":"37","author":"Campos","year":"2021","journal-title":"IEEE Trans. Robot."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1004","DOI":"10.1109\/TRO.2018.2853729","article-title":"VINS-Mono: A Robust and Versatile Monocular Visual-Inertial State Estimator","volume":"34","author":"Qin","year":"2018","journal-title":"IEEE Trans. Robot."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Forster, C., Pizzoli, M., and Scaramuzza, D. (June, January 31). SVO: Fast semi-direct monocular visual odometry. Proceedings of the 2014 IEEE International Conference on Robotics and Automation (ICRA), Hong Kong, China.","DOI":"10.1109\/ICRA.2014.6906584"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Chengqi, D., Kaitao, Q., and Rong, X. (2019, January 13\u201315). Comparative Study of Deep Learning Based Features in SLAM. Proceedings of the 2019 4th Asia-Pacific Conference on Intelligent Robot Systems (ACIRS), Nagoya, Japan.","DOI":"10.1109\/ACIRS.2019.8935995"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"3400","DOI":"10.1109\/ACCESS.2022.3140292","article-title":"DeepFeat: Robust Large-Scale Multi-Features Outdoor Localization in LTE Networks Using Deep Learning","volume":"10","author":"Mohamed","year":"2022","journal-title":"IEEE Access"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Leibe, B., Matas, J., Sebe, N., and Welling, M. (2016, January 11\u201314). LIFT: Learned Invariant Feature Transform. Proceedings of the Computer Vision\u2014ECCV 2016, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46478-7"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Ballester, I., Font\u00e1n, A., Civera, J., Strobl, K.H., and Triebel, R. (June, January 30). DOT: Dynamic Object Tracking for Visual SLAM. Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), Xi\u2019an, China.","DOI":"10.1109\/ICRA48506.2021.9561452"},{"key":"ref_10","unstructured":"Li, H., Li, J., Wei, H., Liu, Z., Zhan, Z., and Ren, Q. (2022). Slim-neck by GSConv: A better design paradigm of detector architectures for autonomous vehicles. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Xie, Y., Tang, Y., Tang, G., and Hoff, W. (2021, January 10\u201315). Learning To Find Good Correspondences Of Multiple Objects. Proceedings of the 2020 25th International Conference on Pattern Recognition (ICPR), Milan, Italy.","DOI":"10.1109\/ICPR48806.2021.9413319"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Dusmanu, M., Rocco, I., Pajdla, T., Pollefeys, M., Sivic, J., Torii, A., and Sattler, T. (2019, January 16\u201320). D2-Net: A Trainable CNN for Joint Description and Detection of Local Features. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00828"},{"key":"ref_13","unstructured":"Revaud, J., Weinzaepfel, P., Souza, C.D., and Humenberger, M. (2019, January 8\u201314). R2D2: Repeatable and Reliable Detector and Descriptor. Proceedings of the 33rd International Conference on Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"DeTone, D., Malisiewicz, T., and Rabinovich, A. (2018, January 18\u201322). SuperPoint: Self-Supervised Interest Point Detection and Description. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00060"},{"key":"ref_15","first-page":"3505","article-title":"GCNv2: Efficient Correspondence Prediction for Real-Time SLAM","volume":"4","author":"Tang","year":"2019","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Clark, R., Wang, S., Wen, H., Markham, A., and Trigoni, N. (2017, January 4\u20139). VINet: Visual-Inertial Odometry as a Sequence-to-Sequence Learning Problem. Proceedings of the AAAI Conference on Artificial Intelligence, San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.11215"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Zhang, R., and Zhang, X. (2023). Geometric Constraint-Based and Improved YOLOv5 Semantic SLAM for Dynamic Scenes. ISPRS Int. J. Geo-Inf., 12.","DOI":"10.3390\/ijgi12060211"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Zhang, X., Zhang, R., and Wang, X. (2022). Visual SLAM Mapping Based on YOLOv5 in Dynamic Scenes. Appl. Sci., 12.","DOI":"10.3390\/app122211548"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1565","DOI":"10.1109\/TRO.2016.2609395","article-title":"Effective Background Model-Based RGB-D Dense Visual Odometry in a Dynamic Environment","volume":"32","author":"Kim","year":"2016","journal-title":"IEEE Trans. Robot."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Kerl, C., Sturm, J., and Cremers, D. (2013, January 3\u20137). Dense visual SLAM for RGB-D cameras. Proceedings of the 2013 IEEE\/RSJ International Conference on Intelligent Robots and Systems, Tokyo, Japan.","DOI":"10.1109\/IROS.2013.6696650"},{"key":"ref_21","first-page":"203","article-title":"Geometric Constraint-Based Visual SLAM Under Dynamic Indoor Environment","volume":"57","author":"Guohao","year":"2021","journal-title":"Comput. Eng. Appl."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zhang, C., Zhang, R., Jin, S., and Yi, X. (2022). PFD-SLAM: A New RGB-D SLAM for Dynamic Indoor Environments Based on Non-Prior Semantic Segmentation. Remote Sens., 14.","DOI":"10.3390\/rs14102445"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Yu, C., Liu, Z., Liu, X.-J., Xie, F., Yang, Y., Wei, Q., and Fei, Q. (2018, January 1\u20135). DS-SLAM: A Semantic Visual SLAM towards Dynamic Environments. Proceedings of the 2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain.","DOI":"10.1109\/IROS.2018.8593691"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Zhong, F., Wang, S., Zhang, Z., Chen, C., and Wang, Y. (2018, January 12\u201315). Detect-SLAM: Making Object Detection and SLAM Mutually Beneficial. Proceedings of the 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Tahoe, NV, USA.","DOI":"10.1109\/WACV.2018.00115"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.-Y., and Berg, A.C. (2016, January 11\u201314). SSD: Single Shot MultiBox Detector. Proceedings of the Computer Vision\u2014ECCV 2016, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"23772","DOI":"10.1109\/ACCESS.2021.3050617","article-title":"RDS-SLAM: Real-Time Dynamic SLAM Using Semantic Segmentation Methods","volume":"9","author":"Liu","year":"2021","journal-title":"IEEE Access"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"128","DOI":"10.1016\/j.ins.2020.12.019","article-title":"DP-SLAM: A visual SLAM with moving probability towards dynamic environments","volume":"556","author":"Li","year":"2020","journal-title":"Inf. Sci."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask R-CNN. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"4076","DOI":"10.1109\/LRA.2018.2860039","article-title":"DynaSLAM: Tracking, Mapping, and Inpainting in Dynamic Scenes","volume":"3","author":"Bescos","year":"2018","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"6011","DOI":"10.1007\/s00521-021-06764-3","article-title":"YOLO-SLAM: A semantic SLAM system towards dynamic environment with geometric constraint","volume":"34","author":"Wu","year":"2022","journal-title":"Neural Comput. Appl."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Wei, S., Wang, S., Li, H., Liu, G., Yang, T., and Liu, C. (2023). A Semantic Information-Based Optimized vSLAM in Indoor Dynamic Environments. Appl. Sci., 13.","DOI":"10.3390\/app13158790"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Wang, X., and Zhang, X. (2023). MCBM-SLAM: An Improved Mask-Region-Convolutional Neural Network-Based Simultaneous Localization and Mapping System for Dynamic Environments. Electronics, 12.","DOI":"10.3390\/electronics12173596"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Xiao, J., Owens, A., and Torralba, A. (2013, January 1\u20138). SUN3D: A Database of Big Spaces Reconstructed Using SfM and Object Labels. Proceedings of the 2013 IEEE International Conference on Computer Vision, Sydney, Australia.","DOI":"10.1109\/ICCV.2013.458"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Wang, C.-Y., Liao, H.-Y.M., Wu, Y.-H., Chen, P.-Y., Hsieh, J.-W., and Yeh, I.-H. (2020, January 14\u201319). CSPNet: A New Backbone that can Enhance Learning Capability of CNN. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Seattle, WA, USA.","DOI":"10.1109\/CVPRW50498.2020.00203"},{"key":"ref_37","unstructured":"Lucas, B.D., and Kanade, T. (1981, January 24\u201328). An Iterative Image Registration Technique with an Application to Stereo Vision. Proceedings of the 7th International Joint Conference on Artificial Intelligence\u2014Volume 2, Vancouver, BC, Canada."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/16\/2\/246\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T13:42:18Z","timestamp":1760103738000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/16\/2\/246"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,1,8]]},"references-count":37,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2024,1]]}},"alternative-id":["rs16020246"],"URL":"https:\/\/doi.org\/10.3390\/rs16020246","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,1,8]]}}}