{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,6]],"date-time":"2026-06-06T19:33:56Z","timestamp":1780774436971,"version":"3.54.1"},"reference-count":31,"publisher":"Emerald","issue":"1","license":[{"start":{"date-parts":[[2024,8,6]],"date-time":"2024-08-06T00:00:00Z","timestamp":1722902400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.emerald.com\/insight\/site-policies"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IR"],"published-print":{"date-parts":[[2025,1,27]]},"abstract":"<jats:sec><jats:title content-type=\"abstract-subheading\">Purpose<\/jats:title>\n<jats:p>This paper proposes a self-supervised monocular depth estimation algorithm under multiple constraints, which can generate the corresponding depth map end-to-end based on RGB images. On this basis, based on the traditional visual simultaneous localisation and mapping (VSLAM) framework, a dynamic object detection framework based on deep learning is introduced, and dynamic objects in the scene are culled during mapping.<\/jats:p>\n<\/jats:sec>\n<jats:sec><jats:title content-type=\"abstract-subheading\">Design\/methodology\/approach<\/jats:title>\n<jats:p>Typical SLAM algorithms or data sets assume a static environment and do not consider the potential consequences of accidentally adding dynamic objects to a 3D map. This shortcoming limits the applicability of VSLAM in many practical cases, such as long-term mapping. In light of the aforementioned considerations, this paper presents a self-supervised monocular depth estimation algorithm based on deep learning. Furthermore, this paper introduces the YOLOv5 dynamic detection framework into the traditional ORBSLAM2 algorithm for the purpose of removing dynamic objects.<\/jats:p>\n<\/jats:sec>\n<jats:sec><jats:title content-type=\"abstract-subheading\">Findings<\/jats:title>\n<jats:p>Compared with Dyna-SLAM, the algorithm proposed in this paper reduces the error by about 13%, and compared with ORB-SLAM2 by about 54.9%. In addition, the algorithm in this paper can process a single frame of image at a speed of 15\u201320 FPS on GeForce RTX 2080s, far exceeding Dyna-SLAM in real-time performance.<\/jats:p>\n<\/jats:sec>\n<jats:sec><jats:title content-type=\"abstract-subheading\">Originality\/value<\/jats:title>\n<jats:p>This paper proposes a VSLAM algorithm that can be applied to dynamic environments. The algorithm consists of a self-supervised monocular depth estimation part under multiple constraints and the introduction of a dynamic object detection framework based on YOLOv5.<\/jats:p>\n<\/jats:sec>","DOI":"10.1108\/ir-04-2024-0166","type":"journal-article","created":{"date-parts":[[2024,8,3]],"date-time":"2024-08-03T06:38:14Z","timestamp":1722667094000},"page":"28-35","source":"Crossref","is-referenced-by-count":5,"title":["Visual SLAM algorithm in dynamic environment based on deep learning"],"prefix":"10.1108","volume":"52","author":[{"given":"Yingjie","family":"Yu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shuai","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xinpeng","family":"Yang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Changzhen","family":"Xu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sen","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wendong","family":"Xiao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"140","published-online":{"date-parts":[[2024,8,6]]},"reference":[{"issue":"4","key":"key2025012405290980600_ref001","doi-asserted-by":"publisher","first-page":"4076","DOI":"10.1109\/LRA.2018.2860039","article-title":"DynaSLAM: tracking, mapping, and inpainting in dynamic scenes","volume":"3","year":"2018","journal-title":"IEEE Robotics and Automation Letters"},{"key":"key2025012405290980600_ref002","article-title":"Unsupervised scale-consistent depth and egomotion learning from monocular video","year":"2019"},{"key":"key2025012405290980600_ref003","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/TIM.2022.3228006","article-title":"SG-SLAM: a real-time RGB-D visual SLAM toward dynamic scenes with semantic and geometric information","volume":"72","year":"2023","journal-title":"IEEE Transactions on Instrumentation and Measurement"},{"issue":"1","key":"key2025012405290980600_ref004","doi-asserted-by":"publisher","first-page":"373","DOI":"10.1109\/TPAMI.2020.3010942","article-title":"RGB-D SLAM in dynamic environments using point correlations","volume":"44","year":"2022","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"7","key":"key2025012405290980600_ref005","article-title":"A review of visual-LiDAR fusion based simultaneous localization and mapping","volume":"20","year":"2020","journal-title":"Sensors (Basel, Switzerland)"},{"key":"key2025012405290980600_ref006","first-page":"400","article-title":"Comparison of various SLAM systems for mobile robot in an indoor environment","year":"2018"},{"key":"key2025012405290980600_ref007","first-page":"6602","article-title":"Unsupervised monocular depth estimation with left-right consistency","volume-title":"2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","year":"2016"},{"key":"key2025012405290980600_ref008","first-page":"3827","article-title":"Digging into self-supervised monocular depth estimation","volume-title":"2019 IEEE\/CVF International Conference on Computer Vision (ICCV)","year":"2018"},{"key":"key2025012405290980600_ref009","unstructured":"Jocher, G. (2020), \u201cYOLOv5 by ultralytics\u201d, Version 7.0. doi:10.5281\/zenodo.3908559, available at: www.github.com\/ultralytics\/yolov5"},{"key":"key2025012405290980600_ref010","doi-asserted-by":"crossref","first-page":"117","DOI":"10.1109\/ITNEC.2019.8729285","article-title":"Review of vision-based simultaneous localization and mapping","volume-title":"2019 IEEE 3rd Information Technology, Networking, Electronic and Automation Control Conference (ITNEC)","year":"2019"},{"key":"key2025012405290980600_ref011","doi-asserted-by":"publisher","first-page":"2340","DOI":"10.1109\/ICASSP43922.2022.9746689","article-title":"Adaptive weighted network with edge enhancement module for monocular self-supervised depth estimation","volume-title":"ICASSP 2022 \u2013 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","year":"2022"},{"key":"key2025012405290980600_ref012","doi-asserted-by":"publisher","first-page":"2285","DOI":"10.1109\/ICIP49359.2023.10221949","article-title":"Self-supervised focus measure fusing for depth estimation from computer-generated holograms","volume-title":"2023 IEEE International Conference on Image Processing (ICIP)","year":"2023"},{"issue":"5","key":"key2025012405290980600_ref013","first-page":"1255","article-title":"ORB-SLAM2: an open-source SLAM system for monocular, stereo, and RGB-D cameras","volume":"33","year":"2016","journal-title":"IEEE Transactions on Robotics"},{"key":"key2025012405290980600_ref014","first-page":"602","article-title":"A review of SLAM techniques and security in autonomous driving","year":"2019"},{"issue":"4","key":"key2025012405290980600_ref015","doi-asserted-by":"publisher","first-page":"11523","DOI":"10.48550\/arXiv.2208.11500","article-title":"DynaVINS: a visual-inertial SLAM for dynamic environments","volume":"7","year":"2022","journal-title":"IEEE Robotics and Automation Letters"},{"key":"key2025012405290980600_ref016","first-page":"209","article-title":"Robust monocular SLAM in dynamic environments","year":"2013"},{"key":"key2025012405290980600_ref017","unstructured":"Yang, L., Kang, B., Huang, Z., Xu, X., Feng, J. and Zhao, H. (2024), \u201cDepth anything: unleashing the power of large-scale unlabeled data\u201d, ArXiv abs\/2401.10891, available at: www.api.semanticscholar.org\/CorpusID:267061016"},{"issue":"1","key":"key2025012405290980600_ref018","doi-asserted-by":"publisher","first-page":"289","DOI":"10.1109\/TRO.2022.3199087","article-title":"Dynam-SLAM: an accurate, robust stereo visual-inertial SLAM method in dynamic environments","volume":"39","year":"2023","journal-title":"IEEE Transactions on Robotics"},{"key":"key2025012405290980600_ref019","doi-asserted-by":"crossref","first-page":"1983","DOI":"10.1109\/CVPR.2018.00212","article-title":"GeoNet: unsupervised learning of dense depth, optical flow and camera pose","volume-title":"2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","year":"2018"},{"key":"key2025012405290980600_ref020","doi-asserted-by":"crossref","first-page":"9633","DOI":"10.1109\/CVPR.2019.00987","article-title":"VITAMIN-E: VIsual tracking and MappINg with extremely dense feature points","volume-title":"2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","year":"2019"},{"key":"key2025012405290980600_ref021","doi-asserted-by":"crossref","first-page":"340","DOI":"10.1109\/CVPR.2018.00043","article-title":"Unsupervised learning of monocular depth estimation and visual odometry with deep feature reconstruction","volume-title":"2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","year":"2018"},{"key":"key2025012405290980600_ref022","doi-asserted-by":"crossref","first-page":"6612","DOI":"10.1109\/CVPR.2017.700","article-title":"Unsupervised learning of depth and ego-motion from video","volume-title":"2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","year":"2017"},{"issue":"2","key":"key2025012405290980600_ref023","doi-asserted-by":"crossref","first-page":"354","DOI":"10.1109\/TPAMI.2012.104","article-title":"CoSLAM: collaborative visual SLAM in dynamic environments","volume":"35","year":"2013","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"key2025012405290980600_ref024","article-title":"Depth map prediction from a single image using a multi-scale deep network","year":"2014","journal-title":"In: Neural Information Processing Systems"},{"key":"key2025012405290980600_ref025","doi-asserted-by":"crossref","first-page":"2002","DOI":"10.1109\/CVPR.2018.00214","article-title":"Deep ordinal regression network for monocular depth estimation","volume-title":"2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","year":"2018"},{"key":"key2025012405290980600_ref026","doi-asserted-by":"crossref","first-page":"2215","DOI":"10.1109\/CVPR.2017.238","article-title":"Semi-supervised deep learning for monocular depth map prediction","volume-title":"2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","year":"2017"},{"issue":"10","key":"key2025012405290980600_ref027","first-page":"2024","article-title":"Learning depth from single monocular images using deep convolutional neural fields","volume":"38","year":"2015","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"key2025012405290980600_ref028","first-page":"12232","article-title":"Competitive collaboration: joint unsupervised learning of depth, camera motion, optical flow and motion segmentation","volume-title":"2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","year":"2018"},{"key":"key2025012405290980600_ref029","first-page":"2022","article-title":"Learning depth from monocular videos using direct methods","volume-title":"2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","year":"2017"},{"key":"key2025012405290980600_ref030","article-title":"Unsupervised learning of geometry with edgeaware depth-normal consistency","volume-title":"AAAI Conference on Artificial Intelligence.","year":"2017"},{"key":"key2025012405290980600_ref031","article-title":"DF-Net: unsupervised joint learning of depth and flow using cross-task consistency","year":"2018"}],"container-title":["Industrial Robot: the international journal of robotics research and application"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/IR-04-2024-0166\/full\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/IR-04-2024-0166\/full\/html","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,24]],"date-time":"2025-07-24T21:38:58Z","timestamp":1753393138000},"score":1,"resource":{"primary":{"URL":"http:\/\/www.emerald.com\/ir\/article\/52\/1\/28-35\/1242402"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,8,6]]},"references-count":31,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,8,6]]},"published-print":{"date-parts":[[2025,1,27]]}},"alternative-id":["10.1108\/IR-04-2024-0166"],"URL":"https:\/\/doi.org\/10.1108\/ir-04-2024-0166","relation":{},"ISSN":["0143-991X","1758-5791"],"issn-type":[{"value":"0143-991X","type":"print"},{"value":"1758-5791","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,8,6]]}}}