{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,13]],"date-time":"2026-03-13T12:07:00Z","timestamp":1773403620362,"version":"3.50.1"},"reference-count":45,"publisher":"MDPI AG","issue":"24","license":[{"start":{"date-parts":[[2022,12,9]],"date-time":"2022-12-09T00:00:00Z","timestamp":1670544000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation of China (NSFC), Essential projects","award":["U2033218"],"award-info":[{"award-number":["U2033218"]}]},{"name":"National Natural Science Foundation of China (NSFC), Essential projects","award":["61831018"],"award-info":[{"award-number":["61831018"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Robust and accurate visual feature tracking is essential for good pose estimation in visual odometry. However, in fast-moving scenes, feature point extraction and matching are unstable because of blurred images and large image disparity. In this paper, we propose an unsupervised monocular visual odometry framework based on a fusion of features extracted from two sources, that is, the optical flow network and the traditional point feature extractor. In the training process, point features are generated for scene images and the outliers of matched point pairs are filtered by FlannMatch. Meanwhile, the optical flow network constrained by the principle of forward\u2013backward flow consistency is used to select another group of corresponding point pairs. The Euclidean distance between the matching points found by FlannMatch and the corresponding point pairs by the flow network is added to the loss function of the flow network. Compared with SURF, the trained flow network shows more robust performance in complicated fast-motion scenarios. Furthermore, we propose the AvgFlow estimation module, which selects one group of the matched point pairs generated by the two methods according to the scene motion. The camera pose is then recovered by Perspective-n-Point (PnP) or the epipolar geometry. Experiments conducted on the KITTI Odometry dataset verify the effectiveness of the trajectory estimation of our approach, especially in fast-moving scenarios.<\/jats:p>","DOI":"10.3390\/s22249647","type":"journal-article","created":{"date-parts":[[2022,12,9]],"date-time":"2022-12-09T03:59:46Z","timestamp":1670558386000},"page":"9647","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Unsupervised Monocular Visual Odometry for Fast-Moving Scenes Based on Optical Flow Network with Feature Point Matching Constraint"],"prefix":"10.3390","volume":"22","author":[{"given":"Yuji","family":"Zhuang","sequence":"first","affiliation":[{"name":"School of Electronic and Electrical Engineering, Shanghai University of Engineering Science, Shanghai 201600, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaoyan","family":"Jiang","sequence":"additional","affiliation":[{"name":"School of Electronic and Electrical Engineering, Shanghai University of Engineering Science, Shanghai 201600, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yongbin","family":"Gao","sequence":"additional","affiliation":[{"name":"School of Electronic and Electrical Engineering, Shanghai University of Engineering Science, Shanghai 201600, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhijun","family":"Fang","sequence":"additional","affiliation":[{"name":"School of Electronic and Electrical Engineering, Shanghai University of Engineering Science, Shanghai 201600, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5256-210X","authenticated-orcid":false,"given":"Hamido","family":"Fujita","sequence":"additional","affiliation":[{"name":"Faculty of Information Technology, HUTECH University, Ho Chi Minh City, Vietnam"},{"name":"i-SOMET Inc., Morioka 020-0104, Japan"},{"name":"Regional Research Center, Iwate Prefectural University, Takizawa 020-0693, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,12,9]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1109\/MRA.2006.1678144","article-title":"Simultaneous localization and mapping: Part I","volume":"13","author":"Bailey","year":"2006","journal-title":"IEEE Robot. Autom. Mag."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"Distinctive image features from scale-invariant key points","volume":"60","author":"Lowe","year":"2003","journal-title":"Int. J. Comput. Vis. IJCV"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Bay, H., Tuytelaars, T., and Gool, L.V. (2006, January 7\u201313). Surf: Speeded up robust features. Proceedings of the European Conference on Computer Vision (ECCV), Graz, Austria.","DOI":"10.1007\/11744023_32"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1255","DOI":"10.1109\/TRO.2017.2705103","article-title":"Orb-slam2: An open-source slam system for monocular, stereo, and rgb-d cameras","volume":"33","year":"2017","journal-title":"IEEE Trans. Robot. TRO"},{"key":"ref_5","unstructured":"Bian, J., Li, Z., Wang, N., Zhan, H., Shen, C., Cheng, M.M., and Reid, I. (2019). Unsupervised scale-consistent depth and ego-motion learning from monocular video. Adv. Neural Inf. Process. Syst. NeurIPS, 32."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Zhao, W., Liu, S., Shu, Y., and Liu, Y.J. (2020, January 14\u201319). Towards better generalization: Joint depth-pose learning without posenet. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR42600.2020.00917"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Hartley, R., and Zisserman, A. (2003). Multiple View Geometry in Computer Vision, Cambridge University Press.","DOI":"10.1017\/CBO9780511811685"},{"key":"ref_8","first-page":"80","article-title":"Visual odometry: Part I: The first 30 years and fundamentals","volume":"18","author":"Davide","year":"2011","journal-title":"IEEE Robot. Autom. Mag."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Klein, G., and Murray, D. (2007, January 13\u201316). Parallel Tracking and Mapping for Small AR Workspaces. Proceedings of the IEEE and ACM International Symposium on Mixed and Augmented Reality, Washington, DC, USA.","DOI":"10.1109\/ISMAR.2007.4538852"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Rublee, E., Rabaud, V., Konolige, K., and Bradski, G. (2011, January 6\u201313). ORB: An efficient alternative to SIFT or SURF. Proceedings of the International Conference on Computer Vision (ICCV), Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126544"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"611","DOI":"10.1109\/TPAMI.2017.2658577","article-title":"Direct sparse odometry","volume":"40","author":"Engel","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell. PAMI"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Zhou, L., Huang, G., Mao, Y., Wang, S., and Kaess, M. (2022, January 23\u201327). EDPLVO: Efficient Direct Point-Line Visual Odometry. Proceedings of the International Conference on Robotics and Automation (ICRA), Philadelphia, PA, USA.","DOI":"10.1109\/ICRA46639.2022.9812133"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Tian, R., Zhang, Y., Zhu, D., Liang, S., Coleman, S., and Kerr, D. (2021, January 23\u201327). Accurate and robust scale recovery for monocular visual odometry based on plane geometry. Proceedings of the International Conference on Robotics and Automation (ICRA), Xi\u2019an, China.","DOI":"10.1109\/ICRA48506.2021.9561215"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"2803","DOI":"10.1109\/LRA.2022.3142900","article-title":"MSC-VO: Exploiting Manhattan and Structural Constraints for Visual Odometry","volume":"7","author":"Ortiz","year":"2022","journal-title":"IEEE Robot. Autom. Lett. RAL"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"15844","DOI":"10.1109\/ACCESS.2018.2810849","article-title":"Improvement of Generalization Ability of Deep CNN via Implicit Regularization in Two-Stage Training Process","volume":"6","author":"Zheng","year":"2018","journal-title":"IEEE Access"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"102048","DOI":"10.1016\/j.media.2021.102048","article-title":"Faster Mean-shift: GPU-accelerated clustering for cosine embedding-based cell segmentation and tracking","volume":"71","author":"Zhao","year":"2021","journal-title":"Med. Image Anal."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Yao, T., Qu, C., Liu, Q., Deng, R., Tian, Y., Xu, J., Jha, A., Bao, S., Zhao, M., and Fogo, A.B. (2021, January 1). Compound figure separation of biomedical images with side loss. Proceedings of the Deep Generative Models, and Data Augmentation, Labelling, and Imperfections: First Workshop, DGM4MICCAI 2021, and First Workshop, DALI 2021, Strasbourg, France.","DOI":"10.1007\/978-3-030-88210-5_16"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"21780","DOI":"10.1109\/JSEN.2022.3197235","article-title":"Pseudo RGB-D Face Recognition","volume":"22","author":"Jin","year":"2022","journal-title":"IEEE Sensors J."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zhou, T., Brown, M., Snavely, N., and Lowe, D.G. (2017, January 21\u201326). Unsupervised learning of depth and ego-motion from video. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.700"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Zhan, H., Garg, R., Weerasekera, C.S., Li, K., Agarwal, H., and Reid, I. (2018, January 18\u201323). Unsupervised learning of monocular depth estimation and visual odometry with deep feature reconstruction. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00043"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Ranjan, A., Jampani, V., Balles, L., Kim, K., Sun, D., Wulff, J., and Black, M.J. (2019, January 15\u201320). Competitive collaboration: Joint unsupervised learning of depth, camera motion, optical flow and motion segmentation. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01252"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Li, S., Wang, X., Cao, Y., Xue, F., Yan, Z., and Zha, H. (2020, January 14\u201319). Self-supervised deep visual odometry with online adaptation. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00637"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Teed, Z., and Deng, J. (2020, January 23\u201328). Raft: Recurrent all-pairs field transforms for optical flow. Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK.","DOI":"10.1007\/978-3-030-58536-5_24"},{"key":"ref_24","unstructured":"Wang, W., Hu, Y., and Scherer, S. (2021, January 8\u201311). Tartanvo: A generalizable learning-based vo. Proceedings of the Conference on Robot Learning (CoRL), London, UK."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Kuo, X.Y., Liu, C., Lin, K.C., and Lee, C.Y. (2020, January 14\u201319). Dynamic attention-based visual odometry. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPRW50498.2020.00026"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Wang, C., Wang, Y.P., and Manocha, D. (2022, January 23\u201327). Motionhint: Self-supervised monocular visual odometry with motion constraints. Proceedings of the International Conference on Robotics and Automation (ICRA), Philadelphia, PA, USA.","DOI":"10.1109\/ICRA46639.2022.9812288"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Yin, Z., and Shi, J. (2018, January 18\u201323). Geonet: Unsupervised learning of dense depth, optical flow and camera pose. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00212"},{"key":"ref_28","first-page":"331","article-title":"Fast approximate nearest neighbors with automatic algorithm configuration","volume":"2","author":"Muja","year":"2009","journal-title":"Int. Conf. Comput. Vis. Theory Appl."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"155","DOI":"10.1007\/s11263-008-0152-6","article-title":"Epnp: An accurate o (n) solution to the pnp problem","volume":"81","author":"Lepetit","year":"2009","journal-title":"Int. J. Comput. Vis. IJCV"},{"key":"ref_30","unstructured":"Godard, C., Mac Aodha, O., Firman, M., and Brostow, G.J. (November, January 27). Digging into self-supervised monocular depth estimation. Proceedings of the International Conference on Computer Vision (ICCV), Seoul, Republic of Korea."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Nekrasov, V., Dharmasiri, T., Spek, A., Drummond, T., Shen, C., and Reid, I. (2019, January 20\u201324). Real-time joint semantic segmentation and depth estimation using asymmetric annotations. Proceedings of the International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada.","DOI":"10.1109\/ICRA.2019.8794220"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Fu, H., Gong, M., Wang, C., Batmanghelich, K., and Tao, D. (2018, January 18\u201323). Deep ordinal regression network for monocular depth estimation. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00214"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1109\/TIP.2003.819861","article-title":"Image quality assessment: From error visibility to structural similarity","volume":"13","author":"Wang","year":"2004","journal-title":"IEEE Trans Image Process TIP"},{"key":"ref_34","unstructured":"Nister, D. (2003, January 16\u201322). An efficient solution to the five-point relative pose problem. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Madison, WI, USA."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"161","DOI":"10.1023\/A:1007941100561","article-title":"Determining the epipolar geometry and its uncertainty: A review","volume":"27","author":"Zhang","year":"1998","journal-title":"Int. J. Comput. Vis. IJCV"},{"key":"ref_36","unstructured":"Hartley, R.I. (1995, January 20\u201323). In defence of the 8-point algorithm. Proceedings of the International Conference on Computer Vision (ICCV), Seoul, Republic of Korea."},{"key":"ref_37","unstructured":"Bian, J.W., Wu, Y.H., Zhao, J., Liu, Y., Zhang, L., Cheng, M.M., and Reid, I. (2019, January 9\u201312). An evaluation of feature matchers for fundamental matrix estimation. Proceedings of the British Machine Vision Conference (BMVC), Cardiff, UK."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Li, S., Wu, X., Cao, Y., and Zha, H. (2021, January 19\u201325). Generalizing to the open world: Deep visual odometry with online adaptation. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR46437.2021.01298"},{"key":"ref_39","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., and Antiga, L. (2019, January 8\u201314). PyTorch: An Imperative Style, High-Performance Deep Learning Library. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada."},{"key":"ref_40","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A Method for Stochastic Optimization. arXiv."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"1231","DOI":"10.1177\/0278364913491297","article-title":"Vision meets robotics: The kitti dataset","volume":"32","author":"Geiger","year":"2013","journal-title":"Int. J. Robot. Res."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Calonder, M., Lepetit, V., Strecha, C., and Fua, P. (2010, January 5\u201311). Brief: Binary robust independent elementary features. Proceedings of the European Conference on Computer Vision (ECCV), Heraklion, Greece.","DOI":"10.1007\/978-3-642-15561-1_56"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Noble, F.K. (2016, January 28\u201330). Comparison of OpenCV\u2019s feature detectors and feature matchers. Proceedings of the International Conference on Mechatronics and Machine Vision in Practice, Nanjing, China.","DOI":"10.1109\/M2VIP.2016.7827292"},{"key":"ref_44","unstructured":"Wang, S., Clark, R., Wen, H., and Trigoni, N. (June, January 29). Deepvo: Towards end-to-end visual odometry with deep recurrent convolutional neural networks. Proceedings of the International Conference on Robotics and Automation (ICRA), Singapore."},{"key":"ref_45","unstructured":"Liang, Z., Wang, Q., and Yu, Y. (October, January 27). Deep Unsupervised Learning Based Visual Odometry with Multi-scale Matching and Latent Feature Constraint. Proceedings of the International Conference on Intelligent Robots and Systems (IROS), Prague, Czech Republic."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/24\/9647\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:36:57Z","timestamp":1760146617000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/24\/9647"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,9]]},"references-count":45,"journal-issue":{"issue":"24","published-online":{"date-parts":[[2022,12]]}},"alternative-id":["s22249647"],"URL":"https:\/\/doi.org\/10.3390\/s22249647","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12,9]]}}}