{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T01:57:03Z","timestamp":1781315823795,"version":"3.54.1"},"reference-count":49,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2023,9,25]],"date-time":"2023-09-25T00:00:00Z","timestamp":1695600000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U2033218, 61831018, 61802253"],"award-info":[{"award-number":["U2033218, 61831018, 61802253"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Shanghai Local Capacity Enhancement","award":["21010501500"],"award-info":[{"award-number":["21010501500"]}]},{"name":"Science and Technology Innovation Action Plan"},{"DOI":"10.13039\/501100003399","name":"Shanghai Science and Technology Commission","doi-asserted-by":"crossref","award":["21DZ1204900"],"award-info":[{"award-number":["21DZ1204900"]}],"id":[{"id":"10.13039\/501100003399","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Chenguang talented program of Shanghai","award":["17CG59"],"award-info":[{"award-number":["17CG59"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,2,29]]},"abstract":"<jats:p>\n            3D perception of depth and ego-motion is of vital importance in intelligent agent and\n            <jats:bold>Human Computer Interaction (HCI)<\/jats:bold>\n            tasks, such as robotics and autonomous driving. There are different kinds of sensors that can directly obtain 3D depth information. However, the commonly used Lidar sensor is expensive, and the effective range of RGB-D cameras is limited. In the field of computer vision, researchers have done a lot of work on 3D perception. While traditional geometric algorithms require a lot of manual features for depth estimation, Deep Learning methods have achieved great success in this field. In this work, we proposed a novel self-supervised method based on\n            <jats:bold>Vision Transformer (ViT)<\/jats:bold>\n            with\n            <jats:bold>Convolutional Neural Network (CNN)<\/jats:bold>\n            architecture, which is referred to as\n            <jats:bold>ViT-Depth<\/jats:bold>\n            . The image reconstruction losses computed by the estimated depth and motion between adjacent frames are treated as supervision signal to establish a self-supervised learning pipeline. This is an effective solution for tasks that need accurate and low-cost 3D perception, such as autonomous driving, robotic navigation, 3D reconstruction, and so on. Our method could leverage both the ability of CNN and Transformer to extract deep features and capture global contextual information. In addition, we propose a cross-frame loss that could constrain photometric error and scale consistency among multi-frames, which lead the training process to be more stable and improve the performance. Extensive experimental results on autonomous driving dataset demonstrate the proposed approach is competitive with the state-of-the-art depth and motion estimation methods.\n          <\/jats:p>","DOI":"10.1145\/3588571","type":"journal-article","created":{"date-parts":[[2023,3,23]],"date-time":"2023-03-23T12:23:57Z","timestamp":1679574237000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Self-Supervised Learning of Depth and Ego-Motion for 3D Perception in Human Computer Interaction"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9580-2641","authenticated-orcid":false,"given":"Shanbao","family":"Qiao","sequence":"first","affiliation":[{"name":"Shanghai University of Engineering Science"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0394-4635","authenticated-orcid":false,"given":"Neal N.","family":"Xiong","sequence":"additional","affiliation":[{"name":"Sul Ross State University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9930-0502","authenticated-orcid":false,"given":"Yongbin","family":"Gao","sequence":"additional","affiliation":[{"name":"Shanghai University of Engineering Science"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8563-5678","authenticated-orcid":false,"given":"Zhijun","family":"Fang","sequence":"additional","affiliation":[{"name":"Shanghai University of Engineering Science"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-4880-9339","authenticated-orcid":false,"given":"Wenjun","family":"Yu","sequence":"additional","affiliation":[{"name":"Shanghai University of Engineering Science"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1558-1656","authenticated-orcid":false,"given":"Juan","family":"Zhang","sequence":"additional","affiliation":[{"name":"Shanghai University of Engineering Science"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1946-576X","authenticated-orcid":false,"given":"Xiaoyan","family":"Jiang","sequence":"additional","affiliation":[{"name":"Shanghai University of Engineering Science"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,9,25]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"1","article-title":"Energy-quality scalable monocular depth estimation on low-power CPUs","volume":"99","author":"Cipolletta Antonio","year":"2021","unstructured":"Antonio Cipolletta, Valentino Peluso, Andrea Calimera, Matteo Poggi, and Stefano Mattoccia. 2021. Energy-quality scalable monocular depth estimation on low-power CPUs. IEEE Internet of Things Journal 99 (2021), 1\u20131.","journal-title":"IEEE Internet of Things Journal"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/JIOT.2018.2872435"},{"key":"e_1_3_1_4_2","article-title":"Attention is all you need","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. NIPS (2017).","journal-title":"NIPS"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2913372"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1177\/0278364913491297"},{"key":"e_1_3_1_7_2","unstructured":"Richard Hartley and Andrew Zisserman. 2000. Multiple View Geometry in Computer Vision . Cambridge University Press ISBN 0-521-62304-9 2000."},{"key":"e_1_3_1_8_2","article-title":"Depth map prediction from a single image using a multi-scale deep network","author":"Eigen David","year":"2014","unstructured":"David Eigen, Christian Puhrsch, and Rob Fergus. 2014. Depth map prediction from a single image using a multi-scale deep network. NIPS (2014).","journal-title":"NIPS"},{"key":"e_1_3_1_9_2","first-page":"3372","article-title":"A two-streamed network for estimating fine-scaled depth maps from single RGB images","author":"Li Jun","year":"2017","unstructured":"Jun Li, Reinhard Klein, and Angela Yao. 2017. A two-streamed network for estimating fine-scaled depth maps from single RGB images. CVPR (2017), 3372\u20133380.","journal-title":"CVPR"},{"key":"e_1_3_1_10_2","article-title":"Unsupervised learning of depth and ego-motion from video","author":"Zhou Tinghui","year":"2017","unstructured":"Tinghui Zhou, Matthew Brown, Noah Snavely, and David G. Lowe. 2017. Unsupervised learning of depth and ego-motion from video. CVPR (2017).","journal-title":"CVPR"},{"key":"e_1_3_1_11_2","doi-asserted-by":"crossref","unstructured":"Chaoyang Wang Jos\u00e9 Miguel Buenaposada Rui Zhu and Simon Lucey. 2018. Learning depth from monocular videos using direct methods. IEEE\/CVF Conference on Computer Vision and Pattern Recognition . 2022\u20132030.","DOI":"10.1109\/CVPR.2018.00216"},{"key":"e_1_3_1_12_2","unstructured":"Zhichao Yin and Jianping Shi. 2018. GeoNet: Unsupervised learning of dense depth optical flow and camera pose. IEEE\/CVF Conference on Computer Vision and Pattern Recognition . 1983\u20131992."},{"key":"e_1_3_1_13_2","doi-asserted-by":"crossref","unstructured":"Vincent Casser Soeren Pirk Reza Mahjourian and Anelia Angelova. 2019. Depth prediction without the sensors: Leveraging structure for unsupervised learning from monocular videos. In Proceedings of the Thirty-third AAAI Conference on Artificial Intelligence AAAI Honolulu Hawaii USA 27 January\u20131 February 2019; AAAI Press: Menlo Park CA USA. 8001\u20138008.","DOI":"10.1609\/aaai.v33i01.33018001"},{"key":"e_1_3_1_14_2","article-title":"Unsupervised monocular depth learning in dynamic scenes","author":"Gordon Ariel","year":"2020","unstructured":"Hanhan Li, Ariel Gordon, Hang Zhao, Vincent Casser, and Anelia Angelova. 2020. Unsupervised monocular depth learning in dynamic scenes, arXiv preprint arXiv: 2010.16404.","journal-title":"arXiv preprint"},{"key":"e_1_3_1_15_2","first-page":"35","article-title":"Unsupervised scale-consistent depth and ego-motion learning from monocular video","volume":"32","author":"Bian Jiawang","year":"2019","unstructured":"Jiawang Bian, Zhichao Li, Naiyan Wang, Huangying Zhan, Chunhua Shen, Ming-Ming Cheng, and Ian Reid. 2019. Unsupervised scale-consistent depth and ego-motion learning from monocular video. Advances in Neural Information Processing Systems (NeurIPS) 32 (2019), 35\u201345.","journal-title":"Advances in Neural Information Processing Systems (NeurIPS)"},{"key":"e_1_3_1_16_2","first-page":"3827","article-title":"Digging into self-supervised monocular depth estimation","author":"Godard Cl\u00e9ment","year":"2019","unstructured":"Cl\u00e9ment Godard, Oisin Mac Aodha, Michael Firman, and Gabriel J. Brostow. 2019. Digging into self-supervised monocular depth estimation. ICCV (2019), 3827\u20133837.","journal-title":"ICCV"},{"key":"e_1_3_1_17_2","article-title":"An image is worth 16x16 words: Transformers for image recognition at scale","author":"Dosovitskiy Alexey","year":"2020","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929.","journal-title":"arXiv preprint"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2020.3020583"},{"key":"e_1_3_1_19_2","first-page":"2002","article-title":"Deep ordinal regression network for monocular depth estimation","author":"Huan Fu","year":"2018","unstructured":"Fu Huan, Mingming Gong, Chaohui Wang, Kayhan Batmanghelich, and Dacheng Tao. 2018. Deep ordinal regression network for monocular depth estimation. CVPR (2018), 2002\u20132011.","journal-title":"CVPR"},{"key":"e_1_3_1_20_2","article-title":"Deep virtual stereo odometry: Leveraging deep depth prediction for monocular direct sparse odometry","author":"Yang Nan","year":"2018","unstructured":"Nan Yang, Rui Wang, Jorg Stuckler, and Daniel Cremers. 2018. Deep virtual stereo odometry: Leveraging deep depth prediction for monocular direct sparse odometry. ECCV.","journal-title":"ECCV"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2015.2505283"},{"key":"e_1_3_1_22_2","unstructured":"Jogendra Nath Kundu Phani Krishna Uppala Anuj Pahuja and R. Venkatesh Babu. 2018. AdaDepth: Unsupervised content congruent adaptation for depth estimation. IEEE\/CVF Conference on Computer Vision and Pattern Recognition . 2656\u20132665."},{"key":"e_1_3_1_23_2","doi-asserted-by":"crossref","unstructured":"Yevhen Kuznietsov Jorg Stuckler and Bastian Leibe 2017. Semi-supervised deep learning for monocular depth map prediction. IEEE Conference on Computer Vision & Pattern Recognition .","DOI":"10.1109\/CVPR.2017.238"},{"key":"e_1_3_1_24_2","doi-asserted-by":"crossref","unstructured":"Ishit Mehta Parikshit Sakurikar and P. J. Narayanan. 2018. Structured adversarial training for unsupervised monocular depth estimation. International Conference on 3D Vision (3DV) . 314\u2013323.","DOI":"10.1109\/3DV.2018.00044"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.12257"},{"key":"e_1_3_1_26_2","doi-asserted-by":"crossref","unstructured":"Reza Mahjourian Martin Wicke and Anelia Angelova. 2018. Unsupervised learning of depth and ego-motion from monocular video using 3D geometric constraints. IEEE\/CVF Conference on Computer Vision and Pattern Recognition . 5667\u20135675.","DOI":"10.1109\/CVPR.2018.00594"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2930258"},{"key":"e_1_3_1_28_2","doi-asserted-by":"crossref","unstructured":"Yue Meng Yongxi Lu Aman Raj Samuel Sunarjo Rui Guo Tara Javidi Gaurav Bansal and Dinesh Bharadia. 2019. SIGNet: Semantic instance aided unsupervised 3D geometry perception. IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 9802\u20139812.","DOI":"10.1109\/CVPR.2019.01004"},{"key":"e_1_3_1_29_2","doi-asserted-by":"crossref","unstructured":"Ariel Gordon Hanhan Li Rico Jonschkowski and Anelia Angelova. 2019. Depth from videos in the wild: Unsupervised monocular depth learning from unknown cameras. IEEE\/CVF International Conference on Computer Vision (ICCV) . 8976\u20138985.","DOI":"10.1109\/ICCV.2019.00907"},{"key":"e_1_3_1_30_2","doi-asserted-by":"crossref","unstructured":"Tianwei Shen Zixin Luo Lei Zhou Hanyu Deng Runze Zhang Tian Fang and Long Quan. 2019. Beyond photometric loss for self-supervised ego-motion estimation. International Conference on Robotics and Automation (ICRA) . 6359\u20136365.","DOI":"10.1109\/ICRA.2019.8793479"},{"key":"e_1_3_1_31_2","unstructured":"Yuliang Zou Zelun Luo and Jia-Bin Huang. 2018. DF-Net: Unsupervised joint learning of depth and flow using cross-task consistency. In Proc. 15th European Conference Munich Germany."},{"key":"e_1_3_1_32_2","first-page":"7627","article-title":"Learning single camera depth estimation using dual-pixels","author":"Garg Rahul","year":"2019","unstructured":"Rahul Garg, Neal Wadhwa, Sameer Ansari, and Jonathan T. Barron. 2019. Learning single camera depth estimation using dual-pixels. ICCV (2019), 7627\u20137636.","journal-title":"ICCV"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-021-01484-6"},{"key":"e_1_3_1_34_2","doi-asserted-by":"crossref","unstructured":"Iro Laina Christian Rupprecht Vasileios Belagiannis Federico Tombari and Nassir Navab. 2016. Deeper depth prediction with fully convolutional residual networks. 2016. Fourth International Conference on 3D Vision (3DV) . 239\u2013248.","DOI":"10.1109\/3DV.2016.32"},{"key":"e_1_3_1_35_2","first-page":"581","article-title":"Guiding monocular depth estimation using depth-attention volume","author":"Huynh Lam","year":"2020","unstructured":"Lam Huynh, Phong Nguyen-Ha, Jiri Matas, Esa Rahtu, and Janne Heikkil\u00e4. 2020. Guiding monocular depth estimation using depth-attention volume. European Conference on Computer Vision (ECCV). 581\u2013597.","journal-title":"European Conference on Computer Vision (ECCV)"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-018-1082-6"},{"key":"e_1_3_1_37_2","article-title":"Unsupervised CNN for single view depth estimation: Geometry to the rescue","author":"Garg Ravi","year":"2016","unstructured":"Ravi Garg, Vijay Kumar BG, Gustavo Carneiro, and Ian Reid. 2016. Unsupervised CNN for single view depth estimation: Geometry to the rescue. ECCV (2016).","journal-title":"ECCV"},{"key":"e_1_3_1_38_2","unstructured":"Kaiming He Georgia Gkioxari Piotr Doll\u00e1r and Ross Girshick. 2017. Mask R-CNN. IEEE International Conference on Computer Vision (ICCV) . 2980\u20132988."},{"key":"e_1_3_1_39_2","first-page":"60","article-title":"A non-local algorithm for image denoising","volume":"2","author":"Buades Antoni","year":"2005","unstructured":"Antoni Buades, Bartomeu Coll, and J.-M. Morel. 2005. A non-local algorithm for image denoising. IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR) 2 (2005), 60\u201365.","journal-title":"IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR)"},{"key":"e_1_3_1_40_2","doi-asserted-by":"crossref","unstructured":"Xiaolong Wang Ross Girshick Abhinav Gupta and Kaiming He. 2018. Non-local neural networks. In Proc. IEEE\/CVF Conf. Comput. Vis. Pattern Recognit . 7794\u20137803.","DOI":"10.1109\/CVPR.2018.00813"},{"key":"e_1_3_1_41_2","article-title":"Training data-efficient image transformers & distillation through attention","author":"Touvron Hugo","year":"2020","unstructured":"Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herv\u00e9 J\u00e9gou. 2020. Training data-efficient image transformers & distillation through attention. arXiv preprint arXiv:2012.12877.","journal-title":"arXiv preprint"},{"key":"e_1_3_1_42_2","article-title":"End-to-end video instance segmentation with transformers","author":"Wang Yuqing","year":"2020","unstructured":"Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia. 2020. End-to-end video instance segmentation with transformers. arXiv preprint arXiv:2011.14503.","journal-title":"arXiv preprint"},{"key":"e_1_3_1_43_2","doi-asserted-by":"crossref","unstructured":"Lin Huang Jianchao Tan Ji Liu and Junsong Yuan. 2020. Hand-transformer: Non-autoregressive structured modeling for 3D hand pose estimation. European Conference on Computer Vision 17\u201333.","DOI":"10.1007\/978-3-030-58595-2_2"},{"key":"e_1_3_1_44_2","unstructured":"Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep residual learning for image recognition. IEEE Conference on Computer Vision and Pattern Recognition (CVPR) ."},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_3_1_46_2","doi-asserted-by":"crossref","unstructured":"Huangying Zhan Ravi Garg Chamara Saroj Weerasekera Kejie Li Harsh Agarwal and Ian Reid. 2018. Unsupervised learning of monocular depth estimation and visual odometry with deep feature reconstruction. IEEE Conference on Computer Vision and Pattern Recognition (CVPR) .","DOI":"10.1109\/CVPR.2018.00043"},{"key":"e_1_3_1_47_2","first-page":"1281","article-title":"D3VO: Deep depth, deep pose and deep uncertainty for monocular visual odometry","author":"Yang Nan","year":"2020","unstructured":"Nan Yang, Lukas von Stumberg, Rui Wang, and Daniel Cremers. 2020. D3VO: Deep depth, deep pose and deep uncertainty for monocular visual odometry. CVPR (2020), 1281\u20131292.","journal-title":"CVPR"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2020.3011067"},{"key":"e_1_3_1_49_2","article-title":"VisioMap: Lightweight 3-D scene reconstruction toward natural indoor localization","author":"Li Feng","year":"2019","unstructured":"Feng Li, Jie Hao, Jin Wang, Jun Luo, Ying He, Dongxiao Yu, and Xiuzhen Cheng. 2019. VisioMap: Lightweight 3-D scene reconstruction toward natural indoor localization. IEEE Internet of Things Journal (2019).","journal-title":"IEEE Internet of Things Journal"},{"key":"e_1_3_1_50_2","first-page":"12232","article-title":"Competitive collaboration: Joint unsupervised learning of depth, camera motion, optical flow and motion segmentation","author":"Anurag Ranjan","year":"2019","unstructured":"Anurag Ranjan, Varun Jampani, Lukas Balles, Kihwan Kim, Deqing Sun, Jonas Wulff, and Michael J. Black. 2019. Competitive collaboration: Joint unsupervised learning of depth, camera motion, optical flow and motion segmentation. CVPR (2019), 12232\u201312241.","journal-title":"CVPR"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3588571","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3588571","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:47:13Z","timestamp":1750178833000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3588571"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,9,25]]},"references-count":49,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,2,29]]}},"alternative-id":["10.1145\/3588571"],"URL":"https:\/\/doi.org\/10.1145\/3588571","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,9,25]]},"assertion":[{"value":"2021-11-10","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-03-15","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-09-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}