{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,7]],"date-time":"2026-03-07T19:14:57Z","timestamp":1772910897499,"version":"3.50.1"},"reference-count":42,"publisher":"MIT Press","issue":"5","content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,4,18]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Visual navigation involves a movable robotic agent striving to reach a point goal (target location) using vision sensory input. While navigation with ideal visibility has seen plenty of success, it becomes challenging in suboptimal visual conditions like poor illumination, where traditional approaches suffer from severe performance degradation. We propose E3VN (echo-enhanced embodied visual navigation) to effectively perceive the surroundings even under poor visibility to mitigate this problem. This is made possible by adopting an echoer that actively perceives the environment via auditory signals. E3VN models the robot agent as playing a cooperative Markov game with that echoer. The action policies of robot and echoer are jointly optimized to maximize the reward in a two-stream actor-critic architecture. During optimization, the reward is also adaptively decomposed into the robot and echoer parts. Our experiments and ablation studies show that E3VN is consistently effective and robust in point goal navigation tasks, especially under nonideal visibility.<\/jats:p>","DOI":"10.1162\/neco_a_01579","type":"journal-article","created":{"date-parts":[[2023,3,21]],"date-time":"2023-03-21T22:26:10Z","timestamp":1679437570000},"page":"958-976","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":16,"title":["Echo-Enhanced Embodied Visual Navigation"],"prefix":"10.1162","volume":"35","author":[{"given":"Yinfeng","family":"Yu","sequence":"first","affiliation":[{"name":"Department of Computer Science and Technology, Tsinghua University, Beijing 100084, China"},{"name":"College of Information Science and Engineering, Xinjiang University, Urumqi 830046, China yyf17@mails.tsinghua.edu.cn"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lele","family":"Cao","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Technology, Tsinghua University, Beijing 100084, China"},{"name":"Motherbrain, EQT, Stockholm 11153, Sweden lele.cao@eqtpartners.com"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fuchun","family":"Sun","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Technology, Tsinghua University, Beijing 100084, China fcsun@mail.tsinghua.edu.cn"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chao","family":"Yang","sequence":"additional","affiliation":[{"name":"Shanghai AI Laboratory, Shanghai 200232, China yangchao@pjlab.org.cn"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Huicheng","family":"Lai","sequence":"additional","affiliation":[{"name":"College of Information Science and Engineering, Xinjiang University, Urumqi 830046, China lai@xju.edu.cn"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenbing","family":"Huang","sequence":"additional","affiliation":[{"name":"Institute for AI Industry Research, Tsinghua University, Beijing 100084, China hwenbing@126.com"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2023,4,18]]},"reference":[{"key":"2023041921580115200_B1","unstructured":"Anderson, P., Chang, A., Chaplot, D. S., Dosovitskiy, A., Gupta, S., Koltun, V., \u2026 Zamir, A. R. (2018). On evaluation of embodied navigation agents. arXiv:1807.06757."},{"key":"2023041921580115200_B2","first-page":"13072","article-title":"Context R-CNN: Long term temporal context for per-camera object detection","volume-title":"Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Beery","year":"2020"},{"key":"2023041921580115200_B3","doi-asserted-by":"crossref","unstructured":"Chang, A., Dai, A., Funkhouser, T., Halber, M., Niessner, M., Savva, M., \u2026 Zhang, Y. (2017). Matterport3D: Learning from RGB-D data in indoor environments. In Proceedings of the International Conference on 3D Vision.","DOI":"10.1109\/3DV.2017.00081"},{"key":"2023041921580115200_B4","article-title":"Learning to explore using active neural SLAM","volume-title":"Proceedings of the 8th International Conference on Learning Representations","author":"Chaplot","year":"2020"},{"key":"2023041921580115200_B5","doi-asserted-by":"crossref","unstructured":"Chen, C., Al-Halah, Z., & Grauman, K. (2021). Semantic audio-visual navigation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (pp. 15516\u201315525).","DOI":"10.1109\/CVPR46437.2021.01526"},{"key":"2023041921580115200_B6","first-page":"17","article-title":"Soundspaces: Audio-visual navigation in 3D environments","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Chen","year":"2020"},{"key":"2023041921580115200_B7","article-title":"Learning to set waypoints for audio-visual navigation","volume-title":"Proceedings of the 9th International Conference on Learning Representations","author":"Chen","year":"2021"},{"key":"2023041921580115200_B8","first-page":"11276","article-title":"Topological planning with transformers for vision-and-language navigation","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Chen","year":"2021"},{"key":"2023041921580115200_B9","unstructured":"Chen, L.-C., Papandreou, G., Schroff, F., & Adam, H. (2017). Rethinking atrous convolution for semantic image segmentation. arXiv:1706.05587."},{"key":"2023041921580115200_B10","first-page":"1581","article-title":"Batvision: Learning to see 3D spatial layout with two ears","volume-title":"Proceedings of the 2020 IEEE International Conference on Robotics and Automation","author":"Christensen","year":"2020"},{"key":"2023041921580115200_B11","article-title":"See, hear, explore: Curiosity via audio-visual association","volume-title":"Advances in neural information processing systems, 33","author":"Dean","year":"2020"},{"issue":"107","key":"2023041921580115200_B12","first-page":"1","article-title":"Beyond English-centric multilingual machine translation","volume":"22","author":"Fan","year":"2021","journal-title":"Journal of Machine Learning Research"},{"key":"2023041921580115200_B13","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2021.114741","article-title":"Polygonal coordinate system: Visualizing high-dimensional data using geometric DR, and a deterministic version of t-SNE","volume":"175","author":"Flexa","year":"2021","journal-title":"Expert Syst. Appl."},{"key":"2023041921580115200_B14","first-page":"10523","article-title":"Finding fallen objects via asynchronous audio-visual integration","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Gan","year":"2022"},{"key":"2023041921580115200_B15","first-page":"9701","article-title":"Look, listen, and act: Towards audio-visual embodied navigation","volume-title":"Proceedings of the 2020 IEEE International Conference on Robotics and Automation","author":"Gan","year":"2020"},{"key":"2023041921580115200_B16","first-page":"658","article-title":"VisualEchoes: Spatial image representation learning through echolocation","volume-title":"Proceedings of the 16th European ECCV Conference","author":"Gao","year":"2020"},{"key":"2023041921580115200_B17","first-page":"1022","article-title":"SplitNet: Sim2Sim and Task2Task transfer for embodied visual navigation","volume-title":"Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision","author":"Gordon","year":"2019"},{"issue":"10","key":"2023041921580115200_B18","doi-asserted-by":"publisher","first-page":"1272","DOI":"10.1109\/TPAMI.2004.88","article-title":"Modeling the space of camera response functions","volume":"26","author":"Grossberg","year":"2004","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"2023041921580115200_B19","first-page":"7272","article-title":"Cognitive mapping and planning for visual navigation","volume-title":"Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition","author":"Gupta","year":"2017"},{"key":"2023041921580115200_B20","first-page":"1643","article-title":"VLN BERT: A recurrent vision-and-language BERT for navigation","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Hong","year":"2021"},{"key":"2023041921580115200_B21","doi-asserted-by":"crossref","unstructured":"Irshad, M. Z., Ma, C., & Kira, Z. (2021). Hierarchical cross-modal agent for robotics vision-and-language navigation. In Proceedings of the IEEE International Conference on Robotics and Automation (pp. 13238\u201313246).","DOI":"10.1109\/ICRA48506.2021.9561806"},{"key":"2023041921580115200_B22","first-page":"2815","article-title":"Differentiable SLAM-net: Learning particle SLAM for visual navigation","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Karkus","year":"2021"},{"key":"2023041921580115200_B23","article-title":"Generative language-grounded policy in vision-and-language navigation with Bayes' rule","volume-title":"Proceedings of the 9th International Conference on Learning Representations","author":"Kurita","year":"2021"},{"key":"2023041921580115200_B24","article-title":"Learning to navigate in complex environments","volume-title":"Proceedings of the 5th International Conference on Learning Representations","author":"Mirowski","year":"2017"},{"issue":"2","key":"2023041921580115200_B25","doi-asserted-by":"publisher","first-page":"683","DOI":"10.1109\/LRA.2020.3048662","article-title":"Embodied visual navigation with automatic curriculum learning in real environments","volume":"6","author":"Morad","year":"2021","journal-title":"IEEE Robotics Autom. Lett."},{"key":"2023041921580115200_B26","first-page":"1183","article-title":"Audio-visual floorplan reconstruction","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Purushwalkam","year":"2021"},{"key":"2023041921580115200_B27","first-page":"13709","article-title":"Co-GAT: A co-interactive graph attention network for joint dialog act recognition and sentiment classification","volume-title":"Proceedings of the Thirty-Fifth AAAI Conference on Artificial Intelligence","author":"Qin","year":"2021"},{"key":"2023041921580115200_B28","first-page":"400","article-title":"Occupancy anticipation for efficient exploration and navigation","volume-title":"Proceedings of the 16th European Conference on Computer Vision","author":"Ramakrishnan","year":"2020"},{"key":"2023041921580115200_B29","first-page":"4292","article-title":"QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning","volume-title":"Proceedings of the 35th International Conference on Machine Learning","author":"Rashid","year":"2018"},{"key":"2023041921580115200_B30","first-page":"9338","article-title":"Habitat: A platform for embodied AI research","volume-title":"Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision","author":"Savva","year":"2019"},{"key":"2023041921580115200_B31","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv:1707.06347."},{"key":"2023041921580115200_B32","unstructured":"Straub, J., Whelan, T., Ma, L., Chen, Y., Wijmans, E., Green, S., Engel, J. J., \u2026 Newcombe, R. (2019). The replica dataset: A digital replica of indoor spaces. arXiv:1906.05797."},{"key":"2023041921580115200_B33","first-page":"2085","article-title":"Value-decomposition networks for cooperative multi-agent learning based on team reward","volume-title":"Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems","author":"Sunehag","year":"2018"},{"issue":"1","key":"2023041921580115200_B34","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3216722","article-title":"CloudNavi: Toward ubiquitous indoor navigation service with 3D point clouds","volume":"15","author":"Teng","year":"2019","journal-title":"ACM Transactions on Sensor Networks"},{"key":"2023041921580115200_B35","doi-asserted-by":"crossref","unstructured":"Tracy, E., & Kottege, N. (2021). CatChatter: Acoustic perception for mobile robots. IEEE Robotics and Automation Letters, 6(4),7209\u20137216.","DOI":"10.1109\/LRA.2021.3094492"},{"key":"2023041921580115200_B36","first-page":"5998","article-title":"Attention is all you need","volume-title":"Advances in neural information processing systems","author":"Vaswani","year":"2017"},{"key":"2023041921580115200_B37","first-page":"8455","article-title":"Structured scene memory for vision-language navigation","author":"Wang","year":"2021"},{"key":"2023041921580115200_B38","doi-asserted-by":"crossref","DOI":"10.1145\/3343031.3350983","article-title":"Progressive Retinex: Mutually reinforced illumination-noise perception network for low-light image enhancement","volume-title":"Proceedings of the 27th ACM International Conference on Multimedia","author":"Wang","year":"2019"},{"key":"2023041921580115200_B39","article-title":"DD-PPO: Learning near-perfect pointgoal navigators from 2.5 billion frames","volume-title":"Proceedings of the 8th International Conference on Learning Representations","author":"Wijmans","year":"2020"},{"key":"2023041921580115200_B40","unstructured":"Ye, J., Batra, D., Wijmans, E., & Das, A. (2020). Auxiliary tasks speed up learning pointgoal navigation. arXiv:2007.04561."},{"key":"2023041921580115200_B41","unstructured":"Yu, Y., Cao, L., Sun, F., Liu, X., & Wang, L. (2022). Pay self-attention to audio-visual navigation. arXiv:2210.01353."},{"key":"2023041921580115200_B42","article-title":"Sound adversarial audio-visual navigation","volume-title":"Proceedings of Tenth International Conference on Learning Representations","author":"Yu","year":"2022"}],"container-title":["Neural Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/35\/5\/958\/2079357\/neco_a_01579.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/35\/5\/958\/2079357\/neco_a_01579.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,4,19]],"date-time":"2023-04-19T21:58:49Z","timestamp":1681941529000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/neco\/article\/35\/5\/958\/115252\/Echo-Enhanced-Embodied-Visual-Navigation"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,4,18]]},"references-count":42,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2023,4,18]]},"published-print":{"date-parts":[[2023,4,18]]}},"URL":"https:\/\/doi.org\/10.1162\/neco_a_01579","relation":{},"ISSN":["0899-7667","1530-888X"],"issn-type":[{"value":"0899-7667","type":"print"},{"value":"1530-888X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2023,5]]},"published":{"date-parts":[[2023,4,18]]}}}