{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T16:09:20Z","timestamp":1784736560885,"version":"3.55.0"},"reference-count":85,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2020,8,31]],"date-time":"2020-08-31T00:00:00Z","timestamp":1598832000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,8,31]],"date-time":"2020-08-31T00:00:00Z","timestamp":1598832000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100010577","name":"Toyota Motor Europe","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100010577","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2021,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The role of robots in society keeps expanding, bringing with it the necessity of interacting and communicating with humans. In order to keep such interaction intuitive, we provide automatic wayfinding based on verbal navigational instructions. Our first contribution is the creation of a large-scale dataset with verbal navigation instructions. To this end, we have developed an interactive visual navigation environment based on Google Street View; we further design an annotation method to highlight mined anchor landmarks and local directions between them in order to help annotators formulate typical, human references to those. The annotation task was crowdsourced on the AMT platform, to construct a new Talk2Nav dataset with 10,\u00a0714 routes. Our second contribution is a new learning method. Inspired by spatial cognition research on the mental conceptualization of navigational instructions, we introduce a soft dual attention mechanism defined over the segmented language instructions to jointly extract two partial instructions\u2014one for matching the next upcoming visual landmark and the other for matching the local directions to the next landmark. On the similar lines, we also introduce spatial memory scheme to encode the local directional transitions. Our work takes advantage of the advance in two lines of research: mental formalization of verbal navigational instructions and training neural network agents for automatic way finding. Extensive experiments show that our method significantly outperforms previous navigation methods. For demo video, dataset and code, please refer to our<jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/www.trace.ethz.ch\/publications\/2019\/talk2nav\/index.html\">project page<\/jats:ext-link>.<\/jats:p>","DOI":"10.1007\/s11263-020-01374-3","type":"journal-article","created":{"date-parts":[[2020,8,31]],"date-time":"2020-08-31T04:03:31Z","timestamp":1598846611000},"page":"246-266","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":46,"title":["Talk2Nav: Long-Range Vision-and-Language Navigation with Dual Attention and Spatial Memory"],"prefix":"10.1007","volume":"129","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5409-0780","authenticated-orcid":false,"given":"Arun Balajee","family":"Vasudevan","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dengxin","family":"Dai","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Luc","family":"Van Gool","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2020,8,31]]},"reference":[{"issue":"1","key":"1374_CR1","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1007\/s11263-016-0966-6","volume":"123","author":"A Agrawal","year":"2017","unstructured":"Agrawal, A., Lu, J., Antol, S., Mitchell, M., Zitnick, C. L., Parikh, D., et al. (2017). Vqa: Visual question answering. International Journal of Computer Vision, 123(1), 4\u201331.","journal-title":"International Journal of Computer Vision"},{"key":"1374_CR2","unstructured":"Anderson, P., Chang, A., Chaplot, D. S., Dosovitskiy, A., Gupta, S., Koltun, V., Kosecka, J., Malik, J., Mottaghi, R., Savva, M., et\u00a0al. (2018). On evaluation of embodied navigation agents. arXiv:1807.06757."},{"key":"1374_CR3","doi-asserted-by":"crossref","unstructured":"Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., & Zhang, L. (2018). Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 6077\u20136086).","DOI":"10.1109\/CVPR.2018.00636"},{"key":"1374_CR4","doi-asserted-by":"crossref","unstructured":"Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., S\u00fcnderhauf, N., Reid, I., Gould, S., & van\u00a0den Hengel, A. (2018). Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In Proceedings of the IEEE conference on computer vision and pattern recognition.","DOI":"10.1109\/CVPR.2018.00387"},{"key":"1374_CR5","doi-asserted-by":"crossref","unstructured":"Andreas, J., Rohrbach, M., Darrell, T., & Klein, D. (2016). Learning to compose neural networks for question answering. arXiv preprint arXiv:1601.01705.","DOI":"10.18653\/v1\/N16-1181"},{"key":"1374_CR6","doi-asserted-by":"crossref","unstructured":"Aneja, J., Deshpande, A., & Schwing, A. G. (2018). Convolutional image captioning. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 5561\u20135570).","DOI":"10.1109\/CVPR.2018.00583"},{"key":"1374_CR7","doi-asserted-by":"crossref","unstructured":"Anne\u00a0Hendricks, L., Wang, O., Shechtman, E., Sivic, J., Darrell, T., & Russell, B. (2017). Localizing moments in video with natural language. In Proceedings of the IEEE international conference on computer vision.","DOI":"10.1109\/ICCV.2017.618"},{"key":"1374_CR8","unstructured":"Bahdanau, D., Cho, K., & Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473."},{"key":"1374_CR9","doi-asserted-by":"crossref","unstructured":"Balajee Vasudevan, A., Dai, D., & Van Gool, L. (2018). Object referring in videos with language and human gaze. In Conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2018.00434"},{"issue":"3","key":"1374_CR10","first-page":"1","volume":"6","author":"EM Bender","year":"2011","unstructured":"Bender, E. M. (2011). On achieving and evaluating language-independence in NLP. Linguistic Issues in Language Technology, 6(3), 1\u201326.","journal-title":"Linguistic Issues in Language Technology"},{"key":"1374_CR11","doi-asserted-by":"publisher","first-page":"587","DOI":"10.1162\/tacl_a_00041","volume":"6","author":"EM Bender","year":"2018","unstructured":"Bender, E. M., & Friedman, B. (2018). Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics, 6, 587\u2013604.","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"1374_CR12","doi-asserted-by":"crossref","unstructured":"Boularias, A., Duvallet, F., Oh, J., & Stentz, A. (2015). Grounding spatial relations for outdoor robot navigation. In IEEE international conference on robotics and automation (ICRA).","DOI":"10.1109\/ICRA.2015.7139457"},{"key":"1374_CR13","doi-asserted-by":"crossref","unstructured":"Brahmbhatt, S., & Hays, J. (2017). Deepnav: Learning to navigate large cities. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 5193\u20135202).","DOI":"10.1109\/CVPR.2017.329"},{"key":"1374_CR14","doi-asserted-by":"crossref","unstructured":"Chen, D. L., & Mooney, R. J. (2011). Learning to interpret natural language navigation instructions from observations. In Twenty-fifth AAAI conference on artificial intelligence.","DOI":"10.1609\/aaai.v25i1.7974"},{"key":"1374_CR15","doi-asserted-by":"crossref","unstructured":"Chen, H., Shur, A., Misra, D., Snavely, N., & Artzi, Y. (2019). Touchdown: Natural language navigation and spatial reasoning in visual street environments. In Proceedings of the IEEE conference on computer vision and pattern recognition.","DOI":"10.1109\/CVPR.2019.01282"},{"key":"1374_CR16","doi-asserted-by":"crossref","unstructured":"Chopra, S. (2005). Learning a similarity metric discriminatively, with application to face verification. In IEEE conference on compter vision and pattern recognition.","DOI":"10.1109\/CVPR.2005.202"},{"key":"1374_CR17","doi-asserted-by":"crossref","unstructured":"Coors, B., Paul\u00a0Condurache, A., & Geiger, A.(2018). Spherenet: Learning spherical representations for detection and classification in omnidirectional images. In Proceedings of the European conference on computer vision (ECCV) (pp. 518\u2013533).","DOI":"10.1007\/978-3-030-01240-3_32"},{"key":"1374_CR18","doi-asserted-by":"crossref","unstructured":"Das, A., Datta, S., Gkioxari, G., Lee, S., Parikh, D., & Batra, D. (2018). Embodied question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR)","DOI":"10.1109\/CVPR.2018.00008"},{"key":"1374_CR19","unstructured":"de\u00a0Vries, H., Shuster, K., Batra, D., Parikh, D., Weston, J., & Kiela, D. (2018). Talk the walk: Navigating new york city through grounded dialogue. arXiv preprint arXiv:1807.03367."},{"key":"1374_CR20","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L. J., Li, K., & Fei-Fei, L. (2009). Imagenet: A large-scale hierarchical image database. In IEEE conference on computer vision and pattern recognition.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"1374_CR21","doi-asserted-by":"crossref","unstructured":"Deruyttere, T., Vandenhende, S., Grujicic, D., Van\u00a0Gool, L., & Moens, M. F. (2019). Talk2car: Taking control of your self-driving car. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP) (pp. 2088\u20132098).","DOI":"10.18653\/v1\/D19-1215"},{"key":"1374_CR22","unstructured":"Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805."},{"key":"1374_CR23","doi-asserted-by":"crossref","unstructured":"Donahue, J., Anne\u00a0Hendricks, L., Guadarrama, S., Rohrbach, M., Venugopalan, S., Saenko, K., & Darrell, T. (2015). Long-term recurrent convolutional networks for visual recognition and description. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 2625\u20132634).","DOI":"10.1109\/CVPR.2015.7298878"},{"key":"1374_CR24","unstructured":"Fried, D., Hu, R., Cirik, V., Rohrbach, A., Andreas, J., Morency, L. P., Berg-Kirkpatrick, T., Saenko, K., Klein, D., & Darrell, T. (2018). Speaker-follower models for vision-and-language navigation. In NIPS."},{"key":"1374_CR25","unstructured":"Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press. Retrieved March 2019 from, http:\/\/www.deeplearningbook.org."},{"key":"1374_CR26","doi-asserted-by":"crossref","unstructured":"Gordon, D., Kembhavi, A., Rastegari, M., Redmon, J., Fox, D., & Farhadi, A. (2018). Iqa: Visual question answering in interactive environments. In Proceedings of the IEEE conference on computer vision and pattern recognition.","DOI":"10.1109\/CVPR.2018.00430"},{"key":"1374_CR27","doi-asserted-by":"crossref","unstructured":"Grabler, F., Agrawala, M., Sumner, R. W., & Pauly, M. (2008). Automatic generation of tourist maps. In ACM SIGGRAPH.","DOI":"10.1145\/1399504.1360699"},{"key":"1374_CR28","unstructured":"Graves, A. (2016). Adaptive computation time for recurrent neural networks. arXiv preprint arXiv:1603.08983."},{"key":"1374_CR29","unstructured":"Graves, A., Wayne, G., & Danihelka, I. (2014). Neural turing machines. arXiv:1410.5401."},{"issue":"7626","key":"1374_CR30","doi-asserted-by":"publisher","first-page":"471","DOI":"10.1038\/nature20101","volume":"538","author":"A Graves","year":"2016","unstructured":"Graves, A., Wayne, G., Reynolds, M., Harley, T., Danihelka, I., Grabska-Barwi\u0144ska, A., et al. (2016). Hybrid computing using a neural network with dynamic external memory. Nature, 538(7626), 471.","journal-title":"Nature"},{"key":"1374_CR31","doi-asserted-by":"crossref","unstructured":"Gupta, S., Tolani, V., Davidson, J., Levine, S., Sukthankar, R., & Malik, J. (2019). Cognitive mapping and planning for visual navigation. In International journal of computer vision.","DOI":"10.1007\/s11263-019-01236-7"},{"key":"1374_CR32","doi-asserted-by":"crossref","unstructured":"Gygli, M., Song, Y., & Cao, L. (2016). Video2gif: Automatic generation of animated gifs from video. In Proceedings of the IEEE conference on computer vision and pattern recognition.","DOI":"10.1109\/CVPR.2016.114"},{"key":"1374_CR33","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition.","DOI":"10.1109\/CVPR.2016.90"},{"key":"1374_CR34","doi-asserted-by":"crossref","unstructured":"Hecker, S., Dai, D., & Van Gool, L. (2018). End-to-end learning of driving models with surround-view cameras and route planners. In European conference on computer vision (ECCV).","DOI":"10.1007\/978-3-030-01234-2_27"},{"key":"1374_CR35","unstructured":"Hecker, S., Dai, D., & Van Gool, L. (2019). Learning accurate, comfortable and human-like driving. In arXiv-1903.10995."},{"key":"1374_CR36","unstructured":"Hermann, K. M., Hill, F., Green, S., Wang, F., Faulkner, R., Soyer, H., Szepesvari, D., Czarnecki, W., Jaderberg, M., Teplyashin, D., Wainwright, M., Apps, C., Hassabis, D., & Blunsom, P. (2017). Grounded language learning in a simulated 3d world. CoRR abs\/1706.06551."},{"key":"1374_CR37","doi-asserted-by":"crossref","unstructured":"Hermann, K. M., Malinowski, M., Mirowski, P., Banki-Horvath, A., Anderson, K., & Hadsell, R. (2019). Learning To follow directions in street view. arXiv e-prints.","DOI":"10.1609\/aaai.v34i07.6849"},{"key":"1374_CR38","unstructured":"Hill, F., Hermann, K. M., Blunsom, P., & Clark, S. (2017). Understanding grounded language learning agents. arXiv e-prints."},{"issue":"8","key":"1374_CR39","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","volume":"9","author":"S Hochreiter","year":"1997","unstructured":"Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735\u20131780.","journal-title":"Neural Computation"},{"issue":"2","key":"1374_CR40","doi-asserted-by":"publisher","first-page":"228","DOI":"10.1016\/j.cognition.2011.06.005","volume":"121","author":"C H\u00f6lscher","year":"2011","unstructured":"H\u00f6lscher, C., Tenbrink, T., & Wiener, J. M. (2011). Would you follow your own route description? Cognitive strategies in urban route planning. Cognition, 121(2), 228\u2013247.","journal-title":"Cognition"},{"key":"1374_CR41","doi-asserted-by":"crossref","unstructured":"Hu, R., Andreas, J., Darrell, T., & Saenko, K. (2018). Explainable neural computation via stack neural module networks. In Proceedings of the European conference on computer vision (ECCV).","DOI":"10.1007\/978-3-030-01234-2_4"},{"key":"1374_CR42","doi-asserted-by":"crossref","unstructured":"Hu, R., Rohrbach, A., Darrell, T., & Saenko, K. (2019). Language-conditioned graph networks for relational reasoning. arXiv preprint arXiv:1905.04405.","DOI":"10.1109\/ICCV.2019.01039"},{"key":"1374_CR43","unstructured":"Hudson, D. A., & Manning, C. D. (2018). Compositional attention networks for machine reasoning. arXiv preprint arXiv:1803.03067."},{"issue":"1","key":"1374_CR44","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1080\/13875868.2011.581773","volume":"12","author":"T Ishikawa","year":"2012","unstructured":"Ishikawa, T., & Nakamura, U. (2012). Landmark selection in the environment: Relationships with object characteristics and sense of direction. Spatial Cognition and Computation, 12(1), 1\u201322.","journal-title":"Spatial Cognition and Computation"},{"key":"1374_CR45","doi-asserted-by":"crossref","unstructured":"Karpathy, A., & Fei-Fei, L. (2015). Deep visual-semantic alignments for generating image descriptions. In Proceedings of the IEEE conference on computer vision and pattern recognition.","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"1374_CR46","doi-asserted-by":"crossref","unstructured":"Ke, L., Li, X., Bisk, Y., Holtzman, A., Gan, Z., Liu, J., Gao, J., Choi, Y., & Srinivasa, S.(2019). Tactical rewind: Self-correction via backtracking in vision-and-language navigation. In: Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 6741\u20136749).","DOI":"10.1109\/CVPR.2019.00690"},{"key":"1374_CR47","doi-asserted-by":"crossref","unstructured":"Khosla, A., An\u00a0An, B., Lim, J. J., & Torralba, A.(2014). Looking beyond the visible scene. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3710\u20133717).","DOI":"10.1109\/CVPR.2014.474"},{"key":"1374_CR48","doi-asserted-by":"crossref","unstructured":"Kim, J., Misu, T., Chen, Y. T., Tawari, A., & Canny, J. (2019). Grounding human-to-vehicle advice for self-driving vehicles. In The IEEE conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2019.01084"},{"key":"1374_CR49","doi-asserted-by":"crossref","unstructured":"Klippel, A., & Winter, S. (2005). Structural salience of landmarks for route directions. In Spatial information theory.","DOI":"10.1007\/11556114_22"},{"issue":"4","key":"1374_CR50","doi-asserted-by":"publisher","first-page":"311","DOI":"10.1016\/j.jvlc.2004.11.004","volume":"16","author":"A Klippel","year":"2005","unstructured":"Klippel, A., Tappe, H., Kulik, L., & Lee, P. U. (2005). Wayfinding choremesa language for modeling conceptual route knowledge. Journal of Visual Languages and Computing, 16(4), 311\u2013329.","journal-title":"Journal of Visual Languages and Computing"},{"key":"1374_CR51","unstructured":"Kumar, A., Gupta, S., Fouhey, D., Levine, S., & Malik, J. (2018). Visual memory for robust path following. In Advances in neural information processing systems (pp. 773\u2013782)."},{"key":"1374_CR52","unstructured":"Language Tool. (2016). Spell-Check API. Retrieved May 2018 from, https:\/\/languagetool.org\/."},{"issue":"11","key":"1374_CR53","doi-asserted-by":"publisher","first-page":"2278","DOI":"10.1109\/5.726791","volume":"86","author":"Y LeCun","year":"1998","unstructured":"LeCun, Y., Bottou, L., Bengio, Y., Haffner, P., et al. (1998). Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11), 2278\u20132324.","journal-title":"Proceedings of the IEEE"},{"key":"1374_CR54","doi-asserted-by":"crossref","unstructured":"Luo, R., Price, B., Cohen, S., & Shakhnarovich, G. (2018). Discriminability objective for training descriptive captions. arXiv preprint arXiv:1803.04376.","DOI":"10.1109\/CVPR.2018.00728"},{"key":"1374_CR55","unstructured":"Ma, C. Y., Lu, J., Wu, Z., AlRegib, G., Kira, Z., Socher, R., & Xiong, C. (2019). Self-monitoring navigation agent via auxiliary progress estimation. arXiv preprint arXiv:1901.03035."},{"key":"1374_CR56","doi-asserted-by":"crossref","unstructured":"Ma, C. Y., Wu, Z., AlRegib, G., Xiong, C., & Kira, Z. (2019). The regretful agent: Heuristic-aided navigation through progress estimation. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 6732\u20136740).","DOI":"10.1109\/CVPR.2019.00689"},{"key":"1374_CR57","unstructured":"Mansimov, E., Parisotto, E., Ba, J. L., & Salakhutdinov, R. (2015). Generating images from captions with attention. arXiv preprint arXiv:1511.02793."},{"key":"1374_CR58","doi-asserted-by":"crossref","unstructured":"Michon, P. E., & Denis, M. (2001). When and why are visual landmarks used in giving directions? In Spatial information theory.","DOI":"10.1007\/3-540-45424-1_20"},{"key":"1374_CR59","unstructured":"Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., & Dean, J. (2013). Distributed representations of words and phrases and their compositionality. In Advances in neural information processing systems (pp. 3111\u20133119)."},{"issue":"1","key":"1374_CR60","doi-asserted-by":"publisher","first-page":"43","DOI":"10.1109\/TITS.2006.889439","volume":"8","author":"A Millonig","year":"2007","unstructured":"Millonig, A., & Schechtner, K. (2007). Developing landmark-based pedestrian-navigation systems. IEEE Transactions on Intelligent Transportation Systems, 8(1), 43\u201349.","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"key":"1374_CR61","unstructured":"Mirowski, P., Grimes, M., Malinowski, M., Hermann, K. M., Anderson, K., Teplyashin, D., Simonyan, K., Kavukcuoglu, K., Zisserman, A., & Hadsell, R. (2018). Learning to navigate in cities without a map. In NIPS."},{"key":"1374_CR62","unstructured":"Mirowski, P. W., Pascanu, R., Viola, F., Soyer, H., Ballard, A. J., Banino, A., Denil, M., Goroshin, R., Sifre, L., Kavukcuoglu, K., Kumaran, D., & Hadsell, R. (2017). Learning to navigate in complex environments. In ICLR."},{"key":"1374_CR63","doi-asserted-by":"crossref","unstructured":"Nguyen, K., Dey, D., Brockett, C., & Dolan, B. (2019). Vision-based navigation with language-based assistance via imitation learning with indirect intervention. In The IEEE conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2019.01281"},{"key":"1374_CR64","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., & Manning, C. (2014). Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) (pp. 1532\u20131543).","DOI":"10.3115\/v1\/D14-1162"},{"key":"1374_CR65","unstructured":"Sutskever, I., Vinyals, O., & Le, Q. V. (2014). Sequence to sequence learning with neural networks. In Advances in neural information processing systems (pp 3104\u20133112)."},{"key":"1374_CR66","doi-asserted-by":"crossref","unstructured":"Thoma, J., Paudel, D. P., Chhatkuli, A., Probst, T., & Gool, L. V. (2019). Mapping, localization and path planning for image-based navigation using visual features and map. In The IEEE conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2019.00756"},{"key":"1374_CR67","doi-asserted-by":"crossref","unstructured":"Tom, A., & Denis, M. (2003). Referring to landmark or street information in route directions: What difference does it make? In International conference on spatial information theory (pp. 362\u2013374). Springer.","DOI":"10.1007\/978-3-540-39923-0_24"},{"issue":"9","key":"1374_CR68","doi-asserted-by":"publisher","first-page":"1213","DOI":"10.1002\/acp.1045","volume":"18","author":"A Tom","year":"2004","unstructured":"Tom, A., & Denis, M. (2004). Language and spatial cognition: Comparing the roles of landmarks and street names in route instructions. Applied Cognitive Psychology: The Official Journal of the Society for Applied Research in Memory and Cognition, 18(9), 1213\u20131230.","journal-title":"Applied Cognitive Psychology: The Official Journal of the Society for Applied Research in Memory and Cognition"},{"key":"1374_CR69","doi-asserted-by":"crossref","unstructured":"Tversky, B., & Lee, P. U. (1999). Pictorial and verbal tools for conveying routes. In C.\u00a0Freksa, & D. M. Mark (eds.) Spatial information theory. Cognitive and computational foundations of geographic information science (pp. 51\u201364).","DOI":"10.1007\/3-540-48384-5_4"},{"key":"1374_CR70","doi-asserted-by":"crossref","unstructured":"Vasudevan, A. B., Dai, D., & Van\u00a0Gool, L. (2018). Object referring in visual scene with spoken language. In 2018 IEEE winter conference on applications of computer vision (WACV) (pp. 1861\u20131870). IEEE.","DOI":"10.1109\/WACV.2018.00206"},{"key":"1374_CR71","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, \u0141., & Polosukhin, I. (2017). Attention is all you need. In Advances in neural information processing systems (pp. 5998\u20136008)."},{"key":"1374_CR72","unstructured":"Vogel, A., & Jurafsky, D. (2010). Learning to follow navigational directions. In Proceedings of the 48th annual meeting of the association for computational linguistics (pp. 806\u2013814). Association for Computational Linguistics."},{"key":"1374_CR73","doi-asserted-by":"crossref","unstructured":"Wang, X., Huang, Q., Celikyilmaz, A., Gao, J., Shen, D., Wang, Y. F., Yang\u00a0Wang, W., & Zhang, L. (2019). Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 6629\u20136638).","DOI":"10.1109\/CVPR.2019.00679"},{"key":"1374_CR74","doi-asserted-by":"crossref","unstructured":"Wang, F., Jiang, M., Qian, C., Yang, S., Li, C., Zhang, H., Wang, X., & Tang, X. (2017). Residual attention network for image classification. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3156\u20133164).","DOI":"10.1109\/CVPR.2017.683"},{"key":"1374_CR75","doi-asserted-by":"crossref","unstructured":"Wang, X., Xiong, W., Wang, H., & Yang\u00a0Wang, W. (2018). Look before you leap: Bridging model-free and model-based reinforcement learning for planned-ahead vision-and-language navigation. In ECCV.","DOI":"10.1007\/978-3-030-01270-0_3"},{"key":"1374_CR76","doi-asserted-by":"crossref","unstructured":"Weissenberg, J., Gygli, M., Riemenschneider, H., & Van\u00a0Gool, L. (2014). Navigation using special buildings as signposts. In Proceedings of the 2nd ACM SIGSPATIAL international workshop on interacting with maps (pp. 8\u201314). ACM.","DOI":"10.1145\/2677068.2677070"},{"key":"1374_CR77","doi-asserted-by":"crossref","unstructured":"Weyand, T., Kostrikov, I., & Philbin, J. (2016). Planet - photo geolocation with convolutional neural networks. In European conference on computer vision (ECCV).","DOI":"10.1007\/978-3-319-46484-8_3"},{"key":"1374_CR78","doi-asserted-by":"crossref","unstructured":"Wortsman, M., Ehsani, K., Rastegari, M., Farhadi, A., & Mottaghi, R. (2019). Learning to learn how to learn: Self-adaptive visual navigation using meta-learning. In The IEEE conference on computer vision and pattern recognition (CVPR).","DOI":"10.1109\/CVPR.2019.00691"},{"key":"1374_CR79","unstructured":"Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., et\u00a0al. (2016). Google\u2019s neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv:1609.08144."},{"key":"1374_CR80","unstructured":"Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., & Bengio, Y. (2015). Show, attend and tell: Neural image caption generation with visual attention. In International conference on machine learning (pp. 2048\u20132057)."},{"key":"1374_CR81","doi-asserted-by":"crossref","unstructured":"Yang, Z., He, X., Gao, J., Deng, L., & Smola, A. (2016). Stacked attention networks for image question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 21\u201329).","DOI":"10.1109\/CVPR.2016.10"},{"key":"1374_CR82","doi-asserted-by":"crossref","unstructured":"Zang, X., Pokle, A., V\u00e1zquez, M., Chen, K., Niebles, J. C., Soto, A., & Savarese, S. (2018). Translating navigation instructions in natural language to a high-level plan for behavioral robot navigation. CoRR.","DOI":"10.18653\/v1\/D18-1286"},{"key":"1374_CR83","doi-asserted-by":"crossref","unstructured":"Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., & Fidler, S. (2015). Aligning books and movies: Towards story-like visual explanations by watching movies and reading books. In Proceedings of the IEEE international conference on computer vision (pp. 19\u201327).","DOI":"10.1109\/ICCV.2015.11"},{"key":"1374_CR84","doi-asserted-by":"crossref","unstructured":"Zhu, Y., Mottaghi, R., Kolve, E., Lim, J. J., Gupta, A., Fei-Fei, L., & Farhadi, A. (2017). Target-driven visual navigation in indoor scenes using deep reinforcement learning. In ICRA.","DOI":"10.1109\/ICRA.2017.7989381"},{"issue":"5","key":"1374_CR85","doi-asserted-by":"publisher","first-page":"739","DOI":"10.3390\/app8050739","volume":"8","author":"X Zhu","year":"2018","unstructured":"Zhu, X., Li, L., Liu, J., Peng, H., & Niu, X. (2018). Captioning transformer with stacked attention modules. Applied Sciences, 8(5), 739.","journal-title":"Applied Sciences"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-020-01374-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-020-01374-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-020-01374-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,11,11]],"date-time":"2022-11-11T05:04:22Z","timestamp":1668143062000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-020-01374-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,8,31]]},"references-count":85,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,1]]}},"alternative-id":["1374"],"URL":"https:\/\/doi.org\/10.1007\/s11263-020-01374-3","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,8,31]]},"assertion":[{"value":"13 August 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 August 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"31 August 2020","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}