{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T23:56:35Z","timestamp":1740182195905,"version":"3.37.3"},"reference-count":51,"publisher":"Oxford University Press (OUP)","issue":"5","license":[{"start":{"date-parts":[[2022,9,1]],"date-time":"2022-09-01T00:00:00Z","timestamp":1661990400000},"content-version":"vor","delay-in-days":1,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"funder":[{"DOI":"10.13039\/100008137","name":"Ministry of Natural Resources","doi-asserted-by":"publisher","award":["KF-2020-05-014"],"award-info":[{"award-number":["KF-2020-05-014"]}],"id":[{"id":"10.13039\/100008137","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["2018YFB1305001"],"award-info":[{"award-number":["2018YFB1305001"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,10,14]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Visual navigation task is to steer an embodied agent finding the given target based on observation. The effective transformer from observation of the agent to visual representation determines the navigation actions and promotes more informed navigation policy. In this work, we propose a spatial sequential transformer network (SSTNet) for learning informative visual representation in deep reinforcement learning. SSTNet is composed by spatial attention probability fused model (SAF) and sequential transformer network (STNet). SAF enforces cross-modal state into visual clues in reinforcement learning. It encodes semantic information about observed objects, as well as spatial information about their location, which jointly exploiting image inter-relations. STNet generates (imagines) the next observations and makes action inference of the aspects most relevant to the target. It decodes the image intra-relations. This way, the agent learns to understand the causality between navigation actions and dynamic changes in observations. SSTNet is conditioned on an auto-regressive model on the desired reward, past states, actions, and knowledge graph. The whole navigation framework considers the local and global visual information, as well as time sequential information. Thus, it allows the agent to navigate towards the sought-after object effectively. We evaluate our model on the AI2THOR framework show that our method attains at least $10\\%$ improvement of average success rate over most state-of-the-art models. Code and datasets can be found in https:\/\/github.com\/zhoukang123\/SDTNet_2022.<\/jats:p>","DOI":"10.1093\/jcde\/qwac084","type":"journal-article","created":{"date-parts":[[2022,9,1]],"date-time":"2022-09-01T00:33:53Z","timestamp":1661992433000},"page":"1866-1878","source":"Crossref","is-referenced-by-count":2,"title":["TransNav: spatial sequential transformer network for visual navigation"],"prefix":"10.1093","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4177-7188","authenticated-orcid":false,"given":"Kang","family":"Zhou","sequence":"first","affiliation":[{"name":"School of Computer Science, Wuhan University , Wuhan 430072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3820-3586","authenticated-orcid":false,"given":"Huyin","family":"Zhang","sequence":"additional","affiliation":[{"name":"Key Laboratory of Urban Land Resources Monitoring and Simulation, Ministry of Natural Resources , Shenzhen 518000, China"},{"name":"School of Computer Science, Wuhan University , Wuhan 430072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fei","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer Science, Wuhan University , Wuhan 430072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2022,8,31]]},"reference":[{"key":"2022110904481092700_bib1","doi-asserted-by":"crossref","first-page":"6485","DOI":"10.24963\/ijcai.2019\/931","article-title":"OpenMarkov, an open-source tool for probabilistic graphical models","volume-title":"Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence","author":"Arias","year":"2019"},{"issue":"7705","key":"2022110904481092700_bib2","doi-asserted-by":"crossref","first-page":"429","DOI":"10.1038\/s41586-018-0102-6","article-title":"Vector-based navigation using grid-like representations in artificial agents","volume":"557","author":"Banino","year":"2018","journal-title":"Nature"},{"year":"2020","author":"Bochkovskiy","article-title":"YOLOv4: Optimal speed and accuracy of object detection","key":"2022110904481092700_bib3"},{"year":"2017","author":"Bruce","article-title":"One-shot reinforcement learning for robot navigation with interactive replay","key":"2022110904481092700_bib4"},{"key":"2022110904481092700_bib5","first-page":"213","article-title":"End-to-end object detection with transformers","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Carion","year":"2020"},{"key":"2022110904481092700_bib6","first-page":"12","article-title":"Glit: Neural architecture search for global and local image transformer","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Chen","year":"2021"},{"key":"2022110904481092700_bib7","first-page":"11276","article-title":"Topological planning with transformers for vision-and-language navigation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen","year":"2021"},{"year":"2021","author":"Chen","article-title":"Decision transformer: Reinforcement learning via sequence modeling","key":"2022110904481092700_bib8"},{"key":"2022110904481092700_bib9","first-page":"19","article-title":"Learning object relation graph and tentative policy for visual navigation","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Du","year":"2020"},{"year":"2021","author":"Du","article-title":"VTNet: Visual transformer network for object goal navigation","key":"2022110904481092700_bib10"},{"key":"2022110904481092700_bib11","first-page":"1126","article-title":"Model-agnostic meta-learning for fast adaptation of deep networks","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Finn","year":"2017"},{"key":"2022110904481092700_bib12","article-title":"Speaker-follower models for vision-and-language navigation","volume":"31","author":"Fried","year":"2018","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2022110904481092700_bib13","first-page":"1634","article-title":"Airbert: In-domain pretraining for vision-and-language navigation","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Guhur","year":"2021"},{"key":"2022110904481092700_bib14","first-page":"13137","article-title":"Towards learning a generic agent for vision-and-language navigation via pre-training","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Hao","year":"2020"},{"issue":"6 Special Issue","key":"2022110904481092700_bib15","first-page":"145","article-title":"A study using machine learning with Ngram model in harmonized system classification","volume":"12","author":"Harsani","year":"2020","journal-title":"Journal of Advanced Research in Dynamical and Control Systems"},{"key":"2022110904481092700_bib16","first-page":"1","article-title":"Self-supervised deep reinforcement learning with generalized computation graphs for robot navigation","volume-title":"Proceedings of the 2018 IEEE International Conference on Robotics and Automation, ICRA 2018","author":"Kahn","year":"2018"},{"year":"2016","author":"Kipf","article-title":"Semi-supervised classification with graph convolutional networks","key":"2022110904481092700_bib17"},{"year":"2017","author":"Kolve","article-title":"AI2-THOR: An interactive 3d environment for visual AI","key":"2022110904481092700_bib18"},{"issue":"1","key":"2022110904481092700_bib19","doi-asserted-by":"crossref","first-page":"32","DOI":"10.1007\/s11263-016-0981-7","article-title":"Visual Genome: Connecting language and vision using crowdsourced dense image annotations","volume":"123","author":"Krishna","year":"2017","journal-title":"International journal of computer vision"},{"key":"2022110904481092700_bib20","first-page":"156","article-title":"Temporal convolutional networks for action segmentation and detection","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Lea","year":"2017"},{"key":"2022110904481092700_bib21","first-page":"1","article-title":"Improving target-driven visual navigation with attention on 3D spatial relationships","volume-title":"Neural Processing Letters","author":"Lyu","year":"2020"},{"key":"2022110904481092700_bib22","first-page":"6732","article-title":"The regretful agent: Heuristic-aided navigation through progress estimation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ma","year":"2019"},{"year":"2020","author":"Moghaddam","article-title":"Utilising Prior Knowledge for Visual Navigation: Distil and Adapt","key":"2022110904481092700_bib23"},{"key":"2022110904481092700_bib24","first-page":"16898","article-title":"Visual navigation with spatial attention","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Mayo","year":"2021"},{"key":"2022110904481092700_bib25","first-page":"1928","article-title":"Asynchronous methods for deep reinforcement learning","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Mnih","year":"2016"},{"key":"2022110904481092700_bib26","first-page":"3733","article-title":"Optimistic agent: Accurate graph-based value estimation for more successful visual navigation","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Moghaddam","year":"2021"},{"key":"2022110904481092700_bib27","first-page":"7487","article-title":"Stabilizing transformers for reinforcement learning","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Parisotto","year":"2020"},{"key":"2022110904481092700_bib28","doi-asserted-by":"crossref","first-page":"1532","DOI":"10.3115\/v1\/D14-1162","article-title":"Glove: Global vectors for word representation","volume-title":"Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Pennington","year":"2014"},{"key":"2022110904481092700_bib29","first-page":"517","article-title":"Learning hierarchical relationships for object-goal navigation","volume-title":"Proceedings of the PMLR Conference on Robot Learning","author":"Qiu","year":"2021"},{"issue":"6","key":"2022110904481092700_bib30","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"39","author":"Ren","year":"2016","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2022110904481092700_bib31","article-title":"Rapid task-solving in novel environments","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Ritter","year":"2020"},{"key":"2022110904481092700_bib32","first-page":"7236","article-title":"Multi-agent actor-critic with hierarchical graph attention network","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Ryu","year":"2020"},{"issue":"4","key":"2022110904481092700_bib33","doi-asserted-by":"crossref","first-page":"2393","DOI":"10.1109\/TII.2019.2936167","article-title":"End-to-end navigation strategy with deep reinforcement learning for mobile robots","volume":"16","author":"Shi","year":"2019","journal-title":"IEEE Transactions on Industrial Informatics"},{"key":"2022110904481092700_bib34","first-page":"1746","article-title":"Semantic scene completion from a single depth image","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Song","year":"2017"},{"key":"2022110904481092700_bib35","doi-asserted-by":"crossref","first-page":"1441","DOI":"10.1145\/3357384.3357895","article-title":"BERT4REC: Sequential recommendation with bidirectional encoder representations from transformer","volume-title":"Proceedings of the 28th ACM International Conference on Information and Knowledge Management","author":"Sun","year":"2019"},{"key":"2022110904481092700_bib36","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/978-1-4615-3618-5_1","article-title":"Introduction: The challenge of reinforcement learning","volume-title":"Reinforcement learning","author":"Sutton","year":"1992"},{"year":"2018","author":"Van\u00a0Hasselt","article-title":"Deep reinforcement learning and the deadly triad","key":"2022110904481092700_bib37"},{"key":"2022110904481092700_bib38","first-page":"5998","article-title":"Attention is all you need","volume-title":"Proceedings of the 31st Conference on Advances in Neural Information Processing Systems (NIPS 2017)","author":"Vaswani","year":"2017"},{"key":"2022110904481092700_bib39","article-title":"Graph attention networks","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Veli\u010dkovi\u0107","year":"2018"},{"issue":"3","key":"2022110904481092700_bib40","doi-asserted-by":"crossref","first-page":"4509","DOI":"10.1109\/LRA.2020.3002198","article-title":"Learning scheduling policies for multi-robot coordination with graph attention networks","volume":"5","author":"Wang","year":"2020","journal-title":"IEEE Robotics and Automation Letters"},{"key":"2022110904481092700_bib41","first-page":"6750","article-title":"Learning to learn how to learn: Self-adaptive visual navigation using meta-learning","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wortsman","year":"2019"},{"key":"2022110904481092700_bib43","first-page":"2769","article-title":"Bayesian relational memory for semantic visual navigation","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Wu","year":"2019"},{"key":"2022110904481092700_bib42","first-page":"10001","article-title":"NeoNav: Improving the generalization of visual navigation via generating next expected observations","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Wu","year":"2020"},{"key":"2022110904481092700_bib44","first-page":"670","article-title":"Graph R-CNN for scene graph generation","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Yang","year":"2018"},{"year":"2018","author":"Yang","article-title":"Visual semantic navigation using scene priors","key":"2022110904481092700_bib45"},{"key":"2022110904481092700_bib46","first-page":"11983","article-title":"Graph transformer networks","volume":"32","author":"Yun","year":"2019","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2022110904481092700_bib47","article-title":"Deep reinforcement learning with relational inductive biases","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Zambaldi","year":"2018"},{"key":"2022110904481092700_bib48","first-page":"2371","article-title":"Deep reinforcement learning with successor features for navigation across similar environments","volume-title":"Proceedings of the 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS)","author":"Zhang","year":"2017"},{"issue":"9","key":"2022110904481092700_bib50","doi-asserted-by":"crossref","first-page":"3469","DOI":"10.1109\/TCSVT.2020.3039522","article-title":"Language-guided navigation via cross-modal grounding and alternate adversarial learning","volume":"31","author":"Zhang","year":"2020","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"2022110904481092700_bib49","first-page":"15130","article-title":"Hierarchical object-to-zone graph for object navigation","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zhang","year":"2021"},{"key":"2022110904481092700_bib51","doi-asserted-by":"crossref","first-page":"3357","DOI":"10.1109\/ICRA.2017.7989381","article-title":"Target-driven visual navigation in indoor scenes using deep reinforcement learning","volume-title":"Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA)","author":"Zhu","year":"2017"}],"container-title":["Journal of Computational Design and Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/jcde\/advance-article-pdf\/doi\/10.1093\/jcde\/qwac084\/45633869\/qwac084.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jcde\/article-pdf\/9\/5\/1866\/46878040\/qwac084.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jcde\/article-pdf\/9\/5\/1866\/46878040\/qwac084.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,11,9]],"date-time":"2022-11-09T04:48:48Z","timestamp":1667969328000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/jcde\/article\/9\/5\/1866\/6679563"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,8,31]]},"references-count":51,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2022,10,14]]}},"URL":"https:\/\/doi.org\/10.1093\/jcde\/qwac084","relation":{},"ISSN":["2288-5048"],"issn-type":[{"type":"electronic","value":"2288-5048"}],"subject":[],"published-other":{"date-parts":[[2022,10]]},"published":{"date-parts":[[2022,8,31]]}}}