{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T17:51:04Z","timestamp":1784137864112,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":31,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,12,6]],"date-time":"2023-12-06T00:00:00Z","timestamp":1701820800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"JSPS KAKENHI","award":["23H03402"],"award-info":[{"award-number":["23H03402"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,12,6]]},"DOI":"10.1145\/3595916.3626450","type":"proceedings-article","created":{"date-parts":[[2024,1,1]],"date-time":"2024-01-01T16:34:41Z","timestamp":1704126881000},"page":"1-7","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Vision-Language Navigation for Quadcopters with Conditional Transformer and Prompt-based Text Rephraser"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9108-9028","authenticated-orcid":false,"given":"Zhe","family":"Chen","sequence":"first","affiliation":[{"name":"Hangzhou Dianzi University, China and University of Yamanashi, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4997-3850","authenticated-orcid":false,"given":"Jiyi","family":"Li","sequence":"additional","affiliation":[{"name":"University of Yamanashi, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7858-6206","authenticated-orcid":false,"given":"Fumiyo","family":"Fukumoto","sequence":"additional","affiliation":[{"name":"University of Yamanashi, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3403-2604","authenticated-orcid":false,"given":"Peng","family":"Liu","sequence":"additional","affiliation":[{"name":"Hangzhou Dianzi University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5466-7351","authenticated-orcid":false,"given":"Yoshimi","family":"Suzuki","sequence":"additional","affiliation":[{"name":"University of Yamanashi, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,1]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00387"},{"key":"e_1_3_2_1_2_1","volume-title":"Proceedings of the Robotics Science and Systems Conference (RSS). http:\/\/arxiv.org\/abs\/1806","author":"Blukis Valts","year":"2018","unstructured":"Valts Blukis , Nataly Brukhim , Andrew Bennett , Ross\u00a0 A. Knepper , and Yoav Artzi . 2018 . Following High-level Navigation Instructions on a Simulated Quadcopter with Imitation Learning . In Proceedings of the Robotics Science and Systems Conference (RSS). http:\/\/arxiv.org\/abs\/1806 .00047 Valts Blukis, Nataly Brukhim, Andrew Bennett, Ross\u00a0A. Knepper, and Yoav Artzi. 2018. Following High-level Navigation Instructions on a Simulated Quadcopter with Imitation Learning. In Proceedings of the Robotics Science and Systems Conference (RSS). http:\/\/arxiv.org\/abs\/1806.00047"},{"key":"e_1_3_2_1_3_1","volume-title":"Proceedings of The 2nd Conference on Robot Learning(Proceedings of Machine Learning Research, Vol.\u00a087)","author":"Blukis Valts","year":"2018","unstructured":"Valts Blukis , Dipendra Misra , Ross\u00a0 A. Knepper , and Yoav Artzi . 2018 . Mapping Navigation Instructions to Continuous Control Actions with Position-Visitation Prediction . In Proceedings of The 2nd Conference on Robot Learning(Proceedings of Machine Learning Research, Vol.\u00a087) , Aude Billard, Anca Dragan, Jan Peters, and Jun Morimoto (Eds.). PMLR, 505\u2013518. https:\/\/proceedings.mlr.press\/v87\/blukis18a.html Valts Blukis, Dipendra Misra, Ross\u00a0A. Knepper, and Yoav Artzi. 2018. Mapping Navigation Instructions to Continuous Control Actions with Position-Visitation Prediction. In Proceedings of The 2nd Conference on Robot Learning(Proceedings of Machine Learning Research, Vol.\u00a087), Aude Billard, Anca Dragan, Jan Peters, and Jun Morimoto (Eds.). PMLR, 505\u2013518. https:\/\/proceedings.mlr.press\/v87\/blukis18a.html"},{"key":"e_1_3_2_1_4_1","volume-title":"Advances in Neural Information Processing Systems, H.\u00a0Larochelle, M.\u00a0Ranzato, R.\u00a0Hadsell, M.F. Balcan, and H.\u00a0Lin (Eds.). Vol.\u00a033. Curran Associates","author":"Brown Tom","year":"1877","unstructured":"Tom Brown , Benjamin Mann , Nick Ryder , Melanie Subbiah , Jared\u00a0 D Kaplan , Prafulla Dhariwal , Arvind Neelakantan , Pranav Shyam , Girish Sastry , Amanda Askell , Sandhini Agarwal , Ariel Herbert-Voss , Gretchen Krueger , Tom Henighan , Rewon Child , Aditya Ramesh , Daniel Ziegler , Jeffrey Wu , Clemens Winter , Chris Hesse , Mark Chen , Eric Sigler , Mateusz Litwin , Scott Gray , Benjamin Chess , Jack Clark , Christopher Berner , Sam McCandlish , Alec Radford , Ilya Sutskever , and Dario Amodei . 2020. Language Models are Few-Shot Learners . In Advances in Neural Information Processing Systems, H.\u00a0Larochelle, M.\u00a0Ranzato, R.\u00a0Hadsell, M.F. Balcan, and H.\u00a0Lin (Eds.). Vol.\u00a033. Curran Associates , Inc ., 1877 \u20131901. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2020\/file\/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared\u00a0D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems, H.\u00a0Larochelle, M.\u00a0Ranzato, R.\u00a0Hadsell, M.F. Balcan, and H.\u00a0Lin (Eds.). Vol.\u00a033. Curran Associates, Inc., 1877\u20131901. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2020\/file\/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf"},{"key":"e_1_3_2_1_5_1","unstructured":"Devendra\u00a0Singh Chaplot Kanthashree\u00a0Mysore Sathyendra Rama\u00a0Kumar Pasumarthi Dheeraj Rajagopal and Ruslan Salakhutdinov. 2018. Gated-Attention Architectures for Task-Oriented Language Grounding. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence and Thirtieth Innovative Applications of Artificial Intelligence Conference and Eighth AAAI Symposium on Educational Advances in Artificial Intelligence (New Orleans Louisiana USA) (AAAI\u201918\/IAAI\u201918\/EAAI\u201918). AAAI Press Article 344 8\u00a0pages.  Devendra\u00a0Singh Chaplot Kanthashree\u00a0Mysore Sathyendra Rama\u00a0Kumar Pasumarthi Dheeraj Rajagopal and Ruslan Salakhutdinov. 2018. Gated-Attention Architectures for Task-Oriented Language Grounding. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence and Thirtieth Innovative Applications of Artificial Intelligence Conference and Eighth AAAI Symposium on Educational Advances in Artificial Intelligence (New Orleans Louisiana USA) (AAAI\u201918\/IAAI\u201918\/EAAI\u201918). AAAI Press Article 344 8\u00a0pages."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01282"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1"},{"key":"e_1_3_2_1_9_1","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV). 1634\u20131643","author":"Guhur Pierre-Louis","year":"2021","unstructured":"Pierre-Louis Guhur , Makarand Tapaswi , Shizhe Chen , Ivan Laptev , and Cordelia Schmid . 2021 . Airbert: In-Domain Pretraining for Vision-and-Language Navigation . In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV). 1634\u20131643 . Pierre-Louis Guhur, Makarand Tapaswi, Shizhe Chen, Ivan Laptev, and Cordelia Schmid. 2021. Airbert: In-Domain Pretraining for Vision-and-Language Navigation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV). 1634\u20131643."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00169"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA48506.2021.9561806"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1"},{"key":"e_1_3_2_1_15_1","volume-title":"Computer Vision \u2013 ECCV","author":"Krantz Jacob","year":"2020","unstructured":"Jacob Krantz , Erik Wijmans , Arjun Majumdar , Dhruv Batra , and Stefan Lee . 2020. Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments . In Computer Vision \u2013 ECCV 2020 , Andrea Vedaldi, Horst Bischof , Thomas Brox, and Jan-Michael Frahm (Eds.). Springer International Publishing , Cham, 104\u2013120. Jacob Krantz, Erik Wijmans, Arjun Majumdar, Dhruv Batra, and Stefan Lee. 2020. Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments. In Computer Vision \u2013 ECCV 2020, Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm (Eds.). Springer International Publishing, Cham, 104\u2013120."},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1"},{"key":"e_1_3_2_1_17_1","volume-title":"Decoupled Weight Decay Regularization. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=Bkg6RiCqY7","author":"Loshchilov Ilya","year":"2019","unstructured":"Ilya Loshchilov and Frank Hutter . 2019 . Decoupled Weight Decay Regularization. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=Bkg6RiCqY7 Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=Bkg6RiCqY7"},{"key":"e_1_3_2_1_18_1","volume-title":"How Does Fine-Tuning Impact Out-of-Distribution Detection for Vision-Language Models?International Journal of Computer Vision","author":"Ming Yifei","year":"2023","unstructured":"Yifei Ming and Yixuan Li. 2023. How Does Fine-Tuning Impact Out-of-Distribution Detection for Vision-Language Models?International Journal of Computer Vision ( 2023 ). Yifei Ming and Yixuan Li. 2023. How Does Fine-Tuning Impact Out-of-Distribution Detection for Vision-Language Models?International Journal of Computer Vision (2023)."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"e_1_3_2_1_24_1","volume-title":"Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics(Proceedings of Machine Learning Research, Vol.\u00a015)","author":"Ross St\u00e9phane","year":"2011","unstructured":"St\u00e9phane Ross , Geoffrey Gordon , and Drew Bagnell . 2011 . A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning . In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics(Proceedings of Machine Learning Research, Vol.\u00a015) , Geoffrey Gordon, David Dunson, and Miroslav Dud\u00edk (Eds.). PMLR, Fort Lauderdale, FL, USA, 627\u2013635. https:\/\/proceedings.mlr.press\/v15\/ross11a.html St\u00e9phane Ross, Geoffrey Gordon, and Drew Bagnell. 2011. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics(Proceedings of Machine Learning Research, Vol.\u00a015), Geoffrey Gordon, David Dunson, and Miroslav Dud\u00edk (Eds.). PMLR, Fort Lauderdale, FL, USA, 627\u2013635. https:\/\/proceedings.mlr.press\/v15\/ross11a.html"},{"key":"e_1_3_2_1_25_1","volume-title":"VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View.","author":"Schumann Raphael","year":"2023","unstructured":"Raphael Schumann , Wanrong Zhu , Weixi Feng , Tsu-Jui Fu , Stefan Riezler , and William\u00a0Yang Wang . 2023 . VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View. (2023). arXiv:2307.06082 Raphael Schumann, Wanrong Zhu, Weixi Feng, Tsu-Jui Fu, Stefan Riezler, and William\u00a0Yang Wang. 2023. VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View. (2023). arXiv:2307.06082"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"crossref","unstructured":"Shital Shah Debadeepta Dey Chris Lovett and Ashish Kapoor. 2017. AirSim: High-Fidelity Visual and Physical Simulation for Autonomous Vehicles. In Field and Service Robotics. https:\/\/arxiv.org\/abs\/1705.05065  Shital Shah Debadeepta Dey Chris Lovett and Ashish Kapoor. 2017. AirSim: High-Fidelity Visual and Physical Simulation for Autonomous Vehicles. In Field and Service Robotics. https:\/\/arxiv.org\/abs\/1705.05065","DOI":"10.1007\/978-3-319-67361-5_40"},{"key":"e_1_3_2_1_27_1","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar Aurelien Rodriguez Armand Joulin Edouard Grave and Guillaume Lample. 2023. LLaMA: Open and Efficient Foundation Language Models. arxiv:2302.13971\u00a0[cs.CL]  Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar Aurelien Rodriguez Armand Joulin Edouard Grave and Guillaume Lample. 2023. LLaMA: Open and Efficient Foundation Language Models. arxiv:2302.13971\u00a0[cs.CL]"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2020.2967597"},{"key":"e_1_3_2_1_30_1","volume-title":"Cross-Lingual Vision-Language Navigation. In International Conference on Learning Representations. http:\/\/arxiv.org\/abs\/1910","author":"Yan An","year":"2019","unstructured":"An Yan , Xin Wang , Jiangtao Feng , Lei Li , and William\u00a0Yang Wang . 2019 . Cross-Lingual Vision-Language Navigation. In International Conference on Learning Representations. http:\/\/arxiv.org\/abs\/1910 .11301 An Yan, Xin Wang, Jiangtao Feng, Lei Li, and William\u00a0Yang Wang. 2019. Cross-Lingual Vision-Language Navigation. In International Conference on Learning Representations. http:\/\/arxiv.org\/abs\/1910.11301"},{"key":"e_1_3_2_1_31_1","unstructured":"Tianjun Zhang Yi Zhang Vibhav Vineet Neel Joshi and Xin Wang. 2023. Controllable Text-to-Image Generation with GPT-4. arxiv:2305.18583\u00a0[cs.CV]  Tianjun Zhang Yi Zhang Vibhav Vineet Neel Joshi and Xin Wang. 2023. Controllable Text-to-Image Generation with GPT-4. arxiv:2305.18583\u00a0[cs.CV]"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1"}],"event":{"name":"MMAsia '23: ACM Multimedia Asia","location":"Tainan Taiwan","acronym":"MMAsia '23","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["ACM Multimedia Asia 2023"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3595916.3626450","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3595916.3626450","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:35:56Z","timestamp":1750178156000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3595916.3626450"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,12,6]]},"references-count":31,"alternative-id":["10.1145\/3595916.3626450","10.1145\/3595916"],"URL":"https:\/\/doi.org\/10.1145\/3595916.3626450","relation":{},"subject":[],"published":{"date-parts":[[2023,12,6]]},"assertion":[{"value":"2024-01-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}