{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:12:39Z","timestamp":1750219959195,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":40,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,8,26]],"date-time":"2022-08-26T00:00:00Z","timestamp":1661472000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,8,26]]},"DOI":"10.1145\/3562007.3562052","type":"proceedings-article","created":{"date-parts":[[2022,10,12]],"date-time":"2022-10-12T22:13:51Z","timestamp":1665612831000},"page":"235-240","source":"Crossref","is-referenced-by-count":0,"title":["Differentiate Visual Features with Guidance Signals for Video Captioning"],"prefix":"10.1145","author":[{"given":"Yifan","family":"Yang","sequence":"first","affiliation":[{"name":"Key Laboratory of Spectral Imaging Technology CAS, Xi\u2019an Institute of Optics and Precision Mechanics, Chinese Academy of Sciences, China and University of Chinese Academy of Sciences, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaoqiang","family":"Lu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Spectral Imaging Technology CAS, Xi\u2019an Institute of Optics and Precision Mechanics, Chinese Academy of Sciences, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,10,12]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.12"},{"key":"e_1_3_2_1_2_1","unstructured":"Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473(2014).  Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473(2014)."},{"key":"e_1_3_2_1_3_1","volume-title":"Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and\/or summarization. 65\u201372","author":"Banerjee Satanjeev","year":"2005","unstructured":"Satanjeev Banerjee and Alon Lavie . 2005 . METEOR: An automatic metric for MT evaluation with improved correlation with human judgments . In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and\/or summarization. 65\u201372 . Satanjeev Banerjee and Alon Lavie. 2005. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and\/or summarization. 65\u201372."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.5555\/2002472.2002497"},{"key":"e_1_3_2_1_6_1","volume-title":"Proceedings of the European conference on computer vision (ECCV). 358\u2013373","author":"Chen Yangyu","year":"2018","unstructured":"Yangyu Chen , Shuhui Wang , Weigang Zhang , and Qingming Huang . 2018 . Less is more: Picking informative frames for video captioning . In Proceedings of the European conference on computer vision (ECCV). 358\u2013373 . Yangyu Chen, Shuhui Wang, Weigang Zhang, and Qingming Huang. 2018. Less is more: Picking informative frames for video captioning. In Proceedings of the European conference on computer vision (ECCV). 358\u2013373."},{"key":"e_1_3_2_1_7_1","volume-title":"Proceedings of the AAAI conference on artificial intelligence, Vol.\u00a032","author":"Cirik Volkan","year":"2018","unstructured":"Volkan Cirik , Taylor Berg-Kirkpatrick , and Louis-Philippe Morency . 2018 . Using syntax to ground referring expressions in natural images . In Proceedings of the AAAI conference on artificial intelligence, Vol.\u00a032 . Volkan Cirik, Taylor Berg-Kirkpatrick, and Louis-Philippe Morency. 2018. Using syntax to ground referring expressions in natural images. In Proceedings of the AAAI conference on artificial intelligence, Vol.\u00a032."},{"key":"e_1_3_2_1_8_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10695\u201310704","author":"Deshpande Aditya","year":"2019","unstructured":"Aditya Deshpande , Jyoti Aneja , Liwei Wang , Alexander\u00a0 G Schwing , and David Forsyth . 2019 . Fast, diverse and accurate image captioning guided by part-of-speech . In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10695\u201310704 . Aditya Deshpande, Jyoti Aneja, Liwei Wang, Alexander\u00a0G Schwing, and David Forsyth. 2019. Fast, diverse and accurate image captioning guided by part-of-speech. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10695\u201310704."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00957"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2911066"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.93"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3351072"},{"key":"e_1_3_2_1_13_1","unstructured":"Eric Jang Shixiang Gu and Ben Poole. 2016. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144(2016).  Eric Jang Shixiang Gu and Ben Poole. 2016. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144(2016)."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1020346032608"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.5555\/2891460.2891535"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.3390\/fi13020055"},{"key":"e_1_3_2_1_17_1","volume-title":"Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74\u201381.","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin . 2004 . Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74\u201381. Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74\u201381."},{"key":"e_1_3_2_1_18_1","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision. 4673\u20134682","author":"Liu Daqing","year":"2019","unstructured":"Daqing Liu , Hanwang Zhang , Feng Wu , and Zheng-Jun Zha . 2019 . Learning to assemble neural module tree networks for visual grounding . In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 4673\u20134682 . Daqing Liu, Hanwang Zhang, Feng Wu, and Zheng-Jun Zha. 2019. Learning to assemble neural module tree networks for visual grounding. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 4673\u20134682."},{"key":"e_1_3_2_1_19_1","volume-title":"2018 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 3782\u20133788","author":"Nguyen Anh","year":"2018","unstructured":"Anh Nguyen , Dimitrios Kanoulas , Luca Muratore , Darwin\u00a0 G Caldwell , and Nikos\u00a0 G Tsagarakis . 2018 . Translating videos to commands for robotic manipulation with deep recurrent neural networks . In 2018 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 3782\u20133788 . Anh Nguyen, Dimitrios Kanoulas, Luca Muratore, Darwin\u00a0G Caldwell, and Nikos\u00a0G Tsagarakis. 2018. Translating videos to commands for robotic manipulation with deep recurrent neural networks. In 2018 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 3782\u20133788."},{"key":"e_1_3_2_1_20_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10870\u201310879","author":"Pan Boxiao","year":"2020","unstructured":"Boxiao Pan , Haoye Cai , De-An Huang , Kuan-Hui Lee , Adrien Gaidon , Ehsan Adeli , and Juan\u00a0Carlos Niebles . 2020 . Spatio-temporal graph for video captioning with knowledge distillation . In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10870\u201310879 . Boxiao Pan, Haoye Cai, De-An Huang, Kuan-Hui Lee, Adrien Gaidon, Ehsan Adeli, and Juan\u00a0Carlos Niebles. 2020. Spatio-temporal graph for video captioning with knowledge distillation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10870\u201310879."},{"key":"e_1_3_2_1_21_1","volume-title":"Proceedings of the 40th annual meeting of the Association for Computational Linguistics. 311\u2013318","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni , Salim Roukos , Todd Ward , and Wei-Jing Zhu . 2002 . Bleu: a method for automatic evaluation of machine translation . In Proceedings of the 40th annual meeting of the Association for Computational Linguistics. 311\u2013318 . Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics. 311\u2013318."},{"key":"e_1_3_2_1_22_1","volume-title":"Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems 28","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren , Kaiming He , Ross Girshick , and Jian Sun . 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems 28 ( 2015 ). Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems 28 (2015)."},{"key":"e_1_3_2_1_23_1","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence, Vol.\u00a035","author":"Ryu Hobin","year":"2021","unstructured":"Hobin Ryu , Sunghun Kang , Haeyong Kang , and Chang\u00a0 D Yoo . 2021 . Semantic grouping network for video captioning . In Proceedings of the AAAI Conference on Artificial Intelligence, Vol.\u00a035 . 2514\u20132522. Hobin Ryu, Sunghun Kang, Haeyong Kang, and Chang\u00a0D Yoo. 2021. Semantic grouping network for video captioning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol.\u00a035. 2514\u20132522."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.5555\/3298023.3298188"},{"key":"e_1_3_2_1_25_1","unstructured":"Ganchao Tan Daqing Liu Meng Wang and Zheng-Jun Zha. 2020. Learning to discretely compose reasoning module networks for video captioning. arXiv preprint arXiv:2007.09049(2020).  Ganchao Tan Daqing Liu Meng Wang and Zheng-Jun Zha. 2020. Learning to discretely compose reasoning module networks for video captioning. arXiv preprint arXiv:2007.09049(2020)."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"crossref","unstructured":"Junjiao Tian and Jean Oh. 2020. Image captioning with compositional neural module networks. arXiv preprint arXiv:2007.05608(2020).  Junjiao Tian and Jean Oh. 2020. Image captioning with compositional neural module networks. arXiv preprint arXiv:2007.05608(2020).","DOI":"10.24963\/ijcai.2019\/496"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2020.107702"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.515"},{"key":"e_1_3_2_1_31_1","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision. 2641\u20132650","author":"Wang Bairui","year":"2019","unstructured":"Bairui Wang , Lin Ma , Wei Zhang , Wenhao Jiang , Jingwen Wang , and Wei Liu . 2019 . Controllable video captioning with pos sequence guidance based on gated fusion network . In Proceedings of the IEEE\/CVF international conference on computer vision. 2641\u20132650 . Bairui Wang, Lin Ma, Wei Zhang, Wenhao Jiang, Jingwen Wang, and Wei Liu. 2019. Controllable video captioning with pos sequence guidance based on gated fusion network. In Proceedings of the IEEE\/CVF international conference on computer vision. 2641\u20132650."},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00795"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.571"},{"key":"e_1_3_2_1_34_1","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision. 2561\u20132569","author":"Yang Tianhao","year":"2019","unstructured":"Tianhao Yang , Zheng-Jun Zha , and Hanwang Zhang . 2019 . Making history matter: History-advantage sequence training for visual dialog . In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 2561\u20132569 . Tianhao Yang, Zheng-Jun Zha, and Hanwang Zhang. 2019. Making history matter: History-advantage sequence training for visual dialog. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 2561\u20132569."},{"key":"e_1_3_2_1_35_1","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision. 4250\u20134260","author":"Yang Xu","year":"2019","unstructured":"Xu Yang , Hanwang Zhang , and Jianfei Cai . 2019 . Learning to collocate neural modules for image captioning . In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 4250\u20134260 . Xu Yang, Hanwang Zhang, and Jianfei Cai. 2019. Learning to collocate neural modules for image captioning. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 4250\u20134260."},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.512"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.347"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2909864"},{"key":"e_1_3_2_1_39_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 8327\u20138336","author":"Zhang Junchao","year":"2019","unstructured":"Junchao Zhang and Yuxin Peng . 2019 . Object-aware aggregation with bidirectional temporal graph for video captioning . In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 8327\u20138336 . Junchao Zhang and Yuxin Peng. 2019. Object-aware aggregation with bidirectional temporal graph for video captioning. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 8327\u20138336."},{"key":"e_1_3_2_1_40_1","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. 13278\u201313288","author":"Zhang Ziqi","year":"2020","unstructured":"Ziqi Zhang , Yaya Shi , Chunfeng Yuan , Bing Li , Peijin Wang , Weiming Hu , and Zheng-Jun Zha . 2020 . Object relational graph with teacher-recommended learning for video captioning . In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. 13278\u201313288 . Ziqi Zhang, Yaya Shi, Chunfeng Yuan, Bing Li, Peijin Wang, Weiming Hu, and Zheng-Jun Zha. 2020. Object relational graph with teacher-recommended learning for video captioning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. 13278\u201313288."},{"key":"e_1_3_2_1_41_1","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. 13096\u201313105","author":"Zheng Qi","year":"2020","unstructured":"Qi Zheng , Chaoyue Wang , and Dacheng Tao . 2020 . Syntax-aware action targeting for video captioning . In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. 13096\u201313105 . Qi Zheng, Chaoyue Wang, and Dacheng Tao. 2020. Syntax-aware action targeting for video captioning. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. 13096\u201313105."}],"event":{"name":"CCRIS'22: 2022 3rd International Conference on Control, Robotics and Intelligent System","acronym":"CCRIS'22","location":"Virtual Event China"},"container-title":["2022 3rd International Conference on Control, Robotics and Intelligent System"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3562007.3562052","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3562007.3562052","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:08Z","timestamp":1750182548000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3562007.3562052"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,8,26]]},"references-count":40,"alternative-id":["10.1145\/3562007.3562052","10.1145\/3562007"],"URL":"https:\/\/doi.org\/10.1145\/3562007.3562052","relation":{},"subject":[],"published":{"date-parts":[[2022,8,26]]}}}