{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,21]],"date-time":"2025-10-21T15:27:19Z","timestamp":1761060439095,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":44,"publisher":"ACM","license":[{"start":{"date-parts":[[2017,10,19]],"date-time":"2017-10-19T00:00:00Z","timestamp":1508371200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2017,10,19]]},"DOI":"10.1145\/3123266.3123391","type":"proceedings-article","created":{"date-parts":[[2017,10,20]],"date-time":"2017-10-20T13:04:26Z","timestamp":1508504666000},"page":"1345-1353","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":32,"title":["Adaptively Attending to Visual Attributes and Linguistic Knowledge for Captioning"],"prefix":"10.1145","author":[{"given":"Yi","family":"Bin","sequence":"first","affiliation":[{"name":"University of Electronic Science and Technology of China, Chengdu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yang","family":"Yang","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, Chengdu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jie","family":"Zhou","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, Chengdu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zi","family":"Huang","sequence":"additional","affiliation":[{"name":"University of Queensland, Brisbane, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Heng Tao","family":"Shen","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, Chengdu, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2017,10,19]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"crossref","unstructured":"Lisa Anne Hendricks Subhashini Venugopalan Marcus Rohrbach Raymond Mooney Kate Saenko and Trevor Darrell. 2016. Deep compositional captioning: Describing novel object categories without paired training data CVPR. 1--10. Lisa Anne Hendricks Subhashini Venugopalan Marcus Rohrbach Raymond Mooney Kate Saenko and Trevor Darrell. 2016. Deep compositional captioning: Describing novel object categories without paired training data CVPR. 1--10.","DOI":"10.1109\/CVPR.2016.8"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.279"},{"key":"e_1_3_2_1_3_1","unstructured":"Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate ICLR. Dzmitry Bahdanau Kyunghyun Cho and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate ICLR."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2016.03.091"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2967258"},{"key":"e_1_3_2_1_6_1","unstructured":"David L Chen and William B Dolan. 2011. Collecting highly parallel data for paraphrase evaluation ACL. 190--200. David L Chen and William B Dolan. 2011. Collecting highly parallel data for paraphrase evaluation ACL. 190--200."},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3077136.3080671"},{"volume-title":"Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555","year":"2014","author":"Chung Junyoung","key":"e_1_3_2_1_8_1"},{"key":"e_1_3_2_1_9_1","unstructured":"Michael Denkowski and Alon Lavie Meteor universal: Language specific translation evaluation for any target language In Proceedings of the Ninth Workshop on Statistical Machine Translation. 376--380. Michael Denkowski and Alon Lavie Meteor universal: Language specific translation evaluation for any target language In Proceedings of the Ninth Workshop on Statistical Machine Translation. 376--380."},{"volume-title":"Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell.","year":"2015","author":"Donahue Jeffrey","key":"e_1_3_2_1_10_1"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2984064"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"crossref","unstructured":"Hao Fang Saurabh Gupta Forrest Iandola Rupesh K Srivastava Li Deng Piotr Doll\u00e1r Jianfeng Gao Xiaodong He Margaret Mitchell John C Platt and others. 2015. From captions to visual concepts and back. In CVPR. 1473--1482. Hao Fang Saurabh Gupta Forrest Iandola Rupesh K Srivastava Li Deng Piotr Doll\u00e1r Jianfeng Gao Xiaodong He Margaret Mitchell John C Platt and others. 2015. From captions to visual concepts and back. In CVPR. 1473--1482.","DOI":"10.1109\/CVPR.2015.7298754"},{"volume-title":"Peter Young, Cyrus Rashtchian, Julia Hockenmaier, and David Forsyth.","year":"2010","author":"Farhadi Ali","key":"e_1_3_2_1_13_1"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2967242"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2017.2717185"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.730558"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2967299"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2984070"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"crossref","unstructured":"Andrej Karpathy and Li Fei-Fei. 2015. Deep visual-semantic alignments for generating image descriptions CVPR. 3128--3137. Andrej Karpathy and Li Fei-Fei. 2015. Deep visual-semantic alignments for generating image descriptions CVPR. 3128--3137.","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"crossref","unstructured":"Niveda Krishnamoorthy Girish Malkarnenkar Raymond J Mooney Kate Saenko and Sergio Guadarrama. 2013. Generating Natural-Language Video Descriptions Using Text-Mined Knowledge AAAI. 541--547. Niveda Krishnamoorthy Girish Malkarnenkar Raymond J Mooney Kate Saenko and Sergio Guadarrama. 2013. Generating Natural-Language Video Descriptions Using Text-Mined Knowledge AAAI. 541--547.","DOI":"10.1609\/aaai.v27i1.8679"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2733373.2806314"},{"volume-title":"Rouge: A package for automatic evaluation of summaries ACL","year":"2004","author":"Lin Chin-Yew","key":"e_1_3_2_1_23_1"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"crossref","unstructured":"Tsung-Yi Lin Michael Maire Serge Belongie James Hays Pietro Perona Deva Ramanan Piotr Doll\u00e1r and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In ECCV. 740--755. Tsung-Yi Lin Michael Maire Serge Belongie James Hays Pietro Perona Deva Ramanan Piotr Doll\u00e1r and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In ECCV. 740--755.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2973831"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2967298"},{"volume-title":"Knowing When to Look: Adaptive Attention via A Visual Sentinel for Image Captioning. arXiv preprint arXiv:1612.01887","year":"2016","author":"Lu Jiasen","key":"e_1_3_2_1_27_1"},{"key":"e_1_3_2_1_28_1","unstructured":"Yingwei Pan Tao Mei Ting Yao Houqiang Li and Yong Rui. 2016. Jointly modeling embedding and translation to bridge video and language CVPR. 4594--4602. Yingwei Pan Tao Mei Ting Yao Houqiang Li and Yong Rui. 2016. Jointly modeling embedding and translation to bridge video and language CVPR. 4594--4602."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"},{"key":"e_1_3_2_1_31_1","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition ICLR. Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition ICLR."},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073445.1073478"},{"volume-title":"Cider: Consensus-based image description evaluation CVPR. 4566--4575.","year":"2015","author":"Vedantam Ramakrishna","key":"e_1_3_2_1_33_1"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.515"},{"volume-title":"Translating videos to natural language using deep recurrent neural networks. arXiv preprint arXiv:1412.4729","year":"2014","author":"Venugopalan Subhashini","key":"e_1_3_2_1_35_1"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2964299"},{"key":"e_1_3_2_1_37_1","unstructured":"Kelvin Xu Jimmy Ba Ryan Kiros Kyunghyun Cho Aaron Courville Ruslan Salakhudinov Rich Zemel and Yoshua Bengio. 2015. Show Attend and Tell: Neural Image Caption Generation with Visual Attention ICML. 2048--2057. Kelvin Xu Jimmy Ba Ryan Kiros Kyunghyun Cho Aaron Courville Ruslan Salakhudinov Rich Zemel and Yoshua Bengio. 2015. Show Attend and Tell: Neural Image Caption Generation with Visual Attention ICML. 2048--2057."},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2964319"},{"volume-title":"Hal Daum\u00e9 III, and Yiannis Aloimonos.","year":"2011","author":"Yang Yezhou","key":"e_1_3_2_1_39_1"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2014.2323014"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.512"},{"key":"e_1_3_2_1_42_1","unstructured":"Quanzeng You Hailin Jin Zhaowen Wang Chen Fang and Jiebo Luo. 2016. Image captioning with semantic attention. In CVPR. 4651--4659. Quanzeng You Hailin Jin Zhaowen Wang Chen Fang and Jiebo Luo. 2016. Image captioning with semantic attention. In CVPR. 4651--4659."},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"crossref","unstructured":"Haonan Yu Jiang Wang Zhiheng Huang Yi Yang and Wei Xu. 2016. Video paragraph captioning using hierarchical recurrent neural networks CVPR. 4584--4593. Haonan Yu Jiang Wang Zhiheng Huang Yi Yang and Wei Xu. 2016. Video paragraph captioning using hierarchical recurrent neural networks CVPR. 4584--4593.","DOI":"10.1109\/CVPR.2016.496"},{"volume-title":"ADADELTA: an adaptive learning rate method. arXiv preprint arXiv:1212.5701","year":"2012","author":"Zeiler Matthew D","key":"e_1_3_2_1_44_1"}],"event":{"name":"MM '17: ACM Multimedia Conference","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Mountain View California USA","acronym":"MM '17"},"container-title":["Proceedings of the 25th ACM international conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3123266.3123391","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3123266.3123391","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,26]],"date-time":"2025-06-26T16:41:55Z","timestamp":1750956115000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3123266.3123391"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,10,19]]},"references-count":44,"alternative-id":["10.1145\/3123266.3123391","10.1145\/3123266"],"URL":"https:\/\/doi.org\/10.1145\/3123266.3123391","relation":{},"subject":[],"published":{"date-parts":[[2017,10,19]]},"assertion":[{"value":"2017-10-19","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}