{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,9]],"date-time":"2025-09-09T21:25:15Z","timestamp":1757453115793,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":39,"publisher":"ACM","license":[{"start":{"date-parts":[[2017,6,6]],"date-time":"2017-06-06T00:00:00Z","timestamp":1496707200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key Research and Development Plan","award":["2016YFB1001202"],"award-info":[{"award-number":["2016YFB1001202"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2017,6,6]]},"DOI":"10.1145\/3078971.3079000","type":"proceedings-article","created":{"date-parts":[[2017,5,25]],"date-time":"2017-05-25T16:27:32Z","timestamp":1495729652000},"page":"5-13","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":15,"title":["Generating Video Descriptions with Topic Guidance"],"prefix":"10.1145","author":[{"given":"Shizhe","family":"Chen","sequence":"first","affiliation":[{"name":"Renmin University of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jia","family":"Chen","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qin","family":"Jin","sequence":"additional","affiliation":[{"name":"Renmin University of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2017,6,6]]},"reference":[{"key":"e_1_3_2_1_1_1","first-page":"2085","volume-title":"ICML","author":"Lebret R\u00e9mi","year":"2015","unstructured":"R\u00e9mi Lebret , Pedro H. O. Pinheiro , and Ronan Collobert . Phrase-based image captioning . In ICML , pages 2085 -- 2094 , 2015 . R\u00e9mi Lebret, Pedro H. O. Pinheiro, and Ronan Collobert. Phrase-based image captioning. In ICML, pages 2085--2094, 2015."},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"e_1_3_2_1_3_1","volume-title":"attend and tell: Neural image caption generation with visual attention. arXiv:1502.03044, 2(3):5","author":"Xu Kelvin","year":"2015","unstructured":"Kelvin Xu , Jimmy Ba , Ryan Kiros , Kyunghyun Cho , Aaron Courville , Ruslan Salakhutdinov , Richard S Zemel , and Yoshua Bengio . Show , attend and tell: Neural image caption generation with visual attention. arXiv:1502.03044, 2(3):5 , 2015 . Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard S Zemel, and Yoshua Bengio. Show, attend and tell: Neural image caption generation with visual attention. arXiv:1502.03044, 2(3):5, 2015."},{"key":"e_1_3_2_1_4_1","volume-title":"Image captioning with semantic attention. arXiv:1603.03925","author":"You Quanzeng","year":"2016","unstructured":"Quanzeng You , Hailin Jin , Zhaowen Wang , Chen Fang , and Jiebo Luo . Image captioning with semantic attention. arXiv:1603.03925 , 2016 . Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo. Image captioning with semantic attention. arXiv:1603.03925, 2016."},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2984065"},{"key":"e_1_3_2_1_6_1","first-page":"2654","volume-title":"NIPS","author":"Ba Lei Jimmy","year":"2013","unstructured":"Lei Jimmy Ba and Rich Caruana . Do deep nets really need to be deep ? NIPS , pages 2654 -- 2662 , 2013 . Lei Jimmy Ba and Rich Caruana. Do deep nets really need to be deep? NIPS, pages 2654--2662, 2013."},{"key":"e_1_3_2_1_7_1","volume-title":"Neural machine translation by jointly learning to align and translate. arXiv:1409.0473","author":"Bahdanau Dzmitry","year":"2014","unstructured":"Dzmitry Bahdanau , Kyunghyun Cho , and Yoshua Bengio . Neural machine translation by jointly learning to align and translate. arXiv:1409.0473 , 2014 . Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Neural machine translation by jointly learning to align and translate. arXiv:1409.0473, 2014."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.61"},{"key":"e_1_3_2_1_9_1","volume-title":"Computer Science","author":"Venugopalan Subhashini","year":"2014","unstructured":"Subhashini Venugopalan , Huijuan Xu , Jeff Donahue , Marcus Rohrbach , Raymond Mooney , and Kate Saenko . Translating videos to natural language using deep recurrent neural networks . Computer Science , 2014 . Subhashini Venugopalan, Huijuan Xu, Jeff Donahue, Marcus Rohrbach, Raymond Mooney, and Kate Saenko. Translating videos to natural language using deep recurrent neural networks. Computer Science, 2014."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.497"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.29"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.127"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.340"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00207"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.571"},{"key":"e_1_3_2_1_16_1","unstructured":"Msr video to language challenge. http:\/\/www.acmmm.org\/2016\/wp-content\/uploads\/2016\/04\/ACMMM16_GC_MSR_Video_to_Language_Updated.pdf.  Msr video to language challenge. http:\/\/www.acmmm.org\/2016\/wp-content\/uploads\/2016\/04\/ACMMM16_GC_MSR_Video_to_Language_Updated.pdf."},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.512"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.515"},{"key":"e_1_3_2_1_19_1","volume-title":"Hierarchical recurrent neural encoder for video representation with application to captioning. arXiv:1511.03476","author":"Pan Pingbo","year":"2015","unstructured":"Pingbo Pan , Zhongwen Xu , Yi Yang , Fei Wu , and Yueting Zhuang . Hierarchical recurrent neural encoder for video representation with application to captioning. arXiv:1511.03476 , 2015 . Pingbo Pan, Zhongwen Xu, Yi Yang, Fei Wu, and Yueting Zhuang. Hierarchical recurrent neural encoder for video representation with application to captioning. arXiv:1511.03476, 2015."},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2911996.2912043"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2984066"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.5555\/944919.944937"},{"key":"e_1_3_2_1_23_1","volume-title":"inception-resnet and the impact of residual connections on learning","author":"Szegedy Christian","year":"2016","unstructured":"Christian Szegedy , Sergey Ioffe , Vincent Vanhoucke , and Alex Alemi . Inception-v4 , inception-resnet and the impact of residual connections on learning . 2016 . Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alex Alemi. Inception-v4, inception-resnet and the impact of residual connections on learning. 2016."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_25_1","volume-title":"Places: An image database for deep scene understanding. arXiv:1610.02055","author":"Zhou Bolei","year":"2016","unstructured":"Bolei Zhou , Aditya Khosla , Agata Lapedriza , Antonio Torralba , and Aude Oliva . Places: An image database for deep scene understanding. arXiv:1610.02055 , 2016 . Bolei Zhou, Aditya Khosla, Agata Lapedriza, Antonio Torralba, and Aude Oliva. Places: An image database for deep scene understanding. arXiv:1610.02055, 2016."},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASSP.1980.1163420"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2014.6853821"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-013-0636-x"},{"key":"e_1_3_2_1_30_1","unstructured":"Ibm watson speech to text api. http:\/\/www.ibm.com\/watson\/developercloud\/speech-to-text.html.  Ibm watson speech to text api. http:\/\/www.ibm.com\/watson\/developercloud\/speech-to-text.html."},{"issue":"7","key":"e_1_3_2_1_31_1","first-page":"38","article-title":"Distilling the knowledge in a neural network","volume":"14","author":"Hinton Geoffrey","year":"2015","unstructured":"Geoffrey Hinton , Oriol Vinyals , and Jeff Dean . Distilling the knowledge in a neural network . Computer Science , 14 ( 7 ): 38 -- 39 , 2015 . Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. Computer Science, 14(7):38--39, 2015.","journal-title":"Computer Science"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.277"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2007.383036"},{"key":"e_1_3_2_1_35_1","volume-title":"Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980","author":"Kingma Diederik","year":"2014","unstructured":"Diederik Kingma and Jimmy Ba . Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 , 2014 . Diederik Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014."},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/W14-3348"},{"key":"e_1_3_2_1_38_1","volume-title":"Text summarization branches out: Proceedings of the ACL-04 workshop","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin . Rouge: A package for automatic evaluation of summaries . In Text summarization branches out: Proceedings of the ACL-04 workshop , volume 8 . Barcelona , Spain , 2004 . Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out: Proceedings of the ACL-04 workshop, volume 8. Barcelona, Spain, 2004."},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299087"}],"event":{"name":"ICMR '17: International Conference on Multimedia Retrieval","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Bucharest Romania","acronym":"ICMR '17"},"container-title":["Proceedings of the 2017 ACM on International Conference on Multimedia Retrieval"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3078971.3079000","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3078971.3079000","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:03:24Z","timestamp":1750215804000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3078971.3079000"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,6,6]]},"references-count":39,"alternative-id":["10.1145\/3078971.3079000","10.1145\/3078971"],"URL":"https:\/\/doi.org\/10.1145\/3078971.3079000","relation":{},"subject":[],"published":{"date-parts":[[2017,6,6]]},"assertion":[{"value":"2017-06-06","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}