{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,12]],"date-time":"2026-07-12T04:58:11Z","timestamp":1783832291851,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":38,"publisher":"ACM","license":[{"start":{"date-parts":[[2017,10,23]],"date-time":"2017-10-23T00:00:00Z","timestamp":1508716800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"ARO","award":["W911NF-15-1-0354"],"award-info":[{"award-number":["W911NF-15-1-0354"]}]},{"name":"DARPA","award":["W31P4Q-16-C-0091"],"award-info":[{"award-number":["W31P4Q-16-C-0091"]}]},{"name":"NSF NRI","award":["1522904"],"award-info":[{"award-number":["1522904"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2017,10,23]]},"DOI":"10.1145\/3126686.3126717","type":"proceedings-article","created":{"date-parts":[[2017,10,23]],"date-time":"2017-10-23T19:20:32Z","timestamp":1508786432000},"page":"305-313","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":58,"title":["Watch What You Just Said"],"prefix":"10.1145","author":[{"given":"Luowei","family":"Zhou","sequence":"first","affiliation":[{"name":"University of Michigan, Ann Arbor, MI, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chenliang","family":"Xu","sequence":"additional","affiliation":[{"name":"University of Rochester, Rochester, NY, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Parker","family":"Koch","sequence":"additional","affiliation":[{"name":"University of Michigan, Ann Arbor, MI, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jason J.","family":"Corso","sequence":"additional","affiliation":[{"name":"University of Michigan, Ann Arbor, MI, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2017,10,23]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Deep Compositional Captioning: Describing Novel Object Categories Without Paired Training Data. In IEEE Conference on Computer Vision and Pattern Recognition.","author":"Hendricks Lisa Anne","year":"2016","unstructured":"Lisa Anne Hendricks , Subhashini Venugopalan , Marcus Rohrbach , Raymond Mooney , Kate Saenko , and Trevor Darrell . 2016 . Deep Compositional Captioning: Describing Novel Object Categories Without Paired Training Data. In IEEE Conference on Computer Vision and Pattern Recognition. Lisa Anne Hendricks, Subhashini Venugopalan, Marcus Rohrbach, Raymond Mooney, Kate Saenko, and Trevor Darrell. 2016. Deep Compositional Captioning: Describing Novel Object Categories Without Paired Training Data. In IEEE Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_1_2_1","volume-title":"International Conference on Learning Representations.","author":"Bahdanau Dzmitry","year":"2015","unstructured":"Dzmitry Bahdanau , Kyunghyun Cho , and Yoshua Bengio . 2015 . Neural machine translation by jointly learning to align and translate . In International Conference on Learning Representations. Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In International Conference on Learning Representations."},{"key":"e_1_3_2_1_3_1","volume-title":"Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio.","author":"Cho Kyunghyun","year":"2014","unstructured":"Kyunghyun Cho , Bart Van Merri\u00ebnboer , Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 . Learning phrase representations using RNN encoder-decoder for statistical machine translation. In Empirical Methods in Natural Language Processing . Kyunghyun Cho, Bart Van Merri\u00ebnboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning phrase representations using RNN encoder-decoder for statistical machine translation. In Empirical Methods in Natural Language Processing."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/2433396.2433456"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1146\/annurev.ne.18.030195.001205"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298878"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"crossref","unstructured":"Desmond Elliott and Frank Keller. 2013. Image Description using Visual Dependency Representations. In Empirical Methods in Natural Language Processing. Desmond Elliott and Frank Keller. 2013. Image Description using Visual Dependency Representations. In Empirical Methods in Natural Language Processing.","DOI":"10.18653\/v1\/D13-1128"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298754"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/1888089.1888092"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10593-2_35"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.337"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.277"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.494"},{"key":"e_1_3_2_1_16_1","unstructured":"Andrej Karpathy. 2015. neuraltalk2. https:\/\/github.com\/karpathy\/neuraltalk2. (2015). Andrej Karpathy. 2015. neuraltalk2. https:\/\/github.com\/karpathy\/neuraltalk2. (2015)."},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"e_1_3_2_1_18_1","volume-title":"Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980","author":"Kingma Diederik","year":"2014","unstructured":"Diederik Kingma and Jimmy Ba . 2014 . Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014). Diederik Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2012.162"},{"key":"e_1_3_2_1_20_1","volume-title":"Collective generation of natural image descriptions","author":"Kuznetsova Polina","unstructured":"Polina Kuznetsova , Vicente Ordonez , Alexander C. Berg , Tamara L. Berg , and Yejin Choi . 2012. Collective generation of natural image descriptions . In Association for Computational Linguistics . Polina Kuznetsova, Vicente Ordonez, Alexander C. Berg, Tamara L. Berg, and Yejin Choi. 2012. Collective generation of natural image descriptions. In Association for Computational Linguistics."},{"key":"e_1_3_2_1_21_1","unstructured":"Siming Li Girish Kulkarni Tamara L. Berg Alexander C. Berg and Yejin Choi. 2011. Composing simple image descriptions using web-scale n-grams. In Computational Natural Language Learning. Siming Li Girish Kulkarni Tamara L. Berg Alexander C. Berg and Yejin Choi. 2011. Composing simple image descriptions using web-scale n-grams. In Computational Natural Language Learning."},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2984069"},{"key":"e_1_3_2_1_23_1","volume-title":"European Conference on Computer Vision.","author":"Lin Tsung-Yi","unstructured":"Tsung-Yi Lin , Michael Maire , Serge Belongie , James Hays , Pietro Perona , Deva Ramanan , Piotr Doll\u00e1r , and C. Lawrence Zitnick . 2014. Microsoft coco: Common objects in context . In European Conference on Computer Vision. Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C. Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In European Conference on Computer Vision."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46475-6_17"},{"key":"e_1_3_2_1_25_1","volume-title":"International Conference on Learning Representations.","author":"Mao Junhua","year":"2015","unstructured":"Junhua Mao , Wei Xu , Yi Yang , Jiang Wang , Zhiheng Huang , and Alan Yuille . 2015 . Deep captioning with multimodal recurrent neural networks (m-rnn) . In International Conference on Learning Representations. Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, Zhiheng Huang, and Alan Yuille. 2015. Deep captioning with multimodal recurrent neural networks (m-rnn). In International Conference on Learning Representations."},{"key":"e_1_3_2_1_26_1","volume-title":"Midge: Generating image descriptions from computer vision detections. In EACL. Citeseer.","author":"Mitchell Margaret","year":"2012","unstructured":"Margaret Mitchell , Xufeng Han , Jesse Dodge , Alyssa Mensch , Amit Goyal , Alex Berg , Kota Yamaguchi , Tamara Berg , Karl Stratos , and Hal Daum\u00e9 III. 2012 . Midge: Generating image descriptions from computer vision detections. In EACL. Citeseer. Margaret Mitchell, Xufeng Han, Jesse Dodge, Alyssa Mensch, Amit Goyal, Alex Berg, Kota Yamaguchi, Tamara Berg, Karl Stratos, and Hal Daum\u00e9 III. 2012. Midge: Generating image descriptions from computer vision detections. In EACL. Citeseer."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_2_1_28_1","volume-title":"International Conference on Learning Representations.","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman . 2015 . Very deep convolutional networks for large-scale image recognition . In International Conference on Learning Representations. Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations."},{"key":"e_1_3_2_1_29_1","volume-title":"Le","author":"Sutskever Ilya","year":"2014","unstructured":"Ilya Sutskever , Oriol Vinyals , and Quoc V . Le . 2014 . Sequence to Sequence learning with neural networks. In Advances in Neural Information Processing Systems . Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to Sequence learning with neural networks. In Advances in Neural Information Processing Systems."},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2016.2587640"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2964299"},{"key":"e_1_3_2_1_34_1","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition.","author":"Wu Qi","unstructured":"Qi Wu , Chunhua Shen , Lingqiao Liu , Anthony Dick , and Anton van den Hengel. 2016. What value do explicit high level concepts have in vision to language problems? In IEEE Conference on Computer Vision and Pattern Recognition. Qi Wu, Chunhua Shen, Lingqiao Liu, Anthony Dick, and Anton van den Hengel. 2016. What value do explicit high level concepts have in vision to language problems? In IEEE Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_3_2_1_35_1","volume-title":"International Conference on Machine Learning.","author":"Xu Kelvin","year":"2015","unstructured":"Kelvin Xu , Jimmy Ba , Ryan Kiros , Kyunghyun Cho , Aaron Courville , Ruslan Salakhutdinov , Richard S. Zemel , and Yoshua Bengio . 2015 . Show, attend and tell: Neural image caption generation with visual attention . In International Conference on Machine Learning. Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio. 2015. Show, attend and tell: Neural image caption generation with visual attention. In International Conference on Machine Learning."},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.5555\/2145432.2145484"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2964288"},{"key":"e_1_3_2_1_38_1","volume-title":"Image Captioning With Semantic Attention. In IEEE Conference on Computer Vision and Pattern Recognition.","author":"You Quanzeng","year":"2016","unstructured":"Quanzeng You , Hailin Jin , Zhaowen Wang , Chen Fang , and Jiebo Luo . 2016 . Image Captioning With Semantic Attention. In IEEE Conference on Computer Vision and Pattern Recognition. Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo. 2016. Image Captioning With Semantic Attention. In IEEE Conference on Computer Vision and Pattern Recognition."}],"event":{"name":"MM '17: ACM Multimedia Conference","location":"Mountain View California USA","acronym":"MM '17","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the on Thematic Workshops of ACM Multimedia 2017"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3126686.3126717","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3126686.3126717","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3126686.3126717","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,26]],"date-time":"2025-06-26T17:48:26Z","timestamp":1750960106000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3126686.3126717"}},"subtitle":["Image Captioning with Text-Conditional Attention"],"short-title":[],"issued":{"date-parts":[[2017,10,23]]},"references-count":38,"alternative-id":["10.1145\/3126686.3126717","10.1145\/3126686"],"URL":"https:\/\/doi.org\/10.1145\/3126686.3126717","relation":{},"subject":[],"published":{"date-parts":[[2017,10,23]]},"assertion":[{"value":"2017-10-23","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}