{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,21]],"date-time":"2025-10-21T15:39:11Z","timestamp":1761061151932,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":42,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T00:00:00Z","timestamp":1602460800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Natural Science Foundation of Guangdong Province in China","award":["No.2019B1515120049"],"award-info":[{"award-number":["No.2019B1515120049"]}]},{"name":"National Natural Science Foundation of China","award":["No.U1705262; No.61772443; No.61572410; No.61802324; No.617021"],"award-info":[{"award-number":["No.U1705262; No.61772443; No.61572410; No.61802324; No.617021"]}]},{"name":"National Key R&D Program","award":["No.2017YFC0113000; No.2016YFB1001503"],"award-info":[{"award-number":["No.2017YFC0113000; No.2016YFB1001503"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,10,12]]},"DOI":"10.1145\/3394171.3414009","type":"proceedings-article","created":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T12:26:53Z","timestamp":1602505613000},"page":"4226-4234","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Attacking Image Captioning Towards Accuracy-Preserving Target Words Removal"],"prefix":"10.1145","author":[{"given":"Jiayi","family":"Ji","sequence":"first","affiliation":[{"name":"Xiamen University, Xiamen City, Fujian Province, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaoshuai","family":"Sun","sequence":"additional","affiliation":[{"name":"Xiamen University, Xiamen City, Fujian Province, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yiyi","family":"Zhou","sequence":"additional","affiliation":[{"name":"Xiamen University, Xiamen City, Fujian Province, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rongrong","family":"Ji","sequence":"additional","affiliation":[{"name":"Xiamen University, Xiamen City, Fujian Province, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fuhai","family":"Chen","sequence":"additional","affiliation":[{"name":"Xiamen University, Xiamen City, Fujian Province, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianzhuang","family":"Liu","sequence":"additional","affiliation":[{"name":"Noah's Ark Lab, Huawei Technologies, Beijing City, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qi","family":"Tian","sequence":"additional","affiliation":[{"name":"Huawei Cloud BU, Huawei Technologies, Beijing City, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,10,12]]},"reference":[{"volume-title":"Spice: Semantic propositional image caption evaluation. In ECCV.","year":"2016","author":"Anderson Peter","key":"e_1_3_2_2_1_1"},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"crossref","unstructured":"Peter Anderson Xiaodong He Chris Buehler Damien Teney Mark Johnson Stephen Gould and Lei Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In CVPR.  Peter Anderson Xiaodong He Chris Buehler Damien Teney Mark Johnson Stephen Gould and Lei Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In CVPR.","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"crossref","unstructured":"Jyoti Aneja Harsh Agrawal Dhruv Batra and Alexander Schwing. 2019. Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning. In ICCV.  Jyoti Aneja Harsh Agrawal Dhruv Batra and Alexander Schwing. 2019. Sequential Latent Spaces for Modeling the Intention During Diverse Image Captioning. In ICCV.","DOI":"10.1109\/ICCV.2019.00436"},{"volume-title":"Vqa: Visual question answering. In ICCV.","year":"2015","author":"Antol Stanislaw","key":"e_1_3_2_2_4_1"},{"volume-title":"METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In ACL (Workshops).","year":"2005","author":"Banerjee Satanjeev","key":"e_1_3_2_2_5_1"},{"key":"e_1_3_2_2_6_1","unstructured":"Samy Bengio Oriol Vinyals Navdeep Jaitly and Noam Shazeer. 2015. Scheduled sampling for sequence prediction with recurrent neural networks. In NeurIPS.  Samy Bengio Oriol Vinyals Navdeep Jaitly and Noam Shazeer. 2015. Scheduled sampling for sequence prediction with recurrent neural networks. In NeurIPS."},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"crossref","unstructured":"Nicholas Carlini and David Wagner. 2017. Towards evaluating the robustness of neural networks. In IEEE SSP.  Nicholas Carlini and David Wagner. 2017. Towards evaluating the robustness of neural networks. In IEEE SSP.","DOI":"10.1109\/SP.2017.49"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"crossref","unstructured":"Hongge Chen Huan Zhang Pin-Yu Chen Jinfeng Yi and Cho-Jui Hsieh. 2018. Attacking Visual Language Grounding with Adversarial Examples: A Case Study on Neural Image Captioning. In ACL.  Hongge Chen Huan Zhang Pin-Yu Chen Jinfeng Yi and Cho-Jui Hsieh. 2018. Attacking Visual Language Grounding with Adversarial Examples: A Case Study on Neural Image Captioning. In ACL.","DOI":"10.18653\/v1\/P18-1241"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"crossref","unstructured":"Shizhe Chen Qin Jin Peng Wang and Qi Wu. 2020. Say As You Wish: Fine-grained Control of Image Caption Generation with Abstract Scene Graphs. arXiv preprint (2020).  Shizhe Chen Qin Jin Peng Wang and Qi Wu. 2020. Say As You Wish: Fine-grained Control of Image Caption Generation with Abstract Scene Graphs. arXiv preprint (2020).","DOI":"10.1109\/CVPR42600.2020.00998"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"crossref","unstructured":"Marcella Cornia Lorenzo Baraldi and Rita Cucchiara. 2019. Show control and tell: a framework for generating controllable and grounded captions. In CVPR.  Marcella Cornia Lorenzo Baraldi and Rita Cucchiara. 2019. Show control and tell: a framework for generating controllable and grounded captions. In CVPR.","DOI":"10.1109\/CVPR.2019.00850"},{"volume-title":"Imagenet: A large-scale hierarchical image database. In CVPR.","year":"2009","author":"Deng Jia","key":"e_1_3_2_2_11_1"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"crossref","unstructured":"Aditya Deshpande Jyoti Aneja Liwei Wang Alexander G Schwing and David Forsyth. 2019. Fast diverse and accurate image captioning guided by part-of-speech. In CVPR.  Aditya Deshpande Jyoti Aneja Liwei Wang Alexander G Schwing and David Forsyth. 2019. Fast diverse and accurate image captioning guided by part-of-speech. In CVPR.","DOI":"10.1109\/CVPR.2019.01095"},{"key":"e_1_3_2_2_13_1","unstructured":"Harris Drucker Christopher JC Burges Linda Kaufman Alex J Smola and Vladimir Vapnik. 1997. Support vector regression machines. In NeurIPS.  Harris Drucker Christopher JC Burges Linda Kaufman Alex J Smola and Vladimir Vapnik. 1997. Support vector regression machines. In NeurIPS."},{"key":"e_1_3_2_2_14_1","unstructured":"Lianli Gao Xiangpeng Li Jingkuan Song and Heng Tao Shen. 2019. Hierarchical LSTMs with adaptive attention for visual captioning. IEEE TPAMI (2019).  Lianli Gao Xiangpeng Li Jingkuan Song and Heng Tao Shen. 2019. Hierarchical LSTMs with adaptive attention for visual captioning. IEEE TPAMI (2019)."},{"key":"e_1_3_2_2_15_1","doi-asserted-by":"crossref","unstructured":"Ross Girshick Jeff Donahue Trevor Darrell and Jitendra Malik. 2014. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR.  Ross Girshick Jeff Donahue Trevor Darrell and Jitendra Malik. 2014. Rich feature hierarchies for accurate object detection and semantic segmentation. In CVPR.","DOI":"10.1109\/CVPR.2014.81"},{"key":"e_1_3_2_2_16_1","unstructured":"Ian J Goodfellow Jonathon Shlens and Christian Szegedy. 2014. Explaining and harnessing adversarial examples. arXiv preprint (2014).  Ian J Goodfellow Jonathon Shlens and Christian Szegedy. 2014. Explaining and harnessing adversarial examples. arXiv preprint (2014)."},{"key":"e_1_3_2_2_17_1","unstructured":"Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep residual learning for image recognition. In CVPR.  Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep residual learning for image recognition. In CVPR."},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"crossref","unstructured":"Lun Huang Wenmin Wang Jie Chen and Xiao-Yong Wei. 2019 a. Attention on attention for image captioning. In CVPR.  Lun Huang Wenmin Wang Jie Chen and Xiao-Yong Wei. 2019 a. Attention on attention for image captioning. In CVPR.","DOI":"10.1109\/ICCV.2019.00473"},{"key":"e_1_3_2_2_19_1","unstructured":"Lun Huang Wenmin Wang Yaxian Xia and Jie Chen. 2019 b. Adaptively Aligned Image Captioning via Adaptive Attention Time. In NeurIPS.  Lun Huang Wenmin Wang Yaxian Xia and Jie Chen. 2019 b. Adaptively Aligned Image Captioning via Adaptive Attention Time. In NeurIPS."},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"crossref","unstructured":"Radu Tudor Ionescu Bogdan Alexe Marius Leordeanu Marius Popescu Dim Papadopoulos and Vittorio Ferrari. 2016. How hard can it be? Estimating the difficulty of visual search in an image. In CVPR.  Radu Tudor Ionescu Bogdan Alexe Marius Leordeanu Marius Popescu Dim Papadopoulos and Vittorio Ferrari. 2016. How hard can it be? Estimating the difficulty of visual search in an image. In CVPR.","DOI":"10.1109\/CVPR.2016.237"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"crossref","unstructured":"Andrej Karpathy and Li Fei-Fei. 2015. Deep visual-semantic alignments for generating image descriptions. In CVPR.  Andrej Karpathy and Li Fei-Fei. 2015. Deep visual-semantic alignments for generating image descriptions. In CVPR.","DOI":"10.1109\/CVPR.2015.7298932"},{"volume-title":"Adam: A method for stochastic optimization. arXiv preprint","year":"2014","author":"Kingma Diederik P","key":"e_1_3_2_2_22_1"},{"key":"e_1_3_2_2_23_1","unstructured":"Alex Krizhevsky Ilya Sutskever and Geoffrey E Hinton. 2012. ImageNet Classification with Deep Convolutional Neural Networks. In NeurIPS.  Alex Krizhevsky Ilya Sutskever and Geoffrey E Hinton. 2012. ImageNet Classification with Deep Convolutional Neural Networks. In NeurIPS."},{"key":"e_1_3_2_2_24_1","unstructured":"Alexey Kurakin Ian Goodfellow and Samy Bengio. 2016. Adversarial examples in the physical world. arXiv preprint (2016).  Alexey Kurakin Ian Goodfellow and Samy Bengio. 2016. Adversarial examples in the physical world. arXiv preprint (2016)."},{"volume-title":"Rouge: A package for automatic evaluation of summaries. In ACL (Workshops).","year":"2004","author":"Lin Chin-Yew","key":"e_1_3_2_2_25_1"},{"key":"e_1_3_2_2_26_1","unstructured":"Tsung-Yi Lin Michael Maire Serge Belongie James Hays Pietro Perona Deva Ramanan Piotr Doll\u00e1r and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In ECCV.  Tsung-Yi Lin Michael Maire Serge Belongie James Hays Pietro Perona Deva Ramanan Piotr Doll\u00e1r and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In ECCV."},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"crossref","unstructured":"Jonathan Long Evan Shelhamer and Trevor Darrell. 2015. Fully convolutional networks for semantic segmentation. In CVPR.  Jonathan Long Evan Shelhamer and Trevor Darrell. 2015. Fully convolutional networks for semantic segmentation. In CVPR.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"e_1_3_2_2_28_1","unstructured":"Jiasen Lu Caiming Xiong Devi Parikh and Richard Socher. 2017. Knowing when to look: Adaptive attention via a visual sentinel for image captioning. In CVPR.  Jiasen Lu Caiming Xiong Devi Parikh and Richard Socher. 2017. Knowing when to look: Adaptive attention via a visual sentinel for image captioning. In CVPR."},{"volume-title":"Semstyle: Learning to generate stylised image captions using unaligned text. In CVPR.","year":"2018","author":"Mathews Alexander","key":"e_1_3_2_2_29_1"},{"volume-title":"Senticap: Generating image descriptions with sentiments. In AAAI.","year":"2016","author":"Mathews Alexander Patrick","key":"e_1_3_2_2_30_1"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"crossref","unstructured":"Seyed-Mohsen Moosavi-Dezfooli Alhussein Fawzi and Pascal Frossard. 2016. Deepfool: a simple and accurate method to fool deep neural networks. In CVPR.  Seyed-Mohsen Moosavi-Dezfooli Alhussein Fawzi and Pascal Frossard. 2016. Deepfool: a simple and accurate method to fool deep neural networks. In CVPR.","DOI":"10.1109\/CVPR.2016.282"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"crossref","unstructured":"Kishore Papineni Salim Roukos Todd Ward and Wei-Jing Zhu. 2002. BLEU: a method for automatic evaluation of machine translation. In ACL.  Kishore Papineni Salim Roukos Todd Ward and Wei-Jing Zhu. 2002. BLEU: a method for automatic evaluation of machine translation. In ACL.","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_3_2_2_33_1","unstructured":"Shaoqing Ren Kaiming He Ross Girshick and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. In NeurIPS.  Shaoqing Ren Kaiming He Ross Girshick and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. In NeurIPS."},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"crossref","unstructured":"Steven J Rennie Etienne Marcheret Youssef Mroueh Jerret Ross and Vaibhava Goel. 2017. Self-critical sequence training for image captioning. In CVPR.  Steven J Rennie Etienne Marcheret Youssef Mroueh Jerret Ross and Vaibhava Goel. 2017. Self-critical sequence training for image captioning. In CVPR.","DOI":"10.1109\/CVPR.2017.131"},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"crossref","unstructured":"Kurt Shuster Samuel Humeau Hexiang Hu Antoine Bordes and Jason Weston. 2019. Engaging image captioning via personality. In CVPR.  Kurt Shuster Samuel Humeau Hexiang Hu Antoine Bordes and Jason Weston. 2019. Engaging image captioning via personality. In CVPR.","DOI":"10.1109\/CVPR.2019.01280"},{"key":"e_1_3_2_2_36_1","unstructured":"Christian Szegedy Wojciech Zaremba Ilya Sutskever Joan Bruna Dumitru Erhan Ian Goodfellow and Rob Fergus. 2013. Intriguing properties of neural networks. arXiv preprint (2013).  Christian Szegedy Wojciech Zaremba Ilya Sutskever Joan Bruna Dumitru Erhan Ian Goodfellow and Rob Fergus. 2013. Intriguing properties of neural networks. arXiv preprint (2013)."},{"volume-title":"Cider: Consensus-based image description evaluation. In CVPR.","year":"2015","author":"Vedantam Ramakrishna","key":"e_1_3_2_2_37_1"},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"crossref","unstructured":"Ashwin K Vijayakumar Michael Cogswell Ramprasaath R Selvaraju Qing Sun Stefan Lee David Crandall and Dhruv Batra. 2018. Diverse beam search for improved description of complex scenes. In AAAI.  Ashwin K Vijayakumar Michael Cogswell Ramprasaath R Selvaraju Qing Sun Stefan Lee David Crandall and Dhruv Batra. 2018. Diverse beam search for improved description of complex scenes. In AAAI.","DOI":"10.1609\/aaai.v32i1.12340"},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"crossref","unstructured":"Oriol Vinyals Alexander Toshev Samy Bengio and Dumitru Erhan. 2015. Show and tell: A neural image caption generator. In CVPR.  Oriol Vinyals Alexander Toshev Samy Bengio and Dumitru Erhan. 2015. Show and tell: A neural image caption generator. In CVPR.","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"e_1_3_2_2_40_1","unstructured":"Cihang Xie Jianyu Wang Zhishuai Zhang Yuyin Zhou Lingxi Xie and Alan Yuille. 2017. Adversarial examples for semantic segmentation and object detection. In ICCV.  Cihang Xie Jianyu Wang Zhishuai Zhang Yuyin Zhou Lingxi Xie and Alan Yuille. 2017. Adversarial examples for semantic segmentation and object detection. In ICCV."},{"key":"e_1_3_2_2_41_1","unstructured":"Kelvin Xu Jimmy Ba Ryan Kiros Kyunghyun Cho Aaron Courville Ruslan Salakhudinov Rich Zemel and Yoshua Bengio. 2015. Show attend and tell: Neural image caption generation with visual attention. In ICML.  Kelvin Xu Jimmy Ba Ryan Kiros Kyunghyun Cho Aaron Courville Ruslan Salakhudinov Rich Zemel and Yoshua Bengio. 2015. Show attend and tell: Neural image caption generation with visual attention. In ICML."},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"crossref","unstructured":"Yan Xu Baoyuan Wu Fumin Shen Yanbo Fan Yong Zhang Heng Tao Shen and Wei Liu. 2019. Exact Adversarial Attack to Image Captioning via Structured Output Learning with Latent Variables. In CVPR.  Yan Xu Baoyuan Wu Fumin Shen Yanbo Fan Yong Zhang Heng Tao Shen and Wei Liu. 2019. Exact Adversarial Attack to Image Captioning via Structured Output Learning with Latent Variables. In CVPR.","DOI":"10.1109\/CVPR.2019.00426"}],"event":{"name":"MM '20: The 28th ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Seattle WA USA","acronym":"MM '20"},"container-title":["Proceedings of the 28th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3414009","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3394171.3414009","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:32:07Z","timestamp":1750195927000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3414009"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,12]]},"references-count":42,"alternative-id":["10.1145\/3394171.3414009","10.1145\/3394171"],"URL":"https:\/\/doi.org\/10.1145\/3394171.3414009","relation":{},"subject":[],"published":{"date-parts":[[2020,10,12]]},"assertion":[{"value":"2020-10-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}