{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,23]],"date-time":"2025-08-23T05:21:35Z","timestamp":1755926495495,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":48,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T00:00:00Z","timestamp":1602460800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,10,12]]},"DOI":"10.1145\/3394171.3413885","type":"proceedings-article","created":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T13:10:18Z","timestamp":1602508218000},"page":"1112-1121","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":16,"title":["Structural Semantic Adversarial Active Learning for Image Captioning"],"prefix":"10.1145","author":[{"given":"Beichen","family":"Zhang","sequence":"first","affiliation":[{"name":"University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liang","family":"Li","sequence":"additional","affiliation":[{"name":"Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Li","family":"Su","sequence":"additional","affiliation":[{"name":"University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuhui","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jincan","family":"Deng","sequence":"additional","affiliation":[{"name":"Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zheng-Jun","family":"Zha","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qingming","family":"Huang","sequence":"additional","affiliation":[{"name":"University of Chinese Academy of Sciences &amp; Institute of Computing Technology, Chinese Academy of Sciences, Beijing , China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,10,12]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018","author":"Anderson Peter","year":"2018","unstructured":"Peter Anderson , Xiaodong He , Chris Buehler , Damien Teney , Mark Johnson , Stephen Gould , and Lei Zhang . 2018 . Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018 , Salt Lake City, UT, USA, June 18--22 , 2018. IEEE Computer Society, 6077--6086. Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018. Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18--22, 2018. IEEE Computer Society, 6077--6086."},{"key":"e_1_3_2_2_2_1","volume-title":"Queries and concept learning. Machine learning","author":"Angluin Dana","year":"1988","unstructured":"Dana Angluin . 1988. Queries and concept learning. Machine learning , Vol. 2 , 4 ( 1988 ), 319--342. Dana Angluin. 1988. Queries and concept learning. Machine learning, Vol. 2, 4 (1988), 319--342."},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00976"},{"key":"e_1_3_2_2_4_1","volume-title":"Proceedings of the Twenty-First Annual Conference on Neural Information Processing Systems","author":"Bonilla Edwin V.","year":"2007","unstructured":"Edwin V. Bonilla , Kian Ming Adam Chai , and Christopher K. I. Williams . 2007. Multi-task Gaussian Process Prediction. In Advances in Neural Information Processing Systems 20 , Proceedings of the Twenty-First Annual Conference on Neural Information Processing Systems , Vancouver, British Columbia, Canada, December 3--6 , 2007 . Curran Associates, Inc., 153--160. Edwin V. Bonilla, Kian Ming Adam Chai, and Christopher K. I. Williams. 2007. Multi-task Gaussian Process Prediction. In Advances in Neural Information Processing Systems 20, Proceedings of the Twenty-First Annual Conference on Neural Information Processing Systems, Vancouver, British Columbia, Canada, December 3--6, 2007. Curran Associates, Inc., 153--160."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1016\/B978-1-55860-377-6.50027-X"},{"key":"e_1_3_2_2_6_1","volume-title":"Proceedings of the Ninth Workshop on Statistical Machine Translation, ACL 2014","author":"Michael","year":"2014","unstructured":"Michael J. Denkowski and Alon Lavie. 2014. Meteor Universal: Language Specific Translation Evaluation for Any Target Language . In Proceedings of the Ninth Workshop on Statistical Machine Translation, ACL 2014 , June 26 --27 , 2014 , Baltimore, Maryland, USA. The Association for Computer Linguistics, 376--380. Michael J. Denkowski and Alon Lavie. 2014. Meteor Universal: Language Specific Translation Evaluation for Any Target Language. In Proceedings of the Ninth Workshop on Statistical Machine Translation, ACL 2014, June 26--27, 2014, Baltimore, Maryland, USA. The Association for Computer Linguistics, 376--380."},{"key":"e_1_3_2_2_8_1","volume-title":"Adversarial active learning for deep networks: a margin based approach. arXiv preprint arXiv:1802.09841","author":"Ducoffe Melanie","year":"2018","unstructured":"Melanie Ducoffe and Frederic Precioso . 2018. Adversarial active learning for deep networks: a margin based approach. arXiv preprint arXiv:1802.09841 ( 2018 ). Melanie Ducoffe and Frederic Precioso. 2018. Adversarial active learning for deep networks: a margin based approach. arXiv preprint arXiv:1802.09841 (2018)."},{"key":"e_1_3_2_2_9_1","series-title":"Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL","volume-title":"Long Papers","author":"Fan Zhihao","year":"2019","unstructured":"Zhihao Fan , Zhongyu Wei , Siyuan Wang , and Xuanjing Huang . 2019. Bridging by Word: Image Grounded Vocabulary Construction for Visual Captioning . In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019 , Florence , Italy, July 28- August 2, 2019, Volume 1 : Long Papers . Association for Computational Linguistics , 6514--6524. Zhihao Fan, Zhongyu Wei, Siyuan Wang, and Xuanjing Huang. 2019. Bridging by Word: Image Grounded Vocabulary Construction for Visual Captioning. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers. Association for Computational Linguistics, 6514--6524."},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298754"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10593-2_37"},{"key":"e_1_3_2_2_12_1","volume-title":"international conference on machine learning. 1050--1059","author":"Gal Yarin","year":"2016","unstructured":"Yarin Gal and Zoubin Ghahramani . 2016 . Dropout as a bayesian approximation: Representing model uncertainty in deep learning . In international conference on machine learning. 1050--1059 . Yarin Gal and Zoubin Ghahramani. 2016. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In international conference on machine learning. 1050--1059."},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.5555\/3305381.3305504"},{"key":"e_1_3_2_2_14_1","unstructured":"Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative adversarial nets. In NIPS. 2672--2680.  Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative adversarial nets. In NIPS. 2672--2680."},{"key":"e_1_3_2_2_15_1","volume-title":"Cost-effective active learning for melanoma segmentation. arXiv preprint arXiv:1711.09168","author":"Gorriz Marc","year":"2017","unstructured":"Marc Gorriz , Axel Carlier , Emmanuel Faure , and Xavier Giro-i Nieto . 2017. Cost-effective active learning for melanoma segmentation. arXiv preprint arXiv:1711.09168 ( 2017 ). Marc Gorriz, Axel Carlier, Emmanuel Faure, and Xavier Giro-i Nieto. 2017. Cost-effective active learning for melanoma segmentation. arXiv preprint arXiv:1711.09168 (2017)."},{"key":"e_1_3_2_2_16_1","volume-title":"Deep Residual Learning for Image Recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016","author":"He Kaiming","year":"2016","unstructured":"Kaiming He , Xiangyu Zhang , Shaoqing Ren , and Jian Sun . 2016 . Deep Residual Learning for Image Recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016 , Las Vegas, NV, USA, June 27--30 , 2016. IEEE Computer Society, 770--778. Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27--30, 2016. IEEE Computer Society, 770--778."},{"volume-title":"European Conference on Computer Vision. Springer, 67--84","author":"Joulin Armand","key":"e_1_3_2_2_17_1","unstructured":"Armand Joulin , Laurens van der Maaten, Allan Jabri, and Nicolas Vasilache. 2016. Learning visual features from large weakly supervised data . In European Conference on Computer Vision. Springer, 67--84 . Armand Joulin, Laurens van der Maaten, Allan Jabri, and Nicolas Vasilache. 2016. Learning visual features from large weakly supervised data. In European Conference on Computer Vision. Springer, 67--84."},{"key":"e_1_3_2_2_18_1","volume-title":"Active and continuous exploration with deep neural networks and expected model output changes. arXiv preprint arXiv:1612.06129","author":"Christoph","year":"2016","unstructured":"Christoph K\"ading, Erik Rodner , Alexander Freytag , and Joachim Denzler . 2016. Active and continuous exploration with deep neural networks and expected model output changes. arXiv preprint arXiv:1612.06129 ( 2016 ). Christoph K\"ading, Erik Rodner, Alexander Freytag, and Joachim Denzler. 2016. Active and continuous exploration with deep neural networks and expected model output changes. arXiv preprint arXiv:1612.06129 (2016)."},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSP.2002.1003062"},{"key":"e_1_3_2_2_22_1","volume-title":"Cost-Sensitive Active Learning for Intracranial Hemorrhage Detection. In International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 715--723","author":"Kuo Weicheng","year":"2018","unstructured":"Weicheng Kuo , Christian H\"ane, Esther Yuh , Pratik Mukherjee , and Jitendra Malik . 2018 . Cost-Sensitive Active Learning for Intracranial Hemorrhage Detection. In International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 715--723 . Weicheng Kuo, Christian H\"ane, Esther Yuh, Pratik Mukherjee, and Jitendra Malik. 2018. Cost-Sensitive Active Learning for Intracranial Hemorrhage Detection. In International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 715--723."},{"key":"e_1_3_2_2_23_1","volume-title":"Collective Generation of Natural Image Descriptions. In The 50th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference, July 8--14","volume":"368","author":"Kuznetsova Polina","year":"2012","unstructured":"Polina Kuznetsova , Vicente Ordonez , Alexander C. Berg , Tamara L. Berg , and Yejin Choi . 2012 . Collective Generation of Natural Image Descriptions. In The 50th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference, July 8--14 , 2012, Jeju Island, Korea - Volume 1: Long Papers. The Association for Computer Linguistics, 359-- 368 . Polina Kuznetsova, Vicente Ordonez, Alexander C. Berg, Tamara L. Berg, and Yejin Choi. 2012. Collective Generation of Natural Image Descriptions. In The 50th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference, July 8--14, 2012, Jeju Island, Korea - Volume 1: Long Papers. The Association for Computer Linguistics, 359--368."},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2012.2194993"},{"key":"e_1_3_2_2_25_1","volume-title":"Proceedings of the Fifteenth Conference on Computational Natural Language Learning, CoNLL 2011","author":"Li Siming","year":"2011","unstructured":"Siming Li , Girish Kulkarni , Tamara L. Berg , Alexander C. Berg , and Yejin Choi . [n.d.]. Composing Simple Image Descriptions using Web-scale N-grams . In Proceedings of the Fifteenth Conference on Computational Natural Language Learning, CoNLL 2011 , Portland, Oregon, USA, June 23--24 , 2011 . 220--228. Siming Li, Girish Kulkarni, Tamara L. Berg, Alexander C. Berg, and Yejin Choi. [n.d.]. Composing Simple Image Descriptions using Web-scale N-grams. In Proceedings of the Fifteenth Conference on Computational Natural Language Learning, CoNLL 2011, Portland, Oregon, USA, June 23--24, 2011. 220--228."},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.116"},{"key":"e_1_3_2_2_27_1","volume-title":"In Proceedings of the Workshop on Text Summarization Branches Out (WAS","author":"Lin Chin Yew","year":"2004","unstructured":"Chin Yew Lin . 2004 . ROUGE: A Package for Automatic Evaluation of summaries . In In Proceedings of the Workshop on Text Summarization Branches Out (WAS 2004). Chin Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of summaries. In In Proceedings of the Workshop on Text Summarization Branches Out (WAS 2004)."},{"key":"e_1_3_2_2_28_1","volume-title":"Piotr Doll\u00e1 r, and C. Lawrence Zitnick","author":"Lin Tsung-Yi","year":"2014","unstructured":"Tsung-Yi Lin , Michael Maire , Serge J. Belongie , James Hays , Pietro Perona , Deva Ramanan , Piotr Doll\u00e1 r, and C. Lawrence Zitnick . 2014 . Microsoft COCO: Common Objects in Context. In Computer Vision - ECCV 2014 - 13th European Conference, Zurich, Switzerland, September 6--12, 2014, Proceedings, Part V (Lecture Notes in Computer Science), Vol. 8693 . Springer , 740--755. Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1 r, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common Objects in Context. In Computer Vision - ECCV 2014 - 13th European Conference, Zurich, Switzerland, September 6--12, 2014, Proceedings, Part V (Lecture Notes in Computer Science), Vol. 8693. Springer, 740--755."},{"key":"e_1_3_2_2_29_1","volume-title":"Adaptive Reconstruction Network for Weakly Supervised Referring Expression Grounding. In IEEE\/CVF International Conference on Computer Vision, ICCV. IEEE, 2611--2620","author":"Liu Xuejing","year":"2019","unstructured":"Xuejing Liu , Liang Li , Shuhui Wang , Zheng-Jun Zha , Dechao Meng , and Qingming Huang . 2019 . Adaptive Reconstruction Network for Weakly Supervised Referring Expression Grounding. In IEEE\/CVF International Conference on Computer Vision, ICCV. IEEE, 2611--2620 . Xuejing Liu, Liang Li, Shuhui Wang, Zheng-Jun Zha, Dechao Meng, and Qingming Huang. 2019. Adaptive Reconstruction Network for Weakly Supervised Referring Expression Grounding. In IEEE\/CVF International Conference on Computer Vision, ICCV. IEEE, 2611--2620."},{"key":"e_1_3_2_2_30_1","volume-title":"Neural Baby Talk. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018","author":"Lu Jiasen","year":"2018","unstructured":"Jiasen Lu , Jianwei Yang , Dhruv Batra , and Devi Parikh . 2018 . Neural Baby Talk. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018 , Salt Lake City, UT, USA, June 18--22 , 2018. IEEE Computer Society, 7219--7228. Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2018. Neural Baby Talk. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18--22, 2018. IEEE Computer Society, 7219--7228."},{"volume-title":"Proceedings of the European Conference on Computer Vision (ECCV). 181--196","author":"Mahajan Dhruv","key":"e_1_3_2_2_31_1","unstructured":"Dhruv Mahajan , Ross Girshick , Vignesh Ramanathan , Kaiming He , Manohar Paluri , Yixuan Li , Ashwin Bharambe , and Laurens van der Maaten. 2018. Exploring the limits of weakly supervised pretraining . In Proceedings of the European Conference on Computer Vision (ECCV). 181--196 . Dhruv Mahajan, Ross Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens van der Maaten. 2018. Exploring the limits of weakly supervised pretraining. In Proceedings of the European Conference on Computer Vision (ECCV). 181--196."},{"key":"e_1_3_2_2_32_1","volume-title":"Advances in Neural Information Processing Systems 24: 25th Annual Conference on Neural Information Processing Systems 2011. Proceedings of a meeting held 12--14","author":"Ordonez Vicente","year":"2011","unstructured":"Vicente Ordonez , Girish Kulkarni , and Tamara L. Berg . 2011. Im2Text: Describing Images Using 1 Million Captioned Photographs . In Advances in Neural Information Processing Systems 24: 25th Annual Conference on Neural Information Processing Systems 2011. Proceedings of a meeting held 12--14 December 2011 , Granada, Spain. 1143--1151. Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg. 2011. Im2Text: Describing Images Using 1 Million Captioned Photographs. In Advances in Neural Information Processing Systems 24: 25th Annual Conference on Neural Information Processing Systems 2011. Proceedings of a meeting held 12--14 December 2011, Granada, Spain. 1143--1151."},{"key":"e_1_3_2_2_33_1","volume-title":"Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6--12","author":"Papineni Kishore","year":"2002","unstructured":"Kishore Papineni , Salim Roukos , Todd Ward , and Wei-Jing Zhu . 2002 . Bleu: a Method for Automatic Evaluation of Machine Translation . In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6--12 , 2002, Philadelphia, PA, USA. ACL, 311--318. Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6--12, 2002, Philadelphia, PA, USA. ACL, 311--318."},{"key":"e_1_3_2_2_34_1","volume-title":"Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren , Kaiming He , Ross B. Girshick , and Jian Sun . 2015 . Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015 , December 7 --12 , 2015, Montreal, Quebec, Canada. 91--99. Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2015. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7--12, 2015, Montreal, Quebec, Canada. 91--99."},{"key":"e_1_3_2_2_35_1","volume-title":"Williamstown","author":"Roy Nicholas","year":"2001","unstructured":"Nicholas Roy and Andrew McCallum . 2001. Toward optimal active learning through monte carlo estimation of error reduction. ICML , Williamstown ( 2001 ), 441--448. Nicholas Roy and Andrew McCallum. 2001. Toward optimal active learning through monte carlo estimation of error reduction. ICML, Williamstown (2001), 441--448."},{"key":"e_1_3_2_2_36_1","volume-title":"Active learning for convolutional neural networks: A core-set approach. arXiv preprint arXiv:1708.00489","author":"Sener Ozan","year":"2017","unstructured":"Ozan Sener and Silvio Savarese . 2017. Active learning for convolutional neural networks: A core-set approach. arXiv preprint arXiv:1708.00489 ( 2017 ). Ozan Sener and Silvio Savarese. 2017. Active learning for convolutional neural networks: A core-set approach. arXiv preprint arXiv:1708.00489 (2017)."},{"key":"e_1_3_2_2_38_1","unstructured":"Burr Settles Mark Craven and Soumya Ray. 2008. Multiple-instance active learning. In Advances in neural information processing systems. 1289--1296.  Burr Settles Mark Craven and Soumya Ray. 2008. Multiple-instance active learning. In Advances in neural information processing systems. 1289--1296."},{"key":"e_1_3_2_2_39_1","volume-title":"Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9--15","volume":"97","author":"Shi Weishi","year":"2019","unstructured":"Weishi Shi and Qi Yu. [n.d.]. Fast Direct Search in an Optimally Compressed Continuous Target Space for Efficient Multi-Label Active Learning . In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9--15 June 2019 , Long Beach, California, USA (Proceedings of Machine Learning Research) , Vol. 97 . PMLR, 5769--5778. Weishi Shi and Qi Yu. [n.d.]. Fast Direct Search in an Optimally Compressed Continuous Target Space for Efficient Multi-Label Active Learning. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9--15 June 2019, Long Beach, California, USA (Proceedings of Machine Learning Research), Vol. 97. PMLR, 5769--5778."},{"key":"e_1_3_2_2_40_1","volume-title":"3rd International Conference on Learning Representations, ICLR","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman . 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition . In 3rd International Conference on Learning Representations, ICLR 2015 , San Diego, CA , USA, May 7--9, 2015, Conference Track Proceedings . Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7--9, 2015, Conference Track Proceedings."},{"key":"e_1_3_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00607"},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"e_1_3_2_2_44_1","volume-title":"Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6--11 July 2015 (JMLR Workshop and Conference Proceedings)","volume":"37","author":"Xu Kelvin","year":"2015","unstructured":"Kelvin Xu , Jimmy Ba , Ryan Kiros , Kyunghyun Cho , Aaron C. Courville , Ruslan Salakhutdinov , Richard S. Zemel , and Yoshua Bengio . 2015 . Show, Attend and Tell: Neural Image Caption Generation with Visual Attention . In Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6--11 July 2015 (JMLR Workshop and Conference Proceedings) , Vol. 37 . 2048--2057. Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C. Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio. 2015. Show, Attend and Tell: Neural Image Caption Generation with Visual Attention. In Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6--11 July 2015 (JMLR Workshop and Conference Proceedings), Vol. 37. 2048--2057."},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-66179-7_46"},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2912735"},{"key":"e_1_3_2_2_47_1","volume-title":"Proceedings, Part XIV (Lecture Notes in Computer Science)","volume":"11218","author":"Yao Ting","year":"2018","unstructured":"Ting Yao , Yingwei Pan , Yehao Li , and Tao Mei . 2018 . Exploring Visual Relationship for Image Captioning. In Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, September 8--14, 2018 , Proceedings, Part XIV (Lecture Notes in Computer Science) , Vol. 11218 . Springer, 711--727. Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei. 2018. Exploring Visual Relationship for Image Captioning. In Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, September 8--14, 2018, Proceedings, Part XIV (Lecture Notes in Computer Science), Vol. 11218. Springer, 711--727."},{"key":"e_1_3_2_2_48_1","volume-title":"Boosting Image Captioning with Attributes. In IEEE International Conference on Computer Vision, ICCV 2017","author":"Yao Ting","year":"2017","unstructured":"Ting Yao , Yingwei Pan , Yehao Li , Zhaofan Qiu , and Tao Mei . 2017 . Boosting Image Captioning with Attributes. In IEEE International Conference on Computer Vision, ICCV 2017 , Venice, Italy, October 22--29 , 2017. IEEE Computer Society, 4904--4912. Ting Yao, Yingwei Pan, Yehao Li, Zhaofan Qiu, and Tao Mei. 2017. Boosting Image Captioning with Attributes. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22--29, 2017. IEEE Computer Society, 4904--4912."},{"key":"e_1_3_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00878"},{"volume-title":"A brief introduction to weakly supervised learning. National Science Review","author":"Zhou Zhi Hua","key":"e_1_3_2_2_50_1","unstructured":"Zhi Hua Zhou . [n.d.]. A brief introduction to weakly supervised learning. National Science Review , Vol. v. 5 , 1 ([n.,d.]), 48--57. Zhi Hua Zhou. [n.d.]. A brief introduction to weakly supervised learning. National Science Review, Vol. v.5, 1 ([n.,d.]), 48--57."}],"event":{"name":"MM '20: The 28th ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Seattle WA USA","acronym":"MM '20"},"container-title":["Proceedings of the 28th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3413885","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3394171.3413885","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:32:06Z","timestamp":1750195926000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3413885"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,12]]},"references-count":48,"alternative-id":["10.1145\/3394171.3413885","10.1145\/3394171"],"URL":"https:\/\/doi.org\/10.1145\/3394171.3413885","relation":{},"subject":[],"published":{"date-parts":[[2020,10,12]]},"assertion":[{"value":"2020-10-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}