{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,2]],"date-time":"2026-01-02T07:10:13Z","timestamp":1767337813449,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":41,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,10,17]],"date-time":"2021-10-17T00:00:00Z","timestamp":1634428800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Research Grants Council of the Hong Kong Special Administrative Region, China","award":["Project No. CityU 11215820"],"award-info":[{"award-number":["Project No. CityU 11215820"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,10,17]]},"DOI":"10.1145\/3474085.3475215","type":"proceedings-article","created":{"date-parts":[[2021,10,18]],"date-time":"2021-10-18T05:40:18Z","timestamp":1634535618000},"page":"5020-5028","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":17,"title":["Group-based Distinctive Image Captioning with Memory Attention"],"prefix":"10.1145","author":[{"given":"Jiuniu","family":"Wang","sequence":"first","affiliation":[{"name":"City University of Hong Kong, Aerospace Information Research Institute, Chinese Academy of Sciences, &amp; University of Chinese Academy of Sciences, Hong Kong SAR, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenjia","family":"Xu","sequence":"additional","affiliation":[{"name":"Aerospace Information Research Institute, Chinese Academy of Sciences &amp; University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qingzhong","family":"Wang","sequence":"additional","affiliation":[{"name":"City University of Hong Kong &amp; Baidu Research, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Antoni B.","family":"Chan","sequence":"additional","affiliation":[{"name":"City University of Hong Kong, Hong Kong SAR, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,10,17]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"crossref","unstructured":"Peter Anderson Xiaodong He Chris Buehler Damien Teney Mark Johnson Stephen Gould and Lei Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In CVPR.  Peter Anderson Xiaodong He Chris Buehler Damien Teney Mark Johnson Stephen Gould and Lei Zhang. 2018. Bottom-up and top-down attention for image captioning and visual question answering. In CVPR.","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"crossref","unstructured":"Jyoti Aneja Aditya Deshpande and Alexander G Schwing. 2018. Convolutional image captioning. In CVPR.  Jyoti Aneja Aditya Deshpande and Alexander G Schwing. 2018. Convolutional image captioning. In CVPR.","DOI":"10.1109\/CVPR.2018.00583"},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123275"},{"key":"e_1_3_2_2_4_1","volume-title":"Groupcap: Group-based image captioning with structured relevance and diversity constraints. In CVPR.","author":"Chen Fuhai","year":"2018","unstructured":"Fuhai Chen , Rongrong Ji , Xiaoshuai Sun , Yongjian Wu , and Jinsong Su . 2018 . Groupcap: Group-based image captioning with structured relevance and diversity constraints. In CVPR. Fuhai Chen, Rongrong Ji, Xiaoshuai Sun, Yongjian Wu, and Jinsong Su. 2018. Groupcap: Group-based image captioning with structured relevance and diversity constraints. In CVPR."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"crossref","unstructured":"Long Chen Hanwang Zhang Jun Xiao Liqiang Nie Jian Shao Wei Liu and Tat-Seng Chua. 2017b. SCA-CNN: Spatial and channel-wise attention in convolutional networks for image captioning. In CVPR.  Long Chen Hanwang Zhang Jun Xiao Liqiang Nie Jian Shao Wei Liu and Tat-Seng Chua. 2017b. SCA-CNN: Spatial and channel-wise attention in convolutional networks for image captioning. In CVPR.","DOI":"10.1109\/CVPR.2017.667"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"crossref","unstructured":"Marcella Cornia Matteo Stefanini Lorenzo Baraldi and Rita Cucchiara. 2020. Meshed-memory transformer for image captioning. In CVPR.  Marcella Cornia Matteo Stefanini Lorenzo Baraldi and Rita Cucchiara. 2020. Meshed-memory transformer for image captioning. In CVPR.","DOI":"10.1109\/CVPR42600.2020.01059"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"crossref","unstructured":"Bo Dai Sanja Fidler Raquel Urtasun and Dahua Lin. 2017. Towards diverse and natural image descriptions via a conditional GAN. In ICCV.  Bo Dai Sanja Fidler Raquel Urtasun and Dahua Lin. 2017. Towards diverse and natural image descriptions via a conditional GAN. In ICCV.","DOI":"10.1109\/ICCV.2017.323"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.5555\/3294771.3294857"},{"key":"e_1_3_2_2_9_1","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly etal 2020. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR.  Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR."},{"key":"e_1_3_2_2_10_1","volume-title":"Jamie Ryan Kiros, and Sanja Fidler","author":"Faghri Fartash","year":"2018","unstructured":"Fartash Faghri , David J Fleet , Jamie Ryan Kiros, and Sanja Fidler . 2018 . VSE+: Improving visual-semantic embeddings with hard negatives. In BMVC. Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2018. VSE+: Improving visual-semantic embeddings with hard negatives. In BMVC."},{"key":"e_1_3_2_2_11_1","unstructured":"Longteng Guo Jing Liu Xinxin Zhu Peng Yao Shichen Lu and Hanqing Lu. 2020. Normalized and geometry-aware self-attention network for image captioning. In CVPR.  Longteng Guo Jing Liu Xinxin Zhu Peng Yao Shichen Lu and Hanqing Lu. 2020. Normalized and geometry-aware self-attention network for image captioning. In CVPR."},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"crossref","unstructured":"Lun Huang Wenmin Wang Jie Chen and Xiao-Yong Wei. 2019. Attention on attention for image captioning. In ICCV.  Lun Huang Wenmin Wang Jie Chen and Xiao-Yong Wei. 2019. Attention on attention for image captioning. In ICCV.","DOI":"10.1109\/ICCV.2019.00473"},{"key":"e_1_3_2_2_13_1","volume-title":"Creativity: Generating diverse questions using variational autoencoders. In CVPR.","author":"Jain Unnat","year":"2017","unstructured":"Unnat Jain , Ziyu Zhang , and Alexander G Schwing . 2017 . Creativity: Generating diverse questions using variational autoencoders. In CVPR. Unnat Jain, Ziyu Zhang, and Alexander G Schwing. 2017. Creativity: Generating diverse questions using variational autoencoders. In CVPR."},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"crossref","unstructured":"Andrej Karpathy and Li Fei-Fei. 2015. Deep visual-semantic alignments for generating image descriptions. In CVPR.  Andrej Karpathy and Li Fei-Fei. 2015. Deep visual-semantic alignments for generating image descriptions. In CVPR.","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"e_1_3_2_2_15_1","unstructured":"Guang Li Linchao Zhu Ping Liu and Yi Yang. 2019. Entangled transformer for image captioning. In ICCV.  Guang Li Linchao Zhu Ping Liu and Yi Yang. 2019. Entangled transformer for image captioning. In ICCV."},{"key":"e_1_3_2_2_16_1","unstructured":"Zhuowan Li Quan Tran Long Mai Zhe Lin and Alan L Yuille. 2020. Context-aware group captioning via self-attention and contrastive features. In CVPR.  Zhuowan Li Quan Tran Long Mai Zhe Lin and Alan L Yuille. 2020. Context-aware group captioning via self-attention and contrastive features. In CVPR."},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"crossref","unstructured":"Xihui Liu Hongsheng Li Jing Shao Dapeng Chen and Xiaogang Wang. 2018. Show tell and discriminate: Image captioning by self-retrieval with partially labeled data. In ECCV.  Xihui Liu Hongsheng Li Jing Shao Dapeng Chen and Xiaogang Wang. 2018. Show tell and discriminate: Image captioning by self-retrieval with partially labeled data. In ECCV.","DOI":"10.1007\/978-3-030-01267-0_21"},{"key":"e_1_3_2_2_18_1","unstructured":"Ruotian Luo Brian Price Scott Cohen and Gregory Shakhnarovich. 2018. Discriminability objective for training descriptive captions. In CVPR.  Ruotian Luo Brian Price Scott Cohen and Gregory Shakhnarovich. 2018. Discriminability objective for training descriptive captions. In CVPR."},{"key":"e_1_3_2_2_19_1","volume-title":"ICCV Workshop.","author":"Luo Ruotian","year":"2019","unstructured":"Ruotian Luo and Gregory Shakhnarovich . 2019 . Analysis of diversity-accuracy tradeoff in image captioning . In ICCV Workshop. Ruotian Luo and Gregory Shakhnarovich. 2019. Analysis of diversity-accuracy tradeoff in image captioning. In ICCV Workshop."},{"key":"e_1_3_2_2_20_1","unstructured":"Junhua Mao Wei Xu Yi Yang Jiang Wang Zhiheng Huang and Alan Yuille. 2015. Deep captioning with multimodal recurrent neural networks (m-RNN). In ICLR.  Junhua Mao Wei Xu Yi Yang Jiang Wang Zhiheng Huang and Alan Yuille. 2015. Deep captioning with multimodal recurrent neural networks (m-RNN). In ICLR."},{"key":"e_1_3_2_2_21_1","unstructured":"Yingwei Pan Ting Yao Yehao Li and Tao Mei. 2020. X-linear attention networks for image captioning. In CVPR.  Yingwei Pan Ting Yao Yehao Li and Tao Mei. 2020. X-linear attention networks for image captioning. In CVPR."},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3454294"},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969239.2969250"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"crossref","unstructured":"Steven J Rennie Etienne Marcheret Youssef Mroueh Jerret Ross and Vaibhava Goel. 2017. Self-critical sequence training for image captioning. In CVPR.  Steven J Rennie Etienne Marcheret Youssef Mroueh Jerret Ross and Vaibhava Goel. 2017. Self-critical sequence training for image captioning. In CVPR.","DOI":"10.1109\/CVPR.2017.131"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"crossref","unstructured":"Rakshith Shetty Marcus Rohrbach and Lisa Anne Hendricks. 2017. Speaking the same language: Matching machine to human captions by adversarial Training. In ICCV.  Rakshith Shetty Marcus Rohrbach and Lisa Anne Hendricks. 2017. Speaking the same language: Matching machine to human captions by adversarial Training. In ICCV.","DOI":"10.1109\/ICCV.2017.445"},{"key":"e_1_3_2_2_26_1","unstructured":"Weijie Su Xizhou Zhu Yue Cao Bin Li Lewei Lu Furu Wei and Jifeng Dai. 2020. VL-BERT: Pre-training of generic visual-linguistic representations. In ICLR.  Weijie Su Xizhou Zhu Yue Cao Bin Li Lewei Lu Furu Wei and Jifeng Dai. 2020. VL-BERT: Pre-training of generic visual-linguistic representations. In ICLR."},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"crossref","unstructured":"Ramakrishna Vedantam Samy Bengio Kevin Murphy Devi Parikh and Gal Chechik. 2017. Context-aware captions from context-agnostic supervision. In CVPR.  Ramakrishna Vedantam Samy Bengio Kevin Murphy Devi Parikh and Gal Chechik. 2017. Context-aware captions from context-agnostic supervision. In CVPR.","DOI":"10.1109\/CVPR.2017.120"},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"crossref","unstructured":"Gilad Vered Gal Oren Yuval Atzmon and Gal Chechik. 2019. Joint optimization for cooperative image captioning. In CVPR.  Gilad Vered Gal Oren Yuval Atzmon and Gal Chechik. 2019. Joint optimization for cooperative image captioning. In CVPR.","DOI":"10.1109\/ICCV.2019.00899"},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"crossref","unstructured":"Oriol Vinyals Alexander Toshev Samy Bengio and Dumitru Erhan. 2015. Show and tell: A neural image caption generator. In CVPR.  Oriol Vinyals Alexander Toshev Samy Bengio and Dumitru Erhan. 2015. Show and tell: A neural image caption generator. In CVPR.","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"crossref","unstructured":"Jiuniu Wang Wenjia Xu Qingzhong Wang and Antoni B Chan. 2020 b. Compare and reweight: Distinctive image captioning using similar images sets. In ECCV.  Jiuniu Wang Wenjia Xu Qingzhong Wang and Antoni B Chan. 2020 b. Compare and reweight: Distinctive image captioning using similar images sets. In ECCV.","DOI":"10.1007\/978-3-030-58452-8_22"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295326"},{"key":"e_1_3_2_2_33_1","volume-title":"CVPR Workshop.","author":"Wang Qingzhong","year":"2018","unstructured":"Qingzhong Wang and Antoni B Chan . 2018 a. CNN+CNN: Convolutional decoders for image captioning . In CVPR Workshop. Qingzhong Wang and Antoni B Chan. 2018a. CNN+CNN: Convolutional decoders for image captioning. In CVPR Workshop."},{"key":"e_1_3_2_2_34_1","unstructured":"Qingzhong Wang and Antoni B Chan. 2018b. Gated hierarchical attention for image captioning. In ACCV.  Qingzhong Wang and Antoni B Chan. 2018b. Gated hierarchical attention for image captioning. In ACCV."},{"key":"e_1_3_2_2_35_1","unstructured":"Qingzhong Wang and Antoni B Chan. 2019. Describing like humans: on diversity in image captioning. In CVPR.  Qingzhong Wang and Antoni B Chan. 2019. Describing like humans: on diversity in image captioning. In CVPR."},{"key":"e_1_3_2_2_36_1","unstructured":"Qingzhong Wang and Antoni B Chan. 2020. Towards diverse and accurate image captions via reinforcing determinantal point process. In TPAMI.  Qingzhong Wang and Antoni B Chan. 2020. Towards diverse and accurate image captions via reinforcing determinantal point process. In TPAMI."},{"volume-title":"2020 a. On diversity in image captioning: Metrics and methods","author":"Wang Qingzhong","key":"e_1_3_2_2_37_1","unstructured":"Qingzhong Wang , Jia Wan , and Antoni B Chan . 2020 a. On diversity in image captioning: Metrics and methods . In IEEE TPAMI. Qingzhong Wang, Jia Wan, and Antoni B Chan. 2020 a. On diversity in image captioning: Metrics and methods. In IEEE TPAMI."},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.5555\/3045118.3045336"},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"crossref","unstructured":"Zekun Yang Noa Garcia Chenhui Chu Mayu Otani Yuta Nakashima and Haruo Takemura. 2020. Bert representations for video question answering. In WACV.  Zekun Yang Noa Garcia Chenhui Chu Mayu Otani Yuta Nakashima and Haruo Takemura. 2020. Bert representations for video question answering. In WACV.","DOI":"10.1109\/WACV45572.2020.9093596"},{"key":"e_1_3_2_2_40_1","unstructured":"Linwei Ye Mrigank Rochan Zhi Liu and Yang Wang. 2019. Cross-modal self-attention network for referring image segmentation. In CVPR.  Linwei Ye Mrigank Rochan Zhi Liu and Yang Wang. 2019. Cross-modal self-attention network for referring image segmentation. In CVPR."},{"key":"e_1_3_2_2_41_1","unstructured":"Quanzeng You Hailin Jin Zhaowen Wang Chen Fang and Jiebo Luo. 2016. Image captioning with semantic attention. In CVPR.  Quanzeng You Hailin Jin Zhaowen Wang Chen Fang and Jiebo Luo. 2016. Image captioning with semantic attention. In CVPR."}],"event":{"name":"MM '21: ACM Multimedia Conference","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Virtual Event China","acronym":"MM '21"},"container-title":["Proceedings of the 29th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474085.3475215","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3474085.3475215","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:48:16Z","timestamp":1750193296000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474085.3475215"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,17]]},"references-count":41,"alternative-id":["10.1145\/3474085.3475215","10.1145\/3474085"],"URL":"https:\/\/doi.org\/10.1145\/3474085.3475215","relation":{},"subject":[],"published":{"date-parts":[[2021,10,17]]},"assertion":[{"value":"2021-10-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}