{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,12]],"date-time":"2026-05-12T17:17:24Z","timestamp":1778606244490,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":35,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,10,17]],"date-time":"2021-10-17T00:00:00Z","timestamp":1634428800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"supported by Medical Research Council (MRC) Fellowship","award":["MR\/S003916\/2"],"award-info":[{"award-number":["MR\/S003916\/2"]}]},{"name":"Engineering and Physical Sciences Research Council (EPSRC) Project CRITiCaL: Combatting cRiminals In The CLoud","award":["EP\/M020576\/1"],"award-info":[{"award-number":["EP\/M020576\/1"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,10,17]]},"DOI":"10.1145\/3474085.3475519","type":"proceedings-article","created":{"date-parts":[[2021,10,18]],"date-time":"2021-10-18T06:21:10Z","timestamp":1634538070000},"page":"3556-3564","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":26,"title":["Discriminative Latent Semantic Graph for Video Captioning"],"prefix":"10.1145","author":[{"given":"Yang","family":"Bai","sequence":"first","affiliation":[{"name":"Newcastle University, Newcastle upon Tyne, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Junyan","family":"Wang","sequence":"additional","affiliation":[{"name":"University of New South Wales, Sydney, NSW, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yang","family":"Long","sequence":"additional","affiliation":[{"name":"Durham University, Durham, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bingzhang","family":"Hu","sequence":"additional","affiliation":[{"name":"Hefei CAS Dihuge Automation Co., LTD, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yang","family":"Song","sequence":"additional","affiliation":[{"name":"University of New South Wales, Sydney, NSW, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Maurice","family":"Pagnucco","sequence":"additional","affiliation":[{"name":"University of New South Wales, Sydney, NSW, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yu","family":"Guan","sequence":"additional","affiliation":[{"name":"Newcastle University, Newcastle upon Tyne, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,10,17]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00636"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5555\/2002472.2002497"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"crossref","unstructured":"Yangyu Chen Shuhui Wang Weigang Zhang and Qingming Huang. 2018. Less is more: Picking informative frames for video captioning. In ECCV. 358--373.  Yangyu Chen Shuhui Wang Weigang Zhang and Qingming Huang. 2018. Less is more: Picking informative frames for video captioning. In ECCV. 358--373.","DOI":"10.1007\/978-3-030-01261-8_22"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"crossref","unstructured":"Bo Dai Sanja Fidler Raquel Urtasun and Dahua Lin. 2017. Towards diverse and natural image descriptions via a conditional gan. In ICCV. 2970--2979.  Bo Dai Sanja Fidler Raquel Urtasun and Dahua Lin. 2017. Towards diverse and natural image descriptions via a conditional gan. In ICCV. 2970--2979.","DOI":"10.1109\/ICCV.2017.323"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/W14-3348"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.337"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295327"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"crossref","unstructured":"Jingyi Hou Xinxiao Wu Xiaoxun Zhang Yayun Qi Yunde Jia and Jiebo Luo. 2020. Joint Commonsense and Relation Reasoning for Image and Video Captioning.. In AAAI. 10973--10980.  Jingyi Hou Xinxiao Wu Xiaoxun Zhang Yayun Qi Yunde Jia and Jiebo Luo. 2020. Joint Commonsense and Relation Reasoning for Image and Video Captioning.. In AAAI. 10973--10980.","DOI":"10.1609\/aaai.v34i07.6731"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3351072"},{"key":"e_1_3_2_1_11_1","volume-title":"Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144","author":"Jang Eric","year":"2016","unstructured":"Eric Jang , Shixiang Gu , and Ben Poole . 2016. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144 ( 2016 ). Eric Jang, Shixiang Gu, and Ben Poole. 2016. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144 (2016)."},{"key":"e_1_3_2_1_12_1","volume-title":"Hadamard product for low-rank bilinear pooling. arXiv preprint arXiv:1610.04325","author":"Kim Jin-Hwa","year":"2016","unstructured":"Jin-Hwa Kim , Kyoung-Woon On , Woosang Lim , Jeonghee Kim , Jung-Woo Ha , and Byoung-Tak Zhang . 2016. Hadamard product for low-rank bilinear pooling. arXiv preprint arXiv:1610.04325 ( 2016 ). Jin-Hwa Kim, Kyoung-Woon On, Woosang Lim, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang. 2016. Hadamard product for low-rank bilinear pooling. arXiv preprint arXiv:1610.04325 (2016)."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00099"},{"key":"e_1_3_2_1_14_1","volume-title":"Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74--81.","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin . 2004 . Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74--81. Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74--81."},{"key":"e_1_3_2_1_15_1","unstructured":"Boxiao Pan Haoye Cai De-An Huang Kuan-Hui Lee Adrien Gaidon Ehsan Adeli and Juan Carlos Niebles. 2020. Spatio-Temporal Graph for Video Captioning with Knowledge Distillation. In CPVR. 10870--10879.  Boxiao Pan Haoye Cai De-An Huang Kuan-Hui Lee Adrien Gaidon Ehsan Adeli and Juan Carlos Niebles. 2020. Spatio-Temporal Graph for Video Captioning with Knowledge Distillation. In CPVR. 10870--10879."},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"crossref","unstructured":"Jae Sung Park Marcus Rohrbach Trevor Darrell and Anna Rohrbach. 2019. Adversarial inference for multi-sentence video description. In CVPR. 6598--6608.  Jae Sung Park Marcus Rohrbach Trevor Darrell and Anna Rohrbach. 2019. Adversarial inference for multi-sentence video description. In CVPR. 6598--6608.","DOI":"10.1109\/CVPR.2019.00676"},{"key":"e_1_3_2_1_18_1","unstructured":"Wenjie Pei Jiyuan Zhang Xiangrong Wang Lei Ke Xiaoyong Shen and Yu-Wing Tai. 2019. Memory-attended recurrent network for video captioning. In CVPR. 8347--8356.  Wenjie Pei Jiyuan Zhang Xiangrong Wang Lei Ke Xiaoyong Shen and Yu-Wing Tai. 2019. Memory-attended recurrent network for video captioning. In CVPR. 8347--8356."},{"key":"e_1_3_2_1_19_1","volume-title":"Faster r-cnn: Towards real-time object detection with region proposal networks. arXiv preprint arXiv:1506.01497","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren , Kaiming He , Ross Girshick , and Jian Sun . 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. arXiv preprint arXiv:1506.01497 ( 2015 ). Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. arXiv preprint arXiv:1506.01497 (2015)."},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-11752-2_15"},{"key":"e_1_3_2_1_21_1","volume-title":"Mario Fritz, and Bernt Schiele.","author":"Shetty Rakshith","year":"2017","unstructured":"Rakshith Shetty , Marcus Rohrbach , Lisa Anne Hendricks , Mario Fritz, and Bernt Schiele. 2017 . Speaking the same language: Matching machine to human captions by adversarial training. In ICCV. 4135--4144. Rakshith Shetty, Marcus Rohrbach, Lisa Anne Hendricks, Mario Fritz, and Bernt Schiele. 2017. Speaking the same language: Matching machine to human captions by adversarial training. In ICCV. 4135--4144."},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.5555\/3298023.3298188"},{"key":"e_1_3_2_1_23_1","unstructured":"Ganchao Tan Daqing Liu Wang Meng and Zheng-Jun Zha. 2020. Learning to Discretely Compose Reasoning Module Networks for Video Captioning. In IJCAI-PRICAI.  Ganchao Tan Daqing Liu Wang Meng and Zheng-Jun Zha. 2020. Learning to Discretely Compose Reasoning Module Networks for Video Captioning. In IJCAI-PRICAI."},{"key":"e_1_3_2_1_24_1","volume-title":"Cider: Consensus-based image description evaluation. In CPVR. 4566--4575.","author":"Vedantam Ramakrishna","year":"2015","unstructured":"Ramakrishna Vedantam , C Lawrence Zitnick , and Devi Parikh . 2015 . Cider: Consensus-based image description evaluation. In CPVR. 4566--4575. Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015. Cider: Consensus-based image description evaluation. In CPVR. 4566--4575."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.515"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"crossref","unstructured":"Bairui Wang Lin Ma Wei Zhang and Wei Liu. 2018. Reconstruction network for video captioning. In CVPR. 7622--7631.  Bairui Wang Lin Ma Wei Zhang and Wei Liu. 2018. Reconstruction network for video captioning. In CVPR. 7622--7631.","DOI":"10.1109\/CVPR.2018.00795"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3414064"},{"key":"e_1_3_2_1_28_1","volume-title":"Msr-vtt: A large video description dataset for bridging video and language. In CVPR. 5288--5296.","author":"Xu Jun","year":"2016","unstructured":"Jun Xu , Tao Mei , Ting Yao , and Yong Rui . 2016 . Msr-vtt: A large video description dataset for bridging video and language. In CVPR. 5288--5296. Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016. Msr-vtt: A large video description dataset for bridging video and language. In CVPR. 5288--5296."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2855422"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.512"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58517-4_31"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"crossref","unstructured":"Junchao Zhang and Yuxin Peng. 2019. Object-aware aggregation with bidirectional temporal graph for video captioning. In CPVR. 8327--8336.  Junchao Zhang and Yuxin Peng. 2019. Object-aware aggregation with bidirectional temporal graph for video captioning. In CPVR. 8327--8336.","DOI":"10.1109\/CVPR.2019.00852"},{"key":"e_1_3_2_1_33_1","volume-title":"Latentgnn: Learning efficient non-local relations for visual recognition. In ICML. PMLR, 7374--7383.","author":"Zhang Songyang","year":"2019","unstructured":"Songyang Zhang , Xuming He , and Shipeng Yan . 2019 . Latentgnn: Learning efficient non-local relations for visual recognition. In ICML. PMLR, 7374--7383. Songyang Zhang, Xuming He, and Shipeng Yan. 2019. Latentgnn: Learning efficient non-local relations for visual recognition. In ICML. PMLR, 7374--7383."},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"crossref","unstructured":"Ziqi Zhang Yaya Shi Chunfeng Yuan Bing Li Peijin Wang Weiming Hu and Zheng-Jun Zha. 2020. Object Relational Graph with Teacher-Recommended Learning for Video Captioning. In CVPR. 13278--13288.  Ziqi Zhang Yaya Shi Chunfeng Yuan Bing Li Peijin Wang Weiming Hu and Zheng-Jun Zha. 2020. Object Relational Graph with Teacher-Recommended Learning for Video Captioning. In CVPR. 13278--13288.","DOI":"10.1109\/CVPR42600.2020.01329"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00983"}],"event":{"name":"MM '21: ACM Multimedia Conference","location":"Virtual Event China","acronym":"MM '21","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 29th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474085.3475519","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3474085.3475519","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:49:10Z","timestamp":1750193350000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474085.3475519"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,17]]},"references-count":35,"alternative-id":["10.1145\/3474085.3475519","10.1145\/3474085"],"URL":"https:\/\/doi.org\/10.1145\/3474085.3475519","relation":{},"subject":[],"published":{"date-parts":[[2021,10,17]]},"assertion":[{"value":"2021-10-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}