{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,4]],"date-time":"2025-10-04T08:03:11Z","timestamp":1759564991590,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":77,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,10,17]],"date-time":"2021-10-17T00:00:00Z","timestamp":1634428800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"NSFC","award":["61921003"],"award-info":[{"award-number":["61921003"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,10,17]]},"DOI":"10.1145\/3474085.3475236","type":"proceedings-article","created":{"date-parts":[[2021,10,18]],"date-time":"2021-10-18T21:45:34Z","timestamp":1634593534000},"page":"4892-4901","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":13,"title":["Latent Memory-augmented Graph Transformer for Visual Storytelling"],"prefix":"10.1145","author":[{"given":"Mengshi","family":"Qi","sequence":"first","affiliation":[{"name":"Beijing University of Posts and Telecommunications, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jie","family":"Qin","sequence":"additional","affiliation":[{"name":"Nanjing University of Aeronautics and Astronautics, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Di","family":"Huang","sequence":"additional","affiliation":[{"name":"Beihang University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhiqiang","family":"Shen","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yi","family":"Yang","sequence":"additional","affiliation":[{"name":"University of Technology Sydney, Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiebo","family":"Luo","sequence":"additional","affiliation":[{"name":"University of Rochester, Rochester, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,10,17]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3416305"},{"volume-title":"Spice: Semantic propositional image caption evaluation","year":"2016","author":"Anderson Peter","key":"e_1_3_2_2_2_1"},{"volume-title":"Jamie Ryan Kiros, and Geoffrey E Hinton","year":"2016","author":"Ba Jimmy Lei","key":"e_1_3_2_2_3_1"},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3351037"},{"volume-title":"Microsoft coco captions: Data collection and evaluation server. arXiv preprint arXiv:1504.00325","year":"2015","author":"Chen Xinlei","key":"e_1_3_2_2_5_1"},{"volume-title":"Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu.","year":"2019","author":"Chen Yen-Chun","key":"e_1_3_2_2_6_1"},{"volume-title":"Dzmitry Bahdanau, and Yoshua Bengio.","year":"2014","author":"Cho Kyunghyun","key":"e_1_3_2_2_7_1"},{"volume-title":"Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555","year":"2014","author":"Chung Junyoung","key":"e_1_3_2_2_8_1"},{"volume-title":"Meshed-Memory Transformer for Image Captioning","author":"Cornia Marcella","key":"e_1_3_2_2_9_1"},{"volume-title":"Transformer-xl: Attentive language models beyond a fixed-length context. In ACL .","year":"2019","author":"Dai Zihang","key":"e_1_3_2_2_10_1"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/W14-3348"},{"volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","year":"2018","author":"Devlin Jacob","key":"e_1_3_2_2_12_1"},{"key":"e_1_3_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2017.2729019"},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350943"},{"volume-title":"Deep residual learning for image recognition","author":"He Kaiming","key":"e_1_3_2_2_15_1"},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3455286"},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"crossref","unstructured":"Xudong Hong Rakshith Shetty Asad Sayeed Khushboo Mehra Vera Demberg and Bernt Schiele. 2020. Diverse and Relevant Visual Storytelling with Scene Graph Embeddings. In CNLL .  Xudong Hong Rakshith Shetty Asad Sayeed Khushboo Mehra Vera Demberg and Bernt Schiele. 2020. Diverse and Relevant Visual Storytelling with Scene Graph Embeddings. In CNLL .","DOI":"10.18653\/v1\/2020.conll-1.34"},{"key":"e_1_3_2_2_19_1","unstructured":"Chao-Chun Hsu Zi-Yuan Chen Chi-Yang Hsu Chih-Chia Li Tzu-Yuan Lin Ting-Hao'Kenneth' Huang and Lun-Wei Ku. 2020. Knowledge-Enriched Visual Storytelling. In AAAI .  Chao-Chun Hsu Zi-Yuan Chen Chi-Yang Hsu Chih-Chia Li Tzu-Yuan Lin Ting-Hao'Kenneth' Huang and Lun-Wei Ku. 2020. Knowledge-Enriched Visual Storytelling. In AAAI ."},{"key":"e_1_3_2_2_20_1","unstructured":"Junjie Hu Yu Cheng Zhe Gan Jingjing Liu Jianfeng Gao and Graham Neubig. 2020. What Makes A Good Story? Designing Composite Rewards for Visual Storytelling.. In AAAI .  Junjie Hu Yu Cheng Zhe Gan Jingjing Liu Jianfeng Gao and Graham Neubig. 2020. What Makes A Good Story? Designing Composite Rewards for Visual Storytelling.. In AAAI ."},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3351072"},{"volume-title":"2019 b. Attention on attention for image captioning","author":"Huang Lun","key":"e_1_3_2_2_22_1"},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"crossref","unstructured":"Qiuyuan Huang Zhe Gan Asli Celikyilmaz Dapeng Wu Jianfeng Wang and Xiaodong He. 2019 a. Hierarchically structured reinforcement learning for topically coherent visual story generation. In AAAI .  Qiuyuan Huang Zhe Gan Asli Celikyilmaz Dapeng Wu Jianfeng Wang and Xiaodong He. 2019 a. Hierarchically structured reinforcement learning for topically coherent visual story generation. In AAAI .","DOI":"10.1609\/aaai.v33i01.33018465"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N16-1147"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3414009"},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"crossref","unstructured":"Yunjae Jung Dahun Kim Sanghyun Woo Kyungsu Kim Sungjin Kim and In So Kweon. 2020. Hide-and-Tell: Learning to Bridge Photo Streams for Visual Storytelling. In AAAI .  Yunjae Jung Dahun Kim Sanghyun Woo Kyungsu Kim Sungjin Kim and In So Kweon. 2020. Hide-and-Tell: Learning to Bridge Photo Streams for Visual Storytelling. In AAAI .","DOI":"10.1609\/aaai.v34i07.6780"},{"key":"e_1_3_2_2_27_1","unstructured":"Taehyeong Kim Min-Oh Heo Seonil Son Kyoung-Wha Park and Byoung-Tak Zhang. 2018. Glac net: Glocal attention cascading networks for multi-image cued story generation. In ACL .  Taehyeong Kim Min-Oh Heo Seonil Son Kyoung-Wha Park and Byoung-Tak Zhang. 2018. Glac net: Glocal attention cascading networks for multi-image cued story generation. In ACL ."},{"volume-title":"A hierarchical approach for generating descriptive image paragraphs","author":"Krause Jonathan","key":"e_1_3_2_2_28_1"},{"volume-title":"Dense-captioning events in videos","author":"Krishna Ranjay","key":"e_1_3_2_2_29_1","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2017.83"},{"volume-title":"Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations. https:\/\/arxiv.org\/abs\/1602.07332","year":"2016","author":"Krishna Ranjay","key":"e_1_3_2_2_30_1"},{"volume-title":"Mart: Memory-augmented recurrent transformer for coherent video paragraph captioning. In ACL .","year":"2020","author":"Lei Jie","key":"e_1_3_2_2_31_1"},{"volume-title":"2019 b. Entangled transformer for image captioning","author":"Li Guang","key":"e_1_3_2_2_32_1"},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350918"},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413886"},{"volume-title":"Factorizable net: an efficient subgraph-based framework for scene graph generation","author":"Li Yikang","key":"e_1_3_2_2_35_1"},{"volume-title":"Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74--81.","year":"2004","author":"Lin Chin-Yew","key":"e_1_3_2_2_36_1"},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413985"},{"volume-title":"Fine-tune BERT for extractive summarization. arXiv preprint arXiv:1903.10318","year":"2019","author":"Liu Yang","key":"e_1_3_2_2_38_1"},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.5555\/3298239.3298450"},{"volume-title":"Video captioning with transferred semantic attributes","author":"Pan Yingwei","key":"e_1_3_2_2_41_1"},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073083.1073135"},{"key":"e_1_3_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969239.2969248"},{"volume-title":"Glove: Global vectors for word representation. In EMNLP .","year":"2014","author":"Pennington Jeffrey","key":"e_1_3_2_2_44_1"},{"volume-title":"2019 a. Attentive relational networks for mapping images to scene graphs","author":"Qi Mengshi","key":"e_1_3_2_2_45_1"},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3048680"},{"key":"e_1_3_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3416269"},{"key":"e_1_3_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123311"},{"key":"e_1_3_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3265845.3265851"},{"key":"e_1_3_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2019.2921655"},{"key":"e_1_3_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2019.2894161"},{"key":"e_1_3_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969239.2969250"},{"key":"e_1_3_2_2_53_1","doi-asserted-by":"crossref","unstructured":"Piyush Sharma Nan Ding Sebastian Goodman and Radu Soricut. 2018. Conceptual captions: A cleaned hypernymed image alt-text dataset for automatic image captioning. In ACL .  Piyush Sharma Nan Ding Sebastian Goodman and Radu Soricut. 2018. Conceptual captions: A cleaned hypernymed image alt-text dataset for automatic image captioning. In ACL .","DOI":"10.18653\/v1\/P18-1238"},{"volume-title":"Weakly supervised dense video captioning","author":"Shen Zhiqiang","key":"e_1_3_2_2_54_1"},{"key":"e_1_3_2_2_55_1","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In ICLR .  Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. In ICLR ."},{"key":"e_1_3_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350996"},{"volume-title":"Carl Vondrick, Kevin Murphy, and Cordelia Schmid.","year":"2019","author":"Sun Chen","key":"e_1_3_2_2_57_1"},{"volume-title":"Lxmert: Learning cross-modality encoder representations from transformers. In EMNLP .","year":"2019","author":"Tan Hao","key":"e_1_3_2_2_58_1"},{"key":"e_1_3_2_2_59_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"volume-title":"Cider: Consensus-based image description evaluation","year":"2015","author":"Vedantam Ramakrishna","key":"e_1_3_2_2_60_1"},{"key":"e_1_3_2_2_61_1","unstructured":"Petar Velivc kovi\u0107 Guillem Cucurull Arantxa Casanova Adriana Romero Pietro Lio and Yoshua Bengio. 2018. Graph attention networks. In ICLR .  Petar Velivc kovi\u0107 Guillem Cucurull Arantxa Casanova Adriana Romero Pietro Lio and Yoshua Bengio. 2018. Graph attention networks. In ICLR ."},{"key":"e_1_3_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3416528"},{"volume-title":"Show and tell: A neural image caption generator","author":"Vinyals Oriol","key":"e_1_3_2_2_63_1"},{"key":"e_1_3_2_2_64_1","doi-asserted-by":"crossref","unstructured":"Bairui Wang Lin Ma Wei Zhang Wenhao Jiang and Feng Zhang. 2019. Hierarchical photo-scene encoder for album storytelling. In AAAI .  Bairui Wang Lin Ma Wei Zhang Wenhao Jiang and Feng Zhang. 2019. Hierarchical photo-scene encoder for album storytelling. In AAAI .","DOI":"10.1609\/aaai.v33i01.33018909"},{"key":"e_1_3_2_2_65_1","doi-asserted-by":"crossref","unstructured":"Jing Wang Jianlong Fu Jinhui Tang Zechao Li and Tao Mei. 2018b. Show reward and tell: Automatic generation of narrative paragraph from photo stream by adversarial training. In AAAI .  Jing Wang Jianlong Fu Jinhui Tang Zechao Li and Tao Mei. 2018b. Show reward and tell: Automatic generation of narrative paragraph from photo stream by adversarial training. In AAAI .","DOI":"10.1609\/aaai.v32i1.12318"},{"key":"e_1_3_2_2_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413753"},{"key":"e_1_3_2_2_67_1","doi-asserted-by":"crossref","unstructured":"Ruize Wang Zhongyu Wei Piji Li Qi Zhang and Xuanjing Huang. 2020 b. Storytelling from an Image Stream Using Scene Graphs. In AAAI .  Ruize Wang Zhongyu Wei Piji Li Qi Zhang and Xuanjing Huang. 2020 b. Storytelling from an Image Stream Using Scene Graphs. In AAAI .","DOI":"10.1609\/aaai.v34i05.6455"},{"key":"e_1_3_2_2_68_1","doi-asserted-by":"crossref","unstructured":"Xin Wang Wenhu Chen Yuan-Fang Wang and William Yang Wang. 2018a. No metrics are perfect: Adversarial reward learning for visual storytelling. In ACL .  Xin Wang Wenhu Chen Yuan-Fang Wang and William Yang Wang. 2018a. No metrics are perfect: Adversarial reward learning for visual storytelling. In ACL .","DOI":"10.18653\/v1\/P18-1083"},{"volume-title":"Scene graph generation by iterative message passing","author":"Xu Danfei","key":"e_1_3_2_2_69_1"},{"key":"e_1_3_2_2_70_1","doi-asserted-by":"publisher","DOI":"10.5555\/3367722.3367790"},{"key":"e_1_3_2_2_71_1","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3454804"},{"key":"e_1_3_2_2_72_1","unstructured":"Licheng Yu Mohit Bansal and Tamara L Berg. 2017. Hierarchically-attentive rnn for album summarization and storytelling. In EMNLP .  Licheng Yu Mohit Bansal and Tamara L Berg. 2017. Hierarchically-attentive rnn for album summarization and storytelling. In EMNLP ."},{"key":"e_1_3_2_2_73_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413908"},{"volume-title":"Neural motifs: Scene graph parsing with global context","author":"Zellers Rowan","key":"e_1_3_2_2_74_1"},{"key":"e_1_3_2_2_75_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413885"},{"key":"e_1_3_2_2_76_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413880"},{"volume-title":"End-to-end dense video captioning with masked transformer","author":"Zhou Luowei","key":"e_1_3_2_2_77_1"},{"key":"e_1_3_2_2_78_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350932"}],"event":{"name":"MM '21: ACM Multimedia Conference","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Virtual Event China","acronym":"MM '21"},"container-title":["Proceedings of the 29th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474085.3475236","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3474085.3475236","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:48:16Z","timestamp":1750193296000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474085.3475236"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,17]]},"references-count":77,"alternative-id":["10.1145\/3474085.3475236","10.1145\/3474085"],"URL":"https:\/\/doi.org\/10.1145\/3474085.3475236","relation":{},"subject":[],"published":{"date-parts":[[2021,10,17]]},"assertion":[{"value":"2021-10-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}