{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T05:06:11Z","timestamp":1750309571845,"version":"3.41.0"},"reference-count":47,"publisher":"Association for Computing Machinery (ACM)","issue":"6","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62366021"],"award-info":[{"award-number":["62366021"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2025,6,30]]},"abstract":"<jats:p>\n            Visual storytelling (VIST) involves generating coherent, creative, and vivid narrative for a collection of images. It remains an immense challenge within cross-modal domain. Traditional mainstream storytelling work were less proficient in handling long sequential relationships. Though large-scale visual-language pre-training (VLP) models demonstrated promising prospect on cross-modal tasks. So far they still were not particularly adept at handling tasks involving image sequences. Moreover, the reference descriptions in the available VIST benchmark dataset are short and simplistic, which constrains model\u2019s potential capabilities. Current models struggle to produce truly rich and vivid narratives. Therefore, in this article, we will address these deficiencies, and contribute from both dataset and innovative model aspects. Firstly, by leveraging large language model (LLM), we have constructed a new dataset VIST++ which can enrich vivid narratives for open-ended image sequences. The dataset has potential on providing beneficial support for future model learning. Secondly, we have proposed an innovative auto-regressive story generation model named ReStoryGen. It can be applied to image sequences of varying lengths in an open-ended way. We have performed extensive experiments and evaluations in terms of visual grounding, coherence and non-redundancy. The experiment results have convincingly demonstrated ReStoryGen achieves impressive outcomes through utilizing parameter-efficient instruction-tuning. Related source codes and models are distributed on Github\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"url\" xlink:href=\"https:\/\/github.com\/lixinliu1995\/story_gen\">https:\/\/github.com\/lixinliu1995\/story_gen<\/jats:ext-link>\n            .\n          <\/jats:p>","DOI":"10.1145\/3736759","type":"journal-article","created":{"date-parts":[[2025,5,28]],"date-time":"2025-05-28T09:26:20Z","timestamp":1748424380000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Open-ended Autoregressive Visual Storytelling via Parameter Efficient Instruction Tuning"],"prefix":"10.1145","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-2281-9939","authenticated-orcid":false,"given":"Lixin","family":"Liu","sequence":"first","affiliation":[{"name":"School of Digital Industry, Jiangxi Normal University","place":["Shangrao, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5979-7590","authenticated-orcid":false,"given":"Aiwen","family":"Jiang","sequence":"additional","affiliation":[{"name":"School of Digital Industry, Jiangxi Normal University","place":["Shangrao, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-9217-1817","authenticated-orcid":false,"given":"Yinuo","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Digital Industry, Jiangxi Normal University","place":["Shangrao, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4673-5806","authenticated-orcid":false,"given":"Changhong","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Computer and Infomation Engineering, Jiangxi Normal University","place":["Nanchang, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3603-1875","authenticated-orcid":false,"given":"Qi","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Computer and Information Engineering, Jiangxi Normal University","place":["Nanchang, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8118-9612","authenticated-orcid":false,"given":"Mingwen","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Digital Industry, Jiangxi Normal University","place":["Shangrao, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,6,18]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"5859","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"35","author":"Ammanabrolu Prithviraj","year":"2021","unstructured":"Prithviraj Ammanabrolu, Wesley Cheung, William Broniec, and Mark O. Riedl. 2021. Automated storytelling via causal, commonsense plot ordering. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 5859\u20135867."},{"key":"e_1_3_1_3_2","unstructured":"Jinze Bai Shuai Bai Shusheng Yang Shijie Wang Sinan Tan Peng Wang Junyang Lin Chang Zhou and Jingren Zhou. 2023. Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv:2308.12966. Retrieved from https:\/\/arxiv.org\/abs\/2308.12966"},{"key":"e_1_3_1_4_2","first-page":"999","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"35","author":"Chen Hong","year":"2021","unstructured":"Hong Chen, Yifei Huang, Hiroya Takamura, and Hideki Nakayama. 2021. Commonsense knowledge aware concept selection for diverse and informative visual storytelling. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 999\u20131008."},{"key":"e_1_3_1_5_2","unstructured":"Wei-Lin Chiang Zhuohan Li Zi Lin Ying Sheng Zhanghao Wu Hao Zhang Lianmin Zheng Siyuan Zhuang Yonghao Zhuang Joseph E. Gonzalez et al. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality. (March2023). Retrieved from https:\/\/lmsys.org\/blog\/2023-03-30-vicuna\/"},{"key":"e_1_3_1_6_2","unstructured":"Wenliang Dai Junnan Li Dongxu Li Anthony Meng Huat Tiong Junqi Zhao Weisheng Wang Boyang Li Pascale Fung and Steven Hoi. 2023. InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. In Proceedings of the 37th International Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00675"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00553"},{"key":"e_1_3_1_9_2","first-page":"2790","volume-title":"International Conference on Machine Learning","author":"Houlsby Neil","year":"2019","unstructured":"Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In International Conference on Machine Learning. PMLR, 2790\u20132799."},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6303"},{"key":"e_1_3_1_11_2","unstructured":"Chi-Yang Hsu Yun-Wei Chu Ting-Hao\u2019Kenneth\u2019 Huang and Lun-Wei Ku. 2021. Plot and rework: Modeling storylines for visual storytelling. arXiv:2105.06950. Retrieved from https:\/\/arxiv.org\/abs\/2105.06950"},{"key":"e_1_3_1_12_2","unstructured":"Edward J. Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang Lu Wang and Weizhu Chen. 2022. Lora: Low-rank adaptation of large language models. International Conference on Learning Representations. 1\u201313."},{"key":"e_1_3_1_13_2","first-page":"8465","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"33","author":"Huang Qiuyuan","year":"2019","unstructured":"Qiuyuan Huang, Zhe Gan, Asli Celikyilmaz, Dapeng Wu, Jianfeng Wang, and Xiaodong He. 2019. Hierarchically structured reinforcement learning for topically coherent visual story generation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33. 8465\u20138472."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19827-4_41"},{"key":"e_1_3_1_15_2","doi-asserted-by":"crossref","unstructured":"Yunjae Jung Dahun Kim Sanghyun Woo Kyungsu Kim Sungjin Kim and In So Kweon. 2020. Hide-and-Tell: Learning to Bridge Photo Streams for Visual Storytelling. In Proceedings of the AAAI Conference on Artificial Intelligence 34 (2020) 11213\u201311220.","DOI":"10.1609\/aaai.v34i07.6780"},{"key":"e_1_3_1_16_2","unstructured":"Taehyeong Kim Min-Oh Heo Seonil Son Kyoung-Wha Park and Byoung-Tak Zhang. 2018. Glac net: Glocal attention cascading networks for multi-image cued story generation. arXiv preprintarxiv:1805.10973 (2018)."},{"key":"e_1_3_1_17_2","doi-asserted-by":"crossref","unstructured":"Zhenzhong Lan Mingda Chen Sebastian Goodman Kevin Gimpel Piyush Sharma and Radu Soricut. 2020. Albert: A lite bert for self-supervised learning of language representations. International Conference on Learning Representations. 1\u201317.","DOI":"10.18653\/v1\/2020.repl4nlp-1.3"},{"key":"e_1_3_1_18_2","doi-asserted-by":"crossref","unstructured":"Brian Lester Rami Al-Rfou and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 3045\u20133059.","DOI":"10.18653\/v1\/2021.emnlp-main.243"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2407.07895"},{"key":"e_1_3_1_20_2","unstructured":"Junnan Li Dongxu Li Silvio Savarese and Steven Hoi. 2023. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. In Proceedings of the 40th International Conference on Machine Learning. 19730\u201319742."},{"key":"e_1_3_1_21_2","first-page":"12888","volume-title":"International Conference on Machine Learning","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning. PMLR, 12888\u201312900."},{"key":"e_1_3_1_22_2","article-title":"Knowledge-enriched attention network with group-wise semantic for visual storytelling","author":"Li Tengpeng","year":"2022","unstructured":"Tengpeng Li, Hanli Wang, Bin He, and Chang Wen Chen. 2022. Knowledge-enriched attention network with group-wise semantic for visual storytelling. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 7 (2022), 8634\u20138645.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_23_2","unstructured":"Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics. 4582\u20134597."},{"key":"e_1_3_1_24_2","first-page":"6329","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Yitong","year":"2019","unstructured":"Yitong Li, Zhe Gan, Yelong Shen, Jingjing Liu, Yu Cheng, Yuexin Wu, Lawrence Carin, David Carlson, and Jianfeng Gao. 2019. Storygan: A sequential conditional gan for story visualization. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 6329\u20136338."},{"key":"e_1_3_1_25_2","unstructured":"Zhaojiang Lin Andrea Madotto and Pascale Fung. 2020. Exploring versatile generative language model via parameter-efficient transfer learning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing. 441\u2013459."},{"key":"e_1_3_1_26_2","doi-asserted-by":"crossref","unstructured":"Haotian Liu Chunyuan Li Yuheng Li and Yong Jae Lee. 2023. Improved baselines with visual instruction tuning. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 26286\u201326296.","DOI":"10.1109\/CVPR52733.2024.02484"},{"key":"e_1_3_1_27_2","first-page":"1950","article-title":"Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning","volume":"35","author":"Liu Haokun","year":"2022","unstructured":"Haokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin A. Raffel. 2022. Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning. Advances in Neural Information Processing Systems 35 (2022), 1950\u20131965.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_28_2","unstructured":"Xiao Liu Yanan Zheng Zhengxiao Du Ming Ding Yujie Qian Zhilin Yang and Jie Tang. 2021. GPT Understands Too. (2021). arxiv:cs.CL\/2103.10385"},{"key":"e_1_3_1_29_2","article-title":"Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks","volume":"32","author":"Lu Jiasen","year":"2019","unstructured":"Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in Neural Information Processing Systems 32 (2019), 12\u201323.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N16-1098"},{"key":"e_1_3_1_31_2","article-title":"Expressing an image stream with a sequence of natural sentences","volume":"28","author":"Park Cesc C.","year":"2015","unstructured":"Cesc C. Park and Gunhee Kim. 2015. Expressing an image stream with a sequence of natural sentences. Advances in Neural Information Processing Systems 28 (2015), 73\u201381.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_32_2","first-page":"8748","volume-title":"Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research)","volume":"139","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research), Marina Meila and Tong Zhang (Eds.), Vol. 139. PMLR, 8748\u20138763. Retrieved from https:\/\/proceedings.mlr.press\/v139\/radford21a.html"},{"key":"e_1_3_1_33_2","article-title":"Faster r-cnn: Towards real-time object detection with region proposal networks","volume":"28","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in Neural Information Processing Systems 28 (2015), 91\u201399.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_34_2","first-page":"arXiv\u20131909","article-title":"A hierarchical approach for visual storytelling using image description","author":"Nahian Md Sultan Al","year":"2019","unstructured":"Md Sultan Al Nahian, Tasmia Tasrin, Sagar Gandhi, Ryan Gaines, and Brent Harrison. 2019. A hierarchical approach for visual storytelling using image description. arXiv e-prints (2019), arXiv\u20131909.","journal-title":"arXiv e-prints"},{"key":"e_1_3_1_35_2","unstructured":"Qwen Team. 2024. Qwen2.5: A Party of Foundation Models. (September2024). Retrieved from https:\/\/qwenlm.github.io\/blog\/qwen2.5\/"},{"key":"e_1_3_1_36_2","doi-asserted-by":"crossref","unstructured":"Ting-Hao Huang Francis Ferraro Nasrin Mostafazadeh Ishan Misra Aishwarya Agrawal Jacob Devlin Ross Girshick Xiaodong He Pushmeet Kohli Dhruv Batra C. Lawrence Zitnick Devi Parikh Lucy Vanderwende Michel Galley and Margaret Mitchell. 2016. Visual Storytelling. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 1233\u20131239.","DOI":"10.18653\/v1\/N16-1147"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298935"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-naacl.206"},{"key":"e_1_3_1_39_2","first-page":"9185","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"34","author":"Wang Ruize","year":"2020","unstructured":"Ruize Wang, Zhongyu Wei, Piji Li, Qi Zhang, and Xuanjing Huang. 2020. Storytelling from an image stream using scene graphs. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 34. 9185\u20139192."},{"key":"e_1_3_1_40_2","doi-asserted-by":"crossref","unstructured":"Xin Wang Wenhu Chen Yuan-Fang Wang and William Yang Wang. 2018. No metrics are perfect: Adversarial reward learning for visual storytelling. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics. 899\u2013909.","DOI":"10.18653\/v1\/P18-1083"},{"key":"e_1_3_1_41_2","doi-asserted-by":"crossref","unstructured":"Xiyao Wang Yuhang Zhou Xiaoyu Liu Hongjin Lu Yuancheng Xu Feihong He Jaehong Yoon Taixi Lu Gedas Bertasius Mohit Bansal et\u00a0al. 2024. Mementos: A comprehensive benchmark for multimodal large language model reasoning over image sequences. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. 416\u2013442.","DOI":"10.18653\/v1\/2024.acl-long.25"},{"key":"e_1_3_1_42_2","unstructured":"An Yang Baosong Yang Binyuan Hui Bo Zheng Bowen Yu Chang Zhou Chengpeng Li Chengyuan Li Dayiheng Liu Fei Huang et al. 2024. Qwen2 technical report. arXiv preprintarxiv:2407.10671 (2024)."},{"key":"e_1_3_1_43_2","doi-asserted-by":"crossref","unstructured":"Pengcheng Yang Fuli Luo Peng Chen Lei Li Zhiyi Yin Xiaodong He and Xu Sun. 2019. Knowledgeable storyteller: A commonsense-driven generative model for visual storytelling. In Twenty-Eighth International Joint Conference on Artificial Intelligence. 5356\u20135362.","DOI":"10.24963\/ijcai.2019\/744"},{"key":"e_1_3_1_44_2","unstructured":"Taojiannan Yang Yi Zhu Yusheng Xie Aston Zhang Chen Chen and Mu Li. 2023. Aim: Adapting image models for efficient video action recognition. International Conference on Learning Representations. 1\u201318."},{"key":"e_1_3_1_45_2","first-page":"12658","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201921)","author":"Yu Youngjae","year":"2021","unstructured":"Youngjae Yu, Jiwan Chung, Heeseung Yun, Jongseok Kim, and Gunhee Kim. 2021. Transitional adaptation of pretrained models for visual storytelling. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201921). 12658\u201312668."},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-022-01653-1"},{"key":"e_1_3_1_47_2","unstructured":"Deyao Zhu Jun Chen Kilichbek Haydarov Xiaoqian Shen Wenxuan Zhang and Mohamed Elhoseiny. 2024. Chatgpt asks blip-2 answers: Automatic questioning towards enriched visual descriptions. Transactions on Machine Learning Research 3 1-22 (2024)."},{"key":"e_1_3_1_48_2","unstructured":"Deyao Zhu Jun Chen Xiaoqian Shen Xiang Li and Mohamed Elhoseiny. 2024. MiniGPT-4: Enhancing vision-language understanding with advanced large language models. International Conference on Learning Representations. 1\u201317."}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3736759","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:19:15Z","timestamp":1750295955000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3736759"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,18]]},"references-count":47,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2025,6,30]]}},"alternative-id":["10.1145\/3736759"],"URL":"https:\/\/doi.org\/10.1145\/3736759","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"type":"print","value":"2375-4699"},{"type":"electronic","value":"2375-4702"}],"subject":[],"published":{"date-parts":[[2025,6,18]]},"assertion":[{"value":"2024-07-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-05-15","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-06-18","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}