{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,2]],"date-time":"2025-10-02T09:10:05Z","timestamp":1759396205436,"version":"build-2065373602"},"reference-count":32,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2025,10,2]],"date-time":"2025-10-02T00:00:00Z","timestamp":1759363200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["www.mdpi.com"],"crossmark-restriction":true},"short-container-title":["Symmetry"],"abstract":"<jats:p>With the proliferation of multimedia data, controllable summarization generation has become a key focus in Artificial Intelligence Content Generation. However, many traditional methods lack precise control over output length, often resulting in summaries that are either too verbose or too brief, thus failing to meet diverse user needs. In this paper, we propose a length-customizable approach for multimodal image-text summarization. Our method integrates combinatorial optimization with deep learning to address the length-control challenge. Specifically, we formulate the summarization task as a knapsack optimization problem, enhanced by a greedy algorithm to strictly adhere to user-defined length constraints. Additionally, we introduce a multimodal attention mechanism to ensure balanced and coherent integration of textual and visual information. To further enhance semantic alignment, we employ a cross-modal matching strategy for image selection based on pre-trained vision-language models. Experimental evaluations on the MSMO dataset and validate against baselines like LEAD-3, Seq2Seq, Attention, and Transformer that our method achieves a ROUGE-1 score of 40.52, ROUGE-2 of 16.07, and ROUGE-L of 35.15, outperforming existing length-controllable baselines. Moreover, our approach attains the lowest length variance, confirming its precise adherence to target summary lengths. These results validate the effectiveness of our method in generating high-quality, length-constrained multimodal summaries.<\/jats:p>","DOI":"10.3390\/sym17101629","type":"journal-article","created":{"date-parts":[[2025,10,2]],"date-time":"2025-10-02T08:20:28Z","timestamp":1759393228000},"page":"1629","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Customizable Length Constrained Image-Text Summarization via Knapsack Optimization"],"prefix":"10.3390","volume":"17","author":[{"given":"Xuan","family":"Liu","sequence":"first","affiliation":[{"name":"Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance of MOE, Minzu University of China, Beijing 100081, China"},{"name":"Hainan Li\u2019an International Education Innovation Pilot Zone Management Bureau Information and Technology Innovation Department, Lingshui 572423, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-4478-398X","authenticated-orcid":false,"given":"Xiangyu","family":"Qu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance of MOE, Minzu University of China, Beijing 100081, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yu","family":"Weng","sequence":"additional","affiliation":[{"name":"Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance of MOE, Minzu University of China, Beijing 100081, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6766-0703","authenticated-orcid":false,"given":"Yutong","family":"Gao","sequence":"additional","affiliation":[{"name":"Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance of MOE, Minzu University of China, Beijing 100081, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8522-565X","authenticated-orcid":false,"given":"Zheng","family":"Liu","sequence":"additional","affiliation":[{"name":"Key Laboratory of Ethnic Language Intelligent Analysis and Security Governance of MOE, Minzu University of China, Beijing 100081, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xianggan","family":"Liu","sequence":"additional","affiliation":[{"name":"Hainan Li\u2019an International Education Innovation Pilot Zone Management Bureau Information and Technology Innovation Department, Lingshui 572423, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,10,2]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1483","DOI":"10.1109\/TCSS.2023.3252723","article-title":"Multidocument Aspect Classification for Aspect-Based Abstractive Summarization","volume":"11","author":"Wang","year":"2023","journal-title":"IEEE Trans. Comput. Soc. Syst."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"550","DOI":"10.1109\/TNSM.2019.2906191","article-title":"An Integrated Event Summarization Approach for Complex System Management","volume":"16","author":"Yang","year":"2019","journal-title":"IEEE Trans. Netw. Serv. Manag."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"40311","DOI":"10.1109\/ACCESS.2024.3377463","article-title":"End to End Urdu Abstractive Text Summarization with Dataset and Improvement in Evaluation Metric","volume":"12","author":"Raza","year":"2024","journal-title":"IEEE Access"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Liu, Y., and Lapata, M. (2019, January 3\u20137). Text Summarization with Pretrained Encoders. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China.","DOI":"10.18653\/v1\/D19-1387"},{"key":"ref_5","unstructured":"Zhang, J., Zhao, Y., Saleh, M., and Liu, P. (2020, January 13\u201318). PEGASUS: Pre-Training with Extracted Gap-sentences for Abstractive Summarization. Proceedings of the International Conference on Machine Learning PMLR, Virtual."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"6010","DOI":"10.1109\/TCSS.2024.3384627","article-title":"Unsupervised Video Summarization Based on the Diffusion Model of Feature Fusion","volume":"11","author":"Yu","year":"2024","journal-title":"IEEE Trans. Comput. Soc. Syst."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"4859","DOI":"10.1007\/s11042-023-15274-4","article-title":"A global and local information extraction model incorporating selection mechanism for abstractive text summarization","volume":"82","author":"Li","year":"2024","journal-title":"Multim. Tools Appl."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"996","DOI":"10.1109\/TKDE.2018.2848260","article-title":"Read, watch, listen, and summarize: Multi-modal summarization for asynchronous text, image, audio and video","volume":"31","author":"Li","year":"2018","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Kikuchi, Y., Neubig, G., Sasano, R., Takamura, H., and Okumura, M. (2016, January 1\u20135). Controlling output length in neural encoder-decoders. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Austin, TX, USA.","DOI":"10.18653\/v1\/D16-1140"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Fan, A., Grangier, D., and Auli, M. (2018, January 20). Controllable abstractive summarization. Proceedings of the 2nd Workshop on Neural Machine Translation and Generation, Melbourne, Australia.","DOI":"10.18653\/v1\/W18-2706"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Yu, Z., Wu, Z., Zheng, H., Zhe, X., Fong, J., and Su, W. (2021). LenAtten: An Effective Length Controlling Unit for Text Summarization, ACL\/IJCNLP (Findings).","DOI":"10.18653\/v1\/2021.findings-acl.31"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Liu, Y., Jia, Q., and Zhu, K. (2022, January 22\u201327). Length control in abstractive summarization by pretraining information selection. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, Dublin, Ireland.","DOI":"10.18653\/v1\/2022.acl-long.474"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Chen, P., Wu, S., Chen, Z., Zhang, J., Zhang, X., and Feng, Z. (2023, January 20\u201323). LenANet: A Length-Controllable Attention Network for Source Code Summarization. Proceedings of the International Conference on Neural Information Processing, Changsha, China.","DOI":"10.1007\/978-981-99-8145-8_43"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"102440","DOI":"10.1016\/j.datak.2025.102440","article-title":"Customized long short-term memory architecture for multi-document summarization with improved text feature set","volume":"159","author":"Deo","year":"2025","journal-title":"Data Knowl. Eng."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Urlana, A., Mishra, P., Roy, T., and Mishra, R. (2024). Controllable Text Summarization: Unraveling Challenges, Approaches, and Prospects\u2014A Survey, ACL (Findings).","DOI":"10.18653\/v1\/2024.findings-acl.93"},{"key":"ref_16","unstructured":"Makino, T., Iwakura, T., Takamura, H., and Okumura, M. (August, January 28). Global optimization under length constraint for neural text summarization. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Baek, D., Kim, J., and Lee, H. (2023). VATMAN: Video-Audio-Text Multimodal Abstractive Summarization with Trimodal Hierarchical Multi-Head Attention, ICTC.","DOI":"10.1109\/ICTC58733.2023.10392391"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"109578","DOI":"10.1016\/j.patcog.2023.109578","article-title":"Topic-aware video summarization using multimodal transformer","volume":"140","author":"Zhu","year":"2023","journal-title":"Pattern Recognit."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1553","DOI":"10.1109\/TMM.2013.2267205","article-title":"Multimodal saliency and fusion for movie summarization based on aural, visual, and textual attention","volume":"15","author":"Evangelopoulos","year":"2013","journal-title":"IEEE Trans. Multimed."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Li, H., Zhu, J., Ma, C., Zhang, J., and Zong, C. (2017, January 6\u201310). Multi-modal summarization for asynchronous collection of text, image, audio and video. Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), Singapore.","DOI":"10.18653\/v1\/D17-1114"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zhu, J., Li, H., Liu, T., Zhou, Y., Zhang, J., and Zong, C. (November, January 31). MSMO: Multimodal Summarization with Multimodal Output. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium.","DOI":"10.18653\/v1\/D18-1448"},{"key":"ref_22","first-page":"11757","article-title":"Unims: A unified framework for multimodal summarization with knowledge distillation","volume":"36","author":"Zhang","year":"2022","journal-title":"AAAI Conf. Artif. Intell."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Krubinski, M., and Pecina, P. (2023). MLASK: Multimodal Summarization of Video-Based News Articles, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2023.findings-eacl.67"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Liu, Y., and Lapata, M. (2019). Hierarchical Transformers for Multi-Document Summarization, Association for Computational Linguistics.","DOI":"10.18653\/v1\/P19-1500"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Liu, Y., Luo, Z., and Zhu, K. (November, January 31). Controlling length in abstractive summarization using a convolutional neural network. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium.","DOI":"10.18653\/v1\/D18-1444"},{"key":"ref_26","unstructured":"Zhu, J., Zhou, L., Li, H., Zhang, J., Zhou, Y., and Zong, C. (2017, January 8\u201312). Augmenting neural sentence summarization through extractive summarization. Proceedings of the 6th Conference on Natural Language Processing and Chinese Computing (NLPCC), Dalian, China."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Trabelsi, M., and Uzunalioglu, H. (2023). Absformer: Transformer-Based Model for Unsupervised Multi-Document Abstractive Summarization, ICDAR 2023 Workshops.","DOI":"10.1007\/978-3-031-41501-2_11"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Gao, S., Zhang, Y., Huang, Y., Tan, K., and Yu, Z. (2025). A Mixed-Language Multi-Document News Summarization Dataset and a Graphs-Based Extract-Generate Model, Association for Computational Linguistics. Association for Computational Linguistics, NAACL (Long Papers).","DOI":"10.18653\/v1\/2025.naacl-long.468"},{"key":"ref_29","first-page":"585","article-title":"Innovative abstractive Hindi text summarization model incorporating Bi-LSTM classifier, optimizer and generative AI","volume":"19","author":"Verma","year":"2025","journal-title":"Intell. Decis. Technol."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"959","DOI":"10.1109\/TCSS.2023.3244068","article-title":"Multimodal fake news analysis based on image\u2013text similarity","volume":"11","author":"Zhang","year":"2023","journal-title":"IEEE Trans. Comput. Soc. Syst."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Cho, K., Merri\u00ebnboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014, January 25\u201329). Learning Phrase Representations using RNN Encoder\u2013Decoder for Statistical Machine Translation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar.","DOI":"10.3115\/v1\/D14-1179"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Steffes, B., Rataj, P., Burger, L., and Roth, L. (2023, January 19\u201323). On Evaluating Legal Summaries with ROUGE. Proceedings of the Nineteenth International Conference on Artificial Intelligence and Law, ICAIL, Braga, Portugal.","DOI":"10.1145\/3594536.3595150"}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/10\/1629\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,2]],"date-time":"2025-10-02T08:28:11Z","timestamp":1759393691000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/10\/1629"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,2]]},"references-count":32,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2025,10]]}},"alternative-id":["sym17101629"],"URL":"https:\/\/doi.org\/10.3390\/sym17101629","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,2]]}}}