{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T16:54:52Z","timestamp":1777654492190,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":65,"publisher":"ACM","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62236003, 62306090, 62476071, U24A20328"],"award-info":[{"award-number":["62236003, 62306090, 62476071, U24A20328"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Shenzhen College Stability Support Plan","award":["GXWD20220817144428005"],"award-info":[{"award-number":["GXWD20220817144428005"]}]},{"name":"Natural Science Foundation of Guangdong Province of China","award":["2024A1515010147, 2025A1515011732"],"award-info":[{"award-number":["2024A1515010147, 2025A1515011732"]}]},{"DOI":"10.13039\/501100002858","name":"China Postdoctoral Science Foundation","doi-asserted-by":"publisher","award":["2024M764192"],"award-info":[{"award-number":["2024M764192"]}],"id":[{"id":"10.13039\/501100002858","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,10,27]]},"DOI":"10.1145\/3746027.3755029","type":"proceedings-article","created":{"date-parts":[[2025,10,25]],"date-time":"2025-10-25T05:47:42Z","timestamp":1761371262000},"page":"7653-7662","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-1562-804X","authenticated-orcid":false,"given":"Yibo","family":"Lyu","sequence":"first","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0090-9604","authenticated-orcid":false,"given":"Rui","family":"Shao","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0634-6075","authenticated-orcid":false,"given":"Gongwei","family":"Chen","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-3601-7893","authenticated-orcid":false,"given":"Yijie","family":"Zhu","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology , Shenzhen, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5658-5509","authenticated-orcid":false,"given":"Weili","family":"Guan","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1476-0273","authenticated-orcid":false,"given":"Liqiang","family":"Nie","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology, Shenzhen, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,10,27]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02506"},{"key":"e_1_3_2_1_2_1","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision.","author":"Chen Gongwei","year":"2025","unstructured":"Gongwei Chen, Xurui Zhou, Rui Shao, Yibo Lyu, Kaiwen Zhou, Shuai Wang, Wentao Li, Yinchuan Li, Zhongang Qi, and Liqiang Nie. 2025. Less is More: Empowering GUI Agent with Context-Aware Simplification. In Proceedings of the IEEE\/CVF International Conference on Computer Vision."},{"key":"e_1_3_2_1_3_1","unstructured":"Junya Chen Zhe Gan Xuan Li Qing Guo Liqun Chen Shuyang Gao Tagyoung Chung Yi Xu Belinda Zeng Wenlian Lu et al. 2021a. Simpler faster stronger: Breaking the log-k curse on contrastive learners with flatnce. arXiv preprint arXiv:2107.01152 (2021)."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-acl.137"},{"key":"e_1_3_2_1_5_1","volume-title":"European Conference on Computer Vision. Springer, 19-35","author":"Chen Liang","year":"2024","unstructured":"Liang Chen, Haozhe Zhao, Tianyu Liu, Shuai Bai, Junyang Lin, Chang Zhou, and Baobao Chang. 2024c. An image is worth 1\/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models. In European Conference on Computer Vision. Springer, 19-35."},{"key":"e_1_3_2_1_6_1","volume-title":"International conference on machine learning. PmLR, 1597-1607","author":"Chen Ting","year":"2020","unstructured":"Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International conference on machine learning. PmLR, 1597-1607."},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00950"},{"key":"e_1_3_2_1_8_1","volume-title":"Not all layers of llms are necessary during inference. arXiv preprint arXiv:2403.02181","author":"Fan Siqi","year":"2024","unstructured":"Siqi Fan, Xin Jiang, Xiang Li, Xuying Meng, Peng Han, Shuo Shang, Aixin Sun, Yequan Wang, and Zhongyuan Wang. 2024. Not all layers of llms are necessary during inference. arXiv preprint arXiv:2403.02181 (2024)."},{"key":"e_1_3_2_1_9_1","unstructured":"Tim Fischer Chris Biemann et al. 2024. Large language models are overparameterized text encoders. arXiv preprint arXiv:2410.14578 (2024)."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.552"},{"key":"e_1_3_2_1_11_1","unstructured":"Aaron Grattafiori Abhimanyu Dubey Abhinav Jauhri Abhinav Pandey Abhishek Kadian Ahmad Al-Dahle Aiesha Letman Akhil Mathur Alan Schelten Alex Vaughan et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)."},{"key":"e_1_3_2_1_12_1","volume-title":"The unreasonable ineffectiveness of the deeper layers","author":"Gromov Andrey","year":"2024","unstructured":"Andrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso, and Daniel A Roberts. [n.d.]. The unreasonable ineffectiveness of the deeper layers, 2024. URL https:\/\/arxiv.org\/abs\/2403.17887 ([n.d.])."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"e_1_3_2_1_14_1","first-page":"3","article-title":"Lora: Low-rank adaptation of large language models","volume":"1","author":"Hu Edward J","year":"2022","unstructured":"Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al., 2022. Lora: Low-rank adaptation of large language models. ICLR, Vol. 1, 2 (2022), 3.","journal-title":"ICLR"},{"key":"e_1_3_2_1_15_1","volume-title":"Joint Fusion and Encoding: Advancing Multimodal Retrieval from the Ground Up. arXiv preprint arXiv:2502.20008","author":"Huang Lang","year":"2025","unstructured":"Lang Huang, Qiyu Wu, Zhongtao Miao, and Toshihiko Yamasaki. 2025. Joint Fusion and Encoding: Advancing Multimodal Retrieval from the Ground Up. arXiv preprint arXiv:2502.20008 (2025)."},{"key":"e_1_3_2_1_16_1","volume-title":"International conference on machine learning. PMLR, 4904-4916","author":"Jia Chao","year":"2021","unstructured":"Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. In International conference on machine learning. PMLR, 4904-4916."},{"key":"e_1_3_2_1_17_1","volume-title":"Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al.","author":"Jiang Albert Q","year":"2024","unstructured":"Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al., 2024a. Mixtral of experts. arXiv preprint arXiv:2401.04088 (2024)."},{"key":"e_1_3_2_1_18_1","volume-title":"E5-v: Universal embeddings with multimodal large language models. arXiv preprint arXiv:2407.12580","author":"Jiang Ting","year":"2024","unstructured":"Ting Jiang, Minghui Song, Zihan Zhang, Haizhen Huang, Weiwei Deng, Feng Sun, Qi Zhang, Deqing Wang, and Fuzhen Zhuang. 2024b. E5-v: Universal embeddings with multimodal large language models. arXiv preprint arXiv:2407.12580 (2024)."},{"key":"e_1_3_2_1_19_1","volume-title":"What's in the Image? A Deep-Dive into the Vision of Vision Language Models. arXiv preprint arXiv:2411.17491","author":"Kaduri Omri","year":"2024","unstructured":"Omri Kaduri, Shai Bagon, and Tali Dekel. 2024. What's in the Image? A Deep-Dive into the Vision of Vision Language Models. arXiv preprint arXiv:2411.17491 (2024)."},{"key":"e_1_3_2_1_20_1","volume-title":"Noe Pion, Philippe Weinzaepfel, and Diane Larlus.","author":"Kalantidis Yannis","year":"2020","unstructured":"Yannis Kalantidis, Mert Bulent Sariyildiz, Noe Pion, Philippe Weinzaepfel, and Diane Larlus. 2020. Hard negative mixing for contrastive learning. Advances in neural information processing systems, Vol. 33 (2020), 21798-21809."},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01225-0_13"},{"key":"e_1_3_2_1_22_1","volume-title":"Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326","author":"Li Bo","year":"2024","unstructured":"Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, et al., 2024b. Llava-onevision: Easy visual task transfer. arXiv preprint arXiv:2408.03326 (2024)."},{"key":"e_1_3_2_1_23_1","volume-title":"Forty-second International Conference on Machine Learning.","author":"Li Hao","unstructured":"Hao Li, Qi Lv, Rui Shao, Xiang Deng, Yinchuan Li, Jianye HAO, and Liqiang Nie. [n.d.]. STAR: Learning Diverse Robot Skill Abstractions through Rotation-Augmented Vector Quantization. In Forty-second International Conference on Machine Learning."},{"key":"e_1_3_2_1_24_1","volume-title":"International conference on machine learning. PMLR","author":"Li Junnan","year":"2023","unstructured":"Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023a. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning. PMLR, 19730-19742."},{"key":"e_1_3_2_1_25_1","volume-title":"International conference on machine learning. PMLR, 12888-12900","author":"Li Junnan","year":"2022","unstructured":"Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International conference on machine learning. PMLR, 12888-12900."},{"key":"e_1_3_2_1_26_1","volume-title":"Align before fuse: Vision and language representation learning with momentum distillation. Advances in neural information processing systems","author":"Li Junnan","year":"2021","unstructured":"Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021. Align before fuse: Vision and language representation learning with momentum distillation. Advances in neural information processing systems, Vol. 34 (2021), 9694-9705."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.00308"},{"key":"e_1_3_2_1_28_1","volume-title":"Kaiwen Zhou, Jiqian Dong, Kaiyang Guo, Xiu Li, Zhitang Chen, et al.","author":"Li Yinchuan","year":"2025","unstructured":"Yinchuan Li, Xinyu Shao, Jianping Zhang, Haozhi Wang, Leo Maxime Brunswic, Kaiwen Zhou, Jiqian Dong, Kaiyang Guo, Xiu Li, Zhitang Chen, et al., 2025b. Generative models in decision making: A survey. arXiv preprint arXiv:2502.17100 (2025)."},{"key":"e_1_3_2_1_29_1","first-page":"49881","article-title":"Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks","volume":"37","author":"Li Zaijing","year":"2024","unstructured":"Zaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen, Dongmei Jiang, and Liqiang Nie. 2024a. Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks. In Advances in Neural Information Processing Systems, Vol. 37. 49881-49913.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.00845"},{"key":"e_1_3_2_1_31_1","volume-title":"Towards general text embeddings with multi-stage contrastive learning. arXiv preprint arXiv:2308.03281","author":"Li Zehan","year":"2023","unstructured":"Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. 2023b. Towards general text embeddings with multi-stage contrastive learning. arXiv preprint arXiv:2308.03281 (2023)."},{"key":"e_1_3_2_1_32_1","volume-title":"MM-EMBED: UNIVERSAL MULTIMODAL RETRIEVAL WITH MULTIMODAL LLMS. In The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=i45NQb2iKO","author":"Lin Sheng-Chieh","year":"2025","unstructured":"Sheng-Chieh Lin, Chankyu Lee, Mohammad Shoeybi, Jimmy Lin, Bryan Catanzaro, and Wei Ping. 2025. MM-EMBED: UNIVERSAL MULTIMODAL RETRIEVAL WITH MULTIMODAL LLMS. In The Thirteenth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=i45NQb2iKO"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.emnlp-main.987"},{"key":"e_1_3_2_1_34_1","volume-title":"Visual instruction tuning. Advances in neural information processing systems","author":"Liu Haotian","year":"2023","unstructured":"Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruction tuning. Advances in neural information processing systems, Vol. 36 (2023), 34892-34916."},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.00380"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2022.07.028"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.eacl-main.148"},{"key":"e_1_3_2_1_38_1","volume-title":"Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748","author":"van den Oord Aaron","year":"2018","unstructured":"Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748 (2018)."},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01820"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.emnlp-main.435"},{"key":"e_1_3_2_1_41_1","volume-title":"International conference on machine learning. PmLR, 8748-8763","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning. PmLR, 8748-8763."},{"key":"e_1_3_2_1_42_1","volume-title":"CONTRASTIVE LEARNING WITH HARD NEGATIVE SAMPLES. In International Conference on Learning Representations (ICLR).","author":"Robinson Joshua","year":"2021","unstructured":"Joshua Robinson, Ching-Yao Chuang, Suvrit Sra, and Stefanie Jegelka. 2021. CONTRASTIVE LEARNING WITH HARD NEGATIVE SAMPLES. In International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_1_43_1","volume-title":"Antoine Chassang, Carlo Gatta, and Yoshua Bengio.","author":"Romero Adriana","year":"2014","unstructured":"Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2014. Fitnets: Hints for thin deep nets. arXiv preprint arXiv:1412.6550 (2014)."},{"key":"e_1_3_2_1_44_1","volume-title":"Lightweight safety classification using pruned language models. arXiv preprint arXiv:2412.13435","author":"Sawtell Mason","year":"2024","unstructured":"Mason Sawtell, Tula Masterman, Sandi Besen, and Jim Brown. 2024. Lightweight safety classification using pruned language models. arXiv preprint arXiv:2412.13435 (2024)."},{"key":"e_1_3_2_1_45_1","volume-title":"Yong Jae Lee, and Yan Yan","author":"Shang Yuzhang","year":"2024","unstructured":"Yuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee, and Yan Yan. 2024. Llava-prumerge: Adaptive token reduction for efficient large multimodal models. arXiv preprint arXiv:2403.15388 (2024)."},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01026"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00667"},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3367749"},{"key":"e_1_3_2_1_49_1","first-page":"42048","article-title":"MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models","volume":"37","author":"Shen Leyang","year":"2024","unstructured":"Leyang Shen, Gongwei Chen, Rui Shao, Weili Guan, and Liqiang Nie. 2024. MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models. In Advances in Neural Information Processing Systems, Vol. 37. 42048-42070.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_50_1","unstructured":"Gemini Team Rohan Anil Sebastian Borgeaud Jean-Baptiste Alayrac Jiahui Yu Radu Soricut Johan Schalkwyk Andrew M Dai Anja Hauth Katie Millican et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805 (2023)."},{"key":"e_1_3_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.acl-long.642"},{"key":"e_1_3_2_1_52_1","unstructured":"Peng Wang Shuai Bai Sinan Tan Shijie Wang Zhihao Fan Jinze Bai Keqin Chen Xuejing Liu Jialin Wang Wenbin Ge et al. 2024a. Qwen2-vl: Enhancing vision-language model's perception of the world at any resolution. arXiv preprint arXiv:2409.12191 (2024)."},{"key":"e_1_3_2_1_53_1","volume-title":"European Conference on Computer Vision. Springer, 387-404","author":"Wei Cong","year":"2024","unstructured":"Cong Wei, Yang Chen, Haonan Chen, Hexiang Hu, Ge Zhang, Jie Fu, Alan Ritter, and Wenhu Chen. 2024. Uniir: Training and benchmarking universal multimodal information retrievers. In European Conference on Computer Vision. Springer, 387-404."},{"key":"e_1_3_2_1_54_1","volume-title":"Efficient vision-language models by summarizing visual tokens into compact registers. arXiv preprint arXiv:2410.14072","author":"Wen Yuxin","year":"2024","unstructured":"Yuxin Wen, Qingqing Cao, Qichen Fu, Sachin Mehta, and Mahyar Najibi. 2024. Efficient vision-language models by summarizing visual tokens into compact registers. arXiv preprint arXiv:2410.14072 (2024)."},{"key":"e_1_3_2_1_55_1","volume-title":"GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent. In Annual Meeting of the Association for Computational Linguistics (ACL).","author":"Xie Bin","year":"2025","unstructured":"Bin Xie, Rui Shao, Gongwei Chen, Kaiwen Zhou, Yinchuan Li, Jie Liu, Min Zhang, and Liqiang Nie. 2025. GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent. In Annual Meeting of the Association for Computational Linguistics (ACL)."},{"key":"e_1_3_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-emnlp.372"},{"key":"e_1_3_2_1_57_1","volume-title":"Voco-llama: Towards vision compression with large language models. arXiv preprint arXiv:2406.12275","author":"Ye Xubing","year":"2024","unstructured":"Xubing Ye, Yukang Gan, Xiaoke Huang, Yixiao Ge, Ying Shan, and Yansong Tang. 2024. Voco-llama: Towards vision compression with large language models. arXiv preprint arXiv:2406.12275 (2024)."},{"key":"e_1_3_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01100"},{"key":"e_1_3_2_1_59_1","volume-title":"Token-level correlation-guided compression for efficient multimodal document understanding. arXiv preprint arXiv:2407.14439","author":"Zhang Renshan","year":"2024","unstructured":"Renshan Zhang, Yibo Lyu, Rui Shao, Gongwei Chen, Weili Guan, and Liqiang Nie. 2024. Token-level correlation-guided compression for efficient multimodal document understanding. arXiv preprint arXiv:2407.14439 (2024)."},{"key":"e_1_3_2_1_60_1","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV).","author":"Zhang Renshan","year":"2025","unstructured":"Renshan Zhang, Rui Shao, Gongwei Chen, Miao Zhang, Kaiwen Zhou, Weili Guan, and Liqiang Nie. 2025c. FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)."},{"key":"e_1_3_2_1_61_1","volume-title":"LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token. arXiv preprint arXiv:2501.03895","author":"Zhang Shaolei","year":"2025","unstructured":"Shaolei Zhang, Qingkai Fang, Zhe Yang, and Yang Feng. 2025a. LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token. arXiv preprint arXiv:2501.03895 (2025)."},{"key":"e_1_3_2_1_62_1","volume-title":"Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics","author":"Zhang Xiaofeng","year":"2025","unstructured":"Xiaofeng Zhang, Yihao Quan, Chen Shen, Xiaosong Yuan, Shaotian Yan, Liang Xie, Wenxiao Wang, Chaochen Gu, Hao Tang, and Jieping Ye. 2025b. From Redundancy to Relevance: Enhancing Explainability in Multimodal Large Language Models. Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (2025)."},{"key":"e_1_3_2_1_63_1","volume-title":"FinerCut: Finer-grained Interpretable Layer Pruning for Large Language Models. In Workshop on Machine Learning and Compression, NeurIPS","author":"Zhang Yang","year":"2024","unstructured":"Yang Zhang, Yawei Li, Xinpeng Wang, Qianli Shen, Barbara Plank, Bernd Bischl, Mina Rezaei, and Kenji Kawaguchi. [n.d.]. FinerCut: Finer-grained Interpretable Layer Pruning for Large Language Models. In Workshop on Machine Learning and Compression, NeurIPS 2024."},{"key":"e_1_3_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_42"},{"key":"e_1_3_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/3746027.3754549"}],"event":{"name":"MM '25: The 33rd ACM International Conference on Multimedia","location":"Dublin Ireland","acronym":"MM '25","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 33rd ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3746027.3755029","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,9]],"date-time":"2025-12-09T19:16:09Z","timestamp":1765307769000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3746027.3755029"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,27]]},"references-count":65,"alternative-id":["10.1145\/3746027.3755029","10.1145\/3746027"],"URL":"https:\/\/doi.org\/10.1145\/3746027.3755029","relation":{},"subject":[],"published":{"date-parts":[[2025,10,27]]},"assertion":[{"value":"2025-10-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}