{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T17:03:45Z","timestamp":1784567025638,"version":"3.55.0"},"reference-count":71,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2026,6,1]],"date-time":"2026-06-01T00:00:00Z","timestamp":1780272000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T00:00:00Z","timestamp":1780963200000},"content-version":"vor","delay-in-days":8,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100005713","name":"Technische Universit\u00e4t M\u00fcnchen","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100005713","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2026,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>While Large Vision Language Models (LVLMs) demonstrate impressive capabilities, their substantial computational and memory requirements pose deployment challenges on resource-constrained edge devices. Current parameter reduction techniques primarily involve training LVLMs from small language models, but these methods offer limited flexibility and remain computationally intensive. We study a complementary route: compressing existing LVLMs by applying structured pruning to the language model backbone, followed by lightweight recovery training. Specifically, we investigate two structural pruning paradigms: layerwise and widthwise pruning, and pair them with supervised finetuning and knowledge distillation on logits and hidden states. Additionally, we assess the feasibility of conducting recovery training with only a small fraction of the available data. Our results show that widthwise pruning generally maintains better performance in low-resource scenarios, where computational resources are limited or there is insufficient finetuning data. As for the recovery training, finetuning only the multimodal projector is sufficient at small compression levels. Furthermore, a combination of supervised finetuning and hidden-state distillation yields optimal recovery across various pruning levels. Notably, effective recovery can be achieved using just 5% of the original data, while retaining over 95% of the original performance. Through empirical study on three representative LVLM families ranging from 3B to 7B parameters, this study offers actionable insights for practitioners to compress LVLMs without extensive computation resources or sufficient data.<\/jats:p>","DOI":"10.1007\/s11263-026-02865-5","type":"journal-article","created":{"date-parts":[[2026,6,9]],"date-time":"2026-06-09T17:12:15Z","timestamp":1781025135000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency"],"prefix":"10.1007","volume":"134","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-6937-5668","authenticated-orcid":false,"given":"Yiran","family":"Huang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lukas","family":"Thede","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Massimiliano","family":"Mancini","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenjia","family":"Xu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zeynep","family":"Akata","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,6,9]]},"reference":[{"key":"2865_CR1","unstructured":"Bo, L., Peiyuan, Z., Kaichen, Z., Fanyi, P., Xinrun, D., Yuhao, D., Haotian, L., Yuanhan, Z., Ge, Z., Chunyuan, L., & Ziwei, L. (2024). LMMs-Eval: Accelerating the Development of Large Multimoal Models. Zenodo. https:\/\/github.com\/EvolvingLMMs-Lab\/lmms-eval."},{"key":"2865_CR2","doi-asserted-by":"crossref","unstructured":"Cai, Y., Zhang, J., He, H., He, X., Tong, A., Gan, Z., Wang, C., Xue, Z., Liu, Y., & Bai, X. (2025). Llava-kd: A framework of distilling multimodal large language models. Proceedings of the IEEE\/CVF international conference on computer vision, (pp. 239\u2013249).","DOI":"10.1109\/ICCV51701.2025.00030"},{"key":"2865_CR3","unstructured":"Chen, X., Hu, Y., Zhang, J., Wang, Y., Li, C., & Chen, H. (2024). Streamlining redundant layers to compress large language models. The thirteenth international conference on learning representations."},{"key":"2865_CR4","doi-asserted-by":"crossref","unstructured":"Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Zhong, M., Zhang, Q., Zhu, X., Lu, L., and others. (2024). Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, (pp. 24185\u201324198)","DOI":"10.1109\/CVPR52733.2024.02283"},{"key":"2865_CR5","unstructured":"Chu, X., Qiao, L., Lin, X., Xu, S., Yang, Y., Hu, Y., Wei, F., Zhang, X., Zhang, B., Wei, X., and others. (2023). Mobilevlm: A fast, reproducible and strong vision language assistant for mobile devices. arXiv:2312.16886"},{"key":"2865_CR6","unstructured":"Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., and others. (2021). Training verifiers to solve math word problems. arXiv:2110.14168."},{"key":"2865_CR7","unstructured":"Dery, L., Kolawole, S., Kagy, J.-F., Smith, V., Neubig, G., & Talwalkar, A. (2024). Everybody prune now: Structured pruning of LLMs with only forward passes. arXiv:2402.05406."},{"key":"2865_CR8","doi-asserted-by":"publisher","first-page":"30318","DOI":"10.52202\/068431-2198","volume":"35","author":"T Dettmers","year":"2022","unstructured":"Dettmers, T., Lewis, M., Belkada, Y., & Zettlemoyer, L. (2022). Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale. Advances in Neural Information Processing Systems, 35, 30318\u201330332.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2865_CR9","unstructured":"Dong, X., Chen, S., & Pan, S.J. (2017). Learning to prune deep neural networks via layer-wise optimal brain surgeon. arXiv:1705.07565."},{"key":"2865_CR10","unstructured":"Fan, A., Grave, E., & Joulin, A. (2019). Reducing transformer depth on demand with structured dropout. arXiv:1909.11556."},{"key":"2865_CR11","doi-asserted-by":"crossref","unstructured":"Fang, G., Ma, X., Song, M., Mi, M.B., & Wang, X. (2023). Depgraph: Towards any structural pruning. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, (pp. 16091\u201316101) .","DOI":"10.1109\/CVPR52729.2023.01544"},{"key":"2865_CR12","doi-asserted-by":"crossref","unstructured":"Farina, M., Mancini, M., Cunegatti, E., Liu, G., Iacca, G., & Ricci, E. (2024). Multiflow: Shifting towards task-agnostic vision-language pruning. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, (pp. 16185\u201316195).","DOI":"10.1109\/CVPR52733.2024.01532"},{"key":"2865_CR13","unstructured":"Frankle, J., & Carbin, M. (2019). The lottery ticket hypothesis: Finding sparse, trainable neural networks. arXiv:1803.03635."},{"key":"2865_CR14","unstructured":"Frantar, E., & Alistarh, D. (2023). Sparsegpt: Massive language models can be accurately pruned in one-shot. International conference on machine learning, PMLR. (pp. 10323\u201310337)."},{"issue":"6","key":"2865_CR15","doi-asserted-by":"publisher","first-page":"1789","DOI":"10.1007\/s11263-021-01453-z","volume":"129","author":"J Gou","year":"2021","unstructured":"Gou, J., Yu, B., Maybank, S. J., & Tao, D. (2021). Knowledge distillation: A survey. International Journal of Computer Vision, 129(6), 1789\u20131819. https:\/\/doi.org\/10.1007\/s11263-021-01453-z","journal-title":"International Journal of Computer Vision"},{"key":"2865_CR16","unstructured":"Gu, Y., Dong, L., Wei, F., & Huang, M. (2023). Knowledge distillation of large language models. arXiv:2306.08543."},{"key":"2865_CR17","unstructured":"He, M., Liu, Y., Wu, B., Yuan, J., Wang, Y., Huang, T., & Zhao, B. (2024). Efficient multimodal learning from data-centric perspective. arXiv:2402.11530."},{"key":"2865_CR18","unstructured":"Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the knowledge in a neural network. arXiv:1503.02531."},{"key":"2865_CR19","first-page":"1","volume":"23","author":"T Hoefler","year":"2021","unstructured":"Hoefler, T., Alistarh, D., Ben-Nun, T., Dryden, N., & Peste, A. (2021). Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks. Journal of Machine Learning Research, 23, 1\u2013124.","journal-title":"Journal of Machine Learning Research"},{"key":"2865_CR20","unstructured":"Hoffmann, J., Agnihotri, S., Saikia, T., & Brox, T. (2021). Towards improving robustness of compressed cnns. ICML workshop on uncertainty and robustness in deep learning (UDL), vol. 4."},{"key":"2865_CR21","unstructured":"Holtzman, A., Buys, J., Du, L., Forbes, M., & Choi, Y. (2019). The curious case of neural text degeneration. arXiv:1904.09751."},{"key":"2865_CR22","doi-asserted-by":"crossref","unstructured":"Hsieh, C.-Y., Li, C.-L., Yeh, C.-K., Nakhost, H., Fujii, Y., Ratner, A., Krishna, R., Lee, C.-Y., & Pfister, T. (2023). Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes. arXiv:2305.02301","DOI":"10.18653\/v1\/2023.findings-acl.507"},{"key":"2865_CR23","unstructured":"Hu, E.J., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., and others. (2021). Lora: Low-rank adaptation of large language models. International conference on learning representations."},{"key":"2865_CR24","doi-asserted-by":"crossref","unstructured":"Huang, Y., Thede, L., Mancini, M., Xu, W., & Akata, Z. (2025). Investigating structural pruning and recovery techniques for compressing multimodal large language models: An empirical study. DAGM german conference on pattern recognition, Springer. (pp. 320\u2013336).","DOI":"10.1007\/978-3-032-12840-9_21"},{"key":"2865_CR25","doi-asserted-by":"crossref","unstructured":"Hudson, D.A., & Manning, C.D. (2019). Gqa: A new dataset for real-world visual reasoning and compositional question answering. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, (pp. 6700\u20136709).","DOI":"10.1109\/CVPR.2019.00686"},{"key":"2865_CR26","doi-asserted-by":"crossref","unstructured":"Jiang, F. (2024). Identifying and mitigating vulnerabilities in llm-integrated applications. Master\u2019s thesis, University of Washington.","DOI":"10.1145\/3634737.3659433"},{"key":"2865_CR27","doi-asserted-by":"crossref","unstructured":"Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., & Liu, Q. (2020). TinyBERT: Distilling BERT for natural language understanding. arXiv:1909.10351.","DOI":"10.18653\/v1\/2020.findings-emnlp.372"},{"key":"2865_CR28","unstructured":"Karamcheti, S., Nair, S., Balakrishna, A., Liang, P., Kollar, T., & Sadigh, D. (2024). Prismatic vlms: Investigating the design space of visually-conditioned language models. arXiv:2402.07865."},{"key":"2865_CR29","doi-asserted-by":"crossref","unstructured":"Kembhavi, A., Salvato, M., Kolve, E., Seo, M., Hajishirzi, H., & Farhadi, A. (2016). A diagram is worth a dozen images. European conference on computer vision, Springer. (pp. 235\u2013251).","DOI":"10.1007\/978-3-319-46493-0_15"},{"key":"2865_CR30","unstructured":"Kim, J., Kim, K., Seo, S., & Park, C. (2025). Compodistill: Attention distillation for compositional reasoning in multimodal llms. arXiv:2510.12184."},{"key":"2865_CR31","unstructured":"Lee, N., Ajanthan, T., Gould, S., & Torr, P.H.S. (2020). A signal propagation perspective for pruning neural networks at initialization. arXiv:1906.06307."},{"key":"2865_CR32","doi-asserted-by":"crossref","unstructured":"Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W.X., & Wen, J.-R. (2023). Evaluating object hallucination in large vision-language models. arXiv:2305.10355.","DOI":"10.18653\/v1\/2023.emnlp-main.20"},{"key":"2865_CR33","unstructured":"Li, H., Kadav, A., Durdanovic, I., Samet, H., & Graf, H.P. (2017). Pruning filters for efficient convnets. International conference on learning representations."},{"key":"2865_CR34","unstructured":"Liang, K.J., Hao, W., Shen, D., Zhou, Y., Chen, W., Chen, C., & Carin, L. (2021). MixKD: Towards efficient distillation of large-scale language models. arXiv:2011.00593."},{"key":"2865_CR35","doi-asserted-by":"crossref","unstructured":"Liu, H., Li, C., Li, Y. & Lee, Y.J. (2024). Improved baselines with visual instruction tuning. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, (pp. 26296\u201326306).","DOI":"10.1109\/CVPR52733.2024.02484"},{"key":"2865_CR36","unstructured":"Liu, L., Zhang, S., Kuang, Z., Zhou, A., Xue, J.-H., Wang, X., Chen, Y., Yang, W., Liao, Q., & Zhang, W. (2021). Group fisher pruning for practical network compression. arXiv:2108.00708"},{"key":"2865_CR37","unstructured":"Liu, Z., Zhao, C., Iandola, F., Lai, C., Tian, Y., Fedorov, I., Xiong, Y., Chang, E., Shi, Y., Krishnamoorthi, R., and others. (2024). Mobilellm: Optimizing sub-billion parameter language models for on-device use cases. Forty-first international conference on machine learning."},{"key":"2865_CR38","unstructured":"Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., & Gao, J. (2023). Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. arXiv:2310.02255"},{"key":"2865_CR39","doi-asserted-by":"publisher","first-page":"2507","DOI":"10.52202\/068431-0182","volume":"35","author":"P Lu","year":"2022","unstructured":"Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.-W., Zhu, S.-C., Tafjord, O., Clark, P., & Kalyan, A. (2022). Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems, 35, 2507\u20132521.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2865_CR40","doi-asserted-by":"publisher","first-page":"21702","DOI":"10.52202\/075280-0950","volume":"36","author":"X Ma","year":"2023","unstructured":"Ma, X., Fang, G., & Wang, X. (2023). Llm-pruner: On the structural pruning of large language models. Advances in neural information processing systems, 36, 21702\u201321720.","journal-title":"Advances in neural information processing systems"},{"key":"2865_CR41","doi-asserted-by":"crossref","unstructured":"Mathew, M., Karatzas, D., & Jawahar, C.(2021). Docvqa: A dataset for vqa on document images. Proceedings of the IEEE\/CVF winter conference on applications of computer vision, (pp. 2200\u20132209).","DOI":"10.1109\/WACV48630.2021.00225"},{"key":"2865_CR42","unstructured":"McCarley, J., Chakravarti, R., & Sil, A. (2019). Structured pruning of a bert-based question answering model. arXiv:1910.06360."},{"key":"2865_CR43","doi-asserted-by":"crossref","unstructured":"Men, X., Xu, M., Zhang, Q., Wang, B., Lin, H., Lu, Y., Han, X., & Chen, W. (2024). Shortgpt: Layers in large language models are more redundant than you expect. arXiv:2403.03853","DOI":"10.18653\/v1\/2025.findings-acl.1035"},{"key":"2865_CR44","unstructured":"Michel, P., Levy, O., & Neubig, G. (2019). Are sixteen heads really better than one? Advances in neural information processing systems,32."},{"key":"2865_CR45","doi-asserted-by":"publisher","first-page":"41076","DOI":"10.52202\/079017-1299","volume":"37","author":"S Muralidharan","year":"2024","unstructured":"Muralidharan, S., Turuvekere Sreenivas, S., Joshi, R., Chochowski, M., Patwary, M., Shoeybi, M., Catanzaro, B., Kautz, J., & Molchanov, P. (2024). Compact language models via pruning and knowledge distillation. Advances in Neural Information Processing Systems, 37, 41076\u201341102.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2865_CR46","unstructured":"Park, S., Lee, J., Mo, S., & Shin, J. (2020). Lookahead: A far-sighted alternative of magnitude-based pruning. arXiv:2002.04809."},{"key":"2865_CR47","unstructured":"Popp, N., Metzen, J.H., & Hein, M. (2024). Zero-shot distillation for image encoders: How to make effective use of synthetic data. arXiv:2404.16637."},{"key":"2865_CR48","doi-asserted-by":"publisher","first-page":"101429","DOI":"10.1016\/j.csl.2022.101429","volume":"77","author":"H Sajjad","year":"2023","unstructured":"Sajjad, H., Dalvi, F., Durrani, N., & Nakov, P. (2023). On the effect of dropping layers of pre-trained transformer models. Computer Speech & Language, 77, 101429.","journal-title":"Computer Speech & Language"},{"key":"2865_CR49","unstructured":"Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2020). DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter . arXiv:1910.01108."},{"key":"2865_CR50","unstructured":"Sanh, V., Wolf, T., & Rush, A.M. (2020). Movement pruning: Adaptive sparsity by fine-tuning. arXiv:2005.07683."},{"key":"2865_CR51","doi-asserted-by":"crossref","unstructured":"Shang, Y., Cai, M., Xu, B., Lee, Y.J., & Yan, Y. (2025). Llava-prumerge: Adaptive token reduction for efficient large multimodal models. Proceedings of the IEEE\/CVF international conference on computer vision, (pp. 22857\u201322867).","DOI":"10.1109\/ICCV51701.2025.02122"},{"key":"2865_CR52","unstructured":"Song, J., Oh, K., Kim, T., Kim, H., Kim, Y., and others. (2024). Sleb: Streamlining llms through redundancy verification and elimination of transformer blocks. Forty-first international conference on machine learning."},{"key":"2865_CR53","unstructured":"Sreenivas, S.T., Muralidharan, S., Joshi, R., Chochowski, M., Mahabaleshwarkar, A.S., Shen, G., Zeng, J., Chen, Z., Suhara, Y., Diao, S., and others. (2024). Llm pruning and distillation in practice: The minitron approach. arXiv:2408.11796."},{"key":"2865_CR54","doi-asserted-by":"crossref","unstructured":"Sun, S., Cheng, Y., Gan, Z., & Liu, J. (2019). Patient knowledge distillation for BERT model compression. arXiv:1908.09355.","DOI":"10.18653\/v1\/D19-1441"},{"key":"2865_CR55","unstructured":"Sun, M., Liu, Z., Bair, A., & Kolter, J.Z. (2024). A simple and effective pruning approach for large language models. The twelfth international conference on learning representations."},{"key":"2865_CR56","doi-asserted-by":"crossref","unstructured":"Tang, Z., Ma, Z., Wang, S., Li, Z., Zhang, L., Zhao, H., Li, Y., & Wang, Q. (2025). Covipal: Layer-wise contextualized visual token pruning for large vision-language models. Findings of the association for computational linguistics: EMNLP 2025, (pp. 20701\u201320714).","DOI":"10.18653\/v1\/2025.findings-emnlp.1127"},{"key":"2865_CR57","unstructured":"Telgarsky, M. (2016). Benefits of depth in neural networks. Conference on learning theory (COLT), (pp. 1517\u20131539)."},{"key":"2865_CR58","unstructured":"Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozi\u00e8re, B., Goyal, N., Hambro, E., Azhar, F., and others. (2023). Llama: Open and efficient foundation language models. arXiv:2302.13971"},{"key":"2865_CR59","doi-asserted-by":"crossref","unstructured":"Voita, E., Talbot, D., Moiseev, F., Sennrich, R., & Titov, I. (2019). Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. Proceedings of the 57th annual meeting of the association for computational linguistics, (pp. 5797\u20135808).","DOI":"10.18653\/v1\/P19-1580"},{"key":"2865_CR60","doi-asserted-by":"crossref","unstructured":"Wang, W., Bao, H., Huang, S., Dong, L., & Wei, F. (2021). MiniLMv2: Multi-head self-attention relation distillation for compressing pretrained transformers. arXiv:2012.15828.","DOI":"10.18653\/v1\/2021.findings-acl.188"},{"key":"2865_CR61","unstructured":"Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., & Zhou, M. (2020). MiniLM: Deep self-attention distillation for task-agnostic compression of pre-trained transformers. arXiv:2002.10957."},{"key":"2865_CR62","unstructured":"Wang, H., Xu, Y., Xu, Z., Gao, J., Liu, Y., Hu, W., Wang, K., & Zhang, Z. (2025). Autoprune: Each complexity deserves a pruning policy. arXiv:2509.23931."},{"key":"2865_CR63","unstructured":"Xia, M., Gao, T., Zeng, Z., & Chen, D. (2024). Sheared LLaMA: Accelerating language model pre-training via structured pruning. arXiv:2310.06694."},{"key":"2865_CR64","doi-asserted-by":"crossref","unstructured":"Yang, C., An, Z., Huang, L., Bi, J., Yu, X., Yang, H., Diao, B., & Xu, Y. (2024). Clip-kd: An empirical study of clip model distillation. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, (pp. 15952\u201315962).","DOI":"10.1109\/CVPR52733.2024.01510"},{"key":"2865_CR65","doi-asserted-by":"crossref","unstructured":"Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., & Chen, E. (2023). A survey on multimodal large language models. arXiv:2306.13549.","DOI":"10.1093\/nsr\/nwae403"},{"key":"2865_CR66","unstructured":"You, Z., Yan, K., Ye, J., Ma, M., & Wang, P. (2019). Gate decorator: Global filter pruning method for accelerating deep convolutional neural networks. arXiv:1909.08174."},{"key":"2865_CR67","doi-asserted-by":"crossref","unstructured":"Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., Wei, C., Yu, B., Yuan, R., Sun, R., Yin, M., Zheng, B., Yang, Z., Liu, Y., Huang, W., Sun, H., Su, Y., & Chen, W. (2024). Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. Proceedings of CVPR.","DOI":"10.1109\/CVPR52733.2024.00913"},{"issue":"67","key":"2865_CR68","first-page":"1","volume":"23","author":"C Zhang","year":"2022","unstructured":"Zhang, C., Bengio, S., & Singer, Y. (2022). Are all layers created equal? Journal of Machine Learning Research, 23(67), 1\u201328.","journal-title":"Journal of Machine Learning Research"},{"key":"2865_CR69","unstructured":"Zhou, A., Ma, Y., Zhu, J., Liu, J., Zhang, Z., Yuan, K., Sun, W., & Li, H. (2021). Learning n: M fine-grained structured sparse neural networks from scratch. International conference on learning representations."},{"key":"2865_CR70","unstructured":"Zhu, M., Zhu, Y., Liu, X., Liu, N., Xu, Z., Shen, C., Peng, Y., Ou, Z., Feng, F., & Tang, J. (2024). A comprehensive overhaul of multimodal assistant with small language models. arXiv:2403.06199"},{"key":"2865_CR71","doi-asserted-by":"crossref","unstructured":"Zhu, Y., Zhu, M., Liu, N., Xu, Z., & Peng, Y. (2024). Llava-phi: Efficient multi-modal assistant with small language model. Proceedings of the 1st international workshop on efficient multimedia computing under limited, (pp. 18\u201322).","DOI":"10.1145\/3688863.3689575"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-026-02865-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-026-02865-5","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-026-02865-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T16:16:32Z","timestamp":1784564192000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-026-02865-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6]]},"references-count":71,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2026,6]]}},"alternative-id":["2865"],"URL":"https:\/\/doi.org\/10.1007\/s11263-026-02865-5","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6]]},"assertion":[{"value":"15 January 2026","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 April 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 June 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"313"}}