{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T23:13:07Z","timestamp":1780614787206,"version":"3.54.1"},"reference-count":128,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2025,8,28]],"date-time":"2025-08-28T00:00:00Z","timestamp":1756339200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>With the exponential growth of multimodal data, the limitations of traditional unimodal models in cross-modal understanding and complex scenario reasoning have become increasingly evident. Built upon the foundation of Large Language Models (LLMs), Multimodal Large Language Models (MLLMs) retain strong reasoning abilities and demonstrate unique capabilities in multimodal understanding. This survey provides a comprehensive overview of the current research landscape of MLLMs. It systematically analyzes mainstream model architectures, training, fine-tuning strategies, and task classifications, while offering a structured account of evaluation methodologies. Beyond synthesis, the paper highlights emerging trends that aim for balanced integration across modalities, tasks, and components, and critically examines key challenges together with potential solutions. The survey specifically emphasizes recent reasoning-oriented MLLMs, with a focus on DeepSeek-R1, analyzing their design paradigms and contributions from the perspective of symmetric reasoning capabilities. Overall, this work offers a comprehensive overview of cutting-edge advancements and lays a foundation for the future development of MLLMs, especially those guided by symmetry principles.<\/jats:p>","DOI":"10.3390\/sym17091400","type":"journal-article","created":{"date-parts":[[2025,8,28]],"date-time":"2025-08-28T07:43:16Z","timestamp":1756366996000},"page":"1400","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Symmetry-Aware Advances in Multimodal Large Language Models: Architectures, Training, and Evaluation"],"prefix":"10.3390","volume":"17","author":[{"given":"Xinran","family":"Liu","sequence":"first","affiliation":[{"name":"Department of Computer Science and Technology, Sichuan University, Chengdu 610207, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haojie","family":"Liu","sequence":"additional","affiliation":[{"name":"College of Control Science and Engineering, Zhejiang University, Hangzhou 310027, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,8,28]]},"reference":[{"key":"ref_1","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 1). Learning Transferable Visual Models from Natural Language Supervision. Proceedings of the 38th International Conference on Machine Learning, PMLR, Vienna, Austria."},{"key":"ref_2","first-page":"23716","article-title":"Flamingo: A Visual Language Model for Few-Shot Learning","volume":"35","author":"Alayrac","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_3","unstructured":"Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q., Sung, Y.-H., Li, Z., and Duerig, T. (2021, January 18\u201324). Scaling up Visual and Vision-Language Representation Learning with Noisy Text Supervision. Proceedings of the International Conference on Machine Learning, PMLR, Vienna, Austria."},{"key":"ref_4","unstructured":"Wang, P., Yang, A., Men, R., Lin, J., Bai, S., Li, Z., Ma, J., Zhou, C., Zhou, J., and Yang, H. (2022, January 25\u201327). Ofa: Unifying Architectures, Tasks, and Modalities through a Simple Sequence-to-Sequence Learning Framework. Proceedings of the International Conference on Machine Learning, PMLR, Baltimore, MD, USA."},{"key":"ref_5","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Attention Is All You Need. arXiv."},{"key":"ref_6","first-page":"32897","article-title":"Vlmo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts","volume":"35","author":"Bao","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_7","unstructured":"Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., and Bi, X. (2025). Deepseek-R1: Incentivizing Reasoning Capability in Llms via Reinforcement Learning. arXiv."},{"key":"ref_8","unstructured":"Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., Zhao, C., Deng, C., Zhang, C., and Ruan, C. (2024). DeepSeek-V3 Technical Report. arXiv."},{"key":"ref_9","unstructured":"Zhang, B., Li, K., Cheng, Z., Hu, Z., Yuan, Y., Chen, G., Leng, S., Jiang, Y., Zhang, H., and Li, X. (2025). VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding 2025. arXiv."},{"key":"ref_10","unstructured":"Fu, C., Lin, H., Wang, X., Zhang, Y.-F., Shen, Y., Liu, X., Cao, H., Long, Z., Gao, H., and Li, K. (2025). VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction. arXiv."},{"key":"ref_11","unstructured":"Kim, W., Son, B., and Kim, I. (2021, January 18\u201324). ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision. Proceedings of the 38th International Conference on Machine Learning (ICML 2021), Virtual Event."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"76","DOI":"10.1109\/MNET.2024.3421306","article-title":"Mobile-Llama: Instruction Fine-Tuning Open-Source Llm for Network Analysis in 5g Networks","volume":"38","author":"Kan","year":"2024","journal-title":"IEEE Netw."},{"key":"ref_13","first-page":"34892","article-title":"Visual Instruction Tuning","volume":"36","author":"Liu","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Yuan, Y., Li, W., Liu, J., Tang, D., Luo, X., Qin, C., Zhang, L., and Zhu, J. (2024, January 16\u201322). Osprey: Pixel Understanding with Visual Instruction Tuning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.02664"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Cherti, M., Beaumont, R., Wightman, R., Wortsman, M., Ilharco, G., Gordon, C., Schuhmann, C., Schmidt, L., and Jitsev, J. (2023, January 17\u201324). Reproducible Scaling Laws for Contrastive Language-Image Learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00276"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Li, Z., Yang, B., Liu, Q., Ma, Z., Zhang, S., Yang, J., Sun, Y., Liu, Y., and Bai, X. (2024, January 16\u201322). Monkey: Image Resolution and Text Label Are Important Things for Large Multi-Modal Models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.02527"},{"key":"ref_17","first-page":"18090","article-title":"Pengi: An Audio Language Model for Audio Tasks","volume":"36","author":"Deshmukh","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Elizalde, B., Deshmukh, S., Al Ismail, M., and Wang, H. (2023, January 4\u201310). Clap Learning Audio Concepts from Natural Language Supervision. Proceedings of the ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes, Greece.","DOI":"10.1109\/ICASSP49357.2023.10095889"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Girdhar, R., El-Nouby, A., Liu, Z., Singh, M., Alwala, K.V., Joulin, A., and Misra, I. (2023, January 17\u201324). Imagebind: One Embedding Space to Bind Them All. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01457"},{"key":"ref_20","first-page":"1877","article-title":"Language Models Are Few-Shot Learners","volume":"33","author":"Brown","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_21","first-page":"1","article-title":"Scaling Instruction-Finetuned Language Models","volume":"25","author":"Chung","year":"2024","journal-title":"J. Mach. Learn. Res."},{"key":"ref_22","unstructured":"Li, J., Li, D., Savarese, S., and Hoi, S. (2023, January 23\u201329). Blip-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models. Proceedings of the International Conference on Machine Learning, PMLR, Honolulu, HI, USA."},{"key":"ref_23","unstructured":"Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozi\u00e8re, B., Goyal, N., Hambro, E., and Azhar, F. (2023). Llama: Open and Efficient Foundation Language Models. arXiv."},{"key":"ref_24","unstructured":"Gu, A., and Dao, T. (2023). Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv."},{"key":"ref_25","unstructured":"Zhang, R., Han, J., Liu, C., Zhou, A., Lu, P., Qiao, Y., Li, H., and Gao, P. (2024, January 7\u201311). LLaMA-Adapter: Efficient Fine-Tuning of Large Language Models with Zero-Initialized Attention. Proceedings of the Twelfth International Conference on Learning Representations, Vienna, Austria."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Liu, H., Li, C., Li, Y., and Lee, Y.J. (2024, January 16\u201322). Improved Baselines with Visual Instruction Tuning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.02484"},{"key":"ref_27","first-page":"121475","article-title":"Cogvlm: Visual Expert for Pretrained Language Models","volume":"37","author":"Wang","year":"2024","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Chen, S., Chen, X., Zhang, C., Li, M., Yu, G., Fei, H., Zhu, H., Fan, J., and Chen, T. (2024, January 17\u201321). Ll3da: Visual Interactive Instruction Tuning for Omni-3d Understanding Reasoning and Planning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.02496"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Lin, J., Yin, H., Ping, W., Molchanov, P., Shoeybi, M., and Han, S. (2024, January 17\u201321). VILA: On Pre-training for Visual Language Models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.02520"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Chen, G., Shen, L., Shao, R., Deng, X., and Nie, L. (2024, January 17\u201321). Lion: Empowering Multimodal Large Language Model with Dual-Level Visual Knowledge. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.02506"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Hu, W., Xu, Y., Li, Y., Li, W., Chen, Z., and Tu, Z. (2024, January 20\u201327). Bliva: A Simple Multimodal Llm for Better Handling of Text-Rich Visual Questions. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2024), Vancouver, BC, Canada.","DOI":"10.1609\/aaai.v38i3.27999"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Ye, Q., Xu, H., Ye, J., Yan, M., Hu, A., Liu, H., Qian, Q., Zhang, J., and Huang, F. (2024, January 17\u201321). mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01239"},{"key":"ref_33","unstructured":"Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., and Ge, W. (2024). Qwen2-VL: Enhancing Vision-Language Model\u2019s Perception of the World at Any Resolution. arXiv."},{"key":"ref_34","unstructured":"Jaegle, A., Gimeno, F., Brock, A., Vinyals, O., Zisserman, A., and Carreira, J. (2021, January 18\u201324). Perceiver: General Perception with Iterative Attention. Proceedings of the International Conference on Machine Learning (ICML 2021), Virtual Event."},{"key":"ref_35","unstructured":"Chen, D., Liu, J., Dai, W., and Wang, B. (2024, January 20\u201327). Visual Instruction Tuning with Polite Flamingo. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2024), Vancouver, BC, Canada."},{"key":"ref_36","first-page":"71683","article-title":"Obelics: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents","volume":"36","author":"Saulnier","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_37","unstructured":"Li, K., He, Y., Wang, Y., Li, Y., Wang, W., Luo, P., Wang, Y., Wang, L., and Qiao, Y. (2024). VideoChat: Chat-Centric Video Understanding. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Van Gansbeke, W., Vandenhende, S., Georgoulis, S., Proesmans, M., and Van Gool, L. (2020). SCAN: Learning to Classify Images Without Labels. Computer Vision\u2014ECCV 2020, Springer International Publishing. Lecture Notes in Computer Science.","DOI":"10.1007\/978-3-030-58607-2_16"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Sharma, P., Ding, N., Goodman, S., and Soricut, R. (2018, January 15\u201320). Conceptual Captions: A Cleaned, Hypernymed, Image Alt-Text Dataset For Automatic Image Captioning. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Melbourne, Australia.","DOI":"10.18653\/v1\/P18-1238"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Perez, E., Strub, F., de Vries, H., Dumoulin, V., and Courville, A. (2018, January 2\u20137). FiLM: Visual Reasoning with a General Conditioning Layer. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2018), New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11671"},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"104","DOI":"10.1007\/978-3-030-58577-8_7","article-title":"UNITER: UNiversal Image-TExt Representation Learning","volume":"Volume 12375","author":"Vedaldi","year":"2020","journal-title":"Computer Vision\u2014ECCV 2020"},{"key":"ref_42","unstructured":"Huang, Z., Zeng, Z., Liu, B., Fu, D., and Fu, J. (2020). Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers. arXiv."},{"key":"ref_43","unstructured":"Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., and Shi, Y. (2024). mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality. arXiv."},{"key":"ref_44","unstructured":"Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M. (2023). MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. arXiv."},{"key":"ref_45","first-page":"49250","article-title":"Instructblip: Towards General-Purpose Vision-Language Models with Instruction Tuning","volume":"36","author":"Dai","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"7543","DOI":"10.1109\/TPAMI.2025.3571946","article-title":"Otter: A Multi-Modal Model with in-Context Instruction Tuning","volume":"47","author":"Li","year":"2025","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_47","unstructured":"Chen, K., Zhang, Z., Zeng, W., Zhang, R., Zhu, F., and Zhao, R. (2023). Shikra: Unleashing Multimodal LLM\u2019s Referential Dialogue Magic. arXiv."},{"key":"ref_48","unstructured":"Chen, J., Zhu, D., Shen, X., Li, X., Liu, Z., Zhang, P., Krishnamoorthi, R., Chandra, V., Xiong, Y., and Elhoseiny, M. (2023). MiniGPT-v2: Large Language Model as a Unified Interface for Vision-Language Multi-Task Learning. arXiv."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Chen, X., Djolonga, J., Padlewski, P., Mustafa, B., Changpinyo, S., Wu, J., Ruiz, C.R., Goodman, S., Wang, X., and Tay, Y. (2023). PaLI-X: On Scaling up a Multilingual Vision and Language Model. arXiv.","DOI":"10.1109\/CVPR52733.2024.01368"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Zhao, L., Yu, E., Ge, Z., Yang, J., Wei, H., Zhou, H., Sun, J., Peng, Y., Dong, R., and Han, C. (2023). ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning. arXiv.","DOI":"10.24963\/ijcai.2024\/193"},{"key":"ref_51","unstructured":"Zheng, K., He, X., and Wang, X.E. (2023). MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens. arXiv."},{"key":"ref_52","unstructured":"Jin, Y., Xu, K., Xu, K., Chen, L., Liao, C., Tan, J., Huang, Q., Chen, B., Lei, C., and Liu, A. (2023). Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization. arXiv."},{"key":"ref_53","first-page":"21487","article-title":"Generating Images with Multimodal Language Models","volume":"36","author":"Koh","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_54","first-page":"28541","article-title":"Llava-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day","volume":"36","author":"Li","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_55","first-page":"61501","article-title":"Visionllm: Large Language Model Is Also an Open-Ended Decoder for Vision-Centric Tasks","volume":"36","author":"Wang","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_56","unstructured":"Koh, J.Y., Salakhutdinov, R., and Fried, D. (2023, January 23\u201329). Grounding Language Models to Images for Multimodal Inputs and Outputs. Proceedings of the International Conference on Machine Learning (ICML 2023), Honolulu, HI, USA."},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O.K., Singhal, S., and Som, S. (2023, January 18\u201322). Image as a Foreign Language: BEiT Pretraining for Vision and Vision-Language Tasks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2023), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01838"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Chen, L., Li, J., Dong, X., Zhang, P., He, C., Wang, J., Zhao, F., and Lin, D. (2025). ShareGPT4V: Improving Large Multi-Modal Models with Better Captions. Lecture Notes in Computer Science, Springer Nature Switzerland.","DOI":"10.1007\/978-3-031-72643-9_22"},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., and Parikh, D. (2015, January 7\u201313). VQA: Visual Question Answering. Proceedings of the IEEE International Conference on Computer Vision (ICCV 2015), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.279"},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Karpathy, A., and Fei-Fei, L. (2015, January 7\u201312). Deep Visual-Semantic Alignments for Generating Image Descriptions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2015), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298932"},{"key":"ref_61","unstructured":"Huynh, N.D., Bouadjenek, M.R., Razzak, I., Hacid, H., and Aryal, S. (2025). SVLA: A Unified Speech-Vision-Language Assistant with Multimodal Reasoning and Speech Generation. arXiv."},{"key":"ref_62","doi-asserted-by":"crossref","unstructured":"Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N.A., Khashabi, D., and Hajishirzi, H. (2023). Self-Instruct: Aligning Language Models with Self-Generated Instructions. arXiv.","DOI":"10.18653\/v1\/2023.acl-long.754"},{"key":"ref_63","doi-asserted-by":"crossref","unstructured":"Zeng, Y., Zhang, H., Zheng, J., Xia, J., Wei, G., Wei, Y., Zhang, Y., Kong, T., and Song, R. (2024, January 16\u201321). What Matters in Training a Gpt4-Style Language Model with Multimodal Inputs?. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT 2024), Mexico City, Mexico.","DOI":"10.18653\/v1\/2024.naacl-long.440"},{"key":"ref_64","unstructured":"Yue, T., Guo, L., Tang, Y., Zhao, Z., Zhu, X., Huang, H., and Liu, J. (2025). LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation. arXiv."},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Chen, P.-Y. (2024, January 20\u201327). Model Reprogramming: Resource-Efficient Cross-Domain Machine Learning. Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2024), Honolulu, HI, USA.","DOI":"10.1609\/aaai.v38i20.30267"},{"key":"ref_66","doi-asserted-by":"crossref","first-page":"208","DOI":"10.1016\/j.aiopen.2023.08.012","article-title":"GPT Understands, Too","volume":"5","author":"Liu","year":"2024","journal-title":"AI Open"},{"key":"ref_67","first-page":"3","article-title":"Lora: Low-Rank Adaptation of Large Language Models","volume":"1","author":"Hu","year":"2022","journal-title":"ICLR"},{"key":"ref_68","first-page":"10088","article-title":"Qlora: Efficient Finetuning of Quantized Llms","volume":"36","author":"Dettmers","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_69","doi-asserted-by":"crossref","unstructured":"Sun, Z., Shen, S., Cao, S., Liu, H., Li, C., Shen, Y., Gan, C., Gui, L.-Y., Wang, Y.-X., and Yang, Y. (2023). Aligning Large Multimodal Models with Factually Augmented RLHF. arXiv.","DOI":"10.18653\/v1\/2024.findings-acl.775"},{"key":"ref_70","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An Image Is Worth 16 \u00d7 16 Words: Transformers for Image Recognition at Scale. arXiv."},{"key":"ref_71","doi-asserted-by":"crossref","unstructured":"Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R. (2019, January 16\u201320). OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2019), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00331"},{"key":"ref_72","doi-asserted-by":"crossref","unstructured":"Zellers, R., Bisk, Y., Farhadi, A., and Choi, Y. (2019, January 16\u201320). From Recognition to Cognition: Visual Commonsense Reasoning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2019), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00688"},{"key":"ref_73","doi-asserted-by":"crossref","unstructured":"Masry, A., Long, D.X., Tan, J.Q., Joty, S., and Hoque, E. (2022). ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning. arXiv.","DOI":"10.18653\/v1\/2022.findings-acl.177"},{"key":"ref_74","first-page":"216","article-title":"MMBench: Is Your Multi-Modal Model an All-Around Player?","volume":"Volume 15064","author":"Leonardis","year":"2025","journal-title":"Computer Vision\u2014ECCV 2024"},{"key":"ref_75","first-page":"27056","article-title":"Are We on the Right Way for Evaluating Large Vision-Language Models?","volume":"37","author":"Chen","year":"2024","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_76","doi-asserted-by":"crossref","unstructured":"Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., and Sun, Y. (2024, January 17\u201321). MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert Agi. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00913"},{"key":"ref_77","unstructured":"Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J. (2024). MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts. arXiv."},{"key":"ref_78","unstructured":"Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., and Wang, L. (2024). MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities. arXiv."},{"key":"ref_79","doi-asserted-by":"crossref","unstructured":"He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. (2020, January 14\u201319). Momentum Contrast for Unsupervised Visual Representation Learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2020), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"ref_80","doi-asserted-by":"crossref","unstructured":"Cho, J., Zala, A., and Bansal, M. (2023, January 2\u20136). Dall-Eval: Probing the Reasoning Skills and Social Biases of Text-to-Image Generation Models. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV 2023), Paris, France.","DOI":"10.1109\/ICCV51070.2023.00283"},{"key":"ref_81","doi-asserted-by":"crossref","unstructured":"Bakr, E.M., Sun, P., Shen, X., Khan, F.F., Li, L.E., and Elhoseiny, M. (2023, January 2\u20136). Hrs-Bench: Holistic, Reliable and Scalable Benchmark for Text-to-Image Models. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV 2023), Paris, France.","DOI":"10.1109\/ICCV51070.2023.01834"},{"key":"ref_82","first-page":"52132","article-title":"Geneval: An Object-Focused Framework for Evaluating Text-to-Image Alignment","volume":"36","author":"Ghosh","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_83","unstructured":"Hu, X., Wang, R., Fang, Y., Fu, B., Cheng, P., and Yu, G. (2024). ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment. arXiv."},{"key":"ref_84","doi-asserted-by":"crossref","unstructured":"Wang, S., Saharia, C., Montgomery, C., Pont-Tuset, J., Noy, S., Pellegrini, S., Onoe, Y., Laszlo, S., Fleet, D.J., and Soricut, R. (2023, January 18\u201322). Imagen Editor and Editbench: Advancing and Evaluating Text-Guided Image Inpainting. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2023), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01761"},{"key":"ref_85","first-page":"86004","article-title":"Conceptmix: A Compositional Image Generation Benchmark with Controllable Difficulty","volume":"37","author":"Wu","year":"2024","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_86","doi-asserted-by":"crossref","unstructured":"Sheynin, S., Polyak, A., Singer, U., Kirstain, Y., Zohar, A., Ashual, O., Parikh, D., and Taigman, Y. (2024, January 17\u201321). Emu Edit: Precise Image Editing via Recognition and Generation Tasks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00847"},{"key":"ref_87","unstructured":"Basu, S., Saberi, M., Bhardwaj, S., Chegini, A.M., Massiceti, D., Sanjabi, M., Hu, S.X., and Feizi, S. (2023). EditVal: Benchmarking Diffusion Based Text-Guided Image Editing Methods. arXiv."},{"key":"ref_88","doi-asserted-by":"crossref","unstructured":"Huang, Y., Xie, L., Wang, X., Yuan, Z., Cun, X., Ge, Y., Zhou, J., Dong, C., Huang, R., and Zhang, R. (2024, January 17\u201321). Smartedit: Exploring Complex Instruction-Based Image Editing with Multimodal Large Language Models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00799"},{"key":"ref_89","doi-asserted-by":"crossref","unstructured":"Yu, Q., Chow, W., Yue, Z., Pan, K., Wu, Y., Wan, X., Li, J., Tang, S., Zhang, H., and Zhuang, Y. (2025, January 11\u201315). AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea. Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR 2025), Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.02433"},{"key":"ref_90","unstructured":"Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. (2017). Gans Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium. arXiv."},{"key":"ref_91","unstructured":"Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X. (2016). Improved Techniques for Training Gans. arXiv."},{"key":"ref_92","unstructured":"Li, J., Li, D., Xiong, C., and Hoi, S. (2022, January 17\u201323). BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation. Proceedings of the International Conference on Machine Learning (ICML 2022), Baltimore, MD, USA."},{"key":"ref_93","doi-asserted-by":"crossref","unstructured":"Rasheed, H., Maaz, M., Shaji, S., Shaker, A., Khan, S., Cholakkal, H., Anwer, R.M., Xing, E., Yang, M.-H., and Khan, F.S. (2024, January 17\u201321). GlaMM: Pixel Grounding Large Multimodal Model. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01236"},{"key":"ref_94","doi-asserted-by":"crossref","unstructured":"Pramanick, S., Han, G., Hou, R., Nag, S., Lim, S.-N., Ballas, N., Wang, Q., Chellappa, R., and Almahairi, A. (2024, January 17\u201321). Jack of All Tasks Master of Many: Designing General-Purpose Coarse-to-Fine Vision-Language Model. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01335"},{"key":"ref_95","unstructured":"Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. (2022). Hierarchical Text-Conditional Image Generation with Clip Latents. arXiv."},{"key":"ref_96","doi-asserted-by":"crossref","unstructured":"Brooks, T., Holynski, A., and Efros, A.A. (2023, January 18\u201322). InstructPix2Pix: Learning to Follow Image Editing Instructions. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2023), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01764"},{"key":"ref_97","doi-asserted-by":"crossref","unstructured":"Kubota, K., Togo, R., Maeda, K., Ogawa, T., and Haseyama, M. (November, January 29). MLLM-Based Automatic Exploration of Editing Prompt for High Engagement Image Generation. Proceedings of the 2024 IEEE 13th Global Conference on Consumer Electronics (GCCE), Kitakyushu, Japan.","DOI":"10.1109\/GCCE62371.2024.10760859"},{"key":"ref_98","unstructured":"Wu, S., Fei, H., Qu, L., Ji, W., and Chua, T.-S. (2024, January 21\u201327). NExT-GPT: Any-to-Any Multimodal LLM. Proceedings of the Forty-first International Conference on Machine Learning (ICML 2024), Vancouver, BC, Canada."},{"key":"ref_99","doi-asserted-by":"crossref","first-page":"2720","DOI":"10.1109\/TBME.2018.2814538","article-title":"Medical Image Synthesis with Deep Convolutional Adversarial Networks","volume":"65","author":"Nie","year":"2018","journal-title":"IEEE Trans. Biomed. Eng."},{"key":"ref_100","doi-asserted-by":"crossref","unstructured":"Yoshino, K., Wakimoto, K., Nishimura, Y., and Nakamura, S. (2021). Caption Generation of Robot Behaviors Based on Unsupervised Learning of Action Segments. Conversational Dialogue Systems for the Next Decade, Springer. Lecture Notes in Electrical Engineering.","DOI":"10.1007\/978-981-15-8395-7_17"},{"key":"ref_101","doi-asserted-by":"crossref","unstructured":"Vedantam, R., Lawrence Zitnick, C., and Parikh, D. (2015, January 7\u201312). CIDEr: Consensus-based Image Description Evaluation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2015), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"ref_102","doi-asserted-by":"crossref","unstructured":"Li, B., Ge, Y., Ge, Y., Wang, G., Wang, R., Zhang, R., and Shan, Y. (2024, January 17\u201321). Seed-Bench: Benchmarking Multimodal Large Language Models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01263"},{"key":"ref_103","unstructured":"Yao, R., Zhang, B., Huang, J., Long, X., Zhang, Y., Zou, T., Wu, Y., Su, S., Xu, Y., and Zeng, W. (2025). LENS: Multi-Level Evaluation of Multimodal Reasoning with Large Language Models. arXiv."},{"key":"ref_104","doi-asserted-by":"crossref","unstructured":"Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W.X., and Wen, J.-R. (2023). Evaluating Object Hallucination in Large Vision-Language Models. arXiv.","DOI":"10.18653\/v1\/2023.emnlp-main.20"},{"key":"ref_105","first-page":"71995","article-title":"Gpt4tools: Teaching Large Language Model to Use Tools via Self-Instruction","volume":"36","author":"Yang","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_106","doi-asserted-by":"crossref","first-page":"220105","DOI":"10.1007\/s11432-024-4251-x","article-title":"Woodpecker: Hallucination Correction for Multimodal Large Language Models","volume":"67","author":"Yin","year":"2024","journal-title":"Sci. China Inf. Sci."},{"key":"ref_107","unstructured":"Raza, S., Narayanan, A., Khazaie, V.R., Vayani, A., Chettiar, M.S., Singh, A., Shah, M., and Pandya, D. (2025). HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation. arXiv."},{"key":"ref_108","doi-asserted-by":"crossref","unstructured":"Yu, T., Yao, Y., Zhang, H., He, T., Han, Y., Cui, G., Hu, J., Liu, Z., Zheng, H.-T., and Sun, M. (2024, January 17\u201321). RLHF-V: Towards Trustworthy Mllms via Behavior Alignment from Fine-Grained Correctional Human Feedback. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01310"},{"key":"ref_109","unstructured":"Liu, F., Lin, K., Li, L., Wang, J., Yacoob, Y., and Wang, L. (2023). Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning. arXiv, Available online: https:\/\/arxiv.org\/abs\/2306.14565."},{"key":"ref_110","doi-asserted-by":"crossref","unstructured":"Leng, S., Zhang, H., Chen, G., Li, X., Lu, S., Miao, C., and Bing, L. (2024, January 17\u201321). Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01316"},{"key":"ref_111","doi-asserted-by":"crossref","unstructured":"Jiang, C., Xu, H., Dong, M., Chen, J., Ye, W., Yan, M., Ye, Q., Zhang, J., Huang, F., and Zhang, S. (2024, January 17\u201321). Hallucination Augmented Contrastive Learning for Multimodal Large Language Model. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.02553"},{"key":"ref_112","first-page":"9694","article-title":"Align before Fuse: Vision and Language Representation Learning with Momentum Distillation","volume":"34","author":"Li","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_113","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks","volume":"39","author":"Ren","year":"2016","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_114","doi-asserted-by":"crossref","unstructured":"Schramowski, P., Brack, M., Deiseroth, B., and Kersting, K. (2023, January 18\u201322). Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2023), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.02157"},{"key":"ref_115","doi-asserted-by":"crossref","unstructured":"Poppi, S., Poppi, T., Cocchi, F., Cornia, M., Baraldi, L., and Cucchiara, R. (2025). Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models. Computer Vision\u2014ECCV 2024, Springer Nature. Lecture Notes in Computer Science.","DOI":"10.1007\/978-3-031-73668-1_20"},{"key":"ref_116","unstructured":"You, H., Zhang, H., Gan, Z., Du, X., Zhang, B., Wang, Z., Cao, L., Chang, S.-F., and Yang, Y. (2023). Ferret: Refer and Ground Anything Anywhere at Any Granularity. arXiv."},{"key":"ref_117","doi-asserted-by":"crossref","unstructured":"Lai, X., Tian, Z., Chen, Y., Li, Y., Yuan, Y., Liu, S., and Jia, J. (2024, January 17\u201321). Lisa: Reasoning Segmentation via Large Language Model. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00915"},{"key":"ref_118","unstructured":"Figure AI (2024). Helix: A Vision-Language-Action Model for Generalist Humanoid Control. Fig. AI News, Available online: https:\/\/www.figure.ai\/news\/helix."},{"key":"ref_119","unstructured":"NVIDIA, Bjorck, J., Casta\u00f1eda, F., Cherniadev, N., Da, X., Ding, R., Fan, L.J., Fang, Y., Fox, D., and Hu, F. (2025). GR00T N1: An Open Foundation Model for Generalist Humanoid Robots. arXiv."},{"key":"ref_120","unstructured":"Team, G.R., Abeyruwan, S., Ainslie, J., Alayrac, J.-B., Arenas, M.G., Armstrong, T., Balakrishna, A., Baruch, R., Bauza, M., and Blokzijl, M. (2025). Gemini Robotics: Bringing AI into the Physical World. arXiv."},{"key":"ref_121","doi-asserted-by":"crossref","unstructured":"Hong, W., Wang, W., Lv, Q., Xu, J., Yu, W., Ji, J., Wang, Y., Wang, Z., Dong, Y., and Ding, M. (2024, January 17\u201321). Cogagent: A Visual Language Model for Gui Agents. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01354"},{"key":"ref_122","doi-asserted-by":"crossref","unstructured":"Zhang, D., Li, S., Zhang, X., Zhan, J., Wang, P., Zhou, Y., and Qiu, X. (2023). SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities. arXiv.","DOI":"10.18653\/v1\/2023.findings-emnlp.1055"},{"key":"ref_123","unstructured":"Su, Y., Lan, T., Li, H., Xu, J., Wang, Y., and Cai, D. (2023). PandaGPT: One Model To Instruction-Follow Them All. arXiv."},{"key":"ref_124","unstructured":"Han, J., Zhang, R., Shao, W., Gao, P., Xu, P., Xiao, H., Zhang, K., Liu, C., Wen, S., and Guo, Z. (2023). ImageBind-LLM: Multi-Modality Instruction Tuning. arXiv."},{"key":"ref_125","doi-asserted-by":"crossref","unstructured":"Tang, Z., Yang, Z., Khademi, M., Liu, Y., Zhu, C., and Bansal, M. (2024, January 17\u201321). CoDi-2: In-Context Interleaved and Interactive Any-to-Any Generation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.02589"},{"key":"ref_126","unstructured":"Xie, Z., and Wu, C. (2024). Mini-Omni2: Towards Open-Source GPT-4o with Vision, Speech and Duplex Capabilities. arXiv."},{"key":"ref_127","doi-asserted-by":"crossref","unstructured":"Ge, Z., Huang, H., Zhou, M., Li, J., Wang, G., Tang, S., and Zhuang, Y. (2024, January 28). WorldGPT: Empowering LLM as Multimodal World Model. Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne, VIC, Australia.","DOI":"10.1145\/3664647.3681488"},{"key":"ref_128","unstructured":"Wang, W., Xie, J., Hu, C., Zou, H., Fan, J., Tong, W., Wen, Y., Wu, S., Deng, H., and Li, Z. (2023). DriveMLM: Aligning Multi-Modal Large Language Models with Behavioral Planning States for Autonomous Driving. arXiv."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/9\/1400\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T18:34:08Z","timestamp":1760034848000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/9\/1400"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,28]]},"references-count":128,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2025,9]]}},"alternative-id":["sym17091400"],"URL":"https:\/\/doi.org\/10.3390\/sym17091400","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,28]]}}}