{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,24]],"date-time":"2025-09-24T00:14:55Z","timestamp":1758672895344,"version":"3.44.0"},"publisher-location":"California","reference-count":0,"publisher":"International Joint Conferences on Artificial Intelligence Organization","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,9]]},"abstract":"<jats:p>In the latest advancements in multimodal learning, effectively addressing the spatial and semantic losses of visual data after encoding remains a critical challenge. This is because the performance of large multimodal models is positively correlated with the coupling between visual encoders and large language models. Existing approaches often face issues such as vector gaps or semantic disparities, resulting in information loss during the propagation process. To address these issues, we propose MAGE (Multimodal Alignment and Generation Enhancement), a novel framework that bridges the semantic spaces of vision and text through an innovative alignment mechanism. By introducing the Intelligent Alignment Network (IAN), MAGE achieves dimensional and semantic alignment. To reduce the gap between synonymous heterogeneous data, we employ a training strategy that combines cross-entropy and mean squared error, significantly enhancing the alignment effect. Moreover, to enhance MAGE\u2019s \u201cAny-to-Any\u201d capability, we developed a fine-tuning dataset for multimodal tool-calling instructions to expand the model\u2019s output capability boundaries. Finally, our proposed multimodal large model architecture, MAGE, achieved significantly better performance compared to similar works across various evaluation benchmarks, including MME, MMBench, and SEED. Complete code and appendix are available at: https:\/\/github.com\/GTCOM-NLP\/MAGE<\/jats:p>","DOI":"10.24963\/ijcai.2025\/107","type":"proceedings-article","created":{"date-parts":[[2025,9,19]],"date-time":"2025-09-19T08:10:40Z","timestamp":1758269440000},"page":"954-962","source":"Crossref","is-referenced-by-count":0,"title":["MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces"],"prefix":"10.24963","author":[{"given":"Shaojun","family":"E","sequence":"first","affiliation":[{"name":"Global Tone Communication Technology Co., Ltd., Beijing, China"},{"name":"School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuchen","family":"Yang","sequence":"additional","affiliation":[{"name":"Faculty of computing, Harbin Institute of Technology, Harbin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiaheng","family":"Wu","sequence":"additional","affiliation":[{"name":"Faculty of computing, Harbin Institute of Technology, Harbin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yan","family":"Zhang","sequence":"additional","affiliation":[{"name":"Global Tone Communication Technology Co., Ltd., Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tiejun","family":"Zhao","sequence":"additional","affiliation":[{"name":"Faculty of computing, Harbin Institute of Technology, Harbin, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ziyan","family":"Chen","sequence":"additional","affiliation":[{"name":"Global Tone Communication Technology Co., Ltd., Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"10584","event":{"number":"34","sponsor":["International Joint Conferences on Artificial Intelligence Organization (IJCAI)"],"acronym":"IJCAI-2025","name":"Thirty-Fourth International Joint Conference on Artificial Intelligence {IJCAI-25}","start":{"date-parts":[[2025,8,16]]},"theme":"Artificial Intelligence","location":"Montreal, Canada","end":{"date-parts":[[2025,8,22]]}},"container-title":["Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence"],"original-title":[],"deposited":{"date-parts":[[2025,9,23]],"date-time":"2025-09-23T11:33:00Z","timestamp":1758627180000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.ijcai.org\/proceedings\/2025\/107"}},"subtitle":[],"proceedings-subject":"Artificial Intelligence Research Articles","short-title":[],"issued":{"date-parts":[[2025,9]]},"references-count":0,"URL":"https:\/\/doi.org\/10.24963\/ijcai.2025\/107","relation":{},"subject":[],"published":{"date-parts":[[2025,9]]}}}