{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,24]],"date-time":"2025-09-24T00:14:42Z","timestamp":1758672882228,"version":"3.44.0"},"publisher-location":"California","reference-count":0,"publisher":"International Joint Conferences on Artificial Intelligence Organization","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025,9]]},"abstract":"<jats:p>We proposes a novel visual tokenizer by combining high-level semantic tokens and low-level pixel tokens to represent images, aiming to address the challenges of image-to-sequence conversion for Large Language Models (LLMs). Existing visual tokenizers, such as VQ-VAE and diffusion-based models, either struggle with token explosion as image resolution increases or fail to capture detailed structural information. Our method introduces a dual-token system: high-level semantic tokens capture the main content of the image, while low-level pixel tokens preserve structural details. By integrating these tokens in a hybrid architecture, we leverage a VQ-VAE branch to generate low-resolution guidance and a diffusion process to reconstruct high-resolution images with both semantic coherence and structural accuracy. This approach significantly reduces the number of required tokens and enhances image reconstruction quality, offering an efficient solution for tasks like image generation and understanding based on LLMs.<\/jats:p>","DOI":"10.24963\/ijcai.2025\/167","type":"proceedings-article","created":{"date-parts":[[2025,9,19]],"date-time":"2025-09-19T08:10:40Z","timestamp":1758269440000},"page":"1494-1502","source":"Crossref","is-referenced-by-count":0,"title":["A Dual Stream Visual Tokenizer for LLM Image Generation"],"prefix":"10.24963","author":[{"given":"Yongqian","family":"Li","sequence":"first","affiliation":[{"name":"School of Computer Science, National Engineering Research Center for Multimedia Software and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yong","family":"Luo","sequence":"additional","affiliation":[{"name":"School of Computer Science, National Engineering Research Center for Multimedia Software and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiantao","family":"Cai","sequence":"additional","affiliation":[{"name":"School of Computer Science, National Engineering Research Center for Multimedia Software and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zheng","family":"He","sequence":"additional","affiliation":[{"name":"School of Computer Science, National Engineering Research Center for Multimedia Software and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhennan","family":"Meng","sequence":"additional","affiliation":[{"name":"Mobvoi Innovation Technology Company Limited"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nidong","family":"Wang","sequence":"additional","affiliation":[{"name":"Mobvoi Innovation Technology Company Limited"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yunlin","family":"Chen","sequence":"additional","affiliation":[{"name":"Mobvoi Innovation Technology Company Limited"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhifei","family":"Li","sequence":"additional","affiliation":[{"name":"Mobvoi Innovation Technology Company Limited"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"10584","event":{"number":"34","sponsor":["International Joint Conferences on Artificial Intelligence Organization (IJCAI)"],"acronym":"IJCAI-2025","name":"Thirty-Fourth International Joint Conference on Artificial Intelligence {IJCAI-25}","start":{"date-parts":[[2025,8,16]]},"theme":"Artificial Intelligence","location":"Montreal, Canada","end":{"date-parts":[[2025,8,22]]}},"container-title":["Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence"],"original-title":[],"deposited":{"date-parts":[[2025,9,23]],"date-time":"2025-09-23T11:33:11Z","timestamp":1758627191000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.ijcai.org\/proceedings\/2025\/167"}},"subtitle":[],"proceedings-subject":"Artificial Intelligence Research Articles","short-title":[],"issued":{"date-parts":[[2025,9]]},"references-count":0,"URL":"https:\/\/doi.org\/10.24963\/ijcai.2025\/167","relation":{},"subject":[],"published":{"date-parts":[[2025,9]]}}}