{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,14]],"date-time":"2026-06-14T07:11:16Z","timestamp":1781421076761,"version":"3.54.1"},"reference-count":77,"publisher":"Association for Computing Machinery (ACM)","issue":"11","funder":[{"DOI":"10.13039\/501100012166","name":"National Key R & D Program of China","doi-asserted-by":"crossref","award":["2023YFB4502804"],"award-info":[{"award-number":["2023YFB4502804"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Science Fund for Distinguished Young Scholars","award":["62025603"],"award-info":[{"award-number":["62025603"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["No. U21B2037, No. U22B2051, No. 62072389"],"award-info":[{"award-number":["No. U21B2037, No. U22B2051, No. 62072389"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Natural Science Foundation of Fujian Province of China","award":["No.2021J01002, No.2022J06001"],"award-info":[{"award-number":["No.2021J01002, No.2022J06001"]}]},{"DOI":"10.13039\/501100002858","name":"China Postdoctoral Science Foundation","doi-asserted-by":"crossref","award":["No. 2023M732948"],"award-info":[{"award-number":["No. 2023M732948"]}],"id":[{"id":"10.13039\/501100002858","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Key Research and Development Program of China","award":["No.2023YFB4502804"],"award-info":[{"award-number":["No.2023YFB4502804"]}]},{"name":"National Natural Science Fund for Young Scholars of China","award":["No. 62302411"],"award-info":[{"award-number":["No. 62302411"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,11,30]]},"abstract":"<jats:p>\n                    In recent times, automatic text-to-3D content creation has made significant progress, driven by the development of pretrained 2D diffusion models. Existing text-to-3D methods typically optimize the 3D representation to ensure that the rendered image aligns well with the given text, as evaluated by the pretrained 2D diffusion model. Nevertheless, a substantial domain gap exists between 2D images and 3D assets, primarily attributed to variations in camera-related attributes and the exclusive presence of foreground objects. Consequently, employing 2D diffusion models directly for optimizing 3D representations may lead to suboptimal outcomes. To address this issue, we present X-Dreamer, a novel approach for high-quality text-to-3D content creation that effectively bridges the gap between text-to-2D and text-to-3D synthesis. The key components of X-Dreamer are two innovative designs: Camera-Guided Low-Rank Adaptation (CG-LoRA) and Attention-Mask Alignment (AMA) Loss. CG-LoRA dynamically incorporates camera information into the pretrained diffusion models by employing camera-dependent generation for trainable parameters. This integration makes the 2D diffusion model camera-sensitive. AMA loss guides the attention map of the pretrained diffusion model using the binary mask of the 3D object, prioritizing the creation of the foreground object. This module ensures that the model focuses on generating accurate and detailed foreground objects. Extensive evaluations demonstrate the effectiveness of our proposed method compared to existing text-to-3D approaches. Our project webpage:\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"url\" xlink:href=\"https:\/\/anonymous-11111.github.io\/\">https:\/\/anonymous-11111.github.io\/<\/jats:ext-link>\n                    . Our code is available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"url\" xlink:href=\"https:\/\/github.com\/xmu-xiaoma666\/X-Dreamer\">https:\/\/github.com\/xmu-xiaoma666\/X-Dreamer<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3687475","type":"journal-article","created":{"date-parts":[[2024,8,28]],"date-time":"2024-08-28T12:28:13Z","timestamp":1724848093000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Creating High-Quality 3D Content by Bridging the Gap between Text-to-2D and Text-to-3D Generation"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8744-3423","authenticated-orcid":false,"given":"Yiwei","family":"Ma","sequence":"first","affiliation":[{"name":"Xiamen University, Xiamen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-7689-6803","authenticated-orcid":false,"given":"Yijun","family":"Fan","sequence":"additional","affiliation":[{"name":"Xiamen University, Xiamen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9956-6308","authenticated-orcid":false,"given":"Jiayi","family":"Ji","sequence":"additional","affiliation":[{"name":"Xiamen University, Xiamen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-0289-9672","authenticated-orcid":false,"given":"Haowei","family":"Wang","sequence":"additional","affiliation":[{"name":"Xiamen University, Xiamen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3025-0938","authenticated-orcid":false,"given":"Haibing","family":"Yin","sequence":"additional","affiliation":[{"name":"Hangzhou Dianzi University, Hangzhou, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3912-9306","authenticated-orcid":false,"given":"Xiaoshuai","family":"Sun","sequence":"additional","affiliation":[{"name":"Xiamen University, Xiamen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9163-2932","authenticated-orcid":false,"given":"Rongrong","family":"Ji","sequence":"additional","affiliation":[{"name":"Xiamen University, Xiamen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,11,10]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","unstructured":"Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia Leoni Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat Red Avila Igor Babuschkin Suchir Balaji Valerie Balcom Paul Baltescu Haiming Bao Mohammad Bavarian Jeff Belgum Irwan Bello Jake Berdine Gabriel Bernadett-Shapiro Christopher Berner Lenny Bogdonoff Oleg Boiko Madelaine Boyd Anna-Luisa Brakman Greg Brockman Tim Brooks Miles Brundage Kevin Button Trevor Cai Rosie Campbell Andrew Cann Brittany Carey Chelsea Carlson Rory Carmichael Brooke Chan Che Chang Fotis Chantzis Derek Chen Sully Chen Ruby Chen Jason Chen Mark Chen Ben Chess Chester Cho Casey Chu Hyung Won Chung Dave Cummings Jeremiah Currier Yunxing Dai Cory Decareaux Thomas Degry Noah Deutsch Damien Deville Arka Dhar David Dohan Steve Dowling Sheila Dunning Adrien Ecoffet Atty Eleti Tyna Eloundou David Farhi Liam Fedus Niko Felix Sim\u00f3n Posada Fishman Juston Forte Isabella Fulford Leo Gao Elie Georges Christian Gibson Vik Goel Tarun Gogineni Gabriel Goh Rapha Gontijo-Lopes Jonathan Gordon Morgan Grafstein Scott Gray Ryan Greene Joshua Gross Shixiang Shane Gu Yufei Guo Chris Hallacy Jesse Han Jeff Harris Yuchen He Mike Heaton Johannes Heidecke Chris Hesse Alan Hickey Wade Hickey Peter Hoeschele Brandon Houghton Kenny Hsu Shengli Hu Xin Hu Joost Huizinga Shantanu Jain and Shawn Jain. 2023. Gpt-4 technical report. arXiv:2303.08774. Retrieved from 10.48550\/arXiv.2303.08774","DOI":"10.48550\/arXiv.2303.08774"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","unstructured":"Sudarshan Babu Richard Liu Avery Zhou Michael Maire Greg Shakhnarovich and Rana Hanocka. 2023. Hyperfields: Towards zero-shot generation of nerfs from text. arXiv:2310.17075. Retrieved from 10.48550\/arXiv.2310.17075","DOI":"10.48550\/arXiv.2310.17075"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","unstructured":"Yogesh Balaji Seungjun Nah Xun Huang Arash Vahdat Jiaming Song Karsten Kreis Miika Aittala Timo Aila Samuli Laine Bryan Catanzaro Tero Karras and Ming-Yu Liu. 2022. ediffi: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv:2211.01324. Retrieved from 10.48550\/arXiv.2211.01324","DOI":"10.48550\/arXiv.2211.01324"},{"key":"e_1_3_2_5_2","unstructured":"Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared D. Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell Sandhini Agarwal Ariel Herbert-Voss Gretchen Krueger Tom Henighan Rewon Child Aditya Ramesh Daniel M. Ziegler Jeffrey Wu Clemens Winter Christopher Hesse Mark Chen Eric Sigler Mateusz Litwin Scott Gray Benjamin Chess Jack Clark Christopher Berner Sam McCandlish Alec Radford Ilya Sutskever and Dario Amodei. 2020. Language models are few-shot learners. In Proceedings of the 34th International Conference on Neural Information Processing Systems 1877\u20131901."},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00385"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612524"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","unstructured":"Rui Chen Yongwei Chen Ningxin Jiao and Kui Jia. 2023. Fantasia3d: Disentangling geometry and appearance for high-quality text-to-3d content creation. arXiv:2303.13873. Retrieved from 10.48550\/arXiv.2303.13873","DOI":"10.48550\/arXiv.2303.13873"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612489"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","unstructured":"Zilong Chen Feng Wang and Huaping Liu. 2023. Text-to-3D using Gaussian splatting. arXiv:2309.16585. Retrieved from 10.48550\/arXiv.2309.16585","DOI":"10.48550\/arXiv.2309.16585"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","unstructured":"Zhe Chen Jiannan Wu Wenhai Wang Weijie Su Guo Chen Sen Xing Zhong Muyan Qinglong Zhang Xizhou Zhu Lewei Lu Bin Li Ping Luo Tong Lu Yu Qiao and Jifeng Dai. 2023. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. arXiv:2312.14238. Retrieved from 10.48550\/arXiv.2312.14238","DOI":"10.48550\/arXiv.2312.14238"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","unstructured":"Ziluo Ding Hao Luo Ke Li Junpeng Yue Tiejun Huang and Zongqing Lu. 2023. Clip4mc: An rl-friendly vision-language model for minecraft. arXiv:2303.10571. Retrieved from 10.48550\/arXiv.2303.10571","DOI":"10.48550\/arXiv.2303.10571"},{"key":"e_1_3_2_13_2","first-page":"7641","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Fei Hao","year":"2024","unstructured":"Hao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang, and Tat-Seng Chua. 2024. Dysen-VDM: Empowering dynamics-aware text-to-video diffusion with LLMs. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 7641\u20137653."},{"key":"e_1_3_2_14_2","first-page":"6373","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Fei Hao","year":"2024","unstructured":"Hao Fei, Shengqiong Wu, Wei Ji, Hanwang Zhang, Meishan Zhang, Mong-Li Lee, and Wynne Hsu. 2024. Video-of-thought: Step-by-step video reasoning from perception to cognition. In Proceedings of the International Conference on Machine Learning, 6373\u20136391."},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","unstructured":"Hao Fei Shengqiong Wu Hanwang Zhang Tat-Seng Chua and Shuicheng Yan. 2024. VITRON: A unified pixel-level vision LLM for understanding generating segmenting editing. Computing Research Repository (2024). DOI: 10.13140\/RG.2.2.12584.58887","DOI":"10.13140\/RG.2.2.12584.58887"},{"key":"e_1_3_2_16_2","first-page":"1","article-title":"Enhancing video-language representations with structural spatio-temporal alignment","author":"Fei Hao","year":"2024","unstructured":"Hao Fei, Shengqiong Wu, Meishan Zhang, Min Zhang, Tat-Seng Chua, and Shuicheng Yan. 2024. Enhancing video-language representations with structural spatio-temporal alignment. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024), 1\u201318.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_17_2","unstructured":"Yuan-Chen Guo Ying-Tian Liu Ruizhi Shao Christian Laforte Vikram Voleti Guan Luo Chia-Hao Chen Zi-Xin Zou Chen Wang Yan-Pei Cao and Song-Hai Zhang. 2023. threestudio: A Unified Framework for 3D Content Generation. Retrieved from https:\/\/github.com\/threestudio-project\/threestudio"},{"key":"e_1_3_2_18_2","unstructured":"Jonathan Ho Ajay Jain and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. In Proceedings of the 34th International Conference on Neural Information Processing Systems 6840\u20136851."},{"key":"e_1_3_2_19_2","unstructured":"Edward J Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang Lu Wang and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv:2106.09685."},{"key":"e_1_3_2_20_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Hu Edward J","year":"2022","unstructured":"Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=nZeVKeeFYf9"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612022"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","unstructured":"Yukun Huang Jianan Wang Ailing Zeng He Cao Xianbiao Qi Yukai Shi Zheng-Jun Zha and Lei Zhang. 2023. DreamWaltz: Make a scene with complex 3D animatable avatars. arXiv:2305.12529. Retrieved from 10.48550\/arXiv.2305.12529","DOI":"10.48550\/arXiv.2305.12529"},{"key":"e_1_3_2_23_2","doi-asserted-by":"crossref","unstructured":"Yukun Huang Jianan Wang Ailing Zeng He Cao Xianbiao Qi Yukai Shi Zheng-Jun Zha and Lei Zhang. 2024. DreamWaltz: Make a scene with complex 3D animatable avatars. In Proceedings of the International Conference onNeural Information Processing Systems 1\u201319.","DOI":"10.1109\/TPAMI.2025.3586284"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00094"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3611789"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","unstructured":"Ruixiang Jiang Can Wang Jingbo Zhang Menglei Chai Mingming He Dongdong Chen and Jing Liao. 2023. AvatarCraft: Transforming text into neural human avatars with parameterized shape and pose control. arXiv:2303.17606. Retrieved from 10.48550\/arXiv.2303.17606","DOI":"10.48550\/arXiv.2303.17606"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00192"},{"key":"e_1_3_2_28_2","first-page":"30923","volume-title":"Proceedings of the 36 International Conference on Neural Information Processing Systems","author":"Lei Jiabao","year":"2022","unstructured":"Jiabao Lei, Yabin Zhang, and Kui Jia. 2022. Tango: Text-driven photorealistic and robust 3D stylization via lighting decomposition. In Proceedings of the 36 International Conference on Neural Information Processing Systems, 30923\u201330936."},{"key":"e_1_3_2_29_2","first-page":"28541","volume-title":"Proceedings of the 37th International Conference onNeural Information Processing Systems","author":"Li Chunyuan","year":"2024","unstructured":"Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. 2024b. LLaVa-med: Training a large language-and-vision assistant for biomedicine in one day. In Proceedings of the 37th International Conference onNeural Information Processing Systems, 28541\u201328564."},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01216"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","unstructured":"Weiyu Li Rui Chen Xuelin Chen and Ping Tan. 2023. Sweetdreamer: Aligning geometric priors in 2d diffusion for consistent text-to-3D. arXiv:2310.02596. Retrieved from 10.48550\/arXiv.2310.02596","DOI":"10.48550\/arXiv.2310.02596"},{"key":"e_1_3_2_32_2","first-page":"3279","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"38","author":"Li Yuhan","year":"2024","unstructured":"Yuhan Li, Yishun Dou, Yue Shi, Yu Lei, Xuanhong Chen, Yi Zhang, Peng Zhou, and Bingbing Ni. 2024a. Focaldreamer: Text-driven 3D editing via focal-fusion assembly. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, 3279\u20133287."},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","unstructured":"Zhang Li Biao Yang Qiang Liu Zhiyin Ma Shuo Zhang Jingxu Yang Yabo Sun Yuliang Liu and Xiang Bai. 2023. Monkey: Image resolution and text label are important things for large multi-modal models. arXiv:2311.06607. Retrieved from 10.48550\/arXiv.2311.06607","DOI":"10.48550\/arXiv.2311.06607"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00037"},{"key":"e_1_3_2_35_2","unstructured":"Haotian Liu Chunyuan Li Qingyang Wu and Yong Jae Lee. 2024. Visual instruction tuning. In Proceedings of the 37th International Conference onNeural Information Processing Systems 34892\u201334916."},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00853"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01645"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612451"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","unstructured":"Gen Luo Yiyi Zhou Yuxin Zhang Xiawu Zheng Xiaoshuai Sun and Rongrong Ji. 2024. Feast your eyes: Mixture-of-resolution adaptation for multimodal large language models. arXiv:2403.03003. Retrieved from 10.48550\/arXiv.2403.03003","DOI":"10.48550\/arXiv.2403.03003"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2023.109420"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3547910"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00258"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1145\/2343483.2343493"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01218"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01313"},{"key":"e_1_3_2_46_2","first-page":"99","volume-title":"Communications of the ACM","volume":"65","author":"Mildenhall Ben","year":"2021","unstructured":"Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. 2021. Nerf: Representing scenes as neural radiance fields for view synthesis. Communications of the ACM 65, 1 (2021), 99\u2013106."},{"key":"e_1_3_2_47_2","first-page":"1","volume-title":"Proceedings of the SIGGRAPH Asia 2022 Conference Papers","author":"Khalid Nasir Mohammad","year":"2022","unstructured":"Nasir Mohammad Khalid, Tianhao Xie, Eugene Belilovsky, and Tiberiu Popa. 2022. Clip-mesh: Generating textured meshes from text using pretrained image-text models. In Proceedings of the SIGGRAPH Asia 2022 Conference Papers, 1\u20138."},{"key":"e_1_3_2_48_2","first-page":"1","volume-title":"ACM Transactions on Graphics","volume":"41","author":"M\u00fcller Thomas","year":"2022","unstructured":"Thomas M\u00fcller, Alex Evans, Christoph Schied, and Alexander Keller. 2022. Instant neural graphics primitives with a multiresolution hash encoding. ACM Transactions on Graphics 41, 4 (2022), 1\u201315."},{"key":"e_1_3_2_49_2","first-page":"8026","volume-title":"Proceedings of the 33rd International Conference onNeural Information Processing Systems","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas K\u00f6pf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. Pytorch: An imperative style, high-performance deep learning library. In Proceedings of the 33rd International Conference onNeural Information Processing Systems, 8026\u20138037."},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","unstructured":"Ben Poole Ajay Jain Jonathan T. Barron and Ben Mildenhall. 2022. Dreamfusion: Text-to-3d using 2d diffusion. arXiv:2209.14988. Retrieved from 10.48550\/arXiv.2209.14988","DOI":"10.48550\/arXiv.2209.14988"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612012"},{"key":"e_1_3_2_52_2","first-page":"8748","volume-title":"International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning. PMLR, 8748\u20138763."},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00223"},{"key":"e_1_3_2_54_2","unstructured":"Aditya Ramesh Prafulla Dhariwal Alex Nichol Casey Chu and Mark Chen. 2022. Hierarchical text-conditional image generation with clip latents. arXiv:2204.06125."},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3588432.3591503"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00985"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.5555\/3600270.3602913"},{"key":"e_1_3_2_59_2","first-page":"25278","article-title":"Laion-5b: An open large-scale dataset for training next generation image-text models","author":"Schuhmann Christoph","year":"2022","unstructured":"Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. 2022. Laion-5b: An open large-scale dataset for training next generation image-text models. In Proceedings of the 36th International Conference on Neural Information Processing Systems, 25278\u201325294.","journal-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00046"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","unstructured":"Junyoung Seo Wooseok Jang Min-Seop Kwak Jaehoon Ko Hyeonsu Kim Junho Kim Jin-Hwa Kim Jiyoung Lee and Seungryong Kim. 2023. Let 2d diffusion model know 3d-consistency for robust text-to-3d generation. arXiv:2303.07937. Retrieved from 10.48550\/arXiv.2303.07937","DOI":"10.48550\/arXiv.2303.07937"},{"key":"e_1_3_2_62_2","first-page":"6087","volume-title":"Proceedings of the 35th International Conference on Neural Information Processing Systems","author":"Shen Tianchang","year":"2021","unstructured":"Tianchang Shen, Jun Gao, Kangxue Yin, Ming-Yu Liu, and Sanja Fidler. 2021. Deep marching tetrahedra: A hybrid representation for high-resolution 3d shape synthesis. In Proceedings of the 35th International Conference on Neural Information Processing Systems, 6087\u20136101."},{"key":"e_1_3_2_63_2","first-page":"2256","volume-title":"International Conference on Machine Learning","author":"Sohl-Dickstein Jascha","year":"2015","unstructured":"Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning. PMLR, 2256\u20132265."},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","unstructured":"Yang Song Jascha Sohl-Dickstein Diederik P. Kingma Abhishek Kumar Stefano Ermon and Ben Poole. 2020. Score-based generative modeling through stochastic differential equations. arXiv:2011.13456. Retrieved from 10.48550\/arXiv.2011.13456","DOI":"10.48550\/arXiv.2011.13456"},{"key":"e_1_3_2_65_2","volume-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence","author":"Sun Gan","year":"2024","unstructured":"Gan Sun, Wenqi Liang, Jiahua Dong, Jun Li, Zhengming Ding, and Yang Cong. 2024. Create your world: Lifelong text-to-image diffusion. IEEE Transactions on Pattern Analysis and Machine Intelligence 46, 9 (2024), 6454\u20136470."},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","unstructured":"Junshu Tang Tengfei Wang Bo Zhang Ting Zhang Ran Yi Lizhuang Ma and Dong Chen. 2023. Make-it-3d: High-fidelity 3d creation from a single image with diffusion prior. arXiv:2303.14184. Retrieved from 10.48550\/arXiv.2303.14184","DOI":"10.48550\/arXiv.2303.14184"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1364\/JOSA.57.001105"},{"key":"e_1_3_2_68_2","unstructured":"Patrick von Platen Suraj Patil Anton Lozhkov Pedro Cuenca Nathan Lambert Kashif Rasul Mishig Davaadorj and Thomas Wolf. 2022. Diffusers: State-of-the-art diffusion models. Retrieved from https:\/\/github.com\/huggingface\/diffusers"},{"key":"e_1_3_2_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01214"},{"key":"e_1_3_2_70_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612490"},{"key":"e_1_3_2_71_2","first-page":"1","volume-title":"Proceedings of the 37th International Conference onNeural Information Processing Systems","author":"Wang Wenhai","year":"2024","unstructured":"Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, and Jifeng Dai. 2024. VisionLLM: Large language model is also an open-ended decoder for vision-centric tasks. In Proceedings of the 37th International Conference onNeural Information Processing Systems, 1\u201313."},{"key":"e_1_3_2_72_2","doi-asserted-by":"publisher","unstructured":"Zhengyi Wang Cheng Lu Yikai Wang Fan Bao Chongxuan Li Hang Su and Jun Zhu. 2023. ProlificDreamer: High-fidelity and diverse text-to-3D generation with variational score distillation. arXiv:2305.16213. Retrieved from 10.48550\/arXiv.2305.16213","DOI":"10.48550\/arXiv.2305.16213"},{"key":"e_1_3_2_73_2","first-page":"16805","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wei Jiacheng","year":"2023","unstructured":"Jiacheng Wei, Hao Wang, Jiashi Feng, Guosheng Lin, and Kim-Hui Yap. 2023. TAPS3D: Text-guided 3D textured shape generation from pseudo supervision. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 16805\u201316815."},{"key":"e_1_3_2_74_2","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Wu Shengqiong","year":"2024","unstructured":"Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. 2024. NExT-GPT: Any-to-any multimodal LLM. In Proceedings of the International Conference on Machine Learning."},{"key":"e_1_3_2_75_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.02003"},{"key":"e_1_3_2_76_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612363"},{"key":"e_1_3_2_77_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612232"},{"key":"e_1_3_2_78_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00355"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3687475","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,10]],"date-time":"2025-11-10T14:51:25Z","timestamp":1762786285000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3687475"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,10]]},"references-count":77,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2025,11,30]]}},"alternative-id":["10.1145\/3687475"],"URL":"https:\/\/doi.org\/10.1145\/3687475","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,10]]},"assertion":[{"value":"2024-06-14","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-07-26","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}