{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,15]],"date-time":"2026-01-15T20:37:38Z","timestamp":1768509458986,"version":"3.49.0"},"reference-count":42,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2026,1,11]],"date-time":"2026-01-11T00:00:00Z","timestamp":1768089600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100008535","name":"Nanning Normal University","doi-asserted-by":"publisher","award":["602021239078,602021239375"],"award-info":[{"award-number":["602021239078,602021239375"]}],"id":[{"id":"10.13039\/501100008535","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>While text-to-image and customized generation methods demonstrate strong capabilities in single-image generation, they fall short in supporting immersive applications that require coherent 360\u00b0 panoramas. Conversely, existing panorama generation models lack customization capabilities. In panoramic scenes, reference objects often appear as minor background elements and may be multiple in number, while reference images across different views exhibit weak correlations. To address these challenges, we propose a diffusion-based framework for customized multi-view image generation. Our approach introduces a decoupled feature injection mechanism within a dual-UNet architecture to handle weakly correlated reference images, effectively integrating spatial information by concurrently feeding both reference images and noise into the denoising branch. A hybrid attention mechanism enables deep fusion of reference features and multi-view representations. Furthermore, a data augmentation strategy facilitates viewpoint-adaptive pose adjustments, and panoramic coordinates are employed to guide multi-view attention. The experimental results demonstrate our model\u2019s effectiveness in generating coherent, high-quality customized multi-view images.<\/jats:p>","DOI":"10.3390\/jimaging12010040","type":"journal-article","created":{"date-parts":[[2026,1,12]],"date-time":"2026-01-12T08:20:37Z","timestamp":1768206037000},"page":"40","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["A Dual-UNet Diffusion Framework for Personalized Panoramic Generation"],"prefix":"10.3390","volume":"12","author":[{"given":"Jing","family":"Shen","sequence":"first","affiliation":[{"name":"School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China"},{"name":"State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Leigang","family":"Huo","sequence":"additional","affiliation":[{"name":"Guangxi Key Lab of Human-Machine Interaction and Intelligent Decision, Nanning Normal University, Nanning 530100, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6748-6709","authenticated-orcid":false,"given":"Chunlei","family":"Huo","sequence":"additional","affiliation":[{"name":"School of Information Engineering, Capital Normal University, Beijing 100048, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2089-9733","authenticated-orcid":false,"given":"Shiming","family":"Xiang","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China"},{"name":"State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,1,11]]},"reference":[{"key":"ref_1","unstructured":"Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., M\u00fcller, J., Penna, J., and Rombach, R. (2023). Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv."},{"key":"ref_2","first-page":"8","article-title":"Improving image generation with better captions","volume":"2","author":"Betker","year":"2023","journal-title":"Comput. Sci."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Mokady, R., Hertz, A., Aberman, K., Pritch, Y., and Cohen-Or, D. (2023, January 17\u201324). Null-text inversion for editing real images using guided diffusion models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00585"},{"key":"ref_4","first-page":"36479","article-title":"Photorealistic text-to-image diffusion models with deep language understanding","volume":"35","author":"Saharia","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022, January 18\u201324). High-resolution image synthesis with latent diffusion models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Karras, T., Laine, S., and Aila, T. (2019, January 15\u201320). A style-based generator architecture for generative adversarial networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00453"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Song, Y., Yang, P., Ci, H., and Shou, M.Z. (2024). IDProtector: An Adversarial Noise Encoder to Protect Against ID-Preserving Image Generation. arXiv.","DOI":"10.1109\/CVPR52734.2025.00287"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Ma, J., Liang, J., Chen, C., and Lu, H. (August, January 27). Subject-diffusion: Open domain personalized text-to-image generation without test-time fine-tuning. Proceedings of the ACM SIGGRAPH 2024 Conference Papers, Denver, CO, USA.","DOI":"10.1145\/3641519.3657469"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Zhang, H., Duan, Z., Wang, X., Chen, Y., and Zhang, Y. (2025). EliGen: Entity-Level Controlled Image Generation with Regional Attention. arXiv.","DOI":"10.1145\/3743093.3771013"},{"key":"ref_10","unstructured":"He, J., Tuo, Y., Chen, B., Zhong, C., Geng, Y., and Bo, L. (2025). AnyStory: Towards Unified Single and Multiple Subject Personalization in Text-to-Image Generation. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K. (2023, January 17\u201324). Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.02155"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Kumari, N., Zhang, B., Zhang, R., Shechtman, E., and Zhu, J.Y. (2023, January 17\u201324). Multi-Concept Customization of Text-to-Image Diffusion. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00192"},{"key":"ref_13","first-page":"1","article-title":"Neural rendering in a room: Amodal 3D understanding and free-viewpoint rendering for the closed scene composed of pre-captured objects","volume":"41","author":"Yang","year":"2022","journal-title":"ACM Trans. Graph. (TOG)"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Yang, B., Dong, W., Ma, L., Hu, W., Liu, X., Cui, Z., and Ma, Y. (2024, January 16\u201321). Dreamspace: Dreaming your room space with text-driven panoramic texture propagation. Proceedings of the 2024 IEEE Conference Virtual Reality and 3D User Interfaces (VR), Orlando, FL, USA.","DOI":"10.1109\/VR58804.2024.00085"},{"key":"ref_15","unstructured":"Tang, S., Zhang, F., Chen, J., Wang, P., and Yasutaka, F. (2023). MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware Diffusion. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zhang, C., Wu, Q., Gambardella, C.C., Huang, X., Phung, D., Ouyang, W., and Cai, J. (2024, January 16\u201322). Taming stable diffusion for text to 360 panorama image generation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00607"},{"key":"ref_17","unstructured":"Wu, T., Zheng, C., and Cham, T.J. (2023). Panodiffusion: 360-degree panorama outpainting via diffusion. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Li, Z., Li, Z., Cui, Z., Pollefeys, M., and Oswald, M.R. (2024, January 16\u201322). Sat2scene: 3d urban scene generation from satellite images with diffusion. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00682"},{"key":"ref_19","first-page":"1","article-title":"Text2light: Zero-shot text-driven hdr panorama generation","volume":"41","author":"Chen","year":"2022","journal-title":"ACM Trans. Graph. (TOG)"},{"key":"ref_20","unstructured":"Ni, C., Wang, X., Zhu, Z., Wang, W., Li, H., Zhao, G., Li, J., Qin, W., Huang, G., and Mei, W. (2025). Wonderturbo: Generating interactive 3d world in 0.72 s. arXiv."},{"key":"ref_21","unstructured":"Gao, R., Chen, K., Li, Z., Hong, L., Li, Z., and Xu, Q. (2024). Magicdrive3d: Controllable 3d generation for any-view rendering in street scenes. arXiv."},{"key":"ref_22","unstructured":"Huang, Z., Guo, Y.C., Wang, H., Yi, R., Ma, L., Cao, Y.P., and Sheng, L. (2024). Mv-adapter: Multi-view consistent image generation made easy. arXiv."},{"key":"ref_23","unstructured":"Liu, Y., Lin, C., Zeng, Z., Long, X., Liu, L., Komura, T., and Wang, W. (2023). Syncdreamer: Generating multiview-consistent images from a single-view image. arXiv."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Long, X., Guo, Y.C., Lin, C., Liu, Y., Dou, Z., Liu, L., Ma, Y., Zhang, S.H., Habermann, M., and Theobalt, C. (2024, January 7\u201321). Wonder3d: Single image to 3d using cross-domain diffusion. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00951"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Huang, Z., Wen, H., Dong, J., Wang, Y., Li, Y., Chen, X., Cao, Y.P., Liang, D., Qiao, Y., and Dai, B. (2024, January 7\u201321). Epidiff: Enhancing multi-view synthesis via localized epipolar-constrained diffusion. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00934"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Tang, S., Chen, J., Wang, D., Tang, C., Zhang, F., Fan, Y., Chandra, V., Furukawa, Y., and Ranjan, R. (October, January 29). MVDiffusion++: A Dense High-Resolution Multi-view Diffusion Model for Single or Sparse-View 3D Object Reconstruction. Proceedings of the Computer Vision\u2014ECCV 2024, Milan, Italy.","DOI":"10.1007\/978-3-031-72640-8_10"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Yuan, X., Tang, S., Li, K., Yuille, A., and Wang, P. (2024). CamFreeDiff: Camera-free Image to Panorama Generation with Diffusion Model. arXiv.","DOI":"10.1109\/CVPR52734.2025.01530"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Wang, J., Chen, Z., Ling, J., Xie, R., and Song, L. (2023). 360-degree panorama generation from few unregistered nfov images. arXiv.","DOI":"10.1145\/3581783.3612508"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Lu, Z., Hu, K., Wang, C., Bai, L., and Wang, Z. (2024, January 20\u201327). Autoregressive omni-aware outpainting for open-vocabulary 360-degree image generation. Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada.","DOI":"10.1609\/aaai.v38i13.29332"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Yang, B., Gu, S., Zhang, B., Zhang, T., Chen, X., Sun, X., Chen, D., and Wen, F. (2023, January 17\u201324). Paint by example: Exemplar-based image editing with diffusion models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01763"},{"key":"ref_31","unstructured":"Ye, H., Zhang, J., Liu, S., Han, X., and Yang, W. (2023). Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Chen, X., Huang, L., Liu, Y., Shen, Y., Zhao, D., and Zhao, H. (2024, January 16\u201322). Anydoor: Zero-shot object-level image customization. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00630"},{"key":"ref_33","first-page":"6000","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_34","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning transferable visual models from natural language supervision. Proceedings of the International Conference on Machine Learning, Virtual."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Zhang, L., Rao, A., and Agrawala, M. (2023, January 2\u20136). Adding conditional control to text-to-image diffusion models. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.00355"},{"key":"ref_36","unstructured":"Kingma, D.P., and Welling, M. (2022). Auto-Encoding Variational Bayes. arXiv."},{"key":"ref_37","first-page":"55975","article-title":"Era3d: High-resolution multiview diffusion using efficient row-wise attention","volume":"37","author":"Li","year":"2024","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Chang, A., Dai, A., Funkhouser, T., Halber, M., Niessner, M., Savva, M., Song, S., Zeng, A., and Zhang, Y. (2017). Matterport3d: Learning from rgb-d data in indoor environments. arXiv.","DOI":"10.1109\/3DV.2017.00081"},{"key":"ref_39","unstructured":"Li, J., Li, D., Savarese, S., and Hoi, S. (2023, January 23\u201329). Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. Proceedings of the International Conference on Machine Learning, Honolulu, HI, USA."},{"key":"ref_40","unstructured":"Jocher, G., Chaurasia, A., and Qiu, J. (2025, March 05). Ultralytics YOLOv8. Available online: https:\/\/github.com\/ultralytics\/ultralytics."},{"key":"ref_41","first-page":"6629","article-title":"Gans trained by a two time-scale update rule converge to a local nash equilibrium","volume":"30","author":"Heusel","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_42","first-page":"2234","article-title":"Improved techniques for training gans","volume":"29","author":"Salimans","year":"2016","journal-title":"Adv. Neural Inf. Process. Syst."}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/1\/40\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,15]],"date-time":"2026-01-15T05:33:45Z","timestamp":1768455225000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/1\/40"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,11]]},"references-count":42,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,1]]}},"alternative-id":["jimaging12010040"],"URL":"https:\/\/doi.org\/10.3390\/jimaging12010040","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,11]]}}}