{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T18:53:46Z","timestamp":1782413626245,"version":"3.54.5"},"reference-count":106,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2025,11,6]],"date-time":"2025-11-06T00:00:00Z","timestamp":1762387200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Recent deep learning-based remote sensing analysis models often struggle with performance degradation due to domain shifts caused by illumination variations (clear to overcast), changing atmospheric conditions (clear to foggy, dusty), and physical scene changes (clear to snowy). Addressing domain shift in aerial image segmentation is challenging due to limited training data availability, including costly data collection and annotation. We propose Multi-Weather DomainShifter, a comprehensive multi-weather domain transfer system that augments single-domain images into various weather conditions without additional laborious annotation, coordinated by a large language model (LLM) agent. Specifically, we utilize Unreal Engine to construct a synthetic dataset featuring images captured under diverse conditions such as overcast, foggy, and dusty settings. We then propose a latent space style transfer model that generates alternate domain versions based on real aerial datasets. Additionally, we present a multi-modal snowy scene diffusion model with LLM-assisted scene descriptors to add snowy elements into scenes. Multi-weather DomainShifter integrates these two approaches into a tool library and leverages the agent for tool selection and execution. Extensive experiments on the ISPRS Vaihingen and Potsdam dataset demonstrate that domain shift caused by weather change in aerial image-leads to significant performance drops, then verify our proposal\u2019s capacity to adapt models to perform well in shifted domains while maintaining their effectiveness in the original domain.<\/jats:p>","DOI":"10.3390\/jimaging11110395","type":"journal-article","created":{"date-parts":[[2025,11,7]],"date-time":"2025-11-07T09:14:45Z","timestamp":1762506885000},"page":"395","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Multi-Weather DomainShifter: A Comprehensive Multi-Weather Transfer LLM Agent for Handling Domain Shift in Aerial Image Processing"],"prefix":"10.3390","volume":"11","author":[{"given":"Yubo","family":"Wang","sequence":"first","affiliation":[{"name":"Department of Modern Mechanical Engineering, Waseda University, Tokyo 169-8555, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ruijia","family":"Wen","sequence":"additional","affiliation":[{"name":"Department of Modern Mechanical Engineering, Waseda University, Tokyo 169-8555, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hiroyuki","family":"Ishii","sequence":"additional","affiliation":[{"name":"Department of Modern Mechanical Engineering, Waseda University, Tokyo 169-8555, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jun","family":"Ohya","sequence":"additional","affiliation":[{"name":"Department of Modern Mechanical Engineering, Waseda University, Tokyo 169-8555, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,11,6]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"101009","DOI":"10.1016\/j.aei.2019.101009","article-title":"Convolutional neural networks for object detection in aerial imagery for disaster response and recovery","volume":"43","author":"Pi","year":"2020","journal-title":"Adv. Eng. Inform."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Wang, Y., Wang, Z., Nakano, Y., Nishimatsu, K., Hasegawa, K., and Ohya, J. (2022, January 26\u201329). Context Enhanced Traffic Segmentation: Traffic jam and road surface segmentation from aerial image. Proceedings of the 2022 IEEE 14th Image, Video, and Multidimensional Signal Processing Workshop (IVMSP), Nafplio, Greece.","DOI":"10.1109\/IVMSP54334.2022.9816350"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"105586","DOI":"10.1016\/j.envsoft.2022.105586","article-title":"V-FloodNet: A video segmentation system for urban flood detection and quantification","volume":"160","author":"Liang","year":"2023","journal-title":"Environ. Model. Softw."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Li, X., He, H., Li, X., Li, D., Cheng, G., Shi, J., Weng, L., Tong, Y., and Lin, Z. (2021, January 19\u201325). Pointflow: Flowing semantics through points for aerial image segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Virtual.","DOI":"10.1109\/CVPR46437.2021.00420"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Wang, Y., Wang, Z., Nakano, Y., Hasegawa, K., Ishii, H., and Ohya, J. (2024, January 24\u201326). MAC: Multi-Scales Attention Cascade for Aerial Image Segmentation. Proceedings of the 13th International Conference on Pattern Recognition Applications and Methods, ICPRAM 2024, Science and Technology Publications, Lda, Rome, Italy.","DOI":"10.5220\/0012343500003654"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Toker, A., Eisenberger, M., Cremers, D., and Leal-Taix\u00e9, L. (2024, January 17\u201321). Satsynth: Augmenting image-mask pairs through diffusion models for aerial semantic segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.02615"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Dai, D., and Van Gool, L. (2018, January 4\u20137). Dark model adaptation: Semantic image segmentation from daytime to nighttime. Proceedings of the 2018 21st International Conference on Intelligent Transportation Systems (ITSC), Maui, HI, USA.","DOI":"10.1109\/ITSC.2018.8569387"},{"key":"ref_8","unstructured":"Michaelis, C., Mitzkus, B., Geirhos, R., Rusak, E., Bringmann, O., Ecker, A.S., Bethge, M., and Brendel, W. (2019). Benchmarking robustness in object detection: Autonomous driving when winter is coming. arXiv."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Sun, T., Segu, M., Postels, J., Wang, Y., Van Gool, L., Schiele, B., Tombari, F., and Yu, F. (2022, January 19\u201324). SHIFT: A synthetic driving dataset for continuous multi-task domain adaptation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.02068"},{"key":"ref_10","unstructured":"International Society for Photogrammetry and Remote Sensing (ISPRS) (2025, July 21). ISPRS 2D Semantic Labeling Contest. Available online: https:\/\/www.isprs.org\/resources\/datasets\/benchmarks\/UrbanSemLab\/semantic-labeling.aspx."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"256","DOI":"10.1016\/j.isprsjprs.2013.10.004","article-title":"Results of the ISPRS benchmark on urban object detection and 3D building reconstruction","volume":"93","author":"Rottensteiner","year":"2014","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 11\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Virtual.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Denver, CO, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_14","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Xiao, T., Liu, Y., Zhou, B., Jiang, Y., and Sun, J. (2018, January 8\u201314). Unified perceptual parsing for scene understanding. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01228-1_26"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Chen, L.C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H. (2018, January 8\u201314). Encoder-decoder with atrous separable convolution for semantic image segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Fu, J., Liu, J., Tian, H., Li, Y., Bao, Y., Fang, Z., and Lu, H. (2019, January 15\u201320). Dual attention network for scene segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00326"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Kirillov, A., Wu, Y., He, K., and Girshick, R. (2020, January 16\u201318). Pointrend: Image segmentation as rendering. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00982"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully convolutional networks for semantic segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Strudel, R., Garcia, R., Laptev, I., and Schmid, C. (2021, January 10\u201317). Segmenter: Transformer for semantic segmentation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Virtual.","DOI":"10.1109\/ICCV48922.2021.00717"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. (2017, January 21\u201326). Pyramid scene parsing network. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.660"},{"key":"ref_22","unstructured":"Waqas Zamir, S., Arora, A., Gupta, A., Khan, S., Sun, G., Shahbaz Khan, F., Zhu, F., Shao, L., Xia, G.S., and Bai, X. (2019, January 16\u201320). isaid: A large-scale dataset for instance segmentation in aerial images. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, Long Beach, CA, USA."},{"key":"ref_23","unstructured":"Brown, T.B. (2020). Language models are few-shot learners. arXiv."},{"key":"ref_24","unstructured":"Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., Altman, S., and Anadkat, S. (2023). Gpt-4 technical report. arXiv."},{"key":"ref_25","unstructured":"Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.A., Lacroix, T., Rozi\u00e8re, B., Goyal, N., Hambro, E., and Azhar, F. (2023). Llama: Open and efficient foundation language models. arXiv."},{"key":"ref_26","unstructured":"Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., and Bhosale, S. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv."},{"key":"ref_27","unstructured":"Bai, J., Bai, S., Chu, Y., Cui, Z., Dang, K., Deng, X., Fan, Y., Ge, W., Han, Y., and Huang, F. (2023). Qwen technical report. arXiv."},{"key":"ref_28","unstructured":"Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., and Lv, C. (2025). Qwen3 technical report. arXiv."},{"key":"ref_29","unstructured":"Liu, A., Feng, B., Xue, B., Wang, B., Wu, B., Lu, C., Zhao, C., Deng, C., Zhang, C., and Ruan, C. (2024). Deepseek-v3 technical report. arXiv."},{"key":"ref_30","unstructured":"Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., and Bi, X. (2025). Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022, January 18\u201322). High-resolution image synthesis with latent diffusion models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"ref_32","unstructured":"Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., M\u00fcller, J., Penna, J., and Rombach, R. (2023). Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv."},{"key":"ref_33","unstructured":"Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. (2022). Hierarchical text-conditional image generation with clip latents. arXiv."},{"key":"ref_34","first-page":"8","article-title":"Improving image generation with better captions","volume":"2","author":"Betker","year":"2023","journal-title":"Comput. Sci."},{"key":"ref_35","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014). Generative adversarial nets. Advances in Neural Information Processing Systems, Curran Associates, Inc.. Available online: https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2014\/hash\/f033ed80deb0234979a61f95710dbe25-Abstract.html."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Zhu, J.Y., Park, T., Isola, P., and Efros, A.A. (2017, January 22\u201329). Unpaired image-to-image translation using cycle-consistent adversarial networks. Proceedings of the IEEE International Conference On Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.244"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Zhang, H., Xu, T., Li, H., Zhang, S., Wang, X., Huang, X., and Metaxas, D.N. (2017, January 22\u201329). Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks. Proceedings of the IEEE internAtional Conference On Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.629"},{"key":"ref_38","unstructured":"Brock, A. (2018). Large Scale GAN Training for High Fidelity Natural Image Synthesis. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Karras, T., Laine, S., and Aila, T. (2019, January 15\u201320). A style-based generator architecture for generative adversarial networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00453"},{"key":"ref_40","unstructured":"Zhang, H., Goodfellow, I., Metaxas, D., and Odena, A. (2019, January 9\u201315). Self-attention generative adversarial networks. Proceedings of the International Conference on Machine Learning, PMLR, Long Beach, CA, USA."},{"key":"ref_41","unstructured":"Kingma, D.P. (2013). Auto-encoding variational bayes. arXiv."},{"key":"ref_42","first-page":"19667","article-title":"NVAE: A deep hierarchical variational autoencoder","volume":"33","author":"Vahdat","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_43","unstructured":"Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. (2015, January 7\u20139). Deep Unsupervised Learning using Nonequilibrium Thermodynamics. Proceedings of the 32nd International Conference on Machine Learning, Lille, France."},{"key":"ref_44","first-page":"6840","article-title":"Denoising diffusion probabilistic models","volume":"33","author":"Ho","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Zhang, L., Rao, A., and Agrawala, M. (2023, January 2\u20136). Adding conditional control to text-to-image diffusion models. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.00355"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Luo, Z., Gustafsson, F.K., Zhao, Z., Sj\u00f6lund, J., and Sch\u00f6n, T.B. (2023, January 18\u201322). Refusion: Enabling large-size realistic image restoration with latent-space diffusion models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPRW59228.2023.00169"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Li, T., Chang, H., Mishra, S., Zhang, H., Katabi, D., and Krishnan, D. (2023, January 18\u201322). Mage: Masked generative encoder to unify representation learning and image synthesis. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00213"},{"key":"ref_48","unstructured":"Khanna, S., Liu, P., Zhou, L., Meng, C., Rombach, R., Burke, M., Lobell, D., and Ermon, S. (2023). Diffusionsat: A generative foundation model for satellite imagery. arXiv."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Peebles, W., and Xie, S. (2023, January 2\u20136). Scalable diffusion models with transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.00387"},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"5737","DOI":"10.1109\/TIP.2023.3323799","article-title":"Txt2Img-MHN: Remote sensing image generation from text using modern Hopfield networks","volume":"32","author":"Xu","year":"2023","journal-title":"IEEE Trans. Image Process."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Sastry, S., Khanal, S., Dhakal, A., and Jacobs, N. (2024, January 17\u201321). Geosynth: Contextually-aware high-resolution satellite image synthesis. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPRW63382.2024.00051"},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Gatys, L.A., Ecker, A.S., and Bethge, M. (2016, January 27\u201330). Image style transfer using convolutional neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.265"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Li, C., and Wand, M. (2016, January 11\u201314). Precomputed real-time texture synthesis with markovian generative adversarial networks. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46487-9_43"},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Deng, Y., Tang, F., Dong, W., Ma, C., Pan, X., Wang, L., and Xu, C. (2022, January 18\u201324). Stytr2: Image style transfer with transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01104"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Brooks, T., Holynski, A., and Efros, A.A. (2023, January 18\u201322). InstructPix2Pix: Learning To Follow Image Editing Instructions. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01764"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Wang, Z., Zhao, L., and Xing, W. (2023, January 2\u20136). StyleDiffusion: Controllable Disentangled Style Transfer via Diffusion Models. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Paris, France.","DOI":"10.1109\/ICCV51070.2023.00706"},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Huang, N., Tang, F., Huang, H., Ma, C., Dong, W., and Xu, C. (2023, January 18\u201322). Inversion-based style transfer with diffusion models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00978"},{"key":"ref_58","first-page":"66860","article-title":"Styledrop: Text-to-image synthesis of any style","volume":"36","author":"Sohn","year":"2024","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Chung, J., Hyun, S., and Heo, J.P. (2024, January 17\u201321). Style Injection in Diffusion: A Training-free Approach for Adapting Large-scale Diffusion Models for Style Transfer. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Denver, CO, USA.","DOI":"10.1109\/CVPR52733.2024.00840"},{"key":"ref_60","first-page":"38154","article-title":"Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face","volume":"36","author":"Shen","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_61","unstructured":"Qin, J., Wu, J., Chen, W., Ren, Y., Li, H., Wu, H., Xiao, X., Wang, R., and Wen, S. (2024). Diffusiongpt: Llm-driven text-to-image generation system. arXiv."},{"key":"ref_62","unstructured":"Liu, Z., He, Y., Wang, W., Wang, W., Wang, Y., Chen, S., Zhang, Q., Yang, Y., Li, Q., and Yu, J. (2023). Internchat: Solving vision-centric tasks by interacting with chatbots beyond language. arXiv."},{"key":"ref_63","unstructured":"Wang, Z., Xie, E., Li, A., Wang, Z., Liu, X., and Li, Z. (2024). Divide and conquer: Language models can plan and self-correct for compositional text-to-image generation. arXiv."},{"key":"ref_64","first-page":"128374","article-title":"Genartist: Multimodal llm as an agent for unified image generation and editing","volume":"37","author":"Wang","year":"2024","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_65","unstructured":"Unreal, E. (2025, July 21). Unreal Engine. Available online: https:\/\/www.unrealengine.com\/en-us."},{"key":"ref_66","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-net: Convolutional networks for biomedical image segmentation. Proceedings of the International Conference on Medical image Computing and Computer-Assisted Intervention, Munich, Germany.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_67","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"Segnet: A deep convolutional encoder-decoder architecture for image segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell. (PAMI)"},{"key":"ref_68","doi-asserted-by":"crossref","unstructured":"Lin, G., Milan, A., Shen, C., and Reid, I. (2017, January 21\u201326). Refinenet: Multi-path refinement networks for high-resolution semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.549"},{"key":"ref_69","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs","volume":"40","author":"Chen","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_70","unstructured":"Chen, L.C., Papandreou, G., Schroff, F., and Adam, H. (2017). Rethinking atrous convolution for semantic image segmentation. arXiv."},{"key":"ref_71","first-page":"1140","article-title":"Segnext: Rethinking convolutional attention design for semantic segmentation","volume":"35","author":"Guo","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_72","doi-asserted-by":"crossref","unstructured":"Zheng, S., Lu, J., Zhao, H., Zhu, X., Luo, Z., Wang, Y., Fu, Y., Feng, J., Xiang, T., and Torr, P.H. (2021, January 20\u201325). Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR46437.2021.00681"},{"key":"ref_73","first-page":"12077","article-title":"SegFormer: Simple and efficient design for semantic segmentation with transformers","volume":"34","author":"Xie","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_74","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_75","doi-asserted-by":"crossref","unstructured":"Zheng, Z., Zhong, Y., Wang, J., and Ma, A. (2020, January 13\u201319). Foreground-aware relation network for geospatial object segmentation in high spatial resolution remote sensing imagery. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00415"},{"key":"ref_76","doi-asserted-by":"crossref","unstructured":"Xia, G.S., Bai, X., Ding, J., Zhu, Z., Belongie, S., Luo, J., Datcu, M., Pelillo, M., and Zhang, L. (2018, January 18\u201322). DOTA: A large-scale dataset for object detection in aerial images. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00418"},{"key":"ref_77","doi-asserted-by":"crossref","unstructured":"Johnson, J., Alahi, A., and Fei-Fei, L. (2016, January 11\u201314). Perceptual losses for real-time style transfer and super-resolution. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46475-6_43"},{"key":"ref_78","doi-asserted-by":"crossref","unstructured":"Huang, X., and Belongie, S. (2017, January 22\u201329). Arbitrary style transfer in real-time with adaptive instance normalization. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.167"},{"key":"ref_79","first-page":"26561","article-title":"Artistic style transfer with internal-external learning and contrastive learning","volume":"34","author":"Chen","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_80","first-page":"3","article-title":"Lora: Low-rank adaptation of large language models","volume":"1","author":"Hu","year":"2022","journal-title":"ICLR"},{"key":"ref_81","unstructured":"Shah, V., Ruiz, N., Cole, F., Lu, E., Lazebnik, S., Li, Y., and Jampani, V. (October, January 28). Ziplora: Any subject in any style by effectively merging loras. Proceedings of the European Conference on Computer Vision, Milano, Italy."},{"key":"ref_82","unstructured":"Liu, C., Shah, V., Cui, A., and Lazebnik, S. (2024). Unziplora: Separating content and style from a single image. arXiv."},{"key":"ref_83","doi-asserted-by":"crossref","unstructured":"Jones, M., Wang, S.Y., Kumari, N., Bau, D., and Zhu, J.Y. (2024, January 3\u20136). Customizing text-to-image models with a single image pair. Proceedings of the SIGGRAPH Asia 2024 Conference Papers, Tokyo, Japan.","DOI":"10.1145\/3680528.3687642"},{"key":"ref_84","unstructured":"Frenkel, Y., Vinker, Y., Shamir, A., and Cohen-Or, D. (October, January 28). Implicit style-content separation using b-lora. Proceedings of the European Conference on Computer Vision, Milano, Italy."},{"key":"ref_85","unstructured":"Chen, B., Zhao, B., Xie, H., Cai, Y., Li, Q., and Mao, X. (2025). Consislora: Enhancing content and style consistency for lora-based style transfer. arXiv."},{"key":"ref_86","doi-asserted-by":"crossref","first-page":"151","DOI":"10.1007\/s10994-009-5152-4","article-title":"A theory of learning from different domains","volume":"79","author":"Blitzer","year":"2010","journal-title":"Mach. Learn."},{"key":"ref_87","doi-asserted-by":"crossref","unstructured":"Khosla, A., Zhou, T., Malisiewicz, T., Efros, A.A., and Torralba, A. (2012, January 7\u201313). Undoing the damage of dataset bias. Proceedings of the European Conference on Computer Vision, Florence, Italy.","DOI":"10.1007\/978-3-642-33718-5_12"},{"key":"ref_88","unstructured":"Muandet, K., Balduzzi, D., and Sch\u00f6lkopf, B. (2013, January 16\u201321). Domain generalization via invariant feature representation. Proceedings of the International Conference on Machine Learning, PMLR, Atlanta, GA, USA."},{"key":"ref_89","doi-asserted-by":"crossref","unstructured":"Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P. (2017, January 24\u201328). Domain randomization for transferring deep neural networks from simulation to the real world. Proceedings of the 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada.","DOI":"10.1109\/IROS.2017.8202133"},{"key":"ref_90","first-page":"5339","article-title":"Generalizing to unseen domains via adversarial data augmentation","volume":"31","author":"Volpi","year":"2018","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_91","doi-asserted-by":"crossref","unstructured":"Tzeng, E., Hoffman, J., Saenko, K., and Darrell, T. (2017, January 21\u201326). Adversarial discriminative domain adaptation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.316"},{"key":"ref_92","doi-asserted-by":"crossref","first-page":"135","DOI":"10.1016\/j.neucom.2018.05.083","article-title":"Deep visual domain adaptation: A survey","volume":"312","author":"Wang","year":"2018","journal-title":"Neurocomputing"},{"key":"ref_93","doi-asserted-by":"crossref","unstructured":"Farahani, A., Voghoei, S., Rasheed, K., and Arabnia, H.R. (2021). A brief review of domain adaptation. Advances in Data Science and Information Engineering: Proceedings from ICDATA 2020 and IKE 2020, Springer.","DOI":"10.1007\/978-3-030-71704-9_65"},{"key":"ref_94","unstructured":"Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. (2023, January 1\u20135). React: Synergizing reasoning and acting in language models. Proceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda."},{"key":"ref_95","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_96","first-page":"439","article-title":"Implicit reparameterization gradients","volume":"31","author":"Figurnov","year":"2018","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_97","unstructured":"Song, J., Meng, C., and Ermon, S. (2020). Denoising diffusion implicit models. arXiv."},{"key":"ref_98","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning transferable visual models from natural language supervision. Proceedings of the International Conference on Machine Learning, PmLR, Virtual."},{"key":"ref_99","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_100","unstructured":"Loshchilov, I., and Hutter, F. (2017). Decoupled weight decay regularization. arXiv."},{"key":"ref_101","unstructured":"OpenMMLab (2025, October 24). MMSegmentation: OpenMMLab Semantic Segmentation Toolbox and Benchmark. Available online: https:\/\/github.com\/open-mmlab\/mmsegmentation."},{"key":"ref_102","unstructured":"Cordts, M., Omran, M., Ramos, S., Scharw\u00e4chter, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B. (2015, January 7\u201312). The cityscapes dataset. Proceedings of the CVPR Workshop on the Future of Datasets in Vision, Boston, MA, USA."},{"key":"ref_103","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_104","unstructured":"Zhu, L. (2025, October 25). THOP: PyTorch-OpCounter. Available online: https:\/\/pypi.org\/project\/thop\/."},{"key":"ref_105","unstructured":"Wang, J., Zheng, Z., Ma, A., Lu, X., and Zhong, Y. (2021). LoveDA: A remote sensing land-cover dataset for domain adaptive semantic segmentation. arXiv."},{"key":"ref_106","unstructured":"Anthropic (2025, October 26). Model Context Protocol: Getting Started. Available online: https:\/\/modelcontextprotocol.io\/docs\/getting-started\/intro."}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/11\/11\/395\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,7]],"date-time":"2025-11-07T10:03:39Z","timestamp":1762509819000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/11\/11\/395"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,6]]},"references-count":106,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2025,11]]}},"alternative-id":["jimaging11110395"],"URL":"https:\/\/doi.org\/10.3390\/jimaging11110395","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,11,6]]}}}