{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,18]],"date-time":"2026-04-18T14:45:37Z","timestamp":1776523537064,"version":"3.51.2"},"reference-count":36,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2025,5,22]],"date-time":"2025-05-22T00:00:00Z","timestamp":1747872000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100004837","name":"Ministerio de Ciencia, Innovaci\u00f3n y Universidades of the Spanish Government","doi-asserted-by":"publisher","award":["PID2021-125051OB-I00"],"award-info":[{"award-number":["PID2021-125051OB-I00"]}],"id":[{"id":"10.13039\/501100004837","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004837","name":"Ministerio de Ciencia, Innovaci\u00f3n y Universidades of the Spanish Government","doi-asserted-by":"publisher","award":["TEC-2024\/COM-322"],"award-info":[{"award-number":["TEC-2024\/COM-322"]}],"id":[{"id":"10.13039\/501100004837","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Regional Government of Madrid in Spain","award":["PID2021-125051OB-I00"],"award-info":[{"award-number":["PID2021-125051OB-I00"]}]},{"name":"Regional Government of Madrid in Spain","award":["TEC-2024\/COM-322"],"award-info":[{"award-number":["TEC-2024\/COM-322"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Semantic segmentation is a computer vision task where classification is performed at the pixel level. Due to this, the process of labeling images for semantic segmentation is time-consuming and expensive. To mitigate this cost there has been a surge in the use of synthetically generated data\u2014usually created using simulators or videogames\u2014which, in combination with domain adaptation methods, can effectively learn how to segment real data. Still, these datasets have a particular limitation: due to their closed-set nature, it is not possible to include novel classes without modifying the tool used to generate them, which is often not public. Concurrently, generative models have made remarkable progress, particularly with the introduction of diffusion models, enabling the creation of high-quality images from text prompts without additional supervision. In this work, we propose an unsupervised pipeline that leverages Stable Diffusion and Segment Anything Module to generate class examples with an associated segmentation mask, and a method to integrate generated cutouts for novel classes in semantic segmentation datasets, all with minimal user input. Our approach aims to improve the performance of unsupervised domain adaptation methods by introducing novel samples into the training data without modifications to the underlying algorithms. With our methods, we show how models can not only effectively learn how to segment novel classes, with an average performance of 51% intersection over union for novel classes, but also reduce errors for other, already existing classes, reaching a higher performance level overall.<\/jats:p>","DOI":"10.3390\/jimaging11060172","type":"journal-article","created":{"date-parts":[[2025,5,22]],"date-time":"2025-05-22T04:32:37Z","timestamp":1747888357000},"page":"172","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Unsupervised Class Generation to Expand Semantic Segmentation Datasets"],"prefix":"10.3390","volume":"11","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6610-1566","authenticated-orcid":false,"given":"Javier","family":"Montalvo","sequence":"first","affiliation":[{"name":"Video Processing and Understanding Lab, Escuela Polit\u00e9cnica Superior, Universidad Aut\u00f3noma de Madrid, 28049 Madrid, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1705-3972","authenticated-orcid":false,"given":"\u00c1lvaro","family":"Garc\u00eda-Mart\u00edn","sequence":"additional","affiliation":[{"name":"Video Processing and Understanding Lab, Escuela Polit\u00e9cnica Superior, Universidad Aut\u00f3noma de Madrid, 28049 Madrid, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7199-698X","authenticated-orcid":false,"given":"Pablo","family":"Carballeira","sequence":"additional","affiliation":[{"name":"Video Processing and Understanding Lab, Escuela Polit\u00e9cnica Superior, Universidad Aut\u00f3noma de Madrid, 28049 Madrid, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4999-2851","authenticated-orcid":false,"given":"Juan C.","family":"SanMiguel","sequence":"additional","affiliation":[{"name":"Video Processing and Understanding Lab, Escuela Polit\u00e9cnica Superior, Universidad Aut\u00f3noma de Madrid, 28049 Madrid, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,5,22]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B. (2016, January 27\u201330). The Cityscapes Dataset for Semantic Urban Scene Understanding. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.350"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Ros, G., Sellart, L., Materzynska, J., Vazquez, D., and Lopez, A.M. (2016, January 27\u201330). The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.352"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Richter, S.R., Vineet, V., Roth, S., and Koltun, V. (2016, January 11\u201314). Playing for Data: Ground Truth from Computer Games. Proceedings of the IEEE European Conference Computer Vision (ECCV), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46475-6_7"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"54296","DOI":"10.1109\/ACCESS.2023.3277785","article-title":"Survey on Unsupervised Domain Adaptation for Semantic Segmentation for Visual Perception in Automated Driving","volume":"11","author":"Schwonberg","year":"2023","journal-title":"IEEE Access"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022). High-Resolution Image Synthesis with Latent Diffusion Models. arXiv.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Jia, Y., Hoyer, L., Huang, S., Wang, T., Van Gool, L., Schindler, K., and Obukhov, A. (2023). DGInStyle: Domain-Generalizable Semantic Segmentation with Image Diffusion Models and Stylized Semantic Control. arXiv.","DOI":"10.1007\/978-3-031-72933-1_6"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Hoyer, L., Dai, D., and Van Gool, L. (2022, January 18\u201324). DAFormer: Improving Network Architectures and Training Strategies for Domain-Adaptive Semantic Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00969"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A.C., and Lo, W.Y. (2023, January 2\u20136). Segment anything. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"ref_9","unstructured":"Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., and Koltun, V. (2017, January 13\u201315). CARLA: An Open Urban Driving Simulator. Proceedings of the 1st Annual Conference on Robot Learning, Mountain View, CA, USA."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Rong, G., Shin, B.H., Tabatabaee, H., Lu, Q., Lemke, S., Mo\u017eeiko, M., Boise, E., Uhm, G., Gerow, M., and Mehta, S. (2020, January 20\u201323). Lgsvl simulator: A high fidelity simulator for autonomous driving. Proceedings of the 2020 IEEE 23rd International Conference on Intelligent Transportation Systems (ITSC), Rhodes, Greece.","DOI":"10.1109\/ITSC45102.2020.9294422"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Shah, S., Dey, D., Lovett, C., and Kapoor, A. (2017, January 12\u201315). Airsim: High-fidelity visual and physical simulation for autonomous vehicles. Proceedings of the Field and Service Robotics: Results of the 11th International Conference, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-67361-5_40"},{"key":"ref_12","unstructured":"Xiao, A., Huang, J., Guan, D., Zhan, F., and Lu, S. (March, January 22). Transfer learning from synthetic to real lidar point cloud for semantic segmentation. Proceedings of the AAAI Conference on Artificial Intelligence, Virtual."},{"key":"ref_13","first-page":"54683","article-title":"Datasetdm: Synthesizing data with perception annotations using diffusion models","volume":"36","author":"Wu","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Hoyer, L., Dai, D., and Van Gool, L. (2022, January 23\u201327). HRDA: Context-Aware High-Resolution Domain-Adaptive Semantic Segmentation. Proceedings of the IEEE European Conference on Computer Vision (ECCV), Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-20056-4_22"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Hoyer, L., Dai, D., Wang, H., and Van Gool, L. (2023, January 17\u201324). MIC: Masked Image Consistency for Context-Enhanced Domain Adaptation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01128"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Tranheden, W., Olsson, V., Pinto, J., and Svensson, L. (2020, January 1\u20135). DACS: Domain Adaptation via Cross-domain Mixed Sampling. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV), Snowmass Village, CO, USA.","DOI":"10.1109\/WACV48630.2021.00142"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"101248","DOI":"10.1109\/ACCESS.2022.3205414","article-title":"A Novel Unsupervised Domain Adaption Method for Depth-Guided Semantic Segmentation Using Coarse-to-Fine Alignment","volume":"10","author":"Nam","year":"2022","journal-title":"IEEE Access"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1966","DOI":"10.1109\/TITS.2023.3314680","article-title":"Self-Supervised Adversarial Learning for Domain Adaptation of Pavement Distress Classification","volume":"25","author":"Wu","year":"2024","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"20217","DOI":"10.1109\/TITS.2022.3176397","article-title":"ParaUDA: Invariant Feature Learning with Auxiliary Synthetic Samples for Unsupervised Domain Adaptation","volume":"23","author":"Zhang","year":"2022","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Gao, L., Zhang, J., Zhang, L., and Tao, D. (2021, January 20\u201324). DSP: Dual Soft-Paste for Unsupervised Domain Adaptive Semantic Segmentation. Proceedings of the ACM International Conference on Multimedia (MM), Chengdu, China.","DOI":"10.1145\/3474085.3475186"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"2420","DOI":"10.1109\/TMM.2019.2953375","article-title":"Weighted and Class-Specific Maximum Mean Discrepancy for Unsupervised Domain Adaptation","volume":"22","author":"Yan","year":"2019","journal-title":"IEEE Trans. Multimed."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Yan, H., Ding, Y., Li, P., Wang, Q., Xu, Y., and Zuo, W. (2017, January 21\u201326). Mind the Class Weight Bias: Weighted Maximum Mean Discrepancy for Unsupervised Domain Adaptation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.107"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Fan, Q., Shen, X., Ying, S., and Du, S. (2024, January 24\u201327). OTCLDA: Optimal Transport and Contrastive Learning for Domain Adaptive Semantic Segmentation. Proceedings of the IEEE Transactions on Intelligent Transportation Systems, Edmonton, AB, Canada.","DOI":"10.1109\/TITS.2024.3399399"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"106172","DOI":"10.1016\/j.engappai.2023.106172","article-title":"Visual Domain Adaptation through Locality Information","volume":"123","author":"Devika","year":"2023","journal-title":"Eng. Appl. Artif. Intell."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Tsai, Y.H., Hung, W.C., Schulter, S., Sohn, K., Yang, M.H., and Chandraker, M. (2018, January 18\u201323). Learning to Adapt Structured Output Space for Semantic Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00780"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Vu, T.H., Jain, H., Bucher, M., Cord, M., and P\u00e9rez, P. (2019, January 15\u201320). ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00262"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Marcos-Manch\u00f3n, P., Alcover-Couso, R., SanMiguel, J.C., and Mart\u00ednez, J.M. (2024, January 16\u201322). Open-Vocabulary Attention Maps with Token Optimization for Semantic Segmentation in Diffusion Models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00883"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Tang, R., Liu, L., Pandey, A., Jiang, Z., Yang, G., Kumar, K., Stenetorp, P., Lin, J., and Ture, F. (2022). What the daam: Interpreting stable diffusion using cross attention. arXiv.","DOI":"10.18653\/v1\/2023.acl-long.310"},{"key":"ref_29","unstructured":"Kr\u00e4henb\u00fchl, P., and Koltun, V. (2011). Efficient inference in fully connected crfs with gaussian edge potentials. Adv. Neural Inf. Process. Syst., 24."},{"key":"ref_30","first-page":"301","article-title":"The third criterion: Compactness as a procedural safeguard against partisan gerrymandering","volume":"9","author":"Polsby","year":"1991","journal-title":"Yale L. Pol\u2019y Rev."},{"key":"ref_31","first-page":"179","article-title":"A method of assigning numerical and percentage values to the degree of roundness of sand grains","volume":"1","author":"Cox","year":"1927","journal-title":"J. Paleontol."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Wang, Z., Guo, S., Shang, X., and Ye, X. (2023, January 6\u20139). Pseudo-label Assisted Optimization of Multi-branch Network for Cross-domain Person Re-identification. Proceedings of the IEEE International Conference on Mechatronics and Automation (ICMA), Harbin, China.","DOI":"10.1109\/ICMA57826.2023.10215862"},{"key":"ref_33","unstructured":"Montalvo, J., Alcover-Couso, R., Carballeira, P., Garc\u00eda-Mart\u00edn, \u00c1., SanMiguel, J.C., and Escudero-Vi\u00f1olo, M. (2024). Leveraging Contrastive Learning for Semantic Segmentation with Consistent Labels Across Varying Appearances. arXiv."},{"key":"ref_34","unstructured":"Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., and Bhosale, S. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","article-title":"The PASCAL Visual Object Classes (VOC) challenge","volume":"88","author":"Everingham","year":"2010","journal-title":"Int. J. Comput. Vis."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P. (2017, January 24\u201328). Domain randomization for transferring deep neural networks from simulation to the real world. Proceedings of the 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada.","DOI":"10.1109\/IROS.2017.8202133"}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/11\/6\/172\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T17:38:03Z","timestamp":1760031483000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/11\/6\/172"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,5,22]]},"references-count":36,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2025,6]]}},"alternative-id":["jimaging11060172"],"URL":"https:\/\/doi.org\/10.3390\/jimaging11060172","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,5,22]]}}}