{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T15:38:17Z","timestamp":1778081897273,"version":"3.51.4"},"reference-count":83,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2025,1,24]],"date-time":"2025-01-24T00:00:00Z","timestamp":1737676800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"JST CREST","award":["JPMJCR20D3"],"award-info":[{"award-number":["JPMJCR20D3"]}]},{"name":"JST CREST","award":["JPMJFR216O"],"award-info":[{"award-number":["JPMJFR216O"]}]},{"name":"JST CREST","award":["JP22K12091"],"award-info":[{"award-number":["JP22K12091"]}]},{"name":"JST CREST","award":["JP23H00497"],"award-info":[{"award-number":["JP23H00497"]}]},{"name":"JST CREST","award":["JPMJSP2138"],"award-info":[{"award-number":["JPMJSP2138"]}]},{"name":"JST FOREST","award":["JPMJCR20D3"],"award-info":[{"award-number":["JPMJCR20D3"]}]},{"name":"JST FOREST","award":["JPMJFR216O"],"award-info":[{"award-number":["JPMJFR216O"]}]},{"name":"JST FOREST","award":["JP22K12091"],"award-info":[{"award-number":["JP22K12091"]}]},{"name":"JST FOREST","award":["JP23H00497"],"award-info":[{"award-number":["JP23H00497"]}]},{"name":"JST FOREST","award":["JPMJSP2138"],"award-info":[{"award-number":["JPMJSP2138"]}]},{"name":"JSPS KAKENHI","award":["JPMJCR20D3"],"award-info":[{"award-number":["JPMJCR20D3"]}]},{"name":"JSPS KAKENHI","award":["JPMJFR216O"],"award-info":[{"award-number":["JPMJFR216O"]}]},{"name":"JSPS KAKENHI","award":["JP22K12091"],"award-info":[{"award-number":["JP22K12091"]}]},{"name":"JSPS KAKENHI","award":["JP23H00497"],"award-info":[{"award-number":["JP23H00497"]}]},{"name":"JSPS KAKENHI","award":["JPMJSP2138"],"award-info":[{"award-number":["JPMJSP2138"]}]},{"name":"JST SPRING","award":["JPMJCR20D3"],"award-info":[{"award-number":["JPMJCR20D3"]}]},{"name":"JST SPRING","award":["JPMJFR216O"],"award-info":[{"award-number":["JPMJFR216O"]}]},{"name":"JST SPRING","award":["JP22K12091"],"award-info":[{"award-number":["JP22K12091"]}]},{"name":"JST SPRING","award":["JP23H00497"],"award-info":[{"award-number":["JP23H00497"]}]},{"name":"JST SPRING","award":["JPMJSP2138"],"award-info":[{"award-number":["JPMJSP2138"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Social biases in generative models have gained increasing attention. This paper proposes an automatic evaluation protocol for text-to-image generation, examining how gender bias originates and perpetuates in the generation process of Stable Diffusion. Using triplet prompts that vary by gender indicators, we trace presentations at several stages of the generation process and explore dependencies between prompts and images. Our findings reveal the bias persists throughout all internal stages of the generating process and manifests in the entire images. For instance, differences in object presence, such as different instruments and outfit preferences, are observed across genders and extend to overall image layouts. Moreover, our experiments demonstrate that neutral prompts tend to produce images more closely aligned with those from masculine prompts than with their female counterparts. We also investigate prompt-image dependencies to further understand how bias is embedded in the generated content. Finally, we offer recommendations for developers and users to mitigate this effect in text-to-image generation.<\/jats:p>","DOI":"10.3390\/jimaging11020035","type":"journal-article","created":{"date-parts":[[2025,1,24]],"date-time":"2025-01-24T11:50:47Z","timestamp":1737719447000},"page":"35","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Revealing Gender Bias from Prompt to Image in Stable Diffusion"],"prefix":"10.3390","volume":"11","author":[{"given":"Yankun","family":"Wu","sequence":"first","affiliation":[{"name":"D3 Center, Osaka University, Suita 565-0871, Osaka, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuta","family":"Nakashima","sequence":"additional","affiliation":[{"name":"D3 Center, Osaka University, Suita 565-0871, Osaka, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Noa","family":"Garcia","sequence":"additional","affiliation":[{"name":"D3 Center, Osaka University, Suita 565-0871, Osaka, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,1,24]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022, January 18\u201324). High-resolution image synthesis with latent diffusion models. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"ref_2","unstructured":"Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. (2022). Hierarchical text-conditional image generation with CLIP latents. arXiv."},{"key":"ref_3","unstructured":"Birhane, A., Prabhu, V.U., and Kahembwe, E. (2021). Multimodal datasets: Misogyny, pornography, and malignant stereotypes. arXiv."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Garcia, N., Hirota, Y., Wu, Y., and Nakashima, Y. (2023, January 17\u201324). Uncurated Image-Text Datasets: Shedding Light on Demographic Bias. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00672"},{"key":"ref_5","unstructured":"Birhane, A., Han, S., Boddeti, V., and Luccioni, S. (2023, January 10\u201316). Into the LAION\u2019s Den: Investigating Hate in Multimodal Datasets. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), New Orleans, LA, USA."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Bianchi, F., Kalluri, P., Durmus, E., Ladhak, F., Cheng, M., Nozza, D., Hashimoto, T., Jurafsky, D., Zou, J., and Caliskan, A. (2023, January 12\u201315). Easily accessible text-to-image generation amplifies demographic stereotypes at large scale. Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT), Chicago, IL, USA.","DOI":"10.1145\/3593013.3594095"},{"key":"ref_7","unstructured":"Luccioni, A.S., Akiki, C., Mitchell, M., and Jernite, Y. (2023, January 10\u201316). Stable bias: Analyzing societal representations in diffusion models. Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS 2023), New Orleans, LA, USA."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Ungless, E., Ross, B., and Lauscher, A. (2023, January 9\u201314). Stereotypes and Smut: The (Mis) representation of Non-cisgender Identities by Text-to-Image Models. Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), Toronto, ON, Canada.","DOI":"10.18653\/v1\/2023.findings-acl.502"},{"key":"ref_9","unstructured":"Carlini, N., Hayes, J., Nasr, M., Jagielski, M., Sehwag, V., Tramer, F., Balle, B., Ippolito, D., and Wallace, E. (2023, January 9\u201311). Extracting training data from diffusion models. Proceedings of the USENIX Security Symposium, Anaheim, CA, USA."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Katirai, A., Garcia, N., Ide, K., Nakashima, Y., and Kishimoto, A. (2023). Situating the social issues of image generation models in the model life cycle: A sociotechnical approach. arXiv.","DOI":"10.1007\/s43681-024-00517-3"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Somepalli, G., Singla, V., Goldblum, M., Geiping, J., and Goldstein, T. (2023, January 17\u201324). Diffusion art or digital forgery? investigating data replication in diffusion models. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00586"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Wang, S.Y., Efros, A.A., Zhu, J.Y., and Zhang, R. (2023, January 2\u20133). Evaluating Data Attribution for Text-to-Image Models. Proceedings of the International Conference on Computer Vision (ICCV), Paris, France.","DOI":"10.1109\/ICCV51070.2023.00661"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Cho, J., Zala, A., and Bansal, M. (2023, January 2\u20133). Dall-Eval: Probing the reasoning skills and social biases of text-to-image generation models. Proceedings of the International Conference on Computer Vision (ICCV), Paris, France.","DOI":"10.1109\/ICCV51070.2023.00283"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Wang, J., Liu, X.G., Di, Z., Liu, Y., and Wang, X.E. (2023, January 9\u201314). T2IAT: Measuring Valence and Stereotypical Biases in Text-to-Image Generation. Proceedings of the Association for Computational Linguistics (ACL), Toronto, ON, Canada.","DOI":"10.18653\/v1\/2023.findings-acl.160"},{"key":"ref_15","unstructured":"Seshadri, P., Singh, S., and Elazar, Y. (2023). The Bias Amplification Paradox in Text-to-Image Generation. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Naik, R., and Nushi, B. (2023, January 8\u201310). Social Biases through the Text-to-Image Generation Lens. Proceedings of the Conference on AI, Ethics, and Society (AIES),  Montreal, QC, Canada.","DOI":"10.1145\/3600211.3604711"},{"key":"ref_17","unstructured":"Lin, A., Paes, L.M., Tanneru, S.H., Srinivas, S., and Lakkaraju, H. (2023). Word-Level Explanations for Analyzing Bias in Text-to-Image Models. arXiv."},{"key":"ref_18","unstructured":"Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., and Askell, A. (2020, January 6\u201312). Language models are few-shot learners. Proceedings of the Conference on Neural Information Processing Systems (NeurlPS), Vancouver, BC, Canada."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Wu, Y., Nakashima, Y., and Garcia, N. (2024, January 21\u201323). Stable diffusion exposed: Gender bias from prompt to image. Proceedings of the Conference on AI, Ethics, and Society (AIES), San Jose, CA, USA.","DOI":"10.1609\/aies.v7i1.31754"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"139","DOI":"10.1145\/3422622","article-title":"Generative adversarial networks","volume":"63","author":"Goodfellow","year":"2020","journal-title":"Commun. ACM"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Tao, M., Tang, H., Wu, F., Jing, X.Y., Bao, B.K., and Xu, C. (2022, January 18\u201324). DF-GAN: A simple and effective baseline for text-to-image synthesis. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01602"},{"key":"ref_22","unstructured":"Reed, S., Akata, Z., Yan, X., Logeswaran, L., Schiele, B., and Lee, H. (2016, January 19\u201324). Generative adversarial text to image synthesis. Proceedings of the International Conference on Machine Learning (ICML), New York City, NY, USA."},{"key":"ref_23","unstructured":"Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. (2021, January 18\u201324). Zero-shot text-to-image generation. Proceedings of the International Conference on Machine Learning (ICML), Online."},{"key":"ref_24","unstructured":"Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., and Yang, H. (2021, January 6\u201314). CogView: Mastering Text-to-Image Generation via Transformers. Proceedings of the Advances in Neural Information Processing Systems (NeurlPS), Online."},{"key":"ref_25","unstructured":"Ding, M., Zheng, W., Hong, W., and Tang, J. (December, January 28). CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers. Proceedings of the Advances in Neural Information Processing Systems (NeurlPS), New Orleans, LA, USA."},{"key":"ref_26","unstructured":"Yu, J., Xu, Y., Koh, J.Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., and Ayan, B.K. (2022). Scaling Autoregressive Models for Content-Rich Text-to-Image Generation. arXiv."},{"key":"ref_27","unstructured":"Ho, J., Jain, A., and Abbeel, P. (2020, January 6\u201312). Denoising diffusion probabilistic models. Proceedings of the Neural Information Processing Systems (NeurlPS), Vancouver, BC, Canada."},{"key":"ref_28","unstructured":"Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E.L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., and Salimans, T. (December, January 28). Photorealistic text-to-image diffusion models with deep language understanding. Proceedings of the Neural Information Processing Systems (NeurlPS), New Orleans, LA, USA."},{"key":"ref_29","unstructured":"Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., and Cohen-or, D. (2023, January 1\u20135). Prompt-to-Prompt Image Editing with Cross-Attention Control. Proceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Lu, S., Liu, Y., and Kong, A.W.K. (2023, January 2\u20136). TF-ICON: Diffusion-Based Training-Free Cross-Domain Image Composition. Proceedings of the International Conference on Computer Vision (ICCV), Paris, France.","DOI":"10.1109\/ICCV51070.2023.00218"},{"key":"ref_31","unstructured":"Epstein, D., Jabri, A., Poole, B., Efros, A.A., and Holynski, A. (2023, January 10\u201316). Diffusion self-guidance for controllable image generation. Proceedings of the Advances in Neural Information Processing Systems (NeurlPS), New Orleans, LA, USA."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Gandikota, R., Materzynska, J., Fiotto-Kaufman, J., and Bau, D. (2023, January 2\u20136). Erasing concepts from diffusion models. Proceedings of the International Conference on Computer Vision (ICCV), Paris, France.","DOI":"10.1109\/ICCV51070.2023.00230"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Gandikota, R., Orgad, H., Belinkov, Y., Materzy\u0144ska, J., and Bau, D. (2023). Unified concept editing in diffusion models. arXiv.","DOI":"10.1109\/WACV57701.2024.00503"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Tang, R., Liu, L., Pandey, A., Jiang, Z., Yang, G., Kumar, K., Stenetorp, P., Lin, J., and Ture, F. (2023, January 9\u201314). What the DAAM: Interpreting stable diffusion using cross attention. Proceedings of the Association for Computational Linguistics (ACL), Toronto, ON, Canada.","DOI":"10.18653\/v1\/2023.acl-long.310"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Wu, W., Zhao, Y., Shou, M.Z., Zhou, H., and Shen, C. (2023, January 2\u20136). Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using diffusion models. Proceedings of the ICCV, Paris, France.","DOI":"10.1109\/ICCV51070.2023.00117"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Pnvr, K., Singh, B., Ghosh, P., Siddiquie, B., and Jacobs, D. (2023, January 2\u20136). LD-ZNet: A Latent Diffusion Approach for Text-Based Image Segmentation. Proceedings of the International Conference on Computer Vision (ICCV), Paris, France.","DOI":"10.1109\/ICCV51070.2023.00384"},{"key":"ref_37","unstructured":"Mandal, A., Leavy, S., and Little, S. (2023). Multimodal Composite Association Score: Measuring Gender Bias in Generative Multimodal Models. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Berg, H., Hall, S., Bhalgat, Y., Kirk, H., Shtedritski, A., and Bain, M. (2022, January 20\u201323). A Prompt Array Keeps the Bias Away: Debiasing Vision-Language Models with Adversarial Learning. Proceedings of the Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (AACL-IJNCLP), Taipei, Taiwan.","DOI":"10.18653\/v1\/2022.aacl-main.61"},{"key":"ref_39","unstructured":"Mannering, H. (2023). Analysing Gender Bias in Text-to-Image Models using Object Detection. arXiv."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Jiang, L., Turk, G., and Yang, D. (2023). Auditing gender presentation differences in text-to-image models. arXiv.","DOI":"10.1145\/3689904.3694710"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Wolfe, R., and Caliskan, A. (2022, January 1\u20133). American== white in multimodal language-and-image AI. Proceedings of the Conference on Artificial Intelligence, Ethics, and Society (AIES), Oxford, UK.","DOI":"10.1145\/3514094.3534136"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Wolfe, R., Yang, Y., Howe, B., and Caliskan, A. (2023, January 12\u201315). Contrastive language-vision ai models pretrained on web-scraped multimodal data exhibit sexual objectification bias. Proceedings of the Conference on Fairness, Accountability, and Transparency (FAccT), Chicago, IL, USA.","DOI":"10.1145\/3593013.3594072"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Zhang, C., Chen, X., Chai, S., Wu, C.H., Lagun, D., Beeler, T., and De la Torre, F. (2023, January 17\u201324). ITI-GEN: Inclusive Text-to-Image Generation. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/ICCV51070.2023.00367"},{"key":"ref_44","unstructured":"Hall, M., Ross, C., Williams, A., Carion, N., Drozdzal, M., and Soriano, A.R. (2023). DIG In: Evaluating Disparities in Image Generations with Indicators for Geographic Diversity. arXiv."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Basu, A., Babu, R.V., and Pruthi, D. (2023, January 4\u20136). Inspecting the Geographical Representativeness of Images from Text-to-Image Models. Proceedings of the International Conference on Computer Vision (ICCV), Paris, France.","DOI":"10.1109\/ICCV51070.2023.00474"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Liu, Z., Schaldenbrand, P., Okogwu, B.C., Peng, W., Yun, Y., Hundt, A., Kim, J., and Oh, J. (2024, January 17\u201321). SCoFT: Self-Contrastive Fine-Tuning for Equitable Image Generation. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01029"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Bakr, E.M., Sun, P., Shen, X., Khan, F.F., Li, L.E., and Elhoseiny, M. (2023, January 2\u20133). HRS-Bench: Holistic, reliable and scalable benchmark for text-to-image models. Proceedings of the International Conference on Computer Vision (ICCV), Paris, France.","DOI":"10.1109\/ICCV51070.2023.01834"},{"key":"ref_48","unstructured":"Teo, C., Abdollahzadeh, M., and Cheung, N.M.M. (2023, January 10\u201316). On measuring fairness in generative models. Proceedings of the Advances in Neural Information Processing Systems (NeurlPS), New Orleans, LA, USA."},{"key":"ref_49","unstructured":"Lee, T., Yasunaga, M., Meng, C., Mai, Y., Park, J.S., Gupta, A., Zhang, Y., Narayanan, D., Teufel, H.B., and Bellagente, M. (2023, January 10\u201316). Holistic Evaluation of Text-to-Image Models. Proceedings of the NeurlPS Datasets and Benchmarks Track, New Orleans, LA, USA."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Chinchure, A., Shukla, P., Bhatt, G., Salij, K., Hosanagar, K., Sigal, L., and Turk, M. (2023). TIBET: Identifying and Evaluating Biases in Text-to-Image Generative Models. arXiv.","DOI":"10.1007\/978-3-031-72986-7_25"},{"key":"ref_51","unstructured":"Friedrich, F., H\u00e4mmerl, K., Schramowski, P., Libovicky, J., Kersting, K., and Fraser, A. (2024). Multilingual Text-to-Image Generation Magnifies Gender Stereotypes and Prompt Engineering May Not Help You. arXiv."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Sathe, A., Jain, P., and Sitaram, S. (2024). A unified framework and dataset for assessing gender bias in vision-language models. arXiv.","DOI":"10.18653\/v1\/2024.findings-emnlp.66"},{"key":"ref_53","unstructured":"Luo, H., Huang, H., Deng, Z., Liu, X., Chen, R., and Liu, Z. (2024). BIGbench: A Unified Benchmark for Social Bias in Text-to-Image Generative Models Based on Multi-modal LLM. arXiv."},{"key":"ref_54","unstructured":"Chen, M., Liu, Y., Yi, J., Xu, C., Lai, Q., Wang, H., Ho, T.Y., and Xu, Q. (2024). Evaluating text-to-image generative models: An empirical study on human image synthesis. arXiv."},{"key":"ref_55","unstructured":"Wan, Y., and Chang, K.W. (2024). The Male CEO and the Female Assistant: Probing Gender Biases in Text-To-Image Models Through Paired Stereotype Test. arXiv."},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Wang, W., Bai, H., Huang, J.t., Wan, Y., Yuan, Y., Qiu, H., Peng, N., and Lyu, M.R. (2024). New Job, New Gender? Measuring the Social Bias in Image Generation Models. arXiv.","DOI":"10.1145\/3664647.3681433"},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"D\u2019Inc\u00e0, M., Peruzzo, E., Mancini, M., Xu, D., Goel, V., Xu, X., Wang, Z., Shi, H., and Sebe, N. (2024, January 17\u201321). OpenBias: Open-set Bias Detection in Text-to-Image Generative Models. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01162"},{"key":"ref_58","unstructured":"Mandal, A., Leavy, S., and Little, S. (October, January 29). Generated Bias: Auditing Internal Bias Dynamics of Text-To-Image Generative Models. Proceedings of the European Conference on Computer Vision (ECCV), Milan, Italy."},{"key":"ref_59","unstructured":"Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., and Chen, X. (2016, January 5\u201310). Improved techniques for training gans. Proceedings of the Advances in Neural Information Processing Systems (NeurlPS), Barcelona, Spain."},{"key":"ref_60","unstructured":"Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. (2017, January 4\u20139). GANs trained by a two time-scale update rule converge to a local nash equilibrium. Proceedings of the Advances in Neural Information Processing Systems (NeurlPS), Long Beach, CA, USA."},{"key":"ref_61","doi-asserted-by":"crossref","unstructured":"Vedantam, R., Lawrence Zitnick, C., and Parikh, D. (2015, January 7\u201312). CIDEr: Consensus-based image description evaluation. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299087"},{"key":"ref_62","unstructured":"Papineni, K., Roukos, S., Ward, T., and Zhu, W.J. (2002, January 6\u201312). BLEU: A method for automatic evaluation of machine translation. Proceedings of the Association for Computational Linguistics (ACL), Philadelphia, PA, USA."},{"key":"ref_63","doi-asserted-by":"crossref","unstructured":"Otani, M., Togashi, R., Sawai, Y., Ishigami, R., Nakashima, Y., Rahtu, E., Heikkil\u00e4, J., and Satoh, S. (2023, January 17\u201324). Toward verifiable and reproducible human evaluation for text-to-image generation. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01372"},{"key":"ref_64","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft COCO: Common objects in context. Proceedings of the European Conference on Computer Vision (ECCV), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Wang, Z.J., Montoya, E., Munechika, D., Yang, H., Hoover, B., and Chau, D.H. (2023, January 9\u201314). DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models. Proceedings of the Association for Computational Linguistics (ACL), Toronto, ON, Canada.","DOI":"10.18653\/v1\/2023.acl-long.51"},{"key":"ref_66","doi-asserted-by":"crossref","unstructured":"Young, P., Lai, A., Hodosh, M., and Hockenmaier, J. (2014, January 22\u201327). From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Proceedings of the Association for Computational Linguistics (ACL), Baltimore, MD, USA.","DOI":"10.1162\/tacl_a_00166"},{"key":"ref_67","doi-asserted-by":"crossref","unstructured":"Sharma, P., Ding, N., Goodman, S., and Soricut, R. (2018, January 15\u201320). Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset for Automatic Image Captioning. Proceedings of the Association for Computational Linguistics (ACL), Melbourne, VIC, Australia.","DOI":"10.18653\/v1\/P18-1238"},{"key":"ref_68","doi-asserted-by":"crossref","unstructured":"Sidorov, O., Hu, R., Rohrbach, M., and Singh, A. (2020, January 23\u201328). TextCaps: A dataset for image captioning with reading comprehension. Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK.","DOI":"10.1007\/978-3-030-58536-5_44"},{"key":"ref_69","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-Net: Convolutional networks for biomedical image segmentation. Proceedings of the Medical Image Computing and Computer-Assisted Intervention (MICCAI), Munich, Germany.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_70","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_71","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning transferable visual models from natural language supervision. Proceedings of the International Conference on Machine Learning (ICML), Online."},{"key":"ref_72","doi-asserted-by":"crossref","unstructured":"Caron, M., Touvron, H., Misra, I., J\u00e9gou, H., Mairal, J., Bojanowski, P., and Joulin, A. (2021, January 10\u201317). Emerging properties in self-supervised vision transformers. Proceedings of the International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"ref_73","doi-asserted-by":"crossref","unstructured":"Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K. (2023, January 17\u201324). DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.02155"},{"key":"ref_74","unstructured":"Ren, T., Liu, S., Zeng, A., Lin, J., Li, K., Cao, H., Chen, J., Huang, X., Chen, Y., and Yan, F. (2024). Grounded sam: Assembling open-world models for diverse visual tasks. arXiv."},{"key":"ref_75","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Huang, X., Ma, J., Li, Z., Luo, Z., Xie, Y., Qin, Y., Luo, T., Li, Y., and Liu, S. (2023). Recognize Anything: A Strong Image Tagging Model. arXiv.","DOI":"10.1109\/CVPRW63382.2024.00179"},{"key":"ref_76","doi-asserted-by":"crossref","unstructured":"Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., Li, C., Yang, J., Su, H., and Zhu, J. (2023). Grounding DINO: Marrying dino with grounded pre-training for open-set object detection. arXiv.","DOI":"10.1007\/978-3-031-72970-6_3"},{"key":"ref_77","doi-asserted-by":"crossref","unstructured":"Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A.C., and Lo, W.Y. (2023, January 2\u20133). Segment Anything. Proceedings of the International Conference on Computer Vision (ICCV), Paris, France.","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"ref_78","doi-asserted-by":"crossref","unstructured":"Zhao, J., Wang, T., Yatskar, M., Ordonez, V., and Chang, K.W. (2017, January 9\u201311). Men Also Like Shopping: Reducing Gender Bias Amplification using Corpus-level Constraints. Proceedings of the Empirical Methods in Natural Language Processing (EMNLP), Copenhagen, Denmark.","DOI":"10.18653\/v1\/D17-1323"},{"key":"ref_79","doi-asserted-by":"crossref","unstructured":"Hirota, Y., Nakashima, Y., and Garcia, N. (2022, January 21\u201324). Gender and racial bias in visual question answering datasets. Proceedings of the Conference on Fairness, Accountability, and Transparency (FAccT), Seoul, Republic of Korea.","DOI":"10.1145\/3531146.3533184"},{"key":"ref_80","unstructured":"Bird, S., Klein, E., and Loper, E. (2009). Natural Language Processing with Python: Analyzing Text with the Natural Language Toolkit, O\u2019Reilly Media Inc."},{"key":"ref_81","doi-asserted-by":"crossref","unstructured":"Wolfe, R., and Caliskan, A. (2022, January 21\u201324). Markedness in visual semantic AI. Proceedings of the Conference on Fairness, Accountability, and Transparency (FAccT), Seoul, Republic of Korea.","DOI":"10.1145\/3531146.3533183"},{"key":"ref_82","unstructured":"Agarwal, S., Krueger, G., Clark, J., Radford, A., Kim, J.W., and Brundage, M. (2021). Evaluating CLIP: Towards characterization of broader capabilities and downstream implications. arXiv."},{"key":"ref_83","doi-asserted-by":"crossref","unstructured":"Wolfe, R., Banaji, M.R., and Caliskan, A. (2022, January 21\u201324). Evidence for hypodescent in visual semantic AI. Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, Seoul, Republic of Korea.","DOI":"10.1145\/3531146.3533185"}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/11\/2\/35\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,8]],"date-time":"2025-10-08T10:35:52Z","timestamp":1759919752000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/11\/2\/35"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,24]]},"references-count":83,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2025,2]]}},"alternative-id":["jimaging11020035"],"URL":"https:\/\/doi.org\/10.3390\/jimaging11020035","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1,24]]}}}