{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,5]],"date-time":"2026-07-05T05:47:21Z","timestamp":1783230441220,"version":"3.54.6"},"reference-count":29,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2021,3,2]],"date-time":"2021-03-02T00:00:00Z","timestamp":1614643200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,3,2]],"date-time":"2021-03-02T00:00:00Z","timestamp":1614643200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2021,5]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>We present two new metrics for evaluating generative models in the class-conditional image generation setting. These metrics are obtained by generalizing the two most popular unconditional metrics: the Inception Score (IS) and the Fr\u00e9chet Inception Distance (FID). A theoretical analysis shows the motivation behind each proposed metric and links the novel metrics to their unconditional counterparts. The link takes the form of a product in the case of IS or an upper bound in the FID case. We provide an extensive empirical evaluation, comparing the metrics to their unconditional variants and to other metrics, and utilize them to analyze existing generative models, thus providing additional insights about their performance, from unlearned classes to mode collapse.<\/jats:p>","DOI":"10.1007\/s11263-020-01424-w","type":"journal-article","created":{"date-parts":[[2021,3,2]],"date-time":"2021-03-02T09:03:47Z","timestamp":1614675827000},"page":"1712-1731","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":49,"title":["Evaluation Metrics for Conditional Image Generation"],"prefix":"10.1007","volume":"129","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0524-7689","authenticated-orcid":false,"given":"Yaniv","family":"Benny","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tomer","family":"Galanti","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sagie","family":"Benaim","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lior","family":"Wolf","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,3,2]]},"reference":[{"key":"1424_CR1","unstructured":"Arjovsky, M., Chintala, S., & Bottou, L. (2017). Wasserstein generative adversarial networks. In D.\u00a0Precup, & Y. W. Teh (Eds.), Proceedings of the 34th international conference on machine learning, proceedings of machine learning research (Vol.\u00a070, pp. 214\u2013223). PMLR, International Convention Centre, Sydney, Australia."},{"key":"1424_CR2","unstructured":"Bi\u0144kowski, M., Sutherland, D. J., Arbel, M., & Gretton, A. (2018). Demystifying mmd gans. arXiv preprint arXiv:1801.01401."},{"key":"1424_CR3","unstructured":"Brock, A., Donahue, J., & Simonyan, K. (2019). Large scale GAN training for high fidelity natural image synthesis. In International conference on learning representations. https:\/\/openreview.net\/forum?id=B1xsqj09Fm"},{"key":"1424_CR4","unstructured":"Chen, X., Duan, Y., Houthooft, R., Schulman, J., Sutskever, I., & Abbeel, P. (2016). Infogan: interpretable representation learning by information maximizing generative adversarial nets. In Advances in neural information processing systems (pp. 2172\u20132180)."},{"issue":"3","key":"1424_CR5","doi-asserted-by":"publisher","first-page":"450","DOI":"10.1016\/0047-259X(82)90077-X","volume":"12","author":"DC Dowson","year":"1982","unstructured":"Dowson, D. C., & Landau, B. V. (1982). The fr\u00e9chet distance between multivariate normal distributions. Journal of Multivariate Analysis, 12(3), 450\u2013455.","journal-title":"Journal of Multivariate Analysis"},{"key":"1424_CR6","unstructured":"Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial nets. In Proceedings of the 27th international conference on neural information processing systems - Volume 2, NIPS\u201914 (pp. 2672\u20132680). Cambridge, MA: MIT Press."},{"key":"1424_CR7","unstructured":"Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., & Courville, A. (2017). Improved training of wasserstein gans. In Proceedings of the 31st international conference on neural information processing systems, NIPS\u201917 (pp. 5769\u20135779). Red Hook, NY: Curran Associates Inc."},{"key":"1424_CR8","unstructured":"Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., & Hochreiter, S. (2017). Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in neural information processing systems (pp. 6626\u20136637)."},{"key":"1424_CR9","doi-asserted-by":"crossref","unstructured":"Huang, X., Liu, M. Y., Belongie, S., & Kautz, J. (2018). Multimodal unsupervised image-to-image translation. In Proceedings of the European conference on computer vision (ECCV) (pp. 172\u2013189).","DOI":"10.1007\/978-3-030-01219-9_11"},{"key":"1424_CR10","doi-asserted-by":"crossref","unstructured":"Isola, P., Zhu, J. Y., Zhou, T., & Efros, A. A. (2017). Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1125\u20131134).","DOI":"10.1109\/CVPR.2017.632"},{"key":"1424_CR11","doi-asserted-by":"crossref","unstructured":"Johnson, J., Gupta, A., & Fei-Fei, L. (2018). Image generation from scene graphs. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1219\u20131228).","DOI":"10.1109\/CVPR.2018.00133"},{"key":"1424_CR12","unstructured":"Karras, T., Aila, T., Laine, S., & Lehtinen, J. (2018). Progressive growing of GANs for improved quality, stability, and variation. In International conference on learning representations. https:\/\/openreview.net\/forum?id=Hk99zCeAb."},{"key":"1424_CR13","doi-asserted-by":"crossref","unstructured":"Karras, T., Laine, S., & Aila, T. (2019). A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 4401\u20134410).","DOI":"10.1109\/CVPR.2019.00453"},{"key":"1424_CR14","doi-asserted-by":"crossref","unstructured":"Karras, T., Laine, S., Aittala, M., Hellsten, J., Lehtinen, J., & Aila, T. (2019). Analyzing and improving the image quality of StyleGAN. CoRR, abs\/1912.04958.","DOI":"10.1109\/CVPR42600.2020.00813"},{"key":"1424_CR15","unstructured":"Krizhevsky, A., Nair, V., & Hinton, G. (2010). Cifar-10 (Canadian Institute for Advanced Research). http:\/\/www.cs.toronto.edu\/~kriz\/cifar.html."},{"key":"1424_CR16","unstructured":"LeCun, Y., & Cortes, C. (2010). MNIST handwritten digit database. http:\/\/yann.lecun.com\/exdb\/mnist\/."},{"key":"1424_CR17","unstructured":"Mirza, M., & Osindero, S. (2014). Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784."},{"key":"1424_CR18","unstructured":"Miyato, T., & Koyama, M. (2018). cgans with projection discriminator. arXiv preprint arXiv:1802.05637."},{"key":"1424_CR19","unstructured":"Odena, A. (2016). Semi-supervised learning with generative adversarial networks. arXiv preprint arXiv:1606.01583."},{"key":"1424_CR20","unstructured":"Odena, A., Olah, C., & Shlens, J. (2017). Conditional image synthesis with auxiliary classifier gans. In Proceedings of the 34th international conference on machine learning-Volume 70 (pp. 2642\u20132651). JMLR.org."},{"key":"1424_CR21","unstructured":"Ravuri, S., & Vinyals, O. (2019). Classification accuracy score for conditional generative models. In Advances in neural information processing systems (pp. 12268\u201312279)."},{"issue":"3","key":"1424_CR22","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","volume":"115","author":"O Russakovsky","year":"2015","unstructured":"Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., et al. (2015). Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 115(3), 211\u2013252.","journal-title":"International Journal of Computer Vision"},{"key":"1424_CR23","unstructured":"Salimans, T., Goodfellow, I., Zaremba, W., Cheung, V., Radford, A., & Chen, X. (2016). Improved techniques for training gans. In Advances in neural information processing systems (pp. 2234\u20132242)."},{"key":"1424_CR24","unstructured":"Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556."},{"key":"1424_CR25","doi-asserted-by":"crossref","unstructured":"Singh, K.K., Ojha, U., & Lee, Y.J. (2019). Finegan: Unsupervised hierarchical disentanglement for fine-grained object generation and discovery. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 6490\u20136499).","DOI":"10.1109\/CVPR.2019.00665"},{"key":"1424_CR26","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., & Wojna, Z. (2015). Rethinking the inception architecture for computer vision. In 2016 IEEE conference on computer vision and pattern recognition (CVPR) (pp. 2818\u20132826).","DOI":"10.1109\/CVPR.2016.308"},{"key":"1424_CR27","doi-asserted-by":"crossref","unstructured":"Xu, T., Zhang, P., Huang, Q., Zhang, H., Gan, Z., Huang, X., & He, X. (2018). Attngan: Fine-grained text to image generation with attentional generative adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1316\u20131324).","DOI":"10.1109\/CVPR.2018.00143"},{"key":"1424_CR28","doi-asserted-by":"crossref","unstructured":"Zhang, H., Xu, T., Li, H., Zhang, S., Wang, X., Huang, X., & Metaxas, D. N. (2017). Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks. In Proceedings of the IEEE international conference on computer vision (pp. 5907\u20135915).","DOI":"10.1109\/ICCV.2017.629"},{"key":"1424_CR29","doi-asserted-by":"crossref","unstructured":"Zhu, J. Y., Park, T., Isola, P., & Efros, A. A. (2017). Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE international conference on computer vision (pp. 2223\u20132232).","DOI":"10.1109\/ICCV.2017.244"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-020-01424-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-020-01424-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-020-01424-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,5,5]],"date-time":"2021-05-05T18:18:24Z","timestamp":1620238704000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-020-01424-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,3,2]]},"references-count":29,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2021,5]]}},"alternative-id":["1424"],"URL":"https:\/\/doi.org\/10.1007\/s11263-020-01424-w","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,3,2]]},"assertion":[{"value":"23 April 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"23 December 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 March 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}