{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T11:40:40Z","timestamp":1783510840621,"version":"3.55.0"},"reference-count":47,"publisher":"American Association for the Advancement of Science (AAAS)","content-domain":{"domain":["spj.science.org"],"crossmark-restriction":true},"short-container-title":["Intell Comput"],"published-print":{"date-parts":[[2023,1]]},"abstract":"<jats:p>Synthesizing vivid images with descriptive texts is gradually emerging as a frontier cross-domain generation task. However, it is obviously inadequate to generate the high-quality image with one single sentence accurately due to the information asymmetry between modalities, which needs external knowledge to balance the process. Moreover, the limited description of the entities in the sentence cannot guarantee the semantic consistency between text and generated image, causing the deficiency of details in foreground and background. Here, we propose a commonsense-driven generative adversarial network to generate photo-realistic images depending on entity-related commonsense knowledge. Commonsense-driven generative adversarial network contains 2 key commonsense-based modules: (a) Entity semantic augment is designed to enhance entity semantics with common sense for abating the information asymmetry, and (b) adaptive entity refinement is used to generate the high-resolution image guided by various commonsense knowledges in multistage for keeping text-image consistency. We demonstrated extensive synthetic cases on the widely used CUB-birds (Caltech-UCSD Birds-200-2011) dataset, where our model achieves competitive results compared to the other state-of-the-art models.<\/jats:p>","DOI":"10.34133\/icomputing.0017","type":"journal-article","created":{"date-parts":[[2023,1,20]],"date-time":"2023-01-20T14:37:55Z","timestamp":1674225475000},"update-policy":"https:\/\/doi.org\/10.34133\/aaas_crossmark_01","source":"Crossref","is-referenced-by-count":6,"title":["CD-GAN: Commonsense-Driven Generative Adversarial Network with Hierarchical Refinement for Text-to-Image Synthesis"],"prefix":"10.34133","volume":"2","author":[{"given":"Guokai","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Electrical and Information Engineering, Tianjin University, Tianjin 300072, China."},{"name":"The Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, Hefei 230088, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ning","family":"Xu","sequence":"additional","affiliation":[{"name":"School of Electrical and Information Engineering, Tianjin University, Tianjin 300072, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chenggang","family":"Yan","sequence":"additional","affiliation":[{"name":"The Institute of Information and Control, Hangzhou Dianzi University, Hangzhou 310018, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bolun","family":"Zheng","sequence":"additional","affiliation":[{"name":"The Institute of Information and Control, Hangzhou Dianzi University, Hangzhou 310018, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yulong","family":"Duan","sequence":"additional","affiliation":[{"name":"The 30th Research Institute of CETC, Chengdu, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bo","family":"Lv","sequence":"additional","affiliation":[{"name":"The 30th Research Institute of CETC, Chengdu, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"An-An","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Electrical and Information Engineering, Tianjin University, Tianjin 300072, China."},{"name":"The Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, Hefei 230088, China."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"221","published-online":{"date-parts":[[2023,2,24]]},"reference":[{"key":"e_1_3_3_2_2","doi-asserted-by":"crossref","first-page":"358","DOI":"10.1016\/j.neucom.2019.12.086","article-title":"Efficient discrete supervised hashing for large-scale cross-modal retrieval","volume":"385","author":"Yao T","year":"2020","unstructured":"Yao T, Han Y, Wang R, Kong X, Yan L, Fu H, Tian Q. Efficient discrete supervised hashing for large-scale cross-modal retrieval. Neurocomputing. 2020;385:358\u2013367.","journal-title":"Neurocomputing"},{"issue":"3","key":"e_1_3_3_3_2","doi-asserted-by":"crossref","first-page":"45","DOI":"10.1109\/MIS.2022.3169884","article-title":"An orthogonal subspace decomposition method for cross-modal retrieval","volume":"37","author":"Zeng Z","year":"2022","unstructured":"Zeng Z, Xu N, Mao W, Zeng D. An orthogonal subspace decomposition method for cross-modal retrieval. IEEE Intell Syst. 2022;37(3):45\u201353.","journal-title":"IEEE Intell Syst"},{"issue":"4","key":"e_1_3_3_4_2","doi-asserted-by":"crossref","first-page":"541","DOI":"10.1162\/neco.1989.1.4.541","article-title":"Backpropagation applied to handwritten zip code recognition","volume":"1","author":"LeCun Y","year":"1989","unstructured":"LeCun Y, Boser BE, Denker JS, Henderson D, Howard RE, Hubbard WE, Jackel LD. Backpropagation applied to handwritten zip code recognition. Neural Comput. 1989;1(4):541\u2013551.","journal-title":"Neural Comput"},{"issue":"2","key":"e_1_3_3_5_2","doi-asserted-by":"crossref","first-page":"179","DOI":"10.1207\/s15516709cog1402_1","article-title":"Finding structure in time","volume":"14","author":"Elman JL","year":"1990","unstructured":"Elman JL. Finding structure in time. Cogn Sci. 1990;14(2):179\u2013211.","journal-title":"Cogn Sci"},{"issue":"8","key":"e_1_3_3_6_2","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter S","year":"1997","unstructured":"Hochreiter S, Schmidhuber J. Long short-term memory. Neural Comput. 1997;9(8):1735\u20131780.","journal-title":"Neural Comput"},{"key":"e_1_3_3_7_2","unstructured":"Radford A Kim JW Hallacy C Ramesh A Goh G Agarwal S Sastry G. Learning transferable visual models from natural language supervision. Paper presented at: Proceedings of the 38th International Conference on Machine Learning; 2021 Jul 18\u201324; online. p. 8748\u20138763."},{"key":"e_1_3_3_8_2","unstructured":"Lu J Batra D Parikh D Lee S. ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Paper presented at: Proceedings of the 33rd international conference on neural information processing systems; 2019 Dec 3\u20138; Vancouver Canada. p. 13\u201323."},{"key":"e_1_3_3_9_2","doi-asserted-by":"crossref","unstructured":"Alberti C Ling J Collins M Reitter D. Fusion of detected objects in text for visual question answering. Paper presented at: 2019 Conference on Empirical Methods in Natural Language Processing and 9th International Joint Conference on Natural Language Processing; 2019 Nov 3\u20137; Hong Kong China. p. 2131\u20132140.","DOI":"10.18653\/v1\/D19-1219"},{"issue":"1","key":"e_1_3_3_10_2","doi-asserted-by":"crossref","first-page":"607","DOI":"10.1609\/aaai.v36i1.19940","article-title":"Attention-aligned transformer for image captioning","volume":"36","author":"Fei Z","year":"2022","unstructured":"Fei Z. Attention-aligned transformer for image captioning. AAAI. 2022;36(1):607\u2013615.","journal-title":"AAAI"},{"issue":"3","key":"e_1_3_3_11_2","doi-asserted-by":"crossref","first-page":"1552","DOI":"10.1109\/TPAMI.2020.3021209","article-title":"Semantic object accuracy for generative text-to-image synthesis","volume":"44","author":"Hinz T","year":"2022","unstructured":"Hinz T, Heinrich S, Wermter S. Semantic object accuracy for generative text-to-image synthesis. IEEE Trans Pattern Anal Mach Intell. 2022;44(3):1552\u20131565.","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"e_1_3_3_12_2","unstructured":"Goodfellow I Pouget-Abadie J Mirza M Xu B Warde-Farley D Ozair S Courville A Bengio Y. Generative adversarial nets. In: Advances in neural information processing systems . Curran Associates Inc.; 2014. p. 2672\u20132680."},{"key":"e_1_3_3_13_2","doi-asserted-by":"crossref","unstructured":"Niu G Li B Zhang Y Pu S. CAKE: A scalable commonsense-aware framework for multi-view knowledge graph completion. Paper presented at: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics; 2022 May 22\u201327; Dublin Ireland. p. 2867\u20132877.","DOI":"10.18653\/v1\/2022.acl-long.205"},{"key":"e_1_3_3_14_2","doi-asserted-by":"crossref","unstructured":"Speer R Chin J Havasi C. ConceptNet 5.5: An open multilingual graph of general knowledge. Paper presented at: Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence; 2017 Feb 4\u20139; San Francisco USA; p. 4444\u20134451.","DOI":"10.1609\/aaai.v31i1.11164"},{"key":"e_1_3_3_15_2","doi-asserted-by":"crossref","unstructured":"Qiao T Zhang J Xu D Tao D. Mirrorgan: Learning text-to-image generation by redescription. Paper presented at: 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2019; Long Beach Ca. p. 1505\u20131514.","DOI":"10.1109\/CVPR.2019.00160"},{"key":"e_1_3_3_16_2","doi-asserted-by":"crossref","unstructured":"Zhang H Xu T Li H. Zhang S Wang X Huang X Metaxas D. StackGAN: Text to photo-realistic image synthesis with stacked generative adversarial networks. Paper presented at: 2017 IEEE International Conference on Computer Vision (ICCV); 2017 Oct 22\u201329; Venice Italy. p. 5908\u20135916.","DOI":"10.1109\/ICCV.2017.629"},{"issue":"8","key":"e_1_3_3_17_2","doi-asserted-by":"crossref","first-page":"1947","DOI":"10.1109\/TPAMI.2018.2856256","article-title":"StackGAN++: Realistic image synthesis with stacked generative adversarial networks","volume":"41","author":"Zhang H","year":"2019","unstructured":"Zhang H, Xu T, Li H, Zhang S, Wang X, Huang X, Metaxas DN. StackGAN++: Realistic image synthesis with stacked generative adversarial networks. IEEE Trans Pattern Anal Mach Intell. 2019;41(8):1947\u20131962.","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"e_1_3_3_18_2","unstructured":"Wah C Branson S Welinder P Perona P Belongie S. The caltech-UCSD birds-200-2011 dataset. California Institute of Technology; 2011."},{"key":"e_1_3_3_19_2","unstructured":"Kingma DP Welling M. Auto-encoding variational bayes. Paper presented at: Proceedings of the 2nd International Conference on Learning Representations; 2014 Apr 14\u201316; Banff Canada."},{"key":"e_1_3_3_20_2","unstructured":"Kingma DP Mohamed S Rezende DJ Welling M. Semi-supervised learning with deep generative models. Paper presented at: Proceedings of the 27th International Conference on Neural Information Processing Systems; 2014 Dec 7\u201314; Montreal Canada. p. 3581\u20133589."},{"key":"e_1_3_3_21_2","unstructured":"Goodfellow IJ Pouget-Abadie J Mirza M Xu B Warde-Farley D Ozair S Courville A Bengio Y. Generative adversarial nets. In: Advances in neural information processing systems . Curran Associates Inc.; 2014. p. 2672\u20132680."},{"key":"e_1_3_3_22_2","unstructured":"Odena A Olah C Shlens J. Conditional image synthesis with auxiliary classifier Gans. Paper presented at: Proceedings of the 34th International Conference on Machine Learning ; 2013 June 16\u201321; Atlanta USA. p. 2642\u20132651."},{"key":"e_1_3_3_23_2","unstructured":"Reed SE Akata Z Yan X Logeswaran L Schiele B Lee H. Generative adversarial text to image synthesis. In: Balcan M and Weinberger KQ editors. Proceedings of the 33rd international conference on machine learning . PMLR; 2016. p. 1060\u20131069."},{"key":"e_1_3_3_24_2","unstructured":"Devlin J Chang MW Lee K Toutanova K. BERT: Pre-training of deep bidirectional transformers for language understanding. Paper presented at: Proceedings of NAACL-HLT; 2019 Jun 2\u20137; Minneapolis Minnesota."},{"key":"e_1_3_3_25_2","doi-asserted-by":"crossref","unstructured":"Pavllo D Lucchi A Hofmann T. Controlling style and semantics in weakly-supervised image generation. In: Computer vision \u2013 ECCV . Springer; 2020. p. 482\u2013499.","DOI":"10.1007\/978-3-030-58539-6_29"},{"key":"e_1_3_3_26_2","unstructured":"Vaswani A Shazeer N Parmar N Uszkoreit J Jones L Gomez AN Kaiser L Polosukhin I. Attention is all you need. Paper presented at: 31st Conference on Neural Information Processing Systems; 2017 Dec 4\u20139; Long Beach (CA) USA. p. 5998\u20136008."},{"key":"e_1_3_3_27_2","doi-asserted-by":"crossref","unstructured":"Xu T Zhang P Huang Q Zhang H Gan Z Huang X He X. Attngan: Fine-grained text to image generation with attentional generative adversarial networks. Paper presented at: 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition; 2018 Jun 18\u201323; Salt Lake City UT. p. 1316\u20131324.","DOI":"10.1109\/CVPR.2018.00143"},{"key":"e_1_3_3_28_2","unstructured":"Huang W Xu RYD Oppermann I. Realistic image generation using region-phrase attention. Paper presented at: Proceedings of the Eleventh Asian Conference on Machine Learning; 2017 Nov 15\u201317; Seoul Korea. p. 284\u2013299."},{"key":"e_1_3_3_29_2","doi-asserted-by":"crossref","unstructured":"Tan H Liu X Li X Zhang Y Yin B. Semantics-enhanced adversarial nets for text-to-image synthesis. Paper presented at: 2019 IEEE\/CVF International Conference on Computer Vision (ICCV); 2019 Oct 27\u2013Nov 2; Seoul Korea (South). p. 10500\u201310509.","DOI":"10.1109\/ICCV.2019.01060"},{"key":"e_1_3_3_30_2","unstructured":"Li B Qi X Lukasiewicz T Torr PHS. Controllable text-to-image generation. Paper presented at: 33rd Conference on Neural Information Processing Systems; 2019. p. 2063\u20132073."},{"key":"e_1_3_3_31_2","doi-asserted-by":"crossref","unstructured":"Niu T Feng F Li L Wang X. Image synthesis from locally related texts. Paper presented at: Proceedings of the 2020 international conference on multimedia retrieval; 2020 Jun 8\u201311; Dublin Ireland. p. 145\u2013153.","DOI":"10.1145\/3372278.3390684"},{"key":"e_1_3_3_32_2","doi-asserted-by":"crossref","unstructured":"Johnson J Gupta A Fei-Fei L. Image generation from scene graphs. Paper presented at: Proceedings of the IEEE conference on computer vision and pattern recognition; 2018 Jun 18\u201322; Salt Lake City (UT) USA. p. 1219\u20131228.","DOI":"10.1109\/CVPR.2018.00133"},{"key":"e_1_3_3_33_2","doi-asserted-by":"crossref","unstructured":"Bao J Chen D Wen F Li H Hua G. CVAE-GAN: Fine-grained image generation through asymmetric training. Paper presented at: 2017 IEEE International Conference on Computer Vision (ICCV); 2017 Oct 22\u201329; Venice Italy. p. 2764\u20132773.","DOI":"10.1109\/ICCV.2017.299"},{"key":"e_1_3_3_34_2","unstructured":"Huang H Li Z He R Sun Z Tan T. Introvae: Introspective variational autoencoders for photographic image synthesis. Paper presented at: 32nd Conference on Neural Information Processing Systems; 2018. p. 52\u201363."},{"key":"e_1_3_3_35_2","unstructured":"Huang H Li Z He R Sun Z Tan T. IntroVAE: Introspective variational autoencoders for photographic image synthesis. Paper presented at: Proceedings of the 31st International Conference on Neural Information Processing Systems; 2018 Nov 2\u20137; Montreal Canada. p. 52\u201363."},{"key":"e_1_3_3_36_2","unstructured":"Ho J Jain A Abbeel P. Denoising diffusion probabilistic models. Paper presented at: Proceedings of the 34th International Conference on Neural Information Processing Systems; 2020 Dec 6\u201312."},{"key":"e_1_3_3_37_2","unstructured":"Ramesh A Pavlov M Goh G Gray S Voss C Radford A Chen M Sutskever I. Zero-shot text-to-image generation. Paper presented at: Proceedings of the 38th International Conference on Machine Learning; 2021 Jul 18\u201324; online. p. 8821\u20138831."},{"key":"e_1_3_3_38_2","doi-asserted-by":"crossref","unstructured":"Saharia C Chan W Saxena S Li L Whang J. Photorealistic text-to-image diffusion models with deep language understanding. arXiv. 2022. https:\/\/doi.org\/10.48550\/arXiv.2205.11487","DOI":"10.1145\/3528233.3530757"},{"key":"e_1_3_3_39_2","doi-asserted-by":"crossref","unstructured":"Li J Fei H Liu J Wu S Zhang M Teng C Ji D Li F. Unified named entity recognition as word-word relation classification. Paper presented at: The Thirty-Sixth AAAI Conference on Artificial Intelligence; 2022 Feb 22\u2013Mar 1; Vancouver Canada. p. 10965\u201310973.","DOI":"10.1609\/aaai.v36i10.21344"},{"key":"e_1_3_3_40_2","doi-asserted-by":"crossref","unstructured":"Lin BY Chen X Chen J Ren X. Kagnet: Knowledge-aware graph networks for commonsense reasoning. In: Inui K Jiang J Ng V Wan X editors. Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP) . Association for Computational Linguistics; 2019. p. 2829\u20132839.","DOI":"10.18653\/v1\/D19-1282"},{"key":"e_1_3_3_41_2","doi-asserted-by":"crossref","unstructured":"Karras T Laine S Aila T. A style-based generator architecture for generative adversarial networks. Paper presented at: 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2021 Jun 15\u201320; Long Beach CA. p. 4217\u20134228.","DOI":"10.1109\/TPAMI.2020.2970919"},{"key":"e_1_3_3_42_2","doi-asserted-by":"crossref","unstructured":"Zhu M Pan P Chen W Yang Y. DM-GAN: Dynamic memory generative adversarial networks for text-to-image synthesis. Paper presented at: 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2019 Jun 15\u201320; Long Beach CA. p. 5802\u20135810.","DOI":"10.1109\/CVPR.2019.00595"},{"key":"e_1_3_3_43_2","doi-asserted-by":"crossref","unstructured":"Zhang C Peng Y. Stacking VAE and GAN for context-aware text-to-image generation. Paper presented at: IEEE Fourth International Conference on Multimedia Big Data; 2018 Sep 13\u201316; Xi'an China. p. 1\u20135.","DOI":"10.1109\/BigMM.2018.8499439"},{"key":"e_1_3_3_44_2","doi-asserted-by":"crossref","unstructured":"Ruan S Zhang Y Zhang K Fan Y Tang F Liu Q Chen E. DAE-GAN: Dynamic aspect-aware GAN for text-to-image synthesis. Paper presented at: 2021 IEEE\/CVF International Conference on Computer Vision (ICCV); 2021 Oct 10\u201317; Montreal QC Canada. p. 13940\u201313949.","DOI":"10.1109\/ICCV48922.2021.01370"},{"key":"e_1_3_3_45_2","doi-asserted-by":"crossref","unstructured":"Tao M Tang H Wu F Jing X Bao B-K Xu C. DF-GAN: A simple and effective baseline for text-to-image synthesis. Paper presented at: 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2022 Jun 18\u201324; New Orleans LA. p. 16494\u201316504.","DOI":"10.1109\/CVPR52688.2022.01602"},{"key":"e_1_3_3_46_2","unstructured":"Salimans T Goodfellow IJ Zaremba W Cheung V Radford A Chen X. Improved techniques for training GANs. In: Advances in neural information processing systems . Curran Associates Inc.; 2016. p. 2226\u20132234."},{"key":"e_1_3_3_47_2","unstructured":"Heusel M Ramsauer H Unterthiner T Nessler B Hochreiter S. GANs trained by a two time-scale update rule converge to a local nash equilibrium. Paper presented at: Proceedings of the 31st International Conference on Neural Information Processing Systems; 2017 Dec 4\u20139; Long Beach (CA) USA. p. 6626\u20136637."},{"key":"e_1_3_3_48_2","unstructured":"Reed SE Akata Z Mohan S Tenka S Schiele B Lee H. Learning what and where to draw. In: Lee DD Sugiyama M von Luxburg U Guyon I Garnett R editors. Proceedings of the 30th international conference on neural information processing systems . Curran Associates Inc.; 2016. p. 217\u2013225."}],"container-title":["Intelligent Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/spj.science.org\/doi\/pdf\/10.34133\/icomputing.0017","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,1,9]],"date-time":"2024-01-09T14:11:42Z","timestamp":1704809502000},"score":1,"resource":{"primary":{"URL":"https:\/\/spj.science.org\/doi\/10.34133\/icomputing.0017"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1]]},"references-count":47,"alternative-id":["10.34133\/icomputing.0017"],"URL":"https:\/\/doi.org\/10.34133\/icomputing.0017","relation":{},"ISSN":["2771-5892"],"issn-type":[{"value":"2771-5892","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1]]},"assertion":[{"value":"2022-10-10","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-01-12","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-24","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"0017"}}