{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T16:08:46Z","timestamp":1783094926857,"version":"3.54.6"},"reference-count":57,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2024,5,16]],"date-time":"2024-05-16T00:00:00Z","timestamp":1715817600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62325206, 61936005, 62206132"],"award-info":[{"award-number":["62325206, 61936005, 62206132"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Key Research and Development Program of Jiangsu Province","award":["BE2023016-4"],"award-info":[{"award-number":["BE2023016-4"]}]},{"name":"Natural Science Research Start-up Foundation of Recruiting Talents of Nanjing University of Posts and Telecommunications","award":["NY222113"],"award-info":[{"award-number":["NY222113"]}]},{"name":"Postgraduate Research & Practice Innovation Program of Jiangsu Province","award":["KYCX23_1023"],"award-info":[{"award-number":["KYCX23_1023"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,7,31]]},"abstract":"<jats:p>\n            Text-to-Image synthesis aims to generate an accurate and semantically consistent image from a given text description. However, it is difficult for existing generative methods to generate semantically complete images from a single piece of text. Some works try to expand the input text to multiple captions via retrieving similar descriptions of the input text from the training set but still fail to fill in missing image semantics. In this article, we propose a GAN-based approach to Imagine, Select, and Fuse for Text-to-image synthesis, named ISF-GAN. The proposed ISF-GAN contains Imagine Stage and Select and Fuse Stage to solve the above problems. First, the Imagine Stage proposes a text completion and enrichment module. This module guides a GPT-based model to enrich the text expression beyond the original dataset. Second, the Select and Fuse Stage selects qualified text descriptions and then introduces a cross-modal attentional mechanism to interact these different sentence embeddings with the image features at different scales. In short, our proposed model enriches the input text information for completing missing semantics and introduces a cross-modal attentional mechanism to maximize the utilization of enriched text information to generate semantically consistent images. Experimental results on CUB, Oxford-102, and CelebA-HQ datasets prove the effectiveness and superiority of the proposed network. Code is available at\n            <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/Feilingg\/ISF-GAN\">https:\/\/github.com\/Feilingg\/ISF-GAN<\/jats:ext-link>\n          <\/jats:p>","DOI":"10.1145\/3650033","type":"journal-article","created":{"date-parts":[[2024,3,28]],"date-time":"2024-03-28T13:13:12Z","timestamp":1711631592000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["ISF-GAN: Imagine, Select, and Fuse with GPT-Based Text Enrichment for Text-to-Image Synthesis"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2547-4646","authenticated-orcid":false,"given":"Yefei","family":"Sheng","sequence":"first","affiliation":[{"name":"Nanjing University of Posts and Telecommunications, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4662-7170","authenticated-orcid":false,"given":"Ming","family":"Tao","sequence":"additional","affiliation":[{"name":"Nanjing University of Posts and Telecommunications, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8662-9488","authenticated-orcid":false,"given":"Jie","family":"Wang","sequence":"additional","affiliation":[{"name":"Nanjing University of Posts and Telecommunications, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5956-831X","authenticated-orcid":false,"given":"Bing-Kun","family":"Bao*","sequence":"additional","affiliation":[{"name":"Nanjing University of Posts and Telecommunications, Nanjing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,5,16]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"crossref","unstructured":"Nikolich Alexandr Osliakova Irina Kudinova Tatyana Kappusheva Inessa and Puchkova Arina. 2021. Fine-tuning GPT-3 for Russian text summarization. Data Science and Intelligent Systems: Proceedings of 5th Computational Methods in Systems and Software 2 (2021) 748\u2013757.","DOI":"10.1007\/978-3-030-90321-3_61"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3137605"},{"key":"e_1_3_1_4_2","doi-asserted-by":"crossref","unstructured":"Jun Cheng and Fuxiang Wu. 2021. RiFeGAN2: Rich feature generation for text-to-image synthesis from constrained prior knowledge. IEEE Transactions on Circuits and Systems for Video Technology 32 8 (2021) 5187\u20135200.","DOI":"10.1109\/TCSVT.2021.3136857"},{"key":"e_1_3_1_5_2","first-page":"10911","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Cheng Jun","year":"2020","unstructured":"Jun Cheng, Fuxiang Wu, Yanling Tian, Lei Wang, and Dapeng Tao. 2020. RiFeGAN: Rich feature generation for text-to-image synthesis from prior knowledge. In IEEE Conference on Computer Vision and Pattern Recognition. 10911\u201310920."},{"key":"e_1_3_1_6_2","doi-asserted-by":"crossref","unstructured":"Zijun Deng Xiangteng He and Yuxin Peng. 2023. LFR-GAN: Local feature refinement based generative adversarial network for text-to-image generation. ACM Transactions on Multimedia Computing Communications and Applications 19 6 (2023) 1\u201318.","DOI":"10.1145\/3589002"},{"key":"e_1_3_1_7_2","unstructured":"Ming Ding Zhuoyi Yang Wenyi Hong Wendi Zheng Chang Zhou Da Yin Junyang Lin Xu Zou Zhou Shao Hongxia Yang and J. Tang. 2021. CogView: Mastering text-to-image generation via transformers. Advances in Neural Information Processing Systems 34 (2021) 19822\u201319835."},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3075997"},{"key":"e_1_3_1_9_2","first-page":"8312","volume-title":"AAAI Conference on Artificial Intelligence","volume":"33","author":"Gao Lianli","year":"2019","unstructured":"Lianli Gao, Daiyuan Chen, Jingkuan Song, Xing Xu, Dongxiang Zhang, and Heng Tao Shen. 2019. Perceptual pyramid adversarial networks for text-to-image synthesis. In AAAI Conference on Artificial Intelligence, Vol. 33. 8312\u20138319."},{"key":"e_1_3_1_10_2","first-page":"10696","volume-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Gu Shuyang","year":"2022","unstructured":"Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo. 2022. Vector quantized diffusion model for text-to-image synthesis. In IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10696\u201310706."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco_a_01060"},{"key":"e_1_3_1_12_2","article-title":"GANs trained by a two time-scale update rule converge to a local Nash equilibrium","volume":"30","author":"Heusel Martin","year":"2017","unstructured":"Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. Adv. Neural Inf. Process. Syst. 30 (2017).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_1_13_2","first-page":"159","volume-title":"IEEE 8th International Conference on Big Data Analytics (ICBDA\u201923)","author":"Huang Pingda","year":"2023","unstructured":"Pingda Huang, Yedan Liu, Chunjiang Fu, and Liang Zhao. 2023. Multi-semantic fusion generative adversarial network for text-to-image generation. In IEEE 8th International Conference on Big Data Analytics (ICBDA\u201923). IEEE, 159\u2013164."},{"key":"e_1_3_1_14_2","first-page":"358","volume-title":"IEEE Winter Conference on Applications of Computer Vision","author":"Joseph K. J.","year":"2019","unstructured":"K. J. Joseph, Arghya Pal, Sailaja Rajanala, and Vineeth N. Balasubramanian. 2019. C4Synth: Cross-caption cycle-consistent text-to-image synthesis. In IEEE Winter Conference on Applications of Computer Vision. 358\u2013366."},{"key":"e_1_3_1_15_2","first-page":"852","article-title":"Alias-free generative adversarial networks","volume":"34","author":"Karras Tero","year":"2021","unstructured":"Tero Karras, Miika Aittala, Samuli Laine, Erik H\u00e4rk\u00f6nen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2021. Alias-free generative adversarial networks. Adv. Neural Inf. Process. Syst. 34 (2021), 852\u2013863.","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_1_16_2","article-title":"Controllable text-to-image generation","volume":"32","author":"Li Bowen","year":"2019","unstructured":"Bowen Li, Xiaojuan Qi, Thomas Lukasiewicz, and Philip Torr. 2019. Controllable text-to-image generation. Adv. Neural Inf. Process. Syst. 32 (2019).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_1_17_2","unstructured":"Mingjie Li Po-Yao Huang Xiaojun Chang Junjie Hu Yi Yang and Alex Hauptmann. 2022. Video pivoting unsupervised multi-modal machine translation. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 3 (2022) 3918\u20133932."},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.2972856"},{"key":"e_1_3_1_19_2","unstructured":"Wenbo Li Pengchuan Zhang Lei Zhang Qiuyuan Huang Xiaodong He Siwei Lyu and Jianfeng Gao. 2019. Object-driven text-to-image synthesis via adversarial training. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. (2019) 12174\u201312182."},{"key":"e_1_3_1_20_2","doi-asserted-by":"crossref","unstructured":"Tsung-Yi Lin Michael Maire Serge Belongie James Hays Pietro Perona Deva Ramanan Piotr Doll\u00e1r and C. Lawrence Zitnick. 2014. Microsoft COCO: Common objects in context. (2014) 740\u2013755.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_3_1_21_2","first-page":"2082","volume-title":"AAAI Conference on Artificial Intelligence","volume":"35","author":"Liu Bingchen","year":"2021","unstructured":"Bingchen Liu, Kunpeng Song, Yizhe Zhu, Gerard de Melo, and Ahmed Elgammal. 2021. TIME: Text and image mutual-translation adversarial networks. In AAAI Conference on Artificial Intelligence, Vol. 35. 2082\u20132090."},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2022.109750"},{"key":"e_1_3_1_23_2","article-title":"Semi-supervised sequence tagging with bidirectional language models","author":"Peters Matthew E.","year":"2017","unstructured":"Matthew E. Peters, Waleed Ammar, Chandra Bhagavatula, and Russell Power. 2017. Semi-supervised sequence tagging with bidirectional language models. arXiv preprint arXiv:1705.00108 (2017).","journal-title":"arXiv preprint arXiv:1705.00108"},{"key":"e_1_3_1_24_2","article-title":"Learn, imagine and create: Text-to-image generation from prior knowledge","volume":"32","author":"Qiao Tingting","year":"2019","unstructured":"Tingting Qiao, Jing Zhang, Duanqing Xu, and Dacheng Tao. 2019. Learn, imagine and create: Text-to-image generation from prior knowledge. Adv. Neural Inf. Process. Syst. 32 (2019).","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_1_25_2","doi-asserted-by":"crossref","unstructured":"Tingting Qiao Jing Zhang Duanqing Xu and Dacheng Tao. 2019. MirrorGAN: Learning text-to-image generation by redescription. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition. (2019) 1505\u20131514.","DOI":"10.1109\/CVPR.2019.00160"},{"key":"e_1_3_1_26_2","first-page":"8748","volume-title":"International Conference on Machine Learning","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, and I. Sutskever. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning. PMLR. 8748\u20138763."},{"key":"e_1_3_1_27_2","unstructured":"Alec Radford Karthik Narasimhan Tim Salimans Ilya Sutskever et\u00a0al. 2018. Improving language understanding by generative pre-training. (2018)."},{"key":"e_1_3_1_28_2","article-title":"Hierarchical text-conditional image generation with clip latents","author":"Ramesh Aditya","year":"2022","unstructured":"Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125 (2022).","journal-title":"arXiv preprint arXiv:2204.06125"},{"key":"e_1_3_1_29_2","unstructured":"Scott Reed Zeynep Akata Xinchen Yan Lajanugen Logeswaran Bernt Schiele and Honglak Lee. 2016. Generative adversarial text to image synthesis. International Conference on Machine Learning. PMLR. (2016) 1060\u20131069."},{"key":"e_1_3_1_30_2","first-page":"217","article-title":"Learning what and where to draw","volume":"29","author":"Reed Scott E.","year":"2016","unstructured":"Scott E. Reed, Zeynep Akata, Santosh Mohan, Samuel Tenka, Bernt Schiele, and Honglak Lee. 2016. Learning what and where to draw. Adv. Neural Inf. Process. Syst. 29 (2016), 217\u2013225.","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_1_31_2","article-title":"Semi-supervised multitask learning for sequence labeling","author":"Rei Marek","year":"2017","unstructured":"Marek Rei. 2017. Semi-supervised multitask learning for sequence labeling. arXiv preprint arXiv:1704.07156 (2017).","journal-title":"arXiv preprint arXiv:1704.07156"},{"key":"e_1_3_1_32_2","first-page":"1213","volume-title":"Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Rodriguez Juan","year":"2022","unstructured":"Juan Rodriguez, Todd Hay, David Gros, Zain Shamsi, and Ravi Srinivasan. 2022. Cross-domain detection of GPT-2-generated technical text. In Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 1213\u20131233."},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_3_1_34_2","first-page":"2234","article-title":"Improved techniques for training GANs","volume":"29","author":"Salimans Tim","year":"2016","unstructured":"Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. 2016. Improved techniques for training GANs. Adv. Neural Inf. Process. Syst. 29 (2016), 2234\u20132242.","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_3_1_35_2","unstructured":"Bjarne Sievers. 2020. Question answering for comparative questions with GPT-2.Conference and Labs of the Evaluation Forum. (2020)."},{"key":"e_1_3_1_36_2","first-page":"2290","volume-title":"ACM International Conference on Multimedia","author":"Sun Jianxin","year":"2021","unstructured":"Jianxin Sun, Qi Li, Weining Wang, Jian Zhao, and Zhenan Sun. 2021. Multi-caption text-to-face synthesis: Dataset and algorithm. In ACM International Conference on Multimedia. 2290\u20132298."},{"key":"e_1_3_1_37_2","doi-asserted-by":"crossref","unstructured":"Christian Szegedy Vincent Vanhoucke Sergey Ioffe Jon Shlens and Zbigniew Wojna. 2016. Rethinking the inception architecture for computer vision. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. (2016) 2818\u20132826.","DOI":"10.1109\/CVPR.2016.308"},{"key":"e_1_3_1_38_2","first-page":"10501","volume-title":"IEEE International Conference on Computer Vision","author":"Tan Hongchen","year":"2019","unstructured":"Hongchen Tan, Xiuping Liu, Xin Li, Yi Zhang, and Baocai Yin. 2019. Semantics-enhanced adversarial nets for text-to-image synthesis. In IEEE International Conference on Computer Vision. 10501\u201310510."},{"key":"e_1_3_1_39_2","doi-asserted-by":"crossref","unstructured":"Hongchen Tan Baocai Yin Kun Wei Xiuping Liu and Xin Li. 2023. ALR-GAN: Adaptive layout refinement for text-to-image synthesis. Trans. Multi. 25 (2023) 8620\u20138631.","DOI":"10.1109\/TMM.2023.3238554"},{"key":"e_1_3_1_40_2","article-title":"DF-GAN: Deep fusion generative adversarial networks for text-to-image synthesis","author":"Tao Ming","year":"2020","unstructured":"Ming Tao, Hao Tang, Songsong Wu, Nicu Sebe, Xiao-Yuan Jing, Fei Wu, and Bingkun Bao. 2020. DF-GAN: Deep fusion generative adversarial networks for text-to-image synthesis. arXiv preprint arXiv:2008.05865 (2020).","journal-title":"arXiv preprint arXiv:2008.05865"},{"key":"e_1_3_1_41_2","article-title":"Zero-shot video captioning with evolving pseudo-tokens","author":"Tewel Yoad","year":"2022","unstructured":"Yoad Tewel, Yoav Shalev, Roy Nadler, Idan Schwartz, and Lior Wolf. 2022. Zero-shot video captioning with evolving pseudo-tokens. arXiv preprint arXiv:2207.11100 (2022).","journal-title":"arXiv preprint arXiv:2207.11100"},{"key":"e_1_3_1_42_2","unstructured":"Catherine Wah Steve Branson Peter Welinder Pietro Perona and Serge Belongie. 2011. The caltech-ucsd birds-200-2011 dataset."},{"key":"e_1_3_1_43_2","doi-asserted-by":"crossref","unstructured":"Jun Wen Risheng Liu Nenggan Zheng Qian Zheng Zhefeng Gong and Junsong Yuan. 2019. Exploiting local feature patterns for unsupervised domain adaptation. Proceedings of the AAAI Conference on Artificial Intelligence. 33 1 (2019) 5401\u20135408.","DOI":"10.1609\/aaai.v33i01.33015401"},{"key":"e_1_3_1_44_2","first-page":"2256","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Xia Weihao","year":"2021","unstructured":"Weihao Xia, Yujiu Yang, Jing-Hao Xue, and Baoyuan Wu. 2021. TediGAN: Text-guided diverse face image generation and manipulation. In IEEE Conference on Computer Vision and Pattern Recognition. 2256\u20132265."},{"key":"e_1_3_1_45_2","doi-asserted-by":"crossref","unstructured":"Tao Xu Pengchuan Zhang Qiuyuan Huang Han Zhang Zhe Gan Xiaolei Huang and Xiaodong He. 2018. AttnGAN: Fine-grained text to image generation with attentional generative adversarial networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. (2018) 1316\u20131324.","DOI":"10.1109\/CVPR.2018.00143"},{"issue":"12","key":"e_1_3_1_46_2","first-page":"9733","article-title":"Zeronas: Differentiable generative adversarial networks search for zero-shot learning","volume":"44","author":"Yan Caixia","year":"2021","unstructured":"Caixia Yan, Xiaojun Chang, Zhihui Li, Weili Guan, Zongyuan Ge, Lei Zhu, and Qinghua Zheng. 2021. Zeronas: Differentiable generative adversarial networks search for zero-shot learning. IEEE Trans. Pattern Anal. Mach. Intell. 44, 12 (2021), 9733\u20139740.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3055062"},{"key":"e_1_3_1_48_2","first-page":"3081","volume-title":"AAAI Conference on Artificial Intelligence","volume":"36","author":"Yang Zhengyuan","year":"2022","unstructured":"Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Yumao Lu, Zicheng Liu, and Lijuan Wang. 2022. An empirical study of GPT-3 for few-shot knowledge-based VQA. In AAAI Conference on Artificial Intelligence, Vol. 36. 3081\u20133089."},{"key":"e_1_3_1_49_2","first-page":"2327","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Yin Guojun","year":"2019","unstructured":"Guojun Yin, Bin Liu, Lu Sheng, Nenghai Yu, Xiaogang Wang, and Jing Shao. 2019. Semantics disentangling for text-to-image generation. In IEEE Conference on Computer Vision and Pattern Recognition. 2327\u20132336."},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","unstructured":"Bowen Yuan Yefei Sheng Bing-Kun Bao Yi-Ping Phoebe Chen and Changsheng Xu. 2024. Semantic distance adversarial learning for text-to-image synthesis. In IEEE Transactions on Multimedia 26 (2024) 1255\u20131266. DOI:10.1109\/TMM.2023.3278992","DOI":"10.1109\/TMM.2023.3278992"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2951463"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2856256"},{"key":"e_1_3_1_53_2","doi-asserted-by":"crossref","unstructured":"Han Zhang Tao Xu Hongsheng Li Shaoting Zhang Xiaogang Wang Xiaolei Huang and Dimitris N. Metaxas. 2017. StackGAN: Text to photo-realistic image synthesis with stacked generative adversarial networks. Proceedings of the IEEE International Conference on Computer Vision. (2017) 5907\u20135915.","DOI":"10.1109\/ICCV.2017.629"},{"key":"e_1_3_1_54_2","unstructured":"Lingling Zhang Xiaojun Chang Jun Liu Minnan Luo Zhihui Li Lina Yao and Alex Hauptmann. 2022. TN-ZSTAD: Transferable network for zero-shot temporal activity detection. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 3 (2022) 3848\u20133861."},{"key":"e_1_3_1_55_2","article-title":"OptGAN: Optimizing and interpreting the latent space of the conditional text-to-image GANs","author":"Zhang Zhenxing","year":"2022","unstructured":"Zhenxing Zhang and Lambert Schomaker. 2022. OptGAN: Optimizing and interpreting the latent space of the conditional text-to-image GANs. arXiv preprint arXiv:2202.12929 (2022).","journal-title":"arXiv preprint arXiv:2202.12929"},{"key":"e_1_3_1_56_2","first-page":"6199","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhang Zizhao","year":"2018","unstructured":"Zizhao Zhang, Yuanpu Xie, and Lin Yang. 2018. Photographic text-to-image synthesis with a hierarchically-nested adversarial network. In IEEE Conference on Computer Vision and Pattern Recognition. 6199\u20136208."},{"key":"e_1_3_1_57_2","article-title":"LAFITE: Towards language-free training for text-to-image generation","author":"Zhou Yufan","year":"2021","unstructured":"Yufan Zhou, Ruiyi Zhang, Changyou Chen, Chunyuan Li, Chris Tensmeyer, Tong Yu, Jiuxiang Gu, Jinhui Xu, and Tong Sun. 2021. LAFITE: Towards language-free training for text-to-image generation. arXiv preprint arXiv:2111.13792 (2021).","journal-title":"arXiv preprint arXiv:2111.13792"},{"key":"e_1_3_1_58_2","unstructured":"Minfeng Zhu Pingbo Pan Wei Chen and Yi Yang. 2019. DM-GAN: Dynamic memory generative adversarial networks for text-to-image synthesis. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. (2019) 5802\u20135810."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3650033","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3650033","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:03:43Z","timestamp":1750291423000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3650033"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,16]]},"references-count":57,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2024,7,31]]}},"alternative-id":["10.1145\/3650033"],"URL":"https:\/\/doi.org\/10.1145\/3650033","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,5,16]]},"assertion":[{"value":"2023-08-06","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-02-16","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-05-16","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}