{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:14:58Z","timestamp":1750220098130,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":31,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,10,19]],"date-time":"2021-10-19T00:00:00Z","timestamp":1634601600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key Research and Development Plan of China","award":["2017YFD0400101"],"award-info":[{"award-number":["2017YFD0400101"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,10,19]]},"DOI":"10.1145\/3487075.3487158","type":"proceedings-article","created":{"date-parts":[[2021,12,7]],"date-time":"2021-12-07T20:35:46Z","timestamp":1638909346000},"page":"1-5","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Text Pared into Scene Graph for Diverse Image Generation"],"prefix":"10.1145","author":[{"given":"Yonghua","family":"Zhu","sequence":"first","affiliation":[{"name":"Shanghai Film Academy, Shanghai University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jieyu","family":"Huang","sequence":"additional","affiliation":[{"name":"Shanghai Film Academy, Shanghai University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ning","family":"Ge","sequence":"additional","affiliation":[{"name":"Shanghai Film Academy, Shanghai University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yunwen","family":"Zhu","sequence":"additional","affiliation":[{"name":"Shanghai Film Academy, Shanghai University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Binghui","family":"Zheng","sequence":"additional","affiliation":[{"name":"Shanghai Film Academy, Shanghai University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenjun","family":"Zhang","sequence":"additional","affiliation":[{"name":"Shanghai Film Academy, Shanghai University, China and Information Technology Academy, Shanghai Jian Qiao University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,12,7]]},"reference":[{"key":"e_1_3_2_1_1_1","first-page":"1060","article-title":"Generative adversarial text to image synthesis[C]\/\/International Conference on Machine Learning","author":"Reed S","year":"2016","unstructured":"Reed S , Akata Z , Yan X , ( 2016 ). Generative adversarial text to image synthesis[C]\/\/International Conference on Machine Learning . PMLR , 1060 - 1069 . Reed S, Akata Z, Yan X, (2016). Generative adversarial text to image synthesis[C]\/\/International Conference on Machine Learning. PMLR, 1060-1069.","journal-title":"PMLR"},{"key":"e_1_3_2_1_2_1","volume-title":"Image generation from scene graphs[C]\/\/Proceedings of the IEEE conference on computer vision and pattern recognition, 1219-1228","author":"Johnson J","year":"2018","unstructured":"Johnson J , Gupta A , Fei-Fei L ( 2018 ). Image generation from scene graphs[C]\/\/Proceedings of the IEEE conference on computer vision and pattern recognition, 1219-1228 .. Johnson J, Gupta A, Fei-Fei L (2018). Image generation from scene graphs[C]\/\/Proceedings of the IEEE conference on computer vision and pattern recognition, 1219-1228.."},{"key":"e_1_3_2_1_3_1","first-page":"3948","article-title":"Pastegan: A semi-parametric method to generate image from scene graph[J]","volume":"32","author":"Li Y","year":"2019","unstructured":"Li Y , Ma T , Bai Y , ( 2019 ). Pastegan: A semi-parametric method to generate image from scene graph[J] . Advances in Neural Information Processing Systems , 32 : 3948 - 3958 .. Li Y, Ma T, Bai Y, (2019). Pastegan: A semi-parametric method to generate image from scene graph[J]. Advances in Neural Information Processing Systems, 32: 3948-3958..","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_4_1","volume-title":"Specifying object attributes and relations in interactive scene generation[C]\/\/Proceedings of the IEEE\/CVF International Conference on Computer Vision, 4561-4569","author":"Ashual O","year":"2019","unstructured":"Ashual O , Wolf L ( 2019 ). Specifying object attributes and relations in interactive scene generation[C]\/\/Proceedings of the IEEE\/CVF International Conference on Computer Vision, 4561-4569 . Ashual O, Wolf L (2019). Specifying object attributes and relations in interactive scene generation[C]\/\/Proceedings of the IEEE\/CVF International Conference on Computer Vision, 4561-4569."},{"key":"e_1_3_2_1_5_1","volume-title":"Image generation from layout[C]\/\/Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 8584-8593","author":"Zhao B","year":"2019","unstructured":"Zhao B , Meng L , Yin W , ( 2019 ). Image generation from layout[C]\/\/Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 8584-8593 . Zhao B, Meng L, Yin W, (2019). Image generation from layout[C]\/\/Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 8584-8593."},{"key":"e_1_3_2_1_6_1","volume-title":"Image synthesis from reconfigurable layout and style[C]\/\/Proceedings of the IEEE\/CVF International Conference on Computer Vision, 10531-10540","author":"Sun W","year":"2019","unstructured":"Sun W , Wu T ( 2019 ). Image synthesis from reconfigurable layout and style[C]\/\/Proceedings of the IEEE\/CVF International Conference on Computer Vision, 10531-10540 . Sun W, Wu T (2019). Image synthesis from reconfigurable layout and style[C]\/\/Proceedings of the IEEE\/CVF International Conference on Computer Vision, 10531-10540."},{"key":"e_1_3_2_1_7_1","volume-title":"Object-centric image generation from layouts[J]. arXiv preprint arXiv:2003.07449, 1(2): 4","author":"Sylvain T","year":"2020","unstructured":"Sylvain T , Zhang P , Bengio Y , ( 2020 ). Object-centric image generation from layouts[J]. arXiv preprint arXiv:2003.07449, 1(2): 4 .. Sylvain T, Zhang P, Bengio Y, (2020). Object-centric image generation from layouts[J]. arXiv preprint arXiv:2003.07449, 1(2): 4.."},{"key":"e_1_3_2_1_8_1","volume-title":"Text2scene: Generating compositional scenes from textual descriptions[C]\/\/Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 6710-6719","author":"Tan F","year":"2019","unstructured":"Tan F , Feng S , Ordonez V ( 2019 ). Text2scene: Generating compositional scenes from textual descriptions[C]\/\/Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 6710-6719 Tan F, Feng S, Ordonez V (2019). Text2scene: Generating compositional scenes from textual descriptions[C]\/\/Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 6710-6719"},{"key":"e_1_3_2_1_9_1","volume-title":"Distributed representations of words and phrases and their compositionality[C]\/\/Advances in neural information processing systems, 3111-3119","author":"Mikolov T","year":"2013","unstructured":"Mikolov T , Sutskever I , Chen K , ( 2013 ). Distributed representations of words and phrases and their compositionality[C]\/\/Advances in neural information processing systems, 3111-3119 . Mikolov T, Sutskever I, Chen K, (2013). Distributed representations of words and phrases and their compositionality[C]\/\/Advances in neural information processing systems, 3111-3119."},{"key":"e_1_3_2_1_10_1","volume-title":"Learning visual relation priors for image-text matching and image captioning with neural scene graph generators[J]. arXiv preprint arXiv:1909. 09953","author":"Lee K H","year":"2019","unstructured":"Lee K H , Palangi H , Chen X , ( 2019 ). Learning visual relation priors for image-text matching and image captioning with neural scene graph generators[J]. arXiv preprint arXiv:1909. 09953 . Lee K H, Palangi H, Chen X, (2019). Learning visual relation priors for image-text matching and image captioning with neural scene graph generators[J]. arXiv preprint arXiv:1909. 09953."},{"key":"e_1_3_2_1_11_1","volume-title":"Scene graph generation from objects, phrases and region captions[C]\/\/Proceedings of the IEEE international conference on computer vision, 1261-1270","author":"Li Y","year":"2017","unstructured":"Li Y , Ouyang W , Zhou B , ( 2017 ). Scene graph generation from objects, phrases and region captions[C]\/\/Proceedings of the IEEE international conference on computer vision, 1261-1270 . Li Y, Ouyang W, Zhou B, (2017). Scene graph generation from objects, phrases and region captions[C]\/\/Proceedings of the IEEE international conference on computer vision, 1261-1270."},{"key":"e_1_3_2_1_12_1","volume-title":"Adversarial learning of semantic relevance in text to image synthesis[C]\/\/Proceedings of the AAAI Conference on Artificial Intelligence, 33(01): 3272-3279","author":"Cha M","year":"2019","unstructured":"Cha M , Gwon Y L , Kung H T ( 2019 ). Adversarial learning of semantic relevance in text to image synthesis[C]\/\/Proceedings of the AAAI Conference on Artificial Intelligence, 33(01): 3272-3279 . Cha M, Gwon Y L, Kung H T (2019). Adversarial learning of semantic relevance in text to image synthesis[C]\/\/Proceedings of the AAAI Conference on Artificial Intelligence, 33(01): 3272-3279."},{"key":"e_1_3_2_1_13_1","volume-title":"Generating semantically precise scene graphs from textual descriptions for improved image retrieval[C]\/\/Proceedings of the fourth workshop on vision and language, 70-80","author":"Schuster S","year":"2015","unstructured":"Schuster S , Krishna R , Chang A , ( 2015 ). Generating semantically precise scene graphs from textual descriptions for improved image retrieval[C]\/\/Proceedings of the fourth workshop on vision and language, 70-80 . Schuster S, Krishna R, Chang A, (2015). Generating semantically precise scene graphs from textual descriptions for improved image retrieval[C]\/\/Proceedings of the fourth workshop on vision and language, 70-80."},{"key":"e_1_3_2_1_14_1","volume-title":"Generative adversarial nets[J]. Advances in neural information processing systems, 27","author":"Goodfellow I","year":"2014","unstructured":"Goodfellow I , Pouget-Abadie J , Mirza M , ( 2014 ). Generative adversarial nets[J]. Advances in neural information processing systems, 27 . Goodfellow I, Pouget-Abadie J, Mirza M, (2014). Generative adversarial nets[J]. Advances in neural information processing systems, 27."},{"key":"e_1_3_2_1_15_1","volume-title":"Conditional generative adversarial nets[J]. arXiv preprint arXiv:1411. 1784","author":"Mirza M","year":"2014","unstructured":"Mirza M , Osindero S ( 2014 ). Conditional generative adversarial nets[J]. arXiv preprint arXiv:1411. 1784 . Mirza M, Osindero S (2014). Conditional generative adversarial nets[J]. arXiv preprint arXiv:1411. 1784."},{"key":"e_1_3_2_1_16_1","volume-title":"Learning what and where to draw[J]. Advances in neural information processing systems, 29: 217-225","author":"Reed S E","year":"2016","unstructured":"Reed S E , Akata Z , Mohan S , ( 2016 ). Learning what and where to draw[J]. Advances in neural information processing systems, 29: 217-225 .. Reed S E, Akata Z, Mohan S, (2016). Learning what and where to draw[J]. Advances in neural information processing systems, 29: 217-225.."},{"key":"e_1_3_2_1_17_1","volume-title":"Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks[C]\/\/Proceedings of the IEEE international conference on computer vision, 5907-5915","author":"Zhang H","year":"2017","unstructured":"Zhang H , Xu T , Li H , ( 2017 ). Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks[C]\/\/Proceedings of the IEEE international conference on computer vision, 5907-5915 . Zhang H, Xu T, Li H, (2017). Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks[C]\/\/Proceedings of the IEEE international conference on computer vision, 5907-5915."},{"key":"e_1_3_2_1_18_1","volume-title":"Stackgan++: Realistic image synthesis with stacked generative adversarial networks[J]","author":"Zhang H","year":"2018","unstructured":"Zhang H , Xu T , Li H , ( 2018 ). Stackgan++: Realistic image synthesis with stacked generative adversarial networks[J] . IEEE transactions on pattern analysis and machine intelligence, 41(8): 1947-1962. Zhang H, Xu T, Li H, (2018). Stackgan++: Realistic image synthesis with stacked generative adversarial networks[J]. IEEE transactions on pattern analysis and machine intelligence, 41(8): 1947-1962."},{"key":"e_1_3_2_1_19_1","volume-title":"Attngan: Fine-grained text to image generation with attentional generative adversarial networks[C]\/\/Proceedings of the IEEE conference on computer vision and pattern recognition, 1316-1324","author":"Xu T","year":"2018","unstructured":"Xu T , Zhang P , Huang Q , ( 2018 ). Attngan: Fine-grained text to image generation with attentional generative adversarial networks[C]\/\/Proceedings of the IEEE conference on computer vision and pattern recognition, 1316-1324 . Xu T, Zhang P, Huang Q, (2018). Attngan: Fine-grained text to image generation with attentional generative adversarial networks[C]\/\/Proceedings of the IEEE conference on computer vision and pattern recognition, 1316-1324."},{"key":"e_1_3_2_1_20_1","volume-title":"Photographic text-to-image synthesis with a hierarchically-nested adversarial network[C]\/\/Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6199-6208","author":"Zhang Z","year":"2018","unstructured":"Zhang Z , Xie Y , Yang L ( 2018 ). Photographic text-to-image synthesis with a hierarchically-nested adversarial network[C]\/\/Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6199-6208 . Zhang Z, Xie Y, Yang L (2018). Photographic text-to-image synthesis with a hierarchically-nested adversarial network[C]\/\/Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6199-6208."},{"key":"e_1_3_2_1_21_1","first-page":"2642","article-title":"Conditional image synthesis with auxiliary classifier gans[C]\/\/International conference on machine learning","author":"Odena A","year":"2017","unstructured":"Odena A , Olah C , Shlens J ( 2017 ). Conditional image synthesis with auxiliary classifier gans[C]\/\/International conference on machine learning . PMLR , 2642 - 2651 . Odena A, Olah C, Shlens J (2017). Conditional image synthesis with auxiliary classifier gans[C]\/\/International conference on machine learning. PMLR, 2642-2651.","journal-title":"PMLR"},{"key":"e_1_3_2_1_22_1","volume-title":"Spice: Semantic propositional image caption evaluation[C]\/\/European conference on computer vision","author":"Anderson P","year":"2016","unstructured":"Anderson P , Fernando B , Johnson M , ( 2016 ). Spice: Semantic propositional image caption evaluation[C]\/\/European conference on computer vision . Springer , Cham , 382-398. Anderson P, Fernando B, Johnson M, (2016). Spice: Semantic propositional image caption evaluation[C]\/\/European conference on computer vision. Springer, Cham, 382-398."},{"key":"e_1_3_2_1_23_1","volume-title":"Semantic image synthesis with spatially-adaptive normalization[C]\/\/Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2337-2346","author":"Park T","year":"2019","unstructured":"Park T , Liu M Y , Wang T C , ( 2019 ). Semantic image synthesis with spatially-adaptive normalization[C]\/\/Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2337-2346 . Park T, Liu M Y, Wang T C, (2019). Semantic image synthesis with spatially-adaptive normalization[C]\/\/Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 2337-2346."},{"key":"e_1_3_2_1_24_1","volume-title":"Gans trained by a two time-scale update rule converge to a local nash equilibrium[J]. Advances in neural information processing systems, 30","author":"Heusel M","year":"2017","unstructured":"Heusel M , Ramsauer H , Unterthiner T , ( 2017 ). Gans trained by a two time-scale update rule converge to a local nash equilibrium[J]. Advances in neural information processing systems, 30 . Heusel M, Ramsauer H, Unterthiner T, (2017). Gans trained by a two time-scale update rule converge to a local nash equilibrium[J]. Advances in neural information processing systems, 30."},{"key":"e_1_3_2_1_25_1","volume-title":"Image-to-image translation with conditional adversarial networks[C]\/\/Proceedings of the IEEE conference on computer vision and pattern recognition, 1125-1134","author":"Isola P","year":"2017","unstructured":"Isola P , Zhu J Y , Zhou T , ( 2017 ). Image-to-image translation with conditional adversarial networks[C]\/\/Proceedings of the IEEE conference on computer vision and pattern recognition, 1125-1134 . Isola P, Zhu J Y, Zhou T, (2017). Image-to-image translation with conditional adversarial networks[C]\/\/Proceedings of the IEEE conference on computer vision and pattern recognition, 1125-1134."},{"key":"e_1_3_2_1_26_1","volume-title":"Perceptual losses for real-time style transfer and super-resolution[C]\/\/European conference on computer vision","author":"Johnson J","year":"2016","unstructured":"Johnson J , Alahi A , Fei-Fei L ( 2016 ). Perceptual losses for real-time style transfer and super-resolution[C]\/\/European conference on computer vision . Springer , Cham , 694-711. Johnson J, Alahi A, Fei-Fei L (2016). Perceptual losses for real-time style transfer and super-resolution[C]\/\/European conference on computer vision. Springer, Cham, 694-711."},{"key":"e_1_3_2_1_27_1","volume-title":"Microsoft coco: Common objects in context[C]\/\/European conference on computer vision","author":"Lin T Y","year":"2014","unstructured":"Lin T Y , Maire M , Belongie S , ( 2014 ). Microsoft coco: Common objects in context[C]\/\/European conference on computer vision . Springer , Cham , 740-755. Lin T Y, Maire M, Belongie S, (2014). Microsoft coco: Common objects in context[C]\/\/European conference on computer vision. Springer, Cham, 740-755."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_28_1","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_3_2_1_29_1","volume-title":"Improved techniques for training gans[J]. Advances in neural information processing systems, 29: 2234-2242","author":"Salimans T","year":"2016","unstructured":"Salimans T , Goodfellow I , Zaremba W , ( 2016 ). Improved techniques for training gans[J]. Advances in neural information processing systems, 29: 2234-2242 Salimans T, Goodfellow I, Zaremba W, (2016). Improved techniques for training gans[J]. Advances in neural information processing systems, 29: 2234-2242"},{"key":"e_1_3_2_1_30_1","volume-title":"Very deep convolutional networks for large-scale image recognition[J]. arXiv preprint arXiv:1409.1556","author":"Simonyan K","year":"2014","unstructured":"Simonyan K , Zisserman A. ( 2014 ). Very deep convolutional networks for large-scale image recognition[J]. arXiv preprint arXiv:1409.1556 . Simonyan K, Zisserman A. (2014). Very deep convolutional networks for large-scale image recognition[J]. arXiv preprint arXiv:1409.1556."},{"key":"e_1_3_2_1_31_1","volume-title":"The unreasonable effectiveness of deep features as a perceptual metric[C]\/\/Proceedings of the IEEE conference on computer vision and pattern recognition, 586-595","author":"Zhang R","year":"2018","unstructured":"Zhang R , Isola P , Efros A A , ( 2018 ). The unreasonable effectiveness of deep features as a perceptual metric[C]\/\/Proceedings of the IEEE conference on computer vision and pattern recognition, 586-595 . Zhang R, Isola P, Efros A A, (2018). The unreasonable effectiveness of deep features as a perceptual metric[C]\/\/Proceedings of the IEEE conference on computer vision and pattern recognition, 586-595."}],"event":{"acronym":"CSAE 2021","name":"CSAE 2021: The 5th International Conference on Computer Science and Application Engineering","location":"Sanya China"},"container-title":["Proceedings of the 5th International Conference on Computer Science and Application Engineering"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3487075.3487158","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3487075.3487158","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:10:10Z","timestamp":1750183810000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3487075.3487158"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,19]]},"references-count":31,"alternative-id":["10.1145\/3487075.3487158","10.1145\/3487075"],"URL":"https:\/\/doi.org\/10.1145\/3487075.3487158","relation":{},"subject":[],"published":{"date-parts":[[2021,10,19]]},"assertion":[{"value":"2021-12-07","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}