{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,10]],"date-time":"2025-12-10T08:59:39Z","timestamp":1765357179582,"version":"3.41.0"},"reference-count":55,"publisher":"Association for Computing Machinery (ACM)","issue":"2s","license":[{"start":{"date-parts":[[2023,2,17]],"date-time":"2023-02-17T00:00:00Z","timestamp":1676592000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Key-Area Research and Development Program of Guangdong Province, China","award":["2020B010165004, 2020B010166003"],"award-info":[{"award-number":["2020B010165004, 2020B010166003"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61772206, 61972162"],"award-info":[{"award-number":["61772206, 61972162"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Guangdong International Science and Technology Cooperation Project","award":["2021A0505030009"],"award-info":[{"award-number":["2021A0505030009"]}]},{"DOI":"10.13039\/501100003453","name":"Guangdong Natural Science Foundation","doi-asserted-by":"crossref","award":["2021A1515012625"],"award-info":[{"award-number":["2021A1515012625"]}],"id":[{"id":"10.13039\/501100003453","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Guangzhou Basic and Applied Research Project","award":["202102021074"],"award-info":[{"award-number":["202102021074"]}]},{"name":"CCF-Tencent Open Research fund","award":["RAGR20210114"],"award-info":[{"award-number":["RAGR20210114"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,6,30]]},"abstract":"<jats:p>Person Image Synthesis aims at transferring the appearance of the source person image into a target pose. Existing methods cannot handle large pose variations and therefore suffer from two critical problems: (1) synthesis distortion due to the entanglement of pose and appearance information among different body components and (2) failure in preserving original semantics (e.g., the same outfit). In this article, we explicitly address these two problems by proposing a Pose- and Attribute-consistent Person Image Synthesis Network (PAC-GAN). To reduce pose and appearance matching ambiguity, we propose a component-wise transferring model consisting of two stages. The former stage focuses only on synthesizing target poses, while the latter renders target appearances by explicitly transferring the appearance information from the source image to the target image in a component-wise manner. In this way, source-target matching ambiguity is eliminated due to the component-wise disentanglement of pose and appearance synthesis. Second, to maintain attribute consistency, we represent the input image as an attribute vector and impose a high-level semantic constraint using this vector to regularize the target synthesis. Extensive experimental results on the DeepFashion dataset demonstrate the superiority of our method over the state of the art, especially for maintaining pose and attribute consistencies under large pose variations.<\/jats:p>","DOI":"10.1145\/3554739","type":"journal-article","created":{"date-parts":[[2022,8,4]],"date-time":"2022-08-04T11:59:37Z","timestamp":1659614377000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["Pose- and Attribute-consistent Person Image Synthesis"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4281-6214","authenticated-orcid":false,"given":"Cheng","family":"Xu","sequence":"first","affiliation":[{"name":"South China University of Technology, Guangzhou, Guangdong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3987-7315","authenticated-orcid":false,"given":"Zejun","family":"Chen","sequence":"additional","affiliation":[{"name":"South China University of Technology, Guangzhou, Guangdong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0172-5981","authenticated-orcid":false,"given":"Jiajie","family":"Mai","sequence":"additional","affiliation":[{"name":"King\u2019s College, London, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8006-3663","authenticated-orcid":false,"given":"Xuemiao","family":"Xu","sequence":"additional","affiliation":[{"name":"South China University of Technology, Guangzhou, Guangdong, China, State Key Laboratory of Subtropical Building Science, Guangzhou, Guangdong, China, Ministry of Education Key Laboratory of Big Data and Intelligent Robot, Guangzhou, Guangdong, China, and Guangdong Provincial Key Lab of Computational Intelligence and Cyberspace Information, Guangzhou, Guangdong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3802-4644","authenticated-orcid":false,"given":"Shengfeng","family":"He","sequence":"additional","affiliation":[{"name":"South China University of Technology, Guangzhou, Guangdong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,2,17]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"8340","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Balakrishnan Guha","year":"2018","unstructured":"Guha Balakrishnan, Amy Zhao, Adrian V. Dalca, Fredo Durand, and John Guttag. 2018. Synthesizing images of humans in unseen poses. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 8340\u20138348."},{"key":"e_1_3_1_3_2","first-page":"7291","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Cao Zhe","year":"2017","unstructured":"Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh. 2017. Realtime multi-person 2D pose estimation using part affinity fields. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 7291\u20137299."},{"key":"e_1_3_1_4_2","volume-title":"Advances in neural information processing systems (NeurIPS)","author":"Dong Haoye","year":"2018","unstructured":"Haoye Dong, Xiaodan Liang, Ke Gong, Hanjiang Lai, Jia Zhu, and Jian Yin. 2018. Soft-gated warping-GAN for pose-guided person image synthesis. In Advances in neural information processing systems (NeurIPS)."},{"key":"e_1_3_1_5_2","first-page":"1161","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Dong Haoye","year":"2019","unstructured":"Haoye Dong, Xiaodan Liang, Xiaohui Shen, Bowen Wu, Bing-Cheng Chen, and Jian Yin. 2019. FW-GAN: Flow-navigated warping GAN for video virtual try-on. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV). 1161\u20131170."},{"key":"e_1_3_1_6_2","first-page":"8857","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Esser Patrick","year":"2018","unstructured":"Patrick Esser, Ekaterina Sutter, and Bj\u00f6rn Ommer. 2018. A variational u-net for conditional appearance and shape generation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 8857\u20138866."},{"key":"e_1_3_1_7_2","article-title":"Recapture as you want","author":"Gao Chen","year":"2020","unstructured":"Chen Gao, Si Liu, Ran He, Shuicheng Yan, and Bo Li. 2020. Recapture as you want. arXiv preprint arXiv:2006.01435.","journal-title":"arXiv preprint arXiv:2006.01435"},{"key":"e_1_3_1_8_2","first-page":"3370","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV)","author":"Ge Pu","year":"2021","unstructured":"Pu Ge, Qiushi Huang, Wei Xiang, Xue Jing, Yule Li, Yiyong Li, and Zhun Sun. 2021. Focus and retain: Complement the broken pose in human image synthesis. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV). 3370\u20133379."},{"key":"e_1_3_1_9_2","first-page":"932","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Gong Ke","year":"2017","unstructured":"Ke Gong, Xiaodan Liang, Dongyu Zhang, Xiaohui Shen, and Liang Lin. 2017. Look into person: Self-supervised structure-sensitive learning and a new benchmark for human parsing. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 932\u2013940."},{"key":"e_1_3_1_10_2","first-page":"770","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"He Kaiming","year":"2016","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 770\u2013778."},{"key":"e_1_3_1_11_2","volume-title":"Advances in neural information processing systems (NeurIPS)","author":"Heusel Martin","year":"2017","unstructured":"Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. In Advances in neural information processing systems (NeurIPS)."},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(81)90024-2"},{"key":"e_1_3_1_13_2","article-title":"Generating person images with appearance-aware pose stylizer","author":"Huang Siyu","year":"2020","unstructured":"Siyu Huang, Haoyi Xiong, Zhi-Qi Cheng, Qingzhong Wang, Xingran Zhou, Bihan Wen, Jun Huan, and Dejing Dou. 2020. Generating person images with appearance-aware pose stylizer. arXiv preprint arXiv:2007.09077.","journal-title":"arXiv preprint arXiv:2007.09077"},{"key":"e_1_3_1_14_2","first-page":"1501","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Huang Xun","year":"2017","unstructured":"Xun Huang and Serge Belongie. 2017. Arbitrary style transfer in real-time with adaptive instance normalization. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV). 1501\u20131510."},{"key":"e_1_3_1_15_2","first-page":"1125","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Isola Phillip","year":"2017","unstructured":"Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A. Efros. 2017. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 1125\u20131134."},{"key":"e_1_3_1_16_2","first-page":"2287","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV) Workshops (ICCV Workshops)","author":"Jetchev Nikolay","year":"2017","unstructured":"Nikolay Jetchev and Urs Bergmann. 2017. The conditional analogy GAN: Swapping fashion articles on people images. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV) Workshops (ICCV Workshops). 2287\u20132292."},{"key":"e_1_3_1_17_2","article-title":"Rethinking of pedestrian attribute recognition: Realistic datasets with efficient method","author":"Jia Jian","year":"2020","unstructured":"Jian Jia, Houjing Huang, Wenjie Yang, Xiaotang Chen, and Kaiqi Huang. 2020. Rethinking of pedestrian attribute recognition: Realistic datasets with efficient method. arXiv preprint arXiv:2005.11909.","journal-title":"arXiv preprint arXiv:2005.11909"},{"key":"e_1_3_1_18_2","first-page":"4401","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Karras Tero","year":"2019","unstructured":"Tero Karras, Samuli Laine, and Timo Aila. 2019. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 4401\u20134410."},{"key":"e_1_3_1_19_2","article-title":"Adam: A method for stochastic optimization","author":"Kingma Diederik P.","year":"2014","unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.","journal-title":"arXiv preprint arXiv:1412.6980"},{"key":"e_1_3_1_20_2","article-title":"Auto-encoding variational Bayes","author":"Kingma Diederik P.","year":"2013","unstructured":"Diederik P. Kingma and Max Welling. 2013. Auto-encoding variational Bayes. arXiv preprint arXiv:1312.6114.","journal-title":"arXiv preprint arXiv:1312.6114"},{"key":"e_1_3_1_21_2","first-page":"439","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV)","author":"Lathuili\u00e8re St\u00e9phane","year":"2020","unstructured":"St\u00e9phane Lathuili\u00e8re, Enver Sangineto, Aliaksandr Siarohin, and Nicu Sebe. 2020. Attention-based fusion for multi-source human image generation. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV). 439\u2013448."},{"key":"e_1_3_1_22_2","article-title":"A richly annotated dataset for pedestrian attribute recognition","author":"Li Dangwei","year":"2016","unstructured":"Dangwei Li, Zhang Zhang, Xiaotang Chen, Haibin Ling, and Kaiqi Huang. 2016. A richly annotated dataset for pedestrian attribute recognition. arXiv preprint arXiv:1603.07054.","journal-title":"arXiv preprint arXiv:1603.07054"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3029455"},{"key":"e_1_3_1_24_2","first-page":"3693","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Li Yining","year":"2019","unstructured":"Yining Li, Chen Huang, and Chen Change Loy. 2019. Dense intrinsic appearance flow for human pose transfer. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 3693\u20133702."},{"key":"e_1_3_1_25_2","first-page":"5904","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Liu Wen","year":"2019","unstructured":"Wen Liu, Zhixin Piao, Jie Min, Wenhan Luo, Lin Ma, and Shenghua Gao. 2019. Liquid warping GAN: A unified framework for human motion imitation, appearance transfer and novel view synthesis. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 5904\u20135913."},{"key":"e_1_3_1_26_2","first-page":"1096","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Liu Ziwei","year":"2016","unstructured":"Ziwei Liu, Ping Luo, Shi Qiu, Xiaogang Wang, and Xiaoou Tang. 2016. Deepfashion: Powering robust clothes recognition and retrieval with rich annotations. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 1096\u20131104."},{"key":"e_1_3_1_27_2","first-page":"3431","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Long Jonathan","year":"2015","unstructured":"Jonathan Long, Evan Shelhamer, and Trevor Darrell. 2015. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 3431\u20133440."},{"key":"e_1_3_1_28_2","volume-title":"Advances in neural information processing systems (NeurIPS)","author":"Ma Liqian","year":"2017","unstructured":"Liqian Ma, Xu Jia, Qianru Sun, Bernt Schiele, Tinne Tuytelaars, and Luc Van Gool. 2017. Pose guided person image generation. In Advances in neural information processing systems (NeurIPS)."},{"key":"e_1_3_1_29_2","first-page":"99","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Ma Liqian","year":"2018","unstructured":"Liqian Ma, Qianru Sun, Stamatios Georgoulis, Luc Van Gool, Bernt Schiele, and Mario Fritz. 2018. Disentangled person image generation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 99\u2013108."},{"key":"e_1_3_1_30_2","first-page":"5084","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Men Yifang","year":"2020","unstructured":"Yifang Men, Yiming Mao, Yuning Jiang, Wei-Ying Ma, and Zhouhui Lian. 2020. Controllable person image synthesis with attribute-decomposed GAN. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 5084\u20135093."},{"key":"e_1_3_1_31_2","article-title":"Conditional generative adversarial nets","author":"Mirza Mehdi","year":"2014","unstructured":"Mehdi Mirza and Simon Osindero. 2014. Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784.","journal-title":"arXiv preprint arXiv:1411.1784"},{"key":"e_1_3_1_32_2","first-page":"123","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Neverova Natalia","year":"2018","unstructured":"Natalia Neverova, Riza Alp Guler, and Iasonas Kokkinos. 2018. Dense pose transfer. In Proceedings of the European Conference on Computer Vision (ECCV). 123\u2013138."},{"key":"e_1_3_1_33_2","first-page":"2337","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Park Taesung","year":"2019","unstructured":"Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu. 2019. Semantic image synthesis with spatially-adaptive normalization. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2337\u20132346."},{"key":"e_1_3_1_34_2","volume-title":"Advances in neural information processing systems (NeurIPS)","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et\u00a0al. 2019. Pytorch: An imperative style, high-performance deep learning library. In Advances in neural information processing systems (NeurIPS)."},{"key":"e_1_3_1_35_2","first-page":"8620","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Pumarola Albert","year":"2018","unstructured":"Albert Pumarola, Antonio Agudo, Alberto Sanfeliu, and Francesc Moreno-Noguer. 2018. Unsupervised person image synthesis in arbitrary poses. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 8620\u20138628."},{"key":"e_1_3_1_36_2","first-page":"7690","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Ren Yurui","year":"2020","unstructured":"Yurui Ren, Xiaoming Yu, Junming Chen, Thomas H. Li, and Ge Li. 2020. Deep image spatial transformation for person image generation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 7690\u20137699."},{"key":"e_1_3_1_37_2","first-page":"234","volume-title":"International Conference on Medical image computing and computer-assisted intervention (MICCAI)","author":"Ronneberger Olaf","year":"2015","unstructured":"Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention (MICCAI). 234\u2013241."},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_3_1_39_2","volume-title":"Advances in neural information processing systems (NeurIPS)","author":"Salimans Tim","year":"2016","unstructured":"Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. 2016. Improved techniques for training GANs. In Advances in neural information processing systems (NeurIPS)."},{"key":"e_1_3_1_40_2","first-page":"3408","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Siarohin Aliaksandr","year":"2018","unstructured":"Aliaksandr Siarohin, Enver Sangineto, St\u00e9phane Lathuiliere, and Nicu Sebe. 2018. Deformable GANs for pose-based human image generation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 3408\u20133416."},{"key":"e_1_3_1_41_2","article-title":"Very deep convolutional networks for large-scale image recognition","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556.","journal-title":"arXiv preprint arXiv:1409.1556"},{"key":"e_1_3_1_42_2","first-page":"2357","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Song Sijie","year":"2019","unstructured":"Sijie Song, Wei Zhang, Jiaying Liu, and Tao Mei. 2019. Unsupervised person image generation with semantic parsing transformation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2357\u20132366."},{"key":"e_1_3_1_43_2","article-title":"Bipartite graph reasoning GANs for person image generation","author":"Tang Hao","year":"2020","unstructured":"Hao Tang, Song Bai, Philip H. S. Torr, and Nicu Sebe. 2020. Bipartite graph reasoning GANs for person image generation. arXiv preprint arXiv:2008.04381.","journal-title":"arXiv preprint arXiv:2008.04381"},{"key":"e_1_3_1_44_2","first-page":"717","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Tang Hao","year":"2020","unstructured":"Hao Tang, Song Bai, Li Zhang, Philip H. S. Torr, and Nicu Sebe. 2020. Xinggan for person image generation. In Proceedings of the European Conference on Computer Vision (ECCV). 717\u2013734."},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350980"},{"key":"e_1_3_1_46_2","first-page":"3332","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Walker Jacob","year":"2017","unstructured":"Jacob Walker, Kenneth Marino, Abhinav Gupta, and Martial Hebert. 2017. The pose knows: Video forecasting by generating pose futures. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV). 3332\u20133341."},{"key":"e_1_3_1_47_2","first-page":"8798","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Wang Ting-Chun","year":"2018","unstructured":"Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Andrew Tao, Jan Kautz, and Bryan Catanzaro. 2018. High-resolution image synthesis and semantic manipulation with conditional GANs. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 8798\u20138807."},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_3_1_49_2","first-page":"1","volume-title":"IEEE International Conference on Multimedia and Expo (ICME)","author":"Yang Lingbo","year":"2020","unstructured":"Lingbo Yang, Pan Wang, Xinfeng Zhang, Shanshe Wang, Zhanning Gao, Peiran Ren, Xuansong Xie, Siwei Ma, and Wen Gao. 2020. Region-adaptive texture enhancement for detailed person image synthesis. In IEEE International Conference on Multimedia and Expo (ICME). 1\u20136."},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2017.2703078"},{"key":"e_1_3_1_51_2","first-page":"586","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhang Richard","year":"2018","unstructured":"Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. 2018. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 586\u2013595."},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2914575"},{"key":"e_1_3_1_53_2","first-page":"286","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Zhou Tinghui","year":"2016","unstructured":"Tinghui Zhou, Shubham Tulsiani, Weilun Sun, Jitendra Malik, and Alexei A. Efros. 2016. View synthesis by appearance flow. In Proceedings of the European Conference on Computer Vision (ECCV). 286\u2013301."},{"key":"e_1_3_1_54_2","first-page":"2223","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)","author":"Zhu Jun-Yan","year":"2017","unstructured":"Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros. 2017. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV). 2223\u20132232."},{"key":"e_1_3_1_55_2","first-page":"5104","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhu Peihao","year":"2020","unstructured":"Peihao Zhu, Rameen Abdal, Yipeng Qin, and Peter Wonka. 2020. Sean: Image synthesis with semantic region-adaptive normalization. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 5104\u20135113."},{"key":"e_1_3_1_56_2","first-page":"2347","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhu Zhen","year":"2019","unstructured":"Zhen Zhu, Tengteng Huang, Baoguang Shi, Miao Yu, Bofei Wang, and Xiang Bai. 2019. Progressive pose attention transfer for person image generation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2347\u20132356."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3554739","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3554739","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T17:49:29Z","timestamp":1750182569000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3554739"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,17]]},"references-count":55,"journal-issue":{"issue":"2s","published-print":{"date-parts":[[2023,6,30]]}},"alternative-id":["10.1145\/3554739"],"URL":"https:\/\/doi.org\/10.1145\/3554739","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2023,2,17]]},"assertion":[{"value":"2021-11-09","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-07-19","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}